Patrick Robertson
37eac64442
Remove desc
2025-03-11 17:10:44 +00:00
Patrick Robertson
2d87935042
Start on opentimestamps enricher
2025-02-11 14:54:46 +00:00
erinhmclark
c8cd7ea63c
Merge branch 'load_modules' into add_module_tests
...
# Conflicts:
# src/auto_archiver/modules/telethon_extractor/telethon_extractor.py
2025-02-11 13:08:08 +00:00
msramalho
977618b4ce
doc: adds note about telethon vs telegram extractors
2025-02-11 13:04:59 +00:00
msramalho
d90d3cec28
fix telethon_extractor setup
2025-02-11 13:03:18 +00:00
msramalho
977f06c37a
renames api_db property for clarity
2025-02-11 12:56:33 +00:00
msramalho
5c59029221
updates api_db for new API endpoint
2025-02-11 12:53:58 +00:00
msramalho
4eeb39477c
improves gsheetdb feedback on retrieve sheet failure
2025-02-11 12:53:46 +00:00
msramalho
6fdd5f0e66
fix cases of single : vs :: in entrypoint
2025-02-11 12:53:12 +00:00
msramalho
e6594ad3dc
merge result into cached results for context preservation
2025-02-11 12:52:42 +00:00
msramalho
7309cd32e7
fix: context to be updated on Metadata.merge
2025-02-11 12:51:17 +00:00
Patrick Robertson
ed81dcdaf0
Remove dangling 'b = ' from config.py
2025-02-10 23:07:03 +00:00
erinhmclark
8d894066f2
Merge branch 'load_modules' into add_module_tests
...
# Conflicts:
# src/auto_archiver/modules/gsheet_feeder/gsheet_feeder.py
# src/auto_archiver/utils/misc.py
2025-02-10 19:00:05 +00:00
erinhmclark
3dae2337a1
remove cdn_url check before storage.
2025-02-10 18:56:46 +00:00
erinhmclark
e97ccf8a73
Separate setup() and module_setup().
2025-02-10 18:07:47 +00:00
erinhmclark
2c3d1f591f
Separate setup() and module_setup().
2025-02-10 17:25:15 +00:00
msramalho
12f14cccc9
fixes gsheet feeder<->db connection via context.
2025-02-10 16:58:35 +00:00
msramalho
ab6cf52533
fixes bad hash initialization
2025-02-10 16:45:28 +00:00
erinhmclark
c4bb667cec
Merge branch 'load_modules' into add_module_tests
...
# Conflicts:
# src/auto_archiver/modules/s3_storage/s3_storage.py
# src/auto_archiver/utils/gsheet.py
# src/auto_archiver/utils/misc.py
2025-02-10 16:17:08 +00:00
erinhmclark
f311621e58
Small fixes.
...
Add timestamp helper method.
2025-02-10 15:57:42 +00:00
msramalho
15abf686b1
decouples s3_storage from hash_enricher
2025-02-10 15:48:54 +00:00
msramalho
8fb3dc754b
fixing telethon extractor to use default entrypoint
2025-02-10 14:59:51 +00:00
msramalho
7c848046e8
adds better info about wrong/missing modules
2025-02-10 14:59:32 +00:00
Patrick Robertson
74207d7821
Implementation tests for auto-archiver
2025-02-10 13:27:11 +01:00
Patrick Robertson
e9dd321dcd
Fix setting cli_feeder as default feeder on clean install
2025-02-10 13:06:24 +01:00
Patrick Robertson
1fad37fd93
Remove blank file
2025-02-07 23:08:30 +01:00
Patrick Robertson
63aba6ad39
Fix sphinx-autoapi imports
2025-02-07 21:54:49 +01:00
erinhmclark
950624dd4b
Fix S3 storage to media in whisper_enricher.py.
2025-02-07 20:26:00 +00:00
erinhmclark
2920cf685f
Small fixes to whisper_enricher.py.
2025-02-07 12:35:40 +00:00
erinhmclark
e9ad1e1b85
Pass media to storage cdn_call
2025-02-06 22:01:55 +00:00
erinhmclark
266c7a14e6
Context related fixes, some more tests.
2025-02-06 16:53:00 +00:00
erinhmclark
67504a683e
Merge branch 'load_modules' into add_module_tests
2025-02-06 10:13:37 +00:00
erinhmclark
5b0bad832f
Updated test, test metadata
2025-02-06 10:11:56 +00:00
Patrick Robertson
a506f2a88f
Clarify that an extractor's method can also return False if no valid data was found
2025-02-06 10:20:05 +01:00
Patrick Robertson
6ab8fd2ee4
Tidy up setting modules as Orchestrator attributes on startup.
...
Don't override the values in config['steps'] – the config should be left as is
2025-02-06 10:20:05 +01:00
erinhmclark
52542812dc
Merge tests from version with context.
2025-02-05 16:42:58 +00:00
Patrick Robertson
48abb5e66b
Remove dangling screenshot_enricher file. Moved to modules/screenshot_enricher
2025-02-04 18:16:03 +01:00
Patrick Robertson
91ca325fd5
Update yt-dlp to latest version + remove code no longer needed from bluesky dropin
2025-02-04 17:46:46 +01:00
Patrick Robertson
0633e17998
Close the facebook 'login' window if it's there - to allow for proper screenshots
2025-02-04 14:18:46 +01:00
Patrick Robertson
034197a81f
Fix typos in csv feeder docs (in manifest)
2025-02-04 13:40:07 +01:00
Patrick Robertson
78e6418249
Unit tests for csv feeder + fix some bugs
2025-02-04 13:37:26 +01:00
Patrick Robertson
b301f60ea3
Fix using validators set in __manifest__.py
...
E.g. you can use the validator 'is_file' to check if a config is a valid file
2025-02-04 13:37:26 +01:00
Patrick Robertson
a873e56b87
Remove old csv_feeder file - now inside a module
2025-02-04 12:57:35 +01:00
Patrick Robertson
72b5ea9ab6
Restore headless arg
2025-02-03 17:40:40 +01:00
Patrick Robertson
c574b694ed
Set up screenshot enricher to use authentication/cookies
2025-02-03 17:25:59 +01:00
Patrick Robertson
7ec328ab40
Remove cookie options from generic_extractor - it now uses 'authentication' global settings :D
2025-02-03 16:04:36 +01:00
Patrick Robertson
7a2be5a0da
Add cookie extraction to 'authentication' options, get generic_extractor working using this info
2025-02-03 16:03:07 +01:00
Patrick Robertson
9c9e9b370e
Remove lingering reference to ArchivingContext
2025-02-03 16:02:38 +01:00
Patrick Robertson
9a8c94b641
Fix getting/setting folder context for metadata
2025-02-03 16:02:17 +01:00
Patrick Robertson
c25d5cae84
Remove ArchivingContext completely
...
Context for a specific url/item is now passed around via the metadata (metadata.set_context('key', 'val') and metadata.get_context('key', default='something')
The only other thing that was passed around in ArchivingContext was the storage info, which is already accessible now via self.config
2025-01-30 17:50:54 +01:00