Erin Clark
613ba0c05d
Merge pull request #262 from bellingcat/generic_extractor_args
...
Add flexible extractor_args to generic_extractor.py
This allows users to pass any of the options listed [here](https://github.com/yt-dlp/yt-dlp/blob/master/README.md#extractor-arguments ) to yt-dlp extractor_args.
example usage:
```
generic_extractor:
facebook_cookie:
...
extractor_args:
youtube:
player_client: web,tv
generic:
is_live: true
```
2025-03-20 15:38:20 +00:00
Patrick Robertson
0a5ba3385e
Fix small bug in twitter dropin
...
- previously the 'content' was being set to a json dump of the tweet, it should be set to full_text
2025-03-20 18:55:22 +04:00
erinhmclark
2921061fde
Add flexible extractor_args to generic_extractor.py
2025-03-19 19:19:28 +00:00
erinhmclark
a577228465
Update generic_extractor.py for general/ youtube extraction.
2025-03-18 21:10:06 +00:00
Patrick Robertson
d03ecdb037
Standardise parse dates to get_datetime_from_str
2025-03-18 10:22:58 +00:00
Patrick Robertson
89e387030d
Tests for suitable URLs for tikwm
2025-03-18 10:04:03 +00:00
Patrick Robertson
8ec053ed1b
Refactor the dropin 'is_suitable' method + fix tikwm implementation
...
Makes it easier to maintain/understand.
2025-03-18 09:14:14 +00:00
Patrick Robertson
3d4056ef70
Merge pull request #223 from bellingcat/facebook_extractor
...
Create facebook dropin - working for images + text.
2025-03-17 12:45:05 +00:00
Patrick Robertson
0765640bff
Fix up tiktok dropin for slightly modified generic_extractor format
2025-03-17 10:31:22 +00:00
Patrick Robertson
06b1f4c0ca
Fix lingering merge conflict issues
2025-03-17 10:12:55 +00:00
Patrick Robertson
59b910ec30
Merge main
2025-03-17 10:05:11 +00:00
Patrick Robertson
7e360240bf
Copy ytdlp code into AA project - seems like ytdlp won't be merged anytime soon
2025-03-17 09:57:05 +00:00
Patrick Robertson
42162c5e3f
Various docs improvements based on Friday Office Hours discussion
2025-03-17 09:23:43 +00:00
Patrick Robertson
a8e5585e6c
github format
2025-03-14 12:52:01 +00:00
Patrick Robertson
19715c8ec2
Merge branch 'main' into webdriver-cookies
2025-03-14 12:44:48 +00:00
erinhmclark
72f48f0147
Fix merge conflicts.
2025-03-14 12:11:24 +00:00
erinhmclark
846474a4e2
Merge branch 'main' into linting_etc
2025-03-14 10:50:13 +00:00
Patrick Robertson
f504d2e304
Merge branch 'main' into webdriver-cookies
2025-03-14 09:37:12 +00:00
msramalho
4d67dce4c8
minor log fix
2025-03-13 19:24:05 +00:00
Patrick Robertson
f6b13327f0
Tweaks and additional debug logging
2025-03-13 17:41:41 +00:00
Patrick Robertson
589c834047
Fix parsing ytdlp args - we should first run them through the parse_options method
2025-03-13 17:41:40 +00:00
Patrick Robertson
10ceb7aa15
Move tikwm extractor into a droping for the generic extractor
2025-03-13 15:59:42 +00:00
erinhmclark
ca44a40b88
Ruff fix on src.
2025-03-10 19:03:45 +00:00
erinhmclark
85abe1837a
Ruff format with defaults.
2025-03-10 18:44:54 +00:00
Patrick Robertson
503ba3d1c1
Add note on auto updates to readme
2025-03-07 14:46:50 +00:00
Patrick Robertson
2c5e138263
Add a note on disabling the auto-update for yt-dlp
2025-03-07 11:44:24 +00:00
Patrick Robertson
478f0b2171
Tidy-ups to auto-updating code
2025-03-07 09:59:18 +00:00
Patrick Robertson
358884c5d1
Fix unit tests for yt-dlp update
2025-03-04 17:04:23 +00:00
Patrick Robertson
0eb112431b
Auto-update yt-dlp based on generic_extractor.ytdlp_update_interval (default=5 days)
2025-03-04 16:43:46 +00:00
erinhmclark
8124bb831d
Merge branch 'main' into small_issues
...
# Conflicts:
# src/auto_archiver/core/base_module.py
# src/auto_archiver/utils/misc.py
2025-02-26 13:19:49 +00:00
erinhmclark
9bc6dd5c3c
Add set_content into generic_extractor.py.
2025-02-25 20:07:00 +00:00
Patrick Robertson
f8e846d59a
Create facebook dropin - working for images + text. CAVEAT: only gets the first ~100 chars of the post at the moment
2025-02-25 11:44:35 +00:00
Patrick Robertson
7dde8d609d
Merge main
2025-02-20 10:29:57 +00:00
erinhmclark
a8ffb19325
Fix auth key name for cookies_from_browser.
2025-02-19 10:40:54 +00:00
Patrick Robertson
222a94563f
WIP: Docs tidyups+add howto on logging and authentication
...
(Authentication is WIP)
2025-02-19 10:37:04 +00:00
Patrick Robertson
ea728a7a97
TODO on facebook dropin not working
2025-02-11 15:56:12 +00:00
msramalho
91f1ebf7b3
fix temp for yandex new shortlink
2025-02-11 15:23:16 +00:00
msramalho
5478ed3860
bsky fix media fetching
2025-02-11 15:02:00 +00:00
msramalho
47d1dc9d47
typing warnings fixed
2025-02-11 15:01:37 +00:00
Patrick Robertson
63aba6ad39
Fix sphinx-autoapi imports
2025-02-07 21:54:49 +01:00
Patrick Robertson
91ca325fd5
Update yt-dlp to latest version + remove code no longer needed from bluesky dropin
2025-02-04 17:46:46 +01:00
Patrick Robertson
c574b694ed
Set up screenshot enricher to use authentication/cookies
2025-02-03 17:25:59 +01:00
Patrick Robertson
7ec328ab40
Remove cookie options from generic_extractor - it now uses 'authentication' global settings :D
2025-02-03 16:04:36 +01:00
Patrick Robertson
7a2be5a0da
Add cookie extraction to 'authentication' options, get generic_extractor working using this info
2025-02-03 16:03:07 +01:00
Patrick Robertson
c25d5cae84
Remove ArchivingContext completely
...
Context for a specific url/item is now passed around via the metadata (metadata.set_context('key', 'val') and metadata.get_context('key', default='something')
The only other thing that was passed around in ArchivingContext was the storage info, which is already accessible now via self.config
2025-01-30 17:50:54 +01:00
Patrick Robertson
d6b4b7a932
Further cleanup
...
* Removes (partly) the ArchivingOrchestrator
* Removes the cli_feeder module, and makes it the 'default', allowing you to pass URLs directly on the command line, without having to use the cumbersome --cli_feeder.urls. Just do auto-archiver https://my.url.com
* More unit tests
* Improved error handling
2025-01-30 16:44:40 +01:00
Patrick Robertson
7a4871db6b
Fix up unit tests for new structure
2025-01-28 14:40:12 +01:00
Patrick Robertson
1d2a1d4db7
Allow framework for config settings that should not be stored in config (e.g. cli_feeder.urls
...
Use 'do_not_store': True in the config settings to apply this. Also: fix up generic archiver dropins loading + local_storage defaults (same as what's in example orchestration)
2025-01-28 11:14:12 +01:00
erinhmclark
e1a9373336
Refactoring for new config setup
2025-01-27 19:03:02 +00:00
Patrick Robertson
7fd95866a1
Further fixes/changes to loading 'types' for config + manifest edits
2025-01-27 11:48:04 +01:00