diff --git a/TOOLING.md b/TOOLING.md index 10f17a6..dd4ec40 100644 --- a/TOOLING.md +++ b/TOOLING.md @@ -58,6 +58,53 @@ Long runs: `setsid nohup ~/gdl/gdl-cron.sh profiles >/dev/null 2>&1 &` and `tail -f ~/gdl/logs/$(ls -1t ~/gdl/logs | head -1)`. Nothing reaches the archive until a run finishes, so killing one midway is safe. +## Safety net: a global gallery-dl config + +`~/.config/gallery-dl/config.json` on `mattellite` exists so that a **plain +`gallery-dl ` typed by hand** — for a quick manual check, outside +`gdl-sync.py` entirely — still gets the hand-paced caution settings instead of +gallery-dl's own faster defaults. It is loaded automatically; nothing needs to +reference it. `gdl-sync.py`'s own `--sleep-request`/`--sleep`/`--rate` flags +still override it as normal — this is only a floor for when nobody passed any. + +```json +{ + "extractor": { + "instagram": { + "api": "rest", + "cookies": ["chrome", "/home/matt/.config/google-chrome-devtools"], + "sleep-request": [12.0, 20.0], + "sleep": [5.0, 10.0], + "sleep-429": 120.0, + "retries": 8, + "videos": true + } + }, + "downloader": { + "http": { + "rate": "500K", + "retries": 8 + } + } +} +``` + +Every value here mirrors `gdl-cron.sh`'s own hand-paced defaults (see its +`SLEEP_REQUEST`/`SLEEP`/`RATE` comments) and `build_config()` in +`gdl-sync.py` — `api: rest` matters most: the graphql backend issues one +request PER POST for every video and carousel, the pattern that got this +account banned once already. `cookies` here is the config-file equivalent of +`--cookies-from-browser`, so a bare `gallery-dl ` is already +authenticated as the archive account, not anonymous. + +Verified with `gallery_dl.config.load()` + `config.get(...)` (zero live +requests) — every value above loads correctly with no `--config` flag passed. + +This file is **not tracked in the repo** — like `artms.db` and the state +files, it is host-local runtime config, and it embeds the same real +Chrome-profile path already documented above. Recreate it by hand (or from +this section) after a fresh `mattellite` setup. + ## Remotes | remote | what goes there | @@ -482,6 +529,72 @@ The documented escalation is CDN 429 → 400 on a stories/highlights endpoint suspension. It has now run twice, and both times the 400 was the last warning before something worse. +## Reels: the API is blocked, scrape by scrolling instead + +As of 2026-08-26/27, `gallery-dl`'s dedicated reels extractor fails on every +profile with `HTTP redirect to home page` — confirmed hours apart, with a +freshly-warmed session and a correct `X-IG-WWW-Claim` header (that was the +first suspect; ruled out by tracing the raw HTTP exchange). It is not a +scraping-warning interstitial — the reels tab loads completely normally in a +real, already-signed-in browser — so this is Meta blocking the specific +`/api/v1/clips/user/` endpoint gallery-dl calls, not an account-health issue. +Posts, highlights and stories are unaffected; only reels breaks. + +`scripts/reels-scrape.py` works around it by never calling that endpoint: +it drives the same signed-in Chrome via its loopback CDP port (`:9222`, +already exposed for MCP automation), scrolls the reels tab like a person +would, and scrapes `/reel//` links out of the rendered page. It only +finds shortcodes — nothing is downloaded until step 2 — and dedupes against +the archive first, so re-running it costs nothing for reels already held. + +```sh +ssh mattellite + +PROFILE=someuser + +# 1. Scrape by scrolling; dedupe against the archive; write new URLs to a file. +# Needs gallery-dl's own pipx venv python -- that's where websocket-client +# (the one extra dependency this needs) got injected. +~/.local/share/pipx/venvs/gallery-dl/bin/python3 ~/gdl/reels-scrape.py \ + --profile "$PROFILE" \ + --index https://instaarchive.ergosteur.com \ + --out ~/gdl/"$PROFILE"-reels.txt + +# 2. Fetch whatever's new -- dry run first, same as any other gdl-sync.py call. +cd ~/gdl && PATH=$HOME/.local/bin:$PATH ./gdl-sync.py \ + --publish agentapi@10.20.28.200:/volume1/rslsync/sync/Instagram-archive/archives/ \ + --staging ~/gdl/staging-reels-scraped \ + --sleep-request 12 20 --sleep 5 10 --rate 500K \ + --post-urls-file ~/gdl/"$PROFILE"-reels.txt --dry-run +# ...then swap --dry-run for --execute once the plan looks right. + +# 3. Clean up the scratch files -- reels-scrape.py's --out file and gdl-sync.py's +# own staging directory are both left behind on purpose (same reasoning as +# everywhere else here: nothing is silently deleted). +rm -rf ~/gdl/staging-reels-scraped ~/gdl/staging-reels-scraped.gdl-config.json \ + ~/gdl/"$PROFILE"-reels.txt +``` + +Verified end to end against `zindoriyam` on 2026-08-27: 26 reels found on the +page, 16 already archived (correctly skipped), 10 new ones fetched and +published cleanly. + +Notes: + +- `reels-scrape.py` fails fast if Chrome/CDP isn't up + (`curl -s http://localhost:9222/json/version` to check first). +- If it finds nothing new, `gdl-sync.py --post-urls-file` refuses to run + ("no usable URLs") rather than doing nothing quietly — expected when a + profile is already caught up. +- `--max-idle-rounds` (default 3) and `--scroll-pause MIN MAX` (default + `2.0 3.5`) are worth raising for an unusually large or slow-loading reels + tab; the default stops once 3 consecutive scrolls find nothing new. +- **Not wired into `gdl-cron.sh` or the timers.** It drives your actual + browser session rather than firing a background request, and it is slower + by design (real scrolling, not an API call) — both good reasons not to run + it unattended without deciding that deliberately. Today it is a per-profile, + by-hand tool only. + ## What changed on 2026-08-20 One session, three separate pieces of work. Recorded because the reasons are