docs: document reels-scrape.py, and add a global gallery-dl safety net

Two additions:

- A "Reels" section in TOOLING.md with the exact two-step command
  sequence for scraping and fetching reels by hand, since it's a
  separate manual tool, not wired into gdl-cron.sh.
- ~/.config/gallery-dl/config.json on mattellite, so a plain
  `gallery-dl <url>` typed by hand -- outside gdl-sync.py entirely --
  still gets the hand-paced sleep/rate/api settings instead of
  gallery-dl's own faster defaults. Documented in TOOLING.md since the
  file itself isn't tracked (host-local, like artms.db).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011qAds5qr7nZRq5R4yAuxUk
This commit is contained in:
2026-08-27 14:41:49 -04:00
co-authored by Claude Sonnet 5
parent 2f46123022
commit 784ba43bbc
+113
View File
@@ -58,6 +58,53 @@ Long runs: `setsid nohup ~/gdl/gdl-cron.sh profiles >/dev/null 2>&1 &` and
`tail -f ~/gdl/logs/$(ls -1t ~/gdl/logs | head -1)`. Nothing reaches the
archive until a run finishes, so killing one midway is safe.
## Safety net: a global gallery-dl config
`~/.config/gallery-dl/config.json` on `mattellite` exists so that a **plain
`gallery-dl <url>` typed by hand** — for a quick manual check, outside
`gdl-sync.py` entirely — still gets the hand-paced caution settings instead of
gallery-dl's own faster defaults. It is loaded automatically; nothing needs to
reference it. `gdl-sync.py`'s own `--sleep-request`/`--sleep`/`--rate` flags
still override it as normal — this is only a floor for when nobody passed any.
```json
{
"extractor": {
"instagram": {
"api": "rest",
"cookies": ["chrome", "/home/matt/.config/google-chrome-devtools"],
"sleep-request": [12.0, 20.0],
"sleep": [5.0, 10.0],
"sleep-429": 120.0,
"retries": 8,
"videos": true
}
},
"downloader": {
"http": {
"rate": "500K",
"retries": 8
}
}
}
```
Every value here mirrors `gdl-cron.sh`'s own hand-paced defaults (see its
`SLEEP_REQUEST`/`SLEEP`/`RATE` comments) and `build_config()` in
`gdl-sync.py``api: rest` matters most: the graphql backend issues one
request PER POST for every video and carousel, the pattern that got this
account banned once already. `cookies` here is the config-file equivalent of
`--cookies-from-browser`, so a bare `gallery-dl <url>` is already
authenticated as the archive account, not anonymous.
Verified with `gallery_dl.config.load()` + `config.get(...)` (zero live
requests) — every value above loads correctly with no `--config` flag passed.
This file is **not tracked in the repo** — like `artms.db` and the state
files, it is host-local runtime config, and it embeds the same real
Chrome-profile path already documented above. Recreate it by hand (or from
this section) after a fresh `mattellite` setup.
## Remotes
| remote | what goes there |
@@ -482,6 +529,72 @@ The documented escalation is CDN 429 → 400 on a stories/highlights endpoint
suspension. It has now run twice, and both times the 400 was the last warning
before something worse.
## Reels: the API is blocked, scrape by scrolling instead
As of 2026-08-26/27, `gallery-dl`'s dedicated reels extractor fails on every
profile with `HTTP redirect to home page` — confirmed hours apart, with a
freshly-warmed session and a correct `X-IG-WWW-Claim` header (that was the
first suspect; ruled out by tracing the raw HTTP exchange). It is not a
scraping-warning interstitial — the reels tab loads completely normally in a
real, already-signed-in browser — so this is Meta blocking the specific
`/api/v1/clips/user/` endpoint gallery-dl calls, not an account-health issue.
Posts, highlights and stories are unaffected; only reels breaks.
`scripts/reels-scrape.py` works around it by never calling that endpoint:
it drives the same signed-in Chrome via its loopback CDP port (`:9222`,
already exposed for MCP automation), scrolls the reels tab like a person
would, and scrapes `/reel/<code>/` links out of the rendered page. It only
finds shortcodes — nothing is downloaded until step 2 — and dedupes against
the archive first, so re-running it costs nothing for reels already held.
```sh
ssh mattellite
PROFILE=someuser
# 1. Scrape by scrolling; dedupe against the archive; write new URLs to a file.
# Needs gallery-dl's own pipx venv python -- that's where websocket-client
# (the one extra dependency this needs) got injected.
~/.local/share/pipx/venvs/gallery-dl/bin/python3 ~/gdl/reels-scrape.py \
--profile "$PROFILE" \
--index https://instaarchive.ergosteur.com \
--out ~/gdl/"$PROFILE"-reels.txt
# 2. Fetch whatever's new -- dry run first, same as any other gdl-sync.py call.
cd ~/gdl && PATH=$HOME/.local/bin:$PATH ./gdl-sync.py \
--publish agentapi@10.20.28.200:/volume1/rslsync/sync/Instagram-archive/archives/ \
--staging ~/gdl/staging-reels-scraped \
--sleep-request 12 20 --sleep 5 10 --rate 500K \
--post-urls-file ~/gdl/"$PROFILE"-reels.txt --dry-run
# ...then swap --dry-run for --execute once the plan looks right.
# 3. Clean up the scratch files -- reels-scrape.py's --out file and gdl-sync.py's
# own staging directory are both left behind on purpose (same reasoning as
# everywhere else here: nothing is silently deleted).
rm -rf ~/gdl/staging-reels-scraped ~/gdl/staging-reels-scraped.gdl-config.json \
~/gdl/"$PROFILE"-reels.txt
```
Verified end to end against `zindoriyam` on 2026-08-27: 26 reels found on the
page, 16 already archived (correctly skipped), 10 new ones fetched and
published cleanly.
Notes:
- `reels-scrape.py` fails fast if Chrome/CDP isn't up
(`curl -s http://localhost:9222/json/version` to check first).
- If it finds nothing new, `gdl-sync.py --post-urls-file` refuses to run
("no usable URLs") rather than doing nothing quietly — expected when a
profile is already caught up.
- `--max-idle-rounds` (default 3) and `--scroll-pause MIN MAX` (default
`2.0 3.5`) are worth raising for an unusually large or slow-loading reels
tab; the default stops once 3 consecutive scrolls find nothing new.
- **Not wired into `gdl-cron.sh` or the timers.** It drives your actual
browser session rather than firing a background request, and it is slower
by design (real scrolling, not an API call) — both good reasons not to run
it unattended without deciding that deliberately. Today it is a per-profile,
by-hand tool only.
## What changed on 2026-08-20
One session, three separate pieces of work. Recorded because the reasons are