feat: add reels-sync.sh, and dedupe reels-scrape.py against the whole archive

reels-sync.sh is the single-command version of the two-step pipeline:
scrape a profile's reels tab, then fetch and publish whatever's new,
with the same hand-paced settings gdl-cron.sh uses. Takes a bare
username or a full profile URL. Exits clean without touching
gdl-sync.py at all when a profile has nothing new.

Also fixes a real inefficiency in reels-scrape.py's dedup, found by
running the new script twice in a row: checking only the scraped
profile's own directories missed that a shortcode already existed
under its true owner elsewhere in the archive (reposts/collabs by
other tracked accounts), so 9 already-held reels got re-fetched for no
reason. Shortcodes are globally unique, so dedup now checks every
archived profile's listing -- all local requests to the viewer's own
API, never instagram.com, so this costs nothing on the budget that
actually matters. See TOOLING.md for the full story.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011qAds5qr7nZRq5R4yAuxUk
This commit is contained in:
2026-08-27 14:54:39 -04:00
co-authored by Claude Sonnet 5
parent 784ba43bbc
commit e0566f06ac
3 changed files with 180 additions and 20 deletions
+51 -17
View File
@@ -544,12 +544,40 @@ Posts, highlights and stories are unaffected; only reels breaks.
it drives the same signed-in Chrome via its loopback CDP port (`:9222`,
already exposed for MCP automation), scrolls the reels tab like a person
would, and scrapes `/reel/<code>/` links out of the rendered page. It only
finds shortcodes — nothing is downloaded until step 2 — and dedupes against
the archive first, so re-running it costs nothing for reels already held.
finds shortcodes — nothing is downloaded until the fetch step — and dedupes
against the **whole archive**, not just the profile being scraped (see why
below), so re-running it costs nothing for reels already held anywhere.
### The easy way: `reels-sync.sh`
```sh
ssh mattellite
~/gdl/reels-sync.sh zindoriyam
# or a full URL: ~/gdl/reels-sync.sh https://www.instagram.com/zindoriyam/
```
Runs both steps (scrape, then fetch + publish whatever's new) with the same
hand-paced settings `gdl-cron.sh` uses, logs to `~/gdl/logs/reels-<profile>-*`,
and exits 0 with "no new reels" printed when a profile is already caught up
`gdl-sync.py` never even gets invoked in that case. Override pacing the
same way as `gdl-cron.sh` (`GDL_SLEEP_REQUEST`, `GDL_SLEEP`, `GDL_RATE`), plus
`GDL_SCROLL_PAUSE` and `GDL_MAX_IDLE_ROUNDS` for the scrape step. Staging
(`~/gdl/staging-reels-<profile>`) and the scraped URL list
(`~/gdl/<profile>-reels.txt`) are wiped at the START of the next run, not
after — left behind for inspection, same as `gdl-cron.sh`'s `staging-*`.
Verified end to end against `zindoriyam` on 2026-08-27: 26 reels found on the
page, 10 new ones fetched and published cleanly on the first run.
**Not wired into `gdl-cron.sh` or the timers.** Both scripts drive your
actual browser session rather than firing a background request, and are
slower by design (real scrolling, not an API call) — both good reasons not
to run this unattended without deciding that deliberately. Today it is a
per-profile, by-hand tool only.
### What it's doing, if you want to run the two steps separately
```sh
PROFILE=someuser
# 1. Scrape by scrolling; dedupe against the archive; write new URLs to a file.
@@ -567,18 +595,8 @@ cd ~/gdl && PATH=$HOME/.local/bin:$PATH ./gdl-sync.py \
--sleep-request 12 20 --sleep 5 10 --rate 500K \
--post-urls-file ~/gdl/"$PROFILE"-reels.txt --dry-run
# ...then swap --dry-run for --execute once the plan looks right.
# 3. Clean up the scratch files -- reels-scrape.py's --out file and gdl-sync.py's
# own staging directory are both left behind on purpose (same reasoning as
# everywhere else here: nothing is silently deleted).
rm -rf ~/gdl/staging-reels-scraped ~/gdl/staging-reels-scraped.gdl-config.json \
~/gdl/"$PROFILE"-reels.txt
```
Verified end to end against `zindoriyam` on 2026-08-27: 26 reels found on the
page, 16 already archived (correctly skipped), 10 new ones fetched and
published cleanly.
Notes:
- `reels-scrape.py` fails fast if Chrome/CDP isn't up
@@ -589,11 +607,27 @@ Notes:
- `--max-idle-rounds` (default 3) and `--scroll-pause MIN MAX` (default
`2.0 3.5`) are worth raising for an unusually large or slow-loading reels
tab; the default stops once 3 consecutive scrolls find nothing new.
- **Not wired into `gdl-cron.sh` or the timers.** It drives your actual
browser session rather than firing a background request, and it is slower
by design (real scrolling, not an API call) — both good reasons not to run
it unattended without deciding that deliberately. Today it is a per-profile,
by-hand tool only.
### Why dedup checks the whole archive, not just the scraped profile
First version deduped only against the scraped profile's own directories.
Re-running it on `zindoriyam` minutes after a successful fetch found "9 new"
reels again — all reposts/collabs originally by `0ct0ber19` and
`official_artms` (also tracked profiles). The *old*, broken direct-reels-tab
fetch had left orphaned sidecar-only remnants (`.json`/`.txt`, no media) for
them, misfiled under `zindoriyam - reels/` with the true owner's name baked
into the filename stem — `index_existing()` correctly does not count a
sidecar-only entry as "held", so they looked new. `gdl-sync.py` re-fetched
them, filed them correctly under `{username}` (the post's *true* owner, per
its own metadata) — where they already existed from that profile's own
regular sync — and `rsync --ignore-existing` silently skipped every one, so
nothing was lost or duplicated. But 9 Instagram requests were spent finding
that out. Since a shortcode is globally unique, `reels-scrape.py` now checks
every archived profile's listing, not just the one being scraped — all local
requests to the viewer's own API, never to `instagram.com`, so checking all
16 profiles costs nothing on the budget that actually matters. Confirmed
fixed the same day: re-running against `zindoriyam` immediately afterward
found "26 already archived, 0 new" and exited clean.
## What changed on 2026-08-20