feat: add reels-sync.sh, and dedupe reels-scrape.py against the whole archive
reels-sync.sh is the single-command version of the two-step pipeline: scrape a profile's reels tab, then fetch and publish whatever's new, with the same hand-paced settings gdl-cron.sh uses. Takes a bare username or a full profile URL. Exits clean without touching gdl-sync.py at all when a profile has nothing new. Also fixes a real inefficiency in reels-scrape.py's dedup, found by running the new script twice in a row: checking only the scraped profile's own directories missed that a shortcode already existed under its true owner elsewhere in the archive (reposts/collabs by other tracked accounts), so 9 already-held reels got re-fetched for no reason. Shortcodes are globally unique, so dedup now checks every archived profile's listing -- all local requests to the viewer's own API, never instagram.com, so this costs nothing on the budget that actually matters. See TOOLING.md for the full story. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011qAds5qr7nZRq5R4yAuxUk
This commit is contained in:
+51
-17
@@ -544,12 +544,40 @@ Posts, highlights and stories are unaffected; only reels breaks.
|
||||
it drives the same signed-in Chrome via its loopback CDP port (`:9222`,
|
||||
already exposed for MCP automation), scrolls the reels tab like a person
|
||||
would, and scrapes `/reel/<code>/` links out of the rendered page. It only
|
||||
finds shortcodes — nothing is downloaded until step 2 — and dedupes against
|
||||
the archive first, so re-running it costs nothing for reels already held.
|
||||
finds shortcodes — nothing is downloaded until the fetch step — and dedupes
|
||||
against the **whole archive**, not just the profile being scraped (see why
|
||||
below), so re-running it costs nothing for reels already held anywhere.
|
||||
|
||||
### The easy way: `reels-sync.sh`
|
||||
|
||||
```sh
|
||||
ssh mattellite
|
||||
~/gdl/reels-sync.sh zindoriyam
|
||||
# or a full URL: ~/gdl/reels-sync.sh https://www.instagram.com/zindoriyam/
|
||||
```
|
||||
|
||||
Runs both steps (scrape, then fetch + publish whatever's new) with the same
|
||||
hand-paced settings `gdl-cron.sh` uses, logs to `~/gdl/logs/reels-<profile>-*`,
|
||||
and exits 0 with "no new reels" printed when a profile is already caught up
|
||||
— `gdl-sync.py` never even gets invoked in that case. Override pacing the
|
||||
same way as `gdl-cron.sh` (`GDL_SLEEP_REQUEST`, `GDL_SLEEP`, `GDL_RATE`), plus
|
||||
`GDL_SCROLL_PAUSE` and `GDL_MAX_IDLE_ROUNDS` for the scrape step. Staging
|
||||
(`~/gdl/staging-reels-<profile>`) and the scraped URL list
|
||||
(`~/gdl/<profile>-reels.txt`) are wiped at the START of the next run, not
|
||||
after — left behind for inspection, same as `gdl-cron.sh`'s `staging-*`.
|
||||
|
||||
Verified end to end against `zindoriyam` on 2026-08-27: 26 reels found on the
|
||||
page, 10 new ones fetched and published cleanly on the first run.
|
||||
|
||||
**Not wired into `gdl-cron.sh` or the timers.** Both scripts drive your
|
||||
actual browser session rather than firing a background request, and are
|
||||
slower by design (real scrolling, not an API call) — both good reasons not
|
||||
to run this unattended without deciding that deliberately. Today it is a
|
||||
per-profile, by-hand tool only.
|
||||
|
||||
### What it's doing, if you want to run the two steps separately
|
||||
|
||||
```sh
|
||||
PROFILE=someuser
|
||||
|
||||
# 1. Scrape by scrolling; dedupe against the archive; write new URLs to a file.
|
||||
@@ -567,18 +595,8 @@ cd ~/gdl && PATH=$HOME/.local/bin:$PATH ./gdl-sync.py \
|
||||
--sleep-request 12 20 --sleep 5 10 --rate 500K \
|
||||
--post-urls-file ~/gdl/"$PROFILE"-reels.txt --dry-run
|
||||
# ...then swap --dry-run for --execute once the plan looks right.
|
||||
|
||||
# 3. Clean up the scratch files -- reels-scrape.py's --out file and gdl-sync.py's
|
||||
# own staging directory are both left behind on purpose (same reasoning as
|
||||
# everywhere else here: nothing is silently deleted).
|
||||
rm -rf ~/gdl/staging-reels-scraped ~/gdl/staging-reels-scraped.gdl-config.json \
|
||||
~/gdl/"$PROFILE"-reels.txt
|
||||
```
|
||||
|
||||
Verified end to end against `zindoriyam` on 2026-08-27: 26 reels found on the
|
||||
page, 16 already archived (correctly skipped), 10 new ones fetched and
|
||||
published cleanly.
|
||||
|
||||
Notes:
|
||||
|
||||
- `reels-scrape.py` fails fast if Chrome/CDP isn't up
|
||||
@@ -589,11 +607,27 @@ Notes:
|
||||
- `--max-idle-rounds` (default 3) and `--scroll-pause MIN MAX` (default
|
||||
`2.0 3.5`) are worth raising for an unusually large or slow-loading reels
|
||||
tab; the default stops once 3 consecutive scrolls find nothing new.
|
||||
- **Not wired into `gdl-cron.sh` or the timers.** It drives your actual
|
||||
browser session rather than firing a background request, and it is slower
|
||||
by design (real scrolling, not an API call) — both good reasons not to run
|
||||
it unattended without deciding that deliberately. Today it is a per-profile,
|
||||
by-hand tool only.
|
||||
|
||||
### Why dedup checks the whole archive, not just the scraped profile
|
||||
|
||||
First version deduped only against the scraped profile's own directories.
|
||||
Re-running it on `zindoriyam` minutes after a successful fetch found "9 new"
|
||||
reels again — all reposts/collabs originally by `0ct0ber19` and
|
||||
`official_artms` (also tracked profiles). The *old*, broken direct-reels-tab
|
||||
fetch had left orphaned sidecar-only remnants (`.json`/`.txt`, no media) for
|
||||
them, misfiled under `zindoriyam - reels/` with the true owner's name baked
|
||||
into the filename stem — `index_existing()` correctly does not count a
|
||||
sidecar-only entry as "held", so they looked new. `gdl-sync.py` re-fetched
|
||||
them, filed them correctly under `{username}` (the post's *true* owner, per
|
||||
its own metadata) — where they already existed from that profile's own
|
||||
regular sync — and `rsync --ignore-existing` silently skipped every one, so
|
||||
nothing was lost or duplicated. But 9 Instagram requests were spent finding
|
||||
that out. Since a shortcode is globally unique, `reels-scrape.py` now checks
|
||||
every archived profile's listing, not just the one being scraped — all local
|
||||
requests to the viewer's own API, never to `instagram.com`, so checking all
|
||||
16 profiles costs nothing on the budget that actually matters. Confirmed
|
||||
fixed the same day: re-running against `zindoriyam` immediately afterward
|
||||
found "26 already archived, 0 new" and exited clean.
|
||||
|
||||
## What changed on 2026-08-20
|
||||
|
||||
|
||||
Reference in New Issue
Block a user