diff --git a/TOOLING.md b/TOOLING.md index 881030c..56a506d 100644 --- a/TOOLING.md +++ b/TOOLING.md @@ -118,12 +118,47 @@ PATH=$HOME/.local/bin:$PATH ./gdl-sync.py \ --publish agentapi@10.20.28.200:/volume1/rslsync/sync/Instagram-archive/archives/ \ --archive-db ~/gdl/artms.db \ --urls-file ~/gdl/artms_account_links.txt \ + --sleep-request 12 20 --sleep 5 10 --rate 500K \ --abort 50 --dry-run # swap for --execute when the plan looks right ``` +**Do not omit the pacing flags.** `gdl-sync.py`'s own defaults are 6-10s / 3-6s +/ 1M, which is roughly half the caution the wrapper applies. A hand-run that +leaves them off is *less* careful than the automation — which is backwards, and +was true of this very example until 2026-08-22. Prefer `gdl-cron.sh`; it +carries them for you. + `PATH` matters: `gallery-dl` is a pipx install in `~/.local/bin`, which is not on cron's PATH and not on a non-login shell's either. +Useful variations: + +```sh +--profile 0ct0ber19 --profile kimxxlip # instead of --urls-file +--only stories # one surface +--max-sources 4 # hard ceiling on what one run touches +--force # ignore --min-interval AND the probe + # cache; almost never what you want +``` + +### Long runs + +A sync runs for minutes to hours at this pacing, so detach it rather than +holding an ssh session open: + +```sh +cd ~/gdl && setsid nohup ./gdl-cron.sh profiles > /dev/null 2>&1 & +tail -f ~/gdl/logs/$(ls -1t ~/gdl/logs | head -1) +``` + +**Do not kill it with `pkill -f ` over ssh.** The pattern matches your +own `ssh` command line, so you kill your own shell and the sync survives — this +happened twice on 2026-08-22. Use `pkill -x chrome` style exact-name matches, +or kill the pid: `pgrep -f 'only posts,reels' | tail -1`. + +Nothing reaches the archive until the run finishes: `gdl-sync.py` publishes +once, at the end, so a run killed midway leaves the archive untouched. + ### The three modes Renamed on 2026-08-22. `full` was misleading — it is the abort-*limited* run — @@ -342,6 +377,74 @@ Note the gap this leaves: **the scheduled `profiles` mode still uses the default the timers, or the automation will be less careful than the hand runs that followed a warning. +## When a run fails + +Worked through on 2026-08-22. Do these **in order** — the first two cost no +Instagram requests, and the third costs one page view. + +### 1. Read the error, not the exit code + +```sh +L=~/gdl/logs/$(ls -1t ~/gdl/logs | head -1); grep -vE '^(profiles|surfaces|sources|pacing|staging|publish|read )' "$L" | head -40 +``` + +Two results are **not** failures, despite how they look: + +| looks like | actually | +|---|---| +| `rsync error: some files/attrs were not transferred (code 23)` and `done; 1 step(s) failed` | the `chown` to `rslsync` failing because the ssh user is not root. Data landed. Confirm by re-running the rsync with `--dry-run`: an empty file list means everything arrived | +| a source reporting `No results` | that profile simply has no active story / no highlights | + +### 2. Check the cookies before concluding you are blocked + +Free, and it separates a local fault from a server-side one: + +```sh +PATH=$HOME/.local/bin:$PATH gallery-dl \ + --cookies-from-browser "chrome:/home/matt/.config/google-chrome-devtools" \ + --cookies-export /tmp/ck.txt +grep instagram /tmp/ck.txt | awk '{print $6, length($7)}' # names + value lengths +rm -f /tmp/ck.txt +``` + +A healthy `sessionid` is ~77 chars, printable ASCII, colon-delimited, and +unexpired; `ds_user_id` should be an 11-digit number. Garbage or non-printable +values mean Chrome's cookie decryption failed locally — not that the account is +in trouble. **Never print the values into a transcript or a bug report.** + +### 3. Look at the account in the browser + +The session lives in a dedicated Chrome on `mattellite`, on display `:1` with +CDP on 9222, run by a systemd unit: + +```sh +export XDG_RUNTIME_DIR=/run/user/$(id -u) +systemctl --user status chrome-devtools.service +systemctl --user start chrome-devtools.service # it has Restart=no +``` + +Then VNC to `:1` and open `instagram.com`. What you are looking for: + +- **`/accounts/scraping_warning/`** — *"We suspect automated behaviour on your + account"*. This is what `400 Bad Request` on `/api/v1/feed/reels_media/` + actually means; the session is in a challenge state, not banned. Dismissing + it in the browser restores API access immediately, verified 2026-08-22. +- a checkpoint or login page — the session is gone; re-log in the browser. +- a normal feed — the fault is elsewhere. + +**Do not dismiss a warning and immediately resume.** The dismissal fixes the +symptom. The behaviour that caused it is the thing to change. + +### 4. Stop, if it was a 429 or a repeated 400 + +```sh +systemctl --user disable --now gdl-sync@stories.timer gdl-sync@profiles.timer gdl-sync@full-sweep.timer +``` + +The documented escalation is CDN 429 → 400 on a stories/highlights endpoint → +suspension. It has now run twice, and both times the 400 was the last warning +before something worse. + ## What changed on 2026-08-20 One session, three separate pieces of work. Recorded because the reasons are