The pause section predicted official_artms and 0ct0ber19 stories would expire uncollected. They did not: both surfaces were fetched by hand the next day at roughly double the configured caution, with 0 400s and 0 429s. 16 story media and 26 posts/reels media, 85 files into the archive. That confirms the 400s were the challenge state rather than a block -- once the interstitial was dismissed, the same endpoints served normally. It does not retire the warning, and the section says so. Two hand-paced runs are not evidence the old cadence was safe. It also flags the gap that matters for whenever the timers go back on: the scheduled "full" mode still runs at the default 6-10s pacing, so the automation would be less careful than the manual runs that followed a warning. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UXfdJu7QhSJLr47K7koTDF
415 lines
19 KiB
Markdown
415 lines
19 KiB
Markdown
# Tooling branch
|
||
|
||
> [!CAUTION]
|
||
> ## This branch must not be pushed to GitHub
|
||
>
|
||
> `tooling` is the only branch that still contains the archive-fetching
|
||
> scripts and their docs, and those name things `main` was rewritten to
|
||
> remove:
|
||
>
|
||
> - the fetch host's **public IP** (`docs/gallery-dl.md`)
|
||
> - the browser profile the session cookie is read from
|
||
> - the NAS archive path
|
||
> - the **list of Instagram accounts being archived**
|
||
>
|
||
> On 2026-08-20 `main`'s entire history was rewritten with `git filter-repo`,
|
||
> the GitHub repo was deleted and recreated, and 22 container images were
|
||
> pruned from ghcr — all to get exactly this material out of public view.
|
||
> **One push of this branch to GitHub undoes all of it.**
|
||
>
|
||
> A second cleanup would be harder than the first: after a force-push the old
|
||
> commits stayed reachable by raw SHA, and only deleting the repository
|
||
> outright removed them.
|
||
|
||
## Remotes
|
||
|
||
| remote | what goes there |
|
||
|---|---|
|
||
| `origin` → gitea | everything: `main`, `tooling`, tags, backups |
|
||
| `github` | **`main` and the current release tag only** — it exists to run the CI/CD image build |
|
||
|
||
The other 22 release tags stay on gitea. Pushing them all to GitHub triggers
|
||
one container build per tag, because each tag carries its own workflow file.
|
||
|
||
## Guards — recreate these after a fresh clone
|
||
|
||
Neither guard is versioned, so a new clone has **no protection at all**:
|
||
|
||
```sh
|
||
git config remote.github.push refs/heads/main:refs/heads/main
|
||
|
||
cat > .git/hooks/pre-push <<'HOOK'
|
||
#!/bin/sh
|
||
remote_url="$2"
|
||
case "$remote_url" in *github.com*) ;; *) exit 0 ;; esac
|
||
while read -r _ _ remote_ref _; do
|
||
[ -z "$remote_ref" ] && continue
|
||
case "$remote_ref" in
|
||
refs/heads/main|refs/tags/*) ;;
|
||
*) echo "pre-push: refusing to push '$remote_ref' to GitHub." >&2; exit 1 ;;
|
||
esac
|
||
done
|
||
exit 0
|
||
HOOK
|
||
chmod +x .git/hooks/pre-push
|
||
```
|
||
|
||
Decide on the **remote** ref, not the local one: a delete push sends
|
||
`(delete)` as the local ref, and an earlier version of this hook rejected
|
||
every deletion because of it.
|
||
|
||
## What lives here
|
||
|
||
| path | what it is |
|
||
|---|---|
|
||
| `scripts/gdl-sync.py` | the gallery-dl fetcher; replaced JD2 for the ARTMS profiles |
|
||
| `scripts/gdl-cron.sh` | unattended wrapper: `stories` \| `full` \| `sweep` |
|
||
| `scripts/systemd/` | the timers actually installed on the fetch host |
|
||
| `scripts/test_gdl_sync.py` | its tests |
|
||
| `scripts/jd2-sync.ts` | JDownloader `.crawljob` generator, still used elsewhere |
|
||
| `docs/gallery-dl.md` | the measurements behind every option in the fetcher — **read before changing pacing** |
|
||
| `docs/jdownloader.md` | the older JD2 flow |
|
||
| `docs/artms-instagram-accounts.txt` | the profile list passed to `--urls-file` |
|
||
|
||
## Commands
|
||
|
||
```sh
|
||
# fetch: always --dry-run first; it prints the plan and the publish step
|
||
./scripts/gdl-sync.py --index https://instaarchive.ergosteur.com \
|
||
--staging <dir> --publish <user>@<nas>:<archives> \
|
||
--archive-db <db> --urls-file artms_account_links.txt --abort 50 --dry-run
|
||
|
||
# crawljobs (no npm script — package.json is kept identical to main)
|
||
npx tsx scripts/jd2-sync.ts --archives <dir> --dry-run
|
||
```
|
||
|
||
Run the fetcher from the host whose public IP matches the browser the cookie
|
||
came from. `--abort 50` is the routine setting; omit it for a full sweep that
|
||
also catches edited carousels.
|
||
|
||
## Why there is no CLAUDE.md entry for any of this
|
||
|
||
`CLAUDE.md`, `README.md` and `package.json` are kept **byte-identical** to
|
||
`main` so that merging `main` into `tooling` never conflicts. The earlier
|
||
attempt put tooling notes in `CLAUDE.md` and a `jd2` script in `package.json`;
|
||
because `main` had *deleted* those lines, every merge re-applied the deletion.
|
||
Keep branch-specific documentation in this file, which `main` does not have.
|
||
|
||
## Running it by hand
|
||
|
||
Everything happens on **`mattellite`** — the fetch host whose public IP matches
|
||
the browser the cookie came from. Running it anywhere else is what
|
||
session-hijack detection looks for.
|
||
|
||
```sh
|
||
ssh mattellite
|
||
~/gdl/gdl-cron.sh full # or: stories | sweep
|
||
```
|
||
|
||
That is the whole thing: it wipes staging, fetches, and publishes straight to
|
||
the NAS. To drive `gdl-sync.py` directly instead — always `--dry-run` first,
|
||
which prints the plan and the exact rsync that would touch the archive:
|
||
|
||
```sh
|
||
cd ~/gdl
|
||
PATH=$HOME/.local/bin:$PATH ./gdl-sync.py \
|
||
--index https://instaarchive.ergosteur.com \
|
||
--staging ~/gdl/staging-manual \
|
||
--publish agentapi@10.20.28.200:/volume1/rslsync/sync/Instagram-archive/archives/ \
|
||
--archive-db ~/gdl/artms.db \
|
||
--urls-file ~/gdl/artms_account_links.txt \
|
||
--abort 50 --dry-run # swap for --execute when the plan looks right
|
||
```
|
||
|
||
`PATH` matters: `gallery-dl` is a pipx install in `~/.local/bin`, which is not
|
||
on cron's PATH and not on a non-login shell's either.
|
||
|
||
### The three modes
|
||
|
||
| mode | cadence | cost | why |
|
||
|---|---|---|---|
|
||
| `stories` | daily | ~6 requests | stories expire in 24h and **cannot be backfilled**; this is the only run that loses content if skipped |
|
||
| `full` | monthly | ~40-60 requests | every surface, `--abort 50` — stops enumerating once it reaches content already held |
|
||
| `sweep` | quarterly | ~420 requests | no abort; the **only** run that notices carousels edited after we archived them (test case 15) |
|
||
|
||
The skip-archive means an infrequent `full` costs barely more than a frequent
|
||
one — it only fetches what is new. Frequency buys freshness, not completeness,
|
||
except for stories.
|
||
|
||
## Scheduling — installed on `mattellite`
|
||
|
||
systemd **user** timers, running as `matt`, with lingering enabled so they fire
|
||
without a login session:
|
||
|
||
```sh
|
||
loginctl show-user matt --property=Linger # Linger=yes
|
||
systemctl --user list-timers 'gdl-sync@*'
|
||
```
|
||
|
||
| unit | schedule | next fire (as installed) |
|
||
|---|---|---|
|
||
| `gdl-sync@stories.timer` | daily 09:00 | 09:36:45 — the delay is the randomisation working |
|
||
| `gdl-sync@full.timer` | 3rd of each month, 04:00 | 04:37:44 |
|
||
| `gdl-sync@sweep.timer` | 7th of Jan/Apr/Jul/Oct, 04:00 | 04:42:39 |
|
||
|
||
Unit files are version-controlled in `scripts/systemd/` and installed to
|
||
`~/.config/systemd/user/`. One templated service, `gdl-sync@.service`, takes
|
||
the mode as its instance name and runs `gdl-cron.sh %i`.
|
||
|
||
Three settings are load-bearing:
|
||
|
||
- **`RandomizedDelaySec=45m`** — a job firing at exactly 09:00 daily is
|
||
obviously a machine, and the entire safety model is about not looking like
|
||
one. This is why the table above shows 09:36 rather than 09:00.
|
||
- **`Persistent=true`** — catch up a run missed because the host was off.
|
||
cron silently skips, and a skipped `stories` run is content gone for good.
|
||
- **`TimeoutStartSec=infinity`** — a sweep can run for hours at this pacing.
|
||
The default 90s would kill it mid-fetch.
|
||
|
||
Operating them:
|
||
|
||
```sh
|
||
export XDG_RUNTIME_DIR=/run/user/$(id -u) # needed over non-interactive ssh
|
||
systemctl --user start gdl-sync@stories.service # run one now
|
||
systemctl --user status gdl-sync@full.timer
|
||
journalctl --user -u 'gdl-sync@*' -n 50
|
||
systemctl --user disable --now gdl-sync@sweep.timer # stop one
|
||
```
|
||
|
||
`systemctl --user` fails with "Failed to connect to bus" over ssh unless
|
||
`XDG_RUNTIME_DIR` is set. Note also that **month names are not valid in
|
||
`OnCalendar`'s date field** — `Jan,Apr,Jul,Oct-07` is rejected outright, hence
|
||
`*-01,04,07,10-07`. Check any change with `systemd-analyze calendar '<expr>'`
|
||
before installing it.
|
||
|
||
### cron, if you ever prefer it
|
||
|
||
```cron
|
||
17 9 * * * sleep $(shuf -i 0-2700 -n1); $HOME/gdl/gdl-cron.sh stories
|
||
43 4 3 * * sleep $(shuf -i 0-2700 -n1); $HOME/gdl/gdl-cron.sh full
|
||
11 4 7 1,4,7,10 * sleep $(shuf -i 0-2700 -n1); $HOME/gdl/gdl-cron.sh sweep
|
||
```
|
||
|
||
cron runs `/bin/sh`, so `$RANDOM` does not exist — hence `shuf`. And `%` in a
|
||
crontab line means newline unless escaped, so avoid it entirely. cron has no
|
||
equivalent of `Persistent=true`.
|
||
|
||
## Verifying a run
|
||
|
||
The first unattended run is **2026-08-21, around 09:36** (09:00 plus the
|
||
randomised delay). Nothing below costs an Instagram request — every check is
|
||
against the journal, local logs, or our own viewer's API.
|
||
|
||
```sh
|
||
# 1. did it run, and did it exit 0?
|
||
ssh mattellite
|
||
export XDG_RUNTIME_DIR=/run/user/$(id -u)
|
||
systemctl --user list-timers 'gdl-sync@*' # LAST/PASSED columns
|
||
journalctl --user -u 'gdl-sync@stories.service' --since yesterday --no-pager
|
||
|
||
# 2. what did it actually fetch? (sidecars vastly outnumber media -- count media)
|
||
ls -1t ~/gdl/logs | head -3
|
||
grep -E '^(==>| FAILED|done;)' ~/gdl/logs/stories-*.log | tail -20
|
||
grep -ic 429 ~/gdl/logs/stories-*.log # MUST be 0 -- see below
|
||
```
|
||
|
||
**A `429` from `scontent-*.cdninstagram.com` ends the session, it is not a
|
||
pacing knob to tune.** The warning order last time was CDN 429 → `400` on the
|
||
highlights endpoint → suspension. If a run logs one, disable the timers and
|
||
stop for the day:
|
||
|
||
```sh
|
||
systemctl --user disable --now gdl-sync@stories.timer gdl-sync@full.timer gdl-sync@sweep.timer
|
||
```
|
||
|
||
Then confirm the archive actually grew, from the workstation:
|
||
|
||
```sh
|
||
# 3. did the publish land? compare against yesterday's counts
|
||
for u in 0ct0ber19 kimxxlip withaseul cher_ryppo zindoriyam official_artms; do
|
||
n=$(curl -s "https://instaarchive.ergosteur.com/api/archives/$u/files" | python3 -c 'import json,sys; d=json.load(sys.stdin); print(len(d if isinstance(d,list) else d["files"]))')
|
||
printf '%-18s %s\n' "$u" "$n"
|
||
done
|
||
```
|
||
|
||
Counts after the 2026-08-20 run, to diff against:
|
||
|
||
| profile | files |
|
||
|---|---:|
|
||
| 0ct0ber19 | 3151 |
|
||
| official_artms | 6714 |
|
||
| cher_ryppo | 3031 |
|
||
| kimxxlip | 3019 |
|
||
| zindoriyam | 2253 |
|
||
| withaseul | 1760 |
|
||
|
||
A stories-only run adds few files and often **none** — profiles frequently have
|
||
no active story. "0 new" is a normal result, not a failure. `fileCount` in
|
||
`/api/archives` is stale by design; use the per-profile `/files` listing.
|
||
|
||
## 2026-08-21 — scraping warning, automation stopped
|
||
|
||
**Instagram flagged the account.** Not suspended: an interstitial at
|
||
`/accounts/scraping_warning/` reading *"We suspect automated behaviour on your
|
||
account"*. It was dismissed in the browser and the account is healthy — feed
|
||
loads, still signed in. **All three timers are disabled.** Do not re-enable
|
||
them without deciding the cadence question below.
|
||
|
||
How it unfolded, because each step misled in a different way:
|
||
|
||
1. **09:13** the daily timer fired, exited 0 in one second, logged
|
||
`done; 0 step(s) failed` — and fetched nothing. Yesterday's manual run was
|
||
18.9–19.2h earlier, just under the `--min-interval 20` floor, so all six
|
||
sources were skipped. **A silent no-op on the one surface that cannot be
|
||
backfilled, reported as success.**
|
||
2. **20:28** `chrome-devtools.service` was OOM-killed (5.1 GB peak, ~1w3d CPU).
|
||
Unrelated to the above, and it does not break fetching — gallery-dl reads
|
||
the cookie *file*, not a live browser — but it meant no browser was running
|
||
to notice anything was wrong.
|
||
3. **23:19** a manual recovery run passed the floor (33h) and every source
|
||
failed with `400 Bad Request` on
|
||
`/api/v1/feed/reels_media/?reel_ids=…`. Six identical failures across six
|
||
profiles is not a per-profile fault.
|
||
4. Cookies were **exported and checked before assuming a block**:
|
||
`sessionid` 77 chars, printable, colon-delimited, 360 days to expiry;
|
||
`ds_user_id` present. Decryption was fine, so the fault was server-side.
|
||
This check costs no Instagram requests and should always come first.
|
||
5. The browser then showed the interstitial. The 400s were the challenge
|
||
state, not a ban.
|
||
|
||
### What has to change before automation is re-enabled
|
||
|
||
- **The `--min-interval` floor silently defeats the daily job.** Any manual run
|
||
in the preceding 20h makes the scheduled one a no-op. The floor exists to
|
||
stop an *aborted restart* re-enumerating profiles — a minutes-to-hours
|
||
concern — and a stories fetch is one request per profile. `stories` should
|
||
use something like `--min-interval 8`, not 20.
|
||
- **A skipped stories run must be loud.** `0 to sync, 6 skipped` currently
|
||
exits 0 and looks identical to success. On this surface a skip is a real
|
||
loss, and it should be visible in the journal without reading the log.
|
||
- **Reconsider the daily cadence itself.** A job hitting story endpoints for
|
||
six profiles every morning is the most machine-like thing here, randomised
|
||
delay or not, and it is what was flagged. Every-few-days, or on-demand, may
|
||
be the honest answer even though stories will be missed.
|
||
- **`chrome-devtools.service` has `Restart=no`** and died silently for three
|
||
hours. It needs `Restart=on-failure` and probably a `MemoryMax=`, or it will
|
||
be dead the next time the cookie needs refreshing.
|
||
|
||
### 2026-08-22 — caught up by hand, cleanly
|
||
|
||
Both surfaces were fetched manually the next day, on the owner's call, at
|
||
**roughly double the configured caution**: `--sleep-request 12 20`,
|
||
`--sleep 5 10`, `--rate 500K`, versus the defaults of 6-10 / 3-6 / 1M.
|
||
|
||
| run | result |
|
||
|---|---|
|
||
| stories, 6 profiles | 16 media, +21 files, **0 400s, 0 429s** |
|
||
| posts+reels, 12 sources, `--abort 50` | 26 media, +64 files, **0 400s, 0 429s** |
|
||
|
||
So the 400s really were the challenge state and nothing more: once the
|
||
interstitial was dismissed in the browser, the same endpoints served normally.
|
||
The stories that looked lost — `official_artms`, `0ct0ber19`, `cher_ryppo`,
|
||
`kimxxlip`, `zindoriyam` — were all captured before expiry.
|
||
|
||
**This does not retire the warning.** Two hand-paced runs a day later are not
|
||
evidence that the previous cadence was safe; they are evidence that the
|
||
account still works. What actually changed the request cost is `--abort 50`:
|
||
twelve sources across six profiles, `official_artms` included at 1829 posts
|
||
and 781 reels, finished in minutes for a few dozen requests where the old
|
||
behaviour would have spent ~400.
|
||
|
||
Note the gap this leaves: **the scheduled `full` mode still uses the default
|
||
6-10s pacing**, not the 12-20s used here. Reconcile that before re-enabling
|
||
the timers, or the automation will be less careful than the hand runs that
|
||
followed a warning.
|
||
|
||
## What changed on 2026-08-20
|
||
|
||
One session, three separate pieces of work. Recorded because the reasons are
|
||
not recoverable from the diffs.
|
||
|
||
**The sync run.** First incremental fetch in four days: 184 new media, 299
|
||
files published, 0 failures, 0 CDN 429s. 20 story items, which are the part
|
||
that could not have been recovered later. Cost about half what it would have,
|
||
because the archive DB was already seeded and the state file was primed by hand
|
||
so no probe passes ran.
|
||
|
||
**`--abort 50`.** The skip-archive suppresses *downloads*, which spends the
|
||
CDN; it does nothing about the *listing pass*, which spends `instagram.com` and
|
||
scales with how big a profile is rather than how much is new. Measured from
|
||
sidecar write times mid-run: 3 new posts took ~100s each, the other 2272 were
|
||
written in one second. Enumerating `cher_ryppo` fell from 2151 posts to 7.
|
||
|
||
**The repo split.** `main` is public and now carries none of the fetching
|
||
tooling, no host details, and no real account names — its entire history was
|
||
rewritten, the GitHub repo deleted and recreated to clear force-push residue,
|
||
and 22 container images pruned from ghcr because the server bundle had been
|
||
shipping source comments naming real accounts. This branch holds everything
|
||
that was removed. See the caution at the top.
|
||
|
||
**Automation.** mattellite got a key on the NAS, closing the last manual step,
|
||
and three systemd timers now run the sync unattended.
|
||
|
||
## Outstanding
|
||
|
||
State as of 2026-08-20, after the sync run and the repo split. Nothing here is
|
||
broken; these are decisions not yet made and cleanups not yet done.
|
||
|
||
### Fetching
|
||
|
||
- **The first unattended run has not happened yet** — 2026-08-21 ~09:36. Until
|
||
it has, the timers are unproven in the one condition that matters: firing
|
||
with nobody watching. Check it with "Verifying a run" above; the 20h floor
|
||
means a manual run beforehand would make the automatic one a no-op.
|
||
- `mattellite`'s `~/.ssh/id_ed25519.pub` is in the NAS's `authorized_keys` for
|
||
`agentapi` (added 2026-08-20, alongside the workstation's existing key), so
|
||
the fetch host publishes straight to the archive and no `sshpass` step is
|
||
needed. **That key is what makes the timers work** — remove it and every
|
||
scheduled run will fetch successfully and then fail at publish.
|
||
|
||
- **5.3 GB of stale staging on `mattellite`** — `~/gdl/staging` and `~/gdl/out`
|
||
(2.2 GB each, from the 2026-08-17 run) and `~/gdl/staging-0820` /
|
||
`~/gdl/out-0820` (446 MB each, from 2026-08-20). Every file in all four was
|
||
verified present in the live archive, so they are safe to delete. 46 GB free,
|
||
so there is no urgency — but nothing will clean them up on its own.
|
||
- **No daily stories run is scheduled.** Stories expire in 24h and cannot be
|
||
backfilled, so this is the *only* surface where waiting loses content
|
||
permanently. `--only stories` never seeds and costs roughly six requests for
|
||
all six profiles. When scheduling it, randomise the minute and avoid the hour
|
||
boundary: a job firing at exactly 09:00 daily is obviously a machine.
|
||
- **`--abort 50` is opt-in and nothing uses it yet.** It is the right setting
|
||
for routine runs — it cut a 2151-post profile to 7 enumerated posts — but it
|
||
stops noticing **edited carousels** (test case 15), which only a full
|
||
enumeration finds. A full-sweep cadence has not been decided; quarterly was
|
||
suggested and never agreed.
|
||
- **The `seeded` flags in `<db>.state.json` were hand-written**, reconstructed
|
||
from the 2026-08-17 log rather than derived from the archive DB. They assert
|
||
"the skip-archive already knows this source". If `artms.db` is ever rebuilt,
|
||
moved or lost, **clear the state file too** — otherwise those sources will
|
||
never re-seed and a fetch into empty staging re-downloads everything.
|
||
- **`~/gdl/gdl-sync.py` on the fetch host is a copy, not a checkout.** It
|
||
currently matches this branch (`96e5694e…`), but nothing keeps them in sync;
|
||
`scp` it after any change and re-check the hash.
|
||
- The 2026-08-20 run is split across two logs — `artms-run3.log` (12 sources,
|
||
no abort) and `artms-run4.log` (12 sources, `--abort 50`) — because it was
|
||
stopped midway to pick up the new flag.
|
||
|
||
### Repo and infrastructure
|
||
|
||
- **The `pre-rewrite-*` branches on gitea hold the unredacted history** — real
|
||
account names, the fetch host's IP, and the tooling, as it was before the
|
||
rewrite. They are deliberate backups. Decide whether they expire; the
|
||
`pre-push` hook does cover them (it allows only `main` and tags to GitHub).
|
||
- **The `pre-rewrite-full.bundle` backup is in a session scratchpad** and will
|
||
be deleted with it. If a durable backup outside gitea is wanted, move it now.
|
||
- **Only one container image exists.** 22 versions were pruned, so rolling back
|
||
to an older release means checking out its tag from gitea and pushing that
|
||
tag to GitHub to rebuild it — the old images are gone, not archived.
|
||
- **CI warns that the Node 20 actions are deprecated.** `actions/checkout@v4`,
|
||
`docker/login-action@v3`, `docker/metadata-action@v5` and
|
||
`docker/build-push-action@v5` are being forced onto Node 24. They work today;
|
||
bump when convenient.
|
||
- **`review-fixes`** on gitea is a stale v1.3.0-era branch, never merged,
|
||
published only because the whole local repo was pushed. Probably deletable.
|
||
- GitHub Actions run history was lost when the repo was recreated. Cosmetic.
|