Nothing here is broken — these are decisions not made and cleanups not done,
written down before the session's context is lost.
The two that can actually cost something: no daily stories run is scheduled,
and stories are the one surface that cannot be backfilled; and the `seeded`
flags in the state file were reconstructed by hand from a log rather than
derived from the archive DB, so losing artms.db without also clearing the
state file would leave those sources permanently unseeded and re-download
everything.
Also corrects "Scanner work (not done yet)", which shipped in 53b1f80 —
the three .json shapes are told apart structurally in gallery-dl-sidecar.ts
and isReel comes from the sidecar's type. That file lives on main: it parses
archives at display time and is viewer code, not tooling.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UXfdJu7QhSJLr47K7koTDF
6.7 KiB
Tooling branch
Caution
This branch must not be pushed to GitHub
toolingis the only branch that still contains the archive-fetching scripts and their docs, and those name thingsmainwas rewritten to remove:
- the fetch host's public IP (
docs/gallery-dl.md)- the browser profile the session cookie is read from
- the NAS archive path
- the list of Instagram accounts being archived
On 2026-08-20
main's entire history was rewritten withgit filter-repo, the GitHub repo was deleted and recreated, and 22 container images were pruned from ghcr — all to get exactly this material out of public view. One push of this branch to GitHub undoes all of it.A second cleanup would be harder than the first: after a force-push the old commits stayed reachable by raw SHA, and only deleting the repository outright removed them.
Remotes
| remote | what goes there |
|---|---|
origin → gitea |
everything: main, tooling, tags, backups |
github |
main and the current release tag only — it exists to run the CI/CD image build |
The other 22 release tags stay on gitea. Pushing them all to GitHub triggers one container build per tag, because each tag carries its own workflow file.
Guards — recreate these after a fresh clone
Neither guard is versioned, so a new clone has no protection at all:
git config remote.github.push refs/heads/main:refs/heads/main
cat > .git/hooks/pre-push <<'HOOK'
#!/bin/sh
remote_url="$2"
case "$remote_url" in *github.com*) ;; *) exit 0 ;; esac
while read -r _ _ remote_ref _; do
[ -z "$remote_ref" ] && continue
case "$remote_ref" in
refs/heads/main|refs/tags/*) ;;
*) echo "pre-push: refusing to push '$remote_ref' to GitHub." >&2; exit 1 ;;
esac
done
exit 0
HOOK
chmod +x .git/hooks/pre-push
Decide on the remote ref, not the local one: a delete push sends
(delete) as the local ref, and an earlier version of this hook rejected
every deletion because of it.
What lives here
| path | what it is |
|---|---|
scripts/gdl-sync.py |
the gallery-dl fetcher; replaced JD2 for the ARTMS profiles |
scripts/test_gdl_sync.py |
its tests |
scripts/jd2-sync.ts |
JDownloader .crawljob generator, still used elsewhere |
docs/gallery-dl.md |
the measurements behind every option in the fetcher — read before changing pacing |
docs/jdownloader.md |
the older JD2 flow |
docs/artms-instagram-accounts.txt |
the profile list passed to --urls-file |
Commands
# fetch: always --dry-run first; it prints the plan and the publish step
./scripts/gdl-sync.py --index https://instaarchive.ergosteur.com \
--staging <dir> --publish <user>@<nas>:<archives> \
--archive-db <db> --urls-file artms_account_links.txt --abort 50 --dry-run
# crawljobs (no npm script — package.json is kept identical to main)
npx tsx scripts/jd2-sync.ts --archives <dir> --dry-run
Run the fetcher from the host whose public IP matches the browser the cookie
came from. --abort 50 is the routine setting; omit it for a full sweep that
also catches edited carousels.
Why there is no CLAUDE.md entry for any of this
CLAUDE.md, README.md and package.json are kept byte-identical to
main so that merging main into tooling never conflicts. The earlier
attempt put tooling notes in CLAUDE.md and a jd2 script in package.json;
because main had deleted those lines, every merge re-applied the deletion.
Keep branch-specific documentation in this file, which main does not have.
Outstanding
State as of 2026-08-20, after the sync run and the repo split. Nothing here is broken; these are decisions not yet made and cleanups not yet done.
Fetching
- 5.3 GB of stale staging on
mattellite—~/gdl/stagingand~/gdl/out(2.2 GB each, from the 2026-08-17 run) and~/gdl/staging-0820/~/gdl/out-0820(446 MB each, from 2026-08-20). Every file in all four was verified present in the live archive, so they are safe to delete. 46 GB free, so there is no urgency — but nothing will clean them up on its own. - No daily stories run is scheduled. Stories expire in 24h and cannot be
backfilled, so this is the only surface where waiting loses content
permanently.
--only storiesnever seeds and costs roughly six requests for all six profiles. When scheduling it, randomise the minute and avoid the hour boundary: a job firing at exactly 09:00 daily is obviously a machine. --abort 50is opt-in and nothing uses it yet. It is the right setting for routine runs — it cut a 2151-post profile to 7 enumerated posts — but it stops noticing edited carousels (test case 15), which only a full enumeration finds. A full-sweep cadence has not been decided; quarterly was suggested and never agreed.- The
seededflags in<db>.state.jsonwere hand-written, reconstructed from the 2026-08-17 log rather than derived from the archive DB. They assert "the skip-archive already knows this source". Ifartms.dbis ever rebuilt, moved or lost, clear the state file too — otherwise those sources will never re-seed and a fetch into empty staging re-downloads everything. ~/gdl/gdl-sync.pyon the fetch host is a copy, not a checkout. It currently matches this branch (96e5694e…), but nothing keeps them in sync;scpit after any change and re-check the hash.- The 2026-08-20 run is split across two logs —
artms-run3.log(12 sources, no abort) andartms-run4.log(12 sources,--abort 50) — because it was stopped midway to pick up the new flag.
Repo and infrastructure
- The
pre-rewrite-*branches on gitea hold the unredacted history — real account names, the fetch host's IP, and the tooling, as it was before the rewrite. They are deliberate backups. Decide whether they expire; thepre-pushhook does cover them (it allows onlymainand tags to GitHub). - The
pre-rewrite-full.bundlebackup is in a session scratchpad and will be deleted with it. If a durable backup outside gitea is wanted, move it now. - Only one container image exists. 22 versions were pruned, so rolling back to an older release means checking out its tag from gitea and pushing that tag to GitHub to rebuild it — the old images are gone, not archived.
- CI warns that the Node 20 actions are deprecated.
actions/checkout@v4,docker/login-action@v3,docker/metadata-action@v5anddocker/build-push-action@v5are being forced onto Node 24. They work today; bump when convenient. review-fixeson gitea is a stale v1.3.0-era branch, never merged, published only because the whole local repo was pushed. Probably deletable.- GitHub Actions run history was lost when the repo was recreated. Cosmetic.