gdl-sync.py always needed a fistful of paths on the command line, and the 2026-08-20 run published to a local directory and then rsynced to the NAS by hand -- fine interactively, impossible from cron. gdl-cron.sh takes one argument (stories|full|sweep) and does the whole thing. Staging is wiped every run deliberately: what we already hold is decided by the skip-archive, never by what happens to be sitting in staging, so starting empty is correct and keeps the publish to just the new files. Documents the three cadences and what each is for -- notably that only `sweep` notices carousels edited after we archived them, and that only `stories` loses content if it is skipped, since stories cannot be backfilled. Cron notes the two things that would silently break it: cron runs /bin/sh so $RANDOM does not exist (use shuf), and `%` in a crontab line means newline. systemd is offered as the better option because RandomizedDelaySec and Persistent=true are built in -- a host asleep at 09:00 silently skips a cron job, and a skipped stories run is content gone for good. The publish step is still blocked on mattellite having no key on the NAS; that is recorded under Outstanding as the one thing between this and unattended. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UXfdJu7QhSJLr47K7koTDF
231 lines
10 KiB
Markdown
231 lines
10 KiB
Markdown
# Tooling branch
|
|
|
|
> [!CAUTION]
|
|
> ## This branch must not be pushed to GitHub
|
|
>
|
|
> `tooling` is the only branch that still contains the archive-fetching
|
|
> scripts and their docs, and those name things `main` was rewritten to
|
|
> remove:
|
|
>
|
|
> - the fetch host's **public IP** (`docs/gallery-dl.md`)
|
|
> - the browser profile the session cookie is read from
|
|
> - the NAS archive path
|
|
> - the **list of Instagram accounts being archived**
|
|
>
|
|
> On 2026-08-20 `main`'s entire history was rewritten with `git filter-repo`,
|
|
> the GitHub repo was deleted and recreated, and 22 container images were
|
|
> pruned from ghcr — all to get exactly this material out of public view.
|
|
> **One push of this branch to GitHub undoes all of it.**
|
|
>
|
|
> A second cleanup would be harder than the first: after a force-push the old
|
|
> commits stayed reachable by raw SHA, and only deleting the repository
|
|
> outright removed them.
|
|
|
|
## Remotes
|
|
|
|
| remote | what goes there |
|
|
|---|---|
|
|
| `origin` → gitea | everything: `main`, `tooling`, tags, backups |
|
|
| `github` | **`main` and the current release tag only** — it exists to run the CI/CD image build |
|
|
|
|
The other 22 release tags stay on gitea. Pushing them all to GitHub triggers
|
|
one container build per tag, because each tag carries its own workflow file.
|
|
|
|
## Guards — recreate these after a fresh clone
|
|
|
|
Neither guard is versioned, so a new clone has **no protection at all**:
|
|
|
|
```sh
|
|
git config remote.github.push refs/heads/main:refs/heads/main
|
|
|
|
cat > .git/hooks/pre-push <<'HOOK'
|
|
#!/bin/sh
|
|
remote_url="$2"
|
|
case "$remote_url" in *github.com*) ;; *) exit 0 ;; esac
|
|
while read -r _ _ remote_ref _; do
|
|
[ -z "$remote_ref" ] && continue
|
|
case "$remote_ref" in
|
|
refs/heads/main|refs/tags/*) ;;
|
|
*) echo "pre-push: refusing to push '$remote_ref' to GitHub." >&2; exit 1 ;;
|
|
esac
|
|
done
|
|
exit 0
|
|
HOOK
|
|
chmod +x .git/hooks/pre-push
|
|
```
|
|
|
|
Decide on the **remote** ref, not the local one: a delete push sends
|
|
`(delete)` as the local ref, and an earlier version of this hook rejected
|
|
every deletion because of it.
|
|
|
|
## What lives here
|
|
|
|
| path | what it is |
|
|
|---|---|
|
|
| `scripts/gdl-sync.py` | the gallery-dl fetcher; replaced JD2 for the ARTMS profiles |
|
|
| `scripts/test_gdl_sync.py` | its tests |
|
|
| `scripts/jd2-sync.ts` | JDownloader `.crawljob` generator, still used elsewhere |
|
|
| `docs/gallery-dl.md` | the measurements behind every option in the fetcher — **read before changing pacing** |
|
|
| `docs/jdownloader.md` | the older JD2 flow |
|
|
| `docs/artms-instagram-accounts.txt` | the profile list passed to `--urls-file` |
|
|
|
|
## Commands
|
|
|
|
```sh
|
|
# fetch: always --dry-run first; it prints the plan and the publish step
|
|
./scripts/gdl-sync.py --index https://instaarchive.ergosteur.com \
|
|
--staging <dir> --publish <user>@<nas>:<archives> \
|
|
--archive-db <db> --urls-file artms_account_links.txt --abort 50 --dry-run
|
|
|
|
# crawljobs (no npm script — package.json is kept identical to main)
|
|
npx tsx scripts/jd2-sync.ts --archives <dir> --dry-run
|
|
```
|
|
|
|
Run the fetcher from the host whose public IP matches the browser the cookie
|
|
came from. `--abort 50` is the routine setting; omit it for a full sweep that
|
|
also catches edited carousels.
|
|
|
|
## Why there is no CLAUDE.md entry for any of this
|
|
|
|
`CLAUDE.md`, `README.md` and `package.json` are kept **byte-identical** to
|
|
`main` so that merging `main` into `tooling` never conflicts. The earlier
|
|
attempt put tooling notes in `CLAUDE.md` and a `jd2` script in `package.json`;
|
|
because `main` had *deleted* those lines, every merge re-applied the deletion.
|
|
Keep branch-specific documentation in this file, which `main` does not have.
|
|
|
|
## Running it by hand
|
|
|
|
Everything happens on **`mattellite`** — the fetch host whose public IP matches
|
|
the browser the cookie came from. Running it anywhere else is what
|
|
session-hijack detection looks for.
|
|
|
|
```sh
|
|
ssh mattellite
|
|
~/gdl/gdl-cron.sh full # or: stories | sweep
|
|
```
|
|
|
|
That is the whole thing: it wipes staging, fetches, and publishes straight to
|
|
the NAS. To drive `gdl-sync.py` directly instead — always `--dry-run` first,
|
|
which prints the plan and the exact rsync that would touch the archive:
|
|
|
|
```sh
|
|
cd ~/gdl
|
|
PATH=$HOME/.local/bin:$PATH ./gdl-sync.py \
|
|
--index https://instaarchive.ergosteur.com \
|
|
--staging ~/gdl/staging-manual \
|
|
--publish agentapi@10.20.28.200:/volume1/rslsync/sync/Instagram-archive/archives/ \
|
|
--archive-db ~/gdl/artms.db \
|
|
--urls-file ~/gdl/artms_account_links.txt \
|
|
--abort 50 --dry-run # swap for --execute when the plan looks right
|
|
```
|
|
|
|
`PATH` matters: `gallery-dl` is a pipx install in `~/.local/bin`, which is not
|
|
on cron's PATH and not on a non-login shell's either.
|
|
|
|
### The three modes
|
|
|
|
| mode | cadence | cost | why |
|
|
|---|---|---|---|
|
|
| `stories` | daily | ~6 requests | stories expire in 24h and **cannot be backfilled**; this is the only run that loses content if skipped |
|
|
| `full` | monthly | ~40-60 requests | every surface, `--abort 50` — stops enumerating once it reaches content already held |
|
|
| `sweep` | quarterly | ~420 requests | no abort; the **only** run that notices carousels edited after we archived them (test case 15) |
|
|
|
|
The skip-archive means an infrequent `full` costs barely more than a frequent
|
|
one — it only fetches what is new. Frequency buys freshness, not completeness,
|
|
except for stories.
|
|
|
|
## Cron
|
|
|
|
```sh
|
|
crontab -e
|
|
```
|
|
|
|
```cron
|
|
# Randomise the start: a job firing at exactly 09:00 every day is obviously a
|
|
# machine. `shuf` rather than $RANDOM, because cron runs /bin/sh, not bash --
|
|
# and avoid `%` in a crontab line, where it means "newline" unless escaped.
|
|
17 9 * * * sleep $(shuf -i 0-2700 -n1); $HOME/gdl/gdl-cron.sh stories
|
|
43 4 3 * * sleep $(shuf -i 0-2700 -n1); $HOME/gdl/gdl-cron.sh full
|
|
11 4 7 1,4,7,10 * sleep $(shuf -i 0-2700 -n1); $HOME/gdl/gdl-cron.sh sweep
|
|
```
|
|
|
|
The wrapper exits non-zero when the sync does, so cron will mail you on
|
|
failure; it also tees every run to `~/gdl/logs/` and keeps the last 30.
|
|
|
|
**systemd is the better fit if you want it**, because the randomisation is
|
|
built in rather than a `sleep`:
|
|
|
|
```ini
|
|
# ~/.config/systemd/user/gdl-stories.timer
|
|
[Timer]
|
|
OnCalendar=*-*-* 09:00:00
|
|
RandomizedDelaySec=45m
|
|
Persistent=true # catches up after the host is off, which cron will not
|
|
```
|
|
|
|
with `loginctl enable-linger matt` so user timers run without a session.
|
|
`Persistent=true` matters here: a laptop asleep at 09:00 silently skips a cron
|
|
job, and a skipped stories run is content gone for good.
|
|
|
|
## Outstanding
|
|
|
|
State as of 2026-08-20, after the sync run and the repo split. Nothing here is
|
|
broken; these are decisions not yet made and cleanups not yet done.
|
|
|
|
### Fetching
|
|
|
|
- **The fetch host cannot yet publish to the NAS.** `gdl-sync.py --publish`
|
|
accepts `user@host:/path` and shells out to plain `rsync`, but `mattellite`
|
|
has no key on the NAS (`agentapi@10.20.28.200: Permission denied`). The
|
|
2026-08-20 run worked around it by publishing to a local directory and
|
|
rsyncing to the NAS by hand with `sshpass` from the workstation. **Until
|
|
mattellite's `~/.ssh/id_ed25519.pub` is in the NAS's `authorized_keys`, any
|
|
cron job will fetch successfully and then fail at the publish step.** That is
|
|
the single thing standing between this and being fully unattended.
|
|
|
|
- **5.3 GB of stale staging on `mattellite`** — `~/gdl/staging` and `~/gdl/out`
|
|
(2.2 GB each, from the 2026-08-17 run) and `~/gdl/staging-0820` /
|
|
`~/gdl/out-0820` (446 MB each, from 2026-08-20). Every file in all four was
|
|
verified present in the live archive, so they are safe to delete. 46 GB free,
|
|
so there is no urgency — but nothing will clean them up on its own.
|
|
- **No daily stories run is scheduled.** Stories expire in 24h and cannot be
|
|
backfilled, so this is the *only* surface where waiting loses content
|
|
permanently. `--only stories` never seeds and costs roughly six requests for
|
|
all six profiles. When scheduling it, randomise the minute and avoid the hour
|
|
boundary: a job firing at exactly 09:00 daily is obviously a machine.
|
|
- **`--abort 50` is opt-in and nothing uses it yet.** It is the right setting
|
|
for routine runs — it cut a 2151-post profile to 7 enumerated posts — but it
|
|
stops noticing **edited carousels** (test case 15), which only a full
|
|
enumeration finds. A full-sweep cadence has not been decided; quarterly was
|
|
suggested and never agreed.
|
|
- **The `seeded` flags in `<db>.state.json` were hand-written**, reconstructed
|
|
from the 2026-08-17 log rather than derived from the archive DB. They assert
|
|
"the skip-archive already knows this source". If `artms.db` is ever rebuilt,
|
|
moved or lost, **clear the state file too** — otherwise those sources will
|
|
never re-seed and a fetch into empty staging re-downloads everything.
|
|
- **`~/gdl/gdl-sync.py` on the fetch host is a copy, not a checkout.** It
|
|
currently matches this branch (`96e5694e…`), but nothing keeps them in sync;
|
|
`scp` it after any change and re-check the hash.
|
|
- The 2026-08-20 run is split across two logs — `artms-run3.log` (12 sources,
|
|
no abort) and `artms-run4.log` (12 sources, `--abort 50`) — because it was
|
|
stopped midway to pick up the new flag.
|
|
|
|
### Repo and infrastructure
|
|
|
|
- **The `pre-rewrite-*` branches on gitea hold the unredacted history** — real
|
|
account names, the fetch host's IP, and the tooling, as it was before the
|
|
rewrite. They are deliberate backups. Decide whether they expire; the
|
|
`pre-push` hook does cover them (it allows only `main` and tags to GitHub).
|
|
- **The `pre-rewrite-full.bundle` backup is in a session scratchpad** and will
|
|
be deleted with it. If a durable backup outside gitea is wanted, move it now.
|
|
- **Only one container image exists.** 22 versions were pruned, so rolling back
|
|
to an older release means checking out its tag from gitea and pushing that
|
|
tag to GitHub to rebuild it — the old images are gone, not archived.
|
|
- **CI warns that the Node 20 actions are deprecated.** `actions/checkout@v4`,
|
|
`docker/login-action@v3`, `docker/metadata-action@v5` and
|
|
`docker/build-push-action@v5` are being forced onto Node 24. They work today;
|
|
bump when convenient.
|
|
- **`review-fixes`** on gitea is a stale v1.3.0-era branch, never merged,
|
|
published only because the whole local repo was pushed. Probably deletable.
|
|
- GitHub Actions run history was lost when the repo was recreated. Cosmetic.
|