diff --git a/TOOLING.md b/TOOLING.md index 5765c46..490343d 100644 --- a/TOOLING.md +++ b/TOOLING.md @@ -232,15 +232,15 @@ for u in 0ct0ber19 kimxxlip withaseul cher_ryppo zindoriyam official_artms; do done ``` -Counts after the 2026-08-20 run, to diff against: +Counts after the 2026-08-22 runs, to diff against: | profile | files | |---|---:| -| 0ct0ber19 | 3151 | -| official_artms | 6714 | -| cher_ryppo | 3031 | -| kimxxlip | 3019 | -| zindoriyam | 2253 | +| 0ct0ber19 | 3160 | +| official_artms | 6760 | +| cher_ryppo | 3056 | +| kimxxlip | 3021 | +| zindoriyam | 2256 | | withaseul | 1760 | A stories-only run adds few files and often **none** — profiles frequently have @@ -357,31 +357,36 @@ broken; these are decisions not yet made and cleanups not yet done. ### Fetching -- **The first unattended run has not happened yet** — 2026-08-21 ~09:36. Until - it has, the timers are unproven in the one condition that matters: firing - with nobody watching. Check it with "Verifying a run" above; the 20h floor - means a manual run beforehand would make the automatic one a no-op. +- **The timers are DISABLED** after the 2026-08-21 scraping warning; see that + section. Their one unattended firing did the wrong thing — it skipped every + source on the 20h floor and reported success — so re-enabling should wait on + the four fixes listed there, not just on the account settling. - `mattellite`'s `~/.ssh/id_ed25519.pub` is in the NAS's `authorized_keys` for `agentapi` (added 2026-08-20, alongside the workstation's existing key), so the fetch host publishes straight to the archive and no `sshpass` step is needed. **That key is what makes the timers work** — remove it and every scheduled run will fetch successfully and then fail at publish. -- **5.3 GB of stale staging on `mattellite`** — `~/gdl/staging` and `~/gdl/out` - (2.2 GB each, from the 2026-08-17 run) and `~/gdl/staging-0820` / - `~/gdl/out-0820` (446 MB each, from 2026-08-20). Every file in all four was - verified present in the live archive, so they are safe to delete. 46 GB free, - so there is no urgency — but nothing will clean them up on its own. -- **No daily stories run is scheduled.** Stories expire in 24h and cannot be - backfilled, so this is the *only* surface where waiting loses content - permanently. `--only stories` never seeds and costs roughly six requests for - all six profiles. When scheduling it, randomise the minute and avoid the hour - boundary: a job firing at exactly 09:00 daily is obviously a machine. -- **`--abort 50` is opt-in and nothing uses it yet.** It is the right setting - for routine runs — it cut a 2151-post profile to 7 enumerated posts — but it - stops noticing **edited carousels** (test case 15), which only a full - enumeration finds. A full-sweep cadence has not been decided; quarterly was - suggested and never agreed. +- **5.4 GB of stale staging on `mattellite`** (`~/gdl`, 44 GB free): `staging` + and `out` at 2.2 GB each from 2026-08-17, `staging-0820` / `out-0820` at + 446 MB each, plus `staging-full` (78 MB) and `staging-stories` (66 MB) from + 2026-08-22. All of it was verified published, so all of it is safe to delete. + Staging is wiped per-run by `gdl-cron.sh`, but the `out-*` publish targets and + anything created by a direct `gdl-sync.py` call are never cleaned up. +- **Stories currently depend on someone remembering.** The daily timer exists + but is disabled, so the one surface that cannot be backfilled has no + automation behind it. Every day nobody runs `gdl-cron.sh stories` is a day + of stories gone. That is the central unresolved tension: the cadence that + protects stories is also the most machine-like pattern here. +- **The scheduled modes have not been reconciled with the pacing used by + hand.** The 2026-08-22 runs used `--sleep-request 12 20 --rate 500K`; + `gdl-cron.sh` still passes the defaults (6-10s, 1M). Re-enabling the timers + as they stand would make the automation *less* careful than the manual runs + that followed a warning. +- **`--abort 50` is opt-in.** `gdl-cron.sh full` passes it and the manual runs + used it; `sweep` deliberately does not. It stops noticing **edited + carousels** (test case 15), which only a full enumeration finds — which is + what `sweep` is for. The quarterly cadence was proposed and never agreed. - **The `seeded` flags in `.state.json` were hand-written**, reconstructed from the 2026-08-17 log rather than derived from the archive DB. They assert "the skip-archive already knows this source". If `artms.db` is ever rebuilt,