docs: bring the outstanding list back in line with reality

Four claims went stale over two days and would each have misled someone
reading this cold:

- "the first unattended run has not happened yet" -- it has, and it did the
  wrong thing: skipped every source on the 20h floor and reported success.
- "no daily stories run is scheduled" -- one exists, but is disabled after the
  scraping warning, which is a materially different situation. Stories now
  depend on someone remembering, and every day nobody runs it is a day gone.
- "--abort 50 is opt-in and nothing uses it yet" -- gdl-cron.sh full passes it
  and both manual runs used it. sweep deliberately does not, which is the
  point of sweep.
- the diff baseline in "Verifying a run" was pre-2026-08-22 and would have
  made a correct run look like it had lost files.

Adds the gap that matters most for re-enabling: the scheduled modes still use
the default 6-10s pacing while the hand runs after the warning used 12-20s, so
turning the timers back on as they stand would make the automation less
careful than the humans were.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UXfdJu7QhSJLr47K7koTDF
This commit is contained in:
2026-08-22 11:15:15 -04:00
co-authored by Claude Opus 5
parent 627cedd3a4
commit 3a7dc749c7
+30 -25
View File
@@ -232,15 +232,15 @@ for u in 0ct0ber19 kimxxlip withaseul cher_ryppo zindoriyam official_artms; do
done done
``` ```
Counts after the 2026-08-20 run, to diff against: Counts after the 2026-08-22 runs, to diff against:
| profile | files | | profile | files |
|---|---:| |---|---:|
| 0ct0ber19 | 3151 | | 0ct0ber19 | 3160 |
| official_artms | 6714 | | official_artms | 6760 |
| cher_ryppo | 3031 | | cher_ryppo | 3056 |
| kimxxlip | 3019 | | kimxxlip | 3021 |
| zindoriyam | 2253 | | zindoriyam | 2256 |
| withaseul | 1760 | | withaseul | 1760 |
A stories-only run adds few files and often **none** — profiles frequently have A stories-only run adds few files and often **none** — profiles frequently have
@@ -357,31 +357,36 @@ broken; these are decisions not yet made and cleanups not yet done.
### Fetching ### Fetching
- **The first unattended run has not happened yet** — 2026-08-21 ~09:36. Until - **The timers are DISABLED** after the 2026-08-21 scraping warning; see that
it has, the timers are unproven in the one condition that matters: firing section. Their one unattended firing did the wrong thing — it skipped every
with nobody watching. Check it with "Verifying a run" above; the 20h floor source on the 20h floor and reported success — so re-enabling should wait on
means a manual run beforehand would make the automatic one a no-op. the four fixes listed there, not just on the account settling.
- `mattellite`'s `~/.ssh/id_ed25519.pub` is in the NAS's `authorized_keys` for - `mattellite`'s `~/.ssh/id_ed25519.pub` is in the NAS's `authorized_keys` for
`agentapi` (added 2026-08-20, alongside the workstation's existing key), so `agentapi` (added 2026-08-20, alongside the workstation's existing key), so
the fetch host publishes straight to the archive and no `sshpass` step is the fetch host publishes straight to the archive and no `sshpass` step is
needed. **That key is what makes the timers work** — remove it and every needed. **That key is what makes the timers work** — remove it and every
scheduled run will fetch successfully and then fail at publish. scheduled run will fetch successfully and then fail at publish.
- **5.3 GB of stale staging on `mattellite`** `~/gdl/staging` and `~/gdl/out` - **5.4 GB of stale staging on `mattellite`** (`~/gdl`, 44 GB free): `staging`
(2.2 GB each, from the 2026-08-17 run) and `~/gdl/staging-0820` / and `out` at 2.2 GB each from 2026-08-17, `staging-0820` / `out-0820` at
`~/gdl/out-0820` (446 MB each, from 2026-08-20). Every file in all four was 446 MB each, plus `staging-full` (78 MB) and `staging-stories` (66 MB) from
verified present in the live archive, so they are safe to delete. 46 GB free, 2026-08-22. All of it was verified published, so all of it is safe to delete.
so there is no urgency — but nothing will clean them up on its own. Staging is wiped per-run by `gdl-cron.sh`, but the `out-*` publish targets and
- **No daily stories run is scheduled.** Stories expire in 24h and cannot be anything created by a direct `gdl-sync.py` call are never cleaned up.
backfilled, so this is the *only* surface where waiting loses content - **Stories currently depend on someone remembering.** The daily timer exists
permanently. `--only stories` never seeds and costs roughly six requests for but is disabled, so the one surface that cannot be backfilled has no
all six profiles. When scheduling it, randomise the minute and avoid the hour automation behind it. Every day nobody runs `gdl-cron.sh stories` is a day
boundary: a job firing at exactly 09:00 daily is obviously a machine. of stories gone. That is the central unresolved tension: the cadence that
- **`--abort 50` is opt-in and nothing uses it yet.** It is the right setting protects stories is also the most machine-like pattern here.
for routine runs — it cut a 2151-post profile to 7 enumerated posts — but it - **The scheduled modes have not been reconciled with the pacing used by
stops noticing **edited carousels** (test case 15), which only a full hand.** The 2026-08-22 runs used `--sleep-request 12 20 --rate 500K`;
enumeration finds. A full-sweep cadence has not been decided; quarterly was `gdl-cron.sh` still passes the defaults (6-10s, 1M). Re-enabling the timers
suggested and never agreed. as they stand would make the automation *less* careful than the manual runs
that followed a warning.
- **`--abort 50` is opt-in.** `gdl-cron.sh full` passes it and the manual runs
used it; `sweep` deliberately does not. It stops noticing **edited
carousels** (test case 15), which only a full enumeration finds — which is
what `sweep` is for. The quarterly cadence was proposed and never agreed.
- **The `seeded` flags in `<db>.state.json` were hand-written**, reconstructed - **The `seeded` flags in `<db>.state.json` were hand-written**, reconstructed
from the 2026-08-17 log rather than derived from the archive DB. They assert from the 2026-08-17 log rather than derived from the archive DB. They assert
"the skip-archive already knows this source". If `artms.db` is ever rebuilt, "the skip-archive already knows this source". If `artms.db` is ever rebuilt,