docs: bring the outstanding list back in line with reality
Four claims went stale over two days and would each have misled someone reading this cold: - "the first unattended run has not happened yet" -- it has, and it did the wrong thing: skipped every source on the 20h floor and reported success. - "no daily stories run is scheduled" -- one exists, but is disabled after the scraping warning, which is a materially different situation. Stories now depend on someone remembering, and every day nobody runs it is a day gone. - "--abort 50 is opt-in and nothing uses it yet" -- gdl-cron.sh full passes it and both manual runs used it. sweep deliberately does not, which is the point of sweep. - the diff baseline in "Verifying a run" was pre-2026-08-22 and would have made a correct run look like it had lost files. Adds the gap that matters most for re-enabling: the scheduled modes still use the default 6-10s pacing while the hand runs after the warning used 12-20s, so turning the timers back on as they stand would make the automation less careful than the humans were. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UXfdJu7QhSJLr47K7koTDF
This commit is contained in:
+30
-25
@@ -232,15 +232,15 @@ for u in 0ct0ber19 kimxxlip withaseul cher_ryppo zindoriyam official_artms; do
|
|||||||
done
|
done
|
||||||
```
|
```
|
||||||
|
|
||||||
Counts after the 2026-08-20 run, to diff against:
|
Counts after the 2026-08-22 runs, to diff against:
|
||||||
|
|
||||||
| profile | files |
|
| profile | files |
|
||||||
|---|---:|
|
|---|---:|
|
||||||
| 0ct0ber19 | 3151 |
|
| 0ct0ber19 | 3160 |
|
||||||
| official_artms | 6714 |
|
| official_artms | 6760 |
|
||||||
| cher_ryppo | 3031 |
|
| cher_ryppo | 3056 |
|
||||||
| kimxxlip | 3019 |
|
| kimxxlip | 3021 |
|
||||||
| zindoriyam | 2253 |
|
| zindoriyam | 2256 |
|
||||||
| withaseul | 1760 |
|
| withaseul | 1760 |
|
||||||
|
|
||||||
A stories-only run adds few files and often **none** — profiles frequently have
|
A stories-only run adds few files and often **none** — profiles frequently have
|
||||||
@@ -357,31 +357,36 @@ broken; these are decisions not yet made and cleanups not yet done.
|
|||||||
|
|
||||||
### Fetching
|
### Fetching
|
||||||
|
|
||||||
- **The first unattended run has not happened yet** — 2026-08-21 ~09:36. Until
|
- **The timers are DISABLED** after the 2026-08-21 scraping warning; see that
|
||||||
it has, the timers are unproven in the one condition that matters: firing
|
section. Their one unattended firing did the wrong thing — it skipped every
|
||||||
with nobody watching. Check it with "Verifying a run" above; the 20h floor
|
source on the 20h floor and reported success — so re-enabling should wait on
|
||||||
means a manual run beforehand would make the automatic one a no-op.
|
the four fixes listed there, not just on the account settling.
|
||||||
- `mattellite`'s `~/.ssh/id_ed25519.pub` is in the NAS's `authorized_keys` for
|
- `mattellite`'s `~/.ssh/id_ed25519.pub` is in the NAS's `authorized_keys` for
|
||||||
`agentapi` (added 2026-08-20, alongside the workstation's existing key), so
|
`agentapi` (added 2026-08-20, alongside the workstation's existing key), so
|
||||||
the fetch host publishes straight to the archive and no `sshpass` step is
|
the fetch host publishes straight to the archive and no `sshpass` step is
|
||||||
needed. **That key is what makes the timers work** — remove it and every
|
needed. **That key is what makes the timers work** — remove it and every
|
||||||
scheduled run will fetch successfully and then fail at publish.
|
scheduled run will fetch successfully and then fail at publish.
|
||||||
|
|
||||||
- **5.3 GB of stale staging on `mattellite`** — `~/gdl/staging` and `~/gdl/out`
|
- **5.4 GB of stale staging on `mattellite`** (`~/gdl`, 44 GB free): `staging`
|
||||||
(2.2 GB each, from the 2026-08-17 run) and `~/gdl/staging-0820` /
|
and `out` at 2.2 GB each from 2026-08-17, `staging-0820` / `out-0820` at
|
||||||
`~/gdl/out-0820` (446 MB each, from 2026-08-20). Every file in all four was
|
446 MB each, plus `staging-full` (78 MB) and `staging-stories` (66 MB) from
|
||||||
verified present in the live archive, so they are safe to delete. 46 GB free,
|
2026-08-22. All of it was verified published, so all of it is safe to delete.
|
||||||
so there is no urgency — but nothing will clean them up on its own.
|
Staging is wiped per-run by `gdl-cron.sh`, but the `out-*` publish targets and
|
||||||
- **No daily stories run is scheduled.** Stories expire in 24h and cannot be
|
anything created by a direct `gdl-sync.py` call are never cleaned up.
|
||||||
backfilled, so this is the *only* surface where waiting loses content
|
- **Stories currently depend on someone remembering.** The daily timer exists
|
||||||
permanently. `--only stories` never seeds and costs roughly six requests for
|
but is disabled, so the one surface that cannot be backfilled has no
|
||||||
all six profiles. When scheduling it, randomise the minute and avoid the hour
|
automation behind it. Every day nobody runs `gdl-cron.sh stories` is a day
|
||||||
boundary: a job firing at exactly 09:00 daily is obviously a machine.
|
of stories gone. That is the central unresolved tension: the cadence that
|
||||||
- **`--abort 50` is opt-in and nothing uses it yet.** It is the right setting
|
protects stories is also the most machine-like pattern here.
|
||||||
for routine runs — it cut a 2151-post profile to 7 enumerated posts — but it
|
- **The scheduled modes have not been reconciled with the pacing used by
|
||||||
stops noticing **edited carousels** (test case 15), which only a full
|
hand.** The 2026-08-22 runs used `--sleep-request 12 20 --rate 500K`;
|
||||||
enumeration finds. A full-sweep cadence has not been decided; quarterly was
|
`gdl-cron.sh` still passes the defaults (6-10s, 1M). Re-enabling the timers
|
||||||
suggested and never agreed.
|
as they stand would make the automation *less* careful than the manual runs
|
||||||
|
that followed a warning.
|
||||||
|
- **`--abort 50` is opt-in.** `gdl-cron.sh full` passes it and the manual runs
|
||||||
|
used it; `sweep` deliberately does not. It stops noticing **edited
|
||||||
|
carousels** (test case 15), which only a full enumeration finds — which is
|
||||||
|
what `sweep` is for. The quarterly cadence was proposed and never agreed.
|
||||||
- **The `seeded` flags in `<db>.state.json` were hand-written**, reconstructed
|
- **The `seeded` flags in `<db>.state.json` were hand-written**, reconstructed
|
||||||
from the 2026-08-17 log rather than derived from the archive DB. They assert
|
from the 2026-08-17 log rather than derived from the archive DB. They assert
|
||||||
"the skip-archive already knows this source". If `artms.db` is ever rebuilt,
|
"the skip-archive already knows this source". If `artms.db` is ever rebuilt,
|
||||||
|
|||||||
Reference in New Issue
Block a user