The account was suspended on 2026-08-17 for "spam", during the session
that built this tooling. Both fetching docs were confidently wrong about
what the risk was, so both now carry the correction.
docs/jdownloader.md said the ban vector is instagram.com requests, which
is right, and implied that meant per-post metadata fetching, which is
only part of it. The suspension came from read-only verification:
automated browser scrolling to enumerate profile grids (~18 paginated
loads per profile, done twice on one after a selector bug), repeated
--simulate and -j passes over the same profiles, per-post /p/ fetches
while testing filename formats, and an aborted sync that re-ran every
listing pass before dying. None of that produced a file, and together it
rivalled the real sync for request count.
The rules that follow are in docs/gallery-dl.md: verify against the
archive rather than the live site, count read-only work against the same
budget, treat the first CDN 429 as the end of the session rather than a
pacing knob, and cache probe_live so a restart does not re-enumerate
everything. The warning order was CDN 429, then 400 on the highlights
tray, then suspension.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Script comments should describe the script. Where a run has got to is
project state, so it belongs in docs/gallery-dl.md, which now records the
live withaseul publish, the file-ownership caveat and where the profile
list lives.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Wires up the staging -> rsync step and replaces --archives with --index
(a listing source: local path or the viewer's API), --staging and
--publish, so the fetch host needs no copy of the archive.
rsync runs --ignore-existing with no --delete. That is a safety property
rather than an optimisation: the archive deliberately outlives Instagram,
so publishing must only ever add. It runs once at the end so a profile
that fails midway never reaches the archive half-written.
Exercised end-to-end against withaseul across all four surfaces,
publishing to a scratch directory. Seeding worked as designed (915 of 984
post items and 28 of 34 reel items already held), stories and highlights
returned no results cleanly, and the collab-reel case landed correctly:
"withaseul - reels" holds files owned by cher_ryppo, 0ct0ber19 and
official_artms, each with the owner in the filename and the crawl scope
as the directory.
The first run drew '429 Too Many Requests' from the CDN at 3M with 1-3s
sleeps and lost two videos. That is the tolerant surface complaining, so
the defaults are now 1M, 6-10s between requests, 3-6s between downloads,
sleep-429 of 120s and 8 retries. Re-running recovered both videos with
zero failures and zero 429s. Installing yt-dlp on the fetch host also
matters: without it DASH videos fall back to a progressive URL, which is
what the rate limiting hit hardest.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
gallery-dl skips already-held media either by file existence -- which
requires the archive mounted where it writes -- or by a sqlite
skip-archive, which requires nothing on disk. Using the latter lets the
fetch host write to local disk and rsync afterwards, avoiding tens of
thousands of small writes over CIFS and keeping a mid-sync failure from
leaving partial files on the live Resilio share.
The key is archive_prefix + archive_fmt: the literal "instagram" plus the
per-media numeric pk. Verified against a real run -- a 3-image carousel
produced 3 rows and a re-run skipped every media file.
media_id is absent from our filenames, so the DB cannot be built from
names alone, but the listing pass we already make maps every live item to
its media_id, and a file listing says which we hold. Seeding therefore
costs no extra Instagram requests and no archive content -- the listing
GET /api/archives/:name/files already serves is enough.
Measured on 0ct0ber19: 2275 live items, 2248 seeded, 27 left to fetch --
exactly the media of the two posts added since the last crawl.
The trap worth the comment it carries: posts and reels are filed under
post_shortcode, while stories and highlights use the per-item shortcode
(post_shortcode there is the containing reel's id, shared by every item).
Matching on the wrong field seeded 5 of 2275 rather than failing loudly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every claim in docs/gallery-dl.md was measured against the live site and
the archive rather than taken from documentation, because two of the
assumptions turned out to be wrong.
The safety model is the reason the config looks the way it does.
gallery-dl has two API backends: the graphql one issues a request PER
POST for every video and carousel -- the pattern that got this account
banned via Instaloader -- while the default rest one paginates listings
at 30-50 items and carries carousel_media, video_versions and
product_type inline. A 300-post profile costs ~10 requests.
Findings worth recording:
- JD2 stamped filenames in desktop LOCAL time (US Eastern), not UTC.
Across 212 comparable posts: UTC 19 mismatches, UTC-5 10, UTC-4 zero.
{date:Olocal/%Y-%m-%d} reproduces it; the trailing separator must be
omitted or it lands in the strftime format.
- A profile's reels tab returns collab reels owned by OTHER accounts, so
the directory must be forced with -D. JD2 did the same: chuuo3o and
official_artms filenames sit inside "0ct0ber19 - reels".
- Stories and highlights need per-item {shortcode}; {post_shortcode} is
the reel's id and is shared by every item. {date} is per-item, verified
on a 154-item highlight with distinct times.
- gallery-dl reproduces JD2's caption .txt exactly, including writing
nothing for an empty caption and omitting the trailing newline.
- The json sidecar needs `include`, not `fields`; `fields` silently does
nothing in mode:json and leaks audio_user blobs. It yields `type`
(post/reel) -- Instagram's own flag, which can retire the lone-video
heuristic once the scanner reads it.
Naming differences between the two tools are cosmetic: EXPORT_RE already
makes the index optional and parseInt normalises zero-padding, so a mixed
archive parses identically. Tests pin that down.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two failures in this session were expensive because nothing recorded them:
- The CSP must keep 'wasm-unsafe-eval' and connect-src data:, because the xz
decompressor for Instaloader sidecars is WebAssembly embedded as a data: URL.
Removing either breaks decoding with a bare "Failed to fetch" and no stack,
and the visible symptom is silent metadata loss rather than an error.
- The service worker precaches index.html with its headers, so a server-only
change never reaches installed clients. The version compiled into the client
is what forces the precache to turn over each release; it is load-bearing,
not decoration.
Also notes that the Vite dev server sends none of these headers, so CSP and PWA
behaviour must be verified against a built dist/ served by server.js, and that
the live archive root is now the archives/ subdirectory.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
Covers why fetching goes through JDownloader rather than Instaloader (the
instagram.com vs CDN split, and what the metadata gap actually costs), the
settings that matter, cookie handling, the two-URL workflow, jd2-sync usage,
the expected on-disk layout, and what to do when something breaks.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7