Wires up the staging -> rsync step and replaces --archives with --index (a listing source: local path or the viewer's API), --staging and --publish, so the fetch host needs no copy of the archive. rsync runs --ignore-existing with no --delete. That is a safety property rather than an optimisation: the archive deliberately outlives Instagram, so publishing must only ever add. It runs once at the end so a profile that fails midway never reaches the archive half-written. Exercised end-to-end against withaseul across all four surfaces, publishing to a scratch directory. Seeding worked as designed (915 of 984 post items and 28 of 34 reel items already held), stories and highlights returned no results cleanly, and the collab-reel case landed correctly: "withaseul - reels" holds files owned by cher_ryppo, 0ct0ber19 and official_artms, each with the owner in the filename and the crawl scope as the directory. The first run drew '429 Too Many Requests' from the CDN at 3M with 1-3s sleeps and lost two videos. That is the tolerant surface complaining, so the defaults are now 1M, 6-10s between requests, 3-6s between downloads, sleep-429 of 120s and 8 retries. Re-running recovered both videos with zero failures and zero 429s. Installing yt-dlp on the fetch host also matters: without it DASH videos fall back to a progressive URL, which is what the rate limiting hit hardest. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
15 KiB
gallery-dl — a CLI replacement for JDownloader2
Status: design + verified config. scripts/gdl-sync.py is a skeleton; no
profile has been migrated yet.
Everything below was measured against the live site and the real archive on 2026-08-16, not inferred from documentation.
Why gallery-dl and not a hand-rolled script
The hard parts of fetching Instagram are pagination, cookie handling, CDN URL
expiry and resumption. gallery-dl already has all of them, plus extractors that
map 1:1 onto our sidecar directory layout (posts, reels, stories,
highlights). Rolling our own would mean reimplementing the ban-sensitive part
by hand.
The safety model — read this before changing any option
The ban vector is requests to instagram.com, not bandwidth. See
docs/jdownloader.md for the history; Instaloader got this account banned by
asking instagram.com a question per post.
gallery-dl has two API backends and the difference is exactly that vector:
if self.config("api") == "graphql":
self.api = InstagramGraphqlAPI(self) # per-post api.media() for every
else: # video and every carousel
self.api = InstagramRestAPI(self) # <- default, listing-only
The REST backend paginates at count: 30 (feed) / page_size: 50 (clips), and
those responses already carry carousel_media, image_versions2,
video_versions and product_type. No per-post request. A 300-post
profile costs roughly 10 requests to instagram.com.
Rules, in order of importance:
"api": "rest"always. Nevergraphql. This is the whole ballgame.- Never enable
metadata-style options that trigger extra calls. If a field is not already in the listing response, it is not worth a request. - Pace it.
"sleep-request": [4.0, 7.0]— a randomised gap, not a fixed one. Also"sleep": [1.0, 3.0]between downloads. - Cap the download rate (
downloader.http.rate) so the CDN side looks like a person, not a mirror. - Run from the same public IP as the browser the cookie came from. At time
of writing that is
mattellite(66.23.52.196); the dev workstation is a different public IP and using the cookie from there is precisely what session-hijack detection looks for. - No programmatic login, ever. gallery-dl's username/password path is
disabled upstream anyway; use
--cookies-from-browser.
Do not add proxy rotation, fingerprint spoofing or account rotation. Throttling and request-avoidance are welcome; evasion is not.
Cookies
The logged-in Chrome on mattellite runs with a non-default profile:
--user-data-dir=/home/matt/.config/google-chrome-devtools
so the cookie flag is:
--cookies-from-browser "chrome:/home/matt/.config/google-chrome-devtools"
Plain --cookies-from-browser chrome fails with "Unable to find chrome cookies
database" because it looks in ~/.config/google-chrome/.
Anonymous access is not a viable fallback: it serves lower-resolution media,
caps profile pagination at 12 posts, and returns AuthRequired for stories and
highlights.
Output format
The viewer's parser is the contract, not JD2's exact bytes. EXPORT_RE in
src/lib/archive-patterns.ts accepts all of these, and normalises the index
with parseInt, so JD2 and gallery-dl naming interoperate:
"… - CrORBIcJJbM.mp4" -> postId=CrORBIcJJbM index=1
"… - CrORBIcJJbM - 1.mp4" -> postId=CrORBIcJJbM index=1
"… - C53YPQzp7Wj - 09.jpg" -> postId=C53YPQzp7Wj index=9
That means zero-padding and the presence/absence of - N on single-media posts
are cosmetic. Don't spend effort forcing them.
Directory layout
| kind | directory | note |
|---|---|---|
| posts | <user> |
|
| reels | <user> - reels |
|
| stories | story - <user> |
|
| highlights | story highlights - <user> - <title> |
Force the directory with -D; never use {username} for it. A profile's
reels tab returns collab reels owned by other accounts — /0ct0ber19/reels/
served 6 reels owned by official_artms and 1 by chuuo3o. With
{username} those would scatter into official_artms - reels/. JD2 got this
right and the archive proves it: chuuo3o and official_artms filenames sit
inside 0ct0ber19 - reels/.
So: owner in the filename, crawl scope in the directory.
Filenames
"filename": {
"sidecar_shortcode and count >= 10":
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode} - {num:02}.{extension}",
"sidecar_shortcode":
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode} - {num}.{extension}",
"":
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.{extension}"
}
sidecar_shortcode is set only when the post is a carousel, so it is the
carousel discriminator. Conditions are evaluated in order, first match wins
(path.py:265).
Stories and highlights use the per-item {shortcode}, not {post_shortcode}
(which is the reel's id, shared by every item in it):
"{date:Olocal/%Y-%m-%d}_{username} - {shortcode}.{extension}"
{date} on a story/highlight file is the per-item taken_at
(instagram.py:337 prefers item["taken_at"]), verified on a 154-item
highlight whose items carried distinct times while post_date stayed pinned to
the reel. Highlights therefore gain real dates — today they fall back to
directory mtime.
The timezone is not UTC
JD2 stamped filenames in desktop local time (US Eastern). Measured across 212 comparable posts:
| model | mismatches |
|---|---|
| UTC | 19 |
| UTC−5 (EST) | 10 |
| UTC−4 (EDT) | 0 |
| America/New_York (DST-aware) | 0 |
{date:Olocal/%Y-%m-%d} uses the machine's local zone with per-timestamp DST
awareness, which reproduces it — mattellite is America/Toronto, the same
offsets. Note the trailing / must be omitted: Olocal/%Y-%m-%d/ puts the
separator into the strftime format and it sanitises to an underscore, giving
2026-08-15__0ct0ber19.
If the sync ever moves to a host in another timezone, set an explicit
{date:O-4/…} or the dates will silently shift for ~9% of posts.
Caption sidecars
JD2 writes one .txt per post, named without the index, containing the caption
with no trailing newline, and writes nothing when the caption is empty
(measured: 197 of 217 posts, 86 of 86 reels, 0 of 10 stories, 0 of 16
highlights). gallery-dl reproduces this exactly with the default
"empty": false:
{ "name": "metadata", "event": "post", "mode": "custom",
"content-format": "{description}", "extension": "txt",
"filename": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.txt" }
"event": "post" is what makes it one file per post rather than per media file.
Metadata sidecar (new — JD2 had no equivalent)
{ "name": "metadata", "event": "post", "mode": "json",
"filename": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.json",
"include": ["post_shortcode","post_id","type","date","post_date","username",
"fullname","owner_id","description","count","likes","post_url",
"sidecar_shortcode"] }
Use include, not fields — fields is for mode: custom and silently
does nothing here, leaving audio_user blobs (including another user's profile
picture URL) in the output.
The payoff is type, which is Instagram's own classification:
{ "post_shortcode": "DbdG9L9jU4m", "type": "post", "count": 2 } // feed video
{ "post_shortcode": "Db-lNCoib9m", "type": "reel", "count": 1 } // real reel
This is the product_type: "clips" signal, delivered free in the listing
response. It is the authoritative answer to "is this a reel", and would let the
viewer retire the lone-video heuristic in src/lib/post-tabs.ts — see
"Scanner work" below.
type is only populated by listing extractors. Extracting a single
/p/<shortcode>/ URL leaves it null. Sync always uses listing URLs, so this
only matters when testing by hand.
Incremental sync — why the fetch host needs no copy of the archive
gallery-dl can skip already-held media two ways, and the difference decides whether the fetcher needs the archive mounted:
- By file existence (default). Needs the destination to already contain the files, so it only works if the archive is mounted where gallery-dl writes.
- By skip-archive (
--download-archive). A sqlite DB of ids. Needs nothing on disk.
We use the second, so the fetch host can write to local disk and rsync afterwards. That avoids writing tens of thousands of small files over CIFS, and keeps a mid-sync failure from leaving partial files on the live Resilio share.
The key is archive_prefix + archive_fmt, which for this extractor is the
literal instagram plus the per-media numeric pk (instagram.py:25,
job.py:713-719). Verified: a 3-image carousel produced
instagram3079387627521318672
instagram3079387627521429433
instagram3079387627529716672
and a second run skipped every media file, rewriting only the idempotent
.txt/.json sidecars.
Seeding. media_id is not in our filenames, so the DB cannot be built from
names alone — but one listing pass (the pass we make anyway) maps every live
item to its media_id, and the archive's file listing says which we already
hold. No extra Instagram requests, and no archive content — a listing is
enough, which GET /api/archives/:name/files already serves.
Measured on 0ct0ber19: 2275 live media items, 2248 seeded from the existing
listing, 27 left to download — precisely the media of the two posts added
since the last crawl.
The one trap, which silently seeds almost nothing if you get it backwards:
| surface | filed under | why |
|---|---|---|
| posts, reels | post_shortcode |
carousel children each have their own shortcode, which never appears in a filename |
| stories, highlights | shortcode (per item) |
post_shortcode is the containing reel's id, shared by every item |
live_key() encodes this. Matching on the wrong field seeded 5 of 2275.
Publishing
The fetch host stages to local disk and rsyncs afterwards. rsync --ignore-existing is not an optimisation but the safety property: the archive
deliberately outlives Instagram, so publishing must only ever add. No
--delete, and nothing already present is overwritten — including sidecars,
which are rewritten every run and would otherwise churn the synced share.
Publishing happens once at the end of a run, so a profile that fails midway never reaches the archive half-written.
Verified run
withaseul, all four surfaces, staged locally and published to a scratch
directory (never the live archive):
==> withaseul / posts seeded 915 of 984 live items
==> withaseul / reels seeded 28 of 34 live items
==> withaseul / stories no results (none active)
==> withaseul / highlights no results
Output landed correctly, including the collab-reel case — withaseul - reels
contains 53 files owned by withaseul, 10 by cher_ryppo, 3 by 0ct0ber19
and 2 by official_artms, all with the owner in the filename and the crawl
scope as the directory.
The CDN rate-limits, and the first run tripped it
At rate: 3M with sleep: [1.0, 3.0], scontent-*.cdninstagram.com returned
429 Too Many Requests and two videos were lost (gallery-dl retried, then
gave up with exit 4). This is the tolerant surface complaining, which is a
clear signal the pacing was too aggressive.
Defaults are now:
| option | value |
|---|---|
--rate |
1M |
--sleep-request |
6–10 s |
--sleep |
3–6 s |
sleep-429 |
120 s |
retries (extractor and downloader) |
8 |
Re-running with those recovered both videos and produced 0 failures and 0 429s. Do not raise them for speed; an archive sync has no deadline.
yt-dlp is worth installing
Without it, gallery-dl logs Cannot import yt-dlp or youtube-dl and falls back
to a progressive URL for DASH videos. The fallback mostly works but is what the
429s hit hardest. pipx install yt-dlp on the fetch host.
Known quirks
countis not the emitted file count. For 135 of 214 posts it was exactly one higher than the number of files written. This makes thecount >= 10padding condition mis-pad a handful of 9-item posts (10 of 214 measured). Since the parser normalises the index, this is cosmetic — but it means a re-fetch over an existing JD2 tree writes- 01.jpgbeside an existing- 1.jpg.- Carousels get edited. Two posts had a different media count live than on disk. Padding width follows the count at download time, so a grown carousel produces mixed widths — the archive already contains one such post from JD2.
- Highlights already have two naming styles on disk, and every undated file has a dated twin. The scanner dedupes by index so they render once; it is wasted disk, not a display bug.
Scanner work (not done yet)
useArchiveScanner currently treats any .json in the tree as a possible
manifest. Adding gallery-dl sidecars needs it to distinguish three things:
- Instagram export manifests (
posts_1.json) — existing path. - Instaloader
.json.xz— existing path, GraphQL node shape. - gallery-dl
.json— new, flat shape, identified by havingpost_shortcode+typeat the top level.
Once (3) is read, source/isStory and the reel flag should come from type
rather than from the directory and the lone-video heuristic.
Test cases
Real subjects, all present in the archive today. See
scripts/gdl-sync.py --selftest for the harness.
| # | case | shortcode | expected |
|---|---|---|---|
| 1 | single image | CwcXnQhOqFG |
one .jpg, no index |
| 2 | single feed video | DbdG9L9jU4m |
one .mp4, type: post |
| 3 | carousel, images only | Cq8LrxSJAJE |
- 1 … - 3 |
| 4 | carousel, image + video | CtohvHxLnWO |
- 1.jpg … - 4.mp4, no .txt |
| 5 | carousel of exactly 9 | Cv2Hb_brx_N |
1-digit index |
| 6 | carousel of 10+ | CzM8Uf6B6H_ |
2-digit index - 01 … - 10 |
| 7 | reel shown on the posts grid | C8FHM6EJl15 |
in <user>, type: reel |
| 8 | reel on the reels tab | Db-lNCoib9m |
in <user> - reels, type: reel |
| 9 | collab reel (other owner) | DYcZOb0h6Sv |
dir 0ct0ber19 - reels, filename chuuo3o |
| 10 | story | live only | story - <user>, per-item shortcode + date |
| 11 | story highlight | C-IImhvpFuk |
story highlights - <user> - <title> |
| 12 | highlight, unicode title | Drawheeing⠀ |
trailing U+2800 preserved in dirname |
| 13 | empty caption | CrdsY5CrSsO |
media written, .txt absent |
| 14 | deleted post | C0TgI7sphfZ |
on disk, absent live — must not be removed |
| 15 | edited carousel | C7zG7-jJMlq |
18 on disk, 8 live — must not be removed |
| 16 | pinned posts | 0ct0ber19 |
3 pinned, returned out of date order |
| 17 | profile avatar | 0ct0ber19.jpg |
base dir, undated |
Cases 14–16 are reconciliation, not naming: a sync must never delete, since the archive deliberately outlives Instagram.
Not covered, decide before relying on them: the /reposts/ tab (0ct0ber19
has one) and /tagged/. Neither is fetched today.