Files
instaarchive-viewer/docs/gallery-dl.md
T
ergosteurandClaude Opus 5 4fee8b1dfe feat: seed gallery-dl's skip-archive so the fetcher needs no archive copy
gallery-dl skips already-held media either by file existence -- which
requires the archive mounted where it writes -- or by a sqlite
skip-archive, which requires nothing on disk. Using the latter lets the
fetch host write to local disk and rsync afterwards, avoiding tens of
thousands of small writes over CIFS and keeping a mid-sync failure from
leaving partial files on the live Resilio share.

The key is archive_prefix + archive_fmt: the literal "instagram" plus the
per-media numeric pk. Verified against a real run -- a 3-image carousel
produced 3 rows and a re-run skipped every media file.

media_id is absent from our filenames, so the DB cannot be built from
names alone, but the listing pass we already make maps every live item to
its media_id, and a file listing says which we hold. Seeding therefore
costs no extra Instagram requests and no archive content -- the listing
GET /api/archives/:name/files already serves is enough.

Measured on 0ct0ber19: 2275 live items, 2248 seeded, 27 left to fetch --
exactly the media of the two posts added since the last crawl.

The trap worth the comment it carries: posts and reels are filed under
post_shortcode, while stories and highlights use the per-item shortcode
(post_shortcode there is the containing reel's id, shared by every item).
Matching on the wrong field seeded 5 of 2275 rather than failing loudly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 21:10:37 -04:00

13 KiB
Raw Blame History

gallery-dl — a CLI replacement for JDownloader2

Status: design + verified config. scripts/gdl-sync.py is a skeleton; no profile has been migrated yet.

Everything below was measured against the live site and the real archive on 2026-08-16, not inferred from documentation.

The hard parts of fetching Instagram are pagination, cookie handling, CDN URL expiry and resumption. gallery-dl already has all of them, plus extractors that map 1:1 onto our sidecar directory layout (posts, reels, stories, highlights). Rolling our own would mean reimplementing the ban-sensitive part by hand.

The safety model — read this before changing any option

The ban vector is requests to instagram.com, not bandwidth. See docs/jdownloader.md for the history; Instaloader got this account banned by asking instagram.com a question per post.

gallery-dl has two API backends and the difference is exactly that vector:

if self.config("api") == "graphql":
    self.api = InstagramGraphqlAPI(self)   # per-post api.media() for every
else:                                      # video and every carousel
    self.api = InstagramRestAPI(self)      # <- default, listing-only

The REST backend paginates at count: 30 (feed) / page_size: 50 (clips), and those responses already carry carousel_media, image_versions2, video_versions and product_type. No per-post request. A 300-post profile costs roughly 10 requests to instagram.com.

Rules, in order of importance:

  1. "api": "rest" always. Never graphql. This is the whole ballgame.
  2. Never enable metadata-style options that trigger extra calls. If a field is not already in the listing response, it is not worth a request.
  3. Pace it. "sleep-request": [4.0, 7.0] — a randomised gap, not a fixed one. Also "sleep": [1.0, 3.0] between downloads.
  4. Cap the download rate (downloader.http.rate) so the CDN side looks like a person, not a mirror.
  5. Run from the same public IP as the browser the cookie came from. At time of writing that is mattellite (66.23.52.196); the dev workstation is a different public IP and using the cookie from there is precisely what session-hijack detection looks for.
  6. No programmatic login, ever. gallery-dl's username/password path is disabled upstream anyway; use --cookies-from-browser.

Do not add proxy rotation, fingerprint spoofing or account rotation. Throttling and request-avoidance are welcome; evasion is not.

Cookies

The logged-in Chrome on mattellite runs with a non-default profile:

--user-data-dir=/home/matt/.config/google-chrome-devtools

so the cookie flag is:

--cookies-from-browser "chrome:/home/matt/.config/google-chrome-devtools"

Plain --cookies-from-browser chrome fails with "Unable to find chrome cookies database" because it looks in ~/.config/google-chrome/.

Anonymous access is not a viable fallback: it serves lower-resolution media, caps profile pagination at 12 posts, and returns AuthRequired for stories and highlights.

Output format

The viewer's parser is the contract, not JD2's exact bytes. EXPORT_RE in src/lib/archive-patterns.ts accepts all of these, and normalises the index with parseInt, so JD2 and gallery-dl naming interoperate:

"… - CrORBIcJJbM.mp4"      -> postId=CrORBIcJJbM  index=1
"… - CrORBIcJJbM - 1.mp4"  -> postId=CrORBIcJJbM  index=1
"… - C53YPQzp7Wj - 09.jpg" -> postId=C53YPQzp7Wj  index=9

That means zero-padding and the presence/absence of - N on single-media posts are cosmetic. Don't spend effort forcing them.

Directory layout

kind directory note
posts <user>
reels <user> - reels
stories story - <user>
highlights story highlights - <user> - <title>

Force the directory with -D; never use {username} for it. A profile's reels tab returns collab reels owned by other accounts/0ct0ber19/reels/ served 6 reels owned by official_artms and 1 by chuuo3o. With {username} those would scatter into official_artms - reels/. JD2 got this right and the archive proves it: chuuo3o and official_artms filenames sit inside 0ct0ber19 - reels/.

So: owner in the filename, crawl scope in the directory.

Filenames

"filename": {
  "sidecar_shortcode and count >= 10":
      "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode} - {num:02}.{extension}",
  "sidecar_shortcode":
      "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode} - {num}.{extension}",
  "":
      "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.{extension}"
}

sidecar_shortcode is set only when the post is a carousel, so it is the carousel discriminator. Conditions are evaluated in order, first match wins (path.py:265).

Stories and highlights use the per-item {shortcode}, not {post_shortcode} (which is the reel's id, shared by every item in it):

"{date:Olocal/%Y-%m-%d}_{username} - {shortcode}.{extension}"

{date} on a story/highlight file is the per-item taken_at (instagram.py:337 prefers item["taken_at"]), verified on a 154-item highlight whose items carried distinct times while post_date stayed pinned to the reel. Highlights therefore gain real dates — today they fall back to directory mtime.

The timezone is not UTC

JD2 stamped filenames in desktop local time (US Eastern). Measured across 212 comparable posts:

model mismatches
UTC 19
UTC5 (EST) 10
UTC4 (EDT) 0
America/New_York (DST-aware) 0

{date:Olocal/%Y-%m-%d} uses the machine's local zone with per-timestamp DST awareness, which reproduces it — mattellite is America/Toronto, the same offsets. Note the trailing / must be omitted: Olocal/%Y-%m-%d/ puts the separator into the strftime format and it sanitises to an underscore, giving 2026-08-15__0ct0ber19.

If the sync ever moves to a host in another timezone, set an explicit {date:O-4/…} or the dates will silently shift for ~9% of posts.

Caption sidecars

JD2 writes one .txt per post, named without the index, containing the caption with no trailing newline, and writes nothing when the caption is empty (measured: 197 of 217 posts, 86 of 86 reels, 0 of 10 stories, 0 of 16 highlights). gallery-dl reproduces this exactly with the default "empty": false:

{ "name": "metadata", "event": "post", "mode": "custom",
  "content-format": "{description}", "extension": "txt",
  "filename": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.txt" }

"event": "post" is what makes it one file per post rather than per media file.

Metadata sidecar (new — JD2 had no equivalent)

{ "name": "metadata", "event": "post", "mode": "json",
  "filename": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.json",
  "include": ["post_shortcode","post_id","type","date","post_date","username",
              "fullname","owner_id","description","count","likes","post_url",
              "sidecar_shortcode"] }

Use include, not fieldsfields is for mode: custom and silently does nothing here, leaving audio_user blobs (including another user's profile picture URL) in the output.

The payoff is type, which is Instagram's own classification:

{ "post_shortcode": "DbdG9L9jU4m", "type": "post",  "count": 2 }   // feed video
{ "post_shortcode": "Db-lNCoib9m", "type": "reel",  "count": 1 }   // real reel

This is the product_type: "clips" signal, delivered free in the listing response. It is the authoritative answer to "is this a reel", and would let the viewer retire the lone-video heuristic in src/lib/post-tabs.ts — see "Scanner work" below.

type is only populated by listing extractors. Extracting a single /p/<shortcode>/ URL leaves it null. Sync always uses listing URLs, so this only matters when testing by hand.

Incremental sync — why the fetch host needs no copy of the archive

gallery-dl can skip already-held media two ways, and the difference decides whether the fetcher needs the archive mounted:

  • By file existence (default). Needs the destination to already contain the files, so it only works if the archive is mounted where gallery-dl writes.
  • By skip-archive (--download-archive). A sqlite DB of ids. Needs nothing on disk.

We use the second, so the fetch host can write to local disk and rsync afterwards. That avoids writing tens of thousands of small files over CIFS, and keeps a mid-sync failure from leaving partial files on the live Resilio share.

The key is archive_prefix + archive_fmt, which for this extractor is the literal instagram plus the per-media numeric pk (instagram.py:25, job.py:713-719). Verified: a 3-image carousel produced

instagram3079387627521318672
instagram3079387627521429433
instagram3079387627529716672

and a second run skipped every media file, rewriting only the idempotent .txt/.json sidecars.

Seeding. media_id is not in our filenames, so the DB cannot be built from names alone — but one listing pass (the pass we make anyway) maps every live item to its media_id, and the archive's file listing says which we already hold. No extra Instagram requests, and no archive content — a listing is enough, which GET /api/archives/:name/files already serves.

Measured on 0ct0ber19: 2275 live media items, 2248 seeded from the existing listing, 27 left to download — precisely the media of the two posts added since the last crawl.

The one trap, which silently seeds almost nothing if you get it backwards:

surface filed under why
posts, reels post_shortcode carousel children each have their own shortcode, which never appears in a filename
stories, highlights shortcode (per item) post_shortcode is the containing reel's id, shared by every item

live_key() encodes this. Matching on the wrong field seeded 5 of 2275.

Known quirks

  • count is not the emitted file count. For 135 of 214 posts it was exactly one higher than the number of files written. This makes the count >= 10 padding condition mis-pad a handful of 9-item posts (10 of 214 measured). Since the parser normalises the index, this is cosmetic — but it means a re-fetch over an existing JD2 tree writes - 01.jpg beside an existing - 1.jpg.
  • Carousels get edited. Two posts had a different media count live than on disk. Padding width follows the count at download time, so a grown carousel produces mixed widths — the archive already contains one such post from JD2.
  • Highlights already have two naming styles on disk, and every undated file has a dated twin. The scanner dedupes by index so they render once; it is wasted disk, not a display bug.

Scanner work (not done yet)

useArchiveScanner currently treats any .json in the tree as a possible manifest. Adding gallery-dl sidecars needs it to distinguish three things:

  1. Instagram export manifests (posts_1.json) — existing path.
  2. Instaloader .json.xz — existing path, GraphQL node shape.
  3. gallery-dl .json — new, flat shape, identified by having post_shortcode + type at the top level.

Once (3) is read, source/isStory and the reel flag should come from type rather than from the directory and the lone-video heuristic.

Test cases

Real subjects, all present in the archive today. See scripts/gdl-sync.py --selftest for the harness.

# case shortcode expected
1 single image CwcXnQhOqFG one .jpg, no index
2 single feed video DbdG9L9jU4m one .mp4, type: post
3 carousel, images only Cq8LrxSJAJE - 1 … - 3
4 carousel, image + video CtohvHxLnWO - 1.jpg … - 4.mp4, no .txt
5 carousel of exactly 9 Cv2Hb_brx_N 1-digit index
6 carousel of 10+ CzM8Uf6B6H_ 2-digit index - 01 … - 10
7 reel shown on the posts grid C8FHM6EJl15 in <user>, type: reel
8 reel on the reels tab Db-lNCoib9m in <user> - reels, type: reel
9 collab reel (other owner) DYcZOb0h6Sv dir 0ct0ber19 - reels, filename chuuo3o
10 story live only story - <user>, per-item shortcode + date
11 story highlight C-IImhvpFuk story highlights - <user> - <title>
12 highlight, unicode title Drawheeing trailing U+2800 preserved in dirname
13 empty caption CrdsY5CrSsO media written, .txt absent
14 deleted post C0TgI7sphfZ on disk, absent live — must not be removed
15 edited carousel C7zG7-jJMlq 18 on disk, 8 live — must not be removed
16 pinned posts 0ct0ber19 3 pinned, returned out of date order
17 profile avatar 0ct0ber19.jpg base dir, undated

Cases 1416 are reconciliation, not naming: a sync must never delete, since the archive deliberately outlives Instagram.

Not covered, decide before relying on them: the /reposts/ tab (0ct0ber19 has one) and /tagged/. Neither is fetched today.