# gallery-dl — a CLI replacement for JDownloader2 Status: **design + verified config.** `scripts/gdl-sync.py` is a skeleton; no profile has been migrated yet. Everything below was measured against the live site and the real archive on 2026-08-16, not inferred from documentation. ## Why gallery-dl and not a hand-rolled script The hard parts of fetching Instagram are pagination, cookie handling, CDN URL expiry and resumption. gallery-dl already has all of them, plus extractors that map 1:1 onto our sidecar directory layout (`posts`, `reels`, `stories`, `highlights`). Rolling our own would mean reimplementing the ban-sensitive part by hand. ## The safety model — read this before changing any option The ban vector is **requests to `instagram.com`**, not bandwidth. See `docs/jdownloader.md` for the history; Instaloader got this account banned by asking `instagram.com` a question *per post*. gallery-dl has two API backends and the difference is exactly that vector: ```python if self.config("api") == "graphql": self.api = InstagramGraphqlAPI(self) # per-post api.media() for every else: # video and every carousel self.api = InstagramRestAPI(self) # <- default, listing-only ``` The REST backend paginates at `count: 30` (feed) / `page_size: 50` (clips), and those responses already carry `carousel_media`, `image_versions2`, `video_versions` and `product_type`. **No per-post request.** A 300-post profile costs roughly 10 requests to `instagram.com`. Rules, in order of importance: 1. **`"api": "rest"` always.** Never `graphql`. This is the whole ballgame. 2. **Never enable `metadata`-style options that trigger extra calls.** If a field is not already in the listing response, it is not worth a request. 3. **Pace it.** `"sleep-request": [4.0, 7.0]` — a randomised gap, not a fixed one. Also `"sleep": [1.0, 3.0]` between downloads. 4. **Cap the download rate** (`downloader.http.rate`) so the CDN side looks like a person, not a mirror. 5. **Run from the same public IP as the browser the cookie came from.** At time of writing that is `mattellite` (`66.23.52.196`); the dev workstation is a *different* public IP and using the cookie from there is precisely what session-hijack detection looks for. 6. **No programmatic login, ever.** gallery-dl's username/password path is disabled upstream anyway; use `--cookies-from-browser`. Do not add proxy rotation, fingerprint spoofing or account rotation. Throttling and request-avoidance are welcome; evasion is not. ### Cookies The logged-in Chrome on `mattellite` runs with a non-default profile: ``` --user-data-dir=/home/matt/.config/google-chrome-devtools ``` so the cookie flag is: ``` --cookies-from-browser "chrome:/home/matt/.config/google-chrome-devtools" ``` Plain `--cookies-from-browser chrome` fails with "Unable to find chrome cookies database" because it looks in `~/.config/google-chrome/`. Anonymous access is **not** a viable fallback: it serves lower-resolution media, caps profile pagination at 12 posts, and returns `AuthRequired` for stories and highlights. ## Output format The viewer's parser is the contract, not JD2's exact bytes. `EXPORT_RE` in `src/lib/archive-patterns.ts` accepts all of these, and normalises the index with `parseInt`, so **JD2 and gallery-dl naming interoperate**: ``` "… - CrORBIcJJbM.mp4" -> postId=CrORBIcJJbM index=1 "… - CrORBIcJJbM - 1.mp4" -> postId=CrORBIcJJbM index=1 "… - C53YPQzp7Wj - 09.jpg" -> postId=C53YPQzp7Wj index=9 ``` That means zero-padding and the presence/absence of ` - N` on single-media posts are cosmetic. Don't spend effort forcing them. ### Directory layout | kind | directory | note | |---|---|---| | posts | `` | | | reels | ` - reels` | | | stories | `story - ` | | | highlights | `story highlights - - ` | | **Force the directory with `-D`; never use `{username}` for it.** A profile's reels tab returns *collab reels owned by other accounts* — `/0ct0ber19/reels/` served 6 reels owned by `official_artms` and 1 by `chuuo3o`. With `{username}` those would scatter into `official_artms - reels/`. JD2 got this right and the archive proves it: `chuuo3o` and `official_artms` filenames sit inside `0ct0ber19 - reels/`. So: **owner in the filename, crawl scope in the directory.** ### Filenames ```jsonc "filename": { "sidecar_shortcode and count >= 10": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode} - {num:02}.{extension}", "sidecar_shortcode": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode} - {num}.{extension}", "": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.{extension}" } ``` `sidecar_shortcode` is set only when the post is a carousel, so it is the carousel discriminator. Conditions are evaluated in order, first match wins (`path.py:265`). Stories and highlights use the per-item `{shortcode}`, not `{post_shortcode}` (which is the *reel's* id, shared by every item in it): ``` "{date:Olocal/%Y-%m-%d}_{username} - {shortcode}.{extension}" ``` `{date}` on a story/highlight file is the **per-item** `taken_at` (`instagram.py:337` prefers `item["taken_at"]`), verified on a 154-item highlight whose items carried distinct times while `post_date` stayed pinned to the reel. Highlights therefore gain real dates — today they fall back to directory mtime. ### The timezone is not UTC JD2 stamped filenames in **desktop local time (US Eastern)**. Measured across 212 comparable posts: | model | mismatches | |---|---:| | UTC | 19 | | UTC−5 (EST) | 10 | | UTC−4 (EDT) | **0** | | America/New_York (DST-aware) | **0** | `{date:Olocal/%Y-%m-%d}` uses the machine's local zone with per-timestamp DST awareness, which reproduces it — `mattellite` is `America/Toronto`, the same offsets. Note the **trailing `/` must be omitted**: `Olocal/%Y-%m-%d/` puts the separator into the strftime format and it sanitises to an underscore, giving `2026-08-15__0ct0ber19`. If the sync ever moves to a host in another timezone, set an explicit `{date:O-4/…}` or the dates will silently shift for ~9% of posts. ### Caption sidecars JD2 writes one `.txt` per post, named without the index, containing the caption with **no trailing newline**, and writes nothing when the caption is empty (measured: 197 of 217 posts, 86 of 86 reels, 0 of 10 stories, 0 of 16 highlights). gallery-dl reproduces this exactly with the default `"empty": false`: ```jsonc { "name": "metadata", "event": "post", "mode": "custom", "content-format": "{description}", "extension": "txt", "filename": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.txt" } ``` `"event": "post"` is what makes it one file per post rather than per media file. ### Metadata sidecar (new — JD2 had no equivalent) ```jsonc { "name": "metadata", "event": "post", "mode": "json", "filename": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.json", "include": ["post_shortcode","post_id","type","date","post_date","username", "fullname","owner_id","description","count","likes","post_url", "sidecar_shortcode"] } ``` Use **`include`**, not `fields` — `fields` is for `mode: custom` and silently does nothing here, leaving `audio_user` blobs (including another user's profile picture URL) in the output. The payoff is `type`, which is Instagram's own classification: ```json { "post_shortcode": "DbdG9L9jU4m", "type": "post", "count": 2 } // feed video { "post_shortcode": "Db-lNCoib9m", "type": "reel", "count": 1 } // real reel ``` This is the `product_type: "clips"` signal, delivered free in the listing response. It is the authoritative answer to "is this a reel", and would let the viewer retire the lone-video heuristic in `src/lib/post-tabs.ts` — see "Scanner work" below. **`type` is only populated by listing extractors.** Extracting a single `/p/<shortcode>/` URL leaves it `null`. Sync always uses listing URLs, so this only matters when testing by hand. ## Incremental sync — why the fetch host needs no copy of the archive gallery-dl can skip already-held media two ways, and the difference decides whether the fetcher needs the archive mounted: - **By file existence** (default). Needs the destination to already contain the files, so it only works if the archive is mounted where gallery-dl writes. - **By skip-archive** (`--download-archive`). A sqlite DB of ids. Needs nothing on disk. We use the second, so the fetch host can write to **local disk and rsync afterwards**. That avoids writing tens of thousands of small files over CIFS, and keeps a mid-sync failure from leaving partial files on the live Resilio share. The key is `archive_prefix + archive_fmt`, which for this extractor is the literal `instagram` plus the per-media numeric pk (`instagram.py:25`, `job.py:713-719`). Verified: a 3-image carousel produced ``` instagram3079387627521318672 instagram3079387627521429433 instagram3079387627529716672 ``` and a second run skipped every media file, rewriting only the idempotent `.txt`/`.json` sidecars. **Seeding.** `media_id` is not in our filenames, so the DB cannot be built from names alone — but one listing pass (the pass we make anyway) maps every live item to its `media_id`, and the archive's *file listing* says which we already hold. No extra Instagram requests, and no archive content — a listing is enough, which `GET /api/archives/:name/files` already serves. Measured on `0ct0ber19`: 2275 live media items, 2248 seeded from the existing listing, **27 left to download** — precisely the media of the two posts added since the last crawl. The one trap, which silently seeds almost nothing if you get it backwards: | surface | filed under | why | |---|---|---| | posts, reels | `post_shortcode` | carousel children each have their own `shortcode`, which never appears in a filename | | stories, highlights | `shortcode` (per item) | `post_shortcode` is the containing reel's id, shared by every item | `live_key()` encodes this. Matching on the wrong field seeded 5 of 2275. ## Publishing The fetch host stages to local disk and rsyncs afterwards. `rsync --ignore-existing` is not an optimisation but the safety property: the archive deliberately outlives Instagram, so publishing must only ever **add**. No `--delete`, and nothing already present is overwritten — including sidecars, which are rewritten every run and would otherwise churn the synced share. Publishing happens once at the end of a run, so a profile that fails midway never reaches the archive half-written. ## Status In use. `withaseul` has been fetched and published to the live archive — 322 files added (74 media, 241 `.json`, 7 `.txt`), nothing overwritten or deleted. Of the 74 new media files, **zero** duplicated media already held under a different name, which is the check that says JD2 and gallery-dl naming really do converge. Published files land owned by the SSH user rather than `rslsync`. The viewer reads them fine (world-readable), but Resilio does not own what it syncs; worth a `chown` if that ever matters. The profiles to fetch live in `artms_account_links.txt` at the archive root, passed with `--urls-file`. ## Verified run `withaseul`, all four surfaces, staged locally and published to a scratch directory before the live publish above: ``` ==> withaseul / posts seeded 915 of 984 live items ==> withaseul / reels seeded 28 of 34 live items ==> withaseul / stories no results (none active) ==> withaseul / highlights no results ``` Output landed correctly, including the collab-reel case — `withaseul - reels` contains 53 files owned by `withaseul`, 10 by `cher_ryppo`, 3 by `0ct0ber19` and 2 by `official_artms`, all with the owner in the filename and the crawl scope as the directory. ### The CDN rate-limits, and the first run tripped it At `rate: 3M` with `sleep: [1.0, 3.0]`, `scontent-*.cdninstagram.com` returned **`429 Too Many Requests`** and two videos were lost (gallery-dl retried, then gave up with exit 4). This is the *tolerant* surface complaining, which is a clear signal the pacing was too aggressive. Defaults are now: | option | value | |---|---| | `--rate` | `1M` | | `--sleep-request` | 6–10 s | | `--sleep` | 3–6 s | | `sleep-429` | 120 s | | `retries` (extractor and downloader) | 8 | Re-running with those recovered both videos and produced **0 failures and 0 429s**. Do not raise them for speed; an archive sync has no deadline. ### yt-dlp is worth installing Without it, gallery-dl logs `Cannot import yt-dlp or youtube-dl` and falls back to a progressive URL for DASH videos. The fallback mostly works but is what the 429s hit hardest. `pipx install yt-dlp` on the fetch host. ## Known quirks - **`count` is not the emitted file count.** For 135 of 214 posts it was exactly one higher than the number of files written. This makes the `count >= 10` padding condition mis-pad a handful of 9-item posts (10 of 214 measured). Since the parser normalises the index, this is cosmetic — but it means a re-fetch over an existing JD2 tree writes `- 01.jpg` beside an existing `- 1.jpg`. - **Carousels get edited.** Two posts had a different media count live than on disk. Padding width follows the count *at download time*, so a grown carousel produces mixed widths — the archive already contains one such post from JD2. - **Highlights already have two naming styles on disk**, and every undated file has a dated twin. The scanner dedupes by index so they render once; it is wasted disk, not a display bug. ## Scanner work (not done yet) `useArchiveScanner` currently treats any `.json` in the tree as a possible manifest. Adding gallery-dl sidecars needs it to distinguish three things: 1. Instagram export manifests (`posts_1.json`) — existing path. 2. Instaloader `.json.xz` — existing path, GraphQL node shape. 3. gallery-dl `.json` — new, flat shape, identified by having `post_shortcode` + `type` at the top level. Once (3) is read, `source`/`isStory` and the reel flag should come from `type` rather than from the directory and the lone-video heuristic. ## Test cases Real subjects, all present in the archive today. See `scripts/gdl-sync.py --selftest` for the harness. | # | case | shortcode | expected | |---|---|---|---| | 1 | single image | `CwcXnQhOqFG` | one `.jpg`, no index | | 2 | single feed video | `DbdG9L9jU4m` | one `.mp4`, `type: post` | | 3 | carousel, images only | `Cq8LrxSJAJE` | `- 1 … - 3` | | 4 | carousel, image + video | `CtohvHxLnWO` | `- 1.jpg … - 4.mp4`, **no `.txt`** | | 5 | carousel of exactly 9 | `Cv2Hb_brx_N` | 1-digit index | | 6 | carousel of 10+ | `CzM8Uf6B6H_` | 2-digit index `- 01 … - 10` | | 7 | reel shown on the posts grid | `C8FHM6EJl15` | in `<user>`, `type: reel` | | 8 | reel on the reels tab | `Db-lNCoib9m` | in `<user> - reels`, `type: reel` | | 9 | collab reel (other owner) | `DYcZOb0h6Sv` | dir `0ct0ber19 - reels`, filename `chuuo3o` | | 10 | story | live only | `story - <user>`, per-item shortcode + date | | 11 | story highlight | `C-IImhvpFuk` | `story highlights - <user> - <title>` | | 12 | highlight, unicode title | `Drawheeing⠀` | trailing U+2800 preserved in dirname | | 13 | empty caption | `CrdsY5CrSsO` | media written, `.txt` absent | | 14 | deleted post | `C0TgI7sphfZ` | on disk, absent live — must not be removed | | 15 | edited carousel | `C7zG7-jJMlq` | 18 on disk, 8 live — must not be removed | | 16 | pinned posts | `0ct0ber19` | 3 pinned, returned out of date order | | 17 | profile avatar | `0ct0ber19.jpg` | base dir, undated | Cases 14–16 are reconciliation, not naming: **a sync must never delete**, since the archive deliberately outlives Instagram. Not covered, decide before relying on them: the `/reposts/` tab (`0ct0ber19` has one) and `/tagged/`. Neither is fetched today.