gallery-dl skips already-held media either by file existence -- which requires the archive mounted where it writes -- or by a sqlite skip-archive, which requires nothing on disk. Using the latter lets the fetch host write to local disk and rsync afterwards, avoiding tens of thousands of small writes over CIFS and keeping a mid-sync failure from leaving partial files on the live Resilio share. The key is archive_prefix + archive_fmt: the literal "instagram" plus the per-media numeric pk. Verified against a real run -- a 3-image carousel produced 3 rows and a re-run skipped every media file. media_id is absent from our filenames, so the DB cannot be built from names alone, but the listing pass we already make maps every live item to its media_id, and a file listing says which we hold. Seeding therefore costs no extra Instagram requests and no archive content -- the listing GET /api/archives/:name/files already serves is enough. Measured on 0ct0ber19: 2275 live items, 2248 seeded, 27 left to fetch -- exactly the media of the two posts added since the last crawl. The trap worth the comment it carries: posts and reels are filed under post_shortcode, while stories and highlights use the per-item shortcode (post_shortcode there is the containing reel's id, shared by every item). Matching on the wrong field seeded 5 of 2275 rather than failing loudly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
311 lines
13 KiB
Markdown
311 lines
13 KiB
Markdown
# gallery-dl — a CLI replacement for JDownloader2
|
||
|
||
Status: **design + verified config.** `scripts/gdl-sync.py` is a skeleton; no
|
||
profile has been migrated yet.
|
||
|
||
Everything below was measured against the live site and the real archive on
|
||
2026-08-16, not inferred from documentation.
|
||
|
||
## Why gallery-dl and not a hand-rolled script
|
||
|
||
The hard parts of fetching Instagram are pagination, cookie handling, CDN URL
|
||
expiry and resumption. gallery-dl already has all of them, plus extractors that
|
||
map 1:1 onto our sidecar directory layout (`posts`, `reels`, `stories`,
|
||
`highlights`). Rolling our own would mean reimplementing the ban-sensitive part
|
||
by hand.
|
||
|
||
## The safety model — read this before changing any option
|
||
|
||
The ban vector is **requests to `instagram.com`**, not bandwidth. See
|
||
`docs/jdownloader.md` for the history; Instaloader got this account banned by
|
||
asking `instagram.com` a question *per post*.
|
||
|
||
gallery-dl has two API backends and the difference is exactly that vector:
|
||
|
||
```python
|
||
if self.config("api") == "graphql":
|
||
self.api = InstagramGraphqlAPI(self) # per-post api.media() for every
|
||
else: # video and every carousel
|
||
self.api = InstagramRestAPI(self) # <- default, listing-only
|
||
```
|
||
|
||
The REST backend paginates at `count: 30` (feed) / `page_size: 50` (clips), and
|
||
those responses already carry `carousel_media`, `image_versions2`,
|
||
`video_versions` and `product_type`. **No per-post request.** A 300-post
|
||
profile costs roughly 10 requests to `instagram.com`.
|
||
|
||
Rules, in order of importance:
|
||
|
||
1. **`"api": "rest"` always.** Never `graphql`. This is the whole ballgame.
|
||
2. **Never enable `metadata`-style options that trigger extra calls.** If a
|
||
field is not already in the listing response, it is not worth a request.
|
||
3. **Pace it.** `"sleep-request": [4.0, 7.0]` — a randomised gap, not a fixed
|
||
one. Also `"sleep": [1.0, 3.0]` between downloads.
|
||
4. **Cap the download rate** (`downloader.http.rate`) so the CDN side looks like
|
||
a person, not a mirror.
|
||
5. **Run from the same public IP as the browser the cookie came from.** At time
|
||
of writing that is `mattellite` (`66.23.52.196`); the dev workstation is a
|
||
*different* public IP and using the cookie from there is precisely what
|
||
session-hijack detection looks for.
|
||
6. **No programmatic login, ever.** gallery-dl's username/password path is
|
||
disabled upstream anyway; use `--cookies-from-browser`.
|
||
|
||
Do not add proxy rotation, fingerprint spoofing or account rotation. Throttling
|
||
and request-avoidance are welcome; evasion is not.
|
||
|
||
### Cookies
|
||
|
||
The logged-in Chrome on `mattellite` runs with a non-default profile:
|
||
|
||
```
|
||
--user-data-dir=/home/matt/.config/google-chrome-devtools
|
||
```
|
||
|
||
so the cookie flag is:
|
||
|
||
```
|
||
--cookies-from-browser "chrome:/home/matt/.config/google-chrome-devtools"
|
||
```
|
||
|
||
Plain `--cookies-from-browser chrome` fails with "Unable to find chrome cookies
|
||
database" because it looks in `~/.config/google-chrome/`.
|
||
|
||
Anonymous access is **not** a viable fallback: it serves lower-resolution media,
|
||
caps profile pagination at 12 posts, and returns `AuthRequired` for stories and
|
||
highlights.
|
||
|
||
## Output format
|
||
|
||
The viewer's parser is the contract, not JD2's exact bytes. `EXPORT_RE` in
|
||
`src/lib/archive-patterns.ts` accepts all of these, and normalises the index
|
||
with `parseInt`, so **JD2 and gallery-dl naming interoperate**:
|
||
|
||
```
|
||
"… - CrORBIcJJbM.mp4" -> postId=CrORBIcJJbM index=1
|
||
"… - CrORBIcJJbM - 1.mp4" -> postId=CrORBIcJJbM index=1
|
||
"… - C53YPQzp7Wj - 09.jpg" -> postId=C53YPQzp7Wj index=9
|
||
```
|
||
|
||
That means zero-padding and the presence/absence of ` - N` on single-media posts
|
||
are cosmetic. Don't spend effort forcing them.
|
||
|
||
### Directory layout
|
||
|
||
| kind | directory | note |
|
||
|---|---|---|
|
||
| posts | `<user>` | |
|
||
| reels | `<user> - reels` | |
|
||
| stories | `story - <user>` | |
|
||
| highlights | `story highlights - <user> - <title>` | |
|
||
|
||
**Force the directory with `-D`; never use `{username}` for it.** A profile's
|
||
reels tab returns *collab reels owned by other accounts* — `/0ct0ber19/reels/`
|
||
served 6 reels owned by `official_artms` and 1 by `chuuo3o`. With
|
||
`{username}` those would scatter into `official_artms - reels/`. JD2 got this
|
||
right and the archive proves it: `chuuo3o` and `official_artms` filenames sit
|
||
inside `0ct0ber19 - reels/`.
|
||
|
||
So: **owner in the filename, crawl scope in the directory.**
|
||
|
||
### Filenames
|
||
|
||
```jsonc
|
||
"filename": {
|
||
"sidecar_shortcode and count >= 10":
|
||
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode} - {num:02}.{extension}",
|
||
"sidecar_shortcode":
|
||
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode} - {num}.{extension}",
|
||
"":
|
||
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.{extension}"
|
||
}
|
||
```
|
||
|
||
`sidecar_shortcode` is set only when the post is a carousel, so it is the
|
||
carousel discriminator. Conditions are evaluated in order, first match wins
|
||
(`path.py:265`).
|
||
|
||
Stories and highlights use the per-item `{shortcode}`, not `{post_shortcode}`
|
||
(which is the *reel's* id, shared by every item in it):
|
||
|
||
```
|
||
"{date:Olocal/%Y-%m-%d}_{username} - {shortcode}.{extension}"
|
||
```
|
||
|
||
`{date}` on a story/highlight file is the **per-item** `taken_at`
|
||
(`instagram.py:337` prefers `item["taken_at"]`), verified on a 154-item
|
||
highlight whose items carried distinct times while `post_date` stayed pinned to
|
||
the reel. Highlights therefore gain real dates — today they fall back to
|
||
directory mtime.
|
||
|
||
### The timezone is not UTC
|
||
|
||
JD2 stamped filenames in **desktop local time (US Eastern)**. Measured across
|
||
212 comparable posts:
|
||
|
||
| model | mismatches |
|
||
|---|---:|
|
||
| UTC | 19 |
|
||
| UTC−5 (EST) | 10 |
|
||
| UTC−4 (EDT) | **0** |
|
||
| America/New_York (DST-aware) | **0** |
|
||
|
||
`{date:Olocal/%Y-%m-%d}` uses the machine's local zone with per-timestamp DST
|
||
awareness, which reproduces it — `mattellite` is `America/Toronto`, the same
|
||
offsets. Note the **trailing `/` must be omitted**: `Olocal/%Y-%m-%d/` puts the
|
||
separator into the strftime format and it sanitises to an underscore, giving
|
||
`2026-08-15__0ct0ber19`.
|
||
|
||
If the sync ever moves to a host in another timezone, set an explicit
|
||
`{date:O-4/…}` or the dates will silently shift for ~9% of posts.
|
||
|
||
### Caption sidecars
|
||
|
||
JD2 writes one `.txt` per post, named without the index, containing the caption
|
||
with **no trailing newline**, and writes nothing when the caption is empty
|
||
(measured: 197 of 217 posts, 86 of 86 reels, 0 of 10 stories, 0 of 16
|
||
highlights). gallery-dl reproduces this exactly with the default
|
||
`"empty": false`:
|
||
|
||
```jsonc
|
||
{ "name": "metadata", "event": "post", "mode": "custom",
|
||
"content-format": "{description}", "extension": "txt",
|
||
"filename": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.txt" }
|
||
```
|
||
|
||
`"event": "post"` is what makes it one file per post rather than per media file.
|
||
|
||
### Metadata sidecar (new — JD2 had no equivalent)
|
||
|
||
```jsonc
|
||
{ "name": "metadata", "event": "post", "mode": "json",
|
||
"filename": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.json",
|
||
"include": ["post_shortcode","post_id","type","date","post_date","username",
|
||
"fullname","owner_id","description","count","likes","post_url",
|
||
"sidecar_shortcode"] }
|
||
```
|
||
|
||
Use **`include`**, not `fields` — `fields` is for `mode: custom` and silently
|
||
does nothing here, leaving `audio_user` blobs (including another user's profile
|
||
picture URL) in the output.
|
||
|
||
The payoff is `type`, which is Instagram's own classification:
|
||
|
||
```json
|
||
{ "post_shortcode": "DbdG9L9jU4m", "type": "post", "count": 2 } // feed video
|
||
{ "post_shortcode": "Db-lNCoib9m", "type": "reel", "count": 1 } // real reel
|
||
```
|
||
|
||
This is the `product_type: "clips"` signal, delivered free in the listing
|
||
response. It is the authoritative answer to "is this a reel", and would let the
|
||
viewer retire the lone-video heuristic in `src/lib/post-tabs.ts` — see
|
||
"Scanner work" below.
|
||
|
||
**`type` is only populated by listing extractors.** Extracting a single
|
||
`/p/<shortcode>/` URL leaves it `null`. Sync always uses listing URLs, so this
|
||
only matters when testing by hand.
|
||
|
||
## Incremental sync — why the fetch host needs no copy of the archive
|
||
|
||
gallery-dl can skip already-held media two ways, and the difference decides
|
||
whether the fetcher needs the archive mounted:
|
||
|
||
- **By file existence** (default). Needs the destination to already contain the
|
||
files, so it only works if the archive is mounted where gallery-dl writes.
|
||
- **By skip-archive** (`--download-archive`). A sqlite DB of ids. Needs nothing
|
||
on disk.
|
||
|
||
We use the second, so the fetch host can write to **local disk and rsync
|
||
afterwards**. That avoids writing tens of thousands of small files over CIFS,
|
||
and keeps a mid-sync failure from leaving partial files on the live Resilio
|
||
share.
|
||
|
||
The key is `archive_prefix + archive_fmt`, which for this extractor is the
|
||
literal `instagram` plus the per-media numeric pk (`instagram.py:25`,
|
||
`job.py:713-719`). Verified: a 3-image carousel produced
|
||
|
||
```
|
||
instagram3079387627521318672
|
||
instagram3079387627521429433
|
||
instagram3079387627529716672
|
||
```
|
||
|
||
and a second run skipped every media file, rewriting only the idempotent
|
||
`.txt`/`.json` sidecars.
|
||
|
||
**Seeding.** `media_id` is not in our filenames, so the DB cannot be built from
|
||
names alone — but one listing pass (the pass we make anyway) maps every live
|
||
item to its `media_id`, and the archive's *file listing* says which we already
|
||
hold. No extra Instagram requests, and no archive content — a listing is
|
||
enough, which `GET /api/archives/:name/files` already serves.
|
||
|
||
Measured on `0ct0ber19`: 2275 live media items, 2248 seeded from the existing
|
||
listing, **27 left to download** — precisely the media of the two posts added
|
||
since the last crawl.
|
||
|
||
The one trap, which silently seeds almost nothing if you get it backwards:
|
||
|
||
| surface | filed under | why |
|
||
|---|---|---|
|
||
| posts, reels | `post_shortcode` | carousel children each have their own `shortcode`, which never appears in a filename |
|
||
| stories, highlights | `shortcode` (per item) | `post_shortcode` is the containing reel's id, shared by every item |
|
||
|
||
`live_key()` encodes this. Matching on the wrong field seeded 5 of 2275.
|
||
|
||
## Known quirks
|
||
|
||
- **`count` is not the emitted file count.** For 135 of 214 posts it was exactly
|
||
one higher than the number of files written. This makes the `count >= 10`
|
||
padding condition mis-pad a handful of 9-item posts (10 of 214 measured). Since
|
||
the parser normalises the index, this is cosmetic — but it means a re-fetch
|
||
over an existing JD2 tree writes `- 01.jpg` beside an existing `- 1.jpg`.
|
||
- **Carousels get edited.** Two posts had a different media count live than on
|
||
disk. Padding width follows the count *at download time*, so a grown carousel
|
||
produces mixed widths — the archive already contains one such post from JD2.
|
||
- **Highlights already have two naming styles on disk**, and every undated file
|
||
has a dated twin. The scanner dedupes by index so they render once; it is
|
||
wasted disk, not a display bug.
|
||
|
||
## Scanner work (not done yet)
|
||
|
||
`useArchiveScanner` currently treats any `.json` in the tree as a possible
|
||
manifest. Adding gallery-dl sidecars needs it to distinguish three things:
|
||
|
||
1. Instagram export manifests (`posts_1.json`) — existing path.
|
||
2. Instaloader `.json.xz` — existing path, GraphQL node shape.
|
||
3. gallery-dl `.json` — new, flat shape, identified by having
|
||
`post_shortcode` + `type` at the top level.
|
||
|
||
Once (3) is read, `source`/`isStory` and the reel flag should come from `type`
|
||
rather than from the directory and the lone-video heuristic.
|
||
|
||
## Test cases
|
||
|
||
Real subjects, all present in the archive today. See
|
||
`scripts/gdl-sync.py --selftest` for the harness.
|
||
|
||
| # | case | shortcode | expected |
|
||
|---|---|---|---|
|
||
| 1 | single image | `CwcXnQhOqFG` | one `.jpg`, no index |
|
||
| 2 | single feed video | `DbdG9L9jU4m` | one `.mp4`, `type: post` |
|
||
| 3 | carousel, images only | `Cq8LrxSJAJE` | `- 1 … - 3` |
|
||
| 4 | carousel, image + video | `CtohvHxLnWO` | `- 1.jpg … - 4.mp4`, **no `.txt`** |
|
||
| 5 | carousel of exactly 9 | `Cv2Hb_brx_N` | 1-digit index |
|
||
| 6 | carousel of 10+ | `CzM8Uf6B6H_` | 2-digit index `- 01 … - 10` |
|
||
| 7 | reel shown on the posts grid | `C8FHM6EJl15` | in `<user>`, `type: reel` |
|
||
| 8 | reel on the reels tab | `Db-lNCoib9m` | in `<user> - reels`, `type: reel` |
|
||
| 9 | collab reel (other owner) | `DYcZOb0h6Sv` | dir `0ct0ber19 - reels`, filename `chuuo3o` |
|
||
| 10 | story | live only | `story - <user>`, per-item shortcode + date |
|
||
| 11 | story highlight | `C-IImhvpFuk` | `story highlights - <user> - <title>` |
|
||
| 12 | highlight, unicode title | `Drawheeing⠀` | trailing U+2800 preserved in dirname |
|
||
| 13 | empty caption | `CrdsY5CrSsO` | media written, `.txt` absent |
|
||
| 14 | deleted post | `C0TgI7sphfZ` | on disk, absent live — must not be removed |
|
||
| 15 | edited carousel | `C7zG7-jJMlq` | 18 on disk, 8 live — must not be removed |
|
||
| 16 | pinned posts | `0ct0ber19` | 3 pinned, returned out of date order |
|
||
| 17 | profile avatar | `0ct0ber19.jpg` | base dir, undated |
|
||
|
||
Cases 14–16 are reconciliation, not naming: **a sync must never delete**, since
|
||
the archive deliberately outlives Instagram.
|
||
|
||
Not covered, decide before relying on them: the `/reposts/` tab (`0ct0ber19`
|
||
has one) and `/tagged/`. Neither is fetched today.
|