The scripts and docs for fetching from Instagram now live here rather than on main, which is the branch published to GitHub. They carry things that do not belong in a public repo: the fetch host's public IP, the browser profile path the cookie is read from, the NAS archive path, and the list of accounts being archived. This branch is a superset of main — the viewer plus the tooling — so it can take main's changes by merging, and the npm script and CLAUDE.md entries that reference the tooling live here where the files actually exist. Restored with the sync work from the 2026-08-20 run already in place: the --abort flag, the corrected yt-dlp install advice, and the measurements behind both. Note that main's history was rewritten to strip these paths, so the tooling's own per-file history does not exist on this branch. It is preserved on gitea as pre-rewrite-20260820 and pre-rewrite-tooling-20260820. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UXfdJu7QhSJLr47K7koTDF
606 lines
26 KiB
Markdown
606 lines
26 KiB
Markdown
# gallery-dl — a CLI replacement for JDownloader2
|
||
|
||
Status: **in production.** All six ARTMS profiles are synced with
|
||
`scripts/gdl-sync.py`; JD2 is no longer used for them.
|
||
|
||
Everything below was measured against the live site and the real archive on
|
||
2026-08-16 and 2026-08-20, not inferred from documentation.
|
||
|
||
## Why gallery-dl and not a hand-rolled script
|
||
|
||
The hard parts of fetching Instagram are pagination, cookie handling, CDN URL
|
||
expiry and resumption. gallery-dl already has all of them, plus extractors that
|
||
map 1:1 onto our sidecar directory layout (`posts`, `reels`, `stories`,
|
||
`highlights`). Rolling our own would mean reimplementing the ban-sensitive part
|
||
by hand.
|
||
|
||
## The account was suspended on 2026-08-17 — read this first
|
||
|
||
The account used for all of the below was suspended the same day this tooling
|
||
was built, for "activity that doesn't follow our Community Standards on spam".
|
||
The fetching was not the expensive part. **Verification was.**
|
||
|
||
**It was restored, and synced normally again on 2026-08-20** — a full run
|
||
across all six profiles with 0 failures and 0 CDN 429s. That is not evidence
|
||
the limits were imagined; it is one data point on a restored account that has
|
||
been treated carefully since. Everything below still applies, and the budget is
|
||
still per session rather than per command.
|
||
|
||
What was actually spent against `instagram.com` in a few hours, from one
|
||
session and one IP:
|
||
|
||
| activity | rough requests | downloaded |
|
||
|---|---:|---|
|
||
| enumerating a profile grid by scrolling it in an automated browser | ~18 pages | nothing |
|
||
| the same profile again, after a bug in the scraping selector | ~18 pages | nothing |
|
||
| a Reels tab enumerated the same way | ~9 pages | nothing |
|
||
| full `-j` metadata dumps of one profile, twice | ~16 pages | nothing |
|
||
| `--simulate` runs over the same profile, three times | ~24 pages | nothing |
|
||
| single-post `/p/<code>/` fetches while testing filename formats | ~8 | a handful |
|
||
| an aborted sync that re-ran every listing pass before dying | ~40 pages | ~270 MB |
|
||
| the real sync, 24 sources across 6 profiles | ~150 pages | 2.2 GB |
|
||
|
||
The two rows that actually mattered to the archive are the last one and part of
|
||
the second-to-last. **Everything above them produced no files at all**, and
|
||
together they were a comparable number of requests.
|
||
|
||
The warnings arrived in this order and were each rationalised:
|
||
|
||
1. `429 Too Many Requests` from `scontent-*.cdninstagram.com`, losing two
|
||
videos. Treated as a pacing problem — pacing was lowered and the run
|
||
continued.
|
||
2. `400 Bad Request` from `/api/v1/highlights/<id>/highlights_tray/`, on an
|
||
endpoint that had worked hours earlier. Correctly read as a possible block;
|
||
requests stopped.
|
||
3. Suspension.
|
||
|
||
**Treat the first CDN 429 as a stop signal for the session, not a tuning
|
||
parameter.** It is the tolerant surface complaining; if that surface is
|
||
complaining, the rate-limited one has been unhappy for a while.
|
||
|
||
### Rules that follow from this
|
||
|
||
- **Count verification requests against the same budget as fetching.** A
|
||
`--simulate`, a `-j` dump and a browser scroll all hit `instagram.com` and
|
||
download nothing. Being read-only does not make them free; it makes them
|
||
invisible, which is worse.
|
||
- **Never enumerate the live site with an automated browser.** Scrolling a
|
||
214-post grid is ~18 paginated GraphQL loads at machine speed with no dwell
|
||
time between them. It is the most obviously non-human thing in this whole
|
||
document, and it was done here twice on one profile.
|
||
- **Verify against the archive, not against Instagram.** Every naming, dating
|
||
and classification question answered in this file could have been answered
|
||
from files already on disk plus a single listing pass.
|
||
- **`probe_live` is not cached, so every restart re-enumerates everything.**
|
||
The aborted run cost a full duplicate set of listing passes for five
|
||
profiles. Cache probe output to disk before running anything twice.
|
||
- **Budget per session, not per command.** Nothing in the tooling knows what
|
||
the last command spent.
|
||
|
||
### For a replacement account
|
||
|
||
- Let it exist and be used normally for a while before pointing any tool at it.
|
||
- Keep the cookie on one machine and one public IP, as before.
|
||
- Start with a single small profile and stop for the day afterwards.
|
||
- Prefer Instagram's own "Download a copy" export where possible: it is
|
||
first-party, costs no scraping requests, and carries the metadata this whole
|
||
document works around not having.
|
||
|
||
## The safety model — read this before changing any option
|
||
|
||
The ban vector is **requests to `instagram.com`**, not bandwidth. See
|
||
`docs/jdownloader.md` for the history; Instaloader got this account banned by
|
||
asking `instagram.com` a question *per post*.
|
||
|
||
gallery-dl has two API backends and the difference is exactly that vector:
|
||
|
||
```python
|
||
if self.config("api") == "graphql":
|
||
self.api = InstagramGraphqlAPI(self) # per-post api.media() for every
|
||
else: # video and every carousel
|
||
self.api = InstagramRestAPI(self) # <- default, listing-only
|
||
```
|
||
|
||
The REST backend paginates at `count: 30` (feed) / `page_size: 50` (clips), and
|
||
those responses already carry `carousel_media`, `image_versions2`,
|
||
`video_versions` and `product_type`. **No per-post request.** A 300-post
|
||
profile costs roughly 10 requests to `instagram.com`.
|
||
|
||
Rules, in order of importance:
|
||
|
||
1. **`"api": "rest"` always.** Never `graphql`. This is the whole ballgame.
|
||
2. **Never enable `metadata`-style options that trigger extra calls.** If a
|
||
field is not already in the listing response, it is not worth a request.
|
||
3. **Pace it.** `"sleep-request": [4.0, 7.0]` — a randomised gap, not a fixed
|
||
one. Also `"sleep": [1.0, 3.0]` between downloads.
|
||
4. **Cap the download rate** (`downloader.http.rate`) so the CDN side looks like
|
||
a person, not a mirror.
|
||
5. **Run from the same public IP as the browser the cookie came from.** At time
|
||
of writing that is `mattellite` (`66.23.52.196`); the dev workstation is a
|
||
*different* public IP and using the cookie from there is precisely what
|
||
session-hijack detection looks for.
|
||
6. **No programmatic login, ever.** gallery-dl's username/password path is
|
||
disabled upstream anyway; use `--cookies-from-browser`.
|
||
|
||
Do not add proxy rotation, fingerprint spoofing or account rotation. Throttling
|
||
and request-avoidance are welcome; evasion is not.
|
||
|
||
### Cookies
|
||
|
||
The logged-in Chrome on `mattellite` runs with a non-default profile:
|
||
|
||
```
|
||
--user-data-dir=/home/matt/.config/google-chrome-devtools
|
||
```
|
||
|
||
so the cookie flag is:
|
||
|
||
```
|
||
--cookies-from-browser "chrome:/home/matt/.config/google-chrome-devtools"
|
||
```
|
||
|
||
Plain `--cookies-from-browser chrome` fails with "Unable to find chrome cookies
|
||
database" because it looks in `~/.config/google-chrome/`.
|
||
|
||
Anonymous access is **not** a viable fallback: it serves lower-resolution media,
|
||
caps profile pagination at 12 posts, and returns `AuthRequired` for stories and
|
||
highlights.
|
||
|
||
## Output format
|
||
|
||
The viewer's parser is the contract, not JD2's exact bytes. `EXPORT_RE` in
|
||
`src/lib/archive-patterns.ts` accepts all of these, and normalises the index
|
||
with `parseInt`, so **JD2 and gallery-dl naming interoperate**:
|
||
|
||
```
|
||
"… - CrORBIcJJbM.mp4" -> postId=CrORBIcJJbM index=1
|
||
"… - CrORBIcJJbM - 1.mp4" -> postId=CrORBIcJJbM index=1
|
||
"… - C53YPQzp7Wj - 09.jpg" -> postId=C53YPQzp7Wj index=9
|
||
```
|
||
|
||
That means zero-padding and the presence/absence of ` - N` on single-media posts
|
||
are cosmetic. Don't spend effort forcing them.
|
||
|
||
### Directory layout
|
||
|
||
| kind | directory | note |
|
||
|---|---|---|
|
||
| posts | `<user>` | |
|
||
| reels | `<user> - reels` | |
|
||
| stories | `story - <user>` | |
|
||
| highlights | `story highlights - <user> - <title>` | |
|
||
|
||
**Force the directory with `-D`; never use `{username}` for it.** A profile's
|
||
reels tab returns *collab reels owned by other accounts* — `/4utumn07/reels/`
|
||
served 6 reels owned by `official_band` and 1 by `jayo3o`. With
|
||
`{username}` those would scatter into `official_band - reels/`. JD2 got this
|
||
right and the archive proves it: `jayo3o` and `official_band` filenames sit
|
||
inside `4utumn07 - reels/`.
|
||
|
||
So: **owner in the filename, crawl scope in the directory.**
|
||
|
||
### Filenames
|
||
|
||
```jsonc
|
||
"filename": {
|
||
"sidecar_shortcode and count >= 10":
|
||
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode} - {num:02}.{extension}",
|
||
"sidecar_shortcode":
|
||
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode} - {num}.{extension}",
|
||
"":
|
||
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.{extension}"
|
||
}
|
||
```
|
||
|
||
`sidecar_shortcode` is set only when the post is a carousel, so it is the
|
||
carousel discriminator. Conditions are evaluated in order, first match wins
|
||
(`path.py:265`).
|
||
|
||
Stories and highlights use the per-item `{shortcode}`, not `{post_shortcode}`
|
||
(which is the *reel's* id, shared by every item in it):
|
||
|
||
```
|
||
"{date:Olocal/%Y-%m-%d}_{username} - {shortcode}.{extension}"
|
||
```
|
||
|
||
`{date}` on a story/highlight file is the **per-item** `taken_at`
|
||
(`instagram.py:337` prefers `item["taken_at"]`), verified on a 154-item
|
||
highlight whose items carried distinct times while `post_date` stayed pinned to
|
||
the reel. Highlights therefore gain real dates — today they fall back to
|
||
directory mtime.
|
||
|
||
### The timezone is not UTC
|
||
|
||
JD2 stamped filenames in **desktop local time (US Eastern)**. Measured across
|
||
212 comparable posts:
|
||
|
||
| model | mismatches |
|
||
|---|---:|
|
||
| UTC | 19 |
|
||
| UTC−5 (EST) | 10 |
|
||
| UTC−4 (EDT) | **0** |
|
||
| America/New_York (DST-aware) | **0** |
|
||
|
||
`{date:Olocal/%Y-%m-%d}` uses the machine's local zone with per-timestamp DST
|
||
awareness, which reproduces it — `mattellite` is `America/Toronto`, the same
|
||
offsets. Note the **trailing `/` must be omitted**: `Olocal/%Y-%m-%d/` puts the
|
||
separator into the strftime format and it sanitises to an underscore, giving
|
||
`2026-08-15__4utumn07`.
|
||
|
||
If the sync ever moves to a host in another timezone, set an explicit
|
||
`{date:O-4/…}` or the dates will silently shift for ~9% of posts.
|
||
|
||
### Caption sidecars
|
||
|
||
JD2 writes one `.txt` per post, named without the index, containing the caption
|
||
with **no trailing newline**, and writes nothing when the caption is empty
|
||
(measured: 197 of 217 posts, 86 of 86 reels, 0 of 10 stories, 0 of 16
|
||
highlights). gallery-dl reproduces this exactly with the default
|
||
`"empty": false`:
|
||
|
||
```jsonc
|
||
{ "name": "metadata", "event": "post", "mode": "custom",
|
||
"content-format": "{description}", "extension": "txt",
|
||
"filename": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.txt" }
|
||
```
|
||
|
||
`"event": "post"` is what makes it one file per post rather than per media file.
|
||
|
||
### Metadata sidecar (new — JD2 had no equivalent)
|
||
|
||
```jsonc
|
||
{ "name": "metadata", "event": "post", "mode": "json",
|
||
"filename": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.json",
|
||
"include": ["post_shortcode","post_id","type","date","post_date","username",
|
||
"fullname","owner_id","description","count","likes","post_url",
|
||
"sidecar_shortcode"] }
|
||
```
|
||
|
||
Use **`include`**, not `fields` — `fields` is for `mode: custom` and silently
|
||
does nothing here, leaving `audio_user` blobs (including another user's profile
|
||
picture URL) in the output.
|
||
|
||
The payoff is `type`, which is Instagram's own classification:
|
||
|
||
```json
|
||
{ "post_shortcode": "DbdG9L9jU4m", "type": "post", "count": 2 } // feed video
|
||
{ "post_shortcode": "Db-lNCoib9m", "type": "reel", "count": 1 } // real reel
|
||
```
|
||
|
||
This is the `product_type: "clips"` signal, delivered free in the listing
|
||
response. It is the authoritative answer to "is this a reel", and would let the
|
||
viewer retire the lone-video heuristic in `src/lib/post-tabs.ts` — see
|
||
"Scanner work" below.
|
||
|
||
**`type` is only populated by listing extractors.** Extracting a single
|
||
`/p/<shortcode>/` URL leaves it `null`. Sync always uses listing URLs, so this
|
||
only matters when testing by hand.
|
||
|
||
## Cadence, and the budget that enforces it
|
||
|
||
**Monthly for everything, daily for stories only.** Stories expire in 24h and
|
||
cannot be backfilled, so they are the one surface where missing a day means
|
||
losing the content permanently. Everything else can wait — the skip-archive
|
||
means an infrequent full sync costs barely more than a frequent one, because it
|
||
only fetches what is new.
|
||
|
||
```
|
||
# monthly, everything
|
||
gdl-sync.py --index <viewer-url> --staging ~/gdl/staging \
|
||
--publish <user>@<nas>:<archives> --archive-db ~/gdl/artms.db \
|
||
--urls-file artms_account_links.txt --execute
|
||
|
||
# daily, stories only -- one request per profile
|
||
gdl-sync.py ... --only stories --execute
|
||
```
|
||
|
||
A stories-only run is one source per profile and **never seeds**, because a
|
||
story cannot be in the archive before it is fetched; probing would double the
|
||
cost of the cheapest surface for no benefit. Six profiles is a handful of
|
||
requests.
|
||
|
||
When scheduling it, **randomise the minute and avoid the hour boundary**. A job
|
||
that fires at exactly 09:00 every day is a machine; one that fires somewhere in
|
||
a window looks like someone opening the app.
|
||
|
||
The tool now refuses to repeat itself:
|
||
|
||
| flag | default | what it prevents |
|
||
|---|---|---|
|
||
| `--min-interval` | 20h | re-fetching a source touched recently — the aborted-restart case that re-enumerated five profiles |
|
||
| `--probe-ttl` | 24h | paying for a listing pass twice within a run cycle |
|
||
| `--max-sources` | off | a runaway list touching more than intended |
|
||
| `--force` | off | (escape hatch: ignores both guards) |
|
||
|
||
State lives beside the archive DB as `<db>.state.json`, recording per source
|
||
when it was seeded and last fetched. **Seeding is a one-time bootstrap**: after
|
||
the first successful sync the archive DB records everything gallery-dl has
|
||
seen, so the source is never probed again. That is the single biggest saving
|
||
here — a second full sync costs roughly half what the first did.
|
||
|
||
## Incremental sync — why the fetch host needs no copy of the archive
|
||
|
||
gallery-dl can skip already-held media two ways, and the difference decides
|
||
whether the fetcher needs the archive mounted:
|
||
|
||
- **By file existence** (default). Needs the destination to already contain the
|
||
files, so it only works if the archive is mounted where gallery-dl writes.
|
||
- **By skip-archive** (`--download-archive`). A sqlite DB of ids. Needs nothing
|
||
on disk.
|
||
|
||
We use the second, so the fetch host can write to **local disk and rsync
|
||
afterwards**. That avoids writing tens of thousands of small files over CIFS,
|
||
and keeps a mid-sync failure from leaving partial files on the live Resilio
|
||
share.
|
||
|
||
The key is `archive_prefix + archive_fmt`, which for this extractor is the
|
||
literal `instagram` plus the per-media numeric pk (`instagram.py:25`,
|
||
`job.py:713-719`). Verified: a 3-image carousel produced
|
||
|
||
```
|
||
instagram3079387627521318672
|
||
instagram3079387627521429433
|
||
instagram3079387627529716672
|
||
```
|
||
|
||
and a second run skipped every media file, rewriting only the idempotent
|
||
`.txt`/`.json` sidecars.
|
||
|
||
**Seeding.** `media_id` is not in our filenames, so the DB cannot be built from
|
||
names alone — but one listing pass (the pass we make anyway) maps every live
|
||
item to its `media_id`, and the archive's *file listing* says which we already
|
||
hold. No extra Instagram requests, and no archive content — a listing is
|
||
enough, which `GET /api/archives/:name/files` already serves.
|
||
|
||
Measured on `4utumn07`: 2275 live media items, 2248 seeded from the existing
|
||
listing, **27 left to download** — precisely the media of the two posts added
|
||
since the last crawl.
|
||
|
||
The one trap, which silently seeds almost nothing if you get it backwards:
|
||
|
||
| surface | filed under | why |
|
||
|---|---|---|
|
||
| posts, reels | `post_shortcode` | carousel children each have their own `shortcode`, which never appears in a filename |
|
||
| stories, highlights | `shortcode` (per item) | `post_shortcode` is the containing reel's id, shared by every item |
|
||
|
||
`live_key()` encodes this. Matching on the wrong field seeded 5 of 2275.
|
||
|
||
### The skip-archive saves the CDN, not `instagram.com`
|
||
|
||
Worth being exact about, because the two costs land on different surfaces and
|
||
only one of them bans accounts:
|
||
|
||
| what | which surface | scales with |
|
||
|---|---|---|
|
||
| downloading media | `scontent-*.cdninstagram.com` | how much is **new** |
|
||
| enumerating the profile to find it | `instagram.com` | how **big** the profile is |
|
||
|
||
The skip-archive suppresses the first. It does nothing about the second, so a
|
||
2275-post profile costs ~76 pages of pagination every run, forever, whether it
|
||
has three new posts or none. Seeding (above) saved a *second* full pass, not
|
||
the first.
|
||
|
||
Measured on the 2026-08-20 run, from sidecar write times in staging — free,
|
||
since the run was paying for the listing anyway:
|
||
|
||
```
|
||
1787248852 2026-08-19 … DcOeoVxkthi new, +0s
|
||
1787248944 2026-08-18 … DcLpfoJCZtp new, +92s
|
||
1787249058 2026-08-17 … DcIlGbxCUk0 new, +114s
|
||
1787249162 2026-07-24 … DbKr1TxlPSX ┐ all one second: nothing
|
||
1787249162 2026-08-15 … DcD-FdBCYGm ┘ downloaded, sidecars only
|
||
```
|
||
|
||
Three posts took ~100s each; the remaining 2272 were enumeration with nothing
|
||
to show for it.
|
||
|
||
**Pinned posts do not break early abort.** Test case 16 previously claimed
|
||
`4utumn07` returns its 3 pinned posts out of date order — that is true of the
|
||
*web grid*, but the REST `/posts/` listing came back strictly
|
||
reverse-chronological, newest first, no hoisting. That matters because
|
||
front-loaded old posts are the one thing that would make `skip: abort:N`
|
||
dangerous: it would trip on them and abort before reaching anything new.
|
||
|
||
So `skip: abort:N` is viable, and cuts ~420 requests per run to ~40-60:
|
||
|
||
| surface | live items | pages | with `abort:50` |
|
||
|---|---:|---:|---:|
|
||
| posts, 6 profiles | 11,248 | ~377 | ~12 |
|
||
| reels, 6 profiles | 1,080 | ~24 | ~8 |
|
||
| stories + highlights | — | ~20 | ~20 |
|
||
|
||
N counts consecutive skipped **files**, not posts, so it must clear the largest
|
||
already-held carousel — `DcD-FdBCYGm` alone is 22 media. 50 is comfortable; 5
|
||
would not be.
|
||
|
||
**The tradeoff is edited carousels.** Test case 15 is a post that gained items
|
||
after we archived it, and only a full enumeration finds those. Suggested
|
||
policy: `abort:50` for routine runs, a full sweep occasionally.
|
||
|
||
Measured the same day, resuming a stopped run with `--abort 50`:
|
||
|
||
| source | live items | enumerated |
|
||
|---|---:|---:|
|
||
| `dawn_petal` posts | 2,151 | **7** |
|
||
| `dawn_petal` reels | 92 | 53 |
|
||
|
||
One page instead of 72, and every new post was still caught. The 7 is roughly
|
||
3 new posts plus 4 already-held carousels making up the 50 skipped files.
|
||
Reels need 53 because they are single-media, so 50 consecutive skips really is
|
||
50 reels — another reminder that N counts files, and that the same N behaves
|
||
very differently on a carousel-heavy surface than on a reels tab.
|
||
|
||
## Publishing
|
||
|
||
The fetch host stages to local disk and rsyncs afterwards. `rsync
|
||
--ignore-existing` is not an optimisation but the safety property: the archive
|
||
deliberately outlives Instagram, so publishing must only ever **add**. No
|
||
`--delete`, and nothing already present is overwritten — including sidecars,
|
||
which are rewritten every run and would otherwise churn the synced share.
|
||
|
||
Publishing happens once at the end of a run, so a profile that fails midway
|
||
never reaches the archive half-written.
|
||
|
||
## Status
|
||
|
||
In use for all six ARTMS profiles.
|
||
|
||
`crimsonfox` first — 322 files added (74 media, 241 `.json`, 7 `.txt`), nothing
|
||
overwritten or deleted. Of the 74 new media, **zero** duplicated media already
|
||
held under a different name, which is the check that says JD2 and gallery-dl
|
||
naming really do converge.
|
||
|
||
**2026-08-20**, the first full incremental sync, four days after the previous
|
||
one. 184 new media, 299 files published, 0 failures and **0 CDN 429s**:
|
||
|
||
| profile | posts | reels | stories | files added |
|
||
|---|---:|---:|---:|---:|
|
||
| 4utumn07 | 58 | 2 | 4 | +77 |
|
||
| official_band | 12 | — | 2 | +85 |
|
||
| dawn_petal | 41 | 1 | 8 | +63 |
|
||
| emberwood | 23 | — | 4 | +35 |
|
||
| bluejay | 16 | — | 2 | +23 |
|
||
| crimsonfox | 10 | — | — | +16 |
|
||
|
||
The 20 story items are the part that could not have been recovered later.
|
||
|
||
Two things made it cheap, and both are worth keeping:
|
||
|
||
- The archive DB was already seeded from the previous run, so `--min-interval`
|
||
and the recorded `seeded` state meant **no probe passes at all**. A state
|
||
file has to exist for this; if one is missing after a manual run, write it
|
||
rather than letting the tool re-seed 24 sources.
|
||
- `--abort 50` (see above) cut the remaining listing cost by roughly 85%.
|
||
|
||
The run was deliberately **stopped and resumed** halfway to pick up `--abort`.
|
||
That is safe precisely because of the state file: the 12 finished sources were
|
||
already marked `fetched`, so the 20h floor skipped them and only the remaining
|
||
12 re-ran. Stopping a run is cheap now; it was not before.
|
||
|
||
Published files land owned by the SSH user rather than `rslsync`. The viewer
|
||
reads them fine (world-readable), but Resilio does not own what it syncs; worth
|
||
a `chown` if that ever matters. This also makes **`rsync` exit 23**
|
||
("some files/attrs were not transferred") the *normal* outcome of a publish —
|
||
it is the failed `chown`, not lost data. Confirm by re-running the same rsync
|
||
with `--dry-run`: an empty file list means everything arrived.
|
||
|
||
The profiles to fetch live in `artms_account_links.txt` at the archive root,
|
||
passed with `--urls-file`.
|
||
|
||
## Verified run
|
||
|
||
`crimsonfox`, all four surfaces, staged locally and published to a scratch
|
||
directory before the live publish above:
|
||
|
||
```
|
||
==> crimsonfox / posts seeded 915 of 984 live items
|
||
==> crimsonfox / reels seeded 28 of 34 live items
|
||
==> crimsonfox / stories no results (none active)
|
||
==> crimsonfox / highlights no results
|
||
```
|
||
|
||
Output landed correctly, including the collab-reel case — `crimsonfox - reels`
|
||
contains 53 files owned by `crimsonfox`, 10 by `dawn_petal`, 3 by `4utumn07`
|
||
and 2 by `official_band`, all with the owner in the filename and the crawl
|
||
scope as the directory.
|
||
|
||
### The CDN rate-limits, and the first run tripped it
|
||
|
||
At `rate: 3M` with `sleep: [1.0, 3.0]`, `scontent-*.cdninstagram.com` returned
|
||
**`429 Too Many Requests`** and two videos were lost (gallery-dl retried, then
|
||
gave up with exit 4). This is the *tolerant* surface complaining, which is a
|
||
clear signal the pacing was too aggressive.
|
||
|
||
Defaults are now:
|
||
|
||
| option | value |
|
||
|---|---|
|
||
| `--rate` | `1M` |
|
||
| `--sleep-request` | 6–10 s |
|
||
| `--sleep` | 3–6 s |
|
||
| `sleep-429` | 120 s |
|
||
| `retries` (extractor and downloader) | 8 |
|
||
|
||
Re-running with those recovered both videos and produced **0 failures and 0
|
||
429s**. Do not raise them for speed; an archive sync has no deadline.
|
||
|
||
### yt-dlp is worth installing
|
||
|
||
Without it, gallery-dl logs `Cannot import yt-dlp or youtube-dl` and falls back
|
||
to a progressive URL for DASH videos. The fallback mostly works but is what the
|
||
429s hit hardest.
|
||
|
||
**`pipx install yt-dlp` does not work** — it was the advice here until
|
||
2026-08-20, and it is wrong. It gives yt-dlp its own venv, so the binary lands
|
||
on `PATH` while gallery-dl, in a *different* venv, still cannot `import yt_dlp`.
|
||
The symptom is that everything looks installed and the log keeps saying
|
||
`Cannot import yt-dlp`. gallery-dl needs it importable, not runnable:
|
||
|
||
```sh
|
||
pipx inject gallery-dl yt-dlp
|
||
```
|
||
|
||
Verify by asking gallery-dl's own interpreter, not the shell:
|
||
|
||
```sh
|
||
/home/matt/.local/share/pipx/venvs/gallery-dl/bin/python -c 'import yt_dlp'
|
||
```
|
||
|
||
## Known quirks
|
||
|
||
- **`count` is not the emitted file count.** For 135 of 214 posts it was exactly
|
||
one higher than the number of files written. This makes the `count >= 10`
|
||
padding condition mis-pad a handful of 9-item posts (10 of 214 measured). Since
|
||
the parser normalises the index, this is cosmetic — but it means a re-fetch
|
||
over an existing JD2 tree writes `- 01.jpg` beside an existing `- 1.jpg`.
|
||
- **Carousels get edited.** Two posts had a different media count live than on
|
||
disk. Padding width follows the count *at download time*, so a grown carousel
|
||
produces mixed widths — the archive already contains one such post from JD2.
|
||
- **Highlights already have two naming styles on disk**, and every undated file
|
||
has a dated twin. The scanner dedupes by index so they render once; it is
|
||
wasted disk, not a display bug.
|
||
|
||
## Scanner work (not done yet)
|
||
|
||
`useArchiveScanner` currently treats any `.json` in the tree as a possible
|
||
manifest. Adding gallery-dl sidecars needs it to distinguish three things:
|
||
|
||
1. Instagram export manifests (`posts_1.json`) — existing path.
|
||
2. Instaloader `.json.xz` — existing path, GraphQL node shape.
|
||
3. gallery-dl `.json` — new, flat shape, identified by having
|
||
`post_shortcode` + `type` at the top level.
|
||
|
||
Once (3) is read, `source`/`isStory` and the reel flag should come from `type`
|
||
rather than from the directory and the lone-video heuristic.
|
||
|
||
## Test cases
|
||
|
||
Real subjects, all present in the archive today. See
|
||
`scripts/gdl-sync.py --selftest` for the harness.
|
||
|
||
| # | case | shortcode | expected |
|
||
|---|---|---|---|
|
||
| 1 | single image | `CwcXnQhOqFG` | one `.jpg`, no index |
|
||
| 2 | single feed video | `DbdG9L9jU4m` | one `.mp4`, `type: post` |
|
||
| 3 | carousel, images only | `Cq8LrxSJAJE` | `- 1 … - 3` |
|
||
| 4 | carousel, image + video | `CtohvHxLnWO` | `- 1.jpg … - 4.mp4`, **no `.txt`** |
|
||
| 5 | carousel of exactly 9 | `Cv2Hb_brx_N` | 1-digit index |
|
||
| 6 | carousel of 10+ | `CzM8Uf6B6H_` | 2-digit index `- 01 … - 10` |
|
||
| 7 | reel shown on the posts grid | `C8FHM6EJl15` | in `<user>`, `type: reel` |
|
||
| 8 | reel on the reels tab | `Db-lNCoib9m` | in `<user> - reels`, `type: reel` |
|
||
| 9 | collab reel (other owner) | `DYcZOb0h6Sv` | dir `4utumn07 - reels`, filename `jayo3o` |
|
||
| 10 | story | live only | `story - <user>`, per-item shortcode + date |
|
||
| 11 | story highlight | `C-IImhvpFuk` | `story highlights - <user> - <title>` |
|
||
| 12 | highlight, unicode title | `Sketching⠀` | trailing U+2800 preserved in dirname |
|
||
| 13 | empty caption | `CrdsY5CrSsO` | media written, `.txt` absent |
|
||
| 14 | deleted post | `C0TgI7sphfZ` | on disk, absent live — must not be removed |
|
||
| 15 | edited carousel | `C7zG7-jJMlq` | 18 on disk, 8 live — must not be removed |
|
||
| 16 | pinned posts | `4utumn07` | REST listing is strictly reverse-chronological; see below |
|
||
| 17 | profile avatar | `4utumn07.jpg` | base dir, undated |
|
||
|
||
Cases 14–16 are reconciliation, not naming: **a sync must never delete**, since
|
||
the archive deliberately outlives Instagram.
|
||
|
||
Not covered, decide before relying on them: the `/reposts/` tab (`4utumn07`
|
||
has one) and `/tagged/`. Neither is fetched today.
|