Nothing in this tool had any memory: every invocation started from zero and would happily re-enumerate a profile it had listed minutes earlier. That is what suspended the account -- the listing passes, not the downloads -- and an aborted run re-enumerating five profiles on restart was a large part of the bill. Three changes, in order of how much they save: - Seeding is now a one-time bootstrap per source. After the first successful sync the archive DB records everything gallery-dl has seen, so the source is never probed again. A second full sync costs roughly half what the first did. - Stories never seed at all. A story cannot be in the archive before it is fetched, so there is nothing to seed from, and probing would double the cost of the cheapest surface we have. - A source fetched within --min-interval (20h) is refused, and listing results are cached for --probe-ttl (24h), so a restart mid-run is free rather than a repeat. --force overrides both. --only replaces --no-stories and takes any subset of the surfaces, which is what makes a daily stories-only run possible: one source per profile, no seeding, a handful of requests. Everything else stays monthly. Tested with stdlib unittest -- no new dependencies, and it runs anywhere the sync does. The cases include the aborted-restart scenario, which now plans zero work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
488 lines
21 KiB
Markdown
488 lines
21 KiB
Markdown
# gallery-dl — a CLI replacement for JDownloader2
|
||
|
||
Status: **design + verified config.** `scripts/gdl-sync.py` is a skeleton; no
|
||
profile has been migrated yet.
|
||
|
||
Everything below was measured against the live site and the real archive on
|
||
2026-08-16, not inferred from documentation.
|
||
|
||
## Why gallery-dl and not a hand-rolled script
|
||
|
||
The hard parts of fetching Instagram are pagination, cookie handling, CDN URL
|
||
expiry and resumption. gallery-dl already has all of them, plus extractors that
|
||
map 1:1 onto our sidecar directory layout (`posts`, `reels`, `stories`,
|
||
`highlights`). Rolling our own would mean reimplementing the ban-sensitive part
|
||
by hand.
|
||
|
||
## The account was suspended on 2026-08-17 — read this first
|
||
|
||
The account used for all of the below was suspended the same day this tooling
|
||
was built, for "activity that doesn't follow our Community Standards on spam".
|
||
The fetching was not the expensive part. **Verification was.**
|
||
|
||
What was actually spent against `instagram.com` in a few hours, from one
|
||
session and one IP:
|
||
|
||
| activity | rough requests | downloaded |
|
||
|---|---:|---|
|
||
| enumerating a profile grid by scrolling it in an automated browser | ~18 pages | nothing |
|
||
| the same profile again, after a bug in the scraping selector | ~18 pages | nothing |
|
||
| a Reels tab enumerated the same way | ~9 pages | nothing |
|
||
| full `-j` metadata dumps of one profile, twice | ~16 pages | nothing |
|
||
| `--simulate` runs over the same profile, three times | ~24 pages | nothing |
|
||
| single-post `/p/<code>/` fetches while testing filename formats | ~8 | a handful |
|
||
| an aborted sync that re-ran every listing pass before dying | ~40 pages | ~270 MB |
|
||
| the real sync, 24 sources across 6 profiles | ~150 pages | 2.2 GB |
|
||
|
||
The two rows that actually mattered to the archive are the last one and part of
|
||
the second-to-last. **Everything above them produced no files at all**, and
|
||
together they were a comparable number of requests.
|
||
|
||
The warnings arrived in this order and were each rationalised:
|
||
|
||
1. `429 Too Many Requests` from `scontent-*.cdninstagram.com`, losing two
|
||
videos. Treated as a pacing problem — pacing was lowered and the run
|
||
continued.
|
||
2. `400 Bad Request` from `/api/v1/highlights/<id>/highlights_tray/`, on an
|
||
endpoint that had worked hours earlier. Correctly read as a possible block;
|
||
requests stopped.
|
||
3. Suspension.
|
||
|
||
**Treat the first CDN 429 as a stop signal for the session, not a tuning
|
||
parameter.** It is the tolerant surface complaining; if that surface is
|
||
complaining, the rate-limited one has been unhappy for a while.
|
||
|
||
### Rules that follow from this
|
||
|
||
- **Count verification requests against the same budget as fetching.** A
|
||
`--simulate`, a `-j` dump and a browser scroll all hit `instagram.com` and
|
||
download nothing. Being read-only does not make them free; it makes them
|
||
invisible, which is worse.
|
||
- **Never enumerate the live site with an automated browser.** Scrolling a
|
||
214-post grid is ~18 paginated GraphQL loads at machine speed with no dwell
|
||
time between them. It is the most obviously non-human thing in this whole
|
||
document, and it was done here twice on one profile.
|
||
- **Verify against the archive, not against Instagram.** Every naming, dating
|
||
and classification question answered in this file could have been answered
|
||
from files already on disk plus a single listing pass.
|
||
- **`probe_live` is not cached, so every restart re-enumerates everything.**
|
||
The aborted run cost a full duplicate set of listing passes for five
|
||
profiles. Cache probe output to disk before running anything twice.
|
||
- **Budget per session, not per command.** Nothing in the tooling knows what
|
||
the last command spent.
|
||
|
||
### For a replacement account
|
||
|
||
- Let it exist and be used normally for a while before pointing any tool at it.
|
||
- Keep the cookie on one machine and one public IP, as before.
|
||
- Start with a single small profile and stop for the day afterwards.
|
||
- Prefer Instagram's own "Download a copy" export where possible: it is
|
||
first-party, costs no scraping requests, and carries the metadata this whole
|
||
document works around not having.
|
||
|
||
## The safety model — read this before changing any option
|
||
|
||
The ban vector is **requests to `instagram.com`**, not bandwidth. See
|
||
`docs/jdownloader.md` for the history; Instaloader got this account banned by
|
||
asking `instagram.com` a question *per post*.
|
||
|
||
gallery-dl has two API backends and the difference is exactly that vector:
|
||
|
||
```python
|
||
if self.config("api") == "graphql":
|
||
self.api = InstagramGraphqlAPI(self) # per-post api.media() for every
|
||
else: # video and every carousel
|
||
self.api = InstagramRestAPI(self) # <- default, listing-only
|
||
```
|
||
|
||
The REST backend paginates at `count: 30` (feed) / `page_size: 50` (clips), and
|
||
those responses already carry `carousel_media`, `image_versions2`,
|
||
`video_versions` and `product_type`. **No per-post request.** A 300-post
|
||
profile costs roughly 10 requests to `instagram.com`.
|
||
|
||
Rules, in order of importance:
|
||
|
||
1. **`"api": "rest"` always.** Never `graphql`. This is the whole ballgame.
|
||
2. **Never enable `metadata`-style options that trigger extra calls.** If a
|
||
field is not already in the listing response, it is not worth a request.
|
||
3. **Pace it.** `"sleep-request": [4.0, 7.0]` — a randomised gap, not a fixed
|
||
one. Also `"sleep": [1.0, 3.0]` between downloads.
|
||
4. **Cap the download rate** (`downloader.http.rate`) so the CDN side looks like
|
||
a person, not a mirror.
|
||
5. **Run from the same public IP as the browser the cookie came from.** At time
|
||
of writing that is `mattellite` (`66.23.52.196`); the dev workstation is a
|
||
*different* public IP and using the cookie from there is precisely what
|
||
session-hijack detection looks for.
|
||
6. **No programmatic login, ever.** gallery-dl's username/password path is
|
||
disabled upstream anyway; use `--cookies-from-browser`.
|
||
|
||
Do not add proxy rotation, fingerprint spoofing or account rotation. Throttling
|
||
and request-avoidance are welcome; evasion is not.
|
||
|
||
### Cookies
|
||
|
||
The logged-in Chrome on `mattellite` runs with a non-default profile:
|
||
|
||
```
|
||
--user-data-dir=/home/matt/.config/google-chrome-devtools
|
||
```
|
||
|
||
so the cookie flag is:
|
||
|
||
```
|
||
--cookies-from-browser "chrome:/home/matt/.config/google-chrome-devtools"
|
||
```
|
||
|
||
Plain `--cookies-from-browser chrome` fails with "Unable to find chrome cookies
|
||
database" because it looks in `~/.config/google-chrome/`.
|
||
|
||
Anonymous access is **not** a viable fallback: it serves lower-resolution media,
|
||
caps profile pagination at 12 posts, and returns `AuthRequired` for stories and
|
||
highlights.
|
||
|
||
## Output format
|
||
|
||
The viewer's parser is the contract, not JD2's exact bytes. `EXPORT_RE` in
|
||
`src/lib/archive-patterns.ts` accepts all of these, and normalises the index
|
||
with `parseInt`, so **JD2 and gallery-dl naming interoperate**:
|
||
|
||
```
|
||
"… - CrORBIcJJbM.mp4" -> postId=CrORBIcJJbM index=1
|
||
"… - CrORBIcJJbM - 1.mp4" -> postId=CrORBIcJJbM index=1
|
||
"… - C53YPQzp7Wj - 09.jpg" -> postId=C53YPQzp7Wj index=9
|
||
```
|
||
|
||
That means zero-padding and the presence/absence of ` - N` on single-media posts
|
||
are cosmetic. Don't spend effort forcing them.
|
||
|
||
### Directory layout
|
||
|
||
| kind | directory | note |
|
||
|---|---|---|
|
||
| posts | `<user>` | |
|
||
| reels | `<user> - reels` | |
|
||
| stories | `story - <user>` | |
|
||
| highlights | `story highlights - <user> - <title>` | |
|
||
|
||
**Force the directory with `-D`; never use `{username}` for it.** A profile's
|
||
reels tab returns *collab reels owned by other accounts* — `/0ct0ber19/reels/`
|
||
served 6 reels owned by `official_artms` and 1 by `chuuo3o`. With
|
||
`{username}` those would scatter into `official_artms - reels/`. JD2 got this
|
||
right and the archive proves it: `chuuo3o` and `official_artms` filenames sit
|
||
inside `0ct0ber19 - reels/`.
|
||
|
||
So: **owner in the filename, crawl scope in the directory.**
|
||
|
||
### Filenames
|
||
|
||
```jsonc
|
||
"filename": {
|
||
"sidecar_shortcode and count >= 10":
|
||
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode} - {num:02}.{extension}",
|
||
"sidecar_shortcode":
|
||
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode} - {num}.{extension}",
|
||
"":
|
||
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.{extension}"
|
||
}
|
||
```
|
||
|
||
`sidecar_shortcode` is set only when the post is a carousel, so it is the
|
||
carousel discriminator. Conditions are evaluated in order, first match wins
|
||
(`path.py:265`).
|
||
|
||
Stories and highlights use the per-item `{shortcode}`, not `{post_shortcode}`
|
||
(which is the *reel's* id, shared by every item in it):
|
||
|
||
```
|
||
"{date:Olocal/%Y-%m-%d}_{username} - {shortcode}.{extension}"
|
||
```
|
||
|
||
`{date}` on a story/highlight file is the **per-item** `taken_at`
|
||
(`instagram.py:337` prefers `item["taken_at"]`), verified on a 154-item
|
||
highlight whose items carried distinct times while `post_date` stayed pinned to
|
||
the reel. Highlights therefore gain real dates — today they fall back to
|
||
directory mtime.
|
||
|
||
### The timezone is not UTC
|
||
|
||
JD2 stamped filenames in **desktop local time (US Eastern)**. Measured across
|
||
212 comparable posts:
|
||
|
||
| model | mismatches |
|
||
|---|---:|
|
||
| UTC | 19 |
|
||
| UTC−5 (EST) | 10 |
|
||
| UTC−4 (EDT) | **0** |
|
||
| America/New_York (DST-aware) | **0** |
|
||
|
||
`{date:Olocal/%Y-%m-%d}` uses the machine's local zone with per-timestamp DST
|
||
awareness, which reproduces it — `mattellite` is `America/Toronto`, the same
|
||
offsets. Note the **trailing `/` must be omitted**: `Olocal/%Y-%m-%d/` puts the
|
||
separator into the strftime format and it sanitises to an underscore, giving
|
||
`2026-08-15__0ct0ber19`.
|
||
|
||
If the sync ever moves to a host in another timezone, set an explicit
|
||
`{date:O-4/…}` or the dates will silently shift for ~9% of posts.
|
||
|
||
### Caption sidecars
|
||
|
||
JD2 writes one `.txt` per post, named without the index, containing the caption
|
||
with **no trailing newline**, and writes nothing when the caption is empty
|
||
(measured: 197 of 217 posts, 86 of 86 reels, 0 of 10 stories, 0 of 16
|
||
highlights). gallery-dl reproduces this exactly with the default
|
||
`"empty": false`:
|
||
|
||
```jsonc
|
||
{ "name": "metadata", "event": "post", "mode": "custom",
|
||
"content-format": "{description}", "extension": "txt",
|
||
"filename": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.txt" }
|
||
```
|
||
|
||
`"event": "post"` is what makes it one file per post rather than per media file.
|
||
|
||
### Metadata sidecar (new — JD2 had no equivalent)
|
||
|
||
```jsonc
|
||
{ "name": "metadata", "event": "post", "mode": "json",
|
||
"filename": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.json",
|
||
"include": ["post_shortcode","post_id","type","date","post_date","username",
|
||
"fullname","owner_id","description","count","likes","post_url",
|
||
"sidecar_shortcode"] }
|
||
```
|
||
|
||
Use **`include`**, not `fields` — `fields` is for `mode: custom` and silently
|
||
does nothing here, leaving `audio_user` blobs (including another user's profile
|
||
picture URL) in the output.
|
||
|
||
The payoff is `type`, which is Instagram's own classification:
|
||
|
||
```json
|
||
{ "post_shortcode": "DbdG9L9jU4m", "type": "post", "count": 2 } // feed video
|
||
{ "post_shortcode": "Db-lNCoib9m", "type": "reel", "count": 1 } // real reel
|
||
```
|
||
|
||
This is the `product_type: "clips"` signal, delivered free in the listing
|
||
response. It is the authoritative answer to "is this a reel", and would let the
|
||
viewer retire the lone-video heuristic in `src/lib/post-tabs.ts` — see
|
||
"Scanner work" below.
|
||
|
||
**`type` is only populated by listing extractors.** Extracting a single
|
||
`/p/<shortcode>/` URL leaves it `null`. Sync always uses listing URLs, so this
|
||
only matters when testing by hand.
|
||
|
||
## Cadence, and the budget that enforces it
|
||
|
||
**Monthly for everything, daily for stories only.** Stories expire in 24h and
|
||
cannot be backfilled, so they are the one surface where missing a day means
|
||
losing the content permanently. Everything else can wait — the skip-archive
|
||
means an infrequent full sync costs barely more than a frequent one, because it
|
||
only fetches what is new.
|
||
|
||
```
|
||
# monthly, everything
|
||
gdl-sync.py --index <viewer-url> --staging ~/gdl/staging \
|
||
--publish <user>@<nas>:<archives> --archive-db ~/gdl/artms.db \
|
||
--urls-file artms_account_links.txt --execute
|
||
|
||
# daily, stories only -- one request per profile
|
||
gdl-sync.py ... --only stories --execute
|
||
```
|
||
|
||
A stories-only run is one source per profile and **never seeds**, because a
|
||
story cannot be in the archive before it is fetched; probing would double the
|
||
cost of the cheapest surface for no benefit. Six profiles is a handful of
|
||
requests.
|
||
|
||
When scheduling it, **randomise the minute and avoid the hour boundary**. A job
|
||
that fires at exactly 09:00 every day is a machine; one that fires somewhere in
|
||
a window looks like someone opening the app.
|
||
|
||
The tool now refuses to repeat itself:
|
||
|
||
| flag | default | what it prevents |
|
||
|---|---|---|
|
||
| `--min-interval` | 20h | re-fetching a source touched recently — the aborted-restart case that re-enumerated five profiles |
|
||
| `--probe-ttl` | 24h | paying for a listing pass twice within a run cycle |
|
||
| `--max-sources` | off | a runaway list touching more than intended |
|
||
| `--force` | off | (escape hatch: ignores both guards) |
|
||
|
||
State lives beside the archive DB as `<db>.state.json`, recording per source
|
||
when it was seeded and last fetched. **Seeding is a one-time bootstrap**: after
|
||
the first successful sync the archive DB records everything gallery-dl has
|
||
seen, so the source is never probed again. That is the single biggest saving
|
||
here — a second full sync costs roughly half what the first did.
|
||
|
||
## Incremental sync — why the fetch host needs no copy of the archive
|
||
|
||
gallery-dl can skip already-held media two ways, and the difference decides
|
||
whether the fetcher needs the archive mounted:
|
||
|
||
- **By file existence** (default). Needs the destination to already contain the
|
||
files, so it only works if the archive is mounted where gallery-dl writes.
|
||
- **By skip-archive** (`--download-archive`). A sqlite DB of ids. Needs nothing
|
||
on disk.
|
||
|
||
We use the second, so the fetch host can write to **local disk and rsync
|
||
afterwards**. That avoids writing tens of thousands of small files over CIFS,
|
||
and keeps a mid-sync failure from leaving partial files on the live Resilio
|
||
share.
|
||
|
||
The key is `archive_prefix + archive_fmt`, which for this extractor is the
|
||
literal `instagram` plus the per-media numeric pk (`instagram.py:25`,
|
||
`job.py:713-719`). Verified: a 3-image carousel produced
|
||
|
||
```
|
||
instagram3079387627521318672
|
||
instagram3079387627521429433
|
||
instagram3079387627529716672
|
||
```
|
||
|
||
and a second run skipped every media file, rewriting only the idempotent
|
||
`.txt`/`.json` sidecars.
|
||
|
||
**Seeding.** `media_id` is not in our filenames, so the DB cannot be built from
|
||
names alone — but one listing pass (the pass we make anyway) maps every live
|
||
item to its `media_id`, and the archive's *file listing* says which we already
|
||
hold. No extra Instagram requests, and no archive content — a listing is
|
||
enough, which `GET /api/archives/:name/files` already serves.
|
||
|
||
Measured on `0ct0ber19`: 2275 live media items, 2248 seeded from the existing
|
||
listing, **27 left to download** — precisely the media of the two posts added
|
||
since the last crawl.
|
||
|
||
The one trap, which silently seeds almost nothing if you get it backwards:
|
||
|
||
| surface | filed under | why |
|
||
|---|---|---|
|
||
| posts, reels | `post_shortcode` | carousel children each have their own `shortcode`, which never appears in a filename |
|
||
| stories, highlights | `shortcode` (per item) | `post_shortcode` is the containing reel's id, shared by every item |
|
||
|
||
`live_key()` encodes this. Matching on the wrong field seeded 5 of 2275.
|
||
|
||
## Publishing
|
||
|
||
The fetch host stages to local disk and rsyncs afterwards. `rsync
|
||
--ignore-existing` is not an optimisation but the safety property: the archive
|
||
deliberately outlives Instagram, so publishing must only ever **add**. No
|
||
`--delete`, and nothing already present is overwritten — including sidecars,
|
||
which are rewritten every run and would otherwise churn the synced share.
|
||
|
||
Publishing happens once at the end of a run, so a profile that fails midway
|
||
never reaches the archive half-written.
|
||
|
||
## Status
|
||
|
||
In use. `withaseul` has been fetched and published to the live archive — 322
|
||
files added (74 media, 241 `.json`, 7 `.txt`), nothing overwritten or deleted.
|
||
Of the 74 new media files, **zero** duplicated media already held under a
|
||
different name, which is the check that says JD2 and gallery-dl naming really
|
||
do converge.
|
||
|
||
Published files land owned by the SSH user rather than `rslsync`. The viewer
|
||
reads them fine (world-readable), but Resilio does not own what it syncs; worth
|
||
a `chown` if that ever matters.
|
||
|
||
The profiles to fetch live in `artms_account_links.txt` at the archive root,
|
||
passed with `--urls-file`.
|
||
|
||
## Verified run
|
||
|
||
`withaseul`, all four surfaces, staged locally and published to a scratch
|
||
directory before the live publish above:
|
||
|
||
```
|
||
==> withaseul / posts seeded 915 of 984 live items
|
||
==> withaseul / reels seeded 28 of 34 live items
|
||
==> withaseul / stories no results (none active)
|
||
==> withaseul / highlights no results
|
||
```
|
||
|
||
Output landed correctly, including the collab-reel case — `withaseul - reels`
|
||
contains 53 files owned by `withaseul`, 10 by `cher_ryppo`, 3 by `0ct0ber19`
|
||
and 2 by `official_artms`, all with the owner in the filename and the crawl
|
||
scope as the directory.
|
||
|
||
### The CDN rate-limits, and the first run tripped it
|
||
|
||
At `rate: 3M` with `sleep: [1.0, 3.0]`, `scontent-*.cdninstagram.com` returned
|
||
**`429 Too Many Requests`** and two videos were lost (gallery-dl retried, then
|
||
gave up with exit 4). This is the *tolerant* surface complaining, which is a
|
||
clear signal the pacing was too aggressive.
|
||
|
||
Defaults are now:
|
||
|
||
| option | value |
|
||
|---|---|
|
||
| `--rate` | `1M` |
|
||
| `--sleep-request` | 6–10 s |
|
||
| `--sleep` | 3–6 s |
|
||
| `sleep-429` | 120 s |
|
||
| `retries` (extractor and downloader) | 8 |
|
||
|
||
Re-running with those recovered both videos and produced **0 failures and 0
|
||
429s**. Do not raise them for speed; an archive sync has no deadline.
|
||
|
||
### yt-dlp is worth installing
|
||
|
||
Without it, gallery-dl logs `Cannot import yt-dlp or youtube-dl` and falls back
|
||
to a progressive URL for DASH videos. The fallback mostly works but is what the
|
||
429s hit hardest. `pipx install yt-dlp` on the fetch host.
|
||
|
||
## Known quirks
|
||
|
||
- **`count` is not the emitted file count.** For 135 of 214 posts it was exactly
|
||
one higher than the number of files written. This makes the `count >= 10`
|
||
padding condition mis-pad a handful of 9-item posts (10 of 214 measured). Since
|
||
the parser normalises the index, this is cosmetic — but it means a re-fetch
|
||
over an existing JD2 tree writes `- 01.jpg` beside an existing `- 1.jpg`.
|
||
- **Carousels get edited.** Two posts had a different media count live than on
|
||
disk. Padding width follows the count *at download time*, so a grown carousel
|
||
produces mixed widths — the archive already contains one such post from JD2.
|
||
- **Highlights already have two naming styles on disk**, and every undated file
|
||
has a dated twin. The scanner dedupes by index so they render once; it is
|
||
wasted disk, not a display bug.
|
||
|
||
## Scanner work (not done yet)
|
||
|
||
`useArchiveScanner` currently treats any `.json` in the tree as a possible
|
||
manifest. Adding gallery-dl sidecars needs it to distinguish three things:
|
||
|
||
1. Instagram export manifests (`posts_1.json`) — existing path.
|
||
2. Instaloader `.json.xz` — existing path, GraphQL node shape.
|
||
3. gallery-dl `.json` — new, flat shape, identified by having
|
||
`post_shortcode` + `type` at the top level.
|
||
|
||
Once (3) is read, `source`/`isStory` and the reel flag should come from `type`
|
||
rather than from the directory and the lone-video heuristic.
|
||
|
||
## Test cases
|
||
|
||
Real subjects, all present in the archive today. See
|
||
`scripts/gdl-sync.py --selftest` for the harness.
|
||
|
||
| # | case | shortcode | expected |
|
||
|---|---|---|---|
|
||
| 1 | single image | `CwcXnQhOqFG` | one `.jpg`, no index |
|
||
| 2 | single feed video | `DbdG9L9jU4m` | one `.mp4`, `type: post` |
|
||
| 3 | carousel, images only | `Cq8LrxSJAJE` | `- 1 … - 3` |
|
||
| 4 | carousel, image + video | `CtohvHxLnWO` | `- 1.jpg … - 4.mp4`, **no `.txt`** |
|
||
| 5 | carousel of exactly 9 | `Cv2Hb_brx_N` | 1-digit index |
|
||
| 6 | carousel of 10+ | `CzM8Uf6B6H_` | 2-digit index `- 01 … - 10` |
|
||
| 7 | reel shown on the posts grid | `C8FHM6EJl15` | in `<user>`, `type: reel` |
|
||
| 8 | reel on the reels tab | `Db-lNCoib9m` | in `<user> - reels`, `type: reel` |
|
||
| 9 | collab reel (other owner) | `DYcZOb0h6Sv` | dir `0ct0ber19 - reels`, filename `chuuo3o` |
|
||
| 10 | story | live only | `story - <user>`, per-item shortcode + date |
|
||
| 11 | story highlight | `C-IImhvpFuk` | `story highlights - <user> - <title>` |
|
||
| 12 | highlight, unicode title | `Drawheeing⠀` | trailing U+2800 preserved in dirname |
|
||
| 13 | empty caption | `CrdsY5CrSsO` | media written, `.txt` absent |
|
||
| 14 | deleted post | `C0TgI7sphfZ` | on disk, absent live — must not be removed |
|
||
| 15 | edited carousel | `C7zG7-jJMlq` | 18 on disk, 8 live — must not be removed |
|
||
| 16 | pinned posts | `0ct0ber19` | 3 pinned, returned out of date order |
|
||
| 17 | profile avatar | `0ct0ber19.jpg` | base dir, undated |
|
||
|
||
Cases 14–16 are reconciliation, not naming: **a sync must never delete**, since
|
||
the archive deliberately outlives Instagram.
|
||
|
||
Not covered, decide before relying on them: the `/reposts/` tab (`0ct0ber19`
|
||
has one) and `/tagged/`. Neither is fetched today.
|