Compare commits
62
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
31f7ece973 | ||
|
|
07df53acde | ||
|
|
c81f18276f | ||
|
|
4ce5cf048e | ||
|
|
7a085d2272 | ||
|
|
087304243e | ||
|
|
e14dbf6ec8 | ||
|
|
3e29976866 | ||
|
|
0ac1ef951f | ||
|
|
ee66b87bff | ||
|
|
68568ef855 | ||
|
|
3a34e4359e | ||
|
|
4fee8b1dfe | ||
|
|
7b63ba6a76 | ||
|
|
b84113af18 | ||
|
|
abf22eb50f | ||
|
|
e6d874a5d9 | ||
|
|
ff45b6d898 | ||
|
|
30ff8de1e8 | ||
|
|
9560f85515 | ||
|
|
f0c054946f | ||
|
|
42eb598d1c | ||
|
|
aac753ced9 | ||
|
|
9173504190 | ||
|
|
1340d85945 | ||
|
|
c5b0a5cb5f | ||
|
|
ae0f075855 | ||
|
|
3146f896c3 | ||
|
|
ed078c7b46 | ||
|
|
55b0752b05 | ||
|
|
53703cd7cd | ||
|
|
97f5d19ce4 | ||
|
|
1b4aba54d6 | ||
|
|
b30285fe70 | ||
|
|
20209bcad5 | ||
|
|
74902234b3 | ||
|
|
d62bddc3aa | ||
|
|
42c13ea106 | ||
|
|
a4e9ce16a7 | ||
|
|
4f89a69ee3 | ||
|
|
147dcdf2f1 | ||
|
|
d4e20d9b98 | ||
|
|
ec8c771733 | ||
|
|
3784e8729b | ||
|
|
69d62eaa5c | ||
|
|
767f9c508b | ||
|
|
103ce6f207 | ||
|
|
5267dab236 | ||
|
|
ebf2bf660a | ||
|
|
c0f3523a9c | ||
|
|
6f5021638c | ||
|
|
67f7750157 | ||
|
|
9e306eb85e | ||
|
|
b2da08d52d | ||
|
|
d7c13ecc19 | ||
|
|
d396b356be | ||
|
|
e23dfe4474 | ||
|
|
41e7c5e206 | ||
|
|
f685eaebd7 | ||
|
|
47e44ec5e9 | ||
|
|
cd7dc5f981 | ||
|
|
a724e5bc87 |
@@ -9,3 +9,5 @@ coverage/
|
||||
!.env.example
|
||||
_sample-archives
|
||||
_gemini-plans
|
||||
__pycache__/
|
||||
*.pyc
|
||||
|
||||
@@ -37,10 +37,10 @@ Loading is unified behind the `ArchiveFile` interface (`src/types/index.ts`, imp
|
||||
An archive root holds one directory per profile plus *sidecars* that belong to it:
|
||||
|
||||
```
|
||||
4utumn07 -> posts (base)
|
||||
4utumn07 - reels -> reels
|
||||
story - 4utumn07 -> stories
|
||||
story highlights - 4utumn07 - Sunstory -> highlight "Sunstory"
|
||||
0ct0ber19 -> posts (base)
|
||||
0ct0ber19 - reels -> reels
|
||||
story - 0ct0ber19 -> stories
|
||||
story highlights - 0ct0ber19 - Heestory -> highlight "Heestory"
|
||||
```
|
||||
|
||||
`src/lib/archive-grouping.ts` (shared by server and tests) folds these into a single profile with a `sources` list. Sidecars never appear as standalone archives. Each file the server returns carries its `kind`, so the client routes posts / reels / story ring / highlight circles without re-deriving naming rules.
|
||||
@@ -64,6 +64,18 @@ story highlights - 4utumn07 - Sunstory -> highlight "Sunstory"
|
||||
|
||||
Results are cached to IndexedDB. Media records store a stable `path`; **`url` is not persistable** for local archives because blob URLs die with the document.
|
||||
|
||||
Three different JSON shapes turn up as `.json`, so they are told apart structurally, not by filename (`src/lib/gallery-dl-sidecar.ts`):
|
||||
|
||||
| shape | marker |
|
||||
|---|---|
|
||||
| Instagram export manifest | top-level `media` array |
|
||||
| Instaloader `.json.xz` | GraphQL node under `node` / `__typename` |
|
||||
| gallery-dl sidecar | flat, `post_shortcode` + `type`, none of the above |
|
||||
|
||||
The gallery-dl sidecar is the only source that states what a post *is*: its `type` (`post` / `reel` / `story` / `highlight`) is Instagram's own classification, so `post.isReel` set from it beats every fallback in `post-tabs.ts`. This matters — of the 781 items in `official_artms - reels`, the sidecars say only **360 are reels**; the other 421 are ordinary feed videos the clips endpoint returns via `include_feed_video`. Directory-based classification counted all 781.
|
||||
|
||||
**Dates are ranked, not last-write-wins** (`src/lib/post-dates.ts`): sidecar (what Instagram reported) beats filename (what the fetcher wrote) beats mtime (when the file hit disk, and unrelated to when it was posted). Ties keep the incumbent. Several files describe one post and they are scanned in directory order, not in order of trustworthiness, so without the ranking the date was decided by whichever file came first. Only JDownloader highlights fall to mtime at all — `parseArchiveFilename` flags those via `dateFromMtime`.
|
||||
|
||||
### Cache and local-archive persistence (`src/lib/archive-cache.ts`)
|
||||
|
||||
IndexedDB keys are namespaced (`archive:`, `thumb:`, `handle:`) so listing archives does not deserialize every cached thumbnail blob, and thumbnails are scoped per archive to avoid cross-archive collisions.
|
||||
@@ -74,6 +86,23 @@ Restoring an archive **rehydrates URLs from `path`**: server archives rebuild HT
|
||||
|
||||
Images over 1MiB are downscaled in a Web Worker via `OffscreenCanvas`. The queue is **serial on purpose** — decoding several 50MP+ images at once OOMs the tab. `requestThumbnail` must keep a stable identity (it reads cache state through a ref), or every completed thumbnail re-runs the effect in all mounted thumbnails.
|
||||
|
||||
### Profile tabs (`src/lib/post-tabs.ts`)
|
||||
|
||||
The grid holds **everything**, reels included, and the Reels tab is a *filtered view* of that same set. Only the Reels tab filters. The tabs were mutually exclusive until v1.7.0, which hid a lot: 1100 of `loonatheworld`'s 3813 posts and 533 of `for.heejin`'s 1225 never appeared in the grid at all.
|
||||
|
||||
This *approximates* Instagram rather than matching it. Instagram's grid includes a reel only if the creator shared it to feed — a per-post choice, measured live on 2026-08-16: `official_artms` had 21 reels in its first 34 grid tiles, `0ct0ber19` just 1 in 214. That flag appears nowhere in an archive (JD2 stores no metadata, and Instaloader's `product_type` says what a post *is*, not whether it was shared to feed), so showing everything is the closest reachable behaviour. Instagram's "N posts" counter equals its grid, which is why the header counts `postsForTab(allPosts, 'posts')` and not `allPosts` — the raw list still holds both copies of a double-fetched post.
|
||||
|
||||
When checking the live site, note that grid reels link to `/reel/<code>/` while ordinary posts link to `/<user>/p/<code>/`. Matching only `/p/` silently drops every reel, which once produced a confident and completely wrong conclusion that Instagram never shows reels in the grid.
|
||||
|
||||
Deciding *what is a reel* has no good answer for most archives. Instagram's own marker is `product_type` on the post's GraphQL node (`clips` = reel, `feed` = ordinary feed video, `igtv`, `story`) — `__typename` is `GraphVideo` for all three, and aspect ratio does not separate them either. But:
|
||||
|
||||
- Only Instaloader archives carry that metadata, and only newer captures. A survey of `gibiofficial` found `product_type` on 1101 of 5919 sidecars, and just **2** posts marked `clips`.
|
||||
- JDownloader archives carry none at all — media plus a `.txt` holding the bare caption.
|
||||
|
||||
So the viewer believes a `- reels` sidecar directory when one exists, and otherwise falls back to treating a lone video as a reel. **The fallback is a guess**: it cannot tell a reel from a feed video or an old IGTV upload, and it misses videos inside carousels.
|
||||
|
||||
`dedupePostCopies` exists because the JDownloader flow crawls the profile URL and the `/reels/` URL separately (the profile page misses some reels), so the two overlap and a reel can land on disk twice. Those become two posts with distinct directory-scoped ids, which the grid would otherwise render side by side. It dedupes by shortcode, preferring the reels-source copy. It is only safe over `allPosts` — stories and highlights are excluded there, and a shortcode may legitimately appear in both a profile and a highlight.
|
||||
|
||||
### URL state (`src/App.tsx`, `src/lib/routing.ts`)
|
||||
|
||||
Paths mirror Instagram: `/<archive>/`, `/<archive>/reels/`, `/<archive>/p/<shortcode>/`. The old `?a=&t=&p=` form is still parsed for existing links but never written. Reserved prefixes (`api`, `archives`, `assets`…) can't be mistaken for a profile name.
|
||||
@@ -95,11 +124,40 @@ Below `md`, opening a post renders a scrolling feed page rather than the modal (
|
||||
Serves `/api/archives`, `/api/archives/:name/files`, static `/archives`, and the built SPA. Notes:
|
||||
|
||||
- Express decodes route params **after** segment matching, so `..%2f` reaches the handler as `../`. All user-supplied archive names go through `resolveArchivePath`.
|
||||
- Sets CSP and related security headers. The CSP allows `blob:`/`data:` for media and `unsafe-inline` styles (the animation library sets inline styles); scripts stay same-origin only.
|
||||
- `os.userInfo()` throws for a UID with no `/etc/passwd` entry, which is what `--user 1234:1234` produces — use `describeUser()`.
|
||||
|
||||
### CSP: do not tighten `script-src` or `connect-src` without testing xz
|
||||
|
||||
The xz decompressor for Instaloader `.json.xz` sidecars is **WebAssembly**, embedded as a `data:` URL the library fetches at startup. The policy must keep:
|
||||
|
||||
```
|
||||
script-src 'self' 'wasm-unsafe-eval' // compile wasm, without allowing eval() of JS
|
||||
connect-src 'self' data: // fetch the embedded module
|
||||
```
|
||||
|
||||
Removing either breaks decoding with a bare `TypeError: Failed to fetch` **and no stack** — it surfaces through `new Response(stream).json()`, so it reads like a network fault rather than a policy block. The visible symptom is not an error page: archives silently lose captions, story flags and all profile metadata (follower counts, bio, name). This shipped broken for several releases.
|
||||
|
||||
To check quickly, run in the page console:
|
||||
|
||||
```js
|
||||
await fetch('data:application/wasm;base64,AGFzbQEAAAA=') // connect-src
|
||||
await WebAssembly.instantiate(Uint8Array.of(0,97,115,109,1,0,0,0)) // script-src
|
||||
```
|
||||
|
||||
**The Vite dev server does not send these headers**, so anything CSP-related is invisible in `npm run dev`. Verify security-header and PWA behaviour by building and serving `dist/` through `server.js`, not against the dev server.
|
||||
|
||||
### PWA: server-only changes do not reach installed clients
|
||||
|
||||
The service worker precaches `index.html` **together with its response headers**. A change that touches only the server (a CSP fix, a new header) leaves the client build byte-identical, so the precache manifest and `sw.js` are unchanged, the worker never updates, and installed clients keep replaying the old shell with the old headers — indefinitely.
|
||||
|
||||
`vite.config.ts` therefore compiles the package version into the client via `define: { __APP_VERSION__ }`, and `App.tsx` renders it in the footer. That is **load-bearing**: it makes every release change the bundle hash → `index.html` → its precache revision → `sw.js`, which is what browsers byte-compare to decide whether to update. Don't remove it as dead weight.
|
||||
|
||||
To recover a client stuck on an old shell: unregister the service worker, delete its caches, reload.
|
||||
|
||||
### Deployment
|
||||
|
||||
Live archives live in `<share>/Instagram-archive/archives/` — one directory per profile plus sidecars. Directories that are not Instagram profiles (tool output, exports from other services) sit *outside* that folder so they never reach the viewer.
|
||||
|
||||
Container runs as non-root. The image defaults to `node`, but the archive share must be *listable* by that UID — a mode-711 share owned by another account needs `user: "<uid>:<gid>"` in compose. Mount a volume at `/cache` so the index survives restarts.
|
||||
|
||||
### PWA / build quirks
|
||||
|
||||
@@ -0,0 +1,6 @@
|
||||
https://www.instagram.com/0ct0ber19/
|
||||
https://www.instagram.com/kimxxlip/
|
||||
https://www.instagram.com/withaseul/
|
||||
https://www.instagram.com/cher_ryppo/
|
||||
https://www.instagram.com/zindoriyam/
|
||||
https://www.instagram.com/official_artms/
|
||||
@@ -0,0 +1,487 @@
|
||||
# gallery-dl — a CLI replacement for JDownloader2
|
||||
|
||||
Status: **design + verified config.** `scripts/gdl-sync.py` is a skeleton; no
|
||||
profile has been migrated yet.
|
||||
|
||||
Everything below was measured against the live site and the real archive on
|
||||
2026-08-16, not inferred from documentation.
|
||||
|
||||
## Why gallery-dl and not a hand-rolled script
|
||||
|
||||
The hard parts of fetching Instagram are pagination, cookie handling, CDN URL
|
||||
expiry and resumption. gallery-dl already has all of them, plus extractors that
|
||||
map 1:1 onto our sidecar directory layout (`posts`, `reels`, `stories`,
|
||||
`highlights`). Rolling our own would mean reimplementing the ban-sensitive part
|
||||
by hand.
|
||||
|
||||
## The account was suspended on 2026-08-17 — read this first
|
||||
|
||||
The account used for all of the below was suspended the same day this tooling
|
||||
was built, for "activity that doesn't follow our Community Standards on spam".
|
||||
The fetching was not the expensive part. **Verification was.**
|
||||
|
||||
What was actually spent against `instagram.com` in a few hours, from one
|
||||
session and one IP:
|
||||
|
||||
| activity | rough requests | downloaded |
|
||||
|---|---:|---|
|
||||
| enumerating a profile grid by scrolling it in an automated browser | ~18 pages | nothing |
|
||||
| the same profile again, after a bug in the scraping selector | ~18 pages | nothing |
|
||||
| a Reels tab enumerated the same way | ~9 pages | nothing |
|
||||
| full `-j` metadata dumps of one profile, twice | ~16 pages | nothing |
|
||||
| `--simulate` runs over the same profile, three times | ~24 pages | nothing |
|
||||
| single-post `/p/<code>/` fetches while testing filename formats | ~8 | a handful |
|
||||
| an aborted sync that re-ran every listing pass before dying | ~40 pages | ~270 MB |
|
||||
| the real sync, 24 sources across 6 profiles | ~150 pages | 2.2 GB |
|
||||
|
||||
The two rows that actually mattered to the archive are the last one and part of
|
||||
the second-to-last. **Everything above them produced no files at all**, and
|
||||
together they were a comparable number of requests.
|
||||
|
||||
The warnings arrived in this order and were each rationalised:
|
||||
|
||||
1. `429 Too Many Requests` from `scontent-*.cdninstagram.com`, losing two
|
||||
videos. Treated as a pacing problem — pacing was lowered and the run
|
||||
continued.
|
||||
2. `400 Bad Request` from `/api/v1/highlights/<id>/highlights_tray/`, on an
|
||||
endpoint that had worked hours earlier. Correctly read as a possible block;
|
||||
requests stopped.
|
||||
3. Suspension.
|
||||
|
||||
**Treat the first CDN 429 as a stop signal for the session, not a tuning
|
||||
parameter.** It is the tolerant surface complaining; if that surface is
|
||||
complaining, the rate-limited one has been unhappy for a while.
|
||||
|
||||
### Rules that follow from this
|
||||
|
||||
- **Count verification requests against the same budget as fetching.** A
|
||||
`--simulate`, a `-j` dump and a browser scroll all hit `instagram.com` and
|
||||
download nothing. Being read-only does not make them free; it makes them
|
||||
invisible, which is worse.
|
||||
- **Never enumerate the live site with an automated browser.** Scrolling a
|
||||
214-post grid is ~18 paginated GraphQL loads at machine speed with no dwell
|
||||
time between them. It is the most obviously non-human thing in this whole
|
||||
document, and it was done here twice on one profile.
|
||||
- **Verify against the archive, not against Instagram.** Every naming, dating
|
||||
and classification question answered in this file could have been answered
|
||||
from files already on disk plus a single listing pass.
|
||||
- **`probe_live` is not cached, so every restart re-enumerates everything.**
|
||||
The aborted run cost a full duplicate set of listing passes for five
|
||||
profiles. Cache probe output to disk before running anything twice.
|
||||
- **Budget per session, not per command.** Nothing in the tooling knows what
|
||||
the last command spent.
|
||||
|
||||
### For a replacement account
|
||||
|
||||
- Let it exist and be used normally for a while before pointing any tool at it.
|
||||
- Keep the cookie on one machine and one public IP, as before.
|
||||
- Start with a single small profile and stop for the day afterwards.
|
||||
- Prefer Instagram's own "Download a copy" export where possible: it is
|
||||
first-party, costs no scraping requests, and carries the metadata this whole
|
||||
document works around not having.
|
||||
|
||||
## The safety model — read this before changing any option
|
||||
|
||||
The ban vector is **requests to `instagram.com`**, not bandwidth. See
|
||||
`docs/jdownloader.md` for the history; Instaloader got this account banned by
|
||||
asking `instagram.com` a question *per post*.
|
||||
|
||||
gallery-dl has two API backends and the difference is exactly that vector:
|
||||
|
||||
```python
|
||||
if self.config("api") == "graphql":
|
||||
self.api = InstagramGraphqlAPI(self) # per-post api.media() for every
|
||||
else: # video and every carousel
|
||||
self.api = InstagramRestAPI(self) # <- default, listing-only
|
||||
```
|
||||
|
||||
The REST backend paginates at `count: 30` (feed) / `page_size: 50` (clips), and
|
||||
those responses already carry `carousel_media`, `image_versions2`,
|
||||
`video_versions` and `product_type`. **No per-post request.** A 300-post
|
||||
profile costs roughly 10 requests to `instagram.com`.
|
||||
|
||||
Rules, in order of importance:
|
||||
|
||||
1. **`"api": "rest"` always.** Never `graphql`. This is the whole ballgame.
|
||||
2. **Never enable `metadata`-style options that trigger extra calls.** If a
|
||||
field is not already in the listing response, it is not worth a request.
|
||||
3. **Pace it.** `"sleep-request": [4.0, 7.0]` — a randomised gap, not a fixed
|
||||
one. Also `"sleep": [1.0, 3.0]` between downloads.
|
||||
4. **Cap the download rate** (`downloader.http.rate`) so the CDN side looks like
|
||||
a person, not a mirror.
|
||||
5. **Run from the same public IP as the browser the cookie came from.** At time
|
||||
of writing that is `mattellite` (`66.23.52.196`); the dev workstation is a
|
||||
*different* public IP and using the cookie from there is precisely what
|
||||
session-hijack detection looks for.
|
||||
6. **No programmatic login, ever.** gallery-dl's username/password path is
|
||||
disabled upstream anyway; use `--cookies-from-browser`.
|
||||
|
||||
Do not add proxy rotation, fingerprint spoofing or account rotation. Throttling
|
||||
and request-avoidance are welcome; evasion is not.
|
||||
|
||||
### Cookies
|
||||
|
||||
The logged-in Chrome on `mattellite` runs with a non-default profile:
|
||||
|
||||
```
|
||||
--user-data-dir=/home/matt/.config/google-chrome-devtools
|
||||
```
|
||||
|
||||
so the cookie flag is:
|
||||
|
||||
```
|
||||
--cookies-from-browser "chrome:/home/matt/.config/google-chrome-devtools"
|
||||
```
|
||||
|
||||
Plain `--cookies-from-browser chrome` fails with "Unable to find chrome cookies
|
||||
database" because it looks in `~/.config/google-chrome/`.
|
||||
|
||||
Anonymous access is **not** a viable fallback: it serves lower-resolution media,
|
||||
caps profile pagination at 12 posts, and returns `AuthRequired` for stories and
|
||||
highlights.
|
||||
|
||||
## Output format
|
||||
|
||||
The viewer's parser is the contract, not JD2's exact bytes. `EXPORT_RE` in
|
||||
`src/lib/archive-patterns.ts` accepts all of these, and normalises the index
|
||||
with `parseInt`, so **JD2 and gallery-dl naming interoperate**:
|
||||
|
||||
```
|
||||
"… - CrORBIcJJbM.mp4" -> postId=CrORBIcJJbM index=1
|
||||
"… - CrORBIcJJbM - 1.mp4" -> postId=CrORBIcJJbM index=1
|
||||
"… - C53YPQzp7Wj - 09.jpg" -> postId=C53YPQzp7Wj index=9
|
||||
```
|
||||
|
||||
That means zero-padding and the presence/absence of ` - N` on single-media posts
|
||||
are cosmetic. Don't spend effort forcing them.
|
||||
|
||||
### Directory layout
|
||||
|
||||
| kind | directory | note |
|
||||
|---|---|---|
|
||||
| posts | `<user>` | |
|
||||
| reels | `<user> - reels` | |
|
||||
| stories | `story - <user>` | |
|
||||
| highlights | `story highlights - <user> - <title>` | |
|
||||
|
||||
**Force the directory with `-D`; never use `{username}` for it.** A profile's
|
||||
reels tab returns *collab reels owned by other accounts* — `/0ct0ber19/reels/`
|
||||
served 6 reels owned by `official_artms` and 1 by `chuuo3o`. With
|
||||
`{username}` those would scatter into `official_artms - reels/`. JD2 got this
|
||||
right and the archive proves it: `chuuo3o` and `official_artms` filenames sit
|
||||
inside `0ct0ber19 - reels/`.
|
||||
|
||||
So: **owner in the filename, crawl scope in the directory.**
|
||||
|
||||
### Filenames
|
||||
|
||||
```jsonc
|
||||
"filename": {
|
||||
"sidecar_shortcode and count >= 10":
|
||||
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode} - {num:02}.{extension}",
|
||||
"sidecar_shortcode":
|
||||
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode} - {num}.{extension}",
|
||||
"":
|
||||
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.{extension}"
|
||||
}
|
||||
```
|
||||
|
||||
`sidecar_shortcode` is set only when the post is a carousel, so it is the
|
||||
carousel discriminator. Conditions are evaluated in order, first match wins
|
||||
(`path.py:265`).
|
||||
|
||||
Stories and highlights use the per-item `{shortcode}`, not `{post_shortcode}`
|
||||
(which is the *reel's* id, shared by every item in it):
|
||||
|
||||
```
|
||||
"{date:Olocal/%Y-%m-%d}_{username} - {shortcode}.{extension}"
|
||||
```
|
||||
|
||||
`{date}` on a story/highlight file is the **per-item** `taken_at`
|
||||
(`instagram.py:337` prefers `item["taken_at"]`), verified on a 154-item
|
||||
highlight whose items carried distinct times while `post_date` stayed pinned to
|
||||
the reel. Highlights therefore gain real dates — today they fall back to
|
||||
directory mtime.
|
||||
|
||||
### The timezone is not UTC
|
||||
|
||||
JD2 stamped filenames in **desktop local time (US Eastern)**. Measured across
|
||||
212 comparable posts:
|
||||
|
||||
| model | mismatches |
|
||||
|---|---:|
|
||||
| UTC | 19 |
|
||||
| UTC−5 (EST) | 10 |
|
||||
| UTC−4 (EDT) | **0** |
|
||||
| America/New_York (DST-aware) | **0** |
|
||||
|
||||
`{date:Olocal/%Y-%m-%d}` uses the machine's local zone with per-timestamp DST
|
||||
awareness, which reproduces it — `mattellite` is `America/Toronto`, the same
|
||||
offsets. Note the **trailing `/` must be omitted**: `Olocal/%Y-%m-%d/` puts the
|
||||
separator into the strftime format and it sanitises to an underscore, giving
|
||||
`2026-08-15__0ct0ber19`.
|
||||
|
||||
If the sync ever moves to a host in another timezone, set an explicit
|
||||
`{date:O-4/…}` or the dates will silently shift for ~9% of posts.
|
||||
|
||||
### Caption sidecars
|
||||
|
||||
JD2 writes one `.txt` per post, named without the index, containing the caption
|
||||
with **no trailing newline**, and writes nothing when the caption is empty
|
||||
(measured: 197 of 217 posts, 86 of 86 reels, 0 of 10 stories, 0 of 16
|
||||
highlights). gallery-dl reproduces this exactly with the default
|
||||
`"empty": false`:
|
||||
|
||||
```jsonc
|
||||
{ "name": "metadata", "event": "post", "mode": "custom",
|
||||
"content-format": "{description}", "extension": "txt",
|
||||
"filename": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.txt" }
|
||||
```
|
||||
|
||||
`"event": "post"` is what makes it one file per post rather than per media file.
|
||||
|
||||
### Metadata sidecar (new — JD2 had no equivalent)
|
||||
|
||||
```jsonc
|
||||
{ "name": "metadata", "event": "post", "mode": "json",
|
||||
"filename": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.json",
|
||||
"include": ["post_shortcode","post_id","type","date","post_date","username",
|
||||
"fullname","owner_id","description","count","likes","post_url",
|
||||
"sidecar_shortcode"] }
|
||||
```
|
||||
|
||||
Use **`include`**, not `fields` — `fields` is for `mode: custom` and silently
|
||||
does nothing here, leaving `audio_user` blobs (including another user's profile
|
||||
picture URL) in the output.
|
||||
|
||||
The payoff is `type`, which is Instagram's own classification:
|
||||
|
||||
```json
|
||||
{ "post_shortcode": "DbdG9L9jU4m", "type": "post", "count": 2 } // feed video
|
||||
{ "post_shortcode": "Db-lNCoib9m", "type": "reel", "count": 1 } // real reel
|
||||
```
|
||||
|
||||
This is the `product_type: "clips"` signal, delivered free in the listing
|
||||
response. It is the authoritative answer to "is this a reel", and would let the
|
||||
viewer retire the lone-video heuristic in `src/lib/post-tabs.ts` — see
|
||||
"Scanner work" below.
|
||||
|
||||
**`type` is only populated by listing extractors.** Extracting a single
|
||||
`/p/<shortcode>/` URL leaves it `null`. Sync always uses listing URLs, so this
|
||||
only matters when testing by hand.
|
||||
|
||||
## Cadence, and the budget that enforces it
|
||||
|
||||
**Monthly for everything, daily for stories only.** Stories expire in 24h and
|
||||
cannot be backfilled, so they are the one surface where missing a day means
|
||||
losing the content permanently. Everything else can wait — the skip-archive
|
||||
means an infrequent full sync costs barely more than a frequent one, because it
|
||||
only fetches what is new.
|
||||
|
||||
```
|
||||
# monthly, everything
|
||||
gdl-sync.py --index <viewer-url> --staging ~/gdl/staging \
|
||||
--publish <user>@<nas>:<archives> --archive-db ~/gdl/artms.db \
|
||||
--urls-file artms_account_links.txt --execute
|
||||
|
||||
# daily, stories only -- one request per profile
|
||||
gdl-sync.py ... --only stories --execute
|
||||
```
|
||||
|
||||
A stories-only run is one source per profile and **never seeds**, because a
|
||||
story cannot be in the archive before it is fetched; probing would double the
|
||||
cost of the cheapest surface for no benefit. Six profiles is a handful of
|
||||
requests.
|
||||
|
||||
When scheduling it, **randomise the minute and avoid the hour boundary**. A job
|
||||
that fires at exactly 09:00 every day is a machine; one that fires somewhere in
|
||||
a window looks like someone opening the app.
|
||||
|
||||
The tool now refuses to repeat itself:
|
||||
|
||||
| flag | default | what it prevents |
|
||||
|---|---|---|
|
||||
| `--min-interval` | 20h | re-fetching a source touched recently — the aborted-restart case that re-enumerated five profiles |
|
||||
| `--probe-ttl` | 24h | paying for a listing pass twice within a run cycle |
|
||||
| `--max-sources` | off | a runaway list touching more than intended |
|
||||
| `--force` | off | (escape hatch: ignores both guards) |
|
||||
|
||||
State lives beside the archive DB as `<db>.state.json`, recording per source
|
||||
when it was seeded and last fetched. **Seeding is a one-time bootstrap**: after
|
||||
the first successful sync the archive DB records everything gallery-dl has
|
||||
seen, so the source is never probed again. That is the single biggest saving
|
||||
here — a second full sync costs roughly half what the first did.
|
||||
|
||||
## Incremental sync — why the fetch host needs no copy of the archive
|
||||
|
||||
gallery-dl can skip already-held media two ways, and the difference decides
|
||||
whether the fetcher needs the archive mounted:
|
||||
|
||||
- **By file existence** (default). Needs the destination to already contain the
|
||||
files, so it only works if the archive is mounted where gallery-dl writes.
|
||||
- **By skip-archive** (`--download-archive`). A sqlite DB of ids. Needs nothing
|
||||
on disk.
|
||||
|
||||
We use the second, so the fetch host can write to **local disk and rsync
|
||||
afterwards**. That avoids writing tens of thousands of small files over CIFS,
|
||||
and keeps a mid-sync failure from leaving partial files on the live Resilio
|
||||
share.
|
||||
|
||||
The key is `archive_prefix + archive_fmt`, which for this extractor is the
|
||||
literal `instagram` plus the per-media numeric pk (`instagram.py:25`,
|
||||
`job.py:713-719`). Verified: a 3-image carousel produced
|
||||
|
||||
```
|
||||
instagram3079387627521318672
|
||||
instagram3079387627521429433
|
||||
instagram3079387627529716672
|
||||
```
|
||||
|
||||
and a second run skipped every media file, rewriting only the idempotent
|
||||
`.txt`/`.json` sidecars.
|
||||
|
||||
**Seeding.** `media_id` is not in our filenames, so the DB cannot be built from
|
||||
names alone — but one listing pass (the pass we make anyway) maps every live
|
||||
item to its `media_id`, and the archive's *file listing* says which we already
|
||||
hold. No extra Instagram requests, and no archive content — a listing is
|
||||
enough, which `GET /api/archives/:name/files` already serves.
|
||||
|
||||
Measured on `0ct0ber19`: 2275 live media items, 2248 seeded from the existing
|
||||
listing, **27 left to download** — precisely the media of the two posts added
|
||||
since the last crawl.
|
||||
|
||||
The one trap, which silently seeds almost nothing if you get it backwards:
|
||||
|
||||
| surface | filed under | why |
|
||||
|---|---|---|
|
||||
| posts, reels | `post_shortcode` | carousel children each have their own `shortcode`, which never appears in a filename |
|
||||
| stories, highlights | `shortcode` (per item) | `post_shortcode` is the containing reel's id, shared by every item |
|
||||
|
||||
`live_key()` encodes this. Matching on the wrong field seeded 5 of 2275.
|
||||
|
||||
## Publishing
|
||||
|
||||
The fetch host stages to local disk and rsyncs afterwards. `rsync
|
||||
--ignore-existing` is not an optimisation but the safety property: the archive
|
||||
deliberately outlives Instagram, so publishing must only ever **add**. No
|
||||
`--delete`, and nothing already present is overwritten — including sidecars,
|
||||
which are rewritten every run and would otherwise churn the synced share.
|
||||
|
||||
Publishing happens once at the end of a run, so a profile that fails midway
|
||||
never reaches the archive half-written.
|
||||
|
||||
## Status
|
||||
|
||||
In use. `withaseul` has been fetched and published to the live archive — 322
|
||||
files added (74 media, 241 `.json`, 7 `.txt`), nothing overwritten or deleted.
|
||||
Of the 74 new media files, **zero** duplicated media already held under a
|
||||
different name, which is the check that says JD2 and gallery-dl naming really
|
||||
do converge.
|
||||
|
||||
Published files land owned by the SSH user rather than `rslsync`. The viewer
|
||||
reads them fine (world-readable), but Resilio does not own what it syncs; worth
|
||||
a `chown` if that ever matters.
|
||||
|
||||
The profiles to fetch live in `artms_account_links.txt` at the archive root,
|
||||
passed with `--urls-file`.
|
||||
|
||||
## Verified run
|
||||
|
||||
`withaseul`, all four surfaces, staged locally and published to a scratch
|
||||
directory before the live publish above:
|
||||
|
||||
```
|
||||
==> withaseul / posts seeded 915 of 984 live items
|
||||
==> withaseul / reels seeded 28 of 34 live items
|
||||
==> withaseul / stories no results (none active)
|
||||
==> withaseul / highlights no results
|
||||
```
|
||||
|
||||
Output landed correctly, including the collab-reel case — `withaseul - reels`
|
||||
contains 53 files owned by `withaseul`, 10 by `cher_ryppo`, 3 by `0ct0ber19`
|
||||
and 2 by `official_artms`, all with the owner in the filename and the crawl
|
||||
scope as the directory.
|
||||
|
||||
### The CDN rate-limits, and the first run tripped it
|
||||
|
||||
At `rate: 3M` with `sleep: [1.0, 3.0]`, `scontent-*.cdninstagram.com` returned
|
||||
**`429 Too Many Requests`** and two videos were lost (gallery-dl retried, then
|
||||
gave up with exit 4). This is the *tolerant* surface complaining, which is a
|
||||
clear signal the pacing was too aggressive.
|
||||
|
||||
Defaults are now:
|
||||
|
||||
| option | value |
|
||||
|---|---|
|
||||
| `--rate` | `1M` |
|
||||
| `--sleep-request` | 6–10 s |
|
||||
| `--sleep` | 3–6 s |
|
||||
| `sleep-429` | 120 s |
|
||||
| `retries` (extractor and downloader) | 8 |
|
||||
|
||||
Re-running with those recovered both videos and produced **0 failures and 0
|
||||
429s**. Do not raise them for speed; an archive sync has no deadline.
|
||||
|
||||
### yt-dlp is worth installing
|
||||
|
||||
Without it, gallery-dl logs `Cannot import yt-dlp or youtube-dl` and falls back
|
||||
to a progressive URL for DASH videos. The fallback mostly works but is what the
|
||||
429s hit hardest. `pipx install yt-dlp` on the fetch host.
|
||||
|
||||
## Known quirks
|
||||
|
||||
- **`count` is not the emitted file count.** For 135 of 214 posts it was exactly
|
||||
one higher than the number of files written. This makes the `count >= 10`
|
||||
padding condition mis-pad a handful of 9-item posts (10 of 214 measured). Since
|
||||
the parser normalises the index, this is cosmetic — but it means a re-fetch
|
||||
over an existing JD2 tree writes `- 01.jpg` beside an existing `- 1.jpg`.
|
||||
- **Carousels get edited.** Two posts had a different media count live than on
|
||||
disk. Padding width follows the count *at download time*, so a grown carousel
|
||||
produces mixed widths — the archive already contains one such post from JD2.
|
||||
- **Highlights already have two naming styles on disk**, and every undated file
|
||||
has a dated twin. The scanner dedupes by index so they render once; it is
|
||||
wasted disk, not a display bug.
|
||||
|
||||
## Scanner work (not done yet)
|
||||
|
||||
`useArchiveScanner` currently treats any `.json` in the tree as a possible
|
||||
manifest. Adding gallery-dl sidecars needs it to distinguish three things:
|
||||
|
||||
1. Instagram export manifests (`posts_1.json`) — existing path.
|
||||
2. Instaloader `.json.xz` — existing path, GraphQL node shape.
|
||||
3. gallery-dl `.json` — new, flat shape, identified by having
|
||||
`post_shortcode` + `type` at the top level.
|
||||
|
||||
Once (3) is read, `source`/`isStory` and the reel flag should come from `type`
|
||||
rather than from the directory and the lone-video heuristic.
|
||||
|
||||
## Test cases
|
||||
|
||||
Real subjects, all present in the archive today. See
|
||||
`scripts/gdl-sync.py --selftest` for the harness.
|
||||
|
||||
| # | case | shortcode | expected |
|
||||
|---|---|---|---|
|
||||
| 1 | single image | `CwcXnQhOqFG` | one `.jpg`, no index |
|
||||
| 2 | single feed video | `DbdG9L9jU4m` | one `.mp4`, `type: post` |
|
||||
| 3 | carousel, images only | `Cq8LrxSJAJE` | `- 1 … - 3` |
|
||||
| 4 | carousel, image + video | `CtohvHxLnWO` | `- 1.jpg … - 4.mp4`, **no `.txt`** |
|
||||
| 5 | carousel of exactly 9 | `Cv2Hb_brx_N` | 1-digit index |
|
||||
| 6 | carousel of 10+ | `CzM8Uf6B6H_` | 2-digit index `- 01 … - 10` |
|
||||
| 7 | reel shown on the posts grid | `C8FHM6EJl15` | in `<user>`, `type: reel` |
|
||||
| 8 | reel on the reels tab | `Db-lNCoib9m` | in `<user> - reels`, `type: reel` |
|
||||
| 9 | collab reel (other owner) | `DYcZOb0h6Sv` | dir `0ct0ber19 - reels`, filename `chuuo3o` |
|
||||
| 10 | story | live only | `story - <user>`, per-item shortcode + date |
|
||||
| 11 | story highlight | `C-IImhvpFuk` | `story highlights - <user> - <title>` |
|
||||
| 12 | highlight, unicode title | `Drawheeing⠀` | trailing U+2800 preserved in dirname |
|
||||
| 13 | empty caption | `CrdsY5CrSsO` | media written, `.txt` absent |
|
||||
| 14 | deleted post | `C0TgI7sphfZ` | on disk, absent live — must not be removed |
|
||||
| 15 | edited carousel | `C7zG7-jJMlq` | 18 on disk, 8 live — must not be removed |
|
||||
| 16 | pinned posts | `0ct0ber19` | 3 pinned, returned out of date order |
|
||||
| 17 | profile avatar | `0ct0ber19.jpg` | base dir, undated |
|
||||
|
||||
Cases 14–16 are reconciliation, not naming: **a sync must never delete**, since
|
||||
the archive deliberately outlives Instagram.
|
||||
|
||||
Not covered, decide before relying on them: the `/reposts/` tab (`0ct0ber19`
|
||||
has one) and `/tagged/`. Neither is fetched today.
|
||||
@@ -0,0 +1,181 @@
|
||||
# JDownloader2 — archive fetching quick reference
|
||||
|
||||
How content gets into this archive, and why the setup is shaped the way it is.
|
||||
|
||||
## Why JDownloader and not Instaloader
|
||||
|
||||
There are two surfaces, and they're treated very differently:
|
||||
|
||||
| Surface | What hits it | Risk |
|
||||
|---|---|---|
|
||||
| `instagram.com` | profile pages, GraphQL/API metadata | Tied to your session, heavily rate-limited. **This is where bans come from.** |
|
||||
| `scontent*.cdninstagram.com` | the actual media | Signed URLs, CDN-served, tolerant. Mostly a bandwidth question. |
|
||||
|
||||
JDownloader does nearly all its work on the CDN. Instaloader's value — the rich
|
||||
`.json.xz` metadata — comes from asking `instagram.com` a question *per post*.
|
||||
|
||||
Concretely, from this archive: `rivvsofficial` has 188 post-metadata files, so
|
||||
backfilling it cost 188 API requests for one 605-file profile. That's the ban
|
||||
vector. Downloading the 238 photos was never the problem.
|
||||
|
||||
Instaloader got this account banned once. JDownloader with throttling did not.
|
||||
|
||||
> **The account was suspended anyway, on 2026-08-17, for "spam".** Not by
|
||||
> JDownloader, and not by downloading. It was suspended during a day of
|
||||
> *building and verifying* the gallery-dl replacement — automated browser
|
||||
> scrolling to enumerate profile grids, repeated `--simulate` and `-j` metadata
|
||||
> passes, and one aborted sync that re-ran every listing pass before dying.
|
||||
>
|
||||
> The framing above is right about which surface is dangerous and wrong about
|
||||
> what reaches it. **Every read of `instagram.com` counts, including the ones
|
||||
> that download nothing** — and read-only work is easy not to count precisely
|
||||
> because it leaves no files behind. See the post-mortem at the top of
|
||||
> `docs/gallery-dl.md`.
|
||||
>
|
||||
> The rule that would have prevented it: *verify against the archive, never
|
||||
> against the live site*, and treat the first CDN `429` as the end of the
|
||||
> session rather than a pacing knob.
|
||||
|
||||
### What the metadata gap actually costs
|
||||
|
||||
Comparing a JDownloader profile against an Instaloader one:
|
||||
|
||||
| | JDownloader | Instaloader |
|
||||
|---|---|---|
|
||||
| Media | ✅ | ✅ |
|
||||
| Captions (`.txt`) | ✅ | ✅ |
|
||||
| Dates (from filenames) | ✅ | ✅ |
|
||||
| Bio / full name | ❌ | ✅ |
|
||||
| Follower counts | ❌ | ✅ |
|
||||
| External URL | ❌ | ✅ |
|
||||
|
||||
Captions already work — the viewer reads the `.txt` sidecars. Everything missing
|
||||
lives in a *single* profile-level record, not the per-post ones. That's why
|
||||
JDownloader-sourced profiles show "0 followers" and a placeholder bio.
|
||||
|
||||
Not worth extra requests. If you ever want it, the zero-request option is a
|
||||
hand-written `profile.json` sidecar (not implemented yet — ask).
|
||||
|
||||
## Settings that matter
|
||||
|
||||
**Chunks per download → 1.** The single most important one. JDownloader splits
|
||||
each file into multiple ranged requests by default; that `Range` pattern looks
|
||||
nothing like a browser or the app. One chunk = one sequential GET per file.
|
||||
`jd2-sync` sets `chunks=1` per job, so no global change is needed — but set it
|
||||
globally too if you ever add links by hand.
|
||||
|
||||
**Max simultaneous downloads → 2–3**, connections-per-host low. Concurrency is
|
||||
what turns "a user" into a statistic.
|
||||
|
||||
**Leave reconnect / IP-change features off.** A mid-session IP change on a live
|
||||
cookie is a *stronger* anomaly signal than the request rate you'd be avoiding.
|
||||
|
||||
## The cookie
|
||||
|
||||
Exported manually from a real browser session. This is the right approach — no
|
||||
programmatic login anywhere, which is the thing that actually gets flagged.
|
||||
|
||||
- Use it from the **same public IP** as the browser it came from. A cookie used
|
||||
from a different network is what session-hijack detection looks for.
|
||||
- When it expires, **re-export from the browser**. Never add a login step to a tool.
|
||||
- It's a full account credential. Keep it off the NAS share and out of the repo.
|
||||
|
||||
## Workflow
|
||||
|
||||
Two URLs per profile, because the profile grid misses some reels:
|
||||
|
||||
```
|
||||
https://www.instagram.com/<user>/
|
||||
https://www.instagram.com/<user>/reels/
|
||||
```
|
||||
|
||||
They overlap slightly — a reel caught by both lands in each directory and shows
|
||||
up twice in the viewer. That's correct and matches Instagram, which also shows
|
||||
reels in the profile grid *and* the Reels tab.
|
||||
|
||||
## Generating jobs
|
||||
|
||||
Instead of pasting URLs and setting output folders by hand:
|
||||
|
||||
```bash
|
||||
npm run jd2 -- --archives /volume1/rslsync/sync/Instagram-archive/archives --dry-run
|
||||
```
|
||||
|
||||
Review, then write it into JDownloader's folder-watch directory:
|
||||
|
||||
```bash
|
||||
npm run jd2 -- --archives /volume1/rslsync/sync/Instagram-archive/archives \
|
||||
--out ~/.jd2/folderwatch
|
||||
```
|
||||
|
||||
JDownloader runs on the desktop while the archive lives on the NAS, so tell it
|
||||
the path *it* sees:
|
||||
|
||||
```bash
|
||||
npm run jd2 -- --archives /mnt/nas/Instagram-archive/archives \
|
||||
--download-base 'Z:\Instagram-archive\archives' \
|
||||
--out ~/.jd2/folderwatch
|
||||
```
|
||||
|
||||
| Flag | Purpose |
|
||||
|---|---|
|
||||
| `--archives <dir>` | Archive root to scan (or `$ARCHIVES_DIR`) |
|
||||
| `--out <dir>` | JDownloader folder-watch directory |
|
||||
| `--download-base <dir>` | Root path as JDownloader sees it (Windows paths fine) |
|
||||
| `--user <name>` | Just this profile (repeatable) |
|
||||
| `--skip <name>` | Never emit jobs for this directory (repeatable) |
|
||||
| `--chunks <n>` | Connections per file (default 1) |
|
||||
| `--auto-start` | Start immediately instead of parking in LinkGrabber |
|
||||
| `--all-reels` | Emit a reels job even where no reels directory exists |
|
||||
| `--dry-run` | Print instead of writing |
|
||||
|
||||
Defaults are deliberately conservative: `chunks=1`, and links park in the
|
||||
LinkGrabber for review rather than auto-starting.
|
||||
|
||||
Only posts and reels are emitted. Highlight URLs need a numeric id and story
|
||||
URLs expire, so those stay manual.
|
||||
|
||||
Directories that aren't Instagram profiles are skipped by username shape
|
||||
(letters, digits, dots, underscores, ≤30 chars) — pointing a crawl at those
|
||||
spends `instagram.com` requests to be told the profile doesn't exist. For names
|
||||
that *look* like usernames but aren't, use `--skip` or a `.jd2ignore` file in
|
||||
the archive root, one name per line.
|
||||
|
||||
Format reference: `src/org/jdownloader/extensions/folderwatchV2/explain.txt`.
|
||||
JDownloader develops on SVN — read it via the daily mirror at
|
||||
<https://github.com/mycodedoesnotcompile2/jdownloader_mirror> (`svn_trunk/`),
|
||||
not one of the abandoned GitHub copies.
|
||||
|
||||
## Expected layout
|
||||
|
||||
Everything downloads into `<archives>/`, one directory per source:
|
||||
|
||||
```
|
||||
archives/
|
||||
0ct0ber19/ posts
|
||||
0ct0ber19 - reels/ reels
|
||||
story - 0ct0ber19/ stories
|
||||
story highlights - 0ct0ber19 - Heestory/ a highlight
|
||||
```
|
||||
|
||||
Non-archive directories (tool output, exports from elsewhere) live *outside*
|
||||
`archives/` so they never reach the viewer.
|
||||
|
||||
The server picks up changes automatically — its index is keyed on directory
|
||||
mtime, so a new file invalidates only that directory.
|
||||
|
||||
## If something goes wrong
|
||||
|
||||
**429 / rate limited** — stop for hours, not seconds. Retrying into a limit is
|
||||
what converts a soft throttle into something worse.
|
||||
|
||||
**Cookie stops working** — re-export from the browser. Don't add a login step.
|
||||
|
||||
**Files land in the wrong folder** — a Packagizer rule is overriding the job.
|
||||
Generated jobs set `overwritePackagizerEnabled=TRUE` to prevent this; check that
|
||||
rules aren't set to run after it.
|
||||
|
||||
**Viewer doesn't show new posts** — check the file is in the right directory and
|
||||
matches the naming pattern (`YYYY-MM-DD_<user> - <shortcode>[ - NN].<ext>`).
|
||||
The index refreshes on directory mtime, so a genuinely new file is picked up on
|
||||
the next request.
|
||||
Generated
+2
-2
@@ -1,12 +1,12 @@
|
||||
{
|
||||
"name": "instaarchive-viewer",
|
||||
"version": "1.6.2",
|
||||
"version": "1.8.0",
|
||||
"lockfileVersion": 3,
|
||||
"requires": true,
|
||||
"packages": {
|
||||
"": {
|
||||
"name": "instaarchive-viewer",
|
||||
"version": "1.6.2",
|
||||
"version": "1.8.0",
|
||||
"dependencies": {
|
||||
"@tailwindcss/vite": "^4.1.14",
|
||||
"@vitejs/plugin-react": "^5.0.4",
|
||||
|
||||
+1
-1
@@ -1,7 +1,7 @@
|
||||
{
|
||||
"name": "instaarchive-viewer",
|
||||
"private": true,
|
||||
"version": "1.6.2",
|
||||
"version": "1.8.0",
|
||||
"type": "module",
|
||||
"scripts": {
|
||||
"dev": "vite --port=3000 --host=0.0.0.0",
|
||||
|
||||
Executable
+817
@@ -0,0 +1,817 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Fetch Instagram profiles into the archive layout using gallery-dl.
|
||||
|
||||
The CLI replacement for the JDownloader2 workflow. See docs/gallery-dl.md for
|
||||
the measurements behind every choice here — especially the safety model, which
|
||||
is the reason this script exists in this shape rather than a simpler one.
|
||||
|
||||
The fetch host needs no copy of the archive. It stages locally and rsyncs
|
||||
afterwards; what it already holds is learned from a *file listing* alone
|
||||
(`--index`), which the viewer's own API serves.
|
||||
|
||||
Usage:
|
||||
./scripts/gdl-sync.py --index https://instaarchive.ergosteur.com \\
|
||||
--staging /var/tmp/gdl --publish user@host:/path/to/archives \\
|
||||
--urls-file artms_account_links.txt --dry-run
|
||||
|
||||
# ...then swap --dry-run for --execute. --index also accepts a local path,
|
||||
# and --profile / --all work instead of --urls-file.
|
||||
|
||||
Always --dry-run first: it prints the plan, and the publish step it reports is
|
||||
the one that would touch the archive.
|
||||
|
||||
Run it from the host whose public IP matches the browser the cookie came from;
|
||||
using the cookie from elsewhere is what session-hijack detection looks for.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import datetime as dt
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
from dataclasses import dataclass, field
|
||||
from pathlib import Path
|
||||
|
||||
# --------------------------------------------------------------------------
|
||||
# Archive layout
|
||||
# --------------------------------------------------------------------------
|
||||
|
||||
# Mirrors src/lib/archive-grouping.ts. Instagram usernames cannot contain
|
||||
# spaces, which is what makes the username separable from a highlight title.
|
||||
RE_HIGHLIGHT = re.compile(r"^story highlights - ([^ ]+) - (.+)$")
|
||||
RE_STORIES = re.compile(r"^story - ([^ ]+)$")
|
||||
RE_REELS = re.compile(r"^([^ ]+) - reels$")
|
||||
|
||||
DATE_FMT = "{date:Olocal/%Y-%m-%d}"
|
||||
"""Local-time date. JD2 stamped US Eastern, NOT UTC (0/212 mismatches vs 19 for
|
||||
UTC). `Olocal` is DST-aware per timestamp. The trailing separator must be
|
||||
omitted or it lands in the strftime format and sanitises to an underscore."""
|
||||
|
||||
POST_STEM = DATE_FMT + "_{username} - {post_shortcode}"
|
||||
ITEM_STEM = DATE_FMT + "_{username} - {shortcode}"
|
||||
|
||||
|
||||
@dataclass
|
||||
class Source:
|
||||
"""One gallery-dl invocation: a URL fetched into a specific directory."""
|
||||
|
||||
kind: str # posts | reels | stories | highlights
|
||||
url: str
|
||||
directory: str # relative to the archives root
|
||||
subcategory: str # gallery-dl config key
|
||||
title: str | None = None # highlight title, when known
|
||||
|
||||
|
||||
@dataclass
|
||||
class Profile:
|
||||
user: str
|
||||
existing: dict[str, str] = field(default_factory=dict) # kind -> dirname
|
||||
|
||||
def sources(self, kinds: set[str]) -> list[Source]:
|
||||
u = self.user
|
||||
base = f"https://www.instagram.com/{u}"
|
||||
all_sources = [
|
||||
Source("posts", f"{base}/posts/", u, "posts"),
|
||||
Source("reels", f"{base}/reels/", f"{u} - reels", "reels"),
|
||||
# Stories expire after 24h, so these can only ever be captured
|
||||
# live. There is no backfill and no re-fetch -- which is why they
|
||||
# are the one surface worth visiting daily.
|
||||
Source("stories", f"https://www.instagram.com/stories/{u}/",
|
||||
f"story - {u}", "stories"),
|
||||
# Highlight directories embed the title, which gallery-dl only
|
||||
# learns mid-extraction -- so this one source fans out into many
|
||||
# directories and is handled with a directory format string.
|
||||
Source("highlights", f"{base}/highlights", "", "highlights"),
|
||||
]
|
||||
return [s for s in all_sources if s.kind in kinds]
|
||||
|
||||
|
||||
def scan_archives(root: Path) -> dict[str, Profile]:
|
||||
"""Group existing directories into profiles, as the server does."""
|
||||
profiles: dict[str, Profile] = {}
|
||||
|
||||
def get(user: str) -> Profile:
|
||||
return profiles.setdefault(user, Profile(user))
|
||||
|
||||
for entry in sorted(os.listdir(root)):
|
||||
if not (root / entry).is_dir() or entry.startswith("."):
|
||||
continue
|
||||
if m := RE_HIGHLIGHT.match(entry):
|
||||
get(m.group(1)).existing.setdefault("highlights", entry)
|
||||
elif m := RE_STORIES.match(entry):
|
||||
get(m.group(1)).existing["stories"] = entry
|
||||
elif m := RE_REELS.match(entry):
|
||||
get(m.group(1)).existing["reels"] = entry
|
||||
else:
|
||||
get(entry).existing["posts"] = entry
|
||||
return profiles
|
||||
|
||||
|
||||
RE_PROFILE_URL = re.compile(
|
||||
r"^(?:https?://)?(?:www\.)?instagram\.com/(?P<user>[^/?#\s]+)/?", re.I)
|
||||
|
||||
# Path segments that are Instagram features, not profiles. A line like
|
||||
# ".../p/ABC123/" names a post, and treating "p" as a username would silently
|
||||
# sync nothing under a nonsense directory.
|
||||
RESERVED_SEGMENTS = {
|
||||
"p", "reel", "reels", "stories", "explore", "accounts", "direct",
|
||||
"tv", "s", "invites", "challenge", "about", "developer",
|
||||
}
|
||||
|
||||
|
||||
def read_urls_file(path: Path) -> list[str]:
|
||||
"""
|
||||
Read profile URLs (or bare usernames) from a file, one per line.
|
||||
|
||||
Written for hand-maintained lists: blank lines are skipped, `#` starts a
|
||||
comment, and either a full URL or a bare username works. Order is kept and
|
||||
duplicates dropped, so a list can be appended to without care.
|
||||
"""
|
||||
users: list[str] = []
|
||||
seen: set[str] = set()
|
||||
|
||||
for lineno, raw in enumerate(path.read_text().splitlines(), 1):
|
||||
line = raw.split("#", 1)[0].strip()
|
||||
if not line:
|
||||
continue
|
||||
|
||||
m = RE_PROFILE_URL.match(line)
|
||||
user = m.group("user") if m else line.strip("/")
|
||||
|
||||
if not user or "/" in user or " " in user:
|
||||
print(f"{path}:{lineno}: cannot read a username from {raw.strip()!r}",
|
||||
file=sys.stderr)
|
||||
continue
|
||||
if user.lower() in RESERVED_SEGMENTS:
|
||||
print(f"{path}:{lineno}: {user!r} is an Instagram path, not a "
|
||||
f"profile — skipping", file=sys.stderr)
|
||||
continue
|
||||
if user in seen:
|
||||
continue
|
||||
|
||||
seen.add(user)
|
||||
users.append(user)
|
||||
|
||||
return users
|
||||
|
||||
|
||||
class ArchiveIndex:
|
||||
"""
|
||||
What the archive already holds, as filenames only.
|
||||
|
||||
Deliberately never reads file *contents*, so the fetch host does not need a
|
||||
copy of the archive — it can stage locally and rsync afterwards. Backed
|
||||
either by a local directory or by the viewer's own API, which already
|
||||
serves exactly this listing and is the cheaper option when the archive
|
||||
lives on network storage (a full walk there took ~52s).
|
||||
"""
|
||||
|
||||
def __init__(self, source: str):
|
||||
self.remote = source.startswith(("http://", "https://"))
|
||||
self.source = source.rstrip("/") if self.remote else None
|
||||
self.root = None if self.remote else Path(source)
|
||||
if self.root and not self.root.is_dir():
|
||||
raise SystemExit(f"archive index not found: {source}")
|
||||
self._cache: dict[str, list[str]] = {}
|
||||
|
||||
def _get(self, path: str):
|
||||
from urllib.request import urlopen
|
||||
with urlopen(f"{self.source}{path}", timeout=60) as resp:
|
||||
return json.load(resp)
|
||||
|
||||
def profiles(self) -> set[str]:
|
||||
if self.remote:
|
||||
return {a["name"] for a in self._get("/api/archives")}
|
||||
return set(scan_archives(self.root))
|
||||
|
||||
def listing(self, user: str) -> list[str]:
|
||||
"""Every filename belonging to a profile, across all its sidecars."""
|
||||
if user in self._cache:
|
||||
return self._cache[user]
|
||||
|
||||
names: list[str] = []
|
||||
if self.remote:
|
||||
try:
|
||||
data = self._get(f"/api/archives/{user}/files")
|
||||
except Exception:
|
||||
data = []
|
||||
files = data if isinstance(data, list) else data.get("files", [])
|
||||
names = [f["path"] for f in files]
|
||||
else:
|
||||
prof = scan_archives(self.root).get(user)
|
||||
for dirname in (prof.existing.values() if prof else ()):
|
||||
d = self.root / dirname
|
||||
if d.is_dir():
|
||||
names += [f"{dirname}/{n}" for n in os.listdir(d)]
|
||||
|
||||
self._cache[user] = names
|
||||
return names
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------
|
||||
# gallery-dl configuration
|
||||
# --------------------------------------------------------------------------
|
||||
|
||||
def build_config(rate: str, sleep_request: list[float],
|
||||
sleep: list[float]) -> dict:
|
||||
"""
|
||||
The config is generated rather than checked in so the safety-critical
|
||||
options cannot drift out of sync with the docs.
|
||||
|
||||
`api: rest` is the single most important line in this file. The graphql
|
||||
backend issues one request PER POST for every video and carousel, which is
|
||||
the pattern that got this account banned once already.
|
||||
"""
|
||||
caption_pp = {
|
||||
"name": "metadata",
|
||||
"event": "post",
|
||||
"mode": "custom",
|
||||
"content-format": "{description}",
|
||||
"extension": "txt",
|
||||
# JD2 wrote no .txt when the caption was empty; "empty": false (the
|
||||
# default) reproduces that.
|
||||
}
|
||||
meta_pp = {
|
||||
"name": "metadata",
|
||||
"event": "post",
|
||||
"mode": "json",
|
||||
# `include`, NOT `fields` -- `fields` applies to mode:custom and
|
||||
# silently does nothing here, dumping audio_user blobs that contain
|
||||
# unrelated users' profile picture URLs.
|
||||
"include": [
|
||||
"post_shortcode", "post_id", "type", "date", "post_date",
|
||||
"username", "fullname", "owner_id", "description", "count",
|
||||
"likes", "post_url", "sidecar_shortcode",
|
||||
],
|
||||
}
|
||||
|
||||
def post_like(stem: str) -> dict:
|
||||
"""Naming for surfaces whose unit is a post (posts, reels)."""
|
||||
return {
|
||||
# `sidecar_shortcode` is set only for carousels, so it is the
|
||||
# carousel discriminator. First matching condition wins.
|
||||
"filename": {
|
||||
"sidecar_shortcode and count >= 10":
|
||||
stem + " - {num:02}.{extension}",
|
||||
"sidecar_shortcode":
|
||||
stem + " - {num}.{extension}",
|
||||
"":
|
||||
stem + ".{extension}",
|
||||
},
|
||||
"postprocessors": [
|
||||
{**caption_pp, "filename": stem + ".txt"},
|
||||
{**meta_pp, "filename": stem + ".json"},
|
||||
],
|
||||
}
|
||||
|
||||
def item_like(stem: str) -> dict:
|
||||
"""
|
||||
Naming for surfaces whose unit is an item inside a reel (stories,
|
||||
highlights). `{shortcode}` is per item; `{post_shortcode}` is the
|
||||
reel's id and is shared by every item in it.
|
||||
|
||||
The media filename uses the per-item shortcode, but the sidecar cannot:
|
||||
it runs at `event: post`, where the kwdict describes the *reel* and has
|
||||
no `shortcode` at all -- which silently formatted as the literal
|
||||
"None", producing one "<date>_<user> - None.json" per reel. It is keyed
|
||||
by `post_shortcode` instead, and is genuinely reel-level data (the
|
||||
reel's own date and item count); per-item dates live in the media
|
||||
filenames, which is the more precise source anyway.
|
||||
"""
|
||||
return {
|
||||
"filename": stem + ".{extension}",
|
||||
"postprocessors": [
|
||||
{**meta_pp,
|
||||
"filename": DATE_FMT + "_{username} - {post_shortcode}.json"},
|
||||
],
|
||||
}
|
||||
|
||||
return {
|
||||
"extractor": {
|
||||
"base-directory": ".",
|
||||
"instagram": {
|
||||
"api": "rest", # never "graphql" -- see docstring
|
||||
"sleep-request": sleep_request,
|
||||
"sleep": sleep,
|
||||
# The CDN does rate-limit: a first run at 3M/1-3s drew
|
||||
# '429 Too Many Requests' from scontent-*.cdninstagram.com and
|
||||
# lost two videos. Back off hard rather than retry fast.
|
||||
"sleep-429": 120.0,
|
||||
"retries": 8,
|
||||
"videos": True,
|
||||
"include": "", # never "all"; sources are explicit
|
||||
# Directory is forced per-invocation with -D, because a reels
|
||||
# tab returns collab reels owned by OTHER accounts and
|
||||
# {username} would scatter them into the wrong profile.
|
||||
"directory": [],
|
||||
"posts": post_like(POST_STEM),
|
||||
"reels": post_like(POST_STEM),
|
||||
"stories": item_like(ITEM_STEM),
|
||||
"highlights": {
|
||||
**item_like(ITEM_STEM),
|
||||
# The only surface that must derive its own directory,
|
||||
# since the title is not known until extraction.
|
||||
"directory": ["story highlights - {username} - {highlight_title}"],
|
||||
},
|
||||
},
|
||||
},
|
||||
# `retries` here is the CDN-side counterpart to sleep-429 above.
|
||||
"downloader": {"http": {"rate": rate, "retries": 8}},
|
||||
"output": {"mode": "null"},
|
||||
}
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------
|
||||
# Planning and execution
|
||||
# --------------------------------------------------------------------------
|
||||
|
||||
def gdl_command(src: Source, staging: Path, config: Path, cookies: str,
|
||||
archive_db: Path | None) -> list[str]:
|
||||
cmd = [
|
||||
"gallery-dl",
|
||||
"--config", str(config),
|
||||
"--cookies-from-browser", cookies,
|
||||
]
|
||||
if archive_db:
|
||||
# Without a seeded skip-archive, staging is empty and every file is
|
||||
# re-downloaded; see seed_archive_db.
|
||||
cmd += ["--download-archive", str(archive_db)]
|
||||
# Forced destination -- never `{username}` -- because a reels tab returns
|
||||
# collab reels owned by other accounts, which would otherwise be filed
|
||||
# under the wrong profile. Highlights are the exception: their directory
|
||||
# embeds a title only known mid-extraction, so the config formats it.
|
||||
dest = staging if src.subcategory == "highlights" else staging / src.directory
|
||||
cmd += ["--destination", str(dest)]
|
||||
cmd.append(src.url)
|
||||
return cmd
|
||||
|
||||
|
||||
# gallery-dl keys its skip-archive on `archive_prefix + archive_fmt`, which for
|
||||
# this extractor is the literal "instagram" followed by the per-media numeric
|
||||
# pk (`instagram.py:25`, `job.py:713-719`). Verified against a real run: a
|
||||
# 3-image carousel produced 3 rows, one per item.
|
||||
ARCHIVE_KEY = "instagram{}".format
|
||||
ARCHIVE_SCHEMA = "CREATE TABLE IF NOT EXISTS archive (entry TEXT PRIMARY KEY)"
|
||||
|
||||
RE_ARCHIVED = re.compile(
|
||||
r"^(\d{4}-\d{2}-\d{2})_(.+?) - ([A-Za-z0-9_-]+?)(?: - (\d+))?\.(\w+)$")
|
||||
NON_MEDIA = {"txt", "json"}
|
||||
|
||||
|
||||
def index_existing(listing: list[str]) -> set[tuple[str, int]]:
|
||||
"""
|
||||
Reduce a flat list of filenames to the (shortcode, index) pairs already
|
||||
held. Only names matter — never the bytes — which is what lets the sync run
|
||||
on a host that has no copy of the archive.
|
||||
"""
|
||||
have: set[tuple[str, int]] = set()
|
||||
for name in listing:
|
||||
m = RE_ARCHIVED.match(name.rsplit("/", 1)[-1])
|
||||
if not m or m.group(5).lower() in NON_MEDIA:
|
||||
continue
|
||||
# An absent index means a single-media post, which is index 1 — the
|
||||
# same normalisation the viewer's EXPORT_RE applies.
|
||||
have.add((m.group(3), int(m.group(4) or 1)))
|
||||
return have
|
||||
|
||||
|
||||
def live_key(item: dict, kind: str) -> tuple[str, int]:
|
||||
"""
|
||||
The (shortcode, index) a live item *would* be filed under, mirroring the
|
||||
filename template exactly.
|
||||
|
||||
The two surfaces disagree about which shortcode identifies a file, and
|
||||
getting this wrong silently seeds almost nothing:
|
||||
|
||||
posts/reels filed under {post_shortcode} — for a carousel, each
|
||||
child item ALSO has its own `shortcode`, which is not
|
||||
what appears in the filename.
|
||||
stories/highlights filed under the per-item {shortcode}, because
|
||||
`post_shortcode` there is the containing reel's id and
|
||||
is shared by every item in it.
|
||||
"""
|
||||
if kind in ("stories", "highlights"):
|
||||
return (item.get("shortcode"), 1)
|
||||
return (item.get("post_shortcode"), item.get("num"))
|
||||
|
||||
|
||||
def seed_archive_db(db: Path, existing: set[tuple[str, int]],
|
||||
live: list[dict], kind: str) -> int:
|
||||
"""
|
||||
Mark everything already held as downloaded, so a fetch into an empty
|
||||
directory pulls only what is missing.
|
||||
|
||||
`live` is the metadata of one listing pass — the pass we have to make
|
||||
anyway — each entry carrying at least `media_id` plus the shortcode fields
|
||||
`live_key` needs. Seeding costs no additional Instagram requests, and needs
|
||||
only a *listing* of the archive, never its contents.
|
||||
"""
|
||||
import sqlite3
|
||||
|
||||
db.parent.mkdir(parents=True, exist_ok=True)
|
||||
con = sqlite3.connect(db)
|
||||
con.execute(ARCHIVE_SCHEMA)
|
||||
rows = [
|
||||
(ARCHIVE_KEY(item["media_id"]),)
|
||||
for item in live
|
||||
if live_key(item, kind) in existing
|
||||
]
|
||||
con.executemany("INSERT OR IGNORE INTO archive (entry) VALUES (?)", rows)
|
||||
con.commit()
|
||||
con.close()
|
||||
return len(rows)
|
||||
|
||||
|
||||
def probe_live(src: Source, config: Path, cookies: str) -> list[dict]:
|
||||
"""
|
||||
One metadata-only listing pass. `sleep` is forced to 0 because it otherwise
|
||||
applies per *file* even with no download — 2275 files at 1-3s each is over
|
||||
an hour for a single profile.
|
||||
"""
|
||||
out = subprocess.run(
|
||||
["gallery-dl", "-j", "--config", str(config),
|
||||
"--cookies-from-browser", cookies, "-o", "sleep=0", src.url],
|
||||
capture_output=True, text=True, check=True,
|
||||
)
|
||||
items: list[dict] = []
|
||||
|
||||
def walk(node):
|
||||
if isinstance(node, dict):
|
||||
if "media_id" in node and "shortcode" in node:
|
||||
items.append(node)
|
||||
for value in node.values():
|
||||
walk(value)
|
||||
elif isinstance(node, list):
|
||||
for value in node:
|
||||
walk(value)
|
||||
|
||||
walk(json.loads(out.stdout))
|
||||
return items
|
||||
|
||||
|
||||
ALL_KINDS = ("posts", "reels", "stories", "highlights")
|
||||
|
||||
# Stories cannot be backfilled and expire in 24h, so a run that only wants
|
||||
# stories is both cheap and the one worth scheduling daily.
|
||||
STORIES_ONLY = {"stories"}
|
||||
|
||||
|
||||
class SyncState:
|
||||
"""
|
||||
What has already been spent against `instagram.com`.
|
||||
|
||||
Exists because nothing else in this tool has any memory: every invocation
|
||||
used to start from zero and happily re-enumerate profiles it had listed
|
||||
minutes earlier. That is what suspended the account — the listing passes,
|
||||
not the downloads.
|
||||
|
||||
Two facts are tracked per source:
|
||||
|
||||
seeded the skip-archive has been primed from the archive listing.
|
||||
This is a ONE-TIME bootstrap: afterwards the archive DB records
|
||||
every item gallery-dl has seen, so the source never needs
|
||||
probing again. This is the single biggest request saving here.
|
||||
fetched when it was last downloaded, so a re-run soon after is refused
|
||||
rather than silently repeating the whole pass.
|
||||
"""
|
||||
|
||||
VERSION = 1
|
||||
|
||||
def __init__(self, path: Path):
|
||||
self.path = path
|
||||
self.data = {"version": self.VERSION, "sources": {}}
|
||||
if path.is_file():
|
||||
try:
|
||||
loaded = json.loads(path.read_text())
|
||||
if loaded.get("version") == self.VERSION:
|
||||
self.data = loaded
|
||||
except Exception:
|
||||
pass # a corrupt state file must never block a sync
|
||||
|
||||
def _entry(self, url: str) -> dict:
|
||||
return self.data.setdefault("sources", {}).setdefault(url, {})
|
||||
|
||||
def needs_seed(self, url: str) -> bool:
|
||||
return not self._entry(url).get("seeded")
|
||||
|
||||
def mark_seeded(self, url: str, stamp: str) -> None:
|
||||
self._entry(url)["seeded"] = stamp
|
||||
|
||||
def last_fetch(self, url: str) -> str | None:
|
||||
return self._entry(url).get("fetched")
|
||||
|
||||
def mark_fetched(self, url: str, stamp: str) -> None:
|
||||
self._entry(url)["fetched"] = stamp
|
||||
|
||||
def save(self) -> None:
|
||||
self.path.parent.mkdir(parents=True, exist_ok=True)
|
||||
self.path.write_text(json.dumps(self.data, indent=1, sort_keys=True))
|
||||
|
||||
|
||||
def hours_since(stamp: str | None, now: float) -> float:
|
||||
"""Hours between an ISO stamp and `now`; infinite when never."""
|
||||
if not stamp:
|
||||
return float("inf")
|
||||
try:
|
||||
then = dt.datetime.fromisoformat(stamp)
|
||||
except ValueError:
|
||||
return float("inf")
|
||||
if then.tzinfo is None:
|
||||
then = then.replace(tzinfo=dt.timezone.utc)
|
||||
return (now - then.timestamp()) / 3600.0
|
||||
|
||||
|
||||
def plan_source(src: Source, state: SyncState, now: float,
|
||||
min_interval: float) -> tuple[bool, bool, str]:
|
||||
"""
|
||||
Decide what a source needs: (fetch, seed, reason).
|
||||
|
||||
Seeding is skipped once done, and skipped entirely for stories — a story
|
||||
cannot exist in the archive before it is fetched, so there is nothing to
|
||||
seed from, and probing would double the request cost of the cheapest
|
||||
surface we have.
|
||||
"""
|
||||
since = hours_since(state.last_fetch(src.url), now)
|
||||
if since < min_interval:
|
||||
return (False, False, f"fetched {since:.1f}h ago, under the "
|
||||
f"{min_interval:g}h floor")
|
||||
if src.kind == "stories":
|
||||
return (True, False, "stories: no seed needed")
|
||||
if state.needs_seed(src.url):
|
||||
return (True, True, "first run: seeding from the archive listing")
|
||||
return (True, False, "already seeded; the skip-archive knows what we hold")
|
||||
|
||||
|
||||
class ProbeCache:
|
||||
"""
|
||||
Listing-pass results, kept so an interrupted run does not pay for them
|
||||
twice. Yesterday an aborted sync re-enumerated five profiles on restart.
|
||||
"""
|
||||
|
||||
def __init__(self, path: Path, ttl_hours: float):
|
||||
self.path = path
|
||||
self.ttl = ttl_hours
|
||||
self.data: dict = {}
|
||||
if path.is_file():
|
||||
try:
|
||||
self.data = json.loads(path.read_text())
|
||||
except Exception:
|
||||
self.data = {}
|
||||
|
||||
def get(self, url: str, now: float) -> list[dict] | None:
|
||||
entry = self.data.get(url)
|
||||
if not entry or hours_since(entry.get("at"), now) > self.ttl:
|
||||
return None
|
||||
return entry.get("items")
|
||||
|
||||
def put(self, url: str, items: list[dict], stamp: str) -> None:
|
||||
# Only the fields seeding needs, so the cache stays small.
|
||||
self.data[url] = {"at": stamp, "items": [
|
||||
{k: i.get(k) for k in ("shortcode", "post_shortcode", "num", "media_id")}
|
||||
for i in items
|
||||
]}
|
||||
|
||||
def save(self) -> None:
|
||||
self.path.parent.mkdir(parents=True, exist_ok=True)
|
||||
self.path.write_text(json.dumps(self.data))
|
||||
|
||||
|
||||
def rsync_command(staging: Path, dest: str, dry_run: bool) -> list[str]:
|
||||
"""
|
||||
Publish a staging tree into the archive.
|
||||
|
||||
`--ignore-existing` is not an optimisation, it is the safety property: the
|
||||
archive deliberately outlives Instagram (posts exist here that Instagram no
|
||||
longer serves), so publishing must only ever *add*. No `--delete`, and
|
||||
nothing already present is overwritten — including sidecars, which get
|
||||
rewritten on every run and would otherwise churn the synced share.
|
||||
|
||||
`dest` may be a local path or any rsync destination (`user@host:/path`),
|
||||
because the archive usually is not writable from the fetch host.
|
||||
"""
|
||||
cmd = ["rsync", "-a", "--ignore-existing", "--partial", "--info=stats2",
|
||||
# Belt and braces: the config lives outside staging, but nothing
|
||||
# resembling tooling output should ever reach the archive. Archive
|
||||
# sidecars are always "<date>_<user> - <code>.json", so none of
|
||||
# these can match real content.
|
||||
"--exclude", "gdl-sync*.json",
|
||||
"--exclude", "*.gdl-config.json",
|
||||
"--exclude", ".gdl-*",
|
||||
"--exclude", "*.sqlite", "--exclude", "*.db"]
|
||||
if dry_run:
|
||||
cmd.append("--dry-run")
|
||||
# Trailing slash: copy the *contents* of staging into dest.
|
||||
cmd += [f"{staging}/", dest if dest.endswith("/") else dest + "/"]
|
||||
return cmd
|
||||
|
||||
|
||||
def publish(staging: Path, dest: str, dry_run: bool) -> int:
|
||||
if not any(staging.iterdir()):
|
||||
print(" nothing staged; skipping publish")
|
||||
return 0
|
||||
cmd = rsync_command(staging, dest, dry_run)
|
||||
print(" " + " ".join(cmd))
|
||||
return subprocess.run(cmd).returncode
|
||||
|
||||
|
||||
def main() -> int:
|
||||
# A sync runs for hours and is normally watched through a redirected log,
|
||||
# where Python's block buffering would withhold progress until it happened
|
||||
# to flush -- and the gallery-dl subprocesses write to the same descriptor
|
||||
# unbuffered, so the log would also interleave out of order.
|
||||
sys.stdout.reconfigure(line_buffering=True)
|
||||
sys.stderr.reconfigure(line_buffering=True)
|
||||
|
||||
ap = argparse.ArgumentParser(description=__doc__,
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter)
|
||||
ap.add_argument("--index", required=True,
|
||||
help="existing archive listing: a local root, or the "
|
||||
"viewer's base URL (only a FILE LISTING is needed, "
|
||||
"never the contents)")
|
||||
ap.add_argument("--publish", required=True,
|
||||
help="rsync destination for fetched files; a local path or "
|
||||
"user@host:/path")
|
||||
ap.add_argument("--staging", type=Path, required=True,
|
||||
help="local scratch directory gallery-dl writes into")
|
||||
g = ap.add_mutually_exclusive_group(required=True)
|
||||
g.add_argument("--profile", action="append", default=[],
|
||||
help="profile to sync; repeatable")
|
||||
g.add_argument("--all", action="store_true", help="every profile on disk")
|
||||
g.add_argument("--urls-file", type=Path,
|
||||
help="file of Instagram profile URLs or usernames, one per "
|
||||
"line; # comments and blank lines allowed")
|
||||
ap.add_argument("--cookies", default="chrome:/home/matt/.config/google-chrome-devtools",
|
||||
help="gallery-dl --cookies-from-browser value")
|
||||
ap.add_argument("--archive-db", type=Path, default=None,
|
||||
help="gallery-dl skip-archive sqlite path")
|
||||
ap.add_argument("--rate", default="1M", help="per-download rate cap")
|
||||
ap.add_argument("--sleep-request", nargs=2, type=float, default=[6.0, 10.0],
|
||||
metavar=("MIN", "MAX"))
|
||||
ap.add_argument("--sleep", nargs=2, type=float, default=[3.0, 6.0],
|
||||
metavar=("MIN", "MAX"))
|
||||
ap.add_argument("--only", default=",".join(ALL_KINDS),
|
||||
help="comma-separated surfaces to sync: "
|
||||
"posts,reels,stories,highlights. Use --only stories "
|
||||
"for the cheap daily run.")
|
||||
ap.add_argument("--min-interval", type=float, default=20.0, metavar="HOURS",
|
||||
help="refuse to re-fetch a source touched more recently "
|
||||
"than this (default 20h); the guard that makes a "
|
||||
"restart cheap instead of a repeat")
|
||||
ap.add_argument("--max-sources", type=int, default=0, metavar="N",
|
||||
help="hard ceiling on sources touched in one run "
|
||||
"(0 = no limit)")
|
||||
ap.add_argument("--probe-ttl", type=float, default=24.0, metavar="HOURS",
|
||||
help="reuse cached listing results younger than this")
|
||||
ap.add_argument("--force", action="store_true",
|
||||
help="ignore --min-interval and the probe cache")
|
||||
mode = ap.add_mutually_exclusive_group()
|
||||
mode.add_argument("--dry-run", action="store_true", default=True,
|
||||
help="print the plan and the config; default")
|
||||
mode.add_argument("--execute", action="store_true",
|
||||
help="actually run gallery-dl")
|
||||
args = ap.parse_args()
|
||||
|
||||
if not shutil.which("gallery-dl"):
|
||||
print("gallery-dl not on PATH", file=sys.stderr)
|
||||
return 2
|
||||
if not shutil.which("rsync"):
|
||||
print("rsync not on PATH", file=sys.stderr)
|
||||
return 2
|
||||
|
||||
index = ArchiveIndex(args.index)
|
||||
names = index.profiles()
|
||||
if args.urls_file:
|
||||
if not args.urls_file.is_file():
|
||||
print(f"urls file not found: {args.urls_file}", file=sys.stderr)
|
||||
return 2
|
||||
wanted = read_urls_file(args.urls_file)
|
||||
if not wanted:
|
||||
print(f"no usable profiles in {args.urls_file}", file=sys.stderr)
|
||||
return 2
|
||||
print(f"read {len(wanted)} profile(s) from {args.urls_file}")
|
||||
selected = [Profile(p) for p in wanted]
|
||||
elif args.profile:
|
||||
for p in args.profile:
|
||||
if p not in names:
|
||||
print(f"note: {p} is not in the index yet; it will be created")
|
||||
selected = [Profile(p) for p in args.profile]
|
||||
else:
|
||||
selected = [Profile(p) for p in sorted(names)]
|
||||
|
||||
config = build_config(args.rate, list(args.sleep_request), list(args.sleep))
|
||||
args.staging.mkdir(parents=True, exist_ok=True)
|
||||
# Deliberately a SIBLING of the staging directory, not inside it: staging is
|
||||
# rsynced wholesale into the archive, and a dry run caught this file being
|
||||
# published to the archive root.
|
||||
config_path = args.staging.parent / f"{args.staging.name}.gdl-config.json"
|
||||
|
||||
kinds = {k.strip() for k in args.only.split(",") if k.strip()}
|
||||
unknown = kinds - set(ALL_KINDS)
|
||||
if unknown:
|
||||
print(f"unknown surface(s): {', '.join(sorted(unknown))}", file=sys.stderr)
|
||||
return 2
|
||||
|
||||
state_path = (args.archive_db.with_suffix(".state.json") if args.archive_db
|
||||
else args.staging.parent / f"{args.staging.name}.state.json")
|
||||
state = SyncState(state_path)
|
||||
now = dt.datetime.now(dt.timezone.utc)
|
||||
now_ts, stamp = now.timestamp(), now.isoformat()
|
||||
min_interval = 0.0 if args.force else args.min_interval
|
||||
|
||||
plan: list[tuple[Profile, Source, bool]] = []
|
||||
skipped = 0
|
||||
for prof in selected:
|
||||
for src in prof.sources(kinds):
|
||||
fetch, seed, reason = plan_source(src, state, now_ts, min_interval)
|
||||
if not fetch:
|
||||
skipped += 1
|
||||
print(f" skip {prof.user}/{src.kind}: {reason}")
|
||||
continue
|
||||
if args.max_sources and len(plan) >= args.max_sources:
|
||||
skipped += 1
|
||||
continue
|
||||
plan.append((prof, src, seed))
|
||||
|
||||
print(f"profiles : {len(selected)}")
|
||||
print(f"surfaces : {','.join(k for k in ALL_KINDS if k in kinds)}")
|
||||
print(f"sources : {len(plan)} to sync, {skipped} skipped")
|
||||
print(f"pacing : {args.sleep_request[0]}-{args.sleep_request[1]}s between "
|
||||
f"requests, rate cap {args.rate}")
|
||||
print(f"staging : {args.staging}")
|
||||
print(f"publish : {args.publish}")
|
||||
print()
|
||||
|
||||
if not args.execute:
|
||||
for prof, src, seed in plan:
|
||||
dest = src.directory or "(per-highlight)"
|
||||
note = " [will seed]" if seed else ""
|
||||
print(f" {prof.user:<20} {src.kind:<11} -> {dest}{note}")
|
||||
print()
|
||||
print(" " + " ".join(rsync_command(args.staging, args.publish, True)))
|
||||
print("\ndry run; nothing fetched. pass --execute to run.")
|
||||
return 0
|
||||
|
||||
config_path.write_text(json.dumps(config, indent=2))
|
||||
probes = ProbeCache(state_path.with_suffix(".probes.json"),
|
||||
0.0 if args.force else args.probe_ttl)
|
||||
failures = 0
|
||||
|
||||
for prof, src, seed in plan:
|
||||
print(f"==> {prof.user} / {src.kind}")
|
||||
stage_dir = args.staging / (src.directory or ".")
|
||||
stage_dir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
# Prime the skip-archive from what the archive already holds, so
|
||||
# fetching into an empty staging directory pulls only what is missing.
|
||||
# Done once per source, ever: afterwards the archive DB records
|
||||
# everything gallery-dl has seen and no listing pass is needed.
|
||||
if seed and args.archive_db:
|
||||
try:
|
||||
live = probes.get(src.url, now_ts)
|
||||
if live is None:
|
||||
live = probe_live(src, config_path, args.cookies)
|
||||
probes.put(src.url, live, stamp)
|
||||
probes.save()
|
||||
else:
|
||||
print(f" reusing {len(live)} cached listing items")
|
||||
held = index_existing(index.listing(prof.user))
|
||||
seeded = seed_archive_db(args.archive_db, held, live,
|
||||
src.subcategory)
|
||||
print(f" seeded {seeded} of {len(live)} live items")
|
||||
state.mark_seeded(src.url, stamp)
|
||||
state.save()
|
||||
except subprocess.CalledProcessError as exc:
|
||||
failures += 1
|
||||
print(f" probe FAILED: {exc}", file=sys.stderr)
|
||||
continue
|
||||
|
||||
cmd = gdl_command(src, args.staging, config_path, args.cookies,
|
||||
args.archive_db)
|
||||
result = subprocess.run(cmd)
|
||||
if result.returncode != 0:
|
||||
failures += 1
|
||||
# Keep going: one private or renamed profile must not abort the run.
|
||||
print(f" FAILED (exit {result.returncode})", file=sys.stderr)
|
||||
else:
|
||||
# Recorded even for an empty fetch: the request was still spent.
|
||||
state.mark_fetched(src.url, stamp)
|
||||
state.save()
|
||||
|
||||
# Publish once, at the end, so a partially-fetched profile never reaches
|
||||
# the archive mid-run. Only ever adds -- see rsync_command.
|
||||
print("\n==> publish")
|
||||
if publish(args.staging, args.publish, dry_run=False) != 0:
|
||||
failures += 1
|
||||
|
||||
print(f"\ndone; {failures} step(s) failed")
|
||||
return 1 if failures else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,277 @@
|
||||
/**
|
||||
* Generate JDownloader2 .crawljob files for the archives on disk.
|
||||
*
|
||||
* The manual flow is: paste a profile URL into JDownloader, paste the /reels
|
||||
* URL separately (the profile page misses some reels), and set the output
|
||||
* folder by hand — times however many profiles you keep. This emits one
|
||||
* crawljob per source with the folder already pointed at the right directory,
|
||||
* so JDownloader's folder-watch picks the whole batch up at once.
|
||||
*
|
||||
* Profiles and their sidecar directories are derived with the same grouping
|
||||
* logic the server uses, so the output folders always match what the viewer
|
||||
* expects to find.
|
||||
*
|
||||
* Only posts and reels are emitted. Story and highlight URLs can't be rebuilt
|
||||
* from a directory name — highlights need their numeric id and stories expire —
|
||||
* so those stay manual.
|
||||
*
|
||||
* Crawljob format verified against JDownloader's own docs for the extension:
|
||||
* src/org/jdownloader/extensions/folderwatchV2/explain.txt. JDownloader
|
||||
* develops on SVN; read it via the daily mirror at
|
||||
* https://github.com/mycodedoesnotcompile2/jdownloader_mirror (svn_trunk/),
|
||||
* not one of the abandoned GitHub copies — several are a decade stale.
|
||||
*
|
||||
* Entries are separated by `->NEW ENTRY<-` and any property may be omitted.
|
||||
* There is also a `setBeforePackagizerEnabled` companion to
|
||||
* `overwritePackagizerEnabled`, if the Packagizer ever needs to see these
|
||||
* values before they're applied.
|
||||
*
|
||||
* Usage:
|
||||
* npx tsx scripts/jd2-sync.ts --archives <dir> [options]
|
||||
*
|
||||
* --archives <dir> Archive root to scan (default: $ARCHIVES_DIR)
|
||||
* --out <dir> JDownloader folder-watch directory to write into
|
||||
* --download-base <dir> Root path as *JDownloader* sees it, when it runs on
|
||||
* a different machine than this script (e.g. a mapped
|
||||
* drive). Defaults to --archives.
|
||||
* --user <name> Only this profile (repeatable)
|
||||
* --skip <name> Never emit jobs for this directory (repeatable).
|
||||
* Also read from a `.jd2ignore` file in the archive
|
||||
* root, one name per line.
|
||||
* --chunks <n> Connections per file (default 1: multi-chunk ranged
|
||||
* requests are the one CDN pattern that doesn't look
|
||||
* like a browser)
|
||||
* --auto-start Start downloads immediately instead of parking them
|
||||
* in the LinkGrabber for review
|
||||
* --all-reels Emit a reels job even where no reels directory
|
||||
* exists yet
|
||||
* --dry-run Print the crawljob instead of writing it
|
||||
*/
|
||||
|
||||
import fs from 'fs';
|
||||
import path from 'path';
|
||||
import { groupArchiveDirectories, ArchiveSource } from '../src/lib/archive-grouping.js';
|
||||
|
||||
interface Options {
|
||||
archives: string;
|
||||
out: string | null;
|
||||
downloadBase: string;
|
||||
users: string[];
|
||||
skip: Set<string>;
|
||||
chunks: number;
|
||||
autoStart: boolean;
|
||||
allReels: boolean;
|
||||
dryRun: boolean;
|
||||
}
|
||||
|
||||
const parseArgs = (argv: string[]): Options => {
|
||||
const opts: Options = {
|
||||
archives: process.env.ARCHIVES_DIR ?? '',
|
||||
out: null,
|
||||
downloadBase: '',
|
||||
users: [],
|
||||
skip: new Set(),
|
||||
chunks: 1,
|
||||
autoStart: false,
|
||||
allReels: false,
|
||||
dryRun: false,
|
||||
};
|
||||
|
||||
for (let i = 0; i < argv.length; i++) {
|
||||
const arg = argv[i];
|
||||
const next = () => argv[++i];
|
||||
switch (arg) {
|
||||
case '--archives': opts.archives = path.resolve(next()); break;
|
||||
case '--out': opts.out = path.resolve(next()); break;
|
||||
case '--download-base': opts.downloadBase = next(); break;
|
||||
case '--user': opts.users.push(next()); break;
|
||||
case '--skip': opts.skip.add(next()); break;
|
||||
case '--chunks': opts.chunks = parseInt(next(), 10); break;
|
||||
case '--auto-start': opts.autoStart = true; break;
|
||||
case '--all-reels': opts.allReels = true; break;
|
||||
case '--dry-run': opts.dryRun = true; break;
|
||||
case '--help': case '-h': printUsage(); process.exit(0);
|
||||
default:
|
||||
console.error(`Unknown argument: ${arg}`);
|
||||
process.exit(1);
|
||||
}
|
||||
}
|
||||
|
||||
if (!opts.archives) {
|
||||
console.error('No archive root. Pass --archives <dir> or set ARCHIVES_DIR.');
|
||||
process.exit(1);
|
||||
}
|
||||
if (!opts.downloadBase) opts.downloadBase = opts.archives;
|
||||
if (!opts.out && !opts.dryRun) {
|
||||
console.error('No destination. Pass --out <folder-watch dir>, or --dry-run to preview.');
|
||||
process.exit(1);
|
||||
}
|
||||
return opts;
|
||||
};
|
||||
|
||||
const printUsage = () => {
|
||||
const header = readHeaderComment();
|
||||
console.log(header);
|
||||
};
|
||||
|
||||
/** Print the usage block from this file's own header comment. */
|
||||
const readHeaderComment = () => {
|
||||
try {
|
||||
const self = fs.readFileSync(new URL(import.meta.url), 'utf8');
|
||||
const usage = self.slice(self.indexOf(' * Usage:'), self.indexOf(' */'));
|
||||
return usage.split('\n').map(l => l.replace(/^ \* ?/, '')).join('\n');
|
||||
} catch {
|
||||
return 'See the comment at the top of scripts/jd2-sync.ts';
|
||||
}
|
||||
};
|
||||
|
||||
/**
|
||||
* JDownloader escapes nothing in crawljob values, so a stray newline would
|
||||
* silently split a property. Paths with spaces are fine as-is.
|
||||
*/
|
||||
const sanitise = (value: string) => value.replace(/[\r\n]+/g, ' ').trim();
|
||||
|
||||
/**
|
||||
* Instagram usernames are 1–30 characters of letters, digits, dots and
|
||||
* underscores. Archive roots also collect directories that aren't profiles at
|
||||
* all — tool output, exports from other services — and pointing a crawl at
|
||||
* those spends requests on instagram.com to be told the profile doesn't exist.
|
||||
* That's the exact traffic worth not spending.
|
||||
*/
|
||||
const USERNAME_RE = /^[A-Za-z0-9._]{1,30}$/;
|
||||
|
||||
/** Directory names to skip, from `.jd2ignore` in the archive root. */
|
||||
const readIgnoreFile = (archives: string): string[] => {
|
||||
try {
|
||||
return fs.readFileSync(path.join(archives, '.jd2ignore'), 'utf8')
|
||||
.split('\n').map(l => l.trim()).filter(l => l && !l.startsWith('#'));
|
||||
} catch {
|
||||
return [];
|
||||
}
|
||||
};
|
||||
|
||||
interface Job {
|
||||
user: string;
|
||||
kind: 'posts' | 'reels';
|
||||
url: string;
|
||||
packageName: string;
|
||||
downloadFolder: string;
|
||||
fileCount: number | null;
|
||||
}
|
||||
|
||||
const buildJobs = (opts: Options): Job[] => {
|
||||
const dirNames = fs.readdirSync(opts.archives, { withFileTypes: true })
|
||||
.filter(e => e.isDirectory() && !/^[.@_]/.test(e.name))
|
||||
.map(e => e.name);
|
||||
|
||||
const groups = groupArchiveDirectories(dirNames);
|
||||
const jobs: Job[] = [];
|
||||
const skipped: string[] = [];
|
||||
|
||||
for (const name of readIgnoreFile(opts.archives)) opts.skip.add(name);
|
||||
|
||||
const countFiles = (dir: string): number | null => {
|
||||
try {
|
||||
return fs.readdirSync(path.join(opts.archives, dir)).length;
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
};
|
||||
|
||||
// JDownloader must be given the path *it* can see, which differs from the
|
||||
// scan path whenever the archive lives on a share.
|
||||
const downloadFolderFor = (dir: string) =>
|
||||
opts.downloadBase.includes('\\')
|
||||
? `${opts.downloadBase.replace(/\\$/, '')}\\${dir}`
|
||||
: path.posix.join(opts.downloadBase, dir);
|
||||
|
||||
for (const [user, sources] of [...groups].sort(([a], [b]) => a.localeCompare(b))) {
|
||||
if (opts.users.length && !opts.users.includes(user)) continue;
|
||||
|
||||
if (opts.skip.has(user)) { skipped.push(`${user} (ignored)`); continue; }
|
||||
if (!USERNAME_RE.test(user)) { skipped.push(`${user} (not a username)`); continue; }
|
||||
|
||||
const has = (kind: ArchiveSource['kind']) => sources.find(s => s.kind === kind);
|
||||
const base = has('posts');
|
||||
if (!base) continue; // sidecar-only group: nothing sensible to point a URL at
|
||||
|
||||
jobs.push({
|
||||
user, kind: 'posts',
|
||||
url: `https://www.instagram.com/${encodeURIComponent(user)}/`,
|
||||
packageName: base.dir,
|
||||
downloadFolder: downloadFolderFor(base.dir),
|
||||
fileCount: countFiles(base.dir),
|
||||
});
|
||||
|
||||
const reels = has('reels');
|
||||
if (reels || opts.allReels) {
|
||||
const dir = reels?.dir ?? `${user} - reels`;
|
||||
jobs.push({
|
||||
user, kind: 'reels',
|
||||
url: `https://www.instagram.com/${encodeURIComponent(user)}/reels/`,
|
||||
packageName: dir,
|
||||
downloadFolder: downloadFolderFor(dir),
|
||||
fileCount: reels ? countFiles(dir) : null,
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
if (skipped.length) {
|
||||
console.error(`Skipped ${skipped.length} director${skipped.length === 1 ? 'y' : 'ies'}:`);
|
||||
for (const s of skipped) console.error(` - ${s}`);
|
||||
console.error('');
|
||||
}
|
||||
|
||||
return jobs;
|
||||
};
|
||||
|
||||
const renderCrawljob = (jobs: Job[], opts: Options): string =>
|
||||
jobs.map(job => [
|
||||
`text=${sanitise(job.url)}`,
|
||||
`packageName=${sanitise(job.packageName)}`,
|
||||
`downloadFolder=${sanitise(job.downloadFolder)}`,
|
||||
`chunks=${opts.chunks}`,
|
||||
// Without this a Packagizer rule can override downloadFolder and scatter
|
||||
// files away from the directory the viewer reads.
|
||||
'overwritePackagizerEnabled=TRUE',
|
||||
`autoStart=${opts.autoStart ? 'TRUE' : 'FALSE'}`,
|
||||
`autoConfirm=${opts.autoStart ? 'TRUE' : 'FALSE'}`,
|
||||
'enabled=TRUE',
|
||||
`comment=instaarchive jd2-sync (${job.kind})`,
|
||||
].join('\n')).join('\n->NEW ENTRY<-\n');
|
||||
|
||||
const main = () => {
|
||||
const opts = parseArgs(process.argv.slice(2));
|
||||
const jobs = buildJobs(opts);
|
||||
|
||||
if (!jobs.length) {
|
||||
console.error('No profiles matched.');
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
console.error(`Archive root : ${opts.archives}`);
|
||||
console.error(`JD sees root : ${opts.downloadBase}`);
|
||||
console.error(`Jobs : ${jobs.length} (${new Set(jobs.map(j => j.user)).size} profiles)\n`);
|
||||
for (const job of jobs) {
|
||||
const count = job.fileCount === null ? 'new' : `${job.fileCount} files`;
|
||||
console.error(` ${job.kind.padEnd(5)} ${job.user.padEnd(24)} -> ${job.packageName} (${count})`);
|
||||
}
|
||||
console.error('');
|
||||
|
||||
const body = renderCrawljob(jobs, opts);
|
||||
|
||||
if (opts.dryRun || !opts.out) {
|
||||
console.log(body);
|
||||
return;
|
||||
}
|
||||
|
||||
fs.mkdirSync(opts.out, { recursive: true });
|
||||
const file = path.join(opts.out, `instaarchive-${new Date().toISOString().replace(/[:.]/g, '-')}.crawljob`);
|
||||
fs.writeFileSync(file, body, 'utf8');
|
||||
console.error(`Wrote ${file}`);
|
||||
console.error(opts.autoStart
|
||||
? 'Downloads will start automatically.'
|
||||
: 'Links land in the LinkGrabber for review; start them when ready.');
|
||||
};
|
||||
|
||||
main();
|
||||
@@ -0,0 +1,231 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Tests for the request-budget logic in gdl-sync.py.
|
||||
|
||||
python3 -m unittest discover -s scripts -p 'test_*.py'
|
||||
|
||||
Deliberately stdlib-only, so it runs anywhere the sync itself runs. What is
|
||||
covered here is the part that decides whether to spend a request — the part
|
||||
whose absence got the archive's Instagram account suspended.
|
||||
"""
|
||||
|
||||
import datetime as dt
|
||||
import importlib.util
|
||||
import json
|
||||
import sys
|
||||
import tempfile
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
_spec = importlib.util.spec_from_file_location(
|
||||
"gdl_sync", Path(__file__).with_name("gdl-sync.py"))
|
||||
gdl = importlib.util.module_from_spec(_spec)
|
||||
sys.modules["gdl_sync"] = gdl
|
||||
_spec.loader.exec_module(gdl)
|
||||
|
||||
NOW = dt.datetime(2026, 8, 18, 12, 0, tzinfo=dt.timezone.utc)
|
||||
NOW_TS = NOW.timestamp()
|
||||
|
||||
|
||||
def ago(hours: float) -> str:
|
||||
return (NOW - dt.timedelta(hours=hours)).isoformat()
|
||||
|
||||
|
||||
class SourceSelection(unittest.TestCase):
|
||||
def test_only_stories_is_a_single_cheap_source(self):
|
||||
srcs = gdl.Profile("u").sources(gdl.STORIES_ONLY)
|
||||
self.assertEqual([s.kind for s in srcs], ["stories"])
|
||||
self.assertEqual(srcs[0].directory, "story - u")
|
||||
|
||||
def test_full_sync_covers_every_surface(self):
|
||||
srcs = gdl.Profile("u").sources(set(gdl.ALL_KINDS))
|
||||
self.assertEqual([s.kind for s in srcs], list(gdl.ALL_KINDS))
|
||||
|
||||
def test_reels_and_stories_go_to_their_own_directories(self):
|
||||
by_kind = {s.kind: s for s in gdl.Profile("u").sources(set(gdl.ALL_KINDS))}
|
||||
self.assertEqual(by_kind["posts"].directory, "u")
|
||||
self.assertEqual(by_kind["reels"].directory, "u - reels")
|
||||
# Highlights derive their directory from the title mid-extraction.
|
||||
self.assertEqual(by_kind["highlights"].directory, "")
|
||||
|
||||
|
||||
class PlanSource(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self.tmp = tempfile.TemporaryDirectory()
|
||||
self.state = gdl.SyncState(Path(self.tmp.name) / "state.json")
|
||||
self.posts = gdl.Profile("u").sources({"posts"})[0]
|
||||
self.stories = gdl.Profile("u").sources({"stories"})[0]
|
||||
|
||||
def tearDown(self):
|
||||
self.tmp.cleanup()
|
||||
|
||||
def test_first_run_seeds(self):
|
||||
fetch, seed, _ = gdl.plan_source(self.posts, self.state, NOW_TS, 20)
|
||||
self.assertTrue(fetch)
|
||||
self.assertTrue(seed)
|
||||
|
||||
def test_seeding_happens_only_once(self):
|
||||
self.state.mark_seeded(self.posts.url, ago(720))
|
||||
fetch, seed, reason = gdl.plan_source(self.posts, self.state, NOW_TS, 20)
|
||||
self.assertTrue(fetch)
|
||||
self.assertFalse(seed, "a seeded source must never be re-probed")
|
||||
self.assertIn("already seeded", reason)
|
||||
|
||||
def test_stories_never_seed(self):
|
||||
# A story cannot be in the archive before it is fetched, so probing
|
||||
# would double the cost of the cheapest surface for no benefit.
|
||||
_, seed, reason = gdl.plan_source(self.stories, self.state, NOW_TS, 20)
|
||||
self.assertFalse(seed)
|
||||
self.assertIn("no seed", reason)
|
||||
|
||||
def test_recent_fetch_is_refused(self):
|
||||
self.state.mark_fetched(self.posts.url, ago(3))
|
||||
fetch, _, reason = gdl.plan_source(self.posts, self.state, NOW_TS, 20)
|
||||
self.assertFalse(fetch)
|
||||
self.assertIn("under the", reason)
|
||||
|
||||
def test_an_old_fetch_is_allowed_again(self):
|
||||
self.state.mark_fetched(self.posts.url, ago(30))
|
||||
fetch, _, _ = gdl.plan_source(self.posts, self.state, NOW_TS, 20)
|
||||
self.assertTrue(fetch)
|
||||
|
||||
def test_daily_stories_pass_a_20h_floor(self):
|
||||
# The cadence this is built for: once a day, every day.
|
||||
self.state.mark_fetched(self.stories.url, ago(24))
|
||||
fetch, _, _ = gdl.plan_source(self.stories, self.state, NOW_TS, 20)
|
||||
self.assertTrue(fetch)
|
||||
|
||||
def test_force_disables_the_floor(self):
|
||||
self.state.mark_fetched(self.posts.url, ago(1))
|
||||
fetch, _, _ = gdl.plan_source(self.posts, self.state, NOW_TS, 0.0)
|
||||
self.assertTrue(fetch)
|
||||
|
||||
def test_the_aborted_run_scenario(self):
|
||||
"""
|
||||
Yesterday's failure: a run died mid-way and the restart re-enumerated
|
||||
every profile. Seeded-but-not-fetched must not re-probe.
|
||||
"""
|
||||
self.state.mark_seeded(self.posts.url, ago(0.5))
|
||||
fetch, seed, _ = gdl.plan_source(self.posts, self.state, NOW_TS, 20)
|
||||
self.assertTrue(fetch, "the fetch still needs to happen")
|
||||
self.assertFalse(seed, "but the listing pass must not be paid for twice")
|
||||
|
||||
|
||||
class StatePersistence(unittest.TestCase):
|
||||
def test_state_survives_a_reload(self):
|
||||
with tempfile.TemporaryDirectory() as d:
|
||||
path = Path(d) / "state.json"
|
||||
a = gdl.SyncState(path)
|
||||
a.mark_seeded("https://x/", ago(1))
|
||||
a.mark_fetched("https://x/", ago(1))
|
||||
a.save()
|
||||
b = gdl.SyncState(path)
|
||||
self.assertFalse(b.needs_seed("https://x/"))
|
||||
self.assertEqual(b.last_fetch("https://x/"), ago(1))
|
||||
|
||||
def test_a_corrupt_state_file_never_blocks_a_sync(self):
|
||||
with tempfile.TemporaryDirectory() as d:
|
||||
path = Path(d) / "state.json"
|
||||
path.write_text("{ not json")
|
||||
self.assertTrue(gdl.SyncState(path).needs_seed("https://x/"))
|
||||
|
||||
|
||||
class ProbeCaching(unittest.TestCase):
|
||||
def test_fresh_entries_are_reused_and_stale_ones_are_not(self):
|
||||
with tempfile.TemporaryDirectory() as d:
|
||||
cache = gdl.ProbeCache(Path(d) / "p.json", ttl_hours=24)
|
||||
cache.put("https://x/", [{"shortcode": "A", "post_shortcode": "A",
|
||||
"num": 1, "media_id": "1"}], ago(1))
|
||||
self.assertEqual(len(cache.get("https://x/", NOW_TS)), 1)
|
||||
|
||||
cache.put("https://y/", [{"shortcode": "B", "post_shortcode": "B",
|
||||
"num": 1, "media_id": "2"}], ago(48))
|
||||
self.assertIsNone(cache.get("https://y/", NOW_TS))
|
||||
|
||||
def test_cache_keeps_only_the_fields_seeding_needs(self):
|
||||
with tempfile.TemporaryDirectory() as d:
|
||||
path = Path(d) / "p.json"
|
||||
cache = gdl.ProbeCache(path, ttl_hours=24)
|
||||
cache.put("https://x/", [{"shortcode": "A", "post_shortcode": "A",
|
||||
"num": 1, "media_id": "1",
|
||||
"description": "x" * 5000}], ago(0))
|
||||
cache.save()
|
||||
self.assertNotIn("description", path.read_text())
|
||||
|
||||
def test_a_miss_is_reported_rather_than_guessed(self):
|
||||
with tempfile.TemporaryDirectory() as d:
|
||||
cache = gdl.ProbeCache(Path(d) / "p.json", ttl_hours=24)
|
||||
self.assertIsNone(cache.get("https://never-seen/", NOW_TS))
|
||||
|
||||
|
||||
class Seeding(unittest.TestCase):
|
||||
"""The bug that seeded 5 of 2275: matching the wrong shortcode field."""
|
||||
|
||||
def test_posts_are_keyed_by_post_shortcode(self):
|
||||
item = {"shortcode": "childcode", "post_shortcode": "POSTCODE",
|
||||
"num": 2, "media_id": "9"}
|
||||
self.assertEqual(gdl.live_key(item, "posts"), ("POSTCODE", 2))
|
||||
|
||||
def test_stories_are_keyed_by_the_per_item_shortcode(self):
|
||||
item = {"shortcode": "ITEMCODE", "post_shortcode": "reelid",
|
||||
"num": 3, "media_id": "9"}
|
||||
self.assertEqual(gdl.live_key(item, "stories"), ("ITEMCODE", 1))
|
||||
self.assertEqual(gdl.live_key(item, "highlights"), ("ITEMCODE", 1))
|
||||
|
||||
def test_index_existing_normalises_a_missing_index_to_one(self):
|
||||
held = gdl.index_existing([
|
||||
"u/2023-04-19_u - ABC.mp4",
|
||||
"u/2023-04-12_u - DEF - 3.jpg",
|
||||
"u/2023-04-12_u - DEF.txt", # sidecars are not media
|
||||
"u/2023-04-12_u - DEF.json",
|
||||
])
|
||||
self.assertEqual(held, {("ABC", 1), ("DEF", 3)})
|
||||
|
||||
def test_seeding_marks_only_what_is_already_held(self):
|
||||
with tempfile.TemporaryDirectory() as d:
|
||||
db = Path(d) / "a.db"
|
||||
live = [
|
||||
{"post_shortcode": "HELD", "shortcode": "x", "num": 1, "media_id": "11"},
|
||||
{"post_shortcode": "NEW", "shortcode": "y", "num": 1, "media_id": "22"},
|
||||
]
|
||||
n = gdl.seed_archive_db(db, {("HELD", 1)}, live, "posts")
|
||||
self.assertEqual(n, 1)
|
||||
import sqlite3
|
||||
rows = {r[0] for r in sqlite3.connect(db).execute(
|
||||
"SELECT entry FROM archive")}
|
||||
self.assertEqual(rows, {"instagram11"})
|
||||
|
||||
|
||||
class Publishing(unittest.TestCase):
|
||||
def test_publish_only_ever_adds(self):
|
||||
cmd = gdl.rsync_command(Path("/stage"), "host:/archives", dry_run=False)
|
||||
self.assertIn("--ignore-existing", cmd)
|
||||
self.assertNotIn("--delete", cmd)
|
||||
|
||||
def test_tooling_files_are_excluded_from_the_archive(self):
|
||||
cmd = " ".join(gdl.rsync_command(Path("/stage"), "/dest", dry_run=True))
|
||||
for pattern in ("gdl-sync*.json", "*.db"):
|
||||
self.assertIn(pattern, cmd)
|
||||
self.assertIn("--dry-run", cmd)
|
||||
|
||||
|
||||
class UrlsFile(unittest.TestCase):
|
||||
def test_reads_every_form_a_person_might_paste(self):
|
||||
with tempfile.TemporaryDirectory() as d:
|
||||
p = Path(d) / "urls.txt"
|
||||
p.write_text(
|
||||
"# comment\n"
|
||||
"https://www.instagram.com/a/\n"
|
||||
"https://instagram.com/b\n"
|
||||
"www.instagram.com/c/\n"
|
||||
"d\n"
|
||||
" e # trailing\n"
|
||||
"\n"
|
||||
"https://www.instagram.com/a/\n" # duplicate
|
||||
"https://www.instagram.com/p/ABC123/\n" # a post, not a profile
|
||||
"not a username\n")
|
||||
self.assertEqual(gdl.read_urls_file(p), ["a", "b", "c", "d", "e"])
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main(verbosity=2)
|
||||
+12
-13
@@ -19,6 +19,7 @@ import { motion, AnimatePresence } from 'motion/react';
|
||||
import { cn } from './lib/utils';
|
||||
import { PRESS, prefersReducedMotion } from './lib/motion';
|
||||
import { buildPath, findPostBySlug, parseRoute, postSlug, tabForSource } from './lib/routing';
|
||||
import { postsForTab } from './lib/post-tabs';
|
||||
import { LocalArchiveFile, RemoteArchiveFile } from './lib/archive-files';
|
||||
import {
|
||||
deleteCachedArchive,
|
||||
@@ -172,20 +173,18 @@ export default function App() {
|
||||
const clearCache = async (name: string) => { await deleteCachedArchive(name); await refreshCachedArchives(); };
|
||||
|
||||
/**
|
||||
* Archives with a `- reels` sidecar directory say outright which posts are
|
||||
* reels; only fall back to the "lone video" heuristic for archives that have
|
||||
* no such directory.
|
||||
* The grid shows everything, reels included, and the Reels tab is a filtered
|
||||
* view of the same set — see src/lib/post-tabs.ts for the reel test and for
|
||||
* why a reel can arrive on disk twice.
|
||||
*/
|
||||
const hasReelSource = useMemo(() => allPosts.some(p => p.source === 'reels'), [allPosts]);
|
||||
const isReel = useCallback((p: Post) => (
|
||||
hasReelSource ? p.source === 'reels' : p.media.length === 1 && p.media[0].type === 'video'
|
||||
), [hasReelSource]);
|
||||
const filteredPosts = useMemo(() => postsForTab(allPosts, activeTab), [allPosts, activeTab]);
|
||||
|
||||
const filteredPosts = useMemo(() => {
|
||||
if (activeTab === 'reels') return allPosts.filter(isReel);
|
||||
if (activeTab === 'posts') return allPosts.filter(p => !isReel(p));
|
||||
return [];
|
||||
}, [allPosts, activeTab, isReel]);
|
||||
/**
|
||||
* Instagram's "N posts" counter equals what its grid holds, so count the grid
|
||||
* rather than `allPosts` — the raw list still holds both copies of any post
|
||||
* fetched into two directories.
|
||||
*/
|
||||
const totalPosts = useMemo(() => postsForTab(allPosts, 'posts').length, [allPosts]);
|
||||
|
||||
/** Story highlights, grouped into the circles shown under the bio. */
|
||||
const highlightGroups = useMemo(() => {
|
||||
@@ -527,7 +526,7 @@ export default function App() {
|
||||
{allProfilePics.length > 1 && <button onClick={cycleProfilePic} className="bg-gray-100 hover:bg-gray-200 px-4 py-1.5 rounded-lg text-sm font-semibold transition-colors flex items-center gap-2 text-black"><Layers size={16} />Next Profile Pic</button>}
|
||||
</div>
|
||||
</div>
|
||||
<div className="flex justify-center md:justify-start gap-10 text-sm md:text-base text-black"><div><span className="font-semibold text-black/80 text-black">{allPosts.length}</span> posts</div><div><span className="font-semibold text-black/80 text-black">{(followerCount || 0).toLocaleString()}</span> followers</div><div><span className="font-semibold text-black/80 text-black">{(followingCount || 0).toLocaleString()}</span> following</div></div>
|
||||
<div className="flex justify-center md:justify-start gap-10 text-sm md:text-base text-black"><div><span className="font-semibold text-black/80 text-black">{totalPosts.toLocaleString()}</span> posts</div><div><span className="font-semibold text-black/80 text-black">{(followerCount || 0).toLocaleString()}</span> followers</div><div><span className="font-semibold text-black/80 text-black">{(followingCount || 0).toLocaleString()}</span> following</div></div>
|
||||
<div className="space-y-1 text-black/80 text-black"><div className="font-semibold text-black">{fullName || `@${username}`}</div><div className="text-gray-600 whitespace-pre-wrap max-w-sm mx-auto md:mx-0 text-sm md:text-base text-black">{bio || 'Archived profile viewer for local files.'}</div>{externalUrl && <a href={externalUrl} target="_blank" rel="noopener noreferrer" className="text-blue-900 font-semibold text-sm block hover:underline truncate max-w-[250px] text-black">{externalUrl.replace(/^https?:\/\/(www\.)?/, '')}</a>}</div>
|
||||
</div>
|
||||
</header>
|
||||
|
||||
@@ -4,6 +4,8 @@ import { XzReadableStream } from 'xz-decompress';
|
||||
import { ArchiveFile, CacheData, Post, ServerArchive } from '../types';
|
||||
import { setCachedArchive, getDirectoryHandle } from '../lib/archive-cache';
|
||||
import { parseArchiveFilename, scopedPostId, EXPORT_RE, INSTALOADER_RE } from '../lib/archive-patterns';
|
||||
import { isGalleryDlSidecar, sidecarDate, sidecarIsReel } from '../lib/gallery-dl-sidecar';
|
||||
import { DateSource, shouldReplaceDate } from '../lib/post-dates';
|
||||
|
||||
const hasDirectoryHandle = async (name: string) => Boolean(await getDirectoryHandle(name));
|
||||
|
||||
@@ -142,6 +144,17 @@ export const useArchiveScanner = (
|
||||
|
||||
try {
|
||||
const postsMap = new Map<string, Partial<Post>>();
|
||||
/**
|
||||
* Which source supplied each post's date, so a better one can replace it.
|
||||
* Sidecar beats filename beats mtime — see src/lib/post-dates.ts.
|
||||
*/
|
||||
const dateSources = new Map<string, DateSource>();
|
||||
const applyDate = (postId: string, post: Partial<Post>, date: string, source: DateSource) => {
|
||||
const current = post.date ? { date: post.date, source: dateSources.get(postId) ?? 'mtime' } : undefined;
|
||||
if (!shouldReplaceDate(current, { date, source })) return;
|
||||
post.date = date;
|
||||
dateSources.set(postId, source);
|
||||
};
|
||||
const mediaFilesMap = new Map<string, ArchiveFile>();
|
||||
const discoveredProfilePics: { name: string, url: string }[] = [];
|
||||
const allImageFiles: ArchiveFile[] = [];
|
||||
@@ -300,13 +313,26 @@ export const useArchiveScanner = (
|
||||
}
|
||||
else if (isStory) post.isStory = true;
|
||||
|
||||
// Files describing one post are scanned in directory order, not in
|
||||
// order of trustworthiness, so every date goes through the ranking
|
||||
// in post-dates.ts rather than last-write-wins.
|
||||
applyDate(postId, post, date, parsed.dateFromMtime ? 'mtime' : 'filename');
|
||||
|
||||
const lowerExt = ext.toLowerCase();
|
||||
if (lowerExt === 'txt') {
|
||||
try { post.caption = await file.text(); } catch(e) {}
|
||||
} else if (lowerExt === 'json' || lowerName.endsWith('.json.xz')) {
|
||||
try {
|
||||
const data = lowerName.endsWith('.xz') ? await parseXZFile(file) : JSON.parse(await file.text());
|
||||
if (data) {
|
||||
if (isGalleryDlSidecar(data)) {
|
||||
// The only format that states what a post is rather than
|
||||
// leaving it to be inferred from filenames.
|
||||
if (data.description) post.caption = data.description;
|
||||
const reel = sidecarIsReel(data);
|
||||
if (reel !== undefined) post.isReel = reel;
|
||||
if (data.type === 'story') post.isStory = true;
|
||||
applyDate(postId, post, sidecarDate(data), 'sidecar');
|
||||
} else if (data) {
|
||||
const node = data.node || data; const iphone = node.iphone_struct || {};
|
||||
const captionText = node.edge_media_to_caption?.edges?.[0]?.node?.text || node.caption?.text || iphone.caption?.text || '';
|
||||
if (captionText) post.caption = captionText;
|
||||
|
||||
@@ -3,39 +3,39 @@ import { classifyDirectory, groupArchiveDirectories } from './archive-grouping';
|
||||
|
||||
describe('classifyDirectory', () => {
|
||||
it('treats a bare profile directory as the base', () => {
|
||||
expect(classifyDirectory('4utumn07')).toEqual({
|
||||
owner: '4utumn07',
|
||||
source: { kind: 'posts', dir: '4utumn07' },
|
||||
expect(classifyDirectory('0ct0ber19')).toEqual({
|
||||
owner: '0ct0ber19',
|
||||
source: { kind: 'posts', dir: '0ct0ber19' },
|
||||
});
|
||||
});
|
||||
|
||||
it('recognises a reels sidecar', () => {
|
||||
expect(classifyDirectory('4utumn07 - reels')).toEqual({
|
||||
owner: '4utumn07',
|
||||
source: { kind: 'reels', dir: '4utumn07 - reels' },
|
||||
expect(classifyDirectory('0ct0ber19 - reels')).toEqual({
|
||||
owner: '0ct0ber19',
|
||||
source: { kind: 'reels', dir: '0ct0ber19 - reels' },
|
||||
});
|
||||
});
|
||||
|
||||
it('recognises a stories sidecar', () => {
|
||||
expect(classifyDirectory('story - dawn_petal')).toEqual({
|
||||
owner: 'dawn_petal',
|
||||
source: { kind: 'stories', dir: 'story - dawn_petal' },
|
||||
expect(classifyDirectory('story - cher_ryppo')).toEqual({
|
||||
owner: 'cher_ryppo',
|
||||
source: { kind: 'stories', dir: 'story - cher_ryppo' },
|
||||
});
|
||||
});
|
||||
|
||||
it('splits highlight owner from title', () => {
|
||||
const { owner, source } = classifyDirectory('story highlights - 4utumn07 - Sunstory');
|
||||
expect(owner).toBe('4utumn07');
|
||||
const { owner, source } = classifyDirectory('story highlights - 0ct0ber19 - Heestory');
|
||||
expect(owner).toBe('0ct0ber19');
|
||||
expect(source.kind).toBe('highlight');
|
||||
expect(source.title).toBe('Sunstory');
|
||||
expect(source.title).toBe('Heestory');
|
||||
});
|
||||
|
||||
it.each([
|
||||
['story highlights - theoldlyricmuseinsta - 💙1999-2005 era', 'theoldlyricmuseinsta', '💙1999-2005 era'],
|
||||
['story highlights - member_theworld - [Bracket]', 'member_theworld', '[Bracket]'],
|
||||
['story highlights - official_band - Tour Schedule', 'official_band', 'Tour Schedule'],
|
||||
['story highlights - 4utumn07 - Sketching⠀', '4utumn07', 'Sketching⠀'],
|
||||
['story highlights - official_band - A.B.C', 'official_band', 'A.B.C'],
|
||||
['story highlights - theoldtaylorswiftinsta - 💙2014-1989 era', 'theoldtaylorswiftinsta', '💙2014-1989 era'],
|
||||
['story highlights - heejin_theworld - [Dall]', 'heejin_theworld', '[Dall]'],
|
||||
['story highlights - official_artms - Cosmo Schedule', 'official_artms', 'Cosmo Schedule'],
|
||||
['story highlights - 0ct0ber19 - Drawheeing⠀', '0ct0ber19', 'Drawheeing⠀'],
|
||||
['story highlights - official_artms - G.C.I', 'official_artms', 'G.C.I'],
|
||||
])('handles real-world title %s', (dir, owner, title) => {
|
||||
const result = classifyDirectory(dir);
|
||||
expect(result.owner).toBe(owner);
|
||||
@@ -57,25 +57,25 @@ describe('classifyDirectory', () => {
|
||||
|
||||
describe('groupArchiveDirectories', () => {
|
||||
const dirs = [
|
||||
'4utumn07',
|
||||
'4utumn07 - reels',
|
||||
'story - 4utumn07',
|
||||
'story highlights - 4utumn07 - Sunstory',
|
||||
'story highlights - 4utumn07 - Sketching⠀',
|
||||
'kestrelsings',
|
||||
'0ct0ber19',
|
||||
'0ct0ber19 - reels',
|
||||
'story - 0ct0ber19',
|
||||
'story highlights - 0ct0ber19 - Heestory',
|
||||
'story highlights - 0ct0ber19 - Drawheeing⠀',
|
||||
'carlyraejepsen',
|
||||
];
|
||||
|
||||
it('folds sidecars into their base profile', () => {
|
||||
const groups = groupArchiveDirectories(dirs);
|
||||
expect([...groups.keys()].sort()).toEqual(['4utumn07', 'kestrelsings']);
|
||||
expect(groups.get('4utumn07')).toHaveLength(5);
|
||||
expect(groups.get('kestrelsings')).toHaveLength(1);
|
||||
expect([...groups.keys()].sort()).toEqual(['0ct0ber19', 'carlyraejepsen']);
|
||||
expect(groups.get('0ct0ber19')).toHaveLength(5);
|
||||
expect(groups.get('carlyraejepsen')).toHaveLength(1);
|
||||
});
|
||||
|
||||
it('orders sources posts, reels, stories, then highlights by title', () => {
|
||||
const sources = groupArchiveDirectories(dirs).get('4utumn07')!;
|
||||
const sources = groupArchiveDirectories(dirs).get('0ct0ber19')!;
|
||||
expect(sources.map(s => s.kind)).toEqual(['posts', 'reels', 'stories', 'highlight', 'highlight']);
|
||||
expect(sources.slice(3).map(s => s.title)).toEqual(['Sketching⠀', 'Sunstory']);
|
||||
expect(sources.slice(3).map(s => s.title)).toEqual(['Drawheeing⠀', 'Heestory']);
|
||||
});
|
||||
|
||||
it('still groups a sidecar whose base profile is missing', () => {
|
||||
|
||||
@@ -18,10 +18,10 @@ export interface ArchiveSource {
|
||||
/**
|
||||
* Sidecar directories sit next to the profile directory they belong to:
|
||||
*
|
||||
* 4utumn07 -> posts (base)
|
||||
* 4utumn07 - reels -> reels
|
||||
* story - 4utumn07 -> stories
|
||||
* story highlights - 4utumn07 - Sunstory -> highlight "Sunstory"
|
||||
* 0ct0ber19 -> posts (base)
|
||||
* 0ct0ber19 - reels -> reels
|
||||
* story - 0ct0ber19 -> stories
|
||||
* story highlights - 0ct0ber19 - Heestory -> highlight "Heestory"
|
||||
*
|
||||
* Instagram usernames cannot contain spaces, so matching the username as a
|
||||
* run of non-space characters reliably separates it from a highlight title
|
||||
|
||||
@@ -14,11 +14,11 @@ describe('isSystemDirectory', () => {
|
||||
});
|
||||
|
||||
it.each([
|
||||
'4utumn07',
|
||||
'4utumn07 - reels',
|
||||
'story - dawn_petal',
|
||||
'story highlights - official_band - A.B.C',
|
||||
'story highlights - theoldlyricmuseinsta - 💙1999-2005 era',
|
||||
'0ct0ber19',
|
||||
'0ct0ber19 - reels',
|
||||
'story - cher_ryppo',
|
||||
'story highlights - official_artms - G.C.I',
|
||||
'story highlights - theoldtaylorswiftinsta - 💙2014-1989 era',
|
||||
'Heejin_Bubble heejinmedia',
|
||||
'gallery-dl',
|
||||
'posts',
|
||||
|
||||
@@ -1,20 +1,21 @@
|
||||
import { describe, expect, it } from 'vitest';
|
||||
import { parseArchiveFilename, scopedPostId } from './archive-patterns';
|
||||
import { canonicalItemId, parseArchiveFilename, scopedPostId } from './archive-patterns';
|
||||
|
||||
describe('parseArchiveFilename — Instagram export format', () => {
|
||||
it('parses a single-image post', () => {
|
||||
expect(parseArchiveFilename('2023-04-19_4utumn07 - CrORBIcJJbM.mp4')).toEqual({
|
||||
expect(parseArchiveFilename('2023-04-19_0ct0ber19 - CrORBIcJJbM.mp4')).toEqual({
|
||||
postId: 'CrORBIcJJbM',
|
||||
date: '2023-04-19',
|
||||
username: '4utumn07',
|
||||
username: '0ct0ber19',
|
||||
index: 1,
|
||||
ext: 'mp4',
|
||||
isStory: false,
|
||||
dateFromMtime: false,
|
||||
});
|
||||
});
|
||||
|
||||
it('parses a carousel slide index', () => {
|
||||
const parsed = parseArchiveFilename('2023-04-12_4utumn07 - Cq8LrxSJAJE - 3.jpg');
|
||||
const parsed = parseArchiveFilename('2023-04-12_0ct0ber19 - Cq8LrxSJAJE - 3.jpg');
|
||||
expect(parsed).toMatchObject({ postId: 'Cq8LrxSJAJE', index: 3, ext: 'jpg' });
|
||||
});
|
||||
|
||||
@@ -26,7 +27,7 @@ describe('parseArchiveFilename — Instagram export format', () => {
|
||||
});
|
||||
|
||||
it('parses caption sidecar files', () => {
|
||||
expect(parseArchiveFilename('2023-04-12_4utumn07 - Cq8LrxSJAJE.txt')).toMatchObject({
|
||||
expect(parseArchiveFilename('2023-04-12_0ct0ber19 - Cq8LrxSJAJE.txt')).toMatchObject({
|
||||
postId: 'Cq8LrxSJAJE',
|
||||
ext: 'txt',
|
||||
});
|
||||
@@ -38,8 +39,8 @@ describe('parseArchiveFilename — Instagram export format', () => {
|
||||
|
||||
it('parses the story sidecar layout (date_user - N - shortcode)', () => {
|
||||
// Files in `story - <user>` carry a per-day ordinal before the shortcode.
|
||||
const parsed = parseArchiveFilename('2025-10-26_4utumn07 - 2 - DQRuDx9iW5Q.jpg', 'stories');
|
||||
expect(parsed).toMatchObject({ date: '2025-10-26', username: '4utumn07', ext: 'jpg' });
|
||||
const parsed = parseArchiveFilename('2025-10-26_0ct0ber19 - 2 - DQRuDx9iW5Q.jpg', 'stories');
|
||||
expect(parsed).toMatchObject({ date: '2025-10-26', username: '0ct0ber19', ext: 'jpg' });
|
||||
expect(parsed!.postId).toContain('DQRuDx9iW5Q');
|
||||
});
|
||||
|
||||
@@ -70,9 +71,9 @@ describe('parseArchiveFilename — Instaloader format', () => {
|
||||
|
||||
describe('parseArchiveFilename — story highlights', () => {
|
||||
it('parses the dateless highlight layout', () => {
|
||||
expect(parseArchiveFilename('4utumn07 - C5dQPEYpd9W.mp4', 'highlight')).toMatchObject({
|
||||
expect(parseArchiveFilename('0ct0ber19 - C5dQPEYpd9W.mp4', 'highlight')).toMatchObject({
|
||||
postId: 'C5dQPEYpd9W',
|
||||
username: '4utumn07',
|
||||
username: '0ct0ber19',
|
||||
ext: 'mp4',
|
||||
isStory: false,
|
||||
});
|
||||
@@ -94,7 +95,7 @@ describe('parseArchiveFilename — story highlights', () => {
|
||||
});
|
||||
|
||||
describe('parseArchiveFilename — non-matching files', () => {
|
||||
it.each(['4utumn07.jpg', 'profile_pic.jpg', 'README.md', 'no-separator.png'])(
|
||||
it.each(['0ct0ber19.jpg', 'profile_pic.jpg', 'README.md', 'no-separator.png'])(
|
||||
'returns null for %s',
|
||||
name => expect(parseArchiveFilename(name)).toBeNull(),
|
||||
);
|
||||
@@ -106,8 +107,8 @@ describe('scopedPostId', () => {
|
||||
});
|
||||
|
||||
it('namespaces sidecar ids by directory', () => {
|
||||
expect(scopedPostId('C5dQ', 'highlight', 'story highlights - u - Sunstory'))
|
||||
.toBe('story highlights - u - Sunstory/C5dQ');
|
||||
expect(scopedPostId('C5dQ', 'highlight', 'story highlights - u - Heestory'))
|
||||
.toBe('story highlights - u - Heestory/C5dQ');
|
||||
});
|
||||
|
||||
it('keeps the same shortcode distinct across sources', () => {
|
||||
@@ -116,3 +117,111 @@ describe('scopedPostId', () => {
|
||||
expect(inPosts).not.toBe(inHighlight);
|
||||
});
|
||||
});
|
||||
|
||||
/**
|
||||
* gallery-dl is replacing JDownloader as the fetcher (docs/gallery-dl.md).
|
||||
* Its naming differs cosmetically, and these cases pin down that the two
|
||||
* interoperate so a mixed archive parses identically.
|
||||
*/
|
||||
describe('gallery-dl / JDownloader naming interop', () => {
|
||||
it('treats a single-media post the same with or without an index', () => {
|
||||
const jd2 = parseArchiveFilename('2023-04-19_0ct0ber19 - CrORBIcJJbM.mp4')!;
|
||||
const gdl = parseArchiveFilename('2023-04-19_0ct0ber19 - CrORBIcJJbM - 1.mp4')!;
|
||||
expect(jd2.postId).toBe(gdl.postId);
|
||||
expect(jd2.index).toBe(gdl.index);
|
||||
expect(jd2.index).toBe(1);
|
||||
});
|
||||
|
||||
it('normalises zero-padded carousel indices', () => {
|
||||
// JD2 pads to the width of the media count (10+ items -> "01"), and
|
||||
// gallery-dl's count can be one higher, so the same post may be padded
|
||||
// by one tool and not the other.
|
||||
expect(parseArchiveFilename('2024-04-17_0ct0ber19 - C53YPQzp7Wj - 09.jpg')!.index).toBe(9);
|
||||
expect(parseArchiveFilename('2024-04-17_0ct0ber19 - C53YPQzp7Wj - 9.jpg')!.index).toBe(9);
|
||||
expect(parseArchiveFilename('2023-11-03_0ct0ber19 - CzM8Uf6B6H_ - 01.jpg')!.index).toBe(1);
|
||||
});
|
||||
|
||||
it('reads a gallery-dl story name, which carries a per-item shortcode', () => {
|
||||
const p = parseArchiveFilename('2026-08-16_official_artms - DcF9OyhBJ1H.jpg', 'stories')!;
|
||||
expect(p.postId).toBe('DcF9OyhBJ1H');
|
||||
expect(p.date).toBe('2026-08-16');
|
||||
});
|
||||
|
||||
it('gives a dated highlight a real date instead of the mtime fallback', () => {
|
||||
const mtime = Date.parse('2026-08-17T00:00:00Z');
|
||||
const undated = parseArchiveFilename('0ct0ber19 - C-IImhvpFuk.jpg', 'highlight', mtime)!;
|
||||
const dated = parseArchiveFilename('2024-08-04_0ct0ber19 - C-IImhvpFuk.jpg', 'highlight', mtime)!;
|
||||
// Same item either way, so re-fetching cannot split it into two posts.
|
||||
expect(dated.postId).toBe(undated.postId);
|
||||
expect(undated.date).toBe('2026-08-17');
|
||||
expect(dated.date).toBe('2024-08-04');
|
||||
});
|
||||
});
|
||||
|
||||
/**
|
||||
* Highlights are the only files with no date in the name, so they fall back to
|
||||
* mtime — which is when the file was written, not when it was posted. Callers
|
||||
* need to know the difference to let a real date win.
|
||||
*/
|
||||
describe('dateFromMtime', () => {
|
||||
const mtime = Date.parse('2026-08-17T00:00:00Z');
|
||||
|
||||
it('flags an undated highlight name as mtime-dated', () => {
|
||||
const p = parseArchiveFilename('0ct0ber19 - C-IImhvpFuk.jpg', 'highlight', mtime)!;
|
||||
expect(p.date).toBe('2026-08-17');
|
||||
expect(p.dateFromMtime).toBe(true);
|
||||
});
|
||||
|
||||
it('does not flag a highlight that carries its own date', () => {
|
||||
const p = parseArchiveFilename('2024-08-04_0ct0ber19 - C-IImhvpFuk.jpg', 'highlight', mtime)!;
|
||||
expect(p.date).toBe('2024-08-04');
|
||||
expect(p.dateFromMtime).toBe(false);
|
||||
});
|
||||
|
||||
it('never flags ordinary post or Instaloader names', () => {
|
||||
expect(parseArchiveFilename('2023-04-19_u - ABC.mp4', 'posts', mtime)!.dateFromMtime).toBe(false);
|
||||
expect(parseArchiveFilename('2024-01-01_12-00-00_UTC.jpg', 'posts', mtime)!.dateFromMtime).toBe(false);
|
||||
});
|
||||
|
||||
it('leaves the date empty rather than guessing when no mtime is given', () => {
|
||||
const p = parseArchiveFilename('0ct0ber19 - C-IImhvpFuk.jpg', 'highlight')!;
|
||||
expect(p.date).toBe('');
|
||||
expect(p.dateFromMtime).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
/**
|
||||
* JDownloader wrote story-shaped names for highlights during one period, so
|
||||
* the same item exists under two conventions. They must be one post.
|
||||
*/
|
||||
describe('canonicalItemId', () => {
|
||||
it('collapses the two highlight naming conventions onto one id', () => {
|
||||
const dir = 'story highlights - 0ct0ber19 - Heestory';
|
||||
const undated = parseArchiveFilename('0ct0ber19 - C5dQPEYpd9W.mp4', 'highlight', 1)!;
|
||||
const dated = parseArchiveFilename('2024-04-07_0ct0ber19 - 01 - C5dQPEYpd9W.mp4', 'highlight')!;
|
||||
expect(scopedPostId(dated.postId, 'highlight', dir))
|
||||
.toBe(scopedPostId(undated.postId, 'highlight', dir));
|
||||
});
|
||||
|
||||
it('does the same for stories', () => {
|
||||
const a = parseArchiveFilename('2025-10-26_u - 2 - DQRuDx9iW5Q.jpg', 'stories')!;
|
||||
expect(scopedPostId(a.postId, 'stories', 'story - u')).toBe('story - u/DQRuDx9iW5Q');
|
||||
});
|
||||
|
||||
it('keeps distinct story items distinct', () => {
|
||||
const a = parseArchiveFilename('2026-08-13_u - 1 - Db-UTJcCUUr.mp4', 'stories')!;
|
||||
const b = parseArchiveFilename('2026-08-13_u - 2 - Db-oNJ1CWQ4.mp4', 'stories')!;
|
||||
expect(scopedPostId(a.postId, 'stories', 'story - u'))
|
||||
.not.toBe(scopedPostId(b.postId, 'stories', 'story - u'));
|
||||
});
|
||||
|
||||
it('leaves a shortcode that merely starts with digits alone', () => {
|
||||
expect(canonicalItemId('0ct0ber19')).toBe('0ct0ber19');
|
||||
expect(canonicalItemId('C5dQPEYpd9W')).toBe('C5dQPEYpd9W');
|
||||
expect(canonicalItemId('12345')).toBe('12345');
|
||||
});
|
||||
|
||||
it('does not touch posts, whose ids are permalinks', () => {
|
||||
expect(scopedPostId('01 - ABC', 'posts')).toBe('01 - ABC');
|
||||
});
|
||||
});
|
||||
|
||||
@@ -30,6 +30,15 @@ export interface ParsedFilename {
|
||||
index: number;
|
||||
ext: string;
|
||||
isStory: boolean;
|
||||
/**
|
||||
* True when `date` is the file's mtime rather than anything Instagram said.
|
||||
*
|
||||
* Only highlights fetched by JDownloader lack a date in the filename, and
|
||||
* their mtime is just when the file was written. Callers should let any real
|
||||
* date win over this one — the same item is often also present under a
|
||||
* gallery-dl name that does carry the date.
|
||||
*/
|
||||
dateFromMtime: boolean;
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -54,6 +63,7 @@ export const parseArchiveFilename = (
|
||||
index: indexStr ? parseInt(indexStr, 10) : 1,
|
||||
ext,
|
||||
isStory: Boolean(story),
|
||||
dateFromMtime: false,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -67,6 +77,7 @@ export const parseArchiveFilename = (
|
||||
index: indexStr ? parseInt(indexStr, 10) : 1,
|
||||
ext,
|
||||
isStory: Boolean(story),
|
||||
dateFromMtime: false,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -81,6 +92,7 @@ export const parseArchiveFilename = (
|
||||
index: 1,
|
||||
ext,
|
||||
isStory: false,
|
||||
dateFromMtime: Boolean(mtime),
|
||||
};
|
||||
}
|
||||
}
|
||||
@@ -88,12 +100,36 @@ export const parseArchiveFilename = (
|
||||
return null;
|
||||
};
|
||||
|
||||
/**
|
||||
* A leading per-day ordinal on a story or highlight id: `01 - C5dQPEYpd9W`.
|
||||
*
|
||||
* JDownloader wrote story-shaped names for highlights during one period of its
|
||||
* life, so the same item exists as both `user - CODE.jpg` and
|
||||
* `date_user - 01 - CODE.jpg`. Those parse to different ids and the viewer
|
||||
* shows the item twice. The ordinal carries no information the shortcode does
|
||||
* not — it is a position within a day's stories, and the shortcode is already
|
||||
* unique — so it is dropped.
|
||||
*/
|
||||
const LEADING_ORDINAL = /^\d+ - (?=[A-Za-z0-9_-]+$)/;
|
||||
|
||||
/** Strip the ordinal so both naming conventions land on the same post. */
|
||||
export const canonicalItemId = (postId: string): string =>
|
||||
postId.replace(LEADING_ORDINAL, '');
|
||||
|
||||
/**
|
||||
* Namespace a post ID by its source directory.
|
||||
*
|
||||
* Base-profile IDs are left untouched so existing permalinks keep working;
|
||||
* sidecar IDs are prefixed so a shortcode appearing in both the profile and a
|
||||
* highlight stays two distinct posts.
|
||||
*
|
||||
* Story and highlight ids are canonicalised first, so an item fetched under
|
||||
* two different naming conventions is one post rather than two.
|
||||
*/
|
||||
export const scopedPostId = (postId: string, kind: SourceKind, dir?: string): string =>
|
||||
kind === 'posts' ? postId : `${dir ?? kind}/${postId}`;
|
||||
export const scopedPostId = (postId: string, kind: SourceKind, dir?: string): string => {
|
||||
if (kind === 'posts') return postId;
|
||||
const id = (kind === 'stories' || kind === 'highlight')
|
||||
? canonicalItemId(postId)
|
||||
: postId;
|
||||
return `${dir ?? kind}/${id}`;
|
||||
};
|
||||
|
||||
@@ -0,0 +1,82 @@
|
||||
import { describe, expect, it } from 'vitest';
|
||||
import {
|
||||
GalleryDlSidecar, isGalleryDlSidecar, sidecarDate, sidecarIsReel, sidecarSource,
|
||||
} from './gallery-dl-sidecar';
|
||||
|
||||
// Trimmed from real files published to the archive on 2026-08-16.
|
||||
const REEL: GalleryDlSidecar = {
|
||||
post_shortcode: 'Db-lNCoib9m', post_id: '3962768346034323302', type: 'reel',
|
||||
date: '2026-08-13 11:00:44', post_date: '2026-08-13 11:00:44',
|
||||
username: 'official_artms', fullname: 'Official ARTMS',
|
||||
description: 'Dancing in the spotlight', count: 1, likes: 22914,
|
||||
};
|
||||
const FEED_VIDEO: GalleryDlSidecar = { ...REEL, post_shortcode: 'DbdG9L9jU4m', type: 'post', count: 2 };
|
||||
const HIGHLIGHT: GalleryDlSidecar = {
|
||||
post_shortcode: 'BATVdRZi_3', post_id: '18099435932626935', type: 'highlight',
|
||||
date: '2026-08-08 16:22:09', username: 'official_artms', count: 154,
|
||||
};
|
||||
|
||||
describe('isGalleryDlSidecar', () => {
|
||||
it('accepts a real sidecar', () => {
|
||||
expect(isGalleryDlSidecar(REEL)).toBe(true);
|
||||
expect(isGalleryDlSidecar(HIGHLIGHT)).toBe(true);
|
||||
});
|
||||
|
||||
it('rejects an Instaloader GraphQL payload', () => {
|
||||
expect(isGalleryDlSidecar({ node: { __typename: 'GraphVideo', shortcode: 'x' } })).toBe(false);
|
||||
expect(isGalleryDlSidecar({ __typename: 'GraphImage', post_shortcode: 'x', type: 'post' })).toBe(false);
|
||||
});
|
||||
|
||||
it('rejects an Instagram export manifest', () => {
|
||||
expect(isGalleryDlSidecar({ media: [{ uri: 'a.jpg' }] })).toBe(false);
|
||||
expect(isGalleryDlSidecar([{ media: [] }])).toBe(false);
|
||||
});
|
||||
|
||||
it('rejects junk', () => {
|
||||
for (const v of [null, undefined, 0, '', 'string', {}, { post_shortcode: 'x' }]) {
|
||||
expect(isGalleryDlSidecar(v)).toBe(false);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('sidecarDate', () => {
|
||||
it('takes the day from the timestamp', () => {
|
||||
expect(sidecarDate(REEL)).toBe('2026-08-13');
|
||||
});
|
||||
|
||||
it('falls back to post_date', () => {
|
||||
expect(sidecarDate({ post_shortcode: 'x', post_date: '2024-01-02 03:04:05' })).toBe('2024-01-02');
|
||||
});
|
||||
|
||||
it('returns empty when there is no usable date', () => {
|
||||
expect(sidecarDate({ post_shortcode: 'x' })).toBe('');
|
||||
expect(sidecarDate({ post_shortcode: 'x', date: 'not a date' })).toBe('');
|
||||
});
|
||||
});
|
||||
|
||||
describe('sidecarIsReel', () => {
|
||||
it('distinguishes a reel from an ordinary feed video', () => {
|
||||
// Both are single mp4s -- the lone-video heuristic cannot tell them apart.
|
||||
expect(sidecarIsReel(REEL)).toBe(true);
|
||||
expect(sidecarIsReel(FEED_VIDEO)).toBe(false);
|
||||
});
|
||||
|
||||
it('declines to answer for stories and highlights', () => {
|
||||
expect(sidecarIsReel(HIGHLIGHT)).toBeUndefined();
|
||||
expect(sidecarIsReel({ post_shortcode: 'x', type: 'story' as const })).toBeUndefined();
|
||||
expect(sidecarIsReel({ post_shortcode: 'x' })).toBeUndefined();
|
||||
});
|
||||
});
|
||||
|
||||
describe('sidecarSource', () => {
|
||||
it('maps type onto the archive source kinds', () => {
|
||||
expect(sidecarSource(REEL)).toBe('reels');
|
||||
expect(sidecarSource(FEED_VIDEO)).toBe('posts');
|
||||
expect(sidecarSource(HIGHLIGHT)).toBe('highlight');
|
||||
expect(sidecarSource({ post_shortcode: 'x', type: 'story' as const })).toBe('stories');
|
||||
});
|
||||
|
||||
it('is undefined for an unknown type', () => {
|
||||
expect(sidecarSource({ post_shortcode: 'x' })).toBeUndefined();
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,87 @@
|
||||
import { SourceKind } from '../types';
|
||||
|
||||
/**
|
||||
* gallery-dl `.json` metadata sidecars.
|
||||
*
|
||||
* Written one per post next to the media (see docs/gallery-dl.md). This is the
|
||||
* only source in any archive format that states outright what a post *is* —
|
||||
* `type` is Instagram's own classification, the `product_type: "clips"` signal
|
||||
* carried through the listing response. Everything else the viewer knows about
|
||||
* reels is guesswork from filenames and directory names.
|
||||
*
|
||||
* Deliberately separate from the two older JSON shapes the scanner reads:
|
||||
*
|
||||
* Instagram export `posts_1.json`, an array of entries with `media`
|
||||
* Instaloader `.json.xz`, a GraphQL node under `node`
|
||||
* gallery-dl this — flat, no wrapper
|
||||
*/
|
||||
export interface GalleryDlSidecar {
|
||||
post_shortcode: string;
|
||||
post_id?: string;
|
||||
/** Instagram's own classification of the post. */
|
||||
type?: 'post' | 'reel' | 'story' | 'highlight';
|
||||
/** Local-time "YYYY-MM-DD HH:MM:SS" — gallery-dl is configured to emit local. */
|
||||
date?: string;
|
||||
post_date?: string;
|
||||
username?: string;
|
||||
fullname?: string;
|
||||
description?: string;
|
||||
count?: number;
|
||||
likes?: number;
|
||||
post_url?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* Recognise a gallery-dl sidecar.
|
||||
*
|
||||
* Checked structurally rather than by filename, because the older formats are
|
||||
* also plain `.json`. `node` and `__typename` are what an Instaloader or
|
||||
* export payload carries, and their absence is what makes this shape
|
||||
* unambiguous.
|
||||
*/
|
||||
export const isGalleryDlSidecar = (data: unknown): data is GalleryDlSidecar => {
|
||||
if (!data || typeof data !== 'object' || Array.isArray(data)) return false;
|
||||
const o = data as Record<string, unknown>;
|
||||
return typeof o.post_shortcode === 'string'
|
||||
&& typeof o.type === 'string'
|
||||
&& o.node === undefined
|
||||
&& o.__typename === undefined
|
||||
&& o.media === undefined;
|
||||
};
|
||||
|
||||
/** The ISO date (YYYY-MM-DD) a sidecar reports, or '' if it carries none. */
|
||||
export const sidecarDate = (s: GalleryDlSidecar): string => {
|
||||
const raw = s.date || s.post_date || '';
|
||||
const day = raw.slice(0, 10);
|
||||
return /^\d{4}-\d{2}-\d{2}$/.test(day) ? day : '';
|
||||
};
|
||||
|
||||
/**
|
||||
* Whether the sidecar says this post is a reel.
|
||||
*
|
||||
* Returns undefined rather than false for stories and highlights: those are
|
||||
* neither reels nor grid posts, and answering "no" would let them be counted
|
||||
* as ordinary posts.
|
||||
*/
|
||||
export const sidecarIsReel = (s: GalleryDlSidecar): boolean | undefined => {
|
||||
if (s.type === 'reel') return true;
|
||||
if (s.type === 'post') return false;
|
||||
return undefined;
|
||||
};
|
||||
|
||||
/**
|
||||
* Which source kind the sidecar implies, for cross-checking the directory.
|
||||
*
|
||||
* A reel shared to the profile grid legitimately appears under `posts`, so a
|
||||
* disagreement is not an error — the directory says where the file was
|
||||
* fetched from, `type` says what Instagram considers it.
|
||||
*/
|
||||
export const sidecarSource = (s: GalleryDlSidecar): SourceKind | undefined => {
|
||||
switch (s.type) {
|
||||
case 'reel': return 'reels';
|
||||
case 'post': return 'posts';
|
||||
case 'story': return 'stories';
|
||||
case 'highlight': return 'highlight';
|
||||
default: return undefined;
|
||||
}
|
||||
};
|
||||
@@ -0,0 +1,56 @@
|
||||
import { describe, expect, it } from 'vitest';
|
||||
import { DatedValue, preferDate, shouldReplaceDate } from './post-dates';
|
||||
|
||||
const sidecar: DatedValue = { date: '2024-04-07', source: 'sidecar' };
|
||||
const filename: DatedValue = { date: '2024-04-08', source: 'filename' };
|
||||
const mtime: DatedValue = { date: '2026-08-17', source: 'mtime' };
|
||||
|
||||
describe('date precedence', () => {
|
||||
it('ranks sidecar above filename above mtime', () => {
|
||||
expect(preferDate(mtime, filename)).toEqual(filename);
|
||||
expect(preferDate(filename, sidecar)).toEqual(sidecar);
|
||||
expect(preferDate(mtime, sidecar)).toEqual(sidecar);
|
||||
});
|
||||
|
||||
it('never lets a weaker source overwrite a stronger one', () => {
|
||||
expect(preferDate(sidecar, filename)).toEqual(sidecar);
|
||||
expect(preferDate(sidecar, mtime)).toEqual(sidecar);
|
||||
expect(preferDate(filename, mtime)).toEqual(filename);
|
||||
});
|
||||
|
||||
it('keeps the incumbent on a tie, so scan order cannot flip the date', () => {
|
||||
const other: DatedValue = { date: '2020-01-01', source: 'filename' };
|
||||
expect(preferDate(filename, other)).toEqual(filename);
|
||||
expect(preferDate(other, filename)).toEqual(other);
|
||||
});
|
||||
|
||||
it('accepts anything when nothing is held yet', () => {
|
||||
expect(preferDate(undefined, mtime)).toEqual(mtime);
|
||||
expect(shouldReplaceDate(undefined, mtime)).toBe(true);
|
||||
});
|
||||
|
||||
it('ignores an empty date regardless of source', () => {
|
||||
const empty: DatedValue = { date: '', source: 'sidecar' };
|
||||
expect(shouldReplaceDate(filename, empty)).toBe(false);
|
||||
expect(preferDate(filename, empty)).toEqual(filename);
|
||||
});
|
||||
|
||||
it('replaces a held-but-empty date', () => {
|
||||
const empty: DatedValue = { date: '', source: 'filename' };
|
||||
expect(preferDate(empty, mtime)).toEqual(mtime);
|
||||
});
|
||||
|
||||
it('is order-independent for the full three-source case', () => {
|
||||
const orders = [
|
||||
[mtime, filename, sidecar],
|
||||
[sidecar, mtime, filename],
|
||||
[filename, sidecar, mtime],
|
||||
[mtime, sidecar, filename],
|
||||
];
|
||||
for (const order of orders) {
|
||||
const won = order.reduce<DatedValue | undefined>(
|
||||
(acc, next) => preferDate(acc, next), undefined);
|
||||
expect(won).toEqual(sidecar);
|
||||
}
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,46 @@
|
||||
/**
|
||||
* Where a post's date came from, and which source wins.
|
||||
*
|
||||
* A post is usually described by several files — media, a caption `.txt`, a
|
||||
* `.json` sidecar, sometimes the same item under two naming conventions — and
|
||||
* they are scanned in directory order, not in order of trustworthiness. Without
|
||||
* an explicit ranking the date is decided by whichever file happened to be
|
||||
* reached first.
|
||||
*
|
||||
* Ranked best to worst:
|
||||
*
|
||||
* sidecar what Instagram reported, straight from a gallery-dl `.json`
|
||||
* filename a date the fetcher wrote into the name; correct, but derived
|
||||
* mtime when the file was written to disk — unrelated to when it was
|
||||
* posted, and only ever a last resort for JDownloader highlights,
|
||||
* whose filenames carry no date at all
|
||||
*/
|
||||
export type DateSource = 'sidecar' | 'filename' | 'mtime';
|
||||
|
||||
const RANK: Record<DateSource, number> = { sidecar: 0, filename: 1, mtime: 2 };
|
||||
|
||||
export interface DatedValue {
|
||||
date: string;
|
||||
source: DateSource;
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether `next` should replace the date currently held.
|
||||
*
|
||||
* Ties keep the incumbent, so scanning stays stable: two files of equal
|
||||
* authority cannot flip a post's date back and forth by scan order.
|
||||
*/
|
||||
export const shouldReplaceDate = (
|
||||
current: DatedValue | undefined,
|
||||
next: DatedValue,
|
||||
): boolean => {
|
||||
if (!next.date) return false;
|
||||
if (!current || !current.date) return true;
|
||||
return RANK[next.source] < RANK[current.source];
|
||||
};
|
||||
|
||||
/** Apply `next` if it outranks `current`, otherwise keep what we have. */
|
||||
export const preferDate = (
|
||||
current: DatedValue | undefined,
|
||||
next: DatedValue,
|
||||
): DatedValue => (shouldReplaceDate(current, next) ? next : (current ?? next));
|
||||
@@ -0,0 +1,141 @@
|
||||
import { describe, expect, it } from 'vitest';
|
||||
import { dedupePostCopies, hasReelSource, makeIsReel, postsForTab } from './post-tabs';
|
||||
import { MediaFile, Post } from '../types';
|
||||
|
||||
const media = (type: MediaFile['type'], index = 1): MediaFile => ({
|
||||
name: `f${index}.${type === 'video' ? 'mp4' : 'jpg'}`,
|
||||
path: `d/f${index}`, url: '', type, index,
|
||||
});
|
||||
|
||||
const post = (id: string, opts: Partial<Post> = {}): Post => ({
|
||||
id, date: '2024-01-01', username: 'u', caption: '', media: [media('image')], thumbnail: '', ...opts,
|
||||
});
|
||||
|
||||
const video = (id: string, opts: Partial<Post> = {}) => post(id, { media: [media('video')], ...opts });
|
||||
const carousel = (id: string, opts: Partial<Post> = {}) =>
|
||||
post(id, { media: [media('image', 1), media('video', 2)], ...opts });
|
||||
|
||||
describe('hasReelSource', () => {
|
||||
it('is false for an archive with no reels directory', () => {
|
||||
expect(hasReelSource([post('A'), video('B')])).toBe(false);
|
||||
});
|
||||
|
||||
it('is true once any post came from a reels directory', () => {
|
||||
expect(hasReelSource([post('A'), video('u - reels/B', { source: 'reels' })])).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('makeIsReel', () => {
|
||||
it('believes the reels directory when there is one', () => {
|
||||
const posts = [video('A'), video('u - reels/B', { source: 'reels' })];
|
||||
const isReel = makeIsReel(posts);
|
||||
// A is a lone video too, but the archive states which posts are reels.
|
||||
expect(isReel(posts[0])).toBe(false);
|
||||
expect(isReel(posts[1])).toBe(true);
|
||||
});
|
||||
|
||||
it('falls back to the lone-video heuristic without one', () => {
|
||||
const posts = [post('A'), video('B'), carousel('C')];
|
||||
const isReel = makeIsReel(posts);
|
||||
expect(posts.map(isReel)).toEqual([false, true, false]);
|
||||
});
|
||||
});
|
||||
|
||||
describe('dedupePostCopies', () => {
|
||||
it('leaves distinct posts alone', () => {
|
||||
const posts = [post('A'), video('B')];
|
||||
expect(dedupePostCopies(posts).map(p => p.id)).toEqual(['A', 'B']);
|
||||
});
|
||||
|
||||
it('collapses a reel fetched into both the profile and the reels directory', () => {
|
||||
const posts = [video('B'), video('u - reels/B', { source: 'reels' })];
|
||||
const deduped = dedupePostCopies(posts);
|
||||
expect(deduped).toHaveLength(1);
|
||||
// The reels copy wins, so the survivor is still recognised as a reel.
|
||||
expect(deduped[0].source).toBe('reels');
|
||||
});
|
||||
|
||||
it('picks the reels copy regardless of scan order', () => {
|
||||
const profileCopy = video('B');
|
||||
const reelCopy = video('u - reels/B', { source: 'reels' });
|
||||
expect(dedupePostCopies([profileCopy, reelCopy])[0].source).toBe('reels');
|
||||
expect(dedupePostCopies([reelCopy, profileCopy])[0].source).toBe('reels');
|
||||
});
|
||||
|
||||
it('keeps the position of the first copy seen', () => {
|
||||
const posts = [post('A'), video('B'), post('C'), video('u - reels/B', { source: 'reels' })];
|
||||
expect(dedupePostCopies(posts).map(p => p.id.split('/').pop())).toEqual(['A', 'B', 'C']);
|
||||
});
|
||||
});
|
||||
|
||||
describe('postsForTab', () => {
|
||||
it('shows reels in the profile grid, as Instagram does', () => {
|
||||
const posts = [post('A'), video('u - reels/B', { source: 'reels' })];
|
||||
expect(postsForTab(posts, 'posts').map(p => p.id)).toEqual(['A', 'u - reels/B']);
|
||||
});
|
||||
|
||||
it('shows the same reel in both tabs', () => {
|
||||
const posts = [post('A'), video('u - reels/B', { source: 'reels' })];
|
||||
const inGrid = postsForTab(posts, 'posts').map(p => p.id);
|
||||
const inReels = postsForTab(posts, 'reels').map(p => p.id);
|
||||
expect(inReels).toEqual(['u - reels/B']);
|
||||
expect(inGrid).toContain('u - reels/B');
|
||||
});
|
||||
|
||||
it('shows a duplicated reel once in the grid, not twice', () => {
|
||||
const posts = [post('A'), video('B'), video('u - reels/B', { source: 'reels' })];
|
||||
expect(postsForTab(posts, 'posts')).toHaveLength(2);
|
||||
expect(postsForTab(posts, 'reels')).toHaveLength(1);
|
||||
});
|
||||
|
||||
it('treats lone videos as reels for archives with no reels directory', () => {
|
||||
const posts = [post('A'), video('B'), carousel('C')];
|
||||
expect(postsForTab(posts, 'posts').map(p => p.id)).toEqual(['A', 'B', 'C']);
|
||||
expect(postsForTab(posts, 'reels').map(p => p.id)).toEqual(['B']);
|
||||
});
|
||||
|
||||
it('has nothing saved', () => {
|
||||
expect(postsForTab([post('A')], 'saved')).toEqual([]);
|
||||
});
|
||||
});
|
||||
|
||||
/**
|
||||
* Once an archive carries gallery-dl sidecars, the guesswork above is replaced
|
||||
* by Instagram's own classification. These are the cases the heuristic got
|
||||
* wrong (see docs/gallery-dl.md).
|
||||
*/
|
||||
describe('explicit isReel from a sidecar', () => {
|
||||
it('beats the lone-video heuristic for an ordinary feed video', () => {
|
||||
// A single mp4 that Instagram calls a post, not a reel — indistinguishable
|
||||
// by shape alone.
|
||||
const posts = [video('DbdG9L9jU4m', { isReel: false })];
|
||||
expect(postsForTab(posts, 'reels')).toEqual([]);
|
||||
expect(postsForTab(posts, 'posts')).toHaveLength(1);
|
||||
});
|
||||
|
||||
it('recognises a reel that lives in the profile grid', () => {
|
||||
// Shared to feed, so it sits in the base directory with source 'posts'.
|
||||
const posts = [post('A'), video('C8FHM6EJl15', { source: 'posts', isReel: true })];
|
||||
expect(postsForTab(posts, 'reels').map(p => p.id)).toEqual(['C8FHM6EJl15']);
|
||||
expect(postsForTab(posts, 'posts')).toHaveLength(2);
|
||||
});
|
||||
|
||||
it('beats the directory when both are present', () => {
|
||||
const posts = [
|
||||
video('u - reels/A', { source: 'reels', isReel: false }),
|
||||
video('u - reels/B', { source: 'reels' }),
|
||||
];
|
||||
// A is a feed video that the reels tab happened to return; B is unlabelled
|
||||
// and falls back to its directory.
|
||||
expect(postsForTab(posts, 'reels').map(p => p.id)).toEqual(['u - reels/B']);
|
||||
});
|
||||
|
||||
it('falls back per post, so a mixed archive still works', () => {
|
||||
const posts = [
|
||||
video('labelled', { isReel: true }),
|
||||
video('unlabelled'),
|
||||
carousel('C'),
|
||||
];
|
||||
expect(postsForTab(posts, 'reels').map(p => p.id)).toEqual(['labelled', 'unlabelled']);
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,108 @@
|
||||
import { Post, SourceKind } from '../types';
|
||||
import { Tab } from './routing';
|
||||
|
||||
/**
|
||||
* Which posts each profile tab shows.
|
||||
*
|
||||
* Instagram's profile grid holds everything the account posted — photos,
|
||||
* carousels and reels alike — and the Reels tab is a *filtered view* of that
|
||||
* same set rather than a separate one. So a reel belongs in both tabs, and
|
||||
* only the Reels tab does any filtering.
|
||||
*
|
||||
* Kept pure and separate from App.tsx so the reel heuristic and the
|
||||
* duplicate-copy rules can be tested directly.
|
||||
*/
|
||||
|
||||
/**
|
||||
* The shortcode shared by every copy of a post, regardless of which source
|
||||
* directory it came from. Sidecar ids are directory-scoped
|
||||
* (`0ct0ber19 - reels/Cq8LrxSJAJE`); the trailing segment is the shortcode.
|
||||
*/
|
||||
const shortcode = (post: Post): string => post.id.split('/').pop() ?? post.id;
|
||||
|
||||
/**
|
||||
* True when the archive has a `- reels` sidecar directory, i.e. it states
|
||||
* outright which posts are reels.
|
||||
*/
|
||||
export const hasReelSource = (posts: Post[]): boolean => posts.some(p => p.source === 'reels');
|
||||
|
||||
/**
|
||||
* Build the reel test for an archive.
|
||||
*
|
||||
* Instagram's own marker is `product_type: "clips"` on the post's GraphQL
|
||||
* node, but only Instaloader archives carry that metadata, and only on newer
|
||||
* captures — JDownloader grabs are media plus a caption `.txt` and nothing
|
||||
* else (see docs/jdownloader.md). So:
|
||||
*
|
||||
* - archives with a `- reels` directory are believed outright;
|
||||
* - everything else falls back to treating a lone video as a reel.
|
||||
*
|
||||
* The fallback is a guess: it cannot tell a reel from an ordinary feed video
|
||||
* or an old IGTV upload, all three of which are plain `GraphVideo` nodes
|
||||
* distinguished only by `product_type`.
|
||||
*/
|
||||
export const makeIsReel = (posts: Post[]): ((post: Post) => boolean) => {
|
||||
const guess = hasReelSource(posts)
|
||||
? (post: Post) => post.source === 'reels'
|
||||
: (post: Post) => post.media.length === 1 && post.media[0]?.type === 'video';
|
||||
|
||||
// `isReel` comes from a gallery-dl sidecar and is Instagram's own answer, so
|
||||
// it beats both fallbacks — per post, since an archive is usually a mix of
|
||||
// files fetched before and after sidecars existed.
|
||||
return (post: Post) => post.isReel ?? guess(post);
|
||||
};
|
||||
|
||||
/** Preference order when the same post was fetched into more than one directory. */
|
||||
const SOURCE_RANK: Record<SourceKind, number> = { reels: 0, posts: 1, stories: 2, highlight: 3 };
|
||||
|
||||
const rankOf = (post: Post): number => SOURCE_RANK[post.source ?? 'posts'];
|
||||
|
||||
/**
|
||||
* Collapse copies of one post that were fetched into more than one directory.
|
||||
*
|
||||
* The JDownloader flow crawls a profile URL and its `/reels/` URL separately
|
||||
* because the profile page misses some reels — so the two overlap, and a reel
|
||||
* present in both lands on disk twice. Those become two posts with distinct
|
||||
* directory-scoped ids, which the grid would happily render side by side.
|
||||
*
|
||||
* The reels-source copy wins, so the surviving post still reports
|
||||
* `source: 'reels'` and both the Reels tab and `tabForSource` recognise it.
|
||||
*
|
||||
* Only safe because callers pass the grid's posts, which exclude stories and
|
||||
* highlights — a shortcode may legitimately appear in both the profile and a
|
||||
* highlight, and those must stay distinct.
|
||||
*/
|
||||
export const dedupePostCopies = (posts: Post[]): Post[] => {
|
||||
const winners = new Map<string, Post>();
|
||||
|
||||
for (const post of posts) {
|
||||
const code = shortcode(post);
|
||||
const existing = winners.get(code);
|
||||
if (!existing || rankOf(post) < rankOf(existing)) winners.set(code, post);
|
||||
}
|
||||
|
||||
// Preserve input order, keyed on the winner so ordering does not depend on
|
||||
// which copy happened to be scanned first.
|
||||
const emitted = new Set<string>();
|
||||
const result: Post[] = [];
|
||||
for (const post of posts) {
|
||||
const code = shortcode(post);
|
||||
if (emitted.has(code)) continue;
|
||||
emitted.add(code);
|
||||
result.push(winners.get(code)!);
|
||||
}
|
||||
return result;
|
||||
};
|
||||
|
||||
/**
|
||||
* The posts a tab displays.
|
||||
*
|
||||
* `posts` must already exclude stories and highlights (App passes `allPosts`).
|
||||
*/
|
||||
export const postsForTab = (posts: Post[], tab: Tab): Post[] => {
|
||||
if (tab === 'saved') return [];
|
||||
|
||||
const unique = dedupePostCopies(posts);
|
||||
if (tab === 'posts') return unique;
|
||||
return unique.filter(makeIsReel(posts));
|
||||
};
|
||||
+16
-16
@@ -12,21 +12,21 @@ describe('parseRoute', () => {
|
||||
});
|
||||
|
||||
it('reads a profile', () => {
|
||||
expect(parseRoute('/4utumn07/')).toEqual({ archive: '4utumn07', tab: 'posts', post: null });
|
||||
expect(parseRoute('/0ct0ber19/')).toEqual({ archive: '0ct0ber19', tab: 'posts', post: null });
|
||||
});
|
||||
|
||||
it('reads a profile without a trailing slash', () => {
|
||||
expect(parseRoute('/4utumn07')).toEqual({ archive: '4utumn07', tab: 'posts', post: null });
|
||||
expect(parseRoute('/0ct0ber19')).toEqual({ archive: '0ct0ber19', tab: 'posts', post: null });
|
||||
});
|
||||
|
||||
it('reads a tab', () => {
|
||||
expect(parseRoute('/4utumn07/reels/').tab).toBe('reels');
|
||||
expect(parseRoute('/4utumn07/saved/').tab).toBe('saved');
|
||||
expect(parseRoute('/0ct0ber19/reels/').tab).toBe('reels');
|
||||
expect(parseRoute('/0ct0ber19/saved/').tab).toBe('saved');
|
||||
});
|
||||
|
||||
it('reads a post in Instagram form', () => {
|
||||
expect(parseRoute('/4utumn07/p/Db5tIoRCcvm/')).toEqual({
|
||||
archive: '4utumn07', tab: 'posts', post: 'Db5tIoRCcvm',
|
||||
expect(parseRoute('/0ct0ber19/p/Db5tIoRCcvm/')).toEqual({
|
||||
archive: '0ct0ber19', tab: 'posts', post: 'Db5tIoRCcvm',
|
||||
});
|
||||
});
|
||||
|
||||
@@ -41,8 +41,8 @@ describe('parseRoute', () => {
|
||||
});
|
||||
|
||||
it('still understands the legacy query form', () => {
|
||||
expect(parseRoute('/', '?a=4utumn07&t=reels&p=ABC')).toEqual({
|
||||
archive: '4utumn07', tab: 'reels', post: 'ABC',
|
||||
expect(parseRoute('/', '?a=0ct0ber19&t=reels&p=ABC')).toEqual({
|
||||
archive: '0ct0ber19', tab: 'reels', post: 'ABC',
|
||||
});
|
||||
});
|
||||
|
||||
@@ -54,9 +54,9 @@ describe('parseRoute', () => {
|
||||
describe('buildPath', () => {
|
||||
it.each([
|
||||
[{ archive: null, tab: 'posts', post: null }, '/'],
|
||||
[{ archive: '4utumn07', tab: 'posts', post: null }, '/4utumn07/'],
|
||||
[{ archive: '4utumn07', tab: 'reels', post: null }, '/4utumn07/reels/'],
|
||||
[{ archive: '4utumn07', tab: 'posts', post: 'Db5tIoRCcvm' }, '/4utumn07/p/Db5tIoRCcvm/'],
|
||||
[{ archive: '0ct0ber19', tab: 'posts', post: null }, '/0ct0ber19/'],
|
||||
[{ archive: '0ct0ber19', tab: 'reels', post: null }, '/0ct0ber19/reels/'],
|
||||
[{ archive: '0ct0ber19', tab: 'posts', post: 'Db5tIoRCcvm' }, '/0ct0ber19/p/Db5tIoRCcvm/'],
|
||||
] as const)('builds %j', (route, expected) => {
|
||||
expect(buildPath(route as any)).toBe(expected);
|
||||
});
|
||||
@@ -71,8 +71,8 @@ describe('buildPath', () => {
|
||||
|
||||
it('round-trips through parseRoute', () => {
|
||||
for (const route of [
|
||||
{ archive: '4utumn07', tab: 'posts' as const, post: null },
|
||||
{ archive: '4utumn07', tab: 'reels' as const, post: null },
|
||||
{ archive: '0ct0ber19', tab: 'posts' as const, post: null },
|
||||
{ archive: '0ct0ber19', tab: 'reels' as const, post: null },
|
||||
{ archive: 'Heejin_Bubble heejinmedia', tab: 'posts' as const, post: null },
|
||||
]) {
|
||||
expect(parseRoute(buildPath(route))).toEqual(route);
|
||||
@@ -86,12 +86,12 @@ describe('postSlug / findPostBySlug', () => {
|
||||
});
|
||||
|
||||
it('strips the sidecar directory from the slug', () => {
|
||||
expect(postSlug(post('story highlights - u - Sunstory/C5dQPEYpd9W'))).toBe('C5dQPEYpd9W');
|
||||
expect(postSlug(post('story highlights - u - Heestory/C5dQPEYpd9W'))).toBe('C5dQPEYpd9W');
|
||||
});
|
||||
|
||||
it('resolves a slug back to its post', () => {
|
||||
const posts = [post('AAA'), post('4utumn07 - reels/BBB', 'reels')];
|
||||
expect(findPostBySlug(posts, 'BBB')?.id).toBe('4utumn07 - reels/BBB');
|
||||
const posts = [post('AAA'), post('0ct0ber19 - reels/BBB', 'reels')];
|
||||
expect(findPostBySlug(posts, 'BBB')?.id).toBe('0ct0ber19 - reels/BBB');
|
||||
expect(findPostBySlug(posts, 'AAA')?.id).toBe('AAA');
|
||||
});
|
||||
|
||||
|
||||
@@ -37,6 +37,12 @@ export interface Post {
|
||||
isStory?: boolean;
|
||||
/** Defaults to 'posts' for archives without sidecar directories. */
|
||||
source?: SourceKind;
|
||||
/**
|
||||
* Instagram's own answer to "is this a reel", from a gallery-dl `.json`
|
||||
* sidecar. Undefined when the archive carries no such sidecar, which is when
|
||||
* the viewer has to fall back to guessing — see src/lib/post-tabs.ts.
|
||||
*/
|
||||
isReel?: boolean;
|
||||
/** Highlight this post belongs to, for source === 'highlight'. */
|
||||
highlightTitle?: string;
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user