Compare commits

..
Author SHA1 Message Date
ergosteurandClaude Opus 5 71cfd29f36 docs: record the first full incremental sync, and what made it cheap
All six ARTMS profiles, four days after the previous run: 184 new media, 299
files published, 0 failures and 0 CDN 429s. The header still claimed the
script was a skeleton with nothing migrated, which stopped being true a while
ago.

Notes the two things that made it cheap — an already-seeded archive DB, so no
probe passes ran at all, and --abort 50 — and that the run was deliberately
stopped and resumed midway to pick up the new flag. That is only safe because
the state file had already marked the finished sources as fetched, so the 20h
floor skipped them. Stopping a run used to mean repeating it.

Also records that rsync exit 23 is the NORMAL outcome of a publish: the SSH
user cannot chown to rslsync, so attrs fail while data arrives intact. The
check is a --dry-run re-run returning an empty file list, which is what
verified this one (+299 files, matching rsync exactly).

The account was restored; noting that without pretending it licenses more
traffic.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UXfdJu7QhSJLr47K7koTDF
2026-08-20 14:36:32 -04:00
ergosteurandClaude Opus 5 a83da461c1 feat: stop enumerating a profile once it reaches what we already hold
The skip-archive suppresses downloads, which spends the CDN. It does nothing
about the listing pass, which spends `instagram.com` — the surface that
actually bans accounts — and that cost scales with how BIG a profile is, not
with how much of it is new. A 2275-post profile paid ~76 pages every run to
discover three new posts. Seeding saved the second full pass, never the first.

Measured from sidecar write times during today's run, free because the run was
paying for the listing anyway: three new posts took ~100s each, and the other
2272 were written in a single second — enumeration with nothing to show for it.

`--abort N` passes gallery-dl's `skip: abort:N`, stopping the extractor after N
consecutive already-archived files. Resuming a stopped run with `--abort 50`
enumerated 7 posts of cher_ryppo's 2151 and still caught every new one.

Three things make this safe, and all of them are load-bearing:

- N counts FILES, not posts, so it has to clear the largest already-held
  carousel — one post in this archive is 22 media. A reels tab needs 50 actual
  reels for the same threshold, since those are single-media.
- It applies to posts and reels only. Stories are always new, and highlight
  items are not ordered in a way that makes early abort safe.
- The REST listing is strictly reverse-chronological. Test case 16 claimed
  0ct0ber19 returns its 3 pinned posts out of date order; that is true of the
  web grid but not of this endpoint, measured today. Front-loaded old posts are
  the one thing that would trip abort before it reached anything new, so the
  correction is what licenses the feature rather than a footnote to it.

Default is 0 — walk everything — because aborting early stops noticing edited
carousels (test case 15), which only a full enumeration finds. Routine runs
want 50; a full sweep is still worth running occasionally.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UXfdJu7QhSJLr47K7koTDF
2026-08-20 14:31:33 -04:00
ergosteurandClaude Opus 5 01702ea24b docs: fix the yt-dlp advice — a separate venv is not an importable one
`pipx install yt-dlp` was the instruction here, and it never worked. It
gives yt-dlp its own venv, so the binary lands on PATH while gallery-dl,
living in a different venv, still cannot `import yt_dlp`. Everything looks
installed and the log keeps saying `Cannot import yt-dlp or youtube-dl` —
which is exactly what the first run logged, and why it fell back to
progressive URLs for DASH video, the path the CDN 429s hit hardest.

`pipx inject gallery-dl yt-dlp` is the fix. Verify it by asking gallery-dl's
own interpreter rather than the shell, since `which yt-dlp` succeeds either
way and is what made the original advice look correct.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UXfdJu7QhSJLr47K7koTDF
2026-08-20 14:31:09 -04:00
ergosteurandClaude Opus 5 07df53acde chore: release 1.8.0
Docker Build and Publish / build-and-push (push) Failing after 13s
Ships the gallery-dl sidecar work to the viewer. The visible change is
reel classification: official_artms' Reels tab drops from 781 items to
360, because the sidecars say the other 421 are ordinary feed videos the
clips endpoint returns via include_feed_video. Directory-based
classification counted them all as reels.

Also in this release: dates ranked by source rather than scan order, and
highlight items no longer appearing twice when the archive holds them
under both JDownloader naming conventions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 13:58:15 -04:00
ergosteurandClaude Opus 5 c81f18276f fix: stop showing a highlight item twice under two naming conventions
JDownloader wrote story-shaped names for highlights during one period of
its life, so the same item exists on disk as both

    0ct0ber19 - C5dQPEYpd9W.mp4
    2024-04-07_0ct0ber19 - 01 - C5dQPEYpd9W.mp4

which parsed to the ids "C5dQPEYpd9W" and "01 - C5dQPEYpd9W" -- two posts
for one item. The leading ordinal is a position within a day's stories
and carries nothing the shortcode does not, so story and highlight ids
drop it. Post ids are untouched, since those are permalinks.

Measured on the two real files, same archive, cache cleared between:
without the fix the profile reads "Heestory - 2 items", with it
"Heestory - 1 item".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 13:57:42 -04:00
ergosteurandClaude Opus 5 4ce5cf048e feat: give the sync a memory, so it stops paying for the same listing twice
Nothing in this tool had any memory: every invocation started from zero
and would happily re-enumerate a profile it had listed minutes earlier.
That is what suspended the account -- the listing passes, not the
downloads -- and an aborted run re-enumerating five profiles on restart
was a large part of the bill.

Three changes, in order of how much they save:

- Seeding is now a one-time bootstrap per source. After the first
  successful sync the archive DB records everything gallery-dl has seen,
  so the source is never probed again. A second full sync costs roughly
  half what the first did.
- Stories never seed at all. A story cannot be in the archive before it
  is fetched, so there is nothing to seed from, and probing would double
  the cost of the cheapest surface we have.
- A source fetched within --min-interval (20h) is refused, and listing
  results are cached for --probe-ttl (24h), so a restart mid-run is free
  rather than a repeat. --force overrides both.

--only replaces --no-stories and takes any subset of the surfaces, which
is what makes a daily stories-only run possible: one source per profile,
no seeding, a handful of requests. Everything else stays monthly.

Tested with stdlib unittest -- no new dependencies, and it runs anywhere
the sync does. The cases include the aborted-restart scenario, which now
plans zero work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 13:53:18 -04:00
ergosteurandClaude Opus 5 7a085d2272 docs: record why the account was suspended — verification, not fetching
The account was suspended on 2026-08-17 for "spam", during the session
that built this tooling. Both fetching docs were confidently wrong about
what the risk was, so both now carry the correction.

docs/jdownloader.md said the ban vector is instagram.com requests, which
is right, and implied that meant per-post metadata fetching, which is
only part of it. The suspension came from read-only verification:
automated browser scrolling to enumerate profile grids (~18 paginated
loads per profile, done twice on one after a selector bug), repeated
--simulate and -j passes over the same profiles, per-post /p/ fetches
while testing filename formats, and an aborted sync that re-ran every
listing pass before dying. None of that produced a file, and together it
rivalled the real sync for request count.

The rules that follow are in docs/gallery-dl.md: verify against the
archive rather than the live site, count read-only work against the same
budget, treat the first CDN 429 as the end of the session rather than a
pacing knob, and cache probe_live so a restart does not re-enumerate
everything. The warning order was CDN 429, then 400 on the highlights
tray, then suspension.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 12:56:00 -04:00
ergosteurandClaude Opus 5 087304243e fix: rank date sources instead of letting scan order decide
The previous commit had this backwards: the sidecar date was only
consulted when the existing date came from an mtime, so a filename date
silently outranked what Instagram itself reported.

The order is sidecar, then filename, then mtime -- metadata first,
mtime last, since mtime is when the file hit disk and says nothing about
when the post was made. Ties keep the incumbent so two equally
authoritative files cannot flip a post's date by scan order.

Extracted to src/lib/post-dates.ts rather than left inline, because the
rule is easy to state and easy to get wrong -- the tests include an
order-independence case that would have caught the original mistake.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 12:50:48 -04:00
ergosteurandClaude Opus 5 e14dbf6ec8 feat: read gallery-dl sidecars for reel type and post dates
The .json sidecars published with the ARTMS fetch were inert: the scanner
fed them through the Instaloader path, where `node.edge_media_to_caption`
and `checkIsStory`'s `product_type` are both absent, so nothing happened.

They are now recognised structurally -- flat, with post_shortcode and
type, and none of the markers the other two JSON shapes carry -- and used
for three things:

- `type` sets post.isReel, which post-tabs prefers over every fallback.
  This is Instagram's own classification and it disagrees with ours a
  lot: of 781 items in "official_artms - reels", the sidecars say only
  360 are reels. The other 421 are feed videos the clips endpoint returns
  via include_feed_video, and the directory-based rule counted them all.
- `description` fills the caption where no .txt exists.
- `date` dates a post whose filename could not.

Also fixes date precedence. Only JDownloader highlights lack a date in
the filename, so parseArchiveFilename now marks those as mtime-derived
and the scanner lets any real date replace them -- previously the date
depended on which file the scan reached first.

Verified against real published files: a directory of three type=post and
three type=reel renders 6 in the grid and exactly the 3 reels in the
Reels tab.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 12:42:54 -04:00
ergosteurandClaude Opus 5 3e29976866 fix: line-buffer sync output so a redirected log shows progress live
A sync runs for hours and is normally watched through a redirected log,
where Python's block buffering withheld the per-source progress lines
until they happened to flush. The gallery-dl subprocesses write to the
same descriptor unbuffered, so the log also interleaved out of order.

Reconfiguring the streams in-process rather than relying on `python3 -u`
means it holds however the script is invoked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 22:48:40 -04:00
ergosteurandClaude Opus 5 0ac1ef951f docs: move project status out of the sync script into the docs
Script comments should describe the script. Where a run has got to is
project state, so it belongs in docs/gallery-dl.md, which now records the
live withaseul publish, the file-ownership caveat and where the profile
list lives.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 22:33:38 -04:00
ergosteurandClaude Opus 5 ee66b87bff feat: drive gdl-sync from a file of profile URLs
Adds --urls-file so a run can reference a hand-maintained list rather
than repeating --profile, which is how this actually gets used: the
ARTMS accounts now live in artms_account_links.txt at the archive root.

The parser takes what a person would paste. Full URLs, scheme-less URLs
and bare usernames all work; blank lines and # comments are ignored and
duplicates dropped, so the list can be appended to carelessly. Lines that
are not profiles are rejected loudly rather than silently syncing
nothing: an Instagram post URL yields the segment "p", which would
otherwise be treated as a username and create a directory called "p".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 22:32:06 -04:00
ergosteurandClaude Opus 5 68568ef855 fix: keep the generated config out of the published tree
A dry-run publish against the real archive caught gdl-sync.config.json
being created in the archive root: it was written into the staging
directory, and staging is rsynced wholesale. It now lives as a sibling of
staging instead, with rsync excludes as a second line of defence.

The dry run is otherwise clean -- 322 files added, 0 deleted, no new
directories -- and confirms the property that matters most: of 74 new
media files, zero duplicate media already held under a different name.
The JD2 and gallery-dl naming really do converge.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 21:43:12 -04:00
ergosteurandClaude Opus 5 3a34e4359e feat: publish fetched media by rsync, and back off from the CDN's 429s
Wires up the staging -> rsync step and replaces --archives with --index
(a listing source: local path or the viewer's API), --staging and
--publish, so the fetch host needs no copy of the archive.

rsync runs --ignore-existing with no --delete. That is a safety property
rather than an optimisation: the archive deliberately outlives Instagram,
so publishing must only ever add. It runs once at the end so a profile
that fails midway never reaches the archive half-written.

Exercised end-to-end against withaseul across all four surfaces,
publishing to a scratch directory. Seeding worked as designed (915 of 984
post items and 28 of 34 reel items already held), stories and highlights
returned no results cleanly, and the collab-reel case landed correctly:
"withaseul - reels" holds files owned by cher_ryppo, 0ct0ber19 and
official_artms, each with the owner in the filename and the crawl scope
as the directory.

The first run drew '429 Too Many Requests' from the CDN at 3M with 1-3s
sleeps and lost two videos. That is the tolerant surface complaining, so
the defaults are now 1M, 6-10s between requests, 3-6s between downloads,
sleep-429 of 120s and 8 retries. Re-running recovered both videos with
zero failures and zero 429s. Installing yt-dlp on the fetch host also
matters: without it DASH videos fall back to a progressive URL, which is
what the rate limiting hit hardest.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 21:37:38 -04:00
ergosteurandClaude Opus 5 4fee8b1dfe feat: seed gallery-dl's skip-archive so the fetcher needs no archive copy
gallery-dl skips already-held media either by file existence -- which
requires the archive mounted where it writes -- or by a sqlite
skip-archive, which requires nothing on disk. Using the latter lets the
fetch host write to local disk and rsync afterwards, avoiding tens of
thousands of small writes over CIFS and keeping a mid-sync failure from
leaving partial files on the live Resilio share.

The key is archive_prefix + archive_fmt: the literal "instagram" plus the
per-media numeric pk. Verified against a real run -- a 3-image carousel
produced 3 rows and a re-run skipped every media file.

media_id is absent from our filenames, so the DB cannot be built from
names alone, but the listing pass we already make maps every live item to
its media_id, and a file listing says which we hold. Seeding therefore
costs no extra Instagram requests and no archive content -- the listing
GET /api/archives/:name/files already serves is enough.

Measured on 0ct0ber19: 2275 live items, 2248 seeded, 27 left to fetch --
exactly the media of the two posts added since the last crawl.

The trap worth the comment it carries: posts and reels are filed under
post_shortcode, while stories and highlights use the per-item shortcode
(post_shortcode there is the containing reel's id, shared by every item).
Matching on the wrong field seeded 5 of 2275 rather than failing loudly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 21:10:37 -04:00
ergosteurandClaude Opus 5 7b63ba6a76 docs: design a gallery-dl replacement for the JDownloader fetcher
Every claim in docs/gallery-dl.md was measured against the live site and
the archive rather than taken from documentation, because two of the
assumptions turned out to be wrong.

The safety model is the reason the config looks the way it does.
gallery-dl has two API backends: the graphql one issues a request PER
POST for every video and carousel -- the pattern that got this account
banned via Instaloader -- while the default rest one paginates listings
at 30-50 items and carries carousel_media, video_versions and
product_type inline. A 300-post profile costs ~10 requests.

Findings worth recording:

- JD2 stamped filenames in desktop LOCAL time (US Eastern), not UTC.
  Across 212 comparable posts: UTC 19 mismatches, UTC-5 10, UTC-4 zero.
  {date:Olocal/%Y-%m-%d} reproduces it; the trailing separator must be
  omitted or it lands in the strftime format.
- A profile's reels tab returns collab reels owned by OTHER accounts, so
  the directory must be forced with -D. JD2 did the same: chuuo3o and
  official_artms filenames sit inside "0ct0ber19 - reels".
- Stories and highlights need per-item {shortcode}; {post_shortcode} is
  the reel's id and is shared by every item. {date} is per-item, verified
  on a 154-item highlight with distinct times.
- gallery-dl reproduces JD2's caption .txt exactly, including writing
  nothing for an empty caption and omitting the trailing newline.
- The json sidecar needs `include`, not `fields`; `fields` silently does
  nothing in mode:json and leaks audio_user blobs. It yields `type`
  (post/reel) -- Instagram's own flag, which can retire the lone-video
  heuristic once the scanner reads it.

Naming differences between the two tools are cosmetic: EXPORT_RE already
makes the index optional and parseInt normalises zero-padding, so a mixed
archive parses identically. Tests pin that down.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 21:04:39 -04:00
ergosteurandClaude Opus 5 b84113af18 fix: count the grid in the post header, and qualify the Instagram claim
Docker Build and Publish / build-and-push (push) Failing after 10s
The header rendered allPosts.length, which is pre-dedupe — 0ct0ber19
showed "303 posts" over a 300-tile grid. Instagram's counter equals its
grid, so count the grid.

Also correct CLAUDE.md. v1.7.0 claimed the grid holds everything "as on
Instagram"; Instagram actually includes a reel in the grid only when the
creator shared it to feed, per post. Measured live: official_artms has
21 reels in its first 34 grid tiles, 0ct0ber19 has 1 in 214. Archives
carry no such flag, so showing everything approximates the behaviour
rather than reproducing it.

Records the DOM trap that caused the wrong reading in the first place:
grid reels link to /reel/<code>/, not /<user>/p/<code>/.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 20:00:11 -04:00
ergosteurandClaude Opus 5 abf22eb50f feat: show reels in the profile grid, as Instagram does
Docker Build and Publish / build-and-push (push) Failing after 10s
The Posts tab filtered reels out, so the grid was not the archive — it
was the archive minus its videos. For `for.heejin` that hid 533 of 1225
posts; for `loonatheworld`, 1100 of 3813. On Instagram the grid holds
everything and the Reels tab is a filtered view of that same set.

Extract the tab logic to src/lib/post-tabs.ts so the reel heuristic is
testable outside the component, and add dedupePostCopies: the jd2 flow
crawls the profile URL and the /reels URL separately because the profile
page misses some reels, so the two overlap and a reel can land on disk
twice. Those are two posts with distinct directory-scoped ids, which the
grid would now render side by side; the reels-source copy wins so the
survivor is still recognised as a reel.

Deciding what *is* a reel stays a guess for most archives. Instagram
marks it with product_type ("clips" vs "feed" vs "igtv" — all three are
GraphVideo, and aspect ratio does not separate them), but only newer
Instaloader captures carry it: 1101 of gibiofficial's 5919 sidecars, and
only 2 marked clips. JDownloader archives carry none, so those still
fall back to treating a lone video as a reel.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 19:06:53 -04:00
ergosteurandClaude Opus 5 e6d874a5d9 docs: record the CSP/wasm and PWA-precache traps in CLAUDE.md
Two failures in this session were expensive because nothing recorded them:

- The CSP must keep 'wasm-unsafe-eval' and connect-src data:, because the xz
  decompressor for Instaloader sidecars is WebAssembly embedded as a data: URL.
  Removing either breaks decoding with a bare "Failed to fetch" and no stack,
  and the visible symptom is silent metadata loss rather than an error.
- The service worker precaches index.html with its headers, so a server-only
  change never reaches installed clients. The version compiled into the client
  is what forces the precache to turn over each release; it is load-bearing,
  not decoration.

Also notes that the Vite dev server sends none of these headers, so CSP and PWA
behaviour must be verified against a built dist/ served by server.js, and that
the live archive root is now the archives/ subdirectory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 18:36:09 -04:00
ergosteurandClaude Opus 5 ff45b6d898 fix: make server header changes reach installed PWA clients
Docker Build and Publish / build-and-push (push) Failing after 10s
The service worker precaches index.html together with its response headers, so
a server-only change never reaches an installed client: the client build is
byte-identical, the precache manifest is unchanged, and the worker has no
reason to update. That is why the CSP fix in 1.6.1 did not reach a browser that
already had the app cached — it kept replaying a cached shell carrying the old,
broken CSP, indefinitely.

The release version is now compiled into the client, which makes every release
change the bundle hash, hence index.html, hence its precache revision, hence
sw.js itself — the bytes browsers compare to decide whether to update. Verified
by bumping only the version: index-DYufL2Fa.js -> index-60N_3d5j.js, with the
new name carried into the sw.js manifest.

It also surfaces in the footer, so the deployed version is visible without
digging through devtools.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 13:16:25 -04:00
ergosteurandClaude Opus 5 30ff8de1e8 fix: allow WebAssembly in the CSP so xz sidecars can be decoded
Docker Build and Publish / build-and-push (push) Failing after 9s
The xz decompressor for Instaloader's .json.xz sidecars is WebAssembly,
embedded as a data: URL that it fetches at startup. The CSP added in 1.3.0
blocked both halves of that:

  fetch('data:application/wasm;...')  -> TypeError: Failed to fetch
  WebAssembly.instantiate(...)        -> CompileError: violates script-src 'self'

The first surfaces through new Response(stream).json() as a bare "Failed to
fetch" with no stack, which reads like a network fault and is why this was
mis-diagnosed twice. Vite's dev server never sends the CSP, so it reproduced
only in production — every Instaloader archive silently lost its captions,
story flags and profile metadata from 1.3.0 onward.

script-src now allows 'wasm-unsafe-eval', which permits WebAssembly compilation
without permitting eval() of JavaScript, and connect-src allows data: for the
embedded module.

Verified against the production bundle: rivvsofficial goes from 188 posts / 0
followers / no stories to 68 posts, 120 stories, 10,337 followers and its real
name, bio and link — with zero decode errors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 12:51:42 -04:00
ergosteurandClaude Opus 5 9560f85515 fix: read xz sidecars by buffer, page carousel with arrows, stop backdrop flash
Docker Build and Publish / build-and-push (push) Failing after 9s
Instaloader metadata was silently lost
Every .json.xz failed with "Failed to fetch" during a scan, though the same URL
fetched fine on its own. RemoteArchiveFile.stream() started a fetch, piped the
body into a TransformStream and returned the readable immediately — nothing
caught a fetch rejection, and the decompressor stops reading at the end of the
xz member, so the response body was never drained or cancelled. Across ~190
sidecars that exhausted the connection pool.

Everything Instaloader archives carry lives in those files, so the failure was
invisible but total. rivvsofficial reported 188 posts, no stories, 0 followers
and a placeholder bio; it now reports 68 posts, 120 stories, 10,337 followers
and the real name, bio and link — 68 + 120 = 188, matching the sidecars exactly
(106 GraphStoryVideo + 14 GraphStoryImage = 120).

These sidecars are a few KB, so they are now read into memory before
decompressing. stream() was left unused by that change and is removed from the
interface and both implementations rather than kept as a trap.

Arrow keys page the carousel
They moved between posts, which contradicted the arrows drawn on the carousel
itself. Arrows now page slides; , and . move between posts, alongside the side
buttons.

Backdrop cross-fade
AnimatePresence had no exit variant, so the outgoing scan backdrop was removed
instantly while its replacement faded in over 1.5s, exposing the pale page
behind it as a white flash. Layers now stack: the outgoing image holds full
opacity until covered, and the 0.4 moved onto the group so overlapping layers
don't darken as they cross. Measured over a real scan: 152 cross-fades with a
layer always opaque, except the opening fade-in where nothing is underneath.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 12:28:26 -04:00
ergosteurandClaude Opus 5 f0c054946f docs: add a JDownloader quick reference
Covers why fetching goes through JDownloader rather than Instaloader (the
instagram.com vs CDN split, and what the metadata gap actually costs), the
settings that matter, cookie handling, the two-URL workflow, jd2-sync usage,
the expected on-disk layout, and what to do when something breaks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 12:02:41 -04:00
ergosteurandClaude Opus 5 42eb598d1c fix: never descend into NAS metadata directories when indexing
Docker Build and Publish / build-and-push (push) Failing after 9s
The archive root was filtered by prefix, but the recursive walk below it was
not, so anything inside a profile directory got indexed. NAS filesystems put
sidecar metadata *inside* every folder rather than only at the share root:
Synology writes @eaDir (thumbnails and indexing data), #recycle holds
deletions, .sync is Resilio state. On the live share those account for 12,516
of 123,023 files.

None currently sit inside a profile directory, so nothing was miscounted yet —
but the moment that share gets indexed for Photos, every generated thumbnail
would be counted as archive media and stat'd one by one over the network, which
is the cost the index exists to avoid.

One isSystemDirectory rule now applies at every level, and the root listing uses
it too instead of keeping a second copy of the pattern.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 11:57:29 -04:00
ergosteurandClaude Opus 5 aac753ced9 feat: generate JDownloader crawljobs from the archives on disk
The manual flow is: paste a profile URL into JDownloader, paste the /reels URL
separately (the profile page misses some reels), set the output folder by hand,
repeat per profile. scripts/jd2-sync.ts emits one crawljob per source with the
folder already pointed at the right directory, so folder-watch picks up the
whole batch at once.

Profiles and sidecars are derived with the same grouping logic the server uses,
so output folders always match what the viewer expects to find. --download-base
maps the path for a JDownloader running on another machine (Windows paths
included), since it typically runs on a desktop against the share.

Directories that aren't Instagram profiles are skipped: an archive root also
collects tool output and exports from other services, and pointing a crawl at
those spends requests on instagram.com to be told the profile doesn't exist —
exactly the traffic worth not spending. Filtering is by username shape, plus
--skip and a .jd2ignore file for names that look like usernames but aren't.

Defaults are conservative: chunks=1, because multi-chunk ranged requests are the
one CDN-side pattern that doesn't resemble a browser, and links park in the
LinkGrabber for review rather than auto-starting.

Only posts and reels are emitted; highlight URLs need a numeric id and story
URLs expire, so those stay manual.

Format verified against JDownloader's own explain.txt for the folderwatch
extension, read from the daily SVN mirror rather than one of the decade-stale
GitHub copies.

Also refreshes CLAUDE.md, whose URL-state section still described the query
parameters replaced in 1.4.0, and documents the mobile feed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 11:41:23 -04:00
ergosteurandClaude Opus 5 9173504190 feat: mobile opens posts as a scrolling feed instead of a modal
Docker Build and Publish / build-and-push (push) Failing after 9s
Tapping a post on a phone now opens a real feed page — header, media, actions,
caption, next post peeking in below — scrolled with the browser's own vertical
scrolling rather than swipe gestures. Desktop keeps the modal, where a centred
sheet with side arrows suits a pointer.

Only a window of posts is mounted: a profile here holds up to 1129 posts and
mounting them all would mean as many full-size images. The window grows in both
directions as you scroll. Growing upwards shifts everything below it, so the
scroll offset is corrected in the same frame, before paint — measured against
the real archive, an anchored post moves exactly one screen per scroll with no
jump.

Only the post crossing the viewport centre plays its video; the rest stay
paused, so a feed of reels doesn't play ten at once. The URL tracks that same
post, so scrolling updates /<archive>/p/<shortcode>/ the way Instagram does,
and the back button returns to the grid with its scroll position intact.

Feed video sizes to the container width rather than its intrinsic size: a
<video> reports 300x150 until metadata loads, which made it render narrow and
then jump to full width. It also gets a taller height ceiling than the modal so
ordinary portrait media fills the width instead of sitting in side bars.

The carousel is extracted into a shared MediaCarousel used by both surfaces, so
horizontal paging behaves identically; touch-action keeps vertical scrolling
passing through to the feed. PostModal loses its now-dead mobile swipe branch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 10:23:33 -04:00
ergosteurandClaude Opus 5 1340d85945 feat: Instagram-shaped URLs, Instagram-shaped gestures, iOS-feel animations
Docker Build and Publish / build-and-push (push) Failing after 9s
Navigation gestures
Horizontal swipe used to advance the carousel and then, on the last slide,
fling you into the next post — one gesture meaning two things. Horizontal is
now carousel-only. On touch, vertical swipe moves between posts (down on the
first post still dismisses, keeping drag-to-close where it can't mean
"previous"). Desktop keeps the arrows outside the modal.

URLs
Permalinks now mirror Instagram:

  /<archive>/                 profile
  /<archive>/reels/           tab
  /<archive>/p/<shortcode>/   post

A post URL carries no tab, as on Instagram; the tab is re-derived from the
post's source, so opening a reel link lands on the Reels tab with next/prev
paging through reels. Sidecar posts keep directory-scoped ids internally but
expose only the shortcode. The old ?a=&t=&p= form is still parsed so existing
links keep working, and reserved prefixes (api, archives, assets…) can never be
mistaken for a profile name.

Animations
Adds a shared motion vocabulary tuned to feel native: critically damped springs
rather than fixed-duration easing, and gestures hand their exit velocity to the
animation so a flick continues instead of restarting. Post transitions animate
along the axis the input implies — vertical for a swipe, horizontal for the
arrows. Modal and story viewer present/dismiss with a scale, tiles and
highlight circles get touch-down feedback, and prefers-reduced-motion is
honoured throughout.

Adds 22 routing tests (58 total).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 10:01:04 -04:00
ergosteurandClaude Opus 5 c5b0a5cb5f fix: keep post nav arrows outside the modal and stop scroll chaining
Docker Build and Publish / build-and-push (push) Failing after 10s
The prev/next arrows are fixed to the viewport edges while the modal grows to
fill the available width, so below roughly 1200px the modal slid underneath
them and a white chevron landed on the white caption panel — invisible until
hovered. The overlay now reserves a horizontal gutter (md:px-16 lg:px-24) so
the arrows always sit outside the modal, and they get a solid white pill with a
dark chevron so they read against anything behind them. Verified clearing the
modal at 768, 1024, 1440 and 1920px.

The caption sidebar shrinks to w-80 at md so the narrower modal doesn't squeeze
the media pane.

Also adds overscroll-contain to the overlay: the wheel previously chained
through to the post grid behind it, scrolling the background while a post was
open.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 09:05:24 -04:00
ergosteurandClaude Opus 5 ae0f075855 feat: fit full-view media to the viewport and play with sound
Docker Build and Publish / build-and-push (push) Failing after 9s
Media in the post modal used w-full/h-auto, so a portrait video or image grew
taller than the screen (a 720x1280 reel rendered 768x1365 in a 786px viewport)
and forced the modal to scroll. Full view now caps height to the viewport minus
the modal's own padding. Video sizes to its own aspect within the cap so a
portrait clip isn't letterboxed edge to edge; images keep filling the modal
width and only gain a height ceiling.

Opening the modal or a story reel is a user gesture, so playback now starts
unmuted and only falls back to muted if the browser actually refuses the
play() promise — previously it always started muted, and the earlier
muted-by-default fix meant a blocked video could stall the story progress bar.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 08:50:49 -04:00
ergosteurandClaude Opus 5 3146f896c3 docs: update CLAUDE.md for the archive index, sidecars and cache rehydration
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 02:14:11 -04:00
ergosteurandClaude Opus 5 ed078c7b46 chore: bump version to 1.3.1
Docker Build and Publish / build-and-push (push) Failing after 9s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 02:08:35 -04:00
ergosteurandClaude Opus 5 55b0752b05 fix: don't crash at boot when running under a UID with no passwd entry
os.userInfo() throws ERR_SYSTEM_ERROR (uv_os_get_passwd) for a UID that has no
/etc/passwd entry, which is exactly what `docker run --user 1234:1234` produces
— the very workaround the README recommends. Combined with the switch to a
non-root image user, this crashed the server on startup for any deployment that
needed a custom UID to read its archives.

Also document the non-root default and the /cache index volume.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 02:08:35 -04:00
ergosteurandClaude Opus 5 53703cd7cd Merge branch 'review-fixes': security, performance and sidecar archive support
Docker Build and Publish / build-and-push (push) Failing after 9s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 01:59:41 -04:00
ergosteurandClaude Opus 5 97f5d19ce4 chore: bump version to 1.3.0
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 01:59:41 -04:00
ergosteurandClaude Opus 5 1b4aba54d6 fix: security, performance and correctness pass; add sidecar archive support
Security
- Fix path traversal in GET /api/archives/:name/files. Express decodes route
  params after segment matching, so `..%2f..%2fetc` escaped ARCHIVES_DIR and
  returned a recursive listing of arbitrary directories.
- Add CSP and baseline security headers; disable x-powered-by.
- Stop baking GEMINI_API_KEY into the client bundle (the SDK was unused).
- Run the container as `node` instead of root.

Performance
- Add a directory-mtime-keyed archive index, warmed in the background and
  persisted. Listing 110k files went from ~52s to ~0.1s; the largest archive
  (24k files) serves in ~0.3s. Per-file stat over CIFS costs ~1.4ms and does
  not parallelise, so it is now done once rather than per request.
- Build media URLs from the File directly instead of
  `new Blob([await file.arrayBuffer()])`, which read every media file fully
  into memory (a 20GB archive tried to become 20GB of resident blobs).
- Track and revoke object URLs; previously none were ever revoked.
- Give `requestThumbnail` a stable identity so a completed thumbnail stops
  re-running the effect in every mounted thumbnail.
- Namespace IndexedDB keys so listing archives no longer deserializes every
  cached thumbnail blob, and thumbnails no longer collide across archives.
- Serve real file sizes: RemoteArchiveFile was constructed with size 0, which
  silently disabled high-res thumbnailing for every server archive.

Correctness
- Local archives cached media as blob: URLs, which die with the document, so
  a cached local archive restored as an archive of broken images. Media now
  carries a stable path and is rehydrated from a persisted directory handle
  (File System Access API), falling back to re-prompting for the folder.
- Fix permalinks: the URL-writing effect erased ?a= on mount before the
  archive list arrived to consume it, so deep links never resolved.
- Make cache invalidation detect nested changes via a directory signature.
- Add an error boundary and tolerate unparseable dates, which previously
  threw a RangeError and blanked the app.
- Default video to muted so autoplay is not blocked by Safari/Firefox.

Features
- Fold sidecar directories into their base profile: `<user> - reels`,
  `story - <user>` and `story highlights - <user> - <title>` now appear as
  reels, the story ring and Instagram-style highlight circles rather than as
  separate archives.

Housekeeping
- Add @types/react; React was previously type-checked against its JavaScript
  source, so `npm run lint` gave almost no type safety on components.
- Vendor fonts and PWA icons locally; the app made third-party CDN requests
  despite advertising offline support and local-only processing.
- Drop unused better-sqlite3 (a native module that broke `npm install`).
- Add vitest with 36 tests over the filename and directory-naming rules.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 01:59:19 -04:00
ergosteur b30285fe70 docs: update documentation for high-res performance and local persistence
Docker Build and Publish / build-and-push (push) Failing after 10s
Key changes:
- Updated README.md and GEMINI.md with details on background thumbnailing and inter-post preloading.
- Documented persistent local archive caching and smart profile fallback features.
- Added dist-server/ to .gitignore.
- Restored missing feature descriptions and troubleshooting tips in README.
2026-03-07 21:59:43 -05:00
ergosteur 20209bcad5 feat: enable persistent local archives and smart profile fallback
Key changes:
- Enabled full metadata caching for local folder archives, allowing them to load instantly from IndexedDB without re-uploading.
- Implemented oldest-image fallback for profiles missing an explicit profile picture.
- Restored folder-name-to-username detection for local archive uploads.
- Optimized scan indexing to track all image files for fallback use.
2026-03-07 21:56:25 -05:00
ergosteur 74902234b3 fix: restore white glass scanning UI and resolve small image blur bug
Key changes:
- Corrected logic in PostThumbnail to prevent blur effects on images smaller than 1MiB.
- Restored the white glass aesthetic to the scanning dashboard with improved contrast and transparency.
- Optimized scanning background transitions to ensure a smooth, flicker-free crossfade.
2026-03-07 21:46:11 -05:00
ergosteur d62bddc3aa perf: implement background thumbnail generation and inter-post preloading
Key changes:
- Added Web Worker for background image thumbnailing with a 1MiB threshold to optimize CPU/memory usage.
- Implemented a serial task queue for memory-safe high-res image processing, preventing OOM crashes.
- Added inter-post preloading in the modal for seamless 'Previous/Next' navigation.
- Refined scanning UI with double-buffering and a dark background to completely eliminate white flashes.
- Renamed project to 'instaarchive-viewer' in package.json.
- Fixed 'Open image in new tab' by denylisting /archives and /api in PWA config.
2026-03-07 21:42:56 -05:00
ergosteur 42c13ea106 chore: bump version to 1.2.0
Docker Build and Publish / build-and-push (push) Failing after 9s
2026-03-07 21:16:39 -05:00
ergosteur a4e9ce16a7 feat: modularize scanner, enhance carousel preloading, and improve PWA updates
Summary of changes:
- Extracted archive scanning logic into a modular 'useArchiveScanner' hook for better maintainability and performance.
- Refined PostModal carousel with intelligent media preloading and smoother, jitter-free transitions.
- Optimized image rendering with 'decoding=async' and removed 'black flashes' between slide changes.
- Updated PWA configuration to 'autoUpdate' with hourly periodic checks for fresh content.
- Fixed several bugs including stories sorting, permalink parameter cleanup, and profile metadata cache restoration.
- Comprehensive updates to documentation (README.md and GEMINI.md) reflecting the new architecture.
2026-03-07 21:16:28 -05:00
ergosteur 4f89a69ee3 fix: optimize scanning performance and resolve zero-post bug
Docker Build and Publish / build-and-push (push) Failing after 10s
2026-03-07 20:25:23 -05:00
ergosteur 147dcdf2f1 fix: restore missing UI handlers and finalize generic parser
Docker Build and Publish / build-and-push (push) Failing after 10s
2026-03-07 20:18:33 -05:00
ergosteur d4e20d9b98 fix: refine dockerignore and bump version to v1.1.4
Docker Build and Publish / build-and-push (push) Failing after 10s
2026-03-07 20:10:40 -05:00
ergosteur ec8c771733 feat: implement permalinks and document PWA cache troubleshooting
Docker Build and Publish / build-and-push (push) Failing after 10s
2026-03-07 05:20:09 -05:00
ergosteur 3784e8729b debug: add verbose logging to permalink synchronization 2026-03-07 05:12:14 -05:00
ergosteur 69d62eaa5c fix: improve permalink comparison and add debug logging 2026-03-07 05:10:43 -05:00
ergosteur 767f9c508b feat: implement permalinks for archives, tabs, and posts 2026-03-07 05:07:04 -05:00
ergosteur 103ce6f207 fix: exhaustive generic parser and implement local archive history 2026-03-07 05:05:08 -05:00
ergosteur 5267dab236 fix: address Docker EACCES errors with better logging and SELinux hints
Docker Build and Publish / build-and-push (push) Failing after 9s
2026-03-07 03:02:02 -05:00
ergosteur ebf2bf660a fix: improve Docker archive discovery and switch to compiled server
Docker Build and Publish / build-and-push (push) Failing after 1m6s
2026-03-07 02:57:19 -05:00
ergosteur c0f3523a9c docs: update README and GEMINI with Docker usage and new features 2026-03-07 02:50:18 -05:00
ergosteur 6f5021638c feat: add Dockerfile and GitHub Actions workflow for GHCR deployment
Docker Build and Publish / build-and-push (push) Canceled after 11s
2026-03-07 02:45:17 -05:00
ergosteur 67f7750157 feat: refine navigation protection to only warn when leaving the app 2026-03-07 02:41:16 -05:00
ergosteur 9e306eb85e feat: add explicit confirmation for back button and refresh in archives 2026-03-07 02:39:39 -05:00
ergosteur b2da08d52d feat: add navigation protection and refine cached badge visibility 2026-03-07 02:37:17 -05:00
ergosteur d7c13ecc19 fix: resolve Firefox media warnings by improving video cleanup 2026-03-07 02:30:46 -05:00
ergosteur d396b356be feat: implement persistent caching, glassy scanning UI, and UI refinements 2026-03-07 02:28:22 -05:00
ergosteur e23dfe4474 feat: implement self-hostable mode with server-side directory scanning 2026-03-07 00:59:31 -05:00
ergosteur 41e7c5e206 docs: update README and GEMINI.md, remove AI Studio boilerplate and .env.example 2026-03-07 00:36:46 -05:00
ergosteur f685eaebd7 feat: enhance story viewer and media playback experience 2026-03-07 00:31:04 -05:00
ergosteur 47e44ec5e9 feat: improve archive parsing, add .json.xz support, and fix profile pic display 2026-03-07 00:03:54 -05:00
ergosteur cd7dc5f981 feat: Initialize InstaArchive PWA project
Sets up a new React PWA project with Vite, Tailwind CSS, and basic PWA features. Includes essential files like README, .gitignore, package.json, and initial app structure.
2026-03-06 22:41:48 -05:00
ergosteurandGitHub a724e5bc87 Initial commit 2026-03-06 22:41:33 -05:00
41 changed files with 3911 additions and 187 deletions
+77 -9
View File
@@ -15,6 +15,9 @@ InstaArchive Viewer is a React 19 + Vite 6 PWA for browsing archived Instagram d
- `npm run lint` — type-check only (`tsc --noEmit`)
- `npm test` / `npm run test:watch` — vitest
- `npx vitest run src/lib/archive-patterns.test.ts` — a single test file
- `npm run jd2 -- --archives <dir> --dry-run` — generate JDownloader `.crawljob`
files for every profile on disk (see `scripts/jd2-sync.ts` and
`docs/jdownloader.md`)
Local development usually needs both `npm run dev` and `npm run server`. Local-folder mode works without the backend; server-mode archives do not.
@@ -34,10 +37,10 @@ Loading is unified behind the `ArchiveFile` interface (`src/types/index.ts`, imp
An archive root holds one directory per profile plus *sidecars* that belong to it:
```
4utumn07 -> posts (base)
4utumn07 - reels -> reels
story - 4utumn07 -> stories
story highlights - 4utumn07 - Sunstory -> highlight "Sunstory"
0ct0ber19 -> posts (base)
0ct0ber19 - reels -> reels
story - 0ct0ber19 -> stories
story highlights - 0ct0ber19 - Heestory -> highlight "Heestory"
```
`src/lib/archive-grouping.ts` (shared by server and tests) folds these into a single profile with a `sources` list. Sidecars never appear as standalone archives. Each file the server returns carries its `kind`, so the client routes posts / reels / story ring / highlight circles without re-deriving naming rules.
@@ -61,6 +64,18 @@ story highlights - 4utumn07 - Sunstory -> highlight "Sunstory"
Results are cached to IndexedDB. Media records store a stable `path`; **`url` is not persistable** for local archives because blob URLs die with the document.
Three different JSON shapes turn up as `.json`, so they are told apart structurally, not by filename (`src/lib/gallery-dl-sidecar.ts`):
| shape | marker |
|---|---|
| Instagram export manifest | top-level `media` array |
| Instaloader `.json.xz` | GraphQL node under `node` / `__typename` |
| gallery-dl sidecar | flat, `post_shortcode` + `type`, none of the above |
The gallery-dl sidecar is the only source that states what a post *is*: its `type` (`post` / `reel` / `story` / `highlight`) is Instagram's own classification, so `post.isReel` set from it beats every fallback in `post-tabs.ts`. This matters — of the 781 items in `official_artms - reels`, the sidecars say only **360 are reels**; the other 421 are ordinary feed videos the clips endpoint returns via `include_feed_video`. Directory-based classification counted all 781.
**Dates are ranked, not last-write-wins** (`src/lib/post-dates.ts`): sidecar (what Instagram reported) beats filename (what the fetcher wrote) beats mtime (when the file hit disk, and unrelated to when it was posted). Ties keep the incumbent. Several files describe one post and they are scanned in directory order, not in order of trustworthiness, so without the ranking the date was decided by whichever file came first. Only JDownloader highlights fall to mtime at all — `parseArchiveFilename` flags those via `dateFromMtime`.
### Cache and local-archive persistence (`src/lib/archive-cache.ts`)
IndexedDB keys are namespaced (`archive:`, `thumb:`, `handle:`) so listing archives does not deserialize every cached thumbnail blob, and thumbnails are scoped per archive to avoid cross-archive collisions.
@@ -71,25 +86,78 @@ Restoring an archive **rehydrates URLs from `path`**: server archives rebuild HT
Images over 1MiB are downscaled in a Web Worker via `OffscreenCanvas`. The queue is **serial on purpose** — decoding several 50MP+ images at once OOMs the tab. `requestThumbnail` must keep a stable identity (it reads cache state through a ref), or every completed thumbnail re-runs the effect in all mounted thumbnails.
### URL state (`src/App.tsx`)
### Profile tabs (`src/lib/post-tabs.ts`)
App state syncs to `?a=` / `?t=` / `?p=`. Two rules, both learned from real bugs:
The grid holds **everything**, reels included, and the Reels tab is a *filtered view* of that same set. Only the Reels tab filters. The tabs were mutually exclusive until v1.7.0, which hid a lot: 1100 of `loonatheworld`'s 3813 posts and 533 of `for.heejin`'s 1225 never appeared in the grid at all.
- The initial query string is captured into a ref on first render; the URL is rewritten from state as soon as anything loads, so reading `window.location` later sees the rewrite, not the user's link.
This *approximates* Instagram rather than matching it. Instagram's grid includes a reel only if the creator shared it to feed — a per-post choice, measured live on 2026-08-16: `official_artms` had 21 reels in its first 34 grid tiles, `0ct0ber19` just 1 in 214. That flag appears nowhere in an archive (JD2 stores no metadata, and Instaloader's `product_type` says what a post *is*, not whether it was shared to feed), so showing everything is the closest reachable behaviour. Instagram's "N posts" counter equals its grid, which is why the header counts `postsForTab(allPosts, 'posts')` and not `allPosts` — the raw list still holds both copies of a double-fetched post.
When checking the live site, note that grid reels link to `/reel/<code>/` while ordinary posts link to `/<user>/p/<code>/`. Matching only `/p/` silently drops every reel, which once produced a confident and completely wrong conclusion that Instagram never shows reels in the grid.
Deciding *what is a reel* has no good answer for most archives. Instagram's own marker is `product_type` on the post's GraphQL node (`clips` = reel, `feed` = ordinary feed video, `igtv`, `story`) — `__typename` is `GraphVideo` for all three, and aspect ratio does not separate them either. But:
- Only Instaloader archives carry that metadata, and only newer captures. A survey of `gibiofficial` found `product_type` on 1101 of 5919 sidecars, and just **2** posts marked `clips`.
- JDownloader archives carry none at all — media plus a `.txt` holding the bare caption.
So the viewer believes a `- reels` sidecar directory when one exists, and otherwise falls back to treating a lone video as a reel. **The fallback is a guess**: it cannot tell a reel from a feed video or an old IGTV upload, and it misses videos inside carousels.
`dedupePostCopies` exists because the JDownloader flow crawls the profile URL and the `/reels/` URL separately (the profile page misses some reels), so the two overlap and a reel can land on disk twice. Those become two posts with distinct directory-scoped ids, which the grid would otherwise render side by side. It dedupes by shortcode, preferring the reels-source copy. It is only safe over `allPosts` — stories and highlights are excluded there, and a shortcode may legitimately appear in both a profile and a highlight.
### URL state (`src/App.tsx`, `src/lib/routing.ts`)
Paths mirror Instagram: `/<archive>/`, `/<archive>/reels/`, `/<archive>/p/<shortcode>/`. The old `?a=&t=&p=` form is still parsed for existing links but never written. Reserved prefixes (`api`, `archives`, `assets`…) can't be mistaken for a profile name.
A post URL carries no tab, as on Instagram — the tab is re-derived from the post's `source`, so a reel link lands on the Reels tab and pages through reels. Sidecar posts keep directory-scoped ids internally but expose only the shortcode.
Three rules, all learned from real bugs:
- The initial route is captured into a ref on first render; the URL is rewritten from state as soon as anything loads, so reading `window.location` later sees the rewrite, not the user's link.
- URL writing is gated on `hasInitialLoaded`, otherwise it erases the deep link before the loader consumes it.
- Deep-link resolution waits on the archive fetch having *settled* (`archivesFetched`), not on `isServerMode`, which is still false while the request is in flight.
Deep-link resolution waits on the archive fetch having *settled*, not on `isServerMode` (which is still false while in flight).
### Mobile feed (`src/components/PostFeed.tsx`)
Below `md`, opening a post renders a scrolling feed page rather than the modal (`useIsMobile` decides). Only a window of posts is mounted; it grows both ways, and prepending corrects `scrollTop` in a `useLayoutEffect` so content doesn't jump. Only the post crossing the viewport centre plays its video and drives the URL. Desktop keeps `PostModal`; both share `MediaCarousel`.
### Backend (`server.ts`)
Serves `/api/archives`, `/api/archives/:name/files`, static `/archives`, and the built SPA. Notes:
- Express decodes route params **after** segment matching, so `..%2f` reaches the handler as `../`. All user-supplied archive names go through `resolveArchivePath`.
- Sets CSP and related security headers. The CSP allows `blob:`/`data:` for media and `unsafe-inline` styles (the animation library sets inline styles); scripts stay same-origin only.
- `os.userInfo()` throws for a UID with no `/etc/passwd` entry, which is what `--user 1234:1234` produces — use `describeUser()`.
### CSP: do not tighten `script-src` or `connect-src` without testing xz
The xz decompressor for Instaloader `.json.xz` sidecars is **WebAssembly**, embedded as a `data:` URL the library fetches at startup. The policy must keep:
```
script-src 'self' 'wasm-unsafe-eval' // compile wasm, without allowing eval() of JS
connect-src 'self' data: // fetch the embedded module
```
Removing either breaks decoding with a bare `TypeError: Failed to fetch` **and no stack** — it surfaces through `new Response(stream).json()`, so it reads like a network fault rather than a policy block. The visible symptom is not an error page: archives silently lose captions, story flags and all profile metadata (follower counts, bio, name). This shipped broken for several releases.
To check quickly, run in the page console:
```js
await fetch('data:application/wasm;base64,AGFzbQEAAAA=') // connect-src
await WebAssembly.instantiate(Uint8Array.of(0,97,115,109,1,0,0,0)) // script-src
```
**The Vite dev server does not send these headers**, so anything CSP-related is invisible in `npm run dev`. Verify security-header and PWA behaviour by building and serving `dist/` through `server.js`, not against the dev server.
### PWA: server-only changes do not reach installed clients
The service worker precaches `index.html` **together with its response headers**. A change that touches only the server (a CSP fix, a new header) leaves the client build byte-identical, so the precache manifest and `sw.js` are unchanged, the worker never updates, and installed clients keep replaying the old shell with the old headers — indefinitely.
`vite.config.ts` therefore compiles the package version into the client via `define: { __APP_VERSION__ }`, and `App.tsx` renders it in the footer. That is **load-bearing**: it makes every release change the bundle hash → `index.html` → its precache revision → `sw.js`, which is what browsers byte-compare to decide whether to update. Don't remove it as dead weight.
To recover a client stuck on an old shell: unregister the service worker, delete its caches, reload.
### Deployment
Live archives live in `<share>/Instagram-archive/archives/` — one directory per profile plus sidecars. Directories that are not Instagram profiles (tool output, exports from other services) sit *outside* that folder so they never reach the viewer.
Container runs as non-root. The image defaults to `node`, but the archive share must be *listable* by that UID — a mode-711 share owned by another account needs `user: "<uid>:<gid>"` in compose. Mount a volume at `/cache` so the index survives restarts.
### PWA / build quirks
+6
View File
@@ -0,0 +1,6 @@
https://www.instagram.com/0ct0ber19/
https://www.instagram.com/kimxxlip/
https://www.instagram.com/withaseul/
https://www.instagram.com/cher_ryppo/
https://www.instagram.com/zindoriyam/
https://www.instagram.com/official_artms/
+605
View File
@@ -0,0 +1,605 @@
# gallery-dl — a CLI replacement for JDownloader2
Status: **in production.** All six ARTMS profiles are synced with
`scripts/gdl-sync.py`; JD2 is no longer used for them.
Everything below was measured against the live site and the real archive on
2026-08-16 and 2026-08-20, not inferred from documentation.
## Why gallery-dl and not a hand-rolled script
The hard parts of fetching Instagram are pagination, cookie handling, CDN URL
expiry and resumption. gallery-dl already has all of them, plus extractors that
map 1:1 onto our sidecar directory layout (`posts`, `reels`, `stories`,
`highlights`). Rolling our own would mean reimplementing the ban-sensitive part
by hand.
## The account was suspended on 2026-08-17 — read this first
The account used for all of the below was suspended the same day this tooling
was built, for "activity that doesn't follow our Community Standards on spam".
The fetching was not the expensive part. **Verification was.**
**It was restored, and synced normally again on 2026-08-20** — a full run
across all six profiles with 0 failures and 0 CDN 429s. That is not evidence
the limits were imagined; it is one data point on a restored account that has
been treated carefully since. Everything below still applies, and the budget is
still per session rather than per command.
What was actually spent against `instagram.com` in a few hours, from one
session and one IP:
| activity | rough requests | downloaded |
|---|---:|---|
| enumerating a profile grid by scrolling it in an automated browser | ~18 pages | nothing |
| the same profile again, after a bug in the scraping selector | ~18 pages | nothing |
| a Reels tab enumerated the same way | ~9 pages | nothing |
| full `-j` metadata dumps of one profile, twice | ~16 pages | nothing |
| `--simulate` runs over the same profile, three times | ~24 pages | nothing |
| single-post `/p/<code>/` fetches while testing filename formats | ~8 | a handful |
| an aborted sync that re-ran every listing pass before dying | ~40 pages | ~270 MB |
| the real sync, 24 sources across 6 profiles | ~150 pages | 2.2 GB |
The two rows that actually mattered to the archive are the last one and part of
the second-to-last. **Everything above them produced no files at all**, and
together they were a comparable number of requests.
The warnings arrived in this order and were each rationalised:
1. `429 Too Many Requests` from `scontent-*.cdninstagram.com`, losing two
videos. Treated as a pacing problem — pacing was lowered and the run
continued.
2. `400 Bad Request` from `/api/v1/highlights/<id>/highlights_tray/`, on an
endpoint that had worked hours earlier. Correctly read as a possible block;
requests stopped.
3. Suspension.
**Treat the first CDN 429 as a stop signal for the session, not a tuning
parameter.** It is the tolerant surface complaining; if that surface is
complaining, the rate-limited one has been unhappy for a while.
### Rules that follow from this
- **Count verification requests against the same budget as fetching.** A
`--simulate`, a `-j` dump and a browser scroll all hit `instagram.com` and
download nothing. Being read-only does not make them free; it makes them
invisible, which is worse.
- **Never enumerate the live site with an automated browser.** Scrolling a
214-post grid is ~18 paginated GraphQL loads at machine speed with no dwell
time between them. It is the most obviously non-human thing in this whole
document, and it was done here twice on one profile.
- **Verify against the archive, not against Instagram.** Every naming, dating
and classification question answered in this file could have been answered
from files already on disk plus a single listing pass.
- **`probe_live` is not cached, so every restart re-enumerates everything.**
The aborted run cost a full duplicate set of listing passes for five
profiles. Cache probe output to disk before running anything twice.
- **Budget per session, not per command.** Nothing in the tooling knows what
the last command spent.
### For a replacement account
- Let it exist and be used normally for a while before pointing any tool at it.
- Keep the cookie on one machine and one public IP, as before.
- Start with a single small profile and stop for the day afterwards.
- Prefer Instagram's own "Download a copy" export where possible: it is
first-party, costs no scraping requests, and carries the metadata this whole
document works around not having.
## The safety model — read this before changing any option
The ban vector is **requests to `instagram.com`**, not bandwidth. See
`docs/jdownloader.md` for the history; Instaloader got this account banned by
asking `instagram.com` a question *per post*.
gallery-dl has two API backends and the difference is exactly that vector:
```python
if self.config("api") == "graphql":
self.api = InstagramGraphqlAPI(self) # per-post api.media() for every
else: # video and every carousel
self.api = InstagramRestAPI(self) # <- default, listing-only
```
The REST backend paginates at `count: 30` (feed) / `page_size: 50` (clips), and
those responses already carry `carousel_media`, `image_versions2`,
`video_versions` and `product_type`. **No per-post request.** A 300-post
profile costs roughly 10 requests to `instagram.com`.
Rules, in order of importance:
1. **`"api": "rest"` always.** Never `graphql`. This is the whole ballgame.
2. **Never enable `metadata`-style options that trigger extra calls.** If a
field is not already in the listing response, it is not worth a request.
3. **Pace it.** `"sleep-request": [4.0, 7.0]` — a randomised gap, not a fixed
one. Also `"sleep": [1.0, 3.0]` between downloads.
4. **Cap the download rate** (`downloader.http.rate`) so the CDN side looks like
a person, not a mirror.
5. **Run from the same public IP as the browser the cookie came from.** At time
of writing that is `mattellite` (`66.23.52.196`); the dev workstation is a
*different* public IP and using the cookie from there is precisely what
session-hijack detection looks for.
6. **No programmatic login, ever.** gallery-dl's username/password path is
disabled upstream anyway; use `--cookies-from-browser`.
Do not add proxy rotation, fingerprint spoofing or account rotation. Throttling
and request-avoidance are welcome; evasion is not.
### Cookies
The logged-in Chrome on `mattellite` runs with a non-default profile:
```
--user-data-dir=/home/matt/.config/google-chrome-devtools
```
so the cookie flag is:
```
--cookies-from-browser "chrome:/home/matt/.config/google-chrome-devtools"
```
Plain `--cookies-from-browser chrome` fails with "Unable to find chrome cookies
database" because it looks in `~/.config/google-chrome/`.
Anonymous access is **not** a viable fallback: it serves lower-resolution media,
caps profile pagination at 12 posts, and returns `AuthRequired` for stories and
highlights.
## Output format
The viewer's parser is the contract, not JD2's exact bytes. `EXPORT_RE` in
`src/lib/archive-patterns.ts` accepts all of these, and normalises the index
with `parseInt`, so **JD2 and gallery-dl naming interoperate**:
```
"… - CrORBIcJJbM.mp4" -> postId=CrORBIcJJbM index=1
"… - CrORBIcJJbM - 1.mp4" -> postId=CrORBIcJJbM index=1
"… - C53YPQzp7Wj - 09.jpg" -> postId=C53YPQzp7Wj index=9
```
That means zero-padding and the presence/absence of ` - N` on single-media posts
are cosmetic. Don't spend effort forcing them.
### Directory layout
| kind | directory | note |
|---|---|---|
| posts | `<user>` | |
| reels | `<user> - reels` | |
| stories | `story - <user>` | |
| highlights | `story highlights - <user> - <title>` | |
**Force the directory with `-D`; never use `{username}` for it.** A profile's
reels tab returns *collab reels owned by other accounts*`/0ct0ber19/reels/`
served 6 reels owned by `official_artms` and 1 by `chuuo3o`. With
`{username}` those would scatter into `official_artms - reels/`. JD2 got this
right and the archive proves it: `chuuo3o` and `official_artms` filenames sit
inside `0ct0ber19 - reels/`.
So: **owner in the filename, crawl scope in the directory.**
### Filenames
```jsonc
"filename": {
"sidecar_shortcode and count >= 10":
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode} - {num:02}.{extension}",
"sidecar_shortcode":
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode} - {num}.{extension}",
"":
"{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.{extension}"
}
```
`sidecar_shortcode` is set only when the post is a carousel, so it is the
carousel discriminator. Conditions are evaluated in order, first match wins
(`path.py:265`).
Stories and highlights use the per-item `{shortcode}`, not `{post_shortcode}`
(which is the *reel's* id, shared by every item in it):
```
"{date:Olocal/%Y-%m-%d}_{username} - {shortcode}.{extension}"
```
`{date}` on a story/highlight file is the **per-item** `taken_at`
(`instagram.py:337` prefers `item["taken_at"]`), verified on a 154-item
highlight whose items carried distinct times while `post_date` stayed pinned to
the reel. Highlights therefore gain real dates — today they fall back to
directory mtime.
### The timezone is not UTC
JD2 stamped filenames in **desktop local time (US Eastern)**. Measured across
212 comparable posts:
| model | mismatches |
|---|---:|
| UTC | 19 |
| UTC5 (EST) | 10 |
| UTC4 (EDT) | **0** |
| America/New_York (DST-aware) | **0** |
`{date:Olocal/%Y-%m-%d}` uses the machine's local zone with per-timestamp DST
awareness, which reproduces it — `mattellite` is `America/Toronto`, the same
offsets. Note the **trailing `/` must be omitted**: `Olocal/%Y-%m-%d/` puts the
separator into the strftime format and it sanitises to an underscore, giving
`2026-08-15__0ct0ber19`.
If the sync ever moves to a host in another timezone, set an explicit
`{date:O-4/…}` or the dates will silently shift for ~9% of posts.
### Caption sidecars
JD2 writes one `.txt` per post, named without the index, containing the caption
with **no trailing newline**, and writes nothing when the caption is empty
(measured: 197 of 217 posts, 86 of 86 reels, 0 of 10 stories, 0 of 16
highlights). gallery-dl reproduces this exactly with the default
`"empty": false`:
```jsonc
{ "name": "metadata", "event": "post", "mode": "custom",
"content-format": "{description}", "extension": "txt",
"filename": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.txt" }
```
`"event": "post"` is what makes it one file per post rather than per media file.
### Metadata sidecar (new — JD2 had no equivalent)
```jsonc
{ "name": "metadata", "event": "post", "mode": "json",
"filename": "{date:Olocal/%Y-%m-%d}_{username} - {post_shortcode}.json",
"include": ["post_shortcode","post_id","type","date","post_date","username",
"fullname","owner_id","description","count","likes","post_url",
"sidecar_shortcode"] }
```
Use **`include`**, not `fields``fields` is for `mode: custom` and silently
does nothing here, leaving `audio_user` blobs (including another user's profile
picture URL) in the output.
The payoff is `type`, which is Instagram's own classification:
```json
{ "post_shortcode": "DbdG9L9jU4m", "type": "post", "count": 2 } // feed video
{ "post_shortcode": "Db-lNCoib9m", "type": "reel", "count": 1 } // real reel
```
This is the `product_type: "clips"` signal, delivered free in the listing
response. It is the authoritative answer to "is this a reel", and would let the
viewer retire the lone-video heuristic in `src/lib/post-tabs.ts` — see
"Scanner work" below.
**`type` is only populated by listing extractors.** Extracting a single
`/p/<shortcode>/` URL leaves it `null`. Sync always uses listing URLs, so this
only matters when testing by hand.
## Cadence, and the budget that enforces it
**Monthly for everything, daily for stories only.** Stories expire in 24h and
cannot be backfilled, so they are the one surface where missing a day means
losing the content permanently. Everything else can wait — the skip-archive
means an infrequent full sync costs barely more than a frequent one, because it
only fetches what is new.
```
# monthly, everything
gdl-sync.py --index <viewer-url> --staging ~/gdl/staging \
--publish <user>@<nas>:<archives> --archive-db ~/gdl/artms.db \
--urls-file artms_account_links.txt --execute
# daily, stories only -- one request per profile
gdl-sync.py ... --only stories --execute
```
A stories-only run is one source per profile and **never seeds**, because a
story cannot be in the archive before it is fetched; probing would double the
cost of the cheapest surface for no benefit. Six profiles is a handful of
requests.
When scheduling it, **randomise the minute and avoid the hour boundary**. A job
that fires at exactly 09:00 every day is a machine; one that fires somewhere in
a window looks like someone opening the app.
The tool now refuses to repeat itself:
| flag | default | what it prevents |
|---|---|---|
| `--min-interval` | 20h | re-fetching a source touched recently — the aborted-restart case that re-enumerated five profiles |
| `--probe-ttl` | 24h | paying for a listing pass twice within a run cycle |
| `--max-sources` | off | a runaway list touching more than intended |
| `--force` | off | (escape hatch: ignores both guards) |
State lives beside the archive DB as `<db>.state.json`, recording per source
when it was seeded and last fetched. **Seeding is a one-time bootstrap**: after
the first successful sync the archive DB records everything gallery-dl has
seen, so the source is never probed again. That is the single biggest saving
here — a second full sync costs roughly half what the first did.
## Incremental sync — why the fetch host needs no copy of the archive
gallery-dl can skip already-held media two ways, and the difference decides
whether the fetcher needs the archive mounted:
- **By file existence** (default). Needs the destination to already contain the
files, so it only works if the archive is mounted where gallery-dl writes.
- **By skip-archive** (`--download-archive`). A sqlite DB of ids. Needs nothing
on disk.
We use the second, so the fetch host can write to **local disk and rsync
afterwards**. That avoids writing tens of thousands of small files over CIFS,
and keeps a mid-sync failure from leaving partial files on the live Resilio
share.
The key is `archive_prefix + archive_fmt`, which for this extractor is the
literal `instagram` plus the per-media numeric pk (`instagram.py:25`,
`job.py:713-719`). Verified: a 3-image carousel produced
```
instagram3079387627521318672
instagram3079387627521429433
instagram3079387627529716672
```
and a second run skipped every media file, rewriting only the idempotent
`.txt`/`.json` sidecars.
**Seeding.** `media_id` is not in our filenames, so the DB cannot be built from
names alone — but one listing pass (the pass we make anyway) maps every live
item to its `media_id`, and the archive's *file listing* says which we already
hold. No extra Instagram requests, and no archive content — a listing is
enough, which `GET /api/archives/:name/files` already serves.
Measured on `0ct0ber19`: 2275 live media items, 2248 seeded from the existing
listing, **27 left to download** — precisely the media of the two posts added
since the last crawl.
The one trap, which silently seeds almost nothing if you get it backwards:
| surface | filed under | why |
|---|---|---|
| posts, reels | `post_shortcode` | carousel children each have their own `shortcode`, which never appears in a filename |
| stories, highlights | `shortcode` (per item) | `post_shortcode` is the containing reel's id, shared by every item |
`live_key()` encodes this. Matching on the wrong field seeded 5 of 2275.
### The skip-archive saves the CDN, not `instagram.com`
Worth being exact about, because the two costs land on different surfaces and
only one of them bans accounts:
| what | which surface | scales with |
|---|---|---|
| downloading media | `scontent-*.cdninstagram.com` | how much is **new** |
| enumerating the profile to find it | `instagram.com` | how **big** the profile is |
The skip-archive suppresses the first. It does nothing about the second, so a
2275-post profile costs ~76 pages of pagination every run, forever, whether it
has three new posts or none. Seeding (above) saved a *second* full pass, not
the first.
Measured on the 2026-08-20 run, from sidecar write times in staging — free,
since the run was paying for the listing anyway:
```
1787248852 2026-08-19 … DcOeoVxkthi new, +0s
1787248944 2026-08-18 … DcLpfoJCZtp new, +92s
1787249058 2026-08-17 … DcIlGbxCUk0 new, +114s
1787249162 2026-07-24 … DbKr1TxlPSX ┐ all one second: nothing
1787249162 2026-08-15 … DcD-FdBCYGm ┘ downloaded, sidecars only
```
Three posts took ~100s each; the remaining 2272 were enumeration with nothing
to show for it.
**Pinned posts do not break early abort.** Test case 16 previously claimed
`0ct0ber19` returns its 3 pinned posts out of date order — that is true of the
*web grid*, but the REST `/posts/` listing came back strictly
reverse-chronological, newest first, no hoisting. That matters because
front-loaded old posts are the one thing that would make `skip: abort:N`
dangerous: it would trip on them and abort before reaching anything new.
So `skip: abort:N` is viable, and cuts ~420 requests per run to ~40-60:
| surface | live items | pages | with `abort:50` |
|---|---:|---:|---:|
| posts, 6 profiles | 11,248 | ~377 | ~12 |
| reels, 6 profiles | 1,080 | ~24 | ~8 |
| stories + highlights | — | ~20 | ~20 |
N counts consecutive skipped **files**, not posts, so it must clear the largest
already-held carousel — `DcD-FdBCYGm` alone is 22 media. 50 is comfortable; 5
would not be.
**The tradeoff is edited carousels.** Test case 15 is a post that gained items
after we archived it, and only a full enumeration finds those. Suggested
policy: `abort:50` for routine runs, a full sweep occasionally.
Measured the same day, resuming a stopped run with `--abort 50`:
| source | live items | enumerated |
|---|---:|---:|
| `cher_ryppo` posts | 2,151 | **7** |
| `cher_ryppo` reels | 92 | 53 |
One page instead of 72, and every new post was still caught. The 7 is roughly
3 new posts plus 4 already-held carousels making up the 50 skipped files.
Reels need 53 because they are single-media, so 50 consecutive skips really is
50 reels — another reminder that N counts files, and that the same N behaves
very differently on a carousel-heavy surface than on a reels tab.
## Publishing
The fetch host stages to local disk and rsyncs afterwards. `rsync
--ignore-existing` is not an optimisation but the safety property: the archive
deliberately outlives Instagram, so publishing must only ever **add**. No
`--delete`, and nothing already present is overwritten — including sidecars,
which are rewritten every run and would otherwise churn the synced share.
Publishing happens once at the end of a run, so a profile that fails midway
never reaches the archive half-written.
## Status
In use for all six ARTMS profiles.
`withaseul` first — 322 files added (74 media, 241 `.json`, 7 `.txt`), nothing
overwritten or deleted. Of the 74 new media, **zero** duplicated media already
held under a different name, which is the check that says JD2 and gallery-dl
naming really do converge.
**2026-08-20**, the first full incremental sync, four days after the previous
one. 184 new media, 299 files published, 0 failures and **0 CDN 429s**:
| profile | posts | reels | stories | files added |
|---|---:|---:|---:|---:|
| 0ct0ber19 | 58 | 2 | 4 | +77 |
| official_artms | 12 | — | 2 | +85 |
| cher_ryppo | 41 | 1 | 8 | +63 |
| zindoriyam | 23 | — | 4 | +35 |
| kimxxlip | 16 | — | 2 | +23 |
| withaseul | 10 | — | — | +16 |
The 20 story items are the part that could not have been recovered later.
Two things made it cheap, and both are worth keeping:
- The archive DB was already seeded from the previous run, so `--min-interval`
and the recorded `seeded` state meant **no probe passes at all**. A state
file has to exist for this; if one is missing after a manual run, write it
rather than letting the tool re-seed 24 sources.
- `--abort 50` (see above) cut the remaining listing cost by roughly 85%.
The run was deliberately **stopped and resumed** halfway to pick up `--abort`.
That is safe precisely because of the state file: the 12 finished sources were
already marked `fetched`, so the 20h floor skipped them and only the remaining
12 re-ran. Stopping a run is cheap now; it was not before.
Published files land owned by the SSH user rather than `rslsync`. The viewer
reads them fine (world-readable), but Resilio does not own what it syncs; worth
a `chown` if that ever matters. This also makes **`rsync` exit 23**
("some files/attrs were not transferred") the *normal* outcome of a publish —
it is the failed `chown`, not lost data. Confirm by re-running the same rsync
with `--dry-run`: an empty file list means everything arrived.
The profiles to fetch live in `artms_account_links.txt` at the archive root,
passed with `--urls-file`.
## Verified run
`withaseul`, all four surfaces, staged locally and published to a scratch
directory before the live publish above:
```
==> withaseul / posts seeded 915 of 984 live items
==> withaseul / reels seeded 28 of 34 live items
==> withaseul / stories no results (none active)
==> withaseul / highlights no results
```
Output landed correctly, including the collab-reel case — `withaseul - reels`
contains 53 files owned by `withaseul`, 10 by `cher_ryppo`, 3 by `0ct0ber19`
and 2 by `official_artms`, all with the owner in the filename and the crawl
scope as the directory.
### The CDN rate-limits, and the first run tripped it
At `rate: 3M` with `sleep: [1.0, 3.0]`, `scontent-*.cdninstagram.com` returned
**`429 Too Many Requests`** and two videos were lost (gallery-dl retried, then
gave up with exit 4). This is the *tolerant* surface complaining, which is a
clear signal the pacing was too aggressive.
Defaults are now:
| option | value |
|---|---|
| `--rate` | `1M` |
| `--sleep-request` | 610 s |
| `--sleep` | 36 s |
| `sleep-429` | 120 s |
| `retries` (extractor and downloader) | 8 |
Re-running with those recovered both videos and produced **0 failures and 0
429s**. Do not raise them for speed; an archive sync has no deadline.
### yt-dlp is worth installing
Without it, gallery-dl logs `Cannot import yt-dlp or youtube-dl` and falls back
to a progressive URL for DASH videos. The fallback mostly works but is what the
429s hit hardest.
**`pipx install yt-dlp` does not work** — it was the advice here until
2026-08-20, and it is wrong. It gives yt-dlp its own venv, so the binary lands
on `PATH` while gallery-dl, in a *different* venv, still cannot `import yt_dlp`.
The symptom is that everything looks installed and the log keeps saying
`Cannot import yt-dlp`. gallery-dl needs it importable, not runnable:
```sh
pipx inject gallery-dl yt-dlp
```
Verify by asking gallery-dl's own interpreter, not the shell:
```sh
/home/matt/.local/share/pipx/venvs/gallery-dl/bin/python -c 'import yt_dlp'
```
## Known quirks
- **`count` is not the emitted file count.** For 135 of 214 posts it was exactly
one higher than the number of files written. This makes the `count >= 10`
padding condition mis-pad a handful of 9-item posts (10 of 214 measured). Since
the parser normalises the index, this is cosmetic — but it means a re-fetch
over an existing JD2 tree writes `- 01.jpg` beside an existing `- 1.jpg`.
- **Carousels get edited.** Two posts had a different media count live than on
disk. Padding width follows the count *at download time*, so a grown carousel
produces mixed widths — the archive already contains one such post from JD2.
- **Highlights already have two naming styles on disk**, and every undated file
has a dated twin. The scanner dedupes by index so they render once; it is
wasted disk, not a display bug.
## Scanner work (not done yet)
`useArchiveScanner` currently treats any `.json` in the tree as a possible
manifest. Adding gallery-dl sidecars needs it to distinguish three things:
1. Instagram export manifests (`posts_1.json`) — existing path.
2. Instaloader `.json.xz` — existing path, GraphQL node shape.
3. gallery-dl `.json` — new, flat shape, identified by having
`post_shortcode` + `type` at the top level.
Once (3) is read, `source`/`isStory` and the reel flag should come from `type`
rather than from the directory and the lone-video heuristic.
## Test cases
Real subjects, all present in the archive today. See
`scripts/gdl-sync.py --selftest` for the harness.
| # | case | shortcode | expected |
|---|---|---|---|
| 1 | single image | `CwcXnQhOqFG` | one `.jpg`, no index |
| 2 | single feed video | `DbdG9L9jU4m` | one `.mp4`, `type: post` |
| 3 | carousel, images only | `Cq8LrxSJAJE` | `- 1 … - 3` |
| 4 | carousel, image + video | `CtohvHxLnWO` | `- 1.jpg … - 4.mp4`, **no `.txt`** |
| 5 | carousel of exactly 9 | `Cv2Hb_brx_N` | 1-digit index |
| 6 | carousel of 10+ | `CzM8Uf6B6H_` | 2-digit index `- 01 … - 10` |
| 7 | reel shown on the posts grid | `C8FHM6EJl15` | in `<user>`, `type: reel` |
| 8 | reel on the reels tab | `Db-lNCoib9m` | in `<user> - reels`, `type: reel` |
| 9 | collab reel (other owner) | `DYcZOb0h6Sv` | dir `0ct0ber19 - reels`, filename `chuuo3o` |
| 10 | story | live only | `story - <user>`, per-item shortcode + date |
| 11 | story highlight | `C-IImhvpFuk` | `story highlights - <user> - <title>` |
| 12 | highlight, unicode title | `Drawheeing` | trailing U+2800 preserved in dirname |
| 13 | empty caption | `CrdsY5CrSsO` | media written, `.txt` absent |
| 14 | deleted post | `C0TgI7sphfZ` | on disk, absent live — must not be removed |
| 15 | edited carousel | `C7zG7-jJMlq` | 18 on disk, 8 live — must not be removed |
| 16 | pinned posts | `0ct0ber19` | REST listing is strictly reverse-chronological; see below |
| 17 | profile avatar | `0ct0ber19.jpg` | base dir, undated |
Cases 1416 are reconciliation, not naming: **a sync must never delete**, since
the archive deliberately outlives Instagram.
Not covered, decide before relying on them: the `/reposts/` tab (`0ct0ber19`
has one) and `/tagged/`. Neither is fetched today.
+181
View File
@@ -0,0 +1,181 @@
# JDownloader2 — archive fetching quick reference
How content gets into this archive, and why the setup is shaped the way it is.
## Why JDownloader and not Instaloader
There are two surfaces, and they're treated very differently:
| Surface | What hits it | Risk |
|---|---|---|
| `instagram.com` | profile pages, GraphQL/API metadata | Tied to your session, heavily rate-limited. **This is where bans come from.** |
| `scontent*.cdninstagram.com` | the actual media | Signed URLs, CDN-served, tolerant. Mostly a bandwidth question. |
JDownloader does nearly all its work on the CDN. Instaloader's value — the rich
`.json.xz` metadata — comes from asking `instagram.com` a question *per post*.
Concretely, from this archive: `rivvsofficial` has 188 post-metadata files, so
backfilling it cost 188 API requests for one 605-file profile. That's the ban
vector. Downloading the 238 photos was never the problem.
Instaloader got this account banned once. JDownloader with throttling did not.
> **The account was suspended anyway, on 2026-08-17, for "spam".** Not by
> JDownloader, and not by downloading. It was suspended during a day of
> *building and verifying* the gallery-dl replacement — automated browser
> scrolling to enumerate profile grids, repeated `--simulate` and `-j` metadata
> passes, and one aborted sync that re-ran every listing pass before dying.
>
> The framing above is right about which surface is dangerous and wrong about
> what reaches it. **Every read of `instagram.com` counts, including the ones
> that download nothing** — and read-only work is easy not to count precisely
> because it leaves no files behind. See the post-mortem at the top of
> `docs/gallery-dl.md`.
>
> The rule that would have prevented it: *verify against the archive, never
> against the live site*, and treat the first CDN `429` as the end of the
> session rather than a pacing knob.
### What the metadata gap actually costs
Comparing a JDownloader profile against an Instaloader one:
| | JDownloader | Instaloader |
|---|---|---|
| Media | ✅ | ✅ |
| Captions (`.txt`) | ✅ | ✅ |
| Dates (from filenames) | ✅ | ✅ |
| Bio / full name | ❌ | ✅ |
| Follower counts | ❌ | ✅ |
| External URL | ❌ | ✅ |
Captions already work — the viewer reads the `.txt` sidecars. Everything missing
lives in a *single* profile-level record, not the per-post ones. That's why
JDownloader-sourced profiles show "0 followers" and a placeholder bio.
Not worth extra requests. If you ever want it, the zero-request option is a
hand-written `profile.json` sidecar (not implemented yet — ask).
## Settings that matter
**Chunks per download → 1.** The single most important one. JDownloader splits
each file into multiple ranged requests by default; that `Range` pattern looks
nothing like a browser or the app. One chunk = one sequential GET per file.
`jd2-sync` sets `chunks=1` per job, so no global change is needed — but set it
globally too if you ever add links by hand.
**Max simultaneous downloads → 23**, connections-per-host low. Concurrency is
what turns "a user" into a statistic.
**Leave reconnect / IP-change features off.** A mid-session IP change on a live
cookie is a *stronger* anomaly signal than the request rate you'd be avoiding.
## The cookie
Exported manually from a real browser session. This is the right approach — no
programmatic login anywhere, which is the thing that actually gets flagged.
- Use it from the **same public IP** as the browser it came from. A cookie used
from a different network is what session-hijack detection looks for.
- When it expires, **re-export from the browser**. Never add a login step to a tool.
- It's a full account credential. Keep it off the NAS share and out of the repo.
## Workflow
Two URLs per profile, because the profile grid misses some reels:
```
https://www.instagram.com/<user>/
https://www.instagram.com/<user>/reels/
```
They overlap slightly — a reel caught by both lands in each directory and shows
up twice in the viewer. That's correct and matches Instagram, which also shows
reels in the profile grid *and* the Reels tab.
## Generating jobs
Instead of pasting URLs and setting output folders by hand:
```bash
npm run jd2 -- --archives /volume1/rslsync/sync/Instagram-archive/archives --dry-run
```
Review, then write it into JDownloader's folder-watch directory:
```bash
npm run jd2 -- --archives /volume1/rslsync/sync/Instagram-archive/archives \
--out ~/.jd2/folderwatch
```
JDownloader runs on the desktop while the archive lives on the NAS, so tell it
the path *it* sees:
```bash
npm run jd2 -- --archives /mnt/nas/Instagram-archive/archives \
--download-base 'Z:\Instagram-archive\archives' \
--out ~/.jd2/folderwatch
```
| Flag | Purpose |
|---|---|
| `--archives <dir>` | Archive root to scan (or `$ARCHIVES_DIR`) |
| `--out <dir>` | JDownloader folder-watch directory |
| `--download-base <dir>` | Root path as JDownloader sees it (Windows paths fine) |
| `--user <name>` | Just this profile (repeatable) |
| `--skip <name>` | Never emit jobs for this directory (repeatable) |
| `--chunks <n>` | Connections per file (default 1) |
| `--auto-start` | Start immediately instead of parking in LinkGrabber |
| `--all-reels` | Emit a reels job even where no reels directory exists |
| `--dry-run` | Print instead of writing |
Defaults are deliberately conservative: `chunks=1`, and links park in the
LinkGrabber for review rather than auto-starting.
Only posts and reels are emitted. Highlight URLs need a numeric id and story
URLs expire, so those stay manual.
Directories that aren't Instagram profiles are skipped by username shape
(letters, digits, dots, underscores, ≤30 chars) — pointing a crawl at those
spends `instagram.com` requests to be told the profile doesn't exist. For names
that *look* like usernames but aren't, use `--skip` or a `.jd2ignore` file in
the archive root, one name per line.
Format reference: `src/org/jdownloader/extensions/folderwatchV2/explain.txt`.
JDownloader develops on SVN — read it via the daily mirror at
<https://github.com/mycodedoesnotcompile2/jdownloader_mirror> (`svn_trunk/`),
not one of the abandoned GitHub copies.
## Expected layout
Everything downloads into `<archives>/`, one directory per source:
```
archives/
0ct0ber19/ posts
0ct0ber19 - reels/ reels
story - 0ct0ber19/ stories
story highlights - 0ct0ber19 - Heestory/ a highlight
```
Non-archive directories (tool output, exports from elsewhere) live *outside*
`archives/` so they never reach the viewer.
The server picks up changes automatically — its index is keyed on directory
mtime, so a new file invalidates only that directory.
## If something goes wrong
**429 / rate limited** — stop for hours, not seconds. Retrying into a limit is
what converts a soft throttle into something worse.
**Cookie stops working** — re-export from the browser. Don't add a login step.
**Files land in the wrong folder** — a Packagizer rule is overriding the job.
Generated jobs set `overwritePackagizerEnabled=TRUE` to prevent this; check that
rules aren't set to run after it.
**Viewer doesn't show new posts** — check the file is in the right directory and
matches the naming pattern (`YYYY-MM-DD_<user> - <shortcode>[ - NN].<ext>`).
The index refreshes on directory mtime, so a genuinely new file is picked up on
the next request.
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "instaarchive-viewer",
"version": "1.3.2",
"version": "1.8.0",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "instaarchive-viewer",
"version": "1.3.2",
"version": "1.8.0",
"dependencies": {
"@tailwindcss/vite": "^4.1.14",
"@vitejs/plugin-react": "^5.0.4",
+3 -2
View File
@@ -1,7 +1,7 @@
{
"name": "instaarchive-viewer",
"private": true,
"version": "1.3.2",
"version": "1.8.0",
"type": "module",
"scripts": {
"dev": "vite --port=3000 --host=0.0.0.0",
@@ -12,7 +12,8 @@
"clean": "rm -rf dist",
"lint": "tsc --noEmit",
"test": "vitest run",
"test:watch": "vitest"
"test:watch": "vitest",
"jd2": "tsx scripts/jd2-sync.ts"
},
"dependencies": {
"@tailwindcss/vite": "^4.1.14",
Binary file not shown.
Binary file not shown.
+843
View File
@@ -0,0 +1,843 @@
#!/usr/bin/env python3
"""
Fetch Instagram profiles into the archive layout using gallery-dl.
The CLI replacement for the JDownloader2 workflow. See docs/gallery-dl.md for
the measurements behind every choice here — especially the safety model, which
is the reason this script exists in this shape rather than a simpler one.
The fetch host needs no copy of the archive. It stages locally and rsyncs
afterwards; what it already holds is learned from a *file listing* alone
(`--index`), which the viewer's own API serves.
Usage:
./scripts/gdl-sync.py --index https://instaarchive.ergosteur.com \\
--staging /var/tmp/gdl --publish user@host:/path/to/archives \\
--urls-file artms_account_links.txt --dry-run
# ...then swap --dry-run for --execute. --index also accepts a local path,
# and --profile / --all work instead of --urls-file.
Always --dry-run first: it prints the plan, and the publish step it reports is
the one that would touch the archive.
Run it from the host whose public IP matches the browser the cookie came from;
using the cookie from elsewhere is what session-hijack detection looks for.
"""
from __future__ import annotations
import argparse
import datetime as dt
import json
import os
import re
import shutil
import subprocess
import sys
from dataclasses import dataclass, field
from pathlib import Path
# --------------------------------------------------------------------------
# Archive layout
# --------------------------------------------------------------------------
# Mirrors src/lib/archive-grouping.ts. Instagram usernames cannot contain
# spaces, which is what makes the username separable from a highlight title.
RE_HIGHLIGHT = re.compile(r"^story highlights - ([^ ]+) - (.+)$")
RE_STORIES = re.compile(r"^story - ([^ ]+)$")
RE_REELS = re.compile(r"^([^ ]+) - reels$")
DATE_FMT = "{date:Olocal/%Y-%m-%d}"
"""Local-time date. JD2 stamped US Eastern, NOT UTC (0/212 mismatches vs 19 for
UTC). `Olocal` is DST-aware per timestamp. The trailing separator must be
omitted or it lands in the strftime format and sanitises to an underscore."""
POST_STEM = DATE_FMT + "_{username} - {post_shortcode}"
ITEM_STEM = DATE_FMT + "_{username} - {shortcode}"
@dataclass
class Source:
"""One gallery-dl invocation: a URL fetched into a specific directory."""
kind: str # posts | reels | stories | highlights
url: str
directory: str # relative to the archives root
subcategory: str # gallery-dl config key
title: str | None = None # highlight title, when known
@dataclass
class Profile:
user: str
existing: dict[str, str] = field(default_factory=dict) # kind -> dirname
def sources(self, kinds: set[str]) -> list[Source]:
u = self.user
base = f"https://www.instagram.com/{u}"
all_sources = [
Source("posts", f"{base}/posts/", u, "posts"),
Source("reels", f"{base}/reels/", f"{u} - reels", "reels"),
# Stories expire after 24h, so these can only ever be captured
# live. There is no backfill and no re-fetch -- which is why they
# are the one surface worth visiting daily.
Source("stories", f"https://www.instagram.com/stories/{u}/",
f"story - {u}", "stories"),
# Highlight directories embed the title, which gallery-dl only
# learns mid-extraction -- so this one source fans out into many
# directories and is handled with a directory format string.
Source("highlights", f"{base}/highlights", "", "highlights"),
]
return [s for s in all_sources if s.kind in kinds]
def scan_archives(root: Path) -> dict[str, Profile]:
"""Group existing directories into profiles, as the server does."""
profiles: dict[str, Profile] = {}
def get(user: str) -> Profile:
return profiles.setdefault(user, Profile(user))
for entry in sorted(os.listdir(root)):
if not (root / entry).is_dir() or entry.startswith("."):
continue
if m := RE_HIGHLIGHT.match(entry):
get(m.group(1)).existing.setdefault("highlights", entry)
elif m := RE_STORIES.match(entry):
get(m.group(1)).existing["stories"] = entry
elif m := RE_REELS.match(entry):
get(m.group(1)).existing["reels"] = entry
else:
get(entry).existing["posts"] = entry
return profiles
RE_PROFILE_URL = re.compile(
r"^(?:https?://)?(?:www\.)?instagram\.com/(?P<user>[^/?#\s]+)/?", re.I)
# Path segments that are Instagram features, not profiles. A line like
# ".../p/ABC123/" names a post, and treating "p" as a username would silently
# sync nothing under a nonsense directory.
RESERVED_SEGMENTS = {
"p", "reel", "reels", "stories", "explore", "accounts", "direct",
"tv", "s", "invites", "challenge", "about", "developer",
}
def read_urls_file(path: Path) -> list[str]:
"""
Read profile URLs (or bare usernames) from a file, one per line.
Written for hand-maintained lists: blank lines are skipped, `#` starts a
comment, and either a full URL or a bare username works. Order is kept and
duplicates dropped, so a list can be appended to without care.
"""
users: list[str] = []
seen: set[str] = set()
for lineno, raw in enumerate(path.read_text().splitlines(), 1):
line = raw.split("#", 1)[0].strip()
if not line:
continue
m = RE_PROFILE_URL.match(line)
user = m.group("user") if m else line.strip("/")
if not user or "/" in user or " " in user:
print(f"{path}:{lineno}: cannot read a username from {raw.strip()!r}",
file=sys.stderr)
continue
if user.lower() in RESERVED_SEGMENTS:
print(f"{path}:{lineno}: {user!r} is an Instagram path, not a "
f"profile — skipping", file=sys.stderr)
continue
if user in seen:
continue
seen.add(user)
users.append(user)
return users
class ArchiveIndex:
"""
What the archive already holds, as filenames only.
Deliberately never reads file *contents*, so the fetch host does not need a
copy of the archive — it can stage locally and rsync afterwards. Backed
either by a local directory or by the viewer's own API, which already
serves exactly this listing and is the cheaper option when the archive
lives on network storage (a full walk there took ~52s).
"""
def __init__(self, source: str):
self.remote = source.startswith(("http://", "https://"))
self.source = source.rstrip("/") if self.remote else None
self.root = None if self.remote else Path(source)
if self.root and not self.root.is_dir():
raise SystemExit(f"archive index not found: {source}")
self._cache: dict[str, list[str]] = {}
def _get(self, path: str):
from urllib.request import urlopen
with urlopen(f"{self.source}{path}", timeout=60) as resp:
return json.load(resp)
def profiles(self) -> set[str]:
if self.remote:
return {a["name"] for a in self._get("/api/archives")}
return set(scan_archives(self.root))
def listing(self, user: str) -> list[str]:
"""Every filename belonging to a profile, across all its sidecars."""
if user in self._cache:
return self._cache[user]
names: list[str] = []
if self.remote:
try:
data = self._get(f"/api/archives/{user}/files")
except Exception:
data = []
files = data if isinstance(data, list) else data.get("files", [])
names = [f["path"] for f in files]
else:
prof = scan_archives(self.root).get(user)
for dirname in (prof.existing.values() if prof else ()):
d = self.root / dirname
if d.is_dir():
names += [f"{dirname}/{n}" for n in os.listdir(d)]
self._cache[user] = names
return names
# --------------------------------------------------------------------------
# gallery-dl configuration
# --------------------------------------------------------------------------
def build_config(rate: str, sleep_request: list[float],
sleep: list[float], abort: int = 0) -> dict:
"""
The config is generated rather than checked in so the safety-critical
options cannot drift out of sync with the docs.
`api: rest` is the single most important line in this file. The graphql
backend issues one request PER POST for every video and carousel, which is
the pattern that got this account banned once already.
"""
caption_pp = {
"name": "metadata",
"event": "post",
"mode": "custom",
"content-format": "{description}",
"extension": "txt",
# JD2 wrote no .txt when the caption was empty; "empty": false (the
# default) reproduces that.
}
meta_pp = {
"name": "metadata",
"event": "post",
"mode": "json",
# `include`, NOT `fields` -- `fields` applies to mode:custom and
# silently does nothing here, dumping audio_user blobs that contain
# unrelated users' profile picture URLs.
"include": [
"post_shortcode", "post_id", "type", "date", "post_date",
"username", "fullname", "owner_id", "description", "count",
"likes", "post_url", "sidecar_shortcode",
],
}
def post_like(stem: str) -> dict:
"""Naming for surfaces whose unit is a post (posts, reels)."""
skip: dict = {}
if abort:
# Stop enumerating once `abort` consecutive files are already in
# the skip-archive. The listing pass -- not the downloading -- is
# what costs `instagram.com` requests, and it otherwise walks the
# whole profile every run to find three new posts.
#
# Safe here only because the REST listing is strictly
# reverse-chronological: the web grid hoists pinned posts to the
# front, but this endpoint does not (measured 2026-08-20), so old
# posts never appear before new ones.
#
# Counted in FILES, not posts, so it must clear the largest
# already-held carousel -- 22 media for one real post in this
# archive. It also means edited carousels (test case 15) stop
# being noticed, so a full sweep is still worth running
# occasionally.
skip["skip"] = f"abort:{abort}"
return {
**skip,
# `sidecar_shortcode` is set only for carousels, so it is the
# carousel discriminator. First matching condition wins.
"filename": {
"sidecar_shortcode and count >= 10":
stem + " - {num:02}.{extension}",
"sidecar_shortcode":
stem + " - {num}.{extension}",
"":
stem + ".{extension}",
},
"postprocessors": [
{**caption_pp, "filename": stem + ".txt"},
{**meta_pp, "filename": stem + ".json"},
],
}
def item_like(stem: str) -> dict:
"""
Naming for surfaces whose unit is an item inside a reel (stories,
highlights). `{shortcode}` is per item; `{post_shortcode}` is the
reel's id and is shared by every item in it.
The media filename uses the per-item shortcode, but the sidecar cannot:
it runs at `event: post`, where the kwdict describes the *reel* and has
no `shortcode` at all -- which silently formatted as the literal
"None", producing one "<date>_<user> - None.json" per reel. It is keyed
by `post_shortcode` instead, and is genuinely reel-level data (the
reel's own date and item count); per-item dates live in the media
filenames, which is the more precise source anyway.
"""
return {
"filename": stem + ".{extension}",
"postprocessors": [
{**meta_pp,
"filename": DATE_FMT + "_{username} - {post_shortcode}.json"},
],
}
return {
"extractor": {
"base-directory": ".",
"instagram": {
"api": "rest", # never "graphql" -- see docstring
"sleep-request": sleep_request,
"sleep": sleep,
# The CDN does rate-limit: a first run at 3M/1-3s drew
# '429 Too Many Requests' from scontent-*.cdninstagram.com and
# lost two videos. Back off hard rather than retry fast.
"sleep-429": 120.0,
"retries": 8,
"videos": True,
"include": "", # never "all"; sources are explicit
# Directory is forced per-invocation with -D, because a reels
# tab returns collab reels owned by OTHER accounts and
# {username} would scatter them into the wrong profile.
"directory": [],
"posts": post_like(POST_STEM),
"reels": post_like(POST_STEM),
"stories": item_like(ITEM_STEM),
"highlights": {
**item_like(ITEM_STEM),
# The only surface that must derive its own directory,
# since the title is not known until extraction.
"directory": ["story highlights - {username} - {highlight_title}"],
},
},
},
# `retries` here is the CDN-side counterpart to sleep-429 above.
"downloader": {"http": {"rate": rate, "retries": 8}},
"output": {"mode": "null"},
}
# --------------------------------------------------------------------------
# Planning and execution
# --------------------------------------------------------------------------
def gdl_command(src: Source, staging: Path, config: Path, cookies: str,
archive_db: Path | None) -> list[str]:
cmd = [
"gallery-dl",
"--config", str(config),
"--cookies-from-browser", cookies,
]
if archive_db:
# Without a seeded skip-archive, staging is empty and every file is
# re-downloaded; see seed_archive_db.
cmd += ["--download-archive", str(archive_db)]
# Forced destination -- never `{username}` -- because a reels tab returns
# collab reels owned by other accounts, which would otherwise be filed
# under the wrong profile. Highlights are the exception: their directory
# embeds a title only known mid-extraction, so the config formats it.
dest = staging if src.subcategory == "highlights" else staging / src.directory
cmd += ["--destination", str(dest)]
cmd.append(src.url)
return cmd
# gallery-dl keys its skip-archive on `archive_prefix + archive_fmt`, which for
# this extractor is the literal "instagram" followed by the per-media numeric
# pk (`instagram.py:25`, `job.py:713-719`). Verified against a real run: a
# 3-image carousel produced 3 rows, one per item.
ARCHIVE_KEY = "instagram{}".format
ARCHIVE_SCHEMA = "CREATE TABLE IF NOT EXISTS archive (entry TEXT PRIMARY KEY)"
RE_ARCHIVED = re.compile(
r"^(\d{4}-\d{2}-\d{2})_(.+?) - ([A-Za-z0-9_-]+?)(?: - (\d+))?\.(\w+)$")
NON_MEDIA = {"txt", "json"}
def index_existing(listing: list[str]) -> set[tuple[str, int]]:
"""
Reduce a flat list of filenames to the (shortcode, index) pairs already
held. Only names matter — never the bytes — which is what lets the sync run
on a host that has no copy of the archive.
"""
have: set[tuple[str, int]] = set()
for name in listing:
m = RE_ARCHIVED.match(name.rsplit("/", 1)[-1])
if not m or m.group(5).lower() in NON_MEDIA:
continue
# An absent index means a single-media post, which is index 1 — the
# same normalisation the viewer's EXPORT_RE applies.
have.add((m.group(3), int(m.group(4) or 1)))
return have
def live_key(item: dict, kind: str) -> tuple[str, int]:
"""
The (shortcode, index) a live item *would* be filed under, mirroring the
filename template exactly.
The two surfaces disagree about which shortcode identifies a file, and
getting this wrong silently seeds almost nothing:
posts/reels filed under {post_shortcode} — for a carousel, each
child item ALSO has its own `shortcode`, which is not
what appears in the filename.
stories/highlights filed under the per-item {shortcode}, because
`post_shortcode` there is the containing reel's id and
is shared by every item in it.
"""
if kind in ("stories", "highlights"):
return (item.get("shortcode"), 1)
return (item.get("post_shortcode"), item.get("num"))
def seed_archive_db(db: Path, existing: set[tuple[str, int]],
live: list[dict], kind: str) -> int:
"""
Mark everything already held as downloaded, so a fetch into an empty
directory pulls only what is missing.
`live` is the metadata of one listing pass — the pass we have to make
anyway — each entry carrying at least `media_id` plus the shortcode fields
`live_key` needs. Seeding costs no additional Instagram requests, and needs
only a *listing* of the archive, never its contents.
"""
import sqlite3
db.parent.mkdir(parents=True, exist_ok=True)
con = sqlite3.connect(db)
con.execute(ARCHIVE_SCHEMA)
rows = [
(ARCHIVE_KEY(item["media_id"]),)
for item in live
if live_key(item, kind) in existing
]
con.executemany("INSERT OR IGNORE INTO archive (entry) VALUES (?)", rows)
con.commit()
con.close()
return len(rows)
def probe_live(src: Source, config: Path, cookies: str) -> list[dict]:
"""
One metadata-only listing pass. `sleep` is forced to 0 because it otherwise
applies per *file* even with no download — 2275 files at 1-3s each is over
an hour for a single profile.
"""
out = subprocess.run(
["gallery-dl", "-j", "--config", str(config),
"--cookies-from-browser", cookies, "-o", "sleep=0", src.url],
capture_output=True, text=True, check=True,
)
items: list[dict] = []
def walk(node):
if isinstance(node, dict):
if "media_id" in node and "shortcode" in node:
items.append(node)
for value in node.values():
walk(value)
elif isinstance(node, list):
for value in node:
walk(value)
walk(json.loads(out.stdout))
return items
ALL_KINDS = ("posts", "reels", "stories", "highlights")
# Stories cannot be backfilled and expire in 24h, so a run that only wants
# stories is both cheap and the one worth scheduling daily.
STORIES_ONLY = {"stories"}
class SyncState:
"""
What has already been spent against `instagram.com`.
Exists because nothing else in this tool has any memory: every invocation
used to start from zero and happily re-enumerate profiles it had listed
minutes earlier. That is what suspended the account — the listing passes,
not the downloads.
Two facts are tracked per source:
seeded the skip-archive has been primed from the archive listing.
This is a ONE-TIME bootstrap: afterwards the archive DB records
every item gallery-dl has seen, so the source never needs
probing again. This is the single biggest request saving here.
fetched when it was last downloaded, so a re-run soon after is refused
rather than silently repeating the whole pass.
"""
VERSION = 1
def __init__(self, path: Path):
self.path = path
self.data = {"version": self.VERSION, "sources": {}}
if path.is_file():
try:
loaded = json.loads(path.read_text())
if loaded.get("version") == self.VERSION:
self.data = loaded
except Exception:
pass # a corrupt state file must never block a sync
def _entry(self, url: str) -> dict:
return self.data.setdefault("sources", {}).setdefault(url, {})
def needs_seed(self, url: str) -> bool:
return not self._entry(url).get("seeded")
def mark_seeded(self, url: str, stamp: str) -> None:
self._entry(url)["seeded"] = stamp
def last_fetch(self, url: str) -> str | None:
return self._entry(url).get("fetched")
def mark_fetched(self, url: str, stamp: str) -> None:
self._entry(url)["fetched"] = stamp
def save(self) -> None:
self.path.parent.mkdir(parents=True, exist_ok=True)
self.path.write_text(json.dumps(self.data, indent=1, sort_keys=True))
def hours_since(stamp: str | None, now: float) -> float:
"""Hours between an ISO stamp and `now`; infinite when never."""
if not stamp:
return float("inf")
try:
then = dt.datetime.fromisoformat(stamp)
except ValueError:
return float("inf")
if then.tzinfo is None:
then = then.replace(tzinfo=dt.timezone.utc)
return (now - then.timestamp()) / 3600.0
def plan_source(src: Source, state: SyncState, now: float,
min_interval: float) -> tuple[bool, bool, str]:
"""
Decide what a source needs: (fetch, seed, reason).
Seeding is skipped once done, and skipped entirely for stories — a story
cannot exist in the archive before it is fetched, so there is nothing to
seed from, and probing would double the request cost of the cheapest
surface we have.
"""
since = hours_since(state.last_fetch(src.url), now)
if since < min_interval:
return (False, False, f"fetched {since:.1f}h ago, under the "
f"{min_interval:g}h floor")
if src.kind == "stories":
return (True, False, "stories: no seed needed")
if state.needs_seed(src.url):
return (True, True, "first run: seeding from the archive listing")
return (True, False, "already seeded; the skip-archive knows what we hold")
class ProbeCache:
"""
Listing-pass results, kept so an interrupted run does not pay for them
twice. Yesterday an aborted sync re-enumerated five profiles on restart.
"""
def __init__(self, path: Path, ttl_hours: float):
self.path = path
self.ttl = ttl_hours
self.data: dict = {}
if path.is_file():
try:
self.data = json.loads(path.read_text())
except Exception:
self.data = {}
def get(self, url: str, now: float) -> list[dict] | None:
entry = self.data.get(url)
if not entry or hours_since(entry.get("at"), now) > self.ttl:
return None
return entry.get("items")
def put(self, url: str, items: list[dict], stamp: str) -> None:
# Only the fields seeding needs, so the cache stays small.
self.data[url] = {"at": stamp, "items": [
{k: i.get(k) for k in ("shortcode", "post_shortcode", "num", "media_id")}
for i in items
]}
def save(self) -> None:
self.path.parent.mkdir(parents=True, exist_ok=True)
self.path.write_text(json.dumps(self.data))
def rsync_command(staging: Path, dest: str, dry_run: bool) -> list[str]:
"""
Publish a staging tree into the archive.
`--ignore-existing` is not an optimisation, it is the safety property: the
archive deliberately outlives Instagram (posts exist here that Instagram no
longer serves), so publishing must only ever *add*. No `--delete`, and
nothing already present is overwritten — including sidecars, which get
rewritten on every run and would otherwise churn the synced share.
`dest` may be a local path or any rsync destination (`user@host:/path`),
because the archive usually is not writable from the fetch host.
"""
cmd = ["rsync", "-a", "--ignore-existing", "--partial", "--info=stats2",
# Belt and braces: the config lives outside staging, but nothing
# resembling tooling output should ever reach the archive. Archive
# sidecars are always "<date>_<user> - <code>.json", so none of
# these can match real content.
"--exclude", "gdl-sync*.json",
"--exclude", "*.gdl-config.json",
"--exclude", ".gdl-*",
"--exclude", "*.sqlite", "--exclude", "*.db"]
if dry_run:
cmd.append("--dry-run")
# Trailing slash: copy the *contents* of staging into dest.
cmd += [f"{staging}/", dest if dest.endswith("/") else dest + "/"]
return cmd
def publish(staging: Path, dest: str, dry_run: bool) -> int:
if not any(staging.iterdir()):
print(" nothing staged; skipping publish")
return 0
cmd = rsync_command(staging, dest, dry_run)
print(" " + " ".join(cmd))
return subprocess.run(cmd).returncode
def main() -> int:
# A sync runs for hours and is normally watched through a redirected log,
# where Python's block buffering would withhold progress until it happened
# to flush -- and the gallery-dl subprocesses write to the same descriptor
# unbuffered, so the log would also interleave out of order.
sys.stdout.reconfigure(line_buffering=True)
sys.stderr.reconfigure(line_buffering=True)
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--index", required=True,
help="existing archive listing: a local root, or the "
"viewer's base URL (only a FILE LISTING is needed, "
"never the contents)")
ap.add_argument("--publish", required=True,
help="rsync destination for fetched files; a local path or "
"user@host:/path")
ap.add_argument("--staging", type=Path, required=True,
help="local scratch directory gallery-dl writes into")
g = ap.add_mutually_exclusive_group(required=True)
g.add_argument("--profile", action="append", default=[],
help="profile to sync; repeatable")
g.add_argument("--all", action="store_true", help="every profile on disk")
g.add_argument("--urls-file", type=Path,
help="file of Instagram profile URLs or usernames, one per "
"line; # comments and blank lines allowed")
ap.add_argument("--cookies", default="chrome:/home/matt/.config/google-chrome-devtools",
help="gallery-dl --cookies-from-browser value")
ap.add_argument("--archive-db", type=Path, default=None,
help="gallery-dl skip-archive sqlite path")
ap.add_argument("--rate", default="1M", help="per-download rate cap")
ap.add_argument("--sleep-request", nargs=2, type=float, default=[6.0, 10.0],
metavar=("MIN", "MAX"))
ap.add_argument("--sleep", nargs=2, type=float, default=[3.0, 6.0],
metavar=("MIN", "MAX"))
ap.add_argument("--only", default=",".join(ALL_KINDS),
help="comma-separated surfaces to sync: "
"posts,reels,stories,highlights. Use --only stories "
"for the cheap daily run.")
ap.add_argument("--min-interval", type=float, default=20.0, metavar="HOURS",
help="refuse to re-fetch a source touched more recently "
"than this (default 20h); the guard that makes a "
"restart cheap instead of a repeat")
ap.add_argument("--max-sources", type=int, default=0, metavar="N",
help="hard ceiling on sources touched in one run "
"(0 = no limit)")
ap.add_argument("--abort", type=int, default=0, metavar="N",
help="stop enumerating posts/reels after N consecutive "
"already-archived FILES (0 = walk everything, the "
"default). 50 is a safe routine value; it cuts the "
"per-run listing cost by roughly 85%%, at the price "
"of no longer noticing edited carousels")
ap.add_argument("--probe-ttl", type=float, default=24.0, metavar="HOURS",
help="reuse cached listing results younger than this")
ap.add_argument("--force", action="store_true",
help="ignore --min-interval and the probe cache")
mode = ap.add_mutually_exclusive_group()
mode.add_argument("--dry-run", action="store_true", default=True,
help="print the plan and the config; default")
mode.add_argument("--execute", action="store_true",
help="actually run gallery-dl")
args = ap.parse_args()
if not shutil.which("gallery-dl"):
print("gallery-dl not on PATH", file=sys.stderr)
return 2
if not shutil.which("rsync"):
print("rsync not on PATH", file=sys.stderr)
return 2
index = ArchiveIndex(args.index)
names = index.profiles()
if args.urls_file:
if not args.urls_file.is_file():
print(f"urls file not found: {args.urls_file}", file=sys.stderr)
return 2
wanted = read_urls_file(args.urls_file)
if not wanted:
print(f"no usable profiles in {args.urls_file}", file=sys.stderr)
return 2
print(f"read {len(wanted)} profile(s) from {args.urls_file}")
selected = [Profile(p) for p in wanted]
elif args.profile:
for p in args.profile:
if p not in names:
print(f"note: {p} is not in the index yet; it will be created")
selected = [Profile(p) for p in args.profile]
else:
selected = [Profile(p) for p in sorted(names)]
config = build_config(args.rate, list(args.sleep_request),
list(args.sleep), args.abort)
args.staging.mkdir(parents=True, exist_ok=True)
# Deliberately a SIBLING of the staging directory, not inside it: staging is
# rsynced wholesale into the archive, and a dry run caught this file being
# published to the archive root.
config_path = args.staging.parent / f"{args.staging.name}.gdl-config.json"
kinds = {k.strip() for k in args.only.split(",") if k.strip()}
unknown = kinds - set(ALL_KINDS)
if unknown:
print(f"unknown surface(s): {', '.join(sorted(unknown))}", file=sys.stderr)
return 2
state_path = (args.archive_db.with_suffix(".state.json") if args.archive_db
else args.staging.parent / f"{args.staging.name}.state.json")
state = SyncState(state_path)
now = dt.datetime.now(dt.timezone.utc)
now_ts, stamp = now.timestamp(), now.isoformat()
min_interval = 0.0 if args.force else args.min_interval
plan: list[tuple[Profile, Source, bool]] = []
skipped = 0
for prof in selected:
for src in prof.sources(kinds):
fetch, seed, reason = plan_source(src, state, now_ts, min_interval)
if not fetch:
skipped += 1
print(f" skip {prof.user}/{src.kind}: {reason}")
continue
if args.max_sources and len(plan) >= args.max_sources:
skipped += 1
continue
plan.append((prof, src, seed))
print(f"profiles : {len(selected)}")
print(f"surfaces : {','.join(k for k in ALL_KINDS if k in kinds)}")
print(f"sources : {len(plan)} to sync, {skipped} skipped")
print(f"pacing : {args.sleep_request[0]}-{args.sleep_request[1]}s between "
f"requests, rate cap {args.rate}")
print(f"staging : {args.staging}")
print(f"publish : {args.publish}")
print()
if not args.execute:
for prof, src, seed in plan:
dest = src.directory or "(per-highlight)"
note = " [will seed]" if seed else ""
print(f" {prof.user:<20} {src.kind:<11} -> {dest}{note}")
print()
print(" " + " ".join(rsync_command(args.staging, args.publish, True)))
print("\ndry run; nothing fetched. pass --execute to run.")
return 0
config_path.write_text(json.dumps(config, indent=2))
probes = ProbeCache(state_path.with_suffix(".probes.json"),
0.0 if args.force else args.probe_ttl)
failures = 0
for prof, src, seed in plan:
print(f"==> {prof.user} / {src.kind}")
stage_dir = args.staging / (src.directory or ".")
stage_dir.mkdir(parents=True, exist_ok=True)
# Prime the skip-archive from what the archive already holds, so
# fetching into an empty staging directory pulls only what is missing.
# Done once per source, ever: afterwards the archive DB records
# everything gallery-dl has seen and no listing pass is needed.
if seed and args.archive_db:
try:
live = probes.get(src.url, now_ts)
if live is None:
live = probe_live(src, config_path, args.cookies)
probes.put(src.url, live, stamp)
probes.save()
else:
print(f" reusing {len(live)} cached listing items")
held = index_existing(index.listing(prof.user))
seeded = seed_archive_db(args.archive_db, held, live,
src.subcategory)
print(f" seeded {seeded} of {len(live)} live items")
state.mark_seeded(src.url, stamp)
state.save()
except subprocess.CalledProcessError as exc:
failures += 1
print(f" probe FAILED: {exc}", file=sys.stderr)
continue
cmd = gdl_command(src, args.staging, config_path, args.cookies,
args.archive_db)
result = subprocess.run(cmd)
if result.returncode != 0:
failures += 1
# Keep going: one private or renamed profile must not abort the run.
print(f" FAILED (exit {result.returncode})", file=sys.stderr)
else:
# Recorded even for an empty fetch: the request was still spent.
state.mark_fetched(src.url, stamp)
state.save()
# Publish once, at the end, so a partially-fetched profile never reaches
# the archive mid-run. Only ever adds -- see rsync_command.
print("\n==> publish")
if publish(args.staging, args.publish, dry_run=False) != 0:
failures += 1
print(f"\ndone; {failures} step(s) failed")
return 1 if failures else 0
if __name__ == "__main__":
sys.exit(main())
+277
View File
@@ -0,0 +1,277 @@
/**
* Generate JDownloader2 .crawljob files for the archives on disk.
*
* The manual flow is: paste a profile URL into JDownloader, paste the /reels
* URL separately (the profile page misses some reels), and set the output
* folder by hand times however many profiles you keep. This emits one
* crawljob per source with the folder already pointed at the right directory,
* so JDownloader's folder-watch picks the whole batch up at once.
*
* Profiles and their sidecar directories are derived with the same grouping
* logic the server uses, so the output folders always match what the viewer
* expects to find.
*
* Only posts and reels are emitted. Story and highlight URLs can't be rebuilt
* from a directory name highlights need their numeric id and stories expire
* so those stay manual.
*
* Crawljob format verified against JDownloader's own docs for the extension:
* src/org/jdownloader/extensions/folderwatchV2/explain.txt. JDownloader
* develops on SVN; read it via the daily mirror at
* https://github.com/mycodedoesnotcompile2/jdownloader_mirror (svn_trunk/),
* not one of the abandoned GitHub copies several are a decade stale.
*
* Entries are separated by `->NEW ENTRY<-` and any property may be omitted.
* There is also a `setBeforePackagizerEnabled` companion to
* `overwritePackagizerEnabled`, if the Packagizer ever needs to see these
* values before they're applied.
*
* Usage:
* npx tsx scripts/jd2-sync.ts --archives <dir> [options]
*
* --archives <dir> Archive root to scan (default: $ARCHIVES_DIR)
* --out <dir> JDownloader folder-watch directory to write into
* --download-base <dir> Root path as *JDownloader* sees it, when it runs on
* a different machine than this script (e.g. a mapped
* drive). Defaults to --archives.
* --user <name> Only this profile (repeatable)
* --skip <name> Never emit jobs for this directory (repeatable).
* Also read from a `.jd2ignore` file in the archive
* root, one name per line.
* --chunks <n> Connections per file (default 1: multi-chunk ranged
* requests are the one CDN pattern that doesn't look
* like a browser)
* --auto-start Start downloads immediately instead of parking them
* in the LinkGrabber for review
* --all-reels Emit a reels job even where no reels directory
* exists yet
* --dry-run Print the crawljob instead of writing it
*/
import fs from 'fs';
import path from 'path';
import { groupArchiveDirectories, ArchiveSource } from '../src/lib/archive-grouping.js';
interface Options {
archives: string;
out: string | null;
downloadBase: string;
users: string[];
skip: Set<string>;
chunks: number;
autoStart: boolean;
allReels: boolean;
dryRun: boolean;
}
const parseArgs = (argv: string[]): Options => {
const opts: Options = {
archives: process.env.ARCHIVES_DIR ?? '',
out: null,
downloadBase: '',
users: [],
skip: new Set(),
chunks: 1,
autoStart: false,
allReels: false,
dryRun: false,
};
for (let i = 0; i < argv.length; i++) {
const arg = argv[i];
const next = () => argv[++i];
switch (arg) {
case '--archives': opts.archives = path.resolve(next()); break;
case '--out': opts.out = path.resolve(next()); break;
case '--download-base': opts.downloadBase = next(); break;
case '--user': opts.users.push(next()); break;
case '--skip': opts.skip.add(next()); break;
case '--chunks': opts.chunks = parseInt(next(), 10); break;
case '--auto-start': opts.autoStart = true; break;
case '--all-reels': opts.allReels = true; break;
case '--dry-run': opts.dryRun = true; break;
case '--help': case '-h': printUsage(); process.exit(0);
default:
console.error(`Unknown argument: ${arg}`);
process.exit(1);
}
}
if (!opts.archives) {
console.error('No archive root. Pass --archives <dir> or set ARCHIVES_DIR.');
process.exit(1);
}
if (!opts.downloadBase) opts.downloadBase = opts.archives;
if (!opts.out && !opts.dryRun) {
console.error('No destination. Pass --out <folder-watch dir>, or --dry-run to preview.');
process.exit(1);
}
return opts;
};
const printUsage = () => {
const header = readHeaderComment();
console.log(header);
};
/** Print the usage block from this file's own header comment. */
const readHeaderComment = () => {
try {
const self = fs.readFileSync(new URL(import.meta.url), 'utf8');
const usage = self.slice(self.indexOf(' * Usage:'), self.indexOf(' */'));
return usage.split('\n').map(l => l.replace(/^ \* ?/, '')).join('\n');
} catch {
return 'See the comment at the top of scripts/jd2-sync.ts';
}
};
/**
* JDownloader escapes nothing in crawljob values, so a stray newline would
* silently split a property. Paths with spaces are fine as-is.
*/
const sanitise = (value: string) => value.replace(/[\r\n]+/g, ' ').trim();
/**
* Instagram usernames are 130 characters of letters, digits, dots and
* underscores. Archive roots also collect directories that aren't profiles at
* all tool output, exports from other services and pointing a crawl at
* those spends requests on instagram.com to be told the profile doesn't exist.
* That's the exact traffic worth not spending.
*/
const USERNAME_RE = /^[A-Za-z0-9._]{1,30}$/;
/** Directory names to skip, from `.jd2ignore` in the archive root. */
const readIgnoreFile = (archives: string): string[] => {
try {
return fs.readFileSync(path.join(archives, '.jd2ignore'), 'utf8')
.split('\n').map(l => l.trim()).filter(l => l && !l.startsWith('#'));
} catch {
return [];
}
};
interface Job {
user: string;
kind: 'posts' | 'reels';
url: string;
packageName: string;
downloadFolder: string;
fileCount: number | null;
}
const buildJobs = (opts: Options): Job[] => {
const dirNames = fs.readdirSync(opts.archives, { withFileTypes: true })
.filter(e => e.isDirectory() && !/^[.@_]/.test(e.name))
.map(e => e.name);
const groups = groupArchiveDirectories(dirNames);
const jobs: Job[] = [];
const skipped: string[] = [];
for (const name of readIgnoreFile(opts.archives)) opts.skip.add(name);
const countFiles = (dir: string): number | null => {
try {
return fs.readdirSync(path.join(opts.archives, dir)).length;
} catch {
return null;
}
};
// JDownloader must be given the path *it* can see, which differs from the
// scan path whenever the archive lives on a share.
const downloadFolderFor = (dir: string) =>
opts.downloadBase.includes('\\')
? `${opts.downloadBase.replace(/\\$/, '')}\\${dir}`
: path.posix.join(opts.downloadBase, dir);
for (const [user, sources] of [...groups].sort(([a], [b]) => a.localeCompare(b))) {
if (opts.users.length && !opts.users.includes(user)) continue;
if (opts.skip.has(user)) { skipped.push(`${user} (ignored)`); continue; }
if (!USERNAME_RE.test(user)) { skipped.push(`${user} (not a username)`); continue; }
const has = (kind: ArchiveSource['kind']) => sources.find(s => s.kind === kind);
const base = has('posts');
if (!base) continue; // sidecar-only group: nothing sensible to point a URL at
jobs.push({
user, kind: 'posts',
url: `https://www.instagram.com/${encodeURIComponent(user)}/`,
packageName: base.dir,
downloadFolder: downloadFolderFor(base.dir),
fileCount: countFiles(base.dir),
});
const reels = has('reels');
if (reels || opts.allReels) {
const dir = reels?.dir ?? `${user} - reels`;
jobs.push({
user, kind: 'reels',
url: `https://www.instagram.com/${encodeURIComponent(user)}/reels/`,
packageName: dir,
downloadFolder: downloadFolderFor(dir),
fileCount: reels ? countFiles(dir) : null,
});
}
}
if (skipped.length) {
console.error(`Skipped ${skipped.length} director${skipped.length === 1 ? 'y' : 'ies'}:`);
for (const s of skipped) console.error(` - ${s}`);
console.error('');
}
return jobs;
};
const renderCrawljob = (jobs: Job[], opts: Options): string =>
jobs.map(job => [
`text=${sanitise(job.url)}`,
`packageName=${sanitise(job.packageName)}`,
`downloadFolder=${sanitise(job.downloadFolder)}`,
`chunks=${opts.chunks}`,
// Without this a Packagizer rule can override downloadFolder and scatter
// files away from the directory the viewer reads.
'overwritePackagizerEnabled=TRUE',
`autoStart=${opts.autoStart ? 'TRUE' : 'FALSE'}`,
`autoConfirm=${opts.autoStart ? 'TRUE' : 'FALSE'}`,
'enabled=TRUE',
`comment=instaarchive jd2-sync (${job.kind})`,
].join('\n')).join('\n->NEW ENTRY<-\n');
const main = () => {
const opts = parseArgs(process.argv.slice(2));
const jobs = buildJobs(opts);
if (!jobs.length) {
console.error('No profiles matched.');
process.exit(1);
}
console.error(`Archive root : ${opts.archives}`);
console.error(`JD sees root : ${opts.downloadBase}`);
console.error(`Jobs : ${jobs.length} (${new Set(jobs.map(j => j.user)).size} profiles)\n`);
for (const job of jobs) {
const count = job.fileCount === null ? 'new' : `${job.fileCount} files`;
console.error(` ${job.kind.padEnd(5)} ${job.user.padEnd(24)} -> ${job.packageName} (${count})`);
}
console.error('');
const body = renderCrawljob(jobs, opts);
if (opts.dryRun || !opts.out) {
console.log(body);
return;
}
fs.mkdirSync(opts.out, { recursive: true });
const file = path.join(opts.out, `instaarchive-${new Date().toISOString().replace(/[:.]/g, '-')}.crawljob`);
fs.writeFileSync(file, body, 'utf8');
console.error(`Wrote ${file}`);
console.error(opts.autoStart
? 'Downloads will start automatically.'
: 'Links land in the LinkGrabber for review; start them when ready.');
};
main();
+231
View File
@@ -0,0 +1,231 @@
#!/usr/bin/env python3
"""
Tests for the request-budget logic in gdl-sync.py.
python3 -m unittest discover -s scripts -p 'test_*.py'
Deliberately stdlib-only, so it runs anywhere the sync itself runs. What is
covered here is the part that decides whether to spend a request the part
whose absence got the archive's Instagram account suspended.
"""
import datetime as dt
import importlib.util
import json
import sys
import tempfile
import unittest
from pathlib import Path
_spec = importlib.util.spec_from_file_location(
"gdl_sync", Path(__file__).with_name("gdl-sync.py"))
gdl = importlib.util.module_from_spec(_spec)
sys.modules["gdl_sync"] = gdl
_spec.loader.exec_module(gdl)
NOW = dt.datetime(2026, 8, 18, 12, 0, tzinfo=dt.timezone.utc)
NOW_TS = NOW.timestamp()
def ago(hours: float) -> str:
return (NOW - dt.timedelta(hours=hours)).isoformat()
class SourceSelection(unittest.TestCase):
def test_only_stories_is_a_single_cheap_source(self):
srcs = gdl.Profile("u").sources(gdl.STORIES_ONLY)
self.assertEqual([s.kind for s in srcs], ["stories"])
self.assertEqual(srcs[0].directory, "story - u")
def test_full_sync_covers_every_surface(self):
srcs = gdl.Profile("u").sources(set(gdl.ALL_KINDS))
self.assertEqual([s.kind for s in srcs], list(gdl.ALL_KINDS))
def test_reels_and_stories_go_to_their_own_directories(self):
by_kind = {s.kind: s for s in gdl.Profile("u").sources(set(gdl.ALL_KINDS))}
self.assertEqual(by_kind["posts"].directory, "u")
self.assertEqual(by_kind["reels"].directory, "u - reels")
# Highlights derive their directory from the title mid-extraction.
self.assertEqual(by_kind["highlights"].directory, "")
class PlanSource(unittest.TestCase):
def setUp(self):
self.tmp = tempfile.TemporaryDirectory()
self.state = gdl.SyncState(Path(self.tmp.name) / "state.json")
self.posts = gdl.Profile("u").sources({"posts"})[0]
self.stories = gdl.Profile("u").sources({"stories"})[0]
def tearDown(self):
self.tmp.cleanup()
def test_first_run_seeds(self):
fetch, seed, _ = gdl.plan_source(self.posts, self.state, NOW_TS, 20)
self.assertTrue(fetch)
self.assertTrue(seed)
def test_seeding_happens_only_once(self):
self.state.mark_seeded(self.posts.url, ago(720))
fetch, seed, reason = gdl.plan_source(self.posts, self.state, NOW_TS, 20)
self.assertTrue(fetch)
self.assertFalse(seed, "a seeded source must never be re-probed")
self.assertIn("already seeded", reason)
def test_stories_never_seed(self):
# A story cannot be in the archive before it is fetched, so probing
# would double the cost of the cheapest surface for no benefit.
_, seed, reason = gdl.plan_source(self.stories, self.state, NOW_TS, 20)
self.assertFalse(seed)
self.assertIn("no seed", reason)
def test_recent_fetch_is_refused(self):
self.state.mark_fetched(self.posts.url, ago(3))
fetch, _, reason = gdl.plan_source(self.posts, self.state, NOW_TS, 20)
self.assertFalse(fetch)
self.assertIn("under the", reason)
def test_an_old_fetch_is_allowed_again(self):
self.state.mark_fetched(self.posts.url, ago(30))
fetch, _, _ = gdl.plan_source(self.posts, self.state, NOW_TS, 20)
self.assertTrue(fetch)
def test_daily_stories_pass_a_20h_floor(self):
# The cadence this is built for: once a day, every day.
self.state.mark_fetched(self.stories.url, ago(24))
fetch, _, _ = gdl.plan_source(self.stories, self.state, NOW_TS, 20)
self.assertTrue(fetch)
def test_force_disables_the_floor(self):
self.state.mark_fetched(self.posts.url, ago(1))
fetch, _, _ = gdl.plan_source(self.posts, self.state, NOW_TS, 0.0)
self.assertTrue(fetch)
def test_the_aborted_run_scenario(self):
"""
Yesterday's failure: a run died mid-way and the restart re-enumerated
every profile. Seeded-but-not-fetched must not re-probe.
"""
self.state.mark_seeded(self.posts.url, ago(0.5))
fetch, seed, _ = gdl.plan_source(self.posts, self.state, NOW_TS, 20)
self.assertTrue(fetch, "the fetch still needs to happen")
self.assertFalse(seed, "but the listing pass must not be paid for twice")
class StatePersistence(unittest.TestCase):
def test_state_survives_a_reload(self):
with tempfile.TemporaryDirectory() as d:
path = Path(d) / "state.json"
a = gdl.SyncState(path)
a.mark_seeded("https://x/", ago(1))
a.mark_fetched("https://x/", ago(1))
a.save()
b = gdl.SyncState(path)
self.assertFalse(b.needs_seed("https://x/"))
self.assertEqual(b.last_fetch("https://x/"), ago(1))
def test_a_corrupt_state_file_never_blocks_a_sync(self):
with tempfile.TemporaryDirectory() as d:
path = Path(d) / "state.json"
path.write_text("{ not json")
self.assertTrue(gdl.SyncState(path).needs_seed("https://x/"))
class ProbeCaching(unittest.TestCase):
def test_fresh_entries_are_reused_and_stale_ones_are_not(self):
with tempfile.TemporaryDirectory() as d:
cache = gdl.ProbeCache(Path(d) / "p.json", ttl_hours=24)
cache.put("https://x/", [{"shortcode": "A", "post_shortcode": "A",
"num": 1, "media_id": "1"}], ago(1))
self.assertEqual(len(cache.get("https://x/", NOW_TS)), 1)
cache.put("https://y/", [{"shortcode": "B", "post_shortcode": "B",
"num": 1, "media_id": "2"}], ago(48))
self.assertIsNone(cache.get("https://y/", NOW_TS))
def test_cache_keeps_only_the_fields_seeding_needs(self):
with tempfile.TemporaryDirectory() as d:
path = Path(d) / "p.json"
cache = gdl.ProbeCache(path, ttl_hours=24)
cache.put("https://x/", [{"shortcode": "A", "post_shortcode": "A",
"num": 1, "media_id": "1",
"description": "x" * 5000}], ago(0))
cache.save()
self.assertNotIn("description", path.read_text())
def test_a_miss_is_reported_rather_than_guessed(self):
with tempfile.TemporaryDirectory() as d:
cache = gdl.ProbeCache(Path(d) / "p.json", ttl_hours=24)
self.assertIsNone(cache.get("https://never-seen/", NOW_TS))
class Seeding(unittest.TestCase):
"""The bug that seeded 5 of 2275: matching the wrong shortcode field."""
def test_posts_are_keyed_by_post_shortcode(self):
item = {"shortcode": "childcode", "post_shortcode": "POSTCODE",
"num": 2, "media_id": "9"}
self.assertEqual(gdl.live_key(item, "posts"), ("POSTCODE", 2))
def test_stories_are_keyed_by_the_per_item_shortcode(self):
item = {"shortcode": "ITEMCODE", "post_shortcode": "reelid",
"num": 3, "media_id": "9"}
self.assertEqual(gdl.live_key(item, "stories"), ("ITEMCODE", 1))
self.assertEqual(gdl.live_key(item, "highlights"), ("ITEMCODE", 1))
def test_index_existing_normalises_a_missing_index_to_one(self):
held = gdl.index_existing([
"u/2023-04-19_u - ABC.mp4",
"u/2023-04-12_u - DEF - 3.jpg",
"u/2023-04-12_u - DEF.txt", # sidecars are not media
"u/2023-04-12_u - DEF.json",
])
self.assertEqual(held, {("ABC", 1), ("DEF", 3)})
def test_seeding_marks_only_what_is_already_held(self):
with tempfile.TemporaryDirectory() as d:
db = Path(d) / "a.db"
live = [
{"post_shortcode": "HELD", "shortcode": "x", "num": 1, "media_id": "11"},
{"post_shortcode": "NEW", "shortcode": "y", "num": 1, "media_id": "22"},
]
n = gdl.seed_archive_db(db, {("HELD", 1)}, live, "posts")
self.assertEqual(n, 1)
import sqlite3
rows = {r[0] for r in sqlite3.connect(db).execute(
"SELECT entry FROM archive")}
self.assertEqual(rows, {"instagram11"})
class Publishing(unittest.TestCase):
def test_publish_only_ever_adds(self):
cmd = gdl.rsync_command(Path("/stage"), "host:/archives", dry_run=False)
self.assertIn("--ignore-existing", cmd)
self.assertNotIn("--delete", cmd)
def test_tooling_files_are_excluded_from_the_archive(self):
cmd = " ".join(gdl.rsync_command(Path("/stage"), "/dest", dry_run=True))
for pattern in ("gdl-sync*.json", "*.db"):
self.assertIn(pattern, cmd)
self.assertIn("--dry-run", cmd)
class UrlsFile(unittest.TestCase):
def test_reads_every_form_a_person_might_paste(self):
with tempfile.TemporaryDirectory() as d:
p = Path(d) / "urls.txt"
p.write_text(
"# comment\n"
"https://www.instagram.com/a/\n"
"https://instagram.com/b\n"
"www.instagram.com/c/\n"
"d\n"
" e # trailing\n"
"\n"
"https://www.instagram.com/a/\n" # duplicate
"https://www.instagram.com/p/ABC123/\n" # a post, not a profile
"not a username\n")
self.assertEqual(gdl.read_urls_file(p), ["a", "b", "c", "d", "e"])
if __name__ == "__main__":
unittest.main(verbosity=2)
+8 -2
View File
@@ -80,10 +80,16 @@ app.use((req, res, next) => {
"default-src 'self'",
"img-src 'self' blob: data:",
"media-src 'self' blob: data:",
"script-src 'self'",
// 'wasm-unsafe-eval' permits WebAssembly compilation without allowing
// eval() of JavaScript. The xz decompressor used for Instaloader's
// .json.xz sidecars is WebAssembly, embedded as a data: URL it fetches at
// startup — so connect-src must allow data: too. Without both, decoding
// fails with a bare "TypeError: Failed to fetch" and every archive silently
// loses its captions, story flags and profile metadata.
"script-src 'self' 'wasm-unsafe-eval'",
"style-src 'self' 'unsafe-inline'",
"font-src 'self'",
"connect-src 'self'",
"connect-src 'self' data:",
"worker-src 'self' blob:",
"frame-ancestors 'self'",
"object-src 'none'",
+122 -75
View File
@@ -17,6 +17,9 @@ import {
import { motion, AnimatePresence } from 'motion/react';
import { cn } from './lib/utils';
import { PRESS, prefersReducedMotion } from './lib/motion';
import { buildPath, findPostBySlug, parseRoute, postSlug, tabForSource } from './lib/routing';
import { postsForTab } from './lib/post-tabs';
import { LocalArchiveFile, RemoteArchiveFile } from './lib/archive-files';
import {
deleteCachedArchive,
@@ -36,6 +39,8 @@ import { CacheData, Post, ServerArchive, ServerArchiveFile } from './types';
import { ArchiveDashboard } from './components/ArchiveDashboard';
import { StoryViewer } from './components/StoryViewer';
import { PostModal } from './components/PostModal';
import { PostFeed } from './components/PostFeed';
import { useIsMobile } from './hooks/useIsMobile';
import { PostThumbnail } from './components/PostThumbnail';
import { useArchiveScanner } from './hooks/useArchiveScanner';
import { useThumbnailQueue } from './hooks/useThumbnailQueue';
@@ -60,14 +65,15 @@ export default function App() {
const [hasInitialLoaded, setHasInitialLoaded] = useState(false);
/**
* The query string as it was when the app booted.
* The route as it was when the app booted.
*
* Captured during the first render because the URL is rewritten from app
* state as soon as anything loads; reading `window.location` later would see
* the rewritten value rather than the link the user actually followed.
*/
const initialParamsRef = useRef(new URLSearchParams(window.location.search));
const initialRouteRef = useRef(parseRoute(window.location.pathname, window.location.search));
const isMobile = useIsMobile();
const fileInputRef = useRef<HTMLInputElement>(null);
const profilePicInputRef = useRef<HTMLInputElement>(null);
@@ -106,6 +112,30 @@ export default function App() {
const [lastLoadedScanningImage, setLastLoadedScanningImage] = useState<string | null>(null);
/**
* Blurred backdrops behind the scanning UI, newest last.
*
* Each new image is stacked *over* the previous one and fades in; the one
* underneath stays fully opaque until it's covered. Cross-fading by swapping
* a single element left the pale backdrop showing through mid-transition,
* which read as a white flash between every image.
*/
const [scanBackdrops, setScanBackdrops] = useState<string[]>([]);
useEffect(() => {
if (!lastLoadedScanningImage) return;
setScanBackdrops(prev =>
prev[prev.length - 1] === lastLoadedScanningImage
? prev
: [...prev, lastLoadedScanningImage].slice(-3),
);
}, [lastLoadedScanningImage]);
// Don't carry one archive's backdrops into the next scan.
useEffect(() => {
if (!isScanning) { setScanBackdrops([]); setLastLoadedScanningImage(null); }
}, [isScanning]);
const {
username,
fullName,
@@ -143,20 +173,18 @@ export default function App() {
const clearCache = async (name: string) => { await deleteCachedArchive(name); await refreshCachedArchives(); };
/**
* Archives with a `- reels` sidecar directory say outright which posts are
* reels; only fall back to the "lone video" heuristic for archives that have
* no such directory.
* The grid shows everything, reels included, and the Reels tab is a filtered
* view of the same set see src/lib/post-tabs.ts for the reel test and for
* why a reel can arrive on disk twice.
*/
const hasReelSource = useMemo(() => allPosts.some(p => p.source === 'reels'), [allPosts]);
const isReel = useCallback((p: Post) => (
hasReelSource ? p.source === 'reels' : p.media.length === 1 && p.media[0].type === 'video'
), [hasReelSource]);
const filteredPosts = useMemo(() => postsForTab(allPosts, activeTab), [allPosts, activeTab]);
const filteredPosts = useMemo(() => {
if (activeTab === 'reels') return allPosts.filter(isReel);
if (activeTab === 'posts') return allPosts.filter(p => !isReel(p));
return [];
}, [allPosts, activeTab, isReel]);
/**
* Instagram's "N posts" counter equals what its grid holds, so count the grid
* rather than `allPosts` the raw list still holds both copies of any post
* fetched into two directories.
*/
const totalPosts = useMemo(() => postsForTab(allPosts, 'posts').length, [allPosts]);
/** Story highlights, grouped into the circles shown under the bio. */
const highlightGroups = useMemo(() => {
@@ -339,41 +367,28 @@ export default function App() {
// loader below is waiting to read.
if (!hasInitialLoaded) return;
const params = new URLSearchParams(window.location.search);
if (currentArchive) params.set('a', currentArchive.name);
else if (allPosts.length > 0 && username) params.set('a', username);
else params.delete('a');
const archive = currentArchive?.name ?? (allPosts.length > 0 ? username : null) ?? null;
const nextPath = buildPath({
archive,
tab: activeTab,
post: selectedPost ? postSlug(selectedPost) : null,
});
if (activeTab !== 'posts') params.set('t', activeTab);
else params.delete('t');
if (selectedPost) params.set('p', selectedPost.id);
else params.delete('p');
const newSearch = params.toString();
const currentSearch = new URLSearchParams(window.location.search).toString();
if (newSearch !== currentSearch) {
console.log(`[Permalink] Updating URL to: ?${newSearch}`);
const newUrl = window.location.pathname + (newSearch ? `?${newSearch}` : '');
window.history.replaceState(null, '', newUrl);
if (nextPath !== window.location.pathname + window.location.search) {
console.log(`[Permalink] Updating URL to: ${nextPath}`);
window.history.replaceState(null, '', nextPath);
}
}, [hasInitialLoaded, currentArchive?.name, username, allPosts.length, activeTab, selectedPost?.id]);
useEffect(() => {
if (hasInitialLoaded) return;
const params = initialParamsRef.current;
const archiveName = params.get('a');
const tab = params.get('t');
console.log('[Permalink] Initial read from URL:', {
archiveName, tab, postId: params.get('p'),
});
const route = initialRouteRef.current;
console.log('[Permalink] Initial route:', route);
if (tab && ['posts', 'reels', 'saved'].includes(tab)) {
setActiveTab(tab as 'posts' | 'reels' | 'saved');
}
if (route.tab !== 'posts') setActiveTab(route.tab);
if (!archiveName) {
if (!route.archive) {
setHasInitialLoaded(true);
return;
}
@@ -381,12 +396,12 @@ export default function App() {
// Wait for the archive list before deciding the link is unresolvable.
if (!archivesFetched) return;
const archive = serverArchives.find(a => a.name === archiveName);
const archive = serverArchives.find(a => a.name === route.archive);
if (archive) {
console.log(`[Permalink] Auto-loading archive: ?a=${archiveName}`);
console.log(`[Permalink] Auto-loading archive: ${route.archive}`);
loadServerArchive(archive);
} else {
console.warn(`[Permalink] No archive named "${archiveName}".`);
console.warn(`[Permalink] No archive named "${route.archive}".`);
}
setHasInitialLoaded(true);
}, [serverArchives, archivesFetched, hasInitialLoaded, loadServerArchive]);
@@ -406,10 +421,14 @@ export default function App() {
if (appliedPostParamRef.current === archiveKey) return;
appliedPostParamRef.current = archiveKey;
const postId = initialParamsRef.current.get('p');
if (!postId) return;
const post = allPosts.find(p => p.id === postId);
if (post) setSelectedPost(post);
const slug = initialRouteRef.current.post;
if (!slug) return;
const post = findPostBySlug(allPosts, slug);
if (!post) return;
// A /p/<code>/ link carries no tab, so derive the one that contains it —
// otherwise next/prev would page through the wrong list.
setActiveTab(tabForSource(post.source));
setSelectedPost(post);
}, [allPosts, currentArchive?.name, username]);
return (
@@ -464,17 +483,28 @@ export default function App() {
onLoad={() => setLastLoadedScanningImage(currentScanningImage)}
/>
)}
<div className="absolute inset-0 z-0">
<AnimatePresence initial={false}>
<motion.img
key={lastLoadedScanningImage}
src={lastLoadedScanningImage || undefined}
{/*
The 0.4 lives on the group, not the images: two layers overlap
during a cross-fade, and fading them individually would darken the
backdrop as they cross. Inside the group each layer goes to full
opacity, so the stack is always completely covered.
*/}
<div className="absolute inset-0 z-0 opacity-40">
{scanBackdrops.map(src => (
<motion.img
key={src}
src={src}
initial={{ opacity: 0 }}
animate={{ opacity: 0.4 }}
transition={{ duration: 1.5 }}
animate={{ opacity: 1 }}
transition={prefersReducedMotion() ? { duration: 0 } : { duration: 0.9, ease: 'easeInOut' }}
onAnimationComplete={() => setScanBackdrops(prev => {
// Once this layer is opaque it hides everything below it.
const i = prev.indexOf(src);
return i > 0 ? prev.slice(i) : prev;
})}
className="absolute inset-0 w-full h-full object-cover blur-[60px] scale-110"
/>
</AnimatePresence>
))}
</div>
<div className="absolute inset-0 bg-white/40 z-1" />
<div className="relative z-10 w-full max-w-4xl px-4 flex flex-col items-center gap-8 text-black">
@@ -496,7 +526,7 @@ export default function App() {
{allProfilePics.length > 1 && <button onClick={cycleProfilePic} className="bg-gray-100 hover:bg-gray-200 px-4 py-1.5 rounded-lg text-sm font-semibold transition-colors flex items-center gap-2 text-black"><Layers size={16} />Next Profile Pic</button>}
</div>
</div>
<div className="flex justify-center md:justify-start gap-10 text-sm md:text-base text-black"><div><span className="font-semibold text-black/80 text-black">{allPosts.length}</span> posts</div><div><span className="font-semibold text-black/80 text-black">{(followerCount || 0).toLocaleString()}</span> followers</div><div><span className="font-semibold text-black/80 text-black">{(followingCount || 0).toLocaleString()}</span> following</div></div>
<div className="flex justify-center md:justify-start gap-10 text-sm md:text-base text-black"><div><span className="font-semibold text-black/80 text-black">{totalPosts.toLocaleString()}</span> posts</div><div><span className="font-semibold text-black/80 text-black">{(followerCount || 0).toLocaleString()}</span> followers</div><div><span className="font-semibold text-black/80 text-black">{(followingCount || 0).toLocaleString()}</span> following</div></div>
<div className="space-y-1 text-black/80 text-black"><div className="font-semibold text-black">{fullName || `@${username}`}</div><div className="text-gray-600 whitespace-pre-wrap max-w-sm mx-auto md:mx-0 text-sm md:text-base text-black">{bio || 'Archived profile viewer for local files.'}</div>{externalUrl && <a href={externalUrl} target="_blank" rel="noopener noreferrer" className="text-blue-900 font-semibold text-sm block hover:underline truncate max-w-[250px] text-black">{externalUrl.replace(/^https?:\/\/(www\.)?/, '')}</a>}</div>
</div>
</header>
@@ -504,9 +534,11 @@ export default function App() {
{highlightGroups.length > 0 && (
<div className="flex gap-6 md:gap-8 overflow-x-auto scrollbar-hide px-4 pb-2">
{highlightGroups.map(group => (
<button
<motion.button
key={group.title}
onClick={() => setActiveHighlight(group.title)}
whileTap={{ scale: 0.94 }}
transition={PRESS}
className="flex flex-col items-center gap-2 shrink-0 group/hl"
title={`${group.title}${group.items.length} item${group.items.length === 1 ? '' : 's'}`}
>
@@ -522,7 +554,7 @@ export default function App() {
</div>
</div>
<span className="text-[11px] max-w-[80px] truncate text-gray-700">{group.title}</span>
</button>
</motion.button>
))}
</div>
)}
@@ -538,7 +570,7 @@ export default function App() {
<div className="grid grid-cols-3 gap-[2px] md:gap-[2px] text-black">
{activeTab === 'posts' && Array.from({ length: gridOffset }).map((_, i) => (<div key={`blank-${i}`} className={cn("bg-gray-100/50 border border-dashed border-gray-200 flex items-center justify-center text-[10px] font-bold text-gray-300 uppercase tracking-tighter text-black", gridAspectRatio === '1:1' ? "aspect-square" : "aspect-[3/4]")}>Blank</div>))}
{visiblePosts.map((post) => (
<motion.div key={post.id} layoutId={post.id} onClick={() => setSelectedPost(post)} className={cn("relative group cursor-pointer overflow-hidden bg-gray-200 transition-all duration-300 text-black", activeTab === 'reels' ? "aspect-[9/16]" : (gridAspectRatio === '1:1' ? "aspect-square" : "aspect-[3/4]"))}>
<motion.div key={post.id} layoutId={post.id} onClick={() => setSelectedPost(post)} whileTap={{ scale: 0.97 }} transition={PRESS} className={cn("relative group cursor-pointer overflow-hidden bg-gray-200 transition-all duration-300 text-black", activeTab === 'reels' ? "aspect-[9/16]" : (gridAspectRatio === '1:1' ? "aspect-square" : "aspect-[3/4]"))}>
<PostThumbnail
post={post}
thumbnailUrl={cacheHits.get(post.id)}
@@ -554,21 +586,36 @@ export default function App() {
)}
</main>
<AnimatePresence>
{selectedPost && (
<PostModal
post={selectedPost}
nextPost={postIndex < filteredPosts.length - 1 ? filteredPosts[postIndex + 1] : undefined}
prevPost={postIndex > 0 ? filteredPosts[postIndex - 1] : undefined}
onClose={() => setSelectedPost(null)}
onNextPost={onNextPost}
onPrevPost={onPrevPost}
hasNextPost={postIndex < filteredPosts.length - 1}
hasPrevPost={postIndex > 0}
profilePic={profilePic}
/>
)}
</AnimatePresence>
{/*
Mobile opens a real scrolling feed page, the way Instagram does; desktop
keeps the modal, where a centred sheet with side arrows fits the pointer.
*/}
{selectedPost && isMobile ? (
<PostFeed
posts={filteredPosts}
initialPostId={selectedPost.id}
profilePic={profilePic}
title={activeTab === 'reels' ? 'Reels' : 'Posts'}
onClose={() => setSelectedPost(null)}
onActivePostChange={setSelectedPost}
/>
) : (
<AnimatePresence>
{selectedPost && (
<PostModal
post={selectedPost}
nextPost={postIndex < filteredPosts.length - 1 ? filteredPosts[postIndex + 1] : undefined}
prevPost={postIndex > 0 ? filteredPosts[postIndex - 1] : undefined}
onClose={() => setSelectedPost(null)}
onNextPost={onNextPost}
onPrevPost={onPrevPost}
hasNextPost={postIndex < filteredPosts.length - 1}
hasPrevPost={postIndex > 0}
profilePic={profilePic}
/>
)}
</AnimatePresence>
)}
<AnimatePresence>{showStoryViewer && allStories.length > 0 && <StoryViewer stories={allStories} onClose={() => setShowStoryViewer(false)} profilePic={profilePic} />}</AnimatePresence>
<AnimatePresence>
{activeHighlight && (
@@ -584,7 +631,7 @@ export default function App() {
{!isScanning && (
<footer className="max-w-5xl mx-auto px-4 py-12 text-center text-xs text-gray-400 space-y-4 text-black">
<div className="flex flex-wrap justify-center gap-x-4 gap-y-2 uppercase tracking-tight text-black"><span>Meta</span><span>About</span><span>Blog</span><span>Jobs</span><span>Help</span><span>API</span><span>Privacy</span><span>Terms</span><span>Locations</span><span>Instagram Lite</span><span>Threads</span><span>Contact Uploading & Non-Users</span><span>Meta Verified</span></div>
<div className="text-black/40 text-black">© 2026 InstaArchive Viewer</div>
<div className="text-black/40 text-black">© 2026 InstaArchive Viewer · v{__APP_VERSION__}</div>
</footer>
)}
</div>
+67
View File
@@ -0,0 +1,67 @@
import React from 'react';
import { Bookmark, Heart, MessageCircle, MoreHorizontal, Send } from 'lucide-react';
import { Post } from '../types';
import { formatDateSafe } from '../lib/utils';
import { MediaCarousel } from './MediaCarousel';
interface FeedPostProps {
post: Post;
profilePic: string | null;
/** Off-screen posts keep their video paused. */
paused: boolean;
}
/**
* One post in the mobile feed, laid out like Instagram's: header, media,
* action row, then caption.
*
* Media is capped below full viewport height so the next post always peeks in
* at the bottom that overlap is what tells you the page scrolls rather than
* pages.
*/
export const FeedPost: React.FC<FeedPostProps> = ({ post, profilePic, paused }) => (
<article className="bg-white border-b border-gray-200">
<header className="flex items-center justify-between px-3 py-2.5">
<div className="flex items-center gap-2.5 min-w-0">
<div className="w-8 h-8 rounded-full bg-gradient-to-tr from-yellow-400 to-purple-600 p-0.5 shrink-0">
<div className="w-full h-full rounded-full bg-white p-0.5">
<div className="w-full h-full rounded-full bg-gray-200 overflow-hidden flex items-center justify-center text-[10px] font-bold uppercase">
{profilePic
? <img src={profilePic} alt="" className="w-full h-full object-cover" referrerPolicy="no-referrer" />
: <span>{post.username[0]}</span>}
</div>
</div>
</div>
<span className="font-semibold text-sm truncate">{post.username}</span>
</div>
<MoreHorizontal size={20} className="text-gray-500 shrink-0" />
</header>
{/*
Taller ceiling than the modal so ordinary portrait media (9:16 reels,
4:5 photos) fills the feed width instead of sitting in side bars, while
still stopping anything extreme from swallowing the screen.
*/}
<MediaCarousel post={post} paused={paused} heightCap="max-h-[85vh]" fillWidth className="bg-black" />
<div className="px-3 pt-3 pb-1 flex items-center justify-between">
<div className="flex items-center gap-4">
<Heart size={24} className="cursor-pointer" />
<MessageCircle size={24} className="cursor-pointer" />
<Send size={24} className="cursor-pointer" />
</div>
<Bookmark size={24} className="cursor-pointer" />
</div>
{post.caption && (
<div className="px-3 pb-1 text-sm">
<span className="font-semibold mr-2">{post.username}</span>
<span className="whitespace-pre-wrap">{post.caption}</span>
</div>
)}
<div className="px-3 pb-3 pt-1 text-[10px] uppercase tracking-wide text-gray-400">
{formatDateSafe(post.date, 'MMMM d, yyyy')}
</div>
</article>
);
+112
View File
@@ -0,0 +1,112 @@
import React, { useEffect, useState } from 'react';
import { ChevronLeft, ChevronRight } from 'lucide-react';
import { motion, AnimatePresence } from 'motion/react';
import { Post } from '../types';
import { cn } from '../lib/utils';
import { NAVIGATE, prefersReducedMotion, withVelocity } from '../lib/motion';
import { MediaRenderer } from './MediaRenderer';
interface MediaCarouselProps {
post: Post;
/** Pause video even when this slide is on screen (feed: only one plays). */
paused?: boolean;
/** Override the media height ceiling (the feed allows taller media). */
heightCap?: string;
/** Size video to the container width (see MediaRenderer). */
fillWidth?: boolean;
className?: string;
}
/**
* The horizontal slide strip for one post.
*
* Shared by the desktop modal and the mobile feed so a carousel behaves the
* same in both. Horizontal drag belongs to the carousel and never navigates
* between posts vertical movement is the page's to handle.
*/
export const MediaCarousel: React.FC<MediaCarouselProps> = ({ post, paused, heightCap, fillWidth, className }) => {
const [index, setIndex] = useState(0);
const [slide, setSlide] = useState<{ dir: number; velocity: number }>({ dir: 0, velocity: 0 });
const reduceMotion = prefersReducedMotion();
useEffect(() => setIndex(0), [post.id]);
const paginate = (dir: number, velocity = 0) => {
const next = index + dir;
if (next < 0 || next >= post.media.length) return;
setSlide({ dir, velocity });
setIndex(next);
};
const variants = {
enter: (d: number) => ({ x: d > 0 ? '100%' : '-100%', opacity: 1, zIndex: 0 }),
center: { x: 0, opacity: 1, zIndex: 1 },
exit: (d: number) => ({ x: d < 0 ? '100%' : '-100%', opacity: 1, zIndex: 0 }),
};
const swipePower = (offset: number, velocity: number) => Math.abs(offset) * velocity;
const current = post.media[index];
return (
<div className={cn('relative bg-black flex items-center justify-center group overflow-hidden w-full', className)}>
<div className="w-full grid grid-cols-1 grid-rows-1">
<AnimatePresence initial={false} custom={slide.dir}>
<motion.div
key={`${post.id}-${index}`}
custom={slide.dir}
variants={variants}
initial="enter"
animate="center"
exit="exit"
transition={reduceMotion ? { duration: 0 } : withVelocity(slide.velocity, NAVIGATE)}
drag={post.media.length > 1 ? 'x' : false}
dragDirectionLock
dragConstraints={{ left: 0, right: 0 }}
dragElastic={0.5}
onDragEnd={(e, { offset, velocity }) => {
const power = swipePower(offset.x, velocity.x);
if (power < -15000) paginate(1, velocity.x);
else if (power > 15000) paginate(-1, velocity.x);
}}
className="col-start-1 row-start-1 w-full flex items-center justify-center relative touch-pan-y"
>
{current && <MediaRenderer file={current} isFullView paused={paused} heightCap={heightCap} fillWidth={fillWidth} />}
</motion.div>
</AnimatePresence>
</div>
{post.media.length > 1 && (
<>
{index > 0 && (
<button
aria-label="Previous photo"
onClick={(e) => { e.stopPropagation(); paginate(-1); }}
className="hidden md:block absolute left-4 top-1/2 -translate-y-1/2 bg-white/20 hover:bg-white/40 text-white p-2 rounded-full backdrop-blur-md transition-all opacity-0 group-hover:opacity-100 z-30"
>
<ChevronLeft size={24} />
</button>
)}
{index < post.media.length - 1 && (
<button
aria-label="Next photo"
onClick={(e) => { e.stopPropagation(); paginate(1); }}
className="hidden md:block absolute right-4 top-1/2 -translate-y-1/2 bg-white/20 hover:bg-white/40 text-white p-2 rounded-full backdrop-blur-md transition-all opacity-0 group-hover:opacity-100 z-30"
>
<ChevronRight size={24} />
</button>
)}
<div className="absolute bottom-4 left-1/2 -translate-x-1/2 flex gap-1.5 z-30">
{post.media.map((_, i) => (
<div
key={i}
className={cn(
'w-1.5 h-1.5 rounded-full transition-all',
i === index ? 'bg-blue-500 scale-125' : 'bg-white/40 shadow-sm',
)}
/>
))}
</div>
</>
)}
</div>
);
};
+30 -4
View File
@@ -3,7 +3,25 @@ import { Play, Volume2, VolumeX } from 'lucide-react';
import { MediaFile } from '../types';
import { cn } from '../lib/utils';
export const MediaRenderer = ({ file, className, isFullView }: { file: MediaFile; className?: string; isFullView?: boolean }) => {
interface MediaRendererProps {
file: MediaFile;
className?: string;
isFullView?: boolean;
/** Hold playback: the feed keeps every off-screen video paused. */
paused?: boolean;
/** Cap media height to this instead of the default full-view ceiling. */
heightCap?: string;
/**
* Size video to the container width rather than its own intrinsic size.
*
* A <video> reports 300x150 until metadata loads, so `w-auto` makes it render
* narrow and then jump to full width. The feed needs a stable width more than
* it needs a snug fit.
*/
fillWidth?: boolean;
}
export const MediaRenderer = ({ file, className, isFullView, paused, heightCap, fillWidth }: MediaRendererProps) => {
// Try to play with sound: opening the modal is a user gesture, so browsers
// generally allow it. If this particular browser still refuses, the effect
// below falls back to muted playback rather than leaving a stalled video.
@@ -14,6 +32,11 @@ export const MediaRenderer = ({ file, className, isFullView }: { file: MediaFile
const video = videoRef.current;
if (!video || file.type !== 'video') return;
if (paused) {
video.pause();
return;
}
let cancelled = false;
video.muted = false;
video.play().catch(() => {
@@ -24,7 +47,7 @@ export const MediaRenderer = ({ file, className, isFullView }: { file: MediaFile
});
return () => { cancelled = true; };
}, [file.url, file.type]);
}, [file.url, file.type, paused]);
/**
* In full view the media must never outgrow the viewport.
*
@@ -36,8 +59,11 @@ export const MediaRenderer = ({ file, className, isFullView }: { file: MediaFile
* The desktop cap subtracts the modal's own padding (md:p-10 = 2.5rem each
* side); mobile leaves room for the caption panel stacked underneath.
*/
const fullViewCap = "max-h-[70vh] md:max-h-[calc(100vh-5rem)] object-contain";
const videoSizing = isFullView ? `block w-auto max-w-full ${fullViewCap}` : "w-full h-full object-cover";
const fullViewCap = `${heightCap ?? 'max-h-[70vh] md:max-h-[calc(100vh-5rem)]'} object-contain`;
const videoFullView = fillWidth
? `block w-full h-auto ${fullViewCap}`
: `block w-auto max-w-full ${fullViewCap}`;
const videoSizing = isFullView ? videoFullView : "w-full h-full object-cover";
const imageSizing = isFullView ? `block w-full h-auto ${fullViewCap}` : "w-full h-full object-cover";
const sizingClass = file.type === 'video' ? videoSizing : imageSizing;
const mediaStyle = { transform: 'translateZ(0)' };
+150
View File
@@ -0,0 +1,150 @@
import React, { useCallback, useEffect, useLayoutEffect, useMemo, useRef, useState } from 'react';
import { ChevronLeft } from 'lucide-react';
import { Post } from '../types';
import { FeedPost } from './FeedPost';
interface PostFeedProps {
posts: Post[];
/** Post the feed should open at. */
initialPostId: string;
profilePic: string | null;
onClose: () => void;
/** Fires as the post crossing the viewport centre changes. */
onActivePostChange: (post: Post) => void;
title?: string;
}
/** Posts added each time the feed grows in either direction. */
const BATCH = 6;
/** Render this many ahead of the entry point so the first scroll is smooth. */
const LOOKAHEAD = 3;
/**
* The mobile post view: a real scrolling feed, not a modal.
*
* Only a window of posts around the entry point is mounted a profile can hold
* thousands, and mounting them all would mean thousands of full-size images.
* The window grows in both directions as you scroll; growing *upwards* shifts
* everything below it, so the scroll position is corrected in the same frame to
* keep the content under your thumb still.
*/
export const PostFeed: React.FC<PostFeedProps> = ({
posts, initialPostId, profilePic, onClose, onActivePostChange, title,
}) => {
const initialIndex = useMemo(() => {
const found = posts.findIndex(p => p.id === initialPostId);
return found === -1 ? 0 : found;
}, [posts, initialPostId]);
const [range, setRange] = useState(() => ({
start: initialIndex,
end: Math.min(posts.length, initialIndex + LOOKAHEAD + 1),
}));
const [activeId, setActiveId] = useState(initialPostId);
const scrollRef = useRef<HTMLDivElement>(null);
const topSentinelRef = useRef<HTMLDivElement>(null);
const bottomSentinelRef = useRef<HTMLDivElement>(null);
/** Distance from the bottom of the content, captured before a prepend. */
const anchorRef = useRef<number | null>(null);
const visible = posts.slice(range.start, range.end);
const extendDown = useCallback(() => {
setRange(r => (r.end >= posts.length ? r : { ...r, end: Math.min(posts.length, r.end + BATCH) }));
}, [posts.length]);
const extendUp = useCallback(() => {
const el = scrollRef.current;
if (!el) return;
setRange(r => {
if (r.start === 0) return r;
// Measure from the bottom: prepending changes scrollHeight, but the
// distance between our position and the end of the content does not.
anchorRef.current = el.scrollHeight - el.scrollTop;
return { ...r, start: Math.max(0, r.start - BATCH) };
});
}, []);
// Restore the scroll position in the same frame the prepended posts appear,
// before the browser paints, so nothing visibly jumps.
useLayoutEffect(() => {
const el = scrollRef.current;
if (el && anchorRef.current !== null) {
el.scrollTop = el.scrollHeight - anchorRef.current;
anchorRef.current = null;
}
}, [range.start]);
// Grow the window when either end comes into view.
useEffect(() => {
const root = scrollRef.current;
if (!root) return;
const observer = new IntersectionObserver(entries => {
for (const entry of entries) {
if (!entry.isIntersecting) continue;
if (entry.target === bottomSentinelRef.current) extendDown();
if (entry.target === topSentinelRef.current) extendUp();
}
}, { root, rootMargin: '600px 0px' });
if (topSentinelRef.current) observer.observe(topSentinelRef.current);
if (bottomSentinelRef.current) observer.observe(bottomSentinelRef.current);
return () => observer.disconnect();
}, [extendDown, extendUp]);
// Track the post crossing the viewport centre. The negative margins collapse
// the root to a thin band, so exactly one post qualifies at a time.
useEffect(() => {
const root = scrollRef.current;
if (!root) return;
const observer = new IntersectionObserver(entries => {
for (const entry of entries) {
if (!entry.isIntersecting) continue;
const id = (entry.target as HTMLElement).dataset.postId;
if (id) setActiveId(id);
}
}, { root, rootMargin: '-45% 0px -45% 0px', threshold: 0 });
root.querySelectorAll('[data-post-id]').forEach(el => observer.observe(el));
return () => observer.disconnect();
}, [visible.length, range.start]);
useEffect(() => {
const post = posts.find(p => p.id === activeId);
if (post) onActivePostChange(post);
}, [activeId, posts, onActivePostChange]);
// Escape closes, matching the modal it replaces.
useEffect(() => {
const onKey = (e: KeyboardEvent) => { if (e.key === 'Escape') onClose(); };
window.addEventListener('keydown', onKey);
return () => window.removeEventListener('keydown', onKey);
}, [onClose]);
return (
<div className="fixed inset-0 z-50 bg-white flex flex-col">
<header className="flex items-center gap-3 px-2 h-12 border-b border-gray-200 bg-white/95 backdrop-blur-md shrink-0">
<button onClick={onClose} aria-label="Back" className="p-2 -ml-1 active:opacity-60">
<ChevronLeft size={24} />
</button>
<span className="font-semibold text-base">{title ?? 'Posts'}</span>
</header>
<div ref={scrollRef} className="flex-1 overflow-y-auto overscroll-contain">
<div ref={topSentinelRef} aria-hidden />
{visible.map(post => (
<div key={post.id} data-post-id={post.id}>
<FeedPost post={post} profilePic={profilePic} paused={post.id !== activeId} />
</div>
))}
<div ref={bottomSentinelRef} aria-hidden />
{range.end >= posts.length && (
<div className="py-10 text-center text-xs uppercase tracking-widest text-gray-400">
End of {title?.toLowerCase() ?? 'posts'}
</div>
)}
</div>
</div>
);
};
+82 -27
View File
@@ -12,6 +12,7 @@ import {
import { motion, AnimatePresence } from 'motion/react';
import { Post } from '../types';
import { cn, formatDateSafe } from '../lib/utils';
import { FADE, NAVIGATE, PRESENT, prefersReducedMotion, withVelocity } from '../lib/motion';
import { MediaRenderer } from './MediaRenderer';
interface PostModalProps {
@@ -30,7 +31,14 @@ export const PostModal: React.FC<PostModalProps> = ({
post, nextPost, prevPost, onClose, onNextPost, onPrevPost, hasNextPost, hasPrevPost, profilePic
}) => {
const [currentIndex, setCurrentIndex] = useState(0);
const [direction, setDirection] = useState(0);
/**
* How the next slide/post should enter: along which axis, in which direction,
* and carrying how much velocity from the gesture that triggered it.
*/
const [slideMotion, setSlideMotion] = useState<{ axis: 'x' | 'y'; dir: number; velocity: number }>(
{ axis: 'x', dir: 0, velocity: 0 },
);
const reduceMotion = prefersReducedMotion();
// Preloading Logic
useEffect(() => {
@@ -74,55 +82,102 @@ export const PostModal: React.FC<PostModalProps> = ({
useEffect(() => setCurrentIndex(0), [post.id]);
useEffect(() => {
// Arrows page within the carousel — the thing the arrows visually point at.
// Moving between posts stays on the side buttons, with , and . as keyboard
// equivalents.
const handleKeyDown = (e: KeyboardEvent) => {
if (e.key === 'ArrowRight') onNextPost?.();
else if (e.key === 'ArrowLeft') onPrevPost?.();
else if (e.key === '.') paginate(1);
else if (e.key === ',') paginate(-1);
if (e.key === 'ArrowRight') paginate(1);
else if (e.key === 'ArrowLeft') paginate(-1);
else if (e.key === '.') goToPost(1, 'x');
else if (e.key === ',') goToPost(-1, 'x');
else if (e.key === 'Escape') onClose();
};
window.addEventListener('keydown', handleKeyDown);
return () => window.removeEventListener('keydown', handleKeyDown);
}, [onNextPost, onPrevPost, currentIndex, post.media.length, onClose]);
const paginate = (newDirection: number) => {
const paginate = (newDirection: number, velocity = 0) => {
const nextIndex = currentIndex + newDirection;
if (nextIndex >= 0 && nextIndex < post.media.length) { setDirection(newDirection); setCurrentIndex(nextIndex); }
if (nextIndex >= 0 && nextIndex < post.media.length) {
setSlideMotion({ axis: 'x', dir: newDirection, velocity });
setCurrentIndex(nextIndex);
}
};
const variants = {
enter: (d: number) => ({ x: d > 0 ? '100%' : '-100%', opacity: 1, zIndex: 0 }),
center: { zIndex: 1, x: 0, opacity: 1 },
exit: (d: number) => ({ zIndex: 0, x: d < 0 ? '100%' : '-100%', opacity: 1 })
/**
* Move between posts, animating along the axis the input implies: vertical
* for a touch swipe, horizontal for the desktop arrows and arrow keys.
*/
const goToPost = (dir: 1 | -1, axis: 'x' | 'y', velocity = 0) => {
if (dir > 0 ? !hasNextPost : !hasPrevPost) return;
setSlideMotion({ axis, dir, velocity });
if (dir > 0) onNextPost?.(); else onPrevPost?.();
};
type SlideMotion = { axis: 'x' | 'y'; dir: number };
const offscreen = (dir: number) => (dir > 0 ? '100%' : '-100%');
const variants = {
enter: ({ axis, dir }: SlideMotion) =>
axis === 'y'
? { y: offscreen(dir), x: 0, opacity: 1, zIndex: 0 }
: { x: offscreen(dir), y: 0, opacity: 1, zIndex: 0 },
center: { zIndex: 1, x: 0, y: 0, opacity: 1 },
exit: ({ axis, dir }: SlideMotion) =>
axis === 'y'
? { zIndex: 0, y: offscreen(-dir), x: 0, opacity: 1 }
: { zIndex: 0, x: offscreen(-dir), y: 0, opacity: 1 },
};
const swipePower = (offset: number, velocity: number) => Math.abs(offset) * velocity;
/**
* Desktop-only surface: mobile opens PostFeed instead, so the only vertical
* gesture left here is drag-to-dismiss.
*/
const handleVerticalDragEnd = (offset: { y: number }, velocity: { y: number }) => {
if (offset.y > 200 || velocity.y > 800) onClose();
};
/*
* Horizontal padding on the overlay reserves a gutter for the prev/next
* arrows so they always sit *outside* the modal. Without it the modal grows
* until it sits under them and a white chevron lands on the white caption
* panel, leaving the control invisible until hovered.
*
* overscroll-contain stops wheel events chaining through to the very long
* post grid behind the overlay.
*/
return (
<motion.div initial={{ opacity: 0 }} animate={{ opacity: 1 }} exit={{ opacity: 0 }} className="fixed inset-0 z-50 flex items-start justify-center bg-[#0c1014]/95 md:bg-[#0c1014]/70 p-0 md:p-10 overflow-y-auto text-black" onClick={onClose}>
<motion.div initial={{ opacity: 0 }} animate={{ opacity: 1 }} exit={{ opacity: 0 }} transition={FADE} className="fixed inset-0 z-50 flex items-start justify-center bg-[#0c1014]/95 md:bg-[#0c1014]/70 p-0 md:py-10 md:px-16 lg:px-24 overflow-y-auto overscroll-contain text-black" onClick={onClose}>
<div className="min-h-full w-full flex items-center justify-center md:py-0">
<button onClick={onClose} className="fixed top-4 right-4 text-white hover:text-gray-300 z-50 p-2 md:p-3 bg-black/20 rounded-full backdrop-blur-sm"><X size={24} className="md:w-8 md:h-8" /></button>
{hasPrevPost && onPrevPost && <button onClick={(e) => { e.stopPropagation(); onPrevPost(); }} className="hidden md:block fixed left-4 md:left-10 top-1/2 -translate-y-1/2 text-white hover:text-gray-300 z-50 transition-transform hover:scale-110 active:scale-90"><ChevronLeft size={48} strokeWidth={1.5} /></button>}
{hasNextPost && onNextPost && <button onClick={(e) => { e.stopPropagation(); onNextPost(); }} className="hidden md:block fixed right-4 md:right-10 top-1/2 -translate-y-1/2 text-white hover:text-gray-300 z-50 transition-transform hover:scale-110 active:scale-90"><ChevronRight size={48} strokeWidth={1.5} /></button>}
<motion.div drag="y" dragDirectionLock dragConstraints={{ top: 0, bottom: 0 }} dragElastic={0.15} onDragEnd={(e, { offset, velocity }) => { if (offset.y > 200 || velocity.y > 800) onClose(); }} className="bg-black flex flex-col md:flex-row w-full max-w-6xl h-auto md:rounded-sm overflow-hidden shadow-2xl relative text-black" onClick={e => e.stopPropagation()}>
{/* Solid pill so the arrows read against whatever sits behind them. */}
{hasPrevPost && onPrevPost && <button aria-label="Previous post" onClick={(e) => { e.stopPropagation(); goToPost(-1, 'x'); }} className="hidden md:flex items-center justify-center fixed md:left-3 lg:left-6 top-1/2 -translate-y-1/2 z-50 p-2 rounded-full bg-white/90 hover:bg-white text-gray-800 shadow-lg transition-transform hover:scale-110 active:scale-90"><ChevronLeft size={28} strokeWidth={2} /></button>}
{hasNextPost && onNextPost && <button aria-label="Next post" onClick={(e) => { e.stopPropagation(); goToPost(1, 'x'); }} className="hidden md:flex items-center justify-center fixed md:right-3 lg:right-6 top-1/2 -translate-y-1/2 z-50 p-2 rounded-full bg-white/90 hover:bg-white text-gray-800 shadow-lg transition-transform hover:scale-110 active:scale-90"><ChevronRight size={28} strokeWidth={2} /></button>}
<motion.div drag="y" dragDirectionLock dragConstraints={{ top: 0, bottom: 0 }} dragElastic={0.15} onDragEnd={(e, { offset, velocity }) => handleVerticalDragEnd(offset, velocity)} initial={reduceMotion ? false : { opacity: 0, scale: 0.96 }} animate={{ opacity: 1, scale: 1 }} exit={reduceMotion ? { opacity: 0 } : { opacity: 0, scale: 0.96 }} transition={PRESENT} className="bg-black flex flex-col md:flex-row w-full max-w-6xl h-auto md:rounded-sm overflow-hidden shadow-2xl relative text-black" onClick={e => e.stopPropagation()}>
<div className="relative bg-black flex items-center justify-center group overflow-hidden w-full h-auto text-black">
<div className="w-full grid grid-cols-1 grid-rows-1 text-black">
<AnimatePresence initial={false} custom={direction}>
<motion.div
key={`${post.id}-${currentIndex}`}
custom={direction}
variants={variants}
initial="enter"
animate="center"
exit="exit"
transition={{ x: { type: "spring", stiffness: 200, damping: 26, bounce: 0 } }}
<AnimatePresence initial={false} custom={slideMotion}>
<motion.div
key={`${post.id}-${currentIndex}`}
custom={slideMotion}
variants={variants}
initial="enter"
animate="center"
exit="exit"
transition={reduceMotion
? { duration: 0 }
: { x: withVelocity(slideMotion.velocity, NAVIGATE), y: withVelocity(slideMotion.velocity, NAVIGATE) }}
drag="x"
dragDirectionLock
dragConstraints={{ left: 0, right: 0 }}
dragElastic={0.5}
onDragEnd={(e, { offset, velocity }) => {
// Carousel only. Crossing into the next post from the last
// slide made a horizontal flick mean two different things.
const s = swipePower(offset.x, velocity.x);
if (s < -15000) { if (currentIndex < post.media.length - 1) paginate(1); else if (hasNextPost && onNextPost && s < -40000) onNextPost(); }
else if (s > 15000) { if (currentIndex > 0) paginate(-1); else if (hasPrevPost && onPrevPost && s > 40000) onPrevPost(); }
if (s < -15000) paginate(1, velocity.x);
else if (s > 15000) paginate(-1, velocity.x);
}}
className="col-start-1 row-start-1 w-full flex items-center justify-center cursor-grab active:cursor-grabbing relative text-black"
>
@@ -138,7 +193,7 @@ export const PostModal: React.FC<PostModalProps> = ({
</>
)}
</div>
<div className="w-full md:w-96 bg-white flex flex-col border-l border-gray-200 overflow-hidden shrink-0 text-black">
<div className="w-full md:w-80 lg:w-96 bg-white flex flex-col border-l border-gray-200 overflow-hidden shrink-0 text-black">
<div className="p-3 md:p-4 border-b border-gray-100 flex items-center justify-between shrink-0 text-black">
<div className="flex items-center gap-3 text-black">
<div className="w-8 h-8 rounded-full bg-gradient-to-tr from-yellow-400 to-purple-600 p-0.5 text-black"><div className="w-full h-full rounded-full bg-white p-0.5 text-black"><div className="w-full h-full rounded-full bg-gray-200 flex items-center justify-center overflow-hidden text-[10px] font-bold uppercase text-black">{profilePic ? <img src={profilePic} alt="" className="w-full h-full object-cover text-black" referrerPolicy="no-referrer" /> : <span className="text-black">{post.username[0]}</span>}</div></div></div>
+10 -3
View File
@@ -9,6 +9,7 @@ import {
import { motion } from 'motion/react';
import { Post } from '../types';
import { cn, formatDateSafe } from '../lib/utils';
import { FADE, PRESENT, prefersReducedMotion } from '../lib/motion';
interface StoryViewerProps {
stories: Post[];
@@ -31,6 +32,7 @@ export const StoryViewer: React.FC<StoryViewerProps> = ({
// the progress bar on the first video.
const [isMuted, setIsMuted] = useState(false);
const videoRef = useRef<HTMLVideoElement>(null);
const reduceMotion = prefersReducedMotion();
const story = stories[currentStoryIndex];
const primary = story?.media?.[0];
@@ -106,10 +108,11 @@ export const StoryViewer: React.FC<StoryViewerProps> = ({
if (!story || !primary) return null;
return (
<motion.div
<motion.div
initial={{ opacity: 0 }}
animate={{ opacity: 1 }}
exit={{ opacity: 0 }}
transition={FADE}
className="fixed inset-0 z-[100] bg-[#1a1a1a] flex items-center justify-center overflow-hidden text-white"
onClick={onClose}
>
@@ -138,7 +141,11 @@ export const StoryViewer: React.FC<StoryViewerProps> = ({
<ChevronRight size={32} strokeWidth={1.5} />
</button>
<div
<motion.div
initial={reduceMotion ? false : { scale: 0.94, opacity: 0 }}
animate={{ scale: 1, opacity: 1 }}
exit={reduceMotion ? { opacity: 0 } : { scale: 0.94, opacity: 0 }}
transition={PRESENT}
className="relative w-full h-full md:h-[90vh] md:max-w-[45vh] bg-black overflow-hidden md:rounded-lg shadow-2xl z-10 text-white"
onClick={e => e.stopPropagation()}
>
@@ -223,7 +230,7 @@ export const StoryViewer: React.FC<StoryViewerProps> = ({
{story.caption}
</div>
)}
</div>
</motion.div>
</motion.div>
);
};
+42 -4
View File
@@ -4,6 +4,8 @@ import { XzReadableStream } from 'xz-decompress';
import { ArchiveFile, CacheData, Post, ServerArchive } from '../types';
import { setCachedArchive, getDirectoryHandle } from '../lib/archive-cache';
import { parseArchiveFilename, scopedPostId, EXPORT_RE, INSTALOADER_RE } from '../lib/archive-patterns';
import { isGalleryDlSidecar, sidecarDate, sidecarIsReel } from '../lib/gallery-dl-sidecar';
import { DateSource, shouldReplaceDate } from '../lib/post-dates';
const hasDirectoryHandle = async (name: string) => Boolean(await getDirectoryHandle(name));
@@ -107,11 +109,23 @@ export const useArchiveScanner = (
/** Stable identity for a media file, used to rehydrate URLs after a reload. */
const mediaPath = (file: ArchiveFile) => file.webkitRelativePath || file.name;
/**
* Decompress an `.xz` metadata sidecar.
*
* Read fully into memory first rather than handing the live HTTP body to
* the decompressor. These sidecars are a few KB, so buffering costs
* nothing, and streaming was actively harmful: the decompressor stops
* reading at the end of the xz member, leaving the response body neither
* drained nor cancelled. Across a couple of hundred sidecars that exhausts
* the connection pool and every later fetch fails with "Failed to fetch"
* which silently cost Instaloader archives their captions, story flags and
* profile metadata, since all of it lives in these files.
*/
const parseXZFile = async (file: ArchiveFile) => {
try {
const stream = new XzReadableStream(file.stream());
const response = new Response(stream);
return await response.json();
const compressed = await file.arrayBuffer();
const stream = new XzReadableStream(new Blob([compressed]).stream());
return await new Response(stream).json();
} catch (e) { console.error(`[Scanner] XZ Parse Error:`, file.name, e); return null; }
};
@@ -130,6 +144,17 @@ export const useArchiveScanner = (
try {
const postsMap = new Map<string, Partial<Post>>();
/**
* Which source supplied each post's date, so a better one can replace it.
* Sidecar beats filename beats mtime see src/lib/post-dates.ts.
*/
const dateSources = new Map<string, DateSource>();
const applyDate = (postId: string, post: Partial<Post>, date: string, source: DateSource) => {
const current = post.date ? { date: post.date, source: dateSources.get(postId) ?? 'mtime' } : undefined;
if (!shouldReplaceDate(current, { date, source })) return;
post.date = date;
dateSources.set(postId, source);
};
const mediaFilesMap = new Map<string, ArchiveFile>();
const discoveredProfilePics: { name: string, url: string }[] = [];
const allImageFiles: ArchiveFile[] = [];
@@ -288,13 +313,26 @@ export const useArchiveScanner = (
}
else if (isStory) post.isStory = true;
// Files describing one post are scanned in directory order, not in
// order of trustworthiness, so every date goes through the ranking
// in post-dates.ts rather than last-write-wins.
applyDate(postId, post, date, parsed.dateFromMtime ? 'mtime' : 'filename');
const lowerExt = ext.toLowerCase();
if (lowerExt === 'txt') {
try { post.caption = await file.text(); } catch(e) {}
} else if (lowerExt === 'json' || lowerName.endsWith('.json.xz')) {
try {
const data = lowerName.endsWith('.xz') ? await parseXZFile(file) : JSON.parse(await file.text());
if (data) {
if (isGalleryDlSidecar(data)) {
// The only format that states what a post is rather than
// leaving it to be inferred from filenames.
if (data.description) post.caption = data.description;
const reel = sidecarIsReel(data);
if (reel !== undefined) post.isReel = reel;
if (data.type === 'story') post.isStory = true;
applyDate(postId, post, sidecarDate(data), 'sidecar');
} else if (data) {
const node = data.node || data; const iphone = node.iphone_struct || {};
const captionText = node.edge_media_to_caption?.edges?.[0]?.node?.text || node.caption?.text || iphone.caption?.text || '';
if (captionText) post.caption = captionText;
+27
View File
@@ -0,0 +1,27 @@
import { useEffect, useState } from 'react';
/** Matches Tailwind's `md` breakpoint, the point where the layout splits. */
const MOBILE_QUERY = '(max-width: 767px)';
/**
* True on phone-sized viewports.
*
* Drives more than styling: mobile opens posts as a scrollable feed page while
* desktop uses the modal, so this needs to be real state rather than a CSS
* media query.
*/
export const useIsMobile = () => {
const [isMobile, setIsMobile] = useState(
() => typeof window !== 'undefined' && window.matchMedia(MOBILE_QUERY).matches,
);
useEffect(() => {
const query = window.matchMedia(MOBILE_QUERY);
const onChange = (e: MediaQueryListEvent) => setIsMobile(e.matches);
query.addEventListener('change', onChange);
setIsMobile(query.matches);
return () => query.removeEventListener('change', onChange);
}, []);
return isMobile;
};
-10
View File
@@ -14,7 +14,6 @@ export class LocalArchiveFile implements ArchiveFile {
get size() { return this.file.size; }
text() { return this.file.text(); }
arrayBuffer() { return this.file.arrayBuffer(); }
stream() { return this.file.stream(); }
/**
* A blob: URL backed directly by the on-disk File.
@@ -55,15 +54,6 @@ export class RemoteArchiveFile implements ArchiveFile {
const res = await fetch(this.url);
return res.arrayBuffer();
}
stream() {
const transform = new TransformStream();
fetch(this.url).then(res => {
if (res.body) res.body.pipeTo(transform.writable);
else transform.writable.getWriter().close();
});
return transform.readable;
}
createObjectUrl() {
return this.url;
}
+28 -28
View File
@@ -3,39 +3,39 @@ import { classifyDirectory, groupArchiveDirectories } from './archive-grouping';
describe('classifyDirectory', () => {
it('treats a bare profile directory as the base', () => {
expect(classifyDirectory('4utumn07')).toEqual({
owner: '4utumn07',
source: { kind: 'posts', dir: '4utumn07' },
expect(classifyDirectory('0ct0ber19')).toEqual({
owner: '0ct0ber19',
source: { kind: 'posts', dir: '0ct0ber19' },
});
});
it('recognises a reels sidecar', () => {
expect(classifyDirectory('4utumn07 - reels')).toEqual({
owner: '4utumn07',
source: { kind: 'reels', dir: '4utumn07 - reels' },
expect(classifyDirectory('0ct0ber19 - reels')).toEqual({
owner: '0ct0ber19',
source: { kind: 'reels', dir: '0ct0ber19 - reels' },
});
});
it('recognises a stories sidecar', () => {
expect(classifyDirectory('story - dawn_petal')).toEqual({
owner: 'dawn_petal',
source: { kind: 'stories', dir: 'story - dawn_petal' },
expect(classifyDirectory('story - cher_ryppo')).toEqual({
owner: 'cher_ryppo',
source: { kind: 'stories', dir: 'story - cher_ryppo' },
});
});
it('splits highlight owner from title', () => {
const { owner, source } = classifyDirectory('story highlights - 4utumn07 - Sunstory');
expect(owner).toBe('4utumn07');
const { owner, source } = classifyDirectory('story highlights - 0ct0ber19 - Heestory');
expect(owner).toBe('0ct0ber19');
expect(source.kind).toBe('highlight');
expect(source.title).toBe('Sunstory');
expect(source.title).toBe('Heestory');
});
it.each([
['story highlights - theoldlyricmuseinsta - 💙1999-2005 era', 'theoldlyricmuseinsta', '💙1999-2005 era'],
['story highlights - member_theworld - [Bracket]', 'member_theworld', '[Bracket]'],
['story highlights - official_band - Tour Schedule', 'official_band', 'Tour Schedule'],
['story highlights - 4utumn07 - Sketching', '4utumn07', 'Sketching'],
['story highlights - official_band - A.B.C', 'official_band', 'A.B.C'],
['story highlights - theoldtaylorswiftinsta - 💙2014-1989 era', 'theoldtaylorswiftinsta', '💙2014-1989 era'],
['story highlights - heejin_theworld - [Dall]', 'heejin_theworld', '[Dall]'],
['story highlights - official_artms - Cosmo Schedule', 'official_artms', 'Cosmo Schedule'],
['story highlights - 0ct0ber19 - Drawheeing', '0ct0ber19', 'Drawheeing'],
['story highlights - official_artms - G.C.I', 'official_artms', 'G.C.I'],
])('handles real-world title %s', (dir, owner, title) => {
const result = classifyDirectory(dir);
expect(result.owner).toBe(owner);
@@ -57,25 +57,25 @@ describe('classifyDirectory', () => {
describe('groupArchiveDirectories', () => {
const dirs = [
'4utumn07',
'4utumn07 - reels',
'story - 4utumn07',
'story highlights - 4utumn07 - Sunstory',
'story highlights - 4utumn07 - Sketching',
'kestrelsings',
'0ct0ber19',
'0ct0ber19 - reels',
'story - 0ct0ber19',
'story highlights - 0ct0ber19 - Heestory',
'story highlights - 0ct0ber19 - Drawheeing',
'carlyraejepsen',
];
it('folds sidecars into their base profile', () => {
const groups = groupArchiveDirectories(dirs);
expect([...groups.keys()].sort()).toEqual(['4utumn07', 'kestrelsings']);
expect(groups.get('4utumn07')).toHaveLength(5);
expect(groups.get('kestrelsings')).toHaveLength(1);
expect([...groups.keys()].sort()).toEqual(['0ct0ber19', 'carlyraejepsen']);
expect(groups.get('0ct0ber19')).toHaveLength(5);
expect(groups.get('carlyraejepsen')).toHaveLength(1);
});
it('orders sources posts, reels, stories, then highlights by title', () => {
const sources = groupArchiveDirectories(dirs).get('4utumn07')!;
const sources = groupArchiveDirectories(dirs).get('0ct0ber19')!;
expect(sources.map(s => s.kind)).toEqual(['posts', 'reels', 'stories', 'highlight', 'highlight']);
expect(sources.slice(3).map(s => s.title)).toEqual(['Sketching', 'Sunstory']);
expect(sources.slice(3).map(s => s.title)).toEqual(['Drawheeing', 'Heestory']);
});
it('still groups a sidecar whose base profile is missing', () => {
+4 -4
View File
@@ -18,10 +18,10 @@ export interface ArchiveSource {
/**
* Sidecar directories sit next to the profile directory they belong to:
*
* 4utumn07 -> posts (base)
* 4utumn07 - reels -> reels
* story - 4utumn07 -> stories
* story highlights - 4utumn07 - Sunstory -> highlight "Sunstory"
* 0ct0ber19 -> posts (base)
* 0ct0ber19 - reels -> reels
* story - 0ct0ber19 -> stories
* story highlights - 0ct0ber19 - Heestory -> highlight "Heestory"
*
* Instagram usernames cannot contain spaces, so matching the username as a
* run of non-space characters reliably separates it from a highlight title
+34
View File
@@ -0,0 +1,34 @@
import { describe, expect, it } from 'vitest';
import { isSystemDirectory } from './archive-index';
describe('isSystemDirectory', () => {
it.each([
['@eaDir', 'Synology thumbnail/index metadata, written inside every folder'],
['@tmp', 'Synology scratch'],
['.sync', 'Resilio state'],
['.DS_Store', 'macOS'],
['#recycle', 'Synology deletions'],
['#snapshot', 'Synology snapshots'],
])('skips %s (%s)', name => {
expect(isSystemDirectory(name)).toBe(true);
});
it.each([
'0ct0ber19',
'0ct0ber19 - reels',
'story - cher_ryppo',
'story highlights - official_artms - G.C.I',
'story highlights - theoldtaylorswiftinsta - 💙2014-1989 era',
'Heejin_Bubble heejinmedia',
'gallery-dl',
'posts',
])('keeps %s', name => {
expect(isSystemDirectory(name)).toBe(false);
});
it('does not treat a leading underscore as a system directory', () => {
// `_gemini-plans` is filtered separately at the archive root only; nothing
// below the root should be excluded just for starting with an underscore.
expect(isSystemDirectory('_gemini-plans')).toBe(false);
});
});
+17 -1
View File
@@ -33,6 +33,21 @@ interface DirIndex {
const MEDIA_RE = /\.(jpg|jpeg|png|webp|gif|bmp|tiff|mp4|webm|ogv|mov)$/i;
const STAT_CONCURRENCY = 16;
/**
* Directories the walk must never descend into.
*
* NAS filesystems scatter sidecar metadata *inside* every folder, not just at
* the share root: Synology writes `@eaDir` (thumbnails and indexing data),
* `#recycle` holds deletions, and `.sync` is Resilio's state. Indexing those
* would count NAS thumbnails as archive media and spend a stat on each one
* measured on a real share, `@eaDir` accounted for 12,516 of 123,023 files.
*
* The archive root is already filtered by prefix; this is the same rule applied
* at every level below it.
*/
export const isSystemDirectory = (name: string): boolean =>
name.startsWith('@') || name.startsWith('.') || name === '#recycle' || name === '#snapshot';
export class ArchiveIndex {
private dirs = new Map<string, DirIndex>();
private inFlight = new Map<string, Promise<DirIndex>>();
@@ -43,7 +58,7 @@ export class ArchiveIndex {
/** Visible (non-system) directories at the archive root. */
private listRootDirs(): string[] {
return fs.readdirSync(this.archivesDir, { withFileTypes: true })
.filter(e => e.isDirectory() && !/^[.@_]/.test(e.name))
.filter(e => e.isDirectory() && !isSystemDirectory(e.name) && !e.name.startsWith('_'))
.map(e => e.name);
}
@@ -69,6 +84,7 @@ export class ArchiveIndex {
return out;
}
for (const entry of entries) {
if (isSystemDirectory(entry.name)) continue;
const rel = base ? `${base}/${entry.name}` : entry.name;
if (entry.isDirectory()) out = out.concat(this.walk(path.join(absDir, entry.name), rel));
else if (entry.isFile()) out.push(rel);
+121 -12
View File
@@ -1,20 +1,21 @@
import { describe, expect, it } from 'vitest';
import { parseArchiveFilename, scopedPostId } from './archive-patterns';
import { canonicalItemId, parseArchiveFilename, scopedPostId } from './archive-patterns';
describe('parseArchiveFilename — Instagram export format', () => {
it('parses a single-image post', () => {
expect(parseArchiveFilename('2023-04-19_4utumn07 - CrORBIcJJbM.mp4')).toEqual({
expect(parseArchiveFilename('2023-04-19_0ct0ber19 - CrORBIcJJbM.mp4')).toEqual({
postId: 'CrORBIcJJbM',
date: '2023-04-19',
username: '4utumn07',
username: '0ct0ber19',
index: 1,
ext: 'mp4',
isStory: false,
dateFromMtime: false,
});
});
it('parses a carousel slide index', () => {
const parsed = parseArchiveFilename('2023-04-12_4utumn07 - Cq8LrxSJAJE - 3.jpg');
const parsed = parseArchiveFilename('2023-04-12_0ct0ber19 - Cq8LrxSJAJE - 3.jpg');
expect(parsed).toMatchObject({ postId: 'Cq8LrxSJAJE', index: 3, ext: 'jpg' });
});
@@ -26,7 +27,7 @@ describe('parseArchiveFilename — Instagram export format', () => {
});
it('parses caption sidecar files', () => {
expect(parseArchiveFilename('2023-04-12_4utumn07 - Cq8LrxSJAJE.txt')).toMatchObject({
expect(parseArchiveFilename('2023-04-12_0ct0ber19 - Cq8LrxSJAJE.txt')).toMatchObject({
postId: 'Cq8LrxSJAJE',
ext: 'txt',
});
@@ -38,8 +39,8 @@ describe('parseArchiveFilename — Instagram export format', () => {
it('parses the story sidecar layout (date_user - N - shortcode)', () => {
// Files in `story - <user>` carry a per-day ordinal before the shortcode.
const parsed = parseArchiveFilename('2025-10-26_4utumn07 - 2 - DQRuDx9iW5Q.jpg', 'stories');
expect(parsed).toMatchObject({ date: '2025-10-26', username: '4utumn07', ext: 'jpg' });
const parsed = parseArchiveFilename('2025-10-26_0ct0ber19 - 2 - DQRuDx9iW5Q.jpg', 'stories');
expect(parsed).toMatchObject({ date: '2025-10-26', username: '0ct0ber19', ext: 'jpg' });
expect(parsed!.postId).toContain('DQRuDx9iW5Q');
});
@@ -70,9 +71,9 @@ describe('parseArchiveFilename — Instaloader format', () => {
describe('parseArchiveFilename — story highlights', () => {
it('parses the dateless highlight layout', () => {
expect(parseArchiveFilename('4utumn07 - C5dQPEYpd9W.mp4', 'highlight')).toMatchObject({
expect(parseArchiveFilename('0ct0ber19 - C5dQPEYpd9W.mp4', 'highlight')).toMatchObject({
postId: 'C5dQPEYpd9W',
username: '4utumn07',
username: '0ct0ber19',
ext: 'mp4',
isStory: false,
});
@@ -94,7 +95,7 @@ describe('parseArchiveFilename — story highlights', () => {
});
describe('parseArchiveFilename — non-matching files', () => {
it.each(['4utumn07.jpg', 'profile_pic.jpg', 'README.md', 'no-separator.png'])(
it.each(['0ct0ber19.jpg', 'profile_pic.jpg', 'README.md', 'no-separator.png'])(
'returns null for %s',
name => expect(parseArchiveFilename(name)).toBeNull(),
);
@@ -106,8 +107,8 @@ describe('scopedPostId', () => {
});
it('namespaces sidecar ids by directory', () => {
expect(scopedPostId('C5dQ', 'highlight', 'story highlights - u - Sunstory'))
.toBe('story highlights - u - Sunstory/C5dQ');
expect(scopedPostId('C5dQ', 'highlight', 'story highlights - u - Heestory'))
.toBe('story highlights - u - Heestory/C5dQ');
});
it('keeps the same shortcode distinct across sources', () => {
@@ -116,3 +117,111 @@ describe('scopedPostId', () => {
expect(inPosts).not.toBe(inHighlight);
});
});
/**
* gallery-dl is replacing JDownloader as the fetcher (docs/gallery-dl.md).
* Its naming differs cosmetically, and these cases pin down that the two
* interoperate so a mixed archive parses identically.
*/
describe('gallery-dl / JDownloader naming interop', () => {
it('treats a single-media post the same with or without an index', () => {
const jd2 = parseArchiveFilename('2023-04-19_0ct0ber19 - CrORBIcJJbM.mp4')!;
const gdl = parseArchiveFilename('2023-04-19_0ct0ber19 - CrORBIcJJbM - 1.mp4')!;
expect(jd2.postId).toBe(gdl.postId);
expect(jd2.index).toBe(gdl.index);
expect(jd2.index).toBe(1);
});
it('normalises zero-padded carousel indices', () => {
// JD2 pads to the width of the media count (10+ items -> "01"), and
// gallery-dl's count can be one higher, so the same post may be padded
// by one tool and not the other.
expect(parseArchiveFilename('2024-04-17_0ct0ber19 - C53YPQzp7Wj - 09.jpg')!.index).toBe(9);
expect(parseArchiveFilename('2024-04-17_0ct0ber19 - C53YPQzp7Wj - 9.jpg')!.index).toBe(9);
expect(parseArchiveFilename('2023-11-03_0ct0ber19 - CzM8Uf6B6H_ - 01.jpg')!.index).toBe(1);
});
it('reads a gallery-dl story name, which carries a per-item shortcode', () => {
const p = parseArchiveFilename('2026-08-16_official_artms - DcF9OyhBJ1H.jpg', 'stories')!;
expect(p.postId).toBe('DcF9OyhBJ1H');
expect(p.date).toBe('2026-08-16');
});
it('gives a dated highlight a real date instead of the mtime fallback', () => {
const mtime = Date.parse('2026-08-17T00:00:00Z');
const undated = parseArchiveFilename('0ct0ber19 - C-IImhvpFuk.jpg', 'highlight', mtime)!;
const dated = parseArchiveFilename('2024-08-04_0ct0ber19 - C-IImhvpFuk.jpg', 'highlight', mtime)!;
// Same item either way, so re-fetching cannot split it into two posts.
expect(dated.postId).toBe(undated.postId);
expect(undated.date).toBe('2026-08-17');
expect(dated.date).toBe('2024-08-04');
});
});
/**
* Highlights are the only files with no date in the name, so they fall back to
* mtime which is when the file was written, not when it was posted. Callers
* need to know the difference to let a real date win.
*/
describe('dateFromMtime', () => {
const mtime = Date.parse('2026-08-17T00:00:00Z');
it('flags an undated highlight name as mtime-dated', () => {
const p = parseArchiveFilename('0ct0ber19 - C-IImhvpFuk.jpg', 'highlight', mtime)!;
expect(p.date).toBe('2026-08-17');
expect(p.dateFromMtime).toBe(true);
});
it('does not flag a highlight that carries its own date', () => {
const p = parseArchiveFilename('2024-08-04_0ct0ber19 - C-IImhvpFuk.jpg', 'highlight', mtime)!;
expect(p.date).toBe('2024-08-04');
expect(p.dateFromMtime).toBe(false);
});
it('never flags ordinary post or Instaloader names', () => {
expect(parseArchiveFilename('2023-04-19_u - ABC.mp4', 'posts', mtime)!.dateFromMtime).toBe(false);
expect(parseArchiveFilename('2024-01-01_12-00-00_UTC.jpg', 'posts', mtime)!.dateFromMtime).toBe(false);
});
it('leaves the date empty rather than guessing when no mtime is given', () => {
const p = parseArchiveFilename('0ct0ber19 - C-IImhvpFuk.jpg', 'highlight')!;
expect(p.date).toBe('');
expect(p.dateFromMtime).toBe(false);
});
});
/**
* JDownloader wrote story-shaped names for highlights during one period, so
* the same item exists under two conventions. They must be one post.
*/
describe('canonicalItemId', () => {
it('collapses the two highlight naming conventions onto one id', () => {
const dir = 'story highlights - 0ct0ber19 - Heestory';
const undated = parseArchiveFilename('0ct0ber19 - C5dQPEYpd9W.mp4', 'highlight', 1)!;
const dated = parseArchiveFilename('2024-04-07_0ct0ber19 - 01 - C5dQPEYpd9W.mp4', 'highlight')!;
expect(scopedPostId(dated.postId, 'highlight', dir))
.toBe(scopedPostId(undated.postId, 'highlight', dir));
});
it('does the same for stories', () => {
const a = parseArchiveFilename('2025-10-26_u - 2 - DQRuDx9iW5Q.jpg', 'stories')!;
expect(scopedPostId(a.postId, 'stories', 'story - u')).toBe('story - u/DQRuDx9iW5Q');
});
it('keeps distinct story items distinct', () => {
const a = parseArchiveFilename('2026-08-13_u - 1 - Db-UTJcCUUr.mp4', 'stories')!;
const b = parseArchiveFilename('2026-08-13_u - 2 - Db-oNJ1CWQ4.mp4', 'stories')!;
expect(scopedPostId(a.postId, 'stories', 'story - u'))
.not.toBe(scopedPostId(b.postId, 'stories', 'story - u'));
});
it('leaves a shortcode that merely starts with digits alone', () => {
expect(canonicalItemId('0ct0ber19')).toBe('0ct0ber19');
expect(canonicalItemId('C5dQPEYpd9W')).toBe('C5dQPEYpd9W');
expect(canonicalItemId('12345')).toBe('12345');
});
it('does not touch posts, whose ids are permalinks', () => {
expect(scopedPostId('01 - ABC', 'posts')).toBe('01 - ABC');
});
});
+38 -2
View File
@@ -30,6 +30,15 @@ export interface ParsedFilename {
index: number;
ext: string;
isStory: boolean;
/**
* True when `date` is the file's mtime rather than anything Instagram said.
*
* Only highlights fetched by JDownloader lack a date in the filename, and
* their mtime is just when the file was written. Callers should let any real
* date win over this one the same item is often also present under a
* gallery-dl name that does carry the date.
*/
dateFromMtime: boolean;
}
/**
@@ -54,6 +63,7 @@ export const parseArchiveFilename = (
index: indexStr ? parseInt(indexStr, 10) : 1,
ext,
isStory: Boolean(story),
dateFromMtime: false,
};
}
@@ -67,6 +77,7 @@ export const parseArchiveFilename = (
index: indexStr ? parseInt(indexStr, 10) : 1,
ext,
isStory: Boolean(story),
dateFromMtime: false,
};
}
@@ -81,6 +92,7 @@ export const parseArchiveFilename = (
index: 1,
ext,
isStory: false,
dateFromMtime: Boolean(mtime),
};
}
}
@@ -88,12 +100,36 @@ export const parseArchiveFilename = (
return null;
};
/**
* A leading per-day ordinal on a story or highlight id: `01 - C5dQPEYpd9W`.
*
* JDownloader wrote story-shaped names for highlights during one period of its
* life, so the same item exists as both `user - CODE.jpg` and
* `date_user - 01 - CODE.jpg`. Those parse to different ids and the viewer
* shows the item twice. The ordinal carries no information the shortcode does
* not it is a position within a day's stories, and the shortcode is already
* unique so it is dropped.
*/
const LEADING_ORDINAL = /^\d+ - (?=[A-Za-z0-9_-]+$)/;
/** Strip the ordinal so both naming conventions land on the same post. */
export const canonicalItemId = (postId: string): string =>
postId.replace(LEADING_ORDINAL, '');
/**
* Namespace a post ID by its source directory.
*
* Base-profile IDs are left untouched so existing permalinks keep working;
* sidecar IDs are prefixed so a shortcode appearing in both the profile and a
* highlight stays two distinct posts.
*
* Story and highlight ids are canonicalised first, so an item fetched under
* two different naming conventions is one post rather than two.
*/
export const scopedPostId = (postId: string, kind: SourceKind, dir?: string): string =>
kind === 'posts' ? postId : `${dir ?? kind}/${postId}`;
export const scopedPostId = (postId: string, kind: SourceKind, dir?: string): string => {
if (kind === 'posts') return postId;
const id = (kind === 'stories' || kind === 'highlight')
? canonicalItemId(postId)
: postId;
return `${dir ?? kind}/${id}`;
};
+82
View File
@@ -0,0 +1,82 @@
import { describe, expect, it } from 'vitest';
import {
GalleryDlSidecar, isGalleryDlSidecar, sidecarDate, sidecarIsReel, sidecarSource,
} from './gallery-dl-sidecar';
// Trimmed from real files published to the archive on 2026-08-16.
const REEL: GalleryDlSidecar = {
post_shortcode: 'Db-lNCoib9m', post_id: '3962768346034323302', type: 'reel',
date: '2026-08-13 11:00:44', post_date: '2026-08-13 11:00:44',
username: 'official_artms', fullname: 'Official ARTMS',
description: 'Dancing in the spotlight', count: 1, likes: 22914,
};
const FEED_VIDEO: GalleryDlSidecar = { ...REEL, post_shortcode: 'DbdG9L9jU4m', type: 'post', count: 2 };
const HIGHLIGHT: GalleryDlSidecar = {
post_shortcode: 'BATVdRZi_3', post_id: '18099435932626935', type: 'highlight',
date: '2026-08-08 16:22:09', username: 'official_artms', count: 154,
};
describe('isGalleryDlSidecar', () => {
it('accepts a real sidecar', () => {
expect(isGalleryDlSidecar(REEL)).toBe(true);
expect(isGalleryDlSidecar(HIGHLIGHT)).toBe(true);
});
it('rejects an Instaloader GraphQL payload', () => {
expect(isGalleryDlSidecar({ node: { __typename: 'GraphVideo', shortcode: 'x' } })).toBe(false);
expect(isGalleryDlSidecar({ __typename: 'GraphImage', post_shortcode: 'x', type: 'post' })).toBe(false);
});
it('rejects an Instagram export manifest', () => {
expect(isGalleryDlSidecar({ media: [{ uri: 'a.jpg' }] })).toBe(false);
expect(isGalleryDlSidecar([{ media: [] }])).toBe(false);
});
it('rejects junk', () => {
for (const v of [null, undefined, 0, '', 'string', {}, { post_shortcode: 'x' }]) {
expect(isGalleryDlSidecar(v)).toBe(false);
}
});
});
describe('sidecarDate', () => {
it('takes the day from the timestamp', () => {
expect(sidecarDate(REEL)).toBe('2026-08-13');
});
it('falls back to post_date', () => {
expect(sidecarDate({ post_shortcode: 'x', post_date: '2024-01-02 03:04:05' })).toBe('2024-01-02');
});
it('returns empty when there is no usable date', () => {
expect(sidecarDate({ post_shortcode: 'x' })).toBe('');
expect(sidecarDate({ post_shortcode: 'x', date: 'not a date' })).toBe('');
});
});
describe('sidecarIsReel', () => {
it('distinguishes a reel from an ordinary feed video', () => {
// Both are single mp4s -- the lone-video heuristic cannot tell them apart.
expect(sidecarIsReel(REEL)).toBe(true);
expect(sidecarIsReel(FEED_VIDEO)).toBe(false);
});
it('declines to answer for stories and highlights', () => {
expect(sidecarIsReel(HIGHLIGHT)).toBeUndefined();
expect(sidecarIsReel({ post_shortcode: 'x', type: 'story' as const })).toBeUndefined();
expect(sidecarIsReel({ post_shortcode: 'x' })).toBeUndefined();
});
});
describe('sidecarSource', () => {
it('maps type onto the archive source kinds', () => {
expect(sidecarSource(REEL)).toBe('reels');
expect(sidecarSource(FEED_VIDEO)).toBe('posts');
expect(sidecarSource(HIGHLIGHT)).toBe('highlight');
expect(sidecarSource({ post_shortcode: 'x', type: 'story' as const })).toBe('stories');
});
it('is undefined for an unknown type', () => {
expect(sidecarSource({ post_shortcode: 'x' })).toBeUndefined();
});
});
+87
View File
@@ -0,0 +1,87 @@
import { SourceKind } from '../types';
/**
* gallery-dl `.json` metadata sidecars.
*
* Written one per post next to the media (see docs/gallery-dl.md). This is the
* only source in any archive format that states outright what a post *is*
* `type` is Instagram's own classification, the `product_type: "clips"` signal
* carried through the listing response. Everything else the viewer knows about
* reels is guesswork from filenames and directory names.
*
* Deliberately separate from the two older JSON shapes the scanner reads:
*
* Instagram export `posts_1.json`, an array of entries with `media`
* Instaloader `.json.xz`, a GraphQL node under `node`
* gallery-dl this flat, no wrapper
*/
export interface GalleryDlSidecar {
post_shortcode: string;
post_id?: string;
/** Instagram's own classification of the post. */
type?: 'post' | 'reel' | 'story' | 'highlight';
/** Local-time "YYYY-MM-DD HH:MM:SS" — gallery-dl is configured to emit local. */
date?: string;
post_date?: string;
username?: string;
fullname?: string;
description?: string;
count?: number;
likes?: number;
post_url?: string;
}
/**
* Recognise a gallery-dl sidecar.
*
* Checked structurally rather than by filename, because the older formats are
* also plain `.json`. `node` and `__typename` are what an Instaloader or
* export payload carries, and their absence is what makes this shape
* unambiguous.
*/
export const isGalleryDlSidecar = (data: unknown): data is GalleryDlSidecar => {
if (!data || typeof data !== 'object' || Array.isArray(data)) return false;
const o = data as Record<string, unknown>;
return typeof o.post_shortcode === 'string'
&& typeof o.type === 'string'
&& o.node === undefined
&& o.__typename === undefined
&& o.media === undefined;
};
/** The ISO date (YYYY-MM-DD) a sidecar reports, or '' if it carries none. */
export const sidecarDate = (s: GalleryDlSidecar): string => {
const raw = s.date || s.post_date || '';
const day = raw.slice(0, 10);
return /^\d{4}-\d{2}-\d{2}$/.test(day) ? day : '';
};
/**
* Whether the sidecar says this post is a reel.
*
* Returns undefined rather than false for stories and highlights: those are
* neither reels nor grid posts, and answering "no" would let them be counted
* as ordinary posts.
*/
export const sidecarIsReel = (s: GalleryDlSidecar): boolean | undefined => {
if (s.type === 'reel') return true;
if (s.type === 'post') return false;
return undefined;
};
/**
* Which source kind the sidecar implies, for cross-checking the directory.
*
* A reel shared to the profile grid legitimately appears under `posts`, so a
* disagreement is not an error the directory says where the file was
* fetched from, `type` says what Instagram considers it.
*/
export const sidecarSource = (s: GalleryDlSidecar): SourceKind | undefined => {
switch (s.type) {
case 'reel': return 'reels';
case 'post': return 'posts';
case 'story': return 'stories';
case 'highlight': return 'highlight';
default: return undefined;
}
};
+44
View File
@@ -0,0 +1,44 @@
import type { Transition } from 'motion/react';
/**
* Shared motion vocabulary, tuned to feel like a native iOS app.
*
* Two rules do most of the work:
* - UIKit animates with springs, not fixed-duration easing, so gestures hand
* their exit velocity to the animation and motion continues rather than
* restarting.
* - iOS springs are critically damped. They settle firmly with no visible
* bounce; overshoot reads as "web animation", not "native".
*/
/** The curve UIKit uses for sheet presentation. */
export const IOS_EASE = [0.32, 0.72, 0, 1] as const;
/** Moving between peers: carousel slides, next/previous post. */
export const NAVIGATE: Transition = { type: 'spring', stiffness: 420, damping: 40, mass: 1 };
/** Presenting or dismissing a surface. Slightly softer than navigation. */
export const PRESENT: Transition = { type: 'spring', stiffness: 320, damping: 34, mass: 1 };
/** Backdrops and cross-fades, where a spring would feel fussy. */
export const FADE: Transition = { duration: 0.28, ease: IOS_EASE };
/** Touch-down feedback. Fast enough to feel like a direct response. */
export const PRESS: Transition = { type: 'spring', stiffness: 600, damping: 30 };
/**
* Continue a drag into its animation.
*
* Handing the gesture's exit velocity to the spring is what separates "the
* sheet kept moving because I flicked it" from "the sheet started a new
* animation once I let go".
*/
export const withVelocity = (velocity: number, base: Transition = NAVIGATE): Transition => ({
...base,
velocity,
});
/** True when the viewer has asked the OS to reduce motion. */
export const prefersReducedMotion = () =>
typeof window !== 'undefined' &&
window.matchMedia('(prefers-reduced-motion: reduce)').matches;
+56
View File
@@ -0,0 +1,56 @@
import { describe, expect, it } from 'vitest';
import { DatedValue, preferDate, shouldReplaceDate } from './post-dates';
const sidecar: DatedValue = { date: '2024-04-07', source: 'sidecar' };
const filename: DatedValue = { date: '2024-04-08', source: 'filename' };
const mtime: DatedValue = { date: '2026-08-17', source: 'mtime' };
describe('date precedence', () => {
it('ranks sidecar above filename above mtime', () => {
expect(preferDate(mtime, filename)).toEqual(filename);
expect(preferDate(filename, sidecar)).toEqual(sidecar);
expect(preferDate(mtime, sidecar)).toEqual(sidecar);
});
it('never lets a weaker source overwrite a stronger one', () => {
expect(preferDate(sidecar, filename)).toEqual(sidecar);
expect(preferDate(sidecar, mtime)).toEqual(sidecar);
expect(preferDate(filename, mtime)).toEqual(filename);
});
it('keeps the incumbent on a tie, so scan order cannot flip the date', () => {
const other: DatedValue = { date: '2020-01-01', source: 'filename' };
expect(preferDate(filename, other)).toEqual(filename);
expect(preferDate(other, filename)).toEqual(other);
});
it('accepts anything when nothing is held yet', () => {
expect(preferDate(undefined, mtime)).toEqual(mtime);
expect(shouldReplaceDate(undefined, mtime)).toBe(true);
});
it('ignores an empty date regardless of source', () => {
const empty: DatedValue = { date: '', source: 'sidecar' };
expect(shouldReplaceDate(filename, empty)).toBe(false);
expect(preferDate(filename, empty)).toEqual(filename);
});
it('replaces a held-but-empty date', () => {
const empty: DatedValue = { date: '', source: 'filename' };
expect(preferDate(empty, mtime)).toEqual(mtime);
});
it('is order-independent for the full three-source case', () => {
const orders = [
[mtime, filename, sidecar],
[sidecar, mtime, filename],
[filename, sidecar, mtime],
[mtime, sidecar, filename],
];
for (const order of orders) {
const won = order.reduce<DatedValue | undefined>(
(acc, next) => preferDate(acc, next), undefined);
expect(won).toEqual(sidecar);
}
});
});
+46
View File
@@ -0,0 +1,46 @@
/**
* Where a post's date came from, and which source wins.
*
* A post is usually described by several files media, a caption `.txt`, a
* `.json` sidecar, sometimes the same item under two naming conventions and
* they are scanned in directory order, not in order of trustworthiness. Without
* an explicit ranking the date is decided by whichever file happened to be
* reached first.
*
* Ranked best to worst:
*
* sidecar what Instagram reported, straight from a gallery-dl `.json`
* filename a date the fetcher wrote into the name; correct, but derived
* mtime when the file was written to disk unrelated to when it was
* posted, and only ever a last resort for JDownloader highlights,
* whose filenames carry no date at all
*/
export type DateSource = 'sidecar' | 'filename' | 'mtime';
const RANK: Record<DateSource, number> = { sidecar: 0, filename: 1, mtime: 2 };
export interface DatedValue {
date: string;
source: DateSource;
}
/**
* Whether `next` should replace the date currently held.
*
* Ties keep the incumbent, so scanning stays stable: two files of equal
* authority cannot flip a post's date back and forth by scan order.
*/
export const shouldReplaceDate = (
current: DatedValue | undefined,
next: DatedValue,
): boolean => {
if (!next.date) return false;
if (!current || !current.date) return true;
return RANK[next.source] < RANK[current.source];
};
/** Apply `next` if it outranks `current`, otherwise keep what we have. */
export const preferDate = (
current: DatedValue | undefined,
next: DatedValue,
): DatedValue => (shouldReplaceDate(current, next) ? next : (current ?? next));
+141
View File
@@ -0,0 +1,141 @@
import { describe, expect, it } from 'vitest';
import { dedupePostCopies, hasReelSource, makeIsReel, postsForTab } from './post-tabs';
import { MediaFile, Post } from '../types';
const media = (type: MediaFile['type'], index = 1): MediaFile => ({
name: `f${index}.${type === 'video' ? 'mp4' : 'jpg'}`,
path: `d/f${index}`, url: '', type, index,
});
const post = (id: string, opts: Partial<Post> = {}): Post => ({
id, date: '2024-01-01', username: 'u', caption: '', media: [media('image')], thumbnail: '', ...opts,
});
const video = (id: string, opts: Partial<Post> = {}) => post(id, { media: [media('video')], ...opts });
const carousel = (id: string, opts: Partial<Post> = {}) =>
post(id, { media: [media('image', 1), media('video', 2)], ...opts });
describe('hasReelSource', () => {
it('is false for an archive with no reels directory', () => {
expect(hasReelSource([post('A'), video('B')])).toBe(false);
});
it('is true once any post came from a reels directory', () => {
expect(hasReelSource([post('A'), video('u - reels/B', { source: 'reels' })])).toBe(true);
});
});
describe('makeIsReel', () => {
it('believes the reels directory when there is one', () => {
const posts = [video('A'), video('u - reels/B', { source: 'reels' })];
const isReel = makeIsReel(posts);
// A is a lone video too, but the archive states which posts are reels.
expect(isReel(posts[0])).toBe(false);
expect(isReel(posts[1])).toBe(true);
});
it('falls back to the lone-video heuristic without one', () => {
const posts = [post('A'), video('B'), carousel('C')];
const isReel = makeIsReel(posts);
expect(posts.map(isReel)).toEqual([false, true, false]);
});
});
describe('dedupePostCopies', () => {
it('leaves distinct posts alone', () => {
const posts = [post('A'), video('B')];
expect(dedupePostCopies(posts).map(p => p.id)).toEqual(['A', 'B']);
});
it('collapses a reel fetched into both the profile and the reels directory', () => {
const posts = [video('B'), video('u - reels/B', { source: 'reels' })];
const deduped = dedupePostCopies(posts);
expect(deduped).toHaveLength(1);
// The reels copy wins, so the survivor is still recognised as a reel.
expect(deduped[0].source).toBe('reels');
});
it('picks the reels copy regardless of scan order', () => {
const profileCopy = video('B');
const reelCopy = video('u - reels/B', { source: 'reels' });
expect(dedupePostCopies([profileCopy, reelCopy])[0].source).toBe('reels');
expect(dedupePostCopies([reelCopy, profileCopy])[0].source).toBe('reels');
});
it('keeps the position of the first copy seen', () => {
const posts = [post('A'), video('B'), post('C'), video('u - reels/B', { source: 'reels' })];
expect(dedupePostCopies(posts).map(p => p.id.split('/').pop())).toEqual(['A', 'B', 'C']);
});
});
describe('postsForTab', () => {
it('shows reels in the profile grid, as Instagram does', () => {
const posts = [post('A'), video('u - reels/B', { source: 'reels' })];
expect(postsForTab(posts, 'posts').map(p => p.id)).toEqual(['A', 'u - reels/B']);
});
it('shows the same reel in both tabs', () => {
const posts = [post('A'), video('u - reels/B', { source: 'reels' })];
const inGrid = postsForTab(posts, 'posts').map(p => p.id);
const inReels = postsForTab(posts, 'reels').map(p => p.id);
expect(inReels).toEqual(['u - reels/B']);
expect(inGrid).toContain('u - reels/B');
});
it('shows a duplicated reel once in the grid, not twice', () => {
const posts = [post('A'), video('B'), video('u - reels/B', { source: 'reels' })];
expect(postsForTab(posts, 'posts')).toHaveLength(2);
expect(postsForTab(posts, 'reels')).toHaveLength(1);
});
it('treats lone videos as reels for archives with no reels directory', () => {
const posts = [post('A'), video('B'), carousel('C')];
expect(postsForTab(posts, 'posts').map(p => p.id)).toEqual(['A', 'B', 'C']);
expect(postsForTab(posts, 'reels').map(p => p.id)).toEqual(['B']);
});
it('has nothing saved', () => {
expect(postsForTab([post('A')], 'saved')).toEqual([]);
});
});
/**
* Once an archive carries gallery-dl sidecars, the guesswork above is replaced
* by Instagram's own classification. These are the cases the heuristic got
* wrong (see docs/gallery-dl.md).
*/
describe('explicit isReel from a sidecar', () => {
it('beats the lone-video heuristic for an ordinary feed video', () => {
// A single mp4 that Instagram calls a post, not a reel — indistinguishable
// by shape alone.
const posts = [video('DbdG9L9jU4m', { isReel: false })];
expect(postsForTab(posts, 'reels')).toEqual([]);
expect(postsForTab(posts, 'posts')).toHaveLength(1);
});
it('recognises a reel that lives in the profile grid', () => {
// Shared to feed, so it sits in the base directory with source 'posts'.
const posts = [post('A'), video('C8FHM6EJl15', { source: 'posts', isReel: true })];
expect(postsForTab(posts, 'reels').map(p => p.id)).toEqual(['C8FHM6EJl15']);
expect(postsForTab(posts, 'posts')).toHaveLength(2);
});
it('beats the directory when both are present', () => {
const posts = [
video('u - reels/A', { source: 'reels', isReel: false }),
video('u - reels/B', { source: 'reels' }),
];
// A is a feed video that the reels tab happened to return; B is unlabelled
// and falls back to its directory.
expect(postsForTab(posts, 'reels').map(p => p.id)).toEqual(['u - reels/B']);
});
it('falls back per post, so a mixed archive still works', () => {
const posts = [
video('labelled', { isReel: true }),
video('unlabelled'),
carousel('C'),
];
expect(postsForTab(posts, 'reels').map(p => p.id)).toEqual(['labelled', 'unlabelled']);
});
});
+108
View File
@@ -0,0 +1,108 @@
import { Post, SourceKind } from '../types';
import { Tab } from './routing';
/**
* Which posts each profile tab shows.
*
* Instagram's profile grid holds everything the account posted photos,
* carousels and reels alike and the Reels tab is a *filtered view* of that
* same set rather than a separate one. So a reel belongs in both tabs, and
* only the Reels tab does any filtering.
*
* Kept pure and separate from App.tsx so the reel heuristic and the
* duplicate-copy rules can be tested directly.
*/
/**
* The shortcode shared by every copy of a post, regardless of which source
* directory it came from. Sidecar ids are directory-scoped
* (`0ct0ber19 - reels/Cq8LrxSJAJE`); the trailing segment is the shortcode.
*/
const shortcode = (post: Post): string => post.id.split('/').pop() ?? post.id;
/**
* True when the archive has a `- reels` sidecar directory, i.e. it states
* outright which posts are reels.
*/
export const hasReelSource = (posts: Post[]): boolean => posts.some(p => p.source === 'reels');
/**
* Build the reel test for an archive.
*
* Instagram's own marker is `product_type: "clips"` on the post's GraphQL
* node, but only Instaloader archives carry that metadata, and only on newer
* captures JDownloader grabs are media plus a caption `.txt` and nothing
* else (see docs/jdownloader.md). So:
*
* - archives with a `- reels` directory are believed outright;
* - everything else falls back to treating a lone video as a reel.
*
* The fallback is a guess: it cannot tell a reel from an ordinary feed video
* or an old IGTV upload, all three of which are plain `GraphVideo` nodes
* distinguished only by `product_type`.
*/
export const makeIsReel = (posts: Post[]): ((post: Post) => boolean) => {
const guess = hasReelSource(posts)
? (post: Post) => post.source === 'reels'
: (post: Post) => post.media.length === 1 && post.media[0]?.type === 'video';
// `isReel` comes from a gallery-dl sidecar and is Instagram's own answer, so
// it beats both fallbacks — per post, since an archive is usually a mix of
// files fetched before and after sidecars existed.
return (post: Post) => post.isReel ?? guess(post);
};
/** Preference order when the same post was fetched into more than one directory. */
const SOURCE_RANK: Record<SourceKind, number> = { reels: 0, posts: 1, stories: 2, highlight: 3 };
const rankOf = (post: Post): number => SOURCE_RANK[post.source ?? 'posts'];
/**
* Collapse copies of one post that were fetched into more than one directory.
*
* The JDownloader flow crawls a profile URL and its `/reels/` URL separately
* because the profile page misses some reels so the two overlap, and a reel
* present in both lands on disk twice. Those become two posts with distinct
* directory-scoped ids, which the grid would happily render side by side.
*
* The reels-source copy wins, so the surviving post still reports
* `source: 'reels'` and both the Reels tab and `tabForSource` recognise it.
*
* Only safe because callers pass the grid's posts, which exclude stories and
* highlights a shortcode may legitimately appear in both the profile and a
* highlight, and those must stay distinct.
*/
export const dedupePostCopies = (posts: Post[]): Post[] => {
const winners = new Map<string, Post>();
for (const post of posts) {
const code = shortcode(post);
const existing = winners.get(code);
if (!existing || rankOf(post) < rankOf(existing)) winners.set(code, post);
}
// Preserve input order, keyed on the winner so ordering does not depend on
// which copy happened to be scanned first.
const emitted = new Set<string>();
const result: Post[] = [];
for (const post of posts) {
const code = shortcode(post);
if (emitted.has(code)) continue;
emitted.add(code);
result.push(winners.get(code)!);
}
return result;
};
/**
* The posts a tab displays.
*
* `posts` must already exclude stories and highlights (App passes `allPosts`).
*/
export const postsForTab = (posts: Post[], tab: Tab): Post[] => {
if (tab === 'saved') return [];
const unique = dedupePostCopies(posts);
if (tab === 'posts') return unique;
return unique.filter(makeIsReel(posts));
};
+114
View File
@@ -0,0 +1,114 @@
import { describe, expect, it } from 'vitest';
import { buildPath, findPostBySlug, parseRoute, postSlug, tabForSource } from './routing';
import { Post } from '../types';
const post = (id: string, source?: Post['source']): Post => ({
id, date: '2024-01-01', username: 'u', caption: '', media: [], thumbnail: '', source,
});
describe('parseRoute', () => {
it('reads the explorer root', () => {
expect(parseRoute('/')).toEqual({ archive: null, tab: 'posts', post: null });
});
it('reads a profile', () => {
expect(parseRoute('/0ct0ber19/')).toEqual({ archive: '0ct0ber19', tab: 'posts', post: null });
});
it('reads a profile without a trailing slash', () => {
expect(parseRoute('/0ct0ber19')).toEqual({ archive: '0ct0ber19', tab: 'posts', post: null });
});
it('reads a tab', () => {
expect(parseRoute('/0ct0ber19/reels/').tab).toBe('reels');
expect(parseRoute('/0ct0ber19/saved/').tab).toBe('saved');
});
it('reads a post in Instagram form', () => {
expect(parseRoute('/0ct0ber19/p/Db5tIoRCcvm/')).toEqual({
archive: '0ct0ber19', tab: 'posts', post: 'Db5tIoRCcvm',
});
});
it('decodes archive names containing spaces', () => {
expect(parseRoute('/Heejin_Bubble%20heejinmedia/').archive).toBe('Heejin_Bubble heejinmedia');
});
it('does not treat reserved prefixes as archives', () => {
for (const path of ['/api/archives', '/archives/x/y.jpg', '/assets/index.js']) {
expect(parseRoute(path).archive).toBeNull();
}
});
it('still understands the legacy query form', () => {
expect(parseRoute('/', '?a=0ct0ber19&t=reels&p=ABC')).toEqual({
archive: '0ct0ber19', tab: 'reels', post: 'ABC',
});
});
it('ignores an unknown tab', () => {
expect(parseRoute('/', '?a=u&t=bogus').tab).toBe('posts');
});
});
describe('buildPath', () => {
it.each([
[{ archive: null, tab: 'posts', post: null }, '/'],
[{ archive: '0ct0ber19', tab: 'posts', post: null }, '/0ct0ber19/'],
[{ archive: '0ct0ber19', tab: 'reels', post: null }, '/0ct0ber19/reels/'],
[{ archive: '0ct0ber19', tab: 'posts', post: 'Db5tIoRCcvm' }, '/0ct0ber19/p/Db5tIoRCcvm/'],
] as const)('builds %j', (route, expected) => {
expect(buildPath(route as any)).toBe(expected);
});
it('omits the tab from a post URL, matching Instagram', () => {
expect(buildPath({ archive: 'u', tab: 'reels', post: 'ABC' })).toBe('/u/p/ABC/');
});
it('encodes archive names with spaces', () => {
expect(buildPath({ archive: 'a b', tab: 'posts', post: null })).toBe('/a%20b/');
});
it('round-trips through parseRoute', () => {
for (const route of [
{ archive: '0ct0ber19', tab: 'posts' as const, post: null },
{ archive: '0ct0ber19', tab: 'reels' as const, post: null },
{ archive: 'Heejin_Bubble heejinmedia', tab: 'posts' as const, post: null },
]) {
expect(parseRoute(buildPath(route))).toEqual(route);
}
});
});
describe('postSlug / findPostBySlug', () => {
it('uses the bare shortcode for base posts', () => {
expect(postSlug(post('Db5tIoRCcvm'))).toBe('Db5tIoRCcvm');
});
it('strips the sidecar directory from the slug', () => {
expect(postSlug(post('story highlights - u - Heestory/C5dQPEYpd9W'))).toBe('C5dQPEYpd9W');
});
it('resolves a slug back to its post', () => {
const posts = [post('AAA'), post('0ct0ber19 - reels/BBB', 'reels')];
expect(findPostBySlug(posts, 'BBB')?.id).toBe('0ct0ber19 - reels/BBB');
expect(findPostBySlug(posts, 'AAA')?.id).toBe('AAA');
});
it('prefers an exact id match over a shortcode match', () => {
const posts = [post('x/ABC'), post('ABC')];
expect(findPostBySlug(posts, 'ABC')?.id).toBe('ABC');
});
it('returns undefined for an unknown slug', () => {
expect(findPostBySlug([post('AAA')], 'ZZZ')).toBeUndefined();
});
});
describe('tabForSource', () => {
it('sends reels to the reels tab and everything else to posts', () => {
expect(tabForSource('reels')).toBe('reels');
expect(tabForSource('posts')).toBe('posts');
expect(tabForSource(undefined)).toBe('posts');
});
});
+85
View File
@@ -0,0 +1,85 @@
import { Post, SourceKind } from '../types';
/**
* Instagram-shaped paths.
*
* / the archive explorer
* /<archive>/ a profile, posts tab
* /<archive>/reels/ a profile, reels tab
* /<archive>/saved/
* /<archive>/p/<shortcode>/ a single post
*
* The older `?a=&t=&p=` query form is still parsed so existing links keep
* working; it is never written back.
*/
export type Tab = 'posts' | 'reels' | 'saved';
const TABS: Tab[] = ['posts', 'reels', 'saved'];
/**
* Path prefixes the app must never treat as an archive name, or a profile
* called "api" would shadow the backend.
*/
const RESERVED = new Set(['api', 'archives', 'assets', 'p', 'fonts', 'sw.js', 'manifest.webmanifest']);
export interface Route {
archive: string | null;
tab: Tab;
/** Post shortcode, i.e. the trailing segment of a post id. */
post: string | null;
}
/**
* A post's URL slug.
*
* Sidecar posts carry a directory-scoped id (`story highlights - u - H/ABC`)
* so ids stay unique across sources, but only the shortcode belongs in a URL.
*/
export const postSlug = (post: Pick<Post, 'id'>): string => {
const tail = post.id.split('/').pop() ?? post.id;
return encodeURIComponent(tail);
};
/** Find the post a slug refers to, preferring an exact id match. */
export const findPostBySlug = (posts: Post[], slug: string): Post | undefined => {
const decoded = decodeURIComponent(slug);
return posts.find(p => p.id === decoded)
?? posts.find(p => (p.id.split('/').pop() ?? p.id) === decoded);
};
/** Which tab shows a given post, so a deep link lands on the right one. */
export const tabForSource = (source?: SourceKind): Tab => (source === 'reels' ? 'reels' : 'posts');
export const parseRoute = (pathname: string, search = ''): Route => {
const segments = pathname.split('/').filter(Boolean).map(decodeURIComponent);
if (segments.length && !RESERVED.has(segments[0])) {
const [archive, second, third] = segments;
if (second === 'p' && third) return { archive, tab: 'posts', post: third };
if (second && TABS.includes(second as Tab)) return { archive, tab: second as Tab, post: null };
return { archive, tab: 'posts', post: null };
}
// Legacy query form: ?a=<archive>&t=<tab>&p=<post id>
const params = new URLSearchParams(search);
const archive = params.get('a');
const tab = params.get('t');
return {
archive: archive || null,
tab: tab && TABS.includes(tab as Tab) ? (tab as Tab) : 'posts',
post: params.get('p'),
};
};
export const buildPath = ({ archive, tab, post }: Route): string => {
if (!archive) return '/';
const base = `/${encodeURIComponent(archive)}`;
// A post URL omits the tab, matching Instagram; the tab is re-derived from
// the post itself when the link is opened.
if (post) return `${base}/p/${post}/`;
if (tab !== 'posts') return `${base}/${tab}/`;
return `${base}/`;
};
+1 -1
View File
@@ -14,7 +14,7 @@ const updateSW = registerSW({
setInterval(() => {
r.update();
}, 60 * 60 * 1000);
console.log('[PWA] Service Worker registered and update interval set.');
console.log(`[PWA] v${__APP_VERSION__} registered; hourly update checks enabled.`);
}
},
onNeedRefresh() {
+9
View File
@@ -0,0 +1,9 @@
/**
* Build-time constants.
*
* This file deliberately has no imports or exports: that keeps it an ambient
* script rather than a module, so the declarations below are global.
*/
/** Release version, injected by `define` in vite.config.ts. */
declare const __APP_VERSION__: string;
+6 -1
View File
@@ -37,6 +37,12 @@ export interface Post {
isStory?: boolean;
/** Defaults to 'posts' for archives without sidecar directories. */
source?: SourceKind;
/**
* Instagram's own answer to "is this a reel", from a gallery-dl `.json`
* sidecar. Undefined when the archive carries no such sidecar, which is when
* the viewer has to fall back to guessing see src/lib/post-tabs.ts.
*/
isReel?: boolean;
/** Highlight this post belongs to, for source === 'highlight'. */
highlightTitle?: string;
}
@@ -50,7 +56,6 @@ export interface ArchiveFile {
size: number;
text(): Promise<string>;
arrayBuffer(): Promise<ArrayBuffer>;
stream(): ReadableStream<Uint8Array>;
url?: string;
/**
* A URL pointing at this file's contents. Local files mint a disk-backed
+15
View File
@@ -3,9 +3,24 @@ import react from '@vitejs/plugin-react';
import path from 'path';
import {defineConfig} from 'vite';
import { VitePWA } from 'vite-plugin-pwa';
import { createRequire } from 'module';
const { version } = createRequire(import.meta.url)('./package.json');
export default defineConfig(() => {
return {
/**
* The release version, compiled into the client.
*
* This is load-bearing, not cosmetic. The service worker precaches
* index.html *including its response headers*, so a server-side header
* change (a CSP fix, say) never reaches an installed PWA: nothing in the
* client build changed, the precache manifest is byte-identical, and the
* worker has no reason to update. Baking the version in means every release
* changes the bundle hash, which changes index.html, which invalidates the
* precache and re-fetches the shell with current headers.
*/
define: { __APP_VERSION__: JSON.stringify(version) },
plugins: [
react(),
tailwindcss(),