Compare commits

..
11 Commits
Author SHA1 Message Date
ergosteurandClaude Sonnet 5 600949e475 fix: make the browser Back button close a post instead of exiting the app
The URL was only ever synced with history.replaceState, so the app
never created any history entries of its own -- Back always went
straight to whatever page was open before this one, no matter where
you were in the app.

Opening a post now pushState's a new entry, matching how Instagram's
own back button behaves: Back closes the post and returns to the grid.
Closing a post any other way (the X button, the modal's own close
handler) consumes that same entry via history.back() instead of piling
a fresh one on top, and a popstate listener re-syncs app state for
both directions. Tab switches and archive loads still use
replaceState, unchanged -- only the post view gets its own step in
history, deliberately, to keep the history stack shallow.

Verified in a real browser: open a post, Back closes it and stays in
the app; Forward reopens it; the X button closes it too, consuming the
same entry rather than leaving a stale one behind.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011qAds5qr7nZRq5R4yAuxUk
2026-08-27 14:36:36 -04:00
ergosteurandClaude Opus 5 26d2d3e379 chore: release 1.8.1
Docker Build and Publish / build-and-push (push) Failing after 11s
First build from the redacted history, and the first that does not ship
source comments in the server bundle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UXfdJu7QhSJLr47K7koTDF
2026-08-20 15:36:17 -04:00
ergosteurandClaude Opus 5 61c2b62141 fix: stop shipping source comments in the server bundle
`tsc` keeps comments by default, so the doc comment in archive-grouping.ts
describing the sidecar layout was emitted into dist-server and copied into the
runtime image. Every published container image on ghcr carries it — verified by
pulling the dist-server layer of :latest and grepping it:

    app/src/lib/archive-grouping.js:10:  *   <user>  -> posts (base)

That comment names real archived accounts, which is exactly what main was
redacted to remove, so the redaction was incomplete while the build kept
re-emitting them. The frontend was never affected: Vite strips comments, and
the 432K dist layer greps clean.

--removeComments takes dist-server from 0 comment lines. docs/ was never at
risk; the multi-stage build copies only dist/ and dist-server/ into runtime.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UXfdJu7QhSJLr47K7koTDF
2026-08-20 15:36:00 -04:00
ergosteurandClaude Opus 5 882296b1c0 chore: move the archive-fetching tooling out of this branch
The fetching scripts and their docs now live on the `tooling` branch, which
is not published to GitHub. This removes the two references that would
otherwise dangle here: the `jd2` npm script and the CLAUDE.md bullet
describing it.

The viewer's own gallery-dl support is untouched and stays here —
src/lib/gallery-dl-sidecar.ts and friends parse sidecars at display time and
are app code, not tooling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UXfdJu7QhSJLr47K7koTDF
2026-08-20 14:52:03 -04:00
ergosteurandClaude Opus 5 4f8b0021c6 chore: stop tracking compiled Python bytecode
Two .pyc files under scripts/__pycache__ were committed at some point and have
been churning ever since — merely importing gdl-sync.py to check a config
rewrites them and dirties the tree, which is how they surfaced.

.gitignore had no Python entries at all, only Node ones. The files stay on
disk; this just untracks them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UXfdJu7QhSJLr47K7koTDF
2026-08-20 14:38:41 -04:00
ergosteurandClaude Opus 5 dcd8f2ef1d chore: release 1.8.0
Docker Build and Publish / build-and-push (push) Failing after 10s
Ships the gallery-dl sidecar work to the viewer. The visible change is
reel classification: official_artms' Reels tab drops from 781 items to
360, because the sidecars say the other 421 are ordinary feed videos the
clips endpoint returns via include_feed_video. Directory-based
classification counted them all as reels.

Also in this release: dates ranked by source rather than scan order, and
highlight items no longer appearing twice when the archive holds them
under both JDownloader naming conventions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 13:58:15 -04:00
ergosteurandClaude Opus 5 5d5dea10c8 fix: stop showing a highlight item twice under two naming conventions
JDownloader wrote story-shaped names for highlights during one period of
its life, so the same item exists on disk as both

    0ct0ber19 - C5dQPEYpd9W.mp4
    2024-04-07_0ct0ber19 - 01 - C5dQPEYpd9W.mp4

which parsed to the ids "C5dQPEYpd9W" and "01 - C5dQPEYpd9W" -- two posts
for one item. The leading ordinal is a position within a day's stories
and carries nothing the shortcode does not, so story and highlight ids
drop it. Post ids are untouched, since those are permalinks.

Measured on the two real files, same archive, cache cleared between:
without the fix the profile reads "Heestory - 2 items", with it
"Heestory - 1 item".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 13:57:42 -04:00
ergosteurandClaude Opus 5 816fa970b5 fix: rank date sources instead of letting scan order decide
The previous commit had this backwards: the sidecar date was only
consulted when the existing date came from an mtime, so a filename date
silently outranked what Instagram itself reported.

The order is sidecar, then filename, then mtime -- metadata first,
mtime last, since mtime is when the file hit disk and says nothing about
when the post was made. Ties keep the incumbent so two equally
authoritative files cannot flip a post's date by scan order.

Extracted to src/lib/post-dates.ts rather than left inline, because the
rule is easy to state and easy to get wrong -- the tests include an
order-independence case that would have caught the original mistake.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 12:50:48 -04:00
ergosteurandClaude Opus 5 53b1f80e1d feat: read gallery-dl sidecars for reel type and post dates
The .json sidecars published with the ARTMS fetch were inert: the scanner
fed them through the Instaloader path, where `node.edge_media_to_caption`
and `checkIsStory`'s `product_type` are both absent, so nothing happened.

They are now recognised structurally -- flat, with post_shortcode and
type, and none of the markers the other two JSON shapes carry -- and used
for three things:

- `type` sets post.isReel, which post-tabs prefers over every fallback.
  This is Instagram's own classification and it disagrees with ours a
  lot: of 781 items in "official_artms - reels", the sidecars say only
  360 are reels. The other 421 are feed videos the clips endpoint returns
  via include_feed_video, and the directory-based rule counted them all.
- `description` fills the caption where no .txt exists.
- `date` dates a post whose filename could not.

Also fixes date precedence. Only JDownloader highlights lack a date in
the filename, so parseArchiveFilename now marks those as mtime-derived
and the scanner lets any real date replace them -- previously the date
depended on which file the scan reached first.

Verified against real published files: a directory of three type=post and
three type=reel renders 6 in the grid and exactly the 3 reels in the
Reels tab.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 12:42:54 -04:00
ergosteurandClaude Opus 5 89bd5346db docs: design a gallery-dl replacement for the JDownloader fetcher
Every claim in docs/gallery-dl.md was measured against the live site and
the archive rather than taken from documentation, because two of the
assumptions turned out to be wrong.

The safety model is the reason the config looks the way it does.
gallery-dl has two API backends: the graphql one issues a request PER
POST for every video and carousel -- the pattern that got this account
banned via Instaloader -- while the default rest one paginates listings
at 30-50 items and carries carousel_media, video_versions and
product_type inline. A 300-post profile costs ~10 requests.

Findings worth recording:

- JD2 stamped filenames in desktop LOCAL time (US Eastern), not UTC.
  Across 212 comparable posts: UTC 19 mismatches, UTC-5 10, UTC-4 zero.
  {date:Olocal/%Y-%m-%d} reproduces it; the trailing separator must be
  omitted or it lands in the strftime format.
- A profile's reels tab returns collab reels owned by OTHER accounts, so
  the directory must be forced with -D. JD2 did the same: chuuo3o and
  official_artms filenames sit inside "0ct0ber19 - reels".
- Stories and highlights need per-item {shortcode}; {post_shortcode} is
  the reel's id and is shared by every item. {date} is per-item, verified
  on a 154-item highlight with distinct times.
- gallery-dl reproduces JD2's caption .txt exactly, including writing
  nothing for an empty caption and omitting the trailing newline.
- The json sidecar needs `include`, not `fields`; `fields` silently does
  nothing in mode:json and leaks audio_user blobs. It yields `type`
  (post/reel) -- Instagram's own flag, which can retire the lone-video
  heuristic once the scanner reads it.

Naming differences between the two tools are cosmetic: EXPORT_RE already
makes the index optional and parseInt normalises zero-padding, so a mixed
archive parses identically. Tests pin that down.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 21:04:39 -04:00
ergosteurandClaude Opus 5 b153a49bbb fix: count the grid in the post header, and qualify the Instagram claim
Docker Build and Publish / build-and-push (push) Failing after 10s
The header rendered allPosts.length, which is pre-dedupe — 0ct0ber19
showed "303 posts" over a 300-tile grid. Instagram's counter equals its
grid, so count the grid.

Also correct CLAUDE.md. v1.7.0 claimed the grid holds everything "as on
Instagram"; Instagram actually includes a reel in the grid only when the
creator shared it to feed, per post. Measured live: official_artms has
21 reels in its first 34 grid tiles, 0ct0ber19 has 1 in 214. Archives
carry no such flag, so showing everything approximates the behaviour
rather than reproducing it.

Records the DOM trap that caused the wrong reading in the first place:
grid reels link to /reel/<code>/, not /<user>/p/<code>/.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 20:00:11 -04:00
15 changed files with 590 additions and 21 deletions
+2
View File
@@ -9,3 +9,5 @@ coverage/
!.env.example
_sample-archives
_gemini-plans
__pycache__/
*.pyc
+17 -4
View File
@@ -15,9 +15,6 @@ InstaArchive Viewer is a React 19 + Vite 6 PWA for browsing archived Instagram d
- `npm run lint` — type-check only (`tsc --noEmit`)
- `npm test` / `npm run test:watch` — vitest
- `npx vitest run src/lib/archive-patterns.test.ts` — a single test file
- `npm run jd2 -- --archives <dir> --dry-run` — generate JDownloader `.crawljob`
files for every profile on disk (see `scripts/jd2-sync.ts` and
`docs/jdownloader.md`)
Local development usually needs both `npm run dev` and `npm run server`. Local-folder mode works without the backend; server-mode archives do not.
@@ -64,6 +61,18 @@ story highlights - 4utumn07 - Sunstory -> highlight "Sunstory"
Results are cached to IndexedDB. Media records store a stable `path`; **`url` is not persistable** for local archives because blob URLs die with the document.
Three different JSON shapes turn up as `.json`, so they are told apart structurally, not by filename (`src/lib/gallery-dl-sidecar.ts`):
| shape | marker |
|---|---|
| Instagram export manifest | top-level `media` array |
| Instaloader `.json.xz` | GraphQL node under `node` / `__typename` |
| gallery-dl sidecar | flat, `post_shortcode` + `type`, none of the above |
The gallery-dl sidecar is the only source that states what a post *is*: its `type` (`post` / `reel` / `story` / `highlight`) is Instagram's own classification, so `post.isReel` set from it beats every fallback in `post-tabs.ts`. This matters — of the 781 items in `official_band - reels`, the sidecars say only **360 are reels**; the other 421 are ordinary feed videos the clips endpoint returns via `include_feed_video`. Directory-based classification counted all 781.
**Dates are ranked, not last-write-wins** (`src/lib/post-dates.ts`): sidecar (what Instagram reported) beats filename (what the fetcher wrote) beats mtime (when the file hit disk, and unrelated to when it was posted). Ties keep the incumbent. Several files describe one post and they are scanned in directory order, not in order of trustworthiness, so without the ranking the date was decided by whichever file came first. Only JDownloader highlights fall to mtime at all — `parseArchiveFilename` flags those via `dateFromMtime`.
### Cache and local-archive persistence (`src/lib/archive-cache.ts`)
IndexedDB keys are namespaced (`archive:`, `thumb:`, `handle:`) so listing archives does not deserialize every cached thumbnail blob, and thumbnails are scoped per archive to avoid cross-archive collisions.
@@ -76,7 +85,11 @@ Images over 1MiB are downscaled in a Web Worker via `OffscreenCanvas`. The queue
### Profile tabs (`src/lib/post-tabs.ts`)
The grid holds **everything**, reels included, and the Reels tab is a *filtered view* of that same set — as on Instagram. Only the Reels tab filters. The tabs were mutually exclusive until v1.7.0, which hid a lot: 1100 of `groupfandom`'s 3813 posts and 533 of `for.member`'s 1225 never appeared in the grid at all.
The grid holds **everything**, reels included, and the Reels tab is a *filtered view* of that same set. Only the Reels tab filters. The tabs were mutually exclusive until v1.7.0, which hid a lot: 1100 of `groupfandom`'s 3813 posts and 533 of `for.member`'s 1225 never appeared in the grid at all.
This *approximates* Instagram rather than matching it. Instagram's grid includes a reel only if the creator shared it to feed — a per-post choice, measured live on 2026-08-16: `official_band` had 21 reels in its first 34 grid tiles, `4utumn07` just 1 in 214. That flag appears nowhere in an archive (JD2 stores no metadata, and Instaloader's `product_type` says what a post *is*, not whether it was shared to feed), so showing everything is the closest reachable behaviour. Instagram's "N posts" counter equals its grid, which is why the header counts `postsForTab(allPosts, 'posts')` and not `allPosts` — the raw list still holds both copies of a double-fetched post.
When checking the live site, note that grid reels link to `/reel/<code>/` while ordinary posts link to `/<user>/p/<code>/`. Matching only `/p/` silently drops every reel, which once produced a confident and completely wrong conclusion that Instagram never shows reels in the grid.
Deciding *what is a reel* has no good answer for most archives. Instagram's own marker is `product_type` on the post's GraphQL node (`clips` = reel, `feed` = ordinary feed video, `igtv`, `story`) — `__typename` is `GraphVideo` for all three, and aspect ratio does not separate them either. But:
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "instaarchive-viewer",
"version": "1.7.0",
"version": "1.8.1",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "instaarchive-viewer",
"version": "1.7.0",
"version": "1.8.1",
"dependencies": {
"@tailwindcss/vite": "^4.1.14",
"@vitejs/plugin-react": "^5.0.4",
+3 -4
View File
@@ -1,19 +1,18 @@
{
"name": "instaarchive-viewer",
"private": true,
"version": "1.7.0",
"version": "1.8.1",
"type": "module",
"scripts": {
"dev": "vite --port=3000 --host=0.0.0.0",
"build": "vite build && npm run build:server",
"build:server": "tsc server.ts --esModuleInterop --module ESNext --target ES2022 --moduleResolution bundler --outDir dist-server",
"build:server": "tsc server.ts --esModuleInterop --module ESNext --target ES2022 --moduleResolution bundler --removeComments --outDir dist-server",
"preview": "vite preview",
"server": "tsx server.ts",
"clean": "rm -rf dist",
"lint": "tsc --noEmit",
"test": "vitest run",
"test:watch": "vitest",
"jd2": "tsx scripts/jd2-sync.ts"
"test:watch": "vitest"
},
"dependencies": {
"@tailwindcss/vite": "^4.1.14",
+63 -2
View File
@@ -73,6 +73,17 @@ export default function App() {
*/
const initialRouteRef = useRef(parseRoute(window.location.pathname, window.location.search));
/**
* Back-button support for the post view. Every other URL change
* (`replaceState`s the tab/archive) is intentionally NOT pushed — only
* opening a post gets its own history entry, matching Instagram's own
* back-button behaviour: Back closes the post instead of leaving the app.
*/
const pushedPostRef = useRef(false);
/** Set while reacting to a popstate, so the URL-sync effect below does not
* try to push/replace/back() again for a change the browser already made. */
const suppressNextSyncRef = useRef(false);
const isMobile = useIsMobile();
const fileInputRef = useRef<HTMLInputElement>(null);
const profilePicInputRef = useRef<HTMLInputElement>(null);
@@ -179,6 +190,13 @@ export default function App() {
*/
const filteredPosts = useMemo(() => postsForTab(allPosts, activeTab), [allPosts, activeTab]);
/**
* Instagram's "N posts" counter equals what its grid holds, so count the grid
* rather than `allPosts` — the raw list still holds both copies of any post
* fetched into two directories.
*/
const totalPosts = useMemo(() => postsForTab(allPosts, 'posts').length, [allPosts]);
/** Story highlights, grouped into the circles shown under the bio. */
const highlightGroups = useMemo(() => {
const groups = new Map<string, Post[]>();
@@ -360,6 +378,13 @@ export default function App() {
// loader below is waiting to read.
if (!hasInitialLoaded) return;
// Consumed exactly once per popstate, regardless of what happens below:
// a popstate-driven change often already matches the URL (the browser
// already moved the pointer), which used to leave this flag stuck true
// and silently no-op the NEXT real close/open until something reset it.
const wasPopState = suppressNextSyncRef.current;
suppressNextSyncRef.current = false;
const archive = currentArchive?.name ?? (allPosts.length > 0 ? username : null) ?? null;
const nextPath = buildPath({
archive,
@@ -367,12 +392,48 @@ export default function App() {
post: selectedPost ? postSlug(selectedPost) : null,
});
if (nextPath !== window.location.pathname + window.location.search) {
if (nextPath === window.location.pathname + window.location.search) return;
if (wasPopState) return; // the browser already navigated; nothing to add
console.log(`[Permalink] Updating URL to: ${nextPath}`);
if (selectedPost && !pushedPostRef.current) {
// Opening a post: push, so Back closes it instead of leaving the app.
window.history.pushState(null, '', nextPath);
pushedPostRef.current = true;
} else if (!selectedPost && pushedPostRef.current) {
// Closing a post that was pushed for: consume that entry rather than
// piling a new one on top of it, so Back still means "one step".
pushedPostRef.current = false;
window.history.back();
} else {
window.history.replaceState(null, '', nextPath);
}
}, [hasInitialLoaded, currentArchive?.name, username, allPosts.length, activeTab, selectedPost?.id]);
/**
* Back/forward support for the post view. Only a post push (above) ever
* creates an entry, so this only ever needs to open or close a post —
* never re-derive the tab or archive, which stayed on replaceState.
*/
useEffect(() => {
const onPopState = () => {
suppressNextSyncRef.current = true;
const route = parseRoute(window.location.pathname, window.location.search);
const post = route.post ? findPostBySlug(allPosts, route.post) : null;
if (post) {
setActiveTab(tabForSource(post.source));
setSelectedPost(post);
pushedPostRef.current = true; // forward navigation can land back here
} else {
setSelectedPost(null);
pushedPostRef.current = false;
}
};
window.addEventListener('popstate', onPopState);
return () => window.removeEventListener('popstate', onPopState);
}, [allPosts]);
useEffect(() => {
if (hasInitialLoaded) return;
@@ -519,7 +580,7 @@ export default function App() {
{allProfilePics.length > 1 && <button onClick={cycleProfilePic} className="bg-gray-100 hover:bg-gray-200 px-4 py-1.5 rounded-lg text-sm font-semibold transition-colors flex items-center gap-2 text-black"><Layers size={16} />Next Profile Pic</button>}
</div>
</div>
<div className="flex justify-center md:justify-start gap-10 text-sm md:text-base text-black"><div><span className="font-semibold text-black/80 text-black">{allPosts.length}</span> posts</div><div><span className="font-semibold text-black/80 text-black">{(followerCount || 0).toLocaleString()}</span> followers</div><div><span className="font-semibold text-black/80 text-black">{(followingCount || 0).toLocaleString()}</span> following</div></div>
<div className="flex justify-center md:justify-start gap-10 text-sm md:text-base text-black"><div><span className="font-semibold text-black/80 text-black">{totalPosts.toLocaleString()}</span> posts</div><div><span className="font-semibold text-black/80 text-black">{(followerCount || 0).toLocaleString()}</span> followers</div><div><span className="font-semibold text-black/80 text-black">{(followingCount || 0).toLocaleString()}</span> following</div></div>
<div className="space-y-1 text-black/80 text-black"><div className="font-semibold text-black">{fullName || `@${username}`}</div><div className="text-gray-600 whitespace-pre-wrap max-w-sm mx-auto md:mx-0 text-sm md:text-base text-black">{bio || 'Archived profile viewer for local files.'}</div>{externalUrl && <a href={externalUrl} target="_blank" rel="noopener noreferrer" className="text-blue-900 font-semibold text-sm block hover:underline truncate max-w-[250px] text-black">{externalUrl.replace(/^https?:\/\/(www\.)?/, '')}</a>}</div>
</div>
</header>
+27 -1
View File
@@ -4,6 +4,8 @@ import { XzReadableStream } from 'xz-decompress';
import { ArchiveFile, CacheData, Post, ServerArchive } from '../types';
import { setCachedArchive, getDirectoryHandle } from '../lib/archive-cache';
import { parseArchiveFilename, scopedPostId, EXPORT_RE, INSTALOADER_RE } from '../lib/archive-patterns';
import { isGalleryDlSidecar, sidecarDate, sidecarIsReel } from '../lib/gallery-dl-sidecar';
import { DateSource, shouldReplaceDate } from '../lib/post-dates';
const hasDirectoryHandle = async (name: string) => Boolean(await getDirectoryHandle(name));
@@ -142,6 +144,17 @@ export const useArchiveScanner = (
try {
const postsMap = new Map<string, Partial<Post>>();
/**
* Which source supplied each post's date, so a better one can replace it.
* Sidecar beats filename beats mtime — see src/lib/post-dates.ts.
*/
const dateSources = new Map<string, DateSource>();
const applyDate = (postId: string, post: Partial<Post>, date: string, source: DateSource) => {
const current = post.date ? { date: post.date, source: dateSources.get(postId) ?? 'mtime' } : undefined;
if (!shouldReplaceDate(current, { date, source })) return;
post.date = date;
dateSources.set(postId, source);
};
const mediaFilesMap = new Map<string, ArchiveFile>();
const discoveredProfilePics: { name: string, url: string }[] = [];
const allImageFiles: ArchiveFile[] = [];
@@ -300,13 +313,26 @@ export const useArchiveScanner = (
}
else if (isStory) post.isStory = true;
// Files describing one post are scanned in directory order, not in
// order of trustworthiness, so every date goes through the ranking
// in post-dates.ts rather than last-write-wins.
applyDate(postId, post, date, parsed.dateFromMtime ? 'mtime' : 'filename');
const lowerExt = ext.toLowerCase();
if (lowerExt === 'txt') {
try { post.caption = await file.text(); } catch(e) {}
} else if (lowerExt === 'json' || lowerName.endsWith('.json.xz')) {
try {
const data = lowerName.endsWith('.xz') ? await parseXZFile(file) : JSON.parse(await file.text());
if (data) {
if (isGalleryDlSidecar(data)) {
// The only format that states what a post is rather than
// leaving it to be inferred from filenames.
if (data.description) post.caption = data.description;
const reel = sidecarIsReel(data);
if (reel !== undefined) post.isReel = reel;
if (data.type === 'story') post.isStory = true;
applyDate(postId, post, sidecarDate(data), 'sidecar');
} else if (data) {
const node = data.node || data; const iphone = node.iphone_struct || {};
const captionText = node.edge_media_to_caption?.edges?.[0]?.node?.text || node.caption?.text || iphone.caption?.text || '';
if (captionText) post.caption = captionText;
+110 -1
View File
@@ -1,5 +1,5 @@
import { describe, expect, it } from 'vitest';
import { parseArchiveFilename, scopedPostId } from './archive-patterns';
import { canonicalItemId, parseArchiveFilename, scopedPostId } from './archive-patterns';
describe('parseArchiveFilename — Instagram export format', () => {
it('parses a single-image post', () => {
@@ -10,6 +10,7 @@ describe('parseArchiveFilename — Instagram export format', () => {
index: 1,
ext: 'mp4',
isStory: false,
dateFromMtime: false,
});
});
@@ -116,3 +117,111 @@ describe('scopedPostId', () => {
expect(inPosts).not.toBe(inHighlight);
});
});
/**
* gallery-dl is replacing JDownloader as the fetcher (docs/gallery-dl.md).
* Its naming differs cosmetically, and these cases pin down that the two
* interoperate so a mixed archive parses identically.
*/
describe('gallery-dl / JDownloader naming interop', () => {
it('treats a single-media post the same with or without an index', () => {
const jd2 = parseArchiveFilename('2023-04-19_4utumn07 - CrORBIcJJbM.mp4')!;
const gdl = parseArchiveFilename('2023-04-19_4utumn07 - CrORBIcJJbM - 1.mp4')!;
expect(jd2.postId).toBe(gdl.postId);
expect(jd2.index).toBe(gdl.index);
expect(jd2.index).toBe(1);
});
it('normalises zero-padded carousel indices', () => {
// JD2 pads to the width of the media count (10+ items -> "01"), and
// gallery-dl's count can be one higher, so the same post may be padded
// by one tool and not the other.
expect(parseArchiveFilename('2024-04-17_4utumn07 - C53YPQzp7Wj - 09.jpg')!.index).toBe(9);
expect(parseArchiveFilename('2024-04-17_4utumn07 - C53YPQzp7Wj - 9.jpg')!.index).toBe(9);
expect(parseArchiveFilename('2023-11-03_4utumn07 - CzM8Uf6B6H_ - 01.jpg')!.index).toBe(1);
});
it('reads a gallery-dl story name, which carries a per-item shortcode', () => {
const p = parseArchiveFilename('2026-08-16_official_band - DcF9OyhBJ1H.jpg', 'stories')!;
expect(p.postId).toBe('DcF9OyhBJ1H');
expect(p.date).toBe('2026-08-16');
});
it('gives a dated highlight a real date instead of the mtime fallback', () => {
const mtime = Date.parse('2026-08-17T00:00:00Z');
const undated = parseArchiveFilename('4utumn07 - C-IImhvpFuk.jpg', 'highlight', mtime)!;
const dated = parseArchiveFilename('2024-08-04_4utumn07 - C-IImhvpFuk.jpg', 'highlight', mtime)!;
// Same item either way, so re-fetching cannot split it into two posts.
expect(dated.postId).toBe(undated.postId);
expect(undated.date).toBe('2026-08-17');
expect(dated.date).toBe('2024-08-04');
});
});
/**
* Highlights are the only files with no date in the name, so they fall back to
* mtime — which is when the file was written, not when it was posted. Callers
* need to know the difference to let a real date win.
*/
describe('dateFromMtime', () => {
const mtime = Date.parse('2026-08-17T00:00:00Z');
it('flags an undated highlight name as mtime-dated', () => {
const p = parseArchiveFilename('4utumn07 - C-IImhvpFuk.jpg', 'highlight', mtime)!;
expect(p.date).toBe('2026-08-17');
expect(p.dateFromMtime).toBe(true);
});
it('does not flag a highlight that carries its own date', () => {
const p = parseArchiveFilename('2024-08-04_4utumn07 - C-IImhvpFuk.jpg', 'highlight', mtime)!;
expect(p.date).toBe('2024-08-04');
expect(p.dateFromMtime).toBe(false);
});
it('never flags ordinary post or Instaloader names', () => {
expect(parseArchiveFilename('2023-04-19_u - ABC.mp4', 'posts', mtime)!.dateFromMtime).toBe(false);
expect(parseArchiveFilename('2024-01-01_12-00-00_UTC.jpg', 'posts', mtime)!.dateFromMtime).toBe(false);
});
it('leaves the date empty rather than guessing when no mtime is given', () => {
const p = parseArchiveFilename('4utumn07 - C-IImhvpFuk.jpg', 'highlight')!;
expect(p.date).toBe('');
expect(p.dateFromMtime).toBe(false);
});
});
/**
* JDownloader wrote story-shaped names for highlights during one period, so
* the same item exists under two conventions. They must be one post.
*/
describe('canonicalItemId', () => {
it('collapses the two highlight naming conventions onto one id', () => {
const dir = 'story highlights - 4utumn07 - Sunstory';
const undated = parseArchiveFilename('4utumn07 - C5dQPEYpd9W.mp4', 'highlight', 1)!;
const dated = parseArchiveFilename('2024-04-07_4utumn07 - 01 - C5dQPEYpd9W.mp4', 'highlight')!;
expect(scopedPostId(dated.postId, 'highlight', dir))
.toBe(scopedPostId(undated.postId, 'highlight', dir));
});
it('does the same for stories', () => {
const a = parseArchiveFilename('2025-10-26_u - 2 - DQRuDx9iW5Q.jpg', 'stories')!;
expect(scopedPostId(a.postId, 'stories', 'story - u')).toBe('story - u/DQRuDx9iW5Q');
});
it('keeps distinct story items distinct', () => {
const a = parseArchiveFilename('2026-08-13_u - 1 - Db-UTJcCUUr.mp4', 'stories')!;
const b = parseArchiveFilename('2026-08-13_u - 2 - Db-oNJ1CWQ4.mp4', 'stories')!;
expect(scopedPostId(a.postId, 'stories', 'story - u'))
.not.toBe(scopedPostId(b.postId, 'stories', 'story - u'));
});
it('leaves a shortcode that merely starts with digits alone', () => {
expect(canonicalItemId('4utumn07')).toBe('4utumn07');
expect(canonicalItemId('C5dQPEYpd9W')).toBe('C5dQPEYpd9W');
expect(canonicalItemId('12345')).toBe('12345');
});
it('does not touch posts, whose ids are permalinks', () => {
expect(scopedPostId('01 - ABC', 'posts')).toBe('01 - ABC');
});
});
+38 -2
View File
@@ -30,6 +30,15 @@ export interface ParsedFilename {
index: number;
ext: string;
isStory: boolean;
/**
* True when `date` is the file's mtime rather than anything Instagram said.
*
* Only highlights fetched by JDownloader lack a date in the filename, and
* their mtime is just when the file was written. Callers should let any real
* date win over this one — the same item is often also present under a
* gallery-dl name that does carry the date.
*/
dateFromMtime: boolean;
}
/**
@@ -54,6 +63,7 @@ export const parseArchiveFilename = (
index: indexStr ? parseInt(indexStr, 10) : 1,
ext,
isStory: Boolean(story),
dateFromMtime: false,
};
}
@@ -67,6 +77,7 @@ export const parseArchiveFilename = (
index: indexStr ? parseInt(indexStr, 10) : 1,
ext,
isStory: Boolean(story),
dateFromMtime: false,
};
}
@@ -81,6 +92,7 @@ export const parseArchiveFilename = (
index: 1,
ext,
isStory: false,
dateFromMtime: Boolean(mtime),
};
}
}
@@ -88,12 +100,36 @@ export const parseArchiveFilename = (
return null;
};
/**
* A leading per-day ordinal on a story or highlight id: `01 - C5dQPEYpd9W`.
*
* JDownloader wrote story-shaped names for highlights during one period of its
* life, so the same item exists as both `user - CODE.jpg` and
* `date_user - 01 - CODE.jpg`. Those parse to different ids and the viewer
* shows the item twice. The ordinal carries no information the shortcode does
* not — it is a position within a day's stories, and the shortcode is already
* unique — so it is dropped.
*/
const LEADING_ORDINAL = /^\d+ - (?=[A-Za-z0-9_-]+$)/;
/** Strip the ordinal so both naming conventions land on the same post. */
export const canonicalItemId = (postId: string): string =>
postId.replace(LEADING_ORDINAL, '');
/**
* Namespace a post ID by its source directory.
*
* Base-profile IDs are left untouched so existing permalinks keep working;
* sidecar IDs are prefixed so a shortcode appearing in both the profile and a
* highlight stays two distinct posts.
*
* Story and highlight ids are canonicalised first, so an item fetched under
* two different naming conventions is one post rather than two.
*/
export const scopedPostId = (postId: string, kind: SourceKind, dir?: string): string =>
kind === 'posts' ? postId : `${dir ?? kind}/${postId}`;
export const scopedPostId = (postId: string, kind: SourceKind, dir?: string): string => {
if (kind === 'posts') return postId;
const id = (kind === 'stories' || kind === 'highlight')
? canonicalItemId(postId)
: postId;
return `${dir ?? kind}/${id}`;
};
+82
View File
@@ -0,0 +1,82 @@
import { describe, expect, it } from 'vitest';
import {
GalleryDlSidecar, isGalleryDlSidecar, sidecarDate, sidecarIsReel, sidecarSource,
} from './gallery-dl-sidecar';
// Trimmed from real files published to the archive on 2026-08-16.
const REEL: GalleryDlSidecar = {
post_shortcode: 'Db-lNCoib9m', post_id: '3962768346034323302', type: 'reel',
date: '2026-08-13 11:00:44', post_date: '2026-08-13 11:00:44',
username: 'official_band', fullname: 'Official ARTMS',
description: 'Dancing in the spotlight', count: 1, likes: 22914,
};
const FEED_VIDEO: GalleryDlSidecar = { ...REEL, post_shortcode: 'DbdG9L9jU4m', type: 'post', count: 2 };
const HIGHLIGHT: GalleryDlSidecar = {
post_shortcode: 'BATVdRZi_3', post_id: '18099435932626935', type: 'highlight',
date: '2026-08-08 16:22:09', username: 'official_band', count: 154,
};
describe('isGalleryDlSidecar', () => {
it('accepts a real sidecar', () => {
expect(isGalleryDlSidecar(REEL)).toBe(true);
expect(isGalleryDlSidecar(HIGHLIGHT)).toBe(true);
});
it('rejects an Instaloader GraphQL payload', () => {
expect(isGalleryDlSidecar({ node: { __typename: 'GraphVideo', shortcode: 'x' } })).toBe(false);
expect(isGalleryDlSidecar({ __typename: 'GraphImage', post_shortcode: 'x', type: 'post' })).toBe(false);
});
it('rejects an Instagram export manifest', () => {
expect(isGalleryDlSidecar({ media: [{ uri: 'a.jpg' }] })).toBe(false);
expect(isGalleryDlSidecar([{ media: [] }])).toBe(false);
});
it('rejects junk', () => {
for (const v of [null, undefined, 0, '', 'string', {}, { post_shortcode: 'x' }]) {
expect(isGalleryDlSidecar(v)).toBe(false);
}
});
});
describe('sidecarDate', () => {
it('takes the day from the timestamp', () => {
expect(sidecarDate(REEL)).toBe('2026-08-13');
});
it('falls back to post_date', () => {
expect(sidecarDate({ post_shortcode: 'x', post_date: '2024-01-02 03:04:05' })).toBe('2024-01-02');
});
it('returns empty when there is no usable date', () => {
expect(sidecarDate({ post_shortcode: 'x' })).toBe('');
expect(sidecarDate({ post_shortcode: 'x', date: 'not a date' })).toBe('');
});
});
describe('sidecarIsReel', () => {
it('distinguishes a reel from an ordinary feed video', () => {
// Both are single mp4s -- the lone-video heuristic cannot tell them apart.
expect(sidecarIsReel(REEL)).toBe(true);
expect(sidecarIsReel(FEED_VIDEO)).toBe(false);
});
it('declines to answer for stories and highlights', () => {
expect(sidecarIsReel(HIGHLIGHT)).toBeUndefined();
expect(sidecarIsReel({ post_shortcode: 'x', type: 'story' as const })).toBeUndefined();
expect(sidecarIsReel({ post_shortcode: 'x' })).toBeUndefined();
});
});
describe('sidecarSource', () => {
it('maps type onto the archive source kinds', () => {
expect(sidecarSource(REEL)).toBe('reels');
expect(sidecarSource(FEED_VIDEO)).toBe('posts');
expect(sidecarSource(HIGHLIGHT)).toBe('highlight');
expect(sidecarSource({ post_shortcode: 'x', type: 'story' as const })).toBe('stories');
});
it('is undefined for an unknown type', () => {
expect(sidecarSource({ post_shortcode: 'x' })).toBeUndefined();
});
});
+87
View File
@@ -0,0 +1,87 @@
import { SourceKind } from '../types';
/**
* gallery-dl `.json` metadata sidecars.
*
* Written one per post next to the media (see docs/gallery-dl.md). This is the
* only source in any archive format that states outright what a post *is* —
* `type` is Instagram's own classification, the `product_type: "clips"` signal
* carried through the listing response. Everything else the viewer knows about
* reels is guesswork from filenames and directory names.
*
* Deliberately separate from the two older JSON shapes the scanner reads:
*
* Instagram export `posts_1.json`, an array of entries with `media`
* Instaloader `.json.xz`, a GraphQL node under `node`
* gallery-dl this — flat, no wrapper
*/
export interface GalleryDlSidecar {
post_shortcode: string;
post_id?: string;
/** Instagram's own classification of the post. */
type?: 'post' | 'reel' | 'story' | 'highlight';
/** Local-time "YYYY-MM-DD HH:MM:SS" — gallery-dl is configured to emit local. */
date?: string;
post_date?: string;
username?: string;
fullname?: string;
description?: string;
count?: number;
likes?: number;
post_url?: string;
}
/**
* Recognise a gallery-dl sidecar.
*
* Checked structurally rather than by filename, because the older formats are
* also plain `.json`. `node` and `__typename` are what an Instaloader or
* export payload carries, and their absence is what makes this shape
* unambiguous.
*/
export const isGalleryDlSidecar = (data: unknown): data is GalleryDlSidecar => {
if (!data || typeof data !== 'object' || Array.isArray(data)) return false;
const o = data as Record<string, unknown>;
return typeof o.post_shortcode === 'string'
&& typeof o.type === 'string'
&& o.node === undefined
&& o.__typename === undefined
&& o.media === undefined;
};
/** The ISO date (YYYY-MM-DD) a sidecar reports, or '' if it carries none. */
export const sidecarDate = (s: GalleryDlSidecar): string => {
const raw = s.date || s.post_date || '';
const day = raw.slice(0, 10);
return /^\d{4}-\d{2}-\d{2}$/.test(day) ? day : '';
};
/**
* Whether the sidecar says this post is a reel.
*
* Returns undefined rather than false for stories and highlights: those are
* neither reels nor grid posts, and answering "no" would let them be counted
* as ordinary posts.
*/
export const sidecarIsReel = (s: GalleryDlSidecar): boolean | undefined => {
if (s.type === 'reel') return true;
if (s.type === 'post') return false;
return undefined;
};
/**
* Which source kind the sidecar implies, for cross-checking the directory.
*
* A reel shared to the profile grid legitimately appears under `posts`, so a
* disagreement is not an error — the directory says where the file was
* fetched from, `type` says what Instagram considers it.
*/
export const sidecarSource = (s: GalleryDlSidecar): SourceKind | undefined => {
switch (s.type) {
case 'reel': return 'reels';
case 'post': return 'posts';
case 'story': return 'stories';
case 'highlight': return 'highlight';
default: return undefined;
}
};
+56
View File
@@ -0,0 +1,56 @@
import { describe, expect, it } from 'vitest';
import { DatedValue, preferDate, shouldReplaceDate } from './post-dates';
const sidecar: DatedValue = { date: '2024-04-07', source: 'sidecar' };
const filename: DatedValue = { date: '2024-04-08', source: 'filename' };
const mtime: DatedValue = { date: '2026-08-17', source: 'mtime' };
describe('date precedence', () => {
it('ranks sidecar above filename above mtime', () => {
expect(preferDate(mtime, filename)).toEqual(filename);
expect(preferDate(filename, sidecar)).toEqual(sidecar);
expect(preferDate(mtime, sidecar)).toEqual(sidecar);
});
it('never lets a weaker source overwrite a stronger one', () => {
expect(preferDate(sidecar, filename)).toEqual(sidecar);
expect(preferDate(sidecar, mtime)).toEqual(sidecar);
expect(preferDate(filename, mtime)).toEqual(filename);
});
it('keeps the incumbent on a tie, so scan order cannot flip the date', () => {
const other: DatedValue = { date: '2020-01-01', source: 'filename' };
expect(preferDate(filename, other)).toEqual(filename);
expect(preferDate(other, filename)).toEqual(other);
});
it('accepts anything when nothing is held yet', () => {
expect(preferDate(undefined, mtime)).toEqual(mtime);
expect(shouldReplaceDate(undefined, mtime)).toBe(true);
});
it('ignores an empty date regardless of source', () => {
const empty: DatedValue = { date: '', source: 'sidecar' };
expect(shouldReplaceDate(filename, empty)).toBe(false);
expect(preferDate(filename, empty)).toEqual(filename);
});
it('replaces a held-but-empty date', () => {
const empty: DatedValue = { date: '', source: 'filename' };
expect(preferDate(empty, mtime)).toEqual(mtime);
});
it('is order-independent for the full three-source case', () => {
const orders = [
[mtime, filename, sidecar],
[sidecar, mtime, filename],
[filename, sidecar, mtime],
[mtime, sidecar, filename],
];
for (const order of orders) {
const won = order.reduce<DatedValue | undefined>(
(acc, next) => preferDate(acc, next), undefined);
expect(won).toEqual(sidecar);
}
});
});
+46
View File
@@ -0,0 +1,46 @@
/**
* Where a post's date came from, and which source wins.
*
* A post is usually described by several files — media, a caption `.txt`, a
* `.json` sidecar, sometimes the same item under two naming conventions — and
* they are scanned in directory order, not in order of trustworthiness. Without
* an explicit ranking the date is decided by whichever file happened to be
* reached first.
*
* Ranked best to worst:
*
* sidecar what Instagram reported, straight from a gallery-dl `.json`
* filename a date the fetcher wrote into the name; correct, but derived
* mtime when the file was written to disk — unrelated to when it was
* posted, and only ever a last resort for JDownloader highlights,
* whose filenames carry no date at all
*/
export type DateSource = 'sidecar' | 'filename' | 'mtime';
const RANK: Record<DateSource, number> = { sidecar: 0, filename: 1, mtime: 2 };
export interface DatedValue {
date: string;
source: DateSource;
}
/**
* Whether `next` should replace the date currently held.
*
* Ties keep the incumbent, so scanning stays stable: two files of equal
* authority cannot flip a post's date back and forth by scan order.
*/
export const shouldReplaceDate = (
current: DatedValue | undefined,
next: DatedValue,
): boolean => {
if (!next.date) return false;
if (!current || !current.date) return true;
return RANK[next.source] < RANK[current.source];
};
/** Apply `next` if it outranks `current`, otherwise keep what we have. */
export const preferDate = (
current: DatedValue | undefined,
next: DatedValue,
): DatedValue => (shouldReplaceDate(current, next) ? next : (current ?? next));
+41
View File
@@ -98,3 +98,44 @@ describe('postsForTab', () => {
expect(postsForTab([post('A')], 'saved')).toEqual([]);
});
});
/**
* Once an archive carries gallery-dl sidecars, the guesswork above is replaced
* by Instagram's own classification. These are the cases the heuristic got
* wrong (see docs/gallery-dl.md).
*/
describe('explicit isReel from a sidecar', () => {
it('beats the lone-video heuristic for an ordinary feed video', () => {
// A single mp4 that Instagram calls a post, not a reel — indistinguishable
// by shape alone.
const posts = [video('DbdG9L9jU4m', { isReel: false })];
expect(postsForTab(posts, 'reels')).toEqual([]);
expect(postsForTab(posts, 'posts')).toHaveLength(1);
});
it('recognises a reel that lives in the profile grid', () => {
// Shared to feed, so it sits in the base directory with source 'posts'.
const posts = [post('A'), video('C8FHM6EJl15', { source: 'posts', isReel: true })];
expect(postsForTab(posts, 'reels').map(p => p.id)).toEqual(['C8FHM6EJl15']);
expect(postsForTab(posts, 'posts')).toHaveLength(2);
});
it('beats the directory when both are present', () => {
const posts = [
video('u - reels/A', { source: 'reels', isReel: false }),
video('u - reels/B', { source: 'reels' }),
];
// A is a feed video that the reels tab happened to return; B is unlabelled
// and falls back to its directory.
expect(postsForTab(posts, 'reels').map(p => p.id)).toEqual(['u - reels/B']);
});
it('falls back per post, so a mixed archive still works', () => {
const posts = [
video('labelled', { isReel: true }),
video('unlabelled'),
carousel('C'),
];
expect(postsForTab(posts, 'reels').map(p => p.id)).toEqual(['labelled', 'unlabelled']);
});
});
+9 -4
View File
@@ -41,11 +41,16 @@ export const hasReelSource = (posts: Post[]): boolean => posts.some(p => p.sourc
* or an old IGTV upload, all three of which are plain `GraphVideo` nodes
* distinguished only by `product_type`.
*/
export const makeIsReel = (posts: Post[]): ((post: Post) => boolean) => (
hasReelSource(posts)
export const makeIsReel = (posts: Post[]): ((post: Post) => boolean) => {
const guess = hasReelSource(posts)
? (post: Post) => post.source === 'reels'
: (post: Post) => post.media.length === 1 && post.media[0]?.type === 'video'
);
: (post: Post) => post.media.length === 1 && post.media[0]?.type === 'video';
// `isReel` comes from a gallery-dl sidecar and is Instagram's own answer, so
// it beats both fallbacks — per post, since an archive is usually a mix of
// files fetched before and after sidecars existed.
return (post: Post) => post.isReel ?? guess(post);
};
/** Preference order when the same post was fetched into more than one directory. */
const SOURCE_RANK: Record<SourceKind, number> = { reels: 0, posts: 1, stories: 2, highlight: 3 };
+6
View File
@@ -37,6 +37,12 @@ export interface Post {
isStory?: boolean;
/** Defaults to 'posts' for archives without sidecar directories. */
source?: SourceKind;
/**
* Instagram's own answer to "is this a reel", from a gallery-dl `.json`
* sidecar. Undefined when the archive carries no such sidecar, which is when
* the viewer has to fall back to guessing — see src/lib/post-tabs.ts.
*/
isReel?: boolean;
/** Highlight this post belongs to, for source === 'highlight'. */
highlightTitle?: string;
}