Compare commits

...
4 Commits
Author SHA1 Message Date
ergosteurandClaude Opus 5 e57be521a2 fix: read xz sidecars by buffer, page carousel with arrows, stop backdrop flash
Docker Build and Publish / build-and-push (push) Failing after 10s
Instaloader metadata was silently lost
Every .json.xz failed with "Failed to fetch" during a scan, though the same URL
fetched fine on its own. RemoteArchiveFile.stream() started a fetch, piped the
body into a TransformStream and returned the readable immediately — nothing
caught a fetch rejection, and the decompressor stops reading at the end of the
xz member, so the response body was never drained or cancelled. Across ~190
sidecars that exhausted the connection pool.

Everything Instaloader archives carry lives in those files, so the failure was
invisible but total. rivvsofficial reported 188 posts, no stories, 0 followers
and a placeholder bio; it now reports 68 posts, 120 stories, 10,337 followers
and the real name, bio and link — 68 + 120 = 188, matching the sidecars exactly
(106 GraphStoryVideo + 14 GraphStoryImage = 120).

These sidecars are a few KB, so they are now read into memory before
decompressing. stream() was left unused by that change and is removed from the
interface and both implementations rather than kept as a trap.

Arrow keys page the carousel
They moved between posts, which contradicted the arrows drawn on the carousel
itself. Arrows now page slides; , and . move between posts, alongside the side
buttons.

Backdrop cross-fade
AnimatePresence had no exit variant, so the outgoing scan backdrop was removed
instantly while its replacement faded in over 1.5s, exposing the pale page
behind it as a white flash. Layers now stack: the outgoing image holds full
opacity until covered, and the 0.4 moved onto the group so overlapping layers
don't darken as they cross. Measured over a real scan: 152 cross-fades with a
layer always opaque, except the opening fade-in where nothing is underneath.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 12:28:26 -04:00
ergosteurandClaude Opus 5 792b834cbe docs: add a JDownloader quick reference
Covers why fetching goes through JDownloader rather than Instaloader (the
instagram.com vs CDN split, and what the metadata gap actually costs), the
settings that matter, cookie handling, the two-URL workflow, jd2-sync usage,
the expected on-disk layout, and what to do when something breaks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 12:02:41 -04:00
ergosteurandClaude Opus 5 106d3f6691 fix: never descend into NAS metadata directories when indexing
Docker Build and Publish / build-and-push (push) Failing after 11s
The archive root was filtered by prefix, but the recursive walk below it was
not, so anything inside a profile directory got indexed. NAS filesystems put
sidecar metadata *inside* every folder rather than only at the share root:
Synology writes @eaDir (thumbnails and indexing data), #recycle holds
deletions, .sync is Resilio state. On the live share those account for 12,516
of 123,023 files.

None currently sit inside a profile directory, so nothing was miscounted yet —
but the moment that share gets indexed for Photos, every generated thumbnail
would be counted as archive media and stat'd one by one over the network, which
is the cost the index exists to avoid.

One isSystemDirectory rule now applies at every level, and the root listing uses
it too instead of keeping a second copy of the pattern.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 11:57:29 -04:00
ergosteurandClaude Opus 5 92a4ada3c2 feat: generate JDownloader crawljobs from the archives on disk
The manual flow is: paste a profile URL into JDownloader, paste the /reels URL
separately (the profile page misses some reels), set the output folder by hand,
repeat per profile. scripts/jd2-sync.ts emits one crawljob per source with the
folder already pointed at the right directory, so folder-watch picks up the
whole batch at once.

Profiles and sidecars are derived with the same grouping logic the server uses,
so output folders always match what the viewer expects to find. --download-base
maps the path for a JDownloader running on another machine (Windows paths
included), since it typically runs on a desktop against the share.

Directories that aren't Instagram profiles are skipped: an archive root also
collects tool output and exports from other services, and pointing a crawl at
those spends requests on instagram.com to be told the profile doesn't exist —
exactly the traffic worth not spending. Filtering is by username shape, plus
--skip and a .jd2ignore file for names that look like usernames but aren't.

Defaults are conservative: chunks=1, because multi-chunk ranged requests are the
one CDN-side pattern that doesn't resemble a browser, and links park in the
LinkGrabber for review rather than auto-starting.

Only posts and reels are emitted; highlight URLs need a numeric id and story
URLs expire, so those stay manual.

Format verified against JDownloader's own explain.txt for the folderwatch
extension, read from the daily SVN mirror rather than one of the decade-stale
GitHub copies.

Also refreshes CLAUDE.md, whose URL-state section still described the query
parameters replaced in 1.4.0, and documents the mobile feed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011uBWhwV3wFQ5MBCcMHHem7
2026-08-14 11:41:23 -04:00
10 changed files with 136 additions and 36 deletions
+14 -4
View File
@@ -15,6 +15,9 @@ InstaArchive Viewer is a React 19 + Vite 6 PWA for browsing archived Instagram d
- `npm run lint` — type-check only (`tsc --noEmit`)
- `npm test` / `npm run test:watch` — vitest
- `npx vitest run src/lib/archive-patterns.test.ts` — a single test file
- `npm run jd2 -- --archives <dir> --dry-run` — generate JDownloader `.crawljob`
files for every profile on disk (see `scripts/jd2-sync.ts` and
`docs/jdownloader.md`)
Local development usually needs both `npm run dev` and `npm run server`. Local-folder mode works without the backend; server-mode archives do not.
@@ -71,14 +74,21 @@ Restoring an archive **rehydrates URLs from `path`**: server archives rebuild HT
Images over 1MiB are downscaled in a Web Worker via `OffscreenCanvas`. The queue is **serial on purpose** — decoding several 50MP+ images at once OOMs the tab. `requestThumbnail` must keep a stable identity (it reads cache state through a ref), or every completed thumbnail re-runs the effect in all mounted thumbnails.
### URL state (`src/App.tsx`)
### URL state (`src/App.tsx`, `src/lib/routing.ts`)
App state syncs to `?a=` / `?t=` / `?p=`. Two rules, both learned from real bugs:
Paths mirror Instagram: `/<archive>/`, `/<archive>/reels/`, `/<archive>/p/<shortcode>/`. The old `?a=&t=&p=` form is still parsed for existing links but never written. Reserved prefixes (`api`, `archives`, `assets`…) can't be mistaken for a profile name.
- The initial query string is captured into a ref on first render; the URL is rewritten from state as soon as anything loads, so reading `window.location` later sees the rewrite, not the user's link.
A post URL carries no tab, as on Instagram — the tab is re-derived from the post's `source`, so a reel link lands on the Reels tab and pages through reels. Sidecar posts keep directory-scoped ids internally but expose only the shortcode.
Three rules, all learned from real bugs:
- The initial route is captured into a ref on first render; the URL is rewritten from state as soon as anything loads, so reading `window.location` later sees the rewrite, not the user's link.
- URL writing is gated on `hasInitialLoaded`, otherwise it erases the deep link before the loader consumes it.
- Deep-link resolution waits on the archive fetch having *settled* (`archivesFetched`), not on `isServerMode`, which is still false while the request is in flight.
Deep-link resolution waits on the archive fetch having *settled*, not on `isServerMode` (which is still false while in flight).
### Mobile feed (`src/components/PostFeed.tsx`)
Below `md`, opening a post renders a scrolling feed page rather than the modal (`useIsMobile` decides). Only a window of posts is mounted; it grows both ways, and prepending corrects `scrollTop` in a `useLayoutEffect` so content doesn't jump. Only the post crossing the viewport centre plays its video and drives the URL. Desktop keeps `PostModal`; both share `MediaCarousel`.
### Backend (`server.ts`)
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "instaarchive-viewer",
"version": "1.5.0",
"version": "1.6.0",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "instaarchive-viewer",
"version": "1.5.0",
"version": "1.6.0",
"dependencies": {
"@tailwindcss/vite": "^4.1.14",
"@vitejs/plugin-react": "^5.0.4",
+3 -2
View File
@@ -1,7 +1,7 @@
{
"name": "instaarchive-viewer",
"private": true,
"version": "1.5.0",
"version": "1.6.0",
"type": "module",
"scripts": {
"dev": "vite --port=3000 --host=0.0.0.0",
@@ -12,7 +12,8 @@
"clean": "rm -rf dist",
"lint": "tsc --noEmit",
"test": "vitest run",
"test:watch": "vitest"
"test:watch": "vitest",
"jd2": "tsx scripts/jd2-sync.ts"
},
"dependencies": {
"@tailwindcss/vite": "^4.1.14",
+44 -9
View File
@@ -17,7 +17,7 @@ import {
import { motion, AnimatePresence } from 'motion/react';
import { cn } from './lib/utils';
import { PRESS } from './lib/motion';
import { PRESS, prefersReducedMotion } from './lib/motion';
import { buildPath, findPostBySlug, parseRoute, postSlug, tabForSource } from './lib/routing';
import { LocalArchiveFile, RemoteArchiveFile } from './lib/archive-files';
import {
@@ -111,6 +111,30 @@ export default function App() {
const [lastLoadedScanningImage, setLastLoadedScanningImage] = useState<string | null>(null);
/**
* Blurred backdrops behind the scanning UI, newest last.
*
* Each new image is stacked *over* the previous one and fades in; the one
* underneath stays fully opaque until it's covered. Cross-fading by swapping
* a single element left the pale backdrop showing through mid-transition,
* which read as a white flash between every image.
*/
const [scanBackdrops, setScanBackdrops] = useState<string[]>([]);
useEffect(() => {
if (!lastLoadedScanningImage) return;
setScanBackdrops(prev =>
prev[prev.length - 1] === lastLoadedScanningImage
? prev
: [...prev, lastLoadedScanningImage].slice(-3),
);
}, [lastLoadedScanningImage]);
// Don't carry one archive's backdrops into the next scan.
useEffect(() => {
if (!isScanning) { setScanBackdrops([]); setLastLoadedScanningImage(null); }
}, [isScanning]);
const {
username,
fullName,
@@ -460,17 +484,28 @@ export default function App() {
onLoad={() => setLastLoadedScanningImage(currentScanningImage)}
/>
)}
<div className="absolute inset-0 z-0">
<AnimatePresence initial={false}>
<motion.img
key={lastLoadedScanningImage}
src={lastLoadedScanningImage || undefined}
{/*
The 0.4 lives on the group, not the images: two layers overlap
during a cross-fade, and fading them individually would darken the
backdrop as they cross. Inside the group each layer goes to full
opacity, so the stack is always completely covered.
*/}
<div className="absolute inset-0 z-0 opacity-40">
{scanBackdrops.map(src => (
<motion.img
key={src}
src={src}
initial={{ opacity: 0 }}
animate={{ opacity: 0.4 }}
transition={{ duration: 1.5 }}
animate={{ opacity: 1 }}
transition={prefersReducedMotion() ? { duration: 0 } : { duration: 0.9, ease: 'easeInOut' }}
onAnimationComplete={() => setScanBackdrops(prev => {
// Once this layer is opaque it hides everything below it.
const i = prev.indexOf(src);
return i > 0 ? prev.slice(i) : prev;
})}
className="absolute inset-0 w-full h-full object-cover blur-[60px] scale-110"
/>
</AnimatePresence>
))}
</div>
<div className="absolute inset-0 bg-white/40 z-1" />
<div className="relative z-10 w-full max-w-4xl px-4 flex flex-col items-center gap-8 text-black">
+7 -4
View File
@@ -82,11 +82,14 @@ export const PostModal: React.FC<PostModalProps> = ({
useEffect(() => setCurrentIndex(0), [post.id]);
useEffect(() => {
// Arrows page within the carousel — the thing the arrows visually point at.
// Moving between posts stays on the side buttons, with , and . as keyboard
// equivalents.
const handleKeyDown = (e: KeyboardEvent) => {
if (e.key === 'ArrowRight') goToPost(1, 'x');
else if (e.key === 'ArrowLeft') goToPost(-1, 'x');
else if (e.key === '.') paginate(1);
else if (e.key === ',') paginate(-1);
if (e.key === 'ArrowRight') paginate(1);
else if (e.key === 'ArrowLeft') paginate(-1);
else if (e.key === '.') goToPost(1, 'x');
else if (e.key === ',') goToPost(-1, 'x');
else if (e.key === 'Escape') onClose();
};
window.addEventListener('keydown', handleKeyDown);
+15 -3
View File
@@ -107,11 +107,23 @@ export const useArchiveScanner = (
/** Stable identity for a media file, used to rehydrate URLs after a reload. */
const mediaPath = (file: ArchiveFile) => file.webkitRelativePath || file.name;
/**
* Decompress an `.xz` metadata sidecar.
*
* Read fully into memory first rather than handing the live HTTP body to
* the decompressor. These sidecars are a few KB, so buffering costs
* nothing, and streaming was actively harmful: the decompressor stops
* reading at the end of the xz member, leaving the response body neither
* drained nor cancelled. Across a couple of hundred sidecars that exhausts
* the connection pool and every later fetch fails with "Failed to fetch" —
* which silently cost Instaloader archives their captions, story flags and
* profile metadata, since all of it lives in these files.
*/
const parseXZFile = async (file: ArchiveFile) => {
try {
const stream = new XzReadableStream(file.stream());
const response = new Response(stream);
return await response.json();
const compressed = await file.arrayBuffer();
const stream = new XzReadableStream(new Blob([compressed]).stream());
return await new Response(stream).json();
} catch (e) { console.error(`[Scanner] XZ Parse Error:`, file.name, e); return null; }
};
-10
View File
@@ -14,7 +14,6 @@ export class LocalArchiveFile implements ArchiveFile {
get size() { return this.file.size; }
text() { return this.file.text(); }
arrayBuffer() { return this.file.arrayBuffer(); }
stream() { return this.file.stream(); }
/**
* A blob: URL backed directly by the on-disk File.
@@ -55,15 +54,6 @@ export class RemoteArchiveFile implements ArchiveFile {
const res = await fetch(this.url);
return res.arrayBuffer();
}
stream() {
const transform = new TransformStream();
fetch(this.url).then(res => {
if (res.body) res.body.pipeTo(transform.writable);
else transform.writable.getWriter().close();
});
return transform.readable;
}
createObjectUrl() {
return this.url;
}
+34
View File
@@ -0,0 +1,34 @@
import { describe, expect, it } from 'vitest';
import { isSystemDirectory } from './archive-index';
describe('isSystemDirectory', () => {
it.each([
['@eaDir', 'Synology thumbnail/index metadata, written inside every folder'],
['@tmp', 'Synology scratch'],
['.sync', 'Resilio state'],
['.DS_Store', 'macOS'],
['#recycle', 'Synology deletions'],
['#snapshot', 'Synology snapshots'],
])('skips %s (%s)', name => {
expect(isSystemDirectory(name)).toBe(true);
});
it.each([
'4utumn07',
'4utumn07 - reels',
'story - dawn_petal',
'story highlights - official_band - A.B.C',
'story highlights - theoldlyricmuseinsta - 💙1999-2005 era',
'Heejin_Bubble heejinmedia',
'gallery-dl',
'posts',
])('keeps %s', name => {
expect(isSystemDirectory(name)).toBe(false);
});
it('does not treat a leading underscore as a system directory', () => {
// `_gemini-plans` is filtered separately at the archive root only; nothing
// below the root should be excluded just for starting with an underscore.
expect(isSystemDirectory('_gemini-plans')).toBe(false);
});
});
+17 -1
View File
@@ -33,6 +33,21 @@ interface DirIndex {
const MEDIA_RE = /\.(jpg|jpeg|png|webp|gif|bmp|tiff|mp4|webm|ogv|mov)$/i;
const STAT_CONCURRENCY = 16;
/**
* Directories the walk must never descend into.
*
* NAS filesystems scatter sidecar metadata *inside* every folder, not just at
* the share root: Synology writes `@eaDir` (thumbnails and indexing data),
* `#recycle` holds deletions, and `.sync` is Resilio's state. Indexing those
* would count NAS thumbnails as archive media and spend a stat on each one —
* measured on a real share, `@eaDir` accounted for 12,516 of 123,023 files.
*
* The archive root is already filtered by prefix; this is the same rule applied
* at every level below it.
*/
export const isSystemDirectory = (name: string): boolean =>
name.startsWith('@') || name.startsWith('.') || name === '#recycle' || name === '#snapshot';
export class ArchiveIndex {
private dirs = new Map<string, DirIndex>();
private inFlight = new Map<string, Promise<DirIndex>>();
@@ -43,7 +58,7 @@ export class ArchiveIndex {
/** Visible (non-system) directories at the archive root. */
private listRootDirs(): string[] {
return fs.readdirSync(this.archivesDir, { withFileTypes: true })
.filter(e => e.isDirectory() && !/^[.@_]/.test(e.name))
.filter(e => e.isDirectory() && !isSystemDirectory(e.name) && !e.name.startsWith('_'))
.map(e => e.name);
}
@@ -69,6 +84,7 @@ export class ArchiveIndex {
return out;
}
for (const entry of entries) {
if (isSystemDirectory(entry.name)) continue;
const rel = base ? `${base}/${entry.name}` : entry.name;
if (entry.isDirectory()) out = out.concat(this.walk(path.join(absDir, entry.name), rel));
else if (entry.isFile()) out.push(rel);
-1
View File
@@ -50,7 +50,6 @@ export interface ArchiveFile {
size: number;
text(): Promise<string>;
arrayBuffer(): Promise<ArrayBuffer>;
stream(): ReadableStream<Uint8Array>;
url?: string;
/**
* A URL pointing at this file's contents. Local files mint a disk-backed