ergosteurandClaude Opus 5 a83da461c1 feat: stop enumerating a profile once it reaches what we already hold
The skip-archive suppresses downloads, which spends the CDN. It does nothing
about the listing pass, which spends `instagram.com` — the surface that
actually bans accounts — and that cost scales with how BIG a profile is, not
with how much of it is new. A 2275-post profile paid ~76 pages every run to
discover three new posts. Seeding saved the second full pass, never the first.

Measured from sidecar write times during today's run, free because the run was
paying for the listing anyway: three new posts took ~100s each, and the other
2272 were written in a single second — enumeration with nothing to show for it.

`--abort N` passes gallery-dl's `skip: abort:N`, stopping the extractor after N
consecutive already-archived files. Resuming a stopped run with `--abort 50`
enumerated 7 posts of cher_ryppo's 2151 and still caught every new one.

Three things make this safe, and all of them are load-bearing:

- N counts FILES, not posts, so it has to clear the largest already-held
  carousel — one post in this archive is 22 media. A reels tab needs 50 actual
  reels for the same threshold, since those are single-media.
- It applies to posts and reels only. Stories are always new, and highlight
  items are not ordered in a way that makes early abort safe.
- The REST listing is strictly reverse-chronological. Test case 16 claimed
  0ct0ber19 returns its 3 pinned posts out of date order; that is true of the
  web grid but not of this endpoint, measured today. Front-loaded old posts are
  the one thing that would trip abort before it reached anything new, so the
  correction is what licenses the feature rather than a footnote to it.

Default is 0 — walk everything — because aborting early stops noticing edited
carousels (test case 15), which only a full enumeration finds. Routine runs
want 50; a full sweep is still worth running occasionally.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UXfdJu7QhSJLr47K7koTDF
2026-08-20 14:31:33 -04:00
2026-08-17 13:58:15 -04:00
2026-08-17 13:58:15 -04:00

InstaArchive Viewer

A high-performance React PWA for browsing archived Instagram data with a native-feeling interface. Supports both official Instagram exports and Instaloader archives.

Features

  • Advanced Carousel: Seamless, zero-latency transitions between slides with intelligent preloading. Navigating between different posts is now near-instant thanks to inter-post background preloading.
  • High-Res Performance: Handles 50MP+ images effortlessly using a background Web Worker and a memory-safe serial processing queue.
  • Persistent Local Caching: Uses IndexedDB to store parsed archives and generated thumbnails. Local folders now load instantly from cache on return visits without needing to re-upload files.
  • Permalinks: State is synchronized with the URL, allowing you to share direct links to archives, tabs, or specific posts. Navigating back to the explorer cleans up URL parameters automatically.
  • Glassy Scanning UI: A refined, translucent white terminal experience with flicker-free, double-buffered dynamic blurred backgrounds.
  • PWA with Auto-Update: Fully offline-capable and installable. Clients automatically receive updates when a new version is deployed to the server.
  • Local Privacy: All processing is done client-side. Even when using the self-hosted version, your media is processed locally in your browser and never uploaded.
  • Smart Fallbacks: Automatically detects usernames from folder names and uses the oldest archive image as a profile picture if one is missing.
  • Customizable Grid: 1:1 or 3:4 aspect ratios with adjustable "bumps" for aesthetic alignment.
  • Story Viewer: Native-like story experience with segmented progress bars, auto-playback, and audio controls.
  • Navigation Protection: Intercepts accidental browser "Back" or "Refresh" actions to protect your current session.

Deployment

The easiest way to run InstaArchive is using Docker.

docker run -d \
  -p 3000:3000 \
  -v /path/to/your/archives:/archives:ro \
  ghcr.io/ergosteur/instaarchive-viewer:latest

Note for Linux/SELinux users: If you see "Permission Denied" in the logs, append ,z to your volume mount: -v /path/to/archives:/archives:ro,z

Docker Compose

Create a compose.yml file:

services:
  instaarchive:
    image: ghcr.io/ergosteur/instaarchive-viewer:latest
    ports:
      - "3000:3000"
    volumes:
      - ./archives:/archives:ro,z # ,z handles SELinux permissions

Troubleshooting Permissions

If the app shows "No Archives Found" and logs EACCES: permission denied:

  1. Check Directory Permissions: Ensure the archive folder is world-readable:
    chmod -R 755 /path/to/archives
    
  2. SELinux (Fedora/RHEL/CentOS): Use the :z flag in your volume mount as shown above.
  3. User Mapping: The container runs as the non-root node user (UID 1000). If your archives are readable only by another account, run as that user instead — the container needs to list the archive directory, so --x (traverse-only) permissions are not enough:
    docker run --user $(stat -c '%u:%g' /path/to/archives) ...
    
    In Compose:
    services:
      instaarchive:
        user: "1234:1234"   # a UID that can read your archives
    

Archive Index

On first start the server walks the archive root once and caches the result, keyed by directory mtime. This matters on network storage: for a 110k-file archive root, listing went from ~52s per request to ~0.1s. Mount a volume at /cache (or set CACHE_DIR) so the index survives restarts, otherwise it is rebuilt on every start.

Supported Archive Structure

Place your archive folders inside the mounted /archives directory. The directory name will be used as the account username.

Example Structure:

archives/
├── wanderlust_explorer/          # Instaloader format
│   ├── 2024-01-01_12-00-00_UTC.jpg
│   ├── 2024-01-01_12-00-00_UTC.json.xz
│   └── wanderlust_explorer_profile_pic.jpg
└── pixel_architect/             # Instagram Export format
    ├── 2023-12-25_pixel_architect - post_123.jpg
    ├── 2023-12-25_pixel_architect - post_123.json
    └── pixel_architect.jpg

Local Development

Prerequisites: Node.js (LTS recommended)

  1. Install dependencies: npm install
  2. Start dev server: npm run dev (Frontend on port 3000)
  3. Start local backend: npm run server (Optional, serves ./_sample-archives on port 3001)
  4. Build production: npm run build (Generates ./dist for frontend and ./dist-server for the API)
S
Description
No description provided
Readme
1.5 MiB
Languages
TypeScript 97.8%
CSS 1.6%
Dockerfile 0.5%
HTML 0.1%