docs: record why the account was suspended — verification, not fetching
The account was suspended on 2026-08-17 for "spam", during the session that built this tooling. Both fetching docs were confidently wrong about what the risk was, so both now carry the correction. docs/jdownloader.md said the ban vector is instagram.com requests, which is right, and implied that meant per-post metadata fetching, which is only part of it. The suspension came from read-only verification: automated browser scrolling to enumerate profile grids (~18 paginated loads per profile, done twice on one after a selector bug), repeated --simulate and -j passes over the same profiles, per-post /p/ fetches while testing filename formats, and an aborted sync that re-ran every listing pass before dying. None of that produced a file, and together it rivalled the real sync for request count. The rules that follow are in docs/gallery-dl.md: verify against the archive rather than the live site, count read-only work against the same budget, treat the first CDN 429 as the end of the session rather than a pacing knob, and cache probe_live so a restart does not re-enumerate everything. The warning order was CDN 429, then 400 on the highlights tray, then suspension. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -14,6 +14,72 @@ map 1:1 onto our sidecar directory layout (`posts`, `reels`, `stories`,
|
||||
`highlights`). Rolling our own would mean reimplementing the ban-sensitive part
|
||||
by hand.
|
||||
|
||||
## The account was suspended on 2026-08-17 — read this first
|
||||
|
||||
The account used for all of the below was suspended the same day this tooling
|
||||
was built, for "activity that doesn't follow our Community Standards on spam".
|
||||
The fetching was not the expensive part. **Verification was.**
|
||||
|
||||
What was actually spent against `instagram.com` in a few hours, from one
|
||||
session and one IP:
|
||||
|
||||
| activity | rough requests | downloaded |
|
||||
|---|---:|---|
|
||||
| enumerating a profile grid by scrolling it in an automated browser | ~18 pages | nothing |
|
||||
| the same profile again, after a bug in the scraping selector | ~18 pages | nothing |
|
||||
| a Reels tab enumerated the same way | ~9 pages | nothing |
|
||||
| full `-j` metadata dumps of one profile, twice | ~16 pages | nothing |
|
||||
| `--simulate` runs over the same profile, three times | ~24 pages | nothing |
|
||||
| single-post `/p/<code>/` fetches while testing filename formats | ~8 | a handful |
|
||||
| an aborted sync that re-ran every listing pass before dying | ~40 pages | ~270 MB |
|
||||
| the real sync, 24 sources across 6 profiles | ~150 pages | 2.2 GB |
|
||||
|
||||
The two rows that actually mattered to the archive are the last one and part of
|
||||
the second-to-last. **Everything above them produced no files at all**, and
|
||||
together they were a comparable number of requests.
|
||||
|
||||
The warnings arrived in this order and were each rationalised:
|
||||
|
||||
1. `429 Too Many Requests` from `scontent-*.cdninstagram.com`, losing two
|
||||
videos. Treated as a pacing problem — pacing was lowered and the run
|
||||
continued.
|
||||
2. `400 Bad Request` from `/api/v1/highlights/<id>/highlights_tray/`, on an
|
||||
endpoint that had worked hours earlier. Correctly read as a possible block;
|
||||
requests stopped.
|
||||
3. Suspension.
|
||||
|
||||
**Treat the first CDN 429 as a stop signal for the session, not a tuning
|
||||
parameter.** It is the tolerant surface complaining; if that surface is
|
||||
complaining, the rate-limited one has been unhappy for a while.
|
||||
|
||||
### Rules that follow from this
|
||||
|
||||
- **Count verification requests against the same budget as fetching.** A
|
||||
`--simulate`, a `-j` dump and a browser scroll all hit `instagram.com` and
|
||||
download nothing. Being read-only does not make them free; it makes them
|
||||
invisible, which is worse.
|
||||
- **Never enumerate the live site with an automated browser.** Scrolling a
|
||||
214-post grid is ~18 paginated GraphQL loads at machine speed with no dwell
|
||||
time between them. It is the most obviously non-human thing in this whole
|
||||
document, and it was done here twice on one profile.
|
||||
- **Verify against the archive, not against Instagram.** Every naming, dating
|
||||
and classification question answered in this file could have been answered
|
||||
from files already on disk plus a single listing pass.
|
||||
- **`probe_live` is not cached, so every restart re-enumerates everything.**
|
||||
The aborted run cost a full duplicate set of listing passes for five
|
||||
profiles. Cache probe output to disk before running anything twice.
|
||||
- **Budget per session, not per command.** Nothing in the tooling knows what
|
||||
the last command spent.
|
||||
|
||||
### For a replacement account
|
||||
|
||||
- Let it exist and be used normally for a while before pointing any tool at it.
|
||||
- Keep the cookie on one machine and one public IP, as before.
|
||||
- Start with a single small profile and stop for the day afterwards.
|
||||
- Prefer Instagram's own "Download a copy" export where possible: it is
|
||||
first-party, costs no scraping requests, and carries the metadata this whole
|
||||
document works around not having.
|
||||
|
||||
## The safety model — read this before changing any option
|
||||
|
||||
The ban vector is **requests to `instagram.com`**, not bandwidth. See
|
||||
|
||||
+17
-1
@@ -18,7 +18,23 @@ Concretely, from this archive: `rivvsofficial` has 188 post-metadata files, so
|
||||
backfilling it cost 188 API requests for one 605-file profile. That's the ban
|
||||
vector. Downloading the 238 photos was never the problem.
|
||||
|
||||
Instaloader got this account banned once. JDownloader with throttling has not.
|
||||
Instaloader got this account banned once. JDownloader with throttling did not.
|
||||
|
||||
> **The account was suspended anyway, on 2026-08-17, for "spam".** Not by
|
||||
> JDownloader, and not by downloading. It was suspended during a day of
|
||||
> *building and verifying* the gallery-dl replacement — automated browser
|
||||
> scrolling to enumerate profile grids, repeated `--simulate` and `-j` metadata
|
||||
> passes, and one aborted sync that re-ran every listing pass before dying.
|
||||
>
|
||||
> The framing above is right about which surface is dangerous and wrong about
|
||||
> what reaches it. **Every read of `instagram.com` counts, including the ones
|
||||
> that download nothing** — and read-only work is easy not to count precisely
|
||||
> because it leaves no files behind. See the post-mortem at the top of
|
||||
> `docs/gallery-dl.md`.
|
||||
>
|
||||
> The rule that would have prevented it: *verify against the archive, never
|
||||
> against the live site*, and treat the first CDN `429` as the end of the
|
||||
> session rather than a pacing knob.
|
||||
|
||||
### What the metadata gap actually costs
|
||||
|
||||
|
||||
Reference in New Issue
Block a user