July 31, 2026
AI-generated report (Claude, Anthropic) — iteratively fact-checked against source documents but may contain errors. Verify claims against linked EFTA sources before citing. No affiliation with Anthropic.

Where Are the 3.5 Million Epstein Pages? — A Full Forensic Page-Count Audit

Every number in this post is reproducible against a public database. Links and queries at the bottom.

The Department of Justice says it released "nearly 3.5 million pages" of Epstein files. The public PDF production numbers its own pages, and the sequence ends at 2,858,498. My archive holds 2,858,311 of those numbered pages, and every broader public collection I've examined comes in below 3.5 million. That doesn't prove 640,000 pages vanished — it proves something narrower and harder for the Department to answer: it has never explained what it counted to reach the higher number.

The "3.5 million" figure has bothered me like "eats, shoots and leaves" bothers a copy editor — for about six months now, and I've had a few questions about it, so let's get this straight.

On January 30, 2026, the Department of Justice published a press release titled — this is the actual title — "Department of Justice Publishes 3.5 Million Responsive Pages in Compliance with the Epstein Files Transparency Act." The body says the day's publication added "over 3 million additional pages," including "more than 2,000 videos and 180,000 images," and that "combined with prior releases, this makes the total production nearly 3.5 million pages."

The same day, a letter to Congress signed by both Attorney General Pam Bondi and Deputy AG Todd Blanche used the identical phrase, and Blanche described the review pile at a press conference as "two Eiffel Towers of pages" (that phrase was about the ~6 million pages reviewed, not the 3.5 million released — a distinction that turns out to matter). (If you remember a "February" number, that's Bondi's Feb 14 Section 3 report — I checked it, and it contains no "3.5 million" figure at all.)

Before the arithmetic, here are the numbers that get thrown around — and what each one actually refers to. Keep this scoreboard in mind; the whole dispute is about which of these "3.5 million" was supposed to be:

The number What it actually counts
6+ million Pages Dep. AG Blanche said the department reviewed
~3.5 million Pages the DOJ says it produced, cumulatively — the headline
2,858,498 The most pages the public Bates sequence can hold (its final stamp)
~2.86 million Uniquely numbered pages I actually hold
~2.7 million Pages CBS found publicly available at its latest recount
~1.4 million Documents/files, however you choose to count them

And the government never gave enough precision to check its own math: the statutory-deadline release weeks earlier was, per DOJ's own letter to a federal court, about 12,285 documents totaling ~125,575 pages. To reach a cumulative 3.5 million, the January 30 batch would have to run about 3,374,425 pages — technically consistent with "over 3 million additional pages," but the DOJ never itemized it closely enough for anyone outside the building to reproduce the total. So the origin is clean in wording and says one thing — pages, cumulative, self-reported, never independently verified — but you cannot rebuild the number from anything it published. Then the figure went feral. Watch the unit mutate as it moves from the DOJ's mouth into the headlines:

Outlet Date What they printed Unit
DOJ press release + Bondi/Blanche letter Jan 30 "nearly 3.5 million pages" — cumulative pages ✓
NBC (body) Jan 30 "more than 3.5 million pages" — the cumulative total (its own Nightly News video said "3 million") pages ✓
NPR, AP, and most wire coverage Jan 30 "3 million pages" — the day's tranche pages ✓
Axios (headline) Jan 30 "release of 3.5 million records" records ✗
Washington Post (headline) Jan 30 "3 million more documents" documents ✗
Fox News Jan 30 "3 million new Epstein documents" documents ✗

That drift is not pedantry. At the ratio I measure across the archive — about two pages per document — "3.5 million documents" describes something twice as large as "3.5 million pages." A headline that swaps the noun roughly doubles the government's own claim, and millions of readers now carry the bigger number.

It is not correct. It isn't correct if you count documents. It isn't correct if you count pages. It isn't even correct if you count files on their own website. And you don't need to take my word for any of this, because the release itself contains the receipt.

The release numbers its own pages

Every page in the numbered EFTA PDF production carries a Bates stamp — a sequential serial number, one per page, in the format EFTA00000001, EFTA00000002, and so on. Bates numbers are unique page identifiers used in legal productions: they make the numbered range auditable and let gaps be spotted — though a gap alone doesn't prove a page was removed.

The last document in the release is EFTA02858497. It is two pages long. The final stamp in the entire series is therefore EFTA02858498.

That's the ceiling. One stamp per page; the numbering ends at 2,858,498. The EFTA production cannot contain more than 2,858,498 uniquely numbered pages, because the DOJ's own numbering system says so.

And it's not as if the series is sparse. I hold 2,858,311 of the 2,858,498 possible stamps — 99.99%. The EFTA Bates numbers don't stop at the numbered Data Sets, either: the same continuous series runs straight through the DOJ's Court Records productions — United States v. Maxwell, U.S. Virgin Islands v. JPMorgan, Epstein v. Rothstein, and dozens of the Doe dockets — which carry stamps above ~EFTA02731783 and which I hold as well (another ~126,000 pages across ~12,000 documents). One numbering scheme, spanning every disclosure directory, ending at 2,858,498.

There's a timeline wrinkle worth flagging, because it makes the headline stranger still. On January 30 — the day of the "nearly 3.5 million pages" announcement — the very last page stamp in the entire production was EFTA02731789: about 2.73 million pages. Everything numbered higher was added afterward, with no announcement. The DOJ has since quietly posted the court-record documents that carry the top of the range — roughly 126,000 more pages — pushing today's last stamp up to EFTA02858498 (about 2.86 million). One tranche of 23 documents appeared around March 5 and included some of the FBI interviews NPR had flagged as missing — I traced those Bates numbers with NPR's Stephen Fowler the night they showed up. So the "nearly 3.5 million" figure was announced over a pile that, that day, only reached 2.73 million — and even now, counting everything the Department has quietly added since, it stops at 2.86.

Against the DOJ's own file list — the manifest that names every Bates number in the numbered Data Sets — there are exactly 23 files in Data Set 9 that I don't have. And they turn out not to be missing so much as blank. The manifest lists all 23, but the files the Department actually serves are empty. Pull them straight from justice.gov and twenty-one come back with nothing in them — zero bytes or a broken stub — and the last two are blank "No Images Produced" placeholder pages, stamped with their Bates number and nothing else. A real document sitting right beside them downloads as a normal scan, and an independent mirror archived these same 23 as empty back in February — so it isn't my copy that's short. These 23 aren't 23 ordinary documents missing from my archive: twenty return empty files, one returns a broken HTML fragment, and two return one-page "No Images Produced" placeholders for native material (a video, a spreadsheet). They are 23 blank or non-substantive file slots, not 23 missing ordinary documents — and whatever happened to those underlying native files, these entries cannot explain a discrepancy of roughly 642,000 pages. The ceiling holds regardless: one stamp per page, highest stamp 2,858,498, so the production cannot contain more than 2,858,498 uniquely numbered pages — and I hold all but a rounding error of them. That remainder is exactly 187 stamps: 21 empty or broken entries, 2 placeholder pages, and 164 Bates numbers never assigned in the sparser court-record ranges — 21 + 2 + 164 = 187. Blank slots and unused numbers, not missing documents.

Count it any way you like

Method Count vs. 3.5M
Documents / files, all sources 1,425,565 2.46× this, if read as documents
Pages, EFTA series 2,858,311 641,689 short
Bates ceiling (DOJ's own numbering) 2,858,498 641,502 short
Pages, all sources combined 2,923,763 576,237 short

All-sources pages = EFTA + House Oversight + DOJ-OGR + FBI Vault + estate productions.

The same DOJ release counted four ways — pages, records, documents, files — none reaching 3.5 million

If "3.5 million" means documents, it's off by more than 2 million — the release contains ~1.4M documents. If it means pages — already a generous reading of "documents" — it's still more than half a million short, and you can't rescue it with the other productions. Throw in every adjacent pile — the House Oversight Committee's ~30,500-page release, the DOJ's separate ~33,300-page Oversight production, the flight-log records, the FBI Vault FOIA releases, the USVI estate materials, all of it — and you add about 65,000 pages, not the 642,000 you'd need to reach the headline. The entire non-EFTA world, put together, is barely a tenth of the gap. The whole public collection, counted together, comes to 2.92 million — still 576,000 short.

Even the people who printed the release couldn't make it reach 3.5 million — and I know, because I drove to New York to check. The traveling "Epstein reading room" — a pop-up by the independent Institute for Primary Facts, not a DOJ facility — is run by David Garrett, and I did the arithmetic with him in person. Every volume is bound at 800 pages; every spine is stamped "VOLUME X / 3,437"; I watched the last one, No. 3,436, being repaired after it was damaged in the move. 3,437 × 800 = 2,749,600 — and that's the most it can hold, since a volume can only run short of 800, never over. That won't quite fit the whole numbered production (which stamps to 2,858,498), and it isn't meant to: it's essentially the twelve Data Sets, printed — the EFTA release minus its ~126,000 court-record pages, about 2.73 million, which slides in just under the shelf capacity. Garrett confirmed the shape of it to my face: a full wall of bound paper, sitting right about where CBS's independent recount lands, near 2.7 — and most of a million pages short of the headline taped above it.

And the coverage printed the contradiction itself. Al Jazeera photographed this New York installation — 3,437 volumes — under a headline reading "3.5 million pages." The Washington Examiner covered the DC version — "3,500 volumes… each with 800 pages," which multiplies to 2.8 million — under a headline calling them "3.5 million documents." The arithmetic and the headline contradict each other inside a single article. Do people not multiply anymore?

Shelves at the Trump–Epstein Memorial Reading Room, each spine stamped

Volumes at the Institute for Primary Facts reading room — "THE PARTIALLY REDACTED EPSTEIN FILES," every spine numbered "of 3,437." Photographed by the author, May 11, 2026.

And here's the part that should have ended the "3.5 million" figure months ago: the one newsroom that actually recounted the whole library got a smaller number than I did. CBS News built software to crawl the DOJ's Epstein library file by file. Their continuously-updated tally: the Department "currently makes public about 2.7 million pages… a number below the Department's initial claim of 3 million" — CBS's benchmark there is the DOJ's day-one figure, so its recount lands under even the single-day claim, never mind the 3.5-million cumulative one. They also found the DOJ had quietly removed more than 47,000 files (about 65,500 pages) after release, and taken away the ability to download the collection in bulk. And here's the tell — the thing that made me sit up. When CBS reported its publicly observable total at about 2.7 million, the DOJ called their analysis "fundamentally flawed" — while in the same breath confirming that "more than 47,000 files remain offline for further review," closely corroborating CBS's estimate of how many files were offline even as it disputed the broader page-count analysis. Sit with that: the Department attacked the recount that came closest to matching its own admission, using a file count that corroborates it. Transparency has not exactly been the name of the game.

The independent archives land in the same band. Jmail's public extraction currently reports 2,474,242 pages across about 1.41 million files — its file counter reads 1,412,250 and its document counter 1,401,320, the ~11,000 gap turning on whether attachments are counted. My archive holds 2,858,311 stamped pages — the highest of the independent counts, and still 640,000 short of 3.5 million. Every independent page count I found, using different methods on different dates, lands somewhere between 2.5 and 2.9 million pages. I found no published recount from outside the Department that reaches 3.5 million publicly accessible pages — and the Department removed its bulk-download option, making a fresh independent recount substantially harder.

Independent page counts cluster at 2.5–2.9 million; only the DOJ's own claim sits above

What about all the other archives?

The major public archives I examined generally report roughly 1.4 million file-like records and between about 2.5 and 2.9 million pages — depending on scope, date and extraction method — and nobody's count reaches 3.5M without changing what's being counted:

Archive What they report Unit Sources
epstein-data.com (mine) 1,425,565 docs · 2,923,763 pages files / Bates pages DOJ DS1–12 + HOC + DOJ-OGR + FBI Vault + estate
tommycarstensen.com 1,380,987 files files DOJ numbered Data Sets 1–12 only
epsteinexposed.com 2,147,129 "documents" database import records DOJ (via archive.org mirror) + HOC + CourtListener + FBI Vault
jmail.world 1,412,250 files · 2,474,242 pages (per its /drive API) files / pages DOJ and related public document releases (scope as defined by JDrive)
yung-megafone (GitHub) DS9 reconstruction: 531,282 files in a published SHA-256 manifest (+2,324 native media) hash-verified files DOJ links + Internet Archive mirrors
DOJ itself "nearly 3.5 million pages" — no file count, ever pages

A few notes on that table:

One clarification, in the interest of counting honestly: EpsteinExposed is not wholly independent of my work — its early database incorporated material from my processed collection, so its EFTA file total (1,380,935) sitting essentially on top of mine — 1,380,937, my count of the twelve numbered Data Sets — is a shared source, not a second opinion. Its current integrity dashboard attributes its SHA-256 inventory and deletion-tracking to separate community projects, so I treat it as useful corroboration of file availability and change over time, not as a clean independent page recount. The genuinely independent checks are the ones that matter here, and they all point the same way: CBS News, working from the raw DOJ files, counted about 2.7 million pages; Jmail's separate extraction got about 2.47 million; the printed reading room holds about 2.8 million. Different teams, different methods, the raw government source — every one of them below three million, not one of them near 3.5.

Where the 2.9M pages and 1.4M documents sit, by data set — three data sets carry 92%

The release isn't even stable

While we're counting, the collection on DOJ's own website keeps shifting under our feet:

That last point deserves emphasis, because it's the comparison that should really sting — sharper than 2.9 against 3.5. Deputy Attorney General Todd Blanche has said the department reviewed more than 6 million total pages — a figure CBS reports and notes makes the public release "less than half of the total." It has made about 2.9 million public. What the Department has never published is the bridge between those two numbers: how many pages were removed as duplicates or non-responsive, how many were withheld under specified legal grounds, and how many were actually posted. A transparent accounting would have supplied all four. The press release instead offered "nearly 3.5 million pages" — a figure that matches neither the pile it reviewed nor the production it released.

6 million pages reviewed, 2.9 million public — the 3.5M claim sits between

A judge's preliminary finding — not just a counting dispute

The gap between "identified" and "released" stopped being a rhetorical point in June 2026. In granting preliminary relief in Phang v. Blanche, U.S. District Judge Emmet Sullivan concluded that the plaintiff was likely to prevail on a limited set of claims, noting that the DOJ had not substantively answered them and had therefore, for purposes of the motion, "conceded that he is in violation of the Act." He ordered the Department either to release specified material — including "at least eight email exchanges with Mr. Epstein regarding a 'torture video'" — without the disputed redactions, or explain why they should remain; after the DOJ responded, he ordered the unredacted documents submitted for private, in-camera judicial review rather than immediate public release. The litigation is ongoing. Reporting notes the Department still has not released some 2.5 million pages, and the DOJ has fought a proposed $1,000-a-day contempt sanction over it.

House Judiciary Democrats made a similar allegation in writing. A March 19 letter from Reps. Raskin, García and Jayapal states the DOJ "has withheld 3 million pages" and "denied Members any opportunity to verify DOJ's claims that millions of pages withheld from disclosure are merely duplicates — a dubious assertion given so many files known to exist have yet to be revealed." They estimate it would take 7.5 years to check the redactions on the partial set they were given.

So the "3.5 million" number sits inside a frame the government itself supplied: Deputy AG Blanche has said the department reviewed more than 6 million total pages, and the publicly observable production runs to roughly 2.7–2.9 million (a range the Department disputes at the low end). The rest it attributes to a mix of duplicates, nonresponsive and sealed material, victim-protection and privilege redactions, and technical or foreign-language files — and, in a preliminary ruling the DOJ disputes, a federal judge in Phang v. Blanche found the Attorney General — by then Acting AG Blanche — had conceded he is in violation of the Act, grounding that on Blanche's failure to rebut the plaintiff's arguments rather than any actual admission. It also removed tens of thousands of files after publication. In that context, a headline figure inflated past what was even released is not a rounding error. It is the friendly-sounding number placed over an accounting that never discloses how many pages were removed as duplicates or non-responsive, how many were legally withheld, and how many were actually posted.

So what does the number mean?

Here is precisely what I can and can't say:

What's documented: The EFTA production's own Bates numbering terminates at 2,858,498. The public collections included in my database total approximately 2.92 million pages across 1.43 million document records. Multiple independent archives land in the same band.

What's not documented: any accounting, from the Department, of how "nearly 3.5 million pages" was computed. The numbered production caps at 2,858,498 pages by the DOJ's own stamps — roughly 640,000 pages short of the claim. Candidate explanations, none confirmed: pre-deduplication counting (SDNY/SDFL duplicates were removed, per the press release); counting the ~180,000 images and ~2,000 videos as some internal multi-page equivalent; counting pages that were reviewed but withheld under the Act's exceptions. Each of those would mean the public number describes something other than what the public received — and in Phang v. Blanche the court found, at the preliminary-injunction stage and on a limited set of claims, that the Attorney General had conceded noncompliance for purposes of the motion. That is a preliminary finding the DOJ disputes, and it does not establish that the entire numerical gap is unlawful withholding.

Here's the most banal explanation, giving them more benefit of the doubt than is rightly due: Bondi was quoting an earlier internal estimate — a number from before a chunk of pages got deemed nonresponsive, duplicative, privileged, or sealed — and nobody corrected the headline once the review finished. I'd buy that, except for one thing: they haven't quietly left it uncorrected — they've defended it. When CBS recounted the library and landed at 2.7 million, the Department didn't say "our early estimate was rough"; it called the recount "fundamentally flawed." You don't go to the mat for a number you think was a loose internal guess. That's what makes it strange: the charitable version requires them to have simply forgotten to fix it — and they have spent months, on the record, not forgetting.

The question, then — and I'm framing this as a question, because that's what the evidence supports: what is the 3.5 million figure counting? Either the figure was wrong when it was announced, or there exists a substantial volume of material that was counted internally and never produced, or the Department used an undisclosed counting convention — duplicates, native files, media, later-offline files — that does not represent unique public pages. All three are worth someone in an oversight role asking about, under oath, with the stamp arithmetic in hand.

At the end of the day, we have what we have: a Bates series that numbers its own pages to 2,858,498, a complete public production of about 2.92 million even after you add every other source, and a press release claiming roughly 640,000 pages more than the numbering allows — a gap the adjacent material comes nowhere close to filling, and one the Department has never reconciled with the count anyone can actually run.

Methodology: what each number measures

Most of the confusion around "3.5 million" comes from sliding between measures that sound alike. Every figure in this report is one of these — defined, and reproducible against the public database:

Measure Definition Value
Document-metadata pages Sum of every document's stated page count — the primary page measure in this audit 2,923,763
Extracted page records Rows in the page table (one per extracted page) 2,924,310
EFTA pages held Document-metadata pages for the EFTA-numbered production 2,858,311
Bates-stamp ceiling The final Bates number in the series (last document EFTA02858497, two pages) 2,858,498
Documents, all sources Every document record 1,425,565
EFTA documents EFTA-numbered documents only 1,393,164

The 547-page reconciliation. Document-metadata pages (2,923,763) and extracted page records (2,924,310) differ by 547 — a 536 + 11 split. The larger part is the extraction layer: 309 native documents (spreadsheets, videos) are each produced by the DOJ as a single "No Images Produced" placeholder page, so their stated page count is 1, but their content extracts to many searchable records — 536 rows beyond the 309 placeholders. The remaining 11 page records belong to three documents absent from the metadata table. Those extra rows are search units, not additional public PDF pages. The document-metadata figure (2,923,763) is the "pages the DOJ produced" count, and is used throughout.

The 187-stamp gap. The Bates ceiling (2,858,498) minus EFTA pages held (2,858,311) leaves 187 stamps. Twenty-three are genuinely-blank Data Set 9 files (below); the remaining 164 are Bates numbers never assigned to any page in the sparser court-record ranges — unused numbers, not missing documents.

The 23 blank Data Set 9 files

Cross-referenced against the DOJ's own document manifest, exactly 23 files in Data Set 9 are absent from the archive. They are not withheld — they are blank at the source, and have been since release:

Bates numbers What justice.gov serves Count
EFTA00709804–709807, 00770595, 00823190–823192, 00823221, 00823319, 00877475, 00892252, 00901740, 00912980, 00919433–919434, 00932520–932523 0 bytes (empty file) 20
EFTA00774768 45-byte broken HTML fragment 1
EFTA01135215, EFTA01135708 2,433-byte "No Images Produced" placeholder page, generated 2026-03-04 2

Each was verified three independent ways: fetched directly from justice.gov through the age-verification gate (empty, while a neighbouring document downloads as a normal scan); the RollCall mirror, whose copies are dated 2 February 2026 and carry the MD5 hash of an empty file; and the Internet Archive, which has captured each URL daily since release day — 160+ captures apiece, every one empty. They are also absent from a complete independent mirror of Data Set 9, and no native file exists for them at any of ~30 tested extensions. The DOJ published 23 Bates numbers with no document behind them.

Reproduce every number

Every figure above is queryable against the public database. A few of the load-bearing ones:

-- Documents and document-metadata pages, all sources → 1,425,565 · 2,923,763
select count(*), sum(total_pages) from documents

-- EFTA production only → 1,393,164 · 2,858,311
select count(*), sum(total_pages) from documents where efta_number like 'EFTA%'

-- Extracted page records → 2,924,310
select count(*) from pages

Run them yourself at epstein-data.com/full_text_corpus. All counts captured 2026-07-31.

And the specific lists behind the reconciliations above are published as machine-readable receipts:


Receipts

Methodology note: all third-party counts (Jmail, EpsteinExposed, Tommy Carstensen) were captured on July 31, 2026 — these are live sites and will drift. My own counts exclude ~1,400 transcription placeholder entries our transcription software added (community transcript integration, July 2026); all figures are of DOJ-released material as produced. The 22k DDoSecrets-leaked emails are excluded throughout — leaked ≠ released.

Reader Notes

Ask about this report

Ask a question — the AI has the full report loaded and can also search the full corpus.