{"database": "search_index", "table": "report_index", "rows": [["audits/PAGE_COUNT_AUDIT", "Where Are the 3.5 Million Epstein Pages? \u2014 A Full Forensic Page-Count Audit", "audits", "PAGE_COUNT_AUDIT.md", "/reports/audits/PAGE_COUNT_AUDIT.html", 10, "[\"EFTA00000001\", \"EFTA00000002\", \"EFTA00003434\", \"EFTA00709804\", \"EFTA00774768\", \"EFTA01135215\", \"EFTA01135708\", \"EFTA02731783\", \"EFTA02858497\", \"EFTA02858498\"]", null, "", "Where Are the 3.5 Million Epstein Pages? \u2014 A Full Forensic Page-Count Audit Every number in this post is reproducible against a public database. Links and queries at the bottom. The Department of Justice says it released \"nearly 3.5 million pages\" of Epstein files. The public PDF production numbers its own pages, and the sequence ends at 2,858,498. My archive holds 2,858,311 of those numbered pages, and every broader public collection I've examined comes in below 3.5 million. That doesn't prove 640,000 pages vanished \u2014 it proves something narrower and harder for the Department to answer: it has never explained what it counted to reach the higher number. The \"3.5 million\" figure has bothered me like \"eats, shoots and leaves\" bothers a copy editor \u2014 for about six months now, and I've had a few questions about it, so let's get this straight. On January 30, 2026, the Department of Justice published a press release titled \u2014 this is the actual title \u2014 \"Department of Justice Publishes 3.5 Million Responsive Pages in Compliance with the Epstein Files Transparency Act.\" The body says the day's publication added \"over 3 million additional pages,\" including \"more than 2,000 videos and 180,000 images,\" and that \"combined with prior releases, this makes the total production nearly 3.5 million pages.\" The same day, a letter to Congress signed by both Attorney General Pam Bondi and Deputy AG Todd Blanche used the identical phrase, and Blanche described the review pile at a press conference as \"two Eiffel Towers of pages\" (that phrase was about the ~6 million pages reviewed, not the 3.5 million released \u2014 a distinction that turns out to matter). (If you remember a \"February\" number, that's Bondi's Feb 14 Section 3 report \u2014 I checked it, and it contains no \"3.5 million\" figure at all.) Before the arithmetic, here are the numbers that get thrown around \u2014 and what each one actually refers to. Keep this scoreboard in mind; the whole dispute is about which of these \"3.5 million\" was supposed to be: The number What it actually counts --- --- 6+ million Pages Dep. AG Blanche said the department reviewed ~3.5 million Pages the DOJ says it produced, cumulatively \u2014 the headline 2,858,498 The most pages the public Bates sequence can hold (its final stamp) ~2.86 million Uniquely numbered pages I actually hold ~2.7 million Pages CBS found publicly available at its latest recount ~1.4 million Documents/files, however you choose to count them And the government never gave enough precision to check its own math: the statutory-deadline release weeks earlier was, per DOJ's own letter to a federal court, about 12,285 documents totaling ~125,575 pages. To reach a cumulative 3.5 million, the January 30 batch would have to run about 3,374,425 pages \u2014 technically consistent with \"over 3 million additional pages,\" but the DOJ never itemized it closely enough for anyone outside the building to reproduce the total. So the origin is clean in wording and says one thing \u2014 pages, cumulative, self-reported, never independently verified \u2014 but you cannot rebuild the number from anything it published. Then the figure went feral. Watch the unit mutate as it moves from the DOJ's mouth into the headlines: Outlet Date What they printed Unit --- --- --- --- DOJ press release + Bondi/Blanche letter Jan 30 \"nearly 3.5 million pages\" \u2014 cumulative pages \u2713 NBC (body) Jan 30 \"more than 3.5 million pages\" \u2014 the cumulative total (its own Nightly News video said \"3 million\") pages \u2713 NPR, AP, and most wire coverage Jan 30 \"3 million pages\" \u2014 the day's tranche pages \u2713 Axios (headline) Jan 30 \"release of 3.5 million records\" records \u2717 Washington Post (headline) Jan 30 \"3 million more documents\" documents \u2717 Fox News Jan 30 \"3 million new Epstein documents\" documents \u2717 That drift is not pedantry. At the ratio I measure across the archive \u2014 about two pages per document \u2014 \"3.5 million documents\" describes something twice as large as \"3.5 million pages.\" A headline that swaps the noun roughly doubles the government's own claim, and millions of readers now carry the bigger number. It is not correct. It isn't correct if you count documents. It isn't correct if you count pages. It isn't even correct if you count files on their own website. And you don't need to take my word for any of this, because the release itself contains the receipt. The release numbers its own pages Every page in the numbered EFTA PDF production carries a Bates stamp \u2014 a sequential serial number, one per page, in the format EFTA00000001, EFTA00000002, and so on. Bates numbers are unique page identifiers used in legal productions: they make the numbered range auditable and let gaps be spotted \u2014 though a gap alone doesn't prove a page was removed. The last document in the release is EFTA02858497. It is two pages long. The final stamp in the entire series is therefore EFTA02858498. That's the ceiling. One stamp per page; the numbering ends at 2,858,498. The EFTA production cannot contain more than 2,858,498 uniquely numbered pages, because the DOJ's own numbering system says so. And it's not as if the series is sparse. I hold 2,858,311 of the 2,858,498 possible stamps \u2014 99.99%. The EFTA Bates numbers don't stop at the numbered Data Sets, either: the same continuous series runs straight through the DOJ's Court Records productions \u2014 United States v. Maxwell, U.S. Virgin Islands v. JPMorgan, Epstein v. Rothstein, and dozens of the Doe dockets \u2014 which carry stamps above ~EFTA02731783 and which I hold as well (another ~126,000 pages across ~12,000 documents). One numbering scheme, spanning every disclosure directory, ending at 2,858,498. There's a timeline wrinkle worth flagging, because it makes the headline stranger still. On January 30 \u2014 the day of the \"nearly 3.5 million pages\" announcement \u2014 the very last page stamp in the entire production was EFTA02731789: about 2.73 million pages. Everything numbered higher was added afterward, with no announcement. The DOJ has since quietly posted the court-record documents that carry the top of the range \u2014 roughly 126,000 more pages \u2014 pushing today's last stamp up to EFTA02858498 (about 2.86 million). One tranche of 23 documents appeared around March 5 and included some of the FBI interviews NPR had flagged as missing \u2014 I traced those Bates numbers with NPR's Stephen Fowler the night they showed up. So the \"nearly 3.5 million\" figure was announced over a pile that, that day, only reached 2.73 million \u2014 and even now, counting everything the Department has quietly added since, it stops at 2.86. Against the DOJ's own file list \u2014 the manifest that names every Bates number in the numbered Data Sets \u2014 there are exactly 23 files in Data Set 9 that I don't have. And they turn out not to be *missing* so much as *blank*. The manifest lists all 23, but the files the Department actually serves are empty. Pull them straight from justice.gov and twenty-one come back with nothing in them \u2014 zero bytes or a broken stub \u2014 and the last two are blank \"No Images Produced\" placeholder pages, stamped with their Bates number and nothing else. A real document sitting right beside them downloads as a normal scan, and an independent mirror archived these same 23 as empty back in February \u2014 so it isn't my copy that's short. These 23 aren't 23 ordinary documents missing from my archive: twenty return empty files, one returns a broken HTML fragment, and two return one-page \"No Images Produced\" placeholders for native material (a video, a spreadsheet). They are 23 blank or non-substantive file slots, not 23 missing ordinary documents \u2014 and whatever happened to those underlying native files, these entries cannot explain a discrepancy of roughly 642,000 pages. The ceiling holds regardless: one stamp per page, highest stamp 2,858,498, so the production cannot contain more than 2,858,498 uniquely numbered pages \u2014 and I hold all but a rounding error of them. That remainder is exactly 187 stamps: 21 empty or broken entries, 2 placeholder pages, and 164 Bates numbers never assigned in the sparser court-record ranges \u2014 21 + 2 + 164 = 187. Blank slots and unused numbers, not missing documents. Count it any way you like Method Count vs. 3.5M --- --- --- Documents / files, all sources 1,425,565 2.46\u00d7 this, if read as documents Pages, EFTA series 2,858,311 641,689 short Bates ceiling (DOJ's own numbering) 2,858,498 641,502 short Pages, all sources combined 2,923,763 576,237 short All-sources pages = EFTA + House Oversight + DOJ-OGR + FBI Vault + estate productions. If \"3.5 million\" means documents, it's off by more than 2 million \u2014 the release contains ~1.4M documents. If it means pages \u2014 already a generous reading of \"documents\" \u2014 it's still more than half a million short, and you can't rescue it with the other productions. Throw in every adjacent pile \u2014 the House Oversight Committee's ~30,500-page release, the DOJ's separate ~33,300-page Oversight production, the flight-log records, the FBI Vault FOIA releases, the USVI estate materials, all of it \u2014 and you add about 65,000 pages, not the 642,000 you'd need to reach the headline. The entire non-EFTA world, put together, is barely a tenth of the gap. The whole public collection, counted together, comes to 2.92 million \u2014 still 576,000 short. Even the people who printed the release couldn't make it reach 3.5 million \u2014 and I know, because I drove to New York to check. The traveling \"Epstein reading room\" \u2014 a pop-up by the independent Institute for Primary Facts, not a DOJ facility \u2014 is run by David Garrett, and I did the arithmetic with him in person. Every volume is bound at 800 pages; every spine is stamped \"VOLUME X / 3,437\"; I watched the last one, No. 3,436, being repaired after it was damaged in the move. 3,437 \u00d7 800 = 2,749,600 \u2014 and that's the most it can hold, since a volume can only run short of 800, never over. That won't quite fit the whole numbered production (which stamps to 2,858,498), and it isn't meant to: it's essentially the twelve Data Sets, printed \u2014 the EFTA release minus its ~126,000 court-record pages, about 2.73 million, which slides in just under the shelf capacity. Garrett confirmed the shape of it to my face: a full wall of bound paper, sitting right about where CBS's independent recount lands, near 2.7 \u2014 and most of a million pages short of the headline taped above it. And the coverage printed the contradiction itself. Al Jazeera photographed this New York installation \u2014 3,437 volumes \u2014 under a headline reading \"3.5 million pages.\" The Washington Examiner covered the DC version \u2014 \"3,500 volumes\u2026 each with 800 pages,\" which multiplies to 2.8 million \u2014 under a headline calling them \"3.5 million documents.\" The arithmetic and the headline contradict each other inside a single article. Do people not multiply anymore? Volumes at the Institute for Primary Facts reading room \u2014 \"THE PARTIALLY REDACTED EPSTEIN FILES,\" every spine numbered \"of 3,437.\" Photographed by the author, May 11, 2026. And here's the part that should have ended the \"3.5 million\" figure months ago: the one newsroom that actually recounted the whole library got a smaller number than I did. CBS News built software to crawl the DOJ's Epstein library file by file. Their continuously-updated tally: the Department \"currently makes public about 2.7 million pages\u2026 a number below the Department's initial claim of 3 million\" \u2014 CBS's benchmark there is the DOJ's day-one figure, so its recount lands under even the single-day claim, never mind the 3.5-million cumulative one. They also found the DOJ had quietly removed more than 47,000 files (about 65,500 pages) after release, and taken away the ability to download the collection in bulk. And here's the tell \u2014 the thing that made me sit up. When CBS reported its publicly observable total at about 2.7 million, the DOJ called their analysis \"fundamentally flawed\" \u2014 while in the same breath confirming that \"more than 47,000 files remain offline for further review,\" closely corroborating CBS's estimate of how many files were offline even as it disputed the broader page-count analysis. Sit with that: the Department attacked the recount that came closest to matching its own admission, using a file count that corroborates it. Transparency has not exactly been the name of the game. The independent archives land in the same band. Jmail's public extraction currently reports 2,474,242 pages across about 1.41 million files \u2014 its file counter reads 1,412,250 and its document counter 1,401,320, the ~11,000 gap turning on whether attachments are counted. My archive holds 2,858,311 stamped pages \u2014 the highest of the independent counts, and still 640,000 short of 3.5 million. Every independent page count I found, using different methods on different dates, lands somewhere between 2.5 and 2.9 million pages. I found no published recount from outside the Department that reaches 3.5 million publicly accessible pages \u2014 and the Department removed its bulk-download option, making a fresh independent recount substantially harder. What about all the other archives? The major public archives I examined generally report roughly 1.4 million file-like records and between about 2.5 and 2.9 million pages \u2014 depending on scope, date and extraction method \u2014 and nobody's count reaches 3.5M without changing what's being counted: Archive What they report Unit Sources --- --- --- --- epstein-data.com (mine) 1,425,565 docs \u00b7 2,923,763 pages files / Bates pages DOJ DS1\u201312 + HOC + DOJ-OGR + FBI Vault + estate tommycarstensen.com 1,380,987 files files DOJ numbered Data Sets 1\u201312 only epsteinexposed.com 2,147,129 \"documents\" database import records DOJ (via archive.org mirror) + HOC + CourtListener + FBI Vault jmail.world 1,412,250 files \u00b7 2,474,242 pages (per its /drive API) files / pages DOJ and related public document releases (scope as defined by JDrive) yung-megafone (GitHub) DS9 reconstruction: 531,282 files in a published SHA-256 manifest (+2,324 native media) hash-verified files DOJ links + Internet Archive mirrors DOJ itself \"nearly 3.5 million pages\" \u2014 no file count, ever pages \u2014 A few notes on that table: Tommy Carstensen reports 1,380,987 files in the twelve numbered Data Sets, and separately catalogs about 11,830 DOJ court-record files. Combine those and his total is 1,392,817 \u2014 within 347 of my 1,393,164 EFTA-numbered document records. The tiny remainder appears to be additional disclosure categories and database treatment, not a missing archive's worth of files. (His site also tracks dead links \u2014 files that now return \"not found\" \u2014 since the DOJ acknowledged taking tens of thousands of files offline after publication, so counts drift with when you crawled.) yung-megafone's Data Set 9 reconstruction is the strongest cross-check of the set, because it is hash-level rather than byte-level: the project publishes a SHA-256 manifest of every file it holds. For Data Set 9 that manifest lists 531,282 PDF documents (plus 2,324 native media files), with only 40 byte-identical duplicates. My own database independently holds 531,284 Data Set 9 documents \u2014 a match to within two files, reached by an entirely different method: my page-by-page extraction against their file checksums. Two independent inventories agreeing at the file level, on the single largest data set, is about as clean a corroboration as this material allows. EpsteinExposed's 2.15M \"documents\" is not a competing count of the DOJ release \u2014 their own open-source code shows it's a tally of database records \u2014 one per imported item \u2014 across DOJ plus court records plus congressional material, imported from an archive.org mirror. And it's not that they've added some other archive's worth of material: by their own sourcing page, the non-DOJ court content is tiny \u2014 a few hundred RECAP filings from CourtListener, a few thousand unsealed court documents, a few hundred congressional records. Their site separately reports roughly 2.15 million document records and 1.76 million indexed email messages (currently 1,758,382). Because the overlap between those two counters isn't publicly reconciled, neither should be treated as a count of original DOJ files \u2014 they count database records, not released documents. Their about-page description of the EFTA release as \"(6+ million pages)\" appears to garble a different DOJ number \u2014 the ~6 million pages the DOJ says it reviewed, most of which were never released. JMail also hosts or collaborates on a separate collection of ~20,900 leaked emails and attachments from Epstein's Yahoo account (13,010 public so far), via DDoSecrets. Those were leaked, not released under the Act, and I haven't included them in my own totals \u2014 nor should JDrive's page counter (which describes its \"new release\" archive) be assumed to contain them without a source-level query. And even if you generously threw them in: twenty-odd thousand emails against a claimed shortfall north of half a million pages doesn't move the needle. If we mean strictly \"things listed at a path on justice.gov/epstein/doj-disclosures,\" that's ~1.4 million files rendering ~2.9 million stamped pages \u2014 and the DOJ's own Data Set 12 listing ends at EFTA02858497.pdf \u2014 the exact last document in my archive (visible on the DOJ's Data Set 12 listing). One clarification, in the interest of counting honestly: EpsteinExposed is not wholly independent of my work \u2014 its early database incorporated material from my processed collection, so its EFTA file total (1,380,935) sitting essentially on top of mine \u2014 1,380,937, my count of the twelve numbered Data Sets \u2014 is a shared source, not a second opinion. Its current integrity dashboard attributes its SHA-256 inventory and deletion-tracking to separate community projects, so I treat it as useful corroboration of file availability and change over time, not as a clean independent page recount. The genuinely independent checks are the ones that matter here, and they all point the same way: CBS News, working from the raw DOJ files, counted about 2.7 million pages; Jmail's separate extraction got about 2.47 million; the printed reading room holds about 2.8 million. Different teams, different methods, the raw government source \u2014 every one of them below three million, not one of them near 3.5. The release isn't even stable While we're counting, the collection on DOJ's own website keeps shifting under our feet: Files move between data sets without notice: a document the DOJ's own file list puts in Data Set 2 (EFTA00003434) now serves only from Data Set 3 \u2014 so a link that returns \"file not found\" is often a relocation, not a deletion, and file counts drift depending on when you crawled The DOJ acknowledged that more than 47,000 files remained offline for further review after release 503 pages produced to the House Oversight Committee (HOUSEOVERSIGHT009974\u2013010476) were never posted publicly EpsteinExposed's integrity dashboard fingerprints the DOJ library and flags files that change or vanish \u2014 independent confirmation that the collection churns (its own live count of missing files is smaller than CBS's 47,000 and fluctuates, so it corroborates the pattern, not the exact number) Per DOJ's own statements, Deputy AG Blanche said the department reviewed more than 6 million total pages, and Reuters reported 5.2 million pages were slated for review \u2014 meaning the released 2.86M stamped pages are less than half of what the Department itself handled That last point deserves emphasis, because it's the comparison that should really sting \u2014 sharper than 2.9 against 3.5. Deputy Attorney General Todd Blanche has said the department reviewed more than 6 million total pages \u2014 a figure CBS reports and notes makes the public release \"less than half of the total.\" It has made about 2.9 million public. What the Department has never published is the bridge between those two numbers: how many pages were removed as duplicates or non-responsive, how many were withheld under specified legal grounds, and how many were actually posted. A transparent accounting would have supplied all four. The press release instead offered \"nearly 3.5 million pages\" \u2014 a figure that matches neither the pile it reviewed nor the production it released. A judge's preliminary finding \u2014 not just a counting dispute The gap between \"identified\" and \"released\" stopped being a rhetorical point in June 2026. In granting preliminary relief in Phang v. Blanche, U.S. District Judge Emmet Sullivan concluded that the plaintiff was likely to prevail on a limited set of claims, noting that the DOJ had not substantively answered them and had therefore, for purposes of the motion, \"conceded that he is in violation of the Act.\" He ordered the Department either to release specified material \u2014 including \"at least eight email exchanges with Mr. Epstein regarding a 'torture video'\" \u2014 without the disputed redactions, or explain why they should remain; after the DOJ responded, he ordered the unredacted documents submitted for private, in-camera judicial review rather than immediate public release. The litigation is ongoing. Reporting notes the Department still has not released some 2.5 million pages, and the DOJ has fought a proposed $1,000-a-day contempt sanction over it. House Judiciary Democrats made a similar allegation in writing. A March 19 letter from Reps. Raskin, Garc\u00eda and Jayapal states the DOJ \"has withheld 3 million pages\" and \"denied Members any opportunity to verify DOJ's claims that millions of pages withheld from disclosure are merely duplicates \u2014 a dubious assertion given so many files known to exist have yet to be revealed.\" They estimate it would take 7.5 years to check the redactions on the partial set they were given. So the \"3.5 million\" number sits inside a frame the government itself supplied: Deputy AG Blanche has said the department reviewed more than 6 million total pages, and the publicly observable production runs to roughly 2.7\u20132.9 million (a range the Department disputes at the low end). The rest it attributes to a mix of duplicates, nonresponsive and sealed material, victim-protection and privilege redactions, and technical or foreign-language files \u2014 and, in a preliminary ruling the DOJ disputes, a federal judge in Phang v. Blanche found the Attorney General \u2014 by then Acting AG Blanche \u2014 had conceded he is in violation of the Act, grounding that on Blanche's failure to rebut the plaintiff's arguments rather than any actual admission. It also removed tens of thousands of files after publication. In that context, a headline figure inflated past what was even released is not a rounding error. It is the friendly-sounding number placed over an accounting that never discloses how many pages were removed as duplicates or non-responsive, how many were legally withheld, and how many were actually posted. So what does the number mean? Here is precisely what I can and can't say: What's documented: The EFTA production's own Bates numbering terminates at 2,858,498. The public collections included in my database total approximately 2.92 million pages across 1.43 million document records. Multiple independent archives land in the same band. What's not documented: any accounting, from the Department, of how \"nearly 3.5 million pages\" was computed. The numbered production caps at 2,858,498 pages by the DOJ's own stamps \u2014 roughly 640,000 pages short of the claim. Candidate explanations, none confirmed: pre-deduplication counting (SDNY/SDFL duplicates were removed, per the press release); counting the ~180,000 images and ~2,000 videos as some internal multi-page equivalent; counting pages that were reviewed but withheld under the Act's exceptions. Each of those would mean the public number describes something other than what the public received \u2014 and in Phang v. Blanche the court found, at the preliminary-injunction stage and on a limited set of claims, that the Attorney General had conceded noncompliance for purposes of the motion. That is a preliminary finding the DOJ disputes, and it does not establish that the entire numerical gap is unlawful withholding. Here's the most banal explanation, giving them more benefit of the doubt than is rightly due: Bondi was quoting an earlier internal estimate \u2014 a number from before a chunk of pages got deemed nonresponsive, duplicative, privileged, or sealed \u2014 and nobody corrected the headline once the review finished. I'd buy that, except for one thing: they haven't quietly left it uncorrected \u2014 they've defended it. When CBS recounted the library and landed at 2.7 million, the Department didn't say \"our early estimate was rough\"; it called the recount \"fundamentally flawed.\" You don't go to the mat for a number you think was a loose internal guess. That's what makes it strange: the charitable version requires them to have simply forgotten to fix it \u2014 and they have spent months, on the record, not forgetting. The question, then \u2014 and I'm framing this as a question, because that's what the evidence supports: what is the 3.5 million figure counting? Either the figure was wrong when it was announced, or there exists a substantial volume of material that was counted internally and never produced, or the Department used an undisclosed counting convention \u2014 duplicates, native files, media, later-offline files \u2014 that does not represent unique public pages. All three are worth someone in an oversight role asking about, under oath, with the stamp arithmetic in hand. At the end of the day, we have what we have: a Bates series that numbers its own pages to 2,858,498, a complete public production of about 2.92 million even after you add every other source, and a press release claiming roughly 640,000 pages more than the numbering allows \u2014 a gap the adjacent material comes nowhere close to filling, and one the Department has never reconciled with the count anyone can actually run. Methodology: what each number measures Most of the confusion around \"3.5 million\" comes from sliding between measures that sound alike. Every figure in this report is one of these \u2014 defined, and reproducible against the public database: Measure Definition Value --- --- --- Document-metadata pages Sum of every document's stated page count \u2014 the primary page measure in this audit 2,923,763 Extracted page records Rows in the page table (one per extracted page) 2,924,310 EFTA pages held Document-metadata pages for the EFTA-numbered production 2,858,311 Bates-stamp ceiling The final Bates number in the series (last document EFTA02858497, two pages) 2,858,498 Documents, all sources Every document record 1,425,565 EFTA documents EFTA-numbered documents only 1,393,164 The 547-page reconciliation. Document-metadata pages (2,923,763) and extracted page records (2,924,310) differ by 547 \u2014 a 536 + 11 split. The larger part is the extraction layer: 309 native documents (spreadsheets, videos) are each produced by the DOJ as a single \"No Images Produced\" placeholder page, so their stated page count is 1, but their content extracts to many searchable records \u2014 536 rows beyond the 309 placeholders. The remaining 11 page records belong to three documents absent from the metadata table. Those extra rows are search units, not additional public PDF pages. The document-metadata figure (2,923,763) is the \"pages the DOJ produced\" count, and is used throughout. The 187-stamp gap. The Bates ceiling (2,858,498) minus EFTA pages held (2,858,311) leaves 187 stamps. Twenty-three are genuinely-blank Data Set 9 files (below); the remaining 164 are Bates numbers never assigned to any page in the sparser court-record ranges \u2014 unused numbers, not missing documents. The 23 blank Data Set 9 files Cross-referenced against the DOJ's own document manifest, exactly 23 files in Data Set 9 are absent from the archive. They are not withheld \u2014 they are blank at the source, and have been since release: Bates numbers What justice.gov serves Count --- --- --- EFTA00709804\u2013709807, 00770595, 00823190\u2013823192, 00823221, 00823319, 00877475, 00892252, 00901740, 00912980, 00919433\u2013919434, 00932520\u2013932523 0 bytes (empty file) 20 EFTA00774768 45-byte broken HTML fragment 1 EFTA01135215, EFTA01135708 2,433-byte \"No Images Produced\" placeholder page, generated 2026-03-04 2 Each was verified three independent ways: fetched directly from justice.gov through the age-verification gate (empty, while a neighbouring document downloads as a normal scan); the RollCall mirror, whose copies are dated 2 February 2026 and carry the MD5 hash of an empty file; and the Internet Archive, which has captured each URL daily since release day \u2014 160+ captures apiece, every one empty. They are also absent from a complete independent mirror of Data Set 9, and no native file exists for them at any of ~30 tested extensions. The DOJ published 23 Bates numbers with no document behind them. Reproduce every number Every figure above is queryable against the public database. A few of the load-bearing ones: -- Documents and document-metadata pages, all sources \u2192 1,425,565 \u00b7 2,923,763 select count(*), sum(total_pages) from documents -- EFTA production only \u2192 1,393,164 \u00b7 2,858,311 select count(*), sum(totalpages) from documents where eftanumber like 'EFTA%' -- Extracted page records \u2192 2,924,310 select count(*) from pages Run them yourself at epstein-data.com/fulltextcorpus. All counts captured 2026-07-31. And the specific lists behind the reconciliations above are published as machine-readable receipts: The 23 blank Data Set 9 slots, with the exact byte size justice.gov serves for each: DS9BLANKSLOTS.csv The 503 House Oversight pages produced but never posted (HOUSEOVERSIGHT009974\u2013010476): HOCUNPOSTED009974-010476.csv The complete native-file inventory (videos, audio, spreadsheets and their placeholder pages \u2014 the source of the 547-page reconciliation): NATIVEFILESCATALOG.csv The page-based gap-detection method behind the 187-stamp accounting: MISSINGEFTAANALYSIS.md --- Receipts Page-level archive: epstein-data.com/fulltextcorpus/pages \u2014 2.9M+ pages, SQL-queryable Last EFTA document: EFTA02858497, visible on DOJ's Data Set 12 listing DOJ reading room: Washington Examiner Reproducible queries: every number above can be re-derived via the SQL API \u2014 e.g. https://epstein-data.com/fulltextcorpus.json?sql=select count(*), sum(total_pages) from documents DOJ press release, Jan 30, 2026: \"\u2026Publishes 3.5 Million Responsive Pages\u2026\" Community mirror/index of DOJ Data Sets 9\u201311: yung-megafone/Epstein-Files ~48K files removed: Yahoo News NPR, DOJ withheld/removed documents naming Trump: npr.org DOJ Inspector General, audit of EFTA compliance (primary): oig.justice.gov \u00b7 NBC coverage DOJ defends withholding 2.5M pages; Blanche \"likely violated\": Spokesman/AP \u00b7 Forbes, Jul 27 5.2M pages to review (Reuters wire): US News / Reuters \u00b7 CNBC The printed exhibit, on-site account: Al Jazeera, \"A paper city\" EpsteinExposed integrity/hash monitoring: epsteinexposed.com/integrity \u00b7 stats Cross-archive counts: tommycarstensen.com/epstein \u00b7 epsteinexposed.com (open-source code) \u00b7 jmail.world \u00b7 DDoSecrets (leak provenance: ~20,900 emails, 13,010 public) Methodology note: all third-party counts (Jmail, EpsteinExposed, Tommy Carstensen) were captured on July 31, 2026 \u2014 these are live sites and will drift. My own counts exclude ~1,400 transcription placeholder entries our transcription software added (community transcript integration, July 2026); all figures are of DOJ-released material as produced. The 22k DDoSecrets-leaked emails are excluded throughout \u2014 leaked \u2260 released."]], "columns": ["slug", "title", "category", "filename", "url", "efta_count", "eftas_json", "persons_json", "summary", "body"], "primary_keys": ["slug"], "primary_key_values": ["audits/PAGE_COUNT_AUDIT"], "units": {}, "query_ms": 4.851039499044418, "source": "Epstein Files Transparency Act (Public Law 119-38) DOJ Production", "source_url": "https://www.justice.gov/epstein", "license": "CC BY-NC-SA 4.0", "license_url": "https://creativecommons.org/licenses/by-nc-sa/4.0/"}