Before anything else about today’s release, one question, which I am going to leave sitting there until the end because the answer is the whole point: how much of it is anyone going to read?
The Department of Justice published its files this morning under the Epstein Files Transparency Act, which was signed in November. The deputy attorney general’s letter describes roughly three million pages, about two thousand videos and some hundred and eighty thousand images, drawn from five groups of cases and investigations. Around two hundred thousand pages were redacted or withheld under claims of privilege. The department’s position is that it has met its legal obligations.
Note in passing that the department’s own headline says three and a half million responsive pages while the letter describes three million published, with two hundred thousand held back. Those numbers do not close, and somebody will have to explain the gap.
Set that aside. This is a ledger of what a release of this size costs, who pays it, and who collects, because the size is not incidental to the disclosure. It is a property of the disclosure.
What it costs to read
Take the three million pages and give each one a single minute. That is three million minutes, which is fifty thousand hours, which at a two thousand hour working year is twenty five person years of continuous reading.
A minute a page is generous for a corpus of this kind, much of which will be routine. So take ten seconds instead, which is long enough to see what a page is and not long enough to understand it. That is still eight thousand three hundred hours, or a little over four person years, to skim.
Now scale it to an actual newsroom. Twenty people, working on nothing else, for a month, is about three thousand two hundred hours. At ten seconds a page that team gets through somewhere near a third of the pages, retaining almost nothing, and has not touched two thousand videos.
No organization is going to spend that. What organizations will actually do is search for names they already suspect, publish what comes back within about seventy two hours, and move on, which means the effective disclosure is not three million pages. It is whatever a keyword search surfaces in the first three days.
Who pays
The reading cost falls on whoever wants to know: newsrooms, congressional staff, academics, and the lawyers acting for people with a direct interest. None of them was compensated for it and none of them chose the volume.
It also falls unevenly, and that is the part with consequences. A large outlet can put twenty people on it. A small one can put one. Volume is therefore a filter on who can participate in scrutiny, and it filters in favor of the largest institutions, which are not always the ones asking the most awkward questions.
Who collects
The releasing institution collects, and what it collects is finality.
Compliance with a disclosure statute is discharged at the moment of publication. Comprehension, if it happens, happens over months. In the gap between those two events sits a period in which the institution can accurately say it has released everything and critics can only say they have not finished looking, which is a much weaker sentence to have to say in public.
This is not an accusation of design. Three million pages is a plausible number for five sets of investigative files, and the redactions protecting victims’ personal information are required by the statute and are the right ones to make. The point is structural: a release large enough to be unreadable produces the same day one outcome whether or not anyone intended it, and an institution that benefits from a structure does not have to have built it.
What determines whether the volume matters
Here the question stops being about the law and becomes an ordinary problem in information retrieval, with three variables that will settle it inside a week.
The first is whether the pages carry a text layer. Scanned paper published as images is not searchable at all, and published with poor optical character recognition it is searchable in a way that quietly fails: a name misread by the software simply does not exist as far as any search is concerned, and the searcher gets no result and no warning that the result is wrong.
The second is deduplication, and it is the variable that cuts the public’s way. Investigative productions are enormously duplicated. The same email appears in the files of every recipient, exhibits are attached repeatedly, drafts accumulate. In corpora of this kind it is normal for a large fraction of the nominal page count to be duplicates or near duplicates, and software can collapse them reliably. The real quantity of distinct material is likely to be a great deal smaller than three million, and nobody knows yet by how much.
The third is metadata. Professional document productions normally ship with a load file: dates, custodians, document boundaries, which attachment belongs to which parent. With those fields the corpus can be sorted, filtered by period and read as a structure. Without them it is three million loose pages in an arbitrary order, and the difference between those two deliverables is not a matter of degree.
The precedent that shows what it takes
The useful comparison is the Panama Papers, which ran to about eleven and a half million documents and roughly two and a half terabytes, several times larger than this.
It was navigable, and it was navigable because a consortium of journalists spent close to a year before publication running optical character recognition across the whole set, indexing it into a full text search system, extracting the entities and building a graph of the relationships between them, and then handing several hundred reporters an interface that let them ask questions of it.
The lesson is not that big leaks work out. It is that the index was the journalism, and it took a year and a substantial technical team before a single story was written. Nobody has built that for today’s release, because it went up this morning.
The question
So: how much of it is anyone going to read?
Directly, page by page, essentially none of it. That much is arithmetic and it was determined the moment the volume was set.
Indirectly, through a good index, potentially most of it, and the index is buildable. The tools are ordinary and the skills are not rare. Somebody will have a searchable copy standing within weeks, and the useful reporting will come from that rather than from today.
Which leaves the thing that cannot be settled now. Whether this was a disclosure or the appearance of one depends entirely on what those three variables turn out to be, and all three are facts about files that have been public for a few hours. If the text layer is clean and the load files are there, the volume was just volume and the record is genuinely open. If the pages are images, or the metadata was stripped, then three million pages was published and very little was disclosed, and the statute was satisfied by an act that defeated it.
And the uncomfortable part is that the announcement is the same either way. A clean, indexed, fully machine readable corpus and a heap of unsearchable scans both produce the identical sentence, three million pages released in compliance with the Act, and the identical headline under it. Nothing about the claim distinguishes them. Nothing about the number distinguishes them. The only thing that does is the files, and the files are the thing almost nobody is going to open.




