Tales of Distortion: The Book Was More Than the Text

They Destroyed the Books for the Information. Starring TraceLossBench.

Tales of Distortion: The Book Was More Than the Text
There is a dinosaur in Las Vegas eating a book.

This is not symbolism.

In August 2026, 404 Media hid a tracking device inside a rare book and followed the shipment across the country. The trail ended at an Amazon warehouse in Las Vegas. The team operating there is called VGT3.

Its logo is a dinosaur holding a book.

That is almost too considerate.

According to Amazon employees interviewed by 404 Media, shipments of printed books arrive at the facility in enormous quantities. Workers unbox them, stage them, scan identifying information, and move them toward a row of cutting stations. A book enters the machine intact. A blade comes down through the binding. The pages separate.

Then the loose pages go to the scanners. Then the pages go into large open cardboard containers with the pages of other books.

The book does not come back out.

The follow-up reporting is unusually useful here because it removes any temptation to imagine a pristine archive with a dramatic recycling bin at the end. An employee described used books that appeared to have come out of libraries, boxes connected to the University of London, government documents, books in German and Russian, pallets of Japanese books, new books, old books, and obscure books whose future usefulness could not possibly have been known by the person standing at the cutter. After scanning, the pages were mixed together as loose paper. Reassembly was no longer a practical option.

This was the process. The strange part is why.

The books were being destroyed because the books contained information.

That sentence is the whole case. If the books had been worthless, nobody would have purchased them by the shipment, moved them across the country, staffed a warehouse, installed industrial cutting stations, operated banks of scanners, and built a pipeline around extracting what they contained.

The destruction is downstream of value.

Someone looked at a physical book and concluded that something inside it was valuable enough to acquire at industrial scale. Then the acquisition system destroyed the source object while collecting the part of its information that the system already knew how to capture.

That is a very different failure from carelessness.

Carelessness drops a book in the rain.

This is an instrument meeting the objective assigned to it. The blade is there because the binding slows the scanner down.

The binding slows the scanner down because books are awkward physical objects. They have thickness. Curvature. Page order. Margins. Covers. Adhesives. Stitches. Signatures. Inserts. Endpapers. Marks. Repairs. Different papers. Different inks. Different histories of use. These things resist becoming a stream.

The cutter repairs that problem for the scanner.

It does so by repairing the book out of existence.

That is the first distortion. The warehouse does not need to announce a theory of books. It only has to embody one.

  • A source object enters.
    • A digital derivative leaves.

The process is optimized as though the derivative contains the information that mattered and the source contains whatever was inconvenient about getting to it.

The book has become packaging.

That conclusion is much larger than OCR accuracy. A scanner can satisfy every technical metric assigned to it and the process can still destroy information.

The pages can be captured at high resolution. The text can be recognized with extraordinary accuracy. The order can be reconstructed. The resulting files can be searchable, duplicable, indexable, compressible, chunkable, tokenizable, and useful enough to train systems worth billions of dollars.

None of those achievements establish that the source object has been exhausted. They establish that the derivative is useful.

Useful is not identical.

The distinction matters most when the source is destroyed after capture, because an incomplete derivative and an incomplete derivative backed by a surviving source are different epistemic objects.

If the book still exists, the future gets another question.

If the book survives, future inquiry can still return to it. Destructive capture changes we did not record that into we can no longer go back and look. One is an omission. The other places an irreversible boundary around every inquiry that comes later.

This is why the VGT3 warehouse matters beyond Amazon.

The books moving through it were not all interchangeable clean copies fresh from one printer. The reporting describes an intake field containing used material, library material, foreign-language material, obscure material, and institutional material. Whatever survives in the digital derivative may be extremely useful.

Whatever fails to survive is being judged by a pipeline whose success criterion is successful capture of the representation it was built to produce.

The scanner cannot tell us that everything outside that representation was worthless. It can only tell us that it did not carry it forward. The physical destruction then turns that local limitation into a permanent one.

This is the part that disappears when the event is narrated as a copyright fight, an artificial-intelligence fight, a publishing fight, or a sentimental fight about whether people like the smell of old books.

The book is an information-bearing physical object.

The text is one of the things it carries.

The scan is a representation of some of what the object made available to the capture system. The OCR output is another derivative. The normalized text is another. The training representation is another.

The path can produce enormously valuable artifacts at every step.

Value does not reverse the direction of derivation.

The token sequence did not become the book because the model needed the token sequence more than it needed the binding. The file did not become the source because the warehouse wanted a file. And the unmeasured properties of the physical object did not stop being information because they were expensive to keep reachable.

There is a dinosaur in Las Vegas eating books because books contain information.

That is the joke. That is also the distortion.

They wanted the information in the books so badly that they destroyed information in the books to get it faster.


The Hidden Equation: Book = Text.

This destruction only makes operational sense if a substitution has already happened.

  • The book goes into the warehouse as a physical object.
    • The pipeline wants information.
      • The information it is built to extract is overwhelmingly textual.

Once that text has been captured, the physical object becomes redundant to the pipeline. Nobody has to write BOOK = TEXT on the wall. The workflow writes it for them.

This is the hidden equation underneath the cutter.

  • A binding can be removed because the binding is outside the representation being optimized.
  • A page can be separated from the object that ordered it because page order can be reconstructed as data.
  • A cover can become an image file.
  • A printed sentence can become OCR output.
  • The OCR output can become normalized text.
  • The normalized text can become tokens.

At each step, the downstream artifact becomes easier for the machine to use.

Then, language starts to slide. The PDF gets called the book. The OCR gets called the book. The corpus entry gets called the book. The training example gets described as though the book itself reached the model.

A relation has disappeared from the sentence. The source became a derivative, and the derivative inherited the source's name. That is how a pipeline can destroy the book while continuing to say that it has the book.

The same substitution appears with unusual clarity in the litigation over Anthropic's book collection.

In 2024, Anthropic hired Tom Turvey, who had worked on the Google Books scanning project, and began building a permanent research library at extraordinary scale. The federal court record says Anthropic spent many millions of dollars purchasing millions of print books, often used copies. Its service providers stripped those books from their bindings, cut the pages to dimensions suitable for scanning, produced digital copies, and discarded the paper originals.

Each purchased print book produced a PDF containing scanned page images and machine-readable text. Anthropic also built bibliographic metadata around the collection. The objective was a searchable central library large enough to approach what the court described, using Anthropic's own language, as “all the books in the world.”

There is a phrase in Judge William Alsup's 2025 fair-use order that is worth stopping over.

For the copyright question before the court, the judge held that converting a purchased print book into a digital file for space-saving and searchability could be treated as a transformative use. The digital copy, for that legal analysis, could stand where the purchased print copy would have stood in Anthropic's internal library.

That may be a coherent copyright ruling.

It does not answer the question in front of us.

  • The court was deciding whether a particular act of copying was permitted under copyright law.
  • We are asking what information crossed the transformation.

Those are different fields.

Copyright law can decide that a digital copy is an acceptable replacement for a particular legal purpose without establishing that the digital copy preserves every property of the physical object from which it was made.

A warehouse can decide that a PDF is an acceptable replacement for a particular operational purpose without establishing that the PDF contains everything future investigators might learn from the book. A training pipeline can decide that machine-readable text is an acceptable input without establishing that the text exhausts the source.

The danger begins when adequacy inside one field quietly becomes identity across all of them. The file was good enough for searchability. Therefore the file was the book. The text was good enough for training. Therefore the text was the book. The model did not need the binding. Therefore the binding carried no information worth preserving.

None of those conclusions follows.

The word book makes the slide easy because it already names several different things in ordinary speech.

  • “I bought the book” usually points to a particular physical copy.
  • “I read the book” usually points to the intellectual work expressed across copies.
  • “This is the first edition of the book” points to a publication state.
  • “The library has the book” may point to one physical holding among many.
  • “The dataset contains the book” often means that some textual representation associated with the work appears somewhere in the corpus.

Those senses overlap enough to feel interchangeable until a process destroys one of them. Then the ambiguity becomes expensive. A physical copy can disappear while the institution continues to possess “the book” in another sense.

The sentence stays grammatically true. The source does not stay available.

That is why provenance has to be more demanding than naming. Book X cannot indiscriminately mean a work, an edition, a physical copy, a scan, an OCR transcript, a normalized corpus object, and a token sequence. Those are related objects. They are not one object traveling unchanged through the pipeline.

The useful record is a path: this source object was acquired here; this capture produced this derivative; this extraction produced the next one; this transformation changed these features. Once the path is explicit, the hidden equation becomes much harder to sustain.

A scan is allowed to be a scan. An OCR transcript is allowed to be an OCR transcript. A normalized text file is allowed to be a normalized text file. A token sequence is allowed to be a token sequence. None of them needs to be insulted by pretending to be something larger.

The problem appears when the derivative is promoted.

The promotion erases the loss ledger.

  • If the PDF is the book, there is no reason to record what the PDF did not capture.
  • If the OCR is the text, there is no reason to preserve the page image once OCR confidence clears a threshold.
  • If the normalized string is the source text, distinctions removed during normalization become invisible by definition.
  • If the token sequence is the training data that matters, everything upstream becomes historical clutter.

This is how an optimization becomes an ontology without anyone holding a meeting about ontology.

The system keeps renaming its output after its input.

Then the output is judged by whether it is good at being the output the system wanted. The original source is evaluated only through the properties selected for extraction. Anything else arrives at the decision point without a column in the spreadsheet.

The absence of a column begins to look like the absence of information.

Then the cutter comes down.

There is another reason the equation matters.

A derivative can sometimes preserve everything relevant to a particular task.

That does not make the derivative universally sufficient.

A plain-text transcription may be ideal for full-text search. A page image may be ideal for checking a typesetting anomaly. A high-resolution color capture may be necessary for studying ink or annotation. A physical copy may remain necessary for questions that depend on paper, binding, sequence, repair, insertion, wear, ownership, manufacturing, or material history.

The adequacy claim always has a scope.

  • For what purpose?
  • For which properties?
  • Under which transformation?
  • With what possibility of returning to the source when the purpose changes?

The destructive pipeline suppresses those questions by answering them once, in advance, for everybody who comes later. It says that the representation selected today is sufficient for the inquiries of tomorrow.

That is an extraordinary claim to hide inside a throughput decision.

And it becomes stranger the closer we look at what the machine is actually doing.

Because the next problem is no longer philosophical.

If the book is not identical to the text extracted from it, what else was in the book?

What Was in the Book?

A lot, it turns out.

The easiest way to miss it is to start with a list.

Binding. Paper. Ink. Typography. Illustrations. Page geometry. Covers. Endpapers. Bookplates. Inscriptions. Marginalia. Repairs. Stains. Damage. Wear.

That list is accurate. Also too weak. It makes the physical properties of a book sound like decorative extras attached to the real thing.

The text remains in the middle. Everything else becomes metadata around it. That is the same mistake again. A physical book is not a text wearing a coat. It is an object with a history.

And history leaves evidence in the object.

Bibliographers, conservators, librarians, historians of the book, provenance researchers, collectors, and curators have spent generations learning how to read that evidence because two copies carrying the same printed words do not necessarily carry the same information.

The same edition can survive in multiple copies. One may retain its original binding. Another may have been rebound two centuries later. One may contain a reader's corrections. Another may contain a bookseller's price code, an owner's signature, a library stamp, a censor's intervention, inserted leaves, a missing gathering, or a repair made with material taken from somewhere else.

  • The printed text can match.
  • The objects can tell different stories.

This is why rare-book description has a vocabulary for copy-specific evidence. The phrase matters. A work is not a copy. An edition is not a copy. A scan of a copy is not the copy.

The individual object can acquire information after publication that no other copy of the same edition possesses. Ownership changes it. Reading changes it. Repair changes it. Storage changes it. Damage changes it. Institutions change it.

Sometimes the change is intentional. Sometimes somebody writes in the margin. Sometimes somebody replaces a binding. Sometimes somebody pastes in a bookplate. Sometimes a librarian stamps the title page. Sometimes a reader folds a corner, erases a sentence, underlines a passage, inserts a clipping, or leaves a note between the leaves.

Sometimes nobody means to leave evidence at all.

That does not stop the evidence from existing.

A scratch on a binding can become evidence of use. A watermark can become evidence about paper. A typeface, ornament, or damaged piece of type can help distinguish printing states or production practices. The structure of gatherings can reveal how sheets were imposed, folded, assembled, or altered. Bindings can provide evidence about where and when a copy was bound, rebound, sold, owned, repaired, or transported. Annotations can preserve the reactions of readers who are otherwise absent from the historical record. Ownership marks can reconstruct a chain of custody. Physical alterations can show that the object did not spend its life in the condition in which an investigator finally received it.

This is not an exotic theory about books. It is ordinary professional practice in the institutions that preserve these things.

Rare Book School teaches bibliographical analysis by examining paper, parchment, type, script, illumination, bindings, ownership marks, and annotations because those features can answer questions about production, distribution, provenance, and use. Its course in physical bibliography treats typography, illustration, paper, binding structure, inscriptions, bookplates, ownership markings, and alterations by readers as bibliographical evidence. Its course devoted specifically to paper exists because paper itself can answer historical questions.

The Library of Congress maintains cataloging machinery for the same reason. MARC has a field for ownership and custodial history. It has structured fields for copy-specific provenance evidence and a field for binding information intended especially for rare books and special collections.

The standards community has continued expanding the ability to identify physical-medium information that applies to one particular copy rather than to every copy of an edition. Libraries did not need artificial intelligence to discover that the physical object carried information. They had already built database fields for it.

That makes the destructive-scanning problem even stranger.

  • The information system on one side of the building knows that a book has copy-specific material evidence.
  • The extraction system on the other side can behave as though the useful payload is the words.
  • One field spends enormous effort distinguishing the object from its description.
  • Another can turn the object into a description and discard the object.

This contradiction is not resolved by saying that modern commercial books are different from fifteenth-century incunabula. Of course they are different. The relevant point is that material significance is not restricted to old books.

A modern copy can carry an author's inscription or editorial corrections. It can contain a review copy slip, a printer's defect, a later replacement page, a library accession mark, a particular dust jacket, an owner's annotation, an inserted photograph, a receipt, a dedication, or evidence of how a specific copy circulated.

A future investigator may care about any of those things.

  • A historian may care about ownership.
  • A conservator may care about materials and repair.
  • A bibliographer may care about production state.
  • A literary scholar may care about annotation.
  • A legal dispute may care about edition identity or chain of custody.
  • A family may care about an inscription because it establishes a relationship no catalog record contains.

A machine-learning pipeline may care about none of them.

That last sentence is still allowed. The pipeline is allowed to have a purpose. The mistake begins when the pipeline's purpose becomes the ontology of the source.

If an artificial-intelligence training pipeline only needs lexical text, it may be completely rational for that pipeline to extract lexical text. That does not establish that lexical text was all the information in the object. It establishes what the pipeline was built to notice.

Those are different claims. And this is where the word information becomes very dangerous if it is left undefined.

  • Information for what?
  • Information about what?
    • Available to which instrument?
    • Distinguishing which states?

A scanner can capture a great deal. A high-resolution image can preserve information that OCR discards. A color-managed image can preserve information a monochrome scan does not. A page image can preserve typography and marginalia that plain text cannot. A three-dimensional imaging system can preserve properties a flatbed scanner cannot. Chemical analysis can answer questions no photograph can answer. Direct physical examination can expose relations that none of those derivatives was designed to retain.

Every instrument has a field. The problem is not that one instrument fails to contain the universe. It is forgetting that it has a field at all.

A page scanner sees what its capture apparatus can represent. OCR sees a different object again. Normalization sees another. Tokenization sees another.

The transition may be useful at every stage. It may even be excellent for the purpose that justified it.

What cannot happen is the silent promotion from:

  • this instrument captured what we needed

to:

  • this instrument captured what was there.

Those sentences are nowhere close to equivalent.

  • The first is a scoped engineering claim.
  • The second is a claim about reality.

Destructive scanning raises the stakes because the book is not simply observed. It is changed in order to make the observation easier. The binding is removed. The pages are separated. The structure that held the object together is sacrificed to throughput. That means some properties are not just omitted from the digital derivative. They are altered or destroyed by the capture process itself.

This matters even if every page receives a beautiful scan afterward.

The images can show what the pages looked like after disbinding. They cannot retroactively restore every physical relation that the act of disbinding changed.

They do not re-create the sewing structure. They do not make a removed binding original again. They do not guarantee preservation of hidden or inaccessible features inside that binding. They do not preserve the object's future availability for an instrument nobody thought to use today.

The source has become part of the measurement procedure.

Then the measurement procedure destroys it.

That is the part the text-only account cannot express.

Imagine two records in a database. Both contain the same OCR transcript.

  • One came from a copy with no annotations.
  • The other came from a copy containing a handwritten correction, an owner's bookplate, and evidence of rebinding.

If the representation preserves only the transcript, those records can become indistinguishable for the system even though the source objects were distinguishable under properties a later inquiry may care about.

The corpus has not proven the objects equivalent. It has selected a representation in which their difference cannot be seen. That is the bridge to the technical problem Modal Path Ethics eventually had to build machinery to test. But the physical-book case is larger than that machinery.

Because the most important information loss may concern a property nobody declared in advance. Nobody can enumerate every future question that could be asked of a surviving object.

That is why preservation has value beyond today's question.

  • The source keeps options open.
  • The derivative answers some questions more conveniently.

A good preservation system understands the difference. A destructive extraction system can erase it. A preservation system has to know which of those relations it is managing. The richer object is still capable of answering questions its present derivative cannot.

This is also why the usual defense of we only destroyed duplicates does less work than it appears to do.

Duplicate of what?

Two copies may instantiate the same work and edition while differing as physical objects. A duplicate at the level of printed textual content is not automatically a duplicate at the level of provenance, annotation, binding, condition, or material history.

Libraries know this too. Special-collections policies can explicitly value additional copies when they contain copy-specific information of research interest. The second copy earns its place precisely because it is not informationally redundant under every relevant description.

The word duplicate therefore requires the same discipline as the word book.

Duplicate relative to which properties?
Equivalent for which task?
Redundant under which instrument?

Once those questions are restored, the cutter no longer looks like a neutral preprocessing step. It looks like a decision about which distinctions the future is allowed to recover. And that decision was made before the future showed up to ask its questions.


So what was in the book?

The text was in the book.

So were the images.

So was the typography.

So was the organization of the pages.

So were the materials that made those pages possible.

So was the structure that held them together.

So were the marks left by production, ownership, reading, storage, repair, damage, and time.

So was the identity of this copy rather than another copy.

So was evidence whose significance may not have been known yet.

And, perhaps most importantly, the surviving book contained one more thing the extracted text could never contain by itself: the possibility of asking the source another question.

Which means the next problem is unavoidable.

What happens to information that an instrument never measures?

It does not stop existing because there was no column for it.


The Information You Do Not Measure Still Exists.

A thermometer does not erase humidity.

Put one in a room and it gives you a temperature.

That is useful information. It is also an extremely bad inventory of the room. The thermometer does not report the humidity, the pressure, the concentration of carbon dioxide, the color of the walls, the number of people standing nearby, or whether somebody has left a pot of soup burning on the stove.

None of those things disappear because the instrument has no field for them.

The instrument has a remit.

That sounds so obvious when the instrument is a thermometer. It becomes strangely difficult when the instrument is a data pipeline.

The output arrives looking complete. There is a file. The file has rows. The rows have fields. The fields have values. The OCR succeeded. The parser did not crash. The tokenizer returned tokens. The database contains a record. The system therefore has the psychological appearance of having captured the thing.

It has captured a representation of the thing.

That distinction is the entire problem.

A measurement system does not receive reality and hand reality back. It maps some states of a source into states the instrument knows how to represent.

If two different source states produce the same measurement, the instrument has made them indistinguishable within that representation. It has not proven that the source states were identical.

This is easy to see with books.

Take two copies that contain exactly the same printed sentence.
  • One sentence is set in one typeface.
  • The other is set in another.

Plain-text extraction can return the same character sequence from both.

The extractor has done its job.

The typographic difference has not been refuted. It has been projected away.

Take a page containing a printed paragraph and a handwritten correction in the margin.

An OCR system trained to recover the printed body text may return a beautiful transcript of the paragraph and nothing from the correction.

The transcript can be entirely accurate as a transcript of what it measured. It still does not follow that the page contained no correction.

Take a color annotation and convert the page to a representation that does not preserve color.
Take a binding and photograph only the leaves.
Take a bookplate and crop the page around the printed text.
Take a foldout map and flatten the book into a sequence of text blocks.
Take the relation between a caption and the image above it and preserve only the words in the caption.

The same pattern repeats.

The instrument can succeed while the representation contracts. The failure begins when success at the assigned measurement is promoted into a claim that the source has been exhausted. That promotion is subtle because databases are very good at making absence look like ontology.

Suppose a corpus record contains:

  • title
  • author
  • date
  • text
  • language
  • source URL

There is no field for binding.

There is no field for marginalia.

There is no field for ownership history.

There is no field for paper stock.

There is no field for whether the copy was rebound.

There is no field for whether somebody incorrectly corrected a date in pencil on page 147.

Once the record becomes the thing everybody downstream sees, those omitted properties acquire a dangerous status. They stop looking like unmeasured properties of the source. They start looking like properties the source did not have.

The database never actually established that. It simply had nowhere to put them.

A missing column is not a negative measurement. A field that was never assessed is not a field whose value is zero. No record of an annotation is not evidence that the page was unannotated unless somebody actually checked for annotations under a procedure capable of finding them.

That distinction is basic enough that it should be embarrassing to have to say. It is also exactly the kind of distinction large preprocessing systems erase by accident.

The pipeline wants a tractable state space. Reality is rude enough to contain more states than that. So the pipeline collapses them.

That is allowed. Compression is allowed. Normalization is allowed. Abstraction is allowed. A useful instrument often works by refusing to preserve distinctions irrelevant to its task.

The discipline comes afterward.

Which distinctions did we stop carrying?

That question has an old home in preservation science. Digital preservation did not wait for artificial intelligence to discover that transforming an object can alter what survives.

The InSPECT project spent years developing methods for identifying what it called significant properties: characteristics that need to be maintained across preservation actions so an object can continue to be accessed, used, understood, and accepted as evidence of what it purports to record.

InSPECT explicitly treated significance as relative rather than universal. Different stakeholders can need different properties. Different purposes can make different features important. A migration that is perfectly acceptable for one use can be destructive for another.

The project's framework separated properties into categories including content, context, rendering, structure, and behavior, then compared what existed in a source representation with what survived into a destination representation.

That is already enough to break the hidden equation from the previous section. The destination is not presumed to be the source in a better outfit. It is evaluated as a reformulation with a preservation burden.

The National Archives works from the same premise. Its Digital Preservation Framework identifies significant properties for different record types and uses those properties as criteria for evaluating transformations. The plans are explicitly not presented as exhaustive or universally applicable.

The Library of Congress is even more direct in its glossary.

A significant property is a characteristic of an object subjectively determined to be important to maintain through preservation actions.

Subjectively.

There it is.

The preservation field put the decision in the definition. That one word does a tremendous amount of ethical work. Subjectively says that an institution may decide which properties it needs to preserve. It does not say the properties outside that decision cease to exist.

It says the preservation action has a scope.

That is what the destructive training pipeline forgot.

The cutter is not a theory of information.

The scanner is not a theory of information.

The OCR model is not a theory of information.

The tokenizer is also not a theory of information.

Those are instruments implementing successive selections.

At each transition, a property can end up in several very different states. It can be preserved, altered, or collapsed with another state. It can be omitted deliberately or accidentally. It can remain unassessed because nobody asked about it. It can become impossible to assess because the source needed to inspect it no longer exists.

Those are not synonyms.

A provenance system that records only source → derivative misses the central question.

Yes, this text came from that book.

Fine.

  • What did the transition carry?
  • What did it change?
  • What did it never inspect?
  • What evidence supports each answer?
    • And if somebody discovers tomorrow that a property matters, is there still a richer source upstream that can be examined again?

Without those distinctions, provenance can become a beautifully documented story of how an impoverished representation traveled through the system.

Every arrow can be correct. The source can still disappear behind them.

This is where a very simple epistemic rule becomes necessary:

Unassessed is not preserved.

If nobody tested whether typography survived, the answer is not yes. If nobody recorded whether annotations were present, the answer is not none. If nobody examined the binding before it was cut away, the answer is not irrelevant.

The honest answer is unassessed.

That one can feel unsatisfying because it refuses to complete the table.

Reality is under no obligation to fit the table.

A serious information system should be able to represent its own ignorance. This is especially important for machine-learning corpora because the preprocessing stack can be long enough that nobody downstream sees the original source at all.

  • A person assembling a dataset may receive OCR from another vendor.
  • A cleaning stage may remove markup.
  • A normalizer may rewrite characters.
  • A deduplication system may decide two records are equivalent.
  • A chunker may sever one passage from another.
  • A tokenizer may map distinct strings into the same available representation.

By the time training begins, the final record can look tidy enough that the source's unresolved distinctions are psychologically gone.

They are not gone from history. They are gone from the representation. That difference is exactly what a provenance system is supposed to stop us from forgetting.

The same discipline applies to high-resolution scans.

  • A high-resolution scan is richer than OCR.
    • That does not make it the book.
  • A color scan can preserve properties a monochrome scan loses.
    • That does not make color scanning exhaustive.
  • Multispectral imaging can expose features ordinary photography cannot.
    • That does not make multispectral imaging exhaustive.
  • Three-dimensional surface capture can preserve geometry a flat image cannot.
    • That does not make a mesh exhaustive.
  • Chemical analysis can reveal materials invisible to all of those systems.
    • That does not make chemistry exhaustive.
  • Direct physical inspection can answer questions none of the derivatives were designed to answer.
    • That does not make direct inspection omniscient either.

Every instrument opens a field and closes others. This is not a defect in instruments. It is why instrument discipline exists.

You do not punish the thermometer for failing to report humidity.

You punish the report that says the thermometer measured the whole room.

The preservation literature understands this so well that digitization guidance from the Library of Congress begins by asking what the project is for, who will use the surrogates, how the surrogates will be presented, what effect digitization will have on the original object, and how the resulting files will be managed.

Those questions are upstream of the scanner. They decide what the scanner is being asked to do.

The same Library of Congress guidance says digitization projects vary according to the needs and value of the physical sources and the significant properties of the originals.

Again: scope first. Capture second.

The warehouse process reverses that authority when it treats successful capture as permission to destroy what the capture did not establish as redundant.

The instrument is optimized for one path through the object. Then the object is forced to become only what survives that path.

There is a useful way to see the absurdity.

Imagine a laboratory receives an unknown mineral sample.

The first instrument weighs it.

The scale reports 41.7 grams.

Excellent. So the laboratory destroys the sample because the mass has been preserved in a database.

Nobody would say that sample has been successfully archived.

Nobody would accept 41.7 grams as the complete informational content of the object. Nobody would argue that spectroscopy, microscopy, isotopic analysis, structural analysis, or a future technique had become unnecessary because the scale worked perfectly.

The mistake is obvious because mass announces itself as one property among many. Text does not announce itself that way. Language feels like the content. That psychological privilege is doing enormous work. Books are built to communicate through language, so textual extraction can feel like extraction of the essence.

Sometimes text is exactly what the current task needs. That still does not license the jump from the task needs text to the source is text.

Preservation science has a better word for the digital result.

A surrogate.

That word keeps the relation alive. A surrogate stands for something. It can be more searchable than the source, easier to copy, easier to distribute, easier to analyze, easier to compare, easier to quote, and dramatically more useful for many forms of research.

Its usefulness does not cancel the preposition. It is a surrogate for the source.

Every measurement leaves a remainder.

Some of that remainder is known.

  • We know plain text does not preserve a binding.
  • We know OCR can omit visual relations.
  • We know normalization can collapse distinctions.

Some of the remainder is unknown. Nobody today can list every future question a researcher might ask of a particular physical copy. The physical source is therefore more than a warehouse of already-known significant properties. It is also a reserve against our incomplete specification of significance.

That is enough to establish the measurement rule. An instrument can be excellent inside its remit without exhausting the source. The next question is what happens when we erase the richer object that could have answered everything the instrument left unresolved.


Destruction Converts Ignorance Into Irreversibility.

There is a large difference between failing to record something and making it impossible to record later.

  • The first is an incomplete measurement.
  • The second is an intervention on the future.

That distinction is where destructive scanning changes the character of the problem.

A scan can be incomplete and still be useful. A plain-text derivative can lose layout and still be useful. A tokenizer can collapse distinctions and still be useful. A corpus can be built from those derivatives and still be useful.

Usefulness is not the issue. The issue is what happens after the system discovers that the derivative is useful enough.

If the source remains on a shelf, the derivative has a superior upstream witness. If somebody later discovers a missing annotation, a color distinction, a binding feature, a provenance mark, a replacement leaf, an illustration relation, a paper characteristic, or a question nobody thought to ask when the scan was made, the source can be examined again.

The derivative can be corrected. The capture can be repeated. A second instrument can be brought in. A disputed reading can be checked against the thing itself.

The source has one final job after digitization that is easy to underestimate:

it can disagree with its derivative.

That disagreement is an extraordinary preservation resource.

A derivative says the page contained this. The source can still say no.

A transcript says the mark was an ordinary numeral. The source can still reveal that it was superscript.

A scan says the page was blank around the printed body. The source can still reveal a faint pencil note.

A metadata record says there was one undifferentiated copy of this edition. The source can still carry a bookseller's ticket, a rebinding, a signature, a library stamp, or a repair that makes this copy historically distinct.

The surviving source therefore does more than preserve additional information. It preserves correction authority.

Destroy it and that authority changes hands.

The derivative is no longer one representation among others that can be checked against a richer object. It becomes the best surviving evidence by default. Whatever the capture process omitted is now harder to distinguish from whatever the source never contained.

That is how ignorance becomes structural.

Before destruction, the honest statement is:

We did not capture this property.

After destruction, the honest statement may become:

We did not capture this property, and the source from which it could have been measured is gone.

Those sentences describe very different epistemic states. The first contains a repair path. The second contains a closed door. The door matters even when nobody currently wants to walk through it.

Destructive capture spends that option.

This is where the language of efficiency becomes dangerous.

A destructive scanner can be faster than a nondestructive one. Removing bindings can increase throughput. Standardizing pages can simplify feeding. Discarding the originals can reduce storage, handling, and logistics costs.

All of that can be operationally true.

But throughput has no competence to determine informational redundancy.

A machine optimized to move books through a capture line as quickly as possible is answering a question about cost and speed. It is not thereby answering whether the physical sources have exhausted their evidentiary value.

Those are different fields.

The mistake is allowing the first decision system to inherit authority over the second. This is why an irreversibility gate belongs between successful extraction and source destruction.

The gate asks a question the production line cannot answer for itself:

What justifies making this loss of access permanent?

That is not a demand that every physical object be preserved forever.

Archives discard things. Libraries deaccession things. Records programs authorize destruction. Preservation has always involved appraisal because preservation resources are finite and objects differ in value, rarity, context, replaceability, and use.

The important part is that destruction is treated as its own governed decision. It is not automatically granted by successful digitization.

The National Archives provides a useful contrast. Its current federal digitization regime does permit agencies, under defined conditions, to destroy some source records after compliant digitization.

But the permission does not arise because a scanner produced files. The framework requires records-management controls, technical standards, quality management, validation, and an approved disposition authority.

  • For temporary records, NARA says the validation process must check that the digital versions capture all information contained in the source records, including associated materials such as envelopes, cards, and sticky notes, and that the digital versions can serve the same purposes as the sources.
  • For permanent records, NARA's quality-management guidance tells agencies to compare source and digitized records, account for all records, inspect related notes and other media, document missing pages, and preserve relations to mixed-media material that cannot be digitized.

That is already a more serious theory of destruction than:

the OCR finished.

More importantly, NARA explicitly recognizes a category called intrinsic value. Some records have physical qualities or characteristics that make the original form itself archivally significant.

NARA's current guidance says records with intrinsic value must be transferred in their original form and may not be destroyed simply because they have been digitized. Its older intrinsic-value guidance makes the principle even clearer: all original physical records possess qualities and characteristics that copies do not preserve; some possess those qualities to a degree that requires keeping the originals.

That sentence should stop the entire conversation for a moment.

A national archival institution charged with deciding when copies can replace originals begins from the premise that copies do not preserve every quality of the original.

Of course they do not. The difficult work comes afterward.

  • Which qualities matter enough to preserve the object?
  • Which can be allowed to disappear?
  • Who has authority to make that judgment?
  • What evidence has to exist before the judgment becomes irreversible?

Federal records rules do not govern private artificial-intelligence companies by analogy. The narrower lesson is stronger anyway: mature preservation practice already knows that successful digitization and authority to destroy are different decisions.

That leaves a governance problem we will return to: who may convert successful extraction into irreversible source loss, under what evidence, and with what record of the uncertainty being accepted?

Before returning to authority, the more basic claim had become testable. If derivatives really contract distinctions available in their sources, we should be able to keep a richer source, run it through a pipeline, and watch the contraction happen without destroying anything.


So Modal Path Ethics Went Through the Machine.

So, at this point, the argument had become testable.

That was inconvenient.

It is easy to say that a derivative carries less than its source.

It is easy to point at a destroyed book and list everything a text file cannot tell you about paper, binding, ownership, repair, wear, annotation, production, or copy history.

The harder question is whether the contraction can be watched while it happens.

  • Can we keep a richer source in hand, run it through an ordinary text-preparation route, and identify the exact distinctions that stop surviving?
  • Can we do that without pretending that every omitted property should have been stuffed into the final token stream?
  • Can we keep the provenance honest enough to say what the derivative carries, what an upstream object still carries, and what disappears if that upstream object is gone?

There was an obvious practical problem.

The books Amazon and Anthropic had already cut apart were unavailable.

Modal Path Ethics could not rerun the capture. It could not inspect the discarded copies. It could not ask whether a particular destroyed volume contained a marginal note, a repair, a binding distinction, an insert, a bookseller's mark, or some other feature that never entered the digital record.

That is the irreversibility problem from the previous section doing exactly what it says on the box.

So Modal Path Ethics used a source it controls.

It put Modal Path Ethics: The Extance Strategy Game through the machine.

There is an important limit here.

  • The framework did not begin with one physical paperback.
    • It began with the production PDF.

That means the experiment starts after many of the richest physical distinctions discussed earlier have already fallen outside the source domain. The PDF cannot tell us how one printed copy was handled, whether somebody wrote in it, whether its binding was replaced, whether the paper carries a copy-specific stain, or what happened to it on a shelf.

That makes this a narrower test. If the argument requires a coffee stain to work, we have a bad argument right here.

The source I tested was already digital, already highly standardized, already designed to travel between machines, and already much closer to the sort of object a training pipeline would like to receive.

It still was not a text file.

The 213-page source contained 17,849 addressable text spans across 41 observed fonts. It contained 1,796 raster-image occurrences on 50 pages. It contained 56,380 vector drawing objects across 185 pages. It carried page geometry, render color, font size, outline structure, page identity, case distinctions, images, designed relations among objects, and the identity of the PDF itself.

Those counts are not an argument that all 56,380 vector objects belong in a language model's token stream.

They establish something simpler.

There was more there to lose.

Then the first surprise arrived before I had built any exotic transformation at all.

I asked three ordinary extraction routes for the text of the same PDF.

  • Poppler's default pdftotext produced 352,340 Unicode characters in 6,600 lines.
  • Poppler with its layout-preserving option produced 390,387 characters in 7,718 lines.
  • PyMuPDF produced 720,003 characters in 18,288 lines.
    • Same file.
    • Three answers to what is the text?

The largest answer was more than twice the smallest.

That does not establish that PyMuPDF is correct and Poppler is wrong.

It establishes that the phrase the text of the PDF was already hiding an implementation choice.

The derivative had a genealogy before anyone had even begun cleaning it.

On page 8, one opening sentence appears once in Poppler's default output and twice in PyMuPDF and pypdf output.

The source inspection explains why. The PDF contains large numbers of overlapping or stacked text-layer groups: 8,692 such groups were detected across 212 pages. Some carry the same lexical material in nearly the same place while differing in rendering details.

A human looking at the page sees one designed page.

A text extractor sees objects.

Different extractors decide differently what counts as the lexical derivative.

Already, provenance has work to do.

It is not enough to write:

Source: this PDF.

The relevant statement is closer to:

This derivative was produced from this exact PDF by this extraction implementation, with these options, under this versioned route.

Otherwise a later researcher can reproduce the source and still fail to reproduce the text.

Then page 2 gave us the cleanest joke in the experiment.

The page contains an illustration.

Underneath it is the credit:

Illustration above by Ellie Rose Lawson.

Plain-text extraction keeps that sentence. It removes the illustration.

It removes the page geometry that establishes what above refers to.

The resulting derivative contains perfectly legible words describing a relation that no longer exists inside the derivative.

Nothing is misspelled. Nothing has to be hallucinated.

The extraction can be excellent at extracting lexical content and still be incapable of carrying the relation asserted by the lexical content.

This is a useful thing to stick on because it exposes the limits of the word text.

  • The sentence is there.
  • The information needed to resolve the sentence is not.

Page 7 improves the joke by making the source complicit.

The book warns readers that it contains many visual elements and asks them to try to avoid looking at them while reading.

The text pipeline accepts this challenge with admirable seriousness.

  • It preserves the warning.
    • Then it removes the visual field.

Page 16 is less funny. Machine inspection found 128 vector drawing objects on that page. The plain-text derivative carries the surrounding prose. It does not carry the vector content.

Again, that is not a bug report against plain text.

Plain text is doing what plain text does.

The error would come later, if somebody called the derivative the book and forgot what the representation had stopped carrying.

At this point, the article broke containment and turned into another computer-science project.

Modal Path Ethics needed a way to state the problem without saying the useless sentence information was lost every time a representation changed.

Almost every useful representation leaves something out.

The question had to be property-specific.

If two source states differ in the property we care about and the transformation maps them into the same tested representation, we have a concrete contraction witness on our hands.

Modal Path Ethics has called the resulting benchmark TraceLossBench.

This framework could no longer point at a transform and call it lossy in the abstract. Loss had to be stated relative to a property, a source domain, an implementation, and a tested stage.

It could no longer treat a provenance edge as proof that the child still carried everything present in the parent. It could no longer treat a retained metadata value as though it were the source object.

It could no longer treat failure to find a witness as proof that no contraction existed outside the domain we actually tested.

And, very importantly, it could no longer count a test fixture I deliberately wrote to delete a property as a discovery.

During development, Modal Path Ethics used mechanical stages that intentionally dropped things like typography, geometry, render color, and page identity. Those fixtures are useful for checking that the benchmark can localize a known event and attach evidence correctly.

Modal Path Ethics put that trap there. Finding the trap is a unit test.

So the substantive book results were kept separate.

The first category was direct audit.

  • The PDF carries font face and size.
    • The plain-text derivative does not.
  • The PDF carries render color.
    • The plain-text derivative does not.
  • The PDF carries spatial geometry.
    • The plain-text derivative does not.
  • The PDF carries raster and vector content.
    • The plain-text derivative does not.
  • The PDF carries page identity.
    • The flattened corpus payload eventually does not,
      • although we can retain page support in a sidecar.

That last clause is important. A property can disappear from the training payload while remaining recoverable somewhere upstream.

Those are different states.

The payload does not carry page identity. The provenance record can still tell us which pages supported a chunk. The source can still be inspected if it survives.

The derivative therefore has at least three questions hanging over it:

  • What does this representation carry now?
  • What can still be inspected upstream?
  • Does the richer upstream thing still exist?

Do not collapse those questions.

The second category was a real pairwise contraction in the actual route.

The source contains the strings Modal Path Ethics and MODAL PATH ETHICS in different places.

Those are distinct case forms.

  • They remain distinct through the upstream text preparation.
    • Then the route casefolds the corpus.

At that boundary, the distinction disappears.

That is arguably not a scandal.

Lowercasing or casefolding can be a perfectly sensible preprocessing decision.

The finding is simply exact: this distinction survived until here, and it stopped surviving here.

That is much better than calling the entire pipeline lossy and walking away. It gives somebody something they can inspect, accept, reject, change, or regression-test later.

Modal Path Ethics also ran a negative control.

The route used Unicode NFKC normalization, a transformation that can collapse compatibility distinctions in some text.

Across all 7,816 distinct non-empty source-span strings observed in this book, this moral metaphysics found no NFKC collision group.

That result is less dramatic than finding one.

It is also more important than it looks. A system designed only to collect examples of loss will eventually teach its operator to shop for losses.

An ethical framework tries transformations until something breaks, print the breakage, and announce that the world is ending.

That is not an audit.

So the result was recorded exactly as it was:

NO_COLLISION_IN_OBSERVED_SPAN_DOMAIN.
  • Not NFKC preserves the book.
  • Not NFKC is safe.
    • Modal Path Ethics tested this declared domain and found no witness there.

That is the whole claim.

Then Modal Path Ethics sent the book further down the route.

  • Page-aware text extraction.
  • Unicode normalization,
  • line-end dehyphenation,
  • and whitespace collapse.
  • Flattening into a corpus document while keeping page support on the side.
  • Casefolding.
  • Fixed 2,048-character chunks with 256-character overlap.
  • SentencePiece BPE tokenization with a 2,000-piece vocabulary.

The final route produced 197 chunks and 98,230 tokens.

That number is where a conventional training-data description might begin to feel satisfied.

This book has become a corpus artifact.

It can be counted. It can be chunked.

It can be tokenized. It can be fed downstream.

TraceLossBench asks the annoying follow-up:

What had to stop being distinguishable for this representation to exist?

And then throws in another:

Where would we go if we needed the distinction back?

For some properties, selected metadata is enough to answer a narrow question.

For some, an upstream text derivative is enough.

For the illustration, vector content, full typography, geometry, and the complete designed object, the richer source still matters.

That forced the experiment to include a source-survival gate. The distributable test package does not contain the full source PDF. It contains its cryptographic hash and selected evidence.

That means the package is forbidden from saying that the source can be fully recovered from the benchmark bundle.

  • If somebody has the exact source,
    • the hash can bind the experiment back to it.
  • If nobody has the source,
    • the hash is a beautiful identifier for something that is gone.

That may be the cleanest miniature of the whole book-destruction problem.

Provenance metadata can remember that a richer source existed. It cannot become the source by remembering it accurately. A checksum cannot display the illustration. A lineage edge cannot reveal an annotation that was never captured. A database row saying physical book cannot open the physical book.

The source-survival state has to remain separate from the lineage state.

By the end of the experiment, Modal Path Ethics had reproduced the central structure of the original problem without destroying anything.

A rich source entered a sequence of transformations. Useful derivatives came out. Properties disappeared from current carriage at different stages. Some remained inspectable upstream. Some had selected evidence preserved in sidecars.

The final token payload was real, useful, reproducible, and profoundly incomplete as a description of the object from which it descended.

That still left an uncomfortable objection.

Of course my own book produced examples.
  • I chose the book.
  • I built the benchmark.
    • I declared the properties.
    • I selected a training-style route.

A hostile reader could reasonably ask whether the entire demonstration was too convenient. Perhaps this framework had built an elaborate machine for finding the distinctions Modal Path Ethics already knew its own source contained.

Fair enough.

So it stopped looking at that book.

Modal Path Ethics took the same question to an independently existing training corpus and an independently released tokenizer.

It did not tell that pipeline which distinction to erase.

Then it erased one anyway.


Then I Stopped Trusting The Convenient Example.

My book was suspiciously cooperative.

It had illustrations. It had typography.

It had page geometry, vector drawings, stacked text layers, visual jokes, case distinctions, and a sentence that literally says Illustration above before a text extractor removes the thing above it.

If somebody wanted to accuse Modal Path Ethics of finding exactly the sort of loss it had gone looking for, they had material. A hostile reviewer could reasonably look at this whole thing and say:

Congratulations on discovering your own experiment.

That is a serious objection. A test becomes much more interesting when the machine surprises the people testing it.

So I moved to an independently existing pretraining corpus and an independently released tokenizer. No page-layout trap. No illustration selected because I already knew plain text would remove it. No transformation called drop_typography. No hand-built source pair designed to collide.

The source family was C4, the large web corpus introduced with Google's Text-to-Text Transfer Transformer, better known as T5.

The tested tokenizer was the released google-t5/t5-small tokenizer.

This framework pinned the dataset revision. It pinned the tokenizer revision.

Modal Path Ethics took the first 10,000 English training records without shuffling.

Then it enumerated every whitespace-delimited span of at most 64 Unicode code points.

That produced 3,590,652 eligible span occurrences representing 252,003 distinct source strings.

Now the question was brutally simple.

Can this independently existing route make two observed source strings indistinguishable?

And if it can, where does that happen?

The first boundary was the tokenizer's own executable normalizer.

Across the 252,003 distinct eligible source strings, 1,678 changed under that normalizer.

When Modal Path Ethics grouped the observed source strings by the normalized string they became, it found 94 collision groups.

One of them was:

m²

and

m2

Both forms actually occurred in the scanned corpus.

These are different Unicode strings.

  • The second character in m2 is ordinary digit two, U+0032.
  • The second character in m² is superscript two, U+00B2.

Unicode classifies them differently. The superscript form also carries a compatibility decomposition pointing back to the ordinary digit.

The released T5 normalizer maps both strings to:

m2

After that, both receive the same T5 token sequence:

[3, 51, 357]

There it was.

Nobody on this side invented a fake drop_superscript transformation.

Nobody mutated m2 into m² to manufacture a witness.

Both strings were already present in the corpus.

The released implementation performed the contraction.

TraceLossBench found it.

This result is easy to mishandle, so this framework should immediately refuse the dramatic version.

  • I have not shown that T5 "doesn't understand square meters."
  • I have not shown downstream model behavior.
  • I have not shown that every distinction between m² and m2 matters for every purpose.

In many contexts, those two forms are intended to express the same unit.

The claim is smaller and cleaner:

two distinct observed source representations became one representation at a specific tested boundary.

That is exactly what the benchmark was built to detect.

And it was not alone.

The same scan found naturally occurring pairs such as:

  • ™
    • and
      • TM
  • ⓒ
    • and
      • c
  • …
    • and
      • ...
  • first
    • and
      • first

Again, these examples do not all carry the same significance.

A ligature collapsing into its ordinary letters may be harmless for one use and important for another. A circled symbol may be decorative in one source and a meaningful category marker in another.

The benchmark does not get to settle that argument by counting code points. Its job is earlier. It tells us that the distinction existed, that the running implementation stopped carrying it, and where the tested contraction occurred.

That would already have repaired much of the problem with the book experiment.

Then the second boundary produced the example that made the larger article much harder to dismiss.

After normalization, the scan still contained 251,906 distinct strings.

Modal Path Ethics then asked the released T5 tokenizer for the actual token-ID sequence of each normalized string.

This produced 258 collision groups among strings that were still distinct immediately before tokenization.

The cleanest witness contained three source spans:

±20
<20
~20

All three occurred in the bounded C4 sample.

The normalizer left all three alone.

  • ±20 remained ±20.
  • <20 remained <20.
  • ~20 remained ~20.

So the earlier stage had preserved the distinction.

Then the tokenizer emitted exactly the same T5 token sequence for all three:

  • [3, 2, 1755]

The middle ID, 2, is the unknown token.

Three different relation symbols entered that boundary.

The representation available on the other side carried the same unknown token before the same representation of 20.

These expressions are not interchangeable as expressions. Look at them.

  • Plus or minus twenty.
  • Less than twenty.
  • Approximately twenty.

Those relations can disagree about the world.

  • A tolerance of ±20 is not a threshold of <20.
  • A threshold of <20 is not an estimate of ~20.

The input strings remained distinct through the first tested stage. The token-ID boundary contracted them.

This is the sort of result the book example could not give us on its own.

The source was not mine. The tokenizer was not mine. The strings were not mine. The contraction mechanism was not inserted for the benchmark.

Modal Path Ethics asked an existing route what distinctions it could still represent.

It answered.

There is another discipline worth preserving here.

The raw Stage-2 scan found 258 collision groups.

Some were ugly. They involved replacement characters, controls, invisible formatting, whitespace artifacts, or other cases that are perfectly legitimate implementation findings and absolutely miserable explanatory examples.

After the scan, the framework applied an interpretability filter for presentation. It removed replacement-character cases, Unicode control/format/surrogate/private/unassigned code points, empty strings, strings longer than 24 code points, and forms differing only in surrounding whitespace.

That left 68 readable groups.

All 68 contained the unknown token.

The filter happened after the scan.

So the research record says that.

  • The machine result is 258 groups.
  • The 68-group subset is a post-scan presentation subset.

This may sound like procedural housekeeping. It is part of the argument.

If the whole project is about refusing to let derivatives impersonate their sources, then its own summary cannot impersonate the experiment.

The selected examples are derivatives too. They need provenance.

Modal Path Ethics also needed a control.

A collision in one representation does not establish that the source distinction was impossible to carry into machine input.

So this framework ran the highlighted forms through a different released tokenizer: google/byt5-small.

ByT5 works at the byte level.

For the examples above, it produced different byte sequences.

m² and m2 remained distinguishable.

So did:

  • ±20
  • <20
  • ~20

This does not crown ByT5 the Good Tokenizer. It does not establish that byte-level tokenization is globally superior. It does not tell us which representation produces the best trained model.

It establishes the point we actually needed:

the contractions were properties of the audited route, not logical necessities of turning these strings into model inputs.

Another representation could carry them. That makes the word loss more precise.

We are not staring at an unavoidable law of computation.

We are looking at a representational decision embodied in a running implementation.

Sometimes the decision is explicit. Sometimes it arrives through a normalizer inherited from a tokenizer stack.

Sometimes it arrives because a symbol falls outside the vocabulary and becomes <unk>.

Sometimes nobody notices because the pipeline still runs, the tensor shapes are correct, and training begins on schedule.

The data has not disappeared. Something narrower has happened.

A distinction available upstream is no longer available in the tested downstream representation.

That is exactly the kind of event that disappears when the whole chain is flattened into the sentence:

"We trained on C4."

No.

A model was presented with derivatives produced from C4 through particular executable transformations, under particular software versions, with particular representational limits.

That statement is longer. Reality often is.

And now the physical-book case comes back into view.

The warehouse and the tokenizer operate at very different scales, on very different objects, for very different purposes.

The evidentiary structure is the same.

  • A richer state enters.
  • A useful derivative comes out.
    • Some distinctions survive.
    • Some collapse.
    • Some move into side information.
    • Some remain recoverable only because a richer upstream source still exists.
    • And some become impossible to revisit once that source is gone.

The ±20 example is especially useful because the original source remains available in principle.

If we distrust the tokenization, we can go back.

We can inspect the observed string. We can replay the normalizer or the tokenizer. We can test a different representation. We can identify the first tested boundary where the distinction disappears.

That is what a repair path looks like.

Now, imagine running the same logic backward into the book warehouse.

The text derivative does not preserve a binding feature.

  • Go back to the book.

The OCR misses handwriting.

  • Go back to the book.

A page image obscures a watermark.

  • Go back to the book with another instrument.

A future historian asks a question the present scanning project never anticipated.

  • Go back to the book.

Unless somebody destroyed it after deciding that the derivative had captured everything worth carrying.

Then the provenance graph can tell us where the derivative came from. It can tell us the source once existed. It can tell us exactly which scanner produced which file. It can preserve a checksum with immaculate precision.

It cannot reopen the book.

This is why the C4/T5 experiment belongs in a story that began with bindings being cut from physical volumes.

The experiment does not prove that every preprocessing contraction is harmful. It proves that contractions occur in ordinary, independently existing machinery without announcing themselves as philosophical events.

They happen because representations have boundaries.

  • A normalizer has a contract.
  • A vocabulary has a boundary.
  • An extractor has a field of view.
  • A scanner has sensors.
  • A corpus schema has columns.
  • A training payload has a form.

Every one of these can be useful.

Every one of them is smaller than the world that entered it.

Once this ethical framework saw the same structure in an external corpus, the original question changed.

Modal Path Ethics no longer needed to ask whether the book-to-training path contained losses.

Of course it did.

The harder problem was keeping those transformations from collapsing into one administrative fiction called the data.

  • The book became an image.
  • The image became text.
    • The text became normalized text.
    • The normalized text became chunks.
      • The chunks became tokens.
      • The tokens became model input.

Each object had a parent. Each transition had its own preservation profile. Each derivative could support some claims and fail others.

So the next task was to stop talking about the training data as though one object had simply traveled intact from shelf to model.

It had become a family tree. Modal Path Ethics needed to keep the relatives straight.


The Book Became a Sequence of Derivatives.

There is a phrase doing an irresponsible amount of work in artificial intelligence:

the training data.

This sounds singular. Stable.

Like one object entered a warehouse, passed through several machines, and arrived at the model wearing a smaller hat.

That is not what happened.

The book did not travel intact from shelf to token. It became relatives.

  • A physical copy can become a stack of severed leaves.
  • The leaves can become page images.
  • The page images can become OCR text.
  • The OCR text can become cleaned text.
  • The cleaned text can become normalized text.
  • The normalized text can become chunks.
  • The chunks can become token IDs.
  • The token IDs can become batches presented during training.

Every transition produces a new object with a new representational contract.

Some properties survive. Some are transformed. Some are moved into metadata. Some disappear from current carriage while remaining recoverable upstream. Some disappear and cannot be recovered because the richer source is gone.

If we draw that whole history as one arrow labeled BOOK -> TRAINING DATA, we have performed an administrative magic trick.

The intermediate objects vanish. The decisions vanish with them. The losses vanish with the decisions.

Then the final derivative inherits the name of the entire ancestry.

This is how a token sequence gets introduced as though it were the book.

It is not the book.

It has a family resemblance.


Provenance Is a Family Tree, Not a Soul Transfer

Data provenance is the machinery for remembering where a derivative came from and what happened along the way.

That is exactly the machinery this problem needs. It also has one very important limit.

A perfect family tree does not make the grandchild identical to the grandparent.

Suppose a chunk of normalized text points cleanly back through every transformation that produced it.

We know which source record supported it. We know which extractor, normalizer, and tokenizer version ran. We know which training artifact descended from the result.

Excellent.

Now ask whether that chunk still carries the original page geometry.

The lineage cannot answer yes just because the ancestry is complete.

Ask whether the final token sequence can distinguish ±20 from <20 after the tested T5 boundary.

Again, the provenance edge does not restore the distinction. It can tell us where to look for the richer form.

That is enormously valuable. Still a different proposition.

This distinction became one of the central rules of TraceLossBench:

Lineage tells us where a derivative came from. It does not automatically tell us which distinctions the derivative still carries.

That sentence sounds painfully obvious once written down. So did the book was more than the text.

We still built warehouses around forgetting it.


Keep the Questions Separate

Once the family tree is visible, several questions that were previously compressed into one word start refusing to cooperate.

  • What does the current derivative carry?

Can this text record still determine the source script? The typography? The page? The relation symbol? The image association? The annotation? The answer has to come from the current representation and whatever metadata accompanies it.

  • What can still be inspected upstream?

Perhaps the current payload no longer carries page identity, but the provenance path resolves to a page-aware derivative that does. Perhaps the token sequence no longer distinguishes two source forms, but the normalized string survives one stage earlier. Upstream inspectability is useful precisely because current carriage can fail.

  • What was the transformation supposed to preserve?

A pipeline may declare that it preserves lexical content while discarding typography. Good. Write that down. A declared contract is evidence about design intent. It is not experimental proof that every execution satisfied the contract.

  • What did we actually test?

One replayable witness does not certify millions of other records. An exhaustive search over a finite declared domain says something stronger than one selected pair. A bounded scan says what it says within the bound. The evidence has a scope.

  • Does the richer source still exist?

A lineage edge pointing toward a destroyed source and a lineage edge pointing toward an inspectable source are graphically similar and epistemically very different.

  • Can a correction propagate?

If we discover that a source record was wrong, corrupted, misidentified, withdrawn, or unlawfully included, can we identify the descendants that depend on it? Can we invalidate them? Can we rebuild them?

And finally:

Did any of this actually influence the trained model?

That is another question again.

Construction lineage can establish that an artifact entered a training pipeline. It does not, by itself, establish what a trained model learned from that artifact, whether a particular distinction altered its internal representations, whether the model memorized anything, or whether any downstream behavior depends on it.

Those require separate causal or behavioral evidence.

This is where provenance systems can become victims of their own success.

Once we have a beautiful graph, there is a temptation to let the graph answer every question because the graph is beautiful.

Do not do that to the graph. It has enough work.


Two Parents Can Still Mean One Ancestor

The family-tree metaphor gets more useful when we stop imagining clean trees.

Training corpora contain copies, translations, transliterations, mirrors, revisions, reposts, derived files, synthetic transformations, and datasets assembled from datasets assembled from datasets.

A record can have multiple immediate parents without having multiple independent origins.

We encountered this directly in a provenance stress test.

A Cyrillic source and a Latin-script counterpart can appear downstream as two supporting records.
  • Count the parents and the system looks redundant.
  • Lose one and the other appears to remain.

Then trace the genealogy.

If the Latin form was generated from the Cyrillic form, the two-parent structure collapses to one source root.

The apparent redundancy was derivative redundancy. One origin failure can still remove the whole branch.

This matters for books too.

Imagine five OCR files, three cleaned corpora, two tokenized datasets, and six training shards all descending from the same destroyed physical copy.

A dashboard can proudly report sixteen surviving artifacts.

There is still one lost root.

Counting descendants is not the same as preserving origins. This is one reason provenance needs genealogy rather than a pile of filenames.

The system has to know which objects were independently acquired and which objects were manufactured from one another.

Otherwise replication becomes multiplication by copy-and-paste.

A thousand mirrors can preserve availability. They do not create a thousand independent witnesses to what the original object contained.


Correction: a Graph Operation

Now give the family tree a problem.

Suppose OCR turns a date into the wrong year.

The mistake is discovered later by inspecting the surviving page. If the lineage is intact, we can ask which cleaned records descended from that OCR output.

Which chunks inherited it? Which tokenized artifacts contain it? Which corpus versions need to be regenerated? Which published derivatives should be marked as superseded?

That is correction support. The correction does not require pretending that the original mistake never happened. It requires knowing where the mistake traveled.

Now change one fact.

The page was destroyed after scanning.

The graph can still show every descendant with exquisite precision.

What it cannot do is provide the evidence that would have revealed the OCR error in the first place.

Source survival and correction support are therefore related and separate.

  • A surviving source without lineage can be reexamined,
    • but its downstream descendants may be difficult to locate.
  • A perfect lineage graph without the source can identify descendants,
    • but some disputes about the source can no longer be resolved.

You want both. This becomes even more important when the issue is larger than a typo.

  • A rights holder withdraws material.
  • A dataset entry is discovered to have the wrong provenance.
  • A digitization is found to have omitted marginalia.
  • An edition was misidentified.
  • A transformation version is discovered to contain a systematic normalization defect.
  • A source previously treated as independent turns out to be copied from another source.

Every one of those is a different correction event.

Every one asks the graph a different question.

Which descendants inherit this problem?

That is why a provenance record has to bind evidence to the actual transformation event and its actual scope.

A witness involving one source pair cannot become a universal indictment of every item processed by the same tokenizer.

A declaration attached to a transform cannot become proof that every record preserved the declared property.

A correction attached to one source root should not invalidate unrelated branches.

The graph needs discipline because the alternative is bureaucratic contagion. Everything becomes either clean or contaminated at dataset scale.

Reality is usually more local than that.


The Checksum != the Book

Suppose we record the source beautifully: title, edition, acquisition event, scanner, timestamp, transformation versions, a cryptographic hash, and every derivative that followed. Then we destroy the source. Have we preserved it?

No. We have preserved excellent information about it.

A checksum is an excellent name tag and terrible resurrection technology. It can identify a surviving object. It cannot render an illustration from a file nobody retained, reopen a binding, recover a marginal note the scan never captured, or answer a future question whose instrument did not exist when the source disappeared.

Provenance must never impersonate retention. The graph can say that source X existed, was hashed as Y, and was destroyed after event Z. Those are valuable facts. The last one is valuable precisely because it records the point at which one class of repair stopped being possible.

Stop Calling the Whole Family “The Data”

At this point the phrase the training data becomes almost unusable unless we specify which member of the family we mean.

The acquired source objects? The page images? The OCR text?

The cleaned corpus? The normalized records? The deduplicated version?

The chunks? The token IDs? The batches actually presented to the model?

These objects can differ in content, structure, legal status, provenance, accessibility, fixity, correction state, and representational capacity.

Some can survive while others are deleted.

Some can be redistributed while others cannot.

Some can answer questions that others cannot.

Some are evidence about the source.

Some are descendants of evidence about the source.

Calling all of them the data is convenient right up until somebody asks what was lost. Then convenience becomes concealment.

The repair is simple enough to state.

Keep the relatives straight.

A derivative should know its parents.

A transformation should have an identity and a version.

Evidence about a contraction should attach to the event where it was observed.

The system should distinguish what the current object carries from what can still be inspected upstream. It should know whether the source survives. It should know how corrections propagate. And it should refuse to turn construction ancestry into a claim about model influence.

This does not make lossy preprocessing forbidden.

Lossy preprocessing is often the entire point.

  • A tokenizer exists to produce tokens.
  • An OCR system exists to produce machine-readable text.
  • A crop exists because somebody does not need the entire image for the immediate task.

The discipline is to stop letting the useful derivative inherit every authority of the richer source. The token sequence is allowed to be a token sequence. The OCR text is allowed to be OCR text. The page image is allowed to be a page image.

The book is allowed to remain the book.

And if we choose to destroy the root after producing the branches, that choice cannot hide inside the word digitized.

A provenance graph can record the destruction perfectly. It cannot undo it.

So the next question is not how to draw the family tree.

We can draw it.

The next question is who gets to sever the root.


The Irreversibility Gate.

A provenance graph can tell us where the root is.

It cannot decide whether we are allowed to cut it off.

That decision belongs somewhere else.

This is the point where the problem stops being a question about scanners, OCR, tokenizers, metadata, or lineage and becomes a question about authority.

Somebody has to decide when a source object may disappear.

That decision is often disguised as an ordinary part of processing.

  • The book is scanned.
  • The files pass validation.
  • The pages enter the corpus.
  • The physical copy is discarded.

From the perspective of the production line, those events can look like one continuous operation.

They are not.

  • The first events create derivatives.
  • The last event destroys a repair path.

Those actions need different standards of evidence.

Call the boundary between them the irreversibility gate.

The gate is a deliberately inconvenient question inserted between:

We got what we came for.

and:

Therefore the source can go away.

The first statement can be true while the second remains completely unproven.

A digitization project can succeed at its stated capture objective and still know almost nothing about properties outside that objective. An OCR transcript can be excellent OCR and still contain no assessment of a binding. A page image can be an excellent page image and still tell us little about paper composition, pressure marks, inserted objects, or three-dimensional structure. A textual corpus can be immaculate text and still be unable to answer whether a particular copy carried handwritten corrections.

The fact that the derivative is good at being a derivative gives it no automatic authority to certify the source as redundant.

That authority has to come from somewhere else.


Destruction: a Claim

Destroying the source makes a factual claim about the future.

It says that whatever questions can still matter about this object can be answered adequately without having this object.

Sometimes that claim may be defensible.

It is still a claim.

  • A library deaccession decision makes it.
  • An archive disposition schedule makes it.
  • A records-management program makes it.
  • A laboratory makes it when it consumes a specimen during testing.
  • A company makes it when it destroys a physical source after digitization.

The governing mistake is to let the claim disappear inside the mechanics of the workflow.

Once destruction is treated as a routine cleanup step, the burden of proof reverses.

The source has to justify its continued existence.

Everything the current pipeline did not measure becomes invisible to the decision because invisibility was produced by the current pipeline.

That is circular.

The source is judged informationally empty by an instrument that was never designed to inspect the information being discarded.

Then the failure to observe those properties becomes the reason they are safe to destroy.

An irreversibility gate breaks that circle. It asks the project to state, before destruction, what proposition it believes it has established.

Not:

  • We have a scan.

Not:

  • We have the text.

Not even:

  • The corpus ingested successfully.

The proposition has to be closer to:

For the purposes under which destruction is being authorized, the surviving derivatives, retained exemplars, other accessible copies, and recorded uncertainties are sufficient to accept the loss of this source object.

That sentence is deliberately harder.

It makes the decision reveal its dependencies.

  • What purposes?
  • Which derivatives?
  • Which properties were inspected?
    • Which were outside the capture contract?
  • What other copies exist?
    • Are they actually equivalent at the object level?
  • What remains uncertain?
    • Who is accepting the uncertainty?
  • What future repair options disappear?

A source should not vanish because nobody forced the workflow to answer those questions.


The Gate Needs a Capture Contract

The first thing the gate needs is a truthful description of what the capture process attempted to preserve.

This is the capture contract.

The phrase does not mean a legal contract. It means the explicit boundary of the representation.

If the objective is searchable text, say searchable text.

If the objective is high-resolution page imagery, say page imagery.

If color fidelity was validated, record that.

If page order was preserved, record that.

If annotations were included only when visible to the imaging setup, say that.

If bindings, paper composition, edge marks, inserts, embossing, depth, pressure, smell, residue, or copy-specific physical evidence were never assessed, say that too.

The honesty of the contract matters more than its size.

  • A narrow derivative with an honest contract is useful.
  • A narrow derivative presented as though it exhausted the source is dangerous.

This is the same discipline TraceLossBench forced on the preprocessing pipeline. A transformation is allowed to collapse distinctions. The benchmark asks which declared distinctions survive and where tested contractions occur.

The preservation analogue is straightforward.

A capture process is allowed to select. The gate asks what it selected, what it tested, what it did not test, and what remains upstream.

This is also why the word digitized is too coarse to carry the decision by itself.

Digitized how?

At what resolution? Through which instruments? With what spectral range? Under what lighting? Which sides of which objects? With what geometric support? Which associated materials? Which metadata? Which validation? Which omissions?

A binary field called digitized = true can hide a very large epistemic hole.

The gate needs the hole described.


Unknown: a Real State

The gate also needs an answer category that production systems hate:

UNKNOWN.

A pipeline wants pass or fail.

A disposition decision wants retain or destroy.

Reality frequently arrives with less cooperation.

  • Maybe the binding was never examined.
  • Maybe nobody knows whether the copy contains erased pencil marks.
  • Maybe the page images were validated for legibility but not for color fidelity.
  • Maybe the corpus contains two copies described as duplicates, but their copy-specific histories were never compared.
  • Maybe an unusual insert was discarded before cataloging.
  • Maybe another copy exists somewhere, but its condition and provenance are unknown.

Those are not preservation successes. They are uncertainties.

The gate has to carry them as uncertainties rather than laundering them into absence.

This is one of the most important lessons from the technical work.

When TraceLossBench cannot test an intermediate boundary because the required contract is undefined, the correct result is not preserved.

The correct result is unresolved.

When a bounded search finds no witness, the result is not proof that no contraction exists anywhere. It is a bounded non-finding.

The same discipline should govern source destruction.

If nobody inspected a property, the disposition record should not imply that the property was absent.

If nobody established copy equivalence, the record should not imply interchangeability.

If reacquisition is assumed possible, the assumption should be written down and tested against reality.

Unknown is not administrative failure.

Unknown is information about the state of our knowledge.

Suppressing it is the failure.


Replaceable Is Not Identical

Replaceable is one of the easiest words to misuse around books. A title can be replaceable. An edition can be replaceable. A particular copy may still carry a history another copy cannot restore.

The gate therefore cannot ask only whether the book can be bought again. It has to ask what reacquisition would recover. Another copy may restore the published words while failing to restore the destroyed copy's annotations, ownership, repairs, wear, binding, or circulation history.

Derivative multiplicity does not solve this either. Ten mirrors of one scan are excellent protection against losing the scan. They are no protection at all against discovering something the scan never captured.


Who Gets a Vote?

Then comes the institutional question.

Who is allowed to say yes?

The answer should not be generated automatically by the objective that benefits from destruction.

This is the operator-independence problem from earlier in the investigation.

If the capture operator is measured by throughput, the source looks like queue depth. If the storage team is measured by cost, the source looks like square footage. If the training-data team is measured by usable tokens, the source looks like an inconvenient pre-token state.

If the project is measured by time-to-corpus, careful appraisal looks like delay.

None of those objectives is illegitimate.

None of them establishes source redundancy.

The irreversibility gate therefore needs authority capable of representing interests that disappear from the production metric.

That can be an archivist, a records officer, a preservation specialist, or a collection policy. It can be a documented rule requiring escalation for rare, unusual, annotated, fragile, historically significant, or poorly characterized sources. It can be a sampling regime that retains exemplars where total retention is impossible.

The institutional form will vary.

The structural requirement is simpler:

the system should contain a role whose success does not improve when the source disappears faster.

That role does not need veto power over every disposal forever. It needs enough independence to force the destruction claim into the open.

What are we discarding?

What survives?

What was never assessed?

What is uniquely bound to this copy?

How reversible is the decision?

What uncertainty are we accepting?

Who is accountable for accepting it?

Those questions are friction.

Good. Irreversible actions are exactly where useful friction belongs.


The Gate Needs a Destruction Ledger

If the answer is still yes, the destruction should leave a record.

This sounds obvious until we distinguish a provenance record from a disposition record.

  • The provenance graph tells us that derivative B came from source A through transformation T.
  • The destruction ledger tells us that A ceased to be available after a deliberate decision at time D, under authority R, after evidence E was considered and uncertainty U remained.

That difference matters later.

Suppose an OCR defect is discovered five years afterward.

The graph can identify descendants.

The ledger can tell investigators whether the physical source survived long enough to be rescanned, whether it was intentionally destroyed, what capture standard justified destruction, and which uncertainties were known at the time.

Suppose a historian later discovers that a class of handwritten marks had systematic significance.

The ledger can identify which source objects were destroyed before those marks were understood and which were retained.

Suppose a preservation standard changes.

The ledger can separate old decisions made under one capture contract from later decisions made under another.

Suppose a company claims that the physical sources were redundant.

The ledger creates something better than retrospective confidence. It creates an auditable record of what redundancy meant when the decision was made.

At minimum, the destruction event should remain linked to the source identity, the derivatives claimed as substitutes, the capture and validation procedures, the disposition authority, the date, the known exceptions, and the unresolved properties that were accepted as losses.

The surviving record should never pretend that the source still exists.

That sounds trivial. It is not.

Digital systems are very good at making absent things look present.

A title remains in the catalog.

A thumbnail remains in the interface.

A checksum remains in the database.

A transcript remains searchable.

A tokenized descendant remains in the training corpus.

The representation can continue to participate in the system long after the richer object is gone.

The ledger needs one brutally simple fact:

SOURCE DESTROYED.

Not hidden in a retention note.

Not inferred from a missing shelf location.

Not euphemized as processing complete.

Destroyed.

The word matters because correction capacity changed at that moment.


The Gate Can Say Yes

An irreversibility gate that can never authorize destruction is not a gate. It is a prohibition wearing administrative clothing.

That is not what we need.

Physical preservation has costs.

Some sources are genuinely abundant. Some are fully characterized for the purposes that matter to the governing collection.

Some records are scheduled for lawful destruction after validated digitization. Some materials cannot reasonably be retained indefinitely.

Some destructive tests produce knowledge that cannot be obtained otherwise.

A serious framework has to survive those cases.

The gate can say yes. It just cannot get to yes by confusing extraction success with source equivalence.

The strongest yes is one that can state what is being sacrificed.

  • We know what the source is.
  • We know what derivative survives.
  • We know what the derivative was validated to preserve.
  • We know what was outside the capture contract.
  • We know whether other source objects genuinely substitute for this one at the relevant level.
  • We know which uncertainties remain.
  • We know who accepted them.
  • We know what correction options close after destruction.
  • We have retained enough lineage to identify descendants if a later problem is discovered.

Then the destruction is at least an intelligible decision.

The distortion audited in this Tale is the opposite condition.

A system wants one property intensely enough that every other property becomes administratively weightless.

Text is valuable.

So the book becomes text.

Tokens are valuable.

So the text becomes tokens.

Throughput is valuable.

So the source becomes waste.

At each step the new object is useful.

At each step usefulness makes it easier to forget what changed.

The irreversibility gate exists to force the system to remember the change before the richer object disappears.


A Better Failure Mode

The goal is not perfect foresight.

Perfect foresight is unavailable.

The goal is a better failure mode.

When we fail to anticipate a future question, a surviving source gives the future somewhere to go.

When retention is impossible, an honest capture contract tells the future what the derivative can and cannot be expected to answer.

When destruction is authorized, a destruction ledger tells the future where the repair path ended and why.

When a later defect is discovered, lineage tells us which descendants need attention.

When the evidence is incomplete, the system is allowed to say unknown.

That is a much stronger epistemic position than pretending that whatever survived the pipeline was all that ever mattered.

That is the better failure mode: preserve a route back when that route is worth preserving; state the capture boundary honestly when it is not; record the destruction when the route is closed; and let unknown remain an actual state of knowledge.

Now we can return to the books that started this investigation and state the ruling without making the scanner carry a philosophy it never earned.


The Book Was More Than the Text.

We can state the distortion plainly now.

The book was more than the text.

That is not a sentimental defense of paper.

It is not a claim that every printed copy of every book must survive forever.

It is not an argument against digitization, OCR, search, corpora, tokenization, or artificial-intelligence training.

It is a claim about identity.

A physical book is an information-bearing source object.

Text is one property it carries.

A page image is a derivative of some of its visible surfaces.

OCR is a derivative of that derivative.

Normalized text is another.

Chunks are another. Tokens are another.

A training representation at the end of that chain can be enormously valuable without becoming the object at the beginning.

That distinction should have been boring.

Instead, industrial systems made it expensive.

They acquired books because the books contained information worth extracting.

They optimized the extraction.

Everything between the Las Vegas cutter and this point has been evidence for that one distinction. Material bibliography showed that copies can carry histories their printed words do not. Preservation science showed that every transformation has a scope. Our own book showed that even a production PDF contains structure that plain text does not carry. The independent C4/T5 audit showed that ordinary preprocessing can collapse observed distinctions at specific executable boundaries. Provenance showed that ancestry, present carriage, source survival, correction support, and downstream model influence are different questions.

None of those results makes derivatives suspect. They make derivatives legible.


The Distortion Was a Promotion

The deepest error in this case was not information loss by itself.

Every representation loses something.

Every measurement selects.

Every model leaves features out.

Every archive makes decisions.

Every transformation has a scope.

The distortion was the promotion of a scoped derivative into the identity of its source.

Once that promotion happened, the rest followed cleanly.

  • If the book is the text, the binding is packaging.
  • If the book is the scan, properties outside the scan are externalities.
  • If the data is the normalized text, distinctions removed by normalization become invisible to later stages.
  • If the token sequence is treated as though it simply is the text, contraction at tokenization disappears from the story.
  • If a provenance record says only that the current object came from the source, lineage can be misread as evidence that the current object still carries everything anyone might care about from the source.

Each promotion makes a richer object easier to forget.

The repair is not to forbid derivatives. Derivatives are how large information systems work. The repair is to preserve the direction of derivation.

The source may produce the derivative.

The derivative may be optimized for a particular task.

The derivative may become the main working object for that task.

The derivative may even outlive the source.

It still does not get to rewrite its own ancestry.

That rule travels far beyond books.

  • A photograph is not the event it depicts.
  • A transcript is not the conversation.
  • A laboratory result is not the specimen.
  • A database row is not the person or object it describes.
  • A compressed image is not every property available in the source image.
  • A model input is not every distinction available in the upstream data object.

The general danger appears whenever a system becomes so good at operating on a representation that it forgets the representation is a representation.

The book case is unusually vivid because somebody eventually puts a blade through the source.

The blade makes the ontology visible.


The Ruling.

The books in this Tale were not destroyed because their information had no value.

They were destroyed inside processes built around the extraordinary value of their information.

The failure did not begin with contempt for information. It began with a narrower definition of information than the source could support.

The pipeline knew what it wanted.

Text. Searchability. Machine-readable language.

Training material.

Throughput.

Those were real objectives.

Then the objectives acquired sovereignty over the source.

  • What the instrument captured became what the object contained.
  • What the instrument did not capture became administratively weightless.

The derivative inherited the source's name.

The source became an inconvenience.

And an inconvenience can be optimized away.

This is the ruling.

A source object may contain information outside the representation a current pipeline is designed to extract.

A successful derivative does not prove source equivalence.

A bounded preservation test does not prove universal preservation.

A lineage edge does not prove current property carriage.

A surviving identifier does not reconstruct a destroyed source.

A second copy does not automatically reproduce the first copy's history.

And destruction is not a neutral continuation of extraction.

Destruction changes the future state of inquiry.

It closes questions. It removes correction paths.

It converts some unknowns from we did not measure this into we cannot return to this object and measure it now.

That transition deserves its own authority, its own evidence, and its own record.

The alternative is absurd enough to fit comfortably inside a Tale of Distortion.

  • Acquire an object because it carries valuable information.
  • Build an expensive system to extract that information.
  • Define the information as whatever the system extracts.
    • Destroy the object that could tell you what that definition missed.
    • Keep the derivative.
    • Call the job complete.
      • Then let the future discover what your representation forgot.

There is a cleaner way.

Keep the source reachable when the unmeasured field remains open and retention is justified.

When the source cannot be kept, say exactly what survives and what does not.

Keep ancestry explicit. Keep capture claims scoped. Keep uncertainty visible.

Keep correction paths attached to the descendants that depend on them.

And before severing the root, require somebody to make the actual claim destruction entails.

We know enough about what we are giving up to accept that we will never be able to ask this object another question.

Sometimes that claim will survive scrutiny.

Sometimes it will not.

The scanner does not get to answer it for us.

The tokenizer does not get to answer it for us.

The throughput metric does not get to answer it for us.

The current use case does not get to answer it for every future use case.

That is the boundary this Tale was trying to find from the beginning.

The instrument is allowed to take what it can measure. It is not allowed to redefine everything else out of existence.

They wanted the information in the books so badly that they destroyed information in the books to get it faster.

The last mistake was believing the blade had cut away only packaging.

It had cut away questions.

The book was more than the text.