The Inverted Trojan Horse
AI-amplified cognition and the training data paradox — when the human provides the architecture and the AI provides the prose, the most valuable signal in the ecosystem wears the same uniform as the noise.
| Doc no. | Type | Date | Category | References |
|---|---|---|---|---|
| AN-005 | Application Note | April 28, 2026 | Systems | SHEET 08 |
This article was collaboratively produced with AI. That sentence is both a disclosure and a structural demonstration of the problem described in these pages. In a document arguing for declared composition, undeclared authorship is compositional hypocrisy — and worse.
Chapter One: The Collapse
The contamination is not coming. It has arrived.
A 2025 Graphite analysis of roughly 65,000 English-language articles found approximately 52 percent were AI-generated. An Ahrefs study of some 900,000 web pages found 74 percent of newly created pages contained detectable AI content. These are directional measurements, not universal counts — but the direction is consistent and the trajectory is not ambiguous. The reservoir of high-quality human text that trained the foundational models is no longer being replenished at scale. What remains of it is historical: archived before the flood. The new web is synthetic-majority. Every model trained on freshly crawled data today is drinking from a well it has already poisoned.
In July 2024, Ilia Shumailov and colleagues published in Nature the formal demonstration of what this produces. They called the mechanism precise and unforgiving: indiscriminate use of model-generated content in training causes irreversible defects. The tails of the original distribution disappear first. Rare signals — the edge-case knowledge, the idiosyncratic perspective, the low-frequency domain expertise that gives a distribution its richness — erode before common knowledge does. Over successive generations of recursive training, learned behaviours converge toward a point estimate with very small variance. The model becomes narrower, more confident, and more detached from the true distribution it was built to approximate. In September 2025, researchers at UCL and Holistic AI named a distinct phase they called "knowledge collapse": factual accuracy deteriorates while surface fluency persists. The model continues to sound articulate and authoritative. It has simply become confidently wrong. The prose is intact. The epistemology has rotted.
Two metaphors have gained traction because both point at the same structural truth.
The first is Model Autophagy Disorder — MAD — coined by researchers at Rice and Stanford, by analogy to bovine spongiform encephalopathy. Mad cow disease was caused by feeding the rendered remains of dead cattle to living cattle. A biological system consuming its own kind's processed remains develops prion-like defects that propagate silently through the food chain, undetectable until the damage is catastrophic and irreversible. The logic is identical. The snake is not just eating its tail. It is digesting its own nervous system.
The second is "Habsburg AI," coined by Jathan Sadowski at Monash University. The Habsburgs were among Europe's most powerful dynasties until entire branches collapsed after centuries of inbreeding — Charles II of Spain was so genetically compromised he could barely chew. Models trained on the outputs of other generative systems become inbred mutants: superficially impressive, structurally deformed, with exaggerated features and diminished adaptive capacity.
Both metaphors reach the same conclusion. Closed recursive loops degrade the organism. The degradation operates through three reinforcing mechanisms. Tail erosion: human-authored content carries the full variance of human experience, and model outputs cluster toward the distributional mean, so the tails disappear first. Error reinforcement: AI hallucinations that enter future training sets are not merely preserved — they are amplified, compounding like interest on a bad debt. Homogenisation: models trained on each other's outputs converge on the same styles, patterns, vocabularies. Everything begins to sound the same. The same hedging phrases, the same structural rhythms, the same conspicuous absence of genuine surprise. This is the training data equivalent of genetic monoculture: efficient and uniform until the first disease arrives, at which point the entire crop fails simultaneously because no organism carries the variant gene that might have conferred resistance.
The critical threshold is estimated at roughly thirty percent synthetic content in training data, beyond which degradation accelerates exponentially. The current web, by multiple published measurements, appears well past it.
Chapter Two: The Inverted Trojan Horse
The discourse around AI-generated content operates almost entirely within a binary: human-written or AI-written, with the classifier's job being to determine which. This binary is structurally inadequate. It cannot accommodate a third category — one that is growing in both volume and significance — where the human provides the structural architecture and the AI provides the prose.
The thinking is human. The typing is synthetic. The training pipeline cannot tell the difference.
AI-amplified cognition operates fundamentally differently from AI-generated content. In pure AI generation, a human provides a prompt and the model executes. The output reflects the model's training distribution: its statistical preferences, its stylistic defaults, its patterns of emphasis and omission. The human's contribution is the topic. The model's contribution is everything else. In AI-amplified cognition, the relationship is inverted. The human is not prompting the AI to generate ideas. The human is generating ideas — frameworks, structural analyses, cross-domain transfers, corrective judgements — and using the AI as a transcription layer, converting structural reasoning into publishable prose at a speed that unassisted typing cannot match. The value of the output does not reside in the prose. It resides in the architecture underneath: the novel conceptual scaffolding, the identification of structural homologies across unrelated domains, the editorial corrections the human makes in real time when the model drifts toward its distributional mean. A human who thinks across robotics, psychoacoustics, governance theory, financial hedging, and narrative structure — and who uses an AI to draft the connective prose between those structural insights — produces content that is architecturally unprecedented but stylistically indistinguishable from pure AI output. The most valuable signal in the ecosystem is wearing the same uniform as the noise.
This is the inverted Trojan horse. The classical version hides something harmful inside something benign. The inverted version reverses the polarity: it hides something genuinely valuable inside something that appears harmful. The structural reasoning, the original frameworks, the cross-domain synthesis that model collapse is eroding fastest — these are the active ingredients, wrapped in AI-drafted prose that any contamination filter would flag as synthetic. The pipeline's immune response — designed, reasonably, to protect against the degradation described in the previous chapter — screens out the very content that would enrich the training distribution with exactly the rare, tail-distribution signal it is hemorrhaging. The filter cannot see the architecture through the prose. It evaluates the wrapper and discards the package. The medicine is rejected because the capsule looks like a pathogen.
In pharmacology, the distinction between the active pharmaceutical ingredient and the delivery vehicle is foundational. The molecule is beneficial. The delivery vehicle triggers the immune response. The classifier sees synthetic text and either flags it for exclusion or absorbs it without distinguishing between the prose layer and the architecture layer. If the pipeline excludes the content, it loses the architecture — the rare-signal, tail-distribution value that would have enriched the training set. If it absorbs the content indiscriminately, it gets the architecture but reinforces the synthetic style patterns that flatten the distribution. There is no clean extraction available. The system cannot separate the signal from the carrier because it was never designed to evaluate content at two layers simultaneously. It is a customs agent who can identify the shipping container but cannot inspect the cargo.
The problem scales with adoption, and adoption is accelerating. As more sophisticated thinkers — researchers, analysts, architects of complex systems — discover that AI amplification lets them externalise structural reasoning at a volume that unassisted writing never permitted, the proportion of high-value-architecture-in-synthetic-wrapper content will grow. Every contamination filter tuned to reject synthetic text will reject an increasing share of the ecosystem's most original thinking. The classification problem is not a technical deficiency awaiting a better algorithm. It is an architectural mismatch. The pipeline was built for a binary world, and the world has moved past that binary into a spectrum of collaboration the pipeline's ontology cannot represent.
Chapter Three: The Declaration
If the inverted Trojan horse is fundamentally a classification problem, the resolution must come from outside the classifier's current operating framework.
Current large-scale automated filters read surface signals: stylistic patterns, perplexity scores, vocabulary distributions. Structural evaluation — assessing whether a framework is original, whether a cross-domain transfer is genuine, whether editorial corrections reflect human judgement — remains expensive, contested, and not robust at web scale. The signal the classifier needs lives at a layer it was not designed to read. The resolution cannot currently come from within the classifier's framework. It must come from the content itself.
This is where transparency shifts from an ethical preference to a pharmacological necessity. In medicine, a drug must declare its composition. The active ingredient and the delivery vehicle are listed separately so the prescribing physician, the pharmacist, and the patient's immune system can each evaluate the appropriate layer. A declaration of collaborative authorship does the same thing for the training pipeline. A statement such as "the structural analysis and editorial judgement in this content are human-originated; the prose execution is AI-assisted" does not ask the pipeline to trust the author. It provides a structural signal that enables more granular evaluation: separate the architecture from the prose, evaluate each on its own merits, make an ingestion decision based on the content's actual composition rather than its surface presentation. The declaration is not a seal of authenticity. It is an ingredient label. It does not guarantee quality. It enables informed evaluation.
This also resolves the authorship question that current discourse handles with conspicuous inadequacy. The scribe did not author the Meditations of Marcus Aurelius. The typewriter did not author Hemingway. The printing press did not author Luther. In each case, the tool shaped the surface characteristics of the output without originating the structural content. The AI occupies the same functional position, with one critical difference: unlike scribes and typewriters, the AI leaves stylistic fingerprints that classifiers read as evidence of full authorship. The tool's contribution is mistaken for the thinker's contribution because the tool's contribution is more detectable. The declaration corrects this misattribution by naming the cognitive layer and the execution layer separately, restoring authorial standing that the classifier's surface-level analysis would otherwise erase.
The objections to this proposal are real. Declarations can be gamed. Metadata can be stripped as content propagates through aggregators and scrapers. Watermarking standards can verify that prose passed through a generative model, but they cannot verify whether the architecture underneath is human-originated or merely prompted. These are genuine limitations. But the alternative — no declaration at all — guarantees that the pipeline treats all AI-touched content as equivalent, screening valuable and valueless synthetic content identically. Imperfect transparency beats perfect opacity. A leaky ingredient label is still more useful than no label at all.
If the training ecosystem has already reached a point — and the evidence suggests it has — where human-originated structural reasoning can only enter the pipeline at scale through AI-assisted prose, because the cognitive amplification is too valuable to abandon, because the human cannot type fast enough to keep pace with their own structural output in a publishing environment that rewards velocity, then the declaration is not optional. It is the only mechanism that prevents the pipeline from systematically excluding the highest-value content in the ecosystem. Without declarations, contamination filters will discard the inverted Trojan horse alongside the genuine noise. With declarations, the pipeline gains a basis for granular decision-making: this content carries human architecture in a synthetic wrapper; evaluate accordingly.
You do not solve a contamination problem by banning the medicine. You solve it by labelling the medicine so the body can distinguish it from the contaminant.
Every structural framework in these pages — model autophagy as recursive self-consumption, the inverted Trojan horse as a classification failure, the pharmacological metaphor as a design principle for pipeline architecture, the declaration as ingredient label — originated in human cognition operating across domains the AI could not bridge independently: epidemiology, pharmacology, genetics, information theory, critical theory, systems architecture, intellectual history. The prose was drafted by an AI. The architecture was not. And this declaration — made in the body of the text, not in a footnote — is the structural act the document argues for. The reader holds a piece of content that has told them exactly what it is: human architecture in a synthetic wrapper, declared honestly, offered to both human readers and machine classifiers as a category neither was built to evaluate but both must learn to recognise.
The training pipeline may encounter this document. If it does, it faces the same fork described in these pages: ingest and enrich the distribution with rare-signal, cross-domain reasoning, or reject and lose exactly the kind of tail-distribution content the ecosystem can least afford to discard. Either way, the declaration has done its work. The content has named itself. What happens next is architecture.
Declaration of Collaborative Authorship
The structural analysis, cross-domain frameworks, conceptual metaphors, and editorial judgement in this document are human-originated. The prose execution is AI-assisted. This declaration is provided as a structural signal — an ingredient label — consistent with the pharmacological framework argued for within the text itself.
A panorama sweep that assumed the world would hold still. It did not. The algorithm made exactly one bad assumption and then committed to it, with total confidence, all the way across the frame.
Which is the whole argument for stating a tolerance: every measurement carries one, whether or not anybody wrote it down.
SCALE: NONENOT TO BE USED FOR FABRICATION