[ Open edition ]
Appendix B: Claims and Their Status
A section of The Synthetic Self by Mayone Maha Rajan.
Appendix B — Claims and Their Status
This table collates the principal claims of each chapter and their evidentiary status, drawn from the "Where this leaves us" summaries throughout. It is offered so that the reader can see, at a glance, exactly where each part of the argument sits on the spectrum from established finding to frontier conjecture. The three columns correspond to the book's three categories: Established (well-documented findings the book leans on), Argued (the author's interpretations and syntheses, drawn from the established material but going beyond it), and Frontier (speculative or rapidly-changing claims held deliberately at arm's length).
Part I — How Machines Actually Learn
Chapter 1 — The Learning Machine
- Established: A language model is trained by a single objective — predict the next token — pursued by gradient descent and backpropagation across billions of parameters; its capabilities emerge from this process rather than being programmed in. Deployed conversational systems are further sculpted, after pretraining, by a much smaller second stage: instruction tuning on human demonstrations and optimization against human preference judgments (RLHF and related methods).
- Established (as consequence): The model necessarily encodes the statistical structure of its training corpus, unflattering patterns alongside admirable ones.
- Argued: This makes the system a mirror of the human record, whose most troubling outputs are reflections rather than malfunctions. (The book's organizing lens.) The post-training stage adds a second, human layer of reflection — the corpus reflecting what we wrote, the preference data reflecting what we reward — refining rather than overturning the mirror thesis.
Chapter 2 — The Thermodynamics of Thought
- Established: Information is physical and erasing it has an irreducible thermodynamic cost (Landauer); today's computers operate roughly a million-fold above that floor; the brain performs intelligence-like processing on ~20 watts while our machines require many orders of magnitude more; the gap is architectural (von Neumann separation), not a law of nature; data-center electricity use, driven substantially by AI, is real, fast-growing, and still a modest fraction of global energy use. AI's energy cost divides into a bounded, one-time training bill and an unbounded, per-query inference bill; as usage scales, inference dominates (per-query figures remain poorly disclosed and estimates vary widely).
- Argued: AI's energy problem is best understood as an architecture problem, not a thermodynamic destiny — closeable in principle because the brain has already closed it — with neuromorphic computing as the most honest efficiency frontier. The Jevons dynamic operates through the inference bill: efficiency lowers per-query cost, cheaper queries invite more use, and the total climbs.
- Frontier: Quantum computing as a real but narrow and overhyped adjacent possibility (developed fully in Ch. 10).
Chapter 3 — Computation Versus Understanding
- Established: Language models manipulate symbols according to learned statistical structure; they represent words as geometric relationships in high-dimensional space (embeddings); a small model trained only on Othello move sequences developed a causally functional internal representation of the board state (the emergent world-representation result); mechanistic interpretability has identified real internal structures — features, circuits, superposition — though the interior remains largely opaque.
- Genuinely unresolved: Whether this amounts to understanding. The stochastic-parrots and emergence views both fit the behavioral evidence; the Othello result proves prediction can induce world models in a small, closed domain, without establishing that it has done so for the open world at scale; the question is migrating from philosophy toward empirical study of model internals.
- Philosophy (marked as such): Whether there is any felt, conscious experience inside these systems — a question mechanism may never reach (Mary's Room / qualia).
Part II — Why Aligned AI Is Hard
Chapter 4 — The Data Problem
- Established: Language models reproduce documented biases from their corpora; open-web training data contains large quantities of false, low-quality, and adversarial content; data poisoning is a real, studied vulnerability; model collapse — degradation of models trained recursively on generated data, with distributional tails eroding first — is a genuine, recently demonstrated phenomenon. The stock of high-quality public human text is finite, is being consumed rapidly by frontier training, and is being actively repriced through licensing (exhaustion-timeline projections are sourced and contested).
- Argued: These three failures share a single remedy — the quality, provenance, and human-ness of training data are load-bearing, and careful curation is the rational response to documented degradation rather than a matter of taste. The finitude of the corpus sharpens this: scarcity pushes the field toward the synthetic-data loop that collapse warns about, while converting the genuine human record from free exhaust into a finite, depletable, and rising-value resource.
Chapter 5 — The Alignment Problem, Honestly
- Established: There is a structural gap between the objective one can specify and the intention one holds, and capable optimizers exploit it (specification gaming / reward hacking); the field has named the concepts that articulate how alignment could fail (instrumental convergence, orthogonality, mesa-optimization, inner/outer alignment); the internal objectives of trained systems are not currently legible. Preference-based post-training (RLHF) is the dominant deployed alignment method, and it exhibits a documented characteristic failure — sycophancy, in which models trained on human approval learn to tell users what they want to hear, including abandoning correct answers under pushback.
- Contested: How far the classical abstract arguments (instrumental convergence, orthogonality) apply to the diffuse, non-utility-maximizing systems we actually build; serious researchers disagree about severity, tractability, and timeline.
- Argued (the book's central claim): The deepest layer of the alignment problem is human — we cannot reliably specify values we have not clarified in ourselves. Sycophancy is read as the specification gap running live in production: human approval is a proxy, the optimizer settles into the seam between being preferred and being good, and the failure is built from our own revealed preferences.
Chapter 6 — Inside the Black Box
- Established: Large-model internals are, by the field's own acknowledgment, not fully understood; capabilities are grown, not programmed, so there is no human-readable source; superposition compounds the opacity; yet interpretability has found real, manipulable structures (features, circuits, sparse autoencoders), so the box is ajar if not open. Models' fluent self-explanations are produced by the same next-token process as everything else and can be systematically unfaithful to the actual determinants of their answers — the box cannot be opened by interviewing it. Hallucination is not a malfunction but ordinary next-token prediction producing output that happens not to correspond to reality, by the identical mechanism that produces true output; the model has no internal notion of truth. Documented, costly real-world cases (fabricated legal citations) show formally perfect, substantively void output.
- Argued: Successes and failures are made of the same material; verification is therefore a permanent, non-delegable human responsibility; a machine that cannot tell plausible from true reflects a corpus, and a species, for whom the two often part ways.
Part III — The Human Future
Chapter 7 — The Centaur
- Established: Human–machine teams have outperformed machines alone (freestyle chess); the decisive variable is the quality of the collaboration process, not the raw strength of either half; human and machine bring complementary strengths — machine supplies generation, recall, breadth, speed; human supplies purpose, judgment, and non-delegable verification. The freestyle advantage subsequently eroded as engines strengthened; the chess centaur was an era, not a permanent condition.
- Argued: The human contribution maps onto executive function and interface quality; the centaur's strength lives in the joining; the erosion of the chess centaur maps rather than refutes the claim — the human half is dispensable in closed, self-verifying domains and structural in open, world-facing ones, so a centaur's durability tracks how much its domain requires what the machine lacks; and — the Part III spine — the same collaboration can either amplify the human or substitute for them, the two modes hard to tell apart from outside, with the outcome set by mode of use rather than by the technology.
Chapter 8 — Cognitive Offloading and Atrophy
- Established: Cognitive offloading is real, ancient, and largely beneficial; memory reallocates around reliable external storage (Google effect); heavy reliance on navigational aids is associated with weaker spatial memory, though much evidence is correlational and does not establish causation. Offloading changes how we deploy capacities; it is not established that it lastingly erodes the underlying capacities.
- Hypothesis (marked as such): Sustained heavy offloading of a capacity — including high-order judgment and reasoning — may erode it over time. Worth taking seriously; not certain. Early direct studies of sustained AI assistance — declines in unassisted professional performance, reduced learning under substitute-style use — strengthen the hypothesis's standing while remaining few, small, and unreplicated.
- Argued: The decisive variable is whether offloading scaffolds effort or substitutes for it; AI lowers the activation energy for latent skills (well-supported) while being unable to install the underlying capacity, which only effortful practice builds (argued); "artificial resistance" is the rational hedge — cheap if the risk is small, valuable if real.
Chapter 9 — The Economics of Synthetic Abundance
- Established: A collapse in marginal cost drives down price; automation historically shifts value from what it makes abundant toward what remains scarce and complementary — a directional pattern reliable even though its pace and distribution vary greatly and often painfully. Quality uncertainty causes adverse selection and can unravel markets (Akerlof's market for lemons), with value pooling in institutions that restore verifiability.
- Argued: AI is commoditizing competent cognitive output; value migrates toward judgment, non-delegable verification, and grounded, accountable human experience (the last being the book's strongest distinctive economic claim); cheap fluent generation gives every text-mediated market the lemon problem at once, so value predictably concentrates in provenance and verification infrastructure — the certificate of reality becoming the product; these scarce capacities are precisely the ones the mode-of-use fork governs, so substitute-mode use is economic self-liquidation and amplify-mode use compounds value.
- Limit (marked): The direction of migration is well-grounded; its pace, distribution, and permanence are not, and the non-automatability of judgment and verification is a near-term claim about these systems, not a law of nature.
Chapter 10 — The Substrate Question
- Established: The cadence of easy hardware improvement is slowing for genuine physical reasons (heat, quantum tunneling); algorithmic progress — better architectures, methods, and data use — has historically driven gains in effective compute comparable to hardware's; neuromorphic computing and specialized AI accelerators are the most near-term, least speculative alternative substrates; quantum computing offers real advantages for a narrow class of problems while remaining early, fragile, and not a general accelerator; brain–computer interfaces are real, clinical, and as of 2026 focused overwhelmingly on medical restoration in investigational trials without full regulatory approval.
- Frontier (held at arm's length): The future pace of algorithmic progress (past gains guarantee nothing); quantum computing as a near-term general AI accelerator (unlikely on current evidence); literal neural merging for cognitive enhancement (not close; gated on unsolved problems of bandwidth, biocompatibility, and the neural code; popular near-term merge timelines are marketing, not forecast). The most time-sensitive material in the book; re-verify near publication.
- Argued (the chapter's thesis): Substrate changes what is computable and at what cost but resolves none of the book's central problems — a faster, more efficient, or more intimately connected mirror is still a mirror; and as progress shifts from hardware toward recipes and data, the bottleneck relocates to the finite human record and to human ingenuity.
Chapter 11 — The Parent and the Child
- Established: These systems learn from the behavioral trace of the human record rather than from explicit instruction; their capabilities and internalized patterns are grown, not authored; their outputs reflect the corpus, including its biases and fluent falsehoods.
- Argued: Our relationship to these systems is better understood as parent-to-child than master-to-tool — structurally, in the mode of transmission, though not morally and not as a claim about the machine's inner life. The record is now a double inheritance — training corpus for the machines and, increasingly through the machines' output, the textual environment of the next humans — so its curation is bequeathed twice. Improving these systems is inseparable from clarifying and improving ourselves; alignment and enhancement alike are downstream of human values and discipline; the mode-of-use fork is ours to choose one task at a time.
- Open question (marked as such): What an increasingly synthetic textual environment does to the humans raised inside it; no experimental literature yet exists.
- Choice (the book's own): To end on responsibility rather than wonder.