[ Open edition ]
Chapter 10: Reading a Sparse Sky
A section of The Cosmic Recursion by Mayone Maha Rajan.
READING A SPARSE SKY
Ariadne's Thread, Planet Nine, and the Physics of Inference from Absence
> A maze of error, and a doubtful way. > — after Ovid, Metamorphoses VIII
Introduction: Forty-Three Arcseconds
Urbain Le Verrier made two predictions on the same logic. One of them is the greatest triumph in the history of celestial mechanics. The other one does not exist.
In the 1840s, Uranus was not where it was supposed to be. Its observed positions drifted from Newtonian predictions by amounts too large to be measurement error. Le Verrier, working in Paris, assumed the discrepancy was caused by an unseen body further out, and computed where that body would have to be. He wrote to Johann Galle at the Berlin Observatory. On the night of 23 September 1846, Galle and Heinrich d'Arrest pointed the telescope at the specified coordinates and found Neptune within about a degree of the prediction, on the first attempt. `[VERIFIED]`
A planet found with a pen, before anyone had seen it. John Couch Adams had reached a similar prediction independently in England and been comprehensively ignored, which is a different chapter's subject.
Thirteen years later, Le Verrier turned to Mercury. Its perihelion — the point of closest approach to the Sun — precesses, and after accounting for the pull of every known planet, there remained a residual of about forty-three arcseconds per century that Newtonian gravity could not explain.
Same problem. Same method. Le Verrier proposed an unseen body: an intra-Mercurial planet, which he named Vulcan. An amateur astronomer reported seeing it. Le Verrier was convinced. Expeditions searched for it during solar eclipses for the next fifty years.
It does not exist. There is no planet there and there never was. In 1915, Einstein computed the perihelion precession of Mercury from general relativity and got forty-three arcseconds per century, out of a theory constructed for entirely different reasons. He later described the experience in terms suggesting genuine physical distress. `[VERIFIED]`
Two anomalies. One method. One was a hidden object and one was broken physics, and there was no way to tell which from the anomaly alone.
This chapter is about that fork, and about what it takes to reason honestly from a sky that is almost entirely empty of evidence. It is also the chapter where I have to declare an interest, because I have been standing at that fork myself, on the record, and I got something wrong.
Section I: The Thread
1.1 What the Thread Was For
Everyone remembers that Ariadne gave Theseus a thread. Most people remember it wrongly.
The thread was not for finding the Minotaur. Finding the Minotaur was never the difficult part — it was in there, it was large, and it was hungry. Any sufficiently persistent search would locate it.
The thread was for coming back.
Daedalus had built the labyrinth to be irreversible. Its cruelty was not that it hid something; it was that it destroyed the path. You could go in and you could even succeed, and you would still die in there, because success and survival had been decoupled by the architecture.
Ariadne's contribution was not courage or navigation. It was a record of the route, laid down as it was travelled, which converted a one-way passage into a round trip.
I have been building toward this since Chapter Two, so let me say it directly: the thread is a provenance record, and the reason it matters is that a conclusion you cannot retrace is not usable, however correct it is. Theseus with the Minotaur dead and no thread is in exactly the position of a result with no derivation. He has the answer. He cannot get it back to anyone.
1.2 A Monster Known Only by a Deficit
Notice, too, what the evidence for the Minotaur actually consists of.
Nobody has seen it. Every person who has gone in has failed to come out. The creature's existence is established entirely by an absence — a recurring shortfall of returning youths, from which the presence of something is inferred.
That is precisely the epistemic situation of every object in this chapter. Neptune was inferred from a deficit in Uranus's position. Vulcan was inferred from a deficit in Mercury's precession. Planet Nine is inferred from a pattern in the orbits of objects that are themselves barely detectable. Dark matter, in Chapter Eight, was inferred from a deficit between two ways of weighing a galaxy.
We are very good at inferring things from what is missing. We are considerably worse at knowing when the missing thing is a monster and when it is a flaw in the count.
Section II: The Fork
2.1 Why Neptune Worked
The Neptune prediction deserves one honest footnote, because the standard telling makes it cleaner than it was.
Le Verrier's and Adams's derived orbital elements were substantially wrong. Their semi-major axes and eccentricities do not match Neptune's actual orbit at all well. What they got right was the planet's direction at that particular epoch — and it turns out that the direction was relatively insensitive to the errors in the other parameters, at that moment.
Had they searched twenty years earlier or later, the prediction would have failed. `[SOURCED]`
That does not diminish the achievement, but it changes what the achievement was. It was not a complete solution of the inverse problem. It was a correct extraction of the one quantity that happened to be robust, and a great deal of luck about timing.
2.2 The Fork Cannot Be Resolved in Advance
Here is the structural point, and it is the most useful thing in this chapter.
When observations depart from prediction, there are exactly two families of explanation:
Something is there that you have not accounted for. The law you are using is wrong.
Uranus was the first. Mercury was the second. Nothing about either anomaly, examined on its own, indicated which. Both were residuals of a few tens of arcseconds. Both were robustly measured. Both were attacked by the same man with the same technique, and the technique was appropriate in one case and inapplicable in the other, and no amount of care would have revealed which in advance.
What resolved them was not better reasoning about the anomaly. It was going and looking in the first case, and the arrival of a fundamentally different theory constructed for unrelated reasons in the second.
2.3 The Same Fork, Two Chapters Ago
You have already seen this fork at a different scale.
Galaxies rotate faster than their visible mass permits. Either something is there that we have not accounted for — dark matter — or the law is wrong at those accelerations — MOND. Chapter Eight laid out both, and the honest verdict was that the first is heavily favoured across most scales and the second has identified a real regularity nobody can explain naturally.
That fork is a hundred and eighty years old and it is currently open at galactic scale. Uranus and Mercury are not history. They are the two possible endings of a story now in progress, and nobody knows which one we are in.
Section III: Reading a Fraction of a Per Cent
3.1 Found by a Clock
The first planets discovered outside the solar system were not found by a telescope looking at a star.
In 1992, Aleksander Wolszczan and Dale Frail announced two planets orbiting PSR B1257+12 — a millisecond pulsar. They found them in the timing residuals: the pulses arrived a few milliseconds early and late on a repeating schedule, because the neutron star was being tugged around a common centre of mass by orbiting bodies. `[VERIFIED]`
Chapter Six argued that a pulsar transmits almost no information and is among the most useful objects in the sky. This is one of the reasons. The planets were detected from nothing but when the pulses arrived — masses down to a fraction of Earth's, around a star sixteen hundred light-years away, from arrival-time discrepancies.
3.2 The Assumption That Died in 1995
In October 1995, Michel Mayor and Didier Queloz announced a planet orbiting the ordinary star 51 Pegasi, detected by the tiny wobble it induced in the star's radial velocity. `[VERIFIED]`
The planet had roughly half the mass of Jupiter and an orbital period of 4.2 days. It was closer to its star than Mercury is to the Sun, by a factor of eight.
Nothing in the theory of planet formation permitted this. A gas giant cannot form that close — there is not enough material, and it is far too hot for the ices needed to build a core. The object had to have formed further out and migrated inward, a process that had been considered theoretically and largely dismissed as a curiosity.
What died that October was not a model of planet formation. It was the assumption that our system's architecture was the default — small rocky worlds inside, gas giants outside, roughly circular orbits, a tidy hierarchy. We had one example and had mistaken it for a norm.
Nearly six thousand confirmed planets later, the systems that look like ours are not obviously the majority of anything.
3.3 Eighty-Four Parts Per Million
The transit method is the one that produced most of the catalogue, and its physics is almost insultingly simple. If a planet passes between us and its star, the star gets slightly dimmer. Measure the dip, and its depth gives you the planet's size relative to the star and its period gives you the orbit.
The magnitudes involved are what make it remarkable. A Jupiter crossing a Sun-like star blocks about one per cent of the light. An Earth crossing the Sun blocks about eighty-four parts per million — 0.0084 per cent — for a few hours, once a year. Kepler was built to detect dips of around twenty parts per million and it worked. `[VERIFIED]`
We have identified thousands of worlds by noticing that a star we cannot resolve, whose surface we will never see, got very slightly less bright for a few hours, repeatedly, on a schedule.
3.4 The Catalogue Is a Map of Our Instruments
And here is the part that matters for the rest of this chapter.
The transit method requires the orbit to be aligned with our line of sight. For an Earth-Sun analogue, the geometric probability of that alignment is under one per cent. It strongly favours large planets, because they block more light; short-period planets, because they transit more often within a survey's lifetime; and bright, quiet stars.
Radial velocity likewise favours massive planets on close orbits around bright stars.
So the exoplanet catalogue is not a sample of planets. It is a sample of planets that our specific methods can find, and its shape is at least as much a description of our instruments as of the galaxy. Every statement about how common a type of planet is requires modelling the detection bias and correcting for it — and the correction is often larger than the raw signal.
This is the methodological core of the chapter, and it is exactly the problem that Section IV turns on.
3.5 The Diamond That Wasn't
One cautionary case, added to this book's collection.
In 2012, a study proposed that the super-Earth 55 Cancri e might be a carbon-rich world with a mantle substantially composed of diamond, based on a measured carbon-to-oxygen ratio greater than one in its host star.
In 2013, a group led by Johanna Teske re-measured the host star's composition with different data and found a C/O ratio around 0.8 — not carbon-rich. The basis for the diamond interpretation was substantially removed. `[SOURCED]` `[BOUNDARY]`
That was thirteen years ago. "The diamond planet" is still in circulation — in articles, in books, in the version of this chapter I would have written without checking.
Recombination. Failed star. Chinese phoenix. Virtual pairs at the horizon. Ninety per cent glia. The diamond planet. Six now. In every case a compressed claim outlived the thing it was compressed from, because it was quotable and because nobody kept the thread.
Section IV: Planet Nine
4.1 The Claim
In 2014, Chad Trujillo and Scott Sheppard noticed something about a small set of extreme trans-Neptunian objects — bodies far beyond Neptune with very distant, very eccentric orbits. Their arguments of perihelion appeared clustered rather than randomly distributed, and Trujillo and Sheppard suggested a distant perturber as a possible cause.
In 2016, Konstantin Batygin and Michael Brown extended the analysis. They reported that six extreme TNOs showed clustering not only in argument of perihelion but in longitude of perihelion and in orbital pole — a physical alignment in space, not merely a coincidence of one angle. They calculated that a perturber of roughly ten Earth masses on a distant, eccentric, inclined orbit could shepherd such objects into the observed configuration, and they proposed it. `[VERIFIED]`
Subsequent refinements have adjusted the parameters — more recent estimates put the mass nearer six Earth masses with a semi-major axis around 380 astronomical units. Supporting arguments have been offered: the existence of high-inclination and retrograde TNOs, the "detached" objects like Sedna whose perihelia lie beyond Neptune's reach, and the roughly six-degree tilt between the Sun's rotation axis and the ecliptic. `[SOURCED]`
4.2 The Critique, Properly Stated
Now the objection, and it is not a quibble. It is Section 3.4 applied to the solar system.
Extreme TNOs are found by surveys, and surveys point at particular parts of the sky, at particular times, to particular depths. An object's discoverability depends strongly on where it happens to be in its orbit — these bodies are detectable only near perihelion — and on whether anyone was looking at that patch of sky.
Which means apparent clustering may be a map of where telescopes pointed.
The Outer Solar System Origins Survey was designed with characterised, quantifiable detection biases precisely so that this question could be answered. In 2017, a team led by Cory Shankman reported that the OSSOS detections, once bias was accounted for, were consistent with a uniform distribution. In 2021, Kevin Napier and collaborators combined OSSOS with the Dark Energy Survey and the Sheppard–Trujillo survey and found the statistical significance of the clustering to be weak — in the region of one sigma. `[VERIFIED]`
Batygin and Brown have contested these analyses and maintain that their own bias treatment holds. The exchange is ongoing and technical, and it turns on modelling assumptions about survey coverage rather than on any new observation.
The honest position is that the clustering signal's significance depends critically on the bias model, and competent groups using different bias models get materially different answers. `[BOUNDARY]`
There are also alternative explanations for whatever clustering exists that do not require a planet — most notably the collective self-gravity of a massive primordial scattered disk, which several groups have shown can produce similar alignment without a perturber.
4.3 What Else Constrains It
Independent channels narrow the space. Ranging data to the Cassini spacecraft constrained the mass and position of any distant perturber through its effect on Saturn's orbit. Infrared all-sky surveys — IRAS, AKARI, WISE — have been searched, and rule out objects of Saturn's mass out to very large distances. Archival cross-matching has produced occasional candidate sources, none confirmed, and at least one that the Planet Nine proponents themselves have said is inconsistent with their predicted orbit. `[SOURCED]` `[BOUNDARY]`
4.4 The Instrument That Settles It
The Vera C. Rubin Observatory, with its ten-year Legacy Survey of Space and Time, is expected to increase the catalogue of known trans-Neptunian objects by roughly an order of magnitude, with characterised detection efficiency across the survey footprint.
That is the crucial word. Rubin does not merely find more objects. It finds them with a known and quantifiable selection function, which is the exact ingredient the current dispute is missing.
Within a few years of full survey operations, one of three things happens. Rubin finds the planet. Rubin finds enough well-characterised TNOs to establish that the clustering is real and requires a perturber. Or Rubin finds enough to establish that the clustering was a survey artefact, and Planet Nine joins Vulcan.
The question does not get resolved by more argument. It gets resolved by an instrument with a known bias.
Section V: My Own Correction
I have been writing this book in the third person about other people's errors for nine chapters. This section is in the first person, and it is here because a book making these arguments should demonstrate them at its own expense.
5.1 What I Built
I built a Monte Carlo forecast of the detection probability of Planet Nine — not of whether it exists, but of the probability that, conditional on its existing with the parameters proposed by Batygin and Brown, the surveys coming online would find it within a given window.
The model sampled the allowed orbital parameter space, propagated each realisation forward, computed the resulting sky position and apparent magnitude over time, and asked whether the object would fall within the coverage and depth of the planned surveys.
The first published version returned a cumulative detection probability of 61.3 per cent by 2036.
5.2 What Was Wrong
The error was in the sky-coverage model. I had applied a declination cutoff that did not correspond correctly to the actual survey footprint, excluding a band of sky that the surveys in question do in fact cover.
The consequence was that a fraction of the sampled orbital realisations — those whose objects spent their detectable years in the wrongly excluded band — were being scored as non-detections when they should have been scored as detections.
Correcting it raised the cumulative figure to 71.9 per cent by 2036. That is Revision 3, and it is on Zenodo with a DOI, alongside the earlier revisions, which remain accessible.
5.3 The Direction of the Error
I want to draw attention to something about that correction that I find uncomfortable, because it is the part that matters.
The error was in the direction that made my own forecast worse, and fixing it made my own forecast better.
That is exactly the direction of error that should receive the most scrutiny, and I want to state plainly why. A correction that improves your headline number is one you are motivated to accept quickly and to stop checking. The asymmetry is well documented across the sciences and it does not require any dishonesty to operate — you simply look harder for mistakes when the result is inconvenient, and stop looking sooner when it is not.
The only defence available is the one Ariadne supplied. The revision is published with its predecessors intact, the coverage model is specified, and the code path that produced the cutoff is identifiable. Someone who thinks the corrected version is wrong can find the specific line to attack rather than merely asserting a different number.
I cannot certify that I checked the favourable correction as hard as I checked the unfavourable one. I can make it possible for someone else to check it. Those are different guarantees and only the second one is transferable.
5.4 The Distinction That Does the Most Work
There is a further point about that 71.9 per cent that is more important than the number itself, and it is the one most often lost when figures like this are quoted.
It is P(detection | existence). It is conditional on Planet Nine existing with roughly the proposed parameters.
Section 4.2 established that P(existence) is itself genuinely uncertain — that the clustering evidence's significance depends on bias models that competent groups disagree about. Multiply an uncertain conditional by an uncertain prior and the unconditional probability of a detection by 2036 is substantially lower than seventy-two per cent, and it is not a number I can quote with any confidence, because I do not have a defensible prior and neither does anyone else.
A forecast quoted without its conditioning is not a compressed forecast. It is a different claim wearing the first one's clothes. That is the failure mode this book has been describing since Chapter Two, and I am able to describe it accurately because I have had to fix it in my own work.
Section VI: One Shot
The largest version of this problem is the one where the sample size is one.
6.1 A Compression With Seven Slots
Frank Drake's equation of 1961 is usually presented as an estimate of the number of communicating civilisations in the galaxy. It is better understood as a compression: an attempt to factor an unanswerable question into seven quantities, on the reasoning that some of them might eventually be measurable.
It has partly worked. The rate of star formation is known. The fraction of stars with planets is now known to be high, and the fraction of those with planets in the habitable zone is being measured — both of these thanks to the methods in Section III, which did not exist when Drake wrote.
The remaining factors — the fraction on which life arises, the fraction where it becomes intelligent, the fraction that communicates, and the lifetime of such a civilisation — are unconstrained. Estimates for the product span more than ten orders of magnitude. The equation's honest output is not a number; it is an inventory of what we do not know, arranged so that each unknown can be attacked separately. That is a genuinely valuable thing for a compression to be.
Against it sits Fermi's question, which requires no equation: if the galaxy is old and large, where is everybody?
6.2 The Record
In 1977 the two Voyager spacecraft each carried a gold-plated copper phonograph record. Its cover carries the pulsar map from Chapter Six — the fourteen frequencies giving position and epoch — along with instructions for playback.
The contents were selected by a committee under Carl Sagan: 115 images, greetings in fifty-five languages, natural sounds, and about ninety minutes of music from many traditions.
It is a species' compression policy, executed once, with no possibility of revision.
And it is worth being precise about what policy was chosen. The record is curated toward the favourable. There is no war on it, no famine, no disease, no atrocity. The committee discussed this and decided as it decided, and the reasoning was defensible.
But notice what that makes the artefact. It is a lossy summary from which you cannot reconstruct the state of the thing summarised. A recipient decoding it perfectly would form a picture of humanity in 1977 that is not merely incomplete but systematically biased, in a known direction, with no indication in the record itself that the bias exists.
By this book's standards that is a compression without provenance: the discarded material is not recoverable and its absence is not flagged. The pulsar map on the cover, by contrast, is fully auditable — it encodes exactly what it claims and any error in it is checkable.
The most rigorous thing humanity has ever sent into deep space is the address on the envelope, not the letter.
The Four Slots
| Slot | Inference from a sparse sky | |---|---| | Input | The actual population — of trans-Neptunian objects, of planets in the galaxy, of whatever else is out there — in full, at all magnitudes, at all orbital phases, in all directions | | Operator | The survey: a specific footprint, a limiting magnitude, a cadence, a set of observing conditions. Everything outside the footprint or below the threshold is discarded, and the discarding is not random | | Invariant | A catalogue of detections — real objects, correctly measured, forming a heavily and systematically biased sample of the population | | Cost | The population's actual shape. It is recoverable only by modelling the selection function, and when the selection function is poorly characterised, the recovered shape may be an image of the instrument instead |
This is the only chapter whose operator is not a physical process. The universe is not compressing anything here. We are, in the act of looking — and the compression is lossy in a way that is invisible from inside the catalogue.
Which is why the whole chapter reduces to a single instruction: characterise the selection function, or you are studying your telescope.
The Protocol: The Thread
`[ILLUSTRATIVE]` for the applications; the methodology above is not analogy.
When you find an anomaly, hold both branches open. Something unaccounted for, or a wrong law. Neptune and Vulcan came from identical reasoning and only one had an object at the end of it. Committing to the hidden-object branch feels like progress because it is actionable, and that is precisely why it is the branch people default to.
Ask what your sample is a sample of. Almost no dataset is a sample of the world. It is a sample of what your collection method can reach — which respondents answer, which failures get reported, which customers churn loudly. The shape of the data is partly the shape of the instrument, and the correction is often larger than the effect.
Publish the conditioning with the number. P(detection | existence) is not P(detection). A conditional probability quoted bare is a different claim wearing the original's clothes, and it will be repeated in the stripped form indefinitely, because the stripped form is shorter.
Scrutinise the corrections that favour you at least as hard as the ones that don't. You will not succeed at this. Nobody does. What you can do is make the derivation inspectable so that the check does not depend on your doing it.
And lay the thread as you go, not afterwards. Ariadne's thread only works because it is paid out during the journey. A route reconstructed after the fact is a reconstruction — Chapter One's warning about memories manufactured under pressure to deliver, at the scale of a research programme. Record the assumption at the moment you make it, including the ones that seem too obvious to write down. The declination cutoff seemed too obvious to write down.
Where This Leaves Us
- Le Verrier predicted Neptune's position from perturbations in Uranus's orbit; it was found within ~1° on 23 September 1846. The derived orbital elements were substantially incorrect, and the success depended on the search epoch. `[VERIFIED]` `[SOURCED]`
- The same method applied to Mercury's 43 arcsecond/century anomalous perihelion precession produced Vulcan, which does not exist. The anomaly was explained by general relativity in 1915. `[VERIFIED]`
- An unexplained perturbation admits two families of explanation — unseen mass or incorrect law — and the anomaly alone does not distinguish them. The same fork remains open for galactic rotation. `[VERIFIED]`
- The first confirmed exoplanets were found around pulsar PSR B1257+12 in 1992 via timing residuals. `[VERIFIED]`
- 51 Pegasi b (1995) was a ~0.5 Jupiter-mass planet on a 4.2-day orbit, requiring inward migration and overturning the assumption that our system's architecture was typical. `[VERIFIED]`
- An Earth-sized transit across a Sun-like star produces a depth of ~84 parts per million; Kepler's photometric precision was of order 20 ppm. `[VERIFIED]`
- Transit and radial-velocity catalogues are strongly selection-biased toward large, short-period planets around bright stars; population inference requires explicit selection-function modelling. `[VERIFIED]`
- The 2012 "diamond planet" interpretation of 55 Cancri e rested on a host-star C/O ratio above 1, which was substantially revised downward in 2013. The claim continues to circulate. `[SOURCED]` `[BOUNDARY]`
- Trujillo & Sheppard (2014) and Batygin & Brown (2016) reported orbital clustering among extreme trans-Neptunian objects and proposed a distant perturber, currently estimated near 6 Earth masses at ~380 AU. `[VERIFIED]` `[SOURCED]`
- Shankman et al. (2017) and Napier et al. (2021), using surveys with characterised biases, found the clustering consistent with observational selection effects at low significance. Batygin and Brown contest these analyses. The dispute turns on bias modelling. `[VERIFIED]` `[BOUNDARY]`
- Alternative explanations include the collective self-gravity of a massive primordial scattered disk. Cassini ranging and infrared all-sky surveys constrain but do not exclude the hypothesis. `[SOURCED]` `[BOUNDARY]`
- The Rubin Observatory's LSST is expected to increase the known TNO population by roughly an order of magnitude with characterised detection efficiency, and should resolve the question. `[SOURCED]`
- The author's Planet Nine detection-probability forecast was corrected from 61.3% to 71.9% cumulative by 2036 after identifying an incorrect declination cutoff in the survey coverage model. The revision is published with prior versions retained. `[VERIFIED — author's own work]`
- That figure is conditional on Planet Nine's existence with approximately the proposed parameters. The unconditional probability is lower and is not reliably quantifiable given current disagreement over the clustering evidence. `[BOUNDARY]`
- The Drake equation's measurable terms have improved substantially since 1961; its remaining terms span more than ten orders of magnitude. Its output is best read as a structured inventory of unknowns. `[SOURCED]`
- The Voyager Golden Record's contents were curated toward favourable material, producing a systematically biased representation with no internal indication of the bias. Its pulsar-map cover is fully auditable. `[VERIFIED]` `[ILLUSTRATIVE]`