Two Bits of Signal

On the second bit a trained system goes without: not how good a reading is, but how far to trust it. A star tracker carries it because the vacuum charges for its absence; a training loop folds it away, and the model cannot tell what it knows from what it doesn't.

A star tracker is a camera that knows the sky by heart. Several times a second it photographs a patch of sky, finds the points of light in the frame, measures the angles between them, and matches the pattern against a catalog of every star it was built to know. From the fit it computes which way the spacecraft is pointed, to a few arcseconds. Finer than a hand can hold. On a clean field the match is instant and unambiguous: the pattern sits in one place in the catalog and nowhere else, and the number it returns is the truth about where the craft is looking.

Then the sun clears the Earth's limb. Glare floods the aperture, blooming across the detector; a charged particle leaves a bright smear the optics read as a star that was never there. The catalog matches anyway. It fits the corrupted pattern to a place in the sky and hands back a number in the same format, to the same precision, in the same flat voice it used a moment ago: a confident, exact attitude that is wrong.

Nothing in the number says which kind it is. The true fix and the false one come back identical: same shape, same precision, the same certainty on their face. A craft that takes the reading at its word slews to meet an orientation that was never there, turning its panels from the sun, pointing its dish at empty space, firing to correct toward a sky it isn't facing. The reading does not carry the one thing the craft most needs to know about it — whether to believe it.

So the engineers do not ask the reading to carry it. They hold a second quantity beside the first, not a better number but a separate one: how far this fix can be trusted. The reading itself is blind to that, so the trust has to come from outside any single reading. Not from the star tracker. From a second eye, and a third. A spacecraft flies more than one sensor on purpose: star trackers, sun sensors that find the bright disk of the sun, gyroscopes that feel the craft turning in the dark. Each answers the same question, which way the craft is pointed, in a different language. When the answers agree, the fix is trustworthy. When the star tracker drifts from what the sun sensor and the gyros report, the disagreement is the trust made visible. The reading that will not line up with the others is the one the glare has fooled.

The craft acts on it. The outlier is outvoted and thrown out. The fix is taken from the sensors that cohere. And sometimes nothing coheres, the star tracker blinded, the gyros drifted, no two instruments agreeing. The craft does not pick one and commit. Rather than move on what it cannot trust, it drops into safe-mode, points itself at the sun by the crudest reckoning it has, and waits for a clean sky. Whether to trust the reading is the second fact. The number itself never carries it. The whole discipline is the refusal to fold the second back into the first.


A language model's training loop was built the other way. A reward model scores each answer: how good it is, how much a person would approve. The policy, the network being trained, adjusts itself to earn a higher score. One number comes back, and the network moves toward it. There is no second sensor beside that number. No sun sensor. No gyro. Nothing else answering the same question in another language, no vote to lose, no fix to throw out. The scores are consumed as truth. The reward model's own glare arrives in the same format as its reliable readings, and the loop swallows both. Nothing asks how far to trust the number, because the loop has no place to keep the answer.

And the model that comes out of that loop performs the star tracker's false match in software. Ask it something just past what its training covered cleanly and it answers in the same flat voice it uses where it knows the ground. It is not lying. It has no slot for the second fact. From inside the model, the glare and the clean field look the same. The confident wrong answer, the hedge, the agreement that bends to whatever the asker wants: these are the false match, surfacing as behavior. The model cannot tell what it knows from what it doesn't. Nothing it was trained for ever made it need to.


Two questions sit under every reading, and they are not one question. Is there a signal? Can the reading be trusted? Each has two answers, so together they make four states. The star tracker shows all four. A clean field is a signal present and trusted. The glare is a signal present and not to be trusted, the match it should throw out. The sun full in the aperture is a confirmed absence. No fix comes back, the craft knows it, and hands the sky to other sensors. The sky it has not turned to is absent another way, nothing seen because nothing was looked at, no ground to trust or doubt. Present and trusted, present and doubted, confirmed empty, never looked. The single number folds all four into one.

On the present row alone, the two could pass for one. A confident fix and a shaky one look like a single number wearing a tight error bar or a loose one, and an error bar is not a second axis, only the wobble on the first. The absence row tells them apart. Confirmed-empty and never-looked carry no signal at all, nothing for an error bar to wrap around. And still they differ. The only thing that differs is whether the craft trusts a nothing. That is the second question speaking where the first has nothing to say, the proof it was never just the first one's margin of error.

A single number cannot hold four states, so the stack reads two of them onto the other two. The sky it never looked at is scored as empty: unseen, so treat it as bad. The glare's false match is taken for a fix: a number arrived, so use it. Both collapses throw the second question away, each into a different failure.


The field has found its way out of this more than once, and every time it is the same move wearing a different name. Conformal prediction gives a system the right to refuse: a way to say not here rather than emit a number it cannot stand behind. The constitutional methods make a model weigh whether the reward it is about to learn from can be trusted, before it is allowed to move the weights. Both keep the second beside the first, the way the craft does.

Conservative offline reinforcement learning reaches the same end by folding instead. It learns from a fixed record of things that already happened. There is no world to step into, no way to try an action and feel it go wrong. Left alone, the network assigns a confident value to actions that never appear in the record. Not because it knows them: because it has no way to register that it does not. The conservative methods add a term to the loss that punishes that confidence, pulling it down wherever the data runs thin: the second fact folded into the first rather than kept beside it. Built by hand. The penalty stands in for the correction the missing world would have delivered, had the agent been able to take the unseen action and watch it fail. Each builds, in software, the punishment for confident error that reality is not there to deliver.


On the spacecraft, reality is there to deliver it. A confident wrong fix is not a bad row in a log. It is the craft turning its body to a sky that was never there, panels swinging off the sun. It is the mission, spent. The vacuum does not score the answer; it charges for it. So the second bit is not a refinement the engineers added because they were wise about uncertainty. It is what flying requires where a confident error is fatal, and reality keeps it aboard by punishing its absence. The training loop has no such reality. A confident wrong answer there is a bad row in a log, and the loop pays nothing for it, so it folds the bit away.

This is the difference between two kinds of network, and it is one bit wide. One aggregates values. For every state, how good. The other aggregates beliefs. For every state, how good, and how far to trust the answer. The value network is what a system can get away with where nothing charges for a confident error. The belief network is what reality extracts where something does. A belief is a value that has been made to carry how far it can be trusted.


The craft comes back around into the glare. The tracker reads the false fix again, confident and exact and wrong. This time it goes nowhere. The other sensors will not back it, so the craft does not move. It drops to safe-mode and waits for a clean sky.

A system with one bit can be wrong. A system with two can tell when it is lost. The spacecraft was built with both, because the vacuum was never going to let it fly on one. The model has, so far, mostly the one. It answers into the glare in the same flat voice it keeps for a clean sky. The difference it cannot hear is exactly the width of what it does not know it does not know.