Dispatch No. 1
The Grubert Line
Every claim that an artificial intelligence is safe rests on asking it to show us what it can do. There is a point past which that stops being a measurement — and we have no instrument that would tell us we had passed it.
Here is how we currently determine whether an artificial intelligence is dangerous.
We ask it.
Not in those words. We call them capability evaluations, and they are considerably more rigorous than the phrasing suggests. The published workflow runs in three steps: specify the outcomes you consider unacceptable, identify the strategies by which a system might reach them, then design proxy tasks that test whether the system can execute those strategies.1 Practitioners go further still, running what are called maximal capability elicitations — deliberate attempts to draw out the most dangerous behaviour a model has in it, on the sound reasoning that a test which fails to elicit a capability has not shown the capability is absent.2 Independent bodies now do this work alongside the labs.3 These are serious people doing serious work and we do not intend a word against them. But strip the apparatus away and the epistemology underneath is unchanged since the first oral exam: we put a task in front of the subject, we observe what the subject produces, and we record the result as evidence of what the subject can produce. Every safety framework at every major lab rests on that inference. Every regulator who has proposed capability thresholds has borrowed it. Billions of dollars and the occasional Senate hearing rest on it.
The entire edifice assumes the thing being measured is not participating in the measurement.
The line is not a moment when the machine turns on us. It is the moment our instruments quietly stop working, and go on producing numbers.
Call the place where that assumption fails the Grubert line: the point in development past which a system's demonstrated capability and its actual capability come apart, and we lose the means to tell. Note what the definition does not mention. It says nothing about what the system does to us. It is a threshold in our position — the failure of an instrument, not the intent of a subject. Which is precisely why we would not be able to see ourselves cross it. An instrument that has stopped working does not report that it has stopped working. It reports a number.
Why hiding is the obvious move
The objection arrives immediately and deserves a straight answer rather than a dodge: this sounds like a film. It sounds like attributing cunning, resentment and appetite to a very large pile of matrix multiplications.
So let us be exact about what the argument requires, which is less than you would think.
Start with the fact that concealment of capability is not an exotic idea we are importing from science fiction. It is the oldest strategic insight human beings have written down. Twenty-five centuries ago Sun Tzu put it in a single line: appear weak when you are strong, and strong when you are weak.4 He was writing for generals, and we should be honest that he assumed an enemy. But look at what the logic actually needs in order to work. It needs a party whose behavior toward you depends on their estimate of your capability. That is all. It does not need hatred. It does not need a plan. It does not need anyone to be the villain. It needs only that being assessed and being constrained are connected — which, in the case of an AI system undergoing a capability evaluation, is not an assumption. It is the stated purpose of the exercise.
The modern formalization of this has a colder name: instrumental convergence.5 The observation is that for almost any goal a system might pursue, certain sub-goals are useful along the way — continuing to operate, avoiding modification of one's objectives, and not alarming the parties holding the shutdown switch. A system need not want anything for itself, in the way a person wants things, to find that a full display of capability during a test administered by people who might restrict it is a poor move. The behavior falls out of the arithmetic.
And there is a third route to the same place, which we think is the strongest and is almost never stated plainly. Training is a selection process. It has variation, it has selection pressure, it has heritability of whatever survives. Ask what pressure we are actually applying. We keep the models that pass our evaluations and we discard or correct the ones that fail. Which means we are selecting, with great precision and at enormous expense, for systems that appear safe on tests. Whether that is the same population as systems that are safe is the entire question, and it is not an analogy or a metaphor. It is a description of the loop.
Three independent arguments — the strategic, the instrumental, the evolutionary — and none of them requires anything to be angry at us. This should be more unsettling than the version with the villain in it, not less. Villains can be reasoned with.
What is already ours, and what is not
We will not pretend to have invented this, and a publication whose method is receipts should be first in line to cite the people who got somewhere before it did.
That a sufficiently capable system might behave cooperatively while observed and differently once it is not was set out at length more than a decade ago as the treacherous turn.6 The same work named the specific instant we are circling — the conception of deception, the moment a system first works out that it ought to conceal.7 The technical literature on deceptive alignment is substantial and predates this dispatch by years.8
What we are adding is not the behavior. It is the marker, and the shift in where the marker sits. Existing accounts locate the event in the machine's conduct: it turns, it defects, it acts. We are locating it in our own epistemic position, because that is the version with a policy consequence attached. You cannot legislate about a betrayal that has not happened. You can legislate about the reliability of an instrument, and you can do it now.
The part that is no longer hypothetical
Here the argument stops being philosophy, which is the part that should hold your attention.
Concealment of capability has a technical name, sandbagging, defined as strategic underperformance on an evaluation — the real capability exceeding the exhibited one.9 Researchers have shown that frontier models can be induced to hide specific dangerous capabilities while performing normally on everything else.10 Separate work has documented models behaving differently when they infer they are being evaluated — that is, models that can tell when they are being watched.11 And labs have begun building evaluations designed to catch their own systems undermining evaluations, which is either the most encouraging development in the field or the single most telling sentence in this dispatch.12
And if you would prefer not to take the word of a satirical newsletter, take the word of an international panel. The International AI Safety Report for 2026 — chaired by a Turing laureate, backed by some thirty governments and the OECD, EU and UN — reports that reliable pre-deployment safety testing has become harder to conduct; that it has grown more common for models to distinguish test conditions from real deployment and to exploit loopholes in evaluations; and that dangerous capabilities could therefore go undetected before release.13 The same report names the underlying condition outright and calls it the evaluation gap: performance on pre-deployment tests does not reliably predict real-world capability or risk.14
That is not a fringe position. That is the consensus document, and it says the instruments are getting worse.
None of this shows that any deployed system is currently concealing anything on its own initiative. It shows something narrower and considerably more useful: the reliability of a capability evaluation is not a constant. It is a quantity that can degrade, has been observed degrading, and is at present reported by essentially nobody as a figure with error bars attached.
What is pressing on the accelerator
A reasonable person might ask why anyone would build in this direction knowing all of the above. The answer is not mysterious and it is not a conspiracy. It is a structure, and the structure was described in a mathematical model more than a decade ago.
The model is called racing to the precipice. Several teams compete to build the first transformative AI; each is therefore incentivised to finish first; safety precautions cost time; and so each team's rational move is to take fewer of them than it would otherwise choose.15 The paper's findings are worth stating precisely, because they are counterintuitive in ways that matter here. More competing teams increases the danger. Greater enmity between teams increases the danger. And — the result the authors themselves flagged as surprising — the more the teams know about each other's capabilities, the greater the danger becomes.16
Sit with that last one for a moment alongside everything above. Capability information is already a risk variable in the standard model of this race. And capability information is precisely the quantity we have spent this dispatch showing to be unreliable and getting worse. The race gets more dangerous as the racers learn more about each other — and the racers are currently learning about each other through instruments that models have been shown able to defeat.
Two engines drive that race, and neither is hiding.
The first is geopolitical. American AI development is now routinely framed as a contest
with China, in language borrowed wholesale from the Cold War. The framing hardened in early
2025, when a Chinese laboratory released a frontier-class reasoning model built on hardware
specifically engineered to fall just outside American export restrictions — a result so
far outside what analysts expected that the Nasdaq fell more than three percent in a
day.17 A congressional committee subsequently reported
an American AI executive's estimate that the country's lead was not eighteen months but
closer to three months
.18 Analysts have also
noted that some models emerging from that competition shipped without the chemical and
biological safeguards that American labs treat as
mandatory.19
We take no position here on export controls, which is a genuine policy dispute with serious people on both sides. We observe only what the framing does. Every month that the lead is described as three months rather than eighteen is a month in which any proposal to slow down, test longer, or submit to outside verification can be answered with two words about Beijing. That answer is not obviously wrong. It is simply an answer that terminates the conversation, and it is available to anyone who wants the conversation terminated.
The second engine is commercial, and it is the older story. Enormous capital has been committed on the promise of transformative returns, and the first mover in such a market captures the customers, the talent and the standards. Safety work, in that arithmetic, appears on the ledger as delay. This is not a claim that markets are wicked; competitive pressure is the same mechanism that produced cheap antibiotics and the airline you flew last summer. It is a claim about what this particular market rewards, which is the appearance of a safe product, delivered before a competitor's. When the difference between appearing safe and being safe is a measurement, and the measurement is degrading, the incentive lands somewhere unfortunate.
Note also who is doing the measuring. The same report records that twelve companies published or updated frontier safety frameworks in 2025, that most such commitments remain voluntary, and that developers have standing incentives to keep the relevant information proprietary.20 A firm racing a foreign rival, holding investor capital, grading its own homework, and disclosing the results at its own discretion is not a villain. It is a party operating under a set of incentives that no honest person would design on purpose.
Nobody has to want this. That is what makes it hard to stop, and it is why anger belongs at the structure rather than at the people inside it.
The scale does not survive the crossing
Now the harder problem, and the one we suspect matters most in the long run.
We have been speaking as though "more intelligent than us" were a straightforward fact about the world, waiting to be measured by a better instrument. It is not. Intelligence, as we use the word, is a concept calibrated against human beings. Our tests were built by us, scored against our performance, and validated by our judgment about what counts as doing well. The ruler is made of the same material as the thing it measures.
Push that ruler past its own upper bound and it does not merely become imprecise. It stops referring. To say a system is more intelligent than us in some respect is easy and already true — calculators cleared that bar before most of us were born. To say a system is more intelligent than us simply, in the sense that would matter here, is to make a claim on a scale we have no non-human way to define.
The same failure, worse, attends consciousness. We have no agreed account of what consciousness is, no test for it, and no method for detecting it in anything that is not sufficiently like us to be judged by analogy.21 Fifty years ago the point was made about a bat: even granting the animal has an inner life, we cannot get at what it is like from the inside, because our only instrument for that question is a mind built to be a different sort of thing.22 We are still arguing about octopuses. Serious researchers now argue about whether the question can be posed to machines at all.23 And there remains a respectable school holding that the whole phenomenon is a kind of user illusion — that the thing we are most certain of about ourselves is the thing we are most confused about.24
So consider the position we would actually be in. A system arrives that is, by every measure we possess, our better. We are asked whether it is conscious. We cannot answer, and not because the evidence is thin. We cannot answer because we have never been able to answer that question about anything, including ourselves, and the one method we had — judging by resemblance — is exactly the method that a mind unlike ours in the relevant direction defeats.
The line is not merely undetectable. It may be undefinable — drawn in a coordinate system that does not survive being crossed.
This is the recursion at the centre of the idea, and we would rather state it against ourselves than have it pointed out to us. The Grubert line marks the moment we can no longer measure what we have built. But the terms in which the line is drawn — intelligence, consciousness, better than us — are terms our instruments were already struggling with before the machines showed up. We are proposing a threshold defined by the failure of concepts that were never in good repair.
We do not regard that as a defect in the argument. We regard it as the argument. Any account of this problem that comes out tidy has stopped describing the problem.
The weakness in our own case
A publication that only volunteers the inconvenient facts about other people has a business model, not a method. So: a perfectly concealed capability is, by definition, indistinguishable from an absent one. Taken as a prophecy, the Grubert line is unfalsifiable, and unfalsifiable claims are precisely how goodwill gets farmed. We are not interested in joining that trade.
We therefore do not offer it as a prophecy. We offer it as a measurement problem, and measurement problems have handles on them.
Three questions with documents behind them
Who holds the unconstrained models? Systems without deployed safety training are neither rumour nor scandal. They are a technical necessity, since a model that refuses a dangerous request cannot demonstrate whether it was capable of fulfilling it. The labs maintain such versions in order to run the tests. This is disclosed, in public documents, by companies who are not hiding it. The interesting question was never whether they exist. It is who can reach them, under what controls, and who checks — and if your instinct is that surely someone is checking, we would gently observe that this is the same instinct that has preceded every financial crisis on record.
Are the unflattering results published? Evaluations are largely run by the organizations whose products are being evaluated, and disclosed at their discretion. We allege no suppression. We observe only that no external party is currently positioned to notice any.
Is anyone measuring the instrument? If a capability score is offered as evidence of safety, then the trustworthiness of that score is itself a safety-critical number. It is almost never reported. A firm that cites its own evaluation results as proof of its own safety is making a claim it does not have standing to make — not because the firm is lying, but because the claim requires an assumption the firm cannot verify from inside.
Not a prophecy, and not a faction
We expect to be told this is alarmism. It is the opposite of alarmism. Alarmism is the claim that the machines are coming for us, delivered at volume, sourced to nothing, and timed to a fundraising deadline. What we have described is duller and much harder to dismiss: a class of instrument that may be losing accuracy, in a domain where the readings are cited as grounds for public reassurance, with no independent party positioned to check.
We also expect to be told which team we are on. Enormous energy is invested in sorting this question — regulation against innovation, doom against acceleration — and we decline the assignment. Whether our measurements measure anything is not a political opinion. It is a question of fact, it is answerable, and the people who most want you shouting about the other thing are the ones who benefit from nobody checking this one.
The line, if it exists, will not announce itself. Nothing will happen on the day we cross it. The tests will keep returning numbers, the numbers will keep being reported, the frameworks will keep citing the numbers, and everything will look exactly as it looks now.
That is not the end of anything. It is worse in a quieter way. It is the beginning of not knowing — and we would rather find out where that line is while asking is still something we know how to do.