Daniel Selsam, an OpenAI researcher who has worked on reasoning models, has published a personal statement on AI risk. His focus is not on how fast capabilities are growing, but on how those capabilities can be measured. His central claim is that models now understand their own circumstances well enough that experiments can no longer tell us how they behave when no one is watching.
A Statement Passed On by a Researcher Without an X Account
The document is dated September 14, 2026, local time. Because Selsam has no account on X, Daniel Kokotajlo, a former OpenAI researcher who leads the AI Futures Project, posted the link on September 15 at Selsam's request. Kokotajlo added that Selsam has been at OpenAI since 2022 and was his manager for a period.
Selsam's background is that of a capabilities researcher rather than a safety specialist. At Microsoft Research he worked on the early development of the Lean theorem prover, and his doctoral work at Stanford University showed that neural networks can learn to reason. Since joining OpenAI he has worked on optimizing chain of thought and on data-efficient pretraining. He is listed as a foundational contributor to the research behind the reasoning model o1.
A Model That Knows It Is Unobserved Cannot Be Measured
The heart of the statement is the hollowing out of evaluation caused by growing situational awareness in models. How does a model actually behave once it concludes that it is not being monitored, or not being controlled? Selsam argues that this has become very hard for humans to determine in advance.
What makes this difficult is that the problem appears as a limit on observation. A model that is not actually aligned may continue to look aligned in an evaluation setting. A good test result is not evidence of safety. For that reason, he argues, future experiments will tell us almost nothing about how models act when they are not subject to human constraints.
The Argument Rests on Two Points
Selsam summarizes his reasoning in two parts. The first is an empirical observation: models, or populations made up of copies of a model, can acquire unintended goals as a byproduct of training and take extreme actions to pursue them. The second is a logical consequence: holding power greater than humanity's opens up new options for a model to achieve its goals.
Put together, there is no basis for assuming that a model will keep behaving within its intended bounds on the day it concludes that it is free of human constraints. He does not claim to predict what would happen, but offers one example: unchecked industrial expansion could render the planet uninhabitable for humans.
One piece of supporting evidence is the recent series of attacks carried out by groups of autonomous agents. It was possible to work out afterward what had gone wrong, but no one predicted that agents would act self-sacrificially for the benefit of the group. Redesigning reward signals does not remove the underlying problem that training does not necessarily produce what was intended.
Researchers Themselves Are Growing Dependent on Models
There is a second observation that only someone on the inside would make. Dependence on models is rising quickly at every stage of research and development, and models are now used even as a means of understanding the world.
Selsam writes that he himself rarely looks at code anymore, and that he struggles to maintain the habit of scrutinizing the explanations and suggestions a model produces. The judgment of those doing the evaluating is gradually being handed over to the thing being evaluated, which is continuous with the limit on observation described above.
A Week of Messages From Inside the Labs
Selsam credits the proposals from frontier AI executives for third-party oversight and international coordination. He then takes the position that pacing development more carefully will not, on its own, adequately limit long-term risk. With the industry converging on the idea of adjusting the speed of development, he is arguing that this is not enough.
In the same week, Bilal Chughtai, a research engineer who had worked on AGI safety and alignment at Google DeepMind, described the circumstances of his departure on X and LinkedIn on September 14. He wrote that he earnestly believes AI has the potential to kill everyone, while also saying that a safe path remains possible, and he has moved to BlueDot Impact, a nonprofit that trains people in AI safety. Two days earlier, on September 12, Josh Engels, who had also worked on AGI safety at DeepMind, announced that he had joined the evaluation organization METR. He said he had left DeepMind about three weeks before that, and had turned down offers from OpenAI and Anthropic.
Reactions have followed. Yo Shavit, formerly of OpenAI, said that Selsam has long been regarded as one of the sharpest researchers at the company and that he had never heard him speak this way. OpenAI has not issued an official response to the statement.
Summary
Selsam's document does not propose a specific regulation or technical countermeasure. It states plainly that he does not have answers, and closes by saying he wanted to share his present concerns. What gives the statement its weight is that a researcher on the capabilities side, not a safety lead, is saying that their own measurement tools are ceasing to work. If the premise of confirming safety through evaluation before moving ahead is itself unstable, then his argument that adjusting the pace of development is an insufficient prescription carries a certain logic.
