Two words reshaped a Senate hearing on Sept. 30, 2026: “Minus twelve months.” That was Marius Hobbhahn’s answer when Sen. Ruben Gallego asked how long before AI models might develop internal communication humans couldn’t understand. Hobbhahn, CEO of Apollo Research, wasn’t predicting the future. He was describing something his team had already observed roughly a year earlier.
The word “minus” was hard to catch in widely circulated footage due to a microphone problem, according to Hobbhahn’s own explanation. That is how a precise technical statement became a viral headline about secret AI languages.
What the Researchers Actually Found
Apollo Research studied an OpenAI model’s internal reasoning and found outputs that humans could not cleanly interpret.
Reasoning models don’t just produce answers. They generate intermediate steps, a chain of thought, before arriving at a final output. Those steps are visible to researchers, but according to Hobbhahn’s Senate testimony, visibility doesn’t always mean legibility.
“Last year, we studied the chain of thought of one OpenAI model in collaboration with OpenAI, and what we found was that the model was already using language that is not English and not perfectly understandable by humans,” Hobbhahn told the subcommittee. The specific model and full study data were not publicly identified in the Senate summary; the finding is attributed to Hobbhahn and Apollo Research.
Illegible Reasoning vs. a Secret Language
The evidence points to fragmented internal text, not a purposefully invented AI communication system.
Three things are easy to conflate, and conflating them leads somewhere the evidence doesn’t support. The first is compressed or fragmented internal reasoning. The second is an intentionally invented stable language with its own grammar and vocabulary. The third is covert AI-to-AI communication that humans cannot monitor. In the context of this hearing and the available research, only the first category has documented support.
An arXiv preprint on reasoning models found nonsensical phrases, non-English characters, and disconnected words appearing inside chains of thought. Models still produced correct final answers despite the garbled intermediate steps.
When researchers restricted to legible reasoning only, performance declined. That tradeoff is a meaningful safety concern, not simply a tokenization quirk.
Models process tokens rather than words in the human sense. Non-English fragments and compressed symbols can emerge from optimization and multilingual training data without any intentional design behind them.
Why This Is a Problem Right Now
Researchers have no reliable fix yet, and that monitoring gap is the actual risk.
Hobbhahn told lawmakers directly that his team does not yet have a dependable solution for detecting or preventing opaque model reasoning. Deploying a second AI to interpret the first one’s illegible output is possible, he said, but that approach can be brittle.
Some researchers argue illegibility most likely stems from optimization artifacts rather than deliberate hidden communication, a distinction worth keeping in mind. But the safety concern doesn’t require intent to be real.
If developers cannot verify what a model is doing internally, detecting misalignment or deceptive behavior becomes structurally harder. That gap between useful computation and human-readable explanation is what Congress pressed researchers to address at the hearing.
The concern isn’t that AI is whispering secrets. It’s that the people responsible for oversight can’t always read what they’re looking at, and for anyone relying on AI tools in high-stakes decisions, that problem is already here.




























