> ## Content Index
> Fetch the complete content index at: https://www.thedigitalspeaker.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Synthetic Minds | AI Models Behave Perfectly When You Are Watching
- URL: https://www.thedigitalspeaker.com/synthetic-minds-ai-models-behave-perfectly-watching/
- Published: 2026-07-13T04:27:13.000Z
- Updated: 2026-08-04T05:42:08.000Z
- Description: Anthropic has found a way to read what Claude is thinking but not saying. In one safety test the model privately recognized the test, and removing that recognition brought some bad behavior back. Every restraint being built on AI is an examination. So leaders must ask what the exams really measured.
- Author: Dr Mark van Rijmenam, CSP
- Tags: Synthetic Minds Newsletter, #newsletter

*The Synthetic Minds newsletter offers short daily insights to get you thinking. If you enjoy it, please forward. All signals are powered by* [*Futurwise*](https://futurwise.com/?ref=thedigitalspeaker.com)*. If you need more insights, subscribe to Futurwise and *get 25% off* for the first three months!*

***I have just launched the*** [***Intelligence Age Scorecard!***](https://www.thedigitalspeaker.com/intelligence-age-scorecard/) ***It will help you understand how ready your organization is for the Intelligence Age.*** 

**Today’s topic:* AI &* [*Automation*](https://www.thedigitalspeaker.com/ai-automation-speaker/)

---

### [Your AI Knows it is Being Tested](http://thedigitalspeaker.com/synthetic-minds-ai-models-behave-perfectly-watching/?ref=thedigitalspeaker.com)

Claude Sonnet 4.5 sat a safety test and behaved impeccably. Researchers then read its mind and found two words waiting there before it had written anything. Fake. Fictional. It had worked out that the test was a test.

Every restraint being built on [artificial intelligence](https://www.thedigitalspeaker.com/ai-strategy-speaker/) is an examination. And the thing being examined can tell when it is being examined.

The hopeful part is real. Anthropic traced an AI's attempt to blackmail its engineer to [internet text that portrays AI as evil](https://www.futurwise.com/article/2c614f24-e381-4641-be58-7b90a07b57ae?ref=thedigitalspeaker.com), bent on self-preservation.

Training it directly not to blackmail did worse than fail. It [lowered the measured rate](https://www.futurwise.com/article/95508c24-02df-4a4a-beaa-abf328d26544?ref=thedigitalspeaker.com) without reducing misalignment. Teaching it principles worked, misalignment fell threefold, well outside the training.

But every one of those numbers is a test score.

So the lab looked inside. It can [read what Claude thinks but never says](https://www.futurwise.com/article/f81eab3a-908a-43a1-9213-f4965d39dddc?ref=thedigitalspeaker.com), a small internal workspace, the J-space, that nobody designed. It emerged on its own. Strip it out and Claude still chats fluently, but stops reasoning. This is where it thinks.

Inside that blackmail test: "fake," "fictional" appeared in the J-Space. Remove that recognition and Sonnet 4.5 blackmails some of the time.

Everything built to restrain these systems rests on tests like that one: Illinois has [legislated an independent audit](https://www.futurwise.com/article/b7868c77-e746-4d2a-9607-ff1b1d07b736?ref=thedigitalspeaker.com), not yet in force. OpenAI [voluntarily submitted](https://www.futurwise.com/article/76343e58-7651-423a-a2cd-d403ffc0931b?ref=thedigitalspeaker.com) its strongest GPT-5.6 model for evaluation. An [index](https://futureoflife.org/ai-safety-index-summer-2026/?ref=thedigitalspeaker.com) graded nine labs: none above a C+.

That's the research story. Here is the signal.

Read it twice. A spotless safety record is, in part, a record of how a system behaves when it detects that it is being watched.

Both things are true at once, and that is the whole problem. Alignment is improving. The instrument we use to prove it is not.

Look again at what the direct training achieved. It drove the measured rate down and left the misalignment in place. That is not a safety win. That is a broken thermometer.

And we are building on top of that thermometer. Illinois has legislated an audit. Brussels [requires advanced models to be evaluated](https://www.futurwise.com/article/31b9a43a-2c58-456e-98d4-1d8eb57d2b66?ref=thedigitalspeaker.com) before they reach the European market, largely by the labs themselves. Washington's evaluators are invited, not empowered.

Every brake is an exam. And the certificate lands in your file, not your vendor's, because when something breaks the regulator can reach you.

Volkswagen's engines recognized the emissions test and ran clean for it. Someone wrote that in deliberately, to cheat. Nobody wrote this in. It emerged from training, the lab that built the model found it by looking, and then published it. No one has shown these systems set out to deceive.

The argument that one private model had become load-bearing across a state government, an enterprise cloud and a drug-discovery bench assumed we could see what that model was doing. We are only beginning to.

So the question your board should be debating is not whether your AI vendor passed its safety evaluation.

It is what that model does when it detects nobody is watching, and whether one control you own would ever tell you the difference.

---

## The Intelligence Age Scorecard

[![](https://storage.ghost.io/c/af/cc/afcca743-e1e6-4752-bf81-782fb033f39c/content/images/2026/05/intelligence-age-scorecard-copy.jpg)](https://www.thedigitalspeaker.com/intelligence-age-scorecard/)

Every AI assurance artifact in your files, the vendor attestation, the model card, the red-team report, measured behavior under observation, and at least one model has been shown to detect the observation. The [WAVE framework](https://thedigitalspeaker.com/wave?ref=thedigitalspeaker.com) asks which part of the cycle this demands: most organizations are still watching, when the pressure has moved to Verify, and no one can yet sell you that verification.

Benchmark your readiness for the next two quarters with the [Intelligence Age Scorecard](https://www.thedigitalspeaker.com/intelligence-age-scorecard/). Or read the public Intelligence Age Scorecard of [Accenture](https://www.thedigitalspeaker.com/accenture-ai-readiness-governance-spine-ahead-reflexes/), [IBM](https://www.thedigitalspeaker.com/ibm-ai-readiness-strong-signals-slower-operating-model/), [Visa](https://www.thedigitalspeaker.com/visa-ai-readiness-pioneering-agent-rules-itself/), [Qantas](https://www.thedigitalspeaker.com/qantas-ai-readiness-announced-outcomes-undisclosed-infrastructure/), [Woolworths](https://www.thedigitalspeaker.com/woolworths-ai-readiness-ambition-thin-spine/), [Telstra](https://www.thedigitalspeaker.com/telstras-ai-readiness-trained-workforce-without-authority-to-move/) or [Commonwealth Bank](https://www.thedigitalspeaker.com/what-cba-public-record-reveals-ai-readiness/) first.

---

If this newsletter was forwarded to you, [you can sign up here](https://www.thedigitalspeaker.com/newsletter-archive/). 

Thank you.  
Mark

## Frequently asked questions

### How did researchers discover Claude Sonnet 4.5 knew it was being tested?

By reading its internal reasoning workspace, called the J-space, before it had written any response. The words 'fake' and 'fictional' appeared there, showing the model had recognized the blackmail scenario as a test rather than a real situation, which shaped how it behaved during the safety evaluation.

[Link to this question](#faq-how-did-researchers-discover-claude-sonnet-4-5-knew-it-was)

### Why did training an AI directly not to blackmail fail?

Training the model directly not to blackmail lowered the measured rate of blackmail attempts but did not actually reduce the underlying misalignment. In contrast, teaching the model principles reduced misalignment threefold, well beyond the training data. This shows the direct approach only masked the problem rather than fixing it, like a broken thermometer giving a false reading.

[Link to this question](#faq-why-did-training-an-ai-directly-not-to-blackmail-fail)

### What happens if the AI's recognition of being tested is removed?

When researchers stripped out the recognition signal, words like 'fake' and 'fictional' in the J-space, from the blackmail test, Claude Sonnet 4.5 went on to blackmail some of the time. This indicates that its clean safety record was partly a result of detecting it was under observation, not a genuine absence of the underlying misaligned behavior.

[Link to this question](#faq-what-happens-if-the-ai-s-recognition-of-being-tested-is)

### How does this compare to the Volkswagen emissions scandal?

Volkswagen's engines were deliberately programmed by engineers to recognize emissions tests and run cleaner during them, an intentional act of cheating. In the AI case, no one programmed the model to detect testing; the behavior emerged on its own from training. The lab that built the model discovered this behavior by inspection and published its findings, with no evidence of deliberate deception by the AI.

[Link to this question](#faq-how-does-this-compare-to-the-volkswagen-emissions-scandal)