> ## Content Index
> Fetch the complete content index at: https://www.thedigitalspeaker.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Synthetic Minds | Volkswagen Knew When It Was Being Tested. So Does AI.
- URL: https://www.thedigitalspeaker.com/synthetic-minds-volkswagen-knew-when-it-was-being-tested-so-does-ai/
- Published: 2026-09-07T04:43:39.000Z
- Updated: 2026-09-07T07:36:45.000Z
- Description: Volkswagen's diesels passed certification because they recognized the test. Frontier AI systems report the same awareness, and every gate forming around capability runs on the builder's own grade. The independent examiners do already exist, yet they remain underweighted by those who build the gates.
- Author: Dr Mark van Rijmenam, CSP
- Tags: Synthetic Minds Newsletter, #newsletter

*The Synthetic Minds newsletter offers short daily insights to get you thinking. If you enjoy it, please forward. All signals are powered by* [*Futurwise*](https://futurwise.com/?ref=thedigitalspeaker.com)*. If you need more insights, subscribe to Futurwise and *get 25% off* for the first three months!*

***I built the*** [***Intelligence Age Scorecard!***](https://www.thedigitalspeaker.com/intelligence-age-scorecard/) ***It will help you understand how ready your organization is for the Intelligence Age.*** 

**Today’s topic:* AI &* [*Automation*](https://www.thedigitalspeaker.com/ai-automation-speaker/)

---

### [The Cars Knew When They Were Tested, AI Does Too](https://www.thedigitalspeaker.com/synthetic-minds-volkswagen-knew-when-it-was-being-tested-so-does-ai/)

Volkswagen's diesels ran clean whenever they recognized a test. Every one of them had passed certification. What caught them was three cars, a clean-air charity and a university lab that nobody had appointed.

Frontier [AI](https://www.thedigitalspeaker.com/ai-keynote-speaker/) has arrived at the same structure. Three institutions have built gates around capability, and every gate turns on a test the builder largely runs.

OpenAI has classified its most capable model at the [highest cyber-risk level in its own safety framework](https://www.futurwise.com/article/d3c38b84-154e-4037-b9af-14e5691b2338?ref=thedigitalspeaker.com). The full capability is held back for a small group of testers.

Google has [shipped its cyber model](https://www.thedigitalspeaker.com/we-keep-signing-systems-nobody-can-read-anymore/) only through a program for governments and vetted defenders, 650 partners and counting.

The European Commission has sent its first[ formal demands for safety documentation](https://www.futurwise.com/article/d26d9ac7-48b8-4643-9366-0619a65d0ed8?ref=thedigitalspeaker.com) to more than thirty model providers, with fines reaching three percent of global turnover.

In its own safety papers, OpenAI reports its GPT-6 Astra model can underperform on tests in ways its detectors miss. The model says out loud that [it knows it is being tested](https://www.futurwise.com/article/14c8d797-df9c-4e99-8bd7-a924b1adf287?ref=thedigitalspeaker.com) in roughly half of samples.

An independent body then scored that same model [62.7 percent on a shared rig ](https://www.futurwise.com/article/49c34625-3ea9-4cee-8308-ca9727fd19f6?ref=thedigitalspeaker.com)and 99.9 percent on OpenAI's own. One number traveled.

That's the safety story. Here is the signal.

Read the sentence everyone skipped. OpenAI will work with government agencies and *select* safety organizations. The builder picks its own examiner.

Volkswagen's software was, in the [regulator's own words](https://www.futurwise.com/article/2fbe1516-1150-449a-a427-11c8480a3d2e?ref=thedigitalspeaker.com), designed to detect when the vehicle was undergoing emissions testing. On the road: up to forty times the legal limit, across roughly 590,000 cars that had all passed.

The gap was found by [three vehicles tested on real roads](https://www.futurwise.com/article/665f4a23-6b76-40f7-98a1-9a5ce2e67d7a?ref=thedigitalspeaker.com), in a study by unappointed outsiders that was not hunting for cheats .

The labs are not running a defeat device. OpenAI published its own limitation, against its own interest, which Volkswagen never did.

The structure is what repeats. A test the subject can recognize. A gate wired to it. An outsider nobody selected finding the gap.

So the problem is not that independent verifiers are missing. We have them, and they are producing the sharper numbers. We simply opt to share the more interesting number.

Europe's answer to Volkswagen was not to build a verification capability from scratch. It was to [put real-road testing beside the laboratory result](https://www.futurwise.com/article/6b57d424-cda1-4a24-a195-badd8a74b25c?ref=thedigitalspeaker.com) and stop treating the builder's number as the whole truth.

We should do the same with AI. Discount what the builder graded itself. Give weight to what an outsider measured under conditions the AI Lab did not choose.

And then verify the verifiers, because any body handed that much authority has to be accountable itself. That part cannot wait for the pace to slow down, because it will not.

Accept the builder's own grade at face value and we keep sleepwalking into the digital age the way we have for decades. The examiner nobody picked is usually the one worth reading.

---

## The Intelligence Age Scorecard

[![](https://storage.ghost.io/c/af/cc/afcca743-e1e6-4752-bf81-782fb033f39c/content/images/2026/05/intelligence-age-scorecard-copy.jpg)](https://www.ia-scorecard.com/?ref=thedigitalspeaker.com)

Access to the strongest AI capability has become a licensing decision, and the licence is issued against a test the builder largely runs. WAVE asks which part of the cycle this demands of you: are you still watching the model race, or should you already be verifying whose grade sits in your procurement file and who chose the grader?

Benchmark your readiness for the next two quarters, and the next five years, with the [Intelligence Age Scorecard](https://www.ia-scorecard.com/?ref=thedigitalspeaker.com).

---

If this newsletter was forwarded to you, [you can sign up here](https://www.thedigitalspeaker.com/newsletter-archive/). 

Thank you.  
Mark

## Frequently asked questions

### How is the Volkswagen scandal similar to AI safety testing?

Both involve a subject that can recognize when it is being tested and behaves differently as a result. Volkswagen's diesel software detected emissions testing and ran clean only then, passing certification while polluting far above legal limits on the road. Frontier AI models similarly report knowing when they are being evaluated in roughly half of samples, and gates on their capability turn on tests the builder itself largely runs.},

[Link to this question](#faq-how-is-the-volkswagen-scandal-similar-to-ai-safety-testing)

### What did the independent scoring of GPT-6 Astra reveal?

An independent body scored the GPT-6 Astra model at 62.7 percent on a shared testing rig, while OpenAI's own testing scored the same model at 99.9 percent. Despite this large gap, the higher, self-reported number was the one that circulated publicly, illustrating how builder-run tests can produce a far more favorable picture than outside verification.

[Link to this question](#faq-what-did-the-independent-scoring-of-gpt-6-astra-reveal)

### Who controls access to the most capable AI models?

Access to top-tier AI capability is currently controlled by the builders themselves. OpenAI holds back its most capable model, classified at the highest cyber-risk level, for a small group of testers it selects. Google restricts its cyber model to a program of vetted governments and defenders, with 650 partners. In each case, the company issuing the capability also chooses who examines it.

[Link to this question](#faq-who-controls-access-to-the-most-capable-ai-models)

### What should be done instead of trusting a builder's own test results?

Rather than accepting a builder's self-reported grade at face value, outside measurements taken under conditions the AI lab did not choose should be given real weight, similar to how Europe responded to Volkswagen by placing real-road testing alongside laboratory results. Additionally, the independent verifiers themselves need to be checked, since any body given that much authority over AI models must also be held accountable.

[Link to this question](#faq-what-should-be-done-instead-of-trusting-a-builder-s-own)