Companies
DeepMind Runs Double-Blind Evals, Co-Scientist Runs the Lab
In one week Google DeepMind ran the first double-blind evaluation of a frontier model, where neither weights nor test prompts are exposed, and published a Co-Scientist that controls lab equipment and grew semiconductor films on the first try.

Google DeepMind spent the same week putting an agent in charge of lab hardware and changing how its own models get measured. The two announcements were separate, but they answer one question: when an AI system says it did science, who checks?
The evaluation where nobody sees the other side's cards
On August 27 DeepMind described what it calls the first double-blind evaluation of a proprietary frontier model. A Gemini Flash Lite model was tested by the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons inside Confidential Space, part of Google Cloud's confidential computing stack. The evaluator never sees the model weights; Google never sees the test prompts. Until now, independent testing forced one side to expose something: either the lab shipped weights out, or the evaluator handed over its questions, and questions that leave the building tend to end up in training data, which is the contamination problem that has been eroding trust in benchmark scores.
The post does not name the benchmarks or publish results; a technical report covers the method. This is a process announcement rather than a scoreboard, and the stated use case is the sensitive end of evaluation: cybersecurity tests and government audits, where neither party can afford to disclose.
Co-Scientist gets its hands on the equipment
The second piece is an arXiv paper (2608.26701, with Duke, Columbia, Texas A&M and Google Research) published August 28. The new Co-Scientist configuration goes beyond generating hypotheses: it plans experiments, writes and runs code, controls instruments, analyzes results and drafts the manuscript. In materials science it interfaced with a semi-automated chemical vapor deposition reactor, adapted growth recipes to the lab's constraints in minutes, and grew monolayer MoS2, MoSe2 and WS2 semiconductors on the first attempt. The engine is Gemini 3 Deep Think.
The evaluation numbers are the part worth keeping. Thirty domain experts produced 450 independent double-blind reviews of 150 autonomously generated papers. With the system's reliability modules switched on, the fabrication rate in generated papers was 4%; with them off, 46%. A safety layer rejected 98.7% of potentially harmful research directions. The paper also lists what still breaks: selective reporting, mismatches between the described method and the code that actually ran, and no ability to predict behavior in entirely new systems. DeepMind's own line: "There is a long journey ahead before AI systems can navigate the physical realities of science."
Why these belong in one story
Read alongside Anthropic's Model Hardware Standard from the same week, the direction is hard to miss: the large labs are moving agents into the physical world and, in parallel, building the infrastructure to prove claims about them without taking the vendor's word. For any R&D organization the practical takeaway is that AI-run experiment loops are now a purchasable capability, not a paper. The 4% figure is the reason to keep the expert reviewers: on 100 papers that is still four with fabricated content, and the double-blind pilot is a reminder that "who measured it" is becoming a question with an engineering answer.
Sources: Google DeepMind, arXiv (2608.26701), The Decoder (double-blind), The Decoder (Co-Scientist)

Written by
Faruk Talmaç
Co-Founder & Editor
Co-founder of YZ Uzman, with 20+ years of experience in web design and software development.