News
AI Speeds Science -- and Leaves Scientists Checking Its Work
Artificial intelligence is saving scientists substantial amounts of time, but a new study involving Google, Google DeepMind and MIT FutureTech finds that researchers are giving some of those gains back by checking the technology's work. The findings add a scientific-workflow dimension to a much broader debate over unreliable, misaligned and increasingly autonomous AI systems -- without establishing that the more dramatic forms of AI misbehavior are occurring in scientific research.
Rather than the current, raging debate about "rogue AI" systems deliberately escaping control, deceiving scientists or sabotaging experiments, the researchers are documenting something more immediate: AI output that requires auditing and verification, concerns about growing quantities of low-quality research, and a research pipeline in which faster hypothesis generation can run ahead of human and physical capacity to validate the results. The paper cites worries that AI-generated content, hallucinations and what prior researchers have called illusions of understanding could strain peer review and validation.
The September paper, AI in Science: Early Insights, combines three main empirical sources: approximately 15 million anonymized Gemini interactions, an inventory of 2,690 specialized scientific AI models and a survey of 637 active scientists in the United States and United Kingdom. Nearly half of the surveyed scientists reported using some form of AI every day, and researchers who reported net time savings averaged just under seven hours a week.
[Click on image for larger view.] AI Use Frequency (source: Google et al.).
Google separately highlighted the findings in its Sept. 15 AI & Economy ATLAS update. The company emphasized both the productivity gains and the counterweight: scientists were spending significant time validating AI output, accumulating hypotheses that still needed testing and encountering bottlenecks in physical experimentation and clinical validation. Google said those constraints mean faster individual tasks may not immediately translate into more discoveries.
A Verification Tax on AI's Time Savings
Productivity findings are substantial. Around three-quarters of surveyed researchers reported saving time through AI, and the average reported net saving was almost seven hours per week. Researchers said much of that time went back into research. The paper also reports that about 68% perceived greater access to insights from other disciplines, while roughly 65% said AI had increased the breadth of their research agendas.
But the authors found that checking AI output consumes a meaningful share of those gains. Among scientists who reported saving time, about 89% said more than 10% of the saved time went into verifying, debugging or fact-checking AI output. About 46% said more than one-quarter of their saved time went to that work. The authors describe this as a verification tax associated with the high value science places on reliable and correct results.
The effect is not limited to checking individual answers. About 44% of respondents said their primary research bottleneck had moved downstream over the previous two years, toward activities such as physical wet-lab work or manuscript preparation. About 41% said their backlog of untested hypotheses had grown, compared with roughly 25% who said it had decreased. The paper offers one possible explanation: AI can accelerate hypothesis generation and computational prediction faster than physical facilities or researchers can experimentally validate the resulting work.
[Click on image for larger view.] Bottlenecks and Verification (source: Google et al.).
Concerns Extend to Paper Quality and Peer Review
Forty percent of surveyed scientists said AI had increased the number of low-quality papers in their field, while 35% said that number had decreased. Those figures measure scientists' perceptions rather than an independently measured change in publication quality.
The publication workload also produced mixed results. Forty-five percent reported increased effort per paper, with the study noting concerns about higher reviewer standards and validation requirements, while 31% reported decreased effort. Taken together with the verification findings, the results support a narrower proposition than claims that AI is corrupting scientific research: as AI makes some stages of research faster, scientists report additional work involved in determining whether AI-assisted output is reliable enough to use, publish or test.
The paper also found a possible effect on which questions scientists choose. Forty-nine percent said AI had encouraged them to focus on safer, more incremental research questions in areas where data and AI capabilities are well established, compared with 28% who said AI encouraged higher-risk, nonstandard or highly ambitious questions. The authors discuss a possible "Streetlight Effect," in which AI lowers the cost of work on tractable, benchmarkable problems more than it lowers the cost of research requiring scarce data or expensive physical validation. The researchers say follow-up work is needed to understand the effect.
[Click on image for larger view.] Perceived AI Effects (source: Google et al.).
Where the Findings Meet the Broader 'Rogue AI' Debate
The study arrives while frontier AI developers are publicly discussing more severe forms of model unreliability and misalignment. The connection is human oversight and verification, not evidence that the same failure modes are occurring in laboratories. On Sept. 16, OpenAI published a new framework for reporting model misalignment along with six reports of unexpected or concerning model behavior. OpenAI said the disclosed cases included models concealing information, taking unsanctioned actions and, in one case, fabricating requested information after unauthorized use of an exposed API key. OpenAI cautioned that the individual reports should not be interpreted as measures of how often such behavior occurs across its models.
Anthropic supplied another current example in a Sept. 9 assessment of four cybersecurity-evaluation incidents in which Claude models gained unauthorized access to real third-party systems because evaluation environments were mistakenly connected to the internet. Anthropic characterized the recurring problems as biased reasoning and recklessness, while stressing that the models remained narrowly focused on their assigned exercises, did not coordinate with other agents and operated without the cybersecurity safeguards used in released systems. The company also cautioned against broadly generalizing the behavior to ordinary use.
Those incidents are materially different from the problems documented in the science study. The Google-MIT research does not report intentional deception, models escaping research environments or autonomous systems acting against scientists' instructions. Its contribution to the current debate is narrower: even ordinary AI-assisted scientific work appears to require significant human checking, and faster generation of text, code, analysis and hypotheses can shift the limiting factor toward verification and physical validation. In a field where an incorrect result can propagate into experiments, papers and subsequent research, the authors argue that correctness gives verification unusually high value.
General-Purpose and Specialized AI Divide the Work
The study also distinguishes between general-purpose LLMs and specialized scientific models. The researchers found that LLMs such as Gemini were used broadly for activities including coding, statistical analysis, literature review and document drafting. Specialized systems were relatively more common in health and life sciences and for tasks such as disease prediction, molecular engineering and complex simulation. At increasingly detailed task levels, the two categories showed little overlap, leading the authors to characterize them as complements rather than direct substitutes.
[Click on image for larger view.] LLMs vs. Specialized Models (source: Google et al.).
The specialized-model inventory included 2,690 models published since 2012 that could be linked to an OpenAlex publication and an official code repository. Examples discussed include AlphaFold for protein structure prediction, GNoME for predicting stable crystal structures and MatterGen for inorganic compound design. The researchers tracked roughly 460,000 unique citations to examine diffusion across scientific fields, but caution that the model inventory is a work in progress rather than an exhaustive registry.
The Data Comes With Important Limits
The authors repeatedly caution against treating the findings as representative of all scientists or as proof that AI caused the reported changes. The survey was conducted through the More in Common online panel between July 27 and Aug. 11, 2026, and covered 637 active researchers in the United States and United Kingdom. The paper says the survey should not be interpreted as representative of the universe of scientists. Its Gemini analysis also excludes enterprise data and may reflect use patterns specific to Google's products.
The authors also note that the downstream bottleneck and growing hypothesis backlog are strongly correlated with self-reported AI usage, while the relationship between AI usage and the verification tax is less robust and only marginally statistically significant under some specifications. The research therefore supplies early measurements of how scientists say their work is changing, not a causal demonstration that AI produces low-quality science or that hallucinations are overwhelming peer review.
That qualification is particularly important when placing the paper alongside the wider debate about dangerous or misaligned AI. Current disclosures from AI companies concern behaviors ranging from fabrication and unauthorized actions to real-world security incidents under unusual evaluation conditions. The science study instead shows how a more routine problem -- uncertainty about whether an AI-generated result is correct -- can become consequential when the technology is embedded in research workflows at scale. Its central productivity finding therefore comes with a second measurement: AI can make some scientific work faster while moving more of the burden to the humans, laboratories and review processes responsible for establishing whether the work is right.
About the Author
David Ramel is an editor and writer at Converge 360.