Safety
Google's AI Overviews spew out millions of false answers per hour, bombshell study reveals https://trib.al/1ao7qB1
A study commissioned by *The New York Times* and conducted by AI startup Oumi tested 4,326 Google searches using the SimpleQA benchmark, finding that Google's AI Overviews were accurate 85% of the...
A study commissioned by The New York Times and conducted by AI startup Oumi tested 4,326 Google searches using the SimpleQA benchmark, finding that Google's AI Overviews were accurate 85% of the time when running on Gemini 2 (tested in October) and 91% accurate after upgrading to Gemini 3 (tested in February). However, applied to Google's five-trillion-plus annual searches, a 9% error rate still produces tens of millions of incorrect answers every hour — or hundreds of thousands every single minute. Compounding the accuracy problem, when the system ran on Gemini 2, 37% of correct answers were "ungrounded" — meaning cited websites didn't fully support the provided information — and after the upgrade to Gemini 3, that figure jumped to 56%, while Google disputed the findings, saying the study had "serious holes" and questioning Oumi's reliance on its own in-house AI model to conduct the analysis.
Related
- GenAI’s popularity has hit a wall. Hard to see how OpenAI is going to make its numbers, and easy to see why they bought a media company. The…
- 1. If true (and it does fit with my perceptions FWIW), this is an amazing and incredibly damning graph 2. Can anyone find the source on whic…
- A different sense of the word “bubble”
- Voice ChatGPT can’t start a timer, but AGI is imminent! 🤦♂️
Source: safety