Can Independent Testing Make AI Safer?

Can Independent Testing Make AI Safer?

More

Summary

Bloomberg Tech interviews Rayan Krishnan, founder and CEO of Vowels AI, an independent organization that benchmarks and evaluates AI models — including unreleased systems from OpenAI and Anthropic — before public deployment. The conversation is framed around Krishnan’s response to Dario Amodei’s frontier-pacing essay and what independent evaluation of frontier models actually looks like in practice.

The centerpiece of the interview is Vowels AI’s RSI (Recursive Self-Improvement) index, which tracks how close frontier models are to being able to autonomously design and train their own successors — essentially, an AI making the next version of itself without human researchers. Krishnan presents trend data showing that current models are strong at executing experiments and operating as engineers, but still lack the intuition to generate genuinely novel research directions. Extrapolating current trajectories, Vowels AI predicts that models will eclipse human AI researchers approximately in August 2027.

Krishnan discusses the structural challenge of independent evaluation at scale: the historical precedent from ratings agencies and auditing firms shows that for-profit evaluators can remain credible, but only if they are strictly prohibited from also selling remediation or training data back to the labs they audit. He frames Anthropic’s embedded-evaluator proposal as a demand driver for services like Vowels AI, while acknowledging that no single evaluation approach will be sufficient — a diversity of methodologies and governance structures will be necessary as the technology approaches more consequential capability thresholds.


📺 Source: Bloomberg Tech · Published September 14, 2026
🏷️ Format: Interview

1 Item

Channels