Inside the Race to Measure Frontier Intelligence

a16z · 2026-09-09 · 39 min
https://www.youtube.com/watch?v=WO9c9qxDxzUVideo summary
Vals CEO Rayan Krishnan argues private evaluations exposed Llama 4’s benchmark gap and can help companies measure AI performance and ROI.
Vals founder and CEO Rayan Krishnan tells a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li that public AI benchmarks can become outdated or easy to game. He cites Llama 4: it underperformed on Vals’s private tests despite strong public benchmark scores. Vals runs fast, distributed evaluations before model releases and is building proxies for recursive self-improvement, since training a model to build its successor is too costly and slow. As agents handle work over hours or days, evaluations need fewer tasks but richer criteria, and benchmarks must keep pace with real-world changes. For enterprises, Krishnan says private code repositories can reveal which coding tools deliver the best performance and value; Vals’s Val Smith product builds coding tests from a company’s GitHub codebase. He recounts a month when his team used about $1.5 million in tokens—10 times employee salaries—and says evaluation can guide tool choice and spending. On policy, he argues governments should set rules while independent evaluators test model capabilities and risks, including cyber and biosecurity, and share evidence with policymakers.
Chapters
- 0:00Why Vals Exists: Independent Evaluations After Public Benchmarks Fell Short in 2024
- 2:52The Llama 4 Disaster: Strong Public Scores Clashed with Vals' Private Results
- 4:05Inside the 6-Hour Pre-Release Window: Distributed Tests and Vals’ Steve System
- 6:00The Limits of Evaluation: Explicit Work Criteria, Fuzzy Norms, and Vals’ No-Training-Data Rule
- 10:00The Recursive Self-Improvement Index: Proxies for Model-Building Work and Fresh Benchmarks
- 13:13Beyond Capability: Long-Running Agent Tests, Complex Criteria, and Enterprise AI Budgets
- 17:52When Token Spend Starts to Eclipse Salary Spend: The Need to Prove ROI
- 19:19Private Repos vs. Public Benchmarks: Valsmith Tests Agents on Company GitHub Code
- 22:40How Vals Uses Vals: A $1.5 Million Token Experiment and Tool Recommendations
- 24:48Policy: Evidence-Gathering and Third-Party Tests to Balance Innovation and Public Interest
- 28:32Alignment, Reward Hacking & Models Gaming the Test: Vals Briefs Government on Model Risks
- 33:30The Geopolitics of Evals: Shared AI Verification for Sovereign AI and RSI
- 37:14What the Benchmarking Landscape Looks Like Next: Vals Expands Cyber Evals to Cloud and Grid Infrastructure
This is a Tier 1 public summary
Whether the chapter key points, section summaries and mind map are public is up to the person who shared it. Want the full analysis?Submit one yourself.
More from this channel
How Cursor Built One of AI’s Fastest-Growing Companiesa16zCursor bet on the human-model interface, grew rapidly without early sales hires, and later expanded through Graphite and founder acquisitions.
Why Top Founders Are Racing Into AI Infrastructurea16za16z’s Machine Age Fund targets AI infrastructure as GPU supply is booked to 2028 and memory demand needs three years of capacity.
Why AI Demand Is Outrunning Compute Supplya16zGavin Baker argues AI compute demand will outpace supply, with sub-one-year infrastructure paybacks and orbital data centers approaching.
What Today’s Best Models Still Can’t Do in Matha16zDaniel Litt praises AI’s Erdős unit-distance result but warns that proofs alone cannot replace mathematical understanding or human curiosity.
Inside Moderna’s Biggest mRNA Test Since COVIDa16zModerna and Merck’s personalized mRNA vaccine beat Keytruda alone in a Phase 3 melanoma trial, after 1,000 cancer-vaccine trials failed.
Related analyses
Braintrust CEO on Where Engineering Actually Matters in AIa16z Deep DivesBraintrust CEO Ankur Goyal says AI engineering should focus on evals and harnesses, while SQL beat Bash in his agent benchmark.
Why Education Has to Changea16zBen Horowitz and Gagan Biyani launch Horowitz Andreessen Academy, a San Francisco school built around AI-era projects and real-world skills.
Mintlify and the Transition From Human Docs to Agent Infrastructurea16z Deep DivesMintlify grew from eight pivots and a two-day prototype into documentation infrastructure serving AI agents and 20 million monthly visitors.
Temporal CEO on AI Agents & The Future of Software | Deep Dives with a16za16z Deep DivesTemporal CEO Samar Abbas says durable execution will underpin long-running AI agents, with cloud handling 150,000 actions per second.
AI Copilots Are a Dead End. Here's What Actually Works | Kavak CEOa16z Deep DivesKavak CEO Carlos García Ottati says AI agents now handle 90–95% of customer interactions after a year of flat growth during the transition.