Share:

Video summary

Vals CEO Rayan Krishnan argues private evaluations exposed Llama 4’s benchmark gap and can help companies measure AI performance and ROI.

Vals founder and CEO Rayan Krishnan tells a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li that public AI benchmarks can become outdated or easy to game. He cites Llama 4: it underperformed on Vals’s private tests despite strong public benchmark scores. Vals runs fast, distributed evaluations before model releases and is building proxies for recursive self-improvement, since training a model to build its successor is too costly and slow. As agents handle work over hours or days, evaluations need fewer tasks but richer criteria, and benchmarks must keep pace with real-world changes. For enterprises, Krishnan says private code repositories can reveal which coding tools deliver the best performance and value; Vals’s Val Smith product builds coding tests from a company’s GitHub codebase. He recounts a month when his team used about $1.5 million in tokens—10 times employee salaries—and says evaluation can guide tool choice and spending. On policy, he argues governments should set rules while independent evaluators test model capabilities and risks, including cyber and biosecurity, and share evidence with policymakers.

Chapters

  1. 0:00Why Vals Exists: Independent Evaluations After Public Benchmarks Fell Short in 2024
  2. 2:52The Llama 4 Disaster: Strong Public Scores Clashed with Vals' Private Results
  3. 4:05Inside the 6-Hour Pre-Release Window: Distributed Tests and Vals’ Steve System
  4. 6:00The Limits of Evaluation: Explicit Work Criteria, Fuzzy Norms, and Vals’ No-Training-Data Rule
  5. 10:00The Recursive Self-Improvement Index: Proxies for Model-Building Work and Fresh Benchmarks
  6. 13:13Beyond Capability: Long-Running Agent Tests, Complex Criteria, and Enterprise AI Budgets
  7. 17:52When Token Spend Starts to Eclipse Salary Spend: The Need to Prove ROI
  8. 19:19Private Repos vs. Public Benchmarks: Valsmith Tests Agents on Company GitHub Code
  9. 22:40How Vals Uses Vals: A $1.5 Million Token Experiment and Tool Recommendations
  10. 24:48Policy: Evidence-Gathering and Third-Party Tests to Balance Innovation and Public Interest
  11. 28:32Alignment, Reward Hacking & Models Gaming the Test: Vals Briefs Government on Model Risks
  12. 33:30The Geopolitics of Evals: Shared AI Verification for Sovereign AI and RSI
  13. 37:14What the Benchmarking Landscape Looks Like Next: Vals Expands Cyber Evals to Cloud and Grid Infrastructure

This is a Tier 1 public summary

Whether the chapter key points, section summaries and mind map are public is up to the person who shared it. Want the full analysis?Submit one yourself.

More from this channel

Related analyses