Braintrust CEO on Where Engineering Actually Matters in AI

a16z Deep Dives · 2026-02-17 · 47 min
https://www.youtube.com/watch?v=vOVo9NQzPZEVideo summary
Braintrust CEO Ankur Goyal says AI engineering should focus on evals and harnesses, while SQL beat Bash in his agent benchmark.
Braintrust CEO Ankur Goyal tells Martin Casado that AI models are nondeterministic, so teams should use evals to test hypotheses and measure whether changes improve quality, speed, or cost. He argues that agent prompts and internals should be easy to replace as models change; the durable engineering is the harness that connects production feedback to testing. Goyal describes a cycle in which closed models make leaps, then open models catch up, noting Chinese models have high token usage but low dollar share. In a Bash-versus-SQL benchmark, Sonnet 4.5 scored 100%, while GLM-5 and MiniMax M2.5 scored 95%; GLM-5 was three times cheaper than Sonnet, and MiniMax M2.5 three times cheaper than GLM-5. He says demand may ultimately be limited by how slowly enterprises can integrate AI, and argues that structured tools such as SQL and strong type specifications can make agents more accurate, efficient, and reliable than simply giving them a computer.
Chapters
- 0:00Introduction: Ankur Goyal’s Impira and Figma evaluation experience
- 3:10What Evals Actually Mean: Hypotheses, NoSQL Lessons, and Systems Reliability
- 8:29The Bitter Lesson: Universal Models, Data Engineering, and Disposable Agent Code
- 11:58Engineering the Harness, Not the Model: Feedback Loops, Evals, and Model Transfer
- 17:47Chinese Models: High Token Usage, Low Dollar Share, and Bash-versus-SQL Results
- 20:37Why Open-Source Models Aren’t Dominating Yet: Delivery Gaps, Cost, and Frontier Cycles
- 27:08Frontier Labs, Infinite Capital, and Scaling Limits: $1B-to-$10B Races and the AGI Question
- 30:38Demand-Side Saturation: ChatGPT, Gemini, Enterprise Adoption, and Token-Path Margins
- 34:53Demand-Side Saturation: Braintrust’s Gigabyte Pricing and Token-Billing Trade-Offs
- 38:47Bash vs. SQL for Agents: SQL Wins Braintrust’s Benchmark, While Type Specs Guide Engineering
This is a Tier 1 public summary
Whether the chapter key points, section summaries and mind map are public is up to the person who shared it. Want the full analysis?Submit one yourself.
More from this channel
To Regulate AI Effectively, Focus on How It’s Useda16z Deep DivesMartin Casado argues AI laws should target harmful uses, while regulatory uncertainty pushes US startups toward Chinese open-source models.
Mintlify and the Transition From Human Docs to Agent Infrastructurea16z Deep DivesMintlify grew from eight pivots and a two-day prototype into documentation infrastructure serving AI agents and 20 million monthly visitors.
Inferact: Building the Infrastructure That Runs Modern AIa16z Deep DivesInferact’s vLLM founders aim to build an open inference layer for models running across 400,000–500,000 GPUs.
How Palantir Scaled: Why the Best Software Is Built Backwardsa16z Deep DivesPalantir architect Akshay Krishnaswamy explains how field engineers turn customer pain into products and scale software backwards.
Temporal CEO on AI Agents & The Future of Software | Deep Dives with a16za16z Deep DivesTemporal CEO Samar Abbas says durable execution will underpin long-running AI agents, with cloud handling 150,000 actions per second.
Related analyses
Why AI Agents Need Context | Deep Dives with a16za16z Deep DivesFivetran CEO George Fraser says AI agents need centralized business data, while warning AI-native rivals could outpace SaaS incumbents.
Open Models Change The Economics of AIY CombinatorOllama CEO Jeffrey Morgan says open models could handle 80–90% of enterprise tokens as cloud usage surged 150x.
Inside the Race to Measure Frontier Intelligencea16zVals CEO Rayan Krishnan argues private evaluations exposed Llama 4’s benchmark gap and can help companies measure AI performance and ROI.
The Next Chapter of AI Is Inside Our Softwarea16zAaron Levie and Martin Casado debate AI regulation, agent cybersecurity, and a software shift toward probabilistic models.
AI Copilots Are a Dead End. Here's What Actually Works | Kavak CEOa16z Deep DivesKavak CEO Carlos García Ottati says AI agents now handle 90–95% of customer interactions after a year of flat growth during the transition.