Braintrust CEO on Where Engineering Actually Matters in AI

a16z Deep Dives · 2026-02-17 · 47 分鐘
https://www.youtube.com/watch?v=vOVo9NQzPZE影片總結
Braintrust’s Ankur Goyal compares AI evals, Chinese models, and Bash versus SQL agents
Braintrust CEO Ankur Goyal argues that AI systems need engineering around the model, not inside its weights. Evals should turn product expectations into measurable hypotheses, combining quantitative results with human review. He says agents should be built “bitter-lesson pilled,” so their harness, context, and workflows can be discarded when models change, while production feedback loops remain rigorously engineered. Chinese models show high token usage but low dollar share: GLM5 and MiniMax M2.5 scored about 95% on Braintrust’s Bash-versus-SQL agent benchmark, with GLM5 three times cheaper than Sonnet 4.5. The benchmark found SQL more accurate, faster, and more token-efficient than brute-force Bash environments. Goyal also describes a recurring cycle in which closed-source frontier models lead, open-source models catch up within roughly three months, and new frontier releases reset the race—unless capital or enterprise adoption becomes the real bottleneck.
章節
- 0:00Introduction: AI Systems, Capital, and Ankur Goyal’s Braintrust Background
- 3:10What Evals Actually Mean: Hypotheses, Measurement, and Scientific Feedback
- 8:29The Bitter Lesson: Universal Function Approximators Versus Engineering Trade-offs
- 11:58Engineering the Harness, Not the Model: Production Feedback and Guardrails
- 17:47Chinese Models: High Token Usage, Low Dollar Share, and 95% Benchmarks
- 20:37Why Open Source Models Aren’t Dominating Yet: APIs, Cost Cycles, and Llama 3.1
- 27:08Frontier Labs, Infinite Capital, and Scaling Limits: Anthropic’s Race Outgrows Its Ecosystem
- 30:38Demand-Side Saturation: Enterprises Can’t Absorb Better Models Yet
- 34:53Demand-Side Saturation: Braintrust Prices Usage by Ingested Gigabytes
- 38:47Bash vs. SQL for Agents: SQL Wins on Accuracy, Speed, and Token Efficiency
這是 Tier 1 公開摘要
每章重點、段落總結、心智圖由分享者控制是否公開。想看完整分析?自己提交一支。
同頻道的其他分析
Mintlify and the Transition From Human Docs to Agent Infrastructurea16z Deep DivesMintlify’s Han Wang and Hahnbee Lee explain eight pivots, a two-day prototype, and docs becoming AI infrastructure.
To Regulate AI Effectively, Focus on How It’s Useda16z Deep DivesMartin Casado argues AI laws should target illegal use, while US regulatory uncertainty pushes startups toward Chinese open-source models.
How Palantir Scaled: Why the Best Software Is Built Backwardsa16z Deep DivesPalantir’s Akshay Krishnaswamy explains FDEs, backward product building, Foundry, and avoiding the consultancy trap.
Inferact: Building the Infrastructure That Runs Modern AIa16z Deep DivesInferact founders explain vLLM’s rise, 500,000-GPU scale, and a universal open-source inference layer
AI Copilots Are a Dead End. Here's What Actually Works | Kavak CEOa16z Deep DivesKavak uses AI agents for 90–95% of customer interactions after a flat 2023, then grows four times.