分享這篇:

影片總結

Braintrust’s Ankur Goyal compares AI evals, Chinese models, and Bash versus SQL agents

Braintrust CEO Ankur Goyal argues that AI systems need engineering around the model, not inside its weights. Evals should turn product expectations into measurable hypotheses, combining quantitative results with human review. He says agents should be built “bitter-lesson pilled,” so their harness, context, and workflows can be discarded when models change, while production feedback loops remain rigorously engineered. Chinese models show high token usage but low dollar share: GLM5 and MiniMax M2.5 scored about 95% on Braintrust’s Bash-versus-SQL agent benchmark, with GLM5 three times cheaper than Sonnet 4.5. The benchmark found SQL more accurate, faster, and more token-efficient than brute-force Bash environments. Goyal also describes a recurring cycle in which closed-source frontier models lead, open-source models catch up within roughly three months, and new frontier releases reset the race—unless capital or enterprise adoption becomes the real bottleneck.

章節

  1. 0:00Introduction: AI Systems, Capital, and Ankur Goyal’s Braintrust Background
  2. 3:10What Evals Actually Mean: Hypotheses, Measurement, and Scientific Feedback
  3. 8:29The Bitter Lesson: Universal Function Approximators Versus Engineering Trade-offs
  4. 11:58Engineering the Harness, Not the Model: Production Feedback and Guardrails
  5. 17:47Chinese Models: High Token Usage, Low Dollar Share, and 95% Benchmarks
  6. 20:37Why Open Source Models Aren’t Dominating Yet: APIs, Cost Cycles, and Llama 3.1
  7. 27:08Frontier Labs, Infinite Capital, and Scaling Limits: Anthropic’s Race Outgrows Its Ecosystem
  8. 30:38Demand-Side Saturation: Enterprises Can’t Absorb Better Models Yet
  9. 34:53Demand-Side Saturation: Braintrust Prices Usage by Ingested Gigabytes
  10. 38:47Bash vs. SQL for Agents: SQL Wins on Accuracy, Speed, and Token Efficiency

這是 Tier 1 公開摘要

每章重點、段落總結、心智圖由分享者控制是否公開。想看完整分析?自己提交一支。

同頻道的其他分析