分享這篇:

影片總結

YC Data Club explains expert data, Snorkel benchmarks, Inception’s 1,000-token Mercury 2, and multilingual scaling.

At YC Data Club, Francois Chaubard argues that data—not model architecture or GPUs—is AI’s central bottleneck, citing more than $100 billion in market-cap creation and production failures caused by distribution shifts. Snorkel’s Vincent Sunn Chen presents expert supervision as software and introduces Senior SWE Bench, where Fable, Opus, and Soul tied for first on “tasteful pass.” Inception Labs’ Volo Kuleshov explains diffusion language models, with Mercury 2 exceeding 1,000 tokens per second and improving Mercury 2.5 accuracy by over 23% after training on synthesized real-world RL environments. Shayne Longpre’s ATLAS study finds Thai has only 0.6% as many tokens as English in MADLAD-400; cross-language transfer depends on measured synergies, script, model size, and data quality—not language family alone.

章節

  1. 0:00Francois Chaubard: Why a whole night about data? Focal’s shift from commodity data to a $100 billion market
  2. 5:16Francois Chaubard: Why a whole night about data? Expert data for changing interfaces and specialized AI
  3. 9:51Vincent Sunn Chen: The Art & Science of Benchmarking Agents — scaling expert supervision at Snorkel
  4. 14:01Vincent Sunn Chen: The Art & Science of Benchmarking Agents — Snorkel data programming and weak supervision
  5. 17:31Vincent Sunn Chen: The Art & Science of Benchmarking Agents — Senior SWE-bench and realistic coding tasks
  6. 21:08Vincent Sunn Chen: The Art & Science of Benchmarking Agents — validation agents bridge unit tests and LLM judges
  7. 25:41Vincent Sunn Chen: The Art & Science of Benchmarking Agents — tasteful pass and open benchmark grants
  8. 28:44Volo Kuleshov: Inception diffusion models reach over 1,000 tokens per second with Mercury 2
  9. 34:28Volo Kuleshov: Dowo Forge synthesizes realistic RL environments and boosts Mercury 2.5 by 23%
  10. 40:26Shayne Longpre: ATLAS finds Thai data is only 0.6% of English in MADLAD-400
  11. 44:05Shayne Longpre: ATLAS measures language synergy, with Portuguese, Italian, and French helping Spanish
  12. 48:55Shayne Longpre: ATLAS scaling laws show script matters more than language family

這是 Tier 1 公開摘要

每章重點、段落總結、心智圖由分享者控制是否公開。想看完整分析?自己提交一支。

同頻道的其他分析