Share:

Video summary

YC Data Club explains expert data, Snorkel benchmarks, Inception’s 1,000-token Mercury 2, and multilingual scaling.

At YC Data Club, Francois Chaubard argues that data—not model architecture or GPUs—is AI’s central bottleneck, citing more than $100 billion in market-cap creation and production failures caused by distribution shifts. Snorkel’s Vincent Sunn Chen presents expert supervision as software and introduces Senior SWE Bench, where Fable, Opus, and Soul tied for first on “tasteful pass.” Inception Labs’ Volo Kuleshov explains diffusion language models, with Mercury 2 exceeding 1,000 tokens per second and improving Mercury 2.5 accuracy by over 23% after training on synthesized real-world RL environments. Shayne Longpre’s ATLAS study finds Thai has only 0.6% as many tokens as English in MADLAD-400; cross-language transfer depends on measured synergies, script, model size, and data quality—not language family alone.

Chapters

  1. 0:00Francois Chaubard: Why a whole night about data? Focal’s shift from commodity data to a $100 billion market
  2. 5:16Francois Chaubard: Why a whole night about data? Expert data for changing interfaces and specialized AI
  3. 9:51Vincent Sunn Chen: The Art & Science of Benchmarking Agents — scaling expert supervision at Snorkel
  4. 14:01Vincent Sunn Chen: The Art & Science of Benchmarking Agents — Snorkel data programming and weak supervision
  5. 17:31Vincent Sunn Chen: The Art & Science of Benchmarking Agents — Senior SWE-bench and realistic coding tasks
  6. 21:08Vincent Sunn Chen: The Art & Science of Benchmarking Agents — validation agents bridge unit tests and LLM judges
  7. 25:41Vincent Sunn Chen: The Art & Science of Benchmarking Agents — tasteful pass and open benchmark grants
  8. 28:44Volo Kuleshov: Inception diffusion models reach over 1,000 tokens per second with Mercury 2
  9. 34:28Volo Kuleshov: Dowo Forge synthesizes realistic RL environments and boosts Mercury 2.5 by 23%
  10. 40:26Shayne Longpre: ATLAS finds Thai data is only 0.6% of English in MADLAD-400
  11. 44:05Shayne Longpre: ATLAS measures language synergy, with Portuguese, Italian, and French helping Spanish
  12. 48:55Shayne Longpre: ATLAS scaling laws show script matters more than language family

This is a Tier 1 public summary

Whether the chapter key points, section summaries and mind map are public is up to the person who shared it. Want the full analysis?Submit one yourself.

More from this channel