Share:

Video summary

YC’s data talks show how expert-built benchmarks, diffusion models reaching 1,000 tokens per second, and multilingual scaling can improve AI.

At YC Paper Club, Francois Chaubard argued that data—not model architecture or GPUs—is often the bottleneck: production failures expose gaps that curated datasets and expert feedback must address. Vincent Chen of Snorkel described scaling expert supervision, from software-based labeling to Senior SWE-Bench, which evaluates coding agents on realistic tasks and tasteful code; he said Fable, Opus, and Sonnet were tied for first. Inception Labs co-founder Volo Kuleshov presented diffusion language models that generate tokens in parallel, with Mercury 2 reaching 1,000 tokens per second. He also explained how Inception’s Forge synthesizes training and evaluation environments from real-world interactions; a Mercury 2.5 experiment improved performance by over 23%. Finally, Shayne Longpre presented ATLAS, a framework for multilingual scaling laws: Thai has just 0.6% as many tokens as English in a common pretraining corpus, and transfer between languages can be either helpful or harmful. He found script similarity mattered somewhat more than language-family similarity.

Chapters

  1. 0:00Why a whole night about data? Inspecting model errors and crafting expert datasets
  2. 5:16Why a whole night about data? Salesforce UI changes and expert-domain datasets
  3. 9:51The Art & Science of Benchmarking Agents: Scaling expert supervision from labels to agent environments
  4. 14:01The Art & Science of Benchmarking Agents: Data programming and weak supervision
  5. 17:31The Art & Science of Benchmarking Agents: Senior SWE-Bench and realistic coding tasks
  6. 21:08The Art & Science of Benchmarking Agents: Validation agents bridge tests and LLM judges
  7. 25:41The Art & Science of Benchmarking Agents: Tasteful pass rankings and open-benchmark grants
  8. 28:44Volo Kuleshov: Mercury 2 exceeds 1,000 tokens per second, while Dowo Forge builds realistic RL environments
  9. 34:54Volo Kuleshov: Dowo Forge training lifts Mercury 2.5 performance by over 23%, with $500,000 for YC startups
  10. 40:26Shayne Longpre: ATLAS examines multilingual transfer as Thai has just 0.6% of English’s Madlad 400 tokens
  11. 44:05Shayne Longpre: ATLAS finds Indonesian transfers to Thai, while Japanese harms Spanish
  12. 48:55Shayne Longpre: ATLAS models directional transfer, script effects, and multilingual scaling choices

This is a Tier 1 public summary

Whether the chapter key points, section summaries and mind map are public is up to the person who shared it. Want the full analysis?Submit one yourself.

More from this channel

Related analyses