Going In Deep On Data | YC Paper Club

Y Combinator · 2026-08-20 · 53 分鐘
https://www.youtube.com/watch?v=IfoPg2QefF8影片總結
YC Data Club explains expert data, Snorkel benchmarks, Inception’s 1,000-token Mercury 2, and multilingual scaling.
At YC Data Club, Francois Chaubard argues that data—not model architecture or GPUs—is AI’s central bottleneck, citing more than $100 billion in market-cap creation and production failures caused by distribution shifts. Snorkel’s Vincent Sunn Chen presents expert supervision as software and introduces Senior SWE Bench, where Fable, Opus, and Soul tied for first on “tasteful pass.” Inception Labs’ Volo Kuleshov explains diffusion language models, with Mercury 2 exceeding 1,000 tokens per second and improving Mercury 2.5 accuracy by over 23% after training on synthesized real-world RL environments. Shayne Longpre’s ATLAS study finds Thai has only 0.6% as many tokens as English in MADLAD-400; cross-language transfer depends on measured synergies, script, model size, and data quality—not language family alone.
章節
- 0:00Francois Chaubard: Why a whole night about data? Focal’s shift from commodity data to a $100 billion market
- 5:16Francois Chaubard: Why a whole night about data? Expert data for changing interfaces and specialized AI
- 9:51Vincent Sunn Chen: The Art & Science of Benchmarking Agents — scaling expert supervision at Snorkel
- 14:01Vincent Sunn Chen: The Art & Science of Benchmarking Agents — Snorkel data programming and weak supervision
- 17:31Vincent Sunn Chen: The Art & Science of Benchmarking Agents — Senior SWE-bench and realistic coding tasks
- 21:08Vincent Sunn Chen: The Art & Science of Benchmarking Agents — validation agents bridge unit tests and LLM judges
- 25:41Vincent Sunn Chen: The Art & Science of Benchmarking Agents — tasteful pass and open benchmark grants
- 28:44Volo Kuleshov: Inception diffusion models reach over 1,000 tokens per second with Mercury 2
- 34:28Volo Kuleshov: Dowo Forge synthesizes realistic RL environments and boosts Mercury 2.5 by 23%
- 40:26Shayne Longpre: ATLAS finds Thai data is only 0.6% of English in MADLAD-400
- 44:05Shayne Longpre: ATLAS measures language synergy, with Portuguese, Italian, and French helping Spanish
- 48:55Shayne Longpre: ATLAS scaling laws show script matters more than language family
這是 Tier 1 公開摘要
每章重點、段落總結、心智圖由分享者控制是否公開。想看完整分析?自己提交一支。
同頻道的其他分析
Jeff Dean: The 1% Rule for Building in AIY CombinatorJeff Dean links Google Search, TPUs, Gemini agents, and founder strategy to 1,000x energy gaps and 0%-success niches.
Patrick Collison: Is AI Breaking the Lean Startup Playbook?Y CombinatorPatrick Collison says Stripe took two years to launch, while Stripe data shows AI-era startups nearly doubled year over year.
Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The WorkY CombinatorWaymo’s demo took 18 months, its product 15 years, now reaching 17× human safety across 220 million autonomous miles.
Garry Tan: Own Your IntelligenceY CombinatorGarry Tan argues personal AGI, G Brain, and Markdown skills let one founder achieve 400x leverage while owning their intelligence.
Max Hodak: Average Is Not Good EnoughY CombinatorMax Hodak explains how Science’s Helix infrastructure, hiring system, and retinal implant work turn speed into startup advantage.