Why Robotics Still Isn't Solved - But Could Be Soon | YC Paper Club

Y Combinator · 2026-08-08 · 84 分鐘
https://www.youtube.com/watch?v=myDCd0hNqQU影片總結
YC researchers say robotics remains unsolved, but memory, self-supervised reasoning, and simulation-trained dexterity are pushing it forward.
A YC Robotics Club host argues that robotics has been declared “solved next year” for a decade, yet teleoperation remains difficult and four barriers persist: the sim-to-real gap, action representation, limited touch sensing, and embodiment drift. Stanford’s Marcel Torne presents MEM, which adds compressed visual and language memory to robot policies so they can track long tasks, avoid burning grilled cheese, and adapt after mistakes. Milan Ganai describes self-supervised bootstrapping that selects concise, action-predictive reasoning, improving policies across manipulation, navigation, and driving, including models from 1B to 30B parameters. Tyler Lum’s Sim2Real trains one dexterous hand-and-arm policy in simulation; it runs at 60 Hz and handles 12 unseen tools and behaviors zero-shot, while Play to Perfect extends the approach to precise assembly. Rerun CEO Nico urges startups to own customer operations end to end, begin with teleoperation and off-the-shelf hardware, and iterate on real data. General Instinct closes with world-action models and infrastructure optimizations that reduce flow-matching inference from 50–100 steps to one or two.
章節
- 0:00Francois Chaubard: AlphaGo, Aloha, and four barriers to solving robotics
- 7:59Marcel Torne: MEM splits robot memory into compressed short- and long-context streams
- 12:04Marcel Torne: MEM enables grilled-cheese timing and recovery from robot mistakes
- 16:32Marcel Torne: MEM uses annotated text in RAM while exploring learned memory
- 20:21Milan Ganai: Why VLAs need embodiment-specific, action-predictive reasoning
- 24:23Milan Ganai: RNB Encore selects useful reasoning across robot tasks
- 29:48Milan Ganai: Action forcing avoids inference-time reasoning latency
- 33:42Tyler Ga Wei Lum: SimToolReal controls a 22-DoF hand and 7-DoF arm at 60 Hz
- 38:18SimToolReal generalizes to 12 unseen tools; Play to Perfect extends it to assembly
- 42:27SimToolReal Q&A: Random pushes teach recovery, but arm twisting remains difficult
- 47:17SimToolReal evaluation: Pose tracking caused roughly 60% of failures
- 51:21Niko West: Start robotics companies with teleoperation and a paying customer
- 56:27Niko West: Build an early teleoperation learning loop and customer-specific evaluation
- 59:41Niko West: Robotics data tooling must handle new designs, multimodal data, and full-stack operations
- 1:03:43Niko West: Rerun’s open-source SDK and teleoperation support lean robotics startups
- 1:08:30Bill Jiao: World action models predict future frames but face DreamZero-scale compute costs
- 1:13:35General Instinct: Distillation cuts world-action sampling to 1–2 steps and 500 ms per chunk
- 1:18:33General Instinct: Future-video dynamics, action sampling, and adaptive key-frame detail
這是 Tier 1 公開摘要
每章重點、段落總結、心智圖由分享者控制是否公開。想看完整分析?自己提交一支。
同頻道的其他分析
Patrick Collison: Is AI Breaking the Lean Startup Playbook?Y CombinatorPatrick Collison says Stripe’s new-business starts are nearly 2x year over year, making this a better time than ever to found a company.
The Case For Data Centers In SpaceY CombinatorStarcloud CEO Philip Johnston explains how an H100 reached orbit and why 88,000 satellites could host space-based AI data centers.
Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The WorkY CombinatorWaymo co-CEO Dmitri Dolgov says the demo took 18 months, but building a safe robotaxi product took 15 years.
Garry Tan: Own Your IntelligenceY CombinatorGarry Tan says personal AGI can make one founder 400x more productive—and argues your AI skills should stay in your own repo.
Max Hodak: Average Is Not Good EnoughY CombinatorScience CEO Max Hodak says Helix and faster iteration drive startup speed, while one retinal implant patient read a 300-page novel.
相關主題的分析
Chelsea Finn: This is the State of the Art in RoboticsY CombinatorChelsea Finn says Physical Intelligence’s PIO7 controls diverse robots out of the box, while reinforcement learning doubled throughput.
OpenAI 創始成員加入 Anthropic:為什麼押注沒人看好的預訓練? | S2E58矽谷輕鬆談 Just Kidding TechAndrej Karpathy 加入 Anthropic 預訓練團隊,可能用 Mythos 與 auto research 推進 AI 自主研究。
Why Physical AI Is the Next Big Opportunity | Deep Dives with a16za16z Deep DivesDiode Computers aims to automate circuit-board design in two years, while Unlimited Industries targets end-to-end construction automation within a decade.
Why Fei-Fei Li Is Betting on Spatial Intelligencea16zFei-Fei Li and World Labs unveil Atlas, using new-view prediction to cut 3D capture from hundreds of images to just three.
Why The Harness Matters More Than The Model | YC Paper ClubY CombinatorYC researchers show harnesses can lift ARC-AGI scores from 30% to 95%, while Prime Agent and QM enable long-running, collaborative agents.