分享這篇:

影片總結

YC researchers show harnesses can lift ARC-AGI scores from 30% to 95%, while Prime Agent and QM enable long-running, collaborative agents.

At YC Harness Club, speakers argue that agent harnesses are not mere scaffolding: the same model weights that score 30% on ARC-AGI can reach 95% with a better setup, and NVIDIA’s AVO reportedly hits 100%. The opening talk traces harnesses from prompts and tool use to reflective, multi-agent, self-improving systems. Prime Agent researcher Seth Karten describes persistent subagents, editable memory and prompts, and a test-time scaling result of 95.5% for Opus; a seven-day Factorio run used 633 agents and 23 million output tokens. Stanford’s OpenJarvis project aims to run personal AI locally, citing up to 800x lower inference cost and cloud models that can optimize on-device setups. YC’s Josh and Rean then explain QM, an open-source assistant for every employee: it connects Slack, company tools, databases, and flexible sandboxes, while using human review for database writes and permissions to prevent sensitive information leaks.

章節

  1. 0:00Why harnesses matter: ARC-AGI rises from 30% to 95%
  2. 4:27Building an auto-researcher by accident: eight ideas on H100 nodes
  3. 7:13A five-minute harness history: from GPT-2 loops to recursive agents
  4. 13:56Self-improving harnesses: DSPy prompts, Darwin code, and online learning
  5. 17:22Tonight’s speakers: Seth Karten, OpenJarvis, and new YC Labs head Josh Rea
  6. 18:35Seth Karten’s Prime Agent: persistent RLM subagents and live harness edits
  7. 21:44Context as L1, L2, and L3 cache: compaction, RAM, and cleanup
  8. 24:51From Turing machine to von Neumann computer: external memory and persistent agents
  9. 28:33Messaging between agents and evaluating long-horizon performance fairly
  10. 30:04ARC-AGI results: Prime Agent reaches 95.5% with Opus
  11. 33:09Emulator Bench and GPU kernels: Game Boy Color tests and a 633-agent Factorio run
  12. 39:21The five primitives of a personal AI stack: OpenJarvis targets fully on-device execution
  13. 42:47Letting cloud models optimize your local stack: tuned OpenJarvis configurations outperform defaults
  14. 43:53800x cheaper than the cloud: OpenJarvis local inference lowers cost and latency
  15. 45:58Josh France and Regan Bell: QM gives YC employees a customizable work agent
  16. 47:29A history of YC's internal agents: from a January 2025 general agent to Slack coding bots
  17. 49:24OpenClaw and a fleet of 50+ Hermes agents exposed the costs of manual management
  18. 51:04Pulling the brain out of the sandbox: QM centralizes state in Postgres
  19. 54:43Letting the agent choose its own sandbox and model at runtime
  20. 57:16The grind tool: goal budgets keep agents working longer
  21. 58:50Agents don't understand social context: fine-grained permissions prevent information leaks

這是 Tier 1 公開摘要

每章重點、段落總結、心智圖由分享者控制是否公開。想看完整分析?自己提交一支。

同頻道的其他分析

相關主題的分析