What Today’s Best Models Still Can’t Do in Math

a16z · 2026-09-01 · 63 分鐘
https://www.youtube.com/watch?v=tQI35CSNB08影片總結
Daniel Litt praises AI’s Erdős unit-distance result but warns that proofs alone cannot replace mathematical understanding or human curiosity.
University of Toronto mathematician Daniel Litt calls AI’s autonomous solution to the Erdős unit-distance problem the most impressive result so far: it imported classical techniques into a new area and inspired mathematicians to find counterexamples to other questions, including the real sum-product conjecture. He says models excel at computation and applying known methods but remain weak at intuition, theory-building and checking arguments at a big-picture level. Litt’s own work with AI is most useful for coding, examples and proving a lemma after he improved its statement. He warns that AI-generated papers and repeated proofs can reward output over understanding, while mathematics depends on diverse researchers pursuing their own questions. Looking ahead, he argues that schools and research institutions should use AI to deepen human thinking, not replace it; even with a three-year-old daughter, he sees learning math as a way to think clearly and understand the world.
章節
- 0:00Intro: Litt’s Standard for Understanding and the AI Unit Distance Result
- 1:00Meet Daniel Litt: The Toronto Mathematician Rethinking AI’s Role
- 2:24The Erdős Unit Distance Problem: A 1960s Technique Opens New Counterexamples
- 6:12What AI Does for Mathematicians: Computation, Known Techniques, and Hinted Theory Building
- 12:12Intuition and Mathematical Practice: Open Problems and Litt’s Algebraic-Geometry Analogy
- 16:41Mathematical Taste and Discovery: Elliptic-Curve Data, AI-Assisted Coding, and New Questions
- 20:26Deep Thinking vs. Pattern Matching: Why Long-Standing Conjectures Need New Ideas
- 26:17Why the Unit Distance Result Was Actually Creative: Litt’s Lemma Revision and Conceptual Proof
- 33:33How Should the Math Community Adapt to AI? Incentives, Human Understanding, and Repeated Proofs
- 38:56How Should the Math Community Adapt to AI? Preserve Diverse Human Interests and Agency
- 43:19Where AI Will Impact Applied Math First: Education, Group Theory, and Student AI Use
- 46:24Taking Advantage of AI Without Losing the Craft: Deepen Understanding, Not Just Output
- 49:21Comparing Anthropic vs OpenAI in Math: Claude Fable’s Reported Rank-30 Elliptic Curve
- 51:34Why Some Labs Have Gone More Secretive: AI’s 800-Page Proof Is Beyond Reliable Verification
- 54:47Why Some Labs Have Gone More Secretive: Harnesses, SQLite in Rust, and Proof-Checking Limits
- 59:40Raising a Mathematician: Sophia Learns to Count, Platonic Solids, and Group Theory
這是 Tier 1 公開摘要
每章重點、段落總結、心智圖由分享者控制是否公開。想看完整分析?自己提交一支。
同頻道的其他分析
How Cursor Built One of AI’s Fastest-Growing Companiesa16zCursor bet on the human-model interface, grew rapidly without early sales hires, and later expanded through Graphite and founder acquisitions.
Why Top Founders Are Racing Into AI Infrastructurea16za16z’s Machine Age Fund targets AI infrastructure as GPU supply is booked to 2028 and memory demand needs three years of capacity.
Why AI Demand Is Outrunning Compute Supplya16zGavin Baker argues AI compute demand will outpace supply, with sub-one-year infrastructure paybacks and orbital data centers approaching.
Inside Moderna’s Biggest mRNA Test Since COVIDa16zModerna and Merck’s personalized mRNA vaccine beat Keytruda alone in a Phase 3 melanoma trial, after 1,000 cancer-vaccine trials failed.
Why AI Agents Could Finally Reinvent the Credit Carda16zAffirm CEO Max Levchin and Alex Rampell trace Affirm’s origin and argue AI agents could reshape credit card payments.
相關主題的分析
進入Anthropic靠「這個」?用AI讓人更累?頂尖AI公司工作甘苦談 ft. 熊 @SiliconCafeChat哈佛姐夢遊矽谷 AliceInSiliconWonderlandAnthropic 工程師熊用 Claude 寫程式、人工審查 Infra 程式碼,也指出 AI 提高產出與工作期待,讓工程師更忙。
Yikes.ColdFusionColdFusion traces Claude’s role in military targeting, Anthropic’s refusal of surveillance terms, and OpenAI’s controversial Pentagon deal.
Claude 最強模型 Fable 5 深入解析:打著安全旗號,其實在搞反競爭? | S2E61矽谷輕鬆談 Just Kidding TechClaude Fable 5 長任務明顯進步,卻頻繁誤降級 Opus 4.8,Anthropic 的安全與透明度引發質疑。
AI巨头们之间的资本混战,到底是个什么情况?小Lin说微軟持有 OpenAI 約 27% 股權,亞馬遜與 Google 轉向雙押 Anthropic,AI 投資也綁定天價算力訂單。
連續押中 Facebook、Uber 成功出場,與矽谷大神聊聊選公司的眼光、創業與 AI feat. Jarsy 創辦人秦漢 | S2E64矽谷輕鬆談 Just Kidding Tech秦漢從 Facebook、Uber 談到 Jarsy,主張先服務用戶、重視合規,也分享 AI 如何改變工程工作。