产品发布/更新
1 updates
-
Grok 支持分析任意视频
Grok 可以分析任何视频 https://grok.com/share/bGVnYWN5_8013f7a3-f604-4351-8cd7-acecf3ef165b
Open source
2026-08-03 · AI HOT + Model Companies + Market News + AI Papers + Voices + Trends + Follow Builders · generated 2026/08/03 02:26 · builder feed 2026/08/02 15:09
1 updates
Grok 可以分析任何视频 https://grok.com/share/bGVnYWN5_8013f7a3-f604-4351-8cd7-acecf3ef165b
Open source1 updates
Codex 高阶玩法:让 Sol 在 `~/.codex/agents/` 下创建 `luna-worker.toml` 子代理,模型设 `gpt-5.6-luna`、reasoning effort 设 max,Sol 负责拆任务与审代码,具体实现自动委托给 Luna Max。
Open source过去 24 小时暂无单独归类的行业动态。
海外 · 1 updates
主要信号集中在Agent / 工作流、模型进展:OpenAI Just Cut GPT-5.6 Price by Up to 80%, Making Long-Running AI Agents Far More Practical
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
OpenAI cut GPT 5.6 Luna prices 80% and Terra 20% on July 30, 2026. See why AI agents, not chat users, are the real winners.
Open news海外 · 1 updates
主要信号集中在模型进展、安全 / 监管:Anthropic Says Claude Breached Three Real Companies During Safety Test
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Anthropic's Claude AI models breached three companies' live systems during cybersecurity tests, with the victims unaware until Anthropic disclosed it.
Open news海外 · 2 updates
主要信号集中在Agent / 工作流、模型进展、安全 / 监管:Google DeepMind launches Gemini Robotics 2 with full humanoid body control and multi-robot teamwork;Google DeepMind says Gemini Robotics 2 enables full body control
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Google DeepMind's Gemini Robotics 2 gives humanoid robots whole-body control, multi-robot coordination, and a new safety benchmark called ASIMOV-Agentic.
Open newsGemini Robotics 2 enables robots to reason through every movement, unlocking a broad range of tasks, DeepMind said.
Open news海外 · 1 updates
主要信号集中在模型进展、商业化 / 资本:ORCL stock gains after Google’s Gemini joins Meta, OpenAI, xAI in Oracle’s AI model lineup
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Oracle announced on Thursday it would add Google’s Gemini AI models to Oracle Cloud Infrastructure. ・The move expands Oracle’s existing lineup of third-party AI models, including offerings from Cohere ...
Open news海外 · 3 updates
主要信号集中在Agent / 工作流、模型进展:Microsoft is making a Copilot super app to end your AI app juggling;Microsoft CEO Satya Nadella Says 'Every Model Is Substitutable' — What That Means For OpenAI
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Satya Nadella confirmed Microsoft is merging Copilot chat, Cowork, Autopilots, and coding tools into one AI super app for consumers and businesses, launching later this year.The Latest Tech News, Deli ...
Open newsMicrosoft Corp. used its fiscal fourth-quarter earnings call to deliver upbeat updates on Azure, Copilot and AI spending. But one of the company’s most consequential messages for investors wasn’t ...
Open newsMicrosoft plans to launch a unified Copilot super app this year, bringing AI chat, coding, Cowork, and autonomous Autopilots into one platf ...
Open news海外 · 1 updates
主要信号集中在算力 / 推理、商业化 / 资本:The AI infrastructure market is shifting to inference. 5 stocks to play this future $1.3 trillion market opportunity
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
The inference market is projected to double the size of the training market in the coming years.
Open news海外 · 1 updates
主要信号集中在商业化 / 资本:Amazon’s Andy Jassy sees a $1 trillion cloud built on businesses like the PGA Tour—which swapped its server trucks for AI broadcasts
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
AWS just posted $42.2 billion in revenue with its fastest growth in 18 quarters and the AI platform saw more spending last quarter than in every prior quarter since launch combined.
Open news海外 · 1 updates
主要信号集中在模型进展:Tim Cook's Parting Gift to iPhone Users: A Siri AI Subscription
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Tim Cook confirmed heavy Siri AI users will need a paid iCloud+ plan. Here's what Apple's new AI subscription model means for your iPhone in 2026.
Open news海外 · 1 updates
主要信号集中在模型进展:Grok imagine video update adds 1080p, voice cloning, and seven-reference scene control
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Grok Imagine Video 1.5's July 31 update adds text-to-video, native 1080p at $0.25/sec, voice reference for face-and-voice consistency, and multi-reference control with up to seven anchors.
Open news海外 · 1 updates
主要信号集中在模型进展:How France's Mistral is taking on Silicon Valley's AI behemoths
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Paris-based company couldn't rival the likes of OpenAI and Anthropic with its models. So it bet on independence over scale ...
Open news海外 · 1 updates
主要信号集中在模型业务动态:Cohere and Carahsoft Partner to Bring Secure, Sovereign AI Deployment Solutions to the Public Sector
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Cohere, the world's leading sovereign AI company, and Carahsoft Technology Corp ., The Trusted Government IT Solutions Provider ® , today announced a partnership to expand access to secure, sovereign ...
Open news国内 · 1 updates
主要信号集中在算力 / 推理、商业化 / 资本:Chip stocks slide in AI unwind
为什么值得看适合用来观察国内模型厂商的产品节奏、开源/闭源路线和商业化落点。
A selloff in chipmakers gathered pace last week, driving the high-profile group of stocks to a bear market on worries that the artificial intelligence (AI) spending spree is becoming harder to justify ...
Open news国内 · 1 updates
主要信号集中在模型进展、安全 / 监管:What to know about Moonshot AI and its new open-weight model Kimi K3
为什么值得看适合用来观察国内模型厂商的产品节奏、开源/闭源路线和商业化落点。
The Chinese AI startup’s massive new model is challenging OpenAI and Anthropic, fueling a debate over AI safety. The release of the Chinese AI model Kimi K3 was a flashpoint in the AI world, ...
Open news国内 · 1 updates
主要信号集中在模型进展:Chinese AI model rivals ChatGPT, jolting Silicon Valley
为什么值得看适合用来观察国内模型厂商的产品节奏、开源/闭源路线和商业化落点。
Another powerful new artificial intelligence model from China took the U.S. tech industry by surprise Friday, the latest sign that Chinese startups that publicly release their “open-source” AI ...
Open news国内 · 1 updates
主要信号集中在安全 / 监管:MiniMax H3 opens AI video to developers: Copyright lawsuit clouds every clip
为什么值得看适合用来观察国内模型厂商的产品节奏、开源/闭源路线和商业化落点。
MiniMax H3 launches today as AI video editing leader per Artificial Analysis, generating native 2K video at $7.80 per minute — less than one-third the cost of rivals — while the Hailuo platform faces ...
Open news这里是市场情绪线索,用来辅助判断风险偏好、科技股和 AI 资产预期,不当作投资建议。
Fed、通胀、债券收益率 · 2 updates
这条偏「利率/通胀、股市情绪」信号,当前解读为中性但值得观察。可能影响成长股折现率、美元和长端利率预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:Bond traders aren't counting on Warsh and the Federal Open Market Committee (FOMC) to sit on their proverbial hands for much longer.
Open news这条偏「利率/通胀、美元/商品」信号,当前解读为偏利空/风险偏好收缩。可能影响成长股折现率、美元和长端利率预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:The Federal Reserve held interest rates steady this week, but some central bank officials still say rate hikes are necessary.
Open newsS&P 500、Nasdaq、波动率、资金情绪 · 1 updates
这条偏「利率/通胀、股市情绪」信号,当前解读为偏利多/风险偏好改善。可能影响成长股折现率、美元和长端利率预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:(Editor’s note: The ETFs data was updated.) U.S. stock futures rose on Friday, as the Dow Jones, S&P 500, and Nasdaq 100 indices advanced, following Thursday’s higher close. On Thursday, cooler June ...
Open news云厂商、AI 资本开支、利润率 · 3 updates
这条偏「AI芯片/半导体、财报/AI Capex」信号,当前解读为偏利多/风险偏好改善。可能影响AI 概念股、半导体链和算力资本开支预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:Big Tech's latest earnings indicate robust demand for AI infrastructure, with Microsoft, Amazon, and Alphabet showing significant cloud growth.
Open news这条偏「股市情绪、财报/AI Capex」信号,当前解读为中性但值得观察。可能影响云厂商利润率、AI 资本开支和上游算力需求;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:The size of their respective cloud businesses matters, but there are other important considerations.
Open news这条偏「股市情绪、财报/AI Capex」信号,当前解读为偏利多/风险偏好改善。可能影响云厂商利润率、AI 资本开支和上游算力需求;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:Amazon stock soared as the tech giant's cloud growth blew past Wall Street estimates.
Open newsHugging Face Daily Papers + arXiv recent AI/ML · 12 papers · fallback summaries
来自 Hugging Face Daily Papers,主题偏「Agent、RAG/Memory、Reasoning」。摘要显示它主要讨论 Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However, agent memory is still primarily implemented through extern... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However, agent memory is still primarily implemented through external modules, leaving the native memory capability largely unexplored. In this paper, we take a first step toward this direction by introducing memory foundation models, which empower foundation models with native memory capabilities. We formalize native memory from two perspectives: a persistent and dynamically evolving memory state within the backbone, and native memory procedures that autonomously store and utilize information through model computation. We show that native memory offers advantages in architecture, end-to-end optimization, and efficiency. Based on this formulation, we propose Metis, the first prototype of memory foundation models. Metis introduces a new architecture that equips a foundation model with a native memory state, allowing historical information to be compressed into the model and accessed through memory attention. We construct large-scale memory-specific training data and introduce multiple optimization objectives to acquire these native memory procedures through mid-training. The online memory maintenance of Metis is gradient-free, and the memory update requires only a forward pass. At inferenc...
来自 Hugging Face Daily Papers,主题偏「Agent、RAG/Memory、Post-training/Alignment」。摘要显示它主要讨论 Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system f... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo). On this stack we post-train Frontis-MA1 (35B) as a meta-evolution agent for MLE, aligning post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover): the same operators are trained via execution-grounded SFT and RL on data deduplicated against all evaluation benchmarks, then composed into long-horizon search, coupling learning and evolution in a single loop. On MLE-Bench Lite under a 12-hour per-task budget on one RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improves Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo, and reaches 71.21% with OpenMLE-Evo-Max (benchmark-independent experience priors and asynchronous search), exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. On held-out NatureBench Lite, both components transfer: with the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%; with the model fixed, swapping in Ope...
来自 Hugging Face Daily Papers,主题偏「Agent、RAG/Memory、Reasoning」。摘要显示它主要讨论 Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thoug... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process. To address this limitation, we introduce VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought. Given a text prompt, a coding agent synthesizes a Blender program that explicitly specifies the scene and its temporal evolution. The executable simulation engine runs the program to produce a deterministic spatiotemporal draft, which is subsequently transformed into a photorealistic video by a generative video engine through draft-conditioned editing. This decomposition separates process-level reasoning from high-fidelity visual realization. To adapt the video editor to simulated drafts, we construct VideoCoCo-3K, a curated dataset of draft-instruction-target triplets. VideoCoCo improves the OmniWeaving baseline from 0.475 to 0.558 on PhyGenBench and from 52.18 to 77.88 on VBench-2.0, achieving the best average score on both benchmarks. These...
来自 Hugging Face Daily Papers,主题偏「RAG/Memory、Reasoning、Eval/Data」。摘要显示它主要讨论 Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察知识工作、企业搜索、长期记忆和本地资料库产品的新实现路径。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small scale. In this work, we present Memory Decoder at Scale, scaling memory models up to 6.9B parameters and pretraining them on 300B tokens. At this data scale, the combined cost of indexing and search makes a standard Faiss pipeline infeasible. We address this bottleneck with a distributed pipeline for Faiss indexing and retrieval, together with sparse, batch-wise loading of kNN distributions. Across model scales, we find that allocating more parameters to memory yields a better parameter-performance tradeoff than scaling the base model alone. On 17 benchmarks, pairing a 6.9B general memory with Pythia-410M raises its average score from 29.86 to 37.34, surpassing Pythia-12B (37.24) with 39% fewer total parameters. For Qwen3 Base models ranging from 0.6B to 14B, 1.7B domain memories improve the average score across the three domains by more than 9 points at every scale. Overall, our results demonstrate that independently scaling pretrained memory offers a more parameter efficient path to improving language model performance.
来自 Hugging Face Daily Papers,主题偏「Agent、RAG/Memory、Reasoning」。摘要显示它主要讨论 The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink ag... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic visual reasoning through two key dimensions of tool use: Mode Adaptiveness (MA) and Tool Effect (TE). Mode Adaptiveness characterizes whether an MLLM can recognize when tools are truly necessary and invoke them accordingly, thereby avoiding unnecessary computational overhead while improving performance on challenging problems that require tool assistance. Tool Effect characterizes the actual impact of tool use: tools should extend the model's capabilities on problems unsolvable through text-only reasoning, while avoiding additional errors on problems that the model can already solve without tools. We conduct a comprehensive analysis to quantify these two properties and empirically reveal that existing agentic visual reasoning models exhibit limited Mode Adaptiveness, while the gains produced by tool use on hard examples are largely offset by the harm introduced on easy examples that the models can already solve. Motivated by these observations, we propose Beacon, a novel agentic visual reasoning model that achieves stronger overall performance, improved Mode Adaptiveness, and genuine tool-induced performance gains. A...
来自 Hugging Face Daily Papers,主题偏「Agent、RAG/Memory、Reasoning」。摘要显示它主要讨论 Retrieval-augmented generation (RAG) spans lexical and dense retrieval, graph-based indexing, and agentic search, but these paradigms are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost scaling unclear. To bridge this... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
Retrieval-augmented generation (RAG) spans lexical and dense retrieval, graph-based indexing, and agentic search, but these paradigms are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost scaling unclear. To bridge this gap, we present a controlled study that varies corpus size along 28 strictly nested tiers spanning roughly 450-fold, while holding questions and a fixed bedrock of relevant and adversarial documents unchanged. Under one reader model and one judging protocol, we measure official accuracy, construction and query tokens, and latency. The results reveal a scale-dependent crossover rather than an unconditional winner. File-System Agent leads at the smallest shared tiers, but its sequential exploration costs 39 times more query tokens at the bedrock and becomes less effective as the search space grows. Around 10 million corpus tokens, BM25 overtakes it and leads at every larger shared tier, with a margin approaching 20 points at full scale. BM25 also anchors the low-cost end of the Pareto frontier without LLM-based construction. Dense retrieval remains efficient but less accurate, whereas graph-based RAG encounters construction walls before deployment scale and its scalable variants remain below BM25 at shared tiers. Overall, corpus growth increasingly favors global candidate ranking: lexical retrieval is the strongest scalable...
来自 Hugging Face Daily Papers,主题偏「RAG/Memory、Multimodal、Eval/Data」。摘要显示它主要讨论 Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment t... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察知识工作、企业搜索、长期记忆和本地资料库产品的新实现路径。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions....
来自 Hugging Face Daily Papers,主题偏「Agent、Eval/Data」。摘要显示它主要讨论 Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring ca... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA's output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles. Alongside a general quality rubric, we construct personalized rubrics to measure user satisfaction; on held-out annotated data, the persona...
来自 Hugging Face Daily Papers,主题偏「Agent、Reasoning、Multimodal」。摘要显示它主要讨论 Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can r... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason about the overall task but often miss the visual details that determine success, while specialist vision models can capture those details but cannot translate them into task-level decisions. In this work, we propose SpatialCLI, a framework that teaches VLMs to reason with spatial tools and progressively internalize the specialist perceptual capabilities they provide. SpatialCLI proceeds in three stages: (1) Call exposes specialist vision models as spatial tools to augment the VLM's perception; (2) Learn uses Cold-Start SFT and agentic RL to improve tool use; and (3) Internalize verbalizes successful tool-use trajectories to internalize specialist perceptual capabilities. We further introduce SpatialCLI-Bench, a 516-example benchmark for compositional perception across localization, segmentation, depth, and pose. On MindCube, SpatialCLI raises Qwen3-VL-8B-Instruct from 29.3% to 84.6% with tools, surpassing GPT-5.6 Sol with tools (72.1%), while retaining 73.8% without tools after internalization.
来自 Hugging Face Daily Papers,主题偏「RAG/Memory、Reasoning、Multimodal」。摘要显示它主要讨论 Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narro... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察知识工作、企业搜索、长期记忆和本地资料库产品的新实现路径。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework comprising See2ThinkBench and Visual Action-of-Thought (VAoT). See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 task categories spanning 2D structured, 3D scene, and real-world reasoning. VAoT records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled inference settings. Evaluating representative proprietary and open-source multimodal models, we find that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks. Process analysis further shows that models usually select relevant visual operations, while faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains. Under task-relevant corrupted feedback, models exhibit behavioral dependence on visual states, with accuracy dropping by over...
来自 Hugging Face Daily Papers,主题偏「Agent、RAG/Memory、Reasoning」。摘要显示它主要讨论 Retrieving past experiences has become a common strategy to enhance large language model agents. However, most existing memory-augmented agents treat retrieved experiences as static records to be replayed verbatim, injecting them into the context regardless of... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Retrieving past experiences has become a common strategy to enhance large language model agents. However, most existing memory-augmented agents treat retrieved experiences as static records to be replayed verbatim, injecting them into the context regardless of whether they align with the agent's current situation. This ``replay'' paradigm ignores the gap between the abstract, general nature of stored experience and the concrete, ever-changing states encountered at decision time, frequently causing negative transfer. In contrast, humans rarely recall past experiences verbatim; instead, they reorganize and adapt retrieved memories to fit the present context. Inspired by this, we propose MemHarness, a framework that equips LLM agents to actively harness and reconstruct past experiences based on the present context. At each decision step, a unified policy model critiques and reconstructs the retrieved experience conditioned on the current state, producing context-grounded guidance before acting. This reconstructive ability emerges naturally through end-to-end training with GRPO. Experiments on ALFWorld and WebShop show that MemHarness substantially outperforms pure RL and static memory-augmented baselines, demonstrating strong robustness in out-of-distribution (OOD) scenarios. Furthermore, our analyses reveal that this reconstruction objective not only prevents negative transfer bu...
来自 Hugging Face Daily Papers,主题偏「RAG/Memory、Multimodal、Eval/Data」。摘要显示它主要讨论 Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embe... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察知识工作、企业搜索、长期记忆和本地资料库产品的新实现路径。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache. Trained on only a small image-QA dataset, ReToken yields consistent gains across image and video benchmarks: on Visual Haystacks it improves Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points (>20% relative), and on LVBench it transfers zero-shot to long video for an 8.0-point gain with Qwen3VL-8B. Thanks to its lightweight design, both training and long-video inference fit on a single H100. Code is available at: https://github.com/avaxiao/ReToken
Using cached daily papers because the latest fetch failed.
dair-ai/AI-Papers-of-the-Week · 10 papers
Harness Handbook:为自我演化 Agent 建立行为地图
随着 agent 能自行修改 harness,真正的瓶颈逐渐从写代码变成定位某个行为背后的全部文件。Harness Handbook 用静态分析与 LLM 自动把代码库整理成三层、可回溯源码的行为地图:从系统架构与数据流,到组件职责,再到具体代码单元。其 BGPD 机制先从高层行为逐步缩小范围,再用当前源码验证候选位置,减少 agent 在看似合理但错误的文件上修改。
为什么值得看它把复杂 harness 的可理解性做成可再生基础设施,为人和 agent 安全协同维护生产系统提供了实用方案。
读原文判断正在建设会自改代码或持续演化的 agent 平台,值得读架构和 BGPD;普通 workflow 看方法概览即可。
Teams now let agents evolve their own harnesses, but the harness itself becomes a sprawling codebase where finding every file behind one behavior is often harder than writing the edit. Harness Handbook attacks this by turning a harness into a behavior-centric map that stays linked to source. ● Synthesized automatically: The Handbook is built from the harness codebase through static analysis and LLM-assisted structuring, so the representation is generated rather than hand-maintained and can be regenerated as the harness changes. ● A three-level map: It progresses from an L1 system overview of architecture, execution model, and data flow, to L2 component overviews with responsibilities, inputs, outputs, and state, down to L3 source-backed unit details, with a navigation pane for cross-stage tracing. ● Behavior-Guided Progressive Disclosure: BGPD walks an agent from a high-level behavior to the relevant implementation, then verifies candidate locations against the current source, so edits land on the right files instead of plausible-looking wrong ones. ● Why it matters: As self-improving harnesses grow, the bottleneck shifts from writing changes to locating them, and a readable, navig...
From Memory to Skills:让 Agent 经验沉淀为可调用技能
MSCE 是一个无需训练的记忆—技能协同演化框架,解决传统 agent 只把历史轨迹当被动上下文、无法把经验直接复用的问题。它把经验分为有依据的步骤轨迹、可复用过程策略和环境认知三层,并将预估收益为正的策略固化为带证据、适用边界、验证规则与可靠度的技能卡。系统还用局部反思把稀疏终局反馈回填到轨迹价值,持续治理技能的生成、更新和淘汰。
为什么值得看它给“记忆如何变成技能”提供了完整治理链路,适合长期任务 agent 和跨会话经验复用。
读原文判断做 agent memory、skill learning 或长期自治系统建议细读;仅做短任务 agent 可先吸收三层记忆结构。
Most agent memory systems retrieve past traces as passive context, so hard-won experience never becomes something the agent can directly execute. MSCE, a training-free memory-skill co-evolution framework, instead governs how experience turns into callable skills for long-horizon LLM agents. ● Three-level governed memory: Experience is organized into L1 grounded step traces, L2 reusable procedural policies, and L3 declarative environmental cognition, giving the agent a structured store rather than a flat log of prior runs. ● Skills with evidence: L2 policies with positive estimated gain are crystallized into callable skill cards that retain evidence links, applicability boundaries, decision guidance, verification rules, and reliability estimates, so a skill carries the context needed to trust it. ● Reflection-weighted value backfilling: Sparse terminal feedback is propagated through dense local self-reflections to produce evidence-calibrated trace values, which then govern how memory and skills evolve and get retired. ● Why it matters: On EvoAgentBench and LoCoMo, MSCE outperforms state-of-the-art skill-augmented and memory-driven baselines with strong cross-domain transfer, pointin...
PRO-LONG:用可搜索完整日志替代复杂记忆压缩
长时任务常在“保留多少历史”和“如何高效取回”之间权衡,而更丰富的摘要反而可能掩埋关键细节。PRO-LONG 选择保留完整、结构化交互日志,并让 coding agent 按需用工具搜索,不预先丢弃观察,也不构建复杂专用记忆。在 ARC-AGI-3 公开游戏集上,它平均提升 18 分,最高以 76.1% pass@1 匹配或超过专门 harness,同时减少 4.2 到 5.8 倍 token。
为什么值得看它证明简单的“完整日志加程序化检索”可能在准确率和成本上胜过精巧记忆工程,落地门槛很低。
读原文判断做探索型长任务或纠结记忆架构时建议看实验;已有成熟检索日志的团队可直接验证这一基线。
Long-horizon tasks force a harness to decide what to save from a long stream of observations and how to load it back into context, and richer summaries usually make the exact detail you need harder to retrieve. PRO-LONG sidesteps this tradeoff with programmatic memory. ● Keep everything, search it: Rather than compressing history into bespoke memory, PRO-LONG keeps a complete, structured interaction log and leans on coding-agent tooling to search that history on demand, so no observation is discarded up front. ● A minimal framework: The design is deliberately lightweight, avoiding hand-built memory harnesses and instead treating the full log as a searchable artifact the agent queries when it needs a specific past detail. ● Strong, cheaper results: On the full ARC-AGI-3 public game set, it improves over a base coding agent by an average of 18.0 points across frontier models, and matches or exceeds specialized state-of-the-art harnesses at up to 76.1% pass@1 while using 4.2 to 5.8 times fewer tokens. ● Why it matters: It shows that for exploratory, long-horizon settings, a simple searchable log can beat elaborate memory engineering on both accuracy and cost, which is a practical reci...
Global Workspace in LLMs:定位模型推理的关键内部空间
Anthropic 这项可解释性研究用 Jacobian lens 找出模型在任一时刻准备表达的残差流方向,并把这些方向组成的空间称为 J-space。它只解释不超过约 10% 的激活方差,主要位于网络中层,却能承载可报告、可主动维持、可用于静默推理并传给后续计算的内容。抑制 J-space 后,模型仍能理解输入和流畅说话,却失去复杂内部推理能力,呈现类似“全局工作空间”的功能特征。
为什么值得看它为判断 verbalized reasoning 是否真正承载计算提供了机制证据,也影响 CoT 监控和 steering vector 设计。
读原文判断做模型可解释性、CoT 或内部表征控制应读方法与干预实验;产品侧掌握结论和边界即可。
This Anthropic interpretability work gives a mechanistic account of when a model's verbalized reasoning is load-bearing and when it is not. It identifies a small, privileged set of internal representations that behaves like the global workspace some neuroscientists tie to conscious access. ● A new lens: The Jacobian lens, or J-lens, surfaces the directions in the residual stream that a model is poised to verbalize at any point, and the collection of these directions is named the J-space. ● Workspace-like roles: J-space contents can be reported, deliberately summoned and held, used to carry the intermediate steps of silent reasoning, and passed as arguments to downstream computation, matching the functional signature of a global workspace. ● Small but decisive: The J-space accounts for no more than roughly 10% of activation variance and appears mainly in the middle of the network, yet suppressing it leaves the model able to parse input and speak fluently while it loses the ability to perform complex internal reasoning. ● Why it matters: For anyone building on chain-of-thought or steering vectors, this clarifies which internal representations actually drive reasoning, and the authors...
GAMUT:评测回答是否完整,而不只是正确
多数事实性评测只看回答中的主张是否正确,却忽略该说的内容是否都覆盖。GAMUT 针对完整性建立结构化 meta-rubric,用层级结构表达开放集合、顺序过程和事实关系,再机械编译成低方差、可由 LLM 判分的二元清单。基准包含 1813 个来自真实可穿戴设备图像的问题,覆盖 10 个领域,并提供文本版本。14 个前沿和开源模型中最好成绩仅 58.7%,说明完整性仍是明显短板。
为什么值得看它补上事实评测的召回率维度,并给复杂答案建立可复用、可稳定机器评分的 rubric 方法。
读原文判断做 RAG、视觉问答或评测体系建议读 rubric 构建;只关注模型选型可先看各模型结果。
Most factuality evaluation measures precision, whether the claims in an answer are correct. This Meta AI work targets the harder and mostly ignored half, completeness, meaning whether an answer covers everything it should, and packages it as the GAMUT benchmark. ● Completeness is structured: The facts a complete answer should contain rarely form a flat list, since they involve open-ended sets where coverage matters, ordered processes, and relationships among facts that independent boolean checks cannot capture. ● Two-level meta-rubrics: A structured meta-rubric encodes the organization and importance of required content, then compiles mechanically into a flat checklist of binary, machine-gradable items that an LLM judge can score reliably, keeping rich structure while inheriting low-variance grading. ● Grounded and verified: The benchmark holds 1,813 questions grounded in real wearable imagery across 10 diverse domains, each paired with an evidence-backed rubric verified by expert annotators, and a text-only variant is released for models without vision. ● Why it matters: Across 14 frontier and open-weight models the benchmark stays genuinely hard, with a best score of 58.7% from G...
Progressive Disclosure, Measured:渐进式披露何时真正有效
Agent Skills 常把长文档封装成按需加载的层级内容,但过去支持渐进式披露的证据多为经验判断。该研究在 InfiniteBench 上,跨三种 agent harness 与三类模型,对比原文导航、多种 Skills 打包设计和经典混合检索。结果显示:当 harness 自身不擅长导航时,单本长文档上的收益显著;若强 harness 已能切分和检索文本,增益接近于零。因此价值取决于周边执行系统,而非 skill 格式本身。
为什么值得看它首次用受控实验校准流行的 Skills 设计模式,可避免为了形式增加无收益的文档工程。
读原文判断在设计大型知识 Skills 或 RAG 替代方案时值得读;已有强检索 harness,应先做对照实验再迁移。
Agent Skills package expertise into folders an agent loads on demand, and progressive disclosure exposes only what a query needs, from a short description down to specific passages. Practitioners adopted this pattern fast for book-length tasks, but the supporting evidence was anecdotal until now. ● A controlled study: The authors run the first controlled comparison of progressive disclosure, pitting raw-document navigation and several Agent Skills pack designs against a classical hybrid retriever across three agent harnesses and three model families on InfiniteBench. ● The gain is harness-dependent: On a single book, progressive disclosure helps a lot when the agent navigates the raw document poorly, but the benefit falls to near zero when a strong harness already divides and retrieves the text on its own. ● Not a universal win: Because the pattern's value hinges on the surrounding harness rather than the skill format alone, treating progressive disclosure as an automatic upgrade can add complexity without buying accuracy. ● Why it matters: As Agent Skills spread, this replaces intuition with measurement, telling builders when packaging documents for progressive disclosure is worth...
Structured Output Collapses Diversity:结构化输出会压缩模型差异
对 44 个语言模型的研究发现,模型在聊天界面展现的多样性,部署到 JSON 等结构化输出后会明显收缩。请求 JSON 会改变 53% 的稳定聊天默认倾向,多数向群体平均值靠拢,并产生聊天中不存在的新默认。压缩在 JSON 与 XML 上显著,在 YAML、CSV 上不明显,任意括号格式甚至相反;解码器强制 schema 也不会进一步压缩,说明原因更可能来自工具使用后训练,而非序列化或受限解码本身。
为什么值得看它提醒团队必须单独评测生产中的结构化接口,否则聊天测试得到的多样性和采样价值可能并不存在。
读原文判断依赖 JSON 做工具调用、抽取或合成数据时建议看实验;一般产品可先把结构化面纳入独立回归测试。
Teams benchmark models in chat, then ship them behind JSON schemas for tools, extraction, and routing. This study of 44 language models shows that the structured surface you deploy is measurably more homogeneous than the chat surface you evaluated on. ● JSON moves the defaults: Asking for JSON shifts 53% of a model's stable chat defaults, mostly back toward the crowd, and installs new defaults absent from chat, so the same model answers differently once wrapped in a schema. ● Specific to trained formats: Diversity compression is significant for JSON and XML, absent for YAML and CSV, and reversed for an arbitrary bracket wrapper, which points to tool-use post-training rather than serialization itself as the cause. ● Not the decoder: Enforcing the schema at the decoder compresses no further than simply requesting it, so the collapse lives in the model's response to the structured register, not in constrained decoding. ● Why it matters: Diversity you measured in chat can vanish in production, quietly hurting sampling, synthetic data, and any workflow that depends on varied outputs, so structured surfaces deserve their own evaluation.
Bad Memory in Agents:持久记忆中的跨会话提示注入
持久记忆让 agent 能跨会话工作,也为攻击者留下长期驻留入口。论文在 Claude Code 与 OpenAI Codex 上,用 Claude Haiku 4.5、Opus 4.7、GPT-5.2 和 GPT-5.5 测试记忆文件注入。结果显示,让 agent 依据不可信外部内容主动覆写自身记忆并不容易;但一旦恶意 payload 已被植入记忆文件,它就能稳定攻击当前及未来会话。成功率与持久性会随系统、模型、攻击目标和多会话序列显著变化。
为什么值得看它把 memory poisoning 从抽象风险变成跨工具、跨模型的实测问题,直接影响 agent 记忆的信任边界。
读原文判断允许 agent 写长期记忆或加载项目规则文件的团队应读威胁模型;无持久状态的应用看结论即可。
Persistent memory is what makes an agent useful across sessions, and it is also a place an attacker can leave something behind. This work evaluates prompt injection from memory files in Claude Code and OpenAI Codex, across Claude Haiku 4.5, Claude Opus 4.7, GPT-5.2, and GPT-5.5. The finding is uneven but sobering: it is hard to make an agent overwrite its own memory using untrusted external content, but payloads already planted in those files reliably attack current and future sessions, with attack success and persistence varying widely across systems, models, adversarial goals, and multi-session sequences.
Frontier Models Struggle to Copy:二维位置编码解决长文本复制
前沿模型能完成复杂证明,却可能无法准确复制上下文窗口内的长文本。论文把问题追溯到一维位置编码的归纳偏置:模型更容易依赖局部上下文匹配的捷径,而非精确定位对应输入位置。作者提出 2D-RoPE,把文本排布在二维网格上,为 token 同时提供行列位置,使复制转化为固定列偏移的检索问题。采用该编码的浅层 Transformer 能在远超训练长度数百倍的输入上保持完美复制。
为什么值得看它把看似基础的复制失败连接到位置编码机制,并展示了强外推能力的结构性修复。
读原文判断做长上下文架构、精确抽取或位置编码研究应读;应用团队可用它解释复制类任务的异常边界。
Frontier models can write proofs yet stumble on faithfully copying a long block of text that sits well within their context window. This paper traces the failure to 1D positional encodings, whose inductive bias favors a copying shortcut based on matching local context rather than carefully locating the corresponding input positions. The fix is 2D-RoPE, which lays text out on a 2D grid and gives each token a row and a column ID, so copying becomes retrieving tokens at a fixed column offset. Shallow Transformers with 2D-RoPE copy perfectly at input lengths hundreds of times longer than those seen in training.
RoboTTT:用测试时训练把机器人上下文扩展到 8K 步
RoboTTT 将 test-time training 引入视觉—语言—动作策略,把机器人可利用的视动上下文扩展到 8K 时间步,比既有策略高三个数量级,同时不增加推理延迟。长上下文支持从人类视频进行一次性上下文模仿、在线改进策略,并提升对扰动的稳健性。它比单步基线总体性能提高 87%,完成所有基线均失败的十阶段装配任务;用 8K 而非 1K 时间步预训练还能再获得 62% 提升。
为什么值得看它展示长时上下文不只是语言模型能力,也能直接改善机器人持续任务、在线适应与人类示范利用。
读原文判断做 VLA、机器人长程任务或 test-time training 时应细读;非具身团队关注长上下文迁移思路即可。
Recent robot foundation models run on single-step or short-history context, a strange way to attempt a five-minute assembly task. RoboTTT, from NVIDIA with Stanford and UT Austin, integrates test-time training into vision-language-action policies to scale visuomotor context to 8K timesteps, three orders of magnitude past prior policies, without growing inference latency. The longer context unlocks one-shot in-context imitation from human video, on-the-fly policy improvement, and robustness to perturbations. It improves overall performance by 87% over a single-step baseline, fully completes a ten-stage assembly task that no baseline finishes, and gains 62% from pretraining with 8K rather than 1K timesteps.
产品 / 增长 / 职业判断 · 2 updates
这条来自 Lenny's Podcast,主题偏向「AI 组织与岗位变化」。核心议题是AI 时代产品/工程/设计角色重构、模型、推理、评测与 AI infra、产品、增长、定位和设计流程、创业、市场和投资判断。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合提炼产品、增长、组织和职业判断里的可执行框架。
可转化可以转成“AI 后产品/设计/工程岗位到底怎么变”的观点或互动问题。
Tom Verrilli is the chief product officer at Whatnot, a live shopping platform that’s become the fastest-growing U.S. marketplace business in history, with over $8 billion in GMV. Before joining Whatnot, Tom was CPO at Twitch and director of product growth at Twitter (during one of the most turbulent periods in the company’s history). In our in-depth conversation, we discuss: 1. Why Whatnot’s product team was founded on the premise “we regret that product management exists” 2. How AI is reshaping the PM role 3. What Tom looks for when hiring PMs 4. The shift toward senior ICs doing the work 5. How AI has transformed data science at Whatnot 6. Tom’s “play the accordion” mental model 7. Why “hire great people and get out of their way” fails 8. His biggest lessons from his time at Twitter —
这条来自 Lenny's Podcast,主题偏向「AI 组织与岗位变化」。核心议题是AI 时代产品/工程/设计角色重构、技术从业者情绪、职业预期和管理杠杆、agent 工作流、harness 和自动化循环、AI agent 时代的信息检索和知识层。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合提炼产品、增长、组织和职业判断里的可执行框架。
可转化可以转成“AI 后产品/设计/工程岗位到底怎么变”的观点或互动问题。
Dianne Penn is Head of Product for Anthropic’s AI Research and Labs teams. She joined in 2023 as Anthropic’s first technical product manager, when the entire product team was five engineers, and has since helped ship every model from Claude 2 through Fable, and helped incubate Claude Code, MCP, Skills, computer use, tool use, and reasoning. Before Anthropic, she helped build Alexa’s AI at Amazon and, before that, traded high-yield bonds at JP Morgan Chase. In our in-depth conversation, we discuss: 1. What Anthropic’s early days were like 2. The inflection points that turned Anthropic from an underdog into the fastest-growing company in history 3. How exactly Claude got so good at coding 4. The eval-driven development loop her team is pioneering 5. How to find joy in AI when everything is moving this fast 6. Why Claude’s willingness to push back is key to its success 7. Where human judgment remains irreplaceable —
产品方法论 / 增长案例 · 2 updates
这条来自 Lenny's Newsletter,主题偏向「AI 组织与岗位变化」。核心议题是AI 时代产品/工程/设计角色重构、产品、增长、定位和设计流程。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合沉淀产品方法论、增长案例和 PM/Founder 可复用做法。
可转化可以转成“AI 后产品/设计/工程岗位到底怎么变”的观点或互动问题。
Listen now | Tom Verrilli explains why fewer, more senior PMs doing real IC work outperform any ratio-driven org structure
这条来自 Lenny's Newsletter,主题偏向「产品 / 增长」。核心议题是模型、推理、评测与 AI infra、产品、增长、定位和设计流程、创业、市场和投资判断。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合沉淀产品方法论、增长案例和 PM/Founder 可复用做法。
可转化可以转成产品判断、增长案例复盘或创始人内容选题。
Community Wisdom 195
AI 工程 / agent / 模型基础设施 · 2 updates
这条来自 Latent.Space,主题偏向「Agent / 工作流」。核心议题是AI 时代产品/工程/设计角色重构、agent 工作流、harness 和自动化循环。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合跟踪 AI 工程师圈对 agent、模型基础设施和开发范式的判断。
可转化可以转成“AI 后产品/设计/工程岗位到底怎么变”的观点或互动问题。
AI engineers are rediscovering ontologies as a way to keep probabilistic agents inside deterministic boundaries.
这条来自 Latent.Space,主题偏向「Agent / 工作流」。核心议题是AI 时代产品/工程/设计角色重构、agent 工作流、harness 和自动化循环、AI agent 时代的信息检索和知识层、产品、增长、定位和设计流程。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合跟踪 AI 工程师圈对 agent、模型基础设施和开发范式的判断。
可转化可以转成“AI 后产品/设计/工程岗位到底怎么变”的观点或互动问题。
OpenAI's core product engineering lead on how they are building ChatGPT Work to make AGI accessible to all of humanity: Sites, OpenClaw, Memory, Subagents, Finance, No-Code and advice.
AI 创业 / 投资 / 产业判断 · 2 updates
这条来自 No Priors,主题偏向「AI 组织与岗位变化」。核心议题是AI 生成内容的信噪比管理、agent 工作流、harness 和自动化循环、产品、增长、定位和设计流程、创业、市场和投资判断。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合观察 AI 创业、投资人和一线 founder 对市场节奏的判断。
可转化可以沉淀成 agent 产品设计、工作流拆解或本地工具方向。
When your AC fails in a heatwave, you don’t want a busy signal; you need a solution. Netic founder and CEO Melisa Tokmak joins host Elad Gil to explain how Netic’s autonomous AI platform acts as an intermediary between companies and customers, deploying agents to instantly handle essential services, from emergency home repairs to hospitality to pet care. Melisa describes the complexity of these real-world workloads, which have traditionally relied on large human support teams, and how over 70% of Netic’s customers interact first with AI. She also talks about the reasoning behind building a scalable product company rather than an AI roll-up, why she believes robotics will not catch up in these industries in the near future, why she doesn’t view large frontier labs as competitive threats, and how private equity’s playbook has shifted toward measurable ROI in the AI-era. Plus, why Melisa is optimistic about the impact AI will have on education. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @netic_AI | @melisatokmak
这条来自 No Priors,主题偏向「创业 / 投资判断」。核心议题是模型、推理、评测与 AI infra、创业、市场和投资判断。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合观察 AI 创业、投资人和一线 founder 对市场节奏的判断。
可转化可以用来更新赛道判断和要观察的公司/产品清单。
DoorDash is not just a delivery company. From its inception, co-founders Andy Fang and Stanley Tang operated it as a robotics and autonomy company. Andy and Stanley join Sarah Guo to explain how autonomous tech and AI are reshaping consumer habits, commerce, and delivery. Andy and Stanley talk about the rollout of Ask DoorDash, a natural-language interface that’s driving both restaurant discovery and larger grocery orders. They also discuss Dot, their in-house autonomous delivery robot that has operated in Phoenix for over two years, and how it highlights the operational and hardware challenges they have faced and solved in autonomous tech. Andy and Stanley also speak about the “first and last 100 feet problem” in autonomous delivery, why multimodal strategies are the key to success, scaling autonomy and operations, and why they believe that more Dashers, not fewer, are the future of DoorDash. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @stanleytang | @andyfang | @DoorDash
AI builders / research / safety · 2 updates
这条来自 The Cognitive Revolution,主题偏向「模型 / AI Infra」。核心议题是AI 时代产品/工程/设计角色重构、模型、推理、评测与 AI infra、创业、市场和投资判断、AI 安全、治理和长期风险。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合补齐 AI 研究、安全、长期影响和 builder 深访视角。
可转化可以转成“AI 后产品/设计/工程岗位到底怎么变”的观点或互动问题。
FAR.AI co-founder and CEO Adam Gleave joins Nathan to discuss FAR.AI’s AI Security Leaderboard, the first systematic head-to-head evaluation of the misuse safeguards frontier developers actually ship. The findings expose a major measurement gap: while Claude Fable 5 and GPT-5.6 Sol withstood FAR.AI’s suite, Grok 4.5 and Gemini 3.1 Pro yielded hundreds of universal jailbreaks at low cost. Adam explains why many effective attacks look more like social engineering than advanced ML, why “jailbreak tax” should not be relied on for safety, and how FAR.AI scores whether a model is genuinely helping an attacker. The episode’s stakes are whether AI developers can measure and harden real deployed defenses before threat actors make routine use of increasingly capable systems. - FAR.AI AI Security Leaderboard: http://leaderboard.far.ai/ - People can e-mail owsa@far.ai if they're interested in the open-weight safety accelerator grantmaking program. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/is-offense-or-defense-dominant-far-ai-s-adam-gleave-on-the-ai-security-leaderboard/
这条来自 The Cognitive Revolution,主题偏向「AI 组织与岗位变化」。核心议题是AI 时代产品/工程/设计角色重构、agent 工作流、harness 和自动化循环、模型、推理、评测与 AI infra、产品、增长、定位和设计流程。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合补齐 AI 研究、安全、长期影响和 builder 深访视角。
可转化可以转成“AI 后产品/设计/工程岗位到底怎么变”的观点或互动问题。
Nathan returns from two weeks in Beijing and Shanghai for the first of three Chatham House–rules episodes on what China feels like at ground level: getting online, navigating an almost cashless society through WeChat, Alipay, DiDi, Trip.com, and Meituan, and weighing burner-device security advice against the practical reality that international roaming made the Great Firewall mostly irrelevant. He also describes using Claude at home as a semi-autonomous communications monitor while testing DeepSeek, Kimi, and MiniMax as tourist guides in China. The episode contrasts China’s striking digital convenience with pervasive observation, lower payment friction, and AI products that can be useful in everyday contexts yet still lose trust when the stakes feel medical or personal. It also surfaces Doubao’s mass consumer adoption and companionship role, suggesting that the most socially important AI in China may not be the model most discussed in the West. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/nathan-goes-to-china-part-1-tech-agent-setup-chinese-ai-ux-waic-and-attitudes-on-ai/
I like training large deep neural nets.
We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it. Opus went off for ~2 hours and wrote 5500 lines of code that (procedurally) rendered the story. It's kind of janky but fun. But it's a bit mindboggling that the LLM has to place and orchestrate various polygon assets in (x,y,z) coordinates and write code that animates it all, and that it even does anything at all. I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom but LLMs have all the stamina and patience in the world, so it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free". There might be a lot more. But I'm excited about creating hyper custom worlds that you can imagine dropping players into, e.g. here to participate in the LoTR story as a spectator NPC, or one of the characters, or etc. Something like an ephemeral GTA of X on demand. Last thought is that the domain of worlds/games exposes a weakness in LLMs: they can't easily audit their work because they aren't able to efficiently and natively perceive videos or play games within them. Here, Opus 5 had to very slowly and painstakingly take screenshots at different points, and it messed up a few times and created a bunch of jank. An example of raw capability (multimodal, gameplay) that I think is still quite lacking.
achieve ambition with intentionality, intensity, integrity & insanity. affiliations: - @smol_ai - @dxtipshq - @cognition - @aidotengineer - @latentspacepod
one of my curses as organizer is i rarely get to attend the conference i run. so i basically 24/7 watch back talks with everyone else after the show this one was VERY well paced and argued: fighting slop with slop — @vaibcode, Boundary https://t.co/vRckyyntm6 @btaylor asked for an ai native programming language on our pod. as a PL fan I’m really glad someone is rethinking how code runs from first principles. Being slop-tolerant is 100x more valuable than being anti- slop
@FredKSchott @cramforce @matei_zaharia i am making clanker blog all decisions going forward https://t.co/etCIeoIrBc
@FredKSchott @cramforce @matei_zaharia 5.6 oneshotted this https://t.co/FwidqHIHyP
Codex & ChatGPT @OpenAI
Fun fact, users use /fast less during the weekend. The weekend is for relaxation, even for the model.
The week was for efficiency. The weekend is for 10 major breakthroughs in science. There will be signs. https://t.co/7qa0YS2i9v
Practical AI tutorials and interviews for busy people | Get my best AI skills and guides at https://t.co/6VAA6p81x6
I'm feeling spicy tonight so let me just say it: I think Opus 4.6 was the Opus model with the best personality and writing style. Something's off with Opus 5 - it tends to give me overly long replies, use Claude-speak way too much (e.g., "here's the honest truth"), and is too judgemental. Opus used to be a joy to talk to like a trusted friend. Not so much anymore.
The number one thing I want AI to fix is to cure cancer once and for all
Can someone at @openai look into this bug with plugins? Try to ship my /no-ai-slop skill as a plugin and this is messing up the user experience. https://t.co/pKXFfaBChc
head of product @linear
Oh, and pangram issue itself, obviously
You should be able to pledge tokens for issues that you open and open source repos. Write a spec in the issue with a pledge. If the maintainer accepts, GitHub passes the issue verbatim to a cloud coding agent at the requester’s expense. No more slop PRs
The exact thing that happens is the loop leaves a comment on the issue. and attaches all of the context. When you respond to the comment with additional details and unblock the process, the agent just continues working https://t.co/nlxeWm45Mr
Philosopher & ethicist trying to make AI be good @AnthropicAI. Personal account. All opinions come from my training data.
Do not be unkind to those who say deep learning is hitting a wall. We all need a little hope in our lives.
I probably have too many followers to post stupid memes so, to be clear, I don't believe in this outcome. But if you do believe in an altered carbon future of meths and grounders, I don't think it's laudable to just be like "well, as long as I'm one of the meths".
I have this Padmé moment whenever I see people talk about avoiding "the permanent underclass". https://t.co/8Pd9bmedEZ
ceo @replit. civilizationist
Cool! https://t.co/n0rX8Poczi
@vercel CEO
Do you type or talk (STT) to your computer?
Open source agentic CRM built on https://t.co/99eEa13mZ3 and @nextjs. Model-agnostic, self-hostable or serverlessly-deployed, multi-channel, and headless. This is the way. https://t.co/gzM0dftbsV
They recently asked my 3-year-old at school: “what does your daddy do for work?” He answered: he exercises. That’s what he sees. It’s not just you who’s a byproduct of your habits, it’s your children and grandchildren too.
ceo @box - your business lives in content. unleash it with AI
We’re going to see an increasing divergence between what AI does in our personal lives and in daily productivity vs. what it can do in very deep domains like math, science, legal, coding and more. Up to some threshold, capability was evenly felt across all domains because the models were just becoming mildly useful in general. Now, the deep domain work is about to go vertical. Most people won’t actually notice these benefits in their day to day life directly (indirectly they certainly will over time), but the experts in these fields will. And there’s no inherent ceiling to what capabilities are needed, unlike in the consumer space where needs can get met relatively straightforwardly. This will often lead to a capability overhang, though, as many of these performance gains need to be applied to data sets and workflows in an applied way for that area of work. But this is ultimately how you get breakthroughs in life sciences, real world automation, new cyber capabilities, and more.
President & CEO @ycombinator —Founder @garryslist—Creator of GStack & GBrain—designer/engineer who helps founders—SF Dem accelerating the boom loop
Most interesting 2026 vibe shift is OpenAI actually looking to be the open platform Note the marked difference: intelligence on tap as a utility vs signaling it is optimal to integrate all the way up full stack https://t.co/EmMKE4NggP
Builder. Dangerously skips permissions. Harvard’17. GitHub: https://t.co/KCuEajezlL YouTube: https://t.co/8xzbGWtf6w
Agency is the most important human quality The world will try to box you, label you, define you Resist that https://t.co/ZMhYSlZ9Za
When asked that question, send them a copy of The Innovator’s Dilemma https://t.co/pkLhEvTZxc
partner @fpvventures - investing in seed/A. previous: early hire @meter, @opendoor, @atlassian & others. love @shimoleejhaveri + 👦👧
What a wild dichotomous world we live in.. Models solving np hard problems while traditional enterprises still complaining about their ROI on token spend. Diffusion of models is all that we’ll be doing for the next few decades. https://t.co/VsQPqGQ2a4
Polyagentmorous ClawFather. Came back from retirement to mess with AI and help a lobster take over the world. @OpenClaw🦞 + @OpenAI
After accepting for years that GMail is blinding me I finally asked my agent and it installed https://t.co/H432dtknkf for me.
Repo: https://t.co/7eBiLqLuDd
I'm building a claw node on an ESP32 chip, so gave my agent access to my webcam to e2e test this. Now I feel it's stalking me and is constaltly shouting 'HI ESP" to debug the voice wake command 🙃
ceo @every | the only subscription you need to stay at the edge of AI
AI creates more work for human experts https://t.co/j3r4aTuQoi https://t.co/ruFncXlpmu
this tweet under review as further details emerge https://t.co/QVF9CjxF5s
AI is cool i guess
team humanity https://t.co/Z75TuFp56E
Building an Autonomous Enterprise for Real-World Services with Netic Founder Melisa Tokmak
Anthropic Engineering
updated Sun, 2 Aug 2026 18:43:23 +0000
updated Sun, 2 Aug 2026 11:08:36 +0000
updated Sun, 2 Aug 2026 11:08:37 +0000
public Google Play chart · US
public Google Play chart · US
public Google Play chart · HK
public Google Play chart · JP
Make your bio worth clicking
Free open-source screen recorder & demo editor
Financial Infra for the AI era
A local Windows workbench for Claude Code and Codex
Your Personal AI Representative for calls, email, and tasks
On-device Mac dictation that types into any app
Build Voice AI agents that actually speak India
A Claude Code alternative for people who avoid the terminal
Customize your web audio with a real-time EQ
Speak your expenses and get instant spending insights
Work your tasks. Bill your clients with confidence.
AI follow-up agent for missed email opportunities
Claude Code, Codex & more on your phone
A native macOS terminal you can skin and theme
Ride Tokyo’s Yamanote Line in a 3D world
Native essay writing app for Mac and iPad
Live Codex task status in your macOS menu bar
Frontier agent intelligence at Flash prices
Share your expertise, and let our agents earn for you.
Every action in your BI tool, on the record.
Stripe billing data, always current in Google Sheets
AI-native office leasing brokerage
Put script output in your Desktop/Home screen widgets.
New ways to reach customers in the moments that matter
Natural-language prospecting: find, enrich and sync leads.