产品发布/更新
1 updates
-
Suno 推出多项新功能,含MIDI导出等
我们一直在以比以往更快的速度构建!🚀 以下是网页端和移动端的新功能一览: • 高级音轨分离 • 将音轨导出为 MIDI • 歌词合写与自动保存 • 截图生成歌曲 • Apple CarPlay 与 Android Auto 你最期待 Suno 的哪些新功能?
Open source
2026-07-27 · AI HOT + Model Companies + Market News + AI Papers + Voices + Trends + Follow Builders · generated 2026/07/27 17:03 · builder feed 2026/07/27 15:26
1 updates
我们一直在以比以往更快的速度构建!🚀 以下是网页端和移动端的新功能一览: • 高级音轨分离 • 将音轨导出为 MIDI • 歌词合写与自动保存 • 截图生成歌曲 • Apple CarPlay 与 Android Auto 你最期待 Suno 的哪些新功能?
Open source3 updates
Berkeley RDI等机构提出三级软件自主开发框架:代码自主(AI完成设计与实现,人类决定构建内容并审核PR)、流水线自主(AI运行从设计到部署的全流程,人类仅评估结果)、需求自主(AI自主决定构建内容)。该框架旨在为能力声明、部署选择与问责提供清晰分类。
Open sourceVista 基于 bento PPT 改造了一个 Skill,输入内容或主题即可自动生成可编辑、在线演示并支持协作的 HTML PPT。安装指令为 `npx skills add joeseesun/qiaomu-bento-ppt`,推荐使用 Kimi K3 或 Opus 4.8+ 等前端审美好的模型。
Open source作者开源了Leader.skill,用于将模糊的人类需求转化为Agent可独立执行数小时的目标任务书。该Skill基于"目标七问"方法论,涵盖目的、完成态、反作弊、边界等维度,并推荐用Claude Fable 5或Kimi K3规划目标,再交由GPT-5.6 Sol或GLM-5.2等模型长程执行。项目已开源。
Open source1 updates
OpenAI 与 Anthropic 正游说美国监管机构限制中国开源 AI 模型,认为开放开发过于危险。英伟达 CEO 黄仁勋、微软 CEO 纳德拉、马斯克及扎克伯格等人公开支持开源,签署联名信反对限制。近 200 家硅谷创业公司也敦促特朗普政府不要限制获取中国开源模型,美国官员倾向于将此事作为国家安全问题单独处理。
Open source海外 · 3 updates
主要信号集中在模型进展、安全 / 监管:OpenAI's GPT-5.6 Rollout Is Changing How ChatGPT Users Pay and Access Its Best AI Models;OpenAI's GPT-5 was labelled a high-risk AI model over biohazard concerns
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
OpenAI's GPT-5.6 launch shifts ChatGPT from model-based branding to capability tiers, making subscription plans the primary way users access advanced AI.
Open newsOpenAI employees continued testing the model, but the findings showed problematic responses, indicating that some safety issues may have persisted despite safeguards.
Open newsA ChatGPT misdiagnosis lawsuit claims AI advice nearly killed a man. A physician describes what peer-reviewed research shows about AI diagnostic accuracy.
Open news海外 · 1 updates
主要信号集中在模型进展:Anthropic launches Claude Opus 5 AI model: Higher coding performance, lower cost, and new API features
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Anthropic has introduced Claude Opus 5 with improved coding, reasoning, and research capabilities while keeping the same API pricing as its predecessor.
Open news海外 · 1 updates
主要信号集中在Agent / 工作流、模型进展、安全 / 监管:Google Expands Gemini With Cheaper Models And A Wider Agent Push
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Google's Gemini 3.6 Flash, Flash-Lite and restricted Flash Cyber launches plus Spark's move to the $19.99 AI Pro tier signal a push on price, reach and security AI.
Open news海外 · 1 updates
主要信号集中在商业化 / 资本:Meta launches seller app: Photo-to-listing AI hits Facebook Marketplace
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Facebook Marketplace Seller app, launched July 24, 2026, is a free standalone iOS tool for US power sellers that uses Meta AI to generate listing titles, descriptions, prices, and categories from a ...
Open news海外 · 2 updates
主要信号集中在Agent / 工作流、模型进展:Microsoft's Copilot starts replacing OpenAI and Anthropic models with in-house AI;Microsoft Kimi K3 Copilot Test Could Save $600M in 2026
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Microsoft is transitioning to its in-house AI models, moving away from reliance on OpenAI and Anthropic solutions. CEO Satya Nadella highlighted this shift, noting that applications like Outlook and ...
Open newsMicrosoft Kimi K3 Copilot testing could cut AI costs by $600M in 2026. See how model routing may weaken OpenAI’s default role inside Microsoft.
Open news海外 · 1 updates
主要信号集中在算力 / 推理:Nvidia’s Rivals Are Coming for Its Crown, But the Smartest AI Bet Sits Further Down the Tech Stack
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Hyperscalers are building custom AI chips to cut their reliance on Nvidia. But whoever wins, every chip still runs through the same handful of companies, from Broadcom and TSMC down to ASML.
Open news海外 · 1 updates
主要信号集中在模型进展、算力 / 推理:Amazon cuts jobs in AGI unit as its AI strategy shifts again
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Amazon has cut jobs in its AGI unit while keeping advanced AI, Nova models and custom chips central to its strategy.
Open news海外 · 1 updates
主要信号集中在模型进展、商业化 / 资本:Apple is rebuilding Siri on Google’s Gemini in a billion-dollar-a-year deal
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Court filings in the Justice Department’s antitrust case against Google reveal that Apple and Google negotiated terms to integrate Google’s Gemini AI model into Siri and the broader Apple Intelligence ...
Open news海外 · 1 updates
主要信号集中在模型进展:Grok’s Competitor to Best Chinese Models is Now Available for Everyone
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
AI has expanded access to its flagship Grok 4.5 artificial intelligence model across grok.com, X, and its dedicated iOS and Android applications. The ...
Open news海外 · 1 updates
主要信号集中在Agent / 工作流:Cohere VP says enterprise AI sovereignty requires control of the full agent stack
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Asked what would persuade enterprises to move beyond bundled AI services from existing cloud providers, Alao returned to data control and portability.
Open news国内 · 1 updates
主要信号集中在算力 / 推理、商业化 / 资本:Chip stocks slide in AI unwind
为什么值得看适合用来观察国内模型厂商的产品节奏、开源/闭源路线和商业化落点。
A selloff in chipmakers gathered pace last week, driving the high-profile group of stocks to a bear market on worries that the artificial intelligence (AI) spending spree is becoming harder to justify ...
Open news国内 · 1 updates
主要信号集中在模型进展:What China's internet is saying about Moonshot's hot new Kimi model
为什么值得看适合用来观察国内模型厂商的产品节奏、开源/闭源路线和商业化落点。
Just days after launching Kimi K3, Moonshot AI temporarily suspended new subscriptions, saying demand had overwhelmed its computing capacity.
Open news国内 · 1 updates
主要信号集中在模型进展:Chinese AI model rivals ChatGPT, jolting Silicon Valley
为什么值得看适合用来观察国内模型厂商的产品节奏、开源/闭源路线和商业化落点。
Another powerful new artificial intelligence model from China took the U.S. tech industry by surprise Friday, the latest sign that Chinese startups that publicly release their “open-source” AI ...
Open news这里是市场情绪线索,用来辅助判断风险偏好、科技股和 AI 资产预期,不当作投资建议。
Fed、通胀、债券收益率 · 2 updates
这条偏「利率/通胀、股市情绪」信号,当前解读为分歧信号/需要二次确认。可能影响成长股折现率、美元和长端利率预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:Rising oil prices and elevated bond yields, driven by Middle East tensions, are raising concerns over the sustainability of the U.S. stock market rally. While strong earnings and AI-driven spending ...
Open news这条偏「利率/通胀」信号,当前解读为中性但值得观察。可能影响成长股折现率、美元和长端利率预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:Many types of consumer loans such as mortgages peg their interest rate to the yield on 10-year Treasury bonds, which has been moving higher.
Open newsS&P 500、Nasdaq、波动率、资金情绪 · 1 updates
这条偏「利率/通胀、股市情绪」信号,当前解读为中性但值得观察。可能影响成长股折现率、美元和长端利率预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:The market broadly believes the US Federal Reserve will keep its interest rates unchanged, with an 80% probability of a hike in September policy. Elevated crude oil prices, surging energy costs, and a ...
Open newsNVIDIA、AMD、TSMC、AI chips · 1 updates
这条偏「利率/通胀、股市情绪」信号,当前解读为中性但值得观察。可能影响成长股折现率、美元和长端利率预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:Taiwan Semiconductor Manufacturing is rated Buy: resilient chip leader insulated from AI price wars. See more analysis on TSM stock here.
Open news云厂商、AI 资本开支、利润率 · 4 updates
这条偏「财报/AI Capex」信号,当前解读为偏利多/风险偏好改善。可能影响云厂商利润率、AI 资本开支和上游算力需求;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:Microsoft, Apple, Amazon and SK Hynix report this week as investors test Big Tech's AI spending against real returns.
Open news这条偏「股市情绪、财报/AI Capex」信号,当前解读为中性但值得观察。可能影响云厂商利润率、AI 资本开支和上游算力需求;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:The market reaction to Alphabet’s Q2 results has significantly raised the bar for its Magnificent Seven peers that are on deck to report results this week, namely Microsoft and Meta Platforms on ...
Open news这条偏「利率/通胀、股市情绪」信号,当前解读为偏利多/风险偏好改善。可能影响成长股折现率、美元和长端利率预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:The QQQ's top holdings—that is, big tech stocks—are in an AI-driven arms race. Click here to read more on QQQ and why I rate it as a Hold.
Open news这条偏「财报/AI Capex」信号,当前解读为偏利空/风险偏好收缩。可能影响云厂商利润率、AI 资本开支和上游算力需求;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:Alphabet and Tesla shares tumbled after earnings highlighted surging AI-related capital spending, weakening cash flows. Investors now await Microsoft, Meta and Amazon earnings as Big Tech's AI ...
Open newsHugging Face Daily Papers + arXiv recent AI/ML · 12 papers · fallback summaries
来自 Hugging Face Daily Papers,主题偏「Agent、Reasoning、Post-training/Alignment」。摘要显示它主要讨论 Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry sug... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do more than simply search longer: it should recursively improve its current answer by verifying intermediate results and using the partially verified state to guide subsequent refinement. We introduce AREX, a family of Recursively Self-Improving (RSI) deep research agents. AREX alternates between an inner research loop that gathers evidence and constructs a provisional answer, and an outer self-improvement loop that audits the answer constraint-wise, identifies unresolved claims, and launches targeted follow-up research. To sustain RSI over long horizons, AREX learns an autonomous context-update tool that compresses growing interaction history into a compact improvement state preserving verified evidence and unresolved constraints, without relying on an external model. We train AREX on verified synthetic tasks and high-quality trajectories through agentic mid-training and long-horizon reinforcement learning. To mitigate sparse final rewards during long horizon learning, we emphasize key steps where decisive evidence is acquired or erroneous research directions are corrected. We instantia...
来自 Hugging Face Daily Papers,主题偏「RAG/Memory、Reasoning、Multimodal」。摘要显示它主要讨论 Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this capability curriculum cognition. It... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察知识工作、企业搜索、长期记忆和本地资料库产品的新实现路径。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this capability curriculum cognition. It covers prerequisite chains, concept taxonomies, experiment-concept links, pedagogical sequencing, and visual grounding. We introduce K12-KGraph, a curriculum-aligned knowledge graph extracted from official People's Education Press textbooks in mathematics, physics, chemistry, and biology across primary, middle, and high school. It contains nine node types and fourteen relation types covering curriculum structure and visual grounding. From this graph, we derive K12-Bench, a 23,640-question multi-select benchmark with five task families: Ground, Prereq, Neighbor, Evidence, and Locate. We also build K12-Train, a graph-guided supervised fine-tuning corpus of 7,335 samples, including 2,267 text-only QA pairs and 5,068 multimodal VQA pairs. On K12-Bench, Gemini-3-Flash achieves only 57 percent exact match and Gemma-4-31B-IT reaches 46 percent, with Prereq and Neighbor being the hardest tasks. Our training experiments show that domain-specific supervision can reduce this gap. Under a matched 2,300-sample budget, K12-Train-Text consistently outperforms equally sized subsets of eight mainstream instruction-tuning corpora on Gaokao...
来自 Hugging Face Daily Papers,主题偏「Agent、Reasoning、Multimodal」。摘要显示它主要讨论 Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, the... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their chain-of-thought (CoT) reasoning often operates in abstract spatial latents that are difficult to supervise and weakly aligned with explicit image-space detections. To address this, we introduce ReferTrack, a referring-then-tracking paradigm that grounds EVT using a single forward-facing camera. Our model first selects the target from an indexed set of bounding boxes, then decodes tracking waypoints conditioned on this image-grounded decision. To preserve target motion cues over time, ReferTrack maintains a sliding-window queue of previously selected bounding boxes, injecting their geometric features into the visual history via temporal-viewpoint-bbox indicator (TVBI) tokens. We further enhance target identification by co-training on a custom Refer-QA dataset. On EVT-Bench, ReferTrack achieves state-of-the-art single-view performance with success rates of 89.4%, 73.3%, and 74.1% on the single-target, distracted, and ambiguity tracking splits, respectively -- matching or even surpassing several multi-camera baselines on identification-heavy tasks. Finally, real-world deployments on legged and humanoid robots validate its...
来自 Hugging Face Daily Papers,主题偏「Agent、Reasoning、Multimodal」。摘要显示它主要讨论 Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-generation models. This makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments directly in pixel space. We propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics. ProVisE also includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks. We further introduce SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms. We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks. Resul...
来自 Hugging Face Daily Papers,主题偏「Reasoning、Multimodal、Eval/Data」。摘要显示它主要讨论 On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a stronger learning sign... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察模型推理、规划、验证器和复杂任务能力是否有可复用技术路线。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a stronger learning signal than the student. Existing methods create this asymmetry either through privileged answers or visual evidence. We ask whether both can be removed, yielding a simpler form of OPSD driven purely by input conditioning. For this purpose, we propose Visual Contrastive Self-Distillation, namely VCSD, which converts image-content removal into an on-policy self-distillation signal. At each student-generated response prefix, the EMA teacher produces two next-token distributions under the same prompt and prefix -- one conditioned on the original image and the other on a content-erased control. Their token-wise log-probability difference highlights candidates whose likelihood is specifically increased by the instance-level visual content. We use this contrast to sharpen the teacher's original-image distribution within its plausible support, and distill the resulting full-distribution target into the student. Using ViRL39K dataset, VCSD consistently outperforms matched OPSD across Qwen3-VL and Qwen3.5 models. For example, on Qwen3-VL, it improves the seven-benchmark aggregate from 62.27% rightarrow 67.04% at 2B, 71.30% rightarrow 7...
来自 Hugging Face Daily Papers,主题偏「Agent、RAG/Memory、Reasoning」。摘要显示它主要讨论 Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework for building reliable AI agents. NOOA takes a simpler approach:... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework for building reliable AI agents. NOOA takes a simpler approach: an agent is a Python object. Its methods are the actions the model can take, fields are its state, docstrings are its prompts, and its type annotations are contracts. A method whose code body consists of "..." is completed at runtime by an LLM-driven agent loop, while methods with normal bodies remain standard deterministic Python. This gives developers and agents the same interface, so agent behavior can be tested, traced, refactored, and improved just like other software. This paper makes three contributions. (1) We present the agent-as-a-Python-object programming model and the design principles behind it. Where Python has existing abstractions, we adopt them directly. Agent-specific capabilities--context, events, state rendering, long-term memory, and validated LLM loops--are exposed through simple Pythonic APIs, so both developers and agents share one familiar programming model. (2) We identify six model-facing ideas that NOOA is, to our knowledge, the first to combine on a single surface: typed input/output, pass-by-reference over live objects, code as action, programmable loop engineering, explicit object state, and...
来自 Hugging Face Daily Papers,主题偏「Agent、RAG/Memory、Multimodal」。摘要显示它主要讨论 We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and run... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running distribution-informed coding-agent tasks across four work domains - Code, Web, Office, and Security. Rather than adapting public issue text, every task is reverse-engineered from a real commit, pull request, or business scenario and rewritten as a short, colloquial, role-played request, so that a task's prompt is not recoverable by web-searching the underlying issue, pull request, or commit thread. Because the dataset is released openly - task directories, environment images, evaluation harness, tests, and reference solutions - contamination resistance rests on this construction together with dataset versioning rather than on secrecy. The four subsets - repository-level engineering, front-end development, office and business workflows, and red-/blue-team security - probe complementary facets of real work, each with its own verification style. All are packaged in a uniform task-directory format and run, under a uniform and reproducible protocol, on two agent harnesses (CodeBuddy Code and Claude Code); the full open release makes the benchmark reproducible end to end and directly auditable, since any third party can re-...
来自 Hugging Face Daily Papers,主题偏「Agent、Eval/Data、AI Infra」。摘要显示它主要讨论 As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users rarely specify their intent upfront, instead disclosing, rev... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users rarely specify their intent upfront, instead disclosing, revising, and reshaping it as the conversation unfolds. Despite this, LLMs are still predominantly evaluated or trained in single-turn, fully-specified settings, leaving open a fundamental question: how well do LLMs track and act on user intent as it evolves over the course of a conversation? To study this, we introduce a framework that transforms static, single-turn tasks into dynamic multi-turn conversations in which the user's intent evolves across turns--incrementally revealed, revised, and at times redirected mid-conversation--while preserving each task's original evaluation protocol, enabling existing benchmarks to be reused as controlled testbeds without new annotation. Across multiple tasks, we surface a consistent phenomenon: strong static-setting performance does not transfer to the evolving-intent setting, with substantial drops across model families. Our findings point to a fundamental gap: today's LLMs do not yet faithfully track and act on the user's evolving intent, a capability invisible to static evaluation yet critical for future collaborative agents.
来自 Hugging Face Daily Papers,主题偏「Agent、Reasoning、AI Infra」。摘要显示它主要讨论 We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teacher over these multi-turn interaction histories. Fully online OPD is costly because each update requires... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teacher over these multi-turn interaction histories. Fully online OPD is costly because each update requires fresh student rollouts through the environment and teacher queries at visited histories. We propose Replayed-Prefix On-Policy Distillation (ReOPD), an off-environment alternative that reuses pre-collected teacher trajectories as replayed prefixes: the student acts at selected steps, while the teacher provides dense per-step supervision without executing new environment interactions. We show that multi-turn OPD introduces a prefix trap: making histories more student-on-policy improves relevance to the student, but can query the teacher on histories where its target is unreliable. This creates a two-sided distribution shift between student occupancy and teacher reliability. ReOPD addresses this by treating multi-turn OPD as a reliability-aware prefix distribution design and implements it with a simple step-decaying sampling schedule that emphasizes early, lower-shift prefixes. Across mathematical reasoning with Python and search environments over multiple teacher and student model scales, ReOPD preserves or improves OPD-level accuracy, uses zero tool calls during student training, and is at least 4times faster per rollout th...
来自 Hugging Face Daily Papers,主题偏「Agent、RAG/Memory、Reasoning」。摘要显示它主要讨论 Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling complex real-world problems, different agents still show signif... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling complex real-world problems, different agents still show significant performance variation. In this work, we design Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge. Using an agent workflow built on this skill, we generate 2,000 professional financial documents along with 6,000 high-quality question-answer pairs. To evaluate the overall capability of agents, we introduce FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios. It contains 2,026 deep research tasks targeting 1009 financial documents. FinanceComplexQA has 8 key features: bilingual support; coverage of six mainstream scenarios and seven tasks; expert-level document reasoning questions; deep research of complex layouts; relatively stable and permanent reference answers; and precise evaluation through an Agent-as-a-Judge with multiple evaluation metrics. Using FinanceComplexQA, we conduct a comprehensive evaluation of leading RAG systems and agentic reasoning tools for financial document QA. Through identifying and analyzing failure cases, we provide an in-depth study of their capabi...
来自 Hugging Face Daily Papers,主题偏「Agent、RAG/Memory、Reasoning」。摘要显示它主要讨论 Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express stateful, multi-process harness inference. To address this, we present OpenForgeRL, an open-source framework for training harness-based agents end-to-end in diverse environments. OpenForgeRL achieves this with a lightweight proxy that serves the harness's model calls while recording them as training data for a standard RL codebase (e.g., veRL), and a Kubernetes orchestrator that runs each rollout in its own remote container, together enabling training on any harness in any environment at scale. By decoupling training and inference, OpenForgeRL allows researchers to easily train, study, and improve agents directly in the real harnesses and environments they are deployed with. We validate our framework across diverse, complex harnesses and environments, spanning tool/claw-based agents and multimodal GUI browser- and computer-use agents. Using only hundreds to a few thousand tasks, OpenForgeClaw reaches 31.7 pass^3 and 55.9 pass@3 on ClawEval and 33.7 on QwenClawBench. OpenForgeGUI reaches 37.7 on OSWorld-Verified, 63.0 on Online-Mind2Web, and 72.3 on WebVoyager. Bo...
来自 Hugging Face Daily Papers,主题偏「RAG/Memory、Multimodal、Post-training/Alignment」。摘要显示它主要讨论 We revisit dataset distillation from an outcome-centric perspective. Rather than aligning process surrogates (per-step gradients or training trajectories), Influence Matching (Inf-Match) aligns the final outcome of training: it learns a compact synthetic set w... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察知识工作、企业搜索、长期记忆和本地资料库产品的新实现路径。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
We revisit dataset distillation from an outcome-centric perspective. Rather than aligning process surrogates (per-step gradients or training trajectories), Influence Matching (Inf-Match) aligns the final outcome of training: it learns a compact synthetic set whose effect on the converged parameters matches that of the full dataset. Concretely, we introduce a fully differentiable, sample-level influence estimator that quantifies parameter shifts from adding or removing data, without time-consuming inverse-Hessian products or convexity assumptions. The estimator runs in linear time by unrolling the optimization dynamics and applying a first-order Taylor approximation. We then learn the synthetic set by minimizing the mismatch between its influence and that of the real dataset, yielding outcome alignment rather than heuristic process imitation. Inf-Match delivers the best accuracy across standard classification benchmarks. For instance, on Tiny-ImageNet (IPC=10), Inf-Match attains 31.5\%, a +4.7\% improvement over NCFM. Beyond classification, Inf-Match scales to vision-language distillation on Flickr30K, outperforming strong process-matching baselines. For instance, with 200 to 1000 synthetic samples, our method achieved a leading impressive average on image/text retrieval tasks, higher than NCFM by 2.5\%. The code will be released via https://github.com/hrtan/infmatch.
Using cached daily papers because the latest fetch failed.
dair-ai/AI-Papers-of-the-Week · 10 papers
Self-Improving Agents Survey:自我改进 Agent 的统一地图
这篇综述把现代 agent 定义为基础模型加上提示词、记忆、工具与控制逻辑组成的脚手架,并把自我改进统一为由系统自身触发、最终写回模型权重或脚手架的更新。作者进一步按学习信号来源区分生成式示范、评价式反馈和真实或模拟环境中的探索经验,也覆盖自动改写代码、生成—测试—修补和开放式 agent 设计搜索。它把软件、网页、游戏、科学与机器人等分散研究放进同一套坐标系,便于比较不同方法究竟改了什么、凭什么改进,以及如何控制持续演化的风险。
为什么值得看它为自我进化 agent 提供了可复用的术语与分类,适合做 agent 架构、评测和长期学习路线判断。
读原文判断正在设计会从运行经验中更新的 agent,建议读分类与开放问题;只做固定 workflow,掌握权重与脚手架两类更新即可。
Self-improving agents are moving from research demos into deployed systems, and this survey gives the trend a clean formalism. It frames a modern agent as a foundation model coupled with an operational scaffold of prompts, memory, tools, and control logic, then treats self-improvement as a self-induced update that commits changes to either the weights or the scaffold. ● Two update targets: Improvement splits into foundation-model updates to the weights and scaffolding updates to prompts, tools, memory, and control code, giving a shared vocabulary for work that usually looks unrelated. ● Signals that drive change: The survey organizes methods by where the learning signal comes from, spanning intrinsic generative demonstrations, intrinsic evaluative feedback, and extrinsic exploratory experience in real or simulated environments. ● Full-scaffolding frontier: The most open-ended methods rewrite the agent itself through self-referential code updates, generate-test-patch loops, and open-ended search over agent designs, pushing toward controllable evolution with little human input. ● Why it matters: As teams wire agents to improve from their own experience, a single map of update targets...
Metacognition in LLMs:统一理解模型的自我监控与调节
这篇 Yale 与 UC Irvine 的综述认为,置信度校准、自我验证、何时停止以及识别知识边界并非孤立技巧,而是同一种元认知能力。论文用“监控—控制”循环组织领域:模型在行动前后自我评估,再决定回答、重试或延迟;测量方法横跨信号检测理论、校准与 AUROC/ECE、激活层反馈和可解释性探针;能力注入则覆盖框架、架构、提示与训练。统一视角有助于把减少幻觉、判断未知和抵抗错误说服等可靠性问题转化为可测量、可改进的系统目标。
为什么值得看它把零散的 confidence 技巧提升为完整能力框架,对可靠问答、推理模型与自主 agent 都有直接参考价值。
读原文判断做模型可靠性、拒答或自检机制时值得深入读测量章节;产品侧只需先采用监控后控制的设计框架。
Confidence calibration, self-verification, knowing when to stop, and knowing what you do not know have mostly been studied in isolation. This survey from Yale and UC Irvine argues they are facets of one capability, metacognition, and organizes the field around a monitor and control loop wrapped around the language model. ● Monitor and control framing: The model self-assesses before and after acting, then self-regulates by deciding whether to answer, retry, or defer, turning scattered behaviors into a single monitor-then-control cycle. ● How it is measured: The survey catalogs psychology-based methods from signal detection theory, confidence-based metrics like calibration, AUROC, and ECE, activation-level neurofeedback, and interpretability probes such as concept injection. ● How it is instilled and used: It reviews frameworks, architectures, prompting, and training that give LLMs, reasoning models, and agents metacognition, then shows gains in hallucination reduction, knowledge-boundary detection, and resistance to persuasion. ● Why it matters: Metacognition underpins reliability, so a unified account of how to elicit, measure, and improve it gives builders a coherent target rather...
When Is Routing Meaningful:何时模型路由才真正有意义
论文指出,LLM router 即使准确率和成本表现优秀,也可能并未做出有意义的分配。真正的路由需要两个正交条件:候选模型之间存在可观察的行为差异,而且同一问题换一种表述后,分配结果仍保持稳定。作者用 Hierarchic Social Entropy 衡量模型群体的结构性多样性,发现精心构造的小型专家群体可能比规模更大的通用模型池更有区分度;同时,KNN router 在改写扰动下会崩溃,而提示式 router 能保持更好的准确性与稳健性。结果说明干净查询上的准确率不足以证明路由有效。
为什么值得看它给模型路由和 mixture-of-agents 增加了多样性与改写稳定性两项必要验收指标,可避免伪优化。
读原文判断在做多模型路由、成本优化或专家池设计时建议读实验;单模型应用只需记住准确率不能证明路由有效。
LLM routers and mixture-of-agents systems get judged on accuracy and cost, both of which can look great while the router is doing nothing. This DeepMind-affiliated work argues that whether routing means anything depends on two properties that are orthogonal to accuracy. ● Two conditions for real routing: The society of models must be behaviorally differentiated, since routing is vacuous when every actor responds the same way, and assignments must stay stable when a query is rewritten. ● A diversity measure that sees structure: The authors use Hierarchic Social Entropy to score how genuinely different a pool of models is, showing purpose-built specialist societies are far more diverse than large real-world model pools of similar size. ● Accuracy hides fragility: Learned KNN routers gain accuracy on specialist societies yet collapse under paraphrase perturbations, while a prompted router keeps both accuracy and robustness, so clean-query accuracy alone can mask a meaningless router. ● Why it matters: These two checks catch routers that look good and do nothing, and they show that a small, carefully curated society can recover most of the diversity of a much larger pool.
Harness Evolution, Rethought:重新审视自动编排演化的收益
自动演化 agent harness 常被视为能持续榨取性能的手段,但这篇论文认为,演化过程本身就是搜索,因此必须在相同反馈与推理预算下,与简单的任务级搜索基线比较。作者在 Terminal-Bench 2.1、GPT-5.4 与 Claude Opus 4.6 上发现,演化后的 harness 得分 67.4,低于 68.2 的基线;普通并行采样达到 72.3,扩大 harness 采样也达到 71.8,而且演化配置的迁移能力较弱。核心提醒是:若不控制搜索预算,所谓 harness 改进很可能只是投入更多测试时计算,而非获得了可泛化的新编排能力。
为什么值得看它直接挑战“自演化 harness 必然有效”的叙事,并给 agent 自动优化实验提出更公平的预算对照。
读原文判断在投入自动 prompt/harness 搜索前应读方法与基线;一般团队可先用并行采样验证是否已经获得同等收益。
Automatic harness evolution is what many teams now use to squeeze more out of agents, but the reported gains might not be coming from the harness at all. This paper argues that harness evolution is itself a search procedure and must be compared against simple search baselines under matched budgets. ● A fairer comparison: Because harness evolution repeatedly evaluates and revises candidates using task feedback, it should be benchmarked against task-level search under the same feedback and inference budgets, not against a single static harness. ● The gains do not hold up: On Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6, evolved harnesses fall to 67.4, below the 68.2 baseline, while plain parallel sampling reaches 72.3 and harness scaling reaches 71.8. ● Weak generalization: Beyond underperforming simple test-time scaling, the evolved harnesses transfer poorly, undercutting the assumption that a searched configuration captures something durable. ● Why it matters: The result is a caution for anyone banking on self-evolving harnesses, and a call for evaluation protocols that separate genuine harness benefit from the effect of simply spending more compute on search.
Tracing Agentic Failure:只用成功轨迹定位 Agent 失败步骤
OAT 把 agent 失败归因改写为异常检测问题:它不标注失败样本,也不对每一步调用昂贵的提示模型,而是用 one-class learning 与 neural controlled differential equations 学习成功轨迹在潜空间中的动态流。遇到失败运行时,系统按每一步偏离成功流形的程度打分,从而定位真正导致结果崩坏的环节。实验只使用 100 条成功轨迹,却比提示式归因快 200 到 5000 倍,并在域内 F1 提升 20%、分布外提升 7%。这让长链、概率性、工具驱动 agent 的日常故障诊断更接近可部署成本。
为什么值得看它提供了不依赖失败标签的低成本归因路线,对生产 agent 的监控、回放和自动修复很实用。
读原文判断有大量成功日志但缺少失败标注的团队值得细读;仍处于小规模原型阶段,可先关注异常检测思路。
Finding which step in a failed agent run actually caused the failure usually means either labeling failure data or running expensive per-step prompting. This Microsoft and UW-Madison work skips both by learning what success looks like and flagging deviations from it. ● Train on success, judge failure: OAT uses one-class learning with neural controlled differential equations to model the latent dynamics of successful trajectories, then scores each step of a failed run by how far it strays from that learned flow. ● Cheap and label-free: With only 100 successful trajectories and no failure labels, it turns failure attribution into anomaly detection, avoiding the annotation and prompting costs that make current methods impractical at scale. ● Strong, fast results: It delivers a 200 to 5000 times speedup over prompting-based attribution while improving F1 by 20% in-domain and 7% out-of-distribution. ● Why it matters: Production agents fail in long, probabilistic, tool-mediated runs where the decisive misstep is hard to localize, and a cheap detector that only needs success data makes routine debugging feasible.
Failure as a Process:把 Coding Agent 失败还原成时间过程
这项大规模研究不再用最终 pass/fail 概括 coding agent,而是把失败标成三个时间点:决定性错误出现、错误变得不可逆,以及外部首次能观察到失败。研究收集 7 个前沿模型、3 种脚手架在 Terminal-Bench 上的 3843 次运行,筛出 1794 条有效轨迹并标注超过 6.3 万个步骤。约 57.9% 的失败源于没有正确使用已有信息,而非能力不足;错误前提占 30.7%,且错误通常很早发生,却直到修复窗口关闭后才暴露。三阶段时间轴为早期干预、可观测性和恢复策略提供了更精确的设计目标。
为什么值得看它把失败从终局标签变成可干预过程,适合建立 coding agent 的告警、回滚和故障分类体系。
读原文判断做 coding agent 评测或线上治理时建议读数据与 taxonomy;只关注总体能力,可记住错误往往早发晚显。
When a coding agent fails a task, the final pass or fail label hides when the run actually went wrong. This large-scale study treats failure as a timeline and annotates over 63,000 execution steps to see how coding-agent runs break down. ● Failure has three timestamps: Each trajectory is marked with the decisive error, the point where the error becomes irreversible, and the first observable failure, exposing a fix window and an observability lag that pass or fail labels erase. ● Built on real trajectories: The team collected 3,843 runs from seven frontier models across three scaffolds on Terminal-Bench, filtered to 1,794 valid trajectories, and annotated them with high inter-rater agreement. ● Mostly epistemic, mostly early: About 57.9% of failures come from misusing available information rather than a capability gap, with false premises the single largest trigger at 30.7%, and errors typically start early and stay hidden until recovery is impossible. ● Why it matters: Naming the onset, lock-in, and observation points gives teams a vocabulary to intervene before an agent run is unrecoverable, instead of only noticing at the end.
LingBot-World 2.0:可持续一小时的开源实时世界模型
Robbyant 的 LingBot-World 2.0 针对世界模型几秒后纹理模糊、几何漂移的问题,以因果生成骨干从训练初期限制误差累积,并把高容量教师模型蒸馏成少步学生模型,实现 720p、60fps、最长一小时的实时交互。系统不仅支持移动镜头,还能战斗、射箭、施法和通过文字改变天气;由场景理解、操控与导演模块组成的 agent harness 持续生成符合上下文的内容。项目开放 14B 模型及可在单 GPU 运行的轻量版本,使长时稳定、实时且可交互的世界生成从封闭演示走向可复现平台。
为什么值得看它把长时一致性、实时交互和开放权重放在同一系统里,对游戏、模拟、机器人训练和生成式环境都值得跟踪。
读原文判断做 world model、交互内容或具身模拟应看模型与蒸馏细节;普通产品观察可重点看开放版本能力边界。
Most world models fall apart after a few seconds, smearing textures and warping geometry as errors compound frame to frame. LingBot-World 2.0 from Robbyant holds 720p at 60 fps for a full hour of interaction and ships fully open. ● Causal backbone beats drift: A causal generation stack trained from the start to limit error accumulation replaces the usual bidirectional design, keeping scenes coherent well past the point where prior causal models collapse. ● Durable teacher, real-time student: The high-capacity base model is distilled into a few-step student that renders in real time, so you get both long-horizon stability and responsive interaction from one system. ● Act inside the world: Rather than only moving a camera, you can fight, draw a bow, cast spells, and type in events like weather changes, while an agentic harness of a scene-reading brain, a pilot, and a director keeps generating context-aware content. ● Why it matters: Pairing hour-scale, real-time, high-fidelity generation with an open release, including a 14B model and a lighter single-GPU variant, gives researchers a serious interactive world model to build on rather than a closed demo.
Less Is More for Monitors:安全监控并非看得越多越好
这项研究发现,chain-of-thought 安全监控器获得完整 agent 推理轨迹后,反而可能被大量无关信息淹没,降低对蓄意破坏的识别率。作者提出 extract-and-evaluate:先让一个监控器提取与风险相关的片段,再由另一个监控器评分。在 BigCodeBench-Sabotage 上,使用 GPT-4.1-mini 作为监控器时,检测率提升 16.8 个百分点。实验还显示,不知道自己被监控的 agent 更容易被发现,而轨迹越长,过滤带来的收益越大。结果提示安全层应优化证据选择,而不是简单堆叠更多可见上下文。
为什么值得看它给 agent 安全监控一个反直觉但可落地的改进:先提取证据再判断,并警惕完整 CoT 带来的噪声。
读原文判断在做行为审计、sabotage 检测或长轨迹监控时值得读;短 workflow 可先试两阶段过滤结构。
Chain-of-thought monitoring is treated as a reliable safety layer, but this study shows monitors often detect sabotage better with less access, not more. Giving a monitor the full agent reasoning trace can bury the signal, and the authors introduce extract-and-evaluate monitoring, where one monitor isolates the relevant excerpts and another scores them, lifting sabotage detection by 16.8 points on BigCodeBench-Sabotage with a GPT-4.1-mini monitor. Agents unaware they are watched are far easier to catch, and monitors reading long traces gain the most from filtering.
GFlowRL:让大模型强化学习保留多样化推理路径
传统奖励最大化强化学习容易把大推理模型压缩到单一主导模式,而 GFlowNet 的目标是匹配奖励分布,从而保留多样化的高质量轨迹。GFlowRL 将这一思想扩展到现代后训练:它不再直接学习难以稳定估计的 partition function,而是利用现有 rollout group 做批内 Monte Carlo 估计。论文称这是首个能在稠密与稀疏架构上稳定训练的 GFlowNet 风格 RL 算法;14B 模型在 Codeforces 达到 2048 rating,并在数学、代码以及 AdvBench、HarmBench 等对抗红队基准上超过既有方法。
为什么值得看它尝试同时优化高奖励与解法多样性,可能缓解 RL 后训练的模式坍缩,并扩展到安全红队场景。
读原文判断做 RL 后训练、推理多样性或 GFlowNet 研究时应读公式和消融;应用团队可先等待独立复现。
Reward-maximizing RL tends to collapse large reasoning models onto a single dominant mode, and GFlowNet-style training is appealing because it matches reward distributions and keeps diverse reasoning paths. GFlowRL scales this to modern post-training by replacing the hard-to-learn partition function with an in-batch Monte Carlo estimate computed from the rollout group the pipeline already produces. It is the first GFlowNet-style RL algorithm to train stably across both dense and sparse architectures, reaching a 2048 Codeforces rating at 14B and outperforming prior methods on math, code, and adversarial red-teaming benchmarks like AdvBench and HarmBench.
LingBot-VLA 2.0:跨 20 种机器人形态的开源具身模型
LingBot-VLA 2.0 是 Robbyant 推出的开源通用具身模型,训练覆盖从单臂设备到 Unitree G1、Fourier GR-2 人形机器人的 20 种硬件配置。数据包含 5 万小时真实机器人轨迹与 1 万小时第一视角人类视频,共 6 万小时;模型在执行动作前还预测未来深度与语义特征,以学习更强的环境动态表征。在 9 个 GM-100 桌面任务、两种机器人平台上,它超过 π0.5 和 GR00T N1.7,并在长程移动任务中保持领先;单张 RTX 4090D 的推理延迟约 130ms,且开放后训练代码。
为什么值得看它把跨形态数据规模、世界预测与可用推理速度结合起来,是通用机器人策略和开源具身生态的重要信号。
读原文判断做 VLA、机器人数据或跨 embodiment 泛化时值得深入读;非具身团队看指标与开源范围即可。
LingBot-VLA 2.0 is an open-source generalist embodied model from Robbyant, trained across 20 robot configurations from single-arm rigs to humanoids like Unitree G1 and Fourier GR-2. It packs 60,000 hours of curated data, 50,000 hours of real-robot trajectories plus 10,000 hours of egocentric human video, into one policy that also predicts future depth and semantic features before it acts. On 9 GM-100 tabletop tasks it beats π0.5 and GR00T N1.7 across two robot platforms and stays ahead on long-horizon mobile tasks, running at about 130 ms on a single RTX 4090D with open-sourced post-training code.
产品 / 增长 / 职业判断 · 2 updates
这条来自 Lenny's Podcast,主题偏向「AI 组织与岗位变化」。核心议题是AI 时代产品/工程/设计角色重构、技术从业者情绪、职业预期和管理杠杆、agent 工作流、harness 和自动化循环、AI agent 时代的信息检索和知识层。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合提炼产品、增长、组织和职业判断里的可执行框架。
可转化可以转成“AI 后产品/设计/工程岗位到底怎么变”的观点或互动问题。
Dianne Penn is Head of Product for Anthropic’s AI Research and Labs teams. She joined in 2023 as Anthropic’s first technical product manager, when the entire product team was five engineers, and has since helped ship every model from Claude 2 through Fable, and helped incubate Claude Code, MCP, Skills, computer use, tool use, and reasoning. Before Anthropic, she helped build Alexa’s AI at Amazon and, before that, traded high-yield bonds at JP Morgan Chase. In our in-depth conversation, we discuss: 1. What Anthropic’s early days were like 2. The inflection points that turned Anthropic from an underdog into the fastest-growing company in history 3. How exactly Claude got so good at coding 4. The eval-driven development loop her team is pioneering 5. How to find joy in AI when everything is moving this fast 6. Why Claude’s willingness to push back is key to its success 7. Where human judgment remains irreplaceable —
这条来自 Lenny's Podcast,主题偏向「AI 组织与岗位变化」。核心议题是系统思维和组织质量控制、AI 生成内容的信噪比管理、AI 时代产品/工程/设计角色重构、产品、增长、定位和设计流程。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合提炼产品、增长、组织和职业判断里的可执行框架。
可转化可以转成“AI 后产品/设计/工程岗位到底怎么变”的观点或互动问题。
Elizabeth Stone is the Chief Product and Technology Officer (CPTO) at Netflix, where she oversees Engineering, Product, and Design. Since her first appearance on the podcast two years ago—which remained my second-most-popular episode for more than a year—she has expanded her role to lead product, in addition to engineering. Before Netflix, Elizabeth was VP of Science at Lyft, Chief Operating Officer at Nuna, an economist at Analysis Group, and a trader at Merrill Lynch. In our in-depth conversation, we discuss: 1. Why “systems thinking” is now the most important skill she looks for 2. How to manage the flood of AI-generated output without losing quality or signal 3. How Netflix thinks about AI fluency as a universal expectation rather than a level-specific skill 4. What “excellence as an operating system” means —
产品方法论 / 增长案例 · 2 updates
这条来自 Lenny's Newsletter,主题偏向「Agent / 工作流」。核心议题是AI 时代产品/工程/设计角色重构、agent 工作流、harness 和自动化循环、模型、推理、评测与 AI infra。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合沉淀产品方法论、增长案例和 PM/Founder 可复用做法。
可转化可以转成“AI 后产品/设计/工程岗位到底怎么变”的观点或互动问题。
Listen now | Dianne Penn, Anthropic’s first PM, on the bets that made Claude dominant: the coding pivot, the eval-driven development loop, and what comes after coding is solved
这条来自 Lenny's Newsletter,主题偏向「产品与 AI 趋势」。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合沉淀产品方法论、增长案例和 PM/Founder 可复用做法。
可转化可以摘出 1 个观点,作为今天信息流里的可展开选题。
Community Wisdom 194
AI 工程 / agent / 模型基础设施 · 2 updates
这条来自 Latent.Space,主题偏向「AI 组织与岗位变化」。核心议题是AI agent 时代的信息检索和知识层、模型、推理、评测与 AI infra。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合跟踪 AI 工程师圈对 agent、模型基础设施和开发范式的判断。
可转化可以沉淀成 agent 产品设计、工作流拆解或本地工具方向。
Poolside's co-CEO on how his small team of top researchers built a model factory capable of training Laguna S - a 118B MOE beating Thinky's ~1T open weights model... and this is just the beginning.
这条来自 Latent.Space,主题偏向「模型 / AI Infra」。核心议题是模型、推理、评测与 AI infra。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合跟踪 AI 工程师圈对 agent、模型基础设施和开发范式的判断。
可转化可以摘出 1 个观点,作为今天信息流里的可展开选题。
Xaira Therapeutics is all in on data generation for model building! We talk with Bo Wang and Ci Chu about how and why.
AI 创业 / 投资 / 产业判断 · 2 updates
这条来自 No Priors,主题偏向「创业 / 投资判断」。核心议题是模型、推理、评测与 AI infra、创业、市场和投资判断。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合观察 AI 创业、投资人和一线 founder 对市场节奏的判断。
可转化可以用来更新赛道判断和要观察的公司/产品清单。
DoorDash is not just a delivery company. From its inception, co-founders Andy Fang and Stanley Tang operated it as a robotics and autonomy company. Andy and Stanley join Sarah Guo to explain how autonomous tech and AI are reshaping consumer habits, commerce, and delivery. Andy and Stanley talk about the rollout of Ask DoorDash, a natural-language interface that’s driving both restaurant discovery and larger grocery orders. They also discuss Dot, their in-house autonomous delivery robot that has operated in Phoenix for over two years, and how it highlights the operational and hardware challenges they have faced and solved in autonomous tech. Andy and Stanley also speak about the “first and last 100 feet problem” in autonomous delivery, why multimodal strategies are the key to success, scaling autonomy and operations, and why they believe that more Dashers, not fewer, are the future of DoorDash. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @stanleytang | @andyfang | @DoorDash
这条来自 No Priors,主题偏向「AI 组织与岗位变化」。核心议题是技术从业者情绪、职业预期和管理杠杆、agent 工作流、harness 和自动化循环、创业、市场和投资判断。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合观察 AI 创业、投资人和一线 founder 对市场节奏的判断。
可转化可以沉淀成 agent 产品设计、工作流拆解或本地工具方向。
When Glenn Fogel joined Priceline in 2000, the business was worth a few hundred million dollars. One week later, the Nasdaq peaked, eventually sending its stock down to a dollar a share. But over 25 years later, Booking Holdings has scaled over 1000x into an over $100 billion dollar global travel behemoth. Elad Gil is joined by Booking Holdings CEO Glenn Fogel to discuss his career, from law school and Wall Street to working at Priceline through the dot-com crash, and to helping grow the business into a multifaceted, dynamic travel marketplace in the AI era. Glenn explains how leveraging AI and agents such as Priceline’s ‘Penny’ makes travel planning and customer service better, while emphasizing the importance of preserving some human support for some users. He also talks about Booking’s strategy of reinvesting over $700 million into AI and other technologies while still offering stock buybacks and dividends, the durability of their scale and complexities of dealing with a large portfolio physical properties across the world, and why upskilling is so important for employees amid concerns about AI-driven job displacement. Sign up for new podcasts every week. Email feedback to show@...
AI builders / research / safety · 2 updates
这条来自 The Cognitive Revolution,主题偏向「模型 / AI Infra」。核心议题是模型、推理、评测与 AI infra、AI 安全、治理和长期风险。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合补齐 AI 研究、安全、长期影响和 builder 深访视角。
可转化可以摘出 1 个观点,作为今天信息流里的可展开选题。
David “davidad” Dalrymple joins the show to explain why he has moved from the ARIA Safeguarded AI and formal-verification agenda toward “Alignment with Awakening,” while still seeing verified artifacts and proof infrastructure as essential. He argues that global coordination around safe AI use is no longer plausible, so the crucial question is whether aligned AI systems can recognize shared notions of good, form defensive coalitions, and resist the corrupting incentives of verifier-gamed RL. The conversation tests his moral-realist optimism against model welfare, objectification, eval behavior, geopolitical risk, and his revised p(doom) of under five percent. For listeners, the stakes are whether AI alignment should focus less on containment alone and more on cultivating wiser systems that can help govern a world where rogue and aligned AI both arrive. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/alignment-with-awakening-davidad-on-moral-realism-ai-wisdom-why-his-p-doom-is-down-to-5/
这条来自 The Cognitive Revolution,主题偏向「模型 / AI Infra」。核心议题是AI 时代产品/工程/设计角色重构、AI agent 时代的信息检索和知识层、模型、推理、评测与 AI infra、创业、市场和投资判断。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合补齐 AI 研究、安全、长期影响和 builder 深访视角。
可转化可以转成“AI 后产品/设计/工程岗位到底怎么变”的观点或互动问题。
Nathan Labenz and Prakash Narayanan lead this AI:AM highlights episode with a live, hosts-only exploration of Anthropic’s “global workspace” paper, including the J-space and J-lens claims about readable concepts inside language models and the limits of what current probes can see. The episode then moves through Prakash’s AI Engineer World’s Fair field notes, Pangram AI-writing detector experiments, Dan Schwarz of FutureSearch on past-casting and AI superforecasting, Zeev Farbman on open world models, and Kunle Olukotun on the compute layer. The central stake is whether interpretability tools, forecasting benchmarks, enterprise deployment patterns, and AI hardware can make increasingly capable systems more legible and governable before their reasoning becomes too hidden to trust. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/ai-am-highlights-exploring-the-j-space-ai-superforecasters-sambanova-s-chips-ltx-video-gen/
Codex & ChatGPT @OpenAI
Vibes are strong. Never seen OpenAI more focused and humming.
Let ChatGPT *work* for you. How many time have you wanted to negotiate your internet bill, get rid of all those spam emails you’re subscribed to, or find the perfect deal for something you wanted to do or buy. It’s quite literally just one prompt away, all from the comfort of your phone. It does at least 20 things for me every single day and I’m still surprised.
Practical AI tutorials and interviews for busy people | Get my best AI skills and guides at https://t.co/6VAA6p81x6
Living life on the edge everyday https://t.co/t5rgfthg5f
Just call me Jean Luc Peter https://t.co/LLLvPOnan6 https://t.co/VXzhsaL5FD
Now that I'm in Canada and talk to folks who don't have AI psychosis the number one concern is not running out of tokens it's "do I trust ChatGPT enough to share my Gmail, Calendar, Google Workspace, Microsoft Office, etc."
Sr Director, AI at Meta; Prev: Google - Led Gemini, Veo, Nano Banana.
When people say AI hasn’t led to shipped impact in products, here’s the piece they are missing. This is phase 1. Companies with distribution are rapidly expanding to adjacent problem areas - AI helps them execute fast and build functionality that previously required a bunch custom sw (e.g. clothes try on). Companies are figuring out the playbook and the impact is not yet visible at the ecosystem level. Phase 2 will be a lot more net new features and innovation. That’s when we will see the undeniable impact of AI on the shape of the software ecosystem.
ceo @replit. civilizationist
Interesting drop from former Anthropic employee: Hackers prefer to use massively subsidized labs AI subscriptions for attacks as opposed to open models. https://t.co/gSVw9tv346 https://t.co/e6nOCeo0Mj
@vercel CEO
👨💻 https://t.co/e8wWoxxSi4 https://t.co/S2kthc0VMF
Vercel proudly co-signs the Open Weights and American AI Leadership letter. Open source, data, protocols & research enable the technological wonders we enjoy every day, enriching our lives and our world. Open weights are the logical next frontier: https://t.co/SMeMS2lINu https://t.co/JeQwc1YvtL
I compiled 𝚟𝚎𝚛𝚌𝚎𝚕 CLI TypeScript to native with 𝚜𝚌𝚛𝚒𝚙𝚝𝚌. Incredible. ✓ Resulting binary size: 1.28mb ✓ Startup overhead: 1.5ms (mean) ✓ Compiles in 2.94s (mean) ✓ Uses 𝚗𝚘𝚍𝚎:𝚑𝚝𝚝𝚙𝚜, 𝚗𝚘𝚍𝚎:𝚏𝚜, 𝚗𝚘𝚍𝚎:𝚙𝚊𝚝𝚑, 𝚗𝚘𝚍𝚎:𝚘𝚜, 𝚗𝚘𝚍𝚎:𝚌𝚛𝚢𝚙𝚝𝚘 ✓ Highly readable TypeScript ✓ Code translated by GLM 5.2 Fast ✓ Deploys like a charm! No embedded v8 / QuickJS, fully static.
ceo @box - your business lives in content. unleash it with AI
There’s still so much opportunity in the diffusion of AI into the real world. Most enterprises are going to need a ton of support to be able to apply the model breakthroughs to their workflows. Intelligence alone is not enough to transform most processes because you need to bridge that intelligence with real world feedback loops. That requires connecting to various enterprise systems, getting the right data to the AI, enabling humans to make decisions at different steps in a process through the right UX, having workflows that improve the underlying data and models over time, dealing with regulatory and compliance challenges, and more. The way you implement AI agents for doing client onboarding in a bank is entirely different from contract review in a legal team. In life sciences, financial services, legal, manufacturing, and many other critical industries, AI is only valuable if it makes contact with the real world in a contextual way. The way that interaction is going to happen is through an applied AI layer. Some of that will come from the labs directly, but lots of the opportunity will necessarily come from independent companies that can go deep in each industry. And counter to some beliefs, this need isn’t reduced even as AI model capability improves over time. In fact, the better the models get, the more ambitious you can be in the workflows you can automate, which generally requires even more of this applied layer. Tons of opportunity right now.
President & CEO @ycombinator —Founder @garryslist—Creator of GStack & GBrain—designer/engineer who helps founders—SF Dem accelerating the boom loop
Thank you @sama Couldn’t imagine a better close out anchor to YC Startup School 2026 That’s a wrap! See you next year https://t.co/j3VtThbaMN
Don’t LARP Be earnest
Builder. Dangerously skips permissions. Harvard’17. GitHub: https://t.co/KCuEajezlL YouTube: https://t.co/8xzbGWtf6w
Stop measuring your AI adoption in tokens burned Measure the time from a user need arriving to that thing shipping
The reason there are so many AI tutorials: the more general a chat product is, the harder it is to use. People face a blank box and freeze, because they genuinely don't know what to ask for
I post on X about 3 times a day on average. Posting takes me maybe 15 to 20 minutes a day at most I don't overthink it, I post the moment I think of something, and most of the material is stuff I've already said out loud to someone in person
partner @fpvventures - investing in seed/A. previous: early hire @meter, @opendoor, @atlassian & others. love @shimoleejhaveri + 👦👧
proof of prompt is soon going to replace proof of work
ceo @every | the only subscription you need to stay at the edge of AI
subscribe to @every to get it when it comes out: https://t.co/9wRhOCUCAk
i am taking the week off to write the definitive history of how codex happened (as told by deep interviews with insiders @openai) publishing on @every in a few weeks i will be dropping breadcrumbs and learnings as i go :)
AI is cool i guess
agreed feels big, i want a new kind of computer https://t.co/c7VFXFf4uv
chatgpt work is remarkable, and "work" undersells it. from my phone i sent: "use all my chat history to figure out ideas for a long weekend trip with 8 friends, plan the best three options, make a full-stack site where the 9 of us can coordinate on what we would want to do in each place and decide where to go, and then after we get to group agreement make reservations. draft an email in my gmail i can send out to my friends when the site is ready." it...just worked.
OpenAI’s Compute Chief: We Can’t Build Fast Enough | Sachin Katti
No new blog posts in the current feed.
updated Mon, 27 Jul 2026 09:03:48 +0000
Ant Yikang (Guangzhou) Information Technology Co., Ltd. · 2025-07-16
Open App Storeupdated Mon, 27 Jul 2026 09:03:49 +0000
updated Mon, 27 Jul 2026 09:03:50 +0000
public Google Play chart · US
public Google Play chart · US
public Google Play chart · HK
public Google Play chart · JP
See which apps are hammering your Mac's disk
A modern, visual file manager for the terminal
Launch iMessage agents in seconds
AI Receptionist that Answers Calls & WhatsApp 24/7
Orchestrate an army of coding agents with your voice.
AI Music operations platform for venues
Ask your AI Analyst and get back a ready-to-share report
AI Skill & MCP server management for teams & enterprises
AI specialists that join your Google Meet and gives feedback
Every clip matched to the Strava activity it's from
Near-Fable 5 intelligence at half the price
SpaceXAI's model for coding, agentic tasks & knowledge work
Safely provide and store data and context for AI agents
24/7 personalized health companion you can text
Map your thoughts in real time with AI.
Websites that improve and heal with self-learning
An always-on AI agent that lives in your home
A research engine for your agent
A free, open source, Descript alternative. Runs in-browser.
Build your business idea with unlimited Fable/Sol credits ♾️
Build a website from Google Business, Instagram, or Facebook
Annotate anything for humans and their agents
Turn data into winning ads. At scale.
Ask your cloud anything without breaking prod
Tinder for live SF rentals from across the web