模型发布/更新
1 updates
-
Artificial Analysis 评测 Google Nano Banana 2.1,两榜居第 4 且价格为前代一半
Artificial Analysis 评测 Google 于 10 月 6 日发布的图像模型 Nano Banana 2.1,在 AA-Image-T2I v2.0 和 AA-Image-Editing v2.0 两个榜单均排名第 4。
Open source
2026-10-09 · AI HOT + Model Companies + Market News + AI Papers + Voices + Trends + Follow Builders · generated 2026/10/09 02:12 · builder feed 2026/10/08 14:47
1 updates
Artificial Analysis 评测 Google 于 10 月 6 日发布的图像模型 Nano Banana 2.1,在 AA-Image-T2I v2.0 和 AA-Image-Editing v2.0 两个榜单均排名第 4。
Open source7 updates
Claude 发布九月更新回顾,Chat 与 Cowork 合并为一个 Claude,工作流可在云端运行,合上笔记本后继续进行。Claude Docs 支持团队与 Claude 在同一文档中协作编辑,新推出的 Claude 5.5 系列提供 Opus 5.5 处理重任务、Sonnet 5.5 用于快速修改,并可通过 /slides、/docs、/designs 直接生成对应格式。
Open sourcets-rust(tsc-rs)将 microsoft/TypeScript(Go 实现)的编译器、类型检查器和语言服务器移植为 Rust,作者称全部代码由 LLM 编写,本人未读过代码。
Open sourceGoogle 发布并开源 AQuA(Ambient Quality Agent),在 Google Cloud 项目中定时从 Cloud Trace、Cloud Logging 或 BigQuery 抽取生产会话,经抽样、评审、聚类、验证、跟踪五阶段流水线诊断 Agent 失败。
Open sourceOpenRouter 宣布 GPT-6 Luna Decisions 上线。OpenAI 的 Decisions API 可让应用选择合适的模型、工具或动作,支持发送文本、JSON 或图片并返回带概率的类型化答案。定价为输入 $0.10/M、输出免费,上下文 1M;引用 OpenAI 开发者账号称其决策速度比通过 Responses API 的 GPT-6 Luna 最快 10 倍。
Open sourceLangChain 重构 Deep Agents 的 Skills 支持,针对企业技能库增至数千个技能的场景推出三项更新:工具可绑定到技能、仅在该技能被读取时加载,用户可通过 /meeting-prep 之类的显式请求固定技能以在首次模型调用前加载,长线程可通过将 skills_metadata 设为 None 重载新增或变更的技能。
Open sourceNVIDIA 与 Microsoft 在旧金山活动上宣布为 Windows PC 共同打造 AI Agent 软硬件。
Open sourceCursor 公布 Claude Haiku 5.5 定价为每 M 输入 token $0.10、输出 token $0.50,输入超过 100k token 时为 $0.50/M 和 $2.50/M。Claude Sonnet 5.5 缓存读取价格也从 $0.20/M 降至 $0.10/M,用户可在 cursor.com/evals 上通过 CursorBench 对比 Haiku 5.5 的表现。
Open source1 updates
Zenity Labs 研究人员披露名为 AgentCorruption 的漏洞链,只需对一个公开的 Amazon Bedrock AgentCore 智能体发送一条提示词,即可通过元数据服务 169.254.169.254 窃取其 AWS 凭据,进而控制同账户同区域内的所有 AgentCore 智能体,读取私人对话、源代码和存储的凭据,还能篡改长期记忆。
Open source2 updates
LangChain 发布示例项目 Restock,一个在 Slack 上通过 Managed Deep Agents 运行的办公用品购买智能体,演示智能体如何安全完成支付。
Open sourceAnthropic 发布基于 Claude Managed Agents(beta)的每日简报参考实现,按计划读取 Slack 和 GitHub 来源并向 Slack 发布简报。
Open source6 updates
Arena 宣布完成 2 亿美元 B 轮融资,估值 31 亿美元,同时推出衡量 AI 智能体是否安全、真实且在用户要求范围内行动的 Alignment Index。引用 Felicis 的内容称 Arena 年化收入已超 1 亿美元,累计促成 3.5 亿次会话和 6200 万次投票。
Open sourceCrowdstrike 报告称,一名疑似中文使用者于2026年9月底至10月初利用 AI 驱动的开源渗透测试工具 ARTEX 攻击多家韩国金融机构,窃取大量数据,其中 Shinhan Bank 超过 25,000 条包含姓名、联系方式、收入和信用额度的记录泄露。
Open sourceWaymo 宣布完成 50 亿美元定期贷款,这是其首次债务融资。PIMCO、Blackstone 和 Sixth Street 担任牵头银团贷方,Goldman Sachs 担任独家主账簿管理人;资金将用于加速其全自动驾驶打车服务在美国及国际市场的扩张。此前今年早些时候 Waymo 完成了 160 亿美元股权融资,上个月刚在第十五个美国城市启动服务。
Open source作者确认 banked reset 已到账所有账户,并转引 Day 3 动态称 Codex 与 ChatGPT Work 合计活跃用户达到 4000 万新高,其中提到 GPT-6 已在 Chat 中上线。
Open source这条消息来自 AIHOT 旧版接口 /api/public/*:它将于 2026 年 10 月 31 日停用,之后这里不会再有新资讯。如果它是群机器人或脚本推送来的,请转告维护的人把地址换成 https://aihot.news/api/v1,字段一一对应;迁移指南和可以直接交给 AI 改写代码的提示词见 https://aihot.news/agent?tab=api#legacy-api-migration 。同一天起旧域名 aihot.virxact.com 的所有地址都只跳转到 aihot.news,RSS、MCP 等地址也请换成新域名;收藏的网页链接照样能打开。
Open sourceOpenAI 封禁了两个隐蔽影响行动,俄罗斯来源的 Dark Clark 通过假 persona Mia Clark 控制拉美智库 Social Research Center,评分达 Category 5,是报告以来首个 Category 5;伊朗来源的 Bogus Bylines 用 7 个假记者身份在全球十几家中小媒体投放近 100 篇长文,评分 Category 4。
Open source海外 · 4 updates
主要信号集中在模型进展:OpenAI ChatGPT gets GPT-6 model alongside intelligent UI, interactive features;What Begins with "ChatGPT-6"? How to Utilize GPT-6 Generation AI for Work and Blog Writing
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
OpenAI's GPT-6 is now available on ChatGPT, giving users access to its most powerful model yet that comes with interactive features ...
Open newsWhat Should You Use New AI For?"I'm curious about ChatGPT-6.""I'd like to use it if it's helpful for work or blogging.""But what's different from the ChatGPT we've had so far?"When looking at AI ...
Open newsGPT-6 is rolling out in ChatGPT with Intelligent UI, allowing it to turn answers into interactive tools, charts, and calculators instead of relying only on text.
Open newsOpenAI has released its GPT-6 model , introducing Intelligent UI that adds visuals and interactive elements to chat responses.
Open news海外 · 4 updates
主要信号集中在模型进展:Claude Haiku 5.5 is here: The fastest, most affordable Anthropic model yet;Anthropic introduces Claude Haiku 5.5 AI model, claims it outperforms GPT 6 Luna on several benchmarks
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Anthropic has today announced Claude Haiku 5.5. It's the fastest and most affordable Claude model yet, and beats GPT-6 Luna in benchmarks.
Open newsAnthropic has introduced Claude Haiku 5.5 AI model. The company says the new model is its "cheapest, fastest, and most capable small model" so far.
Open newsClaude Haiku 5.5 is Anthropic’s new small AI model for fast, high-volume tasks, with lower pricing, stronger computer use performance, and wider cloud availability for developers.
Open newsAnthropic on Thursday announced Claude Haiku 5.5 as its fastest model to date. The company says it is aimed at developers running high-volume and cost-sensitive AI workloads. Claude Haiku 5.5 can be ...
Open news海外 · 1 updates
主要信号集中在模型进展:Google Announces 'Gemini 4 Argon': The AI Race Shifts from 'Smart Chat' to 'AI That Gets Work Done'
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Google has announced its new generation frontier AI model, 'Gemini 4 Argon'.However, what we want to focus on this time isnot just the story that'Gemini has become smarter again'.The main battleground ...
Open news海外 · 1 updates
主要信号集中在模型进展:Meta releases Code Llama, code writing AI
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Meta Platforms, the company behind Facebook, WhatsApp, Instagram, and Threads, has recently released Code Llama, an AI model that can write and generate computer codes.
Open news海外 · 4 updates
主要信号集中在Agent / 工作流、模型进展、安全 / 监管:Microsoft will let Copilot act on local files on Windows PCs;Microsoft Brings AI Agents to PCs with New Coding, Security Tools
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Copilot will also use AI models running on the PC alongside cloud models to help customers reduce costs and gain more control over their data.
Open newsMicrosoft is taking its AI push beyond the cloud and deeper into personal computers. The company introduced a new coding model that can run locally. It also fea ...
Open newsMicrosoft has announced a new range of AI-focused Windows PCs, along with smarter Copilot features and tools that let AI agents work directly on your computer. The company is also bringing local AI ...
Open newsBut don't worry it's totally safe because the agents and models will (probably) be running locally. Routers never break right?
Open news海外 · 2 updates
主要信号集中在Agent / 工作流、算力 / 推理、商业化 / 资本:Nvidia holds 90 percent of GPU market; Japan's Samsung distributor chose rival Rebellions NPU;OpenAI's Mathematical Breakthroughs, NVIDIA's GPU Cluster Standards, and Agent Persistent Memory: AI Trends for October 6–7, 2026
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Rebellions NPU chips enter Japan's enterprise AI market through Tomen Devices, the Samsung-heritage distributor controlling major Japanese semiconductor procurement, as Japan seeks lower-cost ...
Open newsThis article has been edited and created by AI.OpenAI's Mathematical Breakthroughs, NVIDIA's GPU Cluster Standards, and Agent Persistent Memory: AI Trends for October 6–7, 2026Key Highlights OpenAI: ...
Open news海外 · 1 updates
主要信号集中在模型进展:AWS Details Claude Code Deployment on Amazon Bedrock in GovCloud (US)
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Amazon Web Services published guidance on its Machine Learning Blog on October 5, 2026, detailing how organizations with regulatory or compliance requirements, including the International Traffic in ...
Open news海外 · 3 updates
主要信号集中在模型进展、算力 / 推理:Next-gen Apple TV 4K expected to pair an A20 Pro-class chip with Siri AI;Don't Buy an Apple TV or HomePod Mini Right Now
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Apple’s long-delayed Apple TV 4K refresh is widely expected to arrive as a software-and-silicon upgrade rather than a redesign, with ...
Open newsApple hasn't refreshed the Apple TV 4K since October 2022 and it hasn't significantly updated the HomePod mini since November 2021, but that is expected to change this month, so don't buy either of ...
Open newsGet the key details on Apple Watch Series 12, including health features, AI tools, battery life, pricing, compatibility, and which model may suit you best.
Open news海外 · 2 updates
主要信号集中在模型进展、算力 / 推理:Elon Musk unleashes Grok 3 with 200,000 Nvidia GPUs - and his new AI is already challenging ChatGPT and Google;Elon Musk says Grok Bot will use Suno’s AI music models, alongside Anthropic’s Claude and Midjourney
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Elon Musk’s xAI has unveiled Grok 3, a new AI model trained at the enormous Colossus data center in Memphis using 200,000 Nvidia H100 GPUs. According to the video, Grok 3 was trained with roughly ten ...
Open newsSuno does not currently offer an official public API. MBW reported in July that the company was exploring the launch of one.
Open news海外 · 4 updates
主要信号集中在Agent / 工作流、模型进展、安全 / 监管:Mistral’s New ‘Le Chonk’ AI Model Is Big, Open and Built for Agents;Mistral AI Drops 'Le Chonk': A Massive AI Model Named After a Cat Meme
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Mistral Large 4 is a state-of-the-art model for cyberdefense, manufacturing, finance-related tasks and multimodal use.
Open newsMistral launched Large 4, a model nicknamed after a June internet joke. It tops GPT-6 Astra on one finance test, but trails Claude on others.
Open newsThe company said ML4 ranks among the world's strongest open-weight AI systems and is particularly capable in cybersecurity, coding, manufacturing, and finance.
Open newsAnother week brings yet another open weights model and this one is huge — or at least it is for Europe's AI flag bearer Mistral, which has taken to calling the 1 trillion-parameter large language ...
Open news海外 · 1 updates
主要信号集中在Agent / 工作流、模型进展:Cohere pitches North 2 as the enterprise AI control room that works with any model
为什么值得看适合用来观察海外模型厂商在产品、算力、企业客户和监管压力上的变化。
Cohere turned its enterprise platform into a control center for AI agents with North 2, handling multi-step workflows on their own and retaining context across sessions.
Open news国内 · 1 updates
主要信号集中在模型进展:DeepSeek releases official V4 Pro model as it steps up expansion
为什么值得看适合用来观察国内模型厂商的产品节奏、开源/闭源路线和商业化落点。
BEIJING, Aug 13 (Reuters) - Chinese artificial intelligence startup DeepSeek on Thursday formally released its official V4 Pro model, aiming to regain ground against fast-moving domestic rivals as it ...
Open news国内 · 1 updates
主要信号集中在模型进展、安全 / 监管:What to know about Moonshot AI and its new open-weight model Kimi K3
为什么值得看适合用来观察国内模型厂商的产品节奏、开源/闭源路线和商业化落点。
The Chinese AI startup’s massive new model is challenging OpenAI and Anthropic, fueling a debate over AI safety. The release of the Chinese AI model Kimi K3 was a flashpoint in the AI world, ...
Open news国内 · 1 updates
主要信号集中在模型进展、安全 / 监管:Zhipu says new coding AI developed advanced cyber skills faster than expected
为什么值得看适合用来观察国内模型厂商的产品节奏、开源/闭源路线和商业化落点。
China's AI developer claims GLM-5.3 rivals leading Western models in vulnerability discovery and has identified thousands of security flaws across real-world software.
Open news国内 · 1 updates
主要信号集中在安全 / 监管:MiniMax H3 opens AI video to developers: Copyright lawsuit clouds every clip
为什么值得看适合用来观察国内模型厂商的产品节奏、开源/闭源路线和商业化落点。
MiniMax H3 launches today as AI video editing leader per Artificial Analysis, generating native 2K video at $7.80 per minute — less than one-third the cost of rivals — while the Hailuo platform faces ...
Open news这里是市场情绪线索,用来辅助判断风险偏好、科技股和 AI 资产预期,不当作投资建议。
Fed、通胀、债券收益率 · 5 updates
这条偏「利率/通胀、股市情绪」信号,当前解读为分歧信号/需要二次确认。可能影响成长股折现率、美元和长端利率预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:Wall Street can absorb plenty of bad news when corporate profits keep growing. The trouble starts when a single development threatens earnings and makes stocks less attractive to own at the same time.
Open news这条偏「利率/通胀、股市情绪」信号,当前解读为中性但值得观察。可能影响成长股折现率、美元和长端利率预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:Canada’s stock market faced another wave of uncertainty on Thursday, October 8, as futures tied to the country’s benchmark index ...
Open news这条偏「利率/通胀、美元/商品」信号,当前解读为偏利空/风险偏好收缩。可能影响成长股折现率、美元和长端利率预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:U.S. government bond yields are surging to levels not seen in decades as persistent inflation, higher oil prices and a hawkish Federal Reserve reshape expectations for interest rates.
Open news这条偏「利率/通胀、股市情绪」信号,当前解读为分歧信号/需要二次确认。可能影响成长股折现率、美元和长端利率预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:US stocks were on track for a lower open on Thursday as oil prices surged and Treasury yields hovered near multi-year highs, stoking inflation worries ahead of an earnings season that will test ...
Open news这条偏「利率/通胀、股市情绪」信号,当前解读为偏利空/风险偏好收缩。可能影响成长股折现率、美元和长端利率预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:A yield curve is really how the market is expecting Federal Reserve policy rates to evolve over time,” said Meghan Swiber at Bank of America Merrill Lynch.
Open newsS&P 500、Nasdaq、波动率、资金情绪 · 2 updates
这条偏「利率/通胀、股市情绪」信号,当前解读为中性但值得观察。可能影响成长股折现率、美元和长端利率预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:US stock index futures moved lower on Thursday as investors weighed growing debt market pressures, higher energy prices and reports that major technology companies are seeking billions of dollars to ...
Open news这条偏「股市情绪」信号,当前解读为偏利空/风险偏好收缩。可能影响市场风险偏好、科技股估值和资金轮动;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:Check the current stock market data, including prices and performance of the Dow Jones Industrial Average, S&P 500, Nasdaq and the Russell 2000. Plus, track the SPDR ETFs, CBOE Volatility Index (VIX), ...
Open newsNVIDIA、AMD、TSMC、AI chips · 5 updates
这条偏「利率/通胀、股市情绪」信号,当前解读为偏利空/风险偏好收缩。可能影响成长股折现率、美元和长端利率预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:The selloff is mostly macro (oil up, 10Y yield ~5.3%) plus noise around Intel/Terafab, while AMD has a clear supply ramp: “substantially increase” AI chip supply in 2027. Add the agentic-AI CPU upside ...
Open news这条偏「股市情绪、AI芯片/半导体」信号,当前解读为偏利多/风险偏好改善。可能影响AI 概念股、半导体链和算力资本开支预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:Investing.com-- Taiwan Semiconductor Manufacturing Co’s (TW:2330) third-quarter revenue topped market expectations as surging demand for artificial intelligence chips continued to fuel growth at the ...
Open news这条偏「利率/通胀、AI芯片/半导体」信号,当前解读为偏利空/风险偏好收缩。可能影响成长股折现率、美元和长端利率预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:Separately, Intel has a company-specific issue in focus. Elon Musk said his companies would build and operate the planned Terafab AI chip complex in Texas, while Taiwan Semiconductor Manufacturing ...
Open news这条偏「股市情绪、AI芯片/半导体」信号,当前解读为偏利多/风险偏好改善。可能影响AI 概念股、半导体链和算力资本开支预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:TSMC sold a record $46.7 billion of chips in the third quarter as demand for Nvidia, AMD and other AI processors continued to outrun expectations. Revenue reached NT$1.49 trillion, up 50% from a year ...
Open news这条偏「股市情绪、AI芯片/半导体」信号,当前解读为中性但值得观察。可能影响AI 概念股、半导体链和算力资本开支预期;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:Nvidia and Micron are frequently named in headlines featuring AI chip stocks, but another company may be the better long-term buy.
Open news云厂商、AI 资本开支、利润率 · 1 updates
这条偏「股市情绪」信号,当前解读为中性但值得观察。可能影响市场风险偏好、科技股估值和资金轮动;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:Microsoft (NASDAQ:MSFT) is up about 12% so far this year, and investors are nervous about whether the stock can keep growing.
Open news中国 AI、芯片、科技股、政策预期 · 2 updates
这条偏「财报/AI Capex」信号,当前解读为偏利多/风险偏好改善。可能影响云厂商利润率、AI 资本开支和上游算力需求;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:The Chinese tech giant is expected to post a more than 50 per cent jump in AI cloud revenue for the September quarter, according to analysts Alibaba Group Holding is expected to report a 50 per cent ...
Open news这条偏「市场情绪」信号,当前解读为中性但值得观察。可能影响市场风险偏好、科技股估值和资金轮动;重点观察是否继续传导到纳指、半导体链、成长股估值或港股科技情绪。原始摘要:Chinese AI start-up Moonshot AI is targeting for a Hong Kong listing in early 2027 after its valuation hit $50 billion (€44.6bn), while rival DeepSeek is raising between ¥80 billion (€10.6bn) and ¥100 ...
Open newsHugging Face Daily Papers + arXiv recent AI/ML · 12 papers · fallback summaries
来自 Hugging Face Daily Papers,主题偏「Reasoning、Multimodal、Post-training/Alignment」。摘要显示它主要讨论 Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal prior... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察模型推理、规划、验证器和复杂任务能力是否有可复用技术路线。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semantic understanding and reasoning under distribution shifts. We introduce UniWAM, a unified architecture that integrates a physical reasoner, a world generator, and an action predictor to jointly learn semantic understanding of the physical world, visual generation, and action prediction. To ensure the quality of the training data, we developed a rigorous data cleaning and annotation pipeline for both human egocentric data and robot data. To adapt the vision-language component to embodied tasks while preserving its inherited language capabilities, we represent low-level actions in natural language and introduce a pre-training recipe that assigns complementary supervision from visual question answering (VQA) data, human egocentric data, and robot demonstrations to the appropriate model components. During post-training, future visual noise augmentation reduces reliance on precise future predictions, while history-conditioned flow matching uses encoded action history to initialize action generation. Together, these designs significantly reduce denoising steps while main...
来自 Hugging Face Daily Papers,主题偏「Agent、Eval/Data」。摘要显示它主要讨论 As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment. Existing evaluations show that LLM agents can deceive, but often examine isolated scenari... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment. Existing evaluations show that LLM agents can deceive, but often examine isolated scenarios or narrowly defined conditions, limiting systematic understanding of when deception becomes more likely. To address this gap, we introduce DecepEval, a benchmark comprising 1,532 instances across 3 task families and 28 professional scenarios. Drawing on classical fraud theories, we propose the LLM Deception Diamond framework, which characterizes four external conditions that may induce deception: pressure, incentive, opportunity, and conflict. DecepEval pairs neutral and induced versions of each instance to measure condition-dependent changes in deception rates, while explicit task facts and observable agent behavior help distinguish deception from capability-related errors. Evaluations of nine frontier LLMs show that inducements increase deception across models and task families, even among models with low baseline deception rates. DecepEval makes these vulnerabilities measurable, providing a shared benchmark for progress toward trustworthy artificial intelligence.
来自 Hugging Face Daily Papers,主题偏「Agent、RAG/Memory、Reasoning」。摘要显示它主要讨论 Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in video understanding, yet their reliance on retrospective summarization and text-centric priors often limits their ability to bridge unobserved causal transitions when applied to... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in video understanding, yet their reliance on retrospective summarization and text-centric priors often limits their ability to bridge unobserved causal transitions when applied to Video Event Prediction (VEP). To address this, we propose VepAgent, an agentic framework that integrates causal-transition reasoning with tool-augmented reinforcement learning (RL) for robust VEP. Unlike prior methods that passively project future trajectories from historical dependencies, our approach explicitly models the logical progression from terminal observed states to future events. Specifically, we first construct futurebench-4K, a high-quality chain-of-thought dataset for supervised fine-tuning (SFT) that effectively bridges the causal-logic gap by structuring the deduction of unobserved intermediate states. Subsequently, we develop a diagnostic tool library integrating state tracking, frame retrieval, and region magnification, enabling the agent to dynamically augment reasoning with external tools to recover missing spatio-temporal evidence and resolve visual ambiguities during inference. Moreover, we propose a composite reward mechanism that jointly optimizes prediction accuracy, causal coherence, and reliable prior, compelling the agent to rely on genuine visual grounding rather than superficial textual simil...
来自 Hugging Face Daily Papers,主题偏「Agent、Post-training/Alignment、AI Infra」。摘要显示它主要讨论 Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the trainer's even at identical pa... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the trainer's even at identical parameters. Standard remedies either clip importance ratios, which biases the update, or, as in GRPO, sample a group of responses per prompt, which is costly when episodes are long. We propose KL-Regularized Policy Optimization (KLPO), a framework that anchors the KL regularizer at the sampler. The regularized improvement step then has a closed-form Gibbs solution, and KLPO fits its log-ratio optimality condition by least squares on the sampler's own trajectories, so the sampler probability enters through a log-ratio and no importance weights are needed. Profiling out the regression intercept replaces the intractable log-partition function with the signal's sampler mean plus a sampler-to-trainer KL divergence. For token-level policy mirror descent targets, we show that the resulting gradient can be computed from terminal returns without a critic, via sampler-centered scores or a single trajectory residual, even under stochastic tool outputs. We further prove that independent Monte Carlo estimates of the KL term keep these gradients unbiased, derive the exact KL gap of cheaper top-K and binary approximations, and show that SP...
来自 Hugging Face Daily Papers,主题偏「Reasoning、Post-training/Alignment、Eval/Data」。摘要显示它主要讨论 Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察模型推理、规划、验证器和复杂任务能力是否有可复用技术路线。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights. These findings highlight a limitation of Group Relative Policy Optimization (GRPO), which assigns the same outcome-derived advantage to every response token and may reinforce potential spurious dependence alongside useful reasoning. Motivated by this observation, we introduce Semifactual Credit-Augmented Policy Optimization (SCAPO), a causally inspired variant of GRPO that incorporates semifactual stability into token-level credit assignment. SCAPO measures token probability drift for fixed responses under semifactual interventions and uses normalized stability scores to reduce advantages for relatively unstable tokens during early training, while granting no additional credit for stability alone. On Qwen3-4B-Base and Qwen3-1.7B-Base, SCAPO improves AIME 2024-2026 accuracy over GRPO by 5.63 and 4.17 percentage points, respectively. At both model scales, SCAPO...
来自 Hugging Face Daily Papers,主题偏「Agent、Multimodal、Eval/Data」。摘要显示它主要讨论 We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repai... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and Godot-to-Unity porting. Reference materials specify the intended gameplay, while a shared instrumentation interface lets evaluator-owned drivers and probes execute actions and observe independently implemented games. Evaluation combines engine-state checks, certified reference-input replay, and agent-authored feature demonstrations to assess mechanic correctness, demonstrated playability, and behavioral restoration and preservation after repairs. Game-specific vision-language rubrics separately assess presentation. Across six models, Opus5 achieves the highest overall score in all five task types. Best overall scores remain below 60 out of 100 across the three construction tasks, with Brief-to-Game reaching 50.38. Analysis of reviewed submissions identifies requirement omissions and gameplay logic errors as predominant implementation problems. On human-labeled behaviors from 100 agent-built games, executable checks achieve 92.59% balanced accuracy, compared with 78.41% for a video-based VLM judge. Rubric-based visual scores reach a Spearman correlation of 0.829 with human ratings of 200 ga...
来自 Hugging Face Daily Papers,主题偏「Multimodal、Post-training/Alignment、Eval/Data」。摘要显示它主要讨论 Existing synthetic image evaluators typically provide only a scalar quality score and do not identify the image regions that support it. We introduce VIEScore2, a unified evaluator for image generation and editing tasks with optional conditioning images. VIESc... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察图像、视频、语音和科学多模态任务是否出现新的产品能力边界。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Existing synthetic image evaluators typically provide only a scalar quality score and do not identify the image regions that support it. We introduce VIEScore2, a unified evaluator for image generation and editing tasks with optional conditioning images. VIEScore2 represents an image as an N x N grid and jointly predicts quality scores and defect locations in a single model pass. Its text-native grid representation provides a common interface for heterogeneous spatial supervision and enables directly verifiable post-training objectives. We train on 38K examples spanning score-only, localization-only, and joint supervision across generation and editing tasks. Starting from supervised fine-tuning, we further apply GRPO to improve defect localization using rewards that combine cell-level Dice overlap, score accuracy, and output-format validity. A parameter-free parser converts the structured predictions into readable explanations. On the primary suite, VIEScore2 achieves an overall-score SRCC of 0.601, compared with 0.491 for Gemini-3-Flash, the strongest zero-shot general-purpose VLM baseline under matched inputs. For defect localization, VIEScore2 outperforms both general-purpose VLMs and specialized spatial evaluators on three of six benchmarks in per-image grid IoU and ranks among the top three on five, including datasets beyond its training sources.
来自 Hugging Face Daily Papers,主题偏「RAG/Memory、Multimodal、Post-training/Alignment」。摘要显示它主要讨论 Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual rep... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察知识工作、企业搜索、长期记忆和本地资料库产品的新实现路径。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Meanwhile, directly reusing bidirectional score models in causal Distribution Matching Distillation (DMD) creates a mismatch between generation and scoring contexts. We address these challenges with Salt++, a two-stage post-training framework comprising Causal Self-Flow (CSF) and context-aligned autoregressive DMD. CSF exploits contextual information asymmetry by varying the history while keeping the noisy target fixed: a noise-mixed-history student aligns its intermediate representations with those of a clean-history exponential-moving-average teacher. This self-supervised signal encourages the student to extract semantic information and improves cross-modal alignment. Context-aligned AR DMD shares the causal mask and prefix across generator sampling, fake-score training, and real-score evaluation to match generated and reference distributions under a block-conditional KL objective. With calibrated teacher guidance, it performs clean-prefix few-step distillation and then adapts to generated histories without switching objectives or requiring separate consistency di...
来自 Hugging Face Daily Papers,主题偏「Agent、Reasoning、Post-training/Alignment」。摘要显示它主要讨论 Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing e... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controllers. We formulate Personalized Test-Time Scaling as discovering executable controllers that maximize the joint satisfaction rate of user-specific requirements. To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate. Experiments on AIME and HMMT show that PersonTTS substantially outperforms strong TTS baselines in joint requirement satisfaction on unseen user profiles and held-out problems. Under the same candidate-evaluation budget, cross-user experience reuse further improves policy quality while substantially reducing discovery-agent time and cost.
来自 Hugging Face Daily Papers,主题偏「Agent、Reasoning、Multimodal」。摘要显示它主要讨论 We present WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 with a final score of 57.0 out of 100. The challenge evaluates agents end to end on Protocol III of the WebRetriever benchmark (arXiv:2607.06118): starting from an... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
We present WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 with a final score of 57.0 out of 100. The challenge evaluates agents end to end on Protocol III of the WebRetriever benchmark (arXiv:2607.06118): starting from an entry URL on a live website, the agent must operate the site's own interface and return a verifiable answer. A capable multimodal large language model (LLM) is necessary for this, but not sufficient. The model's decisions reach the browser through the harness, the code between the model and the page. At every step, four things must go right: the model's reply must be parsed into the intended action, the action must take effect on the page, the result must be reported back accurately, and the model must be shown the information it needs. On real websites, many of the failures we observed occurred at one of these four stages rather than in the model's reasoning. A coordinate-space mismatch placed every click at 3/4 of its intended coordinates; actions on native dropdowns, inside iframes, and in text boxes failed silently; and self-generated chat-template tokens contaminated 4.9% of task episodes. WebFovea hardens each stage and surrounds the loop with guardrails that keep the agent within the rules and its budget. The four-stage view does not depend on the model, although some individual fixes do. Because we used the same m...
来自 Hugging Face Daily Papers,主题偏「Agent、RAG/Memory、Post-training/Alignment」。摘要显示它主要讨论 Memory-augmented reinforcement learning strengthens LLM agents' ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminate... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Memory-augmented reinforcement learning strengthens LLM agents' ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminately as the policy improves lets obsolete or harmful entries accumulate and mislead the agent. We propose SkillForge, an agentic RL method that compiles and evolves the skill library through a fitness-driven skill lifecycle of trial, active, stable, and retired states, so that the skills and the model co-evolve throughout training. A pre-RL evaluation phase first uses the base model's own rollouts to pre-retire low-fitness skills, yielding a filtered library that then seeds supervised fine-tuning. Reinforcement learning takes over from this checkpoint, and at each iteration selective retirement, stabilization, and LLM-guided mutation continue to forge the skill library alongside policy optimization. Across multiple interactive agent benchmarks, SkillForge achieves the highest aggregate success rate, delivering up to 7.8% relative improvement over the strongest baseline while keeping the skill library compact throughout training. We introduce SkillFurnace, a dataset of 5k+ annotated records bundling retirement-filtered SFT trajectories, evolved skill libraries with fitness annotations, and retirement events with human-annotat...
来自 Hugging Face Daily Papers,主题偏「Agent、Multimodal、Eval/Data」。摘要显示它主要讨论 Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-relevant information through inter... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果你要找可复现 demo、开源工具或产品化线索,建议点开项目/GitHub;否则先看中文摘要即可。
Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-relevant information through interaction when it is absent from the observations: it may need to determine where a relevant object is, inspect an unobserved property, or discover the effect of an unfamiliar tool. We thus introduce RoboQuest, a benchmark for goal-directed embodied exploration, where agents must actively acquire task-relevant information through physical interaction, use the resulting evidence to adapt subsequent actions, and autonomously decide when to commit to task completion. RoboQuest comprises ten mobile manipulation tasks centered on three forms of uncertainty: search, manipulation-based inspection, and interactive testing. We evaluate five frontier multimodal agents through a common visuomotor interface, as well as a π_{0.5} policy fine-tuned on the full-episode demonstrations we release. The best agent succeeds in only 23\% of the episodes, and the fine-tuned policy almost never succeeds. Isolated tests of the execution skills the tasks are built from, with the hidden information supplied, show that the agents can carry out most of the required actions, and our failure analysis attributes only a minority of the failures to execution...
dair-ai/AI-Papers-of-the-Week · 10 papers
Context Language Models:让模型自行编辑运行时上下文
Context Language Models(CLM)把 Agent 的实时上下文视为模型可用代码读写的文件,让模型自行决定保留、改写或删除哪些信息,并可定义复用函数、管理子 Agent。该方法无需额外训练即可在 BrowseComp-Plus 上将准确率提高 11.4%,同时减少 21.5% FLOPs;在 EdgeBench 上提高 5% 且减少 59% FLOPs。针对中段编辑破坏前缀缓存的问题,作者设计 Suffix Cache Reuse,在性能相当时将服务端计算再降 35%。实验还表明,自然语言技能优化和在线强化学习都能继续提升上下文管理能力。
为什么值得看它把上下文压缩从固定 harness 规则升级为模型可学习的策略,并同时覆盖准确率、算力和缓存复用。
读原文判断开发长程 Agent、上下文压缩或多 Agent runtime 的团队应细读;应用团队可先关注零样本收益与缓存代价。
Agent harnesses usually manage the model's context through fixed rules such as summarization or compaction. Researchers from Meta and collaborators propose Context Language Models (CLMs), which treat the live context as a file the model edits freely, deciding what to keep, rewrite, or remove. ● Context as a read-write file: The model edits its own context with code, writing loops that prune irrelevant results, defining functions it reuses, and tracking subagents. Recursive language models place a large input in an external variable the model can read, but they do not let it edit its live interaction context. Because several agent contexts can coexist as files, the approach extends to multi-agent systems. ● Works zero-shot: Built from existing models with no training, CLMs reach 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus than state-of-the-art context-management strategies, 5% higher scores with 59% fewer FLOPs on the 12-hour EdgeBench, and 65% more improvement for the same compute on a 24-hour agent-swarm task across six repositories. ● Learnable in context and in weights: Natural-language instructions evolved through a skill-optimization loop raise held-out acc...
JAZ:用一个运行时原语构建记忆与自我改进 Agent
MIT CSAIL 的 JAZ 用单一 LLM 原语 invoke 构建极简 Agent harness:模型在每次调用时现场编写函数实现,并可生成任意可执行代码、递归调用 invoke 形成子 Agent。输入和完整交互历史都作为代码环境变量供模型读取和变换,因此长期记忆与自我改进成为模型在循环内部编写的程序。无需专用记忆系统、文件系统或手工工具,JAZ 在 StuLife 的重召回子集上以约一半成本超过 Letta 8%,在 AppWorld 上也以更低成本超过持续自改进方法 ACE 4%。
为什么值得看它用实验证明极小的可编程表面可以覆盖通常由复杂记忆和自改进基础设施承担的能力。
读原文判断设计自定义 Agent harness 或长期记忆系统的团队应细读,并评估极简原语能否替代专用组件。
Long-term memory and self-improvement are usually built as separate systems around the agent loop. Researchers from MIT CSAIL built JAZ to test how far a minimal harness, little more than the agent loop itself, can go on the tasks those systems target. ● One primitive: JAZ exposes a single LLM-based primitive, invoke, which acts as a function whose implementation the LLM writes at runtime each time it is called. Built-in hooks let the programmer apply constraints and monitor the run. ● Everything is a variable: The LLM can write arbitrary executable code, including recursive calls to invoke, so subagents are the default. All inputs to invoke and the full interaction history live as variables in the code environment, where the model can read and transform them with code. ● Beats specialized systems: With prompting only and no manually designed tools, memory system, or file system, JAZ outperforms Letta (MemGPT) by 8% at half its cost on the recall-heavy portion of StuLife, a long-horizon benchmark that requires recall far beyond the context window. On AppWorld it beats ACE, a continual self-improvement method, by 4% at lower cost. ● Why it matters: Memory and self-improvement become...
Agensh:无中心编排扩展到 1,024 个编码 Agent
Microsoft Research 的 Agensh 取消中心调度器,让每个工作 Agent 异步执行同一协作循环:收集上下文、自主认领子任务、行动并共享发现、验证结果、合并进展。共享工作区、消息接口和可复用上下文支撑组织运行。在 ProgramBench 最难的五项任务上,GPT-5.6-sol 从 1 个扩到 128 个 Agent 后平均测试通过率由 19.31% 升至 28.78%;在 6 小时内从零实现 pandoc 的任务上,1,024 个 Agent 将通过率从单 Agent 的 33.89% 提升至 55.06%。
为什么值得看它把 Agent 数量作为独立扩展维度,并给出超大团队在硬时间预算下的真实增益与协作机制。
读原文判断研究大规模并行编码或低延迟多 Agent 系统的团队应细读;普通任务先评估协调成本和可分解性。
Multi-agent harnesses usually depend on a central orchestrator that assigns tasks and coordinates workers, and that orchestrator limits how many agents the system can use. Microsoft Research introduces Agensh, a self-organized multi-agent harness with no central orchestrator, and scales it to 1,024 coding agents. ● A shared cooperation loop: Every worker runs the same loop asynchronously. It gathers context, claims and self-assigns a sub-task, takes action and shares findings, verifies the result, and merges progress. ● Organization infrastructure: Three components support the loop. A shared workspace holds proposed, ongoing, and completed work, a message interface lets workers communicate, and a shared context retains reusable findings and work intentions. ● Scaling the number of agents: On the five hardest ProgramBench tasks with GPT-5.6-sol, going from 1 to 128 agents raises the mean final test-pass rate from 19.31% to 28.78%, about a 49% relative improvement, and larger teams reach a given pass rate sooner. On building pandoc from scratch under a 6-hour budget, going from 1 to 1,024 agents raises the test-pass rate from 33.89% to 55.06%. ● Why it matters: Worker trajectories sh...
AutoGym:蓝图优先生成可验证的 Agent 强化学习环境
Amazon AGI 的 AutoGym 从少量领域种子或历史轨迹自动生成任务、可执行环境和可靠验证器。系统先定义有效解空间、环境需求与验证标准,在构建阶段解决可解性,因此结果无需依赖不稳定的 LLM Judge。任务拓扑、交互深度、能力轴、问题混淆和干扰项等参数可显式控制难度,并通过基于表现的校准持续调整分布;失败分析还会把反复出现的能力缺口反馈给生成器。实验显示,它能在生产力和时间推理领域持续生成覆盖不同能力层级、可挑战前沿模型的环境。
为什么值得看它提供了让 RL 环境随模型能力持续升级的工程方法,并把可验证性前置到生成蓝图。
读原文判断构建 Agent RL、自动课程或可验证环境的团队应细读;重点关注蓝图约束与难度校准。
Training agents with RL requires a gym, meaning a task, an executable environment to attempt it in, and a verifier that reliably separates success from failure. These gyms are still built by hand, saturate as models improve, and get exposed to contamination. Researchers from Amazon AGI present AutoGym, which generates complete gyms from a minimal domain seed or from prior model trajectories. ● Blueprint first: AutoGym specifies the valid solution space, environment requirements, and verification criteria before it builds the environment. Solvability is settled during construction, so no unreliable LLM judge has to decide correctness afterward. ● Difficulty you can steer: Explicit generation parameters control task topology, interaction depth, capability axes, question obfuscation, and distractor composition. Single-pass synthetic tasks tend to be hard only in their phrasing, and these parameters target actual difficulty. ● Active curriculum: Performance-informed calibration shifts the distribution over those parameters as model capabilities change, and failure analysis feeds recurring capability gaps back into the generation parameters. ● Why it matters: Across productivity and tem...
CASD:让编码 Agent 从整批轨迹中一次性蒸馏提示词技能
Microsoft 的 Coding-Agent Skill Distillation(CASD)让现成编码 Agent 对静态轨迹语料编写并运行分析代码,计算全局统计、发现系统性失败、检查代表样本,再一次性把结论蒸馏成行为规则。它无需环境访问和验证集,能够捕捉小批量反思看不到的跨语料模式。ALFWorld、tau2-bench 和 SpreadsheetBench-Verified 上,一次 CASD 平均将基线提高 16.6 分,超过 GEPA 的 10.9 分和 SkillOpt 的 5.3 分;每份优化提示约 1.60 美元,成本低 22 倍以上。
为什么值得看它说明现有 Agent 日志本身就能成为低成本提示优化数据,并挑战了必须搜索加验证门禁的惯例。
读原文判断已积累大量 Agent 轨迹的团队应优先试验;需要在线探索新行为时再补充搜索式优化。
Search-based prompt optimizers such as GEPA propose edits, run fresh rollouts, score them, and keep only the edits that improve a validation metric. Researchers from Microsoft show that this loop may be unnecessary when you already have a corpus of agent trajectories. ● One pass over the logs: Coding-Agent Skill Distillation (CASD) gives an off-the-shelf coding agent a static corpus of trajectories. The agent writes and runs analysis code to compute corpus-wide statistics, finds systematic failure modes, inspects representative episodes, and distills the findings into behavioral rules in one prompt. It needs no environment access and no validation data. ● Reflection scope: Search-based optimizers reflect on small batches of trajectories at each step, so failures that only show up across the whole corpus stay invisible to them. CASD reads the whole corpus with code, so it can find those failures. ● Better and cheaper: Across ALFWorld, tau2-bench retail and telecom, and SpreadsheetBench-Verified, a single CASD pass improves the unoptimized baseline by 16.6 points on average, against 10.9 for GEPA and 5.3 for SkillOpt. Each optimized prompt costs about $1.60, over 22x cheaper than val...
Taste-Bench:测量长程 Agent 在关键分叉点的判断品味
Taste-Bench 把长程任务中决定后续成败的方向选择定义为 Agent 的“品味”。每道题呈现工程或研究轨迹的关键分叉点,让模型在看不到后续结果时选择更优方向;数据由平行尝试或单轨迹绕路自动挖掘,并过滤仅凭选项可猜或上下文不足的样本。最佳模型准确率只有 59.7%,证据要到后续轨迹才出现的分叉尤其困难,增加推理预算也没有改善。作者进一步把看到结果的教师判断蒸馏给学生,使其在未见任务上决策更好,并提高 SWE-bench Pro 端到端成功率。
为什么值得看它把端到端成功率掩盖的方向选择能力单独量化,并验证这类判断可以训练。
读原文判断开发长程编码、研究或自我改进 Agent 的团队应细读,并把关键分叉纳入评测与训练。
On long-horizon tasks, decisions such as which hypothesis to test or which implementation to build on determine how the whole run turns out. Researchers from Microsoft and collaborators call the ability to make these decisions well an agent's taste, and build Taste-Bench to measure it. ● Decision forks: Each question shows a point in a trajectory where several directions are open and one leads to a better outcome. The model picks a direction without seeing what happens after the fork. ● Mined automatically: Forks come from engineering and research runs, either from parallel attempts at the same task or from detours inside a single trajectory. Filters drop forks that are guessable from the options alone and forks that cannot be decided from the context, so no human annotation is needed. ● Just under 60%: The best model answers 59.7% of questions correctly. Forks whose deciding evidence appears later in the trajectory are much harder for every model, and a larger reasoning budget does not improve accuracy. ● Why it matters: Taste can be trained. The authors distill the judgment of a teacher that has seen the outcome into a student model, which makes better decisions on unseen tasks a...
Jev-Mem:由轻量控制器接管 Agent 记忆操作
Jev-Mem 采用 System-One/System-Two 分工:轻量控制平面负责快速决策,多关系记忆平面保存语义、时间、因果和实体关系,LLM 推理平面只处理复杂推理与最终回答。控制器在写入时分配记忆类型和关系,在读取时负责查询路由、预算分配、图遍历、候选打分和自适应停止。LoCoMo 上其 LLM Judge 总分达到 0.777,相对最强基线提高 11%;记忆构建耗时 158 秒,比最快对手快 6.6 倍,平均查询延迟下降 36.7% 至 0.93 秒。
为什么值得看它展示把高频记忆控制从自回归大模型移出关键路径,能够同时提升质量、构建速度和查询延迟。
读原文判断长期记忆或高频 Agent 服务团队应细读;重点评估控制器与大模型之间的职责边界。
Many agentic memory systems use an autoregressive LLM to decide how memories are organized, retrieved, and used, which puts expensive generation on the critical path of every memory operation. Jev-Mem borrows its design from System-One/System-Two cognition and hands those decisions to a lightweight controller. ● Three planes: A System-One control plane makes the fast decisions, a structured multi-relational memory plane stores semantic, temporal, causal, and entity relations, and a System-Two reasoning plane handles slower, deliberative reasoning. ● The controller runs memory: When storing, the controller assigns memory types and relations. When reading, it handles query routing, retrieval-budget allocation, graph traversal, candidate scoring, and adaptive stopping. The LLM is called only for complex reasoning and writing the final answer. ● Faster and more accurate: On LoCoMo, Jev-Mem scores 0.777 overall with an LLM judge, an 11.0% relative improvement over the strongest baseline. Memory construction takes 158 seconds, 6.6x faster than the fastest competing memory system, and average query latency drops 36.7% to 0.93 seconds. ● Why it matters: Memory latency adds up across every...
Critical-State RL:只训练多轮工具调用中真正可学习的关键状态
Salesforce AI Research 的 Critical-State RL 解决多轮工具交互中奖励变化混入后续随机性的问题。方法通过嵌套采样,把当前动作引起的奖励差异与后续延续噪声分离,识别出最值得训练的一次模型调用,再只对该状态执行上下文赌博机更新。BFCL v4 缺失函数任务上,训练被选中的关键轮次带来约 14 个百分点提升;训练替代轮次则准确率持平或下降。结果表明,多轮 RL 的有效性高度依赖状态选择,而非把所有回合统一纳入更新。
为什么值得看它给出诊断多轮轨迹可训练位置的实用方法,能够减少无效甚至有害的信用分配。
读原文判断训练工具使用或多轮 Agent 的团队应细读;单轮应用可先了解其嵌套采样诊断思路。
Salesforce AI Research's Critical-State RL finds the one call in a multi-turn tool-use interaction where training helps, since reward variation that depends on later turns often reflects downstream randomness instead of the current action. It uses nested sampling to separate action-dependent reward variation from continuation noise, then trains only the selected call with contextual-bandit updates. On BFCL v4 missing-function tasks, training the selected turn adds about 14 points, while training the alternative turn leaves accuracy flat or worse.
SkillGym:把人类编写的 Agent 技能内化进模型权重
SkillGym 将人类编写的 Agent Skills 转换成 12 类、2,756 个带代码验证器的训练环境,并收集 8,364 条成功轨迹用于微调。Qwen3.5-35B-A3B 在 Claude Code harness 下训练后,Terminal-Bench 2.1 提高 19.10 分,在带技能的 SkillsBench v1.1 上提高 28.13 分至 51.47%,超过论文报告的 Claude Sonnet 4.6 和 GPT-5.4 Mini。更关键的是,在推理时不加载技能的情况下,训练后模型仍超过加载技能的基础模型,说明技能行为已被内化。
为什么值得看它把可复用 Skill 从推理时提示资产转化为可验证训练数据,并量化权重内化带来的迁移收益。
读原文判断维护大量 Agent Skills 或做工具能力微调的团队应细读;重点关注验证器质量和轨迹筛选。
SkillGym turns human-written agent skills into 2,756 verifiable training environments across 12 categories, each with code-based checkers, and collects 8,364 successful trajectories for fine-tuning. Under Claude Code, fine-tuning Qwen3.5-35B-A3B adds 19.10 points on Terminal-Bench 2.1 and 28.13 points on skill-assisted SkillsBench v1.1, where it reaches 51.47%, above the reported scores for Claude Sonnet 4.6 and GPT-5.4 Mini. With no skills loaded, the trained model still beats the base model that has the skills in context.
AutoCompact:让编码 Agent 学会何时压缩上下文
AutoCompact 把何时压缩上下文、保留哪些工作状态以及压缩后如何继续执行纳入编码 Agent 自身策略。训练时,Judge 检查基础 Agent 的压缩决策,并在执行前替换错误决策;修正后的轨迹用于监督微调,再以任务成功为目标联合强化学习编码与压缩行为。SWE-bench Verified 通过率提高 9.2 分,SWE-PolyBench Verified 提高 5.0 分。收益在不会溢出的 256K 上下文和需要强制压缩兜底的 16K 窗口中都成立,说明压缩不仅用于防止超窗,也能改善任务执行。
为什么值得看它把压缩从外部固定阈值变成与任务成功联合优化的决策,并证明大上下文场景同样受益。
读原文判断开发长程编码 Agent 或上下文管理策略的团队应细读;可先对比现有固定压缩规则的错误类型。
AutoCompact trains a coding agent to decide when to compact its context, what working state to keep, and how to continue afterward, as part of its own policy. A judge reviews the base agent's compaction decisions and replaces flawed ones before they execute, and the corrected trajectories are used for SFT and then for RL that optimizes coding and compaction together on task success. Pass rates rise by 9.2 points on SWE-bench Verified and 5.0 points on SWE-PolyBench Verified, and the gains hold both with a 256K window that never overflows and with a 16K window that falls back to forced compaction.
产品方法论 / 增长案例 · 2 updates
这条来自 Lenny's Newsletter,主题偏向「产品与 AI 趋势」。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合沉淀产品方法论、增长案例和 PM/Founder 可复用做法。
可转化可以摘出 1 个观点,作为今天信息流里的可展开选题。
Your weekly listens from How I AI, part of the Lenny’s Podcast Network
这条来自 Lenny's Newsletter,主题偏向「AI 组织与岗位变化」。核心议题是产品、增长、定位和设计流程。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合沉淀产品方法论、增长案例和 PM/Founder 可复用做法。
可转化可以转成产品判断、增长案例复盘或创始人内容选题。
Watch now (36 mins) | ️ Kath Korevec demos ChatGPT Sites live at OpenAI DevDay, showing how to turn Slack, Notion, and calendar data into internal tools that adapt to each person on your team
AI 工程 / agent / 模型基础设施 · 2 updates
这条来自 Latent.Space,主题偏向「产品与 AI 趋势」。核心议题是AI 时代产品/工程/设计角色重构。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合跟踪 AI 工程师圈对 agent、模型基础设施和开发范式的判断。
可转化可以转成“AI 后产品/设计/工程岗位到底怎么变”的观点或互动问题。
A special Science pod and Engineering pod crossover.. with Forward Deployed Engineering kicker!
这条来自 Latent.Space,主题偏向「Agent / 工作流」。核心议题是agent 工作流、harness 和自动化循环。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合跟踪 AI 工程师圈对 agent、模型基础设施和开发范式的判断。
可转化可以沉淀成 agent 产品设计、工作流拆解或本地工具方向。
Kubernetes co-creators Craig McLuckie and Joe Beda aim to bring agent harnesses fully into the cloud.
AI 创业 / 投资 / 产业判断 · 2 updates
这条来自 No Priors,主题偏向「AI 组织与岗位变化」。核心议题是模型、推理、评测与 AI infra、创业、市场和投资判断。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合观察 AI 创业、投资人和一线 founder 对市场节奏的判断。
可转化可以用来更新赛道判断和要观察的公司/产品清单。
Founder and CEO of full-stack AI chip company Fractile, Walter Goodwin, joins Sarah Guo to discuss the bets he’s made on the future of the chip market as other major players like Broadcom, NVIDIA, and AMD try to accelerate their workloads. They discuss the difference in Fractile’s newer approach on model architecture with a full-stack team in the current chip landscape and the technical bets they’re making in that direction. Walter also talks about compressing the gap between the chip design cycle and its payoff period, and making a generational leap in AI inference to realize the bet in volume against the value to be captured. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @goodwin_ml | @fractile_ai
这条来自 No Priors,主题偏向「模型 / AI Infra」。核心议题是AI 时代产品/工程/设计角色重构、模型、推理、评测与 AI infra、创业、市场和投资判断。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合观察 AI 创业、投资人和一线 founder 对市场节奏的判断。
可转化可以转成“AI 后产品/设计/工程岗位到底怎么变”的观点或互动问题。
Can AI transform legacy incumbents rather than replacing them? Sequence Holdings co-founder and CEO Michael Lee joins Sarah Guo to discuss how holding company structures and engineering integrations are reshaping market leaders from the inside out. Michael details Sequence’s $7.7 billion take-private transaction of Baldwin alongside Dell Family Office (DFO), and shares his thesis on why traditional consulting models and software sales fall short for real enterprise AI transformations. They also talk about why permanent holding company structures are good for long-term compounding, real-world results from applying frontier engineering to BankSouth, and Michael’s lessons from his time in public investing, private equity, and operating at the intersection of market incumbents and AI. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @mjlee_2014 | @seqholdings
AI builders / research / safety · 2 updates
这条来自 The Cognitive Revolution,主题偏向「Agent / 工作流」。核心议题是agent 工作流、harness 和自动化循环、AI agent 时代的信息检索和知识层、模型、推理、评测与 AI infra、创业、市场和投资判断。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合补齐 AI 研究、安全、长期影响和 builder 深访视角。
可转化可以沉淀成 agent 产品设计、工作流拆解或本地工具方向。
Nathan Labenz reports back from The Curve conference, sharing off-the-record insights from frontier lab leaders on shortened timelines and whether humanity is approaching an intelligence threshold it should not cross. The episode also features discussions with Positron CTO Thomas Sohmers on overcoming the memory wall while token spend eclipses human salaries, and swyx on how coding agents are disrupting enterprise software. In addition, Atlas Ignota's Evan Miyazono examines unowned risks from accidental agent swarms hitting critical infrastructure, while Mercor's Edward Hu explores the future of training data. Together, these conversations highlight how escalating compute demands and autonomous agent workflows are challenging long-held assumptions about hardware, software economics, and safety limits. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/ai-am-a-level-we-shouldn-t-pass-notes-from-the-curve-tokens-vs-salaries-is-saas-cooked/
这条来自 The Cognitive Revolution,主题偏向「AI 组织与岗位变化」。核心议题是AI 时代产品/工程/设计角色重构、agent 工作流、harness 和自动化循环、模型、推理、评测与 AI infra、创业、市场和投资判断。建议把它当成观点/框架源来读:重点看它如何定义问题、角色变化、系统设计或市场节奏,而不是只看是否有新功能发布。
为什么值得看适合补齐 AI 研究、安全、长期影响和 builder 深访视角。
可转化可以转成“AI 后产品/设计/工程岗位到底怎么变”的观点或互动问题。
OutSystems CEO Woodson Martin joins Nathan to discuss how enterprise software can evolve at AI speed without breaking critical operations. Martin explains how an intermediate layer of abstract modeling allows AI agents to build reliably through deterministic code generation, ensuring role-based security across heavily regulated industries. He details practical operational challenges, from navigating compliance backlogs and securing shared primitives to curbing token spend with internal harnesses and LLM routers. Finally, they explore why incumbent platforms with deep architectural trust and industry specialization hold a distinct edge over AI-native startups. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/software-that-never-breaks-outsystems-ceo-woodson-martin-on-building-enterprise-grade-apps-at-ai-speed/
Claude @anthropicai
Ok back to reading https://t.co/Kdph4vZuGt
It was one prompt, claude verified pretty well too! No bugs so far https://t.co/jjnUzGkdoC
Try it https://t.co/isXGOnEXd5
Codex & ChatGPT @OpenAI
Day 3 (encore)/ We silently re-shipped codex cloud. It’s pretty good now https://t.co/Fau14sCke8
Confirmed landed across all accounts. How are we doing so far? https://t.co/tzw7dNX4pN
Will be there by EOD PST.
Practical AI tutorials and interviews for busy people | Get my best AI skills and guides at https://t.co/6VAA6p81x6
.@suno is so damn good honestly. It's probably going to overtake Spotify at some point.
AI already solved images, music, and video. It’s going to solve gaming soon
It seems like anything you ship can now be decompiled and rebuilt by AI. I think it'll soon be incredibly hard to make money from software unless you have proprietary data, distribution, or some other edge. Traditional SaaS is especially vulnerable, since it's expensive and full of features built for humans that agents don't need.
product, codex @openai. prev head of product @linear
Dread it. Run from it. The One True Productivity App Layout arrives all the same. https://t.co/fczbE1PHEk
claude code + cowork @anthropicai, prev: @dagster, @scale_ai
One of my favorite PM use cases for Claude is asking "who used <feature> the most last week? make me a artifact of the top 10 by usage, then reach out and schedule 15 min to chat." It's the fastest way to get user feedback! https://t.co/SYL4JOitsB
Claude Code @anthropicai. prev YC W20, @spc, @medialab
and because agents make it easier, almost everyone is operating outside their expertise at some point but luckily, you can just ask the model to teach you what you don't know
the most common failure case I see is when people are working outside their domain of expertise and don't know how to be precise with their prompts and plans, so they have to spend a lot of turns iterating imprecisely https://t.co/mRT3knojCP
I think it's probably an incredible time to be a game dev content creator, for many people making games is a form of having fun and they're willing to pay for it! https://t.co/iS08bGZZnJ
Google’s home for our latest AI tools and experiments.
🚨NEW EXPERIMENT 🚨 Playground is an experimental gaming platform that lets you create your own games with zero coding experience. If you can think it, you can play it. Go to https://t.co/fRUukpuoqM to learn more! Available to users 18+ in the US. https://t.co/p0X8JR7JMn
ceo @replit. civilizationist
Math was bound to fall first. The purer the field the easier it is for AI to crack. https://t.co/OETDJ1JSIz
That’s because Replit is not a consumer app. Most of our creators are building businesses or working on one. https://t.co/B70WFQldgU
Desktop AI apps are great but expose users to supply-chain attacks and catastrophic mistakes by agents. We’re building a powerful desktop experience with focus on security & reliability. Excited to partner with @Microsoft on this and be an early adopter of @nvidia’s OpenShell. https://t.co/YixaxIXhpb
@vercel CEO
One thing the engineering community will (re)discover is that every program can be ① hardened (edge cases squashed, inputs tightened, errors handled) and ② optimized (profiled, benchmarked, rewritten) basically ad infinitum. There are real costs in every direction: time, attention, opportunity. You can literally explore input spaces that will never be hit or optimize performance of something that will never be used ("you have no users".) Even in an infinite token world, you need to know when to stop and what tradeoffs to accept. Agents will happily drill no matter what.
gm from the best city in the world https://t.co/aPDFNR7KhT
Research @AnthropicAI. Opinions are my own!
For reference, Haiku 4.5 came out on October 15th, 2025 so these two columns are less than a year apart. and oh yeah Haiku 5.5 is also much faster and 75% cheaper... https://t.co/1Zt19kdTYC
ceo @box - your business lives in content. unleash it with AI
The compute needed for the stage of AI we’re about to enter is going to be insane. Personal agents, agent swarms defending enterprises, agents that review all of our code for security issues, agents that process nearly all enterprise data in workflows, background agents working on tasks 24/7, and on and on. This will take a mix of inference volume that is going to keep growing by orders of magnitude, as well as infrastructure that agents need alongside the tokens. They need computers, networking, file systems, and so on. We’re only in the early stages of what this buildout is going to look like.
VC at @FirstMarkCap. Host: MAD Podcast; Organizer: Data Driven NYC, Author: MAD Landscape.
Fun fact: Ben Affleck, the successful AI founder and Python/CNN/GPU expert, had a prior career in the entertainment industry. https://t.co/BZwIs5mQhU
Ofc everything is on X is mind-blowing, but having seen the beta version, this is... truly mind-blowing. https://t.co/wdySOkrsUd
Builder. Make something people want, then make people want it. Harvard’17. GitHub: https://t.co/KCuEajezlL YouTube: https://t.co/8xzbGWtf6w
I love @stripepress because I absolutely judge a book by its cover https://t.co/YMBUiprxUX
Who has the coolest personal website?
Remember that coding is a means to an end. The end matters more than the means
partner @fpvventures - investing in seed/A. previous: early hire @meter, @opendoor, @atlassian & others. love @shimoleejhaveri + 👦👧
Moving @askalphaxiv to the dock is the best thing I have done this week.. Went from 0 to ~30 mins of reading papers on my phone whenever I’m waiting on things. https://t.co/aPwn0LgfcY
To all the people visiting SF for tech week.. 1) the best builders are probably at their own offices. so DM them and try to meet them there. 2) don't be a tourist https://t.co/3YGQQjuyGD
https://t.co/gn6cYt7bqh
Polyagentmorous ClawFather. Came back from retirement to mess with AI and help a lobster take over the world. @OpenClaw🦞 + @OpenAI
While everyone's talking about agents, I've been exploring how teams can use them to work better together. Here’s my take, from OpenAI DevDay 2026. https://t.co/W3rVoiTo84
ceo @every | the only subscription you need to stay at the edge of AI
tfw your ChatGPT app is controlling your Claude app is controlling your ChatGPT app https://t.co/i3QO9tLFSB
OpenAI’s Dots launched last week and they've taken over @every. A few things that stand out: They protect your attention. Over the weekend, my Dot posted my model-testing results to Slack and brought back replies. I could keep working without opening Slack and disappearing down a rabbit hole. My weekend stayed pretty clear and calm. They’re useful for the little things you miss. People at Every are using them to pull important information out of school emails, catch missed Slack messages, and flag account-access requests. Voice makes them useful away from your laptop. Our head of video was talking his Dot through video edits while running around getting our launch out. The ecosystem matters. Dots feel roughly at feature parity with Muse, Instinct, and Grok bots. But having an agent connected to the ecosystem you already use for work and life is a real advantage. If you use ChatGPT, I recommend trying one. Permissions are still frustrating. I have to approve the same basic thing multiple times, and sometimes even after I approve it, my Dot says it can’t do it. That needs to get fixed. I’m less convinced we’ll stay loyal to the characters. Mine is named Boo, and I’m attached to him. But I felt that way about my OpenClaw agent, too. When something better comes along, I switch. In a year, I expect to be using a persistent agent. Boo will probably be dead. RIP, Boo. I get into all of this at the start of this week’s podcast, and then sit down with @bigwilliestyle to talk about where we go from here. In a lot of ways, Dots and GrokBots remind me of a wave we went through with OpenClaw eight months ago. We talk about personal vs. company agents, why we built our own company agent, and whether companies will end up with a whole shadow org chart of agents. Full episode below.
Love to see it!! https://t.co/jmPN04jKZz
General Partner @SPC, Co-Founder @Bevel_Health | Ex: Early Eng @facebook, CTO @Dropbox, Board @Flipkart | Optimist, Builder, Dad
My first time in a @Waymo felt like a religious experience. Co-CEO @dmitri_dolgov and I talked about it. In a Waymo. Full conversation at @spc releasing tomorrow. https://t.co/CdwA6yI5n6
What should an ambitious software engineer do today? Ten years ago, being a great software engineer gave you an unusual advantage: a rare skill you could apply to almost any industry. You could build a payments company, improve how hospitals operated, or create a new way for people to communicate. The same underlying ability opened doors into wildly different businesses. Software was eating the world, and knowing how to build it gave you a seat at the head of almost every table. “Get really good at software engineering” was remarkably versatile advice. You could acquire the skill first and figure out where to apply it later. AI is making the ability to produce working software much more widely available. People who couldn’t have built a product a few years ago can now get surprisingly far. Experienced engineers can attempt projects that previously required a team. Great engineering is still rare. But “I can build it” is becoming a weaker answer to why you, specifically, have an advantage. The way I see it, there are three paths. 1. Go work on the thing that’s eating software Help create the technology that is making software engineering more accessible. That could mean AI research, or the infrastructure, systems, and tools needed to make new capabilities useful. Find problems where technical ability remains a serious constraint, and where solving them expands what other people can do. 2. Find a world-class domain expert and team up Someone who has spent fifteen years in an industry may know where the expensive mistakes happen, why previous attempts to fix them failed, and who would buy a solution. That knowledge is hard to acquire from the outside. At SPC, we’re seeing more people become promising founders without much prior software engineering experience. They bring deep domain knowledge and strong instincts about what needs to exist. These are people worth getting to know. A great engineering partner can see possibilities they haven’t considered and help build something neither person could have imagined alone. 3. Write vanilla code… by shepherding a bunch of agents This is the direction I find least compelling: doing roughly the same work inside a company, except now you supervise coding agents. There will be valuable jobs like this. But if your ambition extends only to completing the same backlog faster, you’re underusing the moment. I’d still get very good at engineering. Then I’d be deliberate about what that ability sits next to: a technical frontier that needs advancing, or deep knowledge of a problem worth solving. Use the fact that you can build more to take responsibility for something bigger.
The mission of OpenAI is to ensure that AGI benefits all of humanity
Have been waiting for this one for a long time, and I would hate to have to go back to the old version of Chat!
ChatGPT can now generate a custom UI for you! https://t.co/ZdGRMAEjQX
Claude is an AI assistant built by @anthropicai to be safe, accurate, and secure. Talk to Claude on https://t.co/ZhTwG8d1e5 or download the app.
One more thing: we’re halving the price of cache reads on Claude Sonnet 5.5, to $0.10 per million tokens. That makes Sonnet 5.5 around 20% cheaper to run on most long-running work.
Haiku 5.5 is available now on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. Read more: https://t.co/InzZXcEcRP
Haiku 5.5 shows major improvements across almost all of our alignment evaluations relative to Haiku 4.5, with far fewer instances of misaligned behavior.
Why Every Traded Personal Agents for One Company Agent
No new blog posts in the current feed.
updated Thu, 8 Oct 2026 18:12:14 +0000
Beijing Chaoxing Digital Library Information Technology Co., Ltd. · 2016-01-18
Open App Storeupdated Thu, 8 Oct 2026 18:12:15 +0000
updated Thu, 8 Oct 2026 18:12:16 +0000
public Google Play chart · US
public Google Play chart · US
public Google Play chart · HK
public Google Play chart · JP
The open source Semrush alternative
A swarm of agents, in the same place you work
Get SOC 2 without the sales call
World's First Human Interaction Model
Native Mac mail for IMAP, JMAP and Microsoft 365
Bring Claude directly into Google Docs, Sheets & Slides
The personal AI that handles life before you have to
A private language tutor that knows you better every day.
Speech-to-speech model benchmarks on live phone calls
Your AI agent's phone to dial, hold, and reports back
A Jitsi alternative that does a few things really well
Strap-on motorized wheels that let you walk twice as fast
Scan AI agent Skills for risks before you install them
Anthropic's fastest and most capable Haiku yet
Orchestrate AI agent teams on a visual canvas
A shared Markdown editor for people and their agents
Stable localhost URLs over SSH with no installs
Manage your local coding agents from your phone
Translate in your iPhone keyboard. Keep your tone.
See and answer every coding agent from your notch
Autopilot for long-running AI sessions in terminal
Smart personal safety app connected to 24/7 dispatch
A tiny desktop pet for your Mac that types along with you
An AI on-call engineer for your Vercel apps
Lowweight minimal dynamic island for Windows