-
1. KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
arXiv Recent AI/ML
2 天前 · 2026/10/02 01:59
来自 arXiv Recent AI/ML,主题偏「Agent、Post-training/Alignment、Eval/Data」。摘要显示它主要讨论 LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure L... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
AgentPost-training/AlignmentEval/DataAI Infra
Original abstract
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-p...
-
2. Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
arXiv Recent AI/ML
2 天前 · 2026/10/02 01:59
来自 arXiv Recent AI/ML,主题偏「Agent、Multimodal、Eval/Data」。摘要显示它主要讨论 Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous impro... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
AgentMultimodalEval/DataRobotics
Original abstract
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/
-
3. ROWBench: Do Video Models Render What the Program Specifies?
arXiv Recent AI/ML
2 天前 · 2026/10/02 01:59
来自 arXiv Recent AI/ML,主题偏「RAG/Memory、Multimodal、Post-training/Alignment」。摘要显示它主要讨论 Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and interactions remains insufficiently evaluated. Existing bench... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察知识工作、企业搜索、长期记忆和本地资料库产品的新实现路径。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
RAG/MemoryMultimodalPost-training/AlignmentEval/DataRobotics
Original abstract
Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and interactions remains insufficiently evaluated. Existing benchmarks assess visual quality, controllability, and instruction or physical adherence, but rarely test fidelity to fine-grained, program-specified world events. We introduce PROWBench, comprising 170 programmatically constructed episodes and 600 proxy videos covering diverse scenes and interactions. PROWBench logs entity states and timestamped events, including those outside the camera's field of view, as replayable world records, from which it renders synchronized views and proxy representations. This enables generated videos to be checked against the observable consequences of program execution. An extensible framework constructs scenes, controls behaviors, and can render each camera view in different representations, such as coarse 3D, and bounding boxes. The benchmark covers first- and third-person perspectives, with synchronized multi-view observations available for a subset of episodes. Grounded in these records, PROWBench evaluates entity control, long-horizon memory, and, with two VLM-based metrics, Logic-Render Alignment and Interaction Success Rate, adherence to the prescribed timeline and the visual realization of...
-
4. ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
arXiv Recent AI/ML
2 天前 · 2026/10/02 01:59
来自 arXiv Recent AI/ML,主题偏「Agent、RAG/Memory、Eval/Data」。摘要显示它主要讨论 What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
AgentRAG/MemoryEval/DataAI InfraAI4Science
Original abstract
What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author annotation scalable, we build ScholarCatalyst by having 184 lead authors of 207 recent computer science papers label which candidates did or could have advanced their project, each with a detailed rationale. We introduce a retrieval task with author-provided judgments: given an initial research question, retrieve these papers from only the literature available when the project began. Agentic search does no better than embedding retrieval (0.42 vs. 0.48 Recall@20) despite calling that same retriever as a tool. Even an agent built on Claude Fable 5.1, which may have seen the completed papers during training, reaches only 0.51 R@20. These results highlight the need for new training recipes that equip models with expert intuition for searching broad corpora. We envision ScholarCatalyst as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs.
-
5. VISTA: A Visual Harness for Reasoning in an Interactive World
arXiv Recent AI/ML
2 天前 · 2026/10/02 01:59
来自 arXiv Recent AI/ML,主题偏「Agent、RAG/Memory、Reasoning」。摘要显示它主要讨论 We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal mo... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察 agentic RL、工具调用、工作流自动化或软件代理能力是否出现新方法。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
AgentRAG/MemoryReasoningMultimodalEval/Data
Original abstract
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.
-
6. One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars
arXiv Recent AI/ML
2 天前 · 2026/10/02 01:59
来自 arXiv Recent AI/ML,主题偏「RAG/Memory、AI Infra、Code」。摘要显示它主要讨论 3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combin... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察知识工作、企业搜索、长期记忆和本地资料库产品的新实现路径。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
RAG/MemoryAI InfraCode
Original abstract
3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this finding, we introduce GALA (Gaussian Animation via Linear Approximation), a distillation method that replaces per-frame heavy neural decoding with a shallow coefficient predictor and a linear blend. To improve fidelity and reduce memory requirements, we propose to construct the basis using block-local PCA under a rendering-aware metric and a memory budget. Our method learns a shallow MLP network to predict blendshape coefficients and applies to various animation architectures without retraining original models. We validate GALA by accelerating the inference of three distinct avatar models for 3D animation of facial expressions and full-bodies with clothing dynamics. Across these models, our distillation generalizes to held-out identities and reduces CPU animation cost by up to three orders of magnitude while preserving most of the rendering quality. Excellent results of our method confirm the shared linear structure of learned avatar representations and enable highly efficient and accurate animation at frame rates reaching up to 60fps on mobile devices. Project pa...
-
7. TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
arXiv Recent AI/ML
2 天前 · 2026/10/02 01:59
来自 arXiv Recent AI/ML,主题偏「RAG/Memory」。摘要显示它主要讨论 Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abandon first-order gradients, or change the u... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察知识工作、企业搜索、长期记忆和本地资料库产品的新实现路径。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
RAG/Memory
Original abstract
Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abandon first-order gradients, or change the update geometry while retaining dense state. The recently introduced Muon optimizer reduces optimizer memory through matrix-valued updates. Still, its geometry differs from AdamW and can lead to performance degradation when fine-tuning AdamW-pretrained models. To reduce optimizer memory without sacrificing accuracy or computational efficiency in LLM fine-tuning, we propose Ternary Absolute-max Column-wise One-sparse optimizer, or TACO, which follows Muon's operator-norm steepest-descent view but takes the geometric route further. TACO computes the exact steepest-descent direction under a dimension-normalized $1\to1$ operator norm by selecting the sign of the largest magnitude entry in each column of two-dimensional weight matrices. This retains first-order gradients while making optimizer state memory nearly negligible. Our practical TACO optimizer maintains only a small set of low precision gradient components per column, reducing persistent optimizer state by $174\times$ relative to AdamW8bit (from 27.7 GB to 0.16 GB) and peak training memory by $2.9\times$ (from 80.6 GB to 27.5 GB) on OPT-13B, while achieving comparable...
-
8. SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
arXiv Recent AI/ML
2 天前 · 2026/10/02 01:59
来自 arXiv Recent AI/ML,主题偏「RAG/Memory、AI Infra、Code」。摘要显示它主要讨论 High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates gen... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察知识工作、企业搜索、长期记忆和本地资料库产品的新实现路径。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
RAG/MemoryAI InfraCode
Original abstract
High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes. We introduce SILSA, a topology-aware 3D generation framework that represents shapes with compact sliding-window slice latents. Instead of generating expensive voxel tokens, SILSA uses a fixed set of overlapping slices along the three canonical axes, where each token summarizes a local depth window to preserve cross-sectional continuity and support single-stage rectified-flow generation. A Slice VAE encodes oriented surface samples into multi-axis slice latents and reconstructs them with a sparse volumetric decoder, while a Volumetric Anchor Lattice coordinates directional slice streams through a shared 3D workspace. To preserve structural correctness, we introduce slice-level topology supervision that matches persistence diagrams and aligns Betti transitions across neighboring slices. Experiments show that SILSA improves structural fidelity while substantially reducing generation cost. SILSA improves PSNR by $8.7\%$, coverage by $5.96$ absolute points, and Betti error by $9.2\%$ over the strongest baseline, while using $70.0\%$ few...
-
9. FERPO: Forward Entropy-Regularized Policy Optimization
arXiv Recent AI/ML
2 天前 · 2026/10/02 01:59
来自 arXiv Recent AI/ML,主题偏「RAG/Memory、Post-training/Alignment、Eval/Data」。摘要显示它主要讨论 Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily y... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察知识工作、企业搜索、长期记忆和本地资料库产品的新实现路径。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
RAG/MemoryPost-training/AlignmentEval/DataRobotics
Original abstract
Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the actor to this target by minimizing a forward-KL objective, estimated using self-normalized importance sampling (SNIS) with actions drawn from the rollout policy. By limiting the target distribution's deviation from the rollout policy, the KL regularization helps keep these importance weights well behaved. In contrast to reverse-KL objectives, which can favor a subset of the target distribution's modes, the forward-KL objective encourages coverage of multiple high-value modes and thereby promotes exploration. Experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains. Computational b...
-
10. Moore, Escher, Penrose: A Conformal Golden Braid
arXiv Recent AI/ML
2 天前 · 2026/10/02 01:59
来自 arXiv Recent AI/ML,主题偏「Reasoning、Multimodal」。摘要显示它主要讨论 I don't think I have ever done anything as peculiar in my life. Among other things, it shows a young man looking with interest at a print on the wall of an exhibition that features himself. How can this be? Perhaps I am not far removed from Einstein's curved u... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察模型推理、规划、验证器和复杂任务能力是否有可复用技术路线。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
ReasoningMultimodal
Original abstract
I don't think I have ever done anything as peculiar in my life. Among other things, it shows a young man looking with interest at a print on the wall of an exhibition that features himself. How can this be? Perhaps I am not far removed from Einstein's curved universe.'' So wrote M.C. Escher about his 1956 lithograph Print Gallery. Nearly half a century later, a mathematical analysis related its geometry to an untwisted source image through a conformal power map $z \mapsto z^α$, $α\in \mathbb{C}$. Building on this construction, we use a frozen text-to-image diffusion model to generate new self-referential scenes. Prompting alone does not enforce the recursion, while a post-hoc transformation can leave structures poorly connected. Applying the transformation during sampling is also insufficient: the denoiser may "repair" the intended distortion or drift out of the prescribed geometry. We construct a generalized inverse $T^\dagger$ of the non-invertible image transformation $T$, adapted to its recursive constraint. In the idealized formulation, the Penrose identity $TT^\dagger T = T$ makes $TT^\dagger$ an idempotent projection onto geometrically admissible images. Yet denoising only the transformed image remains an out-of-distribution task, even with projection. We therefore braid denoising steps with $T$ and $T^\dagger$: source-space steps develop the untwisted scene, while trans...
-
11. Sphere Encoder 2
arXiv Recent AI/ML
2 天前 · 2026/10/02 01:59
来自 arXiv Recent AI/ML,主题偏「RAG/Memory、Multimodal、Code」。摘要显示它主要讨论 Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of the original formulation that reduce its generation quality. First, random points concentrate near the equato... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察知识工作、企业搜索、长期记忆和本地资料库产品的新实现路径。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
RAG/MemoryMultimodalCode
Original abstract
Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of the original formulation that reduce its generation quality. First, random points concentrate near the equator relative to the pole on an encoded latent, but the training rotation never reaches this region, leaving a gap that limits one-step generation. Second, training for generation with pixel-wise reconstruction loss encourages the decoder to average over plausible images, producing blurry images that lack high-frequency details. We present Sphere Encoder 2 to address both limitations, substantially improving image generation quality while maintaining the speed and simplicity of a autoencoder. Models are released at \href{https://github.com/kaiyuyue/sphere2}{github.com/kaiyuyue/sphere2}.
-
12. Embedding Prediction Helps Image Generation
arXiv Recent AI/ML
2 天前 · 2026/10/02 01:59
来自 arXiv Recent AI/ML,主题偏「Multimodal」。摘要显示它主要讨论 In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) train... 先把它当作时效信号看:判断它是否正在影响 agent、RAG、多模态、post-training、评测或 AI infra 的产品/研究方向。
为什么值得看适合观察图像、视频、语音和科学多模态任务是否出现新的产品能力边界。
读原文判断如果标题正好贴近当前产品方向,值得点开原文看方法和实验设置;泛读时先存为观察项。
Multimodal
Original abstract
In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times256$ study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.