Research

Latest Research

A running feed of newly published AI research papers, pulled daily from arXiv and filtered by hand before it appears here.

Papers

50 items
Keyword
01Coding Agents Aren't Enough! Evaluating an Enterprise Security Brain for Agentic Cloud InvestigationsarXiv cs.AIarXiv:2609.30345v2 Announce Type: cross Abstract: Cloud-security investigation is dominated by population tasks: which identities can read a data store, how many resources fail a control, what is reachable from another account. These resolve against a complete inventory, not a named object, so a partial answer to one is not a partial result but a different one. Coding agents can now be given read-only cloud credentials and asked to investigate directly, which raises the question of what a purpo29 Sept 2026↗02Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and TrainingarXiv cs.AIarXiv:2609.25510v2 Announce Type: replace-cross Abstract: Large language models (LLMs) can improve solutions to verifiable scientific and algorithmic problems by spending additional computation at test time. Recent systems achieve strong results with increasingly elaborate evolutionary search harnesses or by updating model parameters during test-time training. We ask how much of this machinery is necessary. We introduce Hill Sampling, a form of hill-climbing optimization that repeatedly samples28 Sept 2026↗03You've Seen Enough: Quality-Constrained Image Coding for MachinesarXiv cs.AIarXiv:2609.25108v2 Announce Type: replace-cross Abstract: Visual data is increasingly consumed by machine-vision systems rather than by human observers. Image Coding for Machines (ICM) compresses images assuming the main observer is a computer vision application and that the human observer needs to inspect or validate the decisions. Inspired by just-noticeable distortion, we cap human-observed quality at a desired level and devote the remaining bits to machine performance. Specifically, joint c28 Sept 2026↗04Provably Safe Sim-to-Real TransferarXiv cs.AIarXiv:2609.01418v2 Announce Type: replace-cross Abstract: We address safe sim-to-real transfer, in which an agent leverages an imperfect simulator and limited real-world interaction while ensuring safety throughout data collection in the real system. This problem arises in applications such as robotics and healthcare: simulators provide cheap data, but sim-to-real mismatch makes direct transfer unreliable, and collecting real-world data to correct this mismatch must itself be safe. Moreover, de28 Sept 2026↗05SUN: Agentic Robot Policy Learning with Persistent Task ProgramsarXiv cs.AIarXiv:2608.31167v2 Announce Type: replace-cross Abstract: Model-based control can directly execute specified objectives, while learning can amortize such behaviors into reactive policies, making their combination a natural solution to multi-stage manipulation. We introduce Semantically UNified (SUN) Programs, typed executables that compile grounded relations into aligned optimal control objectives, satisfaction predicates, and learning rewards. Our harness, Kuafu, equips a foundation model as a28 Sept 2026↗06Keep the Future, Drop the Rollout: RIFT for World Action ModelsarXiv cs.AIarXiv:2608.11521v3 Announce Type: replace-cross Abstract: World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evolving rollout trajectory or only its future representation. Across four WAMs on 40 simulated robotic manipulation tasks, paired closed-loop interventions show that blocking access to the future cache or reassigning its values changes execution and reduces succe28 Sept 2026↗07Beyond Forecasting: Recasting Volatility Control as a Routing ProblemarXiv cs.AIarXiv:2608.10375v2 Announce Type: replace-cross Abstract: Volatility control converts risk estimates into portfolio exposure, yet existing approaches often rely on a fixed volatility estimator or a pre-defined control rule that may not adapt to changing market conditions. We propose VolRouter, a modular framework that formulates volatility control as state-conditioned routing over estimator-controller pairs. VolRouter first summarizes market conditions into a control-relevant state profile and 28 Sept 2026↗08Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action ModelsarXiv cs.AIarXiv:2608.06994v2 Announce Type: replace-cross Abstract: World Action Models (WAMs) aim to construct a unified architecture capable of understanding world state evolution and guiding to generative motion planning. However, existing visual branches focus on predicting static visual observation, rather than reflecting potential transition information that captures the evolution of world states under motion interactions. This leads to representational entanglement between high-level physical cond28 Sept 2026↗09Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA AdaptersarXiv cs.AIarXiv:2607.29238v2 Announce Type: replace-cross Abstract: InMyStyle is a privacy-first, single-user system that adapts small language models to rewrite AI-edited text towards an individual user's writing style without an instruction prompt at inference. Given a user's documents, it uses multiple local helper LLMs to construct paired training examples and fine-tunes LoRA adapters on Qwen2.5 models ranging from 0.5B to 7B parameters. Length-aware generation budgets and automatic chunking support 28 Sept 2026↗10Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM AgentsarXiv cs.AIarXiv:2607.26865v3 Announce Type: replace-cross Abstract: LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. Yet, when deployed at the edge, they must tightly manage their reasoning budget while remaining reliable and deferring to a cloud-side model only when local uncertainty is too high to act safely. We propose Think Short, Defer Smart (TSDS), a framework that28 Sept 2026↗11CoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned NavigationarXiv cs.AIarXiv:2607.02222v2 Announce Type: replace-cross Abstract: Vision-Language Navigation has increasingly emphasized high-level instruction reasoning, memory, global map construction, and instruction decomposition, while the low-level action representation remains comparatively underexplored. We propose CoFL-S, a low-level vision-language-action framework that predicts a language-conditioned flow field over the robot's local visible sector and generates continuous trajectories by rolling out the pr28 Sept 2026↗12Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) BenchmarkarXiv cs.AIarXiv:2606.30170v2 Announce Type: replace-cross Abstract: Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical datasets. This combination yields strong benchmark metrics but limits transferability to domains structurally distinct from drug discovery. To overcome this limitation and drive discovery toward real, scientifically grounded targets, we introduce the Nanotechnology Molecular Optimization (NMO) Benchmark,28 Sept 2026↗13Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank ExpertsarXiv cs.AIarXiv:2606.14929v2 Announce Type: replace-cross Abstract: Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models. Despite its practical significance, this problem remains poorly understood under realistic conditions like adversarial queries, bandit feedback, and limited observability of models. We formalize embedding model routing as an adversarial contextual linear bandit with low-rank experts, where contexts are queries, actions are28 Sept 2026↗14Genetic Algorithms with Optimization Guided OperatorsarXiv cs.AIarXiv:2606.12279v2 Announce Type: replace-cross Abstract: Recent work in ML applies genetic algorithms at inference time to iteratively improve solutions to optimization problems. The basic mutation and recombination operators involved are qualitatively different from those studied classically. Mutations are no longer random; an ML algorithm mutates a solution with the goal of improving an objective. Similarly, recombination is not based on random collages of parent solutions. Instead, it is an28 Sept 2026↗15Grid-Orch: An LLM-Powered Orchestrator for Distribution Grid Simulation and AnalyticsarXiv cs.AIarXiv:2605.12728v2 Announce Type: replace-cross Abstract: The power distribution engineering workforce faces a projected shortage of up to 1.5 million engineers by 2030, creating urgent demand for more accessible analysis tools. This paper introduces Grid-Orch, a framework that bridges Large Language Models (LLMs) and power system simulation through the Model Context Protocol (MCP), enabling engineers to perform complex distribution analyses via natural language. Using OpenDSS as the reference 28 Sept 2026↗16Topology-Driven Anti-Entanglement Control for Soft RobotsarXiv cs.AIarXiv:2605.05236v2 Announce Type: replace-cross Abstract: In the field of precision manufacturing in complex constrained environments, the role of soft robots is increasingly prominent, and the realization of anti-winding control based on multi-intelligent body reinforcement learning has become a research hotspot. One of the core problems at present is to coordinate multiple robots to complete the unwinding operation in a highly constrained environment. The existing distributed training framewo28 Sept 2026↗17VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI AutomationarXiv cs.AIarXiv:2604.21375v3 Announce Type: replace-cross Abstract: Autonomous GUI agents face two fundamental challenges: early stopping, where agents prematurely declare success without verifiable evidence, and repetitive loops, where agents cycle through the same failing actions without recovery. We present VLAA-GUI, a modular GUI agentic framework built around three integrated components that guide the system on when to Stop, Recover, and Search. First, a mandatory Completeness Verifier enforces UI-o28 Sept 2026↗18Information Aggregation with AI AgentsarXiv cs.AIarXiv:2604.20050v4 Announce Type: replace-cross Abstract: Can Large Language Models (AI agents) aggregate dispersed private information through trading and reason about the knowledge of others by observing price movements? We conduct a controlled experiment where AI agents trade in a prediction market after receiving private signals, across four information structures of increasing complexity. We find that although the median market is effective at aggregating information in the easy informatio28 Sept 2026↗19Scepsy: Serving Agentic Workflows Using Aggregate LLM PipelinesarXiv cs.AIarXiv:2604.15186v2 Announce Type: replace-cross Abstract: Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools. Serving them at a target throughput with low latency is hard because they are written in arbitrary agentic frameworks and their execution times are unpredictable: execution branches, fans out, or recurs in data-dependent ways. Since their LLMs often outnumber the available GPUs, they also oversubscribe GPUs. We describe Scepsy, a28 Sept 2026↗20The Shrinking Lifespan of LLMs in SciencearXiv cs.AIarXiv:2604.07530v3 Announce Type: replace-cross Abstract: Scaling laws describe how language model capabilities grow with compute and data, but say nothing about how long a model matters once released. We introduce time-to-peak and lifespan as measures of model obsolescence and use them to characterize the scientific adoption trajectories of 62 LLMs across more than 108k citing papers (2019-2025), separating active adoption from background citation to recover per-model trajectories that citatio28 Sept 2026↗