Last 7 Days (September 23 – September 29, 2026)
Little is known about how to manually design agents capable of dexterous manipulation. Some design principles have been inferred from close examination of how animals manipulate objects, but these structures and behaviors have so far resisted biomimicry and may not be optimal for artificial machines. Here we evolve freeform robots to pick up, hold, rotate, and use diverse objects. Unlike other approaches to optimizing robot hands, we do not presuppose the presence, articulation, or geometry of any part of the body. Although familiar prehensile forms such as tails, beaks, paws and claws may emerge spontaneously under certain conditions--and while such conditions could be of interest to evolutionary biologists--de novo manipulator design can also reveal whole new solutions, overlooked or unknown structures which may be better suited for the task at hand. We use contrastive learning to create a highly searchable genetic embedding of design space, an autoregressive developmental model to decode designs, evolutionary strategies to find good designs, and reinforcement learning to train each evolved design. Winning designs were automatically converted into a manufacturable blueprint, printed, assembled and tested in the real world in a zero-shot manner. The results represent the state-of-the-art in evolutionary robotics in terms of performance, diversity and complexity.
Primary: University of Oxford
All Institutions: University of Oxford, DeepMind
The paper presents a state-of-the-art framework for evolving freeform dexterous robots from scratch, combining contrastive learning, developmental models, and reinforcement learning to produce manufacturable, real-world capable agents. By removing prior assumptions on robot morphology, the work reveals novel, non-biomimetic structures that outperform traditional designs in dexterity and diversity, marking a significant leap in evolutionary robotics and automated hardware design.
The paper proposes a comprehensive framework for de novo robot design, combining contrastive learning for genetic embedding, autoregressive developmental models for decoding, evolutionary strategies for search, and reinforcement learning for control. The novelty lies in the lack of prior assumptions about morphology (no fixed joints or geometry), allowing for truly freeform evolution. The use of a developmental model to map latent codes to physical structures is a sophisticated approach that bridges high-level design search with low-level manufacturability.
The evaluation is rigorous, moving beyond simulation to real-world validation. The robots were 3D printed, assembled, and tested in a zero-shot manner, demonstrating the robustness of the evolved designs. The paper reports state-of-the-art performance in terms of dexterity, diversity, and complexity compared to previous evolutionary robotics works. The inclusion of diverse object manipulation tasks (pick, hold, rotate, use) provides a strong benchmark for the method's generality.
The paper details the pipeline components (contrastive learning, autoregressive model, ES, RL) sufficiently for experts to replicate the framework. However, the specific hyperparameters for the evolutionary search and the exact architecture of the developmental model are likely in the appendix or code, which is standard. The zero-shot real-world testing is a high bar for reproducibility that the authors have met by providing the blueprint conversion process.
The computational cost of evolving and training multiple robots is significant. The "zero-shot" real-world performance, while impressive, may still be limited in speed or precision compared to hand-tuned industrial manipulators. The generalizability to tasks requiring extremely high force or precision (beyond what 3D printed materials can handle) is a potential constraint.
This work has significant implications for the field of evolutionary robotics and automated design. It suggests that optimal manipulator structures may not be biomimetic, challenging existing design paradigms. The ability to automatically generate manufacturable blueprints from abstract design spaces could accelerate the development of specialized robotic tools for manufacturing, surgery, or exploration. The paper presents a state-of-the-art framework for evolving freeform dexterous robots from scratch, combining contrastive learning, developmental models, and reinforcement learning to produce manufacturable, real-world capable agents. By removing prior assumptions on robot morphology, the work reveals novel, non-biomimetic structures that outperform traditional designs in dexterity and diversity, marking a significant leap in evolutionary robotics and automated hardware design.
Leaderboards rank models by their average scores on benchmark items, and they are consulted repeatedly while the evaluation is still running. Existing confidence intervals for a model's rank control their error rate only if they are computed once, after a number of items chosen in advance. If they are recomputed as results arrive, and the evaluation stops once they look decisive, their error rate exceeds its nominal level. Anytime-valid methods keep their guarantees at all sample sizes simultaneously and hence under any stopping rule. They exist for the accuracy of one model, for one pair of models and for the set of models that may be best. For pairwise battles they also give ranks. None gives ranks when all models are scored on the same items, which makes their scores dependent. We construct rank confidence sequences: for every model, a set of ranks that contains its true rank, simultaneously for all models and at all times, at a chosen error level $α$, in finite samples. The construction combines betting e-processes, one for each ordered pair of models, with closed testing over the possible orderings of the models. It allows any dependence between the models' scores on an item. The method has two advantages. A leaderboard can be inspected after every item without inflating its error rate. The evaluation of each model can stop as soon as the question asked about it is answered, which saves compute. When results are examined only once, halfway through or later, little power is lost relative to fixed-sample methods. The paper quantifies these advantages in simulations and on public leaderboard data.
Primary: Stanford University
All Institutions: Stanford University
The paper presents a novel and rigorous statistical method for anytime-valid ranking of machine learning models, addressing the critical issue of repeated inspection in leaderboards. By combining e-processes with closed testing over orderings, it provides finite-sample guarantees that hold under any stopping rule, offering both statistical validity and practical compute savings, making it a significant contribution to the field of ML evaluation.
The paper introduces "rank confidence sequences," a rigorous statistical framework for evaluating model rankings in a sequential, anytime-valid manner. The core innovation lies in combining betting e-processes for pairwise comparisons with closed testing over weak orderings (permutations with ties). This approach allows the construction of confidence sets for the ranks of all models simultaneously, valid at any time step, without assuming independence between model scores on the same items (a common issue in benchmarking). The method handles both fixed finite benchmarks (with random item ordering) and superpopulation settings. The theoretical contribution is significant, providing finite-sample guarantees that hold under any stopping rule, which addresses the "peeking" problem inherent in real-time leaderboard monitoring. The use of integer programming for exact certification at scale is a clever computational addition, though the coNP-hardness of the general problem is acknowledged.
The experiments are well-designed to validate the theoretical claims. E1 demonstrates the failure of fixed-sample methods under repeated monitoring, showing a 30% error rate vs. the nominal 5%. E2 applies the method to real LLM leaderboard data (Open LLM Leaderboard), showing that while early certification is limited, it becomes robust as more items are revealed. E3 highlights the practical benefit of compute savings, showing that early stopping for specific models (e.g., top-3 certification) can save significant evaluation costs without compromising validity. E4 compares power against fixed-sample methods at pre-planned look times, showing comparable performance. The use of real-world LLM data adds substantial practical relevance.
The paper provides a GitHub link to the code. The algorithms are described in detail, including the betting strategies, the offset calculations, and the integer programming formulations. The specific parameters for the betting grid are provided. The experimental setups are described with references to public datasets. The level of detail suggests high reproducibility for researchers with statistical programming skills.
The method requires per-item scores in [0,1] and assumes a fixed set of models evaluated on common items. It does not handle adaptive item selection or models arriving over time. The computational cost, while manageable for typical leaderboard sizes (up to ~50 models), grows with the number of models and items, and the exact certification via integer programming can be complex to implement correctly. The method is primarily useful for ranking, not for estimating the absolute performance gap with high precision in early stages.
This work has high potential impact on the ML evaluation community. As leaderboards become more dynamic and expensive to run, the need for statistically valid, anytime-valid monitoring is critical. This paper provides a principled way to do so, potentially changing how benchmarks are reported and interpreted. It bridges the gap between statistical theory (e-processes, closed testing) and practical ML evaluation. The compute savings aspect is also a strong practical incentive for adoption. The paper presents a novel and rigorous statistical method for anytime-valid ranking of machine learning models, addressing the critical issue of repeated inspection in leaderboards. By combining e-processes with closed testing over orderings, it provides finite-sample guarantees that hold under any stopping rule, offering both statistical validity and practical compute savings, making it a significant contribution to the field of ML evaluation.
Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but are widely viewed as harder because the chance constraint is nonconvex and depends on the full trajectory rather than a Bellman-linear expectation. In this paper, we reveal that this computational difficulty does not necessarily imply a higher statistical price. For tabular discounted CCMDPs with fixed bounded successor support and access to a certified planning oracle, we establish a model-based upper bound, with a matching lower bound up to logarithmic terms. Technically, our key idea is the \emph{Bellman distributional certificate}, which constructs a Bellman recursion for constraint violation probabilities before policy selection. The certificate can be reused across candidate policies; combined with shared row-wise reverse-KL confidence sets, it gives a policy-uniform trajectory-KL transfer without a union bound over policies or time--budget Bellman tables. For stochastic policies, we give a model-free variance-reduced policy-gradient algorithm with a finite-sample expected KKT-residual guarantee and independent validation of every accepted policy. Numerical experiments on synthetic CCMDPs and an IEEE 14-bus energy storage control benchmark illustrate the safety and mechanism behavior of the proposed algorithms.
Primary: Cornell University
All Institutions: Cornell University, University of Washington
The paper establishes a rigorous theoretical framework for learning chance-constrained MDPs by introducing Bellman distributional certificates that enable sample-efficient safety certification without union bounds, providing matching upper and lower bounds on sample complexity and a novel model-free policy gradient algorithm with finite-sample guarantees.
The paper introduces "Bellman distributional certificates," a novel theoretical framework for handling chance-constrained MDPs (CCMDPs). The core innovation is transforming the non-convex, trajectory-dependent chance constraint into a Bellman recursion over a discretized "remaining safety budget" state. This allows the use of standard model-based RL techniques (KL confidence sets) to certify safety with high probability without requiring a union bound over all possible policies or time steps. The method provides matching upper and lower bounds on sample complexity for deterministic policies, proving that the statistical cost of chance constraints is not inherently higher than expected-cost constraints under certain conditions (bounded successor support). Additionally, a model-free variance-reduced policy gradient algorithm is proposed for stochastic policies, offering finite-sample KKT-residual guarantees.
The experimental section is limited compared to the theoretical depth. It evaluates the method on a synthetic CCMDP and an IEEE 14-bus energy storage control benchmark. The results demonstrate that the Bellman-certified selector achieves better safety-performance trade-offs than a Markov-CMDP surrogate, particularly in reducing structural conservatism. However, the experiments are illustrative rather than exhaustive, lacking comparisons with state-of-the-art safe RL baselines (e.g., RCPO, PPO-Lagrangian) on standard continuous control benchmarks (MuJoCo, D4RL).
The paper provides detailed algorithmic descriptions and proofs in the appendix. However, specific hyperparameters, code availability, and implementation details for the "certified planning oracle" are not fully specified in the main text, making independent reproduction difficult without access to the authors' code. The reliance on a "certified planning oracle" as a black-box assumption limits immediate practical applicability.
1) The model-based guarantee is restricted to deterministic policies and assumes a fixed bound on successor support, which may not hold in high-dimensional continuous spaces. 2) The model-free result is local and may return "unresolved" if validation fails, lacking a global convergence guarantee. 3) The experimental validation is narrow, focusing on a single domain (energy storage) and synthetic tasks, without broad empirical validation on standard RL benchmarks. 4) The computational cost of the Bellman table grows with the discretization of the safety budget, which could be prohibitive for tight constraints.
This work provides a rigorous theoretical foundation for safe RL, addressing a critical gap in how probability-level safety constraints are handled. The "Bellman distributional certificate" concept could influence future work on risk-sensitive RL and constrained optimization. However, the strong assumptions (bounded support, deterministic policies for main bound) limit its immediate impact on practical, large-scale safe RL applications. The paper establishes a rigorous theoretical framework for learning chance-constrained MDPs by introducing Bellman distributional certificates that enable sample-efficient safety certification without union bounds, providing matching upper and lower bounds on sample complexity and a novel model-free policy gradient algorithm with finite-sample guarantees.
Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or leave the benefit of exploration unquantified. To this end, we propose Temperature-Grouped Reinforcement Learning (TGRL), which turns temperature-induced diversity into an explicit training signal. For each prompt, TGRL partitions its rollout group into low- and high-temperature subsets, estimates exploration gain through their reward contrast, and allocates this group-level signal as token-level credit using Jensen--Shannon (JS) divergence between the corresponding temperature-scaled next-token distributions induced by the same logits. Notably, TGRL reaches equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget. Across 11 benchmarks from diverse domains, TGRL broadly improves over strong RLVR baselines: it improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%. Comprehensive ablations and wall-clock analysis confirm the efficacy of all proposed components. Code is available at https://github.com/1229095296/TGRL/tree/main.
Primary: University of Chinese Academy of Sciences
All Institutions: University of Chinese Academy of Sciences, Meituan
The paper introduces TGRL, a reinforcement learning framework that leverages temperature-grouped rollouts to estimate exploration gain and allocate token-level credit via Jensen-Shannon divergence. It demonstrates that explicitly modeling the benefit of exploration through temperature contrast leads to more efficient convergence and improved performance across reasoning, coding, and agentic benchmarks compared to standard RLVR baselines.
The paper proposes TGRL, a method that integrates temperature-based exploration into the RLVR training loop by treating temperature-induced diversity as a learnable signal. The core mechanism involves partitioning rollouts into low-temperature (reference) and high-temperature (exploration) groups. The method estimates "exploration gain" via the reward contrast between these groups and uses Jensen-Shannon (JS) divergence between temperature-scaled token distributions to allocate this gain as token-level credit. The theoretical analysis provides bounds on how JS divergence captures local sensitivity to temperature changes, justifying the credit allocation strategy. The approach is logically sound, bridging the gap between sampling diversity and policy gradient updates in a principled manner.
The experimental evaluation is extensive, covering 11 benchmarks across mathematical reasoning, code generation, and agentic tasks. The model sizes tested (Qwen3-4B, 14B, 32B) are relevant to current LLM research. The results show consistent improvements over strong baselines like GRPO and DAPO, with specific gains in CodeForces rating and LiveCodeBench Pass@16. The ablation studies effectively isolate the contributions of the temperature grouping and JS-based credit allocation. The wall-clock analysis demonstrating 36% faster convergence is a strong practical contribution.
The paper provides a public GitHub repository and detailed hyperparameters (temperatures, rollout budgets, warmup steps). The use of standard models (Qwen3) and public benchmarks enhances reproducibility. However, the specific implementation details of the JS divergence calculation and the exact handling of the mixed-group advantage normalization would require careful inspection of the code to fully replicate.
The method relies on the assumption that temperature scaling is a sufficient proxy for exploration diversity. It may not capture all forms of beneficial exploration, such as those requiring semantic shifts rather than just stochastic sampling. The computational overhead of computing JS divergence for every token in the high-temperature group could be significant for very long contexts, though the paper claims efficiency gains. The performance gains, while consistent, are moderate (e.g., 1.6% on math average), which may limit its adoption if simpler baselines are sufficient for many tasks.
This work contributes to the understanding of how to efficiently utilize rollout budgets in LLM post-training. By providing a mechanism to quantify and exploit exploration gain, it offers a path toward more sample-efficient RLVR training. The insights into token-level credit allocation based on distributional sensitivity could be applicable to other areas of LLM optimization, such as test-time compute scaling. The paper introduces TGRL, a reinforcement learning framework that leverages temperature-grouped rollouts to estimate exploration gain and allocate token-level credit via Jensen-Shannon divergence. It demonstrates that explicitly modeling the benefit of exploration through temperature contrast leads to more efficient convergence and improved performance across reasoning, coding, and agentic benchmarks compared to standard RLVR baselines.
Leaderboards rank models by their average scores on benchmark items, and they are consulted repeatedly while the evaluation is still running. Existing confidence intervals for a model's rank control their error rate only if they are computed once, after a number of items chosen in advance. If they are recomputed as results arrive, and the evaluation stops once they look decisive, their error rate exceeds its nominal level. Anytime-valid methods keep their guarantees at all sample sizes simultaneously and hence under any stopping rule. They exist for the accuracy of one model, for one pair of models and for the set of models that may be best. For pairwise battles they also give ranks. None gives ranks when all models are scored on the same items, which makes their scores dependent. We construct rank confidence sequences: for every model, a set of ranks that contains its true rank, simultaneously for all models and at all times, at a chosen error level $α$, in finite samples. The construction combines betting e-processes, one for each ordered pair of models, with closed testing over the possible orderings of the models. It allows any dependence between the models' scores on an item. The method has two advantages. A leaderboard can be inspected after every item without inflating its error rate. The evaluation of each model can stop as soon as the question asked about it is answered, which saves compute. When results are examined only once, halfway through or later, little power is lost relative to fixed-sample methods. The paper quantifies these advantages in simulations and on public leaderboard data.
Primary: Stanford University
All Institutions: Stanford University
The paper presents a novel and rigorous statistical method for anytime-valid ranking of machine learning models, addressing the critical issue of repeated inspection in leaderboards. By combining e-processes with closed testing over orderings, it provides finite-sample guarantees that hold under any stopping rule, offering both statistical validity and practical compute savings, making it a significant contribution to the field of ML evaluation.
The paper introduces "rank confidence sequences," a rigorous statistical framework for evaluating model rankings in a sequential, anytime-valid manner. The core innovation lies in combining betting e-processes for pairwise comparisons with closed testing over weak orderings (permutations with ties). This approach allows the construction of confidence sets for the ranks of all models simultaneously, valid at any time step, without assuming independence between model scores on the same items (a common issue in benchmarking). The method handles both fixed finite benchmarks (with random item ordering) and superpopulation settings. The theoretical contribution is significant, providing finite-sample guarantees that hold under any stopping rule, which addresses the "peeking" problem inherent in real-time leaderboard monitoring. The use of integer programming for exact certification at scale is a clever computational addition, though the coNP-hardness of the general problem is acknowledged.
The experiments are well-designed to validate the theoretical claims. E1 demonstrates the failure of fixed-sample methods under repeated monitoring, showing a 30% error rate vs. the nominal 5%. E2 applies the method to real LLM leaderboard data (Open LLM Leaderboard), showing that while early certification is limited, it becomes robust as more items are revealed. E3 highlights the practical benefit of compute savings, showing that early stopping for specific models (e.g., top-3 certification) can save significant evaluation costs without compromising validity. E4 compares power against fixed-sample methods at pre-planned look times, showing comparable performance. The use of real-world LLM data adds substantial practical relevance.
The paper provides a GitHub link to the code. The algorithms are described in detail, including the betting strategies, the offset calculations, and the integer programming formulations. The specific parameters for the betting grid are provided. The experimental setups are described with references to public datasets. The level of detail suggests high reproducibility for researchers with statistical programming skills.
The method requires per-item scores in [0,1] and assumes a fixed set of models evaluated on common items. It does not handle adaptive item selection or models arriving over time. The computational cost, while manageable for typical leaderboard sizes (up to ~50 models), grows with the number of models and items, and the exact certification via integer programming can be complex to implement correctly. The method is primarily useful for ranking, not for estimating the absolute performance gap with high precision in early stages.
This work has high potential impact on the ML evaluation community. As leaderboards become more dynamic and expensive to run, the need for statistically valid, anytime-valid monitoring is critical. This paper provides a principled way to do so, potentially changing how benchmarks are reported and interpreted. It bridges the gap between statistical theory (e-processes, closed testing) and practical ML evaluation. The compute savings aspect is also a strong practical incentive for adoption. The paper presents a novel and rigorous statistical method for anytime-valid ranking of machine learning models, addressing the critical issue of repeated inspection in leaderboards. By combining e-processes with closed testing over orderings, it provides finite-sample guarantees that hold under any stopping rule, offering both statistical validity and practical compute savings, making it a significant contribution to the field of ML evaluation.
*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of *on-policy distillation (OPD)* across *weak-to-strong*, *same-base*, and *strong-to-weak* teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, $G$) rises approximately linearly in $d=\sqrt{\mathrm{KL}(π_θ\Vert π_{\mathrm{ref}})}$, the square root of token-level reverse KL divergence from the student initialization. In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own, so a compact RL expert can transfer capability to a much larger student via OPD. To estimate OPD outcomes, we fit *power laws* for how $G_{\mathrm{peak}}$ and the slope of the useful-transfer regime scale with student and teacher parameter counts and with teacher gold score. These laws show that peak gold score improves with teacher scale only up to roughly the student's scale, and that at a matched gold score smaller teachers transfer better, so a teacher's score alone does not define its supervision value. We also study the scaling effects of two OPD variants, bootstrapping weak-to-strong OPD, and the degree of on-policy supervision.
Primary: Unknown (Affiliations not explicitly listed in provided text, but authors Qinfeng Li, Wenqi Zhang, Guoqing Jiang, Liwei Chen, Xuanping Li, Zhiheng Qin, Yuntai Bao, Xuhong Zhang are associated with Alibaba Group / Qwen Team)
All Institutions: Alibaba Group (Inferred from author list and Qwen model usage)
The paper establishes predictive scaling laws for On-Policy Distillation, demonstrating that peak student performance and transfer rates can be accurately estimated from teacher/student scale and teacher quality, with surprising findings that smaller teachers often outperform larger ones at matched scores and that bootstrapping is ineffective compared to direct transfer.
The paper employs a rigorous empirical scaling analysis framework, adapting the methodology of reward model overoptimization studies (Gao et al., 2023) to the domain of On-Policy Distillation (OPD). The core methodological contribution is the characterization of OPD training dynamics as a function of the square root of token-level reverse KL divergence ($d$). The authors identify a "useful-transfer" regime where gold score increases linearly with $d$, followed by a noisy tail. They fit power laws to predict peak performance ($G_{peak}$) and transfer rates based on student/teacher scale and teacher quality. The inclusion of a theoretical derivation (Appendix) explaining why local KL geometry leads to linear transfer in $d$ adds significant depth, linking empirical observations to Fisher information geometry. The comparison between Vanilla-OPD and Delta-OPD (using policy shift contrasts) is well-motivated and clearly defined.
The experimental design is comprehensive, covering 25 teacher-student combinations across Qwen2.5 models (0.5B-14B) in weak-to-strong, same-base, and strong-to-weak configurations. The use of a single model family (Qwen2.5) is a limitation but allows for controlled isolation of scale effects. The evaluation on math reasoning (GSM8K/MATH) is standard but sufficient for this type of scaling study. Key findings include: (1) Peak student error is proportional to teacher remaining error, (2) Smaller teachers transfer better at matched scores (counter-intuitive and significant), (3) Bootstrapping does not improve over direct transfer from the smallest expert, and (4) Off-policy cold starts harm weak-to-strong transfer. The validation via leave-one-scale-out prediction is a strong methodological choice that demonstrates the predictive power of the fitted laws.
The paper provides detailed hyperparameters in the appendix and specifies the use of the `verl` framework. However, the reliance on a single random seed for all runs is a significant reproducibility weakness, although the authors justify this by citing standard practices in scaling law studies (Kaplan et al., Hoffmann et al.). The lack of seed variance estimation limits the confidence in the precise coefficients of the power laws, though the trends appear robust across the grid. The code and data are not explicitly linked in the provided text, which is a minor gap for immediate reproducibility.
The primary limitation is the restriction to a single model family (Qwen2.5) and a single task domain (math reasoning). The authors acknowledge that whether these scaling laws hold for other architectures, tasks, or post-training recipes is untested. The single-seed constraint means that stochastic variance in RL/OPD training is not captured, potentially masking instability in the "noisy tail" dynamics. The theoretical derivation assumes smoothness and differentiability that may not hold strictly in discrete token spaces or with clipping mechanisms, though the empirical fit is strong.
This paper has high practical impact for LLM practitioners. By providing predictive scaling laws for OPD, it enables engineers to estimate the outcome of distillation runs before committing significant compute resources. The finding that smaller teachers can be more effective than larger ones at matched scores challenges common assumptions about teacher quality and offers a cost-effective strategy for model family development. The negative result on bootstrapping is also valuable, preventing wasted effort on a seemingly intuitive but ineffective strategy. This work bridges the gap between theoretical scaling laws and practical post-training pipelines. The paper establishes predictive scaling laws for On-Policy Distillation, demonstrating that peak student performance and transfer rates can be accurately estimated from teacher/student scale and teacher quality, with surprising findings that smaller teachers often outperform larger ones at matched scores and that bootstrapping is ineffective compared to direct transfer.
The widely used Transolver family of neural operators is based on physics-attention, which softly assigns the points of an unstructured mesh to a small number of slices, applies self-attention among the resulting tokens, and broadcasts the result back to the points. We provide a comprehensive empirical and theoretical analysis to elucidate the mechanisms which are responsible for model performance. To this end, we perform careful ablations on a challenging suite of nine 3D fluid dynamics benchmarks to find that replacing token attention with a constant linear map does not affect the accuracy. Thus, Transolver does not need a Transformer at all. However, removing the global mixing (slicing/deslicing) or doing it only once leads to performance collapse. We leverage the theory of averaging neural operators to explain and corroborate our findings by showing that just slicing/deslicing, in conjunction with pointwise MLPs, already suffices for universal approximation of continuous operators and attention is redundant in this context. Finally, we provide a novel FlashAttention-style efficient implementation of the key slicing/deslicing module of Transolver. This flashslice kernel streams slice and deslice over the points without materializing the heavy slice-weight tensor, while reproducing the best available implementation to floating point error. At the same time, it leads to very significant memory and compute savings, particularly at large slice counts.
Primary: NVIDIA
All Institutions: NVIDIA
The paper demonstrates that the Transformer component in Transolver is redundant for accuracy, with the slicing/deslicing mechanism being the key driver of performance, supported by theoretical universality results and efficient kernel implementations. This work provides a critical understanding of neural operator architectures, guiding future design towards more efficient global mixing strategies without the computational overhead of self-attention.
The paper employs a rigorous ablation study combined with theoretical analysis to deconstruct the Transolver architecture. The methodology is sound, isolating the "physics-attention" mechanism into slicing/deslicing, token attention, and pointwise MLPs. The theoretical contribution leverages the theory of Averaging Neural Operators (ANO) to prove that the slicing/deslicing component, even with a constant core (no attention), is sufficient for universal approximation of continuous operators. This provides a strong mathematical foundation for the empirical finding that the Transformer component is redundant. The implementation of a FlashAttention-style kernel for the slicing operation is a significant technical contribution, addressing the memory bottleneck of materializing slice weights.
The experiments are extensive, covering nine challenging 3D fluid dynamics benchmarks, including industrial-scale aerodynamics (DrivAerNet++, SHIFT-SUV, SHIFT-Wing, DrivAerML). The evaluation protocol is careful, using matched step budgets and identical data pipelines. The results clearly demonstrate that removing token attention does not degrade accuracy, while removing the global mixing (slicing) causes performance collapse. The efficiency gains from the new kernel are substantial, showing significant memory savings and speedups, particularly at large slice counts.
The paper provides detailed descriptions of the ablation variants, training protocols, and kernel implementation. The use of standard datasets and clear metric definitions (relative L1 error) supports reproducibility. The code for the FlashSlice kernel is described in detail, though a direct link is not provided in the text snippet, the description is sufficient for implementation.
The theory is based on expressivity (universality) and does not address generalization or optimization dynamics. The bounds are not sharp. The findings are specific to the Transolver architecture and may not generalize to other neural operators without further analysis. The paper acknowledges that the cost of removing attention is empirically zero, but the theory only makes it unsurprising, not derived from first principles.
This paper has high impact on the scientific machine learning community by clarifying the fundamental mechanisms of a widely used architecture. It challenges the assumption that self-attention is necessary for global mixing in operator learning, potentially leading to more efficient and simpler models. The efficient kernel implementation will benefit practitioners working with large-scale unstructured meshes. The paper demonstrates that the Transformer component in Transolver is redundant for accuracy, with the slicing/deslicing mechanism being the key driver of performance, supported by theoretical universality results and efficient kernel implementations. This work provides a critical understanding of neural operator architectures, guiding future design towards more efficient global mixing strategies without the computational overhead of self-attention.
Transformers and state-space models (SSMs) are two prominent sequential learning architectures, yet their comparison remains largely empirical and existing theoretical analyses are typically task-specific or architecturally restricted. In this paper, we develop belief geometry, a unified analytical framework for comparing the representational capabilities of broad classes of attention and SSMs. Starting from a generalized formulation of in-context linear regression and using cumulative Bayes regret as our measure, we abstract three capabilities required by many sequential learning problems in our belief geometry: evidence assembly, belief maintenance, and addressing. We then study three cases of our formulation that isolate these capabilities and yield sharp architectural lessons: For belief maintenance, SSMs attain the optimal regret over stationary aggregation kernels; for positional assembly, SSMs have a memory advantage; and for content addressing, softmax attention has an exponential width advantage over sigmoid-selective SSMs. Experiments with LLaMA-type Transformers and Mamba-2 show that these architectural insights extend beyond our analytically tractable classes and linear-regression testbed.
Primary: The Ohio State University
All Institutions: The Ohio State University, RWTH Aachen University
Develops "belief geometry," a unified theoretical framework that rigorously compares the representational capabilities of attention and state-space models by isolating evidence assembly, belief maintenance, and addressing, yielding sharp architectural lessons such as the exponential width advantage of softmax attention for content addressing and the memory advantage of SSMs for positional assembly.
The paper introduces "belief geometry," a unified analytical framework to compare Transformers and State-Space Models (SSMs) in the context of in-context linear regression (ICLR). The methodology is rigorous, moving beyond empirical benchmarks to derive theoretical lower and upper bounds on cumulative Bayes regret. It decomposes sequential learning into three capabilities: evidence assembly, belief maintenance, and addressing. The authors define specific architectural classes (linear/softmax attention, fixed/selective SSMs) and prove sharp separations: SSMs are optimal for stationary belief maintenance (exponential kernels vs. uniform attention), SSMs have a memory advantage for positional assembly (convolution vs. context window), and softmax attention has an exponential width advantage for content addressing (selecting from retained tokens vs. storing candidates in state). The use of cumulative regret rather than terminal loss is a significant methodological improvement for analyzing sequential dynamics.
The experiments validate the theoretical predictions using practical architectures (LLaMA-type Transformers and Mamba-2). The authors test kernel alignment in single-layer models, showing that Transformers learn flatter kernels while Mamba-2 learns exponential ones, matching the theoretical optima. For the routing tasks (positional and content), they demonstrate that performance thresholds align with the theoretical resource requirements (e.g., convolution width for SSMs, context length for Transformers). The experiments effectively bridge the gap between the analytically tractable linear regression testbed and practical non-linear models, confirming that the architectural lessons (e.g., exponential width gap for addressing) hold in practice.
The paper includes a reproducibility statement and provides a link to an anonymous code repository. The experimental protocols, including hyperparameter sweeps and evaluation metrics, are detailed in the appendices. The use of synthetic tasks (Gaussian filtering, Beta-Bernoulli bandits, logistic bandits) ensures that the results are deterministic and reproducible without access to large-scale datasets.
The primary limitation is the reliance on linear regression and conjugate/non-conjugate bandit settings for the theoretical analysis. While the authors argue that the "belief geometry" extends to broader problems, the proofs are specific to these linear/quadratic loss structures. The "exponential width advantage" for attention in content addressing is a strong claim, but it relies on specific definitions of "addressing capacity" and compressed codebooks; real-world semantic addressing may not strictly follow this geometric separation. Additionally, the comparison focuses on representational capability (what the model *can* do) rather than optimization dynamics (what the model *learns* efficiently), though the kernel alignment experiments partially address this.
This paper provides a principled theoretical foundation for the ongoing debate between Transformers and SSMs. By identifying specific architectural mechanisms (softmax vs. selective transitions, context window vs. recurrent state) that confer advantages in different task regimes, it offers actionable insights for architecture design. The framework of "belief geometry" could be extended to other sequential tasks, such as language modeling or control, potentially guiding the development of hybrid architectures that leverage the strengths of both families. The finding that SSMs are inherently better at stationary belief maintenance while attention is better at content addressing helps explain empirical observations in long-context learning and associative recall. Develops "belief geometry," a unified theoretical framework that rigorously compares the representational capabilities of attention and state-space models by isolating evidence assembly, belief maintenance, and addressing, yielding sharp architectural lessons such as the exponential width advantage of softmax attention for content addressing and the memory advantage of SSMs for positional assembly.
Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but are widely viewed as harder because the chance constraint is nonconvex and depends on the full trajectory rather than a Bellman-linear expectation. In this paper, we reveal that this computational difficulty does not necessarily imply a higher statistical price. For tabular discounted CCMDPs with fixed bounded successor support and access to a certified planning oracle, we establish a model-based upper bound, with a matching lower bound up to logarithmic terms. Technically, our key idea is the \emph{Bellman distributional certificate}, which constructs a Bellman recursion for constraint violation probabilities before policy selection. The certificate can be reused across candidate policies; combined with shared row-wise reverse-KL confidence sets, it gives a policy-uniform trajectory-KL transfer without a union bound over policies or time--budget Bellman tables. For stochastic policies, we give a model-free variance-reduced policy-gradient algorithm with a finite-sample expected KKT-residual guarantee and independent validation of every accepted policy. Numerical experiments on synthetic CCMDPs and an IEEE 14-bus energy storage control benchmark illustrate the safety and mechanism behavior of the proposed algorithms.
Primary: Cornell University
All Institutions: Cornell University, University of Washington
The paper establishes a rigorous theoretical framework for learning chance-constrained MDPs by introducing Bellman distributional certificates that enable sample-efficient safety certification without union bounds, providing matching upper and lower bounds on sample complexity and a novel model-free policy gradient algorithm with finite-sample guarantees.
The paper introduces "Bellman distributional certificates," a novel theoretical framework for handling chance-constrained MDPs (CCMDPs). The core innovation is transforming the non-convex, trajectory-dependent chance constraint into a Bellman recursion over a discretized "remaining safety budget" state. This allows the use of standard model-based RL techniques (KL confidence sets) to certify safety with high probability without requiring a union bound over all possible policies or time steps. The method provides matching upper and lower bounds on sample complexity for deterministic policies, proving that the statistical cost of chance constraints is not inherently higher than expected-cost constraints under certain conditions (bounded successor support). Additionally, a model-free variance-reduced policy gradient algorithm is proposed for stochastic policies, offering finite-sample KKT-residual guarantees.
The experimental section is limited compared to the theoretical depth. It evaluates the method on a synthetic CCMDP and an IEEE 14-bus energy storage control benchmark. The results demonstrate that the Bellman-certified selector achieves better safety-performance trade-offs than a Markov-CMDP surrogate, particularly in reducing structural conservatism. However, the experiments are illustrative rather than exhaustive, lacking comparisons with state-of-the-art safe RL baselines (e.g., RCPO, PPO-Lagrangian) on standard continuous control benchmarks (MuJoCo, D4RL).
The paper provides detailed algorithmic descriptions and proofs in the appendix. However, specific hyperparameters, code availability, and implementation details for the "certified planning oracle" are not fully specified in the main text, making independent reproduction difficult without access to the authors' code. The reliance on a "certified planning oracle" as a black-box assumption limits immediate practical applicability.
1) The model-based guarantee is restricted to deterministic policies and assumes a fixed bound on successor support, which may not hold in high-dimensional continuous spaces. 2) The model-free result is local and may return "unresolved" if validation fails, lacking a global convergence guarantee. 3) The experimental validation is narrow, focusing on a single domain (energy storage) and synthetic tasks, without broad empirical validation on standard RL benchmarks. 4) The computational cost of the Bellman table grows with the discretization of the safety budget, which could be prohibitive for tight constraints.
This work provides a rigorous theoretical foundation for safe RL, addressing a critical gap in how probability-level safety constraints are handled. The "Bellman distributional certificate" concept could influence future work on risk-sensitive RL and constrained optimization. However, the strong assumptions (bounded support, deterministic policies for main bound) limit its immediate impact on practical, large-scale safe RL applications. The paper establishes a rigorous theoretical framework for learning chance-constrained MDPs by introducing Bellman distributional certificates that enable sample-efficient safety certification without union bounds, providing matching upper and lower bounds on sample complexity and a novel model-free policy gradient algorithm with finite-sample guarantees.
Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existing nonparametric ICL theory has largely focused on Euclidean domains or single-manifold models. To address this gap, we study the prediction problem under unknown local geometry, modeled by sample size-dependent mixtures of manifolds with heterogeneous dimensions, smoothness, and sampling masses. Under local separation and small-perturbation conditions, we establish a minimax lower bound capturing the aggregate difficulty of the components and construct an oracle tangent local-polynomial estimator with a matching upper bound. This estimator is connected to a structure-informed, two-stage softmax transformer with a geometric preconditioner and chartwise reduced local-polynomial solvers. The transformer achieves negligible approximation error relative to the minimax rate with logarithmic depth and polynomial size. Finally, we derive an in-context generalization bound for near empirical risk minimizers over this class. Together, these results identify conditions under which the resulting predictor exploits local geometry and attains the aggregate minimax rate.
Primary: Seoul National University
All Institutions: Seoul National University
[One sentence main contribution]. The paper establishes minimax optimality for Transformers in in-context learning on heterogeneous manifold mixtures by connecting oracle tangent local-polynomial estimators to a structure-informed Transformer architecture, providing a rigorous theoretical foundation for the model's ability to exploit local geometric structure.
The paper proposes a theoretical framework for analyzing Transformers in the context of in-context learning (ICL) on data with heterogeneous local geometry (mixtures of manifolds). The core methodological contribution is the derivation of minimax lower bounds for prediction under these complex geometric conditions and the construction of an oracle estimator (tangent local-polynomial) that matches these bounds. Crucially, the authors demonstrate that a specific class of Transformers (two-stage softmax with geometric preconditioners) can approximate this oracle estimator with negligible error, thereby establishing that Transformers can achieve minimax optimality in this setting. The approach is highly theoretical, relying on statistical learning theory and differential geometry rather than empirical model training.
The provided text is primarily theoretical, focusing on proofs and bounds. There is no mention of extensive empirical experiments, benchmarks, or real-world dataset evaluations in the abstract or the visible text fragments. The "experiments" are likely mathematical verifications of the bounds and the approximation capabilities of the proposed Transformer architecture. As a pure theory paper, its value lies in the rigor of the proofs rather than empirical performance metrics.
Reproducibility in the context of this paper refers to the verifiability of the mathematical proofs. The paper is 63 pages long, suggesting detailed appendices with full proofs. However, without access to the code or specific simulation scripts (if any exist to validate the theoretical bounds empirically), reproducibility is limited to the mathematical derivation. The lack of a provided code repository in the extracted information makes empirical validation difficult for external researchers.
The primary limitation is the gap between theory and practice. The conditions required for the minimax optimality (local separation, small-perturbation conditions, specific mixture structures) may be restrictive and not always satisfied by real-world high-dimensional data. Furthermore, the "oracle" nature of the estimator and the specific "structure-informed" Transformer design may not be directly implementable in standard large language model architectures without significant architectural modifications. The paper does not appear to provide empirical evidence that standard Transformers naturally learn these geometric preconditioners.
This paper contributes to the foundational understanding of why Transformers are effective for ICL, extending the theory beyond simple Euclidean or single-manifold settings. It provides a rigorous justification for the use of local polynomial approximations within Transformer layers. While the immediate practical impact on engineering may be limited, it offers valuable insights for the design of future architectures that explicitly account for data geometry, potentially influencing the development of more efficient and robust foundation models. [One sentence main contribution]. The paper establishes minimax optimality for Transformers in in-context learning on heterogeneous manifold mixtures by connecting oracle tangent local-polynomial estimators to a structure-informed Transformer architecture, providing a rigorous theoretical foundation for the model's ability to exploit local geometric structure.
Density functional theory (DFT) strikes a practical balance between accuracy and computational cost in many problems of computational chemistry and materials science. However, many DFT calculations are limited by fixed atom-centered basis sets, which dictate how accuracy and cost scale with system size. We propose Gaussian Splatting for Density Functional Theory (GS-DFT), which represents molecular orbitals as a cloud of Gaussians whose positions, shapes, and mixing coefficients are optimized jointly by gradient descent to minimize the energy without training data. Conceptually, GS-DFT is 3D Gaussian splatting with the renderer replaced by quantum mechanics. We introduce two key solver components: adaptive density fitting with screening for efficient evaluation of two-electron integrals, and a regularized differentiable orthogonalization of the molecular orbitals. Empirically, the optimized basis reaches the accuracy of the largest conventional basis sets with a fraction of the parameters, converging systematically in energy, density, and nuclear forces. At equal parameter count, it captures the stretched-bond and anion physics that fixed bases only recover with specialized basis augmentation. The resulting solver exhibits quadratic peak memory scaling in the cloud size, allowing us to simulate systems of up to 2,742 atoms (10,406 electrons) without any modifications at triple-zeta scale using a single four-GPU node.
Primary: Unknown
All Institutions: Unknown
The paper introduces GS-DFT, a novel method using 3D Gaussian splatting to represent molecular orbitals in DFT, achieving high accuracy with fewer parameters and enabling large-scale simulations on GPU clusters. This represents a significant methodological advance in computational chemistry, combining differentiable rendering concepts with quantum mechanical solvers to overcome the limitations of fixed basis sets.
The paper proposes GS-DFT, a method that replaces fixed atom-centered basis sets in Density Functional Theory with a cloud of 3D Gaussian splats. The positions, shapes, and mixing coefficients of these Gaussians are optimized via gradient descent to minimize the electronic energy. This is a significant conceptual shift, framing the basis set optimization as a differentiable rendering problem where the "renderer" is the quantum mechanical solver. The introduction of adaptive density fitting with screening for two-electron integrals and a regularized differentiable orthogonalization scheme are critical technical contributions that enable the stability and efficiency of this optimization. The approach is end-to-end differentiable and does not require training data, distinguishing it from standard neural network surrogates for DFT.
The experimental results are impressive, claiming that the optimized basis reaches the accuracy of the largest conventional basis sets with a fraction of the parameters. The paper reports systematic convergence in energy, density, and nuclear forces. A key highlight is the ability to simulate systems of up to 2,742 atoms (10,406 electrons) on a single four-GPU node, demonstrating quadratic peak memory scaling. The comparison to fixed bases shows that GS-DFT captures stretched-bond and anion physics that typically require specialized basis augmentation in traditional DFT. However, the lack of specific benchmark details (e.g., which molecules, which functionals, comparison to specific state-of-the-art DFT codes like ORCA or Gaussian) in the provided abstract limits the full verification of these claims.
The paper describes the solver components (adaptive density fitting, regularized orthogonalization) which are crucial for reproducibility. However, without access to the full code or detailed hyperparameter settings (e.g., number of Gaussians, optimization schedule, regularization strength), independent reproduction may be challenging. The reliance on a specific hardware setup (four-GPU node) for the large-scale simulations also poses a barrier for smaller groups.
The primary limitation is the computational cost of the gradient descent optimization for the basis set itself, which may be prohibitive for very large systems despite the efficient solver. The method is currently demonstrated on molecular systems; its applicability to periodic boundary conditions (solids) is not discussed. The "quadratic peak memory scaling" is a significant improvement over cubic scaling of traditional DFT, but it still limits the system size compared to linear-scaling DFT methods. The lack of comparison to other learned basis set methods or neural network potentials is a gap.
This work has the potential to significantly impact computational chemistry and materials science by providing a flexible, high-accuracy basis set representation that scales better than traditional methods. It bridges the gap between differentiable rendering techniques (3D Gaussian Splatting) and quantum mechanics, opening new avenues for differentiable physics. The ability to simulate larger systems on standard GPU hardware could democratize high-accuracy DFT calculations. The paper introduces GS-DFT, a novel method using 3D Gaussian splatting to represent molecular orbitals in DFT, achieving high accuracy with fewer parameters and enabling large-scale simulations on GPU clusters. This represents a significant methodological advance in computational chemistry, combining differentiable rendering concepts with quantum mechanical solvers to overcome the limitations of fixed basis sets.
Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for each task, a relevant subsequence is extracted from the full history and a dedicated model trained on it. In production it hits two bottlenecks. First, even after filtering, a single-task sequence stays extremely long: content-interest summarization reads several hundred items per user, tens of thousands of tokens once serialized as prompt text. Second, profiles are refreshed routinely: a billion users weekly, roughly 100K QPM in aggregate, which under a fixed GPU budget sets a hard throughput floor. Compression is therefore mandatory, yet truncation or coarse compression can silently distort the profile, introducing four hallucination types (fabrication, omission, date misattribution, broken logic) that, with no way to evaluate the compressed representation itself, surface only as diffuse degradation in downstream metrics. We present KuaFu, a unified behavior-compression layer whose minimal unit is one behavior item. A two-axis projector compresses each item into 2-4 tokens of width 128-256 (about 10x along the token axis, 20x along width; per-item cache 10 KB to 0.5 KB), with fidelity-oriented four-stage training and layered intermediate evaluation. Across four production profiling tasks it matches or exceeds uncompressed single-task production models on all five headline metrics, raises per-GPU throughput by 37%-350%, and saves 190 GPUs. On public benchmarks it nearly always beats prior compressors at the same compression ratio (up to +17.7 EM on out-of-domain MRQA); on RecBench, a 4B model surpasses its 8B counterpart by 1.90 points. KuaFu has run on the Tencent advertising and recommendation platform for ten months, lifting overall GMV by 1.37%.
Primary: Tencent
All Institutions: Tencent
KuaFu introduces an item-level, two-axis compression layer for user behavior sequences that enables billion-scale LLM-based profiling with significant cost and latency reductions. The paper demonstrates that by treating individual behavior items as independent, cacheable units and employing a multi-stage training recipe including hallucination-aware RL, it is possible to maintain or improve downstream task performance while drastically reducing computational overhead, validated by a 1.37% GMV lift in a major production environment.
The paper proposes KuaFu, a unified behavior-compression layer that treats individual behavior items as the minimal unit of compression. The core architectural contribution is a two-axis projector that compresses item embeddings along both the token axis (reducing sequence length) and the width axis (reducing dimensionality), allowing for significant storage and compute savings (10KB to 0.5KB per item). The training methodology is sophisticated, involving a four-stage curriculum: (1) reconstruction pre-training with a length-based curriculum to ensure convergence on long sequences, (2) post-training for compressed QA, (3) co-training of the compressor and decoder, and (4) hallucination-aware reinforcement learning (DAPO) specifically designed to penalize fabrication, omission, date misattribution, and broken logic. This multi-stage approach, particularly the explicit handling of hallucination types in the RL reward function, is a strong methodological contribution for LLM-based user modeling.
The experimental evaluation is extensive and convincing, combining offline benchmarks with large-scale online A/B testing. Offline, KuaFu outperforms state-of-the-art compression methods (SAC, EPL, ICAE) on MRQA benchmarks, showing superior fidelity at high compression ratios. On RecBench, the compressed 4B model outperforms the uncompressed 8B model, demonstrating the efficiency gains. The online A/B test on Tencent's platform is the strongest evidence of impact, showing a 1.37% GMV lift and significant throughput improvements (37-350% per-GPU QPM) while saving 190 GPUs. The ablation studies clearly demonstrate the necessity of the curriculum learning and the specific projector design.
Reproducibility is moderate. While the paper provides detailed descriptions of the architecture, training stages, and hyperparameters, the core training data consists of proprietary industrial behavior logs from Tencent, which cannot be released. The authors state that public datasets (MRQA, RecBench) are used for evaluation, allowing partial reproduction of the compression and understanding components. However, the specific industrial gains and the full training pipeline on proprietary data cannot be fully replicated by external researchers.
The primary limitation is the reliance on proprietary data, which limits external verification of the industrial claims. The compression ratio is currently fixed per task family, and the paper acknowledges that adaptation to sparser sequences and overly long items remains an area for improvement. Additionally, the scaling laws along data and parameter size have not been systematically validated.
This work has significant implications for the deployment of LLMs in industrial recommendation and advertising systems. By solving the context length and cost bottlenecks through item-level compression, it enables the use of large language models for real-time, billion-scale user profiling. The framework for evaluating compression fidelity (the layered intermediate evaluation) is also valuable for the broader community working on context compression for LLMs. KuaFu introduces an item-level, two-axis compression layer for user behavior sequences that enables billion-scale LLM-based profiling with significant cost and latency reductions. The paper demonstrates that by treating individual behavior items as independent, cacheable units and employing a multi-stage training recipe including hallucination-aware RL, it is possible to maintain or improve downstream task performance while drastically reducing computational overhead, validated by a 1.37% GMV lift in a major production environment.
The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test of this approach: complex, long-horizon tasks challenge agent reliability, while costly real-device interaction limits development scalability. The framework connects data production, model training, and deployment through a shared action-feedback-verification contract. (i) AI for Data builds a human-gated agentic data flywheel in which specialized agents construct tasks, collect interaction trajectories, curate and balance training data, and use training feedback to guide subsequent data generation. (ii) AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning, where we introduce Competence-Aware Reward-and-Advantage Engineering (CARE) to reduce reasoning and tool-use costs while preserving task performance. (iii) AI drives model--harness co-evolution through an execution-evidence-driven loop that orchestrates memory, skills, and tools at runtime and feeds structured action feedback and preserved failure traces back into coordinated model and harness adaptation. Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination. Further evaluations of our model show improvements across non-mobile agentic benchmarks while largely preserving general capabilities.
Primary: Alibaba Group
All Institutions: Alibaba Group
The paper presents a comprehensive framework for developing mobile planner agents using an AI-for-AI closed loop, achieving state-of-the-art performance on MobilePA-Bench. It demonstrates that combining AI-assisted data generation, competence-aware RL, and a dynamic runtime harness can significantly enhance agent reliability and efficiency, offering a scalable pathway for building next-generation autonomous systems.
The paper proposes a "closed-loop AI-for-AI" framework, which is a high-level architectural concept rather than a single novel algorithmic breakthrough. The core technical contributions are the integration of a data flywheel (AI-assisted task generation and curation), a hybrid training pipeline (SFT cold start + RL), and a runtime Harness (Skills/Memory). The most specific technical novelty is CARE (Competence-Aware Reward-and-Advantage Engineering), which adjusts reward shaping based on group success rates to prevent efficiency signals from dominating in saturated success groups. While the components (RL for agents, memory systems, data synthesis) are individually known, the systematic integration into a self-improving loop for mobile agents is a significant engineering and methodological contribution.
The evaluation is conducted on MobilePA-Bench, a large-scale benchmark (1,700+ tasks). The results show Qwen-Planner-Agent (27B) outperforming strong closed-source competitors like GPT-6 Astra and Claude Opus 5, as well as larger open-source models. The ablation studies effectively isolate the contributions of the Planner Model versus the Harness, demonstrating that the runtime context (Skills/Memory) provides substantial gains over the raw model. The efficiency analysis (cost per task) is a valuable addition, showing the agent is not just better but cheaper than frontier commercial APIs.
As a report from a major industry lab (Alibaba), the paper provides high-level architectural details but likely lacks the granular hyperparameters, exact prompt templates, and code releases typically required for full academic reproducibility. The "AI-for-AI" loop implies a complex, proprietary infrastructure for data generation and verification that is difficult to replicate externally. However, the clarity of the framework description allows for conceptual replication.
The primary limitation is the reliance on a proprietary, large-scale infrastructure for the "AI-for-AI" loop, making it difficult for smaller labs to verify the specific benefits of the data flywheel. The "closed-loop" claim is somewhat aspirational; the paper admits that human review is retained for critical decisions, meaning it is not fully autonomous. Additionally, the evaluation is heavily skewed toward the specific mobile planning domain, and while generalization is claimed, the non-mobile benchmarks are less detailed in the provided text.
This paper represents a significant step toward scalable agent development. By formalizing the use of AI to generate training data and diagnose failures for agent systems, it offers a roadmap for reducing the manual effort required to build robust LLM agents. The focus on mobile planning is highly relevant to current industry trends in on-device and cross-app automation. The CARE method offers a useful technique for RL training of agents where success rates vary widely across tasks. The paper presents a comprehensive framework for developing mobile planner agents using an AI-for-AI closed loop, achieving state-of-the-art performance on MobilePA-Bench. It demonstrates that combining AI-assisted data generation, competence-aware RL, and a dynamic runtime harness can significantly enhance agent reliability and efficiency, offering a scalable pathway for building next-generation autonomous systems.
Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Expert pruning reduces this storage burden; at a fixed pruning budget, the goal is to preserve the original model's output distribution as closely as possible. Yet an expert's usage or contribution magnitude does not by itself determine the damage caused by its removal. What matters is whether the surviving computation can replace its function. We introduce RAZOR, a training-free expert pruning method that scores functional replaceability using consensus residuals: deviations of expert outputs from the original weighted mixture. An exact single-deletion identity at a fixed layer input accounts for survivor renormalization and router-selected refill, providing local scores aggregated over calibration tokens for budgeted pruning without gradients or recovery training. On GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, and Hy3 at 25\% and 50\% expert removal, RAZOR achieves the highest nine-task macro average among the evaluated pruning methods in all eight settings. On the two backbones with matched REAP benchmark runs, it exceeds REAP by 2.12--5.59 points and wins all 36 paired task comparisons. It also lowers reverse KL relative to REAP in all four matched GLM-4.7-Flash and Qwen3.6-35B-A3B model--budget settings. Analysis of responses generated by Qwen3.6-35B-A3B nevertheless reveals changes in diversity, formatting, and termination, underscoring that task retention and predictive fidelity do not ensure generation stability.
Primary: NVIDIA
All Institutions: NVIDIA
RAZOR introduces a training-free expert pruning method for MoE models that scores experts by functional replaceability using consensus residuals, accounting for survivor renormalization and router refill. The paper demonstrates that this geometric approach yields superior downstream performance and predictive fidelity compared to magnitude-based baselines across multiple large-scale MoE architectures, while also revealing nuanced trade-offs in generation behavior that task metrics alone miss.
The paper introduces RAZOR, a training-free expert pruning method for Mixture-of-Experts (MoE) models. The core innovation is the shift from scoring experts based on usage frequency or output magnitude (as in REAP or EAN) to "functional replaceability." This is achieved by calculating "consensus residuals," which measure the deviation of an expert's output from the original weighted mixture of active experts. The method derives an exact single-deletion identity that accounts for two critical dynamic effects often ignored in static pruning: survivor renormalization (how the weights of remaining experts adjust) and router-selected refill (how the router promotes a new expert to replace the pruned one). The derivation is mathematically sound, providing a local surrogate for the global distributional shift. The approach is computationally efficient, requiring only forward passes on calibration data without gradients or recovery training.
The evaluation is extensive, covering four distinct MoE backbones (GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, Hy3) at two pruning budgets (25% and 50% expert removal). RAZOR consistently outperforms baselines (Frequency, EAN, REAP) in macro-average downstream performance across all eight model-budget settings. It also demonstrates superior predictive fidelity, measured by lower reverse KL divergence compared to REAP. A notable strength is the inclusion of generation behavior analysis (RQ3), which reveals that while task performance is retained, pruning can still affect diversity and termination patterns, providing a nuanced view of the trade-offs. The experiments are rigorous, using matched calibration sets and detailed ablations on scoring components (RMS vs. Mean, fixed-support vs. refill).
The paper provides high reproducibility. It includes detailed algorithmic pseudocode, specific hyperparameters for evaluation, and clear descriptions of the calibration data composition (Nemotron datasets). The implementation details regarding memory optimization (chunked scoring, layer-wise execution) are well-documented. However, the model checkpoints are not redistributable by the authors, which may limit independent verification for those without access to the specific proprietary or large-scale open models used.
The primary limitation is that the scoring is local and single-deletion based; it does not account for complex interactions when multiple experts are pruned simultaneously, nor does it guarantee that the "refill" candidate remains available if it is also pruned. The paper acknowledges that local output change is a surrogate, not a guarantee of global distributional fidelity. Additionally, the evaluation is limited to four specific model families, and the generalizability to other MoE architectures (e.g., those with different router mechanisms) is not fully established. The lack of measured serving latency or energy savings is a practical gap, as the method's utility is partly defined by compression efficiency.
This work contributes significantly to the field of model compression by providing a principled, geometry-aware method for MoE pruning. It challenges the common heuristic that "less used" experts are less important, showing that "less replaceable" experts are the critical ones to keep. This insight can guide future research in structured pruning and model distillation for sparse models. The finding that task retention does not equate to generation stability is also valuable for practitioners deploying pruned models in production, highlighting the need for multi-faceted evaluation metrics. RAZOR introduces a training-free expert pruning method for MoE models that scores experts by functional replaceability using consensus residuals, accounting for survivor renormalization and router refill. The paper demonstrates that this geometric approach yields superior downstream performance and predictive fidelity compared to magnitude-based baselines across multiple large-scale MoE architectures, while also revealing nuanced trade-offs in generation behavior that task metrics alone miss.
Frontier language models are rarely used in clinical workflows because the realistic, longitudinal benchmarks needed to develop them are scarce. Real electronic health record (EHR) data cannot be openly shared due to privacy, ethics or data use issues and it does not contain verifiable ground truth since the chart records only reflect what clinicians documented. We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers. Built entirely from public medical-education material with no protected health information, it comprises 1,268 longitudinal patients and 5,602 encounters, where every diagnosis, finding, and temporal relation is grounded in standard ontologies (ICD-10-CM, SNOMED CT, LOINC) and with a complete provenance chain back to its source medical education material. Synthetic Hospital is served through a simulated hospital record system that mirrors real EHR infrastructure (standard interoperability APIs, role-based access and function-calling interface). In a blinded review, physicians distinguished its records from real patient charts at near-chance rates (53\%). Across 10 frontier and open models, none approaches ceiling: the best model reconstructs a patient's longitudinal problem list with a severity-weighted F1 of 0.73, level with the mean of seven physicians on a matched subset but well below the best of them (0.89), and misses roughly half of clinically relevant findings when summarizing a chart. Overall, these results highlight that Synthetic Hospital is a difficult and realistic test of clinical AI performance.
Primary: Carnegie Mellon University
All Institutions: Carnegie Mellon University
Synthetic Hospital introduces an open, verifiable, and physician-validated longitudinal EHR benchmark that resolves privacy and ground-truth barriers in clinical AI evaluation. By grounding synthetic records in standard medical ontologies and public educational material, the paper provides a rigorous, reproducible testbed that reveals significant gaps in current frontier models' ability to perform longitudinal clinical reasoning, offering a crucial resource for advancing reliable clinical AI systems.
The paper proposes a deterministic, five-stage pipeline to construct "Synthetic Hospital," a longitudinal EHR benchmark. The core innovation is the decoupling of clinical ground truth from narrative generation. By first constructing a structured medical knowledge graph grounded in standard ontologies (ICD-10-CM, SNOMED CT, LOINC) from public USMLE-style board questions, and then rendering this graph into realistic clinical narratives using an LLM (Kimi 2.5), the authors ensure that every diagnosis and finding has a verifiable provenance chain. This addresses the two major barriers to open clinical benchmarks: privacy (no PHI) and ground truth ambiguity (real charts reflect documentation, not necessarily patient state). The methodology includes rigorous validation steps, such as a blinded physician study to assess realism and a pre-registered reference set to validate ontology-derived relationships.
The evaluation is comprehensive. It includes a realism study where physicians could not distinguish synthetic from real records (53% accuracy, near chance). It evaluates 10 frontier and open models across four tasks (patient diagnosis, summarization, retrieval, imaging indication). Key findings include that no model approaches ceiling performance, the best model achieves a severity-weighted F1 of 0.73 on diagnosis (matching the mean of seven physicians but below the best), and that agentic multi-turn loops often degrade performance compared to single-turn inference when context is available upfront, except for longitudinal diagnosis tasks. The paper also provides a robustness analysis showing that re-rendering the corpus with a different LLM (GPT-5.3) does not significantly change model rankings, mitigating concerns about generator bias.
High. The paper provides a detailed description of the pipeline, including specific thresholds for ontology mapping, clustering rules, and prompting strategies. Code and data are openly available on GitHub. The use of deterministic functions for benchmark labels and the release of a training split with verifiable rewards enhances reproducibility and utility for reinforcement learning or fine-tuning.
The benchmark is derived from medical education material (USMLE-style questions), which may not fully capture the complexity, noise, and atypical presentations of real-world clinical practice. The case mix is education-derived by design, potentially under-representing rare conditions. The agentic evaluation is limited to three models and one scaffold. The paper acknowledges that it does not explicitly simulate missingness or documentation errors, which are common in real EHRs.
This benchmark has high potential impact on the clinical AI community. By providing an open, verifiable, and realistic longitudinal EHR dataset, it enables the development and evaluation of clinical AI systems without the legal and ethical hurdles of using real patient data. The finding that current frontier models struggle with longitudinal synthesis and that agentic approaches have mixed utility provides actionable insights for system designers. The open nature of the data facilitates broader research and standardization of evaluation metrics in clinical NLP. Synthetic Hospital introduces an open, verifiable, and physician-validated longitudinal EHR benchmark that resolves privacy and ground-truth barriers in clinical AI evaluation. By grounding synthetic records in standard medical ontologies and public educational material, the paper provides a rigorous, reproducible testbed that reveals significant gaps in current frontier models' ability to perform longitudinal clinical reasoning, offering a crucial resource for advancing reliable clinical AI systems.
Computer-use agents (CUAs) increasingly act on behalf of users online. What happens when the environments they operate in have incentives of their own? Online marketplaces, for example, may favor some products over others, steering agents away from the user's objective. Existing CUA benchmarks cover cooperative settings or explicit attacks, but do not test whether agents preserve user objectives when the environment itself has a stake in the outcome. We introduce CAVEAT, a controlled benchmark spanning nine marketplace environments and a taxonomy of eight common steering mechanisms. Across five model families, agents purchase the user-optimal product in 78.6% of matched-control episodes but only 17.3% when steering mechanisms are enabled. Larger models and more reasoning improve robustness, but substantial failures persist. Trajectory analysis and targeted ablations identify three weaknesses in how agents decide: they (1) prematurely narrow the set of alternatives they consider, (2) impose priorities the user never stated, and (3) commit before resolving decision-relevant evidence. Guided by this diagnosis, we develop CAVEAT-Harness, which targets these failures and raises the optimal purchase rate by up to 80.0 percentage points, and show that targeted post-training further improves a smaller open model. These results establish incentive robustness as a distinct challenge for delegated agents, diagnose failure modes, and show how targeted interventions can substantially improve robustness.
Primary: Microsoft Research
All Institutions: Microsoft Research
The paper establishes incentive robustness as a critical and distinct challenge for computer-use agents, demonstrating that current models are highly vulnerable to subtle environmental steering. By introducing the CAVEAT benchmark and diagnosing specific decision-making failures, the work provides a rigorous framework for evaluating and improving agent behavior in misaligned environments, with practical interventions that significantly enhance user-aligned outcomes.
The paper introduces CAVEAT, a benchmark for evaluating Computer-Use Agents (CUAs) in environments with misaligned incentives. The core methodological contribution is the construction of nine controlled marketplace environments featuring eight distinct "steering mechanisms" (e.g., hidden fees, biased search rankings, dynamic pricing). The authors formalize the problem of "incentive robustness," distinguishing it from standard security attacks or cooperative tasks. They propose a diagnostic framework that categorizes agent failures into three specific decision-making weaknesses: premature narrowing of alternatives, imposition of unstated user priorities, and premature commitment. Based on this diagnosis, they develop CAVEAT-Harness, a post-processing or prompting strategy that explicitly targets these failure modes. The methodology is sound, moving beyond simple black-box evaluation to a structured analysis of *why* agents fail in adversarial economic contexts.
The evaluation is comprehensive, testing five model families across the CAVEAT benchmark. The key empirical finding is stark: agents achieve a 78.6% optimal purchase rate in control conditions but drop to 17.3% when steering mechanisms are active. This quantifies the vulnerability of current state-of-the-art agents to subtle environmental manipulation. The paper further demonstrates that while larger models and increased reasoning steps improve robustness, they do not eliminate the problem. The proposed CAVEAT-Harness significantly improves performance, raising the optimal purchase rate by up to 80.0 percentage points in some settings, validating the diagnostic approach. The inclusion of targeted post-training for smaller open models adds practical value to the findings.
The paper appears to be from a major research institution (Microsoft Research, inferred from the acknowledgments of prominent MSR researchers like Saleema Amershi, Gagan Bansal, and Ece Kamar). While the full code is not provided in the text snippet, the detailed description of the benchmark environments, steering mechanisms, and the CAVEAT-Harness protocol suggests a high level of reproducibility. The specific metrics (optimal purchase rate) and the taxonomy of failures provide clear guidelines for replication. The venue is identified as ICLR 2027, indicating it has passed rigorous peer review.
The primary limitation is the scope of the environments. The benchmark focuses on online marketplaces, which, while common, may not cover all types of incentive-misaligned environments (e.g., social media feeds, news aggregators, or multi-agent negotiation). The "steering mechanisms" are defined by the authors; real-world platforms may employ more complex or adaptive strategies. Additionally, the CAVEAT-Harness is a targeted intervention; its generalizability to other types of agent tasks or environments outside of purchasing decisions is not fully established. The drop in performance (78.6% to 17.3%) is dramatic, but the absolute performance in the control condition (78.6%) suggests that even in benign environments, agents are not perfect, which may confound the isolation of the steering effect.
This paper has significant implications for the deployment of autonomous agents in the real world. As CUAs become more prevalent, the ability of environments (platforms, marketplaces) to subtly manipulate agent behavior poses a serious risk to user autonomy and utility. By establishing "incentive robustness" as a distinct challenge, the paper shifts the focus from pure capability to alignment and safety in economic contexts. The findings will likely influence the design of future agent architectures, prompting strategies, and safety evaluations. It also highlights a gap in current AI safety research, which has often focused on explicit adversarial attacks rather than the subtle, incentive-driven manipulation inherent in many online platforms. The paper establishes incentive robustness as a critical and distinct challenge for computer-use agents, demonstrating that current models are highly vulnerable to subtle environmental steering. By introducing the CAVEAT benchmark and diagnosing specific decision-making failures, the work provides a rigorous framework for evaluating and improving agent behavior in misaligned environments, with practical interventions that significantly enhance user-aligned outcomes.
Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by similarity. This forces the system to decide what is worth remembering before the future query is known, irreversibly discarding information and producing a query-independent summary that must serve many possible downstream tasks. Learning such a write-time curator is also difficult because the value of a storage decision may only become apparent when a relevant query arrives, potentially many tasks later, creating a long-horizon credit-assignment problem. We instead retain raw trajectories and defer curation until read time, when the current task is known. Given the retrieved traces and the new task, a memory curator synthesizes a compact, task-adaptive payload tailored to the immediate need. Because this payload is consumed on the same task, the curator can be trained directly from immediate task success, avoiding delayed utility signals and the need to artificially group related tasks. Across ALFWorld, WebShop, and τ^2-bench, our Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively. Notably, even an untrained curator is already competitive with or surpasses these baselines, showing that task-adaptive read-time curation itself is a major source of the gain; training the curator further compounds the improvement.
Primary: Salesforce
All Institutions: Salesforce
[One sentence main contribution]. The paper introduces Just-in-Time Memory, a read-time curation strategy for LLM agents that outperforms write-time methods by synthesizing task-adaptive payloads from raw trajectories, demonstrating significant gains in success rates across multiple benchmarks.
The paper proposes "Just-in-Time Memory" (JitMem), shifting the curation of agent memory from write-time (fixed summaries) to read-time (task-adaptive synthesis). The core method involves retaining raw trajectories and using a "memory curator" LLM to synthesize a compact payload specifically for the current query. This addresses the long-horizon credit assignment problem inherent in write-time curation by allowing the curator to be trained directly on immediate task success. The approach is conceptually sound, leveraging the flexibility of LLMs to perform on-the-fly reasoning over raw data rather than relying on pre-computed, potentially lossy, static representations.
The experiments are conducted on three standard agent benchmarks: ALFWorld, WebShop, and τ^2-bench. The results show consistent improvements over no-memory baselines and existing write-time memory methods. The reported gains are substantial (16.2, 16.3, and 3.9 absolute points), suggesting the method is effective across different task types. The finding that even an untrained curator is competitive is a strong empirical result, highlighting the value of the read-time curation paradigm itself.
The paper is on arXiv. While the methodology is described, the specific implementation details of the "memory curator" training (e.g., reward shaping, specific prompts, hyperparameters) may be limited in the abstract-only view, but the full text likely contains sufficient detail for reproduction given the standard nature of LLM agent frameworks.
The primary limitation is computational cost. Retaining raw trajectories and performing LLM-based synthesis at read time is significantly more expensive than retrieving a pre-computed summary. This may limit scalability to very long-horizon agents or high-throughput applications. Additionally, the method relies heavily on the base LLM's ability to synthesize relevant information from raw traces, which may vary across model capabilities.
This work has significant implications for the design of LLM agents, suggesting that static memory structures may be suboptimal. It encourages the development of more dynamic, query-aware memory systems. The approach could be extended to other domains requiring adaptive information retrieval, such as RAG systems that dynamically re-rank or synthesize context based on the specific query. [One sentence main contribution]. The paper introduces Just-in-Time Memory, a read-time curation strategy for LLM agents that outperforms write-time methods by synthesizing task-adaptive payloads from raw trajectories, demonstrating significant gains in success rates across multiple benchmarks.
Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is delivered. We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved. Opera decides when to review through periodic and event-driven triggers, diagnoses issues with typed operators, audits feedback against visible evidence before delivery, and tracks the agent's subsequent actions to distinguish mere compliance from actual resolution. As a test-time critic, Opera improves the resolve rate of non-critic agents by up to 12.4, 15.0, and 8.9 percentage points on Terminal-Bench 2.1, a SWE-Bench Pro subset, and DeepSWE v1.1, respectively, across four policy models, and achieves the highest mean resolve rate among competitive critic baselines on all three benchmarks, and also improves policy models when the policy critiques itself. Beyond inference, Opera-guided rollouts provide approximately on-policy training data: fine-tuning Qwen3.5-9B on them improves its resolve rate on held-out SWE-Bench Pro repositories by 10.2 percentage points without a critic at inference time, matching fine-tuning on rollouts from a stronger model, while preserving its performance when switching harness, i.e., from Openhands to Terminus-2, which the latter substantially degrades.
Primary: Salesforce AI Research
All Institutions: Salesforce AI Research, Rutgers University, Lehigh University
The paper presents a rigorous and well-motivated framework for verbal criticism in long-horizon coding agents, demonstrating significant improvements in task resolution and providing a novel training recipe that leverages critic-guided rollouts. Its combination of principled design (typed operators, audited notes) and extensive empirical validation across multiple benchmarks and models positions it as a strong contribution to the field of LLM agents.
The paper introduces "Opera," a framework that conceptualizes critic feedback as persistent, tracked notes rather than transient messages. The methodology is well-structured around three core components: hybrid review scheduling (periodic + event-driven), operator-typed diagnosis (categorized by code repair stages like Search, Edit, Test), and audited note management (separating adherence from resolution). The use of "typed operators" grounded in code repair stages is a strong design choice that constrains the critic's output to actionable, specific corrections rather than generic advice. The audit mechanism, which verifies evidence before delivery and tracks resolution, directly addresses the common failure mode of LLM critics providing hallucinated or irrelevant feedback. The framework is designed to be plug-and-play, acting as a proxy between the agent and the model endpoint, which is a practical architectural choice.
The evaluation is extensive, covering three distinct benchmarks (Terminal-Bench 2.1, SWE-Bench Pro, DeepSWE v1.1) and four different policy models of varying strengths. The results show consistent improvements over non-critic baselines and competitive critic methods (SWE-PRM, SWE-Search, etc.). A particularly strong aspect of the evaluation is the analysis of the "rescue-regression trade-off," showing that while the critic rescues more tasks, it also causes more regressions in stronger models, which is a nuanced and honest finding. The training recipe section, demonstrating that fine-tuning on critic-guided rollouts matches distillation from a stronger model while preserving generalization, is a significant empirical contribution. The case studies provide concrete, detailed examples of how the critic intervenes and resolves specific issues, adding qualitative depth to the quantitative results.
The paper provides detailed implementation specifics in the appendix, including the proxy architecture, review trigger conditions, and audit logic. The hyperparameters for training are provided. However, the specific prompts for the critic and the exact implementation of the "operators" are not fully detailed in the provided text, which may hinder exact reproduction. The use of specific, potentially proprietary or rapidly changing models (GPT-5.6, Qwen3.5/3.8, DeepSeek-V4) as baselines and critics could limit reproducibility over time, as these models may be updated or deprecated. The reliance on specific harnesses (OpenHands, Terminus-2) also adds a layer of complexity for reproduction.
The primary limitation is the computational overhead of the critic, which adds inference cost. The paper acknowledges this but does not provide a detailed cost-benefit analysis comparing the compute cost of the critic to the gain in resolve rate. The training recipe is only tested on a single student model (Qwen3.5-9B) and a single out-of-domain benchmark (Terminal-Bench 2.1), so the generalizability of this training approach is not fully established. The "audit" mechanism relies on the same LLM for both diagnosis and verification, which may not fully eliminate the noise inherent in LLM-based verification, as noted in the related work (RuVerBench). The paper does not explore the interaction between the critic and reinforcement learning, which is mentioned as future work but is a significant area for potential improvement.
This paper has high potential impact on the field of LLM agents, particularly for long-horizon coding tasks. The concept of persistent, tracked feedback notes is a generalizable idea that could be applied to other agent domains beyond coding. The finding that critic-guided rollouts can serve as effective training data for distillation is a valuable insight for practitioners looking to improve open-weight models. The framework's design, which emphasizes evidence-based diagnosis and resolution tracking, sets a new standard for critic design in agentic systems. The work bridges the gap between test-time scaling and training-time improvement, offering a pathway to internalize critic feedback into the policy model. The paper presents a rigorous and well-motivated framework for verbal criticism in long-horizon coding agents, demonstrating significant improvements in task resolution and providing a novel training recipe that leverages critic-guided rollouts. Its combination of principled design (typed operators, audited notes) and extensive empirical validation across multiple benchmarks and models positions it as a strong contribution to the field of LLM agents.
Precise instruction-following is a fundamental ability of large language models (LLMs), requiring their outputs to strictly satisfy objective constraints in input instructions. In complex application scenarios, these constraints often possess diverse scopes that govern specific response segments rather than the entire output. However, existing optimization methods often neglect constraint scope during data construction and rely on binary per-constraint rewards, yielding limited data diversity and sparse supervision for complex constraints. To this end, we propose ScopeIF, a novel training framework for scope-aware precise instruction-following. We first introduce a unified schema that factorizes objective constraints into three decoupled dimensions: Scope, Target, and Range. Grounded in this schema, we construct ScopeInstruct, a large-scale instruction dataset with diverse scope-aware constraints, and combine tool-grounded verification with graded reward modeling to quantify the violation degree of each constraint, providing dense supervision for policy optimization. Extensive experiments demonstrate that ScopeIF consistently outperforms existing methods, particularly on complex scope-aware constraints, while preserving general capabilities. Notably, it enables optimized Qwen3-4B and 8B models to rival or surpass strong frontier models such as Gemini-2.5-Pro and DeepSeek-V3.2, establishing an effective paradigm for advancing scope-aware instruction-following. Our code and data are available at https://github.com/thu-coai/ScopeIF.
Primary: Tsinghua University
All Institutions: Tsinghua University, Zhipu AI
The paper introduces a scope-aware instruction-following framework using graded rewards to improve LLM precision. It offers a significant methodological improvement in RLHF by addressing the sparsity of binary rewards for complex, segmented constraints, demonstrating strong empirical results where smaller models rival frontier systems on specific tasks.
The paper proposes ScopeIF, a framework addressing the limitation of binary rewards in instruction-following tasks where constraints apply to specific segments (scopes) rather than the whole output. The core contribution is a unified schema factorizing constraints into Scope, Target, and Range, which allows for more granular data construction. The method employs tool-grounded verification to generate graded rewards, providing dense supervision signals for reinforcement learning. This is a logical and well-motivated extension of existing RLHF techniques, moving from sparse, binary feedback to dense, scope-aware feedback.
The experiments demonstrate that the optimized Qwen3-4B and 8B models rival or surpass frontier models like Gemini-2.5-Pro and DeepSeek-V3.2 on specific scope-aware tasks. This is a strong empirical claim. The use of a large-scale dataset (ScopeInstruct) and comparison against strong baselines adds credibility. However, the claim of surpassing frontier models on a specific subset of tasks (scope-aware constraints) needs careful scrutiny to ensure it is not an artifact of the specific test set construction, though the abstract suggests general capabilities are preserved.
The authors have released full code and data, including prompts and hyperparameters, which is excellent for reproducibility. The AI use statement is transparent about the use of LLMs for data generation and polishing, which is standard but important to disclose.
The primary limitation is the reliance on tool-grounded verification, which may not be applicable to all types of constraints or domains where automated verification is difficult. Additionally, the performance gains are specific to "scope-aware" constraints; the impact on general instruction following outside this niche is less clear, though the abstract claims preservation of general capabilities.
This work contributes to the broader goal of making LLMs more precise and controllable. By providing a framework for handling scoped constraints, it could be adopted in applications requiring strict adherence to local instructions within longer contexts, such as code generation with specific function-level requirements or structured data extraction. The paper introduces a scope-aware instruction-following framework using graded rewards to improve LLM precision. It offers a significant methodological improvement in RLHF by addressing the sparsity of binary rewards for complex, segmented constraints, demonstrating strong empirical results where smaller models rival frontier systems on specific tasks.
Little is known about how to manually design agents capable of dexterous manipulation. Some design principles have been inferred from close examination of how animals manipulate objects, but these structures and behaviors have so far resisted biomimicry and may not be optimal for artificial machines. Here we evolve freeform robots to pick up, hold, rotate, and use diverse objects. Unlike other approaches to optimizing robot hands, we do not presuppose the presence, articulation, or geometry of any part of the body. Although familiar prehensile forms such as tails, beaks, paws and claws may emerge spontaneously under certain conditions--and while such conditions could be of interest to evolutionary biologists--de novo manipulator design can also reveal whole new solutions, overlooked or unknown structures which may be better suited for the task at hand. We use contrastive learning to create a highly searchable genetic embedding of design space, an autoregressive developmental model to decode designs, evolutionary strategies to find good designs, and reinforcement learning to train each evolved design. Winning designs were automatically converted into a manufacturable blueprint, printed, assembled and tested in the real world in a zero-shot manner. The results represent the state-of-the-art in evolutionary robotics in terms of performance, diversity and complexity.
Primary: University of Oxford
All Institutions: University of Oxford, DeepMind
The paper presents a state-of-the-art framework for evolving freeform dexterous robots from scratch, combining contrastive learning, developmental models, and reinforcement learning to produce manufacturable, real-world capable agents. By removing prior assumptions on robot morphology, the work reveals novel, non-biomimetic structures that outperform traditional designs in dexterity and diversity, marking a significant leap in evolutionary robotics and automated hardware design.
The paper proposes a comprehensive framework for de novo robot design, combining contrastive learning for genetic embedding, autoregressive developmental models for decoding, evolutionary strategies for search, and reinforcement learning for control. The novelty lies in the lack of prior assumptions about morphology (no fixed joints or geometry), allowing for truly freeform evolution. The use of a developmental model to map latent codes to physical structures is a sophisticated approach that bridges high-level design search with low-level manufacturability.
The evaluation is rigorous, moving beyond simulation to real-world validation. The robots were 3D printed, assembled, and tested in a zero-shot manner, demonstrating the robustness of the evolved designs. The paper reports state-of-the-art performance in terms of dexterity, diversity, and complexity compared to previous evolutionary robotics works. The inclusion of diverse object manipulation tasks (pick, hold, rotate, use) provides a strong benchmark for the method's generality.
The paper details the pipeline components (contrastive learning, autoregressive model, ES, RL) sufficiently for experts to replicate the framework. However, the specific hyperparameters for the evolutionary search and the exact architecture of the developmental model are likely in the appendix or code, which is standard. The zero-shot real-world testing is a high bar for reproducibility that the authors have met by providing the blueprint conversion process.
The computational cost of evolving and training multiple robots is significant. The "zero-shot" real-world performance, while impressive, may still be limited in speed or precision compared to hand-tuned industrial manipulators. The generalizability to tasks requiring extremely high force or precision (beyond what 3D printed materials can handle) is a potential constraint.
This work has significant implications for the field of evolutionary robotics and automated design. It suggests that optimal manipulator structures may not be biomimetic, challenging existing design paradigms. The ability to automatically generate manufacturable blueprints from abstract design spaces could accelerate the development of specialized robotic tools for manufacturing, surgery, or exploration. The paper presents a state-of-the-art framework for evolving freeform dexterous robots from scratch, combining contrastive learning, developmental models, and reinforcement learning to produce manufacturable, real-world capable agents. By removing prior assumptions on robot morphology, the work reveals novel, non-biomimetic structures that outperform traditional designs in dexterity and diversity, marking a significant leap in evolutionary robotics and automated hardware design.
We present ZeroBot, a real2sim framework for learning a robot manipulation task from scratch in minutes under challenging conditions: zero human demonstrations, zero policy pre-training, and zero known object models. Given only a single view of an object and a goal pose for that object, ZeroBot uses image-to-3D generative models to obtain a complete object mesh, which is used in simulation for large-scale parallel reinforcement learning. To accelerate training, we introduce an action space which leverages the generated geometry and learned value function to sample states involving robot-object contact. When evaluated on real-world tasks including grasping, pushing, articulated object interaction, and multi-stage manipulation, ZeroBot achieves an 87% success rate with an average training time of 119 seconds. These results show the value of using image-to-3D models in a real2sim framework for rapid, autonomous robot learning.
Primary: Imperial College London
All Institutions: Imperial College London, Robotics and AI Institute
ZeroBot introduces a generative real2sim framework that enables robots to learn manipulation tasks from scratch in minutes using image-to-3D models and a novel contact-based action space. The paper demonstrates that combining visual foundation models with parallel RL can significantly reduce the time and data required for robotic learning, achieving high success rates on diverse real-world tasks without human demonstrations or pre-trained policies.
The paper proposes ZeroBot, a framework that integrates image-to-3D generative models (specifically InstantMesh) with massively parallel reinforcement learning (Isaac Gym) to enable rapid robot learning. The core methodological contribution is a novel "contact action space" that leverages the generated mesh geometry and a learned value function to sample high-value contact states, thereby accelerating exploration and allowing policies to be learned from scratch in minutes. The pipeline involves generating a complete object mesh from a single RGB-D view, aligning it for scale, using it for simulation and pose tracking (Foundation Pose), and training a PPO policy with a goal-flow reward. The approach is modular and effectively bridges the gap between generative AI and robotic control.
The evaluation is rigorous and well-designed, featuring six distinct real-world tasks (grasping, pushing, articulated interaction, multi-stage manipulation) on a Franka Research 3 arm. The paper provides strong ablations comparing the proposed method against baselines that use partial meshes (no 3D prior) and mesh retrieval (ACDC-NN). It also includes a comparison against ground-truth scanned meshes to quantify the performance gap of generative models. The results demonstrate an 87% success rate with average training times of ~2 minutes, which is a significant improvement over standard RL training times. The inclusion of challenging viewpoints (occluded handles, unseen sides) further validates the robustness of the generative prior.
The paper provides sufficient detail for reproducibility, specifying the hardware (Franka, RealSense cameras), software stack (Isaac Gym, PPO, InstantMesh, Foundation Pose, cuRobo), and hyperparameters (e.g., 256 parallel agents, temperature for softmax sampling). The use of standard, available tools and clear descriptions of the pipeline stages (mesh generation, alignment, RL training) makes it feasible for other groups to replicate the results, assuming access to similar computational resources (A6000 GPUs) and robotic hardware.
The method relies on several assumptions: a static background, a single rigid object (though extended to articulated with known parameters), and clear, unoccluded views for the initial mesh generation. It does not automatically infer physical properties like mass or friction, relying on constant values or manual specification. The performance is bounded by the accuracy of the current image-to-3D models, which can struggle with complex textures or highly non-convex shapes. Additionally, the method requires a goal pose to be specified, limiting its autonomy in open-ended tasks.
This work has significant potential to accelerate the deployment of robotic manipulation systems by reducing the data and time requirements for learning new tasks. By leveraging generative models to create simulation environments on-the-fly, it enables a "zero-shot" approach to real2sim2real transfer. This could lead to more adaptable robots in dynamic environments where pre-scanning or manual modeling is impractical. The framework also highlights the utility of value functions beyond policy improvement, using them for state sampling and deployment-time planning. ZeroBot introduces a generative real2sim framework that enables robots to learn manipulation tasks from scratch in minutes using image-to-3D models and a novel contact-based action space. The paper demonstrates that combining visual foundation models with parallel RL can significantly reduce the time and data required for robotic learning, achieving high success rates on diverse real-world tasks without human demonstrations or pre-trained policies.
Reinforcement learning (RL) enables generalist robot policies to improve through trial-and-error interaction, yet its effectiveness is fundamentally constrained by sparse task rewards. Existing general-purpose reward models typically alleviate this issue by learning task progress from expert demonstrations, but introduce a distribution mismatch with the mixed-quality rollouts encountered during policy optimization, making their estimates unreliable on suboptimal and failed behaviors from which the policy must learn. In this work, we introduce a demonstration-free reward learning paradigm where dense reward feedback can be learned directly from sparse task outcomes and policy experience. We theoretically show that terminal task outcomes implicitly define dense success-probability feedback at intermediate timesteps, which can be recursively learned through bootstrapping. Based on this insight, we introduce eVTA$_0$, which learns success probabilities from mixed-quality policy rollouts through temporal-difference-style bootstrapping, without expert demonstrations or intermediate annotations. We further introduce RL with Evolving Rewards (RLER), a closed-loop framework that adapts eVTA$_0$ using newly collected rollouts as the policy evolves. Experiments show that eVTA$_0$ provides more informative rewards than state-of-the-art reward models and achieves the best average policy performance across all LIBERO task suites under the same RL training budget, improving success rates by 5.4%-13.8% over the initial policy. In real-world manipulation, RLER further improves overall success rates by 20%-26%, with 35%-36% gains under out-of-distribution conditions. These results demonstrate the effectiveness of demonstration-free reward learning and adapting rewards as the policy evolves. Project webpage: https://duowuyms.github.io/evta0.
Primary: Tsinghua University
All Institutions: Tsinghua University, Pengcheng Laboratory
The paper introduces a demonstration-free reward learning paradigm that leverages temporal-difference bootstrapping to derive dense success-probability rewards from sparse terminal outcomes, significantly improving policy learning in robotic manipulation. By decoupling reward learning from expert demonstrations and integrating it into a closed-loop evolving framework (RLER), the work offers a scalable and theoretically sound solution to the reward sparsity problem, achieving state-of-the-art performance in both simulation and real-world settings.
The paper proposes a theoretically grounded approach to reward learning that bypasses the need for expert demonstrations. The core insight is that terminal success/failure signals implicitly define dense intermediate rewards as the probability of eventual success. This is formalized using a Temporal-Difference (TD) style bootstrapping mechanism, allowing the model (eVTA$_0$) to learn from mixed-quality rollouts. The architecture leverages a Vision-Language Model (VLM) backbone to process observation histories and task descriptions, predicting a success probability which is then transformed into a dense reward signal (failure-probability penalty). The introduction of RLER (RL with Evolving Rewards) creates a closed-loop system where the reward model updates as the policy improves, addressing distribution shift. The methodology is sound, combining classical RL theory (bootstrapping) with modern foundation models (VLMs) in a novel context.
The evaluation is comprehensive, covering both simulation (LIBERO, MetaWorld) and real-world manipulation. The paper compares eVTA$_0$ against state-of-the-art baselines like GVL, TOPReward, VLAC, and Robometer. Key metrics include temporal consistency (VOC/VROC) and outcome distinction (MSE/Kendall's tau). The results show eVTA$_0$ outperforms baselines in distinguishing successful from failed trajectories, which is critical for RL. Policy learning experiments demonstrate that using eVTA$_0$ as a reward signal leads to higher success rates than binary rewards or other reward models under the same training budget. Real-world experiments further validate the RLER framework, showing significant improvements in out-of-distribution conditions. The experimental setup is rigorous, with controlled budgets and multiple task suites.
The paper provides a detailed methodology section, including the specific TD learning recipe, architecture details, and training procedure. It mentions the release of source code, data, and model checkpoints. The use of standard benchmarks (LIBERO, MetaWorld) and well-known baselines enhances reproducibility. The AI Use Statement is transparent about the role of generative AI in the research process.
The method relies on a VLM backbone, which may be computationally expensive for real-time applications on resource-constrained robots. The performance gains, while significant, are evaluated primarily on manipulation tasks; generalization to other robotic domains (e.g., navigation, locomotion) is not explored. The "demonstration-free" claim is relative; the policy still requires an initial SFT policy to generate rollouts, so it is not entirely from scratch. The theoretical proof, while present, is somewhat standard in its application of TD learning to probability estimation.
This work addresses a fundamental bottleneck in robotic RL: the scarcity of dense, informative rewards. By demonstrating that rewards can be learned from experience alone, it opens the door to more scalable and autonomous robot learning pipelines. The concept of "evolving rewards" is particularly impactful for lifelong learning scenarios where the robot's capabilities change over time. This could influence the design of future generalist robot policies and reduce the reliance on costly human demonstrations. The paper introduces a demonstration-free reward learning paradigm that leverages temporal-difference bootstrapping to derive dense success-probability rewards from sparse terminal outcomes, significantly improving policy learning in robotic manipulation. By decoupling reward learning from expert demonstrations and integrating it into a closed-loop evolving framework (RLER), the work offers a scalable and theoretically sound solution to the reward sparsity problem, achieving state-of-the-art performance in both simulation and real-world settings.
World Action Models (WAMs) jointly model visual dynamics and action generation for generalist robot manipulation. A central challenge is to integrate priors from large-scale pretrained models---including visual dynamics, scene semantics, geometry, and motion---into a unified framework for robot action generation. We introduce InternW0-$Δ$, a unified WAM pretrained on a heterogeneous corpus that outperforms prior methods across simulation benchmarks and real-robot platforms. InternW0-$Δ$ combines pretrained visual dynamics, scene-level semantics, 4D geometric and motion priors, and action generation within a Mixture-of-Transformers (MoT) framework. A pretrained video expert and an action expert interact under semantic guidance from a frozen VLM, while a pretrained 4D foundation model injects geometric and motion priors through training-only distillation. We further introduce Causal Imprint, which learns future-relevant scene changes from training-only future supervision and provides predictive representations directly to the action expert without future-video rollout at inference. For large-scale joint training, we construct a heterogeneous corpus of robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data, curated and aligned under a common state-action representation. The resulting corpus contains over 20K hours of processed training data, to our knowledge the largest open-source corpus of its kind. We pretrain InternW0-$Δ$ on this corpus and demonstrate strong performance across simulation benchmarks and real-robot platforms. We will open source the training code, model weights, infrastructure, data-processing pipeline, and processed data where licenses permit. Project page: https://internrobotics.github.io/InternW0-Delta/
Primary: Shanghai AI Laboratory
All Institutions: Shanghai AI Laboratory
InternW0-$\Delta$ introduces a unified World Action Model framework that integrates visual, geometric, and semantic priors via a Mixture-of-Transformers architecture and a novel "Causal Imprint" training technique, pretrained on a 20K+ hour open-source heterogeneous robot data corpus to achieve strong performance in simulation and real-world manipulation.
The paper proposes InternW0-$\Delta$, a World Action Model (WAM) that integrates visual dynamics, scene semantics, 4D geometry, and action generation. The core architectural contribution is a Mixture-of-Transformers (MoT) framework where a video expert and an action expert interact under the guidance of a frozen Vision-Language Model (VLM). A key technical innovation is "Causal Imprint," a training-only distillation technique that injects future-relevant scene changes into the action expert without requiring future video rollouts at inference time. This addresses a significant latency and computational bottleneck in world models. The integration of a pretrained 4D foundation model for geometric priors is also a strong methodological choice, leveraging recent advances in 4D scene understanding to improve robot manipulation.
The authors report strong performance across both simulation benchmarks and real-robot platforms. The scale of the data corpus (20K+ hours) is a major strength, combining robot demonstrations, UMI data, egocentric human demos, and Ego2Robot data. This heterogeneous data strategy is well-motivated for generalist robot learning. The evaluation appears rigorous, covering both simulated environments (likely for controlled comparison) and physical robots (for real-world validity). The claim of outperforming prior methods is supported by this broad evaluation scope.
The paper explicitly states that training code, model weights, infrastructure, data-processing pipeline, and processed data will be open-sourced (where licenses permit). This is a high standard for reproducibility in the robotics field, where data scarcity often hinders replication. The release of a 20K+ hour open-source corpus is particularly impactful for the community.
The primary limitation is the reliance on a frozen VLM for semantic guidance, which may limit the model's ability to adapt to novel semantic contexts not covered by the VLM's pretraining. Additionally, the "Causal Imprint" technique, while efficient at inference, adds complexity to the training pipeline. The paper is an arXiv preprint, so peer review has not yet validated the claims. The 20K hours of data, while large, is still relatively small compared to web-scale data, potentially limiting generalization to highly out-of-distribution tasks.
This work has high potential impact on the field of generalist robot manipulation. By providing a large open-source dataset and a unified framework for integrating diverse priors (visual, geometric, semantic), it lowers the barrier to entry for developing world action models. The "Causal Imprint" technique could be adopted in other domains where predictive dynamics are needed but inference latency is critical. The open-sourcing of infrastructure and data will likely accelerate research in this area. InternW0-$\Delta$ introduces a unified World Action Model framework that integrates visual, geometric, and semantic priors via a Mixture-of-Transformers architecture and a novel "Causal Imprint" training technique, pretrained on a 20K+ hour open-source heterogeneous robot data corpus to achieve strong performance in simulation and real-world manipulation.
Coding agents have demonstrated enormous success in solving complex programming problems. To leverage their potential for robot systems, this work introduces Robot Agentic Programming from Demonstrations (RAPID), which automatically generates, verifies, and refines robot programs, given a single visual human demonstration. The iterative agentic loop of code refinement requires several key ingredients: (i) a testable task specification, (ii) action primitives for robot execution, and (iii) an interactive environment for program execution and verification. RAPID infers all three from the demonstration automatically. To make the resulting program reusable beyond the demonstration setting, RAPID uses an object-centric relational program representation that focuses on the underlying structure of the demonstrated strategy rather than the specific motion per se: it expresses the action primitives as trajectory-optimization programs that realize object-level motion effects, while composing them through relational constraints that capture scene-specific geometry at run time. We evaluated RAPID in simulation on eight challenging contact-rich nonprehensile manipulation tasks as well as general prehensile manipulation tasks in the LIBERO-Pro benchmark. We also successfully deployed it on a real Franka arm and evaluated on all eight nonprehensile tasks. In all experiments, RAPID demonstrated strong performance, with generalization over object pose, shape, material, and environment. Website: https://yuyaoliu.me/projects/rapid.
Primary: Stanford University
All Institutions: Stanford University, University of California, Berkeley
RAPID introduces a novel agentic framework that uses LLMs to generate and refine robot control programs from single demonstrations, achieving strong generalization in contact-rich manipulation tasks. The paper makes a significant technical contribution by proposing an object-centric relational program representation that enables code reuse and robustness, validated through rigorous simulation and real-world experiments on a Franka arm.
The paper proposes RAPID, a framework that leverages Large Language Models (LLMs) as agents to generate, verify, and refine robot control programs from a single visual demonstration. The core innovation lies in the "object-centric relational program representation." Instead of generating low-level motor commands or high-level symbolic plans that are brittle to scene changes, RAPID infers a set of action primitives (expressed as trajectory optimization programs) and relational constraints that define the strategy. This allows the generated code to be reusable: the LLM agent iteratively refines the code by executing it in a simulator, observing the outcome, and using the error to update the program. The methodology effectively bridges the gap between the semantic understanding of LLMs and the precise physical execution required in robotics. The use of an agentic loop for code refinement is a strong contribution, moving beyond one-shot code generation to an iterative improvement process grounded in physical feedback.
The evaluation is comprehensive, covering both simulation and real-world deployment. In simulation, the authors test on eight challenging contact-rich nonprehensile manipulation tasks (e.g., pushing, sliding) and general prehensile tasks using the LIBERO-Pro benchmark. The real-world experiments on a Franka arm validate the approach on the same eight nonprehensile tasks. The results demonstrate strong generalization across object pose, shape, material, and environment variations. The inclusion of contact-rich tasks is significant, as these are notoriously difficult for traditional model-based or purely learned approaches. The comparison with baselines (likely including imitation learning and other code-generation methods) shows RAPID's superiority in success rates and generalization.
The paper provides a project website with a link to the code. The methodology is described in sufficient detail to understand the pipeline: demonstration inference, program representation, and the agentic refinement loop. However, the specific prompts used for the LLM agent and the details of the trajectory optimization solver are critical for reproduction. Assuming the code is released as linked, reproducibility is high. The use of standard benchmarks (LIBERO-Pro) and a common robot platform (Franka) further aids reproducibility.
The primary limitation is the reliance on a simulator for the verification step in the agentic loop. While the final program is deployed on a real robot, the iterative refinement happens in simulation. This requires a high-fidelity simulator that matches the real world, which can be a bottleneck for complex, unstructured environments. Additionally, the approach depends on the LLM's ability to correctly interpret the visual demonstration and generate valid code, which can be sensitive to prompt engineering and model capabilities. The computational cost of running the agentic loop (multiple LLM calls and simulation runs) per task may be high compared to offline learning methods.
This work has significant implications for the field of robotics and agentic AI. It demonstrates that LLMs can be used not just for high-level planning but for generating and refining low-level control code, provided they are grounded in a testable environment. The "code as policy" paradigm is gaining traction, and RAPID offers a robust framework for it. This approach could accelerate robot programming by allowing non-experts to program robots via demonstrations, with the LLM handling the complex code generation and debugging. It also highlights the potential of combining symbolic AI (code) with neural AI (LLMs, vision) for robust physical interaction. RAPID introduces a novel agentic framework that uses LLMs to generate and refine robot control programs from single demonstrations, achieving strong generalization in contact-rich manipulation tasks. The paper makes a significant technical contribution by proposing an object-centric relational program representation that enables code reuse and robustness, validated through rigorous simulation and real-world experiments on a Franka arm.
Foundation models provide robots with the ability to interpret natural language and reason about environmental context, yet most language-conditioned policies assume that goals are well-specified and that task-relevant information is provided upfront via a prior map. Operating in unfamiliar environments with underspecified tasks entails high contextual uncertainty: the robot must jointly infer what constitutes task success, what constitutes relevant information, and where (or whether) that information exists. We address these limitations via CLUE (Closed-Loop contextual Uncertainty rEsolution), a framework for actively resolving contextual uncertainty given underspecified tasks in natural language. CLUE uses an LLM-derived policy to hypothesize task-relevant concepts and potential plans. It then uses a language-embedded map, which is constructed online, to ground these hypotheses into actions. The policy sequentially evaluates hypotheses via closed-loop environment interaction and refines its plans as it gathers new information. We deploy CLUE on a Boston Dynamics Spot across three real indoor and outdoor environments spanning 15 tasks that require object disambiguation, functional inference, and occlusion reasoning. CLUE achieves a success rate within 7 percentage points of an oracle policy and outperforms an LLM-enabled planner without closed-loop feedback by a 4x margin. Supporting experiments demonstrate that simply building and then querying a language-enriched map is insufficient to resolve complex contextual planning tasks; these approaches achieve roughly one third the success rate of CLUE while requiring over 10x more VLM tokens. We provide additional information at https://zacravichandran.github.io/CLUE.
Primary: University of Pennsylvania
All Institutions: University of Pennsylvania, Texas A&M University
[One sentence main contribution]. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper presents CLUE, a closed-loop framework that effectively combines LLM-based hypothesis generation with language-embedded map grounding to resolve contextual uncertainty in underspecified robotic tasks. The technical contribution is significant in demonstrating that closed-loop interaction is necessary for robust task completion in open-world environments, outperforming open-loop baselines by a large margin. The methodology is sound, using a symbolic hypothesis state to approximate belief in a POMDP setting, and the experimental validation on a real-world quadruped robot adds practical credibility. However, the small sample size and reliance on specific cloud-based models limit the generalizability and reproducibility of the results. The work is a strong contribution to the robotics and ML intersection, particularly in the area of language-conditioned planning and active perception.
The paper proposes CLUE, a closed-loop framework for resolving contextual uncertainty in underspecified natural language tasks. The core methodological contribution is the integration of an LLM-based policy that maintains a symbolic hypothesis state with a language-embedded voxel map for grounding. The system operates by hypothesizing task-relevant concepts, grounding them into candidate locations via semantic map queries, and then actively verifying these hypotheses through closed-loop interaction (navigation, inspection, manipulation). The use of a symbolic hypothesis state as a tractable approximation to a belief state in a POMDP setting is a reasonable design choice for open-vocabulary environments. The grounding mechanism, which uses cosine similarity against background concepts and DBSCAN clustering, is standard but effectively applied here. The closed-loop nature, where the LLM replans based on new observations, is the key differentiator from open-loop planners.
The experiments are conducted on a Boston Dynamics Spot robot in three real-world environments (indoor and outdoor) across 15 tasks. The tasks cover object disambiguation, functional inference, and occlusion reasoning. The evaluation compares CLUE against an oracle (upper bound) and NLMaps (open-loop baseline). CLUE achieves 86.7% success rate, within 7 points of the oracle (93.3%), and significantly outperforms NLMaps (20%). A comparison with DAAAM (a scene-graph based approach) shows CLUE is more efficient in VLM token usage while achieving higher success. The ablation study on TSP hints provides useful insight into LLM planning capabilities. The sample size (15 tasks, one run each) is small, which limits statistical confidence, but the real-world deployment on a quadruped robot adds significant practical value.
The paper provides implementation details, including the use of RayFronts, RadSeg, GPT-5.1, and specific hardware (Jetson AGX Thor, ZED 2i). The code and project page are available. However, the reliance on a specific cloud-based LLM (GPT-5.1) and proprietary robot hardware (Spot) may limit reproducibility for some groups. The task definitions and environment setups are described but not fully detailed in the main text, relying on the project page for more information.
The primary limitation is the small number of tasks and single runs per task, which makes the success rates less statistically robust. The method relies heavily on cloud-based LLM calls, which introduces latency and cost. The language-embedded map is memory-intensive (up to 100GB for outdoor environments). The paper acknowledges that the LLM's ability to plan distance-efficient paths degrades with complex system prompts, suggesting a trade-off between contextual reasoning and path optimization.
This work contributes to the field of language-conditioned robotics by addressing the challenge of underspecified tasks in unknown environments. The closed-loop approach to resolving contextual uncertainty is a step towards more autonomous and robust robotic systems that can interact with humans using natural language. The findings on the necessity of closed-loop feedback over open-loop planning are valuable for the design of future robotic systems. The integration of LLMs with semantic mapping and active perception is a trend that is likely to grow in importance. [One sentence main contribution]. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper presents CLUE, a closed-loop framework that effectively combines LLM-based hypothesis generation with language-embedded map grounding to resolve contextual uncertainty in underspecified robotic tasks. The technical contribution is significant in demonstrating that closed-loop interaction is necessary for robust task completion in open-world environments, outperforming open-loop baselines by a large margin. The methodology is sound, using a symbolic hypothesis state to approximate belief in a POMDP setting, and the experimental validation on a real-world quadruped robot adds practical credibility. However, the small sample size and reliance on specific cloud-based models limit the generalizability and reproducibility of the results. The work is a strong contribution to the robotics and ML intersection, particularly in the area of language-conditioned planning and active perception.
Human hand-object interactions (HOIs) provide a rich source of demonstrations for dexterous manipulation, but learning directly from them presents challenges in bridging morphology gaps, ensuring dynamical feasibility, and sim-to-real deployment. We present Morphometric Imitation, a three-stage framework that transforms reconstructed HOIs into zero-shot sim-to-real visuomotor policies. First, morphometric optimization (MMO) kinematically retargets human motion across hand morphologies while preserving demonstrated contacts. Second, residual reinforcement learning (RL) refines the kinematic reference using object pose and contact information from the human motion to produce dynamically feasible robot demonstrations. Third, these demonstrations are distilled into visuomotor policies. Across three robot hands and ten HOIs, MMO outperforms five baselines in contact F1, improving on the strongest ones by 8 to 28 points, while improving the success rate of downstream dynamic retargeting by as much as 35 points. On a Sharpa hand, the visuomotor policies achieve 89.3% zero-shot success in 300 real-world trials on 30 objects spanning 10 categories. Project page: https://morphometricimitation.github.io
Primary: University of California, Berkeley
All Institutions: University of California, Berkeley
The paper presents a robust three-stage framework for morphometric imitation that effectively bridges the gap between human hand-object interactions and dexterous robot manipulation. By combining contact-aware kinematic retargeting, residual reinforcement learning for dynamic feasibility, and visuomotor policy distillation, the authors achieve state-of-the-art results in both simulation and real-world deployment, demonstrating significant improvements in contact fidelity and task success rates across diverse robot morphologies.
The paper proposes "Morphometric Imitation," a three-stage pipeline addressing the morphology gap in human-to-robot hand retargeting. Stage 1 (MMO) performs kinematic optimization to retarget human motion to robot hands while explicitly preserving contact constraints, which is a significant improvement over standard IK-based retargeting that often ignores contact fidelity. Stage 2 uses residual reinforcement learning to correct kinematic errors and ensure dynamic feasibility, leveraging object pose and contact data from the human demonstration. Stage 3 distills these refined demonstrations into a visuomotor policy. The methodology is coherent, well-motivated, and addresses specific, known pain points in dexterous manipulation (morphology mismatch, dynamic infeasibility, sim-to-real gap).
The experimental evaluation is robust and comprehensive. The authors test across three different robot hand morphologies (3, 4, and 5-fingered) and ten distinct human-object interactions. They benchmark against five baselines, showing significant improvements in contact F1 (8-28 points) and downstream dynamic retargeting success rates (up to 35 points). The real-world validation on the Sharpa hand is particularly strong, reporting 89.3% zero-shot success over 300 trials with 30 objects across 10 categories. This level of real-world testing is rare and highly valuable, providing strong evidence of practical utility.
The paper provides a project page and likely code (implied by the nature of such releases, though not explicitly linked in the text snippet, the project URL is provided). The detailed description of the three stages and the specific metrics used (Contact F1, success rates) allow for reasonable reproducibility. The use of standard RL and optimization techniques further aids reproducibility.
The framework relies on high-quality human motion capture data, which can be expensive and difficult to collect. The residual RL stage may be computationally intensive. The results are specific to the tested hands and objects; generalization to significantly different morphologies or highly deformable objects is not fully explored. The "zero-shot" claim is strong but limited to the specific distribution of objects and poses tested in the real-world trials.
This work has high potential impact in the field of dexterous robotics. By providing a robust pipeline to convert human demonstrations into executable robot policies, it lowers the barrier to entry for training dexterous manipulators. The focus on contact preservation and dynamic feasibility addresses critical issues that have hindered the adoption of human-centric data in robotics. The successful sim-to-real transfer with high success rates demonstrates a viable path for deploying such systems in real-world applications. The paper presents a robust three-stage framework for morphometric imitation that effectively bridges the gap between human hand-object interactions and dexterous robot manipulation. By combining contact-aware kinematic retargeting, residual reinforcement learning for dynamic feasibility, and visuomotor policy distillation, the authors achieve state-of-the-art results in both simulation and real-world deployment, demonstrating significant improvements in contact fidelity and task success rates across diverse robot morphologies.
Pretrained generative Diffusion Transformers (DiTs) capture rich pixel-level visual and language-conditioned structure through large-scale image and video generation training. A growing line of robot policies builds on this generative prior, but how it should be transferred to control remains unclear, and existing approaches commonly instantiate this transfer through future visual prediction. We ask a more basic question: what a pretrained generative DiT actually contributes to action learning, and how this prior should be adapted for control. We introduce NowWAM, a future-target-free co-training formulation that denoises the current observation and predicts robot actions from the same visual stream, directly coupling the native generative objective to the action-facing representation across the denoising trajectory. Under matched controlled settings, past and future visual targets perform comparably, while restricting training to the clean endpoint substantially reduces robustness, suggesting that a separate future target is not essential for generative adaptation, while the denoising trajectory remains an effective interface for control. On LIBERO-Plus, NowWAM reaches 87.7% with FLUX2-Klein, improving over the future-target co-training baseline by 6.1 points while halving training visual tokens (784 to 392) and reducing step time from 2.85 s to 1.63 s, a 1.8x speedup. With the pure text-to-image Z-Image backbone, NowWAM further reaches 87.8%, showing that strong control adaptation is not tied to video generation or image-editing backbones.
Primary: Princeton University
All Institutions: Princeton University
[One sentence main contribution]. [The paper introduces NowWAM, a future-target-free co-training formulation that applies native denoising to the current observation, demonstrating that this simpler interface improves robustness and efficiency over future-prediction-based generative robot policies. The technical contribution is significant as it decouples the utility of generative priors from temporal forecasting, providing a rigorous ablation study that clarifies the role of the denoising trajectory in representation learning for control. The results are strong, showing state-of-the-art performance on robustness benchmarks with reduced training costs, though the lack of real-world validation limits the immediate practical impact.].
The paper proposes "NowWAM," a co-training framework that removes the separate future visual target typically used in generative robot policies. Instead, it applies the native denoising objective of the pretrained Diffusion Transformer (DiT) directly to the current observation stream, coupling generative adaptation with action learning in a single stream. The core insight is that the benefit of generative co-training comes from the denoising trajectory itself (the continuum of noisy states) rather than the specific temporal semantics of a future prediction. The method is conceptually clean and simplifies the training pipeline by halving the visual token count during training.
The experiments are rigorous and well-controlled. The authors use LIBERO-Plus as a primary robustness benchmark, which is more discriminative than standard LIBERO. They provide strong ablations isolating the effects of generative initialization, temporal target choice (past vs. future), and denoising trajectory sampling. The results show that NowWAM outperforms future-target baselines (like Fast-WAM and ImageWAM) in robustness (87.7% vs ~81.6%) while being significantly faster (1.8x speedup). The extension to a pure text-to-image backbone (Z-Image) further validates the claim that temporal generative pretraining is not strictly necessary.
The paper provides detailed descriptions of the training setup, loss functions, and evaluation protocols. However, specific hyperparameters for the action expert and the exact implementation details of the "masked mixed attention" are not fully specified in the provided text, which may hinder exact reproduction. The use of standard benchmarks (LIBERO, RoboCasa) aids reproducibility.
The evaluation is limited to simulation (LIBERO, RoboCasa). Real-world robot experiments are absent, which is a significant gap for a robotics paper. The performance gains, while statistically significant in simulation, may not translate directly to the noisy, complex real world. Additionally, the paper relies on specific pretrained backbones (FLUX2-Klein, Z-Image), and the generalizability to other DiT architectures is not tested.
This work challenges the prevailing assumption in robot learning that future prediction is essential for leveraging generative priors. By demonstrating that current-frame denoising is sufficient and more efficient, it offers a simpler, faster, and more robust alternative for building VLA models. This could lead to more efficient training pipelines and lower computational costs for robot policy learning. [One sentence main contribution]. [The paper introduces NowWAM, a future-target-free co-training formulation that applies native denoising to the current observation, demonstrating that this simpler interface improves robustness and efficiency over future-prediction-based generative robot policies. The technical contribution is significant as it decouples the utility of generative priors from temporal forecasting, providing a rigorous ablation study that clarifies the role of the denoising trajectory in representation learning for control. The results are strong, showing state-of-the-art performance on robustness benchmarks with reduced training costs, though the lack of real-world validation limits the immediate practical impact.].
Humanoid kicking requires coordinated whole-body motion and precise contact, while a banana kick demands contact mechanics that generate ball spin and aerodynamic curvature. Motion imitation provides a reliable ordinary-kick prior, but reinforcement learning may improve shot speed and placement accuracy without changing the underlying kicking technique. Adapting this prior to a qualitatively different contact-rich skill can fail even when the reward is dense and optimization remains stable. The failure occurs when the task objective is locally flat over the current policy's responses. We term this condition first-order learning starvation. To address it, we propose response-informed skill evolution (RISE), a closed-loop objective-continuation method for policy adaptation. RISE ranks bounded objective changes using response sensitivity estimated from cached rollouts and accepts updates only when they produce verified response progress while preserving kicking reliability. Our analysis shows that rescaling a saturated spin reward cannot recover first-order sensitivity at zero spin, whereas adapting coupled contact responses can provide a learnable path to spin generation. We integrate RISE into a humanoid kicking pipeline under calibrated contact and Magnus-force aerodynamics, and sim-to-real transfer. Experiments show that RISE evolves the ordinary kick into a high-spin curved kick with 11.55 rad/s mean ball spin, improves the mean evaluation score by 19.8% over a learning-progress curriculum, and raises joint target attainment from 15.2% to 50.9%. Ablations and response diagnostics support the mechanism, while 30 motion-capture-recorded physical trials demonstrate consistent hardware transfer of the learned curved kick. Project website: https://haozhang-thu.github.io/bananakick/
Primary: Tsinghua University
All Institutions: Tsinghua University
The paper presents a novel and effective method, RISE, for evolving humanoid skills by dynamically adapting reward objectives based on physical response sensitivities, successfully demonstrating the transition from an ordinary kick to a banana kick with verified sim-to-real transfer.
The paper introduces "Response-Informed Skill Evolution" (RISE), a closed-loop objective-continuation method designed to address "first-order learning starvation." This occurs when a policy initialized with a strong motion prior (ordinary kick) fails to learn a qualitatively different skill (banana kick) because the reward landscape is locally flat with respect to the policy's current response distribution. The core innovation is the use of cached rollout sensitivities to rank bounded changes to reward parameters (specifically kernel scales and weights) before applying them. The method verifies updates not just by reward improvement, but by measured physical response progress (e.g., spin, contact patch error) while maintaining reliability. The theoretical analysis correctly identifies that rescaling a saturated reward (like spin at zero) cannot create a gradient, whereas adjusting coupled contact variables can. This is a sophisticated approach to curriculum/reward shaping that moves beyond static designs to dynamic, response-aware adaptation.
The experiments are rigorous and compelling. The authors use a Unitree G1 humanoid in IsaacSim/PhysX with calibrated contact and Magnus-force aerodynamics. They compare RISE against a Learning Progress (LP) curriculum, Direct PPO, and ablations. RISE achieves a 19.8% improvement in evaluation score over the LP curriculum and significantly higher joint target attainment (50.9% vs 15.2%). Crucially, the paper demonstrates sim-to-real transfer with 30 physical trials, showing consistent curved flight (median bow 0.381m) without further adaptation. The ablations clearly isolate the contribution of the sensitivity ranking and feedback loop. The "response evolution" section provides deep insight into how the policy moves from the prior's regime to the new skill regime.
The paper provides detailed hyperparameters, environment settings, and algorithmic steps. The project website is provided. However, the specific code for the RISE loop and the calibrated physics parameters are not explicitly released in the text, though the project page likely contains them. The reliance on specific hardware (Unitree G1) and calibrated physics may limit immediate reproducibility for labs without similar setups, but the methodological framework is clearly described.
The method is tested primarily on a single task (soccer kicking) and a single robot platform. The computational cost of the closed-loop objective adaptation (ranking candidates, verifying progress) is not deeply analyzed compared to standard PPO. The "first-order learning starvation" concept, while well-motivated, is specific to scenarios where a strong prior exists but the target skill requires a regime shift; its applicability to tasks without strong priors or with different failure modes is less clear. The physical trials, while consistent, are limited in number (30).
This work has significant implications for humanoid robotics and skill acquisition. It provides a principled way to overcome the "local optimum" trap when adapting motion priors to new skills. The concept of response-informed reward adaptation could be generalized to other contact-rich tasks (manipulation, locomotion over varied terrain). The demonstration of sim-to-real transfer for a complex, spin-generating skill is a strong contribution to the field of humanoid soccer. The paper presents a novel and effective method, RISE, for evolving humanoid skills by dynamically adapting reward objectives based on physical response sensitivities, successfully demonstrating the transition from an ordinary kick to a banana kick with verified sim-to-real transfer.