Week of September 20 – September 27, 2026
Attributing language model outputs to their internal computations is an open problem in interpretability. Existing methods, which use causal interventions, gradients, or learnable masks, either are infeasibly expensive or struggle to identify actual causally-important internal computations. We propose framing attribution as the problem of identifying nested subsets of internal components which minimise a downstream loss. To learn this task, we introduce Matryoshka Attribution (MAttr), a mask learning method that parametrises the mask with a simple differentiable sigmoid top-$k$ operator. We supervise training over all sparsities simultaneously by randomising $k$ over training, resulting in a learned ordering of components by attribution score. MAttr achieves number 1 on the official leaderboard of the Mechanistic Interpretability Benchmark (Mueller et al., 2025); our method identifies sparse and task-transferrable circuits across varying circuit bases. As a practical application, we show that MAttr can be trained with reinforcement learning to identify weight changes responsible for downstream behaviours in LLM finetuning. We train MAttr on refusal judge scores and find that restoring $1\%$ of Llama 3.1 8B Instruct's weights to their base model state is sufficient to remove refusals while maintaining capabilities. We view MAttr as a successful formulation of interpretability into a learnable objective that we can tackle with gradient descent, and encourage future work along these lines.
Primary: Stanford University
All Institutions: Stanford University
Matryoshka Attribution (MAttr) is a novel mask-learning method that uses a differentiable sigmoid top-$k$ operator to learn a single, nested ordering of internal components for attribution, achieving state-of-the-art performance on the Mechanistic Interpretability Benchmark and enabling the identification of specific weight changes responsible for LLM behaviors like refusal through reinforcement learning.
The paper introduces Matryoshka Attribution (MAttr), a mask-learning method that frames attribution as identifying nested subsets of internal components minimizing a downstream loss. The core innovation is the use of a differentiable sigmoid top-$k$ operator to parametrize the mask, allowing the model to learn a single ordering of components that is valid across all sparsity levels (the "Matryoshka" property). This eliminates the need for separate mask training at each sparsity level and avoids the instability and hyperparameter sensitivity associated with hard concrete distributions or straight-through estimators. The method is grounded in causal abstraction, defining "soft interchange interventions" that interpolate between base and source states. A significant theoretical contribution is the demonstration that MAttr recovers Integrated Gradients in the limit of one training step, providing a rigorous bridge between gradient-based and mask-learning attribution methods. The extension to parameter-level attribution via Reinforcement Learning (GRPO) is a novel application, allowing the method to identify weight changes responsible for specific behaviors (e.g., refusal) without requiring differentiable rewards.
The experimental evaluation is extensive and rigorous. MAttr achieves state-of-the-art performance on the Mechanistic Interpretability Benchmark (MIB), significantly outperforming baselines like Integrated Gradients, Interchange Interventions, and other mask-learning methods (DBM, Node Pruning). The authors introduce "MIB+" to test finer-grained bases (MLP neurons, SAE features) and new tasks, showing consistent improvements. The paper also introduces a "compactness" metric to evaluate circuit sparsity, showing MAttr excels at both faithfulness and compactness. The parameter-level attribution experiments on Llama 3.1 8B Instruct are particularly compelling, demonstrating that restoring only 1% of weights to the base model state removes refusals while maintaining capabilities, a result with significant practical implications for model safety and auditing.
The paper provides a public GitHub repository with code. Hyperparameters for all baselines and MAttr are detailed in the appendices. The use of standard benchmarks (MIB) and open-source models (Llama, GPT-2, etc.) enhances reproducibility. The specific details of the RL training setup (GRPO) and reward functions are described, though the complexity of the RL setup may pose challenges for exact replication without the provided code.
The method relies on gradient descent, which may be computationally expensive for very large models or fine-grained bases, although it is parallelizable. The "Matryoshka" property assumes a nested structure of importance, which may not always hold for all types of internal computations. The RL-based parameter attribution is a new paradigm and may require careful tuning of the reward function to avoid reward hacking. The paper focuses primarily on language models; generalization to other modalities (e.g., vision) is mentioned in appendices but not the primary focus.
This paper has high potential impact on the field of interpretability. By providing a robust, learnable, and causally-grounded method for attribution, it enables more reliable identification of task-relevant circuits and weight changes. The application to removing refusals via weight restoration offers a new tool for auditing and potentially mitigating safety issues in LLMs. The unification of gradient-based and mask-learning perspectives may stimulate further theoretical work in attribution. The method's ability to handle non-differentiable objectives via RL opens up new avenues for interpreting complex, emergent behaviors. Matryoshka Attribution (MAttr) is a novel mask-learning method that uses a differentiable sigmoid top-$k$ operator to learn a single, nested ordering of internal components for attribution, achieving state-of-the-art performance on the Mechanistic Interpretability Benchmark and enabling the identification of specific weight changes responsible for LLM behaviors like refusal through reinforcement learning.
X-ray is medicine's most widely used imaging modality, yet remains among its least quantitative. Unlike volumetric modalities like CT or MRI, X-ray collapses 3D anatomy into a 2D projection, causing structures to overlap and anatomical boundaries to be ambiguous, even to experts. As a result, labeling X-ray databases for training general-purpose segmentation systems is impractical, leaving morphometric and functional X-ray analysis confined to narrow anatomical regions and applications. To this end, we present FleXray, a generalist model for anatomical segmentation across the entire body in clinical X-rays. Instead of curating large, manually annotated X-ray datasets, we build a scalable, physics-based generative X-ray data engine. Using existing 3D whole-body CT segmentation datasets and generative image-editing models, we simulate fully-annotated 2D X-rays with diverse appearances, physiological properties, and imaging geometries. Trained on these simulations, FleXray accurately segments 60 anatomical structures across unseen research datasets and in-the-wild X-rays. We further show that FleXray makes X-rays directly amenable to quantitative analysis, enabling automated measurements for disease grading, robust navigation during X-ray-guided interventions, and data-efficient learning of pathological targets. We release the model, code, a full-body X-ray segmentation dataset, and a local, easy-to-use browser-based tool at https://flexray.csail.mit.edu .
Primary: Massachusetts Institute of Technology (MIT)
All Institutions: Massachusetts Institute of Technology (MIT)
FleXray introduces a scalable, physics-based generative data engine that leverages 3D CT data and generative image editing to train a generalist model for full-body X-ray segmentation. The paper demonstrates that synthetic data, when carefully designed to cover anatomical, geometric, and appearance variation, can effectively bridge the gap to real-world clinical X-rays, enabling quantitative analysis and downstream clinical applications with high accuracy and robustness.
The paper proposes a highly effective data-centric approach to solving the problem of scarce, high-quality annotations for X-ray segmentation. The core innovation is a "physics-based generative data engine" that leverages existing 3D CT segmentation datasets to simulate 2D X-rays (Digitally Reconstructed Radiographs, DRRs). The methodology is robust, combining online analytic projections for geometric diversity with a generative image-editing model (Flux) to bridge the domain gap between synthetic DRRs and real clinical X-rays. The use of a language-conditioned diffusion model to add realistic artifacts (scatter, text, detector noise) while preserving anatomical structure is a clever and effective solution to the "sim-to-real" gap. The training objective is well-designed, handling multi-label overlap and partial annotations from heterogeneous sources.
The evaluation is comprehensive and rigorous. The model is tested on eight held-out datasets spanning various body parts (limbs, pelvis, spine, chest) and populations (adult, pediatric). It outperforms strong baselines like PAXray, TotalSegmentator2D, and FluoroSAM. The paper goes beyond standard segmentation metrics (Dice) to demonstrate downstream utility: automated Cobb angle measurement for scoliosis (achieving inter-rater variability levels), improved 2D/3D registration for surgical navigation, and data-efficient fine-tuning for new pathological targets. The ablation studies clearly isolate the contributions of generative enhancement and attenuation randomization, showing where each component is most beneficial.
High. The authors release the model, code, the full-body X-ray segmentation dataset, and a browser-based tool. The data engine components, including the specific prompts used for the generative editor and the filtering strategies, are detailed in the appendix. The use of public datasets (MOOSE, etc.) for training data generation further enhances reproducibility.
The model inherits biases from the CT source data, which may not fully represent all clinical X-ray scenarios (e.g., severe trauma, specific implants not present in CTs). The evaluation, while broad, is limited by the scarcity of public X-ray segmentation datasets with ground truth. The reliance on a large generative model (Flux) for data augmentation adds computational cost to the data preparation phase, though this is a one-time cost.
This work has significant potential to transform clinical X-ray analysis by making it quantitative and automatable. By providing a generalist model that segments 60 anatomical structures, it enables population-scale studies, automated disease grading, and improved surgical navigation. The release of the data engine and dataset will likely accelerate research in medical imaging by providing a scalable way to generate labeled training data for other modalities or tasks. FleXray introduces a scalable, physics-based generative data engine that leverages 3D CT data and generative image editing to train a generalist model for full-body X-ray segmentation. The paper demonstrates that synthetic data, when carefully designed to cover anatomical, geometric, and appearance variation, can effectively bridge the gap to real-world clinical X-rays, enabling quantitative analysis and downstream clinical applications with high accuracy and robustness.
Contact-rich precision insertion is a key manipulation skill in robotic assembly. Tight clearances make insertion more sensitive to alignment errors and prone to collisions and jamming, while variations in geometry and clearance across parts further complicate policy reuse. We present a reinforcement learning framework that trains insertion policies entirely in simulation for direct deployment without real-world demonstrations or policy fine-tuning. By combining target poses with compact three-dimensional fingertip force feedback, the policy learns to search for alignment and correct its motion despite errors in the estimated hole position. A decoupled gated reward coordinates alignment and insertion. Force-signal smoothing and state-independent standard deviations stabilize the learning process. The resulting policies perform real-world insertion across multiple hole geometries with a minimum nominal clearance of 0.02 mm and improve success while reducing peak contact forces under hole-position errors. Cross-clearance and cross-geometry evaluations further confirm policy generalization. The system achieved the first perfect score of 20/20 on ManipulationNet's peg-in-hole benchmark under its Human-in-the-Loop protocol, with fully autonomous insertion motions. A single policy trained only on a simulated hexagonal insertion task achieved an overall success rate of 95.0% across eight unseen real-world insertion tasks. These results show that learning entirely in simulation can yield precision insertion skills that can be deployed directly and reused across real-world tasks. The project website (https://mzhsoul.github.io/InsertAnything/) provides open-source simulation and real-robot experiment scripts, assets, and trained checkpoints.
Primary: University of Chinese Academy of Sciences
All Institutions: University of Chinese Academy of Sciences, Institute of Automation, Chinese Academy of Sciences, PaXini AI (Beijing) Co, School of Artificial Intelligence, State Key Laboratory of Multimodal Artificial Intelligence Systems
The paper presents a robust simulation-to-reality framework for precision robotic insertion that achieves state-of-the-art results on standardized benchmarks. It effectively combines force feedback with reinforcement learning to solve contact-rich manipulation tasks, demonstrating strong generalization capabilities across clearances and geometries without real-world fine-tuning.
The paper proposes a reinforcement learning framework for contact-rich precision insertion that relies entirely on simulation training for direct real-world deployment. The core methodological contributions are the use of compact 3D fingertip force feedback combined with target poses, a decoupled gated reward function that separates planar alignment, yaw alignment, and axial insertion, and specific stabilization techniques (EMA smoothing for force signals and state-independent standard deviations for the policy) to handle noisy contact data. The approach effectively addresses the sim-to-real gap by randomizing observation errors and dynamics, allowing the policy to learn robust correction strategies without real-world fine-tuning.
The experimental evaluation is rigorous and comprehensive. It includes ablation studies on reward design and stabilization techniques, comparisons with traditional control methods (impedance, hybrid force/position), and extensive generalization tests. The highlight is the real-world validation on the ManipulationNet benchmark, achieving a perfect 20/20 score, and the transfer of a single policy to eight unseen industrial tasks with 95% success. The inclusion of tight clearances (down to 0.02 mm) and diverse geometries (circular, square, hexagonal, L-shaped) demonstrates strong practical relevance.
The authors provide open-source simulation scripts, real-robot experiment scripts, assets, and trained checkpoints via the project website. The paper details the specific hardware (Franka Emika, Paxini sensors) and software stack (Isaac Lab, Factory), along with hyperparameters and randomization ranges in the supplementary material, which supports reproducibility for labs with similar robotic setups.
The method requires precise calibration of the target hole pose, which may not be available in all unstructured environments. The reliance on specific tactile sensor hardware (Paxini) limits immediate applicability to robots with different sensing capabilities. The generalization to "unseen" tasks is still within the domain of mechanical insertion/mating, and performance on highly deformable or non-rigid objects is not explored.
This work has significant implications for industrial automation, particularly in assembly tasks requiring high precision. By demonstrating that simulation-trained policies can handle tight clearances and generalize across geometries without real-world data, it reduces the cost and time associated with deploying robotic manipulation skills. The success on the ManipulationNet benchmark sets a new standard for autonomous precision assembly. The paper presents a robust simulation-to-reality framework for precision robotic insertion that achieves state-of-the-art results on standardized benchmarks. It effectively combines force feedback with reinforcement learning to solve contact-rich manipulation tasks, demonstrating strong generalization capabilities across clearances and geometries without real-world fine-tuning.
Leaderboards rank models by their average scores on benchmark items, and they are consulted repeatedly while the evaluation is still running. Existing confidence intervals for a model's rank control their error rate only if they are computed once, after a number of items chosen in advance. If they are recomputed as results arrive, and the evaluation stops once they look decisive, their error rate exceeds its nominal level. Anytime-valid methods keep their guarantees at all sample sizes simultaneously and hence under any stopping rule. They exist for the accuracy of one model, for one pair of models and for the set of models that may be best. For pairwise battles they also give ranks. None gives ranks when all models are scored on the same items, which makes their scores dependent. We construct rank confidence sequences: for every model, a set of ranks that contains its true rank, simultaneously for all models and at all times, at a chosen error level $α$, in finite samples. The construction combines betting e-processes, one for each ordered pair of models, with closed testing over the possible orderings of the models. It allows any dependence between the models' scores on an item. The method has two advantages. A leaderboard can be inspected after every item without inflating its error rate. The evaluation of each model can stop as soon as the question asked about it is answered, which saves compute. When results are examined only once, halfway through or later, little power is lost relative to fixed-sample methods. The paper quantifies these advantages in simulations and on public leaderboard data.
Primary: Stanford University
All Institutions: Stanford University
The paper presents a novel and rigorous statistical method for anytime-valid ranking of machine learning models, addressing the critical issue of repeated inspection in leaderboards. By combining e-processes with closed testing over orderings, it provides finite-sample guarantees that hold under any stopping rule, offering both statistical validity and practical compute savings, making it a significant contribution to the field of ML evaluation.
The paper introduces "rank confidence sequences," a rigorous statistical framework for evaluating model rankings in a sequential, anytime-valid manner. The core innovation lies in combining betting e-processes for pairwise comparisons with closed testing over weak orderings (permutations with ties). This approach allows the construction of confidence sets for the ranks of all models simultaneously, valid at any time step, without assuming independence between model scores on the same items (a common issue in benchmarking). The method handles both fixed finite benchmarks (with random item ordering) and superpopulation settings. The theoretical contribution is significant, providing finite-sample guarantees that hold under any stopping rule, which addresses the "peeking" problem inherent in real-time leaderboard monitoring. The use of integer programming for exact certification at scale is a clever computational addition, though the coNP-hardness of the general problem is acknowledged.
The experiments are well-designed to validate the theoretical claims. E1 demonstrates the failure of fixed-sample methods under repeated monitoring, showing a 30% error rate vs. the nominal 5%. E2 applies the method to real LLM leaderboard data (Open LLM Leaderboard), showing that while early certification is limited, it becomes robust as more items are revealed. E3 highlights the practical benefit of compute savings, showing that early stopping for specific models (e.g., top-3 certification) can save significant evaluation costs without compromising validity. E4 compares power against fixed-sample methods at pre-planned look times, showing comparable performance. The use of real-world LLM data adds substantial practical relevance.
The paper provides a GitHub link to the code. The algorithms are described in detail, including the betting strategies, the offset calculations, and the integer programming formulations. The specific parameters for the betting grid are provided. The experimental setups are described with references to public datasets. The level of detail suggests high reproducibility for researchers with statistical programming skills.
The method requires per-item scores in [0,1] and assumes a fixed set of models evaluated on common items. It does not handle adaptive item selection or models arriving over time. The computational cost, while manageable for typical leaderboard sizes (up to ~50 models), grows with the number of models and items, and the exact certification via integer programming can be complex to implement correctly. The method is primarily useful for ranking, not for estimating the absolute performance gap with high precision in early stages.
This work has high potential impact on the ML evaluation community. As leaderboards become more dynamic and expensive to run, the need for statistically valid, anytime-valid monitoring is critical. This paper provides a principled way to do so, potentially changing how benchmarks are reported and interpreted. It bridges the gap between statistical theory (e-processes, closed testing) and practical ML evaluation. The compute savings aspect is also a strong practical incentive for adoption. The paper presents a novel and rigorous statistical method for anytime-valid ranking of machine learning models, addressing the critical issue of repeated inspection in leaderboards. By combining e-processes with closed testing over orderings, it provides finite-sample guarantees that hold under any stopping rule, offering both statistical validity and practical compute savings, making it a significant contribution to the field of ML evaluation.
*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of *on-policy distillation (OPD)* across *weak-to-strong*, *same-base*, and *strong-to-weak* teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, $G$) rises approximately linearly in $d=\sqrt{\mathrm{KL}(π_θ\Vert π_{\mathrm{ref}})}$, the square root of token-level reverse KL divergence from the student initialization. In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own, so a compact RL expert can transfer capability to a much larger student via OPD. To estimate OPD outcomes, we fit *power laws* for how $G_{\mathrm{peak}}$ and the slope of the useful-transfer regime scale with student and teacher parameter counts and with teacher gold score. These laws show that peak gold score improves with teacher scale only up to roughly the student's scale, and that at a matched gold score smaller teachers transfer better, so a teacher's score alone does not define its supervision value. We also study the scaling effects of two OPD variants, bootstrapping weak-to-strong OPD, and the degree of on-policy supervision.
Primary: Unknown (Affiliations not explicitly listed in provided text, but authors Qinfeng Li, Wenqi Zhang, Guoqing Jiang, Liwei Chen, Xuanping Li, Zhiheng Qin, Yuntai Bao, Xuhong Zhang are associated with Alibaba Group / Qwen Team)
All Institutions: Alibaba Group (Inferred from author list and Qwen model usage)
The paper establishes predictive scaling laws for On-Policy Distillation, demonstrating that peak student performance and transfer rates can be accurately estimated from teacher/student scale and teacher quality, with surprising findings that smaller teachers often outperform larger ones at matched scores and that bootstrapping is ineffective compared to direct transfer.
The paper employs a rigorous empirical scaling analysis framework, adapting the methodology of reward model overoptimization studies (Gao et al., 2023) to the domain of On-Policy Distillation (OPD). The core methodological contribution is the characterization of OPD training dynamics as a function of the square root of token-level reverse KL divergence ($d$). The authors identify a "useful-transfer" regime where gold score increases linearly with $d$, followed by a noisy tail. They fit power laws to predict peak performance ($G_{peak}$) and transfer rates based on student/teacher scale and teacher quality. The inclusion of a theoretical derivation (Appendix) explaining why local KL geometry leads to linear transfer in $d$ adds significant depth, linking empirical observations to Fisher information geometry. The comparison between Vanilla-OPD and Delta-OPD (using policy shift contrasts) is well-motivated and clearly defined.
The experimental design is comprehensive, covering 25 teacher-student combinations across Qwen2.5 models (0.5B-14B) in weak-to-strong, same-base, and strong-to-weak configurations. The use of a single model family (Qwen2.5) is a limitation but allows for controlled isolation of scale effects. The evaluation on math reasoning (GSM8K/MATH) is standard but sufficient for this type of scaling study. Key findings include: (1) Peak student error is proportional to teacher remaining error, (2) Smaller teachers transfer better at matched scores (counter-intuitive and significant), (3) Bootstrapping does not improve over direct transfer from the smallest expert, and (4) Off-policy cold starts harm weak-to-strong transfer. The validation via leave-one-scale-out prediction is a strong methodological choice that demonstrates the predictive power of the fitted laws.
The paper provides detailed hyperparameters in the appendix and specifies the use of the `verl` framework. However, the reliance on a single random seed for all runs is a significant reproducibility weakness, although the authors justify this by citing standard practices in scaling law studies (Kaplan et al., Hoffmann et al.). The lack of seed variance estimation limits the confidence in the precise coefficients of the power laws, though the trends appear robust across the grid. The code and data are not explicitly linked in the provided text, which is a minor gap for immediate reproducibility.
The primary limitation is the restriction to a single model family (Qwen2.5) and a single task domain (math reasoning). The authors acknowledge that whether these scaling laws hold for other architectures, tasks, or post-training recipes is untested. The single-seed constraint means that stochastic variance in RL/OPD training is not captured, potentially masking instability in the "noisy tail" dynamics. The theoretical derivation assumes smoothness and differentiability that may not hold strictly in discrete token spaces or with clipping mechanisms, though the empirical fit is strong.
This paper has high practical impact for LLM practitioners. By providing predictive scaling laws for OPD, it enables engineers to estimate the outcome of distillation runs before committing significant compute resources. The finding that smaller teachers can be more effective than larger ones at matched scores challenges common assumptions about teacher quality and offers a cost-effective strategy for model family development. The negative result on bootstrapping is also valuable, preventing wasted effort on a seemingly intuitive but ineffective strategy. This work bridges the gap between theoretical scaling laws and practical post-training pipelines. The paper establishes predictive scaling laws for On-Policy Distillation, demonstrating that peak student performance and transfer rates can be accurately estimated from teacher/student scale and teacher quality, with surprising findings that smaller teachers often outperform larger ones at matched scores and that bootstrapping is ineffective compared to direct transfer.
The widely used Transolver family of neural operators is based on physics-attention, which softly assigns the points of an unstructured mesh to a small number of slices, applies self-attention among the resulting tokens, and broadcasts the result back to the points. We provide a comprehensive empirical and theoretical analysis to elucidate the mechanisms which are responsible for model performance. To this end, we perform careful ablations on a challenging suite of nine 3D fluid dynamics benchmarks to find that replacing token attention with a constant linear map does not affect the accuracy. Thus, Transolver does not need a Transformer at all. However, removing the global mixing (slicing/deslicing) or doing it only once leads to performance collapse. We leverage the theory of averaging neural operators to explain and corroborate our findings by showing that just slicing/deslicing, in conjunction with pointwise MLPs, already suffices for universal approximation of continuous operators and attention is redundant in this context. Finally, we provide a novel FlashAttention-style efficient implementation of the key slicing/deslicing module of Transolver. This flashslice kernel streams slice and deslice over the points without materializing the heavy slice-weight tensor, while reproducing the best available implementation to floating point error. At the same time, it leads to very significant memory and compute savings, particularly at large slice counts.
Primary: NVIDIA
All Institutions: NVIDIA
The paper demonstrates that the Transformer component in Transolver is redundant for accuracy, with the slicing/deslicing mechanism being the key driver of performance, supported by theoretical universality results and efficient kernel implementations. This work provides a critical understanding of neural operator architectures, guiding future design towards more efficient global mixing strategies without the computational overhead of self-attention.
The paper employs a rigorous ablation study combined with theoretical analysis to deconstruct the Transolver architecture. The methodology is sound, isolating the "physics-attention" mechanism into slicing/deslicing, token attention, and pointwise MLPs. The theoretical contribution leverages the theory of Averaging Neural Operators (ANO) to prove that the slicing/deslicing component, even with a constant core (no attention), is sufficient for universal approximation of continuous operators. This provides a strong mathematical foundation for the empirical finding that the Transformer component is redundant. The implementation of a FlashAttention-style kernel for the slicing operation is a significant technical contribution, addressing the memory bottleneck of materializing slice weights.
The experiments are extensive, covering nine challenging 3D fluid dynamics benchmarks, including industrial-scale aerodynamics (DrivAerNet++, SHIFT-SUV, SHIFT-Wing, DrivAerML). The evaluation protocol is careful, using matched step budgets and identical data pipelines. The results clearly demonstrate that removing token attention does not degrade accuracy, while removing the global mixing (slicing) causes performance collapse. The efficiency gains from the new kernel are substantial, showing significant memory savings and speedups, particularly at large slice counts.
The paper provides detailed descriptions of the ablation variants, training protocols, and kernel implementation. The use of standard datasets and clear metric definitions (relative L1 error) supports reproducibility. The code for the FlashSlice kernel is described in detail, though a direct link is not provided in the text snippet, the description is sufficient for implementation.
The theory is based on expressivity (universality) and does not address generalization or optimization dynamics. The bounds are not sharp. The findings are specific to the Transolver architecture and may not generalize to other neural operators without further analysis. The paper acknowledges that the cost of removing attention is empirically zero, but the theory only makes it unsurprising, not derived from first principles.
This paper has high impact on the scientific machine learning community by clarifying the fundamental mechanisms of a widely used architecture. It challenges the assumption that self-attention is necessary for global mixing in operator learning, potentially leading to more efficient and simpler models. The efficient kernel implementation will benefit practitioners working with large-scale unstructured meshes. The paper demonstrates that the Transformer component in Transolver is redundant for accuracy, with the slicing/deslicing mechanism being the key driver of performance, supported by theoretical universality results and efficient kernel implementations. This work provides a critical understanding of neural operator architectures, guiding future design towards more efficient global mixing strategies without the computational overhead of self-attention.
Transformers and state-space models (SSMs) are two prominent sequential learning architectures, yet their comparison remains largely empirical and existing theoretical analyses are typically task-specific or architecturally restricted. In this paper, we develop belief geometry, a unified analytical framework for comparing the representational capabilities of broad classes of attention and SSMs. Starting from a generalized formulation of in-context linear regression and using cumulative Bayes regret as our measure, we abstract three capabilities required by many sequential learning problems in our belief geometry: evidence assembly, belief maintenance, and addressing. We then study three cases of our formulation that isolate these capabilities and yield sharp architectural lessons: For belief maintenance, SSMs attain the optimal regret over stationary aggregation kernels; for positional assembly, SSMs have a memory advantage; and for content addressing, softmax attention has an exponential width advantage over sigmoid-selective SSMs. Experiments with LLaMA-type Transformers and Mamba-2 show that these architectural insights extend beyond our analytically tractable classes and linear-regression testbed.
Primary: The Ohio State University
All Institutions: The Ohio State University, RWTH Aachen University
Develops "belief geometry," a unified theoretical framework that rigorously compares the representational capabilities of attention and state-space models by isolating evidence assembly, belief maintenance, and addressing, yielding sharp architectural lessons such as the exponential width advantage of softmax attention for content addressing and the memory advantage of SSMs for positional assembly.
The paper introduces "belief geometry," a unified analytical framework to compare Transformers and State-Space Models (SSMs) in the context of in-context linear regression (ICLR). The methodology is rigorous, moving beyond empirical benchmarks to derive theoretical lower and upper bounds on cumulative Bayes regret. It decomposes sequential learning into three capabilities: evidence assembly, belief maintenance, and addressing. The authors define specific architectural classes (linear/softmax attention, fixed/selective SSMs) and prove sharp separations: SSMs are optimal for stationary belief maintenance (exponential kernels vs. uniform attention), SSMs have a memory advantage for positional assembly (convolution vs. context window), and softmax attention has an exponential width advantage for content addressing (selecting from retained tokens vs. storing candidates in state). The use of cumulative regret rather than terminal loss is a significant methodological improvement for analyzing sequential dynamics.
The experiments validate the theoretical predictions using practical architectures (LLaMA-type Transformers and Mamba-2). The authors test kernel alignment in single-layer models, showing that Transformers learn flatter kernels while Mamba-2 learns exponential ones, matching the theoretical optima. For the routing tasks (positional and content), they demonstrate that performance thresholds align with the theoretical resource requirements (e.g., convolution width for SSMs, context length for Transformers). The experiments effectively bridge the gap between the analytically tractable linear regression testbed and practical non-linear models, confirming that the architectural lessons (e.g., exponential width gap for addressing) hold in practice.
The paper includes a reproducibility statement and provides a link to an anonymous code repository. The experimental protocols, including hyperparameter sweeps and evaluation metrics, are detailed in the appendices. The use of synthetic tasks (Gaussian filtering, Beta-Bernoulli bandits, logistic bandits) ensures that the results are deterministic and reproducible without access to large-scale datasets.
The primary limitation is the reliance on linear regression and conjugate/non-conjugate bandit settings for the theoretical analysis. While the authors argue that the "belief geometry" extends to broader problems, the proofs are specific to these linear/quadratic loss structures. The "exponential width advantage" for attention in content addressing is a strong claim, but it relies on specific definitions of "addressing capacity" and compressed codebooks; real-world semantic addressing may not strictly follow this geometric separation. Additionally, the comparison focuses on representational capability (what the model *can* do) rather than optimization dynamics (what the model *learns* efficiently), though the kernel alignment experiments partially address this.
This paper provides a principled theoretical foundation for the ongoing debate between Transformers and SSMs. By identifying specific architectural mechanisms (softmax vs. selective transitions, context window vs. recurrent state) that confer advantages in different task regimes, it offers actionable insights for architecture design. The framework of "belief geometry" could be extended to other sequential tasks, such as language modeling or control, potentially guiding the development of hybrid architectures that leverage the strengths of both families. The finding that SSMs are inherently better at stationary belief maintenance while attention is better at content addressing helps explain empirical observations in long-context learning and associative recall. Develops "belief geometry," a unified theoretical framework that rigorously compares the representational capabilities of attention and state-space models by isolating evidence assembly, belief maintenance, and addressing, yielding sharp architectural lessons such as the exponential width advantage of softmax attention for content addressing and the memory advantage of SSMs for positional assembly.
Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but are widely viewed as harder because the chance constraint is nonconvex and depends on the full trajectory rather than a Bellman-linear expectation. In this paper, we reveal that this computational difficulty does not necessarily imply a higher statistical price. For tabular discounted CCMDPs with fixed bounded successor support and access to a certified planning oracle, we establish a model-based upper bound, with a matching lower bound up to logarithmic terms. Technically, our key idea is the \emph{Bellman distributional certificate}, which constructs a Bellman recursion for constraint violation probabilities before policy selection. The certificate can be reused across candidate policies; combined with shared row-wise reverse-KL confidence sets, it gives a policy-uniform trajectory-KL transfer without a union bound over policies or time--budget Bellman tables. For stochastic policies, we give a model-free variance-reduced policy-gradient algorithm with a finite-sample expected KKT-residual guarantee and independent validation of every accepted policy. Numerical experiments on synthetic CCMDPs and an IEEE 14-bus energy storage control benchmark illustrate the safety and mechanism behavior of the proposed algorithms.
Primary: Cornell University
All Institutions: Cornell University, University of Washington
The paper establishes a rigorous theoretical framework for learning chance-constrained MDPs by introducing Bellman distributional certificates that enable sample-efficient safety certification without union bounds, providing matching upper and lower bounds on sample complexity and a novel model-free policy gradient algorithm with finite-sample guarantees.
The paper introduces "Bellman distributional certificates," a novel theoretical framework for handling chance-constrained MDPs (CCMDPs). The core innovation is transforming the non-convex, trajectory-dependent chance constraint into a Bellman recursion over a discretized "remaining safety budget" state. This allows the use of standard model-based RL techniques (KL confidence sets) to certify safety with high probability without requiring a union bound over all possible policies or time steps. The method provides matching upper and lower bounds on sample complexity for deterministic policies, proving that the statistical cost of chance constraints is not inherently higher than expected-cost constraints under certain conditions (bounded successor support). Additionally, a model-free variance-reduced policy gradient algorithm is proposed for stochastic policies, offering finite-sample KKT-residual guarantees.
The experimental section is limited compared to the theoretical depth. It evaluates the method on a synthetic CCMDP and an IEEE 14-bus energy storage control benchmark. The results demonstrate that the Bellman-certified selector achieves better safety-performance trade-offs than a Markov-CMDP surrogate, particularly in reducing structural conservatism. However, the experiments are illustrative rather than exhaustive, lacking comparisons with state-of-the-art safe RL baselines (e.g., RCPO, PPO-Lagrangian) on standard continuous control benchmarks (MuJoCo, D4RL).
The paper provides detailed algorithmic descriptions and proofs in the appendix. However, specific hyperparameters, code availability, and implementation details for the "certified planning oracle" are not fully specified in the main text, making independent reproduction difficult without access to the authors' code. The reliance on a "certified planning oracle" as a black-box assumption limits immediate practical applicability.
1) The model-based guarantee is restricted to deterministic policies and assumes a fixed bound on successor support, which may not hold in high-dimensional continuous spaces. 2) The model-free result is local and may return "unresolved" if validation fails, lacking a global convergence guarantee. 3) The experimental validation is narrow, focusing on a single domain (energy storage) and synthetic tasks, without broad empirical validation on standard RL benchmarks. 4) The computational cost of the Bellman table grows with the discretization of the safety budget, which could be prohibitive for tight constraints.
This work provides a rigorous theoretical foundation for safe RL, addressing a critical gap in how probability-level safety constraints are handled. The "Bellman distributional certificate" concept could influence future work on risk-sensitive RL and constrained optimization. However, the strong assumptions (bounded support, deterministic policies for main bound) limit its immediate impact on practical, large-scale safe RL applications. The paper establishes a rigorous theoretical framework for learning chance-constrained MDPs by introducing Bellman distributional certificates that enable sample-efficient safety certification without union bounds, providing matching upper and lower bounds on sample complexity and a novel model-free policy gradient algorithm with finite-sample guarantees.
Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existing nonparametric ICL theory has largely focused on Euclidean domains or single-manifold models. To address this gap, we study the prediction problem under unknown local geometry, modeled by sample size-dependent mixtures of manifolds with heterogeneous dimensions, smoothness, and sampling masses. Under local separation and small-perturbation conditions, we establish a minimax lower bound capturing the aggregate difficulty of the components and construct an oracle tangent local-polynomial estimator with a matching upper bound. This estimator is connected to a structure-informed, two-stage softmax transformer with a geometric preconditioner and chartwise reduced local-polynomial solvers. The transformer achieves negligible approximation error relative to the minimax rate with logarithmic depth and polynomial size. Finally, we derive an in-context generalization bound for near empirical risk minimizers over this class. Together, these results identify conditions under which the resulting predictor exploits local geometry and attains the aggregate minimax rate.
Primary: Seoul National University
All Institutions: Seoul National University
[One sentence main contribution]. The paper establishes minimax optimality for Transformers in in-context learning on heterogeneous manifold mixtures by connecting oracle tangent local-polynomial estimators to a structure-informed Transformer architecture, providing a rigorous theoretical foundation for the model's ability to exploit local geometric structure.
The paper proposes a theoretical framework for analyzing Transformers in the context of in-context learning (ICL) on data with heterogeneous local geometry (mixtures of manifolds). The core methodological contribution is the derivation of minimax lower bounds for prediction under these complex geometric conditions and the construction of an oracle estimator (tangent local-polynomial) that matches these bounds. Crucially, the authors demonstrate that a specific class of Transformers (two-stage softmax with geometric preconditioners) can approximate this oracle estimator with negligible error, thereby establishing that Transformers can achieve minimax optimality in this setting. The approach is highly theoretical, relying on statistical learning theory and differential geometry rather than empirical model training.
The provided text is primarily theoretical, focusing on proofs and bounds. There is no mention of extensive empirical experiments, benchmarks, or real-world dataset evaluations in the abstract or the visible text fragments. The "experiments" are likely mathematical verifications of the bounds and the approximation capabilities of the proposed Transformer architecture. As a pure theory paper, its value lies in the rigor of the proofs rather than empirical performance metrics.
Reproducibility in the context of this paper refers to the verifiability of the mathematical proofs. The paper is 63 pages long, suggesting detailed appendices with full proofs. However, without access to the code or specific simulation scripts (if any exist to validate the theoretical bounds empirically), reproducibility is limited to the mathematical derivation. The lack of a provided code repository in the extracted information makes empirical validation difficult for external researchers.
The primary limitation is the gap between theory and practice. The conditions required for the minimax optimality (local separation, small-perturbation conditions, specific mixture structures) may be restrictive and not always satisfied by real-world high-dimensional data. Furthermore, the "oracle" nature of the estimator and the specific "structure-informed" Transformer design may not be directly implementable in standard large language model architectures without significant architectural modifications. The paper does not appear to provide empirical evidence that standard Transformers naturally learn these geometric preconditioners.
This paper contributes to the foundational understanding of why Transformers are effective for ICL, extending the theory beyond simple Euclidean or single-manifold settings. It provides a rigorous justification for the use of local polynomial approximations within Transformer layers. While the immediate practical impact on engineering may be limited, it offers valuable insights for the design of future architectures that explicitly account for data geometry, potentially influencing the development of more efficient and robust foundation models. [One sentence main contribution]. The paper establishes minimax optimality for Transformers in in-context learning on heterogeneous manifold mixtures by connecting oracle tangent local-polynomial estimators to a structure-informed Transformer architecture, providing a rigorous theoretical foundation for the model's ability to exploit local geometric structure.
Density functional theory (DFT) strikes a practical balance between accuracy and computational cost in many problems of computational chemistry and materials science. However, many DFT calculations are limited by fixed atom-centered basis sets, which dictate how accuracy and cost scale with system size. We propose Gaussian Splatting for Density Functional Theory (GS-DFT), which represents molecular orbitals as a cloud of Gaussians whose positions, shapes, and mixing coefficients are optimized jointly by gradient descent to minimize the energy without training data. Conceptually, GS-DFT is 3D Gaussian splatting with the renderer replaced by quantum mechanics. We introduce two key solver components: adaptive density fitting with screening for efficient evaluation of two-electron integrals, and a regularized differentiable orthogonalization of the molecular orbitals. Empirically, the optimized basis reaches the accuracy of the largest conventional basis sets with a fraction of the parameters, converging systematically in energy, density, and nuclear forces. At equal parameter count, it captures the stretched-bond and anion physics that fixed bases only recover with specialized basis augmentation. The resulting solver exhibits quadratic peak memory scaling in the cloud size, allowing us to simulate systems of up to 2,742 atoms (10,406 electrons) without any modifications at triple-zeta scale using a single four-GPU node.
Primary: Unknown
All Institutions: Unknown
The paper introduces GS-DFT, a novel method using 3D Gaussian splatting to represent molecular orbitals in DFT, achieving high accuracy with fewer parameters and enabling large-scale simulations on GPU clusters. This represents a significant methodological advance in computational chemistry, combining differentiable rendering concepts with quantum mechanical solvers to overcome the limitations of fixed basis sets.
The paper proposes GS-DFT, a method that replaces fixed atom-centered basis sets in Density Functional Theory with a cloud of 3D Gaussian splats. The positions, shapes, and mixing coefficients of these Gaussians are optimized via gradient descent to minimize the electronic energy. This is a significant conceptual shift, framing the basis set optimization as a differentiable rendering problem where the "renderer" is the quantum mechanical solver. The introduction of adaptive density fitting with screening for two-electron integrals and a regularized differentiable orthogonalization scheme are critical technical contributions that enable the stability and efficiency of this optimization. The approach is end-to-end differentiable and does not require training data, distinguishing it from standard neural network surrogates for DFT.
The experimental results are impressive, claiming that the optimized basis reaches the accuracy of the largest conventional basis sets with a fraction of the parameters. The paper reports systematic convergence in energy, density, and nuclear forces. A key highlight is the ability to simulate systems of up to 2,742 atoms (10,406 electrons) on a single four-GPU node, demonstrating quadratic peak memory scaling. The comparison to fixed bases shows that GS-DFT captures stretched-bond and anion physics that typically require specialized basis augmentation in traditional DFT. However, the lack of specific benchmark details (e.g., which molecules, which functionals, comparison to specific state-of-the-art DFT codes like ORCA or Gaussian) in the provided abstract limits the full verification of these claims.
The paper describes the solver components (adaptive density fitting, regularized orthogonalization) which are crucial for reproducibility. However, without access to the full code or detailed hyperparameter settings (e.g., number of Gaussians, optimization schedule, regularization strength), independent reproduction may be challenging. The reliance on a specific hardware setup (four-GPU node) for the large-scale simulations also poses a barrier for smaller groups.
The primary limitation is the computational cost of the gradient descent optimization for the basis set itself, which may be prohibitive for very large systems despite the efficient solver. The method is currently demonstrated on molecular systems; its applicability to periodic boundary conditions (solids) is not discussed. The "quadratic peak memory scaling" is a significant improvement over cubic scaling of traditional DFT, but it still limits the system size compared to linear-scaling DFT methods. The lack of comparison to other learned basis set methods or neural network potentials is a gap.
This work has the potential to significantly impact computational chemistry and materials science by providing a flexible, high-accuracy basis set representation that scales better than traditional methods. It bridges the gap between differentiable rendering techniques (3D Gaussian Splatting) and quantum mechanics, opening new avenues for differentiable physics. The ability to simulate larger systems on standard GPU hardware could democratize high-accuracy DFT calculations. The paper introduces GS-DFT, a novel method using 3D Gaussian splatting to represent molecular orbitals in DFT, achieving high accuracy with fewer parameters and enabling large-scale simulations on GPU clusters. This represents a significant methodological advance in computational chemistry, combining differentiable rendering concepts with quantum mechanical solvers to overcome the limitations of fixed basis sets.
Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for each task, a relevant subsequence is extracted from the full history and a dedicated model trained on it. In production it hits two bottlenecks. First, even after filtering, a single-task sequence stays extremely long: content-interest summarization reads several hundred items per user, tens of thousands of tokens once serialized as prompt text. Second, profiles are refreshed routinely: a billion users weekly, roughly 100K QPM in aggregate, which under a fixed GPU budget sets a hard throughput floor. Compression is therefore mandatory, yet truncation or coarse compression can silently distort the profile, introducing four hallucination types (fabrication, omission, date misattribution, broken logic) that, with no way to evaluate the compressed representation itself, surface only as diffuse degradation in downstream metrics. We present KuaFu, a unified behavior-compression layer whose minimal unit is one behavior item. A two-axis projector compresses each item into 2-4 tokens of width 128-256 (about 10x along the token axis, 20x along width; per-item cache 10 KB to 0.5 KB), with fidelity-oriented four-stage training and layered intermediate evaluation. Across four production profiling tasks it matches or exceeds uncompressed single-task production models on all five headline metrics, raises per-GPU throughput by 37%-350%, and saves 190 GPUs. On public benchmarks it nearly always beats prior compressors at the same compression ratio (up to +17.7 EM on out-of-domain MRQA); on RecBench, a 4B model surpasses its 8B counterpart by 1.90 points. KuaFu has run on the Tencent advertising and recommendation platform for ten months, lifting overall GMV by 1.37%.
Primary: Tencent
All Institutions: Tencent
KuaFu introduces an item-level, two-axis compression layer for user behavior sequences that enables billion-scale LLM-based profiling with significant cost and latency reductions. The paper demonstrates that by treating individual behavior items as independent, cacheable units and employing a multi-stage training recipe including hallucination-aware RL, it is possible to maintain or improve downstream task performance while drastically reducing computational overhead, validated by a 1.37% GMV lift in a major production environment.
The paper proposes KuaFu, a unified behavior-compression layer that treats individual behavior items as the minimal unit of compression. The core architectural contribution is a two-axis projector that compresses item embeddings along both the token axis (reducing sequence length) and the width axis (reducing dimensionality), allowing for significant storage and compute savings (10KB to 0.5KB per item). The training methodology is sophisticated, involving a four-stage curriculum: (1) reconstruction pre-training with a length-based curriculum to ensure convergence on long sequences, (2) post-training for compressed QA, (3) co-training of the compressor and decoder, and (4) hallucination-aware reinforcement learning (DAPO) specifically designed to penalize fabrication, omission, date misattribution, and broken logic. This multi-stage approach, particularly the explicit handling of hallucination types in the RL reward function, is a strong methodological contribution for LLM-based user modeling.
The experimental evaluation is extensive and convincing, combining offline benchmarks with large-scale online A/B testing. Offline, KuaFu outperforms state-of-the-art compression methods (SAC, EPL, ICAE) on MRQA benchmarks, showing superior fidelity at high compression ratios. On RecBench, the compressed 4B model outperforms the uncompressed 8B model, demonstrating the efficiency gains. The online A/B test on Tencent's platform is the strongest evidence of impact, showing a 1.37% GMV lift and significant throughput improvements (37-350% per-GPU QPM) while saving 190 GPUs. The ablation studies clearly demonstrate the necessity of the curriculum learning and the specific projector design.
Reproducibility is moderate. While the paper provides detailed descriptions of the architecture, training stages, and hyperparameters, the core training data consists of proprietary industrial behavior logs from Tencent, which cannot be released. The authors state that public datasets (MRQA, RecBench) are used for evaluation, allowing partial reproduction of the compression and understanding components. However, the specific industrial gains and the full training pipeline on proprietary data cannot be fully replicated by external researchers.
The primary limitation is the reliance on proprietary data, which limits external verification of the industrial claims. The compression ratio is currently fixed per task family, and the paper acknowledges that adaptation to sparser sequences and overly long items remains an area for improvement. Additionally, the scaling laws along data and parameter size have not been systematically validated.
This work has significant implications for the deployment of LLMs in industrial recommendation and advertising systems. By solving the context length and cost bottlenecks through item-level compression, it enables the use of large language models for real-time, billion-scale user profiling. The framework for evaluating compression fidelity (the layered intermediate evaluation) is also valuable for the broader community working on context compression for LLMs. KuaFu introduces an item-level, two-axis compression layer for user behavior sequences that enables billion-scale LLM-based profiling with significant cost and latency reductions. The paper demonstrates that by treating individual behavior items as independent, cacheable units and employing a multi-stage training recipe including hallucination-aware RL, it is possible to maintain or improve downstream task performance while drastically reducing computational overhead, validated by a 1.37% GMV lift in a major production environment.
The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test of this approach: complex, long-horizon tasks challenge agent reliability, while costly real-device interaction limits development scalability. The framework connects data production, model training, and deployment through a shared action-feedback-verification contract. (i) AI for Data builds a human-gated agentic data flywheel in which specialized agents construct tasks, collect interaction trajectories, curate and balance training data, and use training feedback to guide subsequent data generation. (ii) AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning, where we introduce Competence-Aware Reward-and-Advantage Engineering (CARE) to reduce reasoning and tool-use costs while preserving task performance. (iii) AI drives model--harness co-evolution through an execution-evidence-driven loop that orchestrates memory, skills, and tools at runtime and feeds structured action feedback and preserved failure traces back into coordinated model and harness adaptation. Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination. Further evaluations of our model show improvements across non-mobile agentic benchmarks while largely preserving general capabilities.
Primary: Alibaba Group
All Institutions: Alibaba Group
The paper presents a comprehensive framework for developing mobile planner agents using an AI-for-AI closed loop, achieving state-of-the-art performance on MobilePA-Bench. It demonstrates that combining AI-assisted data generation, competence-aware RL, and a dynamic runtime harness can significantly enhance agent reliability and efficiency, offering a scalable pathway for building next-generation autonomous systems.
The paper proposes a "closed-loop AI-for-AI" framework, which is a high-level architectural concept rather than a single novel algorithmic breakthrough. The core technical contributions are the integration of a data flywheel (AI-assisted task generation and curation), a hybrid training pipeline (SFT cold start + RL), and a runtime Harness (Skills/Memory). The most specific technical novelty is CARE (Competence-Aware Reward-and-Advantage Engineering), which adjusts reward shaping based on group success rates to prevent efficiency signals from dominating in saturated success groups. While the components (RL for agents, memory systems, data synthesis) are individually known, the systematic integration into a self-improving loop for mobile agents is a significant engineering and methodological contribution.
The evaluation is conducted on MobilePA-Bench, a large-scale benchmark (1,700+ tasks). The results show Qwen-Planner-Agent (27B) outperforming strong closed-source competitors like GPT-6 Astra and Claude Opus 5, as well as larger open-source models. The ablation studies effectively isolate the contributions of the Planner Model versus the Harness, demonstrating that the runtime context (Skills/Memory) provides substantial gains over the raw model. The efficiency analysis (cost per task) is a valuable addition, showing the agent is not just better but cheaper than frontier commercial APIs.
As a report from a major industry lab (Alibaba), the paper provides high-level architectural details but likely lacks the granular hyperparameters, exact prompt templates, and code releases typically required for full academic reproducibility. The "AI-for-AI" loop implies a complex, proprietary infrastructure for data generation and verification that is difficult to replicate externally. However, the clarity of the framework description allows for conceptual replication.
The primary limitation is the reliance on a proprietary, large-scale infrastructure for the "AI-for-AI" loop, making it difficult for smaller labs to verify the specific benefits of the data flywheel. The "closed-loop" claim is somewhat aspirational; the paper admits that human review is retained for critical decisions, meaning it is not fully autonomous. Additionally, the evaluation is heavily skewed toward the specific mobile planning domain, and while generalization is claimed, the non-mobile benchmarks are less detailed in the provided text.
This paper represents a significant step toward scalable agent development. By formalizing the use of AI to generate training data and diagnose failures for agent systems, it offers a roadmap for reducing the manual effort required to build robust LLM agents. The focus on mobile planning is highly relevant to current industry trends in on-device and cross-app automation. The CARE method offers a useful technique for RL training of agents where success rates vary widely across tasks. The paper presents a comprehensive framework for developing mobile planner agents using an AI-for-AI closed loop, achieving state-of-the-art performance on MobilePA-Bench. It demonstrates that combining AI-assisted data generation, competence-aware RL, and a dynamic runtime harness can significantly enhance agent reliability and efficiency, offering a scalable pathway for building next-generation autonomous systems.
Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Expert pruning reduces this storage burden; at a fixed pruning budget, the goal is to preserve the original model's output distribution as closely as possible. Yet an expert's usage or contribution magnitude does not by itself determine the damage caused by its removal. What matters is whether the surviving computation can replace its function. We introduce RAZOR, a training-free expert pruning method that scores functional replaceability using consensus residuals: deviations of expert outputs from the original weighted mixture. An exact single-deletion identity at a fixed layer input accounts for survivor renormalization and router-selected refill, providing local scores aggregated over calibration tokens for budgeted pruning without gradients or recovery training. On GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, and Hy3 at 25\% and 50\% expert removal, RAZOR achieves the highest nine-task macro average among the evaluated pruning methods in all eight settings. On the two backbones with matched REAP benchmark runs, it exceeds REAP by 2.12--5.59 points and wins all 36 paired task comparisons. It also lowers reverse KL relative to REAP in all four matched GLM-4.7-Flash and Qwen3.6-35B-A3B model--budget settings. Analysis of responses generated by Qwen3.6-35B-A3B nevertheless reveals changes in diversity, formatting, and termination, underscoring that task retention and predictive fidelity do not ensure generation stability.
Primary: NVIDIA
All Institutions: NVIDIA
RAZOR introduces a training-free expert pruning method for MoE models that scores experts by functional replaceability using consensus residuals, accounting for survivor renormalization and router refill. The paper demonstrates that this geometric approach yields superior downstream performance and predictive fidelity compared to magnitude-based baselines across multiple large-scale MoE architectures, while also revealing nuanced trade-offs in generation behavior that task metrics alone miss.
The paper introduces RAZOR, a training-free expert pruning method for Mixture-of-Experts (MoE) models. The core innovation is the shift from scoring experts based on usage frequency or output magnitude (as in REAP or EAN) to "functional replaceability." This is achieved by calculating "consensus residuals," which measure the deviation of an expert's output from the original weighted mixture of active experts. The method derives an exact single-deletion identity that accounts for two critical dynamic effects often ignored in static pruning: survivor renormalization (how the weights of remaining experts adjust) and router-selected refill (how the router promotes a new expert to replace the pruned one). The derivation is mathematically sound, providing a local surrogate for the global distributional shift. The approach is computationally efficient, requiring only forward passes on calibration data without gradients or recovery training.
The evaluation is extensive, covering four distinct MoE backbones (GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, Hy3) at two pruning budgets (25% and 50% expert removal). RAZOR consistently outperforms baselines (Frequency, EAN, REAP) in macro-average downstream performance across all eight model-budget settings. It also demonstrates superior predictive fidelity, measured by lower reverse KL divergence compared to REAP. A notable strength is the inclusion of generation behavior analysis (RQ3), which reveals that while task performance is retained, pruning can still affect diversity and termination patterns, providing a nuanced view of the trade-offs. The experiments are rigorous, using matched calibration sets and detailed ablations on scoring components (RMS vs. Mean, fixed-support vs. refill).
The paper provides high reproducibility. It includes detailed algorithmic pseudocode, specific hyperparameters for evaluation, and clear descriptions of the calibration data composition (Nemotron datasets). The implementation details regarding memory optimization (chunked scoring, layer-wise execution) are well-documented. However, the model checkpoints are not redistributable by the authors, which may limit independent verification for those without access to the specific proprietary or large-scale open models used.
The primary limitation is that the scoring is local and single-deletion based; it does not account for complex interactions when multiple experts are pruned simultaneously, nor does it guarantee that the "refill" candidate remains available if it is also pruned. The paper acknowledges that local output change is a surrogate, not a guarantee of global distributional fidelity. Additionally, the evaluation is limited to four specific model families, and the generalizability to other MoE architectures (e.g., those with different router mechanisms) is not fully established. The lack of measured serving latency or energy savings is a practical gap, as the method's utility is partly defined by compression efficiency.
This work contributes significantly to the field of model compression by providing a principled, geometry-aware method for MoE pruning. It challenges the common heuristic that "less used" experts are less important, showing that "less replaceable" experts are the critical ones to keep. This insight can guide future research in structured pruning and model distillation for sparse models. The finding that task retention does not equate to generation stability is also valuable for practitioners deploying pruned models in production, highlighting the need for multi-faceted evaluation metrics. RAZOR introduces a training-free expert pruning method for MoE models that scores experts by functional replaceability using consensus residuals, accounting for survivor renormalization and router refill. The paper demonstrates that this geometric approach yields superior downstream performance and predictive fidelity compared to magnitude-based baselines across multiple large-scale MoE architectures, while also revealing nuanced trade-offs in generation behavior that task metrics alone miss.
Frontier language models are rarely used in clinical workflows because the realistic, longitudinal benchmarks needed to develop them are scarce. Real electronic health record (EHR) data cannot be openly shared due to privacy, ethics or data use issues and it does not contain verifiable ground truth since the chart records only reflect what clinicians documented. We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers. Built entirely from public medical-education material with no protected health information, it comprises 1,268 longitudinal patients and 5,602 encounters, where every diagnosis, finding, and temporal relation is grounded in standard ontologies (ICD-10-CM, SNOMED CT, LOINC) and with a complete provenance chain back to its source medical education material. Synthetic Hospital is served through a simulated hospital record system that mirrors real EHR infrastructure (standard interoperability APIs, role-based access and function-calling interface). In a blinded review, physicians distinguished its records from real patient charts at near-chance rates (53\%). Across 10 frontier and open models, none approaches ceiling: the best model reconstructs a patient's longitudinal problem list with a severity-weighted F1 of 0.73, level with the mean of seven physicians on a matched subset but well below the best of them (0.89), and misses roughly half of clinically relevant findings when summarizing a chart. Overall, these results highlight that Synthetic Hospital is a difficult and realistic test of clinical AI performance.
Primary: Carnegie Mellon University
All Institutions: Carnegie Mellon University
Synthetic Hospital introduces an open, verifiable, and physician-validated longitudinal EHR benchmark that resolves privacy and ground-truth barriers in clinical AI evaluation. By grounding synthetic records in standard medical ontologies and public educational material, the paper provides a rigorous, reproducible testbed that reveals significant gaps in current frontier models' ability to perform longitudinal clinical reasoning, offering a crucial resource for advancing reliable clinical AI systems.
The paper proposes a deterministic, five-stage pipeline to construct "Synthetic Hospital," a longitudinal EHR benchmark. The core innovation is the decoupling of clinical ground truth from narrative generation. By first constructing a structured medical knowledge graph grounded in standard ontologies (ICD-10-CM, SNOMED CT, LOINC) from public USMLE-style board questions, and then rendering this graph into realistic clinical narratives using an LLM (Kimi 2.5), the authors ensure that every diagnosis and finding has a verifiable provenance chain. This addresses the two major barriers to open clinical benchmarks: privacy (no PHI) and ground truth ambiguity (real charts reflect documentation, not necessarily patient state). The methodology includes rigorous validation steps, such as a blinded physician study to assess realism and a pre-registered reference set to validate ontology-derived relationships.
The evaluation is comprehensive. It includes a realism study where physicians could not distinguish synthetic from real records (53% accuracy, near chance). It evaluates 10 frontier and open models across four tasks (patient diagnosis, summarization, retrieval, imaging indication). Key findings include that no model approaches ceiling performance, the best model achieves a severity-weighted F1 of 0.73 on diagnosis (matching the mean of seven physicians but below the best), and that agentic multi-turn loops often degrade performance compared to single-turn inference when context is available upfront, except for longitudinal diagnosis tasks. The paper also provides a robustness analysis showing that re-rendering the corpus with a different LLM (GPT-5.3) does not significantly change model rankings, mitigating concerns about generator bias.
High. The paper provides a detailed description of the pipeline, including specific thresholds for ontology mapping, clustering rules, and prompting strategies. Code and data are openly available on GitHub. The use of deterministic functions for benchmark labels and the release of a training split with verifiable rewards enhances reproducibility and utility for reinforcement learning or fine-tuning.
The benchmark is derived from medical education material (USMLE-style questions), which may not fully capture the complexity, noise, and atypical presentations of real-world clinical practice. The case mix is education-derived by design, potentially under-representing rare conditions. The agentic evaluation is limited to three models and one scaffold. The paper acknowledges that it does not explicitly simulate missingness or documentation errors, which are common in real EHRs.
This benchmark has high potential impact on the clinical AI community. By providing an open, verifiable, and realistic longitudinal EHR dataset, it enables the development and evaluation of clinical AI systems without the legal and ethical hurdles of using real patient data. The finding that current frontier models struggle with longitudinal synthesis and that agentic approaches have mixed utility provides actionable insights for system designers. The open nature of the data facilitates broader research and standardization of evaluation metrics in clinical NLP. Synthetic Hospital introduces an open, verifiable, and physician-validated longitudinal EHR benchmark that resolves privacy and ground-truth barriers in clinical AI evaluation. By grounding synthetic records in standard medical ontologies and public educational material, the paper provides a rigorous, reproducible testbed that reveals significant gaps in current frontier models' ability to perform longitudinal clinical reasoning, offering a crucial resource for advancing reliable clinical AI systems.
Computer-use agents (CUAs) increasingly act on behalf of users online. What happens when the environments they operate in have incentives of their own? Online marketplaces, for example, may favor some products over others, steering agents away from the user's objective. Existing CUA benchmarks cover cooperative settings or explicit attacks, but do not test whether agents preserve user objectives when the environment itself has a stake in the outcome. We introduce CAVEAT, a controlled benchmark spanning nine marketplace environments and a taxonomy of eight common steering mechanisms. Across five model families, agents purchase the user-optimal product in 78.6% of matched-control episodes but only 17.3% when steering mechanisms are enabled. Larger models and more reasoning improve robustness, but substantial failures persist. Trajectory analysis and targeted ablations identify three weaknesses in how agents decide: they (1) prematurely narrow the set of alternatives they consider, (2) impose priorities the user never stated, and (3) commit before resolving decision-relevant evidence. Guided by this diagnosis, we develop CAVEAT-Harness, which targets these failures and raises the optimal purchase rate by up to 80.0 percentage points, and show that targeted post-training further improves a smaller open model. These results establish incentive robustness as a distinct challenge for delegated agents, diagnose failure modes, and show how targeted interventions can substantially improve robustness.
Primary: Microsoft Research
All Institutions: Microsoft Research
The paper establishes incentive robustness as a critical and distinct challenge for computer-use agents, demonstrating that current models are highly vulnerable to subtle environmental steering. By introducing the CAVEAT benchmark and diagnosing specific decision-making failures, the work provides a rigorous framework for evaluating and improving agent behavior in misaligned environments, with practical interventions that significantly enhance user-aligned outcomes.
The paper introduces CAVEAT, a benchmark for evaluating Computer-Use Agents (CUAs) in environments with misaligned incentives. The core methodological contribution is the construction of nine controlled marketplace environments featuring eight distinct "steering mechanisms" (e.g., hidden fees, biased search rankings, dynamic pricing). The authors formalize the problem of "incentive robustness," distinguishing it from standard security attacks or cooperative tasks. They propose a diagnostic framework that categorizes agent failures into three specific decision-making weaknesses: premature narrowing of alternatives, imposition of unstated user priorities, and premature commitment. Based on this diagnosis, they develop CAVEAT-Harness, a post-processing or prompting strategy that explicitly targets these failure modes. The methodology is sound, moving beyond simple black-box evaluation to a structured analysis of *why* agents fail in adversarial economic contexts.
The evaluation is comprehensive, testing five model families across the CAVEAT benchmark. The key empirical finding is stark: agents achieve a 78.6% optimal purchase rate in control conditions but drop to 17.3% when steering mechanisms are active. This quantifies the vulnerability of current state-of-the-art agents to subtle environmental manipulation. The paper further demonstrates that while larger models and increased reasoning steps improve robustness, they do not eliminate the problem. The proposed CAVEAT-Harness significantly improves performance, raising the optimal purchase rate by up to 80.0 percentage points in some settings, validating the diagnostic approach. The inclusion of targeted post-training for smaller open models adds practical value to the findings.
The paper appears to be from a major research institution (Microsoft Research, inferred from the acknowledgments of prominent MSR researchers like Saleema Amershi, Gagan Bansal, and Ece Kamar). While the full code is not provided in the text snippet, the detailed description of the benchmark environments, steering mechanisms, and the CAVEAT-Harness protocol suggests a high level of reproducibility. The specific metrics (optimal purchase rate) and the taxonomy of failures provide clear guidelines for replication. The venue is identified as ICLR 2027, indicating it has passed rigorous peer review.
The primary limitation is the scope of the environments. The benchmark focuses on online marketplaces, which, while common, may not cover all types of incentive-misaligned environments (e.g., social media feeds, news aggregators, or multi-agent negotiation). The "steering mechanisms" are defined by the authors; real-world platforms may employ more complex or adaptive strategies. Additionally, the CAVEAT-Harness is a targeted intervention; its generalizability to other types of agent tasks or environments outside of purchasing decisions is not fully established. The drop in performance (78.6% to 17.3%) is dramatic, but the absolute performance in the control condition (78.6%) suggests that even in benign environments, agents are not perfect, which may confound the isolation of the steering effect.
This paper has significant implications for the deployment of autonomous agents in the real world. As CUAs become more prevalent, the ability of environments (platforms, marketplaces) to subtly manipulate agent behavior poses a serious risk to user autonomy and utility. By establishing "incentive robustness" as a distinct challenge, the paper shifts the focus from pure capability to alignment and safety in economic contexts. The findings will likely influence the design of future agent architectures, prompting strategies, and safety evaluations. It also highlights a gap in current AI safety research, which has often focused on explicit adversarial attacks rather than the subtle, incentive-driven manipulation inherent in many online platforms. The paper establishes incentive robustness as a critical and distinct challenge for computer-use agents, demonstrating that current models are highly vulnerable to subtle environmental steering. By introducing the CAVEAT benchmark and diagnosing specific decision-making failures, the work provides a rigorous framework for evaluating and improving agent behavior in misaligned environments, with practical interventions that significantly enhance user-aligned outcomes.
Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by similarity. This forces the system to decide what is worth remembering before the future query is known, irreversibly discarding information and producing a query-independent summary that must serve many possible downstream tasks. Learning such a write-time curator is also difficult because the value of a storage decision may only become apparent when a relevant query arrives, potentially many tasks later, creating a long-horizon credit-assignment problem. We instead retain raw trajectories and defer curation until read time, when the current task is known. Given the retrieved traces and the new task, a memory curator synthesizes a compact, task-adaptive payload tailored to the immediate need. Because this payload is consumed on the same task, the curator can be trained directly from immediate task success, avoiding delayed utility signals and the need to artificially group related tasks. Across ALFWorld, WebShop, and τ^2-bench, our Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively. Notably, even an untrained curator is already competitive with or surpasses these baselines, showing that task-adaptive read-time curation itself is a major source of the gain; training the curator further compounds the improvement.
Primary: Salesforce
All Institutions: Salesforce
[One sentence main contribution]. The paper introduces Just-in-Time Memory, a read-time curation strategy for LLM agents that outperforms write-time methods by synthesizing task-adaptive payloads from raw trajectories, demonstrating significant gains in success rates across multiple benchmarks.
The paper proposes "Just-in-Time Memory" (JitMem), shifting the curation of agent memory from write-time (fixed summaries) to read-time (task-adaptive synthesis). The core method involves retaining raw trajectories and using a "memory curator" LLM to synthesize a compact payload specifically for the current query. This addresses the long-horizon credit assignment problem inherent in write-time curation by allowing the curator to be trained directly on immediate task success. The approach is conceptually sound, leveraging the flexibility of LLMs to perform on-the-fly reasoning over raw data rather than relying on pre-computed, potentially lossy, static representations.
The experiments are conducted on three standard agent benchmarks: ALFWorld, WebShop, and τ^2-bench. The results show consistent improvements over no-memory baselines and existing write-time memory methods. The reported gains are substantial (16.2, 16.3, and 3.9 absolute points), suggesting the method is effective across different task types. The finding that even an untrained curator is competitive is a strong empirical result, highlighting the value of the read-time curation paradigm itself.
The paper is on arXiv. While the methodology is described, the specific implementation details of the "memory curator" training (e.g., reward shaping, specific prompts, hyperparameters) may be limited in the abstract-only view, but the full text likely contains sufficient detail for reproduction given the standard nature of LLM agent frameworks.
The primary limitation is computational cost. Retaining raw trajectories and performing LLM-based synthesis at read time is significantly more expensive than retrieving a pre-computed summary. This may limit scalability to very long-horizon agents or high-throughput applications. Additionally, the method relies heavily on the base LLM's ability to synthesize relevant information from raw traces, which may vary across model capabilities.
This work has significant implications for the design of LLM agents, suggesting that static memory structures may be suboptimal. It encourages the development of more dynamic, query-aware memory systems. The approach could be extended to other domains requiring adaptive information retrieval, such as RAG systems that dynamically re-rank or synthesize context based on the specific query. [One sentence main contribution]. The paper introduces Just-in-Time Memory, a read-time curation strategy for LLM agents that outperforms write-time methods by synthesizing task-adaptive payloads from raw trajectories, demonstrating significant gains in success rates across multiple benchmarks.
Many scientific datasets, such as molecular conformational ensembles or single-cell tissue measurements, are naturally modeled as meta-distributions: distributions over probability measures on non-Euclidean domains. Existing generative methods largely assume Euclidean geometry and fail to capture this structure. We introduce Riemannian Wasserstein Entropic Flow Matching (RWEFM), a generative framework on the Wasserstein space $\mathcal{P}_2(\mathcal{M})$ of a Riemannian manifold $(\mathcal{M},g)$. RWEFM is trained by regressing a neural vector field onto Riemannian optimal transport velocities, using McCann displacement interpolations as conditional paths. We confirm theoretically that this construction leads to a valid flow matching approach on $\mathcal{P}_2(\mathcal{M})$ and introduce the Riemannian Entropic Map, a GPU-efficient approximation of the optimal transport map on manifolds. Our experiments show that by respecting the intrinsic geometry of the data, RWEFM can generate whole single-cell samples in hyperspherical latent spaces and protein conformational ensembles on the torus. As RWEFM requires only a geodesic distance and a projection operator, it is not restricted to manifolds with closed-form geometry, which we demonstrate by generating distributions on a general triangulated mesh.
Primary: Genentech Inc.
All Institutions: Genentech Inc., Yale University
The paper presents a novel and theoretically sound framework for generative modeling on Riemannian manifolds, effectively extending flow matching to the Wasserstein space of probability measures. By introducing the Riemannian Entropic Map and demonstrating applications to single-cell genomics and protein structures, it offers a significant methodological advance for handling complex, non-Euclidean scientific data.
The paper introduces Riemannian Wasserstein Entropic Flow Matching (RWEFM), a rigorous extension of Flow Matching to the Wasserstein space of probability measures on Riemannian manifolds. The core theoretical contribution is the validation of the flow matching objective in this infinite-dimensional, non-Euclidean setting, utilizing McCann displacement interpolations. A key methodological innovation is the "Riemannian Entropic Map," a GPU-efficient estimator for the optimal transport map that generalizes the Euclidean entropic map. It employs a lift-average-retract procedure (logarithmic map to tangent space, barycentric projection, exponential map back to manifold) which is computationally tractable and theoretically grounded with error bounds. The framework is designed to be geometry-agnostic, requiring only geodesic distances and projection operators, allowing application to complex shapes like triangulated meshes.
The experiments are diverse and scientifically relevant. The authors demonstrate the method on synthetic data (MNIST/EMNIST/KMNIST mapped to sphere, hyperbolic space, and torus) to validate geometric correctness. They apply the method to real-world scientific problems: generating single-cell RNA-seq samples on hyperspherical latent spaces and protein conformational ensembles on the torus. The inclusion of a general triangulated mesh (Stanford Bunny) experiment is particularly strong, as it proves the method's applicability beyond closed-form geometries. The metrics used (1-NN deviation, MMD, Chamfer Distance) are appropriate for distributional comparison.
The paper provides a public GitHub repository with code and tutorials. The appendix contains detailed hyperparameters, network architecture descriptions (self-attention blocks), and explicit formulas for geometric operations on various manifolds. The training procedure is well-documented, including details on noise generation and mini-batch OT coupling. This level of detail supports high reproducibility.
The method relies on the computation of optimal transport plans, which can be computationally expensive for very large point clouds, although the entropic regularization helps. The "sampled map" approximation used in high-dimensional settings (like single-cell data) may introduce bias compared to the true barycentric map. The theoretical guarantees for the Riemannian Entropic Map depend on regularity assumptions that may not hold for all practical datasets.
This work bridges the gap between geometric deep learning and generative modeling for distributional data. It provides a toolkit for scientists working with non-Euclidean data (molecules, cells, climate) to generate realistic samples that respect the underlying geometry. The framework is likely to influence future work in scientific machine learning and optimal transport. The paper presents a novel and theoretically sound framework for generative modeling on Riemannian manifolds, effectively extending flow matching to the Wasserstein space of probability measures. By introducing the Riemannian Entropic Map and demonstrating applications to single-cell genomics and protein structures, it offers a significant methodological advance for handling complex, non-Euclidean scientific data.
Progress in general-purpose video editing depends on constructing large-scale paired supervision and effectively adapting video-generation backbones to instruction-driven editing. Unlike video generation, video editing must execute a requested transformation while preserving unrelated subjects, scene structure, motion, and temporal continuity. We present VideoX-Qwen, an integrated data-construction and model-training framework for general instruction-based video editing. Our scalable production pipeline organizes specialized generation and understanding models into complementary routes for addition, removal, replacement, and attribute editing, followed by quality screening and instruction enrichment. It produces more than 1.2 million directional video-editing records, including over 400,000 records in each major task group, with an automatic acceptance rate of 89%. The resulting corpus provides broad and structured coverage of common editing operations through a unified source-instruction-target interface. We further develop a unified Qwen-Wan editor that combines multimodal semantic conditioning with dense source-video latent guidance. A progressive image-video training strategy aligns the multimodal instruction interface, adapts the video generator to source-conditioned editing, and refines output quality with selected high-resolution data. In a 100-example comparison with UniVideo and Kling O1, VideoX-Qwen achieves the best mean result on nine of eleven reported metrics, including instruction following, editing quality, content preservation, structural and perceptual similarity, and video-distribution quality. Together, the large-scale data-production system and unified training framework provide a practical foundation for more capable instruction-driven video editing.
Primary: Alibaba Group
All Institutions: Alibaba Group, Zhejiang University
The paper presents a scalable data pipeline and a unified training framework for instruction-based video editing, achieving strong results in a limited comparison. While the data construction methodology is a valuable contribution, the small-scale evaluation limits the confidence in the claimed state-of-the-art performance.
The paper proposes VideoX-Qwen, a framework that integrates a large-scale data construction pipeline with a unified model training strategy for instruction-based video editing. The data pipeline is a significant engineering contribution, leveraging specialized generation and understanding models to create 1.2 million paired video editing records across four task types (addition, removal, replacement, attribute editing). The model architecture, a "Qwen-Wan" editor, combines multimodal semantic conditioning (likely using a Qwen-based LLM/VLM) with dense source-video latent guidance (likely using a Wan-based video diffusion model). The training strategy is progressive, starting with image editing to align the instruction interface, then moving to source-conditioned video editing, and finally refining with high-resolution data. This approach is methodologically sound and addresses the core challenge of video editing: executing specific edits while preserving unrelated content and temporal consistency.
The evaluation is limited to a 100-example comparison against two baselines, UniVideo and Kling O1. While the paper claims state-of-the-art performance on 9 out of 11 metrics, the small sample size (100 examples) is a significant weakness for a paper claiming to introduce a "practical foundation" for general video editing. The metrics reported (instruction following, editing quality, content preservation, etc.) are relevant, but the lack of a larger, standardized benchmark or user study limits the strength of the empirical claims. The comparison with Kling O1, a commercial system, is interesting but potentially unfair if the open-source baselines are not equally tuned.
The paper describes the data pipeline and training strategy in detail, which aids reproducibility. However, the reliance on proprietary or large-scale models (Qwen, Wan) and the specific "specialized generation and understanding models" for data creation makes full reproduction difficult for smaller labs. The code and data are not explicitly mentioned as being released in the provided text, which is a gap for a paper of this scale.
The primary limitation is the small scale of the evaluation (100 examples). This is insufficient to robustly claim superiority over strong baselines like Kling O1. Additionally, the paper does not extensively discuss failure cases or the specific types of edits where the model struggles. The data pipeline, while impressive, may be biased towards the types of edits that are easy to generate and verify automatically, potentially missing more complex or nuanced editing scenarios.
The work has significant potential impact by providing a scalable method for generating high-quality video editing data, which is a major bottleneck in the field. The unified framework for instruction-based editing could enable more accessible and flexible video editing tools for non-experts. The integration of LLM-based instruction understanding with diffusion-based video generation is a trend that is likely to be widely adopted. The paper presents a scalable data pipeline and a unified training framework for instruction-based video editing, achieving strong results in a limited comparison. While the data construction methodology is a valuable contribution, the small-scale evaluation limits the confidence in the claimed state-of-the-art performance.
An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 14.1 points on the split it evolves against and up to 4.7 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution. Code is available at https://github.com/google-research/rrsi and project page is https://regularized-rsi.com/.
Primary: Google Research
All Institutions: Google Research, University of Illinois Urbana-Champaign, University of Maryland, University of Washington
The paper introduces a regularization framework for recursive self-improvement of LLM agent harnesses, effectively mitigating overfitting and improving out-of-distribution generalization. By constraining the proposal and selection of harness edits through annealed budgets and critic-based pruning, RRSI achieves significant performance gains and efficiency improvements across diverse agentic tasks, offering a robust and reproducible approach to automated agent evolution.
The paper proposes RRSI, a framework for regularizing the recursive self-improvement (RSI) of LLM agent harnesses. The core innovation lies in applying regularization principles—typically used in model training—to the evolution of agent components (prompts, tools, memory). The method introduces a "proposer" with a temporally annealed budget to limit edit bundling and encourage exploration, and a "selector" equipped with a critic and pruner to filter out noisy, expensive, or trivial changes. This approach directly addresses the overfitting problem observed in prior automated harness evolution methods, where in-distribution gains do not transfer to out-of-distribution tasks. The methodology is well-structured, combining evolutionary search with explicit constraints to favor reusable mechanisms over benchmark-specific hacks.
The evaluation is comprehensive, spanning eight benchmarks across coding, agentic workspace, and engineering design tasks. The paper reports significant improvements: up to 14.1 points on the evolution split and up to 4.7 points on out-of-distribution benchmarks. Crucially, it demonstrates efficiency, reducing policy tokens by 30% compared to unregularized evolution. The comparison against unregularized baselines is strong, and the inclusion of OOD transfer metrics is a significant strength, validating the claim that regularization improves generalization. The use of multiple policy models and domains adds robustness to the findings.
The paper provides a public GitHub repository and a project page, which significantly enhances reproducibility. The detailed description of the proposer and selector mechanisms, along with the specific hyperparameters for the annealing budget and pruning criteria, allows other researchers to replicate the setup. The availability of code for the agent harnesses and the evolution loop is a major plus for the community.
The study is limited to frozen backbone models, meaning it does not address joint optimization of weights and harness. The method relies on a finite evolve set and several hyperparameters, which may require careful tuning for different agent architectures. The authors acknowledge that broader validation on substantially different tool ecosystems and longer-running self-improvement processes is needed. The reliance on a critic model for selection introduces an additional dependency on the quality of that critic.
This work has significant implications for the field of autonomous agents and LLM self-improvement. By demonstrating that regularization can prevent overfitting in harness evolution, it provides a crucial tool for building more robust and generalizable agents. The efficiency gains (fewer tokens) are also practically important for deployment costs. This paper likely to influence future work on automated agent design and self-improvement loops, establishing a new standard for evaluating and constraining such processes. The paper introduces a regularization framework for recursive self-improvement of LLM agent harnesses, effectively mitigating overfitting and improving out-of-distribution generalization. By constraining the proposal and selection of harness edits through annealed budgets and critic-based pruning, RRSI achieves significant performance gains and efficiency improvements across diverse agentic tasks, offering a robust and reproducible approach to automated agent evolution.
Human-feedback alignment has made language models useful assistants and is commonly described as aligning them with humans. However, the responses people prefer from an AI need not be the responses they themselves would give. We distinguish alignment with human preferences from alignment with human behavior, and show that alignment with human preferences can make model behavior less human-like even when both preferences and responses come entirely from humans. We call this the Turing-test gap. We show that preference alignment preserves the human response distribution only under a restrictive condition, and find no consistent evidence that real human preferences satisfy it. Empirically, the loss of human-response likelihood increases with the strength of preference weighting, regardless of its direction, and the gap also appears under standard DPO. These results establish human-likeness as an explicit dimension of alignment rather than something assumed to follow from preference alignment.
Primary: University of Sydney
All Institutions: University of Sydney, University of Oxford, Southeast University
The paper rigorously demonstrates that standard preference alignment methods (like DPO) systematically reduce a model's fidelity to human behavior, even when trained on human data, by introducing a formal "Turing-test gap" and providing empirical evidence that preference weighting moves the model distribution away from the human response distribution.
The paper introduces a clear theoretical distinction between "alignment with human preferences" (optimizing for what humans prefer) and "alignment with human behavior" (optimizing for what humans actually do). It formalizes the "Turing-test gap" as the divergence between these two objectives. The core theoretical contribution is deriving the condition under which preference alignment preserves the human response distribution, showing that it only holds if the preference reward encodes the exact log-density ratio between the reference policy and the human behavior distribution. The authors provide a rigorous mathematical framework using KL-regularized objectives and demonstrate that standard preference weighting (even with human data) systematically moves the model away from the human distribution. The methodology is sound, combining theoretical derivations with controlled empirical experiments to isolate the effect of preference weighting from data composition.
The experiments are well-designed to validate the theoretical claims. The authors use a reweighting experiment on human-written responses (SHP, StackExchange) to show that increasing preference weighting (in either direction) reduces the likelihood of unseen human responses. They also test DPO from different starting points (human-SFT vs. base) to show that DPO further exacerbates the gap when starting from a human-aligned reference. The use of random weight reassignment as a control is a strong methodological choice, proving that the loss of human-likeness is due to the weighting mechanism itself, not just the specific preference labels. The evaluation metrics (NLL, preference margin, energy distance in embedding space) are appropriate for measuring distributional shift.
The paper provides detailed hyperparameters, dataset splits, and training procedures in the appendix. It specifies the exact models used (Qwen2.5-7B-Instruct, Llama-3-8B-Instruct) and datasets (SHP, StackExchange, HH-RLHF, WebGPT). The code for the source classifier and evaluation metrics is described in detail. However, no explicit GitHub link is provided in the text, which slightly hinders immediate reproducibility, though the details are sufficient for a skilled practitioner to replicate the core experiments.
The paper acknowledges that matching the population distribution does not guarantee individual-level human-likeness. It also notes that the experiments are distributional rather than interactive (Turing-test style). The scale of the models (7B/8B) may not fully capture the behavior of frontier models, although the authors argue the effect is likely more pronounced there. The reliance on specific datasets (SHP, StackExchange) limits the generalizability of the empirical findings to other domains.
This paper has significant implications for the field of AI alignment. It challenges the assumption that making models more "helpful" or "preferred" automatically makes them better proxies for human behavior. This is crucial for applications like social simulation, opinion polling, and behavioral experiments where models are used to represent human populations. The findings suggest that current alignment pipelines may be inadvertently reducing the utility of models for these specific tasks. It also raises ethical questions about the risks of models that are too human-like (impersonation, manipulation) versus those that are not. The paper rigorously demonstrates that standard preference alignment methods (like DPO) systematically reduce a model's fidelity to human behavior, even when trained on human data, by introducing a formal "Turing-test gap" and providing empirical evidence that preference weighting moves the model distribution away from the human response distribution.
X-ray is medicine's most widely used imaging modality, yet remains among its least quantitative. Unlike volumetric modalities like CT or MRI, X-ray collapses 3D anatomy into a 2D projection, causing structures to overlap and anatomical boundaries to be ambiguous, even to experts. As a result, labeling X-ray databases for training general-purpose segmentation systems is impractical, leaving morphometric and functional X-ray analysis confined to narrow anatomical regions and applications. To this end, we present FleXray, a generalist model for anatomical segmentation across the entire body in clinical X-rays. Instead of curating large, manually annotated X-ray datasets, we build a scalable, physics-based generative X-ray data engine. Using existing 3D whole-body CT segmentation datasets and generative image-editing models, we simulate fully-annotated 2D X-rays with diverse appearances, physiological properties, and imaging geometries. Trained on these simulations, FleXray accurately segments 60 anatomical structures across unseen research datasets and in-the-wild X-rays. We further show that FleXray makes X-rays directly amenable to quantitative analysis, enabling automated measurements for disease grading, robust navigation during X-ray-guided interventions, and data-efficient learning of pathological targets. We release the model, code, a full-body X-ray segmentation dataset, and a local, easy-to-use browser-based tool at https://flexray.csail.mit.edu .
Primary: Massachusetts Institute of Technology (MIT)
All Institutions: Massachusetts Institute of Technology (MIT)
FleXray introduces a scalable, physics-based generative data engine that leverages 3D CT data and generative image editing to train a generalist model for full-body X-ray segmentation. The paper demonstrates that synthetic data, when carefully designed to cover anatomical, geometric, and appearance variation, can effectively bridge the gap to real-world clinical X-rays, enabling quantitative analysis and downstream clinical applications with high accuracy and robustness.
The paper proposes a highly effective data-centric approach to solving the problem of scarce, high-quality annotations for X-ray segmentation. The core innovation is a "physics-based generative data engine" that leverages existing 3D CT segmentation datasets to simulate 2D X-rays (Digitally Reconstructed Radiographs, DRRs). The methodology is robust, combining online analytic projections for geometric diversity with a generative image-editing model (Flux) to bridge the domain gap between synthetic DRRs and real clinical X-rays. The use of a language-conditioned diffusion model to add realistic artifacts (scatter, text, detector noise) while preserving anatomical structure is a clever and effective solution to the "sim-to-real" gap. The training objective is well-designed, handling multi-label overlap and partial annotations from heterogeneous sources.
The evaluation is comprehensive and rigorous. The model is tested on eight held-out datasets spanning various body parts (limbs, pelvis, spine, chest) and populations (adult, pediatric). It outperforms strong baselines like PAXray, TotalSegmentator2D, and FluoroSAM. The paper goes beyond standard segmentation metrics (Dice) to demonstrate downstream utility: automated Cobb angle measurement for scoliosis (achieving inter-rater variability levels), improved 2D/3D registration for surgical navigation, and data-efficient fine-tuning for new pathological targets. The ablation studies clearly isolate the contributions of generative enhancement and attenuation randomization, showing where each component is most beneficial.
High. The authors release the model, code, the full-body X-ray segmentation dataset, and a browser-based tool. The data engine components, including the specific prompts used for the generative editor and the filtering strategies, are detailed in the appendix. The use of public datasets (MOOSE, etc.) for training data generation further enhances reproducibility.
The model inherits biases from the CT source data, which may not fully represent all clinical X-ray scenarios (e.g., severe trauma, specific implants not present in CTs). The evaluation, while broad, is limited by the scarcity of public X-ray segmentation datasets with ground truth. The reliance on a large generative model (Flux) for data augmentation adds computational cost to the data preparation phase, though this is a one-time cost.
This work has significant potential to transform clinical X-ray analysis by making it quantitative and automatable. By providing a generalist model that segments 60 anatomical structures, it enables population-scale studies, automated disease grading, and improved surgical navigation. The release of the data engine and dataset will likely accelerate research in medical imaging by providing a scalable way to generate labeled training data for other modalities or tasks. FleXray introduces a scalable, physics-based generative data engine that leverages 3D CT data and generative image editing to train a generalist model for full-body X-ray segmentation. The paper demonstrates that synthetic data, when carefully designed to cover anatomical, geometric, and appearance variation, can effectively bridge the gap to real-world clinical X-rays, enabling quantitative analysis and downstream clinical applications with high accuracy and robustness.
Precise instruction-following is a fundamental ability of large language models (LLMs), requiring their outputs to strictly satisfy objective constraints in input instructions. In complex application scenarios, these constraints often possess diverse scopes that govern specific response segments rather than the entire output. However, existing optimization methods often neglect constraint scope during data construction and rely on binary per-constraint rewards, yielding limited data diversity and sparse supervision for complex constraints. To this end, we propose ScopeIF, a novel training framework for scope-aware precise instruction-following. We first introduce a unified schema that factorizes objective constraints into three decoupled dimensions: Scope, Target, and Range. Grounded in this schema, we construct ScopeInstruct, a large-scale instruction dataset with diverse scope-aware constraints, and combine tool-grounded verification with graded reward modeling to quantify the violation degree of each constraint, providing dense supervision for policy optimization. Extensive experiments demonstrate that ScopeIF consistently outperforms existing methods, particularly on complex scope-aware constraints, while preserving general capabilities. Notably, it enables optimized Qwen3-4B and 8B models to rival or surpass strong frontier models such as Gemini-2.5-Pro and DeepSeek-V3.2, establishing an effective paradigm for advancing scope-aware instruction-following. Our code and data are available at https://github.com/thu-coai/ScopeIF.
Primary: Tsinghua University
All Institutions: Tsinghua University, Zhipu AI
The paper introduces a scope-aware instruction-following framework using graded rewards to improve LLM precision. It offers a significant methodological improvement in RLHF by addressing the sparsity of binary rewards for complex, segmented constraints, demonstrating strong empirical results where smaller models rival frontier systems on specific tasks.
The paper proposes ScopeIF, a framework addressing the limitation of binary rewards in instruction-following tasks where constraints apply to specific segments (scopes) rather than the whole output. The core contribution is a unified schema factorizing constraints into Scope, Target, and Range, which allows for more granular data construction. The method employs tool-grounded verification to generate graded rewards, providing dense supervision signals for reinforcement learning. This is a logical and well-motivated extension of existing RLHF techniques, moving from sparse, binary feedback to dense, scope-aware feedback.
The experiments demonstrate that the optimized Qwen3-4B and 8B models rival or surpass frontier models like Gemini-2.5-Pro and DeepSeek-V3.2 on specific scope-aware tasks. This is a strong empirical claim. The use of a large-scale dataset (ScopeInstruct) and comparison against strong baselines adds credibility. However, the claim of surpassing frontier models on a specific subset of tasks (scope-aware constraints) needs careful scrutiny to ensure it is not an artifact of the specific test set construction, though the abstract suggests general capabilities are preserved.
The authors have released full code and data, including prompts and hyperparameters, which is excellent for reproducibility. The AI use statement is transparent about the use of LLMs for data generation and polishing, which is standard but important to disclose.
The primary limitation is the reliance on tool-grounded verification, which may not be applicable to all types of constraints or domains where automated verification is difficult. Additionally, the performance gains are specific to "scope-aware" constraints; the impact on general instruction following outside this niche is less clear, though the abstract claims preservation of general capabilities.
This work contributes to the broader goal of making LLMs more precise and controllable. By providing a framework for handling scoped constraints, it could be adopted in applications requiring strict adherence to local instructions within longer contexts, such as code generation with specific function-level requirements or structured data extraction. The paper introduces a scope-aware instruction-following framework using graded rewards to improve LLM precision. It offers a significant methodological improvement in RLHF by addressing the sparsity of binary rewards for complex, segmented constraints, demonstrating strong empirical results where smaller models rival frontier systems on specific tasks.
Attributing language model outputs to their internal computations is an open problem in interpretability. Existing methods, which use causal interventions, gradients, or learnable masks, either are infeasibly expensive or struggle to identify actual causally-important internal computations. We propose framing attribution as the problem of identifying nested subsets of internal components which minimise a downstream loss. To learn this task, we introduce Matryoshka Attribution (MAttr), a mask learning method that parametrises the mask with a simple differentiable sigmoid top-$k$ operator. We supervise training over all sparsities simultaneously by randomising $k$ over training, resulting in a learned ordering of components by attribution score. MAttr achieves number 1 on the official leaderboard of the Mechanistic Interpretability Benchmark (Mueller et al., 2025); our method identifies sparse and task-transferrable circuits across varying circuit bases. As a practical application, we show that MAttr can be trained with reinforcement learning to identify weight changes responsible for downstream behaviours in LLM finetuning. We train MAttr on refusal judge scores and find that restoring $1\%$ of Llama 3.1 8B Instruct's weights to their base model state is sufficient to remove refusals while maintaining capabilities. We view MAttr as a successful formulation of interpretability into a learnable objective that we can tackle with gradient descent, and encourage future work along these lines.
Primary: Stanford University
All Institutions: Stanford University
Matryoshka Attribution (MAttr) is a novel mask-learning method that uses a differentiable sigmoid top-$k$ operator to learn a single, nested ordering of internal components for attribution, achieving state-of-the-art performance on the Mechanistic Interpretability Benchmark and enabling the identification of specific weight changes responsible for LLM behaviors like refusal through reinforcement learning.
The paper introduces Matryoshka Attribution (MAttr), a mask-learning method that frames attribution as identifying nested subsets of internal components minimizing a downstream loss. The core innovation is the use of a differentiable sigmoid top-$k$ operator to parametrize the mask, allowing the model to learn a single ordering of components that is valid across all sparsity levels (the "Matryoshka" property). This eliminates the need for separate mask training at each sparsity level and avoids the instability and hyperparameter sensitivity associated with hard concrete distributions or straight-through estimators. The method is grounded in causal abstraction, defining "soft interchange interventions" that interpolate between base and source states. A significant theoretical contribution is the demonstration that MAttr recovers Integrated Gradients in the limit of one training step, providing a rigorous bridge between gradient-based and mask-learning attribution methods. The extension to parameter-level attribution via Reinforcement Learning (GRPO) is a novel application, allowing the method to identify weight changes responsible for specific behaviors (e.g., refusal) without requiring differentiable rewards.
The experimental evaluation is extensive and rigorous. MAttr achieves state-of-the-art performance on the Mechanistic Interpretability Benchmark (MIB), significantly outperforming baselines like Integrated Gradients, Interchange Interventions, and other mask-learning methods (DBM, Node Pruning). The authors introduce "MIB+" to test finer-grained bases (MLP neurons, SAE features) and new tasks, showing consistent improvements. The paper also introduces a "compactness" metric to evaluate circuit sparsity, showing MAttr excels at both faithfulness and compactness. The parameter-level attribution experiments on Llama 3.1 8B Instruct are particularly compelling, demonstrating that restoring only 1% of weights to the base model state removes refusals while maintaining capabilities, a result with significant practical implications for model safety and auditing.
The paper provides a public GitHub repository with code. Hyperparameters for all baselines and MAttr are detailed in the appendices. The use of standard benchmarks (MIB) and open-source models (Llama, GPT-2, etc.) enhances reproducibility. The specific details of the RL training setup (GRPO) and reward functions are described, though the complexity of the RL setup may pose challenges for exact replication without the provided code.
The method relies on gradient descent, which may be computationally expensive for very large models or fine-grained bases, although it is parallelizable. The "Matryoshka" property assumes a nested structure of importance, which may not always hold for all types of internal computations. The RL-based parameter attribution is a new paradigm and may require careful tuning of the reward function to avoid reward hacking. The paper focuses primarily on language models; generalization to other modalities (e.g., vision) is mentioned in appendices but not the primary focus.
This paper has high potential impact on the field of interpretability. By providing a robust, learnable, and causally-grounded method for attribution, it enables more reliable identification of task-relevant circuits and weight changes. The application to removing refusals via weight restoration offers a new tool for auditing and potentially mitigating safety issues in LLMs. The unification of gradient-based and mask-learning perspectives may stimulate further theoretical work in attribution. The method's ability to handle non-differentiable objectives via RL opens up new avenues for interpreting complex, emergent behaviors. Matryoshka Attribution (MAttr) is a novel mask-learning method that uses a differentiable sigmoid top-$k$ operator to learn a single, nested ordering of internal components for attribution, achieving state-of-the-art performance on the Mechanistic Interpretability Benchmark and enabling the identification of specific weight changes responsible for LLM behaviors like refusal through reinforcement learning.
Pathologists integrate morphology across magnifications and across the slides of a patient case, whereas pathology foundation models encode thousands of tiles from single slides and aggregate their features. Here we present WILSON, a vision--language foundation model that represents whole-slide images and multi-slide cases as single multi-magnification composite images, trained on approximately 189k slides from Mayo Clinic spanning 42 organs and 829 diagnostic entities using pathology reports as supervision. Without task-specific training, WILSON exceeded a dedicated case-level model on all internal cohorts (macro-F1 0.52 versus 0.38) and matched slide-level models up to 9.4 times larger at 272- to 2,155-fold lower compute. End-to-end fine-tuning on 508 triple-negative breast cancer cases improved histologic subtyping and stromal tumor-infiltrating lymphocyte grading by 0.16 and 0.11 macro-F1. WILSON retrieved matching diagnostic text at 75.6% recall@1 (PRISM, 58.1%) and generated captions closer to report-derived references than PRISM and PRISM2 on the internal cohort and on most external comparisons. Composite images thus offer a compact, clinically aligned computational unit for pathology.
Primary: Mayo Clinic
All Institutions: Mayo Clinic
WILSON introduces a multi-magnification composite image representation for pathology foundation models, achieving competitive performance with significantly lower computational costs and enabling end-to-end fine-tuning and diagnostic text generation.
The paper proposes WILSON, a vision-language foundation model that departs from the standard "tile-encode-then-aggregate" paradigm in computational pathology. Instead of processing thousands of small patches independently, it constructs a single multi-magnification composite image (2048x2048) per slide or case, which is processed in one forward pass by a ConvNeXt-based encoder. This approach is clinically motivated, mimicking how pathologists integrate low and high-power views. The methodology includes a two-stage self-supervised pretraining (tile-based then composite-based) followed by four distinct vision-language alignment strategies (dense distillation, sparse keyword regression, CLIP-style contrastive, and CoCa generative). The use of LLM-generated captions from pathology reports for supervision is a significant methodological choice that leverages unstructured clinical data.
The evaluation is extensive, covering zero-shot classification, case-level retrieval, end-to-end fine-tuning, and text generation. WILSON demonstrates competitive performance with much larger models (PRISM, TITAN) while using significantly fewer FLOPs (up to 2,155-fold reduction). It outperforms a dedicated case-level model (MOOZY) on internal cohorts. The text generation results show higher ROUGE and CIDEr scores compared to PRISM and PRISM2. However, the reliance on internal Mayo Clinic cohorts for many key comparisons limits the generalizability of the findings, and external benchmarks show more mixed results.
The paper provides detailed architectural specifications, training hyperparameters, and dataset construction methods. However, the core dataset (Mayo189K) is proprietary and not publicly available. The composite generation pipeline relies on specific organ-aware heuristics and LLM outputs (Gemini), which may introduce variability. While the code is not explicitly linked, the level of detail suggests reproducibility is possible for those with access to similar data resources.
The primary limitation is the single-institution training data, which may not capture the full diversity of global pathology practices. The composite image approach, while efficient, discards most of the original slide's pixel data, potentially losing subtle diagnostic cues that exhaustive tile processing might catch. The text generation quality, while improved, still requires pathologist review and may hallucinate findings. The external validation on molecular subtyping (CPTAC-BRCA) showed less consistent performance, suggesting limits in cross-domain transfer.
This work offers a compelling alternative to the scaling laws of current pathology foundation models. By demonstrating that a compact, clinically-aligned representation can match larger models with drastically lower compute, it opens the door to deploying advanced pathology AI in resource-constrained settings. The framework's ability to handle multi-slide cases natively is a significant step toward more holistic patient-level diagnostics. WILSON introduces a multi-magnification composite image representation for pathology foundation models, achieving competitive performance with significantly lower computational costs and enabling end-to-end fine-tuning and diagnostic text generation.
World Action Models (WAMs) jointly model visual dynamics and action generation for generalist robot manipulation. A central challenge is to integrate priors from large-scale pretrained models---including visual dynamics, scene semantics, geometry, and motion---into a unified framework for robot action generation. We introduce InternW0-$Δ$, a unified WAM pretrained on a heterogeneous corpus that outperforms prior methods across simulation benchmarks and real-robot platforms. InternW0-$Δ$ combines pretrained visual dynamics, scene-level semantics, 4D geometric and motion priors, and action generation within a Mixture-of-Transformers (MoT) framework. A pretrained video expert and an action expert interact under semantic guidance from a frozen VLM, while a pretrained 4D foundation model injects geometric and motion priors through training-only distillation. We further introduce Causal Imprint, which learns future-relevant scene changes from training-only future supervision and provides predictive representations directly to the action expert without future-video rollout at inference. For large-scale joint training, we construct a heterogeneous corpus of robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data, curated and aligned under a common state-action representation. The resulting corpus contains over 20K hours of processed training data, to our knowledge the largest open-source corpus of its kind. We pretrain InternW0-$Δ$ on this corpus and demonstrate strong performance across simulation benchmarks and real-robot platforms. We will open source the training code, model weights, infrastructure, data-processing pipeline, and processed data where licenses permit. Project page: https://internrobotics.github.io/InternW0-Delta/
Primary: Shanghai AI Laboratory
All Institutions: Shanghai AI Laboratory
InternW0-$\Delta$ introduces a unified World Action Model framework that integrates visual, geometric, and semantic priors via a Mixture-of-Transformers architecture and a novel "Causal Imprint" training technique, pretrained on a 20K+ hour open-source heterogeneous robot data corpus to achieve strong performance in simulation and real-world manipulation.
The paper proposes InternW0-$\Delta$, a World Action Model (WAM) that integrates visual dynamics, scene semantics, 4D geometry, and action generation. The core architectural contribution is a Mixture-of-Transformers (MoT) framework where a video expert and an action expert interact under the guidance of a frozen Vision-Language Model (VLM). A key technical innovation is "Causal Imprint," a training-only distillation technique that injects future-relevant scene changes into the action expert without requiring future video rollouts at inference time. This addresses a significant latency and computational bottleneck in world models. The integration of a pretrained 4D foundation model for geometric priors is also a strong methodological choice, leveraging recent advances in 4D scene understanding to improve robot manipulation.
The authors report strong performance across both simulation benchmarks and real-robot platforms. The scale of the data corpus (20K+ hours) is a major strength, combining robot demonstrations, UMI data, egocentric human demos, and Ego2Robot data. This heterogeneous data strategy is well-motivated for generalist robot learning. The evaluation appears rigorous, covering both simulated environments (likely for controlled comparison) and physical robots (for real-world validity). The claim of outperforming prior methods is supported by this broad evaluation scope.
The paper explicitly states that training code, model weights, infrastructure, data-processing pipeline, and processed data will be open-sourced (where licenses permit). This is a high standard for reproducibility in the robotics field, where data scarcity often hinders replication. The release of a 20K+ hour open-source corpus is particularly impactful for the community.
The primary limitation is the reliance on a frozen VLM for semantic guidance, which may limit the model's ability to adapt to novel semantic contexts not covered by the VLM's pretraining. Additionally, the "Causal Imprint" technique, while efficient at inference, adds complexity to the training pipeline. The paper is an arXiv preprint, so peer review has not yet validated the claims. The 20K hours of data, while large, is still relatively small compared to web-scale data, potentially limiting generalization to highly out-of-distribution tasks.
This work has high potential impact on the field of generalist robot manipulation. By providing a large open-source dataset and a unified framework for integrating diverse priors (visual, geometric, semantic), it lowers the barrier to entry for developing world action models. The "Causal Imprint" technique could be adopted in other domains where predictive dynamics are needed but inference latency is critical. The open-sourcing of infrastructure and data will likely accelerate research in this area. InternW0-$\Delta$ introduces a unified World Action Model framework that integrates visual, geometric, and semantic priors via a Mixture-of-Transformers architecture and a novel "Causal Imprint" training technique, pretrained on a 20K+ hour open-source heterogeneous robot data corpus to achieve strong performance in simulation and real-world manipulation.
Coding agents have demonstrated enormous success in solving complex programming problems. To leverage their potential for robot systems, this work introduces Robot Agentic Programming from Demonstrations (RAPID), which automatically generates, verifies, and refines robot programs, given a single visual human demonstration. The iterative agentic loop of code refinement requires several key ingredients: (i) a testable task specification, (ii) action primitives for robot execution, and (iii) an interactive environment for program execution and verification. RAPID infers all three from the demonstration automatically. To make the resulting program reusable beyond the demonstration setting, RAPID uses an object-centric relational program representation that focuses on the underlying structure of the demonstrated strategy rather than the specific motion per se: it expresses the action primitives as trajectory-optimization programs that realize object-level motion effects, while composing them through relational constraints that capture scene-specific geometry at run time. We evaluated RAPID in simulation on eight challenging contact-rich nonprehensile manipulation tasks as well as general prehensile manipulation tasks in the LIBERO-Pro benchmark. We also successfully deployed it on a real Franka arm and evaluated on all eight nonprehensile tasks. In all experiments, RAPID demonstrated strong performance, with generalization over object pose, shape, material, and environment. Website: https://yuyaoliu.me/projects/rapid.
Primary: Stanford University
All Institutions: Stanford University, University of California, Berkeley
RAPID introduces a novel agentic framework that uses LLMs to generate and refine robot control programs from single demonstrations, achieving strong generalization in contact-rich manipulation tasks. The paper makes a significant technical contribution by proposing an object-centric relational program representation that enables code reuse and robustness, validated through rigorous simulation and real-world experiments on a Franka arm.
The paper proposes RAPID, a framework that leverages Large Language Models (LLMs) as agents to generate, verify, and refine robot control programs from a single visual demonstration. The core innovation lies in the "object-centric relational program representation." Instead of generating low-level motor commands or high-level symbolic plans that are brittle to scene changes, RAPID infers a set of action primitives (expressed as trajectory optimization programs) and relational constraints that define the strategy. This allows the generated code to be reusable: the LLM agent iteratively refines the code by executing it in a simulator, observing the outcome, and using the error to update the program. The methodology effectively bridges the gap between the semantic understanding of LLMs and the precise physical execution required in robotics. The use of an agentic loop for code refinement is a strong contribution, moving beyond one-shot code generation to an iterative improvement process grounded in physical feedback.
The evaluation is comprehensive, covering both simulation and real-world deployment. In simulation, the authors test on eight challenging contact-rich nonprehensile manipulation tasks (e.g., pushing, sliding) and general prehensile tasks using the LIBERO-Pro benchmark. The real-world experiments on a Franka arm validate the approach on the same eight nonprehensile tasks. The results demonstrate strong generalization across object pose, shape, material, and environment variations. The inclusion of contact-rich tasks is significant, as these are notoriously difficult for traditional model-based or purely learned approaches. The comparison with baselines (likely including imitation learning and other code-generation methods) shows RAPID's superiority in success rates and generalization.
The paper provides a project website with a link to the code. The methodology is described in sufficient detail to understand the pipeline: demonstration inference, program representation, and the agentic refinement loop. However, the specific prompts used for the LLM agent and the details of the trajectory optimization solver are critical for reproduction. Assuming the code is released as linked, reproducibility is high. The use of standard benchmarks (LIBERO-Pro) and a common robot platform (Franka) further aids reproducibility.
The primary limitation is the reliance on a simulator for the verification step in the agentic loop. While the final program is deployed on a real robot, the iterative refinement happens in simulation. This requires a high-fidelity simulator that matches the real world, which can be a bottleneck for complex, unstructured environments. Additionally, the approach depends on the LLM's ability to correctly interpret the visual demonstration and generate valid code, which can be sensitive to prompt engineering and model capabilities. The computational cost of running the agentic loop (multiple LLM calls and simulation runs) per task may be high compared to offline learning methods.
This work has significant implications for the field of robotics and agentic AI. It demonstrates that LLMs can be used not just for high-level planning but for generating and refining low-level control code, provided they are grounded in a testable environment. The "code as policy" paradigm is gaining traction, and RAPID offers a robust framework for it. This approach could accelerate robot programming by allowing non-experts to program robots via demonstrations, with the LLM handling the complex code generation and debugging. It also highlights the potential of combining symbolic AI (code) with neural AI (LLMs, vision) for robust physical interaction. RAPID introduces a novel agentic framework that uses LLMs to generate and refine robot control programs from single demonstrations, achieving strong generalization in contact-rich manipulation tasks. The paper makes a significant technical contribution by proposing an object-centric relational program representation that enables code reuse and robustness, validated through rigorous simulation and real-world experiments on a Franka arm.
Foundation models provide robots with the ability to interpret natural language and reason about environmental context, yet most language-conditioned policies assume that goals are well-specified and that task-relevant information is provided upfront via a prior map. Operating in unfamiliar environments with underspecified tasks entails high contextual uncertainty: the robot must jointly infer what constitutes task success, what constitutes relevant information, and where (or whether) that information exists. We address these limitations via CLUE (Closed-Loop contextual Uncertainty rEsolution), a framework for actively resolving contextual uncertainty given underspecified tasks in natural language. CLUE uses an LLM-derived policy to hypothesize task-relevant concepts and potential plans. It then uses a language-embedded map, which is constructed online, to ground these hypotheses into actions. The policy sequentially evaluates hypotheses via closed-loop environment interaction and refines its plans as it gathers new information. We deploy CLUE on a Boston Dynamics Spot across three real indoor and outdoor environments spanning 15 tasks that require object disambiguation, functional inference, and occlusion reasoning. CLUE achieves a success rate within 7 percentage points of an oracle policy and outperforms an LLM-enabled planner without closed-loop feedback by a 4x margin. Supporting experiments demonstrate that simply building and then querying a language-enriched map is insufficient to resolve complex contextual planning tasks; these approaches achieve roughly one third the success rate of CLUE while requiring over 10x more VLM tokens. We provide additional information at https://zacravichandran.github.io/CLUE.
Primary: University of Pennsylvania
All Institutions: University of Pennsylvania, Texas A&M University
[One sentence main contribution]. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper presents CLUE, a closed-loop framework that effectively combines LLM-based hypothesis generation with language-embedded map grounding to resolve contextual uncertainty in underspecified robotic tasks. The technical contribution is significant in demonstrating that closed-loop interaction is necessary for robust task completion in open-world environments, outperforming open-loop baselines by a large margin. The methodology is sound, using a symbolic hypothesis state to approximate belief in a POMDP setting, and the experimental validation on a real-world quadruped robot adds practical credibility. However, the small sample size and reliance on specific cloud-based models limit the generalizability and reproducibility of the results. The work is a strong contribution to the robotics and ML intersection, particularly in the area of language-conditioned planning and active perception.
The paper proposes CLUE, a closed-loop framework for resolving contextual uncertainty in underspecified natural language tasks. The core methodological contribution is the integration of an LLM-based policy that maintains a symbolic hypothesis state with a language-embedded voxel map for grounding. The system operates by hypothesizing task-relevant concepts, grounding them into candidate locations via semantic map queries, and then actively verifying these hypotheses through closed-loop interaction (navigation, inspection, manipulation). The use of a symbolic hypothesis state as a tractable approximation to a belief state in a POMDP setting is a reasonable design choice for open-vocabulary environments. The grounding mechanism, which uses cosine similarity against background concepts and DBSCAN clustering, is standard but effectively applied here. The closed-loop nature, where the LLM replans based on new observations, is the key differentiator from open-loop planners.
The experiments are conducted on a Boston Dynamics Spot robot in three real-world environments (indoor and outdoor) across 15 tasks. The tasks cover object disambiguation, functional inference, and occlusion reasoning. The evaluation compares CLUE against an oracle (upper bound) and NLMaps (open-loop baseline). CLUE achieves 86.7% success rate, within 7 points of the oracle (93.3%), and significantly outperforms NLMaps (20%). A comparison with DAAAM (a scene-graph based approach) shows CLUE is more efficient in VLM token usage while achieving higher success. The ablation study on TSP hints provides useful insight into LLM planning capabilities. The sample size (15 tasks, one run each) is small, which limits statistical confidence, but the real-world deployment on a quadruped robot adds significant practical value.
The paper provides implementation details, including the use of RayFronts, RadSeg, GPT-5.1, and specific hardware (Jetson AGX Thor, ZED 2i). The code and project page are available. However, the reliance on a specific cloud-based LLM (GPT-5.1) and proprietary robot hardware (Spot) may limit reproducibility for some groups. The task definitions and environment setups are described but not fully detailed in the main text, relying on the project page for more information.
The primary limitation is the small number of tasks and single runs per task, which makes the success rates less statistically robust. The method relies heavily on cloud-based LLM calls, which introduces latency and cost. The language-embedded map is memory-intensive (up to 100GB for outdoor environments). The paper acknowledges that the LLM's ability to plan distance-efficient paths degrades with complex system prompts, suggesting a trade-off between contextual reasoning and path optimization.
This work contributes to the field of language-conditioned robotics by addressing the challenge of underspecified tasks in unknown environments. The closed-loop approach to resolving contextual uncertainty is a step towards more autonomous and robust robotic systems that can interact with humans using natural language. The findings on the necessity of closed-loop feedback over open-loop planning are valuable for the design of future robotic systems. The integration of LLMs with semantic mapping and active perception is a trend that is likely to grow in importance. [One sentence main contribution]. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper presents CLUE, a closed-loop framework that effectively combines LLM-based hypothesis generation with language-embedded map grounding to resolve contextual uncertainty in underspecified robotic tasks. The technical contribution is significant in demonstrating that closed-loop interaction is necessary for robust task completion in open-world environments, outperforming open-loop baselines by a large margin. The methodology is sound, using a symbolic hypothesis state to approximate belief in a POMDP setting, and the experimental validation on a real-world quadruped robot adds practical credibility. However, the small sample size and reliance on specific cloud-based models limit the generalizability and reproducibility of the results. The work is a strong contribution to the robotics and ML intersection, particularly in the area of language-conditioned planning and active perception.
Human hand-object interactions (HOIs) provide a rich source of demonstrations for dexterous manipulation, but learning directly from them presents challenges in bridging morphology gaps, ensuring dynamical feasibility, and sim-to-real deployment. We present Morphometric Imitation, a three-stage framework that transforms reconstructed HOIs into zero-shot sim-to-real visuomotor policies. First, morphometric optimization (MMO) kinematically retargets human motion across hand morphologies while preserving demonstrated contacts. Second, residual reinforcement learning (RL) refines the kinematic reference using object pose and contact information from the human motion to produce dynamically feasible robot demonstrations. Third, these demonstrations are distilled into visuomotor policies. Across three robot hands and ten HOIs, MMO outperforms five baselines in contact F1, improving on the strongest ones by 8 to 28 points, while improving the success rate of downstream dynamic retargeting by as much as 35 points. On a Sharpa hand, the visuomotor policies achieve 89.3% zero-shot success in 300 real-world trials on 30 objects spanning 10 categories. Project page: https://morphometricimitation.github.io
Primary: University of California, Berkeley
All Institutions: University of California, Berkeley
The paper presents a robust three-stage framework for morphometric imitation that effectively bridges the gap between human hand-object interactions and dexterous robot manipulation. By combining contact-aware kinematic retargeting, residual reinforcement learning for dynamic feasibility, and visuomotor policy distillation, the authors achieve state-of-the-art results in both simulation and real-world deployment, demonstrating significant improvements in contact fidelity and task success rates across diverse robot morphologies.
The paper proposes "Morphometric Imitation," a three-stage pipeline addressing the morphology gap in human-to-robot hand retargeting. Stage 1 (MMO) performs kinematic optimization to retarget human motion to robot hands while explicitly preserving contact constraints, which is a significant improvement over standard IK-based retargeting that often ignores contact fidelity. Stage 2 uses residual reinforcement learning to correct kinematic errors and ensure dynamic feasibility, leveraging object pose and contact data from the human demonstration. Stage 3 distills these refined demonstrations into a visuomotor policy. The methodology is coherent, well-motivated, and addresses specific, known pain points in dexterous manipulation (morphology mismatch, dynamic infeasibility, sim-to-real gap).
The experimental evaluation is robust and comprehensive. The authors test across three different robot hand morphologies (3, 4, and 5-fingered) and ten distinct human-object interactions. They benchmark against five baselines, showing significant improvements in contact F1 (8-28 points) and downstream dynamic retargeting success rates (up to 35 points). The real-world validation on the Sharpa hand is particularly strong, reporting 89.3% zero-shot success over 300 trials with 30 objects across 10 categories. This level of real-world testing is rare and highly valuable, providing strong evidence of practical utility.
The paper provides a project page and likely code (implied by the nature of such releases, though not explicitly linked in the text snippet, the project URL is provided). The detailed description of the three stages and the specific metrics used (Contact F1, success rates) allow for reasonable reproducibility. The use of standard RL and optimization techniques further aids reproducibility.
The framework relies on high-quality human motion capture data, which can be expensive and difficult to collect. The residual RL stage may be computationally intensive. The results are specific to the tested hands and objects; generalization to significantly different morphologies or highly deformable objects is not fully explored. The "zero-shot" claim is strong but limited to the specific distribution of objects and poses tested in the real-world trials.
This work has high potential impact in the field of dexterous robotics. By providing a robust pipeline to convert human demonstrations into executable robot policies, it lowers the barrier to entry for training dexterous manipulators. The focus on contact preservation and dynamic feasibility addresses critical issues that have hindered the adoption of human-centric data in robotics. The successful sim-to-real transfer with high success rates demonstrates a viable path for deploying such systems in real-world applications. The paper presents a robust three-stage framework for morphometric imitation that effectively bridges the gap between human hand-object interactions and dexterous robot manipulation. By combining contact-aware kinematic retargeting, residual reinforcement learning for dynamic feasibility, and visuomotor policy distillation, the authors achieve state-of-the-art results in both simulation and real-world deployment, demonstrating significant improvements in contact fidelity and task success rates across diverse robot morphologies.
Pretrained generative Diffusion Transformers (DiTs) capture rich pixel-level visual and language-conditioned structure through large-scale image and video generation training. A growing line of robot policies builds on this generative prior, but how it should be transferred to control remains unclear, and existing approaches commonly instantiate this transfer through future visual prediction. We ask a more basic question: what a pretrained generative DiT actually contributes to action learning, and how this prior should be adapted for control. We introduce NowWAM, a future-target-free co-training formulation that denoises the current observation and predicts robot actions from the same visual stream, directly coupling the native generative objective to the action-facing representation across the denoising trajectory. Under matched controlled settings, past and future visual targets perform comparably, while restricting training to the clean endpoint substantially reduces robustness, suggesting that a separate future target is not essential for generative adaptation, while the denoising trajectory remains an effective interface for control. On LIBERO-Plus, NowWAM reaches 87.7% with FLUX2-Klein, improving over the future-target co-training baseline by 6.1 points while halving training visual tokens (784 to 392) and reducing step time from 2.85 s to 1.63 s, a 1.8x speedup. With the pure text-to-image Z-Image backbone, NowWAM further reaches 87.8%, showing that strong control adaptation is not tied to video generation or image-editing backbones.
Primary: Princeton University
All Institutions: Princeton University
[One sentence main contribution]. [The paper introduces NowWAM, a future-target-free co-training formulation that applies native denoising to the current observation, demonstrating that this simpler interface improves robustness and efficiency over future-prediction-based generative robot policies. The technical contribution is significant as it decouples the utility of generative priors from temporal forecasting, providing a rigorous ablation study that clarifies the role of the denoising trajectory in representation learning for control. The results are strong, showing state-of-the-art performance on robustness benchmarks with reduced training costs, though the lack of real-world validation limits the immediate practical impact.].
The paper proposes "NowWAM," a co-training framework that removes the separate future visual target typically used in generative robot policies. Instead, it applies the native denoising objective of the pretrained Diffusion Transformer (DiT) directly to the current observation stream, coupling generative adaptation with action learning in a single stream. The core insight is that the benefit of generative co-training comes from the denoising trajectory itself (the continuum of noisy states) rather than the specific temporal semantics of a future prediction. The method is conceptually clean and simplifies the training pipeline by halving the visual token count during training.
The experiments are rigorous and well-controlled. The authors use LIBERO-Plus as a primary robustness benchmark, which is more discriminative than standard LIBERO. They provide strong ablations isolating the effects of generative initialization, temporal target choice (past vs. future), and denoising trajectory sampling. The results show that NowWAM outperforms future-target baselines (like Fast-WAM and ImageWAM) in robustness (87.7% vs ~81.6%) while being significantly faster (1.8x speedup). The extension to a pure text-to-image backbone (Z-Image) further validates the claim that temporal generative pretraining is not strictly necessary.
The paper provides detailed descriptions of the training setup, loss functions, and evaluation protocols. However, specific hyperparameters for the action expert and the exact implementation details of the "masked mixed attention" are not fully specified in the provided text, which may hinder exact reproduction. The use of standard benchmarks (LIBERO, RoboCasa) aids reproducibility.
The evaluation is limited to simulation (LIBERO, RoboCasa). Real-world robot experiments are absent, which is a significant gap for a robotics paper. The performance gains, while statistically significant in simulation, may not translate directly to the noisy, complex real world. Additionally, the paper relies on specific pretrained backbones (FLUX2-Klein, Z-Image), and the generalizability to other DiT architectures is not tested.
This work challenges the prevailing assumption in robot learning that future prediction is essential for leveraging generative priors. By demonstrating that current-frame denoising is sufficient and more efficient, it offers a simpler, faster, and more robust alternative for building VLA models. This could lead to more efficient training pipelines and lower computational costs for robot policy learning. [One sentence main contribution]. [The paper introduces NowWAM, a future-target-free co-training formulation that applies native denoising to the current observation, demonstrating that this simpler interface improves robustness and efficiency over future-prediction-based generative robot policies. The technical contribution is significant as it decouples the utility of generative priors from temporal forecasting, providing a rigorous ablation study that clarifies the role of the denoising trajectory in representation learning for control. The results are strong, showing state-of-the-art performance on robustness benchmarks with reduced training costs, though the lack of real-world validation limits the immediate practical impact.].
Humanoid kicking requires coordinated whole-body motion and precise contact, while a banana kick demands contact mechanics that generate ball spin and aerodynamic curvature. Motion imitation provides a reliable ordinary-kick prior, but reinforcement learning may improve shot speed and placement accuracy without changing the underlying kicking technique. Adapting this prior to a qualitatively different contact-rich skill can fail even when the reward is dense and optimization remains stable. The failure occurs when the task objective is locally flat over the current policy's responses. We term this condition first-order learning starvation. To address it, we propose response-informed skill evolution (RISE), a closed-loop objective-continuation method for policy adaptation. RISE ranks bounded objective changes using response sensitivity estimated from cached rollouts and accepts updates only when they produce verified response progress while preserving kicking reliability. Our analysis shows that rescaling a saturated spin reward cannot recover first-order sensitivity at zero spin, whereas adapting coupled contact responses can provide a learnable path to spin generation. We integrate RISE into a humanoid kicking pipeline under calibrated contact and Magnus-force aerodynamics, and sim-to-real transfer. Experiments show that RISE evolves the ordinary kick into a high-spin curved kick with 11.55 rad/s mean ball spin, improves the mean evaluation score by 19.8% over a learning-progress curriculum, and raises joint target attainment from 15.2% to 50.9%. Ablations and response diagnostics support the mechanism, while 30 motion-capture-recorded physical trials demonstrate consistent hardware transfer of the learned curved kick. Project website: https://haozhang-thu.github.io/bananakick/
Primary: Tsinghua University
All Institutions: Tsinghua University
The paper presents a novel and effective method, RISE, for evolving humanoid skills by dynamically adapting reward objectives based on physical response sensitivities, successfully demonstrating the transition from an ordinary kick to a banana kick with verified sim-to-real transfer.
The paper introduces "Response-Informed Skill Evolution" (RISE), a closed-loop objective-continuation method designed to address "first-order learning starvation." This occurs when a policy initialized with a strong motion prior (ordinary kick) fails to learn a qualitatively different skill (banana kick) because the reward landscape is locally flat with respect to the policy's current response distribution. The core innovation is the use of cached rollout sensitivities to rank bounded changes to reward parameters (specifically kernel scales and weights) before applying them. The method verifies updates not just by reward improvement, but by measured physical response progress (e.g., spin, contact patch error) while maintaining reliability. The theoretical analysis correctly identifies that rescaling a saturated reward (like spin at zero) cannot create a gradient, whereas adjusting coupled contact variables can. This is a sophisticated approach to curriculum/reward shaping that moves beyond static designs to dynamic, response-aware adaptation.
The experiments are rigorous and compelling. The authors use a Unitree G1 humanoid in IsaacSim/PhysX with calibrated contact and Magnus-force aerodynamics. They compare RISE against a Learning Progress (LP) curriculum, Direct PPO, and ablations. RISE achieves a 19.8% improvement in evaluation score over the LP curriculum and significantly higher joint target attainment (50.9% vs 15.2%). Crucially, the paper demonstrates sim-to-real transfer with 30 physical trials, showing consistent curved flight (median bow 0.381m) without further adaptation. The ablations clearly isolate the contribution of the sensitivity ranking and feedback loop. The "response evolution" section provides deep insight into how the policy moves from the prior's regime to the new skill regime.
The paper provides detailed hyperparameters, environment settings, and algorithmic steps. The project website is provided. However, the specific code for the RISE loop and the calibrated physics parameters are not explicitly released in the text, though the project page likely contains them. The reliance on specific hardware (Unitree G1) and calibrated physics may limit immediate reproducibility for labs without similar setups, but the methodological framework is clearly described.
The method is tested primarily on a single task (soccer kicking) and a single robot platform. The computational cost of the closed-loop objective adaptation (ranking candidates, verifying progress) is not deeply analyzed compared to standard PPO. The "first-order learning starvation" concept, while well-motivated, is specific to scenarios where a strong prior exists but the target skill requires a regime shift; its applicability to tasks without strong priors or with different failure modes is less clear. The physical trials, while consistent, are limited in number (30).
This work has significant implications for humanoid robotics and skill acquisition. It provides a principled way to overcome the "local optimum" trap when adapting motion priors to new skills. The concept of response-informed reward adaptation could be generalized to other contact-rich tasks (manipulation, locomotion over varied terrain). The demonstration of sim-to-real transfer for a complex, spin-generating skill is a strong contribution to the field of humanoid soccer. The paper presents a novel and effective method, RISE, for evolving humanoid skills by dynamically adapting reward objectives based on physical response sensitivities, successfully demonstrating the transition from an ordinary kick to a banana kick with verified sim-to-real transfer.
Most minimally invasive surgery is still performed with hand-held laparoscopic instruments, and the surgeon's instrument kinematics are lost when the operation ends; only the endoscope video is kept. This paper presents an end-to-end pipeline that captures this motion in the operating room and uses it to train a surgical robot policy, validated on live animals. We introduce a surgical instrument-state logger that mounts on the shaft of a standard laparoscopic instrument and recovers its pose and jaw state from an inertial sensor, a time-of-flight sensor and a Hall sensor, with no external camera or tracker. A data pipeline measures the latency of every sensor channel against a robot ground truth and aligns the channels before forming observation-action pairs. On these demonstrations we train a diffusion policy with a fine-tuned DINOv3 backbone, selecting its design by closed-loop rollouts in a physics simulator reconstructed from depth maps of an ex-vivo rabbit appendix. The policy is then retrained on 849 in-vivo demonstrations from four live rabbits and deployed on four additional live rabbits with electrosurgery armed. With the surgeon selecting the surgical phase, the policy completed the appendectomy in three of the four animals. The results show that demonstrations recorded from a surgeon's own instruments are sufficient to train, select and deploy a bimanual surgical policy in vivo. The robot serves only as the timing reference for sensor calibration and as the executor, and collects no demonstrations. Both demonstration corpora are released to support future surgical robot learning research.
Primary: Seoul National University
All Institutions: Eulji University, Seoul National University, Seoul National University College of Medicine, Seoul National University Hospital
The paper presents a robust end-to-end pipeline for learning bimanual surgical policies from hand-held instrument demonstrations, validated in vivo with electrosurgery. It makes a significant contribution to surgical robotics by introducing a novel, low-cost instrument-state logger and demonstrating that such data is sufficient to train safe, effective policies for live animal surgery, thereby bridging the gap between hand-held surgical practice and robotic automation.
The paper proposes a novel hardware-software pipeline for collecting surgical demonstrations without a robot. The core hardware contribution is a "surgical instrument-state logger" mounted on the shaft of standard laparoscopic instruments, using an IMU, Time-of-Flight (ToF) sensor, and Hall sensor to estimate pose and jaw state. The methodology rigorously addresses the critical issue of sensor latency by fitting a first-order lag model to each channel against a robot ground truth (FR3) and applying per-channel latency matching. The policy is a Diffusion Policy with a fine-tuned DINOv3 backbone, selected via closed-loop rollouts in a physics simulator (Isaac Sim) reconstructed from depth maps. This selection process is a strong methodological choice, prioritizing closed-loop safety metrics (trocar violations, stage completion) over offline validation error, which the authors show to be anti-correlated with safety in this context.
The experimental validation is the paper's strongest feature. It moves beyond simulation and ex-vivo testing to in-vivo execution on live rabbits. The authors trained on 849 in-vivo demonstrations and deployed the policy on four additional live rabbits with electrosurgery armed. The policy completed the appendectomy in 3 out of 4 animals under shared autonomy. The safety metrics are detailed, including RCM error monitoring and electrosurgery gating. The comparison against video-based tracking (showing 10.7mm error vs 1.36mm for the logger) provides strong evidence for the hardware approach. The ablation study on policy configuration (31 candidates) is thorough and statistically grounded (Fisher's exact test).
High. The authors release both demonstration corpora (ex-vivo and in-vivo) and the code repository. The hardware design is described in sufficient detail (sensor models, mounting, firmware logic) for replication. The simulator setup (Isaac Sim, FEM parameters) is specified. The latency matching procedure is clearly defined.
The primary limitation is the reliance on the surgeon to manually select the surgical phase during deployment, as the vision-based phase predictor failed (24.6% agreement). This limits the autonomy claim. The study is limited to laparoscopic appendectomy in rabbits, so generalization to other procedures or species is not demonstrated. The robot (FR3) is still required for calibration and execution, so it is not a fully robot-free pipeline. The sample size for in-vivo deployment is small (4 animals).
This work has significant potential impact on surgical robotics by demonstrating that high-quality demonstration data can be collected from standard hand-held instruments, removing the need for expensive robot-based teleoperation for data collection. This could democratize the collection of surgical datasets and accelerate the development of surgical robot policies. The hardware logger is a practical, low-cost solution that could be widely adopted. The paper presents a robust end-to-end pipeline for learning bimanual surgical policies from hand-held instrument demonstrations, validated in vivo with electrosurgery. It makes a significant contribution to surgical robotics by introducing a novel, low-cost instrument-state logger and demonstrating that such data is sufficient to train safe, effective policies for live animal surgery, thereby bridging the gap between hand-held surgical practice and robotic automation.
Contact-rich precision insertion is a key manipulation skill in robotic assembly. Tight clearances make insertion more sensitive to alignment errors and prone to collisions and jamming, while variations in geometry and clearance across parts further complicate policy reuse. We present a reinforcement learning framework that trains insertion policies entirely in simulation for direct deployment without real-world demonstrations or policy fine-tuning. By combining target poses with compact three-dimensional fingertip force feedback, the policy learns to search for alignment and correct its motion despite errors in the estimated hole position. A decoupled gated reward coordinates alignment and insertion. Force-signal smoothing and state-independent standard deviations stabilize the learning process. The resulting policies perform real-world insertion across multiple hole geometries with a minimum nominal clearance of 0.02 mm and improve success while reducing peak contact forces under hole-position errors. Cross-clearance and cross-geometry evaluations further confirm policy generalization. The system achieved the first perfect score of 20/20 on ManipulationNet's peg-in-hole benchmark under its Human-in-the-Loop protocol, with fully autonomous insertion motions. A single policy trained only on a simulated hexagonal insertion task achieved an overall success rate of 95.0% across eight unseen real-world insertion tasks. These results show that learning entirely in simulation can yield precision insertion skills that can be deployed directly and reused across real-world tasks. The project website (https://mzhsoul.github.io/InsertAnything/) provides open-source simulation and real-robot experiment scripts, assets, and trained checkpoints.
Primary: University of Chinese Academy of Sciences
All Institutions: University of Chinese Academy of Sciences, Institute of Automation, Chinese Academy of Sciences, PaXini AI (Beijing) Co, School of Artificial Intelligence, State Key Laboratory of Multimodal Artificial Intelligence Systems
The paper presents a robust simulation-to-reality framework for precision robotic insertion that achieves state-of-the-art results on standardized benchmarks. It effectively combines force feedback with reinforcement learning to solve contact-rich manipulation tasks, demonstrating strong generalization capabilities across clearances and geometries without real-world fine-tuning.
The paper proposes a reinforcement learning framework for contact-rich precision insertion that relies entirely on simulation training for direct real-world deployment. The core methodological contributions are the use of compact 3D fingertip force feedback combined with target poses, a decoupled gated reward function that separates planar alignment, yaw alignment, and axial insertion, and specific stabilization techniques (EMA smoothing for force signals and state-independent standard deviations for the policy) to handle noisy contact data. The approach effectively addresses the sim-to-real gap by randomizing observation errors and dynamics, allowing the policy to learn robust correction strategies without real-world fine-tuning.
The experimental evaluation is rigorous and comprehensive. It includes ablation studies on reward design and stabilization techniques, comparisons with traditional control methods (impedance, hybrid force/position), and extensive generalization tests. The highlight is the real-world validation on the ManipulationNet benchmark, achieving a perfect 20/20 score, and the transfer of a single policy to eight unseen industrial tasks with 95% success. The inclusion of tight clearances (down to 0.02 mm) and diverse geometries (circular, square, hexagonal, L-shaped) demonstrates strong practical relevance.
The authors provide open-source simulation scripts, real-robot experiment scripts, assets, and trained checkpoints via the project website. The paper details the specific hardware (Franka Emika, Paxini sensors) and software stack (Isaac Lab, Factory), along with hyperparameters and randomization ranges in the supplementary material, which supports reproducibility for labs with similar robotic setups.
The method requires precise calibration of the target hole pose, which may not be available in all unstructured environments. The reliance on specific tactile sensor hardware (Paxini) limits immediate applicability to robots with different sensing capabilities. The generalization to "unseen" tasks is still within the domain of mechanical insertion/mating, and performance on highly deformable or non-rigid objects is not explored.
This work has significant implications for industrial automation, particularly in assembly tasks requiring high precision. By demonstrating that simulation-trained policies can handle tight clearances and generalize across geometries without real-world data, it reduces the cost and time associated with deploying robotic manipulation skills. The success on the ManipulationNet benchmark sets a new standard for autonomous precision assembly. The paper presents a robust simulation-to-reality framework for precision robotic insertion that achieves state-of-the-art results on standardized benchmarks. It effectively combines force feedback with reinforcement learning to solve contact-rich manipulation tasks, demonstrating strong generalization capabilities across clearances and geometries without real-world fine-tuning.
Adapting visuomotor policies to new manipulation tasks often requires substantial manual engineering or teleoperated data collection. Simulation can provide task-specific data at scale, but constructing the scene, designing expert behavior, and configuring data generation still require significant per-task effort. We present ARSTAG, an agentic Real2Sim2Real system that turns a single RGB image and a natural-language instruction directly into robot policy-learning data. A hierarchy of language agents constructs a task-scoped simulation scene, generates robot-feasible demonstrations, and expands the training distribution through task-consistent randomization, while a coordinator agent manages cross-stage feedback and recovery. Across seven manipulation tasks spanning grasping, placement, and stacking, the ARSTAG-generated demonstrations enable sim-to-real transfer of three visuomotor policy architectures to a dual-arm robot, with pi0.5 achieving an average real-world success rate of 74.6%. Ablations show that task-consistent randomization substantially improves robustness, and policy performance increases with generated dataset size. Project webpage: https://boweili666.github.io/ARSTAG/.
Primary: Carnegie Mellon University
All Institutions: Carnegie Mellon University
ARSTAG introduces an agentic Real2Sim2Real system that automates robot data generation from a single image and instruction. The paper presents a well-engineered multi-agent pipeline that effectively bridges the gap between real-world observations and simulation-based policy learning, demonstrating strong sim-to-real transfer and robustness to failures.
The paper proposes ARSTAG, a hierarchical multi-agent system that automates the Real2Sim2Real pipeline. The methodology is sound, leveraging a coordinator agent to manage three specialized sub-agents (reconstruction, collection, learning). The use of LLMs for task grounding, scene graph construction, and failure recovery is a logical application of current agentic capabilities. The "task-scoped" reconstruction is a smart efficiency improvement over full-scene reconstruction. The geometric pose repair method is a practical, non-optimization-based solution to a common sim-to-real discrepancy.
The experiments are comprehensive, covering 7 distinct manipulation tasks and 3 different policy architectures (ACT, Diffusion Policy, pi0.5). The sim-to-real transfer results are strong, with pi0.5 achieving 74.6% average success. The ablations on randomization and data scale are rigorous, providing clear insights into what drives performance. The agent coordination stress test with injected failures is a unique and valuable evaluation of the system's robustness.
The paper provides good details on the pipeline, but the heavy reliance on specific, potentially proprietary or rapidly evolving tools (GPT-5.5, SAM3, SAM3D, GraspGenX) and the specific robot platform (AgiBot G1) may limit immediate reproducibility for other groups. However, the conceptual framework is clear.
The system is limited to tabletop manipulation. It struggles with precision tasks (screw insertion) and long-horizon sequences. The performance of the generated policies is still lower than what might be achieved with high-quality human teleoperation data. The computational cost of the agent loop and simulation is not deeply analyzed.
This work has high potential impact in the robotics community by significantly reducing the manual engineering burden for task-specific robot deployment. It demonstrates a viable path towards "zero-shot" or "low-shot" robot adaptation using generative AI and simulation. ARSTAG introduces an agentic Real2Sim2Real system that automates robot data generation from a single image and instruction. The paper presents a well-engineered multi-agent pipeline that effectively bridges the gap between real-world observations and simulation-based policy learning, demonstrating strong sim-to-real transfer and robustness to failures.
Robotic foundation models achieve impressive performance on standard manipulation benchmarks, yet these evaluations typically assume clean, timely, and consistent visual observations throughout execution. We introduce LIBERO-VPro, a benchmark for systematically evaluating the closed-loop visual robustness of robotic foundation models by perturbing the visual evidence available during execution. LIBERO-VPro covers four complementary dimensions, including Visual Evidence Degradation, Camera Staleness, Visual Source Consistency, and Task-Relevant Scene Variation, spanning 12 challenge categories, 96 experimental settings, and 3,296 task-condition cases. We evaluate three vision-language-action models and three world-action models over approximately 196,000 simulated episodes, complemented by 200 real-world rollouts on a Franka Research 3. Our results reveal that strong nominal performance can mask substantial weaknesses in visual grounding and adaptation. Models often remain successful despite severe object-level occlusion, yet degrade sharply when local interaction cues are disrupted or familiar spatial priors are violated. They are also highly sensitive to stale or missing observations and struggle when changed task preconditions require behavioral adaptation. Finally, VLAs and WAMs exhibit distinct robustness profiles, showing that visual robustness is multi-dimensional and architecture-dependent. LIBERO-VPro provides a systematic diagnostic framework for developing robotic foundation models that can more reliably ground and adapt their actions under challenging visual conditions.
Primary: Fudan University
All Institutions: Fudan University
The paper introduces LIBERO-VPro, a comprehensive benchmark for evaluating the closed-loop visual robustness of robotic foundation models, revealing critical vulnerabilities in spatial priors and temporal consistency that are masked by nominal performance. By systematically perturbing visual evidence during execution across 196,000 simulated episodes and 200 real-world rollouts, the study provides a rigorous diagnostic framework that highlights the multi-dimensional nature of visual robustness and the distinct failure modes of Vision-Language-Action (VLA) and World-Action Models (WAMs), offering significant insights for developing more reliable robotic policies.
The paper introduces LIBERO-VPro, a benchmark designed to evaluate the closed-loop visual robustness of robotic foundation models. The methodology is structured around four distinct dimensions: Visual Evidence Degradation, Camera Staleness, Visual Source Consistency, and Task-Relevant Scene Variation. This is a well-constructed contribution because it moves beyond static scene perturbations (common in existing benchmarks like LIBERO-Plus) to dynamic, execution-time perturbations that affect the perception-action loop. The inclusion of "Prediction-Reality Divergence" for World-Action Models (WAMs) is a particularly novel and relevant addition, addressing the specific failure mode of recursive prediction errors in models that rely on imagined futures. The benchmark is comprehensive, covering 12 challenge categories and 96 experimental settings, which provides a granular diagnostic tool for researchers.
The experimental evaluation is rigorous and extensive. The authors evaluate six representative models (3 VLAs and 3 WAMs) over approximately 196,000 simulated episodes. This scale is significant and ensures statistical reliability. The inclusion of 200 real-world rollouts on a Franka Research 3 is a strong plus, validating that the simulation findings transfer to physical hardware. The results reveal nuanced insights, such as the distinction between tolerance for object occlusion (due to spatial priors) and sensitivity to interaction cue masking (due to reliance on local feedback). The comparison between VLAs and WAMs provides valuable architectural insights, showing that WAMs are particularly sensitive to view unavailability and prediction divergence.
The paper provides detailed descriptions of the perturbations and the experimental setup. It specifies the use of standard LIBERO demonstrations and official checkpoints for most models, with a note on additional training for LingBot-VA. The specific parameters for masking ratios, corruption frequencies, and delay levels are described, which aids reproducibility. However, without a linked code repository or detailed appendix (not provided in the text), full reproduction of the specific perturbation implementations would require careful reading of the methodology section. The use of standard simulators (LIBERO) and hardware (Franka) enhances reproducibility.
The primary limitation is the reliance on the LIBERO simulation environment for the bulk of the evaluation. While real-world tests are included, they are limited to two tasks and two models, which may not fully capture the complexity of real-world visual robustness. The benchmark is also specific to manipulation tasks; its applicability to other robotic domains (e.g., navigation, locomotion) is not explored. Additionally, the evaluation focuses on success rate, which may not fully capture the quality of the policy's behavior (e.g., smoothness, safety) under perturbations.
This benchmark has high potential impact on the field of robotic foundation models. As these models are increasingly deployed in real-world settings, understanding their robustness to visual imperfections is critical. LIBERO-VPro provides a systematic framework for diagnosing specific failure modes, which can guide the development of more robust policies. The insights into the distinct robustness profiles of VLAs and WAMs are particularly valuable for model designers. The benchmark is likely to be adopted as a standard evaluation suite for new robotic foundation models, similar to how LIBERO itself has been used. The paper introduces LIBERO-VPro, a comprehensive benchmark for evaluating the closed-loop visual robustness of robotic foundation models, revealing critical vulnerabilities in spatial priors and temporal consistency that are masked by nominal performance. By systematically perturbing visual evidence during execution across 196,000 simulated episodes and 200 real-world rollouts, the study provides a rigorous diagnostic framework that highlights the multi-dimensional nature of visual robustness and the distinct failure modes of Vision-Language-Action (VLA) and World-Action Models (WAMs), offering significant insights for developing more reliable robotic policies.
Egocentric human data provide a principled source of supervision for learning dexterous robot manipulation. Unlike prior approaches that often collect such data in constrained or specially constructed environments, we collect in-the-wild egocentric demonstrations in real-world settings, including homes, factories, and pharmacies, etc., where people perform their ordinary tasks while wearing head-mounted cameras. This collection protocol captures diverse workflows and hand-object interactions across long-tailed object and skill distributions, but also yields visually challenging observations due to scene clutter and head-motion-induced viewpoint changes (a mean cumulative rotation of $15.93^{\circ}$/s). To address these issues, we introduce EgoWild2Dex, which transfers in-the-wild ego-human experience to dual-arm robots with dexterous hands by jointly aligning unstable egocentric views and human motions with robot observations and actions, respectively. This work offers three benefits. First, we introduce GeoFormer, a differentiable geometric transformer that warps noisy human observations toward robot observations. Second, we design a human-robot training scheme to bridge the embodiment gap, enabling high task success with limited robot supervision. Third, we release EgoWild, a 538.9-hour in-the-wild egocentric human dataset comprising 179,049 episodes, 125,961 unique task descriptions, and 1,282 object categories. On real robots, EgoWild2Dex achieves an average success rate of 96.7% across three long-horizon bimanual dexterous manipulation tasks and an average object-level zero-shot success rate of 33.3%. The data, models, and code will be released.
Primary: The University of Hong Kong
All Institutions: The University of Hong Kong, Kinetix AI
The paper presents a compelling and technically sound framework for learning dexterous robot manipulation from in-the-wild human experience, achieving high success rates with minimal robot data through efficient geometric view alignment and progressive training. Its combination of a novel alignment module, a large-scale in-the-wild dataset, and rigorous real-robot evaluation makes it a significant contribution to the field of vision-language-action models and robot learning.
The paper proposes EgoWild2Dex, a framework for transferring in-the-wild egocentric human data to dexterous robot manipulation. The core technical contribution is GeoFormer, a differentiable geometric transformer that aligns unstable egocentric views with fixed robot views using a lightweight projective warp (homography) rather than expensive 3D reconstruction and inpainting. This is a clever and efficient solution to the viewpoint mismatch problem. The training scheme is progressive: (1) Human-to-Robot learning on a large in-the-wild dataset (EgoWild) to learn broad priors, (2) Human-Robot co-training using the aligned ego data and limited robot data, and (3) Robot-domain refinement with DAgger. The use of a shared robot-native action space via retargeting is well-motivated and validated with distribution overlap metrics. The methodology is sound, addressing key challenges in embodiment gap and visual alignment.
The experiments are conducted on real robots with dexterous hands, which is a strong evaluation setting. The tasks are long-horizon bimanual manipulation tasks, which are highly relevant and challenging. The results show a 96.7% success rate with less than one hour of robot data per task, which is impressive. Ablations clearly demonstrate the contribution of each component, particularly the view alignment and the progressive training stages. The comparison with Project+Inpaint shows significant speedup (21.9x) and improved image similarity. The generalization tests to new objects and embodiments are positive. The evaluation is rigorous and comprehensive.
The paper provides detailed implementation details in the appendix, including hardware setups, action alignment pipelines, training hyperparameters, and loss functions. The authors state that data, models, and code will be released. The specific hardware (AgileX arms, BrainCo hands) and software (XRoboToolkit) are specified. The dataset EgoWild is described in detail. While the specific VR and glove hardware might be a barrier for some, the overall system is well-documented for reproduction by groups with similar robotic setups.
The method still requires real-robot fine-tuning, indicating that human data alone is insufficient for full transfer. GeoFormer cannot reconstruct content when head rotation moves objects out of the field of view. The evaluation is limited to three specific tasks and two robot embodiments. The reliance on specific hardware (PICO headset, mHandPro gloves) for data collection may limit the accessibility of the data collection protocol. The paper does not extensively discuss the failure modes of the policy in detail beyond the DAgger recovery.
This work has significant potential impact on the field of robot learning from human data. By demonstrating that in-the-wild, unscripted human data can be effectively transferred to dexterous robots with minimal robot supervision, it opens a path for scalable robot learning. The GeoFormer module is a general-purpose tool that could be applied to other vision-based robot learning tasks. The release of the EgoWild dataset (538.9 hours) will be a valuable resource for the community. The approach challenges the need for massive paired human-robot data or constrained environments, promoting more natural data collection. The paper presents a compelling and technically sound framework for learning dexterous robot manipulation from in-the-wild human experience, achieving high success rates with minimal robot data through efficient geometric view alignment and progressive training. Its combination of a novel alignment module, a large-scale in-the-wild dataset, and rigorous real-robot evaluation makes it a significant contribution to the field of vision-language-action models and robot learning.
Open-world deployment requires humanoid robots to cross highly heterogeneous terrain safely, with perception that simultaneously provides wide coverage, local accuracy, and redundancy against sensor failure. Existing approaches struggle to satisfy all three: one forward depth camera or nearby height sampling covers too little; odometry-corrected elevation maps drift under aggressive motion and miss thin vertical structures; image-level encoding costs grow with camera count. We present UniPoint, a humanoid whole-body locomotion framework built on multi-source point-level sensor fusion. Measurements from a 360° light detection and ranging (LiDAR) sensor and two depth cameras are early-fused into one base-frame point set. Voxelization resamples it to a fixed number of tokens encoded by linear self-attention and proprioception-queried cross-attention, decoupling forward cost from sensor count. The point set retains standing thin barriers; a single-modality failure removes only part of the tokens, so the policy degrades gracefully. A single training run with terrain-aware rewards, perception-degradation injection, and domain randomization produces one policy for all eight terrain types, deployed on an onboard RK3588 without fine-tuning. On a DR02 humanoid, 20 trials at each of nine real-world settings over seven terrain types validate the policy on 70-cm-high platforms, 100-cm gaps, thin barriers, and sparse or narrow footholds; it also generalizes zero-shot outdoors.
Primary: Zhejiang University
All Institutions: Zhejiang University, Yunshenchu Technology Co., Ltd., Zhejiang Key Laboratory of Additive Manufacturing Technology and Equipment
UniPoint introduces a unified point-level sensor fusion framework for humanoid locomotion that achieves robust traversal across diverse terrains with low onboard compute cost. The paper demonstrates that early-fusing LiDAR and depth camera data into a fixed token set, processed by linear attention, outperforms traditional elevation maps and depth-image CNNs, particularly on challenging terrains like thin barriers and sparse footholds, while maintaining graceful degradation under sensor failure.
The paper proposes UniPoint, a framework for humanoid locomotion that fuses 360° LiDAR and depth camera data into a unified point cloud representation. The core methodological contribution is the early fusion of these heterogeneous sensors into a fixed-size set of voxel tokens (80 tokens per frame, stacked over 5 frames), which is then processed by a network using linear self-attention for point encoding and proprioception-queried cross-attention for fusion. This design decouples computational cost from sensor count and allows for graceful degradation if one sensor fails. The training utilizes a unified curriculum across eight terrain types with specific terrain-aware rewards (foot-sole support-integrity scan and slope-aligned velocity decomposition) and robustness injection (perception degradation, domain randomization). The approach is technically sound, leveraging standard RL techniques (PPO) but applying them to a novel perception interface. The use of linear attention to reduce complexity for onboard deployment is a practical and effective engineering choice.
The evaluation is extensive, covering both simulation (Isaac Lab) and real-world deployment on a DR02 humanoid. Simulation results show high success rates (94.6% average) across diverse terrains, outperforming baselines like height sampling and depth-image CNNs, particularly on thin barriers and sparse footholds. Real-world experiments validate the policy on challenging terrains (70-cm platforms, 100-cm gaps, thin barriers) with 20 trials per setting. The paper provides strong ablation studies, demonstrating the importance of the foot-sole scan and slope-aligned velocity decomposition. The comparison with a blind proprioception-only baseline and an elevation-map baseline is rigorous. The demonstration of zero-shot outdoor generalization and single-modality failure resilience (LiDAR occlusion) adds significant weight to the claims.
The paper provides detailed descriptions of the sensor configuration, network architecture, reward functions, and training hyperparameters. The use of standard tools like Isaac Lab and PPO aids reproducibility. However, specific code is not linked in the provided text (only a video), and the exact implementation of the "foot-sole scan" reward and the specific domain randomization ranges are described but would require code access for full replication. The hardware platform (DR02) is specific, which may limit immediate reproducibility for labs without similar hardware.
The sensing range is limited to ~1.3m ahead and 1.1m behind, which is a local perception horizon. The method relies on a fixed token budget, which may limit resolution for very complex or distant terrain features. The paper acknowledges that discrete obstacle traversal is simulation-only. The reliance on a specific humanoid platform (DR02) and onboard compute (RK3588) means the results are tied to this hardware configuration. The "thin barrier" test in the real world involves a freestanding plate that can be pushed over, which is a slightly different failure mode than a rigid wall, though the authors note this.
This work contributes to the field of legged robotics by demonstrating that point-level fusion of LiDAR and depth cameras can provide robust, low-latency perception for humanoid locomotion. The approach of decoupling forward cost from sensor count is valuable for multi-sensor systems. The unified training strategy for multiple terrains reduces the need for per-terrain fine-tuning, which is a significant practical advantage for deployment. The findings on graceful degradation under sensor failure are important for safety-critical applications. UniPoint introduces a unified point-level sensor fusion framework for humanoid locomotion that achieves robust traversal across diverse terrains with low onboard compute cost. The paper demonstrates that early-fusing LiDAR and depth camera data into a fixed token set, processed by linear attention, outperforms traditional elevation maps and depth-image CNNs, particularly on challenging terrains like thin barriers and sparse footholds, while maintaining graceful degradation under sensor failure.
Contact-rich manipulation, requiring robots to regulate not only motion but also how they yield to external forces, has emerged as the next frontier for Vision-Language-Action (VLA) models. However, existing VLAs output purely kinematic commands, degrading performance on real-world contact-rich tasks. In this paper, we introduce CompVLA, a unified VLA framework that jointly predicts motion and stiffness matrix from RGB and language inputs. Our approach augments the conventional architecture with a dedicated Compliance Expert, which outputs time-varying stiffness and virtual displacement profiles executed via geometric impedance control. We demonstrate that CompVLA achieves the highest average success rate across diverse contact-rich tasks, outperforming both vanilla and compliance-aware VLA baselines, with ablations confirming each component is essential.
Primary: Seoul National University
All Institutions: Seoul National University
CompVLA introduces a Compliance Expert to VLA models to jointly predict motion and stiffness for contact-rich manipulation. The paper presents a novel architectural extension to VLA models that addresses the critical limitation of purely kinematic outputs in physical interaction tasks, offering a rigorous and physically grounded solution for next-generation robotic manipulation.
The paper proposes CompVLA, a Vision-Language-Action (VLA) model that extends standard kinematic output spaces to include compliance parameters (stiffness matrix and virtual displacement). The core architectural contribution is the "Compliance Expert," a dedicated module that predicts time-varying impedance control parameters alongside the action trajectory. This is a significant methodological shift from purely kinematic VLA models, addressing the critical gap in contact-rich manipulation where force regulation is as important as position. The integration of geometric impedance control with learned compliance profiles is a sound and physically grounded approach.
The experiments demonstrate that CompVLA achieves the highest average success rate on diverse contact-rich tasks compared to vanilla and compliance-aware baselines. Ablation studies confirm the necessity of the Compliance Expert. However, the provided text is a skeleton (section headers only), so the depth of the experimental validation (e.g., number of tasks, specific metrics, comparison against state-of-the-art non-VLA impedance methods) cannot be fully verified. The claim of outperforming baselines is strong but relies on the unseen detailed results.
The paper mentions a unified framework and specific components (Compliance Expert, geometric impedance control), which suggests a clear implementation path. However, without the full text details on hyperparameters, dataset specifics, and code availability, reproducibility is moderate. The use of standard impedance control laws aids in this regard.
The primary limitation is the reliance on real-world contact-rich tasks, which can be expensive and time-consuming to benchmark. The model's generalization to unseen contact dynamics or different robot embodiments is not explicitly detailed in the abstract. Additionally, the computational overhead of predicting full stiffness matrices in real-time needs to be addressed.
This work has high potential impact in the robotics community by bridging the gap between high-level semantic understanding (VLA) and low-level physical interaction (impedance control). It enables robots to perform delicate tasks like assembly, insertion, and handling deformable objects more robustly. The approach could be extended to other manipulation domains requiring force feedback. CompVLA introduces a Compliance Expert to VLA models to jointly predict motion and stiffness for contact-rich manipulation. The paper presents a novel architectural extension to VLA models that addresses the critical limitation of purely kinematic outputs in physical interaction tasks, offering a rigorous and physically grounded solution for next-generation robotic manipulation.
Large language model (LLM) inference replicas run across tightly coupled GPUs and serve traffic continuously for weeks. Hardware and software failures are therefore inevitable, and one worker failure can disrupt an entire replica. Recovery requires reinitializing the engine, taking minutes even when weights and compilation artifacts are cached. Production deployments overprovision serving capacity to mask this window. We argue that the dominant cost is loss of ready serving capacity, not request progress, so recovery should preserve initialized engine state rather than reconstruct it. We present fast recovery for Dynamo based on this principle. Snapshots capture an initialized engine once and restore it instead of reinitializing it. Analysis of 18 weeks of failures from the Dynamo cluster shows that most failures are device-preserving: the engine process fails while the GPU and its resident allocations remain intact. Our key insight is that independent engine processes can reuse the same GPU-resident state while keeping mutable execution state private. The GPU Memory Service (GMS) decouples device-memory ownership from engine processes, enabling engines to share and reattach surviving allocations without copying them. GMS preserves model weights and shares them read-only between replacement and Shadow Engines, avoiding weight reloads. A second initialized runtime on the same GPUs reduces recovery to promotion. Across four models on vLLM and SGLang, these mechanisms recover a failed replica in under 7 seconds, 13-29 times faster than a warm restart, using a fixed 4-8 GiB of device memory per GPU independent of model size. Replaying the production trace, we estimate they would reclaim 79% of GPU-hours lost to recovery.
Primary: NVIDIA Research
All Institutions: NVIDIA, NVIDIA Research
[One sentence main contribution]. The paper presents a fast recovery mechanism for LLM inference that decouples GPU memory lifetime from engine processes, enabling sub-7-second recovery from failures by reusing resident state and pre-initialized shadow engines. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. This is a high-impact systems paper that addresses a critical bottleneck in production LLM serving: the slow recovery from failures. The key insight that most failures are device-preserving is well-supported by production data and leads to a novel design that leverages CUDA VMM to share model weights across processes. The combination of GMS, Snapshots, and Shadow Engines provides a tiered recovery path that is both fast and robust. The evaluation is thorough, using real-world models and frameworks, and the results are compelling, showing order-of-magnitude improvements in recovery time. The work is highly relevant to the ML systems community and likely to influence future designs of inference serving stacks.
The paper proposes a novel recovery architecture for LLM inference that decouples the lifetime of GPU-resident state from the engine process. The core innovation is the GPU Memory Service (GMS), which uses CUDA Virtual Memory Management (VMM) to allow multiple engine processes to share read-only model weights without copying them. This is combined with "Shadow Engines," which maintain a pre-initialized runtime state on the same GPUs, allowing for near-instant promotion upon failure. The methodology is grounded in a rigorous analysis of 18 weeks of production failure data, which reveals that most failures are "device-preserving" (the GPU remains healthy, only the process fails). This insight drives the design choice to reuse resident state rather than reconstruct it. The approach is technically sophisticated, leveraging low-level CUDA APIs (cuMemCreate, cuMemSetAccess) and CRIU for process checkpointing.
The evaluation is strong and production-relevant. It tests four large models (Qwen3.8-27B, DeepSeek-V4-Flash, GLM-5.2, DeepSeek-V4-Pro) on vLLM and SGLang. The results show a 13-29x speedup in recovery time (from minutes to under 7 seconds) compared to warm restarts. The paper also quantifies the memory overhead (4-8 GiB per GPU) and the impact on serving latency (TTFT/ITL) during recovery. A counterfactual analysis of the production trace estimates that this system would reclaim 79% of GPU-hours lost to recovery. The experiments are well-controlled, comparing various combinations of Snapshots, GMS, and Shadow Engines.
The implementation is open-sourced, which is a significant plus. The paper provides detailed descriptions of the GMS protocol, the shadow engine promotion logic, and the snapshot capture/restore process. However, the system is tightly integrated with NVIDIA's Dynamo stack and specific CUDA features, which may limit reproducibility on non-NVIDIA hardware or older CUDA versions. The code is available, but the complexity of the system (23k lines of code) makes independent reproduction non-trivial.
The primary limitation is the memory overhead of maintaining a Shadow Engine, which reduces the available KV cache for the primary engine. The paper notes that this overhead is fixed (4-8 GiB) and independent of model size, but it is still a significant cost for memory-constrained deployments. Additionally, the system currently does not preserve request state (KV cache contents), so in-flight requests must be replayed. The paper also acknowledges that the failure analysis is based on a single cluster and may not generalize to all deployment scenarios, particularly those with different hardware failure rates.
This work has significant implications for the reliability and cost-efficiency of large-scale LLM inference deployments. By reducing recovery time from minutes to seconds, it enables operators to reduce overprovisioning, leading to substantial cost savings. The techniques developed here (decoupling memory ownership from process lifetime) could be applied to other stateful services or even training systems. The paper also highlights the importance of understanding production failure modes, providing a valuable dataset and analysis for the community. [One sentence main contribution]. The paper presents a fast recovery mechanism for LLM inference that decouples GPU memory lifetime from engine processes, enabling sub-7-second recovery from failures by reusing resident state and pre-initialized shadow engines. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. This is a high-impact systems paper that addresses a critical bottleneck in production LLM serving: the slow recovery from failures. The key insight that most failures are device-preserving is well-supported by production data and leads to a novel design that leverages CUDA VMM to share model weights across processes. The combination of GMS, Snapshots, and Shadow Engines provides a tiered recovery path that is both fast and robust. The evaluation is thorough, using real-world models and frameworks, and the results are compelling, showing order-of-magnitude improvements in recovery time. The work is highly relevant to the ML systems community and likely to influence future designs of inference serving stacks.