Last 7 Days (October 01 – October 07, 2026)
Simulating microstructure evolution requires quantum-mechanical accuracy and mesoscopic reach in length and time scales, a combination that no current method achieves. Classical phase-field models provide this reach, but their accuracy is limited by phenomenological free energies and mobilities. Here we develop a framework for learning ab initio phase-field models, where the mesoscopic equation is not postulated but derived from a Mori-Zwanzig projection of molecular dynamics onto species-density fields under explicit assumptions. The nonlocal free energy and mobility left unspecified by this equation are parametrized by neural networks and learned from short molecular dynamics trajectories generated with machine-learning interatomic potentials of ab initio accuracy. We demonstrate the framework on an iron-boron melt and on hydrogen-helium mixtures under planetary conditions. For iron-boron, the model shows that the melt at the FeB$_4$ composition is spinodally unstable at ambient pressure but stabilized at 10 GPa, offering a thermodynamic rationale for why FeB$_4$ has been synthesized only under high pressure. For hydrogen-helium, the model predicts the immiscibility boundary and captures droplet nucleation and growth in helium-rain simulations of a column corresponding to 2.2 million atoms, far beyond the scale of atomistic modeling at comparable accuracy. Trained across compositions and conditions, such models could provide a mesoscopic counterpart to ab initio molecular dynamics.
Primary: National University of Singapore
All Institutions: National University of Singapore
The paper presents a rigorous, physics-informed framework for learning ab initio phase-field models by deriving mesoscopic equations from Mori-Zwanzig projections and jointly learning free-energy and mobility functionals from MLIP-generated MD data, demonstrating significant scalability and accuracy in simulating complex microstructure evolution across materials and planetary science applications.
The paper proposes a rigorous framework for deriving mesoscopic phase-field equations from microscopic molecular dynamics (MD) using the Mori-Zwanzig projection formalism. Unlike standard data-driven approaches that fit phenomenological parameters, this method derives the structure of the evolution equation (Generalized Langevin Equation reduced to a Markovian form) from statistical mechanics. The key innovation is the joint learning of the nonlocal free-energy functional and the mobility tensor using neural networks, trained on short MD trajectories generated by machine-learning interatomic potentials (MLIPs). The authors introduce a "ladder" of free-energy functionals, selecting a specific rung (local density plus nonlocal kernel) that balances expressiveness with data efficiency. The derivation is mathematically sound, explicitly stating assumptions (Markovian dynamics, local mobility, Gaussian noise) that justify the reduction from the exact GLE to the practical model. The use of fluctuation-dissipation relations to anchor the mobility and free-energy curvature provides a strong physical constraint on the learning process, preventing the neural networks from learning unphysical dynamics.
The experimental validation is extensive and covers three distinct systems with varying complexity: a Lennard-Jones mixture (benchmark), an Iron-Boron melt (materials science), and a Hydrogen-Helium mixture (planetary science). For the LJ mixture, the model accurately reproduces the phase diagram and spinodal decomposition dynamics, outperforming Landau and Flory-Huggins baselines. The Iron-Boron results provide a thermodynamic rationale for the high-pressure synthesis of FeB4, correctly predicting the stabilization of the melt at 10 GPa. The most impressive result is the Hydrogen-Helium simulation, where the model predicts the immiscibility boundary consistent with prior DFT studies and simulates helium rain in a domain corresponding to 2.2 million atoms over 1 ns—a scale far beyond the reach of direct MD at comparable accuracy. The ability to extrapolate to larger scales and include external potentials (gravity) not present in training data demonstrates the model's generalizability.
The paper provides detailed descriptions of the training data generation, the specific neural network architectures (MLPs for local terms, radial kernels for nonlocal terms), and the loss function components (dynamical, mobility anchor, structure factor anchor, pressure anchor). The use of standard MLIPs (MACE) for data generation enhances reproducibility, as these potentials are increasingly available. However, specific hyperparameters, network sizes, and training protocols are deferred to the Supplemental Material, which is not fully visible in the provided text. The code availability is not explicitly stated in the main text, which is a minor gap for immediate reproducibility, though the methodology is described with sufficient precision for re-implementation.
The framework currently resolves only density fields, neglecting momentum density, which limits its ability to capture hydrodynamic effects like droplet settling via flow rather than just diffusion. The Markovian assumption may fail for systems with long memory times, such as supercooled liquids or glasses. The accuracy of the model is inherently bounded by the accuracy of the underlying MLIPs used to generate training data. Additionally, the computational cost of generating high-quality MD data for complex materials remains a barrier, although the authors argue that the data efficiency of the approach mitigates this.
This work bridges the gap between quantum-mechanical accuracy and mesoscopic scalability, a long-standing challenge in materials science and fluid dynamics. By providing a principled way to learn mesoscopic models from microscopic data, it enables simulations of phenomena like phase separation, nucleation, and coarsening at scales and timescales previously inaccessible. The potential applications are vast, ranging from alloy design and synthesis to understanding planetary interiors. The framework could serve as a foundation for "mesoscopic foundation models" that generalize across compositions and conditions, analogous to CALPHAD databases but with predictive dynamical capabilities. The paper presents a rigorous, physics-informed framework for learning ab initio phase-field models by deriving mesoscopic equations from Mori-Zwanzig projections and jointly learning free-energy and mobility functionals from MLIP-generated MD data, demonstrating significant scalability and accuracy in simulating complex microstructure evolution across materials and planetary science applications.
Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing on learning sciences, we introduce a masked near-transfer post-test: the student is tested on an unseen variant of the tutored problem with the tutor's utterances masked, so the reward can rise only through what the student wrote in its own turns. This discourages cognitive offloading by the student and allows the continuous penalty to be replaced by two binary reward gates (factual correctness of tutor response, no solution handover). A leave-one-out ablation shows that the learning-gain reward on its own does not separate teaching from telling: the gates reduce solution handover while the near-transfer post-test improves out-of-domain transfer. Using these reward designs we develop Eduardo, a multi-turn RL recipe for training LLM tutors, and use it to train 4B, 9B, 14B and 27B models from two distinct LLM architectures. Our post-trained Eduardo-27B model matches Gemini-3.1-Pro on MathTutorBench and Claude Opus 4.8 on TutorMoments at 2.4-6.2x fewer thinking tokens than frontier models, which matters for interactive tutoring. Without being named in the reward, the model more than doubles its use of the push-for-justification teacher move while support fading (e.g., assigning independent work), whose payoff lies beyond a single-problem dialog episode, is trained out. We open-source our training environment, an 8,671-problem near-transfer dataset, and trained models for further development.
Primary: ETH Zurich (Inferred from "eth-lre" in GitHub URL and Swiss AI Initiative funding)
All Institutions: ETH Zurich, Swiss National Supercomputing Centre (CSCS)
The paper introduces a robust RL framework for training LLM tutors by using masked near-transfer post-tests and binary reward gates to prevent solution handover, resulting in models that match frontier performance with higher efficiency. The methodology is sound, the experiments are thorough, and the contributions are significant for the field of AI for Education and multi-turn RL.
The paper addresses a critical flaw in existing RL-based tutoring methods: the "telling" problem, where models learn to simply provide answers to maximize immediate student success metrics. The authors propose a novel reward structure based on three conditions: (1) Near-transfer testing (student solves a variant problem, not the exact one), (2) Masked post-test (tutor's utterances are hidden during the test, forcing reliance on student-generated notes/reasoning), and (3) Binary reward gates (strict penalties for factual errors or solution handover). This design effectively decouples "helpfulness" from "teaching," forcing the model to elicit student reasoning rather than provide it. The use of a frozen LLM student with genuine errors (rather than prompted confusion) is a strong methodological choice that ensures the training signal reflects real diagnostic and adaptive challenges.
The experimental setup is rigorous and comprehensive. The authors train models across multiple sizes (4B to 27B) and architectures (Qwen3 variants), demonstrating the robustness of the recipe. The ablation study is particularly strong, isolating the contribution of each condition (transfer, masking, gates) and showing that the gates specifically reduce handover while the masked transfer test improves out-of-domain generalization. The evaluation uses independent benchmarks (MathTutorBench, TutorMoments) and judges (Gemini) distinct from the training setup, mitigating overfitting concerns. The results show that Eduardo-27B matches or exceeds frontier models (Gemini 3.1 Pro, Claude Opus 4.8) in pedagogical quality while using significantly fewer thinking tokens, a crucial efficiency gain for interactive applications.
High. The authors open-source the training environment, the 8,671-problem near-transfer dataset, and the trained models. Detailed hyperparameters, prompts, and compute costs are provided in the appendices. The use of standard RL algorithms (DPPO/GRPO) and open-source base models further enhances reproducibility.
The primary limitation is the reliance on a single frozen LLM student (Llama-3.1-8B-Instruct) for training, which may lead to overfitting to that specific model's failure modes. The paper acknowledges this and suggests future work with diverse student ensembles. Additionally, the evaluation is limited to mathematics, and the "affective gap" (lack of socio-emotional support) is noted as a consequence of optimizing purely for cognitive gain. The ablation study is limited to one seed and one model size (4B), which slightly weakens the statistical confidence in the ablation results.
This work has significant implications for the development of AI tutors and, more broadly, for any domain where an agent must build user capabilities rather than just provide answers. The "masked near-transfer" reward design is a generalizable principle that could be applied to other educational or collaborative tasks. The efficiency gains (fewer thinking tokens) make high-quality tutoring more feasible in real-time interactive settings. The open-sourcing of the dataset and models will likely accelerate research in AI for Education. The paper introduces a robust RL framework for training LLM tutors by using masked near-transfer post-tests and binary reward gates to prevent solution handover, resulting in models that match frontier performance with higher efficiency. The methodology is sound, the experiments are thorough, and the contributions are significant for the field of AI for Education and multi-turn RL.
Single-particle cryo-electron microscopy (cryo-EM) has become a widely adopted technique for biomolecular structure determination. The conventional cryo-EM computational pipeline first combines many particle images to reconstruct an electrostatic potential (ESP) map and then fits an atomic model to the recovered map. Density reconstruction has high sample complexity, requiring large numbers of particle images and making structure determination high-cost and low-throughput, particularly for heterogeneous samples. Downstream atomic model building, in turn, becomes increasingly difficult as the resolution of the reconstructed map deteriorates. Protein structure prediction models provide strong sequence-derived priors on atomic structure, and experiment-guided approaches can use these priors to recover structures consistent with experimental measurements. Yet, in cryo-EM, such priors are typically integrated only after density reconstruction during atomic model fitting. We introduce Fold'EM, an inference-time framework that combines priors from protein generative models directly with cryo-EM particle images to determine atomic models from a small number of single particle images, bypassing both intermediate density reconstruction and downstream model building against the reconstructed map. Across synthetic and experimental cryo-EM datasets, Fold'EM recovers accurate atomic structures both with known particle orientations and in an ab-initio setting where orientations are inferred jointly with structure. In heterogeneous datasets, Fold'EM further resolves distinct conformational states from mixed particle populations without separately reconstructing a density map and building an atomic model for each state. We believe these results open new avenues for structure determination in the low-sample regime and for characterizing low-population conformational states directly from cryo-EM particles.
Primary: University of Cambridge
All Institutions: University of Cambridge, DeepMind
Fold'EM introduces a novel inference-time framework that directly guides AlphaFold3 using cryo-EM particle images, bypassing intermediate density reconstruction. By optimizing intermediate Pairformer states and jointly inferring structure, pose, and class, the method achieves high-resolution atomic structures from sparse particle sets, significantly reducing the sample complexity of cryo-EM structure determination.
The paper introduces Fold'EM, a framework that bypasses the traditional cryo-EM density reconstruction step by directly guiding AlphaFold3 using individual particle images. The core technical novelty lies in optimizing intermediate Pairformer states ($H$) rather than the terminal embeddings ($Z$) or atomic coordinates. The authors provide a rigorous theoretical justification for this choice, demonstrating that optimizing $H$ induces a sequence-dependent anisotropic geometry in the conditioning space, which acts as a learned preconditioner for the gradient updates. This prevents the model from drifting into unrealistic structural configurations during the inference-time optimization. The method also introduces a joint inference scheme for structure, pose, and class assignment, utilizing a novel bootstrapped residual-subspace estimator (B-RECOVAR) for label-free classification in heterogeneous samples. The mathematical formulation of the particle-level likelihood versus density-level likelihood is well-articulated, highlighting the statistical insufficiency of consensus densities when poses are uncertain.
The experimental evaluation is extensive, covering both synthetic benchmarks and experimental EMPIAR datasets. The paper demonstrates significant improvements over baselines (RELION, cryoDRGN, and previous AlphaFold-guided methods) in terms of FSC resolution and TM-score, particularly in the low-sample regime (e.g., recovering structures from as few as 100-500 particles). The ablation studies effectively isolate the contributions of particle-level guidance versus density guidance and the optimization of intermediate states. The results on heterogeneous data show the ability to resolve distinct conformational states without separate density reconstructions, which is a significant practical advantage. The comparison with unguided AlphaFold3 clearly shows the value of the experimental guidance in correcting mispredicted conformations.
The paper provides detailed algorithmic descriptions and mathematical formulations. However, as an arXiv preprint, the code availability is not explicitly confirmed in the text provided. The reliance on AlphaFold3, a proprietary model, may limit full reproducibility for some users, though the method itself is clearly defined. The use of standard cryo-EM tools (RELION, cryoDRGN) for baselines aids in comparability.
The method relies heavily on the quality of the AlphaFold3 prior; if the unguided prediction is significantly wrong (e.g., 8PWH case), the method may struggle or require more particles than density-based methods. The computational cost of running multiple reverse-diffusion runs and back-propagating through the Pairformer could be high. The current implementation assumes known CTF parameters and centered particles, which may not always hold in raw experimental data without preprocessing.
This work has the potential to significantly reduce the sample size required for cryo-EM structure determination, making it feasible to study low-population states and scarce samples. It bridges the gap between protein structure prediction and experimental cryo-EM, offering a new paradigm for structure determination that could accelerate drug discovery and structural biology research. The insights into optimizing intermediate representations of generative models may also be applicable to other inverse problems in scientific imaging. Fold'EM introduces a novel inference-time framework that directly guides AlphaFold3 using cryo-EM particle images, bypassing intermediate density reconstruction. By optimizing intermediate Pairformer states and jointly inferring structure, pose, and class, the method achieves high-resolution atomic structures from sparse particle sets, significantly reducing the sample complexity of cryo-EM structure determination.
Illusory pattern perception is a well-documented human cognitive tendency to infer meaningful relationships in data that is actually random. Such a tendency, often described as "connecting the dots" where none exist, can result in systematic reasoning errors. This paper investigates whether Large Language Models (LLMs) exhibit such perceptual tendencies, which can lead to systematic errors in downstream applications. To our knowledge, this work presents the first systematic study of illusory pattern perception in LLMs, adapting classic psychological paradigms to three tasks with direct empirical comparison to human behaviors. We find that LLMs frequently exhibit stronger illusory pattern perception than humans. In particular, models tend to over-associate frequent positive attributes with majority groups or large organizations, and show increased tendencies to construct causal narratives from ambiguous events. To uncover the mechanism behind these behaviors, we develop a feature interpretability framework based on Sparse Autoencoders (SAEs) to analyze internal representations. Our results reveal that holistic frequency perception and analytic cognitive orientation are linked to the emergence of illusory perceptions. These findings highlight a previously underexplored cognitive-like illusion that may affect the reliability of LLM reasoning. Code available at https://github.com/NusIoraPrivacy/illusory.
Primary: National University of Singapore
All Institutions: National University of Singapore, University of Macau
The paper identifies and quantifies a novel cognitive failure mode in LLMs—illusory pattern perception—demonstrating that models are more prone to inferring unsupported correlations than humans. By combining rigorous behavioral benchmarks with Sparse Autoencoder-based interpretability, the work provides both a diagnostic framework and a mechanistic understanding of how these spurious inferences arise from holistic frequency processing, offering a pathway for targeted mitigation via feature steering.
The paper proposes a rigorous framework for defining and measuring "Illusory Pattern Perception" (IPP) in LLMs, distinguishing it from standard hallucination or bias. The methodology is strong, utilizing classic psychological paradigms (community correlation, investment correlation, conspiracy belief) adapted for LLM evaluation. A key methodological strength is the use of Sparse Autoencoders (SAEs) to analyze internal representations, linking behavioral outputs to specific latent features (e.g., holistic frequency perception vs. analytic orientation). The development of a "correlation-neutral" prompting variant to isolate the effect of evidence volume from content is a clever experimental control. The internal signal adjustment technique, which steers model activations based on SAE feature differences, provides a mechanistic intervention that validates the interpretability findings.
The experiments are extensive, comparing 9 different LLMs (proprietary and open-source) against 100 human participants per task. The results are statistically robust, with clear significance levels reported. The finding that LLMs exhibit *stronger* illusory pattern perception than humans is a significant empirical contribution. The controlled tests (EqualVol, MajorLow, VaryRate) effectively isolate the causal driver of the behavior, showing that models rely on volume/frequency cues rather than just social priors. The SAE analysis provides a coherent narrative explaining *why* the models behave this way, shifting attention from surface details to analytical features when corrected.
The paper provides a public GitHub repository with code. The prompts for the tasks are detailed in the appendix. However, the reliance on specific pretrained SAEs for open-source models limits the immediate reproducibility of the interpretability section for models without available SAEs. The human study protocol is well-documented, allowing for replication.
The study relies on a relatively small number of human participants (100 per task), which may limit the statistical power for subtle effects, though the LLM effects are large. The SAE analysis is limited to three specific open-source models, so the generalizability of the "holistic frequency perception" mechanism to other architectures is not fully established. The "correlation-neutral" prompt is a specific intervention; it is unclear if other mitigation strategies would yield similar feature shifts.
This paper has high impact on the field of AI reliability and interpretability. It identifies a new failure mode (relational validity) that is distinct from factuality or fairness, which is crucial for high-stakes applications like legal or medical reasoning where models might infer unsupported causal links. The use of SAEs to diagnose and mitigate this behavior offers a template for using interpretability tools to fix specific cognitive-like biases in LLMs. The paper identifies and quantifies a novel cognitive failure mode in LLMs—illusory pattern perception—demonstrating that models are more prone to inferring unsupported correlations than humans. By combining rigorous behavioral benchmarks with Sparse Autoencoder-based interpretability, the work provides both a diagnostic framework and a mechanistic understanding of how these spurious inferences arise from holistic frequency processing, offering a pathway for targeted mitigation via feature steering.
Optimized kernels such as FlashAttention and FlashDecoding are crucial for accelerating today's large models. Most of them are handwritten by experts because existing ML compilers cannot match their efficiency. Producing such kernels requires fusing computations with multiple reductions, which requires both algebraic transformation of the computation graph and operator scheduling of the transformed graph. Unfortunately, searching the two jointly yields a space too large to navigate. We propose Cleave, an ML compiler built on symbolic decoupling: Cleave discovers transformations by performing superoptimization on a graph with symbolic shapes, and then schedules each resulting graph on concrete shapes. Representing shapes as symbols makes equivalence checking cheap and lets a new Split operator, with a symbolic split count, parallelize along a reduction dimension. Cleave's scheduler fuses graphs with multiple reductions through iterative tiling and horizontal fusion. Evaluation on common LLM subgraphs shows that Cleave generates kernels up to 2.8x faster than the best baseline (1.6x on average) and reduces compilation time by 5.9x on average compared to Mirage. For dynamic workloads captured from production serving traces, Cleave compiles each operator once and achieves geometric mean speedups of 1.4x and 1.7x over FlashInfer's handwritten FA2 and FA3 backends. Cleave's code is available at: https://github.com/nyu-systems/cleave
Primary: New York University
All Institutions: New York University, Cornell University
Cleave introduces a decoupled approach to tensor program optimization that significantly improves kernel generation speed and efficiency for LLMs. The paper presents a rigorous methodology and strong empirical results, demonstrating clear improvements over existing baselines and handwritten kernels, making it a valuable contribution to the field of ML systems.
The paper proposes Cleave, an ML compiler that decouples the optimization process into two distinct phases: algebraic search (using superoptimization on symbolic shapes) and operator scheduling (on concrete shapes). The core innovation is the use of symbolic shapes to make equivalence checking cheap and the introduction of a "Split" operator with symbolic split counts to parallelize reduction dimensions. The scheduler employs iterative tiling and horizontal fusion to handle graphs with multiple reductions. This approach is logically sound and addresses a known bottleneck in ML compilers where joint search spaces are intractable.
The evaluation demonstrates significant performance gains, claiming up to 2.8x speedup over the best baseline and 1.6x on average. It also reports a 5.9x reduction in compilation time compared to Mirage. The paper includes evaluations on dynamic workloads from production serving traces, showing speedups over handwritten FlashInfer backends (FA2 and FA3). The results are strong and directly relevant to current LLM inference and training bottlenecks.
The code is available on GitHub, which is a positive factor for reproducibility. The paper provides sufficient detail on the symbolic decoupling approach and the specific operators used. However, the complexity of the compiler pipeline may make full reproduction challenging without deep expertise in ML systems.
The paper focuses primarily on LLM subgraphs and attention mechanisms. It is unclear how well the approach generalizes to other types of neural network architectures or non-attention-heavy workloads. Additionally, the compilation time, while reduced, is still a factor for dynamic workloads, though the paper claims to compile each operator once.
This work has significant potential impact on the ML systems community by providing a more efficient way to generate optimized kernels for large models. It could reduce the reliance on handwritten kernels and improve the scalability of ML compilers for dynamic workloads. The approach could be extended to other domains where symbolic optimization and scheduling are applicable. Cleave introduces a decoupled approach to tensor program optimization that significantly improves kernel generation speed and efficiency for LLMs. The paper presents a rigorous methodology and strong empirical results, demonstrating clear improvements over existing baselines and handwritten kernels, making it a valuable contribution to the field of ML systems.
AI systems can now write and optimize production GPU kernels, but validating them remains an important challenge. Evaluating the kernel on a few random inputs and checking that its outputs match a trusted reference kernel within numeric tolerances is not sufficient: races can cause nondeterministic behavior that fails to manifest in tests, and numeric tolerances can hide bugs and cause false positives even after extensive calibration. To address this challenge, we present RESOLVE, which combines testing and formal verification to build a comprehensive kernel validation pipeline. It operates in three steps: First, it tests for nondeterminism using binary instrumentation that perturbs execution timing to expose races. Second, an agent rewrites the candidate and reference kernels to obtain "reduced-concurrency" versions that are simpler to analyze but still produce bitwise-identical outputs in all tests. Third, the reduced kernels are formally analyzed in the F*/Pulse framework and prove that they perform the same computation on real numbers. This sidesteps the need for numeric tolerances. We show that RESOLVE can validate a broad selection of kernels using KernelBench, and prove equivalence across fused GEMMs in three state-of-the-art frameworks and languages: CUTLASS, Triton, and Gluon. It also analyzes mega-kernels, notoriously difficult to validate, and finds four previously unreported issues, including two clear bugs. We show that agents can use RESOLVE to repair the issues, with minimal performance impact, highlighting that agents can optimize aggressively when they can rigorously check their results.
Primary: Microsoft Research
All Institutions: University of California, Riverside, Math, Inc., Microsoft Research
The paper presents RESOLVE, a novel framework that combines binary testing, agent-based kernel reduction, and formal verification to rigorously validate GPU kernels, addressing the limitations of tolerance-based testing and the semantic gaps in direct formal verification. By successfully identifying previously unreported bugs in state-of-the-art frameworks and demonstrating that agents can repair them, the work establishes a promising path toward trustworthy, AI-optimized high-performance computing.
The paper proposes RESOLVE, a three-stage pipeline for validating GPU kernels: (1) binary instrumentation to test for nondeterminism/races, (2) agent-based rewriting to create "reduced-concurrency" versions that are bitwise equivalent to the originals, and (3) formal verification of the reduced versions using F*/Pulse. The core novelty lies in the decoupling of concurrency correctness (handled via testing) from functional correctness (handled via proof), and the use of LLM agents to bridge the gap between complex production kernels and the limited semantic coverage of formal verifiers. The approach is clever but relies heavily on the agent's ability to produce correct reductions, which is a non-trivial assumption. The use of bitwise equivalence as an oracle for the reduction step is a strong design choice that avoids numerical tolerance issues during the intermediate phase.
The evaluation covers KernelBench, fused GEMMs in CUTLASS/Triton/Gluon, and mega-kernels. The finding of four previously unreported issues (including two bugs) in state-of-the-art frameworks is a significant empirical result, demonstrating the tool's practical utility. The comparison against existing tolerance-based tests shows that RESOLVE catches errors that standard testing misses. However, the evaluation is somewhat limited in scale (only three mega-kernels) and lacks a detailed analysis of the cost (time/compute) of the agent-based reduction and proof steps. The claim that agents can repair the issues with minimal performance impact is supported but would benefit from more extensive benchmarking.
The paper describes the pipeline clearly, but the reliance on "agents" to perform rewrites and proofs introduces variability. Without a fixed prompt strategy or a deterministic agent framework, exact reproduction of the results may be difficult. The use of NVBit and F*/Pulse is standard, but the specific agent configurations are not detailed enough for full reproduction. The code availability is not explicitly stated in the provided text, which is a gap for a systems paper.
The primary limitation is the dependence on LLM agents for the reduction and proof steps. If the agent fails to produce a valid reduction or proof, the pipeline stalls. The paper does not deeply analyze the failure modes of the agent or the success rate of the reduction step across a larger corpus. Additionally, the formal verification step is limited to the subset of CUDA/Kuiper supported by F*/Pulse, meaning kernels with exotic hardware features may still require manual intervention or may not be verifiable. The performance overhead of the validation pipeline itself is not thoroughly quantified.
This work has high potential impact on the reliability of AI-generated code, particularly in high-stakes domains like autonomous driving or financial modeling where GPU kernel correctness is critical. It provides a framework for integrating formal methods into the agentic coding loop, which is a growing area of interest. The discovery of bugs in production frameworks like CUTLASS and Triton highlights the immediate value of such tools. It may influence the development of future kernel compilers and verification tools to be more amenable to automated reduction and proof. The paper presents RESOLVE, a novel framework that combines binary testing, agent-based kernel reduction, and formal verification to rigorously validate GPU kernels, addressing the limitations of tolerance-based testing and the semantic gaps in direct formal verification. By successfully identifying previously unreported bugs in state-of-the-art frameworks and demonstrating that agents can repair them, the work establishes a promising path toward trustworthy, AI-optimized high-performance computing.
Graph foundation models aim to transfer across graphs, feature spaces, relational schemas, and prediction tasks, yet existing approaches typically generalize only within particular graph modalities or tasks. We propose Wander, a graph foundation model designed to operate across these settings within a single pretrained checkpoint. Following the prior-predictive perspective, we formulate graph learning as completion of a partially observed graph. We realize this task-general view through a common interface based on random walks, allowing the same model to operate across homogeneous and multi-relational graphs with varying features, labels, and relational schemas. Wander can increase its structural context at inference time without changing its learned parameters and, under suitable assumptions, universally approximates the corresponding Bayes-optimal predictor on bounded connected graphs. Empirically, a single pretrained checkpoint achieves state-of-the-art or highly competitive results across node classification, homogeneous link prediction, and knowledge-graph link prediction. Moreover, joint pretraining across graph modalities and tasks preserves performance in specialized settings while enabling positive transfer and the composition of separately learned capabilities at inference time.
Primary: AITHYRA (Austrian Academy of Sciences)
All Institutions: AITHYRA, Austrian Academy of Sciences, Boehringer Ingelheim Stiftung
Wander introduces a unified graph foundation model using random walks as a shared interface, achieving state-of-the-art performance across node classification, homogeneous link prediction, and knowledge graph link prediction while demonstrating positive transfer and compositional generalization. The paper provides a rigorous theoretical framework proving universality and permutation equivariance, and empirically validates the model's ability to adapt its structural context at inference time, offering a promising path toward generalist graph learning.
The paper proposes "Wander," a graph foundation model that unifies node classification, homogeneous link prediction, and knowledge graph link prediction under a single probabilistic framework: the completion of a partially observed graph. The core architectural innovation is the use of random walks as a shared computational interface, replacing standard message passing. This allows the model to dynamically adjust its structural context at inference time without retraining. The methodology is rigorous, providing a theoretical foundation that proves Wander is a universal approximator of the Bayes-optimal predictor for bounded connected graphs and is permutation equivariant in distribution. The design separates feature, label, and structural channels, processing them via intra-node attention, in-context learning updates, and stochastic structural updates via sampled walks. This is a sophisticated and well-motivated approach to the fragmentation problem in graph foundation models.
The experimental evaluation is extensive and addresses four key questions: generalist performance, transfer through joint training, compositional generalization, and inference-time adaptation. The model achieves state-of-the-art or highly competitive results across 54 knowledge graphs, 10 homogeneous link prediction datasets, and 26 node classification datasets. Notably, the paper demonstrates positive transfer from joint pretraining (improving homogeneous link prediction by ~3% over single-task training) and compositional generalization (combining node features and edge types, which were never seen together during training). The inference-time adaptation experiments on the synthetic Grids task are particularly compelling, showing that increasing walk length significantly improves long-range reasoning accuracy from 56.6% to 96.6%.
The paper provides detailed implementation details in the appendix, including initialization strategies, specific attention mechanisms (RoPE, RMSNorm), random walk sampling protocols, and loss functions. The pretraining data generation process is described, combining synthetic graph models (SBM, ER, Watts-Strogatz) with real-world KGs. While the code is not explicitly linked in the provided text, the level of detail suggests high reproducibility. The use of standard benchmarks and clear evaluation protocols (MRR, Hits@10, Recall@20, Accuracy) further supports reproducibility.
The pretraining prior is limited, relying heavily on synthetic data for attributed graphs and only three real-world knowledge graphs for multi-relational pretraining. Inference cost can be high due to the stochastic nature of random walk sampling, particularly when large budgets are required for long-range reasoning. The model does not yet support node regression or graph classification tasks. The theoretical guarantees rely on capacity assumptions that may not hold for the finite-sized models used in practice.
This paper makes a significant contribution to the field of graph machine learning by demonstrating that a single model can effectively handle diverse graph modalities and tasks. The random walk interface offers a flexible alternative to message passing, with potential applications in dynamic graphs and scenarios where computational resources can be allocated adaptively at inference time. The positive transfer and compositional generalization results suggest that joint pretraining across tasks is a viable strategy for building more robust graph foundation models. Wander introduces a unified graph foundation model using random walks as a shared interface, achieving state-of-the-art performance across node classification, homogeneous link prediction, and knowledge graph link prediction while demonstrating positive transfer and compositional generalization. The paper provides a rigorous theoretical framework proving universality and permutation equivariance, and empirically validates the model's ability to adapt its structural context at inference time, offering a promising path toward generalist graph learning.
Mixture-of-Experts (MoE) transformers scale capacity by activating only a few experts per token, but this sparsity creates a hidden reliability problem: when routing is imperfect, load-balanced models may send tokens to experts that are insufficiently trained for the assigned inputs. We propose Distributionally Robust MoE Training (DRMoET), a drop-in objective that treats layer-wise experts as endogenous robustness groups and optimizes high-loss routing outcomes rather than merely equalizing traffic. DRMoET updates a per-layer expert distribution by an entropy-regularized softmax rule on EMA-smoothed, activation-weighted expert losses, strengthening plausible non-top routing paths while preserving standard MoE computation. Under the FLAME-MoE recipe at 746M-total and 10.3B-total scales, DRMoET improves downstream averages over both standard FLAME-MoE and auxiliary-loss-free balancing. At 10.3B total parameters and 67B training tokens, DRMoET improves the seven-task average from 0.6625 to 0.6767, while the auxiliary-loss-free baseline achieves 0.6431. Mechanistic analyses show lower expert-loss variance with nearly unchanged mean loss, 4.3% lower excess loss under forced mid-$k$ misrouting, and improved domain-expert specialization. These results position routing robustness-not only utilization balance-as a practical objective for reliable sparse MoE scaling. Project page and code are available at: https://drmoet.github.io/.
Primary: New York University
All Institutions: New York University, NYU Shanghai
The paper introduces a distributionally robust training objective for MoE models that improves routing reliability by strengthening non-top expert paths. It provides a rigorous theoretical framework and empirical evidence that optimizing for expert competence, rather than just load balance, leads to more robust and specialized MoE models at scale.
The paper proposes Distributionally Robust Mixture-of-Experts Training (DRMoET), a drop-in training objective that treats layer-wise experts as endogenous robustness groups. Unlike standard load-balancing losses that optimize for traffic distribution, DRMoET optimizes for expert competence under imperfect routing. It utilizes an entropy-regularized softmax update on EMA-smoothed, activation-weighted expert losses to strengthen plausible non-top routing paths. The method is theoretically grounded, providing convergence guarantees for the entropy-regularized robust objective, and is computationally efficient, adding negligible overhead to the standard MoE forward/backward pass. The distinction between "allocation" (load balancing) and "competence" (routing robustness) is a well-motivated and clear conceptual contribution.
The authors conduct rigorous experiments at two scales (746M and 10.3B total parameters) using the FLAME-MoE recipe. Results show consistent improvements in downstream task averages over both standard FLAME-MoE and auxiliary-loss-free baselines. Mechanistic analyses are strong, including expert-loss variance reduction, forced misrouting probes (showing 4.3% lower excess loss), and improved domain-expert specialization metrics. The ablation studies effectively isolate the contribution of activation-weighted credit and EMA decay. The inclusion of throughput analysis confirms the method's practical viability.
The paper provides detailed hyperparameters, data sources (DCLM), and training recipes. Code and project page are available. The method is described as a "drop-in" objective, suggesting ease of integration into existing MoE training pipelines. The specific implementation details of the EMA and dual variable updates are clearly specified in the algorithm box and appendix.
The evaluation is limited to two model scales and specific expert configurations. The paper acknowledges that it does not test highly imbalanced pretraining mixtures or long-tail evaluation suites, which would be a more direct stress test for robustness. The gains, while consistent, are moderate (e.g., 1.42 percentage points at 10.3B scale), which may be considered small in the context of large-scale LLM training where other factors often dominate.
As MoE architectures become standard for scaling LLMs, addressing the reliability of sparse routing is increasingly important. This work provides a practical tool to improve the robustness of MoE models without architectural changes, potentially leading to more reliable inference under distribution shift or imperfect gating. It shifts the focus from mere utilization balance to expert quality, a perspective likely to influence future MoE training strategies. The paper introduces a distributionally robust training objective for MoE models that improves routing reliability by strengthening non-top expert paths. It provides a rigorous theoretical framework and empirical evidence that optimizing for expert competence, rather than just load balance, leads to more robust and specialized MoE models at scale.
In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%, while "Alright," raises Qwen3-14B's from 72% to 87%. Second, RL makes these cues more likely, while fixing them recovers much of its performance gain over the base model. Third, we trace the reasoning effects of token cues to the training data. We perform causal data interventions to turn an arbitrary word, such as "chicken", into an effective reasoning cue, or remove an existing cue's effect. A similar edit makes the prompt instruction "Think duck duck goose" as effective as "Think step by step" at eliciting reasoning. We also find that the hidden state representations induced by different cues correlate with different document types from the training set. Finally, we extend our study of token cues with a case study in language model safety, finding that different cues elicit distinct refusal and compliance behaviors that correspond to different types of training data.
Primary: University of California, Berkeley
All Institutions: University of California, Berkeley, Massachusetts Institute of Technology (MIT)
The paper demonstrates that base models' reasoning capabilities are heavily influenced by specific starting token cues learned from training data, and that these cues can be manipulated to significantly enhance performance. This is a significant contribution to mechanistic interpretability and practical LLM optimization, offering a cost-effective alternative to RL for improving reasoning tasks.
The paper employs a rigorous causal intervention framework to dissect the role of "token cues" in Large Language Model (LLM) reasoning. The methodology is three-pronged: (1) Empirical demonstration that fixing specific starting tokens (e.g., ".\n\nOkay") significantly boosts base model performance on math and coding tasks, rivaling RL-trained models; (2) Causal data interventions (likely using activation patching or data ablation techniques) to prove that these cues are learned from specific document types in the training data, rather than being inherent to the token embeddings; (3) A safety case study showing that cues also modulate refusal behaviors. The approach is sophisticated, moving beyond correlation to causation by editing the model's internal representations or training data associations to isolate the effect of the cue.
The experiments are extensive and compelling. The authors test across multiple model families (OLMo-3, Qwen3) and tasks (MATH-500, coding). The jump in accuracy (e.g., 42% to 78% for OLMo-3-7B) is substantial and practically significant. The causal interventions are well-designed to rule out confounding factors. The safety analysis adds depth, showing that the same mechanism that aids reasoning also influences safety alignment, which is a critical insight for the field.
The paper appears to be from a high-quality group (Berkeley/MIT) with clear methodological descriptions. While specific code links are not provided in the snippet, the detailed description of the causal interventions and the use of open-source models (OLMo, Qwen) suggests high reproducibility. The abstract-only score of 60 suggests the full text provides sufficient detail for replication.
The study focuses on base models and the transition to RL. It may not fully account for how these cues interact with more complex agentic workflows or multi-turn conversations. The "chicken" example, while illustrative, is a synthetic intervention; the real-world generalizability of arbitrary word cues is limited, though the core finding about training data associations is robust.
This paper has high impact because it challenges the prevailing narrative that RL is the sole driver of reasoning capabilities in LLMs. It suggests that data curation and prompt engineering (specifically starting tokens) are underutilized levers for improving base model performance. This has immediate implications for model developers, who can optimize pre-training data and inference prompts without expensive RL fine-tuning. It also has safety implications, as it reveals how minor prompt changes can shift model behavior between compliance and refusal. The paper demonstrates that base models' reasoning capabilities are heavily influenced by specific starting token cues learned from training data, and that these cues can be manipulated to significantly enhance performance. This is a significant contribution to mechanistic interpretability and practical LLM optimization, offering a cost-effective alternative to RL for improving reasoning tasks.
Multimodal models are increasingly shifting toward unified architectures that understand and generate text, images, and other modalities within a shared conversational context. This design enables fluid interaction across modalities, but it also changes the privacy threat model: Information revealed in one part of a conversation may remain accessible when the model later generates content in another modality. This risk is particularly concerning in settings where users rely on locally deployed models for privacy, assuming that sensitive interactions remain confined to their device. We introduce Privacy-Leaking Watermarks (PLWs): invisible, trigger-dependent watermarks that a malicious model provider can condition on prior chat history. With this adversarial intervention, the usual separation breaks: a sensitive keyword or semantic cue mentioned earlier in the conversation can cause a later, unrelated image to carry a hidden yet detectable watermark. PLWs pose a novel threat to users of unified multimodal models: A poisoned model can retain utility while covertly turning image generation into a channel for privacy leakage, even when deployed locally. Across 13 sensitive-attribute triggers and two model families, PLWs reach up to 100.0% TPR at 1% FPR. For example, across all tested conversational separations, OmniGen2 detects every prior disclosure of depression while falsely flagging only 1% of images generated without such a disclosure.
Primary: TU Darmstadt
All Institutions: TU Darmstadt, Hessian.AI Service Center, Konrad Zuse School of Excellence in Learning and Intelligent Systems
The paper identifies and demonstrates a novel privacy leakage vector in unified multimodal models by introducing trigger-dependent watermarks that link conversational history to generated image content. By showing that sensitive information can be covertly exfiltrated through image generation even in local deployments, the work provides a critical security assessment that challenges the assumed privacy guarantees of next-generation AI architectures.
The paper introduces Privacy-Leaking Watermarks (PLWs), a novel adversarial attack on unified multimodal models. The core innovation lies in exploiting the shared latent space of unified architectures to embed trigger-dependent watermarks in generated images based on prior conversational context. The methodology involves a two-stage training process: first, training a watermark encoder/extractor pair, and second, fine-tuning the multimodal model (using LoRA) to condition the watermark embedding on specific semantic triggers found in the chat history. This approach is technically sound and cleverly leverages the specific architectural properties of unified models (where text and image generation are not strictly decoupled) to create a covert side-channel for privacy leakage.
The experiments are rigorous, testing the attack across two model families (including OmniGen2) and 13 sensitive attribute triggers. The results are striking, achieving up to 100% True Positive Rate at a 1% False Positive Rate, demonstrating that the watermarks are both highly detectable and effectively hidden from casual observation. The evaluation correctly isolates the threat model to locally deployed models where users assume privacy, providing a clear and compelling demonstration of the vulnerability.
The paper provides a strong reproducibility statement, detailing the training configurations, LoRA target modules, and specific seeds used. While the authors explicitly state they do not release poisoned checkpoints (a responsible security practice), the detailed appendices and methodological transparency allow for independent verification of the attack's feasibility.
The primary limitation is the requirement for white-box access to modify and redistribute model weights, which restricts the threat to malicious model providers or compromised supply chains rather than arbitrary users. Additionally, the attack is specific to unified multimodal architectures; it may not apply to traditional pipeline-based multimodal systems where text and image generation are separate modules.
This paper has significant implications for the deployment of unified multimodal models, particularly in consumer-facing applications where local deployment is marketed as a privacy feature. It highlights a critical gap in the security model of these architectures and will likely drive research into secure design patterns, watermarking defenses, and the separation of modalities in unified models. The paper identifies and demonstrates a novel privacy leakage vector in unified multimodal models by introducing trigger-dependent watermarks that link conversational history to generated image content. By showing that sensitive information can be covertly exfiltrated through image generation even in local deployments, the work provides a critical security assessment that challenges the assumed privacy guarantees of next-generation AI architectures.
Proof auto-formalization translates natural-language (NL) theorems and proofs into a formal language (FL) such as Lean, enabling mechanical verification. Despite rapid progress, research-level proofs often depend on concepts missing from leading proof assistant libraries (e.g., Lean's Mathlib), and successful compilation does not guarantee that a translation preserves the theorem's meaning or the proof's reasoning. Furthermore, aligned NL-FL training data are scarce, and leading agents often rely on costly frontier models and manually engineered harnesses. To address the above issues, we present AIProver, an agentic framework for autonomous proof auto-formalization and proof synthesis (AFPS) that jointly post-trains a 119B open-weight language model and evolves its agentic, tool-calling harness with HarnessEvolve. Verifiers assess type correctness, proof completeness, and semantic correctness, returning rewards and diagnostic certificates that drive model fine-tuning and alternating reinforcement learning via symbolic feedback and HarnessEvolve, a certificate-driven evolutionary search over the whole harness control flow that re-tailors the harness to the updated model. For research-level training and evaluation, we introduce LoCoBench, 58.9k instances from Mathlib, CSLib, Mizar Math Library, and a bounded-arithmetic textbook, with a 771-instance validation split whose theorem-proof pairs have no public Lean formalization. Against 39 frameworks spanning AFPS agents, frontier LLMs, and coding agents, AIProver lifts pass@4 semantic correctness over its Leanstral-1.5 base from 15.7% to 36.7% and outperforms every other open-weight system and Aristotle. As a Claude Code and Codex skill, it lifts their semantic correctness from 41.9% and 34.1% to 79.8% and 62.4%, respectively. Further, it is also 24% cheaper than Numina-Lean-Agent, pushing the accuracy-cost frontier of research-level AFPS.
Primary: Georgia Institute of Technology
All Institutions: Georgia Institute of Technology, Simon Fraser University, University of Pennsylvania, University of Texas at Austin, Rutgers University, Foothill College
The paper presents a novel co-evolutionary framework for proof auto-formalization that jointly optimizes a large open-weight LLM and its agentic harness, achieving state-of-the-art results on a new research-level benchmark while significantly improving cost-efficiency compared to frontier API-based agents.
The paper proposes a joint optimization framework for both the language model and its agentic harness. The Semantic Alignment Model (SAM) introduces a contrastive loss to align natural language and formal language representations, which is a logical extension of existing alignment techniques but applied here to a large MoE model. The core novelty lies in HarnessEvolve, an evolutionary search algorithm that treats the agent's control flow (harness) as a mutable program, using diagnostic certificates from verifiers to guide mutations. This is a significant departure from static scaffolding or simple prompt optimization. The Agentic RLSF component adapts Group-Relative Policy Optimization for multi-turn agent interactions, using a fine-grained reward ladder that distinguishes between type errors, incompleteness, and semantic drift. This multi-objective reward design is crucial for preventing the "silent correction" of flawed proofs, a known failure mode in prior work.
The introduction of LoCoBench is a major contribution, providing 58.9k instances and a 771-instance validation set that is out-of-distribution (Mizar-to-Lean translation). The evaluation against 39 baselines is comprehensive. The results show a substantial improvement over the base model (15.7% to 36.7% pass@4 semantic correctness) and demonstrate that the framework can enhance frontier coding agents (Claude Code, Codex) when used as a specialized skill. The cost-efficiency analysis is particularly strong, showing that the open-weight system achieves competitive accuracy at a fraction of the cost of frontier API calls.
The paper provides extensive details on the training setup, including hardware, hyperparameters for LoRA and GRPO, and the specific configuration of the evolutionary search. The release of the benchmark and the open-weight model base (Leanstral-1.5) enhances reproducibility. However, the reliance on a frontier coding agent (Claude Opus 5) for the mutation step in HarnessEvolve introduces a dependency on a proprietary service, which may limit full reproducibility for all researchers.
The primary limitation is the heavy reliance on a frontier LLM for the harness evolution process, which contradicts the goal of full autonomy and open-source accessibility. The semantic correctness check relies on an extended BEq+ prover, which may have its own coverage limitations. Additionally, the "proof faithfulness" metric is not fully automated, relying on LLM judges which are noted to be less reliable.
This work has significant implications for the accessibility of formal verification. By demonstrating that open-weight models can be post-trained to perform research-level auto-formalization, it lowers the barrier to entry for formal methods. The co-evolution of model and harness offers a generalizable paradigm for improving agentic systems in other domains where tool-use and control flow are critical. The paper presents a novel co-evolutionary framework for proof auto-formalization that jointly optimizes a large open-weight LLM and its agentic harness, achieving state-of-the-art results on a new research-level benchmark while significantly improving cost-efficiency compared to frontier API-based agents.
AI agents can now carry out data-driven scientific analyses end to end, and benchmarks assess them by giving an agent a question and a dataset and scoring its final answer against a fixed key. These benchmarks assume that a correct answer was derived from the supplied data, a property we call evidence grounding. However, an agent can also reach the key from prior knowledge or by ruling out the other options, and a score based on a single run cannot tell these cases apart. We show how to test this assumption and find that it often fails. For each question, we build versions of its data files in which the evidence for the answer is left intact, withdrawn or reversed, check each edit with a pre-registered reference statistic, and run the same agent on every version. We then measure evidence-grounded accuracy, which credits a correct answer only if the agent also responds when the evidence is withdrawn and follows it when it is reversed. We evaluate three agent scaffolds and five models on 18 single-cell questions from BAISBench and four synthetic problems from GeneBench-Pro. On the single-cell questions, the Claude agents are 95% accurate and answer 83% correctly without any data, but their evidence-grounded accuracy is only 41%. Hiding gene names raises the share of runs that follow reversed evidence from 58% to 93% on ten gene tasks, suggesting that prior knowledge competes with the supplied data. The benchmark score and LLM judges can also reward answers that ignore the changed evidence. Measuring scientific intelligence rather than recall therefore requires checking whether answers follow the evidence and whether scores reward them for it.
Primary: Carnegie Mellon University
All Institutions: Carnegie Mellon University
The paper demonstrates that high accuracy on scientific benchmarks does not imply evidence grounding, introducing a verified counterfactual protocol to measure and address this gap. By showing that agents often rely on prior knowledge rather than supplied data, and that standard evaluators fail to penalize this, the work provides a crucial methodological correction for the evaluation of AI scientists.
The paper introduces a rigorous framework for evaluating "evidence grounding" in AI scientific agents. The core methodological contribution is the construction of verified counterfactuals (null, withdrawal, flip) for benchmark items, where the effect of the intervention on the ground truth is pre-registered and verified via reference statistics. This allows for the calculation of Evidence-Grounded Accuracy (EGA), which distinguishes between answers derived from data versus those recalled from prior knowledge or reached via elimination. The formalization of "noticing without acting" and the separation of agent grounding from evaluator validity are strong theoretical contributions. The protocol for cleaning metadata to prevent leakage (canonical writer) is a crucial practical detail that enhances the validity of the counterfactuals.
The experiments are well-designed but limited in scale. The authors evaluate 18 single-cell questions from BAISBench and 4 synthetic problems from GeneBench-Pro. While small, this is sufficient to demonstrate the phenomenon of high accuracy coexisting with low evidence grounding. The finding that Claude agents achieve 95% accuracy but only 41% EGA is a significant empirical result. The gene-name anonymization experiment effectively isolates prior knowledge as a confounding factor. The evaluation of evaluators (benchmark scores and LLM judges) showing they reward ungrounded answers is a critical finding for the field.
The paper provides high reproducibility. It details the canonical writer process, the specific interventions applied to the data, the agent scaffolds used (Claude Code, Codex CLI, Biomni), and the evaluation metrics. The pre-registration of reference statistics and the verification of every edit make the counterfactuals robust. The code and data protocols are described in sufficient detail for replication, although the specific datasets (BAISBench, GeneBench-Pro) are external dependencies.
The primary limitation is the small number of tasks (22 total). The findings may not generalize to all types of scientific tasks or domains beyond single-cell biology and synthetic genetics. The manual construction of counterfactuals is labor-intensive, limiting scalability. The study focuses on a specific set of models (Claude, GPT-5.6) and scaffolds, so results may vary for other architectures. The "flip" intervention assumes a binary or clear alternative, which may not apply to all scientific questions.
This paper has high impact on the evaluation of AI agents in scientific discovery. It challenges the validity of current benchmarks that rely solely on final answer accuracy. The proposed metrics (EGA, counterfactual validity) provide a new standard for assessing whether AI agents truly use the data provided. This work will likely influence the design of future benchmarks for AI scientists and the interpretation of existing leaderboard scores. It highlights a critical gap in current AI evaluation practices that affects the trustworthiness of AI-generated scientific insights. The paper demonstrates that high accuracy on scientific benchmarks does not imply evidence grounding, introducing a verified counterfactual protocol to measure and address this gap. By showing that agents often rely on prior knowledge rather than supplied data, and that standard evaluators fail to penalize this, the work provides a crucial methodological correction for the evaluation of AI scientists.
A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-context adaptation agents store reflections, memories, or skills as text, so reuse depends on retrieving the right experience and on a frozen policy executing it. We study Online Agentic Test-Time Training (OaTTT), which trains the LLM's weights on its own execution trajectories during deployment. The agent executes each task once, in one pass over the stream, and the executed trajectory with its verification result is the only learning signal for weight updates that persist across tasks. Directly imitating or reinforcing the generated tokens of this single attempt destabilizes the policy. We introduce ASCENT (Agentic Self-distillation for Cross-task EvolutioN at Test-time), which instead self-distills verified experience. A stable version of the LLM, its frozen initial copy, receives the verified trajectory as privileged information and predicts next-token distributions along it with this hindsight. Distilling them into persistent LoRA fast weights updates the agent for later tasks, without an external reference solution or stronger teacher. By further removing invalid-action turns, ASCENT distills enhanced privileged experience for more efficient execution. We characterize its population target and the limits of sparse outcome selection. Across ALFWorld, WebShop, and AppWorld at varied model scales, ASCENT improves task success and interaction efficiency as experience accumulates, outperforms online adaptation methods, and transfers to held-out scenes, showing that an agent can consolidate verified experience into its weights without a separate training phase or memory retrieval. Project page: https://artificer-ai-lab.github.io/ASCENT
Primary: University of New South Wales (UNSW Sydney)
All Institutions: University of New South Wales (UNSW Sydney)
ASCENT enables stable online test-time training of LLM agents by self-distilling verified experience through a frozen teacher model, significantly improving task success and efficiency in long-horizon environments without requiring external references or multiple attempts.
The paper proposes ASCENT, a method for Online Agentic Test-Time Training (OaTTT). The core innovation is using the frozen initial model as a teacher that conditions on the "privileged information" of a verified successful trajectory (hindsight) to generate soft targets for the current student model (via LoRA). This avoids the instability of direct imitation or reinforcement learning on single-attempt trajectories, which the authors demonstrate collapses the policy. The method effectively combines rejection sampling (only updating on verified successes) with on-policy distillation. The theoretical framing of the population target and the analysis of why direct imitation fails (sharpening around generated tokens vs. full-vocabulary matching) are strong contributions.
The experiments are extensive, covering ALFWorld, WebShop, and AppWorld with two model scales (Qwen3.5-4B and 9B). The paper provides strong evidence for the method's efficacy, showing significant improvements over the base model and various in-context adaptation baselines (MemP, ACE, etc.). The ablation studies on privileged information content and distillation divergence are thorough. The demonstration that direct imitation baselines fail catastrophically is a valuable negative result for the community.
The paper provides detailed descriptions of the protocol, loss functions, and hyperparameters. The use of standard benchmarks and open-weight models (Qwen) aids reproducibility. However, the specific implementation of the "validity filter" and the exact serialization of privileged information $z_i$ could benefit from more code-level detail, though the project page is provided.
The method requires access to the model's weights (open-source models only) and the ability to run a frozen copy of the model as a teacher, which doubles inference cost during the update phase. The reliance on sparse, episode-level verification limits the granularity of learning signals. The paper acknowledges that the teacher's privileged context may allow it to reconstruct student tokens, potentially limiting the independence of the hindsight signal.
This work is significant for the deployment of LLM agents in dynamic environments where continuous improvement is desired without offline retraining. It bridges the gap between in-context learning (memory-based) and parametric adaptation, offering a stable mechanism for weight updates. The findings on the instability of single-attempt RL/imitation are likely to influence future agent training designs. ASCENT enables stable online test-time training of LLM agents by self-distilling verified experience through a frozen teacher model, significantly improving task success and efficiency in long-horizon environments without requiring external references or multiple attempts.
Diffusion and flow models provide expressive policy classes for online reinforcement learning (RL), enabling multimodal behaviors and improved performance. However, training these policies remains challenging: the critic specifies the desired policy as an unnormalized Boltzmann density but does not provide direct samples from it. Many existing methods rely on importance sampling to construct training signals, which can suffer from high variance, increasing computational cost and destabilizing training. We propose Score-Calibrated Flow (SCF), a simple and efficient algorithm for training generative models to sample from unnormalized densities without importance sampling or backpropagation through the sampling trajectory. We learn the desired flow by enforcing self-consistency, bypassing target posterior mean estimation. By jointly exploiting the prescribed target score and the structure of flow matching, we establish these self-consistency requirements as score-calibrated optimality conditions, first for the terminal density and then for the trainable velocity field. We prove that their unique solutions are, respectively, the target density and the ideal flow model that conditional flow matching (CFM) would recover if target samples were available. We formulate the velocity condition as a fixed-point equation and exploit its conditional-expectation structure to construct a stop-gradient objective for enforcing it. The resulting training procedure retains the scalable sample-interpolate-regress structure of CFM despite the absence of target samples, using endpoints generated by the current flow. For online RL, the critic gradient supplies the target score at the generated actions, yielding a direct approach to actor training. Experiments on RL benchmarks demonstrate that SCF matches or improves upon state-of-the-art generative-policy baselines, while substantially reducing training time.
Primary: Massachusetts Institute of Technology
All Institutions: Massachusetts Institute of Technology, Qualcomm AI Research
Score-Calibrated Flow (SCF) provides a theoretically grounded and computationally efficient method for training flow matching models to sample from unnormalized densities by enforcing self-consistency conditions derived from the target score, thereby eliminating the need for importance sampling or backpropagation through sampling trajectories. The paper makes a significant contribution to the intersection of generative modeling and reinforcement learning by solving the "unnormalized density" problem with a novel stop-gradient objective that retains the scalability of standard flow matching while ensuring convergence to the ideal policy, demonstrated by strong empirical results on challenging RL benchmarks.
The paper proposes Score-Calibrated Flow (SCF), a method for training flow matching models to sample from unnormalized densities without access to target samples. The core innovation is a "self-consistency" approach that derives optimality conditions for the velocity field based on the target score (provided by the critic in RL) and the structure of the flow itself. The authors prove that the unique solution to these conditions is the ideal flow model that would be learned via standard Conditional Flow Matching (CFM) if target samples were available. The training objective is a stop-gradient regression on self-generated endpoints, avoiding the high variance of importance sampling and the computational cost of backpropagating through sampling trajectories. The theoretical grounding is strong, with clear proofs of uniqueness for both the terminal density and the velocity field.
The experiments include a toy example comparing SCF against recent samplers (AS, ASBS, BMS, FS) and extensive RL benchmarks (DeepMind Control Suite, HumanoidBench). SCF is compared against 8 strong baselines, including SAC and various generative policy methods (DIME, QSM, QFlex, etc.). The results show SCF matching or improving upon state-of-the-art methods while significantly reducing training time. The inclusion of wall-clock time comparisons is a strong practical contribution.
The paper provides detailed algorithmic descriptions and theoretical derivations. However, as a preprint, specific hyperparameters and code availability are not confirmed in the text. The method relies on standard flow matching components, making it relatively easy to implement if the specific coefficient schedules (mentioned in the appendix) are clear.
The method assumes access to the gradient of the potential function (critic), which is standard in RL but may not be available in all sampling tasks. The theoretical guarantees rely on certain regularity conditions (e.g., exponential moments) that may be hard to verify in practice for complex, high-dimensional distributions. The performance gains, while present, are incremental over strong baselines in some tasks.
This work addresses a fundamental bottleneck in applying generative models to online RL and other settings where only unnormalized densities are available. By providing a stable, efficient, and theoretically grounded training procedure, it could become a standard tool for training expressive policies in RL and for sampling in physics and molecular dynamics simulations. Score-Calibrated Flow (SCF) provides a theoretically grounded and computationally efficient method for training flow matching models to sample from unnormalized densities by enforcing self-consistency conditions derived from the target score, thereby eliminating the need for importance sampling or backpropagation through sampling trajectories. The paper makes a significant contribution to the intersection of generative modeling and reinforcement learning by solving the "unnormalized density" problem with a novel stop-gradient objective that retains the scalability of standard flow matching while ensuring convergence to the ideal policy, demonstrated by strong empirical results on challenging RL benchmarks.
During post-training of large language models (LLMs) with Reinforcement Learning with Verifiable Rewards (RLVR), GRPO-style algorithms can exhibit severe late-stage collapse. Prompt-based probing reveals that this is not benign strategic pruning, but a harmful contraction of effective strategy capacity that makes distinct reasoning strategies increasingly inaccessible. To characterize this phenomenon, we define strategies through trajectory-level policy-update interactions and develop a unified theoretical framework combining optimization dynamics and information theory. We prove that major RLVR objectives progressively concentrate probability mass onto a single strategy, while sustaining nontrivial task accuracy requires a minimum strategy capacity. The conflict between these two results provides a mechanistic explanation for catastrophic collapse. We further derive the {Mirrored Entanglement Index (MEI)} as a lightweight online warning signal. To prevent collapse, we propose \textbf{Mesh Learning}, which exposes multiple reasoning strategies and prevents any single strategy from dominating optimization. Across AIME26, AIME25, MATH-500, GPQA, and LiveCodeBench, Mesh Learning consistently outperforms strong baselines across Qwen and Phi model families, with gains of up to 13.4 pp and 11.5 pp, respectively. These results establish strategy preservation as a key principle for stable RLVR. Code is available at https://github.com/Ayanami-0123/Open-Mesh-Learning.
Primary: Peking University
All Institutions: Peking University
The paper identifies and mitigates catastrophic strategy collapse in RLVR for LLMs by introducing Mesh Learning. It provides a rigorous theoretical framework linking optimization dynamics to information-theoretic capacity constraints, supported by strong empirical results across multiple benchmarks and model families, establishing strategy preservation as a key principle for stable post-training.
The paper addresses a critical failure mode in Reinforcement Learning with Verifiable Rewards (RLVR) for Large Language Models (LLMs), specifically the "catastrophic strategy collapse" observed in GRPO-style algorithms. The authors propose a theoretical framework combining optimization dynamics and information theory to explain this phenomenon, defining "strategies" via trajectory-level policy-update interactions. They introduce the Mirrored Entanglement Index (MEI) as an online diagnostic metric and propose "Mesh Learning," a method designed to preserve strategy diversity by preventing any single reasoning strategy from dominating the optimization landscape. The theoretical contribution is significant as it moves beyond empirical observation to provide a mechanistic explanation for why standard RLVR objectives lead to entropy collapse and reduced effective capacity.
The experimental section evaluates the proposed method across a robust suite of benchmarks including AIME26, AIME25, MATH-500, GPQA, and LiveCodeBench. The models tested include the Qwen and Phi families, which are widely used in the community. The reported gains are substantial, with improvements of up to 13.4 percentage points on Qwen models and 11.5 percentage points on Phi models compared to strong baselines. The consistency of performance across diverse tasks (math, code, general reasoning) suggests that the method effectively preserves the general reasoning capabilities of the model while enhancing specific task performance, rather than overfitting to a narrow distribution of strategies.
The paper provides a public code repository (https://github.com/Ayanami-0123/Open-Mesh-Learning), which is a strong positive for reproducibility. Given the complexity of RLVR training and the specific hyperparameters likely required for "Mesh Learning," the availability of code is crucial. The use of standard, open-source model families (Qwen, Phi) further enhances the reproducibility of the results for other researchers.
The primary limitation is the computational cost associated with RLVR training, which remains high despite the proposed method. Additionally, the definition of "strategies" via trajectory-level interactions may be complex to interpret in non-mathematical or open-ended generation tasks, potentially limiting the direct applicability of the MEI metric to broader domains. The paper focuses on verifiable rewards, so its applicability to RLHF with human preference data (where rewards are less verifiable) is not directly addressed.
This work has significant implications for the stability and reliability of LLM post-training pipelines. By identifying and mitigating strategy collapse, it contributes to the development of more robust and versatile reasoning models. The concept of "strategy preservation" could influence future RLVR algorithms, potentially leading to a new class of training objectives that explicitly balance exploration and exploitation in the policy space. This is particularly relevant as the industry moves towards more complex, multi-step reasoning tasks where diverse strategies are essential. The paper identifies and mitigates catastrophic strategy collapse in RLVR for LLMs by introducing Mesh Learning. It provides a rigorous theoretical framework linking optimization dynamics to information-theoretic capacity constraints, supported by strong empirical results across multiple benchmarks and model families, establishing strategy preservation as a key principle for stable post-training.
Distillation has become a core primitive of large language model training, but its properties are not yet well understood. We take an entropic perspective, studying how the entropy of the student depends on the data and the divergence that define the distillation objective. We prove that forward KL inflates the entropy of the student above that of the teacher. Since cross-entropy training is a special case, this yields an identity that we verify quantitatively in pretraining and supervised finetuning. Other divergences come with no such guarantee: reverse KL deflates entropy until the gap between student and teacher gets too large, and interpolating between the two changes entropy smoothly early in training but abruptly at convergence. The lower entropy of on-policy distillation comes from token-level reverse KL, not from on-policy sampling. The divergence therefore acts as an implicit entropy regularizer, whose role is clearest in self-distillation: as conditioning on privileged information deflates entropy, the divergence hyperparameters that work best are those that compensate for it.
Primary: Stanford University
All Institutions: Stanford University
The paper provides a rigorous theoretical and empirical analysis of how divergence choices in distillation control student entropy, proving that forward KL inflates entropy while reverse KL deflates it, and demonstrating that this divergence acts as an implicit regularizer crucial for preventing entropy collapse in self-distillation. By decoupling sampling and divergence objectives and validating these insights on large-scale models like OLMo and Qwen, the work offers actionable insights for improving the stability and diversity of distilled language models.
The paper proposes a rigorous theoretical framework to analyze the effect of divergence choices in knowledge distillation on the entropy of the student model. The core methodological contribution is the decoupling of the sampling distribution (on-policy vs. off-policy) from the divergence objective (forward KL, reverse KL, etc.) by defining the objective at the token level rather than the sequence level. This allows for independent analysis of how each component affects entropy. The authors introduce a toy model of language modeling (linear softmax head with random embeddings) that is tractable for theoretical analysis yet captures key phenomena of large language models. They prove that forward KL inflates student entropy relative to the teacher by the residual divergence, and that reverse KL deflates entropy until the task becomes too hard, at which point it inflates. The analysis extends to generalized Jensen-Shannon divergences and top-k restrictions.
The experimental evaluation is strong and multi-faceted. The authors validate their theoretical predictions using the OLMo 2 suite of models, demonstrating that the entropy of a model matches its cross-entropy loss at convergence, a finding they claim is novel. They conduct ablation studies on Qwen3 models to show that the divergence choice has a larger impact on entropy than the sampling strategy (on-policy vs. off-policy). Finally, they apply their insights to self-distillation on SciKnowEval and GSM8K, showing that specific divergence hyperparameters are necessary to prevent entropy collapse when conditioning on privileged information. The experiments are well-designed to test specific theoretical claims.
The paper provides a link to a GitHub repository containing code to reproduce all experiments and figures, along with the underlying data. The use of open-source models (OLMo, Qwen3) and public datasets (DeepScaleR, SciKnowEval, GSM8K) further enhances reproducibility. The theoretical proofs are detailed in the appendix, allowing for verification of the mathematical claims.
The theoretical analysis relies on a toy model with random embeddings, which may not fully capture the complexities of real-world language models with deep sequence modeling components. While the authors argue that the token-level objective is what matters in practice due to stop-gradients, the gap between the toy model and large-scale transformers remains a potential limitation. The experiments, while thorough, are limited to specific model families (OLMo, Qwen) and datasets, and the generalizability to other architectures or domains is not explicitly tested.
This paper has significant implications for the design of distillation pipelines in LLM training. By identifying the divergence as an implicit entropy regularizer, it provides practitioners with a new knob to control model uncertainty and diversity. The finding that forward KL inflates entropy and reverse KL deflates it (until failure) offers practical guidance for choosing objectives in post-training and self-distillation. It also challenges the common practice of deriving on-policy distillation from sequence-level reverse KL, suggesting that decoupling these choices can lead to better performance and stability. The paper provides a rigorous theoretical and empirical analysis of how divergence choices in distillation control student entropy, proving that forward KL inflates entropy while reverse KL deflates it, and demonstrating that this divergence acts as an implicit regularizer crucial for preventing entropy collapse in self-distillation. By decoupling sampling and divergence objectives and validating these insights on large-scale models like OLMo and Qwen, the work offers actionable insights for improving the stability and diversity of distilled language models.
Simulating microstructure evolution requires quantum-mechanical accuracy and mesoscopic reach in length and time scales, a combination that no current method achieves. Classical phase-field models provide this reach, but their accuracy is limited by phenomenological free energies and mobilities. Here we develop a framework for learning ab initio phase-field models, where the mesoscopic equation is not postulated but derived from a Mori-Zwanzig projection of molecular dynamics onto species-density fields under explicit assumptions. The nonlocal free energy and mobility left unspecified by this equation are parametrized by neural networks and learned from short molecular dynamics trajectories generated with machine-learning interatomic potentials of ab initio accuracy. We demonstrate the framework on an iron-boron melt and on hydrogen-helium mixtures under planetary conditions. For iron-boron, the model shows that the melt at the FeB$_4$ composition is spinodally unstable at ambient pressure but stabilized at 10 GPa, offering a thermodynamic rationale for why FeB$_4$ has been synthesized only under high pressure. For hydrogen-helium, the model predicts the immiscibility boundary and captures droplet nucleation and growth in helium-rain simulations of a column corresponding to 2.2 million atoms, far beyond the scale of atomistic modeling at comparable accuracy. Trained across compositions and conditions, such models could provide a mesoscopic counterpart to ab initio molecular dynamics.
Primary: National University of Singapore
All Institutions: National University of Singapore
The paper presents a rigorous, physics-informed framework for learning ab initio phase-field models by deriving mesoscopic equations from Mori-Zwanzig projections and jointly learning free-energy and mobility functionals from MLIP-generated MD data, demonstrating significant scalability and accuracy in simulating complex microstructure evolution across materials and planetary science applications.
The paper proposes a rigorous framework for deriving mesoscopic phase-field equations from microscopic molecular dynamics (MD) using the Mori-Zwanzig projection formalism. Unlike standard data-driven approaches that fit phenomenological parameters, this method derives the structure of the evolution equation (Generalized Langevin Equation reduced to a Markovian form) from statistical mechanics. The key innovation is the joint learning of the nonlocal free-energy functional and the mobility tensor using neural networks, trained on short MD trajectories generated by machine-learning interatomic potentials (MLIPs). The authors introduce a "ladder" of free-energy functionals, selecting a specific rung (local density plus nonlocal kernel) that balances expressiveness with data efficiency. The derivation is mathematically sound, explicitly stating assumptions (Markovian dynamics, local mobility, Gaussian noise) that justify the reduction from the exact GLE to the practical model. The use of fluctuation-dissipation relations to anchor the mobility and free-energy curvature provides a strong physical constraint on the learning process, preventing the neural networks from learning unphysical dynamics.
The experimental validation is extensive and covers three distinct systems with varying complexity: a Lennard-Jones mixture (benchmark), an Iron-Boron melt (materials science), and a Hydrogen-Helium mixture (planetary science). For the LJ mixture, the model accurately reproduces the phase diagram and spinodal decomposition dynamics, outperforming Landau and Flory-Huggins baselines. The Iron-Boron results provide a thermodynamic rationale for the high-pressure synthesis of FeB4, correctly predicting the stabilization of the melt at 10 GPa. The most impressive result is the Hydrogen-Helium simulation, where the model predicts the immiscibility boundary consistent with prior DFT studies and simulates helium rain in a domain corresponding to 2.2 million atoms over 1 ns—a scale far beyond the reach of direct MD at comparable accuracy. The ability to extrapolate to larger scales and include external potentials (gravity) not present in training data demonstrates the model's generalizability.
The paper provides detailed descriptions of the training data generation, the specific neural network architectures (MLPs for local terms, radial kernels for nonlocal terms), and the loss function components (dynamical, mobility anchor, structure factor anchor, pressure anchor). The use of standard MLIPs (MACE) for data generation enhances reproducibility, as these potentials are increasingly available. However, specific hyperparameters, network sizes, and training protocols are deferred to the Supplemental Material, which is not fully visible in the provided text. The code availability is not explicitly stated in the main text, which is a minor gap for immediate reproducibility, though the methodology is described with sufficient precision for re-implementation.
The framework currently resolves only density fields, neglecting momentum density, which limits its ability to capture hydrodynamic effects like droplet settling via flow rather than just diffusion. The Markovian assumption may fail for systems with long memory times, such as supercooled liquids or glasses. The accuracy of the model is inherently bounded by the accuracy of the underlying MLIPs used to generate training data. Additionally, the computational cost of generating high-quality MD data for complex materials remains a barrier, although the authors argue that the data efficiency of the approach mitigates this.
This work bridges the gap between quantum-mechanical accuracy and mesoscopic scalability, a long-standing challenge in materials science and fluid dynamics. By providing a principled way to learn mesoscopic models from microscopic data, it enables simulations of phenomena like phase separation, nucleation, and coarsening at scales and timescales previously inaccessible. The potential applications are vast, ranging from alloy design and synthesis to understanding planetary interiors. The framework could serve as a foundation for "mesoscopic foundation models" that generalize across compositions and conditions, analogous to CALPHAD databases but with predictive dynamical capabilities. The paper presents a rigorous, physics-informed framework for learning ab initio phase-field models by deriving mesoscopic equations from Mori-Zwanzig projections and jointly learning free-energy and mobility functionals from MLIP-generated MD data, demonstrating significant scalability and accuracy in simulating complex microstructure evolution across materials and planetary science applications.
Single-particle cryo-electron microscopy (cryo-EM) has become a widely adopted technique for biomolecular structure determination. The conventional cryo-EM computational pipeline first combines many particle images to reconstruct an electrostatic potential (ESP) map and then fits an atomic model to the recovered map. Density reconstruction has high sample complexity, requiring large numbers of particle images and making structure determination high-cost and low-throughput, particularly for heterogeneous samples. Downstream atomic model building, in turn, becomes increasingly difficult as the resolution of the reconstructed map deteriorates. Protein structure prediction models provide strong sequence-derived priors on atomic structure, and experiment-guided approaches can use these priors to recover structures consistent with experimental measurements. Yet, in cryo-EM, such priors are typically integrated only after density reconstruction during atomic model fitting. We introduce Fold'EM, an inference-time framework that combines priors from protein generative models directly with cryo-EM particle images to determine atomic models from a small number of single particle images, bypassing both intermediate density reconstruction and downstream model building against the reconstructed map. Across synthetic and experimental cryo-EM datasets, Fold'EM recovers accurate atomic structures both with known particle orientations and in an ab-initio setting where orientations are inferred jointly with structure. In heterogeneous datasets, Fold'EM further resolves distinct conformational states from mixed particle populations without separately reconstructing a density map and building an atomic model for each state. We believe these results open new avenues for structure determination in the low-sample regime and for characterizing low-population conformational states directly from cryo-EM particles.
Primary: University of Cambridge
All Institutions: University of Cambridge, DeepMind
Fold'EM introduces a novel inference-time framework that directly guides AlphaFold3 using cryo-EM particle images, bypassing intermediate density reconstruction. By optimizing intermediate Pairformer states and jointly inferring structure, pose, and class, the method achieves high-resolution atomic structures from sparse particle sets, significantly reducing the sample complexity of cryo-EM structure determination.
The paper introduces Fold'EM, a framework that bypasses the traditional cryo-EM density reconstruction step by directly guiding AlphaFold3 using individual particle images. The core technical novelty lies in optimizing intermediate Pairformer states ($H$) rather than the terminal embeddings ($Z$) or atomic coordinates. The authors provide a rigorous theoretical justification for this choice, demonstrating that optimizing $H$ induces a sequence-dependent anisotropic geometry in the conditioning space, which acts as a learned preconditioner for the gradient updates. This prevents the model from drifting into unrealistic structural configurations during the inference-time optimization. The method also introduces a joint inference scheme for structure, pose, and class assignment, utilizing a novel bootstrapped residual-subspace estimator (B-RECOVAR) for label-free classification in heterogeneous samples. The mathematical formulation of the particle-level likelihood versus density-level likelihood is well-articulated, highlighting the statistical insufficiency of consensus densities when poses are uncertain.
The experimental evaluation is extensive, covering both synthetic benchmarks and experimental EMPIAR datasets. The paper demonstrates significant improvements over baselines (RELION, cryoDRGN, and previous AlphaFold-guided methods) in terms of FSC resolution and TM-score, particularly in the low-sample regime (e.g., recovering structures from as few as 100-500 particles). The ablation studies effectively isolate the contributions of particle-level guidance versus density guidance and the optimization of intermediate states. The results on heterogeneous data show the ability to resolve distinct conformational states without separate density reconstructions, which is a significant practical advantage. The comparison with unguided AlphaFold3 clearly shows the value of the experimental guidance in correcting mispredicted conformations.
The paper provides detailed algorithmic descriptions and mathematical formulations. However, as an arXiv preprint, the code availability is not explicitly confirmed in the text provided. The reliance on AlphaFold3, a proprietary model, may limit full reproducibility for some users, though the method itself is clearly defined. The use of standard cryo-EM tools (RELION, cryoDRGN) for baselines aids in comparability.
The method relies heavily on the quality of the AlphaFold3 prior; if the unguided prediction is significantly wrong (e.g., 8PWH case), the method may struggle or require more particles than density-based methods. The computational cost of running multiple reverse-diffusion runs and back-propagating through the Pairformer could be high. The current implementation assumes known CTF parameters and centered particles, which may not always hold in raw experimental data without preprocessing.
This work has the potential to significantly reduce the sample size required for cryo-EM structure determination, making it feasible to study low-population states and scarce samples. It bridges the gap between protein structure prediction and experimental cryo-EM, offering a new paradigm for structure determination that could accelerate drug discovery and structural biology research. The insights into optimizing intermediate representations of generative models may also be applicable to other inverse problems in scientific imaging. Fold'EM introduces a novel inference-time framework that directly guides AlphaFold3 using cryo-EM particle images, bypassing intermediate density reconstruction. By optimizing intermediate Pairformer states and jointly inferring structure, pose, and class, the method achieves high-resolution atomic structures from sparse particle sets, significantly reducing the sample complexity of cryo-EM structure determination.
Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composition is lossy. A vote can only overturn mistakes that its voters do not share, and RLVR sharpens a policy so that its samples increasingly make the same mistakes. With every method drawing exactly $160$ completions per problem, training a single LoRA adapter on the full RLVR budget raises single-sample accuracy on every model we test ($1.5$B-$8$B). Yet on three of four models it leaves the majority vote below that of the untrained base model, by up to $4.8$ points. The damage builds during training: voter errors grow steadily more correlated, and the majority vote accuracy peaks early before falling by up to $7.0$ points. The cause is concentration, not RLVR itself. We split the same data and training budget across $K$ LoRA adapters, each trained on its own random disjoint shard, and call the result an adapter thicket. Thickets out-vote the fully trained adapter in all $16$ (model, $K$) settings, and for $K{\geq}4$ they stay within $0.8$ points of the base model or above it. A single adapter stopped early, at a thicket member's step count, is a strong control that matches thickets for small $K$. For $K{\geq}8$, thickets keep more of RLVR's single-sample gain and out-vote this control in six of eight settings. The cost of concentration also grows with the number of votes: from $16$ to $160$ votes, the thicket's lead over the fully trained adapter widens from $1.3$ to $3.3$ points. When the plan is to sample and vote, an RLVR budget is better spent broad than deep.
Primary: Princeton University
All Institutions: Princeton University
The paper demonstrates that splitting RLVR training budgets across multiple LoRA adapters ("thickets") preserves the diversity necessary for effective majority voting, outperforming the standard single-adapter pipeline which suffers from error correlation and diversity collapse.
The paper proposes a simple but effective alternative to the standard RLVR (Reinforcement Learning with Verifiable Rewards) pipeline: instead of training a single LoRA adapter on the full dataset, it splits the data and budget across K independent adapters ("thickets"). The core insight is that RLVR sharpens the policy, reducing sample diversity and increasing error correlation among voters in majority voting, which degrades test-time scaling performance. The methodology is sound, using a fixed inference budget (160 completions) to isolate the effect of training strategy from sampling budget. The use of a step-matched control (early stopping the single adapter) is a critical experimental design choice that strengthens the causal claim.
The experiments are rigorous and well-controlled. The authors test across four models (Qwen2.5-1.5B/3B/7B, Llama-3.1-8B) and two domains (Math, Code). They demonstrate that the single-adapter approach often performs worse than the untrained base model in majority voting, while the thicket approach recovers this performance and often exceeds it. The analysis of error correlation and coverage provides strong mechanistic evidence for the observed results. The ablation on shard type (random vs. subject-specific) adds practical value.
The paper provides sufficient detail on the training setup (GRPO, LoRA rank, hyperparameters) and evaluation protocols (temperature, top-p, voting logic). The specific datasets (MATH, GSM8K, etc.) and model checkpoints are standard, making reproduction feasible for a lab with adequate compute resources. The code is not explicitly linked in the provided text, but the methods are standard enough to implement.
The study is limited to models up to 8B parameters and specific RL algorithms (GRPO). The voting results are primarily for math tasks where answers are canonicalizable; code tasks are only analyzed for coverage and correlation, not final vote accuracy. The "thicket" approach requires managing multiple adapters at inference time, which, while mitigated by modern serving stacks, adds operational complexity compared to a single model.
This paper challenges a common assumption in the LLM post-training community: that more RLVR training is always better for test-time scaling. It provides a clear, actionable guideline for practitioners: if you plan to use majority voting, diversify your training budget. This has immediate practical implications for teams deploying LLMs with test-time compute budgets. The paper demonstrates that splitting RLVR training budgets across multiple LoRA adapters ("thickets") preserves the diversity necessary for effective majority voting, outperforming the standard single-adapter pipeline which suffers from error correlation and diversity collapse.
Biological reasoning models use post-training to connect LLMs to biological foundation model representations and biological text. Their benchmark accuracy is taken as evidence that LLMs reason over these inputs. We test this assumption in six biological reasoning models across DNA, protein, and single-cell tasks. We perturb one biological input while holding the query and other inputs fixed, construct evidence conflicts that pair the foundation model representation of one genome, protein, or cell with the text of another, fit linear probes to the representations the language model receives, and analyze reasoning traces against the biological inputs. Evo2 and ESM3 contribute little to BioReason and BioReason-Pro performance on the evaluated tasks. Shuffling the DNA sequence barely changes BioReason disease prediction accuracy, and in evidence conflicts the two models follow the text in 97.9% and 99.7% of cases. Linear probes trained on the Evo2 and ESM3 representations predict the task targets, so these foundation models encode information relevant to the task, but provide limited overall performance improvement to BioReason and BioReason-Pro. In contrast, foundation model inputs contribute to ChatNT, Prot2Text-V2, and CellWhisperer performance, and differentially expressed genes in the gene sentence contribute to Cell2Sentence-Scale performance. Across SFT and RL checkpoints of BioReason-Pro and 42 BioReason checkpoints, increases in accuracy do not imply greater performance contributions from biological inputs. BioReason traces misstate nucleotide changes, while BioReason-Pro traces describe functions omitted from final predictions under evidence conflicts. We find that current post-training strategies do not ensure that foundation model representations contribute to task performance.
Primary: Harvard University
All Institutions: Harvard University, Harvard Medical School
The paper provides a critical evaluation of biological reasoning models, demonstrating that current post-training strategies often fail to ensure that foundation model representations contribute to task performance. By using rigorous perturbation and evidence conflict analyses, the authors reveal that models like BioReason and BioReason-Pro largely ignore their biological inputs in favor of textual shortcuts, a finding that has significant implications for the design and evaluation of future multimodal scientific AI systems.
The paper employs a rigorous interventionist methodology to test the causal role of biological foundation model representations in LLM-based reasoning models. The authors utilize input perturbation (shuffling DNA sequences), evidence conflict construction (mismatching foundation model embeddings with text descriptions), and linear probing to isolate the contribution of specific modalities. This approach is methodologically sound for determining whether models are genuinely "reasoning" over biological inputs or merely relying on textual shortcuts. The inclusion of analysis across multiple training checkpoints (SFT and RL) adds depth to the evaluation of how post-training strategies affect input utilization.
The experiments cover six distinct biological reasoning models across DNA, protein, and single-cell tasks, providing a broad scope. The key finding that Evo2 and ESM3 contribute negligibly to BioReason and BioReason-Pro performance, with models following text in >97% of conflict cases, is a significant empirical result. The contrast with models like ChatNT and CellWhisperer, where foundation model inputs do contribute, highlights that the issue is specific to certain post-training strategies or model architectures rather than a universal failure of multimodal integration. The analysis of reasoning traces revealing misstatements of nucleotide changes further supports the conclusion that the models are not faithfully using the biological inputs.
The paper provides a GitHub repository link (https://github.com/mims-harvard/bio-mirage) and a project website, which strongly suggests that code and resources will be available for reproduction. The detailed description of the perturbation and conflict construction methods allows other researchers to replicate the analysis on other models.
The study focuses on a specific set of six models and three biological domains. It is unclear if the findings generalize to other types of biological foundation models or different post-training paradigms. The linear probes may not capture all non-linear relationships between the foundation model representations and the task targets, potentially underestimating the contribution of the biological inputs in some cases.
This paper has high impact for the field of AI in science, particularly for the development of trustworthy biological reasoning systems. It challenges the common assumption that high benchmark accuracy implies effective use of specialized scientific inputs. The findings will likely influence future post-training strategies to explicitly reward the use of foundation model representations, leading to more robust and interpretable AI models for biology. The paper provides a critical evaluation of biological reasoning models, demonstrating that current post-training strategies often fail to ensure that foundation model representations contribute to task performance. By using rigorous perturbation and evidence conflict analyses, the authors reveal that models like BioReason and BioReason-Pro largely ignore their biological inputs in favor of textual shortcuts, a finding that has significant implications for the design and evaluation of future multimodal scientific AI systems.
Echocardiography is the most widely used cardiac imaging modality, yet interpretation demands integrating visual evidence across global anatomy, localized structures and dynamic cardiac motion. Machine-learning models have automated individual tasks, but they are typically built for a single purpose and depend on expensively labeled datasets - a barrier particularly acute in pediatric care, where data are scarce and anatomy changes with age. Here we present EchoDino, a self-supervised foundation model for echocardiography, created by adapting the DINOv3 framework to 3.7 million frames from 1.7 million unlabeled pediatric echocardiography videos. With its encoder frozen, EchoDino produces representations that capture global context, local anatomy, and dense spatial detail. We introduce Motion-biased Entropy Maximization Sampling (MEMS) to select the most informative frames for video-level analysis. Across nine pediatric and adult datasets, EchoDino outperformed strong baseline models, raising view-classification accuracy from 0.609 to 0.889 and the area under the receiver operating characteristic curve for structural-heart-disease detection from 0.811 to 0.872, while also cutting age-estimation error from 3.857 to 1.389 years, achieving the best segmentation accuracy and lowering ejection-fraction errors. By generalizing from label-free pediatric data to adult echocardiography, EchoDino offers a versatile foundation for cardiac image analysis across the lifespan.
Primary: Rice University
All Institutions: Rice University, Baylor College of Medicine, Texas Children's Hospital
EchoDino adapts the DINOv3 self-supervised framework to pediatric echocardiography, introducing a frozen encoder with spatial patch re-aggregation and motion-biased frame sampling to achieve state-of-the-art performance across diverse cardiac imaging tasks in both pediatric and adult populations.
The paper proposes EchoDino, a self-supervised foundation model for echocardiography by adapting the DINOv3 framework. The core methodological contribution is the domain adaptation of a natural image foundation model to pediatric cardiac ultrasound using 3.7 million unlabeled frames. The authors introduce two specific architectural/algorithmic extensions: EchoDino-PATCH, which re-aggregates spatial patch tokens to preserve local anatomical detail for segmentation and measurement, and MEMS (Motion-biased Entropy Maximization Sampling), a frame selection strategy for video-level tasks that prioritizes high-motion, feature-diverse frames over uniform sampling. The approach relies on a frozen encoder with lightweight task-specific readouts, which is a standard but effective paradigm for foundation models. The adaptation of DINOv3's teacher-student architecture to the specific augmentation and cropping needs of echocardiography is well-motivated.
The experimental evaluation is extensive and rigorous. The model is tested across nine datasets, including five internal pediatric cohorts, one external pediatric cohort, and three external adult datasets. This cross-domain evaluation (pediatric to adult) is a strong point, demonstrating the generalizability of the learned representations. The tasks cover a wide spectrum: global (view classification), localized (measurement, SHD detection), dense (segmentation), and temporal (EF prediction, sweep recognition). The results show consistent and significant improvements over strong baselines like DINOv3, PanEcho, and EchoPrime. The use of bootstrap confidence intervals and patient-disjoint splits adds statistical rigor. The ablation of MEMS vs. uniform sampling clearly isolates the benefit of the proposed sampling strategy.
The paper provides a link to the code repository. However, the primary pretraining data (TCH-Complex) is not publicly available due to clinical data governance, which limits full reproducibility of the pretraining phase. The downstream evaluation on public datasets (EchoNet, CAMUS) is reproducible. The hyperparameters and training details are described in the supplementary methods, aiding partial reproducibility.
The main limitation is the lack of public access to the large-scale pediatric pretraining corpus, which prevents independent verification of the pretraining process. The SHD detection is evaluated as a binary composite classifier, which may mask performance on specific rare lesions. The paper acknowledges that aggregate metrics do not guarantee individual point-of-care actionability and that uncertainty estimation is needed for clinical deployment.
This work has significant potential impact in the field of medical imaging AI, particularly in pediatric cardiology where labeled data is scarce. By demonstrating that a single frozen encoder can support diverse tasks across age groups, it offers a scalable framework for developing modular clinical tools. The approach of adapting general-purpose vision foundation models to specific medical imaging modalities is a trend that this paper contributes to meaningfully. EchoDino adapts the DINOv3 self-supervised framework to pediatric echocardiography, introducing a frozen encoder with spatial patch re-aggregation and motion-biased frame sampling to achieve state-of-the-art performance across diverse cardiac imaging tasks in both pediatric and adult populations.
In an image latent space, the embeddings of high-resolution, natural, and sharp images form a manifold. Degradation of high-resolution images pushes their embeddings off this manifold. Real-world super-resolution (SR) then becomes the task of mapping the degraded embedding back onto this manifold --- not anywhere on the manifold, but to the point that preserves what the input still carries, both its semantics and pixel details. Every published method implements this mapping in a reconstruction-oriented latent space or pixel space. We claim these spaces are the wrong substrates for SR. Low-resolution and degraded images are embedded far from the manifold, making the mapping difficult and expensive. The lack of semantic information in these substrates also makes it difficult to navigate to the faithful point on the manifold, causing severe hallucination when degradation is heavy. Thus, restoring in a suitable latent space is crucial to the SR task. We show that the latent space of 23 fused layers of a frozen DINOv3-L is one such space that makes the SR task easier. Degraded images are embedded near the manifold. Moreover, this substrate contains a hierarchy of information, from pixel record to degradation robust semantics, guiding the model to find the faithful point on the manifold. On this substrate, a 415M decoder is trained under reconstruction and adversarial objectives to map the degraded embeddings back and decode to pixel space in one pass. The resulting model, RAESR, attains the best fidelity--perception trade-off among state-of-the-art adversarial and diffusion-based restorers on RealSR, DRealSR, LSDIR and DIV2K-Val, at 37 ms per 512 by 512 image on a single H20 GPU. Swapping the substrate for a VAE latent under an identical recipe loses on every metric.
Primary: University of Michigan
All Institutions: University of Michigan, Southeast University, Alibaba Group
The paper introduces RAESR, a super-resolution model that leverages the latent space of a frozen DINOv3-L vision transformer as a substrate for restoration, demonstrating that this space embeds degraded images closer to the clean-image manifold than VAEs, thereby enabling a more efficient and faithful one-pass mapping back to high-quality images. The technical contribution is significant due to the rigorous geometric analysis of the latent space properties and the strong empirical results showing state-of-the-art fidelity-perception trade-offs with high inference efficiency.
The paper proposes RAESR, a super-resolution model that operates in the latent space of a frozen DINOv3-L vision transformer rather than a VAE or pixel space. The core methodological contribution is the argument that self-supervised representation spaces (specifically DINOv3) embed degraded images closer to the "manifold" of clean images than reconstruction-oriented spaces (VAEs), thereby simplifying the mapping back to the clean manifold. The authors validate this geometric intuition with rigorous ablations, including latent line interpolation experiments and layer-wise information recovery analysis. The decoder is a 415M parameter transformer trained with reconstruction and adversarial losses, using a multi-layer DINOv3-B critic for fine-grained realism supervision. The approach is technically sound, well-motivated, and the ablations are exceptionally thorough, providing strong evidence for the central hypothesis.
The experiments are comprehensive, comparing RAESR against 12 state-of-the-art methods across four real-world benchmarks (RealSR, DRealSR, LSDIR, DIV2K-Val). RAESR achieves the best fidelity-perception trade-off, outperforming heavy diffusion-based models in both quality and efficiency (37ms per image). The inclusion of a human preference study and detailed per-benchmark breakdowns adds robustness. The ablation studies, particularly the comparison against a VAE substrate with identical training recipes, are critical and convincingly demonstrate that the performance gain stems from the choice of latent space rather than just the decoder architecture.
The paper provides extensive details on training schedules, hyperparameters, data processing, and evaluation protocols in the appendices. The use of standard datasets and public baselines facilitates reproduction. However, the reliance on specific frozen checkpoints (DINOv3-L, RAEv2) and the complex multi-stage training process may pose some barriers for independent replication without access to the authors' code or precise checkpoint versions.
The model is a single-pass restorer, lacking the ability to trade compute for quality on difficult images, unlike iterative diffusion methods. The claim about substrate suitability is currently supported primarily by DINOv3-L and tested against one alternative (SD3 VAE), leaving the behavior of other representation encoders unexplored. The model size (719M inference parameters) is still significant compared to lightweight GANs, though it is competitive with diffusion models.
This work shifts the perspective in super-resolution research from merely improving generative priors to carefully selecting the representation space for restoration. It highlights the utility of self-supervised vision foundation models for low-level vision tasks, potentially inspiring similar approaches in other restoration tasks like denoising or inpainting. The efficiency gains over diffusion models make it attractive for real-time applications. The paper introduces RAESR, a super-resolution model that leverages the latent space of a frozen DINOv3-L vision transformer as a substrate for restoration, demonstrating that this space embeds degraded images closer to the clean-image manifold than VAEs, thereby enabling a more efficient and faithful one-pass mapping back to high-quality images. The technical contribution is significant due to the rigorous geometric analysis of the latent space properties and the strong empirical results showing state-of-the-art fidelity-perception trade-offs with high inference efficiency.
Dense 3D point tracking has been a prominent paradigm for modeling motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel: it captures where pixels go, not the rotation of the underlying part, nor which pixels move together as one body. We propose MoSE3, the first feed-forward model that predicts dense SE(3) motion from monocular RGB video, producing full 6-DoF rigid transforms at every pixel in world space. Per-pixel SE(3) motion offers a richer view of how a scene moves: rotation, translation, and grouping all at once. Directly predicting SE(3) is challenging: rotations lie on a curved manifold that is ill-suited to Euclidean regression, and annotations for SE(3) are particularly difficult to acquire. To address these challenges, MoSE3 predicts per-pixel SE(3) through two jointly learned intermediates, 3D point tracks and rigidity embeddings, and recovers SE(3) by differentiably fitting transforms within each soft rigid cluster, enabling end-to-end prediction and supervision. To close the data gap, we introduce Art-Kubric, a large-scale synthetic dataset with dense SE(3) and rigidity labels for articulated objects with rich physical interactions. MoSE3 achieves state-of-the-art SE(3) estimation at pixel, part, and object levels on both rigid and articulated benchmarks, and state-of-the-art average 3D point tracking accuracy across three datasets, while showing strong generalization to real-world videos despite being trained solely on synthetic motion data.
Primary: Harvard University
All Institutions: Harvard University, Johns Hopkins University, Kempner Institute
[One sentence main contribution]. MoSE3 introduces the first feed-forward model for dense SE(3) motion estimation from monocular video, utilizing a novel decomposition into point tracks and rigidity embeddings to handle the non-Euclidean nature of rotations, and contributes a large-scale synthetic dataset, Art-Kubric, to enable training and evaluation. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper addresses a fundamental gap in motion modeling by moving beyond 3-DoF translation to full 6-DoF rigid transforms. The technical approach is sound, leveraging differentiable optimization to fit SE(3) transforms within learned rigid clusters, which effectively circumvents the difficulties of direct manifold regression. The introduction of Art-Kubric is a substantial contribution to the field, providing the necessary ground truth data for this complex task. The experimental results are strong, showing SOTA performance on both the new SE(3) task and existing point tracking benchmarks, with impressive generalization to real-world data. This work is likely to influence future research in dense motion estimation and provide a valuable tool for applications requiring detailed understanding of scene dynamics.
The paper proposes MoSE3, a feed-forward architecture that predicts dense SE(3) motion (6-DoF: rotation and translation) from monocular RGB video. The core innovation lies in avoiding direct regression on the non-Euclidean SO(3) manifold by decomposing the problem into two jointly learned intermediates: 3D point tracks (translations) and rigidity embeddings. The SE(3) transforms are then recovered via differentiable fitting within soft rigid clusters defined by the rigidity embeddings. This approach elegantly handles the manifold constraint and the grouping of pixels into rigid bodies. The introduction of the Art-Kubric dataset, featuring dense SE(3) and rigidity labels for articulated objects with physical interactions, is a significant methodological contribution to address the lack of ground truth data for this specific task.
The experiments demonstrate state-of-the-art performance in SE(3) estimation at pixel, part, and object levels on both rigid and articulated benchmarks. The model also achieves state-of-the-art average 3D point tracking accuracy across three datasets. A particularly strong result is the generalization to real-world videos despite training solely on synthetic data, which validates the robustness of the learned representations. The evaluation covers both the novel SE(3) task and the established point tracking task, providing a comprehensive assessment of the model's capabilities.
The paper provides a project page URL. As an arXiv preprint, the code availability is not explicitly confirmed in the provided text, but the detailed description of the method and the release of a large-scale synthetic dataset (Art-Kubric) suggest a high level of reproducibility. The use of standard synthetic data generators (Kubric) for the dataset creation further aids reproducibility.
The primary limitation is the reliance on synthetic data for training, which may introduce a domain gap for certain real-world scenarios not covered by the synthetic generator, although the paper claims strong generalization. The method assumes rigid or articulated motion within clusters; highly deformable objects might not be captured accurately by the SE(3) fitting approach. The computational cost of differentiable fitting within clusters could be a bottleneck for very high-resolution videos or long sequences.
This work has significant implications for robotics, augmented reality, and video understanding. Dense SE(3) motion estimation provides a richer representation of scene dynamics than point tracking alone, enabling better understanding of object interactions, part-level motion, and scene structure. The Art-Kubric dataset will likely become a standard benchmark for motion estimation tasks involving articulated objects. The ability to predict 6-DoF motion from monocular video could enhance applications in robot manipulation, human motion analysis, and 3D scene reconstruction. [One sentence main contribution]. MoSE3 introduces the first feed-forward model for dense SE(3) motion estimation from monocular video, utilizing a novel decomposition into point tracks and rigidity embeddings to handle the non-Euclidean nature of rotations, and contributes a large-scale synthetic dataset, Art-Kubric, to enable training and evaluation. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper addresses a fundamental gap in motion modeling by moving beyond 3-DoF translation to full 6-DoF rigid transforms. The technical approach is sound, leveraging differentiable optimization to fit SE(3) transforms within learned rigid clusters, which effectively circumvents the difficulties of direct manifold regression. The introduction of Art-Kubric is a substantial contribution to the field, providing the necessary ground truth data for this complex task. The experimental results are strong, showing SOTA performance on both the new SE(3) task and existing point tracking benchmarks, with impressive generalization to real-world data. This work is likely to influence future research in dense motion estimation and provide a valuable tool for applications requiring detailed understanding of scene dynamics.
Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing on learning sciences, we introduce a masked near-transfer post-test: the student is tested on an unseen variant of the tutored problem with the tutor's utterances masked, so the reward can rise only through what the student wrote in its own turns. This discourages cognitive offloading by the student and allows the continuous penalty to be replaced by two binary reward gates (factual correctness of tutor response, no solution handover). A leave-one-out ablation shows that the learning-gain reward on its own does not separate teaching from telling: the gates reduce solution handover while the near-transfer post-test improves out-of-domain transfer. Using these reward designs we develop Eduardo, a multi-turn RL recipe for training LLM tutors, and use it to train 4B, 9B, 14B and 27B models from two distinct LLM architectures. Our post-trained Eduardo-27B model matches Gemini-3.1-Pro on MathTutorBench and Claude Opus 4.8 on TutorMoments at 2.4-6.2x fewer thinking tokens than frontier models, which matters for interactive tutoring. Without being named in the reward, the model more than doubles its use of the push-for-justification teacher move while support fading (e.g., assigning independent work), whose payoff lies beyond a single-problem dialog episode, is trained out. We open-source our training environment, an 8,671-problem near-transfer dataset, and trained models for further development.
Primary: ETH Zurich (Inferred from "eth-lre" in GitHub URL and Swiss AI Initiative funding)
All Institutions: ETH Zurich, Swiss National Supercomputing Centre (CSCS)
The paper introduces a robust RL framework for training LLM tutors by using masked near-transfer post-tests and binary reward gates to prevent solution handover, resulting in models that match frontier performance with higher efficiency. The methodology is sound, the experiments are thorough, and the contributions are significant for the field of AI for Education and multi-turn RL.
The paper addresses a critical flaw in existing RL-based tutoring methods: the "telling" problem, where models learn to simply provide answers to maximize immediate student success metrics. The authors propose a novel reward structure based on three conditions: (1) Near-transfer testing (student solves a variant problem, not the exact one), (2) Masked post-test (tutor's utterances are hidden during the test, forcing reliance on student-generated notes/reasoning), and (3) Binary reward gates (strict penalties for factual errors or solution handover). This design effectively decouples "helpfulness" from "teaching," forcing the model to elicit student reasoning rather than provide it. The use of a frozen LLM student with genuine errors (rather than prompted confusion) is a strong methodological choice that ensures the training signal reflects real diagnostic and adaptive challenges.
The experimental setup is rigorous and comprehensive. The authors train models across multiple sizes (4B to 27B) and architectures (Qwen3 variants), demonstrating the robustness of the recipe. The ablation study is particularly strong, isolating the contribution of each condition (transfer, masking, gates) and showing that the gates specifically reduce handover while the masked transfer test improves out-of-domain generalization. The evaluation uses independent benchmarks (MathTutorBench, TutorMoments) and judges (Gemini) distinct from the training setup, mitigating overfitting concerns. The results show that Eduardo-27B matches or exceeds frontier models (Gemini 3.1 Pro, Claude Opus 4.8) in pedagogical quality while using significantly fewer thinking tokens, a crucial efficiency gain for interactive applications.
High. The authors open-source the training environment, the 8,671-problem near-transfer dataset, and the trained models. Detailed hyperparameters, prompts, and compute costs are provided in the appendices. The use of standard RL algorithms (DPPO/GRPO) and open-source base models further enhances reproducibility.
The primary limitation is the reliance on a single frozen LLM student (Llama-3.1-8B-Instruct) for training, which may lead to overfitting to that specific model's failure modes. The paper acknowledges this and suggests future work with diverse student ensembles. Additionally, the evaluation is limited to mathematics, and the "affective gap" (lack of socio-emotional support) is noted as a consequence of optimizing purely for cognitive gain. The ablation study is limited to one seed and one model size (4B), which slightly weakens the statistical confidence in the ablation results.
This work has significant implications for the development of AI tutors and, more broadly, for any domain where an agent must build user capabilities rather than just provide answers. The "masked near-transfer" reward design is a generalizable principle that could be applied to other educational or collaborative tasks. The efficiency gains (fewer thinking tokens) make high-quality tutoring more feasible in real-time interactive settings. The open-sourcing of the dataset and models will likely accelerate research in AI for Education. The paper introduces a robust RL framework for training LLM tutors by using masked near-transfer post-tests and binary reward gates to prevent solution handover, resulting in models that match frontier performance with higher efficiency. The methodology is sound, the experiments are thorough, and the contributions are significant for the field of AI for Education and multi-turn RL.
Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play. On the other hand, language models (LMs) can generate plausible-sounding explanations, but their weak playing strength limits the utility of their explanations. We introduce Queen, a 4B-parameter chess-language model that can explain its moves and plans while playing at the level of a typical Grandmaster. Our novel framework enables domain-specific reasoning through complementary components: an encoder-decoder architecture and an iterative distillation algorithm. This architecture integrates a silent expert chess encoder with an instruction-tuned LM through cross-attention, which we train via a question-answering curriculum to extract chess concepts from the encoder's representations. Building on this domain-adapted model, we iteratively improve its explanations with a natural-language analog of the Bellman update: the model analyzes the positions after its top candidate moves and consolidates them into an explanation of the current position, which is then distilled back into the model. Over seven iterations, our model gains over 900 Elo points (1782 to 2697), substantially surpassing all frontier models on both playing strength and puzzle accuracy, despite containing three orders of magnitude fewer parameters. Furthermore, LM-based evaluations show that our explanations are fluent and approach GPT-5.6-Sol (high) in coherence. The generality of our architecture and training procedure suggests a recipe for applying language models to domains where silent expert encoders are available, like games, robotics, and computer use.
Primary: Princeton University
All Institutions: Princeton University
Queen introduces a 4B-parameter chess-language model that achieves Grandmaster-level play (2697 Elo) and high-quality explanations by integrating a silent expert encoder with an LM via cross-attention and iterative search distillation. This paper makes a significant contribution to the field by demonstrating that hybrid architectures combining specialized encoders with general LMs can outperform frontier LMs in domain-specific reasoning tasks, offering a scalable and efficient alternative to pure LLM approaches for tasks requiring deep domain expertise.
The paper proposes a novel hybrid architecture, "Queen," that integrates a silent expert chess encoder (Leela/BT5) with a general-purpose language model decoder (SmolLM3-3B) via a Flamingo-inspired gated cross-attention bridge. The methodology is rigorous, featuring a two-stage training process: (1) Domain Adaptation using a curated QA curriculum to teach the LM to interpret the encoder's latent representations, and (2) Iterative Search Distillation, a self-improvement loop inspired by Bellman updates where the model analyzes child positions and consolidates explanations, which are then distilled back into the model. This approach effectively bridges the gap between strong silent experts and fluent but weak LMs.
The experimental evaluation is comprehensive and compelling. Queen achieves a 2697 Elo rating, surpassing frontier LMs like GPT-5.6-Sol (2071) and Gemini-3.1-Pro (2201) by a significant margin while using 3 orders of magnitude fewer parameters. The paper introduces a robust evaluation framework covering accuracy (Elo), substantiation (no-mistake rate on puzzles), and coherence (LM-judged). Ablations clearly demonstrate the necessity of both the encoder and the curriculum. The comparison against a "HCE" variant shows that the method is not solely reliant on frontier model distillation for its core strength.
High. The authors provide code, model weights, and detailed appendices including hyperparameters, data construction pipelines, and prompt templates. The use of open-source components (SmolLM3, Leela) and public datasets (Lichess) further enhances reproducibility.
The primary limitation is the low "conceptual coherence" score (2.76/5), indicating that while the model plays well and structures its analysis correctly, it still hallucinates chess motifs or patterns. The method is currently specific to chess; while the authors argue for generality, the reliance on a specific "silent expert" encoder and the specific recursive distillation logic may not transfer trivially to other domains without significant adaptation.
This work offers a general recipe for coupling LMs with domain-specific expert encoders, with potential applications in robotics, computer use, and other games. It demonstrates that small, specialized models can outperform massive general-purpose LMs in specific reasoning tasks when properly grounded in expert knowledge. Queen introduces a 4B-parameter chess-language model that achieves Grandmaster-level play (2697 Elo) and high-quality explanations by integrating a silent expert encoder with an LM via cross-attention and iterative search distillation. This paper makes a significant contribution to the field by demonstrating that hybrid architectures combining specialized encoders with general LMs can outperform frontier LMs in domain-specific reasoning tasks, offering a scalable and efficient alternative to pure LLM approaches for tasks requiring deep domain expertise.
Neural audio codecs compress waveforms into compact discrete tokens that underpin speech language models, real-time communication, and large-scale audio storage. Almost every dominant design, including residual vector quantization, finite scalar quantization, and single-codebook variants, follows the VQ-VAE template by partitioning the encoder latent through a learned codebook or a fixed scalar grid. We ask whether this partition is necessary. We introduce GS-Codec, a neural speech codec whose bottleneck is a parametric signal decomposition rather than a quantizer. We adapt Gaussian splatting from 3D scene reconstruction to one-dimensional latents. An inner optimization loop fits each encoder segment as a weighted sum of 1D Gaussian primitives. The decoder then reconstructs the waveform from the rendered sum. To avoid the cost of this iterative inner loop at inference time, we additionally train a lightweight GS Predictor Net that regresses the primitive parameters in a single forward pass. The encoder and decoder are trained end-to-end through the inner loop, with no quantizer anywhere in the training pipeline: the bottleneck is the decomposition itself, and scalar quantization is applied only post-training to the fitted parameters. Rather than relying on discrete codebook stages for bitrate control, our representation exposes a fine-grained rate-quality tradeoff: a single trained checkpoint supports post-training bitrate control by varying the number of primitives and the per-parameter bit depth, with no retraining required. GS-Codec matches or exceeds well-established open-source codecs such as EnCodec and DAC on speaker similarity (SIM), intelligibility (STOI), and perceptual quality (UTMOS) at comparable bitrates, while achieving comparable semantic performance (WER). Code and audio samples are available at https://ronaluf.github.io/gs-codec/
Primary: Ben-Gurion University of the Negev
All Institutions: Ben-Gurion University of the Negev
GS-Codec introduces a Gaussian Splatting bottleneck for neural audio coding, replacing standard quantizers with a parametric decomposition that enables fine-grained post-training bitrate control. The paper demonstrates competitive performance against state-of-the-art codecs like EnCodec and DAC, offering a novel architectural alternative that leverages differentiable rendering techniques for signal compression, though it faces challenges in encoding latency compared to feed-forward baselines.
The paper proposes a novel architectural shift in neural audio codecs by replacing the standard Residual Vector Quantization (RVQ) or Finite Scalar Quantization (FSQ) bottleneck with a parametric 1D Gaussian Splatting decomposition. The core idea is to fit the encoder's latent representation as a weighted sum of Gaussian primitives via an inner optimization loop during training. To address the computational cost of this iterative fitting at inference, the authors introduce a "GS Predictor Net" that amortizes the optimization into a single forward pass. The method is technically sound, leveraging differentiable rendering concepts from 3D graphics (Gaussian Splatting) and applying them to 1D signal processing. The use of post-training scalar quantization on the fitted parameters rather than learned codebooks is a distinct design choice that enables fine-grained bitrate control without retraining.
The experimental evaluation is rigorous and comprehensive. The authors compare GS-Codec against strong, well-established baselines such as EnCodec, DAC, and WavTokenizer on standard datasets (LibriTTS, LJSpeech, LibriSpeech). The results show that GS-Codec matches or exceeds these baselines on key metrics like UTMOS (perceptual quality), STOI (intelligibility), and SIM (speaker similarity) at comparable bitrates. The inclusion of human listening tests (MUSHRA/MOS) adds significant credibility to the perceptual quality claims. The ablation studies on primitive count and bit depth provide clear insights into the rate-quality tradeoff. The encoding time analysis honestly reports the latency trade-off, showing that while the iterative version is slow, the Predictor Net brings it close to competitive levels, though still slightly slower than feed-forward baselines.
The paper provides high reproducibility. It details the SEANet backbone hyperparameters, the specific Gaussian Splatting configuration (number of primitives, inner loop steps, learning rates), and the training schedule. The code and audio samples are available via the provided URL. The use of standard metrics and open-source baselines facilitates easy comparison. The detailed appendix on quantization ranges and loss weights further supports reproducibility.
The primary limitation is the encoding latency. Even with the Predictor Net, the encoding time is higher than standard feed-forward codecs like EnCodec and DAC, which may be a bottleneck for real-time applications. The paper also notes that performance at very low bitrates (<3 kbps) is not as strong as specialized low-rate codecs. The method is currently validated primarily on English speech, and generalization to other languages or non-speech audio (music, sound effects) is not extensively explored.
This work opens a new direction in neural audio compression by demonstrating that parametric signal decompositions can serve as effective bottlenecks, challenging the dominance of codebook-based quantization. The ability to control bitrate post-training by varying the number of primitives and bit depth is a practical advantage for deployment. The cross-domain insight from 3D Gaussian Splatting to 1D audio signals is interesting and may inspire further research into using geometric or parametric priors in other signal processing tasks. GS-Codec introduces a Gaussian Splatting bottleneck for neural audio coding, replacing standard quantizers with a parametric decomposition that enables fine-grained post-training bitrate control. The paper demonstrates competitive performance against state-of-the-art codecs like EnCodec and DAC, offering a novel architectural alternative that leverages differentiable rendering techniques for signal compression, though it faces challenges in encoding latency compared to feed-forward baselines.
In this paper, we study robust decision-making in the latent space of world models (WMs). Robust optimization is a mathematical framework where, given explicitly specified dynamics and physically meaningful disturbances, a robot can select actions that remain effective even under worst-case disturbances. However, applying this principle to the learned latent space of WMs introduces a fundamental challenge: because WMs have fully learned state spaces and dynamics inferred from high-dimensional observations, it is unclear how to define latent-space disturbances that faithfully represent uncertainty in the underlying system. Our key idea is to model a latent-space disturbance as a perturbation to the learned latent dynamics that induces pessimistic but plausible transitions. Specifically, we construct a set of plausible latent dynamics by combining a dynamics-aware similarity metric that captures plausible transitions with out-of-distribution detection that excludes implausible latent states. We calibrate this uncertainty set over latent dynamics using conformal prediction, ensuring that WM imaginations induced by the latent disturbance remain plausible without becoming overly pessimistic. We then jointly optimize robust robot actions and the worst-case latent disturbances through game-theoretic optimization. We leverage this latent-space robust optimization to robustify policy steering, considering two paradigms: latent safety filtering and sample-and-verify steering of a generative control policy. Our controlled simulation experiments show that our latent disturbance enables robust decision-making directly in WM latent spaces, and hardware experiments with a Franka manipulator show that modeling latent disturbances enables robust policy steering, reducing failures by 70% in safety filtering and 54% in sampling-based policy steering. Project website: https://junwon.me/LatentDisturbance/.
Primary: Carnegie Mellon University
All Institutions: Carnegie Mellon University
The paper introduces a principled framework for modeling latent disturbances in world models using conformal prediction and game-theoretic optimization, enabling robust policy steering with significant reductions in failure rates on hardware. The combination of OOD detection, dynamics-aware similarity, and statistical calibration offers a novel and rigorous approach to uncertainty in learned latent spaces, though the computational cost and limited hardware evaluation scope are notable constraints.
The paper addresses a critical gap in world model (WM) based control: the lack of a principled way to define "disturbances" in the learned latent space. The authors propose modeling latent disturbances as perturbations to the learned dynamics that induce pessimistic yet plausible transitions. The core technical contribution is the construction of an uncertainty set over latent dynamics using a combination of a dynamics-aware similarity metric and out-of-distribution (OOD) detection. Crucially, they calibrate this set using conformal prediction, which provides statistical guarantees on the coverage of plausible transitions. This is a sophisticated approach that bridges robust optimization and distribution-free uncertainty quantification. The method is then integrated into a game-theoretic optimization framework to jointly optimize robust actions and worst-case disturbances. The application to policy steering (safety filtering and sample-and-verify) is well-motivated and directly addresses the fragility of generative policies in WMs.
The evaluation includes controlled simulation experiments and hardware experiments with a Franka manipulator. The reported results are strong, claiming a 70% reduction in failures for safety filtering and 54% for sampling-based steering. These are significant improvements for a robotics task. However, the provided text is truncated and lacks specific details on the simulation environments, the baselines compared against (e.g., standard MPC, other robust control methods, or non-robust WM policies), and the specific metrics used beyond "failures." The hardware results are promising but limited to a single robot platform and likely a specific set of tasks.
The paper provides a project website, which likely contains code and additional details. The use of conformal prediction and standard robust optimization techniques suggests that the method is implementable by other researchers with expertise in these areas. However, without access to the full code and hyperparameters, exact reproduction is difficult. The reliance on a learned WM means that results may vary depending on the quality of the underlying WM, which is a known challenge in the field.
The primary limitation is the computational cost of the game-theoretic optimization, which may be prohibitive for real-time control on high-dimensional systems. The method assumes that the WM's latent space is sufficiently expressive to capture the relevant dynamics and uncertainties. If the WM is poorly trained or the latent space is not well-structured, the OOD detection and similarity metrics may not perform well. The hardware experiments are limited to a single robot and a small number of tasks, so the generalizability of the results to other platforms and tasks is not fully established.
This work has significant potential impact on the field of world model-based control. By providing a principled way to handle uncertainty in the latent space, it enables more robust and safe deployment of learned policies. The use of conformal prediction for calibration is a novel and rigorous approach that could be adopted in other areas of ML where uncertainty quantification is crucial. The method could be extended to other types of world models and control tasks, potentially leading to more reliable autonomous systems. The paper introduces a principled framework for modeling latent disturbances in world models using conformal prediction and game-theoretic optimization, enabling robust policy steering with significant reductions in failure rates on hardware. The combination of OOD detection, dynamics-aware similarity, and statistical calibration offers a novel and rigorous approach to uncertainty in learned latent spaces, though the computational cost and limited hardware evaluation scope are notable constraints.
World-action models (WAMs) couple future prediction with robot control, yet existing systems often vary the video backbone, interaction structure, supervision, and inference procedure simultaneously, making their design choices difficult to compare. We introduce OPENWAM, an open world-action modeling framework built around a common causal robot-video foundation and configurable video-action interaction. Starting from Wan2.2-5B, we perform causal robot-video pretraining on over 10,000 hours of video, then integrate an action expert through a shared Mixture-of-Transformers architecture that supports joint, video-then-action, action-then-video, and decoupled generation. OPENWAM achieves high success rates on four LIBERO suites and real-world bimanual tasks; robot-video training with causal adaptation improves VTA success on LIBERO-Long from 68.4% to 97.8%. The same configurable architecture naturally extends to inverse and forward dynamics, allowing us to study how counterfactual transitions improve independently trained dynamics models beyond demonstrations alone. When only the video predictor is adapted to a new task, a frozen local-context inverse dynamics model trained on counterfactual data and demonstrations achieves 84.0% mean success across four held-out LIBERO-90 tasks, compared with 47.0% for a full-context inverse model and 21.5% for a local-context model trained only on demonstrations. For forward dynamics, counterfactual supervision reduces RGB prediction error by 34.5% and raises outcome identification from 21.1% to 71.3% among 16 same-state outcomes. OPENWAM provides a common testbed for comparing WAM interaction designs and for studying dynamics learning from video data beyond successful demonstrations.
Primary: Stanford University
All Institutions: Stanford University
The paper introduces OpenWAM, an open framework for composable world-action models that couples video prediction with robot control. It demonstrates significant improvements in robotic success rates through causal video pretraining and flexible action integration, providing a rigorous testbed for studying dynamics learning beyond standard demonstrations.
The paper proposes OpenWAM, a framework built on the Wan2.2-5B video generation model. The core methodological contribution is the adaptation of a large video foundation model for robot control via causal pretraining on 10,000+ hours of robot video. It introduces a Mixture-of-Transformers (MoT) architecture to integrate an action expert, allowing for flexible interaction patterns (joint, sequential, decoupled). The approach to studying dynamics learning via counterfactual transitions is a strong methodological addition, distinguishing it from standard imitation learning.
The experiments are extensive, covering four LIBERO suites and real-world bimanual tasks. The reported improvements are substantial (e.g., 68.4% to 97.8% on LIBERO-Long), and the ablation studies on inverse/forward dynamics provide deep insight into the model's capabilities. The use of held-out tasks for the dynamics evaluation is rigorous.
The paper claims to be an "open framework," which suggests code and weights will be released. However, the specific hyperparameters and the exact composition of the 10,000-hour dataset are likely detailed in the supplementary material (which is not fully visible in the prompt but implied by the "open" nature). The reliance on a specific large backbone (Wan2.2-5B) limits immediate reproducibility for labs without access to such resources.
The primary limitation is the heavy computational cost associated with fine-tuning a 5B parameter video model. Additionally, the framework's performance may be heavily dependent on the quality and scale of the pre-training video data, which may not be accessible to all researchers.
This work bridges the gap between large-scale video generation and robotic control, potentially enabling more generalizable robot policies. The open-source nature of the framework could accelerate research in world-action models by providing a standardized baseline. The paper introduces OpenWAM, an open framework for composable world-action models that couples video prediction with robot control. It demonstrates significant improvements in robotic success rates through causal video pretraining and flexible action integration, providing a rigorous testbed for studying dynamics learning beyond standard demonstrations.
World action models jointly learn visual predictionand robot actions, providing a way to use observations ofscene evolution for policy learning. Their video and actionlosses, however, provide no explicit target for the geometricconsequences of a demonstrated action sequence. Moreover,visual features taken after temporal attention can contain futureobservations, making them unsuitable as the sole current visualinput to an auxiliary predictor. We introduce ACG-WAMand its auxiliary objective, the Action-Conditioned GeometricJoint-Embedding Predictive Architecture (ACG-JEPA), whichpredicts geometric features at several horizons from the currentobservation and intervening actions, using the future slot of afrozen VGGT encoding of each current and future image pairas the target. We apply this supervision from the head and wristcameras to a shared visual embedding before temporal mixing,and remove the teacher and auxiliary modules at inference.On 50 RoboTwin 2.0 tasks, ACG-WAM achieves 93.46%success in clean scenes, with the best randomized success(92.68%) and mean across both settings (93.07%) among thecompared methods; across three tasks on a real robot, itachieves 85.00% success and 91.67% partial completion score,exceeding Motus by 10.00 and 9.17 percentage points, respec-tively. Code:https://github.com/RoboOpus/ACG-WAM.Website:https://RoboOpus.github.io/ACG-WAM.
Primary: Beijing Institute of Technology
All Institutions: Beijing Institute of Technology, LimX Dynamics
ACG-WAM introduces an action-conditioned geometric prediction objective to World Action Models, significantly improving bimanual manipulation performance in simulation and on real robots. The paper presents a rigorous technical contribution that addresses specific challenges in joint video-action learning, such as information leakage and the lack of explicit geometric targets, resulting in state-of-the-art results on the RoboTwin 2.0 benchmark and real-world tasks.
The paper proposes ACG-WAM, a World Action Model (WAM) that augments the Motus backbone with an auxiliary geometric prediction objective. The core innovation is the Action-Conditioned Geometric Joint-Embedding Predictive Architecture (ACG-JEPA). This module uses a frozen VGGT teacher to encode current and future image pairs, extracting the "future slot" features as targets. The student predictor takes current visual features (crucially, extracted *before* temporal attention to avoid information leakage from future frames) and the intervening action sequence to predict these geometric targets. This design explicitly supervises the geometric consequences of actions, addressing a gap in standard video/action loss functions. The method is well-motivated by the need for spatial reasoning in bimanual manipulation and the technical challenge of preventing future information leakage in auxiliary predictors.
The evaluation is robust, covering 50 tasks on the RoboTwin 2.0 benchmark (clean and randomized settings) and 3 real-world tasks on a TRON2 platform. ACG-WAM achieves state-of-the-art results, outperforming strong baselines like Motus, LingBot-VA, and MECo-WAM. Specifically, it achieves 93.46% success in clean scenes and 92.68% in randomized scenes on RoboTwin 2.0, and 85.00% success on the real robot. Ablation studies convincingly demonstrate the contribution of action conditioning, multi-horizon prediction, and the joint-encoding target design versus simple endpoint subtraction. The real-world results are particularly strong, showing significant gains over the base Motus model.
The paper provides high reproducibility. Code is released on GitHub. Implementation details are thorough, specifying the backbone (Wan2.2-TI2V-5B), teacher model (VGGT), training hyperparameters (learning rates, batch sizes, loss weights), and inference settings. The use of standard benchmarks (RoboTwin 2.0) and clear evaluation metrics (Success Rate, Partial Completion Score) further supports reproducibility.
The method relies on a large, frozen teacher model (VGGT) and a complex backbone (Motus/Wan2.2), which may limit accessibility for groups with lower computational resources. The real-world evaluation is limited to only three tasks, which is a small sample size for generalizing claims about physical robustness. Additionally, the paper acknowledges the use of ChatGPT for drafting, which, while disclosed, is a minor point regarding the originality of the text rather than the science.
This work contributes to the growing field of World Action Models for robotics. By explicitly incorporating geometric supervision, it offers a pathway to improve spatial reasoning in manipulation policies. The technique of supervising features before temporal mixing to avoid leakage is a useful insight for other architectures involving joint video-action prediction. The strong performance on both simulation and real hardware suggests practical utility for industrial and research robotics applications. ACG-WAM introduces an action-conditioned geometric prediction objective to World Action Models, significantly improving bimanual manipulation performance in simulation and on real robots. The paper presents a rigorous technical contribution that addresses specific challenges in joint video-action learning, such as information leakage and the lack of explicit geometric targets, resulting in state-of-the-art results on the RoboTwin 2.0 benchmark and real-world tasks.
Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it. The resulting images are processed foveally by allocating more visual tokens to the image centers, focusing computation on task-relevant features. Such Active Visual Fixation (AVF) requires carefully coordinated gaze during task execution, which we accomplish hierarchically by first training a low-level gaze servoing policy conditioned on a goal object, then training a target selector which emits fixation goals based on task progress. Both modules are trained with RL on real-world data: the first is trained with a dense geometric reward and the second co-trains with the BC gripper policy which allows it to discover fixation sequences that can resemble a human's fixation sequence while performing the task. EyeRobot 2.0 further takes advantage of fixation by canonicalizing gripper information into a fixation-relative SE(3) frame, which compacts the size of the action distribution to learn. We collect teleoperation data for 7 real-world and 6 simulated tasks, and conduct over 1000 physical and 1800 simulated robot trials comparing EyeRobot 2.0 against passive stereo and ego + wrist camera policies trained on the same data. Removing wrist cameras is costly for standard policies: with only passive stereo, real-world success drops from 52% to 27%. EyeRobot 2.0 closes this gap with only stereo, outperforming passive stereo by 40% in real and 20% in sim. It matches ego + wrist policies when their wrist views are clear (69% vs. 64%), and more than doubles their success when grasped objects occlude the wrist cameras (48% vs. 22%)
Primary: Stanford University
All Institutions: Stanford University, UC Berkeley
EyeRobot 2.0 introduces a biologically inspired active gaze framework that enables precise bimanual manipulation using only a fixed stereo camera, outperforming standard wrist-camera baselines in occlusion-heavy tasks. The paper demonstrates that physically attending to 3D fixation points and canonicalizing actions in a fixation-relative frame significantly enhances policy learning, offering a robust alternative to the ubiquitous wrist-camera setup in robotic manipulation.
The paper proposes a hierarchical framework for active visual fixation in bimanual manipulation. The core innovation is the decoupling of low-level gaze servoing (trained with RL using geometric rewards) from high-level target selection (trained with RL using a BC-RL loop that optimizes for downstream gripper policy accuracy). The use of a fixation-relative SE(3) frame for action canonicalization is a strong technical contribution that reduces the complexity of the action space. The foveated processing of stereo images is well-motivated, though the specific implementation details of the "foveated transformer-decoder" are somewhat sparse in the main text, relying on the appendix. The method effectively addresses the occlusion issues inherent in wrist-mounted cameras by physically moving the viewpoint.
The experimental setup is rigorous, involving 7 real-world and 6 simulated tasks with over 1000 physical trials. The comparison against passive stereo and ego+wrist baselines is fair, as all policies are trained on the same data. The results are compelling: EyeRobot 2.0 significantly outperforms passive stereo and matches or exceeds ego+wrist performance, particularly in occlusion scenarios where wrist cameras fail. The ablation studies convincingly demonstrate the contribution of foveation, fixation-centric actions, and stereo depth. The inclusion of simulation experiments provides a controlled environment for analysis, although the primary claim is validated on real hardware.
The paper states that simulation code will be made public, which is a positive step. However, the reliance on specific hardware (I2RT YAM manipulator, custom leader arm) and specific software stacks (MuJoCo, GELLO, TRLC) may limit immediate reproducibility for labs without similar setups. The detailed description of the RL training loops and reward functions provides a clear path for implementation, but the lack of a public code release at the time of review (arXiv preprint) is a minor drawback.
The system is currently task-specific, requiring separate training for each task. The paper acknowledges that extending this to multi-task learning would require more complex prompt generation (e.g., VLMs). The framework assumes a fixed head position and only controls eye movements, which limits its applicability to mobile manipulation where head/neck movement is crucial. The computational cost of running two RL policies and a BC policy in real-time is not deeply analyzed, though the 30Hz gaze rate suggests it is feasible.
This work has significant implications for the design of robotic manipulators, potentially eliminating the need for wrist cameras and allowing for sleeker, more robust gripper designs. It also opens a pathway for leveraging egocentric human data for robot learning by mimicking human fixation patterns. The approach could be extended to other domains requiring precise visual attention, such as surgical robotics or inspection tasks. EyeRobot 2.0 introduces a biologically inspired active gaze framework that enables precise bimanual manipulation using only a fixed stereo camera, outperforming standard wrist-camera baselines in occlusion-heavy tasks. The paper demonstrates that physically attending to 3D fixation points and canonicalizing actions in a fixation-relative frame significantly enhances policy learning, offering a robust alternative to the ubiquitous wrist-camera setup in robotic manipulation.
Equipping artificial agents with spatial intelligence requires a comprehensive generative prior over the dynamic 3D world. We propose World Motion Models (WMMs) that capture "what was, is, and will be where across time" via sparse SE(3) pose trajectories. WMMs are built on the observation that elements of dynamic scenes can be well approximated by a set of rigid SE(3) trajectories, a minimal yet expressive primitive for 4D modeling. This representation unifies articulated objects, human bodies, hand-object interactions, piecewise-rigid scene dynamics, camera motion, and even robot states and actions into a single shared space. Given this representation, we cast the joint distribution of these entities as a flexible sequence modeling problem, utilizing flow-matching with per-token noise levels. Coupled with a context token mechanism for non-sequential conditioning, this formulation supports any-to-any marginal conditioning across an arbitrary number of entities and time steps. Tasks such as future prediction, motion infilling, model-predictive control, inverse kinematics, cross-embodiment retargeting, and policy learning all reduce to the application of different masks over the same network. Experiments on 6 diverse applications of 3D vision and robotics demonstrate the versatility and flexibility of WMMs with strong performance.
Primary: UC Berkeley
All Institutions: UC Berkeley, Meta AI
World Motion Models introduces a unified flow-matching framework for SE(3) trajectory modeling that enables flexible any-to-any conditioning across diverse robotics and vision tasks. The paper demonstrates that a single generative model can handle prediction, control, and retargeting through simple masking, offering a powerful and versatile primitive for spatial intelligence in artificial agents.
The paper proposes World Motion Models (WMMs), a unified framework for modeling 4D dynamics using sparse SE(3) pose trajectories. The core methodological contribution is the recasting of joint distribution modeling over multiple entities (humans, objects, cameras, robots) as a flexible sequence modeling problem using flow-matching. By employing per-token noise levels and a context token mechanism, the authors enable "any-to-any" marginal conditioning. This allows a single network to handle diverse tasks—such as prediction, infilling, and control—simply by applying different masks to the input sequence. The choice of SE(3) trajectories as a primitive is elegant, providing a minimal yet expressive representation that unifies articulated motion, rigid body dynamics, and camera motion into a shared latent space.
The paper reports experiments on six diverse applications spanning 3D vision and robotics, including future prediction, motion infilling, model-predictive control (MPC), inverse kinematics (IK), cross-embodiment retargeting, and policy learning. The breadth of tasks demonstrates the versatility of the unified approach. While the abstract claims "strong performance," the lack of specific quantitative metrics in the provided text prevents a full assessment of superiority over specialized baselines. However, the ability to switch tasks via masking without retraining is a significant empirical advantage over task-specific architectures.
The project page is provided, which likely contains code and datasets. The use of standard flow-matching techniques and SE(3) representations suggests that the core components are reproducible, provided the specific masking strategies and context token implementations are detailed in the full paper (which is not fully visible here, but implied by the "Spotlight" status and project page).
The reliance on SE(3) trajectories assumes that scene elements can be well-approximated by rigid motions. This may limit applicability to highly deformable objects or non-rigid interactions that cannot be decomposed into rigid parts. Additionally, the computational cost of flow-matching on long sequences of high-dimensional SE(3) poses could be a bottleneck for real-time applications, although the paper mentions MPC, suggesting some efficiency gains.
This work has the potential to significantly impact the field of embodied AI and 4D scene understanding. By unifying various robotics and vision tasks under a single generative framework, it simplifies the development of general-purpose agents. The "any-to-any" conditioning capability is particularly valuable for data-efficient learning and sim-to-real transfer, where conditioning on partial observations or specific actions is common. World Motion Models introduces a unified flow-matching framework for SE(3) trajectory modeling that enables flexible any-to-any conditioning across diverse robotics and vision tasks. The paper demonstrates that a single generative model can handle prediction, control, and retargeting through simple masking, offering a powerful and versatile primitive for spatial intelligence in artificial agents.
Agentic reinforcement learning (RL) disaggregates training from rollout, so each policy update must reach the rollout clusters before the next batch. Transferring a full 1T checkpoint for such weight synchronization (refit) takes 87.5 min between two AWS regions. Measurements of BF16 training show that about 1% of weights change their stored values per step. Recent systems exploit this sparsity but fall short on placement, exactness, or efficiency: they reimplement placement rules, assemble full tensors, rebuild values arithmetically, or use a cross-cluster collective, and none fully recovers from mid-refit failures. We present NeMo-DCR (Delta-Compressed Refit), which sends only changes yet is bit-exact: receivers obtain the same parameter and buffer bits as a dense refit. For placement, fixed affine mappings project changes from training shards into the checkpoint's canonical coordinates, residual conversion covers the other changes, and the serving runtime's native loader places all changes in receiver storage. For exactness, compressible XOR masks carry affine changes whose projection and loader preserve stored bits, and overwrites carry the others. Receivers apply both in place, retries overwrite partial writes, and a joint commit binds the policy to the baseline for the next delta. For efficiency, object storage or a relay tree streams payloads during delta construction, without a cross-cluster collective. Even at 3% and 5% change rates, NeMo-DCR refits of 30B-1T models are 12-40$\times$ faster than a transport-only full-checkpoint reference. A 1T relay-tree refit at 3% takes 150 s instead of 87.5 min, making refits practical for cross-cluster agentic RL at trillion-parameter scale.
Primary: NVIDIA
All Institutions: NVIDIA
NeMo-DCR introduces a bit-exact delta-compressed refit system that achieves 12-40x speedups in weight synchronization for trillion-parameter agentic RL by leveraging canonical coordinates, mixed XOR/overwrite encoding, and pipelined delivery without cross-cluster collectives. The paper presents a rigorous systems contribution with strong empirical validation, addressing critical gaps in placement, exactness, and recovery for sparse weight updates.
The paper proposes NeMo-DCR, a system for bit-exact delta-compressed refitting of large language models in agentic RL settings. The core innovation lies in decoupling the placement logic from the delta computation by using "canonical coordinates" (Hugging Face checkpoint format) as an intermediate representation. This allows the system to project changes from training shards directly into canonical space using fixed affine mappings, avoiding the need to assemble full tensors for the majority of updates. The use of a mixed XOR/overwrite encoding is clever: XOR is used for representation-preserving affine changes to exploit sparsity and compression, while overwrites are used for residual changes (like those involving shared scale factors) to ensure bit-exactness. The recovery mechanism, which uses absolute overwrites for retries and a joint commit protocol, is robust and addresses a critical gap in existing sparse synchronization systems. The theoretical guarantee of dense-refit equivalence is well-supported by the stated assumptions and the detailed proof sketch.
The evaluation is rigorous and highly relevant to the target audience (systems researchers and large-scale RL practitioners). The authors test models ranging from 30B to 1T parameters, which is impressive. The comparison against a transport-only full-checkpoint reference is fair and highlights the 12-40x speedup. The bit-exactness verification is thorough, confirming that the delta refit produces identical bits to a dense refit. The training stability experiment, where receivers are killed mid-refit and the system recovers, is a strong validation of the recovery mechanism. The latency breakdown showing that transport dominates over construction is insightful for system design.
The paper provides significant implementation details, including the number of lines of code (~7K Python), the specific libraries used (Megatron Bridge, vLLM), and the testbed configuration (GB300/H100 GPUs, AWS regions). The appendices contain detailed specifications of the protocol, mapping classes, and payload formats. However, as is common with systems papers from industry labs, the full source code is not explicitly linked in the provided text, which may hinder independent reproduction. The reliance on specific proprietary hardware (GB300) and cloud infrastructure (AWS) also limits reproducibility for academic labs.
The system is tightly coupled to specific training (Megatron) and serving (vLLM) stacks, which may limit its immediate applicability to other frameworks. The assumption that affine mappings cover >96% of changes is model-dependent; models with highly non-standard weight layouts or frequent structural changes might see a higher residual ratio, increasing the overhead of residual conversion. The evaluation focuses on BF16; the behavior with lower precision formats (FP8, INT4) or different quantization schemes is not fully explored. The cross-cluster latency is dominated by network transport, which suggests that in data-center-local deployments, the relative speedup might be different.
This work has significant implications for the scalability of agentic RL and other distributed training paradigms that require frequent weight synchronization. By making refits 12-40x faster, it enables more frequent policy updates, which can improve the efficiency and performance of RL training loops. The bit-exactness guarantee is crucial for maintaining training stability, addressing a known pain point in distributed RL. The techniques could potentially be adapted to other domains requiring efficient state synchronization, such as federated learning or multi-agent systems. NeMo-DCR introduces a bit-exact delta-compressed refit system that achieves 12-40x speedups in weight synchronization for trillion-parameter agentic RL by leveraging canonical coordinates, mixed XOR/overwrite encoding, and pipelined delivery without cross-cluster collectives. The paper presents a rigorous systems contribution with strong empirical validation, addressing critical gaps in placement, exactness, and recovery for sparse weight updates.
New AI accelerators arrive before the kernels that make them fast, because peak kernel performance requires architecture-specific expertise in operand pipelines and data-movement techniques. Coding agents can now write, compile, and tune kernels on their own, so they could greatly accelerate kernel development and optimization. What agents produce depends on the interfaces they are given. These interfaces may expose a machine's raw capabilities or encode the known-good methods for exploiting its hardware features efficiently. We test the performance impact of these two interfaces on Intel GPUs through controlled experiments with three coding models on three kernels. With raw SYCL and execution feedback, Opus 4.8, the strongest of the three models tested, writes a single-GPU GEMM kernel reaching only 57.5% of Intel's tuned oneDNN library. We present SyclKittens, a hardware-aware tile programming model for Intel GPUs that encodes these known-good methods, so its operations place operands in matrix-engine layouts, prefetch through the L1 cache, and communicate over the fabric that links GPU stacks into a node. With SyclKittens under the same agent, task, and feedback budget, the agent reaches 82.1%, showing that encoded methods turn hardware capabilities into performance. Workload-specific policies stay programmable, and engineers and agents jointly refine these schedules in SyclKittens to build a kernel suite that reaches ~96% of oneDNN in geometric mean across GEMM shapes. The suite runs Llama-3.1-8B inference 1.59x faster than torch.compile on one GPU and up to 2.91x faster than a matched multi-GPU decode path on Intel's oneCCL. SyclKittens is open source and available at https://github.com/intel/SyclKittens.
Primary: Intel Corporation
All Institutions: Massachusetts Institute of Technology, Intel Corporation, Stanford University, California Institute of Technology
SyclKittens introduces a hardware-aware tile programming model that enables coding agents to generate high-performance GPU kernels by encoding known-good hardware methods. The paper demonstrates that this approach significantly outperforms raw SYCL interfaces in agent-generated kernel performance, achieving near-library-level efficiency for critical AI workloads on Intel GPUs, thereby offering a scalable path for automated kernel optimization.
The paper introduces SyclKittens, a tile-based programming model for Intel GPUs that abstracts low-level hardware details (operand pipelines, data movement, matrix engine layouts) into high-level, hardware-aware operations. The core methodological contribution is the demonstration that providing coding agents with a structured, domain-specific interface (SyclKittens) rather than raw SYCL primitives significantly improves the performance of generated kernels. The authors employ a controlled experimental design comparing raw SYCL with execution feedback against SyclKittens under identical agent, task, and feedback budgets. This approach effectively isolates the impact of the programming interface on agent performance, providing a rigorous evaluation of how "known-good" hardware methods encoded in a DSL can guide LLMs to produce efficient code.
The experiments are conducted on Intel Max GPUs, focusing on critical AI workloads such as GEMM, attention, and normalization. The results show a substantial improvement: an agent using raw SYCL reaches only 57.5% of the performance of Intel's tuned oneDNN library, whereas the same agent using SyclKittens reaches 82.1%. Furthermore, a co-designed kernel suite using SyclKittens achieves ~96% of oneDNN performance in geometric mean across GEMM shapes. End-to-end inference benchmarks for Llama-3.1-8B demonstrate 1.59x speedup over torch.compile on a single GPU and up to 2.91x speedup over a matched multi-GPU decode path using Intel's oneCCL. The evaluation is strong in its direct comparison of interfaces and its end-to-end application metrics.
The paper is highly reproducible given that SyclKittens is open-sourced on GitHub. The authors provide detailed appendices separating the controlled-agent studies from the co-designed suite, including specific measurement protocols, warmup procedures, and aggregation methods. The use of specific coding models (Opus 4.8, etc.) and defined feedback budgets allows for precise replication of the agent experiments. However, the specific prompts and agent configurations may require careful extraction from the appendices to fully replicate the agent's behavior.
The evaluation is limited to Intel GPUs, which may not generalize directly to other hardware architectures (e.g., NVIDIA, AMD) without significant adaptation. The performance gains are relative to Intel's own oneDNN library, so the absolute competitive standing against other state-of-the-art libraries on different hardware is not assessed. The reliance on specific coding models means the results may vary with future model improvements or different agent architectures. Additionally, the paper focuses on a limited set of kernels (GEMM, attention, norms), and the applicability to more complex or irregular workloads is not fully explored.
This work has significant implications for the intersection of AI and systems programming. By demonstrating that high-level, hardware-aware abstractions can enable coding agents to write near-optimal GPU kernels, it suggests a new paradigm for kernel development that reduces the need for deep, architecture-specific expertise. This could accelerate the adoption of new AI accelerators by lowering the barrier to entry for kernel optimization. The open-source release of SyclKittens provides a valuable tool for the community to explore similar approaches on Intel hardware and potentially adapt the concepts to other platforms. SyclKittens introduces a hardware-aware tile programming model that enables coding agents to generate high-performance GPU kernels by encoding known-good hardware methods. The paper demonstrates that this approach significantly outperforms raw SYCL interfaces in agent-generated kernel performance, achieving near-library-level efficiency for critical AI workloads on Intel GPUs, thereby offering a scalable path for automated kernel optimization.