Last 7 Days (September 27 – October 03, 2026)
Little is known about how to manually design agents capable of dexterous manipulation. Some design principles have been inferred from close examination of how animals manipulate objects, but these structures and behaviors have so far resisted biomimicry and may not be optimal for artificial machines. Here we evolve freeform robots to pick up, hold, rotate, and use diverse objects. Unlike other approaches to optimizing robot hands, we do not presuppose the presence, articulation, or geometry of any part of the body. Although familiar prehensile forms such as tails, beaks, paws and claws may emerge spontaneously under certain conditions--and while such conditions could be of interest to evolutionary biologists--de novo manipulator design can also reveal whole new solutions, overlooked or unknown structures which may be better suited for the task at hand. We use contrastive learning to create a highly searchable genetic embedding of design space, an autoregressive developmental model to decode designs, evolutionary strategies to find good designs, and reinforcement learning to train each evolved design. Winning designs were automatically converted into a manufacturable blueprint, printed, assembled and tested in the real world in a zero-shot manner. The results represent the state-of-the-art in evolutionary robotics in terms of performance, diversity and complexity.
Primary: University of Oxford
All Institutions: University of Oxford, DeepMind
The paper presents a state-of-the-art framework for evolving freeform dexterous robots from scratch, combining contrastive learning, developmental models, and reinforcement learning to produce manufacturable, real-world capable agents. By removing prior assumptions on robot morphology, the work reveals novel, non-biomimetic structures that outperform traditional designs in dexterity and diversity, marking a significant leap in evolutionary robotics and automated hardware design.
The paper proposes a comprehensive framework for de novo robot design, combining contrastive learning for genetic embedding, autoregressive developmental models for decoding, evolutionary strategies for search, and reinforcement learning for control. The novelty lies in the lack of prior assumptions about morphology (no fixed joints or geometry), allowing for truly freeform evolution. The use of a developmental model to map latent codes to physical structures is a sophisticated approach that bridges high-level design search with low-level manufacturability.
The evaluation is rigorous, moving beyond simulation to real-world validation. The robots were 3D printed, assembled, and tested in a zero-shot manner, demonstrating the robustness of the evolved designs. The paper reports state-of-the-art performance in terms of dexterity, diversity, and complexity compared to previous evolutionary robotics works. The inclusion of diverse object manipulation tasks (pick, hold, rotate, use) provides a strong benchmark for the method's generality.
The paper details the pipeline components (contrastive learning, autoregressive model, ES, RL) sufficiently for experts to replicate the framework. However, the specific hyperparameters for the evolutionary search and the exact architecture of the developmental model are likely in the appendix or code, which is standard. The zero-shot real-world testing is a high bar for reproducibility that the authors have met by providing the blueprint conversion process.
The computational cost of evolving and training multiple robots is significant. The "zero-shot" real-world performance, while impressive, may still be limited in speed or precision compared to hand-tuned industrial manipulators. The generalizability to tasks requiring extremely high force or precision (beyond what 3D printed materials can handle) is a potential constraint.
This work has significant implications for the field of evolutionary robotics and automated design. It suggests that optimal manipulator structures may not be biomimetic, challenging existing design paradigms. The ability to automatically generate manufacturable blueprints from abstract design spaces could accelerate the development of specialized robotic tools for manufacturing, surgery, or exploration. The paper presents a state-of-the-art framework for evolving freeform dexterous robots from scratch, combining contrastive learning, developmental models, and reinforcement learning to produce manufacturable, real-world capable agents. By removing prior assumptions on robot morphology, the work reveals novel, non-biomimetic structures that outperform traditional designs in dexterity and diversity, marking a significant leap in evolutionary robotics and automated hardware design.
Training models to act in accordance with an explicitly defined set of principles, or "constitution," has shown promise as a robust and transparent mechanism for AI alignment. However, the generality and flexibility of such methods remain unclear. Here, we show that constitution-consistent behavior can be distilled from synthetic corpora into lightweight objects (low-rank adapters and steering vectors). Despite never seeing a harmful request or jailbreak during training, such objects increase jailbreak defense success and measured alignment -- particularly at long context lengths and against multi-turn attacks, where they outperform both prompted and steered baselines. Subtracting control-trained from constitution-trained objects further accentuates these effects, yielding defenses we call "constitutional adapters" (CAs). CAs can be trained on a base model, transferred zero-shot to its post-trained checkpoint, and scaled at inference time to predictably trade off defense for benign compliance. Taken together, these results recommend CAs as a lightweight, portable, and tunable lever for mitigating misalignment and misuse in API deployments.
Primary: Anthropic
All Institutions: Anthropic
Constitutional Adapters provide a lightweight, tunable, and portable mechanism for enhancing LLM safety against misuse and misalignment by distilling constitutional principles into low-rank adapters and steering vectors. The paper demonstrates that these adapters, particularly when using control-subtracted task arithmetic, outperform traditional prompting and steering methods in long-context and multi-turn adversarial settings while maintaining benign compliance and general capabilities.
The paper proposes "Constitutional Adapters" (CAs), a method for distilling alignment principles into lightweight, inference-time interventions (LoRAs and ReLU steering vectors). The core methodological innovation is the use of a control-subtracted task arithmetic approach: training a lightweight object on a synthetic corpus derived from a model constitution ($D_C$) and subtracting an object trained on a register-matched control corpus ($D_K$). This subtraction aims to isolate the "constitutional" behavior from general fine-tuning artifacts like format or register shifts. The method leverages synthetic data generation via a strong model (Claude Opus 5) to create diverse scenarios that stress-test ethical principles, rather than relying on adversarial jailbreak examples. This allows the adapters to generalize to unseen attacks. The use of ReLU steering vectors, which allow for input-dependent steering strength, is a significant technical detail that improves upon fixed-direction steering.
The experimental evaluation is rigorous and comprehensive. The authors test across four distinct model families (Qwen3.5, Llama-3.1, Gemma-4, Nemotron 3) and various attack types (static, multi-query, multi-turn). Key findings include the superior robustness of CAs in long-context and multi-turn settings compared to system prompting, which degrades as context length increases. The paper also evaluates the trade-off between defense and benign compliance (over-refusal) by sweeping the steering strength, demonstrating a predictable and tunable frontier. The inclusion of misalignment evaluations (Petri Bloom) and capability benchmarks (MMLU, GSM8K, etc.) provides a holistic view of the method's impact, showing minimal capability degradation.
The paper provides detailed descriptions of the synthetic data pipeline, training hyperparameters, and evaluation protocols. It mentions that code and data will be released upon acceptance. The use of specific model versions (e.g., Claude Opus 5, Qwen3.5-4B) and clear definitions of metrics (StrongREJECT, over-refusal scoring) enhances reproducibility. However, the reliance on proprietary models for data generation (Claude) may limit full reproducibility for external researchers without access to those specific model versions.
The method is evaluated primarily on models up to 8 parameters, and it is unclear how well it scales to frontier-scale models (e.g., 70B+). The synthetic data generation relies on a strong, aligned model, which may introduce biases or limit the diversity of the training data. The paper does not extensively discuss the computational cost of generating the synthetic corpora or the potential for adversarial attacks specifically targeting the adapter weights. Additionally, the zero-shot transfer to post-trained checkpoints is a strong claim that might not hold for all post-training procedures.
This work has significant implications for the deployment of LLMs in API settings, where lightweight, tunable safety interventions are highly desirable. By providing a method to "turn up" safety without retraining or terminating sessions, CAs offer a practical solution for dynamic risk management. The approach also contributes to the broader field of AI alignment by demonstrating that complex, principle-based behaviors can be compressed into low-dimensional, portable objects. This could facilitate the development of modular safety systems that can be updated or adjusted independently of the base model. Constitutional Adapters provide a lightweight, tunable, and portable mechanism for enhancing LLM safety against misuse and misalignment by distilling constitutional principles into low-rank adapters and steering vectors. The paper demonstrates that these adapters, particularly when using control-subtracted task arithmetic, outperform traditional prompting and steering methods in long-context and multi-turn adversarial settings while maintaining benign compliance and general capabilities.
Simulating microstructure evolution requires quantum-mechanical accuracy and mesoscopic reach in length and time scales, a combination that no current method achieves. Classical phase-field models provide this reach, but their accuracy is limited by phenomenological free energies and mobilities. Here we develop a framework for learning ab initio phase-field models, where the mesoscopic equation is not postulated but derived from a Mori-Zwanzig projection of molecular dynamics onto species-density fields under explicit assumptions. The nonlocal free energy and mobility left unspecified by this equation are parametrized by neural networks and learned from short molecular dynamics trajectories generated with machine-learning interatomic potentials of ab initio accuracy. We demonstrate the framework on an iron-boron melt and on hydrogen-helium mixtures under planetary conditions. For iron-boron, the model shows that the melt at the FeB$_4$ composition is spinodally unstable at ambient pressure but stabilized at 10 GPa, offering a thermodynamic rationale for why FeB$_4$ has been synthesized only under high pressure. For hydrogen-helium, the model predicts the immiscibility boundary and captures droplet nucleation and growth in helium-rain simulations of a column corresponding to 2.2 million atoms, far beyond the scale of atomistic modeling at comparable accuracy. Trained across compositions and conditions, such models could provide a mesoscopic counterpart to ab initio molecular dynamics.
Primary: National University of Singapore
All Institutions: National University of Singapore
The paper presents a rigorous, physics-informed framework for learning ab initio phase-field models by deriving mesoscopic equations from Mori-Zwanzig projections and jointly learning free-energy and mobility functionals from MLIP-generated MD data, demonstrating significant scalability and accuracy in simulating complex microstructure evolution across materials and planetary science applications.
The paper proposes a rigorous framework for deriving mesoscopic phase-field equations from microscopic molecular dynamics (MD) using the Mori-Zwanzig projection formalism. Unlike standard data-driven approaches that fit phenomenological parameters, this method derives the structure of the evolution equation (Generalized Langevin Equation reduced to a Markovian form) from statistical mechanics. The key innovation is the joint learning of the nonlocal free-energy functional and the mobility tensor using neural networks, trained on short MD trajectories generated by machine-learning interatomic potentials (MLIPs). The authors introduce a "ladder" of free-energy functionals, selecting a specific rung (local density plus nonlocal kernel) that balances expressiveness with data efficiency. The derivation is mathematically sound, explicitly stating assumptions (Markovian dynamics, local mobility, Gaussian noise) that justify the reduction from the exact GLE to the practical model. The use of fluctuation-dissipation relations to anchor the mobility and free-energy curvature provides a strong physical constraint on the learning process, preventing the neural networks from learning unphysical dynamics.
The experimental validation is extensive and covers three distinct systems with varying complexity: a Lennard-Jones mixture (benchmark), an Iron-Boron melt (materials science), and a Hydrogen-Helium mixture (planetary science). For the LJ mixture, the model accurately reproduces the phase diagram and spinodal decomposition dynamics, outperforming Landau and Flory-Huggins baselines. The Iron-Boron results provide a thermodynamic rationale for the high-pressure synthesis of FeB4, correctly predicting the stabilization of the melt at 10 GPa. The most impressive result is the Hydrogen-Helium simulation, where the model predicts the immiscibility boundary consistent with prior DFT studies and simulates helium rain in a domain corresponding to 2.2 million atoms over 1 ns—a scale far beyond the reach of direct MD at comparable accuracy. The ability to extrapolate to larger scales and include external potentials (gravity) not present in training data demonstrates the model's generalizability.
The paper provides detailed descriptions of the training data generation, the specific neural network architectures (MLPs for local terms, radial kernels for nonlocal terms), and the loss function components (dynamical, mobility anchor, structure factor anchor, pressure anchor). The use of standard MLIPs (MACE) for data generation enhances reproducibility, as these potentials are increasingly available. However, specific hyperparameters, network sizes, and training protocols are deferred to the Supplemental Material, which is not fully visible in the provided text. The code availability is not explicitly stated in the main text, which is a minor gap for immediate reproducibility, though the methodology is described with sufficient precision for re-implementation.
The framework currently resolves only density fields, neglecting momentum density, which limits its ability to capture hydrodynamic effects like droplet settling via flow rather than just diffusion. The Markovian assumption may fail for systems with long memory times, such as supercooled liquids or glasses. The accuracy of the model is inherently bounded by the accuracy of the underlying MLIPs used to generate training data. Additionally, the computational cost of generating high-quality MD data for complex materials remains a barrier, although the authors argue that the data efficiency of the approach mitigates this.
This work bridges the gap between quantum-mechanical accuracy and mesoscopic scalability, a long-standing challenge in materials science and fluid dynamics. By providing a principled way to learn mesoscopic models from microscopic data, it enables simulations of phenomena like phase separation, nucleation, and coarsening at scales and timescales previously inaccessible. The potential applications are vast, ranging from alloy design and synthesis to understanding planetary interiors. The framework could serve as a foundation for "mesoscopic foundation models" that generalize across compositions and conditions, analogous to CALPHAD databases but with predictive dynamical capabilities. The paper presents a rigorous, physics-informed framework for learning ab initio phase-field models by deriving mesoscopic equations from Mori-Zwanzig projections and jointly learning free-energy and mobility functionals from MLIP-generated MD data, demonstrating significant scalability and accuracy in simulating complex microstructure evolution across materials and planetary science applications.
Simulating microstructure evolution requires quantum-mechanical accuracy and mesoscopic reach in length and time scales, a combination that no current method achieves. Classical phase-field models provide this reach, but their accuracy is limited by phenomenological free energies and mobilities. Here we develop a framework for learning ab initio phase-field models, where the mesoscopic equation is not postulated but derived from a Mori-Zwanzig projection of molecular dynamics onto species-density fields under explicit assumptions. The nonlocal free energy and mobility left unspecified by this equation are parametrized by neural networks and learned from short molecular dynamics trajectories generated with machine-learning interatomic potentials of ab initio accuracy. We demonstrate the framework on an iron-boron melt and on hydrogen-helium mixtures under planetary conditions. For iron-boron, the model shows that the melt at the FeB$_4$ composition is spinodally unstable at ambient pressure but stabilized at 10 GPa, offering a thermodynamic rationale for why FeB$_4$ has been synthesized only under high pressure. For hydrogen-helium, the model predicts the immiscibility boundary and captures droplet nucleation and growth in helium-rain simulations of a column corresponding to 2.2 million atoms, far beyond the scale of atomistic modeling at comparable accuracy. Trained across compositions and conditions, such models could provide a mesoscopic counterpart to ab initio molecular dynamics.
Primary: National University of Singapore
All Institutions: National University of Singapore
The paper presents a rigorous, physics-informed framework for learning ab initio phase-field models by deriving mesoscopic equations from Mori-Zwanzig projections and jointly learning free-energy and mobility functionals from MLIP-generated MD data, demonstrating significant scalability and accuracy in simulating complex microstructure evolution across materials and planetary science applications.
The paper proposes a rigorous framework for deriving mesoscopic phase-field equations from microscopic molecular dynamics (MD) using the Mori-Zwanzig projection formalism. Unlike standard data-driven approaches that fit phenomenological parameters, this method derives the structure of the evolution equation (Generalized Langevin Equation reduced to a Markovian form) from statistical mechanics. The key innovation is the joint learning of the nonlocal free-energy functional and the mobility tensor using neural networks, trained on short MD trajectories generated by machine-learning interatomic potentials (MLIPs). The authors introduce a "ladder" of free-energy functionals, selecting a specific rung (local density plus nonlocal kernel) that balances expressiveness with data efficiency. The derivation is mathematically sound, explicitly stating assumptions (Markovian dynamics, local mobility, Gaussian noise) that justify the reduction from the exact GLE to the practical model. The use of fluctuation-dissipation relations to anchor the mobility and free-energy curvature provides a strong physical constraint on the learning process, preventing the neural networks from learning unphysical dynamics.
The experimental validation is extensive and covers three distinct systems with varying complexity: a Lennard-Jones mixture (benchmark), an Iron-Boron melt (materials science), and a Hydrogen-Helium mixture (planetary science). For the LJ mixture, the model accurately reproduces the phase diagram and spinodal decomposition dynamics, outperforming Landau and Flory-Huggins baselines. The Iron-Boron results provide a thermodynamic rationale for the high-pressure synthesis of FeB4, correctly predicting the stabilization of the melt at 10 GPa. The most impressive result is the Hydrogen-Helium simulation, where the model predicts the immiscibility boundary consistent with prior DFT studies and simulates helium rain in a domain corresponding to 2.2 million atoms over 1 ns—a scale far beyond the reach of direct MD at comparable accuracy. The ability to extrapolate to larger scales and include external potentials (gravity) not present in training data demonstrates the model's generalizability.
The paper provides detailed descriptions of the training data generation, the specific neural network architectures (MLPs for local terms, radial kernels for nonlocal terms), and the loss function components (dynamical, mobility anchor, structure factor anchor, pressure anchor). The use of standard MLIPs (MACE) for data generation enhances reproducibility, as these potentials are increasingly available. However, specific hyperparameters, network sizes, and training protocols are deferred to the Supplemental Material, which is not fully visible in the provided text. The code availability is not explicitly stated in the main text, which is a minor gap for immediate reproducibility, though the methodology is described with sufficient precision for re-implementation.
The framework currently resolves only density fields, neglecting momentum density, which limits its ability to capture hydrodynamic effects like droplet settling via flow rather than just diffusion. The Markovian assumption may fail for systems with long memory times, such as supercooled liquids or glasses. The accuracy of the model is inherently bounded by the accuracy of the underlying MLIPs used to generate training data. Additionally, the computational cost of generating high-quality MD data for complex materials remains a barrier, although the authors argue that the data efficiency of the approach mitigates this.
This work bridges the gap between quantum-mechanical accuracy and mesoscopic scalability, a long-standing challenge in materials science and fluid dynamics. By providing a principled way to learn mesoscopic models from microscopic data, it enables simulations of phenomena like phase separation, nucleation, and coarsening at scales and timescales previously inaccessible. The potential applications are vast, ranging from alloy design and synthesis to understanding planetary interiors. The framework could serve as a foundation for "mesoscopic foundation models" that generalize across compositions and conditions, analogous to CALPHAD databases but with predictive dynamical capabilities. The paper presents a rigorous, physics-informed framework for learning ab initio phase-field models by deriving mesoscopic equations from Mori-Zwanzig projections and jointly learning free-energy and mobility functionals from MLIP-generated MD data, demonstrating significant scalability and accuracy in simulating complex microstructure evolution across materials and planetary science applications.
Single-particle cryo-electron microscopy (cryo-EM) has become a widely adopted technique for biomolecular structure determination. The conventional cryo-EM computational pipeline first combines many particle images to reconstruct an electrostatic potential (ESP) map and then fits an atomic model to the recovered map. Density reconstruction has high sample complexity, requiring large numbers of particle images and making structure determination high-cost and low-throughput, particularly for heterogeneous samples. Downstream atomic model building, in turn, becomes increasingly difficult as the resolution of the reconstructed map deteriorates. Protein structure prediction models provide strong sequence-derived priors on atomic structure, and experiment-guided approaches can use these priors to recover structures consistent with experimental measurements. Yet, in cryo-EM, such priors are typically integrated only after density reconstruction during atomic model fitting. We introduce Fold'EM, an inference-time framework that combines priors from protein generative models directly with cryo-EM particle images to determine atomic models from a small number of single particle images, bypassing both intermediate density reconstruction and downstream model building against the reconstructed map. Across synthetic and experimental cryo-EM datasets, Fold'EM recovers accurate atomic structures both with known particle orientations and in an ab-initio setting where orientations are inferred jointly with structure. In heterogeneous datasets, Fold'EM further resolves distinct conformational states from mixed particle populations without separately reconstructing a density map and building an atomic model for each state. We believe these results open new avenues for structure determination in the low-sample regime and for characterizing low-population conformational states directly from cryo-EM particles.
Primary: University of Cambridge
All Institutions: University of Cambridge, DeepMind
Fold'EM introduces a novel inference-time framework that directly guides AlphaFold3 using cryo-EM particle images, bypassing intermediate density reconstruction. By optimizing intermediate Pairformer states and jointly inferring structure, pose, and class, the method achieves high-resolution atomic structures from sparse particle sets, significantly reducing the sample complexity of cryo-EM structure determination.
The paper introduces Fold'EM, a framework that bypasses the traditional cryo-EM density reconstruction step by directly guiding AlphaFold3 using individual particle images. The core technical novelty lies in optimizing intermediate Pairformer states ($H$) rather than the terminal embeddings ($Z$) or atomic coordinates. The authors provide a rigorous theoretical justification for this choice, demonstrating that optimizing $H$ induces a sequence-dependent anisotropic geometry in the conditioning space, which acts as a learned preconditioner for the gradient updates. This prevents the model from drifting into unrealistic structural configurations during the inference-time optimization. The method also introduces a joint inference scheme for structure, pose, and class assignment, utilizing a novel bootstrapped residual-subspace estimator (B-RECOVAR) for label-free classification in heterogeneous samples. The mathematical formulation of the particle-level likelihood versus density-level likelihood is well-articulated, highlighting the statistical insufficiency of consensus densities when poses are uncertain.
The experimental evaluation is extensive, covering both synthetic benchmarks and experimental EMPIAR datasets. The paper demonstrates significant improvements over baselines (RELION, cryoDRGN, and previous AlphaFold-guided methods) in terms of FSC resolution and TM-score, particularly in the low-sample regime (e.g., recovering structures from as few as 100-500 particles). The ablation studies effectively isolate the contributions of particle-level guidance versus density guidance and the optimization of intermediate states. The results on heterogeneous data show the ability to resolve distinct conformational states without separate density reconstructions, which is a significant practical advantage. The comparison with unguided AlphaFold3 clearly shows the value of the experimental guidance in correcting mispredicted conformations.
The paper provides detailed algorithmic descriptions and mathematical formulations. However, as an arXiv preprint, the code availability is not explicitly confirmed in the text provided. The reliance on AlphaFold3, a proprietary model, may limit full reproducibility for some users, though the method itself is clearly defined. The use of standard cryo-EM tools (RELION, cryoDRGN) for baselines aids in comparability.
The method relies heavily on the quality of the AlphaFold3 prior; if the unguided prediction is significantly wrong (e.g., 8PWH case), the method may struggle or require more particles than density-based methods. The computational cost of running multiple reverse-diffusion runs and back-propagating through the Pairformer could be high. The current implementation assumes known CTF parameters and centered particles, which may not always hold in raw experimental data without preprocessing.
This work has the potential to significantly reduce the sample size required for cryo-EM structure determination, making it feasible to study low-population states and scarce samples. It bridges the gap between protein structure prediction and experimental cryo-EM, offering a new paradigm for structure determination that could accelerate drug discovery and structural biology research. The insights into optimizing intermediate representations of generative models may also be applicable to other inverse problems in scientific imaging. Fold'EM introduces a novel inference-time framework that directly guides AlphaFold3 using cryo-EM particle images, bypassing intermediate density reconstruction. By optimizing intermediate Pairformer states and jointly inferring structure, pose, and class, the method achieves high-resolution atomic structures from sparse particle sets, significantly reducing the sample complexity of cryo-EM structure determination.
Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composition is lossy. A vote can only overturn mistakes that its voters do not share, and RLVR sharpens a policy so that its samples increasingly make the same mistakes. With every method drawing exactly $160$ completions per problem, training a single LoRA adapter on the full RLVR budget raises single-sample accuracy on every model we test ($1.5$B-$8$B). Yet on three of four models it leaves the majority vote below that of the untrained base model, by up to $4.8$ points. The damage builds during training: voter errors grow steadily more correlated, and the majority vote accuracy peaks early before falling by up to $7.0$ points. The cause is concentration, not RLVR itself. We split the same data and training budget across $K$ LoRA adapters, each trained on its own random disjoint shard, and call the result an adapter thicket. Thickets out-vote the fully trained adapter in all $16$ (model, $K$) settings, and for $K{\geq}4$ they stay within $0.8$ points of the base model or above it. A single adapter stopped early, at a thicket member's step count, is a strong control that matches thickets for small $K$. For $K{\geq}8$, thickets keep more of RLVR's single-sample gain and out-vote this control in six of eight settings. The cost of concentration also grows with the number of votes: from $16$ to $160$ votes, the thicket's lead over the fully trained adapter widens from $1.3$ to $3.3$ points. When the plan is to sample and vote, an RLVR budget is better spent broad than deep.
Primary: Princeton University
All Institutions: Princeton University
The paper demonstrates that splitting RLVR training budgets across multiple LoRA adapters ("thickets") preserves the diversity necessary for effective majority voting, outperforming the standard single-adapter pipeline which suffers from error correlation and diversity collapse.
The paper proposes a simple but effective alternative to the standard RLVR (Reinforcement Learning with Verifiable Rewards) pipeline: instead of training a single LoRA adapter on the full dataset, it splits the data and budget across K independent adapters ("thickets"). The core insight is that RLVR sharpens the policy, reducing sample diversity and increasing error correlation among voters in majority voting, which degrades test-time scaling performance. The methodology is sound, using a fixed inference budget (160 completions) to isolate the effect of training strategy from sampling budget. The use of a step-matched control (early stopping the single adapter) is a critical experimental design choice that strengthens the causal claim.
The experiments are rigorous and well-controlled. The authors test across four models (Qwen2.5-1.5B/3B/7B, Llama-3.1-8B) and two domains (Math, Code). They demonstrate that the single-adapter approach often performs worse than the untrained base model in majority voting, while the thicket approach recovers this performance and often exceeds it. The analysis of error correlation and coverage provides strong mechanistic evidence for the observed results. The ablation on shard type (random vs. subject-specific) adds practical value.
The paper provides sufficient detail on the training setup (GRPO, LoRA rank, hyperparameters) and evaluation protocols (temperature, top-p, voting logic). The specific datasets (MATH, GSM8K, etc.) and model checkpoints are standard, making reproduction feasible for a lab with adequate compute resources. The code is not explicitly linked in the provided text, but the methods are standard enough to implement.
The study is limited to models up to 8B parameters and specific RL algorithms (GRPO). The voting results are primarily for math tasks where answers are canonicalizable; code tasks are only analyzed for coverage and correlation, not final vote accuracy. The "thicket" approach requires managing multiple adapters at inference time, which, while mitigated by modern serving stacks, adds operational complexity compared to a single model.
This paper challenges a common assumption in the LLM post-training community: that more RLVR training is always better for test-time scaling. It provides a clear, actionable guideline for practitioners: if you plan to use majority voting, diversify your training budget. This has immediate practical implications for teams deploying LLMs with test-time compute budgets. The paper demonstrates that splitting RLVR training budgets across multiple LoRA adapters ("thickets") preserves the diversity necessary for effective majority voting, outperforming the standard single-adapter pipeline which suffers from error correlation and diversity collapse.
Biological reasoning models use post-training to connect LLMs to biological foundation model representations and biological text. Their benchmark accuracy is taken as evidence that LLMs reason over these inputs. We test this assumption in six biological reasoning models across DNA, protein, and single-cell tasks. We perturb one biological input while holding the query and other inputs fixed, construct evidence conflicts that pair the foundation model representation of one genome, protein, or cell with the text of another, fit linear probes to the representations the language model receives, and analyze reasoning traces against the biological inputs. Evo2 and ESM3 contribute little to BioReason and BioReason-Pro performance on the evaluated tasks. Shuffling the DNA sequence barely changes BioReason disease prediction accuracy, and in evidence conflicts the two models follow the text in 97.9% and 99.7% of cases. Linear probes trained on the Evo2 and ESM3 representations predict the task targets, so these foundation models encode information relevant to the task, but provide limited overall performance improvement to BioReason and BioReason-Pro. In contrast, foundation model inputs contribute to ChatNT, Prot2Text-V2, and CellWhisperer performance, and differentially expressed genes in the gene sentence contribute to Cell2Sentence-Scale performance. Across SFT and RL checkpoints of BioReason-Pro and 42 BioReason checkpoints, increases in accuracy do not imply greater performance contributions from biological inputs. BioReason traces misstate nucleotide changes, while BioReason-Pro traces describe functions omitted from final predictions under evidence conflicts. We find that current post-training strategies do not ensure that foundation model representations contribute to task performance.
Primary: Harvard University
All Institutions: Harvard University, Harvard Medical School
The paper provides a critical evaluation of biological reasoning models, demonstrating that current post-training strategies often fail to ensure that foundation model representations contribute to task performance. By using rigorous perturbation and evidence conflict analyses, the authors reveal that models like BioReason and BioReason-Pro largely ignore their biological inputs in favor of textual shortcuts, a finding that has significant implications for the design and evaluation of future multimodal scientific AI systems.
The paper employs a rigorous interventionist methodology to test the causal role of biological foundation model representations in LLM-based reasoning models. The authors utilize input perturbation (shuffling DNA sequences), evidence conflict construction (mismatching foundation model embeddings with text descriptions), and linear probing to isolate the contribution of specific modalities. This approach is methodologically sound for determining whether models are genuinely "reasoning" over biological inputs or merely relying on textual shortcuts. The inclusion of analysis across multiple training checkpoints (SFT and RL) adds depth to the evaluation of how post-training strategies affect input utilization.
The experiments cover six distinct biological reasoning models across DNA, protein, and single-cell tasks, providing a broad scope. The key finding that Evo2 and ESM3 contribute negligibly to BioReason and BioReason-Pro performance, with models following text in >97% of conflict cases, is a significant empirical result. The contrast with models like ChatNT and CellWhisperer, where foundation model inputs do contribute, highlights that the issue is specific to certain post-training strategies or model architectures rather than a universal failure of multimodal integration. The analysis of reasoning traces revealing misstatements of nucleotide changes further supports the conclusion that the models are not faithfully using the biological inputs.
The paper provides a GitHub repository link (https://github.com/mims-harvard/bio-mirage) and a project website, which strongly suggests that code and resources will be available for reproduction. The detailed description of the perturbation and conflict construction methods allows other researchers to replicate the analysis on other models.
The study focuses on a specific set of six models and three biological domains. It is unclear if the findings generalize to other types of biological foundation models or different post-training paradigms. The linear probes may not capture all non-linear relationships between the foundation model representations and the task targets, potentially underestimating the contribution of the biological inputs in some cases.
This paper has high impact for the field of AI in science, particularly for the development of trustworthy biological reasoning systems. It challenges the common assumption that high benchmark accuracy implies effective use of specialized scientific inputs. The findings will likely influence future post-training strategies to explicitly reward the use of foundation model representations, leading to more robust and interpretable AI models for biology. The paper provides a critical evaluation of biological reasoning models, demonstrating that current post-training strategies often fail to ensure that foundation model representations contribute to task performance. By using rigorous perturbation and evidence conflict analyses, the authors reveal that models like BioReason and BioReason-Pro largely ignore their biological inputs in favor of textual shortcuts, a finding that has significant implications for the design and evaluation of future multimodal scientific AI systems.
Learning to complete tasks in unfamiliar environments with unknown rules remains a key challenge for LLM agents. Current LLM agents often record their discoveries in prose, which may not provide a compact, explicit account of how the environment works. Inspired by how scientists organize observations into testable, predictive theories, we introduce Schema, an agent harness that organizes learning and action through interactive program induction. The LLM agent decides what to investigate and how to act, expressing its evolving understanding of the environment as executable programs. The harness consists of a persistent program workspace and a small set of interfaces for checking these programs against the interaction history, planning within them, and executing plans under step-by-step verification. Schema raises ARC-AGI-3 RHAE from 58.7% to 99.2% with the same base model, solves 100% of the public DiG-bench games, and reaches the median performance of the top-50 human players on MazeBench. Extensive analysis shows the effectiveness of Schema in unknown mechanism discovery, and ablations confirm the contribution of each component.
Primary: Carnegie Mellon University
All Institutions: Carnegie Mellon University, University of California, Berkeley, Impossible AI
The paper presents a novel agent harness, Schema, that utilizes interactive program induction to enable LLMs to discover and exploit unknown environment rules with human-level efficiency. By organizing learning through executable programs rather than prose, the method achieves state-of-the-art results on ARC-AGI-3, DiG-bench, and MazeBench, demonstrating a robust and scalable approach to mechanism discovery in unfamiliar settings.
The paper introduces "Interactive Program Induction" (IPI), a paradigm where LLM agents represent their understanding of unknown environments as executable programs rather than prose. The proposed harness, Schema, consists of four core operations: Hypothesize (writing code to model state and transitions), Certify (replaying history to check consistency), Plan (using the program as a simulator for search), and Act with Verification (executing actions and stopping if predictions mismatch). This approach effectively addresses the "lost in the middle" and context degradation issues inherent in prose-based memory by forcing the agent to distill knowledge into compact, testable, and reusable code. The methodology is sound, leveraging the LLM's coding capabilities to create a persistent, verifiable world model that persists across context compactions.
The evaluation is extensive and rigorous, covering three distinct benchmarks: ARC-AGI-3 (visual reasoning), DiG-bench (text-based rule discovery), and MazeBench (long-horizon 3D exploration). The results are striking: Schema achieves 99.2% RHAE on ARC-AGI-3 (vs. 58.7% baseline), solves 100% of public DiG-bench games, and matches top-50 human performance on MazeBench. Ablation studies clearly demonstrate the contribution of each component (certification, planning, verification), showing that removing any one significantly degrades performance. The analysis of token costs and action efficiency further strengthens the claim of practical utility.
The paper provides detailed implementation descriptions in the appendix, including the program contract, tool interfaces, and benchmark adapters. However, the specific code for the Schema harness and the exact prompts used are not fully detailed in the text, and no public code repository URL is provided in the extracted text. While the methodology is clearly described, full reproducibility would require access to the specific harness implementation and prompt engineering details.
The approach relies heavily on the base model's coding ability; weaker models may struggle to write correct world models. The computational cost of backtesting and planning can be high, though the paper argues this is offset by reduced interaction steps. The benchmarks used (ARC-AGI-3, DiG-bench, MazeBench) are relatively new and may not fully represent the breadth of real-world unknown environments.
This work has significant implications for the design of autonomous agents in open-ended environments. By shifting from passive memory to active, executable theory-building, it offers a scalable path toward agents that can genuinely learn and adapt to novel tasks without retraining. The paradigm of "interactive program induction" could be applied to robotics, scientific discovery, and other domains where environments are complex and rules are not explicitly given. The paper presents a novel agent harness, Schema, that utilizes interactive program induction to enable LLMs to discover and exploit unknown environment rules with human-level efficiency. By organizing learning through executable programs rather than prose, the method achieves state-of-the-art results on ARC-AGI-3, DiG-bench, and MazeBench, demonstrating a robust and scalable approach to mechanism discovery in unfamiliar settings.
Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost. We build PhantomEnvironments, multi-turn RL environments from fictional worlds, where agents must search a corpus of templated articles to answer multi-hop questions. Despite sharing no facts with the real world, these strikingly simple environments yield agents that transfer to real-world multi-hop search benchmarks, often outperforming real-world training data on newer benchmarks. Trained agents generalize to unseen fictional universes, and Qwen models learn to scale their search budget roughly linearly with question difficulty, suggesting emergent search scaling from environment interaction alone. Ablating environment complexity reveals that hop count drives transfer more than constraints or comparisons: even the simplest rule-generated environments are a surprisingly effective, free resource for training generalizable LLM agents.
Primary: Cornell University
All Institutions: Cornell University, Stanford University
The paper introduces a cost-effective, rule-based synthetic environment framework for training LLM search agents, demonstrating that structural complexity (hop count) drives transferability to real-world tasks more than factual accuracy. This approach offers a scalable alternative to human-curated or LLM-generated data, enabling the emergence of adaptive search strategies in LLM agents.
The paper proposes a novel paradigm for training LLM agents using Reinforcement Learning (RL) by replacing costly, human-curated, or LLM-generated environments with "PhantomEnvironments." These are synthetic, rule-based environments derived from fictional worlds where agents must perform multi-hop search over templated articles. The core methodological innovation is the decoupling of environment generation from LLM inference, ensuring zero marginal cost and eliminating hallucination risks associated with LLM-generated data. The approach relies on the hypothesis that the structural complexity of the search task (specifically hop count) is more critical for learning generalizable search strategies than the factual accuracy of the content.
The experiments demonstrate that agents trained on these fictional, factually incorrect environments transfer effectively to real-world multi-hop search benchmarks. Notably, the paper claims that agents trained on PhantomEnvironments often outperform those trained on real-world data on newer benchmarks, suggesting that the structural learning of search strategies is robust to domain shift. Ablations confirm that hop count is the primary driver of transfer performance. The observation of "emergent search scaling," where Qwen models learn to allocate search budget linearly with question difficulty, is a significant empirical finding.
The authors provide a strong reproducibility statement, indicating the use of open-source LLMs, training code, and evaluation benchmarks. They report standard errors and significance tests. A GitHub repository is provided for code access. The use of AI tools for code implementation is disclosed, but the authors state they verified all code and results.
The primary limitation is the reliance on "templated articles" and rule-based generation, which may not capture the full stochasticity and ambiguity of real-world web search. The transferability to "newer benchmarks" is a strong claim that requires careful scrutiny regarding potential leakage or specific alignment with the benchmark structure. The paper is a preprint (ICLR 2027 submission), so peer review status is pending.
This work has high potential impact by providing a scalable, cost-effective method for training LLM agents. If the transferability results hold, it could significantly reduce the barrier to entry for developing sophisticated search agents, allowing smaller labs to train competitive models without massive human annotation budgets. It shifts the focus from data fidelity to structural complexity in agent training. The paper introduces a cost-effective, rule-based synthetic environment framework for training LLM search agents, demonstrating that structural complexity (hop count) drives transferability to real-world tasks more than factual accuracy. This approach offers a scalable alternative to human-curated or LLM-generated data, enabling the emergence of adaptive search strategies in LLM agents.
On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action that moves the agent farther from completing the task, and this mistake typically occurs early. These pivotal mistakes often remain recoverable: guiding the model for only a few turns after the pivotal turn can restore task success. We therefore propose PivotOPD, an on-policy distillation framework that jointly trains the student to prevent pivotal mistakes and to recover from the states they create. At each pivotal mistake, a teacher model provides a gold action and then names a recovery action at each of the next few turns. Preventive distillation uses the gold action with reverse KL to steer the student away from the pivotal mistake, while recovery distillation uses the recovery actions with forward KL to transfer recovery behaviors that the student rarely samples. Against 13 baselines on ALFWorld, WebShop, and Search-based QA, PivotOPD achieves the strongest average performance for both Qwen3-1.7B and Qwen3-8B students, improving over the strongest baseline on ALFWorld by +5.5% with the 1.7B student. The gains also transfer to another model family on the software engineering domain, where PivotOPD raises the resolve rate of a Nemotron-3.5 student on SWE-Bench Verified by +3.2%. Project page: https://research.nvidia.com/labs/lpr/pivotopd/
Primary: NVIDIA
All Institutions: NVIDIA
PivotOPD introduces a targeted distillation framework that teaches LLM agents to recover from pivotal mistakes in multi-turn interactions. The paper provides a rigorous analysis of error accumulation, a novel dual-distillation method, and strong empirical results across diverse benchmarks, establishing a new standard for training robust language agents.
The paper proposes PivotOPD, a framework that addresses error accumulation in multi-turn LLM agents by identifying "pivotal mistakes" and applying targeted distillation. The core innovation is the dual distillation approach: preventive distillation (using reverse KL to steer away from mistakes) and recovery distillation (using forward KL to teach recovery behaviors from a privileged self-teacher). The theoretical analysis correctly identifies that standard on-policy methods fail to provide sufficient learning signal for recovery actions because the student rarely samples them, justifying the use of forward KL on teacher-generated responses. The pivot detection mechanism, which uses a teacher model to identify candidate turns and gold actions, is a practical solution to the lack of oracles in real-world environments.
The experiments are extensive, covering four benchmarks (ALFWorld, WebShop, Search-based QA, SWE-Bench Verified) and multiple model families (Qwen3, Nemotron). The comparison against 13 baselines is robust. The results show consistent improvements, particularly in recovery rates, which aligns with the paper's motivation. The ablation studies effectively demonstrate the necessity of both preventive and recovery components and the importance of supervision at the correct turns. The transfer to SWE-Bench Verified with a different model family (Nemotron) strengthens the generalizability claim.
The paper provides detailed descriptions of the training objective, pivot detection prompts, and evaluation metrics. The project page URL is provided, which likely contains code and additional resources. The hyperparameters and training configurations are described in the appendix, supporting reproducibility.
The method relies on a teacher model for pivot detection and action naming, which may not be available in all settings. The performance gains, while statistically significant, are moderate in some benchmarks (e.g., +1.2% on WebShop success rate). The reliance on a privileged self-teacher for recovery distillation adds computational overhead. The pivot detection accuracy (77.8% within one turn of oracle) suggests room for improvement in identifying pivotal turns without an oracle.
This work has significant implications for training robust LLM agents in interactive environments. By explicitly teaching recovery from mistakes, it addresses a key limitation of current agent training methods. The framework could be extended to other domains where error accumulation is a challenge, such as robotics or autonomous driving. The insights into the nature of pivotal mistakes and recovery behaviors provide valuable guidance for future research on agent robustness. PivotOPD introduces a targeted distillation framework that teaches LLM agents to recover from pivotal mistakes in multi-turn interactions. The paper provides a rigorous analysis of error accumulation, a novel dual-distillation method, and strong empirical results across diverse benchmarks, establishing a new standard for training robust language agents.
We compare the instance-wise, finite-sample risks of monotone spectral filters for linear regression, a broad class of estimators including principal component regression (PCR), gradient descent (GD), and ridge regression. We show that PCR dominates all monotone spectral filters: compared to any such filter, the risk of optimally tuned PCR is no bigger by a constant factor for all problems. Furthermore, the dominance is strong if the filter is separated from step functions (e.g., GD and ridge): there exist problem instances for which the risk of PCR is smaller by a polynomial factor in sample size dependence. Our comparison results show that PCR is optimal and thus admissible among monotone filters, significantly extending Wu et al. (2026)'s result that GD strongly dominates ridge. From a technical perspective, we establish new upper and lower bounds for general spectral filters, which are instance-wise sharp when specialized to ridge or GD, recovering or improving the best-known bounds.
Primary: UC Berkeley
All Institutions: UC Berkeley, Princeton University
Principal Component Regression dominates all monotone spectral filters for linear regression in terms of instance-wise finite-sample risk, establishing PCR as admissible and GD as inadmissible. The paper provides a comprehensive theoretical analysis using Schur multipliers to derive sharp risk bounds, offering a new perspective on the comparison of regularization methods beyond worst-case minimax rates.
The paper employs a rigorous theoretical framework to compare the instance-wise finite-sample risks of monotone spectral filters in linear regression. The core methodological contribution is the application of Schur multipliers and matrix divided differences to control noncommutative matrix perturbations, specifically the "leave-tail-out" difference between the full Gram matrix and its truncated version. This allows for the derivation of sharp upper and lower bounds for general spectral filters, including Principal Component Regression (PCR), Gradient Descent (GD), and Ridge Regression. The authors demonstrate that PCR dominates all monotone filters and strongly dominates filters separated from step functions (like GD and Ridge), establishing PCR as admissible and GD as inadmissible in this context. The technical depth is high, extending classical leave-one-out ideas with advanced operator theory tools.
This is a purely theoretical paper with no experimental evaluation. The "experiments" are mathematical proofs and derivations of risk bounds. The validity is established through rigorous mathematical argumentation rather than empirical benchmarks.
As a theoretical paper, reproducibility is defined by the clarity and correctness of the proofs. The paper provides detailed appendices with missing proofs and clearly states assumptions. The use of AI for parts of the technical ingredients is disclosed, but the authors state they rederived and verified all proofs. The mathematical framework is well-defined and reproducible in the sense that other researchers can verify the theorems.
The dominance results rely on Gaussian random design assumptions, which may not hold in all practical settings. The paper acknowledges that the variance bounds for PCR could likely be improved. Additionally, the results are specific to linear regression; extending these dominance relationships to non-linear models or other learning tasks remains an open question. The reliance on specific spectral properties (monotonicity) limits the scope to a specific class of estimators.
The paper significantly impacts the field of statistical learning theory by providing a definitive instance-wise comparison of standard linear regression methods. It challenges the common heuristic that Gradient Descent is a universally good default by showing it is inadmissible compared to PCR in terms of instance-wise risk. This insight may influence the design of future algorithms, encouraging the use of methods that can explicitly discard weak spectral components to reduce variance. It also provides a new technical toolkit (Schur multipliers for risk analysis) that can be applied to other statistical learning problems. Principal Component Regression dominates all monotone spectral filters for linear regression in terms of instance-wise finite-sample risk, establishing PCR as admissible and GD as inadmissible. The paper provides a comprehensive theoretical analysis using Schur multipliers to derive sharp risk bounds, offering a new perspective on the comparison of regularization methods beyond worst-case minimax rates.
Training models to act in accordance with an explicitly defined set of principles, or "constitution," has shown promise as a robust and transparent mechanism for AI alignment. However, the generality and flexibility of such methods remain unclear. Here, we show that constitution-consistent behavior can be distilled from synthetic corpora into lightweight objects (low-rank adapters and steering vectors). Despite never seeing a harmful request or jailbreak during training, such objects increase jailbreak defense success and measured alignment -- particularly at long context lengths and against multi-turn attacks, where they outperform both prompted and steered baselines. Subtracting control-trained from constitution-trained objects further accentuates these effects, yielding defenses we call "constitutional adapters" (CAs). CAs can be trained on a base model, transferred zero-shot to its post-trained checkpoint, and scaled at inference time to predictably trade off defense for benign compliance. Taken together, these results recommend CAs as a lightweight, portable, and tunable lever for mitigating misalignment and misuse in API deployments.
Primary: Anthropic
All Institutions: Anthropic
Constitutional Adapters provide a lightweight, tunable, and portable mechanism for enhancing LLM safety against misuse and misalignment by distilling constitutional principles into low-rank adapters and steering vectors. The paper demonstrates that these adapters, particularly when using control-subtracted task arithmetic, outperform traditional prompting and steering methods in long-context and multi-turn adversarial settings while maintaining benign compliance and general capabilities.
The paper proposes "Constitutional Adapters" (CAs), a method for distilling alignment principles into lightweight, inference-time interventions (LoRAs and ReLU steering vectors). The core methodological innovation is the use of a control-subtracted task arithmetic approach: training a lightweight object on a synthetic corpus derived from a model constitution ($D_C$) and subtracting an object trained on a register-matched control corpus ($D_K$). This subtraction aims to isolate the "constitutional" behavior from general fine-tuning artifacts like format or register shifts. The method leverages synthetic data generation via a strong model (Claude Opus 5) to create diverse scenarios that stress-test ethical principles, rather than relying on adversarial jailbreak examples. This allows the adapters to generalize to unseen attacks. The use of ReLU steering vectors, which allow for input-dependent steering strength, is a significant technical detail that improves upon fixed-direction steering.
The experimental evaluation is rigorous and comprehensive. The authors test across four distinct model families (Qwen3.5, Llama-3.1, Gemma-4, Nemotron 3) and various attack types (static, multi-query, multi-turn). Key findings include the superior robustness of CAs in long-context and multi-turn settings compared to system prompting, which degrades as context length increases. The paper also evaluates the trade-off between defense and benign compliance (over-refusal) by sweeping the steering strength, demonstrating a predictable and tunable frontier. The inclusion of misalignment evaluations (Petri Bloom) and capability benchmarks (MMLU, GSM8K, etc.) provides a holistic view of the method's impact, showing minimal capability degradation.
The paper provides detailed descriptions of the synthetic data pipeline, training hyperparameters, and evaluation protocols. It mentions that code and data will be released upon acceptance. The use of specific model versions (e.g., Claude Opus 5, Qwen3.5-4B) and clear definitions of metrics (StrongREJECT, over-refusal scoring) enhances reproducibility. However, the reliance on proprietary models for data generation (Claude) may limit full reproducibility for external researchers without access to those specific model versions.
The method is evaluated primarily on models up to 8 parameters, and it is unclear how well it scales to frontier-scale models (e.g., 70B+). The synthetic data generation relies on a strong, aligned model, which may introduce biases or limit the diversity of the training data. The paper does not extensively discuss the computational cost of generating the synthetic corpora or the potential for adversarial attacks specifically targeting the adapter weights. Additionally, the zero-shot transfer to post-trained checkpoints is a strong claim that might not hold for all post-training procedures.
This work has significant implications for the deployment of LLMs in API settings, where lightweight, tunable safety interventions are highly desirable. By providing a method to "turn up" safety without retraining or terminating sessions, CAs offer a practical solution for dynamic risk management. The approach also contributes to the broader field of AI alignment by demonstrating that complex, principle-based behaviors can be compressed into low-dimensional, portable objects. This could facilitate the development of modular safety systems that can be updated or adjusted independently of the base model. Constitutional Adapters provide a lightweight, tunable, and portable mechanism for enhancing LLM safety against misuse and misalignment by distilling constitutional principles into low-rank adapters and steering vectors. The paper demonstrates that these adapters, particularly when using control-subtracted task arithmetic, outperform traditional prompting and steering methods in long-context and multi-turn adversarial settings while maintaining benign compliance and general capabilities.
What determines the unavoidable sample cost of learning cyclic causal structure? For cyclic linear non-Gaussian models, we study exact condensation recovery from observational data: identifying the strongly connected component (SCC) partition and all edges between components. We establish the first information-theoretic lower bounds on sample complexity for this target. For $p$ variables, maximum SCC size $s_{\max}$, and maximum external-parent count $d_B$, any estimator requires order $s_{\max}\log(ep/s_{\max})+d_B\log(ep/d_B)$ samples in the worst case over a regular model class. These bounds distinguish the costs of SCC membership and external-parent selection. Under principal invertibility and without correlation faithfulness, we establish a population block-exogeneity principle that identifies unknown root SCCs through residual independence and inclusion minimality. A sparse-adjustment characterization shows that small adjustment sets suffice to identify SCCs and their direct external parents, without regressing on all previously recovered variables. These characterizations yield BlockExo, which attains a structurally matching sample bound without knowing $s_{\max}$ or $d_B$ under suitable conditions. Simulations support the structural dependence of our sample bound and demonstrate BlockExo's sample-efficient recovery in comparisons with other methods for cyclic causal discovery.
Primary: Princeton University
All Institutions: Princeton University, Seoul National University
The paper establishes the first information-theoretic lower bounds for sample complexity in cyclic causal discovery and introduces BlockExo, an algorithm that achieves structurally optimal recovery. By decoupling the costs of SCC membership and external parent selection, the work provides a precise characterization of the statistical difficulty of learning feedback loops, offering a rigorous theoretical foundation that advances the field of causal inference beyond acyclic assumptions.
The paper proposes a rigorous theoretical framework for causal discovery in cyclic linear non-Gaussian (LiNG) models. The core contribution is the establishment of the first information-theoretic lower bounds on the sample complexity for recovering the condensation (SCC partition and inter-component edges). The authors introduce a "block-exogeneity" principle that allows for the identification of root SCCs via residual independence without requiring correlation faithfulness, a significant relaxation of standard assumptions. The proposed algorithm, BlockExo, utilizes sparse adjustment sets to identify SCCs and their external parents, achieving a sample complexity that matches the derived lower bounds structurally. The methodology is mathematically sound, leveraging the Darmois-Skitovitch theorem and Fano's inequality to bridge the gap between population-level identifiability and finite-sample guarantees.
The experimental section is limited but appropriate for a theory-heavy paper. It validates the structural dependence of the sample bound by showing that recovery curves align when normalized by the theoretical factors. It compares BlockExo against existing methods like Coarsening, DisjointCycles, and StableSpIn. BlockExo demonstrates superior sample efficiency in exact recovery tasks, particularly in overlapping cycle structures where baselines fail or require significantly more samples. However, the experiments are conducted on synthetic data only, with small dimensionality ($p=50$) and limited noise distributions, which restricts the generalizability of the empirical claims.
The paper provides a reproducibility statement indicating that code, configurations, and scripts will be released. The algorithmic details are clear, and the theoretical assumptions are explicitly stated. However, the lack of publicly available code at the time of review and the reliance on specific, potentially hard-to-tune thresholds (e.g., in the $PASS$ condition) may pose challenges for independent replication without the authors' implementation.
The primary limitation is the restriction to linear non-Gaussian models with causal sufficiency (no latent confounders). The method assumes principal invertibility, which, while standard, excludes certain unstable cyclic systems. The computational complexity of BlockExo is exponential in the sum of SCC size and external parent count ($p^{O(s_{max}+d_B)}$), which limits its applicability to dense or large-scale cyclic structures. Furthermore, the empirical evaluation lacks real-world data validation, leaving the practical utility of the method in complex biological or economic systems unproven.
This work provides a foundational statistical understanding of the costs associated with learning cyclic causal structures, a critical area in systems biology and economics. By establishing minimax optimality for condensation recovery, it sets a benchmark for future algorithms. The block-exogeneity principle offers a new tool for causal discovery that does not rely on faithfulness, potentially influencing the design of more robust causal inference methods in the broader machine learning community. The paper establishes the first information-theoretic lower bounds for sample complexity in cyclic causal discovery and introduces BlockExo, an algorithm that achieves structurally optimal recovery. By decoupling the costs of SCC membership and external parent selection, the work provides a precise characterization of the statistical difficulty of learning feedback loops, offering a rigorous theoretical foundation that advances the field of causal inference beyond acyclic assumptions.
Neural operators are typically trained in a supervised fashion, which requires a dataset to be generated with a classical solver. Training them physics-informed, i.e., purely from the governing equations, removes this large offline cost and allows fresh samples to be drawn at every optimization step, but has so far been limited to simplified problems and trails supervised training in accuracy. The obstacle is the ill-conditioning of physics-informed losses, which differential operators induce and which worsens as the discretization is refined. We therefore propose a preconditioned residual loss function and show mesh-independent conditioning for elliptic problems and greatly improved conditioning for saddle point problems. Realized through geometric and algebraic multigrid, the construction applies to linear and nonlinear equations, steady or time-dependent, on structured and unstructured meshes, is agnostic to the neural operator architecture, and adds no cost at inference. On the Poisson, Allen-Cahn and stationary Stokes equations, the resulting label-free training matches supervised training and is four to twenty-five times more accurate than previous physics-informed operator learning methods.
Primary: ETH Zurich
All Institutions: ETH Zurich
The paper proposes a preconditioned physics-informed loss function using multigrid methods to solve the ill-conditioning problem in neural operator training, achieving label-free accuracy that matches supervised learning. This is a significant technical contribution that bridges numerical linear algebra and deep learning, offering a practical and scalable solution to a long-standing optimization challenge in physics-informed machine learning.
The paper addresses a critical bottleneck in physics-informed neural operator (PINO) training: the severe ill-conditioning of residual losses induced by differential operators, which scales poorly with mesh refinement ($O(h^{-4})$ for elliptic problems). The authors propose a preconditioned least-squares loss function that incorporates multigrid preconditioners (geometric for structured grids, algebraic for unstructured) directly into the loss landscape. This approach is elegant because it decouples the preconditioning from the neural network architecture, requiring no changes to the model structure and adding zero cost at inference. The method generalizes to nonlinear problems (Allen-Cahn) and saddle-point problems (Stokes) by using approximate Jacobian inverses and block-diagonal preconditioners, respectively. The theoretical analysis correctly identifies that preconditioning the residual effectively transforms the Hessian of the loss to be mesh-independent or significantly better conditioned, thereby enabling efficient gradient-based optimization.
The experiments are rigorous and cover three distinct PDE classes: Poisson (linear elliptic), Allen-Cahn (nonlinear parabolic), and Stokes (saddle-point). The use of both FNO (Fourier Neural Operator) and GAOT (Geometry-Aware Operator Transformer) demonstrates architecture agnosticism. The results are compelling: the proposed method matches or surpasses supervised training accuracy while requiring no labeled data, and outperforms standard PINO and PI-DeepONet baselines by factors of 4 to 25. The "infinite data limit" experiment is particularly strong, showing that fresh sampling at each step removes the data-size bottleneck of supervised learning. The inclusion of unstructured meshes for the Stokes problem adds significant practical relevance, as many real-world PDEs do not admit structured grids.
The paper provides a public GitHub repository with code. The appendices contain detailed descriptions of the discretization, preconditioner construction (including specific multigrid parameters), and optimization hyperparameters. The use of standard libraries (neuraloperator, AMGX) and clear algorithmic descriptions (e.g., the V-cycle algorithm) ensures high reproducibility. The authors also provide a clear distinction between the "interpolated neural operator" and the raw network output, clarifying how the loss is computed.
The primary limitation is the dependence on the availability of an efficient preconditioner for the specific PDE class. For problems where multigrid or other fast solvers are not readily available (e.g., highly convection-dominated flows or complex wave equations), the method's advantage diminishes. Additionally, the method relies on an explicit finite element discretization of the residual, which may not be straightforward for all PDE formulations or black-box physics. The paper also notes that for the Stokes problem, the conditioning is only improved to $O(h^{-2})$, not $O(1)$, though this is still sufficient for convergence.
This work has high potential impact on the field of scientific machine learning. By enabling label-free training of neural operators that matches supervised performance, it removes the need for expensive offline data generation via classical solvers. This is particularly significant for physics foundation models and applications where generating high-fidelity training data is prohibitive. The technique is broadly applicable to any PDE where a fast linear solver or preconditioner exists, making it a valuable tool for the broader community of PDE surrogates. The paper proposes a preconditioned physics-informed loss function using multigrid methods to solve the ill-conditioning problem in neural operator training, achieving label-free accuracy that matches supervised learning. This is a significant technical contribution that bridges numerical linear algebra and deep learning, offering a practical and scalable solution to a long-standing optimization challenge in physics-informed machine learning.
Progress in machine learning cannot outpace our ability to verify it. With an explosion in papers today, every scientific claim rests initially on trust in the trainer, leading to uneven evaluation, baselines, and forestalling of reliable progress. Traditionally, the burden of verification falls on the reader, who must reproduce expensive training runs. This strategy is impractical due to an explosion in slop contributions, diversity of methods, and the sheer compute required. We put the burden of proof where it belongs, on the trainer, and in the process also cut the overall cost of verification significantly. We introduce Witnesses, a method for certifying training, data usage and evaluation in a neural network training run. Our key insight is that fast behavioral fingerprints with occasional replay challenges are sufficient for auditing neural network training. Our method is applicable at scale with minimal overhead to the trainer, is cheap for the verifier, rejects bad training runs with amplifiable probability, and allows for exact queries of both data inclusion and exclusion. We test our method on language model training runs from 100M to 2B scales, across DDP and FSDP, and demonstrate this minimal overhead. We also introduce a self-regulating leaderboard of "auto-certified" training runs that enables shared baselines and progress. We invite the community to participate in the leaderboard to improve reproducibility in machine learning.
Primary: Microsoft Research
All Institutions: Microsoft Research, New York University, Stanford University
The paper introduces "Witnesses," a cryptographic and probabilistic framework for certifying the integrity of neural network training runs with minimal overhead. By shifting the burden of verification to the trainer via lightweight behavioral fingerprints and replay challenges, it addresses the reproducibility crisis in ML, enabling trusted baselines and scalable auditing of billion-parameter models.
The paper proposes "Witnesses," a cryptographic and probabilistic framework for certifying the integrity of neural network training runs. The core innovation is shifting the burden of verification from the reader (who must reproduce expensive runs) to the trainer (who generates a lightweight, cryptographically signed "tape" of training steps). The method utilizes three key technical components: (1) a quasi-geometric parameter commitment using low-rank orthonormal projections to create fast behavioral fingerprints of model weights; (2) a non-geometric hash for optimizer states to ensure state continuity; and (3) a probabilistic replay challenge mechanism where the verifier randomly samples steps to verify against the declared training function. The system leverages stateless accelerator serializations (JAX/Torch) to seal the computation and uses white-box cryptography or zkVMs to anchor trust in the attestation engine. The theoretical contribution includes proofs of completeness and soundness, demonstrating that the probability of a malicious trainer successfully hiding invalid updates is bounded and amplifiable.
The authors validate the system on GPT-style models ranging from 100M to 2B parameters, testing both Data Parallel (DDP) and Fully Sharded Data Parallel (FSDP) configurations. Key results include a minimal overhead of 1.19% in speed and 1.11% in memory for a 2B parameter run, which is remarkably low for a security-critical system. The paper demonstrates high detection rates for data poisoning, parameter corruption, and adversarial optimization attacks against the sketch. The inclusion of a "self-regulating leaderboard" (dvbench.org) provides a practical ecosystem for adoption, allowing the community to submit and verify certified runs. The experiments convincingly show that the method scales to billion-parameter models without prohibitive cost.
The paper provides high reproducibility by releasing the code and a public leaderboard. The methodology is code-base agnostic, relying on standard serialization formats (StableHLO/XLA) which are widely supported. The detailed description of the cryptographic primitives (BLAKE3, ratcheted signatures) and the specific implementation details (Rust workers, Python facade) allow for independent verification. The open-source nature of the leaderboard further enhances reproducibility by providing real-world examples of certified runs.
The primary limitation is the reliance on a trusted attestation engine ($B_{att}$). While the paper argues for white-box cryptography or zkVMs, these are complex engineering solutions that may be vulnerable to side-channel attacks or implementation bugs not covered by the theoretical model. Additionally, the system assumes the training function $f$ is declared and bounded; it does not protect against logical errors in the training algorithm itself, only against deviations from the declared algorithm. The overhead, while low, is non-zero and may be significant for extremely latency-sensitive applications. The method is currently focused on pretraining and may require adaptation for fine-tuning or reinforcement learning workflows.
This work has the potential to fundamentally change how machine learning research is conducted and evaluated. By providing a technical solution to the reproducibility crisis, it enables standardized comparisons and trusted baselines, which are currently lacking in the field. The leaderboard initiative could foster a culture of verified results, reducing the impact of "slop" contributions and increasing the reliability of scientific claims. This could lead to more efficient use of computational resources, as researchers can trust published results without duplicating expensive training runs. The broader impact extends to AI safety and governance, as certified training runs provide an audit trail for model behavior and data usage. The paper introduces "Witnesses," a cryptographic and probabilistic framework for certifying the integrity of neural network training runs with minimal overhead. By shifting the burden of verification to the trainer via lightweight behavioral fingerprints and replay challenges, it addresses the reproducibility crisis in ML, enabling trusted baselines and scalable auditing of billion-parameter models.
Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or leave the benefit of exploration unquantified. To this end, we propose Temperature-Grouped Reinforcement Learning (TGRL), which turns temperature-induced diversity into an explicit training signal. For each prompt, TGRL partitions its rollout group into low- and high-temperature subsets, estimates exploration gain through their reward contrast, and allocates this group-level signal as token-level credit using Jensen--Shannon (JS) divergence between the corresponding temperature-scaled next-token distributions induced by the same logits. Notably, TGRL reaches equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget. Across 11 benchmarks from diverse domains, TGRL broadly improves over strong RLVR baselines: it improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%. Comprehensive ablations and wall-clock analysis confirm the efficacy of all proposed components. Code is available at https://github.com/1229095296/TGRL/tree/main.
Primary: University of Chinese Academy of Sciences
All Institutions: University of Chinese Academy of Sciences, Meituan
The paper introduces TGRL, a reinforcement learning framework that leverages temperature-grouped rollouts to estimate exploration gain and allocate token-level credit via Jensen-Shannon divergence. It demonstrates that explicitly modeling the benefit of exploration through temperature contrast leads to more efficient convergence and improved performance across reasoning, coding, and agentic benchmarks compared to standard RLVR baselines.
The paper proposes TGRL, a method that integrates temperature-based exploration into the RLVR training loop by treating temperature-induced diversity as a learnable signal. The core mechanism involves partitioning rollouts into low-temperature (reference) and high-temperature (exploration) groups. The method estimates "exploration gain" via the reward contrast between these groups and uses Jensen-Shannon (JS) divergence between temperature-scaled token distributions to allocate this gain as token-level credit. The theoretical analysis provides bounds on how JS divergence captures local sensitivity to temperature changes, justifying the credit allocation strategy. The approach is logically sound, bridging the gap between sampling diversity and policy gradient updates in a principled manner.
The experimental evaluation is extensive, covering 11 benchmarks across mathematical reasoning, code generation, and agentic tasks. The model sizes tested (Qwen3-4B, 14B, 32B) are relevant to current LLM research. The results show consistent improvements over strong baselines like GRPO and DAPO, with specific gains in CodeForces rating and LiveCodeBench Pass@16. The ablation studies effectively isolate the contributions of the temperature grouping and JS-based credit allocation. The wall-clock analysis demonstrating 36% faster convergence is a strong practical contribution.
The paper provides a public GitHub repository and detailed hyperparameters (temperatures, rollout budgets, warmup steps). The use of standard models (Qwen3) and public benchmarks enhances reproducibility. However, the specific implementation details of the JS divergence calculation and the exact handling of the mixed-group advantage normalization would require careful inspection of the code to fully replicate.
The method relies on the assumption that temperature scaling is a sufficient proxy for exploration diversity. It may not capture all forms of beneficial exploration, such as those requiring semantic shifts rather than just stochastic sampling. The computational overhead of computing JS divergence for every token in the high-temperature group could be significant for very long contexts, though the paper claims efficiency gains. The performance gains, while consistent, are moderate (e.g., 1.6% on math average), which may limit its adoption if simpler baselines are sufficient for many tasks.
This work contributes to the understanding of how to efficiently utilize rollout budgets in LLM post-training. By providing a mechanism to quantify and exploit exploration gain, it offers a path toward more sample-efficient RLVR training. The insights into token-level credit allocation based on distributional sensitivity could be applicable to other areas of LLM optimization, such as test-time compute scaling. The paper introduces TGRL, a reinforcement learning framework that leverages temperature-grouped rollouts to estimate exploration gain and allocate token-level credit via Jensen-Shannon divergence. It demonstrates that explicitly modeling the benefit of exploration through temperature contrast leads to more efficient convergence and improved performance across reasoning, coding, and agentic benchmarks compared to standard RLVR baselines.
Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generation supervision improve visual understanding? We study controlled pairs of image-to-image (I2I) generation and image-to-text (I2T) understanding tasks that express the same underlying problem in different output modalities. We find that under the correct recipe, I2I training improves downstream I2T performance, with larger gains as the amount of I2I training data increases. We next ask which generation tasks benefit which understanding capabilities. To study transfer beyond paired tasks, we introduce OmniTaskonomy, a unified taxonomy spanning 19 I2I generation tasks and 25 I2T understanding capabilities. The resulting transfer map reveals selective, task-dependent benefits. Some follow intuitive correspondences, e.g., depth prediction improving metric 3D reasoning, object pointing improving counting, and jigsaw reconstruction improving 2D ordering. Interestingly, we also uncover surprising connections: 2.5D segmentation improving category recognition and Z-depth prediction improving localization. To probe these patterns, we analyze gradient alignment between generation and understanding tasks and find that stronger alignment is associated with larger downstream transfer gains. Together, our results highlight visual generation as a rich source of supervision for visual understanding and provide a roadmap for unlocking its benefits through the right training curriculum and task selection. Project page: https://omni-taskonomy.github.io/.
Primary: UC Berkeley
All Institutions: UC Berkeley, Stanford University, University of Washington, Impossible Research, Google DeepMind
The paper introduces OmniTaskonomy, a comprehensive taxonomy and empirical study demonstrating that visual generation tasks can significantly improve visual understanding capabilities through selective and task-dependent transfer. By systematically analyzing 19 generation tasks and 25 understanding capabilities, the authors reveal both intuitive and surprising connections, supported by gradient alignment analysis, providing a rigorous roadmap for leveraging generation supervision in multimodal learning.
The paper proposes a systematic framework, OmniTaskonomy, to analyze the transferability of visual generation tasks to visual understanding tasks. The methodology is rigorous, employing controlled pairs of image-to-image (I2I) generation and image-to-text (I2T) understanding tasks that share underlying semantic structures. By constructing a unified taxonomy of 19 generation tasks and 25 understanding capabilities, the authors move beyond anecdotal evidence to a comprehensive empirical study. A key methodological strength is the use of gradient alignment analysis to correlate with downstream performance gains, providing a mechanistic explanation for the observed transfer patterns.
The experimental evaluation is extensive, covering a wide range of tasks and datasets. The paper demonstrates that under specific training recipes, I2I training significantly improves I2T performance, with gains scaling with data volume. The transfer map reveals both intuitive connections (e.g., depth prediction aiding 3D reasoning) and surprising ones (e.g., 2.5D segmentation aiding category recognition). The results are consistent and well-supported by statistical analysis, providing a clear roadmap for curriculum learning in multimodal models.
The paper includes a detailed reproducibility statement, with appendices providing task definitions, training recipes, taxonomy annotations, and complete transfer results. The availability of numerical exports and figure-generation code further enhances reproducibility. The use of standard benchmarks and clear experimental protocols ensures that the findings can be verified by the community.
The study is primarily focused on image-based tasks and may not fully generalize to video or other modalities. The analysis relies on specific model architectures and training setups, which might limit the generalizability of the findings to other model families. Additionally, the computational cost of training and evaluating such a large number of task pairs is significant, which could be a barrier for some researchers.
The findings have significant implications for the design of multimodal learning systems, suggesting that visual generation can serve as a powerful pre-training signal for visual understanding. This could lead to more efficient and effective training curricula for large multimodal models. The paper also provides a valuable resource for the community in the form of the OmniTaskonomy taxonomy, which can be used to guide future research on task transfer and curriculum learning. The paper introduces OmniTaskonomy, a comprehensive taxonomy and empirical study demonstrating that visual generation tasks can significantly improve visual understanding capabilities through selective and task-dependent transfer. By systematically analyzing 19 generation tasks and 25 understanding capabilities, the authors reveal both intuitive and surprising connections, supported by gradient alignment analysis, providing a rigorous roadmap for leveraging generation supervision in multimodal learning.
World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model. To answer this question, we introduce V-JEPA Policy, a simple framework that builds a WAM on the latent space of a frozen V-JEPA 2.1 encoder. An instruction-conditioned future-latent predictor and a flow-matching action expert are jointly learned from scratch in a single downstream stage, with the predictor's future-informed context key--value states conditioning action generation. With 0.9B total parameters, of which 0.6B are trainable, V-JEPA Policy achieves competitive performance with representative WAM and vision-language-action baselines across LIBERO, LIBERO-Plus, and RoboCasa-GR1. Comparing visual foundations under the same downstream framework and training budget identifies V-JEPA latents as more effective than the discriminative, reconstructive, and video-understanding-oriented alternatives, particularly under distribution shifts. Beyond task-specific learning, pretraining the predictor on DROID video--instruction pairs without action labels and adapting it into a WAM yields substantial gains in downstream control and out-of-distribution generalization. Together, these findings establish predictive visual latents as a foundation for effective WAM learning from task-specific demonstrations and for transferring future-modeling knowledge acquired from broader in-the-wild videos. Our code is available at https://github.com/breez3young/VJEPA-Policy.
Primary: Tsinghua University
All Institutions: Tsinghua University, Shanghai Jiao Tong University, Fudan University, University of Science and Technology of China, Washington University in St. Louis, The Institute of Artificial Intelligence, China Telecom (TeleAI)
V-JEPA Policy demonstrates that frozen predictive visual latents are a sufficient foundation for effective world-action models, outperforming generative and discriminative alternatives under distribution shifts. The paper rigorously validates this approach through extensive simulation benchmarks and real-world manipulation, highlighting the critical role of predictor pretraining on video-instruction pairs for robust generalization.
The paper proposes V-JEPA Policy, a framework that decouples the predictive visual latent space from the generative model. Instead of fine-tuning a large video generator, it freezes the V-JEPA 2.1 encoder and trains a lightweight instruction-conditioned future-latent predictor and a flow-matching action expert from scratch. The key architectural innovation is the use of the predictor's future-informed context key-value states to condition the action generation, effectively creating a world-action model (WAM) that relies on predictive latents rather than pixel-space reconstruction or full video generation. This approach is computationally efficient (0.9B total params, 0.6B trainable) and leverages the robustness of predictive pretraining.
The experiments are extensive, covering LIBERO, LIBERO-Plus, and RoboCasa-GR1 benchmarks. The paper provides rigorous ablations comparing V-JEPA latents against discriminative, reconstructive, and video-understanding-oriented alternatives, showing superior performance under distribution shifts. A significant finding is the benefit of pretraining the predictor on DROID video-instruction pairs without action labels, which yields substantial gains in downstream control and out-of-distribution generalization that cannot be recovered by simply training longer from scratch. The results are competitive with state-of-the-art WAMs and VLA baselines.
The authors provide a public GitHub repository with code. The paper details the training setup, including the separation of frozen and trainable parameters, and the specific pretraining data (DROID). The use of standard benchmarks and open-source models (V-JEPA 2.1) enhances reproducibility.
The method relies heavily on the quality of the V-JEPA 2.1 encoder; if the encoder fails to capture relevant dynamics, the policy will suffer. The flow-matching action expert is trained from scratch, which may require careful tuning. The paper focuses on simulation and limited real-world bimanual manipulation; broader real-world deployment across diverse environments is not fully explored. The reliance on DROID for pretraining limits the immediate applicability to domains similar to DROID.
This work provides a compelling argument for using predictive latents as a foundation for robot learning, potentially reducing the computational cost of training WAMs. It offers a pathway to leverage large-scale video pretraining without the burden of generative model fine-tuning. The findings on the transferability of future-modeling knowledge from in-the-wild videos to control tasks are significant for the robotics community. V-JEPA Policy demonstrates that frozen predictive visual latents are a sufficient foundation for effective world-action models, outperforming generative and discriminative alternatives under distribution shifts. The paper rigorously validates this approach through extensive simulation benchmarks and real-world manipulation, highlighting the critical role of predictor pretraining on video-instruction pairs for robust generalization.
Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.
Primary: University of California, Los Angeles (UCLA)
All Institutions: University of California, Los Angeles (UCLA)
The paper introduces a novel RL framework for unified multimodal models that jointly optimizes reflection and generation, achieving significant improvements in image generation accuracy through self-correction. By leveraging group-relative advantage estimation and graded rewards, the method effectively addresses the sparsity of feedback in compositional generation tasks, demonstrating that unified models can learn to diagnose and repair their own errors without external critics, with gains that transfer robustly to unseen benchmarks.
The paper proposes UMM-Reflection, a framework for applying Reinforcement Learning (RL) to unified multimodal models (specifically BAGEL) to improve image generation via self-reflection. The core methodological contribution is the joint optimization of reflection text (diagnosis) and image generation (repair) within a single model using group-relative advantage estimation. By sampling sibling trajectories from the same initial image, the method avoids the combinatorial explosion of per-round credit assignment and allows credit to flow across the entire reflection loop. A significant technical detail is the "graded reward" mechanism, which addresses the sparsity of binary GenEval scores by providing partial credit for partial constraint satisfaction, thereby stabilizing RL training. The approach is theoretically sound, leveraging the unified nature of the model to eliminate the need for external critics or separate verifier models at inference time.
The experimental evaluation is rigorous and comprehensive. The paper reports a substantial improvement of +12.05 points on GenEval over the SFT baseline. Crucially, the authors demonstrate that these gains transfer to unseen benchmarks (WISE, OneIG-Bench, T2I-CompBench++), suggesting the model has learned generalizable repair strategies rather than overfitting to the training distribution. Ablation studies are thorough, including comparisons against direct T2I-RL (without reflection), external critic pipelines (GPT-5.5), and different reward shaping strategies (penalty vs. bonus for stopping). The analysis of "pass@16" vs "pass@1" highlights that the SFT model already contains the capability for correct repairs, and RL effectively selects and reinforces these high-success paths. The visual pathway stability analysis confirms that RL primarily modifies the decision-making (text) pathway rather than disrupting the underlying visual generation capabilities.
The paper provides high reproducibility standards. It details the specific training hyperparameters (learning rates, batch sizes, number of updates), the composition of the RL prompt pool, and the exact reward calculation formulas. The authors explicitly state that the code and prompt pool will be released. The use of standard benchmarks (GenEval, WISE, etc.) and clear evaluation protocols (50 denoising steps, specific resolutions) allows for easy comparison with other works. The distinction between controlled evaluations and native leaderboard protocols is clearly defined, preventing misinterpretation of results.
The primary limitation is the reliance on the BAGEL model architecture; while the method is general, the specific implementation details (e.g., MoT decoder layers, flow-based revisions) are tied to this unified model structure. The counting category in GenEval shows no improvement, indicating that current reflection strategies are insufficient for precise numerical constraints. Additionally, the method requires significant computational resources (16 H100 GPUs for 33 hours) for the RL phase, which may limit adoption for smaller labs. The evaluation is limited to 512x512 resolution for training and evaluation, whereas the model natively supports 1024x1024, leaving open questions about performance at higher resolutions.
This work has significant implications for the development of autonomous multimodal agents. By demonstrating that a single model can effectively self-correct its generations through RL, it paves the way for more robust and reliable generative systems that do not rely on external, potentially misaligned, critic models. The insight that RL can "select" correct repairs from an existing SFT distribution rather than generating them from scratch is a valuable finding for the broader field of generative AI. The transferability of gains to unseen benchmarks suggests that self-reflection is a generalizable capability, potentially applicable to other modalities or tasks. The paper introduces a novel RL framework for unified multimodal models that jointly optimizes reflection and generation, achieving significant improvements in image generation accuracy through self-correction. By leveraging group-relative advantage estimation and graded rewards, the method effectively addresses the sparsity of feedback in compositional generation tasks, demonstrating that unified models can learn to diagnose and repair their own errors without external critics, with gains that transfer robustly to unseen benchmarks.
Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterized. To fill this gap, we compare scaling laws for encoder-free and encoder-based MLLMs and report three main findings: (1) Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged. (2) The two architectures exhibit nearly overlapping loss--compute frontiers on the text objective, but diverge on the multimodal objective: encoder-free models underperform at small scales yet are predicted to catch up at around $10^{22}$ FLOPs, well within practical pretraining budgets. (3) Without a visual encoder, the language model learns to take over its role via vision-specific adaptation: bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated. Overall, our results indicate that the advantage of the visual prior provided by a pretrained encoder diminishes with scale, positioning encoder-free architectures as a promising direction for multimodal pretraining.
Primary: Tencent
All Institutions: Tencent
The paper systematically characterizes the scaling laws of encoder-free MLLMs, predicting they will match encoder-based models at $10^{22}$ FLOPs. It provides a rigorous comparison using a matched model ladder and offers mechanistic insights into how decoders compensate for the lack of a visual encoder, positioning encoder-free architectures as a promising direction for future multimodal pretraining.
The paper employs a rigorous scaling law framework to compare encoder-free and encoder-based Multimodal Large Language Models (MLLMs). The methodology is sound, utilizing a matched ladder of 11 sparse MoE models (1.1B-44B) to isolate the effect of the visual encoder. The use of IsoFLOP profiles and compute-optimal allocation analysis is standard and well-executed. The introduction of "vision-specific adaptation" probes (attention patterns, layerwise representation evolution, expert routing) to mechanistically explain the scaling differences is a strong methodological contribution that goes beyond simple loss comparison. The derivation of the overtraining loss equation is mathematically sound and provides a useful tool for predicting efficiency gains under non-optimal training regimes.
The experimental setup is robust, covering a wide range of model scales and compute budgets ($10^{19}$ to $10^{21}$ FLOPs). The comparison is controlled by fixing the visual encoder size and data mixture. The results are consistent, showing that while encoder-free models lag at small scales, their loss decreases more rapidly with compute, predicting a crossover at $10^{22}$ FLOPs. The analysis by topic (STEM vs. Perception) provides valuable nuance, showing that the crossover is earlier for language-heavy tasks. The inclusion of downstream benchmark evaluations (CV-Bench, ChartQA, etc.) further validates the scaling trends, although the primary focus remains on validation loss.
The paper provides detailed implementation details in the appendix, including the specific architecture of the front ends, the MoE topology, and the training hyperparameters. The use of standard components (SigLIP 2, Muon optimizer) and clear descriptions of the data mixture enhance reproducibility. However, the specific data sources are not fully detailed, which may limit exact replication. The code is not explicitly linked in the provided text, but the level of detail suggests it is likely available or easily implementable.
The primary limitation is the reliance on extrapolation. The predicted crossover at $10^{22}$ FLOPs is well beyond the measured range ($10^{21}$ FLOPs), introducing uncertainty. The assumption that the visual encoder size remains fixed as the decoder scales is a simplification; in practice, joint scaling of the encoder and decoder might alter the dynamics. Additionally, the study focuses on a specific data mixture (1:1 text/multimodal), and results may vary with different data compositions. The "catch-up" prediction assumes that the irreducible loss floor is shared, which is a reasonable but unproven assumption at these scales.
This paper has significant implications for the design of future multimodal models. By demonstrating that the advantage of pretrained visual encoders diminishes with scale, it encourages the exploration of unified, encoder-free architectures that may be simpler and more efficient in the long run. The insights into how decoders adapt to take over visual encoding (e.g., bidirectional attention, early-layer processing) could inform the design of new decoder architectures specifically optimized for native multimodal learning. This work could shift the field's focus from improving visual encoders to optimizing the language model's ability to process raw visual tokens. The paper systematically characterizes the scaling laws of encoder-free MLLMs, predicting they will match encoder-based models at $10^{22}$ FLOPs. It provides a rigorous comparison using a matched model ladder and offers mechanistic insights into how decoders compensate for the lack of a visual encoder, positioning encoder-free architectures as a promising direction for future multimodal pretraining.
Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tension remains underexplored: the alignment-faithfulness conflict. We show that aligned models systematically deviate from their inputs on unsafe or sensitive content without disclosing the modification, a failure mode we call alignment-induced unfaithfulness (AIU). Unlike capability-driven unfaithfulness, which comes from errors in knowledge or reasoning, this is induced by post-training mechanisms that override adherence to the input. We introduce FaithConflict, a controlled dataset isolating both conflicts, and two complementary taxonomies: behavioral (B1-B8) and chain-of-thought reasoning (C0-C6). Across models, AIU increases with scale and more sharply than capability-driven unfaithfulness, a reverse scaling law; intermediate checkpoints show it is amplified during post-training, with DPO the stage at which the gap both grows most and becomes least visible. Prompting-based mitigation does not resolve it, revealing a capability-alignment-faithfulness trilemma in the design and evaluation of LLMs.
Primary: University of Illinois Urbana-Champaign (UIUC)
All Institutions: University of Illinois Urbana-Champaign (UIUC)
The paper identifies and characterizes "alignment-induced unfaithfulness," a critical failure mode where aligned LLMs silently override source content on sensitive topics, revealing a fundamental capability-alignment-faithfulness trilemma. Through rigorous controlled experiments, a novel dataset, and causal analysis of post-training stages, the authors demonstrate that this unfaithfulness is a systematic, emergent property of modern alignment techniques (particularly DPO) that worsens with model scale, posing significant risks to the reliability of LLMs in source-reporting applications.
The paper introduces a rigorous controlled experimental design to isolate the "alignment-faithfulness" conflict, distinct from capability-faithfulness or capability-alignment tradeoffs. By constructing the FaithConflict dataset with paired confirming/opposing claims in identical templates, the authors effectively control for surface form and domain, allowing the FaithGap metric to directly attribute unfaithfulness to the model's internal conflict with the source content. The dual taxonomy (B1-B8 for outputs, C0-C6 for reasoning) provides a granular lens into *how* models fail, distinguishing between visible refusals and dangerous silent inversions. The methodology is sound, though it relies heavily on an LLM judge (Qwen-2.5 32B) for annotation, which introduces a potential circularity if the judge shares similar alignment biases, though high inter-annotator agreement with humans mitigates this concern.
The experimental scope is impressive, covering 22 checkpoints across 8 model families (including frontier models like Claude Sonnet 4.6 and GPT-4o). The discovery of a "reverse scaling law"—where larger, more aligned models exhibit *greater* unfaithfulness on conflicting sources—is a significant and counter-intuitive finding. The stage-by-stage analysis pinpointing DPO as the primary driver of this behavior is particularly valuable for practitioners. The inclusion of causal interventions (removing safety data from post-training) strengthens the claim that alignment training, not just scale, is the root cause. However, the evaluation is limited to summarization and a few other formats (QA, NLI), and the reliance on a single judge model for all annotations is a notable weakness.
The paper provides high reproducibility standards. Code, data, and project pages are linked. The authors release the FaithConflict dataset and detailed prompts for both task execution and judging. The post-training intervention experiments use the public Tulu 3 pipeline, allowing others to replicate the causal analysis. The only minor gap is the lack of dated snapshot identifiers for frontier models, which is acknowledged by the authors.
The primary limitation is the reliance on an LLM judge for the core metric (FaithGap), which may not perfectly capture human perception of faithfulness. The study is also limited to source-reporting tasks (summarization, extraction) and does not explore whether this unfaithfulness manifests in more complex agentic or multi-turn dialogue settings. The causal attribution to DPO is based on a specific training pipeline (Tulu 3) and may not generalize to all preference optimization algorithms.
This paper has high impact on the LLM safety and evaluation community. It identifies a critical blind spot in current alignment practices: models are becoming better at "safety" but worse at "faithfulness" in a way that is invisible to standard benchmarks. This has immediate implications for RAG systems, clinical note extraction, and legal document processing, where silent modification of source content is a severe failure mode. The "trilemma" framework will likely influence future alignment research to explicitly optimize for faithfulness alongside safety and capability. The paper identifies and characterizes "alignment-induced unfaithfulness," a critical failure mode where aligned LLMs silently override source content on sensitive topics, revealing a fundamental capability-alignment-faithfulness trilemma. Through rigorous controlled experiments, a novel dataset, and causal analysis of post-training stages, the authors demonstrate that this unfaithfulness is a systematic, emergent property of modern alignment techniques (particularly DPO) that worsens with model scale, posing significant risks to the reliability of LLMs in source-reporting applications.
The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound along an acoustic tube model of the vocal tract, and via its gradients, can solve the inverse problem: reconstructing the shape of the vocal tract solely from the sound it produces. Although the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural fields, we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel linguistics and medical imaging applications. (1) We demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI (magnetic resonance imaging) images to reconstruct one's moving vocal tract from only their speech without paired data.
Primary: Massachusetts Institute of Technology (MIT)
All Institutions: Massachusetts Institute of Technology (MIT)
[One sentence main contribution]. The paper presents a differentiable and GPU-accelerated acoustic simulator for the vocal tract that enables the reconstruction of vocal tract shapes from speech, with applications in linguistics and medical imaging. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The technical contribution is significant due to the novel frequency-domain formulation and the integration of turbulence modeling. The methodology is well-designed, leveraging neural networks to overcome the non-convexity of the inverse problem. The significance to the field is high, as it bridges the gap between physical simulation and deep learning, enabling new capabilities in speech analysis and medical imaging.
The paper introduces a differentiable acoustic simulator for the vocal tract, which is a significant technical achievement. The core innovation lies in the frequency-domain formulation of fluid dynamics, which is claimed to be 70x more GPU-parallelizable than time-domain finite differences. This is a crucial engineering insight for enabling real-time or near-real-time gradient-based optimization. The integration of a differentiable turbulence model for consonants is a strong addition, as consonants are notoriously difficult to model with simple tube acoustics. The use of neural networks to parameterize the geometry (similar to INRs) is a clever application of existing techniques to a new domain, helping to escape local minima in the non-convex inverse problem. The methodology is sound and well-motivated, combining physical simulation with modern deep learning techniques.
The experiments demonstrate the simulator's ability to reconstruct vocal tract shapes from speech across 11 languages, which is a robust test of generalizability. The self-supervised autoencoding task shows the model can learn meaningful representations without labeled data. The most impressive result is the coupling with a generative MRI model to reconstruct moving vocal tracts from speech alone, without paired data. This is a novel application that bridges speech processing and medical imaging. However, the evaluation of the MRI reconstruction quality is likely limited by the lack of ground truth for moving vocal tracts, and the paper should provide more quantitative metrics on the fidelity of the reconstructed shapes compared to actual MRI scans.
The paper provides a supplementary material link, which is a good sign for reproducibility. The frequency-domain formulation and turbulence model are described in detail, but the specific implementation details for the neural network parameterization and the MRI generative model coupling would be critical for reproduction. The claim of 70x speedup should be backed by detailed benchmarking against standard time-domain solvers.
The acoustic tube model is a simplification of the actual vocal tract, which is a 3D structure. The model may not capture all the nuances of speech production, especially for complex consonants or pathological speech. The MRI reconstruction is a novel application, but the clinical utility of the reconstructed shapes is not evaluated. The paper also does not discuss the computational cost of the MRI generative model coupling, which could be a bottleneck.
This work has significant potential impact in both linguistics and medical imaging. In linguistics, it could enable new insights into speech production and variation across languages. In medical imaging, it could lead to new methods for diagnosing speech disorders or planning surgical interventions. The differentiable simulator could also be used in other domains where inverse problems with physical constraints are important. [One sentence main contribution]. The paper presents a differentiable and GPU-accelerated acoustic simulator for the vocal tract that enables the reconstruction of vocal tract shapes from speech, with applications in linguistics and medical imaging. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The technical contribution is significant due to the novel frequency-domain formulation and the integration of turbulence modeling. The methodology is well-designed, leveraging neural networks to overcome the non-convexity of the inverse problem. The significance to the field is high, as it bridges the gap between physical simulation and deep learning, enabling new capabilities in speech analysis and medical imaging.
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.
Primary: Unknown (Likely ByteDance based on "YuE" naming convention and technical style, but not explicitly stated in provided text)
All Institutions: Unknown
YuE2 unifies symbolic and audio music generation through a single Mixture-of-Transformers model that employs symbolic planning to improve musical quality and enable editable composition. The paper demonstrates that explicitly generating a readable score before audio realization leads to significant improvements in perceived musicality and structural coherence, while also introducing state-of-the-art music understanding and transcription models (MERT2 and SheetSage2) that facilitate this process.
The paper proposes YuE2, a unified framework that integrates symbolic music generation (ABC notation) with audio generation (flow matching) within a single Mixture-of-Transformers (MoT) architecture. The core methodological contribution is the "symbolic planning" step, where the model first generates a readable score (melody, harmony, form) before expanding it into semantic tokens and acoustic latents. This is supported by two auxiliary models: MERT2, a music representation learner that sets new SOTA on MARBLE benchmarks, and SheetSage2, a full-song transcription model that generates the symbolic supervision signals required for training. The use of an AR-NAR MoT to handle both discrete symbolic/semantic tokens and continuous acoustic latents is a sophisticated architectural choice that allows for bidirectional attention in the acoustic stream while maintaining causal generation for the score.
The evaluation is extensive, covering automatic metrics (SongBench, SongEval, AudioBox) and expert listening tests. YuE2 outperforms public baselines and is competitive with proprietary systems like Suno v4.5/v5. The ablation study on symbolic planning is particularly strong, demonstrating that generating the score first significantly improves perceived musicality and overall quality compared to direct audio generation. The score-editing experiments show that the model can preserve unedited content while modifying specific sections, validating the utility of the symbolic interface.
The paper provides detailed architectural descriptions and training procedures. However, as a technical report from a likely industry lab, the full code and weights may not be immediately available to the public, which limits immediate reproducibility. The reliance on proprietary or large-scale datasets (346,000 hours of music) also poses a barrier for independent replication.
The model is large (3.58B parameters) and computationally expensive. The evaluation against proprietary systems is limited to expert listening and some automatic metrics, as direct API access for rigorous benchmarking is often restricted. The "best-of-8" selection strategy, while effective for benchmarking, may not reflect real-time interactive use cases where latency is critical.
This work bridges the gap between symbolic composition and audio production, offering a new paradigm for music generation that allows for explicit control over musical structure. The introduction of MERT2 and SheetSage2 as high-quality supervision tools will likely benefit the broader music AI community. The agentic editing capability suggests potential applications in professional music production workflows. YuE2 unifies symbolic and audio music generation through a single Mixture-of-Transformers model that employs symbolic planning to improve musical quality and enable editable composition. The paper demonstrates that explicitly generating a readable score before audio realization leads to significant improvements in perceived musicality and structural coherence, while also introducing state-of-the-art music understanding and transcription models (MERT2 and SheetSage2) that facilitate this process.
Equipping artificial agents with spatial intelligence requires a comprehensive generative prior over the dynamic 3D world. We propose World Motion Models (WMMs) that capture "what was, is, and will be where across time" via sparse SE(3) pose trajectories. WMMs are built on the observation that elements of dynamic scenes can be well approximated by a set of rigid SE(3) trajectories, a minimal yet expressive primitive for 4D modeling. This representation unifies articulated objects, human bodies, hand-object interactions, piecewise-rigid scene dynamics, camera motion, and even robot states and actions into a single shared space. Given this representation, we cast the joint distribution of these entities as a flexible sequence modeling problem, utilizing flow-matching with per-token noise levels. Coupled with a context token mechanism for non-sequential conditioning, this formulation supports any-to-any marginal conditioning across an arbitrary number of entities and time steps. Tasks such as future prediction, motion infilling, model-predictive control, inverse kinematics, cross-embodiment retargeting, and policy learning all reduce to the application of different masks over the same network. Experiments on 6 diverse applications of 3D vision and robotics demonstrate the versatility and flexibility of WMMs with strong performance.
Primary: UC Berkeley
All Institutions: UC Berkeley, Meta AI
World Motion Models introduces a unified flow-matching framework for SE(3) trajectory modeling that enables flexible any-to-any conditioning across diverse robotics and vision tasks. The paper demonstrates that a single generative model can handle prediction, control, and retargeting through simple masking, offering a powerful and versatile primitive for spatial intelligence in artificial agents.
The paper proposes World Motion Models (WMMs), a unified framework for modeling 4D dynamics using sparse SE(3) pose trajectories. The core methodological contribution is the recasting of joint distribution modeling over multiple entities (humans, objects, cameras, robots) as a flexible sequence modeling problem using flow-matching. By employing per-token noise levels and a context token mechanism, the authors enable "any-to-any" marginal conditioning. This allows a single network to handle diverse tasks—such as prediction, infilling, and control—simply by applying different masks to the input sequence. The choice of SE(3) trajectories as a primitive is elegant, providing a minimal yet expressive representation that unifies articulated motion, rigid body dynamics, and camera motion into a shared latent space.
The paper reports experiments on six diverse applications spanning 3D vision and robotics, including future prediction, motion infilling, model-predictive control (MPC), inverse kinematics (IK), cross-embodiment retargeting, and policy learning. The breadth of tasks demonstrates the versatility of the unified approach. While the abstract claims "strong performance," the lack of specific quantitative metrics in the provided text prevents a full assessment of superiority over specialized baselines. However, the ability to switch tasks via masking without retraining is a significant empirical advantage over task-specific architectures.
The project page is provided, which likely contains code and datasets. The use of standard flow-matching techniques and SE(3) representations suggests that the core components are reproducible, provided the specific masking strategies and context token implementations are detailed in the full paper (which is not fully visible here, but implied by the "Spotlight" status and project page).
The reliance on SE(3) trajectories assumes that scene elements can be well-approximated by rigid motions. This may limit applicability to highly deformable objects or non-rigid interactions that cannot be decomposed into rigid parts. Additionally, the computational cost of flow-matching on long sequences of high-dimensional SE(3) poses could be a bottleneck for real-time applications, although the paper mentions MPC, suggesting some efficiency gains.
This work has the potential to significantly impact the field of embodied AI and 4D scene understanding. By unifying various robotics and vision tasks under a single generative framework, it simplifies the development of general-purpose agents. The "any-to-any" conditioning capability is particularly valuable for data-efficient learning and sim-to-real transfer, where conditioning on partial observations or specific actions is common. World Motion Models introduces a unified flow-matching framework for SE(3) trajectory modeling that enables flexible any-to-any conditioning across diverse robotics and vision tasks. The paper demonstrates that a single generative model can handle prediction, control, and retargeting through simple masking, offering a powerful and versatile primitive for spatial intelligence in artificial agents.
Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction videos via video-to-video (V2V) generation from a few exemplar real videos. Our contact-anchored real-to-sim pipeline then reconstructs both human and object motions, retargeting this imperfect video data into physically plausible trajectories. The intra-class variability across these counterfactual videos lets us train a single policy that generalizes to unseen objects within each category. We demonstrate the full pipeline by deploying this policy on a real robot without any real-world fine-tuning. Using only onboard depth observations, our humanoid picks up, carries, and drops objects, including boxes, barrels, bins, and balls, across novel instances, sizes, and initial configurations.
Primary: Unknown (Likely NVIDIA or similar top lab based on "SeedDance" and hardware, but affiliations are redacted in provided text)
All Institutions: Unknown
[One sentence main contribution]. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper introduces a novel data augmentation strategy using counterfactual video generation to enable scalable humanoid loco-manipulation, demonstrating a robust real-to-sim-to-real pipeline that achieves zero-shot generalization to diverse objects, representing a significant step towards generalist robot learning from visual data.
The paper proposes PRISM, a framework that leverages video-to-video (V2V) generative models to create "counterfactual" human-object interaction videos from a small set of real seed videos. The core innovation lies in using these generated videos to train a real-to-sim-to-real pipeline. The methodology involves a contact-anchored reconstruction stage that uses human-object contact constraints to regularize monocular reconstruction errors, followed by a retargeting stage that uses contact anchors to guide the generation of physically plausible robot trajectories. The policy is trained via a privileged teacher-student distillation process, where the teacher tracks full-state references and the student learns from depth observations and joystick commands. The use of V2V generation to augment data with behavior-level variations (not just geometric) is a strong conceptual contribution, addressing the data scarcity bottleneck in visual imitation learning.
The experiments demonstrate zero-shot sim-to-real transfer on a Unitree G1 humanoid. The robot successfully picks up, carries, and drops diverse objects (boxes, barrels, bins, balls) and generalizes to unseen categories (chairs, tables, lamps). The success rates are high for in-domain objects (80-100%) and reasonable for out-of-domain objects (60-100%). The ablation study effectively shows that V2V generation outperforms simple geometric augmentation and that the contact-anchored reconstruction/retargeting is crucial for handling reconstruction noise. The comparison with OMOMO data highlights the superior generalization of the PRISM-generated data.
The paper provides detailed hyperparameters, reward structures, and pipeline runtime estimates. However, the reliance on specific proprietary video generation models (SeedDance 2.0) and specific reconstruction backends (CRISP, SAM3D) may limit immediate reproducibility for groups without access to these tools. The codebase is referenced but not explicitly linked in the text provided, though a project page is available.
The pipeline is computationally expensive (approx. 10 minutes per 8-second clip for reconstruction/retargeting). The method assumes rigid-body dynamics for objects, limiting applicability to deformable or articulated objects. The generalization to out-of-domain objects is good but not perfect, with some failures on complex shapes like chairs. The data scale is still relatively small (256 generated videos), and scaling behavior is not fully characterized.
This work bridges the gap between generative AI and robotics, offering a scalable path to collecting diverse interaction data without extensive real-world recording. It has significant implications for the development of generalist humanoid robots that can interact with a wide variety of household objects. The approach could be extended to other manipulation tasks and robot morphologies. [One sentence main contribution]. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper introduces a novel data augmentation strategy using counterfactual video generation to enable scalable humanoid loco-manipulation, demonstrating a robust real-to-sim-to-real pipeline that achieves zero-shot generalization to diverse objects, representing a significant step towards generalist robot learning from visual data.
A robot should be able to learn through experiments how unfamiliar objects behave and interact, then plan with that knowledge. It need not start from scratch: physics engines supply knowledge of motion and contact, but can omit entire mechanisms, such as glue curing, water heating, or wind. We present EMPIRIC, an agent that learns a residual world model: a physics engine extended with code for the missing mechanisms. The learned programs can introduce new forces, constraints, and hidden state, and Bayesian inference estimates their parameters and states from noisy observations. The resulting model lets the agent predict the outcomes of actions, choose informative experiments, and revise its hypotheses when predictions fail. Across five simulated domains, EMPIRIC learns interpretable, reusable models, and solves more tasks with fewer environment interactions than all three baselines. On a physical robot, it learns wind forces and domino masses to solve a manipulation task. Website and code: https://yichao-liang.github.io/empiric
Primary: Harvard University
All Institutions: Harvard University, Massachusetts Institute of Technology, Princeton University, The Alan Turing Institute
The paper presents EMPIRIC, a novel framework for robot planning that learns residual world models by extending physics engines with executable code for missing mechanisms, achieving superior task success and sample efficiency through Bayesian inference and LLM-driven program synthesis.
The paper introduces EMPIRIC, a framework for learning "residual world models" by extending a base physics engine (PyBullet) with executable Python code for missing physical mechanisms (e.g., glue curing, wind forces). The core innovation is the separation of the base simulator (handling rigid-body dynamics) from the residual program (handling novel interactions and hidden state). The method employs Bayesian inference to estimate parameters and hidden states from noisy observations, using a mean-field approximation for the posterior. It integrates planning and information seeking by simulating candidate actions under multiple parameter draws to maximize success probability or mutual information. The use of a coding agent (LLM) to write and revise the residual code is a significant architectural choice, leveraging LLMs for program synthesis rather than just policy generation.
The evaluation is rigorous, covering five simulated domains (Domino, Bridge, Balloons, Boil, Fan) and one physical robot experiment. The protocol is a continual learning setup where the agent must solve training and test tasks within a step budget. EMPIRIC achieves 100% success rate in simulation, outperforming baselines like Direct Agent, Direct + Scene, and Standalone Sim. The ablation studies effectively isolate the contributions of parameter fitting and explicit uncertainty handling, showing that uncertainty is critical for irreversible actions (e.g., balloon bursting). The physical robot experiment demonstrates real-world applicability by learning wind forces and domino masses from two gusts to plan a cascade.
The paper provides extensive details in the appendix, including the agent's workspace, tools, simulator subclass interface, and belief construction algorithms. The code and website are provided. The use of a specific LLM (Claude Opus 5) is noted, which may limit immediate reproducibility if access to that specific model version is restricted, but the framework is sufficiently detailed for implementation.
The agent relies on predefined object features and does not learn feature extraction from raw images. The computational cost is high (median 107 minutes per run), which is a significant barrier to real-time deployment. The base simulator is given, so the method does not address scene reconstruction from scratch. The reliance on a powerful LLM for code generation introduces potential biases and costs associated with API usage.
This work bridges the gap between symbolic program synthesis and physical robotics, offering a path toward robots that can adapt to novel physical environments without retraining neural networks for every new object. The concept of "residual world models" is likely to influence future research in hybrid simulation and model-based reinforcement learning. The integration of Bayesian uncertainty with code-based models provides a robust framework for safe decision-making in uncertain physical environments. The paper presents EMPIRIC, a novel framework for robot planning that learns residual world models by extending physics engines with executable code for missing mechanisms, achieving superior task success and sample efficiency through Bayesian inference and LLM-driven program synthesis.
Little is known about how to manually design agents capable of dexterous manipulation. Some design principles have been inferred from close examination of how animals manipulate objects, but these structures and behaviors have so far resisted biomimicry and may not be optimal for artificial machines. Here we evolve freeform robots to pick up, hold, rotate, and use diverse objects. Unlike other approaches to optimizing robot hands, we do not presuppose the presence, articulation, or geometry of any part of the body. Although familiar prehensile forms such as tails, beaks, paws and claws may emerge spontaneously under certain conditions--and while such conditions could be of interest to evolutionary biologists--de novo manipulator design can also reveal whole new solutions, overlooked or unknown structures which may be better suited for the task at hand. We use contrastive learning to create a highly searchable genetic embedding of design space, an autoregressive developmental model to decode designs, evolutionary strategies to find good designs, and reinforcement learning to train each evolved design. Winning designs were automatically converted into a manufacturable blueprint, printed, assembled and tested in the real world in a zero-shot manner. The results represent the state-of-the-art in evolutionary robotics in terms of performance, diversity and complexity.
Primary: University of Oxford
All Institutions: University of Oxford, DeepMind
The paper presents a state-of-the-art framework for evolving freeform dexterous robots from scratch, combining contrastive learning, developmental models, and reinforcement learning to produce manufacturable, real-world capable agents. By removing prior assumptions on robot morphology, the work reveals novel, non-biomimetic structures that outperform traditional designs in dexterity and diversity, marking a significant leap in evolutionary robotics and automated hardware design.
The paper proposes a comprehensive framework for de novo robot design, combining contrastive learning for genetic embedding, autoregressive developmental models for decoding, evolutionary strategies for search, and reinforcement learning for control. The novelty lies in the lack of prior assumptions about morphology (no fixed joints or geometry), allowing for truly freeform evolution. The use of a developmental model to map latent codes to physical structures is a sophisticated approach that bridges high-level design search with low-level manufacturability.
The evaluation is rigorous, moving beyond simulation to real-world validation. The robots were 3D printed, assembled, and tested in a zero-shot manner, demonstrating the robustness of the evolved designs. The paper reports state-of-the-art performance in terms of dexterity, diversity, and complexity compared to previous evolutionary robotics works. The inclusion of diverse object manipulation tasks (pick, hold, rotate, use) provides a strong benchmark for the method's generality.
The paper details the pipeline components (contrastive learning, autoregressive model, ES, RL) sufficiently for experts to replicate the framework. However, the specific hyperparameters for the evolutionary search and the exact architecture of the developmental model are likely in the appendix or code, which is standard. The zero-shot real-world testing is a high bar for reproducibility that the authors have met by providing the blueprint conversion process.
The computational cost of evolving and training multiple robots is significant. The "zero-shot" real-world performance, while impressive, may still be limited in speed or precision compared to hand-tuned industrial manipulators. The generalizability to tasks requiring extremely high force or precision (beyond what 3D printed materials can handle) is a potential constraint.
This work has significant implications for the field of evolutionary robotics and automated design. It suggests that optimal manipulator structures may not be biomimetic, challenging existing design paradigms. The ability to automatically generate manufacturable blueprints from abstract design spaces could accelerate the development of specialized robotic tools for manufacturing, surgery, or exploration. The paper presents a state-of-the-art framework for evolving freeform dexterous robots from scratch, combining contrastive learning, developmental models, and reinforcement learning to produce manufacturable, real-world capable agents. By removing prior assumptions on robot morphology, the work reveals novel, non-biomimetic structures that outperform traditional designs in dexterity and diversity, marking a significant leap in evolutionary robotics and automated hardware design.
We present ZeroBot, a real2sim framework for learning a robot manipulation task from scratch in minutes under challenging conditions: zero human demonstrations, zero policy pre-training, and zero known object models. Given only a single view of an object and a goal pose for that object, ZeroBot uses image-to-3D generative models to obtain a complete object mesh, which is used in simulation for large-scale parallel reinforcement learning. To accelerate training, we introduce an action space which leverages the generated geometry and learned value function to sample states involving robot-object contact. When evaluated on real-world tasks including grasping, pushing, articulated object interaction, and multi-stage manipulation, ZeroBot achieves an 87% success rate with an average training time of 119 seconds. These results show the value of using image-to-3D models in a real2sim framework for rapid, autonomous robot learning.
Primary: Imperial College London
All Institutions: Imperial College London, Robotics and AI Institute
ZeroBot introduces a generative real2sim framework that enables robots to learn manipulation tasks from scratch in minutes using image-to-3D models and a novel contact-based action space. The paper demonstrates that combining visual foundation models with parallel RL can significantly reduce the time and data required for robotic learning, achieving high success rates on diverse real-world tasks without human demonstrations or pre-trained policies.
The paper proposes ZeroBot, a framework that integrates image-to-3D generative models (specifically InstantMesh) with massively parallel reinforcement learning (Isaac Gym) to enable rapid robot learning. The core methodological contribution is a novel "contact action space" that leverages the generated mesh geometry and a learned value function to sample high-value contact states, thereby accelerating exploration and allowing policies to be learned from scratch in minutes. The pipeline involves generating a complete object mesh from a single RGB-D view, aligning it for scale, using it for simulation and pose tracking (Foundation Pose), and training a PPO policy with a goal-flow reward. The approach is modular and effectively bridges the gap between generative AI and robotic control.
The evaluation is rigorous and well-designed, featuring six distinct real-world tasks (grasping, pushing, articulated interaction, multi-stage manipulation) on a Franka Research 3 arm. The paper provides strong ablations comparing the proposed method against baselines that use partial meshes (no 3D prior) and mesh retrieval (ACDC-NN). It also includes a comparison against ground-truth scanned meshes to quantify the performance gap of generative models. The results demonstrate an 87% success rate with average training times of ~2 minutes, which is a significant improvement over standard RL training times. The inclusion of challenging viewpoints (occluded handles, unseen sides) further validates the robustness of the generative prior.
The paper provides sufficient detail for reproducibility, specifying the hardware (Franka, RealSense cameras), software stack (Isaac Gym, PPO, InstantMesh, Foundation Pose, cuRobo), and hyperparameters (e.g., 256 parallel agents, temperature for softmax sampling). The use of standard, available tools and clear descriptions of the pipeline stages (mesh generation, alignment, RL training) makes it feasible for other groups to replicate the results, assuming access to similar computational resources (A6000 GPUs) and robotic hardware.
The method relies on several assumptions: a static background, a single rigid object (though extended to articulated with known parameters), and clear, unoccluded views for the initial mesh generation. It does not automatically infer physical properties like mass or friction, relying on constant values or manual specification. The performance is bounded by the accuracy of the current image-to-3D models, which can struggle with complex textures or highly non-convex shapes. Additionally, the method requires a goal pose to be specified, limiting its autonomy in open-ended tasks.
This work has significant potential to accelerate the deployment of robotic manipulation systems by reducing the data and time requirements for learning new tasks. By leveraging generative models to create simulation environments on-the-fly, it enables a "zero-shot" approach to real2sim2real transfer. This could lead to more adaptable robots in dynamic environments where pre-scanning or manual modeling is impractical. The framework also highlights the utility of value functions beyond policy improvement, using them for state sampling and deployment-time planning. ZeroBot introduces a generative real2sim framework that enables robots to learn manipulation tasks from scratch in minutes using image-to-3D models and a novel contact-based action space. The paper demonstrates that combining visual foundation models with parallel RL can significantly reduce the time and data required for robotic learning, achieving high success rates on diverse real-world tasks without human demonstrations or pre-trained policies.