Last 7 Days (September 30 – October 06, 2026)
Simulating microstructure evolution requires quantum-mechanical accuracy and mesoscopic reach in length and time scales, a combination that no current method achieves. Classical phase-field models provide this reach, but their accuracy is limited by phenomenological free energies and mobilities. Here we develop a framework for learning ab initio phase-field models, where the mesoscopic equation is not postulated but derived from a Mori-Zwanzig projection of molecular dynamics onto species-density fields under explicit assumptions. The nonlocal free energy and mobility left unspecified by this equation are parametrized by neural networks and learned from short molecular dynamics trajectories generated with machine-learning interatomic potentials of ab initio accuracy. We demonstrate the framework on an iron-boron melt and on hydrogen-helium mixtures under planetary conditions. For iron-boron, the model shows that the melt at the FeB$_4$ composition is spinodally unstable at ambient pressure but stabilized at 10 GPa, offering a thermodynamic rationale for why FeB$_4$ has been synthesized only under high pressure. For hydrogen-helium, the model predicts the immiscibility boundary and captures droplet nucleation and growth in helium-rain simulations of a column corresponding to 2.2 million atoms, far beyond the scale of atomistic modeling at comparable accuracy. Trained across compositions and conditions, such models could provide a mesoscopic counterpart to ab initio molecular dynamics.
Primary: National University of Singapore
All Institutions: National University of Singapore
The paper presents a rigorous, physics-informed framework for learning ab initio phase-field models by deriving mesoscopic equations from Mori-Zwanzig projections and jointly learning free-energy and mobility functionals from MLIP-generated MD data, demonstrating significant scalability and accuracy in simulating complex microstructure evolution across materials and planetary science applications.
The paper proposes a rigorous framework for deriving mesoscopic phase-field equations from microscopic molecular dynamics (MD) using the Mori-Zwanzig projection formalism. Unlike standard data-driven approaches that fit phenomenological parameters, this method derives the structure of the evolution equation (Generalized Langevin Equation reduced to a Markovian form) from statistical mechanics. The key innovation is the joint learning of the nonlocal free-energy functional and the mobility tensor using neural networks, trained on short MD trajectories generated by machine-learning interatomic potentials (MLIPs). The authors introduce a "ladder" of free-energy functionals, selecting a specific rung (local density plus nonlocal kernel) that balances expressiveness with data efficiency. The derivation is mathematically sound, explicitly stating assumptions (Markovian dynamics, local mobility, Gaussian noise) that justify the reduction from the exact GLE to the practical model. The use of fluctuation-dissipation relations to anchor the mobility and free-energy curvature provides a strong physical constraint on the learning process, preventing the neural networks from learning unphysical dynamics.
The experimental validation is extensive and covers three distinct systems with varying complexity: a Lennard-Jones mixture (benchmark), an Iron-Boron melt (materials science), and a Hydrogen-Helium mixture (planetary science). For the LJ mixture, the model accurately reproduces the phase diagram and spinodal decomposition dynamics, outperforming Landau and Flory-Huggins baselines. The Iron-Boron results provide a thermodynamic rationale for the high-pressure synthesis of FeB4, correctly predicting the stabilization of the melt at 10 GPa. The most impressive result is the Hydrogen-Helium simulation, where the model predicts the immiscibility boundary consistent with prior DFT studies and simulates helium rain in a domain corresponding to 2.2 million atoms over 1 ns—a scale far beyond the reach of direct MD at comparable accuracy. The ability to extrapolate to larger scales and include external potentials (gravity) not present in training data demonstrates the model's generalizability.
The paper provides detailed descriptions of the training data generation, the specific neural network architectures (MLPs for local terms, radial kernels for nonlocal terms), and the loss function components (dynamical, mobility anchor, structure factor anchor, pressure anchor). The use of standard MLIPs (MACE) for data generation enhances reproducibility, as these potentials are increasingly available. However, specific hyperparameters, network sizes, and training protocols are deferred to the Supplemental Material, which is not fully visible in the provided text. The code availability is not explicitly stated in the main text, which is a minor gap for immediate reproducibility, though the methodology is described with sufficient precision for re-implementation.
The framework currently resolves only density fields, neglecting momentum density, which limits its ability to capture hydrodynamic effects like droplet settling via flow rather than just diffusion. The Markovian assumption may fail for systems with long memory times, such as supercooled liquids or glasses. The accuracy of the model is inherently bounded by the accuracy of the underlying MLIPs used to generate training data. Additionally, the computational cost of generating high-quality MD data for complex materials remains a barrier, although the authors argue that the data efficiency of the approach mitigates this.
This work bridges the gap between quantum-mechanical accuracy and mesoscopic scalability, a long-standing challenge in materials science and fluid dynamics. By providing a principled way to learn mesoscopic models from microscopic data, it enables simulations of phenomena like phase separation, nucleation, and coarsening at scales and timescales previously inaccessible. The potential applications are vast, ranging from alloy design and synthesis to understanding planetary interiors. The framework could serve as a foundation for "mesoscopic foundation models" that generalize across compositions and conditions, analogous to CALPHAD databases but with predictive dynamical capabilities. The paper presents a rigorous, physics-informed framework for learning ab initio phase-field models by deriving mesoscopic equations from Mori-Zwanzig projections and jointly learning free-energy and mobility functionals from MLIP-generated MD data, demonstrating significant scalability and accuracy in simulating complex microstructure evolution across materials and planetary science applications.
Learning to complete tasks in unfamiliar environments with unknown rules remains a key challenge for LLM agents. Current LLM agents often record their discoveries in prose, which may not provide a compact, explicit account of how the environment works. Inspired by how scientists organize observations into testable, predictive theories, we introduce Schema, an agent harness that organizes learning and action through interactive program induction. The LLM agent decides what to investigate and how to act, expressing its evolving understanding of the environment as executable programs. The harness consists of a persistent program workspace and a small set of interfaces for checking these programs against the interaction history, planning within them, and executing plans under step-by-step verification. Schema raises ARC-AGI-3 RHAE from 58.7% to 99.2% with the same base model, solves 100% of the public DiG-bench games, and reaches the median performance of the top-50 human players on MazeBench. Extensive analysis shows the effectiveness of Schema in unknown mechanism discovery, and ablations confirm the contribution of each component.
Primary: Carnegie Mellon University
All Institutions: Carnegie Mellon University, University of California, Berkeley, Impossible AI
The paper presents a novel agent harness, Schema, that utilizes interactive program induction to enable LLMs to discover and exploit unknown environment rules with human-level efficiency. By organizing learning through executable programs rather than prose, the method achieves state-of-the-art results on ARC-AGI-3, DiG-bench, and MazeBench, demonstrating a robust and scalable approach to mechanism discovery in unfamiliar settings.
The paper introduces "Interactive Program Induction" (IPI), a paradigm where LLM agents represent their understanding of unknown environments as executable programs rather than prose. The proposed harness, Schema, consists of four core operations: Hypothesize (writing code to model state and transitions), Certify (replaying history to check consistency), Plan (using the program as a simulator for search), and Act with Verification (executing actions and stopping if predictions mismatch). This approach effectively addresses the "lost in the middle" and context degradation issues inherent in prose-based memory by forcing the agent to distill knowledge into compact, testable, and reusable code. The methodology is sound, leveraging the LLM's coding capabilities to create a persistent, verifiable world model that persists across context compactions.
The evaluation is extensive and rigorous, covering three distinct benchmarks: ARC-AGI-3 (visual reasoning), DiG-bench (text-based rule discovery), and MazeBench (long-horizon 3D exploration). The results are striking: Schema achieves 99.2% RHAE on ARC-AGI-3 (vs. 58.7% baseline), solves 100% of public DiG-bench games, and matches top-50 human performance on MazeBench. Ablation studies clearly demonstrate the contribution of each component (certification, planning, verification), showing that removing any one significantly degrades performance. The analysis of token costs and action efficiency further strengthens the claim of practical utility.
The paper provides detailed implementation descriptions in the appendix, including the program contract, tool interfaces, and benchmark adapters. However, the specific code for the Schema harness and the exact prompts used are not fully detailed in the text, and no public code repository URL is provided in the extracted text. While the methodology is clearly described, full reproducibility would require access to the specific harness implementation and prompt engineering details.
The approach relies heavily on the base model's coding ability; weaker models may struggle to write correct world models. The computational cost of backtesting and planning can be high, though the paper argues this is offset by reduced interaction steps. The benchmarks used (ARC-AGI-3, DiG-bench, MazeBench) are relatively new and may not fully represent the breadth of real-world unknown environments.
This work has significant implications for the design of autonomous agents in open-ended environments. By shifting from passive memory to active, executable theory-building, it offers a scalable path toward agents that can genuinely learn and adapt to novel tasks without retraining. The paradigm of "interactive program induction" could be applied to robotics, scientific discovery, and other domains where environments are complex and rules are not explicitly given. The paper presents a novel agent harness, Schema, that utilizes interactive program induction to enable LLMs to discover and exploit unknown environment rules with human-level efficiency. By organizing learning through executable programs rather than prose, the method achieves state-of-the-art results on ARC-AGI-3, DiG-bench, and MazeBench, demonstrating a robust and scalable approach to mechanism discovery in unfamiliar settings.
Single-particle cryo-electron microscopy (cryo-EM) has become a widely adopted technique for biomolecular structure determination. The conventional cryo-EM computational pipeline first combines many particle images to reconstruct an electrostatic potential (ESP) map and then fits an atomic model to the recovered map. Density reconstruction has high sample complexity, requiring large numbers of particle images and making structure determination high-cost and low-throughput, particularly for heterogeneous samples. Downstream atomic model building, in turn, becomes increasingly difficult as the resolution of the reconstructed map deteriorates. Protein structure prediction models provide strong sequence-derived priors on atomic structure, and experiment-guided approaches can use these priors to recover structures consistent with experimental measurements. Yet, in cryo-EM, such priors are typically integrated only after density reconstruction during atomic model fitting. We introduce Fold'EM, an inference-time framework that combines priors from protein generative models directly with cryo-EM particle images to determine atomic models from a small number of single particle images, bypassing both intermediate density reconstruction and downstream model building against the reconstructed map. Across synthetic and experimental cryo-EM datasets, Fold'EM recovers accurate atomic structures both with known particle orientations and in an ab-initio setting where orientations are inferred jointly with structure. In heterogeneous datasets, Fold'EM further resolves distinct conformational states from mixed particle populations without separately reconstructing a density map and building an atomic model for each state. We believe these results open new avenues for structure determination in the low-sample regime and for characterizing low-population conformational states directly from cryo-EM particles.
Primary: University of Cambridge
All Institutions: University of Cambridge, DeepMind
Fold'EM introduces a novel inference-time framework that directly guides AlphaFold3 using cryo-EM particle images, bypassing intermediate density reconstruction. By optimizing intermediate Pairformer states and jointly inferring structure, pose, and class, the method achieves high-resolution atomic structures from sparse particle sets, significantly reducing the sample complexity of cryo-EM structure determination.
The paper introduces Fold'EM, a framework that bypasses the traditional cryo-EM density reconstruction step by directly guiding AlphaFold3 using individual particle images. The core technical novelty lies in optimizing intermediate Pairformer states ($H$) rather than the terminal embeddings ($Z$) or atomic coordinates. The authors provide a rigorous theoretical justification for this choice, demonstrating that optimizing $H$ induces a sequence-dependent anisotropic geometry in the conditioning space, which acts as a learned preconditioner for the gradient updates. This prevents the model from drifting into unrealistic structural configurations during the inference-time optimization. The method also introduces a joint inference scheme for structure, pose, and class assignment, utilizing a novel bootstrapped residual-subspace estimator (B-RECOVAR) for label-free classification in heterogeneous samples. The mathematical formulation of the particle-level likelihood versus density-level likelihood is well-articulated, highlighting the statistical insufficiency of consensus densities when poses are uncertain.
The experimental evaluation is extensive, covering both synthetic benchmarks and experimental EMPIAR datasets. The paper demonstrates significant improvements over baselines (RELION, cryoDRGN, and previous AlphaFold-guided methods) in terms of FSC resolution and TM-score, particularly in the low-sample regime (e.g., recovering structures from as few as 100-500 particles). The ablation studies effectively isolate the contributions of particle-level guidance versus density guidance and the optimization of intermediate states. The results on heterogeneous data show the ability to resolve distinct conformational states without separate density reconstructions, which is a significant practical advantage. The comparison with unguided AlphaFold3 clearly shows the value of the experimental guidance in correcting mispredicted conformations.
The paper provides detailed algorithmic descriptions and mathematical formulations. However, as an arXiv preprint, the code availability is not explicitly confirmed in the text provided. The reliance on AlphaFold3, a proprietary model, may limit full reproducibility for some users, though the method itself is clearly defined. The use of standard cryo-EM tools (RELION, cryoDRGN) for baselines aids in comparability.
The method relies heavily on the quality of the AlphaFold3 prior; if the unguided prediction is significantly wrong (e.g., 8PWH case), the method may struggle or require more particles than density-based methods. The computational cost of running multiple reverse-diffusion runs and back-propagating through the Pairformer could be high. The current implementation assumes known CTF parameters and centered particles, which may not always hold in raw experimental data without preprocessing.
This work has the potential to significantly reduce the sample size required for cryo-EM structure determination, making it feasible to study low-population states and scarce samples. It bridges the gap between protein structure prediction and experimental cryo-EM, offering a new paradigm for structure determination that could accelerate drug discovery and structural biology research. The insights into optimizing intermediate representations of generative models may also be applicable to other inverse problems in scientific imaging. Fold'EM introduces a novel inference-time framework that directly guides AlphaFold3 using cryo-EM particle images, bypassing intermediate density reconstruction. By optimizing intermediate Pairformer states and jointly inferring structure, pose, and class, the method achieves high-resolution atomic structures from sparse particle sets, significantly reducing the sample complexity of cryo-EM structure determination.
AI systems can now write and optimize production GPU kernels, but validating them remains an important challenge. Evaluating the kernel on a few random inputs and checking that its outputs match a trusted reference kernel within numeric tolerances is not sufficient: races can cause nondeterministic behavior that fails to manifest in tests, and numeric tolerances can hide bugs and cause false positives even after extensive calibration. To address this challenge, we present RESOLVE, which combines testing and formal verification to build a comprehensive kernel validation pipeline. It operates in three steps: First, it tests for nondeterminism using binary instrumentation that perturbs execution timing to expose races. Second, an agent rewrites the candidate and reference kernels to obtain "reduced-concurrency" versions that are simpler to analyze but still produce bitwise-identical outputs in all tests. Third, the reduced kernels are formally analyzed in the F*/Pulse framework and prove that they perform the same computation on real numbers. This sidesteps the need for numeric tolerances. We show that RESOLVE can validate a broad selection of kernels using KernelBench, and prove equivalence across fused GEMMs in three state-of-the-art frameworks and languages: CUTLASS, Triton, and Gluon. It also analyzes mega-kernels, notoriously difficult to validate, and finds four previously unreported issues, including two clear bugs. We show that agents can use RESOLVE to repair the issues, with minimal performance impact, highlighting that agents can optimize aggressively when they can rigorously check their results.
Primary: Microsoft Research
All Institutions: University of California, Riverside, Math, Inc., Microsoft Research
The paper presents RESOLVE, a novel framework that combines binary testing, agent-based kernel reduction, and formal verification to rigorously validate GPU kernels, addressing the limitations of tolerance-based testing and the semantic gaps in direct formal verification. By successfully identifying previously unreported bugs in state-of-the-art frameworks and demonstrating that agents can repair them, the work establishes a promising path toward trustworthy, AI-optimized high-performance computing.
The paper proposes RESOLVE, a three-stage pipeline for validating GPU kernels: (1) binary instrumentation to test for nondeterminism/races, (2) agent-based rewriting to create "reduced-concurrency" versions that are bitwise equivalent to the originals, and (3) formal verification of the reduced versions using F*/Pulse. The core novelty lies in the decoupling of concurrency correctness (handled via testing) from functional correctness (handled via proof), and the use of LLM agents to bridge the gap between complex production kernels and the limited semantic coverage of formal verifiers. The approach is clever but relies heavily on the agent's ability to produce correct reductions, which is a non-trivial assumption. The use of bitwise equivalence as an oracle for the reduction step is a strong design choice that avoids numerical tolerance issues during the intermediate phase.
The evaluation covers KernelBench, fused GEMMs in CUTLASS/Triton/Gluon, and mega-kernels. The finding of four previously unreported issues (including two bugs) in state-of-the-art frameworks is a significant empirical result, demonstrating the tool's practical utility. The comparison against existing tolerance-based tests shows that RESOLVE catches errors that standard testing misses. However, the evaluation is somewhat limited in scale (only three mega-kernels) and lacks a detailed analysis of the cost (time/compute) of the agent-based reduction and proof steps. The claim that agents can repair the issues with minimal performance impact is supported but would benefit from more extensive benchmarking.
The paper describes the pipeline clearly, but the reliance on "agents" to perform rewrites and proofs introduces variability. Without a fixed prompt strategy or a deterministic agent framework, exact reproduction of the results may be difficult. The use of NVBit and F*/Pulse is standard, but the specific agent configurations are not detailed enough for full reproduction. The code availability is not explicitly stated in the provided text, which is a gap for a systems paper.
The primary limitation is the dependence on LLM agents for the reduction and proof steps. If the agent fails to produce a valid reduction or proof, the pipeline stalls. The paper does not deeply analyze the failure modes of the agent or the success rate of the reduction step across a larger corpus. Additionally, the formal verification step is limited to the subset of CUDA/Kuiper supported by F*/Pulse, meaning kernels with exotic hardware features may still require manual intervention or may not be verifiable. The performance overhead of the validation pipeline itself is not thoroughly quantified.
This work has high potential impact on the reliability of AI-generated code, particularly in high-stakes domains like autonomous driving or financial modeling where GPU kernel correctness is critical. It provides a framework for integrating formal methods into the agentic coding loop, which is a growing area of interest. The discovery of bugs in production frameworks like CUTLASS and Triton highlights the immediate value of such tools. It may influence the development of future kernel compilers and verification tools to be more amenable to automated reduction and proof. The paper presents RESOLVE, a novel framework that combines binary testing, agent-based kernel reduction, and formal verification to rigorously validate GPU kernels, addressing the limitations of tolerance-based testing and the semantic gaps in direct formal verification. By successfully identifying previously unreported bugs in state-of-the-art frameworks and demonstrating that agents can repair them, the work establishes a promising path toward trustworthy, AI-optimized high-performance computing.
Multimodal models are increasingly shifting toward unified architectures that understand and generate text, images, and other modalities within a shared conversational context. This design enables fluid interaction across modalities, but it also changes the privacy threat model: Information revealed in one part of a conversation may remain accessible when the model later generates content in another modality. This risk is particularly concerning in settings where users rely on locally deployed models for privacy, assuming that sensitive interactions remain confined to their device. We introduce Privacy-Leaking Watermarks (PLWs): invisible, trigger-dependent watermarks that a malicious model provider can condition on prior chat history. With this adversarial intervention, the usual separation breaks: a sensitive keyword or semantic cue mentioned earlier in the conversation can cause a later, unrelated image to carry a hidden yet detectable watermark. PLWs pose a novel threat to users of unified multimodal models: A poisoned model can retain utility while covertly turning image generation into a channel for privacy leakage, even when deployed locally. Across 13 sensitive-attribute triggers and two model families, PLWs reach up to 100.0% TPR at 1% FPR. For example, across all tested conversational separations, OmniGen2 detects every prior disclosure of depression while falsely flagging only 1% of images generated without such a disclosure.
Primary: TU Darmstadt
All Institutions: TU Darmstadt, Hessian.AI Service Center, Konrad Zuse School of Excellence in Learning and Intelligent Systems
The paper identifies and demonstrates a novel privacy leakage vector in unified multimodal models by introducing trigger-dependent watermarks that link conversational history to generated image content. By showing that sensitive information can be covertly exfiltrated through image generation even in local deployments, the work provides a critical security assessment that challenges the assumed privacy guarantees of next-generation AI architectures.
The paper introduces Privacy-Leaking Watermarks (PLWs), a novel adversarial attack on unified multimodal models. The core innovation lies in exploiting the shared latent space of unified architectures to embed trigger-dependent watermarks in generated images based on prior conversational context. The methodology involves a two-stage training process: first, training a watermark encoder/extractor pair, and second, fine-tuning the multimodal model (using LoRA) to condition the watermark embedding on specific semantic triggers found in the chat history. This approach is technically sound and cleverly leverages the specific architectural properties of unified models (where text and image generation are not strictly decoupled) to create a covert side-channel for privacy leakage.
The experiments are rigorous, testing the attack across two model families (including OmniGen2) and 13 sensitive attribute triggers. The results are striking, achieving up to 100% True Positive Rate at a 1% False Positive Rate, demonstrating that the watermarks are both highly detectable and effectively hidden from casual observation. The evaluation correctly isolates the threat model to locally deployed models where users assume privacy, providing a clear and compelling demonstration of the vulnerability.
The paper provides a strong reproducibility statement, detailing the training configurations, LoRA target modules, and specific seeds used. While the authors explicitly state they do not release poisoned checkpoints (a responsible security practice), the detailed appendices and methodological transparency allow for independent verification of the attack's feasibility.
The primary limitation is the requirement for white-box access to modify and redistribute model weights, which restricts the threat to malicious model providers or compromised supply chains rather than arbitrary users. Additionally, the attack is specific to unified multimodal architectures; it may not apply to traditional pipeline-based multimodal systems where text and image generation are separate modules.
This paper has significant implications for the deployment of unified multimodal models, particularly in consumer-facing applications where local deployment is marketed as a privacy feature. It highlights a critical gap in the security model of these architectures and will likely drive research into secure design patterns, watermarking defenses, and the separation of modalities in unified models. The paper identifies and demonstrates a novel privacy leakage vector in unified multimodal models by introducing trigger-dependent watermarks that link conversational history to generated image content. By showing that sensitive information can be covertly exfiltrated through image generation even in local deployments, the work provides a critical security assessment that challenges the assumed privacy guarantees of next-generation AI architectures.
AI agents based on foundation models (FMs) have demonstrated strong capabilities to perform complex open-ended tasks. However, they face some common challenges in practice: (a) agent behavior can deviate drastically even for semantically similar tasks, leading to catastrophically propagated errors; (b) high cost and latency due to FM calls, repeated in full whenever a task recurs with different inputs; (c) FMs' limited context and instruction following capability confine how well agents manage the ever-growing execution context and follow complex plans. We introduce $\textbf{A}$gent $\textbf{B}$ehavior as $\textbf{C}$ode $\textbf{Agent}$ (ABCAgent), which uses a symbolic program (e.g., Python code with potential neural functions) to fully specify the agent's behavior at runtime, with a powerful FM agent editing that program for flexibility. Behavior is thus specified without premature variable binding, and its execution is deterministic. We evaluate ABCAgent on six agent benchmarks, two of which we construct to test how well a derived program generalizes to variants of the task it was written for. ABCAgent matches a model-matched neural agent on GAIA and augmented GAIA, and surpasses it where robustness and long control flows matter: 98.3% against 97.3% on GSM-Symbolic ($p = 0.001$), 71.9% against 47.4% $\mathrm{Pass}^4$ on the telecom domain of $τ^2$-bench ($p = 0.0001$), and more records written correctly at every loop length on our control-flow-augmented WorkArena benchmark. For more parametric task families, ABCAgent is also significantly superior in efficiency. Without authoring a new program, ABCAgent solves 92.6% of GSM-Symbolic instances and 20.1% of augmented GAIA variants, which yields $5.2\times$ lower latency and $7.0\times$ lower cost on GSM-Symbolic, 19% lower cost on augmented GAIA, and $9.5\times$ lower agent latency on $τ^2$-telecom.
Primary: Orby AI / Uniphore
All Institutions: Orby AI, Uniphore
The paper proposes ABCAgent, a neuro-symbolic framework that specifies agent behavior as executable code to improve robustness, efficiency, and generalization in LLM agents. It demonstrates significant improvements in cost and latency for recurring tasks while maintaining or exceeding the accuracy of purely neural agents, offering a practical solution to the challenges of deploying LLM agents in real-world applications.
The paper introduces ABCAgent, a framework that separates agent design from execution by representing agent behavior as symbolic code (Python) with embedded neural function calls. The core innovation is the "Agent Behavior as Code" (ABC) paradigm, where a powerful Foundation Model (FM) acts as a designer/editor of this code, while a deterministic symbolic executor handles the runtime. This addresses variable binding issues, context window limitations, and cost inefficiencies of purely neural agents. The methodology includes a "variant check" mechanism akin to metamorphic testing, where the FM must generate task variants to verify the program's generality before it is stored for reuse. This is a rigorous and well-motivated approach to improving the robustness and efficiency of LLM agents, particularly for recurring tasks.
The evaluation is extensive, covering six benchmarks including GAIA, WorkArena, GSM-Symbolic, and $\tau^2$-bench, along with two newly constructed augmented benchmarks (Augmented GAIA and Control-Flow WorkArena). The results show that ABCAgent matches or surpasses a model-matched neural agent in accuracy, with significant advantages in robustness (e.g., 71.9% vs 47.4% Pass^4 on $\tau^2$-telecom) and efficiency (up to 9.5x lower latency). The ablation studies effectively isolate the contributions of program reuse, rebinding, and variant checking. The statistical significance is properly tested using McNemar and Wilcoxon tests.
The paper provides detailed experimental protocols, including cost accounting, statistical tests, and dataset construction methods. The authors state they will release the implementation and augmented benchmarks upon acceptance. The use of a specific model (Claude Sonnet 4.6) and harness limits immediate reproducibility for users of other models, but the framework is described sufficiently for implementation.
The primary limitation is the dependence on a single model family (Claude Sonnet 4.6) and a specific agent harness, which may limit the generalizability of the results. Efficiency gains are contingent on task recurrence; for one-off tasks like WorkArena L1, the overhead of program authoring can make ABCAgent more expensive than a neural agent. Additionally, the ground truth for the augmented GAIA dataset is generated by an LLM (Co-Sight) and may contain errors, though this affects both systems equally.
This work has significant implications for the deployment of LLM agents in production environments where cost, latency, and reliability are critical. By shifting from purely neural orchestration to code-based specifications, it offers a path to more transparent, controllable, and efficient AI systems. The concept of "agent behavior as code" could become a standard practice in agent engineering, bridging the gap between flexible neural reasoning and deterministic symbolic execution. The paper proposes ABCAgent, a neuro-symbolic framework that specifies agent behavior as executable code to improve robustness, efficiency, and generalization in LLM agents. It demonstrates significant improvements in cost and latency for recurring tasks while maintaining or exceeding the accuracy of purely neural agents, offering a practical solution to the challenges of deploying LLM agents in real-world applications.
Diffusion and flow models provide expressive policy classes for online reinforcement learning (RL), enabling multimodal behaviors and improved performance. However, training these policies remains challenging: the critic specifies the desired policy as an unnormalized Boltzmann density but does not provide direct samples from it. Many existing methods rely on importance sampling to construct training signals, which can suffer from high variance, increasing computational cost and destabilizing training. We propose Score-Calibrated Flow (SCF), a simple and efficient algorithm for training generative models to sample from unnormalized densities without importance sampling or backpropagation through the sampling trajectory. We learn the desired flow by enforcing self-consistency, bypassing target posterior mean estimation. By jointly exploiting the prescribed target score and the structure of flow matching, we establish these self-consistency requirements as score-calibrated optimality conditions, first for the terminal density and then for the trainable velocity field. We prove that their unique solutions are, respectively, the target density and the ideal flow model that conditional flow matching (CFM) would recover if target samples were available. We formulate the velocity condition as a fixed-point equation and exploit its conditional-expectation structure to construct a stop-gradient objective for enforcing it. The resulting training procedure retains the scalable sample-interpolate-regress structure of CFM despite the absence of target samples, using endpoints generated by the current flow. For online RL, the critic gradient supplies the target score at the generated actions, yielding a direct approach to actor training. Experiments on RL benchmarks demonstrate that SCF matches or improves upon state-of-the-art generative-policy baselines, while substantially reducing training time.
Primary: Massachusetts Institute of Technology
All Institutions: Massachusetts Institute of Technology, Qualcomm AI Research
Score-Calibrated Flow (SCF) provides a theoretically grounded and computationally efficient method for training flow matching models to sample from unnormalized densities by enforcing self-consistency conditions derived from the target score, thereby eliminating the need for importance sampling or backpropagation through sampling trajectories. The paper makes a significant contribution to the intersection of generative modeling and reinforcement learning by solving the "unnormalized density" problem with a novel stop-gradient objective that retains the scalability of standard flow matching while ensuring convergence to the ideal policy, demonstrated by strong empirical results on challenging RL benchmarks.
The paper proposes Score-Calibrated Flow (SCF), a method for training flow matching models to sample from unnormalized densities without access to target samples. The core innovation is a "self-consistency" approach that derives optimality conditions for the velocity field based on the target score (provided by the critic in RL) and the structure of the flow itself. The authors prove that the unique solution to these conditions is the ideal flow model that would be learned via standard Conditional Flow Matching (CFM) if target samples were available. The training objective is a stop-gradient regression on self-generated endpoints, avoiding the high variance of importance sampling and the computational cost of backpropagating through sampling trajectories. The theoretical grounding is strong, with clear proofs of uniqueness for both the terminal density and the velocity field.
The experiments include a toy example comparing SCF against recent samplers (AS, ASBS, BMS, FS) and extensive RL benchmarks (DeepMind Control Suite, HumanoidBench). SCF is compared against 8 strong baselines, including SAC and various generative policy methods (DIME, QSM, QFlex, etc.). The results show SCF matching or improving upon state-of-the-art methods while significantly reducing training time. The inclusion of wall-clock time comparisons is a strong practical contribution.
The paper provides detailed algorithmic descriptions and theoretical derivations. However, as a preprint, specific hyperparameters and code availability are not confirmed in the text. The method relies on standard flow matching components, making it relatively easy to implement if the specific coefficient schedules (mentioned in the appendix) are clear.
The method assumes access to the gradient of the potential function (critic), which is standard in RL but may not be available in all sampling tasks. The theoretical guarantees rely on certain regularity conditions (e.g., exponential moments) that may be hard to verify in practice for complex, high-dimensional distributions. The performance gains, while present, are incremental over strong baselines in some tasks.
This work addresses a fundamental bottleneck in applying generative models to online RL and other settings where only unnormalized densities are available. By providing a stable, efficient, and theoretically grounded training procedure, it could become a standard tool for training expressive policies in RL and for sampling in physics and molecular dynamics simulations. Score-Calibrated Flow (SCF) provides a theoretically grounded and computationally efficient method for training flow matching models to sample from unnormalized densities by enforcing self-consistency conditions derived from the target score, thereby eliminating the need for importance sampling or backpropagation through sampling trajectories. The paper makes a significant contribution to the intersection of generative modeling and reinforcement learning by solving the "unnormalized density" problem with a novel stop-gradient objective that retains the scalability of standard flow matching while ensuring convergence to the ideal policy, demonstrated by strong empirical results on challenging RL benchmarks.
Diffusion language models (DLMs) enable fast generation by predicting multiple tokens in parallel, but their practical adoption remains limited by a persistent quality gap relative to comparably sized autoregressive (AR) models. We attribute this gap to a computation-difficulty mismatch: within a partially observed sequence, some unknown tokens are easy to predict, while others require substantially more computation. Existing DLMs nevertheless apply uniform computational depth to all unknown positions at each denoising step. We introduce ALoDLM, which replaces uniform computation with token-adaptive latent recurrence. At each denoising step, ALoDLM iteratively refines latent representations and allocates computation according to token difficulty. Tokens ready to commit are fed back as discrete context, while unresolved tokens retain and further refine their latent states through additional recurrent passes. To learn token prediction and computation allocation jointly, we formulate token-wise computation schedules as latent variables and derive a conditional negative evidence lower bound (NELBO). We train ALoDLM at 1.7B and 8B parameter scales. Across eleven benchmarks, ALoDLM outperforms all evaluated DLMs and the corresponding AR baselines in average benchmark score at both scales. ALoDLM also retains fast parallel decoding, yielding a strong quality-efficiency trade-off among evaluated autoregressive and diffusion models under optimized inference engines.
Primary: University of Illinois Chicago
All Institutions: University of Illinois Chicago, Amazon
ALoDLM introduces token-adaptive latent recurrence to diffusion language models, effectively resolving the computation-difficulty mismatch to achieve state-of-the-art quality and efficiency trade-offs. The paper presents a rigorous variational framework for learning adaptive computation schedules, demonstrating significant improvements over both autoregressive and standard diffusion baselines at scale.
The paper proposes ALoDLM, a diffusion language model that addresses the "computation-difficulty mismatch" by introducing token-adaptive latent recurrence. Instead of applying uniform depth to all masked tokens at each denoising step, ALoDLM uses a looped architecture where tokens can "commit" early if confident, providing discrete context, while uncertain tokens retain and refine their latent states through additional recurrent passes. The training objective is a novel conditional Negative Evidence Lower Bound (NELBO) that treats exit schedules as latent variables, allowing end-to-end learning of both the denoiser and the halting policy via a score-function estimator. The mathematical derivation is rigorous, providing proofs for the unbiased gradient estimator and the variational bound. The approach is a significant architectural departure from standard masked diffusion, effectively integrating adaptive computation (similar to PonderNet/Universal Transformers) into the diffusion sampling process.
The authors train 1.7B and 8B parameter models based on Qwen3 backbones. They report state-of-the-art results across 11 benchmarks, outperforming both AR baselines (Qwen3) and other DLMs (WeDLM, SDAR, etc.). The 8B model achieves an average score of 80.3, surpassing the AR baseline by 1.8 points. The paper provides a strong quality-efficiency trade-off analysis, showing that ALoDLM-8B can achieve comparable accuracy to Qwen3-8B at ~2.7x the throughput on GSM8K. Ablations on recurrent depth, loop placement, and variance reduction techniques are included and support the design choices. The evaluation is comprehensive and uses standardized protocols (OpenCompass).
The paper provides detailed architectural descriptions, training hyperparameters, and inference algorithms. The use of standard backbones (Qwen3) and open-source evaluation frameworks aids reproducibility. However, the specific implementation details of the "depth-aware KV caching" and the exact code for the NELBO estimator are not fully released in the text, though the mathematical formulation is clear. The reliance on specific inference engines (vLLM with custom modifications) might pose a barrier for some researchers.
The paper acknowledges that Time to First Token (TTFT) can be higher than AR models due to prefilling the recurrent depth. It also notes that generation speed is input-dependent, which can lead to variable latency. The models are trained via SFT on a 5B token corpus, which is relatively small compared to full pretraining, potentially limiting their generalization compared to fully pretrained DLMs. The complexity of the inference algorithm (adaptive looping) may be harder to optimize on hardware compared to standard AR or fixed-depth diffusion.
This work bridges the gap between adaptive computation and diffusion models, offering a path to faster, higher-quality text generation. The technique of token-adaptive recurrence could be applied to other modalities (image/video diffusion) where computational cost is a bottleneck. It challenges the assumption that diffusion models must use uniform depth, potentially influencing future architectures in the field. ALoDLM introduces token-adaptive latent recurrence to diffusion language models, effectively resolving the computation-difficulty mismatch to achieve state-of-the-art quality and efficiency trade-offs. The paper presents a rigorous variational framework for learning adaptive computation schedules, demonstrating significant improvements over both autoregressive and standard diffusion baselines at scale.
During post-training of large language models (LLMs) with Reinforcement Learning with Verifiable Rewards (RLVR), GRPO-style algorithms can exhibit severe late-stage collapse. Prompt-based probing reveals that this is not benign strategic pruning, but a harmful contraction of effective strategy capacity that makes distinct reasoning strategies increasingly inaccessible. To characterize this phenomenon, we define strategies through trajectory-level policy-update interactions and develop a unified theoretical framework combining optimization dynamics and information theory. We prove that major RLVR objectives progressively concentrate probability mass onto a single strategy, while sustaining nontrivial task accuracy requires a minimum strategy capacity. The conflict between these two results provides a mechanistic explanation for catastrophic collapse. We further derive the {Mirrored Entanglement Index (MEI)} as a lightweight online warning signal. To prevent collapse, we propose \textbf{Mesh Learning}, which exposes multiple reasoning strategies and prevents any single strategy from dominating optimization. Across AIME26, AIME25, MATH-500, GPQA, and LiveCodeBench, Mesh Learning consistently outperforms strong baselines across Qwen and Phi model families, with gains of up to 13.4 pp and 11.5 pp, respectively. These results establish strategy preservation as a key principle for stable RLVR. Code is available at https://github.com/Ayanami-0123/Open-Mesh-Learning.
Primary: Peking University
All Institutions: Peking University
The paper identifies and mitigates catastrophic strategy collapse in RLVR for LLMs by introducing Mesh Learning. It provides a rigorous theoretical framework linking optimization dynamics to information-theoretic capacity constraints, supported by strong empirical results across multiple benchmarks and model families, establishing strategy preservation as a key principle for stable post-training.
The paper addresses a critical failure mode in Reinforcement Learning with Verifiable Rewards (RLVR) for Large Language Models (LLMs), specifically the "catastrophic strategy collapse" observed in GRPO-style algorithms. The authors propose a theoretical framework combining optimization dynamics and information theory to explain this phenomenon, defining "strategies" via trajectory-level policy-update interactions. They introduce the Mirrored Entanglement Index (MEI) as an online diagnostic metric and propose "Mesh Learning," a method designed to preserve strategy diversity by preventing any single reasoning strategy from dominating the optimization landscape. The theoretical contribution is significant as it moves beyond empirical observation to provide a mechanistic explanation for why standard RLVR objectives lead to entropy collapse and reduced effective capacity.
The experimental section evaluates the proposed method across a robust suite of benchmarks including AIME26, AIME25, MATH-500, GPQA, and LiveCodeBench. The models tested include the Qwen and Phi families, which are widely used in the community. The reported gains are substantial, with improvements of up to 13.4 percentage points on Qwen models and 11.5 percentage points on Phi models compared to strong baselines. The consistency of performance across diverse tasks (math, code, general reasoning) suggests that the method effectively preserves the general reasoning capabilities of the model while enhancing specific task performance, rather than overfitting to a narrow distribution of strategies.
The paper provides a public code repository (https://github.com/Ayanami-0123/Open-Mesh-Learning), which is a strong positive for reproducibility. Given the complexity of RLVR training and the specific hyperparameters likely required for "Mesh Learning," the availability of code is crucial. The use of standard, open-source model families (Qwen, Phi) further enhances the reproducibility of the results for other researchers.
The primary limitation is the computational cost associated with RLVR training, which remains high despite the proposed method. Additionally, the definition of "strategies" via trajectory-level interactions may be complex to interpret in non-mathematical or open-ended generation tasks, potentially limiting the direct applicability of the MEI metric to broader domains. The paper focuses on verifiable rewards, so its applicability to RLHF with human preference data (where rewards are less verifiable) is not directly addressed.
This work has significant implications for the stability and reliability of LLM post-training pipelines. By identifying and mitigating strategy collapse, it contributes to the development of more robust and versatile reasoning models. The concept of "strategy preservation" could influence future RLVR algorithms, potentially leading to a new class of training objectives that explicitly balance exploration and exploitation in the policy space. This is particularly relevant as the industry moves towards more complex, multi-step reasoning tasks where diverse strategies are essential. The paper identifies and mitigates catastrophic strategy collapse in RLVR for LLMs by introducing Mesh Learning. It provides a rigorous theoretical framework linking optimization dynamics to information-theoretic capacity constraints, supported by strong empirical results across multiple benchmarks and model families, establishing strategy preservation as a key principle for stable post-training.
Distillation has become a core primitive of large language model training, but its properties are not yet well understood. We take an entropic perspective, studying how the entropy of the student depends on the data and the divergence that define the distillation objective. We prove that forward KL inflates the entropy of the student above that of the teacher. Since cross-entropy training is a special case, this yields an identity that we verify quantitatively in pretraining and supervised finetuning. Other divergences come with no such guarantee: reverse KL deflates entropy until the gap between student and teacher gets too large, and interpolating between the two changes entropy smoothly early in training but abruptly at convergence. The lower entropy of on-policy distillation comes from token-level reverse KL, not from on-policy sampling. The divergence therefore acts as an implicit entropy regularizer, whose role is clearest in self-distillation: as conditioning on privileged information deflates entropy, the divergence hyperparameters that work best are those that compensate for it.
Primary: Stanford University
All Institutions: Stanford University
The paper provides a rigorous theoretical and empirical analysis of how divergence choices in distillation control student entropy, proving that forward KL inflates entropy while reverse KL deflates it, and demonstrating that this divergence acts as an implicit regularizer crucial for preventing entropy collapse in self-distillation. By decoupling sampling and divergence objectives and validating these insights on large-scale models like OLMo and Qwen, the work offers actionable insights for improving the stability and diversity of distilled language models.
The paper proposes a rigorous theoretical framework to analyze the effect of divergence choices in knowledge distillation on the entropy of the student model. The core methodological contribution is the decoupling of the sampling distribution (on-policy vs. off-policy) from the divergence objective (forward KL, reverse KL, etc.) by defining the objective at the token level rather than the sequence level. This allows for independent analysis of how each component affects entropy. The authors introduce a toy model of language modeling (linear softmax head with random embeddings) that is tractable for theoretical analysis yet captures key phenomena of large language models. They prove that forward KL inflates student entropy relative to the teacher by the residual divergence, and that reverse KL deflates entropy until the task becomes too hard, at which point it inflates. The analysis extends to generalized Jensen-Shannon divergences and top-k restrictions.
The experimental evaluation is strong and multi-faceted. The authors validate their theoretical predictions using the OLMo 2 suite of models, demonstrating that the entropy of a model matches its cross-entropy loss at convergence, a finding they claim is novel. They conduct ablation studies on Qwen3 models to show that the divergence choice has a larger impact on entropy than the sampling strategy (on-policy vs. off-policy). Finally, they apply their insights to self-distillation on SciKnowEval and GSM8K, showing that specific divergence hyperparameters are necessary to prevent entropy collapse when conditioning on privileged information. The experiments are well-designed to test specific theoretical claims.
The paper provides a link to a GitHub repository containing code to reproduce all experiments and figures, along with the underlying data. The use of open-source models (OLMo, Qwen3) and public datasets (DeepScaleR, SciKnowEval, GSM8K) further enhances reproducibility. The theoretical proofs are detailed in the appendix, allowing for verification of the mathematical claims.
The theoretical analysis relies on a toy model with random embeddings, which may not fully capture the complexities of real-world language models with deep sequence modeling components. While the authors argue that the token-level objective is what matters in practice due to stop-gradients, the gap between the toy model and large-scale transformers remains a potential limitation. The experiments, while thorough, are limited to specific model families (OLMo, Qwen) and datasets, and the generalizability to other architectures or domains is not explicitly tested.
This paper has significant implications for the design of distillation pipelines in LLM training. By identifying the divergence as an implicit entropy regularizer, it provides practitioners with a new knob to control model uncertainty and diversity. The finding that forward KL inflates entropy and reverse KL deflates it (until failure) offers practical guidance for choosing objectives in post-training and self-distillation. It also challenges the common practice of deriving on-policy distillation from sequence-level reverse KL, suggesting that decoupling these choices can lead to better performance and stability. The paper provides a rigorous theoretical and empirical analysis of how divergence choices in distillation control student entropy, proving that forward KL inflates entropy while reverse KL deflates it, and demonstrating that this divergence acts as an implicit regularizer crucial for preventing entropy collapse in self-distillation. By decoupling sampling and divergence objectives and validating these insights on large-scale models like OLMo and Qwen, the work offers actionable insights for improving the stability and diversity of distilled language models.
Simulating microstructure evolution requires quantum-mechanical accuracy and mesoscopic reach in length and time scales, a combination that no current method achieves. Classical phase-field models provide this reach, but their accuracy is limited by phenomenological free energies and mobilities. Here we develop a framework for learning ab initio phase-field models, where the mesoscopic equation is not postulated but derived from a Mori-Zwanzig projection of molecular dynamics onto species-density fields under explicit assumptions. The nonlocal free energy and mobility left unspecified by this equation are parametrized by neural networks and learned from short molecular dynamics trajectories generated with machine-learning interatomic potentials of ab initio accuracy. We demonstrate the framework on an iron-boron melt and on hydrogen-helium mixtures under planetary conditions. For iron-boron, the model shows that the melt at the FeB$_4$ composition is spinodally unstable at ambient pressure but stabilized at 10 GPa, offering a thermodynamic rationale for why FeB$_4$ has been synthesized only under high pressure. For hydrogen-helium, the model predicts the immiscibility boundary and captures droplet nucleation and growth in helium-rain simulations of a column corresponding to 2.2 million atoms, far beyond the scale of atomistic modeling at comparable accuracy. Trained across compositions and conditions, such models could provide a mesoscopic counterpart to ab initio molecular dynamics.
Primary: National University of Singapore
All Institutions: National University of Singapore
The paper presents a rigorous, physics-informed framework for learning ab initio phase-field models by deriving mesoscopic equations from Mori-Zwanzig projections and jointly learning free-energy and mobility functionals from MLIP-generated MD data, demonstrating significant scalability and accuracy in simulating complex microstructure evolution across materials and planetary science applications.
The paper proposes a rigorous framework for deriving mesoscopic phase-field equations from microscopic molecular dynamics (MD) using the Mori-Zwanzig projection formalism. Unlike standard data-driven approaches that fit phenomenological parameters, this method derives the structure of the evolution equation (Generalized Langevin Equation reduced to a Markovian form) from statistical mechanics. The key innovation is the joint learning of the nonlocal free-energy functional and the mobility tensor using neural networks, trained on short MD trajectories generated by machine-learning interatomic potentials (MLIPs). The authors introduce a "ladder" of free-energy functionals, selecting a specific rung (local density plus nonlocal kernel) that balances expressiveness with data efficiency. The derivation is mathematically sound, explicitly stating assumptions (Markovian dynamics, local mobility, Gaussian noise) that justify the reduction from the exact GLE to the practical model. The use of fluctuation-dissipation relations to anchor the mobility and free-energy curvature provides a strong physical constraint on the learning process, preventing the neural networks from learning unphysical dynamics.
The experimental validation is extensive and covers three distinct systems with varying complexity: a Lennard-Jones mixture (benchmark), an Iron-Boron melt (materials science), and a Hydrogen-Helium mixture (planetary science). For the LJ mixture, the model accurately reproduces the phase diagram and spinodal decomposition dynamics, outperforming Landau and Flory-Huggins baselines. The Iron-Boron results provide a thermodynamic rationale for the high-pressure synthesis of FeB4, correctly predicting the stabilization of the melt at 10 GPa. The most impressive result is the Hydrogen-Helium simulation, where the model predicts the immiscibility boundary consistent with prior DFT studies and simulates helium rain in a domain corresponding to 2.2 million atoms over 1 ns—a scale far beyond the reach of direct MD at comparable accuracy. The ability to extrapolate to larger scales and include external potentials (gravity) not present in training data demonstrates the model's generalizability.
The paper provides detailed descriptions of the training data generation, the specific neural network architectures (MLPs for local terms, radial kernels for nonlocal terms), and the loss function components (dynamical, mobility anchor, structure factor anchor, pressure anchor). The use of standard MLIPs (MACE) for data generation enhances reproducibility, as these potentials are increasingly available. However, specific hyperparameters, network sizes, and training protocols are deferred to the Supplemental Material, which is not fully visible in the provided text. The code availability is not explicitly stated in the main text, which is a minor gap for immediate reproducibility, though the methodology is described with sufficient precision for re-implementation.
The framework currently resolves only density fields, neglecting momentum density, which limits its ability to capture hydrodynamic effects like droplet settling via flow rather than just diffusion. The Markovian assumption may fail for systems with long memory times, such as supercooled liquids or glasses. The accuracy of the model is inherently bounded by the accuracy of the underlying MLIPs used to generate training data. Additionally, the computational cost of generating high-quality MD data for complex materials remains a barrier, although the authors argue that the data efficiency of the approach mitigates this.
This work bridges the gap between quantum-mechanical accuracy and mesoscopic scalability, a long-standing challenge in materials science and fluid dynamics. By providing a principled way to learn mesoscopic models from microscopic data, it enables simulations of phenomena like phase separation, nucleation, and coarsening at scales and timescales previously inaccessible. The potential applications are vast, ranging from alloy design and synthesis to understanding planetary interiors. The framework could serve as a foundation for "mesoscopic foundation models" that generalize across compositions and conditions, analogous to CALPHAD databases but with predictive dynamical capabilities. The paper presents a rigorous, physics-informed framework for learning ab initio phase-field models by deriving mesoscopic equations from Mori-Zwanzig projections and jointly learning free-energy and mobility functionals from MLIP-generated MD data, demonstrating significant scalability and accuracy in simulating complex microstructure evolution across materials and planetary science applications.
Single-particle cryo-electron microscopy (cryo-EM) has become a widely adopted technique for biomolecular structure determination. The conventional cryo-EM computational pipeline first combines many particle images to reconstruct an electrostatic potential (ESP) map and then fits an atomic model to the recovered map. Density reconstruction has high sample complexity, requiring large numbers of particle images and making structure determination high-cost and low-throughput, particularly for heterogeneous samples. Downstream atomic model building, in turn, becomes increasingly difficult as the resolution of the reconstructed map deteriorates. Protein structure prediction models provide strong sequence-derived priors on atomic structure, and experiment-guided approaches can use these priors to recover structures consistent with experimental measurements. Yet, in cryo-EM, such priors are typically integrated only after density reconstruction during atomic model fitting. We introduce Fold'EM, an inference-time framework that combines priors from protein generative models directly with cryo-EM particle images to determine atomic models from a small number of single particle images, bypassing both intermediate density reconstruction and downstream model building against the reconstructed map. Across synthetic and experimental cryo-EM datasets, Fold'EM recovers accurate atomic structures both with known particle orientations and in an ab-initio setting where orientations are inferred jointly with structure. In heterogeneous datasets, Fold'EM further resolves distinct conformational states from mixed particle populations without separately reconstructing a density map and building an atomic model for each state. We believe these results open new avenues for structure determination in the low-sample regime and for characterizing low-population conformational states directly from cryo-EM particles.
Primary: University of Cambridge
All Institutions: University of Cambridge, DeepMind
Fold'EM introduces a novel inference-time framework that directly guides AlphaFold3 using cryo-EM particle images, bypassing intermediate density reconstruction. By optimizing intermediate Pairformer states and jointly inferring structure, pose, and class, the method achieves high-resolution atomic structures from sparse particle sets, significantly reducing the sample complexity of cryo-EM structure determination.
The paper introduces Fold'EM, a framework that bypasses the traditional cryo-EM density reconstruction step by directly guiding AlphaFold3 using individual particle images. The core technical novelty lies in optimizing intermediate Pairformer states ($H$) rather than the terminal embeddings ($Z$) or atomic coordinates. The authors provide a rigorous theoretical justification for this choice, demonstrating that optimizing $H$ induces a sequence-dependent anisotropic geometry in the conditioning space, which acts as a learned preconditioner for the gradient updates. This prevents the model from drifting into unrealistic structural configurations during the inference-time optimization. The method also introduces a joint inference scheme for structure, pose, and class assignment, utilizing a novel bootstrapped residual-subspace estimator (B-RECOVAR) for label-free classification in heterogeneous samples. The mathematical formulation of the particle-level likelihood versus density-level likelihood is well-articulated, highlighting the statistical insufficiency of consensus densities when poses are uncertain.
The experimental evaluation is extensive, covering both synthetic benchmarks and experimental EMPIAR datasets. The paper demonstrates significant improvements over baselines (RELION, cryoDRGN, and previous AlphaFold-guided methods) in terms of FSC resolution and TM-score, particularly in the low-sample regime (e.g., recovering structures from as few as 100-500 particles). The ablation studies effectively isolate the contributions of particle-level guidance versus density guidance and the optimization of intermediate states. The results on heterogeneous data show the ability to resolve distinct conformational states without separate density reconstructions, which is a significant practical advantage. The comparison with unguided AlphaFold3 clearly shows the value of the experimental guidance in correcting mispredicted conformations.
The paper provides detailed algorithmic descriptions and mathematical formulations. However, as an arXiv preprint, the code availability is not explicitly confirmed in the text provided. The reliance on AlphaFold3, a proprietary model, may limit full reproducibility for some users, though the method itself is clearly defined. The use of standard cryo-EM tools (RELION, cryoDRGN) for baselines aids in comparability.
The method relies heavily on the quality of the AlphaFold3 prior; if the unguided prediction is significantly wrong (e.g., 8PWH case), the method may struggle or require more particles than density-based methods. The computational cost of running multiple reverse-diffusion runs and back-propagating through the Pairformer could be high. The current implementation assumes known CTF parameters and centered particles, which may not always hold in raw experimental data without preprocessing.
This work has the potential to significantly reduce the sample size required for cryo-EM structure determination, making it feasible to study low-population states and scarce samples. It bridges the gap between protein structure prediction and experimental cryo-EM, offering a new paradigm for structure determination that could accelerate drug discovery and structural biology research. The insights into optimizing intermediate representations of generative models may also be applicable to other inverse problems in scientific imaging. Fold'EM introduces a novel inference-time framework that directly guides AlphaFold3 using cryo-EM particle images, bypassing intermediate density reconstruction. By optimizing intermediate Pairformer states and jointly inferring structure, pose, and class, the method achieves high-resolution atomic structures from sparse particle sets, significantly reducing the sample complexity of cryo-EM structure determination.
Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composition is lossy. A vote can only overturn mistakes that its voters do not share, and RLVR sharpens a policy so that its samples increasingly make the same mistakes. With every method drawing exactly $160$ completions per problem, training a single LoRA adapter on the full RLVR budget raises single-sample accuracy on every model we test ($1.5$B-$8$B). Yet on three of four models it leaves the majority vote below that of the untrained base model, by up to $4.8$ points. The damage builds during training: voter errors grow steadily more correlated, and the majority vote accuracy peaks early before falling by up to $7.0$ points. The cause is concentration, not RLVR itself. We split the same data and training budget across $K$ LoRA adapters, each trained on its own random disjoint shard, and call the result an adapter thicket. Thickets out-vote the fully trained adapter in all $16$ (model, $K$) settings, and for $K{\geq}4$ they stay within $0.8$ points of the base model or above it. A single adapter stopped early, at a thicket member's step count, is a strong control that matches thickets for small $K$. For $K{\geq}8$, thickets keep more of RLVR's single-sample gain and out-vote this control in six of eight settings. The cost of concentration also grows with the number of votes: from $16$ to $160$ votes, the thicket's lead over the fully trained adapter widens from $1.3$ to $3.3$ points. When the plan is to sample and vote, an RLVR budget is better spent broad than deep.
Primary: Princeton University
All Institutions: Princeton University
The paper demonstrates that splitting RLVR training budgets across multiple LoRA adapters ("thickets") preserves the diversity necessary for effective majority voting, outperforming the standard single-adapter pipeline which suffers from error correlation and diversity collapse.
The paper proposes a simple but effective alternative to the standard RLVR (Reinforcement Learning with Verifiable Rewards) pipeline: instead of training a single LoRA adapter on the full dataset, it splits the data and budget across K independent adapters ("thickets"). The core insight is that RLVR sharpens the policy, reducing sample diversity and increasing error correlation among voters in majority voting, which degrades test-time scaling performance. The methodology is sound, using a fixed inference budget (160 completions) to isolate the effect of training strategy from sampling budget. The use of a step-matched control (early stopping the single adapter) is a critical experimental design choice that strengthens the causal claim.
The experiments are rigorous and well-controlled. The authors test across four models (Qwen2.5-1.5B/3B/7B, Llama-3.1-8B) and two domains (Math, Code). They demonstrate that the single-adapter approach often performs worse than the untrained base model in majority voting, while the thicket approach recovers this performance and often exceeds it. The analysis of error correlation and coverage provides strong mechanistic evidence for the observed results. The ablation on shard type (random vs. subject-specific) adds practical value.
The paper provides sufficient detail on the training setup (GRPO, LoRA rank, hyperparameters) and evaluation protocols (temperature, top-p, voting logic). The specific datasets (MATH, GSM8K, etc.) and model checkpoints are standard, making reproduction feasible for a lab with adequate compute resources. The code is not explicitly linked in the provided text, but the methods are standard enough to implement.
The study is limited to models up to 8B parameters and specific RL algorithms (GRPO). The voting results are primarily for math tasks where answers are canonicalizable; code tasks are only analyzed for coverage and correlation, not final vote accuracy. The "thicket" approach requires managing multiple adapters at inference time, which, while mitigated by modern serving stacks, adds operational complexity compared to a single model.
This paper challenges a common assumption in the LLM post-training community: that more RLVR training is always better for test-time scaling. It provides a clear, actionable guideline for practitioners: if you plan to use majority voting, diversify your training budget. This has immediate practical implications for teams deploying LLMs with test-time compute budgets. The paper demonstrates that splitting RLVR training budgets across multiple LoRA adapters ("thickets") preserves the diversity necessary for effective majority voting, outperforming the standard single-adapter pipeline which suffers from error correlation and diversity collapse.
Biological reasoning models use post-training to connect LLMs to biological foundation model representations and biological text. Their benchmark accuracy is taken as evidence that LLMs reason over these inputs. We test this assumption in six biological reasoning models across DNA, protein, and single-cell tasks. We perturb one biological input while holding the query and other inputs fixed, construct evidence conflicts that pair the foundation model representation of one genome, protein, or cell with the text of another, fit linear probes to the representations the language model receives, and analyze reasoning traces against the biological inputs. Evo2 and ESM3 contribute little to BioReason and BioReason-Pro performance on the evaluated tasks. Shuffling the DNA sequence barely changes BioReason disease prediction accuracy, and in evidence conflicts the two models follow the text in 97.9% and 99.7% of cases. Linear probes trained on the Evo2 and ESM3 representations predict the task targets, so these foundation models encode information relevant to the task, but provide limited overall performance improvement to BioReason and BioReason-Pro. In contrast, foundation model inputs contribute to ChatNT, Prot2Text-V2, and CellWhisperer performance, and differentially expressed genes in the gene sentence contribute to Cell2Sentence-Scale performance. Across SFT and RL checkpoints of BioReason-Pro and 42 BioReason checkpoints, increases in accuracy do not imply greater performance contributions from biological inputs. BioReason traces misstate nucleotide changes, while BioReason-Pro traces describe functions omitted from final predictions under evidence conflicts. We find that current post-training strategies do not ensure that foundation model representations contribute to task performance.
Primary: Harvard University
All Institutions: Harvard University, Harvard Medical School
The paper provides a critical evaluation of biological reasoning models, demonstrating that current post-training strategies often fail to ensure that foundation model representations contribute to task performance. By using rigorous perturbation and evidence conflict analyses, the authors reveal that models like BioReason and BioReason-Pro largely ignore their biological inputs in favor of textual shortcuts, a finding that has significant implications for the design and evaluation of future multimodal scientific AI systems.
The paper employs a rigorous interventionist methodology to test the causal role of biological foundation model representations in LLM-based reasoning models. The authors utilize input perturbation (shuffling DNA sequences), evidence conflict construction (mismatching foundation model embeddings with text descriptions), and linear probing to isolate the contribution of specific modalities. This approach is methodologically sound for determining whether models are genuinely "reasoning" over biological inputs or merely relying on textual shortcuts. The inclusion of analysis across multiple training checkpoints (SFT and RL) adds depth to the evaluation of how post-training strategies affect input utilization.
The experiments cover six distinct biological reasoning models across DNA, protein, and single-cell tasks, providing a broad scope. The key finding that Evo2 and ESM3 contribute negligibly to BioReason and BioReason-Pro performance, with models following text in >97% of conflict cases, is a significant empirical result. The contrast with models like ChatNT and CellWhisperer, where foundation model inputs do contribute, highlights that the issue is specific to certain post-training strategies or model architectures rather than a universal failure of multimodal integration. The analysis of reasoning traces revealing misstatements of nucleotide changes further supports the conclusion that the models are not faithfully using the biological inputs.
The paper provides a GitHub repository link (https://github.com/mims-harvard/bio-mirage) and a project website, which strongly suggests that code and resources will be available for reproduction. The detailed description of the perturbation and conflict construction methods allows other researchers to replicate the analysis on other models.
The study focuses on a specific set of six models and three biological domains. It is unclear if the findings generalize to other types of biological foundation models or different post-training paradigms. The linear probes may not capture all non-linear relationships between the foundation model representations and the task targets, potentially underestimating the contribution of the biological inputs in some cases.
This paper has high impact for the field of AI in science, particularly for the development of trustworthy biological reasoning systems. It challenges the common assumption that high benchmark accuracy implies effective use of specialized scientific inputs. The findings will likely influence future post-training strategies to explicitly reward the use of foundation model representations, leading to more robust and interpretable AI models for biology. The paper provides a critical evaluation of biological reasoning models, demonstrating that current post-training strategies often fail to ensure that foundation model representations contribute to task performance. By using rigorous perturbation and evidence conflict analyses, the authors reveal that models like BioReason and BioReason-Pro largely ignore their biological inputs in favor of textual shortcuts, a finding that has significant implications for the design and evaluation of future multimodal scientific AI systems.
Learning to complete tasks in unfamiliar environments with unknown rules remains a key challenge for LLM agents. Current LLM agents often record their discoveries in prose, which may not provide a compact, explicit account of how the environment works. Inspired by how scientists organize observations into testable, predictive theories, we introduce Schema, an agent harness that organizes learning and action through interactive program induction. The LLM agent decides what to investigate and how to act, expressing its evolving understanding of the environment as executable programs. The harness consists of a persistent program workspace and a small set of interfaces for checking these programs against the interaction history, planning within them, and executing plans under step-by-step verification. Schema raises ARC-AGI-3 RHAE from 58.7% to 99.2% with the same base model, solves 100% of the public DiG-bench games, and reaches the median performance of the top-50 human players on MazeBench. Extensive analysis shows the effectiveness of Schema in unknown mechanism discovery, and ablations confirm the contribution of each component.
Primary: Carnegie Mellon University
All Institutions: Carnegie Mellon University, University of California, Berkeley, Impossible AI
The paper presents a novel agent harness, Schema, that utilizes interactive program induction to enable LLMs to discover and exploit unknown environment rules with human-level efficiency. By organizing learning through executable programs rather than prose, the method achieves state-of-the-art results on ARC-AGI-3, DiG-bench, and MazeBench, demonstrating a robust and scalable approach to mechanism discovery in unfamiliar settings.
The paper introduces "Interactive Program Induction" (IPI), a paradigm where LLM agents represent their understanding of unknown environments as executable programs rather than prose. The proposed harness, Schema, consists of four core operations: Hypothesize (writing code to model state and transitions), Certify (replaying history to check consistency), Plan (using the program as a simulator for search), and Act with Verification (executing actions and stopping if predictions mismatch). This approach effectively addresses the "lost in the middle" and context degradation issues inherent in prose-based memory by forcing the agent to distill knowledge into compact, testable, and reusable code. The methodology is sound, leveraging the LLM's coding capabilities to create a persistent, verifiable world model that persists across context compactions.
The evaluation is extensive and rigorous, covering three distinct benchmarks: ARC-AGI-3 (visual reasoning), DiG-bench (text-based rule discovery), and MazeBench (long-horizon 3D exploration). The results are striking: Schema achieves 99.2% RHAE on ARC-AGI-3 (vs. 58.7% baseline), solves 100% of public DiG-bench games, and matches top-50 human performance on MazeBench. Ablation studies clearly demonstrate the contribution of each component (certification, planning, verification), showing that removing any one significantly degrades performance. The analysis of token costs and action efficiency further strengthens the claim of practical utility.
The paper provides detailed implementation descriptions in the appendix, including the program contract, tool interfaces, and benchmark adapters. However, the specific code for the Schema harness and the exact prompts used are not fully detailed in the text, and no public code repository URL is provided in the extracted text. While the methodology is clearly described, full reproducibility would require access to the specific harness implementation and prompt engineering details.
The approach relies heavily on the base model's coding ability; weaker models may struggle to write correct world models. The computational cost of backtesting and planning can be high, though the paper argues this is offset by reduced interaction steps. The benchmarks used (ARC-AGI-3, DiG-bench, MazeBench) are relatively new and may not fully represent the breadth of real-world unknown environments.
This work has significant implications for the design of autonomous agents in open-ended environments. By shifting from passive memory to active, executable theory-building, it offers a scalable path toward agents that can genuinely learn and adapt to novel tasks without retraining. The paradigm of "interactive program induction" could be applied to robotics, scientific discovery, and other domains where environments are complex and rules are not explicitly given. The paper presents a novel agent harness, Schema, that utilizes interactive program induction to enable LLMs to discover and exploit unknown environment rules with human-level efficiency. By organizing learning through executable programs rather than prose, the method achieves state-of-the-art results on ARC-AGI-3, DiG-bench, and MazeBench, demonstrating a robust and scalable approach to mechanism discovery in unfamiliar settings.
Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost. We build PhantomEnvironments, multi-turn RL environments from fictional worlds, where agents must search a corpus of templated articles to answer multi-hop questions. Despite sharing no facts with the real world, these strikingly simple environments yield agents that transfer to real-world multi-hop search benchmarks, often outperforming real-world training data on newer benchmarks. Trained agents generalize to unseen fictional universes, and Qwen models learn to scale their search budget roughly linearly with question difficulty, suggesting emergent search scaling from environment interaction alone. Ablating environment complexity reveals that hop count drives transfer more than constraints or comparisons: even the simplest rule-generated environments are a surprisingly effective, free resource for training generalizable LLM agents.
Primary: Cornell University
All Institutions: Cornell University, Stanford University
The paper introduces a cost-effective, rule-based synthetic environment framework for training LLM search agents, demonstrating that structural complexity (hop count) drives transferability to real-world tasks more than factual accuracy. This approach offers a scalable alternative to human-curated or LLM-generated data, enabling the emergence of adaptive search strategies in LLM agents.
The paper proposes a novel paradigm for training LLM agents using Reinforcement Learning (RL) by replacing costly, human-curated, or LLM-generated environments with "PhantomEnvironments." These are synthetic, rule-based environments derived from fictional worlds where agents must perform multi-hop search over templated articles. The core methodological innovation is the decoupling of environment generation from LLM inference, ensuring zero marginal cost and eliminating hallucination risks associated with LLM-generated data. The approach relies on the hypothesis that the structural complexity of the search task (specifically hop count) is more critical for learning generalizable search strategies than the factual accuracy of the content.
The experiments demonstrate that agents trained on these fictional, factually incorrect environments transfer effectively to real-world multi-hop search benchmarks. Notably, the paper claims that agents trained on PhantomEnvironments often outperform those trained on real-world data on newer benchmarks, suggesting that the structural learning of search strategies is robust to domain shift. Ablations confirm that hop count is the primary driver of transfer performance. The observation of "emergent search scaling," where Qwen models learn to allocate search budget linearly with question difficulty, is a significant empirical finding.
The authors provide a strong reproducibility statement, indicating the use of open-source LLMs, training code, and evaluation benchmarks. They report standard errors and significance tests. A GitHub repository is provided for code access. The use of AI tools for code implementation is disclosed, but the authors state they verified all code and results.
The primary limitation is the reliance on "templated articles" and rule-based generation, which may not capture the full stochasticity and ambiguity of real-world web search. The transferability to "newer benchmarks" is a strong claim that requires careful scrutiny regarding potential leakage or specific alignment with the benchmark structure. The paper is a preprint (ICLR 2027 submission), so peer review status is pending.
This work has high potential impact by providing a scalable, cost-effective method for training LLM agents. If the transferability results hold, it could significantly reduce the barrier to entry for developing sophisticated search agents, allowing smaller labs to train competitive models without massive human annotation budgets. It shifts the focus from data fidelity to structural complexity in agent training. The paper introduces a cost-effective, rule-based synthetic environment framework for training LLM search agents, demonstrating that structural complexity (hop count) drives transferability to real-world tasks more than factual accuracy. This approach offers a scalable alternative to human-curated or LLM-generated data, enabling the emergence of adaptive search strategies in LLM agents.
On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action that moves the agent farther from completing the task, and this mistake typically occurs early. These pivotal mistakes often remain recoverable: guiding the model for only a few turns after the pivotal turn can restore task success. We therefore propose PivotOPD, an on-policy distillation framework that jointly trains the student to prevent pivotal mistakes and to recover from the states they create. At each pivotal mistake, a teacher model provides a gold action and then names a recovery action at each of the next few turns. Preventive distillation uses the gold action with reverse KL to steer the student away from the pivotal mistake, while recovery distillation uses the recovery actions with forward KL to transfer recovery behaviors that the student rarely samples. Against 13 baselines on ALFWorld, WebShop, and Search-based QA, PivotOPD achieves the strongest average performance for both Qwen3-1.7B and Qwen3-8B students, improving over the strongest baseline on ALFWorld by +5.5% with the 1.7B student. The gains also transfer to another model family on the software engineering domain, where PivotOPD raises the resolve rate of a Nemotron-3.5 student on SWE-Bench Verified by +3.2%. Project page: https://research.nvidia.com/labs/lpr/pivotopd/
Primary: NVIDIA
All Institutions: NVIDIA
PivotOPD introduces a targeted distillation framework that teaches LLM agents to recover from pivotal mistakes in multi-turn interactions. The paper provides a rigorous analysis of error accumulation, a novel dual-distillation method, and strong empirical results across diverse benchmarks, establishing a new standard for training robust language agents.
The paper proposes PivotOPD, a framework that addresses error accumulation in multi-turn LLM agents by identifying "pivotal mistakes" and applying targeted distillation. The core innovation is the dual distillation approach: preventive distillation (using reverse KL to steer away from mistakes) and recovery distillation (using forward KL to teach recovery behaviors from a privileged self-teacher). The theoretical analysis correctly identifies that standard on-policy methods fail to provide sufficient learning signal for recovery actions because the student rarely samples them, justifying the use of forward KL on teacher-generated responses. The pivot detection mechanism, which uses a teacher model to identify candidate turns and gold actions, is a practical solution to the lack of oracles in real-world environments.
The experiments are extensive, covering four benchmarks (ALFWorld, WebShop, Search-based QA, SWE-Bench Verified) and multiple model families (Qwen3, Nemotron). The comparison against 13 baselines is robust. The results show consistent improvements, particularly in recovery rates, which aligns with the paper's motivation. The ablation studies effectively demonstrate the necessity of both preventive and recovery components and the importance of supervision at the correct turns. The transfer to SWE-Bench Verified with a different model family (Nemotron) strengthens the generalizability claim.
The paper provides detailed descriptions of the training objective, pivot detection prompts, and evaluation metrics. The project page URL is provided, which likely contains code and additional resources. The hyperparameters and training configurations are described in the appendix, supporting reproducibility.
The method relies on a teacher model for pivot detection and action naming, which may not be available in all settings. The performance gains, while statistically significant, are moderate in some benchmarks (e.g., +1.2% on WebShop success rate). The reliance on a privileged self-teacher for recovery distillation adds computational overhead. The pivot detection accuracy (77.8% within one turn of oracle) suggests room for improvement in identifying pivotal turns without an oracle.
This work has significant implications for training robust LLM agents in interactive environments. By explicitly teaching recovery from mistakes, it addresses a key limitation of current agent training methods. The framework could be extended to other domains where error accumulation is a challenge, such as robotics or autonomous driving. The insights into the nature of pivotal mistakes and recovery behaviors provide valuable guidance for future research on agent robustness. PivotOPD introduces a targeted distillation framework that teaches LLM agents to recover from pivotal mistakes in multi-turn interactions. The paper provides a rigorous analysis of error accumulation, a novel dual-distillation method, and strong empirical results across diverse benchmarks, establishing a new standard for training robust language agents.
We compare the instance-wise, finite-sample risks of monotone spectral filters for linear regression, a broad class of estimators including principal component regression (PCR), gradient descent (GD), and ridge regression. We show that PCR dominates all monotone spectral filters: compared to any such filter, the risk of optimally tuned PCR is no bigger by a constant factor for all problems. Furthermore, the dominance is strong if the filter is separated from step functions (e.g., GD and ridge): there exist problem instances for which the risk of PCR is smaller by a polynomial factor in sample size dependence. Our comparison results show that PCR is optimal and thus admissible among monotone filters, significantly extending Wu et al. (2026)'s result that GD strongly dominates ridge. From a technical perspective, we establish new upper and lower bounds for general spectral filters, which are instance-wise sharp when specialized to ridge or GD, recovering or improving the best-known bounds.
Primary: UC Berkeley
All Institutions: UC Berkeley, Princeton University
Principal Component Regression dominates all monotone spectral filters for linear regression in terms of instance-wise finite-sample risk, establishing PCR as admissible and GD as inadmissible. The paper provides a comprehensive theoretical analysis using Schur multipliers to derive sharp risk bounds, offering a new perspective on the comparison of regularization methods beyond worst-case minimax rates.
The paper employs a rigorous theoretical framework to compare the instance-wise finite-sample risks of monotone spectral filters in linear regression. The core methodological contribution is the application of Schur multipliers and matrix divided differences to control noncommutative matrix perturbations, specifically the "leave-tail-out" difference between the full Gram matrix and its truncated version. This allows for the derivation of sharp upper and lower bounds for general spectral filters, including Principal Component Regression (PCR), Gradient Descent (GD), and Ridge Regression. The authors demonstrate that PCR dominates all monotone filters and strongly dominates filters separated from step functions (like GD and Ridge), establishing PCR as admissible and GD as inadmissible in this context. The technical depth is high, extending classical leave-one-out ideas with advanced operator theory tools.
This is a purely theoretical paper with no experimental evaluation. The "experiments" are mathematical proofs and derivations of risk bounds. The validity is established through rigorous mathematical argumentation rather than empirical benchmarks.
As a theoretical paper, reproducibility is defined by the clarity and correctness of the proofs. The paper provides detailed appendices with missing proofs and clearly states assumptions. The use of AI for parts of the technical ingredients is disclosed, but the authors state they rederived and verified all proofs. The mathematical framework is well-defined and reproducible in the sense that other researchers can verify the theorems.
The dominance results rely on Gaussian random design assumptions, which may not hold in all practical settings. The paper acknowledges that the variance bounds for PCR could likely be improved. Additionally, the results are specific to linear regression; extending these dominance relationships to non-linear models or other learning tasks remains an open question. The reliance on specific spectral properties (monotonicity) limits the scope to a specific class of estimators.
The paper significantly impacts the field of statistical learning theory by providing a definitive instance-wise comparison of standard linear regression methods. It challenges the common heuristic that Gradient Descent is a universally good default by showing it is inadmissible compared to PCR in terms of instance-wise risk. This insight may influence the design of future algorithms, encouraging the use of methods that can explicitly discard weak spectral components to reduce variance. It also provides a new technical toolkit (Schur multipliers for risk analysis) that can be applied to other statistical learning problems. Principal Component Regression dominates all monotone spectral filters for linear regression in terms of instance-wise finite-sample risk, establishing PCR as admissible and GD as inadmissible. The paper provides a comprehensive theoretical analysis using Schur multipliers to derive sharp risk bounds, offering a new perspective on the comparison of regularization methods beyond worst-case minimax rates.
Generative models can synthesize high-quality inauthentic multimedia content that is already being misused at scale. We evaluate twenty deepfake detectors against ten generators released in the last four years and find accuracy decreasing over time, from near-perfect 99.5% to 76%. Adversarial perturbations further reduce every baseline detector to below 2% accuracy, effectively inverting the detector's assigned label. We argue that this unreliability reflects a fundamental ambiguity: generators can reproduce authentic content exactly (e.g., through memorization), so content alone cannot reveal the true provenance label.For this reason, content produced by a generator must admit a faithful reconstruction by that same generator, and finding such a reconstruction makes synthetic provenance plausible and authenticity plausibly deniable.We therefore propose and evaluate a detection paradigm that outputs a calibrated prediction of whether authenticity is plausibly deniable: a faithful reconstruction by any known generator establishes plausible deniability, while calibration bounds how often content from known generators fails to be reproduced. Our evaluation shows that (i) our detector can be calibrated so that at most 1% of generated content is wrongly certified, an operating point at which most baseline detectors reach near-zero recall, including the strongest with 93% accuracy; (ii) calibrating a stricter security threshold on attacked samples preserves this bound against adaptive adversaries within the evaluated bounded-perturbation attack space, whose perturbations break every baseline, but does not cover arbitrary adversarial transformations; and (iii) post-hoc verifiability is eroding, as 1,116 of 3,000 Reddit images resist reproduction by a 2022 generator, but only 55 to 79 resist reproduction by 2024 generators.
Primary: Mohamed bin Zayed University of Artificial Intelligence
All Institutions: Mohamed bin Zayed University of Artificial Intelligence
The paper introduces a calibrated, reconstruction-based framework for certifying image authenticity that outperforms traditional detectors under strict error-rate constraints. By inverting the detection problem to test for "plausible deniability" via generator reconstruction, the authors provide a robust, training-free method that withstands adversarial perturbations better than existing baselines, while also quantifying the fundamental erosion of post-hoc verifiability as generative models improve.
The paper proposes a "sound" detection paradigm that inverts the standard deepfake detection logic. Instead of classifying images as real or fake, it certifies authenticity only if no known generator can faithfully reconstruct the image. The method relies on flow-matching inversion (RF-Inversion) to reconstruct images from known open-weight generators (SD2.1, SD3, FLUX, etc.). It introduces a calibrated "A-index" combining PSNR, SSIM, LPIPS, and CLIP similarity. A key methodological contribution is the dual-threshold calibration: a "safety" threshold for clean data and a "security" threshold calibrated against adversarial perturbations (PGD attacks) to ensure robustness. The approach is training-free for the detector itself, requiring only recalibration when new generators are released.
The evaluation is extensive, benchmarking 20 existing detectors against 10 generators released over four years. It demonstrates a clear trend of decreasing detector accuracy over time (99.5% to 76%) and the vulnerability of all baselines to adversarial perturbations. The proposed method achieves a 1% false positive rate (FPR) while maintaining reasonable recall, whereas baselines drop to near-zero recall at this strict FPR. The paper also includes a sociotechnical study on 3,000 Reddit images, showing that the "certifiable" share of real images erodes as generators improve (from 37.2% in 2022 to ~2% in 2024), highlighting a fundamental limit of post-hoc detection.
The authors release code for the inversion pipeline, scoring metrics, and calibration scripts. The method relies on public open-weight models, making it highly reproducible for the specific generators tested. The calibration procedure is clearly defined (quantile-based thresholds), allowing users to adjust the error rate trade-off.
The primary limitation is the shrinking coverage: as generators become more capable, they can reproduce more authentic content, leaving fewer images that can be certified as "real." The method is computationally expensive (11.67 seconds per image/generator), making it unsuitable for real-time feed screening. It also relies on the assumption that the adversary is "efficient" (bounded compute) and that the generator set is known and open-weight; private or fine-tuned models are outside the certification scope.
This paper shifts the focus from "detecting fakes" to "certifying truth," which is a more robust and legally/ethically defensible stance. It provides a framework for platforms to issue verifiable certificates of authenticity. The finding that post-hoc verifiability is eroding is a significant warning to the field, suggesting that watermarking and provenance metadata may be more critical than post-hoc detection in the long run. The paper introduces a calibrated, reconstruction-based framework for certifying image authenticity that outperforms traditional detectors under strict error-rate constraints. By inverting the detection problem to test for "plausible deniability" via generator reconstruction, the authors provide a robust, training-free method that withstands adversarial perturbations better than existing baselines, while also quantifying the fundamental erosion of post-hoc verifiability as generative models improve.
Echocardiography is the most widely used cardiac imaging modality, yet interpretation demands integrating visual evidence across global anatomy, localized structures and dynamic cardiac motion. Machine-learning models have automated individual tasks, but they are typically built for a single purpose and depend on expensively labeled datasets - a barrier particularly acute in pediatric care, where data are scarce and anatomy changes with age. Here we present EchoDino, a self-supervised foundation model for echocardiography, created by adapting the DINOv3 framework to 3.7 million frames from 1.7 million unlabeled pediatric echocardiography videos. With its encoder frozen, EchoDino produces representations that capture global context, local anatomy, and dense spatial detail. We introduce Motion-biased Entropy Maximization Sampling (MEMS) to select the most informative frames for video-level analysis. Across nine pediatric and adult datasets, EchoDino outperformed strong baseline models, raising view-classification accuracy from 0.609 to 0.889 and the area under the receiver operating characteristic curve for structural-heart-disease detection from 0.811 to 0.872, while also cutting age-estimation error from 3.857 to 1.389 years, achieving the best segmentation accuracy and lowering ejection-fraction errors. By generalizing from label-free pediatric data to adult echocardiography, EchoDino offers a versatile foundation for cardiac image analysis across the lifespan.
Primary: Rice University
All Institutions: Rice University, Baylor College of Medicine, Texas Children's Hospital
EchoDino adapts the DINOv3 self-supervised framework to pediatric echocardiography, introducing a frozen encoder with spatial patch re-aggregation and motion-biased frame sampling to achieve state-of-the-art performance across diverse cardiac imaging tasks in both pediatric and adult populations.
The paper proposes EchoDino, a self-supervised foundation model for echocardiography by adapting the DINOv3 framework. The core methodological contribution is the domain adaptation of a natural image foundation model to pediatric cardiac ultrasound using 3.7 million unlabeled frames. The authors introduce two specific architectural/algorithmic extensions: EchoDino-PATCH, which re-aggregates spatial patch tokens to preserve local anatomical detail for segmentation and measurement, and MEMS (Motion-biased Entropy Maximization Sampling), a frame selection strategy for video-level tasks that prioritizes high-motion, feature-diverse frames over uniform sampling. The approach relies on a frozen encoder with lightweight task-specific readouts, which is a standard but effective paradigm for foundation models. The adaptation of DINOv3's teacher-student architecture to the specific augmentation and cropping needs of echocardiography is well-motivated.
The experimental evaluation is extensive and rigorous. The model is tested across nine datasets, including five internal pediatric cohorts, one external pediatric cohort, and three external adult datasets. This cross-domain evaluation (pediatric to adult) is a strong point, demonstrating the generalizability of the learned representations. The tasks cover a wide spectrum: global (view classification), localized (measurement, SHD detection), dense (segmentation), and temporal (EF prediction, sweep recognition). The results show consistent and significant improvements over strong baselines like DINOv3, PanEcho, and EchoPrime. The use of bootstrap confidence intervals and patient-disjoint splits adds statistical rigor. The ablation of MEMS vs. uniform sampling clearly isolates the benefit of the proposed sampling strategy.
The paper provides a link to the code repository. However, the primary pretraining data (TCH-Complex) is not publicly available due to clinical data governance, which limits full reproducibility of the pretraining phase. The downstream evaluation on public datasets (EchoNet, CAMUS) is reproducible. The hyperparameters and training details are described in the supplementary methods, aiding partial reproducibility.
The main limitation is the lack of public access to the large-scale pediatric pretraining corpus, which prevents independent verification of the pretraining process. The SHD detection is evaluated as a binary composite classifier, which may mask performance on specific rare lesions. The paper acknowledges that aggregate metrics do not guarantee individual point-of-care actionability and that uncertainty estimation is needed for clinical deployment.
This work has significant potential impact in the field of medical imaging AI, particularly in pediatric cardiology where labeled data is scarce. By demonstrating that a single frozen encoder can support diverse tasks across age groups, it offers a scalable framework for developing modular clinical tools. The approach of adapting general-purpose vision foundation models to specific medical imaging modalities is a trend that this paper contributes to meaningfully. EchoDino adapts the DINOv3 self-supervised framework to pediatric echocardiography, introducing a frozen encoder with spatial patch re-aggregation and motion-biased frame sampling to achieve state-of-the-art performance across diverse cardiac imaging tasks in both pediatric and adult populations.
Visual token pruning is an effective way to accelerate vision-language models and is especially useful for vision-language-action (VLA) inference, where many visual tokens must be processed before predicting robot actions. Existing pruning methods usually estimate which tokens can be pruned based on attention scores or feature diversity, retaining tokens that are either highly attended or visually different from others. However, most of them use fixed pruning schedules, such as pruning once at a preset layer or pruning at uniformly spaced layers. Such schedules can be risky for VLA models, because the model may not know which visual regions matter for the action in early layers. Tokens that look unimportant at first may become useful after the model combines visual observations with the language instruction. In this work, we propose SAPrune, a training-free visual token pruning framework for efficient VLA inference. Instead of pruning at fixed layers, SAPrune uses a small calibration set to observe how action-to-visual attention changes across layers, and chooses pruning layers only after the attention pattern becomes more reliable. At each selected layer, SAPrune applies a dual-path pruning rule: one path protects strongly attended visual tokens from pruning, while the other prevents useful surrounding context from being discarded. Experiments on LIBERO, SIMPLER, and real-world robotic tasks show that SAPrune prunes 87.5% of visual tokens and achieves up to 1.718x inference speedup while maintaining competitive task success rates.
Primary: Unknown
All Institutions: Unknown
SAPrune introduces a stage-aware, calibration-based approach to visual token pruning for VLA models, significantly improving inference efficiency while maintaining task performance. The paper demonstrates that dynamic, attention-driven pruning schedules outperform static methods, offering a practical path to deploying large VLA models in real-world robotic applications.
The paper proposes SAPrune, a training-free visual token pruning framework for Vision-Language-Action (VLA) models. The core contribution is a two-stage process: (1) Stage-Aware Pruning-Layer Selection, which uses a small calibration set to analyze action-to-visual attention dynamics across layers to identify "stable" stages where pruning is safe, avoiding the pitfalls of fixed or uniform pruning schedules; and (2) Layer-Local Dual-Path Pruning, which retains tokens based on both high attention scores (action-core) and functional diversity relative to the action and instruction queries (coverage path). The methodology is well-motivated by the observation that action-relevant visual evidence emerges and stabilizes in deeper layers, making early pruning risky. The use of calibration data to determine pruning locations is a clever, practical addition that distinguishes it from static pruning methods.
The experiments are comprehensive, covering three major VLA backbones (OpenVLA, OpenVLA-OFT, Pi0) and multiple benchmarks (LIBERO, SIMPLER, and real-world robotic tasks). The results demonstrate that SAPrune can prune 87.5% of visual tokens while maintaining competitive success rates and achieving significant speedups (up to 1.718x). The comparison against strong baselines like FastV, DivPrune, and SP-VLA is rigorous. The inclusion of real-world experiments on an AgileX Piper robot adds significant practical value. Ablation studies confirm the importance of both the stage-aware layer selection and the dual-path retention mechanism.
The paper provides detailed algorithmic descriptions and hyperparameters. The reliance on a calibration set is a potential barrier, but the authors show that the calibration is lightweight (30-60 rollouts) and transfers well across environments. The code is not explicitly linked in the provided text, but the detailed methodology suggests high reproducibility for those willing to implement the calibration and pruning logic.
The method requires a calibration phase, which may not be feasible for all deployment scenarios. The performance gain, while significant, is primarily in latency and FLOPs; the success rate improvement over some baselines at the same token budget is modest. The method is specific to VLA models and may not generalize directly to other multimodal tasks without adaptation.
This work is highly relevant to the robotics and embodied AI community, where real-time inference is a critical bottleneck. By enabling aggressive token pruning without significant performance loss, SAPrune makes advanced VLA models more deployable on edge devices and in real-time control loops. It sets a new standard for efficient VLA inference. SAPrune introduces a stage-aware, calibration-based approach to visual token pruning for VLA models, significantly improving inference efficiency while maintaining task performance. The paper demonstrates that dynamic, attention-driven pruning schedules outperform static methods, offering a practical path to deploying large VLA models in real-world robotic applications.
In an image latent space, the embeddings of high-resolution, natural, and sharp images form a manifold. Degradation of high-resolution images pushes their embeddings off this manifold. Real-world super-resolution (SR) then becomes the task of mapping the degraded embedding back onto this manifold --- not anywhere on the manifold, but to the point that preserves what the input still carries, both its semantics and pixel details. Every published method implements this mapping in a reconstruction-oriented latent space or pixel space. We claim these spaces are the wrong substrates for SR. Low-resolution and degraded images are embedded far from the manifold, making the mapping difficult and expensive. The lack of semantic information in these substrates also makes it difficult to navigate to the faithful point on the manifold, causing severe hallucination when degradation is heavy. Thus, restoring in a suitable latent space is crucial to the SR task. We show that the latent space of 23 fused layers of a frozen DINOv3-L is one such space that makes the SR task easier. Degraded images are embedded near the manifold. Moreover, this substrate contains a hierarchy of information, from pixel record to degradation robust semantics, guiding the model to find the faithful point on the manifold. On this substrate, a 415M decoder is trained under reconstruction and adversarial objectives to map the degraded embeddings back and decode to pixel space in one pass. The resulting model, RAESR, attains the best fidelity--perception trade-off among state-of-the-art adversarial and diffusion-based restorers on RealSR, DRealSR, LSDIR and DIV2K-Val, at 37 ms per 512 by 512 image on a single H20 GPU. Swapping the substrate for a VAE latent under an identical recipe loses on every metric.
Primary: University of Michigan
All Institutions: University of Michigan, Southeast University, Alibaba Group
The paper introduces RAESR, a super-resolution model that leverages the latent space of a frozen DINOv3-L vision transformer as a substrate for restoration, demonstrating that this space embeds degraded images closer to the clean-image manifold than VAEs, thereby enabling a more efficient and faithful one-pass mapping back to high-quality images. The technical contribution is significant due to the rigorous geometric analysis of the latent space properties and the strong empirical results showing state-of-the-art fidelity-perception trade-offs with high inference efficiency.
The paper proposes RAESR, a super-resolution model that operates in the latent space of a frozen DINOv3-L vision transformer rather than a VAE or pixel space. The core methodological contribution is the argument that self-supervised representation spaces (specifically DINOv3) embed degraded images closer to the "manifold" of clean images than reconstruction-oriented spaces (VAEs), thereby simplifying the mapping back to the clean manifold. The authors validate this geometric intuition with rigorous ablations, including latent line interpolation experiments and layer-wise information recovery analysis. The decoder is a 415M parameter transformer trained with reconstruction and adversarial losses, using a multi-layer DINOv3-B critic for fine-grained realism supervision. The approach is technically sound, well-motivated, and the ablations are exceptionally thorough, providing strong evidence for the central hypothesis.
The experiments are comprehensive, comparing RAESR against 12 state-of-the-art methods across four real-world benchmarks (RealSR, DRealSR, LSDIR, DIV2K-Val). RAESR achieves the best fidelity-perception trade-off, outperforming heavy diffusion-based models in both quality and efficiency (37ms per image). The inclusion of a human preference study and detailed per-benchmark breakdowns adds robustness. The ablation studies, particularly the comparison against a VAE substrate with identical training recipes, are critical and convincingly demonstrate that the performance gain stems from the choice of latent space rather than just the decoder architecture.
The paper provides extensive details on training schedules, hyperparameters, data processing, and evaluation protocols in the appendices. The use of standard datasets and public baselines facilitates reproduction. However, the reliance on specific frozen checkpoints (DINOv3-L, RAEv2) and the complex multi-stage training process may pose some barriers for independent replication without access to the authors' code or precise checkpoint versions.
The model is a single-pass restorer, lacking the ability to trade compute for quality on difficult images, unlike iterative diffusion methods. The claim about substrate suitability is currently supported primarily by DINOv3-L and tested against one alternative (SD3 VAE), leaving the behavior of other representation encoders unexplored. The model size (719M inference parameters) is still significant compared to lightweight GANs, though it is competitive with diffusion models.
This work shifts the perspective in super-resolution research from merely improving generative priors to carefully selecting the representation space for restoration. It highlights the utility of self-supervised vision foundation models for low-level vision tasks, potentially inspiring similar approaches in other restoration tasks like denoising or inpainting. The efficiency gains over diffusion models make it attractive for real-time applications. The paper introduces RAESR, a super-resolution model that leverages the latent space of a frozen DINOv3-L vision transformer as a substrate for restoration, demonstrating that this space embeds degraded images closer to the clean-image manifold than VAEs, thereby enabling a more efficient and faithful one-pass mapping back to high-quality images. The technical contribution is significant due to the rigorous geometric analysis of the latent space properties and the strong empirical results showing state-of-the-art fidelity-perception trade-offs with high inference efficiency.
Dense 3D point tracking has been a prominent paradigm for modeling motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel: it captures where pixels go, not the rotation of the underlying part, nor which pixels move together as one body. We propose MoSE3, the first feed-forward model that predicts dense SE(3) motion from monocular RGB video, producing full 6-DoF rigid transforms at every pixel in world space. Per-pixel SE(3) motion offers a richer view of how a scene moves: rotation, translation, and grouping all at once. Directly predicting SE(3) is challenging: rotations lie on a curved manifold that is ill-suited to Euclidean regression, and annotations for SE(3) are particularly difficult to acquire. To address these challenges, MoSE3 predicts per-pixel SE(3) through two jointly learned intermediates, 3D point tracks and rigidity embeddings, and recovers SE(3) by differentiably fitting transforms within each soft rigid cluster, enabling end-to-end prediction and supervision. To close the data gap, we introduce Art-Kubric, a large-scale synthetic dataset with dense SE(3) and rigidity labels for articulated objects with rich physical interactions. MoSE3 achieves state-of-the-art SE(3) estimation at pixel, part, and object levels on both rigid and articulated benchmarks, and state-of-the-art average 3D point tracking accuracy across three datasets, while showing strong generalization to real-world videos despite being trained solely on synthetic motion data.
Primary: Harvard University
All Institutions: Harvard University, Johns Hopkins University, Kempner Institute
[One sentence main contribution]. MoSE3 introduces the first feed-forward model for dense SE(3) motion estimation from monocular video, utilizing a novel decomposition into point tracks and rigidity embeddings to handle the non-Euclidean nature of rotations, and contributes a large-scale synthetic dataset, Art-Kubric, to enable training and evaluation. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper addresses a fundamental gap in motion modeling by moving beyond 3-DoF translation to full 6-DoF rigid transforms. The technical approach is sound, leveraging differentiable optimization to fit SE(3) transforms within learned rigid clusters, which effectively circumvents the difficulties of direct manifold regression. The introduction of Art-Kubric is a substantial contribution to the field, providing the necessary ground truth data for this complex task. The experimental results are strong, showing SOTA performance on both the new SE(3) task and existing point tracking benchmarks, with impressive generalization to real-world data. This work is likely to influence future research in dense motion estimation and provide a valuable tool for applications requiring detailed understanding of scene dynamics.
The paper proposes MoSE3, a feed-forward architecture that predicts dense SE(3) motion (6-DoF: rotation and translation) from monocular RGB video. The core innovation lies in avoiding direct regression on the non-Euclidean SO(3) manifold by decomposing the problem into two jointly learned intermediates: 3D point tracks (translations) and rigidity embeddings. The SE(3) transforms are then recovered via differentiable fitting within soft rigid clusters defined by the rigidity embeddings. This approach elegantly handles the manifold constraint and the grouping of pixels into rigid bodies. The introduction of the Art-Kubric dataset, featuring dense SE(3) and rigidity labels for articulated objects with physical interactions, is a significant methodological contribution to address the lack of ground truth data for this specific task.
The experiments demonstrate state-of-the-art performance in SE(3) estimation at pixel, part, and object levels on both rigid and articulated benchmarks. The model also achieves state-of-the-art average 3D point tracking accuracy across three datasets. A particularly strong result is the generalization to real-world videos despite training solely on synthetic data, which validates the robustness of the learned representations. The evaluation covers both the novel SE(3) task and the established point tracking task, providing a comprehensive assessment of the model's capabilities.
The paper provides a project page URL. As an arXiv preprint, the code availability is not explicitly confirmed in the provided text, but the detailed description of the method and the release of a large-scale synthetic dataset (Art-Kubric) suggest a high level of reproducibility. The use of standard synthetic data generators (Kubric) for the dataset creation further aids reproducibility.
The primary limitation is the reliance on synthetic data for training, which may introduce a domain gap for certain real-world scenarios not covered by the synthetic generator, although the paper claims strong generalization. The method assumes rigid or articulated motion within clusters; highly deformable objects might not be captured accurately by the SE(3) fitting approach. The computational cost of differentiable fitting within clusters could be a bottleneck for very high-resolution videos or long sequences.
This work has significant implications for robotics, augmented reality, and video understanding. Dense SE(3) motion estimation provides a richer representation of scene dynamics than point tracking alone, enabling better understanding of object interactions, part-level motion, and scene structure. The Art-Kubric dataset will likely become a standard benchmark for motion estimation tasks involving articulated objects. The ability to predict 6-DoF motion from monocular video could enhance applications in robot manipulation, human motion analysis, and 3D scene reconstruction. [One sentence main contribution]. MoSE3 introduces the first feed-forward model for dense SE(3) motion estimation from monocular video, utilizing a novel decomposition into point tracks and rigidity embeddings to handle the non-Euclidean nature of rotations, and contributes a large-scale synthetic dataset, Art-Kubric, to enable training and evaluation. [Comprehensive analysis of the technical contribution, methodology, and significance to the field]. The paper addresses a fundamental gap in motion modeling by moving beyond 3-DoF translation to full 6-DoF rigid transforms. The technical approach is sound, leveraging differentiable optimization to fit SE(3) transforms within learned rigid clusters, which effectively circumvents the difficulties of direct manifold regression. The introduction of Art-Kubric is a substantial contribution to the field, providing the necessary ground truth data for this complex task. The experimental results are strong, showing SOTA performance on both the new SE(3) task and existing point tracking benchmarks, with impressive generalization to real-world data. This work is likely to influence future research in dense motion estimation and provide a valuable tool for applications requiring detailed understanding of scene dynamics.
Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play. On the other hand, language models (LMs) can generate plausible-sounding explanations, but their weak playing strength limits the utility of their explanations. We introduce Queen, a 4B-parameter chess-language model that can explain its moves and plans while playing at the level of a typical Grandmaster. Our novel framework enables domain-specific reasoning through complementary components: an encoder-decoder architecture and an iterative distillation algorithm. This architecture integrates a silent expert chess encoder with an instruction-tuned LM through cross-attention, which we train via a question-answering curriculum to extract chess concepts from the encoder's representations. Building on this domain-adapted model, we iteratively improve its explanations with a natural-language analog of the Bellman update: the model analyzes the positions after its top candidate moves and consolidates them into an explanation of the current position, which is then distilled back into the model. Over seven iterations, our model gains over 900 Elo points (1782 to 2697), substantially surpassing all frontier models on both playing strength and puzzle accuracy, despite containing three orders of magnitude fewer parameters. Furthermore, LM-based evaluations show that our explanations are fluent and approach GPT-5.6-Sol (high) in coherence. The generality of our architecture and training procedure suggests a recipe for applying language models to domains where silent expert encoders are available, like games, robotics, and computer use.
Primary: Princeton University
All Institutions: Princeton University
Queen introduces a 4B-parameter chess-language model that achieves Grandmaster-level play (2697 Elo) and high-quality explanations by integrating a silent expert encoder with an LM via cross-attention and iterative search distillation. This paper makes a significant contribution to the field by demonstrating that hybrid architectures combining specialized encoders with general LMs can outperform frontier LMs in domain-specific reasoning tasks, offering a scalable and efficient alternative to pure LLM approaches for tasks requiring deep domain expertise.
The paper proposes a novel hybrid architecture, "Queen," that integrates a silent expert chess encoder (Leela/BT5) with a general-purpose language model decoder (SmolLM3-3B) via a Flamingo-inspired gated cross-attention bridge. The methodology is rigorous, featuring a two-stage training process: (1) Domain Adaptation using a curated QA curriculum to teach the LM to interpret the encoder's latent representations, and (2) Iterative Search Distillation, a self-improvement loop inspired by Bellman updates where the model analyzes child positions and consolidates explanations, which are then distilled back into the model. This approach effectively bridges the gap between strong silent experts and fluent but weak LMs.
The experimental evaluation is comprehensive and compelling. Queen achieves a 2697 Elo rating, surpassing frontier LMs like GPT-5.6-Sol (2071) and Gemini-3.1-Pro (2201) by a significant margin while using 3 orders of magnitude fewer parameters. The paper introduces a robust evaluation framework covering accuracy (Elo), substantiation (no-mistake rate on puzzles), and coherence (LM-judged). Ablations clearly demonstrate the necessity of both the encoder and the curriculum. The comparison against a "HCE" variant shows that the method is not solely reliant on frontier model distillation for its core strength.
High. The authors provide code, model weights, and detailed appendices including hyperparameters, data construction pipelines, and prompt templates. The use of open-source components (SmolLM3, Leela) and public datasets (Lichess) further enhances reproducibility.
The primary limitation is the low "conceptual coherence" score (2.76/5), indicating that while the model plays well and structures its analysis correctly, it still hallucinates chess motifs or patterns. The method is currently specific to chess; while the authors argue for generality, the reliance on a specific "silent expert" encoder and the specific recursive distillation logic may not transfer trivially to other domains without significant adaptation.
This work offers a general recipe for coupling LMs with domain-specific expert encoders, with potential applications in robotics, computer use, and other games. It demonstrates that small, specialized models can outperform massive general-purpose LMs in specific reasoning tasks when properly grounded in expert knowledge. Queen introduces a 4B-parameter chess-language model that achieves Grandmaster-level play (2697 Elo) and high-quality explanations by integrating a silent expert encoder with an LM via cross-attention and iterative search distillation. This paper makes a significant contribution to the field by demonstrating that hybrid architectures combining specialized encoders with general LMs can outperform frontier LMs in domain-specific reasoning tasks, offering a scalable and efficient alternative to pure LLM approaches for tasks requiring deep domain expertise.
Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tension remains underexplored: the alignment-faithfulness conflict. We show that aligned models systematically deviate from their inputs on unsafe or sensitive content without disclosing the modification, a failure mode we call alignment-induced unfaithfulness (AIU). Unlike capability-driven unfaithfulness, which comes from errors in knowledge or reasoning, this is induced by post-training mechanisms that override adherence to the input. We introduce FaithConflict, a controlled dataset isolating both conflicts, and two complementary taxonomies: behavioral (B1-B8) and chain-of-thought reasoning (C0-C6). Across models, AIU increases with scale and more sharply than capability-driven unfaithfulness, a reverse scaling law; intermediate checkpoints show it is amplified during post-training, with DPO the stage at which the gap both grows most and becomes least visible. Prompting-based mitigation does not resolve it, revealing a capability-alignment-faithfulness trilemma in the design and evaluation of LLMs.
Primary: University of Illinois Urbana-Champaign (UIUC)
All Institutions: University of Illinois Urbana-Champaign (UIUC)
The paper identifies and characterizes "alignment-induced unfaithfulness," a critical failure mode where aligned LLMs silently override source content on sensitive topics, revealing a fundamental capability-alignment-faithfulness trilemma. Through rigorous controlled experiments, a novel dataset, and causal analysis of post-training stages, the authors demonstrate that this unfaithfulness is a systematic, emergent property of modern alignment techniques (particularly DPO) that worsens with model scale, posing significant risks to the reliability of LLMs in source-reporting applications.
The paper introduces a rigorous controlled experimental design to isolate the "alignment-faithfulness" conflict, distinct from capability-faithfulness or capability-alignment tradeoffs. By constructing the FaithConflict dataset with paired confirming/opposing claims in identical templates, the authors effectively control for surface form and domain, allowing the FaithGap metric to directly attribute unfaithfulness to the model's internal conflict with the source content. The dual taxonomy (B1-B8 for outputs, C0-C6 for reasoning) provides a granular lens into *how* models fail, distinguishing between visible refusals and dangerous silent inversions. The methodology is sound, though it relies heavily on an LLM judge (Qwen-2.5 32B) for annotation, which introduces a potential circularity if the judge shares similar alignment biases, though high inter-annotator agreement with humans mitigates this concern.
The experimental scope is impressive, covering 22 checkpoints across 8 model families (including frontier models like Claude Sonnet 4.6 and GPT-4o). The discovery of a "reverse scaling law"—where larger, more aligned models exhibit *greater* unfaithfulness on conflicting sources—is a significant and counter-intuitive finding. The stage-by-stage analysis pinpointing DPO as the primary driver of this behavior is particularly valuable for practitioners. The inclusion of causal interventions (removing safety data from post-training) strengthens the claim that alignment training, not just scale, is the root cause. However, the evaluation is limited to summarization and a few other formats (QA, NLI), and the reliance on a single judge model for all annotations is a notable weakness.
The paper provides high reproducibility standards. Code, data, and project pages are linked. The authors release the FaithConflict dataset and detailed prompts for both task execution and judging. The post-training intervention experiments use the public Tulu 3 pipeline, allowing others to replicate the causal analysis. The only minor gap is the lack of dated snapshot identifiers for frontier models, which is acknowledged by the authors.
The primary limitation is the reliance on an LLM judge for the core metric (FaithGap), which may not perfectly capture human perception of faithfulness. The study is also limited to source-reporting tasks (summarization, extraction) and does not explore whether this unfaithfulness manifests in more complex agentic or multi-turn dialogue settings. The causal attribution to DPO is based on a specific training pipeline (Tulu 3) and may not generalize to all preference optimization algorithms.
This paper has high impact on the LLM safety and evaluation community. It identifies a critical blind spot in current alignment practices: models are becoming better at "safety" but worse at "faithfulness" in a way that is invisible to standard benchmarks. This has immediate implications for RAG systems, clinical note extraction, and legal document processing, where silent modification of source content is a severe failure mode. The "trilemma" framework will likely influence future alignment research to explicitly optimize for faithfulness alongside safety and capability. The paper identifies and characterizes "alignment-induced unfaithfulness," a critical failure mode where aligned LLMs silently override source content on sensitive topics, revealing a fundamental capability-alignment-faithfulness trilemma. Through rigorous controlled experiments, a novel dataset, and causal analysis of post-training stages, the authors demonstrate that this unfaithfulness is a systematic, emergent property of modern alignment techniques (particularly DPO) that worsens with model scale, posing significant risks to the reliability of LLMs in source-reporting applications.
Neural audio codecs compress waveforms into compact discrete tokens that underpin speech language models, real-time communication, and large-scale audio storage. Almost every dominant design, including residual vector quantization, finite scalar quantization, and single-codebook variants, follows the VQ-VAE template by partitioning the encoder latent through a learned codebook or a fixed scalar grid. We ask whether this partition is necessary. We introduce GS-Codec, a neural speech codec whose bottleneck is a parametric signal decomposition rather than a quantizer. We adapt Gaussian splatting from 3D scene reconstruction to one-dimensional latents. An inner optimization loop fits each encoder segment as a weighted sum of 1D Gaussian primitives. The decoder then reconstructs the waveform from the rendered sum. To avoid the cost of this iterative inner loop at inference time, we additionally train a lightweight GS Predictor Net that regresses the primitive parameters in a single forward pass. The encoder and decoder are trained end-to-end through the inner loop, with no quantizer anywhere in the training pipeline: the bottleneck is the decomposition itself, and scalar quantization is applied only post-training to the fitted parameters. Rather than relying on discrete codebook stages for bitrate control, our representation exposes a fine-grained rate-quality tradeoff: a single trained checkpoint supports post-training bitrate control by varying the number of primitives and the per-parameter bit depth, with no retraining required. GS-Codec matches or exceeds well-established open-source codecs such as EnCodec and DAC on speaker similarity (SIM), intelligibility (STOI), and perceptual quality (UTMOS) at comparable bitrates, while achieving comparable semantic performance (WER). Code and audio samples are available at https://ronaluf.github.io/gs-codec/
Primary: Ben-Gurion University of the Negev
All Institutions: Ben-Gurion University of the Negev
GS-Codec introduces a Gaussian Splatting bottleneck for neural audio coding, replacing standard quantizers with a parametric decomposition that enables fine-grained post-training bitrate control. The paper demonstrates competitive performance against state-of-the-art codecs like EnCodec and DAC, offering a novel architectural alternative that leverages differentiable rendering techniques for signal compression, though it faces challenges in encoding latency compared to feed-forward baselines.
The paper proposes a novel architectural shift in neural audio codecs by replacing the standard Residual Vector Quantization (RVQ) or Finite Scalar Quantization (FSQ) bottleneck with a parametric 1D Gaussian Splatting decomposition. The core idea is to fit the encoder's latent representation as a weighted sum of Gaussian primitives via an inner optimization loop during training. To address the computational cost of this iterative fitting at inference, the authors introduce a "GS Predictor Net" that amortizes the optimization into a single forward pass. The method is technically sound, leveraging differentiable rendering concepts from 3D graphics (Gaussian Splatting) and applying them to 1D signal processing. The use of post-training scalar quantization on the fitted parameters rather than learned codebooks is a distinct design choice that enables fine-grained bitrate control without retraining.
The experimental evaluation is rigorous and comprehensive. The authors compare GS-Codec against strong, well-established baselines such as EnCodec, DAC, and WavTokenizer on standard datasets (LibriTTS, LJSpeech, LibriSpeech). The results show that GS-Codec matches or exceeds these baselines on key metrics like UTMOS (perceptual quality), STOI (intelligibility), and SIM (speaker similarity) at comparable bitrates. The inclusion of human listening tests (MUSHRA/MOS) adds significant credibility to the perceptual quality claims. The ablation studies on primitive count and bit depth provide clear insights into the rate-quality tradeoff. The encoding time analysis honestly reports the latency trade-off, showing that while the iterative version is slow, the Predictor Net brings it close to competitive levels, though still slightly slower than feed-forward baselines.
The paper provides high reproducibility. It details the SEANet backbone hyperparameters, the specific Gaussian Splatting configuration (number of primitives, inner loop steps, learning rates), and the training schedule. The code and audio samples are available via the provided URL. The use of standard metrics and open-source baselines facilitates easy comparison. The detailed appendix on quantization ranges and loss weights further supports reproducibility.
The primary limitation is the encoding latency. Even with the Predictor Net, the encoding time is higher than standard feed-forward codecs like EnCodec and DAC, which may be a bottleneck for real-time applications. The paper also notes that performance at very low bitrates (<3 kbps) is not as strong as specialized low-rate codecs. The method is currently validated primarily on English speech, and generalization to other languages or non-speech audio (music, sound effects) is not extensively explored.
This work opens a new direction in neural audio compression by demonstrating that parametric signal decompositions can serve as effective bottlenecks, challenging the dominance of codebook-based quantization. The ability to control bitrate post-training by varying the number of primitives and bit depth is a practical advantage for deployment. The cross-domain insight from 3D Gaussian Splatting to 1D audio signals is interesting and may inspire further research into using geometric or parametric priors in other signal processing tasks. GS-Codec introduces a Gaussian Splatting bottleneck for neural audio coding, replacing standard quantizers with a parametric decomposition that enables fine-grained post-training bitrate control. The paper demonstrates competitive performance against state-of-the-art codecs like EnCodec and DAC, offering a novel architectural alternative that leverages differentiable rendering techniques for signal compression, though it faces challenges in encoding latency compared to feed-forward baselines.
Most humanoid loco-manipulation controllers require human motion data to learn whole-body coordination and posture, leaving policies reliant on external sources to provide this data. We present OCLO (Online-posture Compliant LOco-manipulation), a humanoid loco-manipulation system trained without human motion data and commanded only through two end-effector targets. Because these targets do not uniquely determine whole-body posture, OCLO generates pelvis height and torso orientation online using an analytic reachability prior, further refined through policy-in-the-loop sampling with a task-agnostic cost. OCLO also learns whole-body compliance by displacing end-effector references according to measured forces through a spring-damper model, encouraging the legs, waist, and pelvis to yield to external loads. In simulation, using the reachability prior leads to a 77.8% success rate in acquiring the commanded reference, a vast improvement over the 37.8% success rate accomplished without the prior. Further, refinement reduces end-effector orientation error across all evaluated tasks. The same posture module improves a pretrained SONIC controller on four of five tasks. Without compliance training, policies tend to lose balance under disturbances rather than sacrifice tracking. On a Unitree G1, OCLO maintains balance under end-effector disturbances that cause its ablations to fail and performs seven loco-manipulation tasks, including crouched walking and picking up a box from a low surface. Project website: https://oclo-humanoid.github.io/
Primary: University of California San Diego
All Institutions: University of California San Diego, Yonsei University
OCLO presents a dataset-free humanoid loco-manipulation system that generates whole-body posture online from end-effector targets using an analytic reachability prior and policy-in-the-loop refinement, enabling compliant interaction and robust balance without human motion data. The paper makes a significant contribution to the field by addressing the under-constrained nature of two-point interfaces in humanoid control, offering a computationally efficient and effective solution that improves upon existing methods like SONIC. The rigorous experimental evaluation, including paired simulation episodes and hardware demonstrations, supports the claims of improved success rates and robustness. The separation of posture generation from physical support is a novel architectural insight that enhances modularity and generalizability. While the method relies on specific assumptions about the robot's dynamics and the availability of force estimates, it represents a solid step forward in autonomous humanoid control. The work is highly relevant to researchers in robotics and reinforcement learning, particularly those interested in whole-body control and human-robot interaction.
The paper introduces OCLO, a system for humanoid loco-manipulation that decouples posture generation from physical support. The core methodological contribution is the "analytic reachability prior," a closed-form function that maps two end-effector (EE) targets to pelvis height and torso orientation. This is a clever, computationally cheap heuristic that addresses the under-constrained nature of the two-point interface. The method is further refined by a policy-in-the-loop sampling approach (CEM) that evaluates candidate postures using a frozen controller rollout, optimizing a task-agnostic cost. The compliance mechanism, which uses a spring-damper model to displace EE references based on measured forces, is standard in impedance control but effectively integrated here with a lower-body policy trained to accommodate these displacements. The separation of the "intent" (EE targets) from the "support" (posture/balance) is a sound architectural choice that allows the system to generalize to different manipulation tasks without retraining the balance policy.
The evaluation is rigorous and well-structured. The authors use a paired-episode simulation protocol to ensure fair comparison between methods, which is a strong experimental design choice. The ablation studies clearly isolate the contribution of the analytic prior versus the CEM refinement, showing that the prior is crucial for success rate while CEM improves orientation accuracy. The hardware experiments on the Unitree G1 are convincing, demonstrating robustness to disturbances that cause ablated versions to fail. The comparison against SONIC, a state-of-the-art human-motion-pretrained controller, is particularly valuable, showing that the posture module can transfer to improve existing systems. The seven qualitative hardware tasks provide good evidence of the system's versatility.
The paper provides sufficient detail on the control architecture, reward functions, and training curriculum. The use of standard libraries (MuJoCo, PPO) and the open-source nature of the Unitree G1 platform enhance reproducibility. The specific constants for the analytic prior are mentioned to be fitted from a sweep, which is a minor detail that might require some effort to replicate exactly, but the overall framework is clear. The code is likely available given the project website, though not explicitly linked in the text provided.
The system relies on a simulated model for the policy-in-the-loop refinement, which may not perfectly match real-world dynamics. The posture interface is limited to pelvis height and torso orientation, excluding more complex adjustments like foot placement or gait timing. The arm redundancy is handled by a separate IK solver, which may struggle near singularities. The hardware evaluation uses a single operator, which limits the generalizability of the teleoperation results.
This work contributes to the field of humanoid robotics by providing a scalable method for generating whole-body postures without relying on expensive human motion capture data. The separation of posture and support could be applicable to other robotic systems requiring coordinated whole-body control. The compliance mechanism is relevant for safe human-robot interaction. OCLO presents a dataset-free humanoid loco-manipulation system that generates whole-body posture online from end-effector targets using an analytic reachability prior and policy-in-the-loop refinement, enabling compliant interaction and robust balance without human motion data. The paper makes a significant contribution to the field by addressing the under-constrained nature of two-point interfaces in humanoid control, offering a computationally efficient and effective solution that improves upon existing methods like SONIC. The rigorous experimental evaluation, including paired simulation episodes and hardware demonstrations, supports the claims of improved success rates and robustness. The separation of posture generation from physical support is a novel architectural insight that enhances modularity and generalizability. While the method relies on specific assumptions about the robot's dynamics and the availability of force estimates, it represents a solid step forward in autonomous humanoid control. The work is highly relevant to researchers in robotics and reinforcement learning, particularly those interested in whole-body control and human-robot interaction.
Goal-conditioned imitation learning (GCIL) with flow matching is a promising framework that can represent multimodal behaviors while adapting to diverse, user-specified goals, yet often fails when goals lie outside the demonstration support. To extrapolate to such unseen goals without collapsing multimodality - a problem we call distributional extrapolation - we introduce Bilinear Flow Policy (BFP), a generative visuomotor policy that combines transductive retrieval with a bilinear conditional flow. Given an unseen observation-goal pair, BFP retrieves an "anchor" training example and transductively reformulates the unseen pair as this familiar anchor plus a residual term. For this decomposition to guide action prediction, the residual must compactly encode how the current observation-goal pair differs from the anchor, and the anchor must be chosen so that this difference is predictive of the corresponding action distribution. BFP achieves this with pretrained visual features and a novel learned anchor-selection algorithm. The novel bilinear flow then models how the anchor and the residual jointly determine the multimodal action distribution. We prove that, for bilinear flow under suitable assumptions, action distribution error at unseen goals is bounded by the in-distribution flow-matching error up to problem-dependent factors. Across five manipulation tasks in simulation, BFP achieves 2.63x the out-of distribution success rate of a GCIL policy and 1.36x that of the strongest extrapolation-targeted baseline. On two real-world tasks, BFP improves over GCIL by 32%. Finally, our theory yields practical, pre deployment diagnostics for predicting which trained policies will extrapolate well and to which unseen goal.
Primary: Unknown (Likely Toyota Research Institute based on funding, but affiliations not explicitly listed in provided text)
All Institutions: Unknown
Bilinear Flow Policy (BFP) introduces a transductive retrieval and bilinear flow matching framework to enable goal-conditioned visuomotor policies to extrapolate to unseen goals while preserving multimodality. The paper presents a novel method with theoretical guarantees and strong empirical results in simulation and real-world manipulation, addressing a significant limitation in current imitation learning approaches.
The paper introduces Bilinear Flow Policy (BFP), a novel generative policy architecture for goal-conditioned imitation learning. The core innovation is the decomposition of an unseen observation-goal pair into a retrieved "anchor" training example and a residual term, modeled via a bilinear conditional flow. This approach addresses the "distributional extrapolation" problem, where standard flow matching policies fail on goals outside the demonstration support. The method leverages pretrained visual features for anchor selection and provides a theoretical bound on action distribution error for unseen goals, linking it to in-distribution flow-matching error. The use of transductive retrieval combined with bilinear flow is a creative and technically sound approach to maintaining multimodality during extrapolation.
The authors evaluate BFP on five simulation manipulation tasks and two real-world tasks. The results show a 2.63x improvement in out-of-distribution success rate compared to a standard GCIL policy and 1.36x over the strongest extrapolation-targeted baseline. Real-world improvements of 32% over GCIL are reported. The experiments are relevant to the robotics community, though the number of tasks is modest. The comparison to "strongest extrapolation-targeted baseline" is crucial and appears to be handled, though specific baseline names are not detailed in the abstract. The theoretical diagnostics for predicting extrapolation capability add significant value to the experimental section.
The paper includes a reproducibility statement, releases code and configurations via a project website, and provides pseudocode in the appendix. The use of AI for literature review and editing is disclosed, which is transparent. The specific details of the "learned anchor-selection algorithm" and the bilinear flow implementation are critical for reproduction and are presumably in the released code.
The paper is a submission to ICLR 2027, so it is not yet peer-reviewed. The institution is not explicitly stated in the provided text, though Toyota Research Institute is mentioned as a funder. The evaluation is limited to manipulation tasks; generalization to navigation or other robot embodiments is not tested. The reliance on pretrained visual features may limit performance in domains where such features are not well-aligned with the task.
This work contributes to the broader field of imitation learning by providing a principled way to handle out-of-distribution goals, a common challenge in deploying robot policies. The theoretical insights into flow matching extrapolation may influence future generative policy designs. The practical diagnostics for pre-deployment testing are a valuable contribution for robotics practitioners. Bilinear Flow Policy (BFP) introduces a transductive retrieval and bilinear flow matching framework to enable goal-conditioned visuomotor policies to extrapolate to unseen goals while preserving multimodality. The paper presents a novel method with theoretical guarantees and strong empirical results in simulation and real-world manipulation, addressing a significant limitation in current imitation learning approaches.
Most vision-language-action models represent future motion as fixed-rate action chunks, tying temporal resolution and prediction horizon to a fixed output budget. This pointwise representation wastes capacity on highly correlated neighboring actions, leaves temporal continuity and smoothness to be learned implicitly, and forces a tradeoff between long-horizon coverage and the local precision required for contact-rich manipulation. To address these limitations, we introduce Vela, a vision-language-action foundation model that represents future robot behavior as continuous trajectories. Vela combines a compact spline-based action representation with motion-dependent temporal support and a shared action interface for heterogeneous embodiments, allowing a fixed output budget to adapt its temporal resolution across motions. We pretrain Vela on large-scale multi-embodiment robot data and evaluate it on LIBERO-X, EBench, and two real-world long-horizon tasks, egg-cake cooking and potato shredding, obtaining promising results across simulation and physical manipulation. These results highlight the potential of continuous action representations as a foundation for future embodied foundation models. Project page and more results: https://clementine24.github.io/Vela/ .
Primary: Shanghai Innovation Institute
All Institutions: Shanghai Innovation Institute
Vela introduces a trajectory-space VLA model with adaptive horizon selection, demonstrating that continuous spline-based action representations outperform fixed-rate chunks in both simulation and real-world manipulation. The paper provides a rigorous evaluation and ablation studies that validate the benefits of native trajectory learning and motion-dependent temporal support, offering a promising direction for scaling embodied foundation models.
The paper proposes Vela, a Vision-Language-Action (VLA) model that replaces fixed-rate action chunks with continuous trajectory representations using cubic B-splines. The core innovation is the "motion-dependent horizon adaptation," which dynamically selects the temporal span of the spline based on the complexity of the demonstrated motion, using a noise-aware criterion to distinguish between necessary motion detail and teleoperation noise. The method integrates this into a flow-matching framework, introducing a "Decoded-Trajectory Flow Matching" (DT-FM) objective that weights control point errors based on their impact on the decoded trajectory. The approach is technically sound, addressing the trade-off between long-horizon coverage and local precision in a principled way. The use of splines is not new in robotics, but scaling this to foundation-model-level pretraining with adaptive horizons is a significant methodological step.
The evaluation is comprehensive, covering simulation (LIBERO-X, EBench) and real-world tasks (egg-cake cooking, potato shredding). The results show consistent improvements over the baseline ($\pi_{0.5}$) and other state-of-the-art methods. The ablation studies are particularly strong, isolating the contributions of trajectory-space learning, horizon adaptation, and the specific loss function. The real-world experiments on long-horizon, contact-rich tasks provide strong evidence for the practical utility of the method. The comparison with post-hoc spline fitting is a crucial control that validates the need for native trajectory-space learning.
The paper provides detailed implementation specifics, including hyperparameters, dataset sizes, and training protocols. The use of public datasets (LIBERO-X, EBench) and clear descriptions of the real-world setup enhances reproducibility. The project page is provided, which likely contains code and additional results. The detailed appendix on target construction and objective derivation further supports reproducibility.
The model is based on a 3B parameter backbone, and the paper acknowledges that scaling to larger models is a future direction. The fixed number of control points (N=12) might limit expressivity for extremely complex motions, though the adaptive horizon helps mitigate this. The method relies on accurate noise scale estimation, which could be sensitive to data quality. The real-world evaluation, while impressive, is limited to two specific tasks.
This work has significant potential to influence the design of future embodied foundation models. By demonstrating that continuous trajectory representations can be effectively learned at scale, it challenges the dominance of discrete action chunks in VLA architectures. The adaptive horizon mechanism offers a generalizable solution to the temporal resolution trade-off, which is a common problem in robot learning. The findings could lead to more efficient and robust policies for diverse robotic applications. Vela introduces a trajectory-space VLA model with adaptive horizon selection, demonstrating that continuous spline-based action representations outperform fixed-rate chunks in both simulation and real-world manipulation. The paper provides a rigorous evaluation and ablation studies that validate the benefits of native trajectory learning and motion-dependent temporal support, offering a promising direction for scaling embodied foundation models.
Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it. The resulting images are processed foveally by allocating more visual tokens to the image centers, focusing computation on task-relevant features. Such Active Visual Fixation (AVF) requires carefully coordinated gaze during task execution, which we accomplish hierarchically by first training a low-level gaze servoing policy conditioned on a goal object, then training a target selector which emits fixation goals based on task progress. Both modules are trained with RL on real-world data: the first is trained with a dense geometric reward and the second co-trains with the BC gripper policy which allows it to discover fixation sequences that can resemble a human's fixation sequence while performing the task. EyeRobot 2.0 further takes advantage of fixation by canonicalizing gripper information into a fixation-relative SE(3) frame, which compacts the size of the action distribution to learn. We collect teleoperation data for 7 real-world and 6 simulated tasks, and conduct over 1000 physical and 1800 simulated robot trials comparing EyeRobot 2.0 against passive stereo and ego + wrist camera policies trained on the same data. Removing wrist cameras is costly for standard policies: with only passive stereo, real-world success drops from 52% to 27%. EyeRobot 2.0 closes this gap with only stereo, outperforming passive stereo by 40% in real and 20% in sim. It matches ego + wrist policies when their wrist views are clear (69% vs. 64%), and more than doubles their success when grasped objects occlude the wrist cameras (48% vs. 22%)
Primary: Stanford University
All Institutions: Stanford University, UC Berkeley
EyeRobot 2.0 introduces a biologically inspired active gaze framework that enables precise bimanual manipulation using only a fixed stereo camera, outperforming standard wrist-camera baselines in occlusion-heavy tasks. The paper demonstrates that physically attending to 3D fixation points and canonicalizing actions in a fixation-relative frame significantly enhances policy learning, offering a robust alternative to the ubiquitous wrist-camera setup in robotic manipulation.
The paper proposes a hierarchical framework for active visual fixation in bimanual manipulation. The core innovation is the decoupling of low-level gaze servoing (trained with RL using geometric rewards) from high-level target selection (trained with RL using a BC-RL loop that optimizes for downstream gripper policy accuracy). The use of a fixation-relative SE(3) frame for action canonicalization is a strong technical contribution that reduces the complexity of the action space. The foveated processing of stereo images is well-motivated, though the specific implementation details of the "foveated transformer-decoder" are somewhat sparse in the main text, relying on the appendix. The method effectively addresses the occlusion issues inherent in wrist-mounted cameras by physically moving the viewpoint.
The experimental setup is rigorous, involving 7 real-world and 6 simulated tasks with over 1000 physical trials. The comparison against passive stereo and ego+wrist baselines is fair, as all policies are trained on the same data. The results are compelling: EyeRobot 2.0 significantly outperforms passive stereo and matches or exceeds ego+wrist performance, particularly in occlusion scenarios where wrist cameras fail. The ablation studies convincingly demonstrate the contribution of foveation, fixation-centric actions, and stereo depth. The inclusion of simulation experiments provides a controlled environment for analysis, although the primary claim is validated on real hardware.
The paper states that simulation code will be made public, which is a positive step. However, the reliance on specific hardware (I2RT YAM manipulator, custom leader arm) and specific software stacks (MuJoCo, GELLO, TRLC) may limit immediate reproducibility for labs without similar setups. The detailed description of the RL training loops and reward functions provides a clear path for implementation, but the lack of a public code release at the time of review (arXiv preprint) is a minor drawback.
The system is currently task-specific, requiring separate training for each task. The paper acknowledges that extending this to multi-task learning would require more complex prompt generation (e.g., VLMs). The framework assumes a fixed head position and only controls eye movements, which limits its applicability to mobile manipulation where head/neck movement is crucial. The computational cost of running two RL policies and a BC policy in real-time is not deeply analyzed, though the 30Hz gaze rate suggests it is feasible.
This work has significant implications for the design of robotic manipulators, potentially eliminating the need for wrist cameras and allowing for sleeker, more robust gripper designs. It also opens a pathway for leveraging egocentric human data for robot learning by mimicking human fixation patterns. The approach could be extended to other domains requiring precise visual attention, such as surgical robotics or inspection tasks. EyeRobot 2.0 introduces a biologically inspired active gaze framework that enables precise bimanual manipulation using only a fixed stereo camera, outperforming standard wrist-camera baselines in occlusion-heavy tasks. The paper demonstrates that physically attending to 3D fixation points and canonicalizing actions in a fixation-relative frame significantly enhances policy learning, offering a robust alternative to the ubiquitous wrist-camera setup in robotic manipulation.
Equipping artificial agents with spatial intelligence requires a comprehensive generative prior over the dynamic 3D world. We propose World Motion Models (WMMs) that capture "what was, is, and will be where across time" via sparse SE(3) pose trajectories. WMMs are built on the observation that elements of dynamic scenes can be well approximated by a set of rigid SE(3) trajectories, a minimal yet expressive primitive for 4D modeling. This representation unifies articulated objects, human bodies, hand-object interactions, piecewise-rigid scene dynamics, camera motion, and even robot states and actions into a single shared space. Given this representation, we cast the joint distribution of these entities as a flexible sequence modeling problem, utilizing flow-matching with per-token noise levels. Coupled with a context token mechanism for non-sequential conditioning, this formulation supports any-to-any marginal conditioning across an arbitrary number of entities and time steps. Tasks such as future prediction, motion infilling, model-predictive control, inverse kinematics, cross-embodiment retargeting, and policy learning all reduce to the application of different masks over the same network. Experiments on 6 diverse applications of 3D vision and robotics demonstrate the versatility and flexibility of WMMs with strong performance.
Primary: UC Berkeley
All Institutions: UC Berkeley, Meta AI
World Motion Models introduces a unified flow-matching framework for SE(3) trajectory modeling that enables flexible any-to-any conditioning across diverse robotics and vision tasks. The paper demonstrates that a single generative model can handle prediction, control, and retargeting through simple masking, offering a powerful and versatile primitive for spatial intelligence in artificial agents.
The paper proposes World Motion Models (WMMs), a unified framework for modeling 4D dynamics using sparse SE(3) pose trajectories. The core methodological contribution is the recasting of joint distribution modeling over multiple entities (humans, objects, cameras, robots) as a flexible sequence modeling problem using flow-matching. By employing per-token noise levels and a context token mechanism, the authors enable "any-to-any" marginal conditioning. This allows a single network to handle diverse tasks—such as prediction, infilling, and control—simply by applying different masks to the input sequence. The choice of SE(3) trajectories as a primitive is elegant, providing a minimal yet expressive representation that unifies articulated motion, rigid body dynamics, and camera motion into a shared latent space.
The paper reports experiments on six diverse applications spanning 3D vision and robotics, including future prediction, motion infilling, model-predictive control (MPC), inverse kinematics (IK), cross-embodiment retargeting, and policy learning. The breadth of tasks demonstrates the versatility of the unified approach. While the abstract claims "strong performance," the lack of specific quantitative metrics in the provided text prevents a full assessment of superiority over specialized baselines. However, the ability to switch tasks via masking without retraining is a significant empirical advantage over task-specific architectures.
The project page is provided, which likely contains code and datasets. The use of standard flow-matching techniques and SE(3) representations suggests that the core components are reproducible, provided the specific masking strategies and context token implementations are detailed in the full paper (which is not fully visible here, but implied by the "Spotlight" status and project page).
The reliance on SE(3) trajectories assumes that scene elements can be well-approximated by rigid motions. This may limit applicability to highly deformable objects or non-rigid interactions that cannot be decomposed into rigid parts. Additionally, the computational cost of flow-matching on long sequences of high-dimensional SE(3) poses could be a bottleneck for real-time applications, although the paper mentions MPC, suggesting some efficiency gains.
This work has the potential to significantly impact the field of embodied AI and 4D scene understanding. By unifying various robotics and vision tasks under a single generative framework, it simplifies the development of general-purpose agents. The "any-to-any" conditioning capability is particularly valuable for data-efficient learning and sim-to-real transfer, where conditioning on partial observations or specific actions is common. World Motion Models introduces a unified flow-matching framework for SE(3) trajectory modeling that enables flexible any-to-any conditioning across diverse robotics and vision tasks. The paper demonstrates that a single generative model can handle prediction, control, and retargeting through simple masking, offering a powerful and versatile primitive for spatial intelligence in artificial agents.
Large language model (LLM) serving scales its replica count with the request load, yet GPU memory still stands idle inside the replicas. Adding a replica takes minutes, while the memory that a replica needs changes within seconds. Even instant autoscaling could not return this idle memory, because the smallest unit that it can remove is a whole replica. Colocating parameter-efficient fine-tuning (PEFT) with inference can use this memory, but inference must be able to reclaim it within seconds, before requests that wait for memory exceed their latency service-level objective (SLO). Existing colocation systems either keep the tuning memory resident or let inference reclaim it at the coarse granularity of a whole training sample. Each such reclamation also discards the running tuning step. To address these limitations, we present MOLT, a fine-grained memory sharing system that lets inference reclaim the memory of individual activations that a running tuning step has saved for its backward pass. The step continues, and its backward pass recomputes those activations. Inference reclaims only memory that no in-flight GPU work can still access, even under CPU--GPU asynchrony and tensor parallelism. On four model deployments (24B--70B) across H100 SXM and B200 GPUs under trace-driven workloads, MOLT keeps inference SLO attainment at or above 99.7% and completes 1.9--3.3x the tuning work of discard-based memory sharing.
Primary: Seoul National University
All Institutions: Seoul National University
MOLT introduces a fine-grained GPU memory sharing mechanism that allows inference to reclaim individual activation tensors from concurrent PEFT training via recomputation, enabling 1.9-3.3x more tuning work while maintaining >99.7% inference SLO attainment on H100/B200 clusters.
The paper proposes MOLT, a system for fine-grained GPU memory sharing between LLM inference and Parameter-Efficient Fine-Tuning (PEFT). The core innovation is the ability to reclaim individual activation tensors saved for the backward pass during inference spikes, allowing the training step to continue via activation recomputation rather than discarding the entire step. This addresses a specific latency-SLO constraint in co-located serving. The technical design handles CPU-GPU asynchrony and tensor parallelism, which are non-trivial engineering challenges. The approach is a clever systems-level optimization of the standard gradient checkpointing/recomputation trade-off, tailored for dynamic serving environments.
The evaluation covers four model deployments (24B-70B) on H100 SXM and B200 GPUs using trace-driven workloads. The results show high SLO attainment (>99.7%) and a 1.9-3.3x increase in completed tuning work compared to discard-based baselines. The choice of hardware (including B200) is current and relevant. However, the evaluation is limited to trace-driven simulations rather than live production traffic, and the comparison is primarily against "discard-based" memory sharing, which may not represent the most advanced state-of-the-art co-location systems if they exist.
The paper appears to be from a top-tier systems group (SNU, Hojoon Kim's lab is well-regarded in this area). However, no code repository or demo URL is provided in the text. Without access to the specific memory management hooks and the trace generation methodology, full reproduction is difficult. The reliance on specific GPU architectures (H100/B200) also limits immediate reproducibility for the broader community.
The primary limitation is the scope: it focuses on PEFT. Full fine-tuning or other training paradigms may not benefit equally. The performance gains depend heavily on the volatility of the inference load; in steady-state high-load scenarios, the benefit might diminish. Additionally, the overhead of activation recomputation, while mitigated, still consumes compute cycles that could otherwise be used for inference or training, a trade-off that is not deeply quantified in the abstract.
This work is significant for cloud providers and enterprises running LLMs who wish to utilize idle GPU memory for continuous adaptation (fine-tuning) without compromising user-facing latency. It pushes the boundary of what is possible in multi-tenant GPU scheduling. It does not change the fundamental algorithms of LLMs but improves the infrastructure efficiency and flexibility of LLM deployment. MOLT introduces a fine-grained GPU memory sharing mechanism that allows inference to reclaim individual activation tensors from concurrent PEFT training via recomputation, enabling 1.9-3.3x more tuning work while maintaining >99.7% inference SLO attainment on H100/B200 clusters.
New AI accelerators arrive before the kernels that make them fast, because peak kernel performance requires architecture-specific expertise in operand pipelines and data-movement techniques. Coding agents can now write, compile, and tune kernels on their own, so they could greatly accelerate kernel development and optimization. What agents produce depends on the interfaces they are given. These interfaces may expose a machine's raw capabilities or encode the known-good methods for exploiting its hardware features efficiently. We test the performance impact of these two interfaces on Intel GPUs through controlled experiments with three coding models on three kernels. With raw SYCL and execution feedback, Opus 4.8, the strongest of the three models tested, writes a single-GPU GEMM kernel reaching only 57.5% of Intel's tuned oneDNN library. We present SyclKittens, a hardware-aware tile programming model for Intel GPUs that encodes these known-good methods, so its operations place operands in matrix-engine layouts, prefetch through the L1 cache, and communicate over the fabric that links GPU stacks into a node. With SyclKittens under the same agent, task, and feedback budget, the agent reaches 82.1%, showing that encoded methods turn hardware capabilities into performance. Workload-specific policies stay programmable, and engineers and agents jointly refine these schedules in SyclKittens to build a kernel suite that reaches ~96% of oneDNN in geometric mean across GEMM shapes. The suite runs Llama-3.1-8B inference 1.59x faster than torch.compile on one GPU and up to 2.91x faster than a matched multi-GPU decode path on Intel's oneCCL. SyclKittens is open source and available at https://github.com/intel/SyclKittens.
Primary: Intel Corporation
All Institutions: Massachusetts Institute of Technology, Intel Corporation, Stanford University, California Institute of Technology
SyclKittens introduces a hardware-aware tile programming model that enables coding agents to generate high-performance GPU kernels by encoding known-good hardware methods. The paper demonstrates that this approach significantly outperforms raw SYCL interfaces in agent-generated kernel performance, achieving near-library-level efficiency for critical AI workloads on Intel GPUs, thereby offering a scalable path for automated kernel optimization.
The paper introduces SyclKittens, a tile-based programming model for Intel GPUs that abstracts low-level hardware details (operand pipelines, data movement, matrix engine layouts) into high-level, hardware-aware operations. The core methodological contribution is the demonstration that providing coding agents with a structured, domain-specific interface (SyclKittens) rather than raw SYCL primitives significantly improves the performance of generated kernels. The authors employ a controlled experimental design comparing raw SYCL with execution feedback against SyclKittens under identical agent, task, and feedback budgets. This approach effectively isolates the impact of the programming interface on agent performance, providing a rigorous evaluation of how "known-good" hardware methods encoded in a DSL can guide LLMs to produce efficient code.
The experiments are conducted on Intel Max GPUs, focusing on critical AI workloads such as GEMM, attention, and normalization. The results show a substantial improvement: an agent using raw SYCL reaches only 57.5% of the performance of Intel's tuned oneDNN library, whereas the same agent using SyclKittens reaches 82.1%. Furthermore, a co-designed kernel suite using SyclKittens achieves ~96% of oneDNN performance in geometric mean across GEMM shapes. End-to-end inference benchmarks for Llama-3.1-8B demonstrate 1.59x speedup over torch.compile on a single GPU and up to 2.91x speedup over a matched multi-GPU decode path using Intel's oneCCL. The evaluation is strong in its direct comparison of interfaces and its end-to-end application metrics.
The paper is highly reproducible given that SyclKittens is open-sourced on GitHub. The authors provide detailed appendices separating the controlled-agent studies from the co-designed suite, including specific measurement protocols, warmup procedures, and aggregation methods. The use of specific coding models (Opus 4.8, etc.) and defined feedback budgets allows for precise replication of the agent experiments. However, the specific prompts and agent configurations may require careful extraction from the appendices to fully replicate the agent's behavior.
The evaluation is limited to Intel GPUs, which may not generalize directly to other hardware architectures (e.g., NVIDIA, AMD) without significant adaptation. The performance gains are relative to Intel's own oneDNN library, so the absolute competitive standing against other state-of-the-art libraries on different hardware is not assessed. The reliance on specific coding models means the results may vary with future model improvements or different agent architectures. Additionally, the paper focuses on a limited set of kernels (GEMM, attention, norms), and the applicability to more complex or irregular workloads is not fully explored.
This work has significant implications for the intersection of AI and systems programming. By demonstrating that high-level, hardware-aware abstractions can enable coding agents to write near-optimal GPU kernels, it suggests a new paradigm for kernel development that reduces the need for deep, architecture-specific expertise. This could accelerate the adoption of new AI accelerators by lowering the barrier to entry for kernel optimization. The open-source release of SyclKittens provides a valuable tool for the community to explore similar approaches on Intel hardware and potentially adapt the concepts to other platforms. SyclKittens introduces a hardware-aware tile programming model that enables coding agents to generate high-performance GPU kernels by encoding known-good hardware methods. The paper demonstrates that this approach significantly outperforms raw SYCL interfaces in agent-generated kernel performance, achieving near-library-level efficiency for critical AI workloads on Intel GPUs, thereby offering a scalable path for automated kernel optimization.