Last 7 Days (September 22 – September 28, 2026)
Attributing language model outputs to their internal computations is an open problem in interpretability. Existing methods, which use causal interventions, gradients, or learnable masks, either are infeasibly expensive or struggle to identify actual causally-important internal computations. We propose framing attribution as the problem of identifying nested subsets of internal components which minimise a downstream loss. To learn this task, we introduce Matryoshka Attribution (MAttr), a mask learning method that parametrises the mask with a simple differentiable sigmoid top-$k$ operator. We supervise training over all sparsities simultaneously by randomising $k$ over training, resulting in a learned ordering of components by attribution score. MAttr achieves number 1 on the official leaderboard of the Mechanistic Interpretability Benchmark (Mueller et al., 2025); our method identifies sparse and task-transferrable circuits across varying circuit bases. As a practical application, we show that MAttr can be trained with reinforcement learning to identify weight changes responsible for downstream behaviours in LLM finetuning. We train MAttr on refusal judge scores and find that restoring $1\%$ of Llama 3.1 8B Instruct's weights to their base model state is sufficient to remove refusals while maintaining capabilities. We view MAttr as a successful formulation of interpretability into a learnable objective that we can tackle with gradient descent, and encourage future work along these lines.
Primary: Stanford University
All Institutions: Stanford University
Matryoshka Attribution (MAttr) is a novel mask-learning method that uses a differentiable sigmoid top-$k$ operator to learn a single, nested ordering of internal components for attribution, achieving state-of-the-art performance on the Mechanistic Interpretability Benchmark and enabling the identification of specific weight changes responsible for LLM behaviors like refusal through reinforcement learning.
The paper introduces Matryoshka Attribution (MAttr), a mask-learning method that frames attribution as identifying nested subsets of internal components minimizing a downstream loss. The core innovation is the use of a differentiable sigmoid top-$k$ operator to parametrize the mask, allowing the model to learn a single ordering of components that is valid across all sparsity levels (the "Matryoshka" property). This eliminates the need for separate mask training at each sparsity level and avoids the instability and hyperparameter sensitivity associated with hard concrete distributions or straight-through estimators. The method is grounded in causal abstraction, defining "soft interchange interventions" that interpolate between base and source states. A significant theoretical contribution is the demonstration that MAttr recovers Integrated Gradients in the limit of one training step, providing a rigorous bridge between gradient-based and mask-learning attribution methods. The extension to parameter-level attribution via Reinforcement Learning (GRPO) is a novel application, allowing the method to identify weight changes responsible for specific behaviors (e.g., refusal) without requiring differentiable rewards.
The experimental evaluation is extensive and rigorous. MAttr achieves state-of-the-art performance on the Mechanistic Interpretability Benchmark (MIB), significantly outperforming baselines like Integrated Gradients, Interchange Interventions, and other mask-learning methods (DBM, Node Pruning). The authors introduce "MIB+" to test finer-grained bases (MLP neurons, SAE features) and new tasks, showing consistent improvements. The paper also introduces a "compactness" metric to evaluate circuit sparsity, showing MAttr excels at both faithfulness and compactness. The parameter-level attribution experiments on Llama 3.1 8B Instruct are particularly compelling, demonstrating that restoring only 1% of weights to the base model state removes refusals while maintaining capabilities, a result with significant practical implications for model safety and auditing.
The paper provides a public GitHub repository with code. Hyperparameters for all baselines and MAttr are detailed in the appendices. The use of standard benchmarks (MIB) and open-source models (Llama, GPT-2, etc.) enhances reproducibility. The specific details of the RL training setup (GRPO) and reward functions are described, though the complexity of the RL setup may pose challenges for exact replication without the provided code.
The method relies on gradient descent, which may be computationally expensive for very large models or fine-grained bases, although it is parallelizable. The "Matryoshka" property assumes a nested structure of importance, which may not always hold for all types of internal computations. The RL-based parameter attribution is a new paradigm and may require careful tuning of the reward function to avoid reward hacking. The paper focuses primarily on language models; generalization to other modalities (e.g., vision) is mentioned in appendices but not the primary focus.
This paper has high potential impact on the field of interpretability. By providing a robust, learnable, and causally-grounded method for attribution, it enables more reliable identification of task-relevant circuits and weight changes. The application to removing refusals via weight restoration offers a new tool for auditing and potentially mitigating safety issues in LLMs. The unification of gradient-based and mask-learning perspectives may stimulate further theoretical work in attribution. The method's ability to handle non-differentiable objectives via RL opens up new avenues for interpreting complex, emergent behaviors. Matryoshka Attribution (MAttr) is a novel mask-learning method that uses a differentiable sigmoid top-$k$ operator to learn a single, nested ordering of internal components for attribution, achieving state-of-the-art performance on the Mechanistic Interpretability Benchmark and enabling the identification of specific weight changes responsible for LLM behaviors like refusal through reinforcement learning.
X-ray is medicine's most widely used imaging modality, yet remains among its least quantitative. Unlike volumetric modalities like CT or MRI, X-ray collapses 3D anatomy into a 2D projection, causing structures to overlap and anatomical boundaries to be ambiguous, even to experts. As a result, labeling X-ray databases for training general-purpose segmentation systems is impractical, leaving morphometric and functional X-ray analysis confined to narrow anatomical regions and applications. To this end, we present FleXray, a generalist model for anatomical segmentation across the entire body in clinical X-rays. Instead of curating large, manually annotated X-ray datasets, we build a scalable, physics-based generative X-ray data engine. Using existing 3D whole-body CT segmentation datasets and generative image-editing models, we simulate fully-annotated 2D X-rays with diverse appearances, physiological properties, and imaging geometries. Trained on these simulations, FleXray accurately segments 60 anatomical structures across unseen research datasets and in-the-wild X-rays. We further show that FleXray makes X-rays directly amenable to quantitative analysis, enabling automated measurements for disease grading, robust navigation during X-ray-guided interventions, and data-efficient learning of pathological targets. We release the model, code, a full-body X-ray segmentation dataset, and a local, easy-to-use browser-based tool at https://flexray.csail.mit.edu .
Primary: Massachusetts Institute of Technology (MIT)
All Institutions: Massachusetts Institute of Technology (MIT)
FleXray introduces a scalable, physics-based generative data engine that leverages 3D CT data and generative image editing to train a generalist model for full-body X-ray segmentation. The paper demonstrates that synthetic data, when carefully designed to cover anatomical, geometric, and appearance variation, can effectively bridge the gap to real-world clinical X-rays, enabling quantitative analysis and downstream clinical applications with high accuracy and robustness.
The paper proposes a highly effective data-centric approach to solving the problem of scarce, high-quality annotations for X-ray segmentation. The core innovation is a "physics-based generative data engine" that leverages existing 3D CT segmentation datasets to simulate 2D X-rays (Digitally Reconstructed Radiographs, DRRs). The methodology is robust, combining online analytic projections for geometric diversity with a generative image-editing model (Flux) to bridge the domain gap between synthetic DRRs and real clinical X-rays. The use of a language-conditioned diffusion model to add realistic artifacts (scatter, text, detector noise) while preserving anatomical structure is a clever and effective solution to the "sim-to-real" gap. The training objective is well-designed, handling multi-label overlap and partial annotations from heterogeneous sources.
The evaluation is comprehensive and rigorous. The model is tested on eight held-out datasets spanning various body parts (limbs, pelvis, spine, chest) and populations (adult, pediatric). It outperforms strong baselines like PAXray, TotalSegmentator2D, and FluoroSAM. The paper goes beyond standard segmentation metrics (Dice) to demonstrate downstream utility: automated Cobb angle measurement for scoliosis (achieving inter-rater variability levels), improved 2D/3D registration for surgical navigation, and data-efficient fine-tuning for new pathological targets. The ablation studies clearly isolate the contributions of generative enhancement and attenuation randomization, showing where each component is most beneficial.
High. The authors release the model, code, the full-body X-ray segmentation dataset, and a browser-based tool. The data engine components, including the specific prompts used for the generative editor and the filtering strategies, are detailed in the appendix. The use of public datasets (MOOSE, etc.) for training data generation further enhances reproducibility.
The model inherits biases from the CT source data, which may not fully represent all clinical X-ray scenarios (e.g., severe trauma, specific implants not present in CTs). The evaluation, while broad, is limited by the scarcity of public X-ray segmentation datasets with ground truth. The reliance on a large generative model (Flux) for data augmentation adds computational cost to the data preparation phase, though this is a one-time cost.
This work has significant potential to transform clinical X-ray analysis by making it quantitative and automatable. By providing a generalist model that segments 60 anatomical structures, it enables population-scale studies, automated disease grading, and improved surgical navigation. The release of the data engine and dataset will likely accelerate research in medical imaging by providing a scalable way to generate labeled training data for other modalities or tasks. FleXray introduces a scalable, physics-based generative data engine that leverages 3D CT data and generative image editing to train a generalist model for full-body X-ray segmentation. The paper demonstrates that synthetic data, when carefully designed to cover anatomical, geometric, and appearance variation, can effectively bridge the gap to real-world clinical X-rays, enabling quantitative analysis and downstream clinical applications with high accuracy and robustness.
Many scientific datasets, such as molecular conformational ensembles or single-cell tissue measurements, are naturally modeled as meta-distributions: distributions over probability measures on non-Euclidean domains. Existing generative methods largely assume Euclidean geometry and fail to capture this structure. We introduce Riemannian Wasserstein Entropic Flow Matching (RWEFM), a generative framework on the Wasserstein space $\mathcal{P}_2(\mathcal{M})$ of a Riemannian manifold $(\mathcal{M},g)$. RWEFM is trained by regressing a neural vector field onto Riemannian optimal transport velocities, using McCann displacement interpolations as conditional paths. We confirm theoretically that this construction leads to a valid flow matching approach on $\mathcal{P}_2(\mathcal{M})$ and introduce the Riemannian Entropic Map, a GPU-efficient approximation of the optimal transport map on manifolds. Our experiments show that by respecting the intrinsic geometry of the data, RWEFM can generate whole single-cell samples in hyperspherical latent spaces and protein conformational ensembles on the torus. As RWEFM requires only a geodesic distance and a projection operator, it is not restricted to manifolds with closed-form geometry, which we demonstrate by generating distributions on a general triangulated mesh.
Primary: Genentech Inc.
All Institutions: Genentech Inc., Yale University
The paper presents a novel and theoretically sound framework for generative modeling on Riemannian manifolds, effectively extending flow matching to the Wasserstein space of probability measures. By introducing the Riemannian Entropic Map and demonstrating applications to single-cell genomics and protein structures, it offers a significant methodological advance for handling complex, non-Euclidean scientific data.
The paper introduces Riemannian Wasserstein Entropic Flow Matching (RWEFM), a rigorous extension of Flow Matching to the Wasserstein space of probability measures on Riemannian manifolds. The core theoretical contribution is the validation of the flow matching objective in this infinite-dimensional, non-Euclidean setting, utilizing McCann displacement interpolations. A key methodological innovation is the "Riemannian Entropic Map," a GPU-efficient estimator for the optimal transport map that generalizes the Euclidean entropic map. It employs a lift-average-retract procedure (logarithmic map to tangent space, barycentric projection, exponential map back to manifold) which is computationally tractable and theoretically grounded with error bounds. The framework is designed to be geometry-agnostic, requiring only geodesic distances and projection operators, allowing application to complex shapes like triangulated meshes.
The experiments are diverse and scientifically relevant. The authors demonstrate the method on synthetic data (MNIST/EMNIST/KMNIST mapped to sphere, hyperbolic space, and torus) to validate geometric correctness. They apply the method to real-world scientific problems: generating single-cell RNA-seq samples on hyperspherical latent spaces and protein conformational ensembles on the torus. The inclusion of a general triangulated mesh (Stanford Bunny) experiment is particularly strong, as it proves the method's applicability beyond closed-form geometries. The metrics used (1-NN deviation, MMD, Chamfer Distance) are appropriate for distributional comparison.
The paper provides a public GitHub repository with code and tutorials. The appendix contains detailed hyperparameters, network architecture descriptions (self-attention blocks), and explicit formulas for geometric operations on various manifolds. The training procedure is well-documented, including details on noise generation and mini-batch OT coupling. This level of detail supports high reproducibility.
The method relies on the computation of optimal transport plans, which can be computationally expensive for very large point clouds, although the entropic regularization helps. The "sampled map" approximation used in high-dimensional settings (like single-cell data) may introduce bias compared to the true barycentric map. The theoretical guarantees for the Riemannian Entropic Map depend on regularity assumptions that may not hold for all practical datasets.
This work bridges the gap between geometric deep learning and generative modeling for distributional data. It provides a toolkit for scientists working with non-Euclidean data (molecules, cells, climate) to generate realistic samples that respect the underlying geometry. The framework is likely to influence future work in scientific machine learning and optimal transport. The paper presents a novel and theoretically sound framework for generative modeling on Riemannian manifolds, effectively extending flow matching to the Wasserstein space of probability measures. By introducing the Riemannian Entropic Map and demonstrating applications to single-cell genomics and protein structures, it offers a significant methodological advance for handling complex, non-Euclidean scientific data.
The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. Mobile planning offers a demanding test of this approach: complex, long-horizon tasks challenge agent reliability, while costly real-device interaction limits development scalability. The framework connects data production, model training, and deployment through a shared action-feedback-verification contract. (i) AI for Data builds a human-gated agentic data flywheel in which specialized agents construct tasks, collect interaction trajectories, curate and balance training data, and use training feedback to guide subsequent data generation. (ii) AI for Training combines a supervised planning cold start with hybrid-environment online agentic reinforcement learning, where we introduce Competence-Aware Reward-and-Advantage Engineering (CARE) to reduce reasoning and tool-use costs while preserving task performance. (iii) AI drives model--harness co-evolution through an execution-evidence-driven loop that orchestrates memory, skills, and tools at runtime and feeds structured action feedback and preserved failure traces back into coordinated model and harness adaptation. Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination. Further evaluations of our model show improvements across non-mobile agentic benchmarks while largely preserving general capabilities.
Primary: Alibaba Group
All Institutions: Alibaba Group
The paper presents a comprehensive framework for developing mobile planner agents using an AI-for-AI closed loop, achieving state-of-the-art performance on MobilePA-Bench. It demonstrates that combining AI-assisted data generation, competence-aware RL, and a dynamic runtime harness can significantly enhance agent reliability and efficiency, offering a scalable pathway for building next-generation autonomous systems.
The paper proposes a "closed-loop AI-for-AI" framework, which is a high-level architectural concept rather than a single novel algorithmic breakthrough. The core technical contributions are the integration of a data flywheel (AI-assisted task generation and curation), a hybrid training pipeline (SFT cold start + RL), and a runtime Harness (Skills/Memory). The most specific technical novelty is CARE (Competence-Aware Reward-and-Advantage Engineering), which adjusts reward shaping based on group success rates to prevent efficiency signals from dominating in saturated success groups. While the components (RL for agents, memory systems, data synthesis) are individually known, the systematic integration into a self-improving loop for mobile agents is a significant engineering and methodological contribution.
The evaluation is conducted on MobilePA-Bench, a large-scale benchmark (1,700+ tasks). The results show Qwen-Planner-Agent (27B) outperforming strong closed-source competitors like GPT-6 Astra and Claude Opus 5, as well as larger open-source models. The ablation studies effectively isolate the contributions of the Planner Model versus the Harness, demonstrating that the runtime context (Skills/Memory) provides substantial gains over the raw model. The efficiency analysis (cost per task) is a valuable addition, showing the agent is not just better but cheaper than frontier commercial APIs.
As a report from a major industry lab (Alibaba), the paper provides high-level architectural details but likely lacks the granular hyperparameters, exact prompt templates, and code releases typically required for full academic reproducibility. The "AI-for-AI" loop implies a complex, proprietary infrastructure for data generation and verification that is difficult to replicate externally. However, the clarity of the framework description allows for conceptual replication.
The primary limitation is the reliance on a proprietary, large-scale infrastructure for the "AI-for-AI" loop, making it difficult for smaller labs to verify the specific benefits of the data flywheel. The "closed-loop" claim is somewhat aspirational; the paper admits that human review is retained for critical decisions, meaning it is not fully autonomous. Additionally, the evaluation is heavily skewed toward the specific mobile planning domain, and while generalization is claimed, the non-mobile benchmarks are less detailed in the provided text.
This paper represents a significant step toward scalable agent development. By formalizing the use of AI to generate training data and diagnose failures for agent systems, it offers a roadmap for reducing the manual effort required to build robust LLM agents. The focus on mobile planning is highly relevant to current industry trends in on-device and cross-app automation. The CARE method offers a useful technique for RL training of agents where success rates vary widely across tasks. The paper presents a comprehensive framework for developing mobile planner agents using an AI-for-AI closed loop, achieving state-of-the-art performance on MobilePA-Bench. It demonstrates that combining AI-assisted data generation, competence-aware RL, and a dynamic runtime harness can significantly enhance agent reliability and efficiency, offering a scalable pathway for building next-generation autonomous systems.
We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate. Building such a teammate requires combining two difficult capabilities: it must perceive and respond to a constantly changing game world under strict latency constraints while interacting naturally with players, keeping its speech synchronized with its actions. Ally therefore combines agentic tool use with real-time game control. A language-model agent uses a controlled interface to inspect game information, interpret player speech, maintain context, decide what to say, and issue high-level action choices that steer a faster control layer for movement, combat, and recovery. Because the player's and Ally's speech and actions continually shape each other and the course of the match, training requires data from actual gameplay. We therefore collect data across nearly 39k sessions in which real players play alongside Ally, recording gameplay, player speech, agent decisions, tool use, actions, and player feedback, and use these records for iterative training. To evaluate teammate quality, we use player feedback and preference comparisons to identify gaps between offline evaluations and player preferences, and iteratively refine the evaluation criteria. Deploying Ally in live service further requires low-latency on-device execution and safeguards for player-facing communication, which we address through model compression, context compaction, targeted safety training, runtime guardrails, and memory redaction. During the live service, we surveyed players in 141 countries. Among respondents whose play with Ally was confirmed in game records, positive responses exceeded negative responses by 25.1 percentage points when asked whether they would recommend Ally, with players describing Ally not only as a tool but also as a teammate or companion.
Primary: Tencent (inferred from PUBG context and typical industry affiliation for such large-scale deployment papers, though not explicitly stated in the provided snippet)
All Institutions: Unknown (based on provided text)
The paper presents a large-scale deployment of an LLM-based embodied agent in a real-time game, demonstrating the practical challenges and solutions for integrating reasoning and control. While the engineering effort and user satisfaction metrics are significant, the lack of rigorous offline benchmarking and reproducibility limits its technical impact on the broader ML research community.
The paper proposes a hybrid architecture for an embodied conversational agent in a complex, real-time game environment (PUBG). The core methodological contribution is the separation of a high-level Language Model (LM) agent responsible for reasoning, context maintenance, and high-level decision-making, from a low-level control layer responsible for fast, precise movement and combat actions. This "agentic tool use" approach allows the LM to interact with the game state via a controlled interface rather than generating raw motor commands, which is a practical and effective solution to the latency and precision constraints of real-time gaming. The system integrates speech recognition, natural language processing, and game control loops. The training methodology relies on iterative refinement using data collected from nearly 39,000 real gameplay sessions, which is a significant engineering effort. The inclusion of safety mechanisms, such as model compression, context compaction, and runtime guardrails, demonstrates a mature approach to deploying LLMs in live, player-facing services.
The evaluation is primarily driven by live service deployment metrics rather than standard offline benchmarks. The authors report player feedback and preference comparisons, noting that positive responses exceeded negative responses by 25.1 percentage points in a survey of players across 141 countries. While this provides strong evidence of user satisfaction and practical utility, it lacks the rigorous quantitative comparison against baseline agents (e.g., traditional rule-based bots or non-conversational AI) that is typical in academic ML papers. The paper identifies gaps between offline evaluations and player preferences, which is an insightful finding, but the specific metrics used to quantify "teammate quality" are subjective and difficult to reproduce or compare across different systems.
Reproducibility is low for external researchers. The system relies on proprietary game assets, specific game APIs, and a large-scale data collection pipeline involving real players. The details of the "controlled interface" and the specific training data splits are mentioned but not fully detailed in a way that would allow independent replication. The model compression and safety training techniques are described at a high level, but specific hyperparameters and architectural choices for the LLM and control layers are likely proprietary or too complex to fully disclose.
The primary limitation is the lack of generalizability and standard benchmarking. The results are specific to PUBG and the particular implementation of the Ally agent. The evaluation relies heavily on subjective player surveys, which can be influenced by novelty effects and do not strictly measure objective performance improvements (e.g., win rate, kill/death ratio) compared to non-conversational AI. The paper does not provide a detailed breakdown of failure modes or specific scenarios where the agent struggles, beyond general mentions of latency and safety constraints.
This paper demonstrates the feasibility of deploying sophisticated LLM-based agents in real-time, high-stakes interactive environments. It highlights the challenges of integrating slow, reasoning-heavy models with fast, reactive control systems, a problem relevant to robotics and other embodied AI applications. The focus on safety, latency, and user experience in live deployment offers valuable insights for industry practitioners looking to integrate LLMs into consumer-facing products. However, its impact on the broader ML research community is limited by the lack of open-source code, standard benchmarks, and generalizable methodological contributions beyond the specific game context. The paper presents a large-scale deployment of an LLM-based embodied agent in a real-time game, demonstrating the practical challenges and solutions for integrating reasoning and control. While the engineering effort and user satisfaction metrics are significant, the lack of rigorous offline benchmarking and reproducibility limits its technical impact on the broader ML research community.
Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfold across multi-shot sequences. In this paper, we present WanPE, a 397B-parameter prompt enhancement model trained on 1.05M real-world videos to master director-level cinematic planning. WanPE formulates shot-level cinematic plans via video-grounded reverse construction and employs Semantic-Consistency GRPO (SC-GRPO) to faithfully preserve user requirements across shots and over time. To benchmark this capability, we curate WanPEval, a human-annotated testbed covering durations from 5 to 30 seconds across varying intent granularities, supported by approximately 11K blind pairwise assessments. When powering Wan3.0's video generator, WanPE-397B boosts human preference over raw user prompts by 10.66-18.84 points at 5-15 seconds and by a dramatic 50.86 points in the 30-second arena. Ablation studies show that reverse construction demonstrates clear superiority over forward rewriting, while SC-GRPO robustly preserves semantic fidelity across model scales. Ultimately, WanPE leads all evaluated commercial offerings at 5-15 seconds and remains competitive with Seedance 2.5 at 30 seconds.
Primary: Alibaba Group
All Institutions: Alibaba Group, Zhejiang University
WanPE introduces a 397B-parameter prompt enhancement model and a new benchmark for cinematic text-to-video generation, demonstrating that structured, reverse-constructed prompts significantly improve video generation fidelity and human preference. The paper presents a rigorous industrial-scale approach to prompt engineering, leveraging large-scale video data and a novel reinforcement learning objective (SC-GRPO) to ensure semantic consistency, thereby establishing a new standard for evaluating and enhancing the textual conditioning of modern video generators.
The paper proposes WanPE, a 397B-parameter Large Language Model specifically fine-tuned for prompt enhancement in text-to-video generation. The core methodological contribution is the formulation of "shot-level cinematic plans" via video-grounded reverse construction, where the model learns to decompose a high-level user intent into a detailed, temporally consistent screenplay. A key technical innovation is the use of Semantic-Consistency GRPO (SC-GRPO), a reinforcement learning objective designed to preserve the semantic fidelity of the original user prompt across the generated multi-shot sequence. This addresses a common failure mode in LLM-based prompt rewriting where the enhanced prompt drifts from the user's original intent. The scale of the model (397B) and the training data (1.05M real-world videos) are significant, indicating a heavy industrial investment in this specific sub-task of the video generation pipeline.
The authors introduce WanPEval, a human-annotated benchmark covering video durations from 5 to 30 seconds with varying intent granularities. The evaluation relies on approximately 11K blind pairwise assessments, which is a robust methodology for measuring human preference in generative tasks. The results show substantial improvements in human preference scores (10.66-50.86 points) over raw prompts when powering the Wan3.0 video generator. The ablation studies effectively isolate the contributions of reverse construction versus forward rewriting and the impact of SC-GRPO, providing clear evidence for the necessity of these components. The comparison against commercial offerings (including Seedance 2.5) adds practical relevance, though the specific metrics for these comparisons are not detailed in the abstract.
Reproducibility is severely limited by the scale of the model (397B parameters) and the proprietary nature of the training data (1.05M real-world videos) and the base video generator (Wan3.0). While the methodology is described, the specific hyperparameters, data curation details, and code for the SC-GRPO implementation are likely not fully public, making independent verification difficult for the broader research community. The benchmark WanPEval is a positive step, but access to the full dataset and annotation guidelines is not confirmed in the provided text.
The primary limitation is the lack of open-source availability for the model weights and training pipeline, which restricts the ability of other researchers to build upon or verify the results independently. The evaluation is heavily dependent on the specific video generator (Wan3.0); it is unclear how well the enhanced prompts transfer to other video generation models. Additionally, the computational cost of using a 397B model for prompt enhancement at inference time may be prohibitive for real-time applications, although this is not explicitly discussed in the abstract.
This work highlights the growing importance of the "prompt engineering" layer in complex generative AI systems. By treating prompt enhancement as a distinct, trainable task with its own model architecture and training objectives, the paper sets a precedent for modularizing the text-to-video pipeline. The introduction of a dedicated benchmark for cinematic prompt quality (WanPEval) will likely influence how the community evaluates the textual conditioning of video models. The findings suggest that significant gains in video generation quality can be achieved not just by scaling the video model itself, but by improving the quality and structure of the textual input. WanPE introduces a 397B-parameter prompt enhancement model and a new benchmark for cinematic text-to-video generation, demonstrating that structured, reverse-constructed prompts significantly improve video generation fidelity and human preference. The paper presents a rigorous industrial-scale approach to prompt engineering, leveraging large-scale video data and a novel reinforcement learning objective (SC-GRPO) to ensure semantic consistency, thereby establishing a new standard for evaluating and enhancing the textual conditioning of modern video generators.
While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model outputs a superposition of the individual next-token distributions. We term this the Superposition Linearity Hypothesis. We provide evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training; in fact, we observe that it tends to diminish as pretraining progresses. However, we demonstrate that linearity can be substantially restored through lightweight fine-tuning, significantly reducing the divergence between the predicted next-token distribution and the average of the individual next-token distributions. Finally, we introduce a guided decoding procedure that disentangles superposed outputs, enabling the simultaneous generation of two coherent continuations from a single forward pass.
Primary: unknown
All Institutions: unknown
The paper demonstrates that Transformers retain intrinsic linear superposition properties that degrade during pre-training but can be restored via lightweight fine-tuning, enabling a proof-of-concept for parallel decoding of two text streams from a single forward pass. While the empirical findings on the linearity of the residual stream are interesting and well-supported, the practical impact is limited by significant degradation in single-stream performance and modest gains in parallel decoding accuracy, suggesting that this approach is not yet a viable alternative to standard batching or specialized multiplexing architectures.
The paper proposes the "Superposition Linearity Hypothesis," arguing that standard Transformers exhibit intrinsic linear superposition properties when processing averaged input embeddings. The methodology involves three main components: (1) empirical validation of this linearity in pre-trained models (Pythia, Llama, Qwen, etc.) using rank analysis and distributional divergence metrics (KL, JS, Wasserstein); (2) a self-distillation fine-tuning procedure to restore degraded linearity, minimizing the KL divergence between the mixed-input output and the average of independent outputs; and (3) a "Joint Contrastive" decoding strategy using a small auxiliary model to disentangle the superposed streams. The approach is sound but relies on a specific definition of linearity (element-wise averaging of embeddings) that may not generalize to all forms of input mixing. The fine-tuning objective is straightforward but effective in the tested regime.
The experiments are extensive, covering multiple model families and sizes. The use of LAMBADA for content-position evaluation and FineWeb for general analysis is appropriate. The results show a clear degradation of linearity during pre-training and its restoration via fine-tuning. However, the practical utility is limited by the significant drop in single-stream performance (e.g., LAMBADA accuracy drops from 0.544 to 0.357 on Pythia-2.8B) and the modest gains in parallel decoding accuracy (0.43 vs 0.54 single-stream baseline on Llama-3.2-3B). The throughput claims are theoretical; actual latency benefits are not rigorously benchmarked against standard batching optimizations.
The paper provides detailed hyperparameters, dataset sources, and training costs. The code for the specific fine-tuning and decoding procedures is not explicitly linked in the text provided, but the methodology is described with sufficient detail for replication. The use of standard datasets (FineWeb, LAMBADA, TinyStories) aids reproducibility.
The primary limitation is the trade-off between superposition fidelity and single-stream quality. The fine-tuned models perform significantly worse on standard language modeling tasks. Additionally, the decoding accuracy for the two-stream setup remains lower than single-stream baselines, questioning the practical viability of this approach for high-quality generation. The context length is limited to 128-512 tokens, which is short for modern LLMs. The "intrinsic" nature of the linearity is challenged by the fact that it degrades during training, suggesting it is a byproduct of initialization rather than a robust architectural feature.
The work contributes to the understanding of Transformer internal representations and offers a potential path for efficient inference via stream multiplexing. However, the current performance trade-offs limit its immediate adoption. It provides a valuable baseline for future work on linear superposition and parallel decoding in LLMs. The paper demonstrates that Transformers retain intrinsic linear superposition properties that degrade during pre-training but can be restored via lightweight fine-tuning, enabling a proof-of-concept for parallel decoding of two text streams from a single forward pass. While the empirical findings on the linearity of the residual stream are interesting and well-supported, the practical impact is limited by significant degradation in single-stream performance and modest gains in parallel decoding accuracy, suggesting that this approach is not yet a viable alternative to standard batching or specialized multiplexing architectures.
General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirectly, to predict constraints or write programs, or give them a view of the scene rather than a world in which to act. We present World Action Agent (WAA), a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace. The workspace has three properties. Contact views, selected automatically from the scene geometry, present the scene around the current interaction. Action rehearsal turns each action into an editable proposal that the agent, alone or through an Imagination Agent, previews and revises against planning feedback before execution. In-view correction closes the loop between observation, rehearsal, and low-level execution, letting the agent remove residual offsets in the view where it observes them. Through the same workspace, WAA acquires embodied procedural knowledge in two ways: it evolves multimodal skills from expert videos and human teaching under evidence-based review and consults them through a Skill Agent, and its interaction traces train smaller VLMs to pilot the same harness. On LIBERO-Pro, WAA with skills evolved only from LIBERO-90 reaches a state-of-the-art 75.6% average success, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone; the same skills remain effective on robosuite without further learning. Fine-tuning Qwen3.5-9B on harness traces raises its out-of-domain success from 1.7% to 43.3%.
Primary: Unknown
All Institutions: Unknown
The paper introduces a multi-agent VLM framework for robot manipulation that achieves state-of-the-art results on LIBERO-Pro by integrating visual action workspaces with skill evolution. While the methodology is innovative and the reported results are strong, the lack of full text visibility in the prompt prevents a complete assessment of experimental rigor, resulting in a solid but not exceptional score.
The paper proposes World Action Agent (WAA), a multi-agent framework that leverages Vision-Language Models (VLMs) for robot manipulation. The core innovation is the "visual action workspace," which integrates three components: contact views (automatic scene cropping around interaction points), action rehearsal (editable action proposals previewed by an Imagination Agent), and in-view correction (closing the loop between observation and low-level execution). The system also incorporates a Skill Agent that evolves procedural knowledge from expert videos and human teaching. The methodology is logically sound and addresses a known limitation of VLMs in robotics: the lack of a grounded "world" for action planning. The use of an "Imagination Agent" to preview actions before execution is a creative application of VLM capabilities.
The paper reports state-of-the-art results on the LIBERO-Pro benchmark, achieving 75.6% average success using skills evolved from LIBERO-90. It claims outperformance over end-to-end Vision-Language-Action (VLA) models, code-as-policy agents, and a visual-harness baseline. Additionally, it demonstrates transferability to robosuite without further learning and shows significant improvement in out-of-domain success for a fine-tuned Qwen3.5-9B model (from 1.7% to 43.3%). However, the full paper text provided in the prompt is truncated (only section headers are visible), so the rigor of the ablation studies, statistical significance, and detailed error analysis cannot be fully verified. The reported numbers are strong, but the lack of visible detailed experimental breakdown in the provided text limits full confidence in the robustness of the claims.
The paper mentions specific models (Qwen3.5-9B) and benchmarks (LIBERO-Pro, LIBERO-90, robosuite), which aids reproducibility. However, without access to the full methods section detailing the specific prompts, agent architectures, and training procedures for the skill evolution, exact reproduction may be challenging. The reliance on "expert videos" and "human teaching" for skill acquisition introduces variability that is difficult to standardize.
The primary limitation is the reliance on a specific VLM backbone (Qwen3.5-9B) and the complexity of the multi-agent setup, which may introduce latency and computational overhead. The "action rehearsal" step, while innovative, adds a layer of inference that could slow down real-time manipulation. Furthermore, the generalization beyond the tested benchmarks (LIBERO, robosuite) to real-world, unstructured environments is not explicitly demonstrated in the abstract.
This work contributes to the trend of using VLMs as high-level planners in robotics. The concept of "action rehearsal" could be applicable to other embodied AI tasks, such as autonomous driving or human-robot interaction, where previewing actions before execution is critical. The ability to evolve skills from videos suggests a path toward more data-efficient robot learning. The paper introduces a multi-agent VLM framework for robot manipulation that achieves state-of-the-art results on LIBERO-Pro by integrating visual action workspaces with skill evolution. While the methodology is innovative and the reported results are strong, the lack of full text visibility in the prompt prevents a complete assessment of experimental rigor, resulting in a solid but not exceptional score.
Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by similarity. This forces the system to decide what is worth remembering before the future query is known, irreversibly discarding information and producing a query-independent summary that must serve many possible downstream tasks. Learning such a write-time curator is also difficult because the value of a storage decision may only become apparent when a relevant query arrives, potentially many tasks later, creating a long-horizon credit-assignment problem. We instead retain raw trajectories and defer curation until read time, when the current task is known. Given the retrieved traces and the new task, a memory curator synthesizes a compact, task-adaptive payload tailored to the immediate need. Because this payload is consumed on the same task, the curator can be trained directly from immediate task success, avoiding delayed utility signals and the need to artificially group related tasks. Across ALFWorld, WebShop, and τ^2-bench, our Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively. Notably, even an untrained curator is already competitive with or surpasses these baselines, showing that task-adaptive read-time curation itself is a major source of the gain; training the curator further compounds the improvement.
Primary: Salesforce
All Institutions: Salesforce
[One sentence main contribution]. The paper introduces Just-in-Time Memory, a read-time curation strategy for LLM agents that outperforms write-time methods by synthesizing task-adaptive payloads from raw trajectories, demonstrating significant gains in success rates across multiple benchmarks.
The paper proposes "Just-in-Time Memory" (JitMem), shifting the curation of agent memory from write-time (fixed summaries) to read-time (task-adaptive synthesis). The core method involves retaining raw trajectories and using a "memory curator" LLM to synthesize a compact payload specifically for the current query. This addresses the long-horizon credit assignment problem inherent in write-time curation by allowing the curator to be trained directly on immediate task success. The approach is conceptually sound, leveraging the flexibility of LLMs to perform on-the-fly reasoning over raw data rather than relying on pre-computed, potentially lossy, static representations.
The experiments are conducted on three standard agent benchmarks: ALFWorld, WebShop, and τ^2-bench. The results show consistent improvements over no-memory baselines and existing write-time memory methods. The reported gains are substantial (16.2, 16.3, and 3.9 absolute points), suggesting the method is effective across different task types. The finding that even an untrained curator is competitive is a strong empirical result, highlighting the value of the read-time curation paradigm itself.
The paper is on arXiv. While the methodology is described, the specific implementation details of the "memory curator" training (e.g., reward shaping, specific prompts, hyperparameters) may be limited in the abstract-only view, but the full text likely contains sufficient detail for reproduction given the standard nature of LLM agent frameworks.
The primary limitation is computational cost. Retaining raw trajectories and performing LLM-based synthesis at read time is significantly more expensive than retrieving a pre-computed summary. This may limit scalability to very long-horizon agents or high-throughput applications. Additionally, the method relies heavily on the base LLM's ability to synthesize relevant information from raw traces, which may vary across model capabilities.
This work has significant implications for the design of LLM agents, suggesting that static memory structures may be suboptimal. It encourages the development of more dynamic, query-aware memory systems. The approach could be extended to other domains requiring adaptive information retrieval, such as RAG systems that dynamically re-rank or synthesize context based on the specific query. [One sentence main contribution]. The paper introduces Just-in-Time Memory, a read-time curation strategy for LLM agents that outperforms write-time methods by synthesizing task-adaptive payloads from raw trajectories, demonstrating significant gains in success rates across multiple benchmarks.
Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low extension cost. Existing generation pipelines commonly construct an environment before defining its outcome rule or annotating its trajectories, leaving dynamics and evaluation to be aligned post hoc. VHD-Play reverses this dependency by sampling and solving a mathematical model before a corpus-grounded setter renders its decision process as stateful tools. The executable dynamics and trajectory-scoring reference are inherited from the same solved model. The pipeline produces 3,300 diverse agentic environments at a cost of a few cents each. Training Qwen3.6-35B-A3B on three families raises its mean agentic score from 0.204 to 0.815 in a five-family diagnostic. Gains also appear on held-out instances from all three training families and eight unseen mechanism families, then extend beyond the generated substrate to external benchmarks for general function calling, travel planning, and 365-day e-commerce. On E-Commerce Bench, the trained checkpoint completes every run without bankruptcy and exceeds Qwen3.7-Max. We compare written-out problems with stateful versions that reveal or hide their parameters. The comparison shows that most of the learnable gap lies in stateful interaction rather than underlying problem solving. A frozen 35B setter realizes larger environments, and scale-matched training retains gains as mechanism size and horizon grow, indicating the potential for an evolving training substrate.
Primary: Georgia Institute of Technology
All Institutions: Georgia Institute of Technology
The paper introduces VHD-Play, a pipeline that generates verifiable agentic RL environments by solving mathematical models before rendering them as stateful tools, significantly improving agent performance and generalization. The approach offers a scalable and cost-effective solution to the environment generation bottleneck in agentic RL, with strong empirical evidence of its efficacy across diverse tasks and benchmarks.
The paper proposes VHD-Play, a pipeline that inverts the standard environment generation process. Instead of defining an environment and then trying to align rewards, it first samples and solves a mathematical model (mechanism), then uses a "setter" LLM to render this solved model into stateful tools and decision processes. This ensures that the dynamics and the scoring reference are inherently consistent because they derive from the same solved instance. The methodology is clever in its dependency inversion, addressing a common pain point in agentic RL where reward hacking or misaligned dynamics occur. The use of a frozen 35B setter to generate larger environments is a practical architectural choice that balances cost and capability.
The experiments are extensive, generating 3,300 environments at low cost. The evaluation shows a significant improvement in agentic scores (0.204 to 0.815) for Qwen3.6-35B-A3B. Crucially, the paper demonstrates generalization to held-out instances and unseen mechanism families, as well as external benchmarks like E-Commerce Bench, where the model outperforms Qwen3.7-Max. The ablation comparing written-out problems vs. stateful interactions is insightful, highlighting that the gain comes from learning stateful interaction rather than just solving the underlying math.
The paper mentions specific model sizes (Qwen3.6-35B-A3B) and costs ("a few cents each"), which aids in understanding the scale. However, without a provided code repository or detailed hyperparameter settings for the "setter" and the RL training loop, full reproducibility is moderate. The reliance on specific proprietary or recent open-source models (Qwen3 series) may limit immediate reproducibility for labs without access to these specific checkpoints.
The primary limitation is the reliance on the quality of the initial mathematical model sampling. If the solver fails or the model is ill-posed, the environment generation fails. Additionally, the "few cents" cost claim, while impressive, may not scale linearly to much larger, more complex real-world scenarios without significant engineering overhead. The generalization to external benchmarks is promising but limited to specific domains (e-commerce, travel).
This work has high potential impact on the field of agentic AI. By providing a scalable, low-cost method to generate diverse, verifiable environments, it lowers the barrier to entry for training robust agents. The insight that stateful interaction is the key learnable gap is valuable for future agent design. The ability to generate thousands of environments cheaply could accelerate research in long-horizon planning and decision-making. The paper introduces VHD-Play, a pipeline that generates verifiable agentic RL environments by solving mathematical models before rendering them as stateful tools, significantly improving agent performance and generalization. The approach offers a scalable and cost-effective solution to the environment generation bottleneck in agentic RL, with strong empirical evidence of its efficacy across diverse tasks and benchmarks.
Many scientific datasets, such as molecular conformational ensembles or single-cell tissue measurements, are naturally modeled as meta-distributions: distributions over probability measures on non-Euclidean domains. Existing generative methods largely assume Euclidean geometry and fail to capture this structure. We introduce Riemannian Wasserstein Entropic Flow Matching (RWEFM), a generative framework on the Wasserstein space $\mathcal{P}_2(\mathcal{M})$ of a Riemannian manifold $(\mathcal{M},g)$. RWEFM is trained by regressing a neural vector field onto Riemannian optimal transport velocities, using McCann displacement interpolations as conditional paths. We confirm theoretically that this construction leads to a valid flow matching approach on $\mathcal{P}_2(\mathcal{M})$ and introduce the Riemannian Entropic Map, a GPU-efficient approximation of the optimal transport map on manifolds. Our experiments show that by respecting the intrinsic geometry of the data, RWEFM can generate whole single-cell samples in hyperspherical latent spaces and protein conformational ensembles on the torus. As RWEFM requires only a geodesic distance and a projection operator, it is not restricted to manifolds with closed-form geometry, which we demonstrate by generating distributions on a general triangulated mesh.
Primary: Genentech Inc.
All Institutions: Genentech Inc., Yale University
The paper presents a novel and theoretically sound framework for generative modeling on Riemannian manifolds, effectively extending flow matching to the Wasserstein space of probability measures. By introducing the Riemannian Entropic Map and demonstrating applications to single-cell genomics and protein structures, it offers a significant methodological advance for handling complex, non-Euclidean scientific data.
The paper introduces Riemannian Wasserstein Entropic Flow Matching (RWEFM), a rigorous extension of Flow Matching to the Wasserstein space of probability measures on Riemannian manifolds. The core theoretical contribution is the validation of the flow matching objective in this infinite-dimensional, non-Euclidean setting, utilizing McCann displacement interpolations. A key methodological innovation is the "Riemannian Entropic Map," a GPU-efficient estimator for the optimal transport map that generalizes the Euclidean entropic map. It employs a lift-average-retract procedure (logarithmic map to tangent space, barycentric projection, exponential map back to manifold) which is computationally tractable and theoretically grounded with error bounds. The framework is designed to be geometry-agnostic, requiring only geodesic distances and projection operators, allowing application to complex shapes like triangulated meshes.
The experiments are diverse and scientifically relevant. The authors demonstrate the method on synthetic data (MNIST/EMNIST/KMNIST mapped to sphere, hyperbolic space, and torus) to validate geometric correctness. They apply the method to real-world scientific problems: generating single-cell RNA-seq samples on hyperspherical latent spaces and protein conformational ensembles on the torus. The inclusion of a general triangulated mesh (Stanford Bunny) experiment is particularly strong, as it proves the method's applicability beyond closed-form geometries. The metrics used (1-NN deviation, MMD, Chamfer Distance) are appropriate for distributional comparison.
The paper provides a public GitHub repository with code and tutorials. The appendix contains detailed hyperparameters, network architecture descriptions (self-attention blocks), and explicit formulas for geometric operations on various manifolds. The training procedure is well-documented, including details on noise generation and mini-batch OT coupling. This level of detail supports high reproducibility.
The method relies on the computation of optimal transport plans, which can be computationally expensive for very large point clouds, although the entropic regularization helps. The "sampled map" approximation used in high-dimensional settings (like single-cell data) may introduce bias compared to the true barycentric map. The theoretical guarantees for the Riemannian Entropic Map depend on regularity assumptions that may not hold for all practical datasets.
This work bridges the gap between geometric deep learning and generative modeling for distributional data. It provides a toolkit for scientists working with non-Euclidean data (molecules, cells, climate) to generate realistic samples that respect the underlying geometry. The framework is likely to influence future work in scientific machine learning and optimal transport. The paper presents a novel and theoretically sound framework for generative modeling on Riemannian manifolds, effectively extending flow matching to the Wasserstein space of probability measures. By introducing the Riemannian Entropic Map and demonstrating applications to single-cell genomics and protein structures, it offers a significant methodological advance for handling complex, non-Euclidean scientific data.
Progress in general-purpose video editing depends on constructing large-scale paired supervision and effectively adapting video-generation backbones to instruction-driven editing. Unlike video generation, video editing must execute a requested transformation while preserving unrelated subjects, scene structure, motion, and temporal continuity. We present VideoX-Qwen, an integrated data-construction and model-training framework for general instruction-based video editing. Our scalable production pipeline organizes specialized generation and understanding models into complementary routes for addition, removal, replacement, and attribute editing, followed by quality screening and instruction enrichment. It produces more than 1.2 million directional video-editing records, including over 400,000 records in each major task group, with an automatic acceptance rate of 89%. The resulting corpus provides broad and structured coverage of common editing operations through a unified source-instruction-target interface. We further develop a unified Qwen-Wan editor that combines multimodal semantic conditioning with dense source-video latent guidance. A progressive image-video training strategy aligns the multimodal instruction interface, adapts the video generator to source-conditioned editing, and refines output quality with selected high-resolution data. In a 100-example comparison with UniVideo and Kling O1, VideoX-Qwen achieves the best mean result on nine of eleven reported metrics, including instruction following, editing quality, content preservation, structural and perceptual similarity, and video-distribution quality. Together, the large-scale data-production system and unified training framework provide a practical foundation for more capable instruction-driven video editing.
Primary: Alibaba Group
All Institutions: Alibaba Group, Zhejiang University
The paper presents a scalable data pipeline and a unified training framework for instruction-based video editing, achieving strong results in a limited comparison. While the data construction methodology is a valuable contribution, the small-scale evaluation limits the confidence in the claimed state-of-the-art performance.
The paper proposes VideoX-Qwen, a framework that integrates a large-scale data construction pipeline with a unified model training strategy for instruction-based video editing. The data pipeline is a significant engineering contribution, leveraging specialized generation and understanding models to create 1.2 million paired video editing records across four task types (addition, removal, replacement, attribute editing). The model architecture, a "Qwen-Wan" editor, combines multimodal semantic conditioning (likely using a Qwen-based LLM/VLM) with dense source-video latent guidance (likely using a Wan-based video diffusion model). The training strategy is progressive, starting with image editing to align the instruction interface, then moving to source-conditioned video editing, and finally refining with high-resolution data. This approach is methodologically sound and addresses the core challenge of video editing: executing specific edits while preserving unrelated content and temporal consistency.
The evaluation is limited to a 100-example comparison against two baselines, UniVideo and Kling O1. While the paper claims state-of-the-art performance on 9 out of 11 metrics, the small sample size (100 examples) is a significant weakness for a paper claiming to introduce a "practical foundation" for general video editing. The metrics reported (instruction following, editing quality, content preservation, etc.) are relevant, but the lack of a larger, standardized benchmark or user study limits the strength of the empirical claims. The comparison with Kling O1, a commercial system, is interesting but potentially unfair if the open-source baselines are not equally tuned.
The paper describes the data pipeline and training strategy in detail, which aids reproducibility. However, the reliance on proprietary or large-scale models (Qwen, Wan) and the specific "specialized generation and understanding models" for data creation makes full reproduction difficult for smaller labs. The code and data are not explicitly mentioned as being released in the provided text, which is a gap for a paper of this scale.
The primary limitation is the small scale of the evaluation (100 examples). This is insufficient to robustly claim superiority over strong baselines like Kling O1. Additionally, the paper does not extensively discuss failure cases or the specific types of edits where the model struggles. The data pipeline, while impressive, may be biased towards the types of edits that are easy to generate and verify automatically, potentially missing more complex or nuanced editing scenarios.
The work has significant potential impact by providing a scalable method for generating high-quality video editing data, which is a major bottleneck in the field. The unified framework for instruction-based editing could enable more accessible and flexible video editing tools for non-experts. The integration of LLM-based instruction understanding with diffusion-based video generation is a trend that is likely to be widely adopted. The paper presents a scalable data pipeline and a unified training framework for instruction-based video editing, achieving strong results in a limited comparison. While the data construction methodology is a valuable contribution, the small-scale evaluation limits the confidence in the claimed state-of-the-art performance.
X-ray is medicine's most widely used imaging modality, yet remains among its least quantitative. Unlike volumetric modalities like CT or MRI, X-ray collapses 3D anatomy into a 2D projection, causing structures to overlap and anatomical boundaries to be ambiguous, even to experts. As a result, labeling X-ray databases for training general-purpose segmentation systems is impractical, leaving morphometric and functional X-ray analysis confined to narrow anatomical regions and applications. To this end, we present FleXray, a generalist model for anatomical segmentation across the entire body in clinical X-rays. Instead of curating large, manually annotated X-ray datasets, we build a scalable, physics-based generative X-ray data engine. Using existing 3D whole-body CT segmentation datasets and generative image-editing models, we simulate fully-annotated 2D X-rays with diverse appearances, physiological properties, and imaging geometries. Trained on these simulations, FleXray accurately segments 60 anatomical structures across unseen research datasets and in-the-wild X-rays. We further show that FleXray makes X-rays directly amenable to quantitative analysis, enabling automated measurements for disease grading, robust navigation during X-ray-guided interventions, and data-efficient learning of pathological targets. We release the model, code, a full-body X-ray segmentation dataset, and a local, easy-to-use browser-based tool at https://flexray.csail.mit.edu .
Primary: Massachusetts Institute of Technology (MIT)
All Institutions: Massachusetts Institute of Technology (MIT)
FleXray introduces a scalable, physics-based generative data engine that leverages 3D CT data and generative image editing to train a generalist model for full-body X-ray segmentation. The paper demonstrates that synthetic data, when carefully designed to cover anatomical, geometric, and appearance variation, can effectively bridge the gap to real-world clinical X-rays, enabling quantitative analysis and downstream clinical applications with high accuracy and robustness.
The paper proposes a highly effective data-centric approach to solving the problem of scarce, high-quality annotations for X-ray segmentation. The core innovation is a "physics-based generative data engine" that leverages existing 3D CT segmentation datasets to simulate 2D X-rays (Digitally Reconstructed Radiographs, DRRs). The methodology is robust, combining online analytic projections for geometric diversity with a generative image-editing model (Flux) to bridge the domain gap between synthetic DRRs and real clinical X-rays. The use of a language-conditioned diffusion model to add realistic artifacts (scatter, text, detector noise) while preserving anatomical structure is a clever and effective solution to the "sim-to-real" gap. The training objective is well-designed, handling multi-label overlap and partial annotations from heterogeneous sources.
The evaluation is comprehensive and rigorous. The model is tested on eight held-out datasets spanning various body parts (limbs, pelvis, spine, chest) and populations (adult, pediatric). It outperforms strong baselines like PAXray, TotalSegmentator2D, and FluoroSAM. The paper goes beyond standard segmentation metrics (Dice) to demonstrate downstream utility: automated Cobb angle measurement for scoliosis (achieving inter-rater variability levels), improved 2D/3D registration for surgical navigation, and data-efficient fine-tuning for new pathological targets. The ablation studies clearly isolate the contributions of generative enhancement and attenuation randomization, showing where each component is most beneficial.
High. The authors release the model, code, the full-body X-ray segmentation dataset, and a browser-based tool. The data engine components, including the specific prompts used for the generative editor and the filtering strategies, are detailed in the appendix. The use of public datasets (MOOSE, etc.) for training data generation further enhances reproducibility.
The model inherits biases from the CT source data, which may not fully represent all clinical X-ray scenarios (e.g., severe trauma, specific implants not present in CTs). The evaluation, while broad, is limited by the scarcity of public X-ray segmentation datasets with ground truth. The reliance on a large generative model (Flux) for data augmentation adds computational cost to the data preparation phase, though this is a one-time cost.
This work has significant potential to transform clinical X-ray analysis by making it quantitative and automatable. By providing a generalist model that segments 60 anatomical structures, it enables population-scale studies, automated disease grading, and improved surgical navigation. The release of the data engine and dataset will likely accelerate research in medical imaging by providing a scalable way to generate labeled training data for other modalities or tasks. FleXray introduces a scalable, physics-based generative data engine that leverages 3D CT data and generative image editing to train a generalist model for full-body X-ray segmentation. The paper demonstrates that synthetic data, when carefully designed to cover anatomical, geometric, and appearance variation, can effectively bridge the gap to real-world clinical X-rays, enabling quantitative analysis and downstream clinical applications with high accuracy and robustness.
Attributing language model outputs to their internal computations is an open problem in interpretability. Existing methods, which use causal interventions, gradients, or learnable masks, either are infeasibly expensive or struggle to identify actual causally-important internal computations. We propose framing attribution as the problem of identifying nested subsets of internal components which minimise a downstream loss. To learn this task, we introduce Matryoshka Attribution (MAttr), a mask learning method that parametrises the mask with a simple differentiable sigmoid top-$k$ operator. We supervise training over all sparsities simultaneously by randomising $k$ over training, resulting in a learned ordering of components by attribution score. MAttr achieves number 1 on the official leaderboard of the Mechanistic Interpretability Benchmark (Mueller et al., 2025); our method identifies sparse and task-transferrable circuits across varying circuit bases. As a practical application, we show that MAttr can be trained with reinforcement learning to identify weight changes responsible for downstream behaviours in LLM finetuning. We train MAttr on refusal judge scores and find that restoring $1\%$ of Llama 3.1 8B Instruct's weights to their base model state is sufficient to remove refusals while maintaining capabilities. We view MAttr as a successful formulation of interpretability into a learnable objective that we can tackle with gradient descent, and encourage future work along these lines.
Primary: Stanford University
All Institutions: Stanford University
Matryoshka Attribution (MAttr) is a novel mask-learning method that uses a differentiable sigmoid top-$k$ operator to learn a single, nested ordering of internal components for attribution, achieving state-of-the-art performance on the Mechanistic Interpretability Benchmark and enabling the identification of specific weight changes responsible for LLM behaviors like refusal through reinforcement learning.
The paper introduces Matryoshka Attribution (MAttr), a mask-learning method that frames attribution as identifying nested subsets of internal components minimizing a downstream loss. The core innovation is the use of a differentiable sigmoid top-$k$ operator to parametrize the mask, allowing the model to learn a single ordering of components that is valid across all sparsity levels (the "Matryoshka" property). This eliminates the need for separate mask training at each sparsity level and avoids the instability and hyperparameter sensitivity associated with hard concrete distributions or straight-through estimators. The method is grounded in causal abstraction, defining "soft interchange interventions" that interpolate between base and source states. A significant theoretical contribution is the demonstration that MAttr recovers Integrated Gradients in the limit of one training step, providing a rigorous bridge between gradient-based and mask-learning attribution methods. The extension to parameter-level attribution via Reinforcement Learning (GRPO) is a novel application, allowing the method to identify weight changes responsible for specific behaviors (e.g., refusal) without requiring differentiable rewards.
The experimental evaluation is extensive and rigorous. MAttr achieves state-of-the-art performance on the Mechanistic Interpretability Benchmark (MIB), significantly outperforming baselines like Integrated Gradients, Interchange Interventions, and other mask-learning methods (DBM, Node Pruning). The authors introduce "MIB+" to test finer-grained bases (MLP neurons, SAE features) and new tasks, showing consistent improvements. The paper also introduces a "compactness" metric to evaluate circuit sparsity, showing MAttr excels at both faithfulness and compactness. The parameter-level attribution experiments on Llama 3.1 8B Instruct are particularly compelling, demonstrating that restoring only 1% of weights to the base model state removes refusals while maintaining capabilities, a result with significant practical implications for model safety and auditing.
The paper provides a public GitHub repository with code. Hyperparameters for all baselines and MAttr are detailed in the appendices. The use of standard benchmarks (MIB) and open-source models (Llama, GPT-2, etc.) enhances reproducibility. The specific details of the RL training setup (GRPO) and reward functions are described, though the complexity of the RL setup may pose challenges for exact replication without the provided code.
The method relies on gradient descent, which may be computationally expensive for very large models or fine-grained bases, although it is parallelizable. The "Matryoshka" property assumes a nested structure of importance, which may not always hold for all types of internal computations. The RL-based parameter attribution is a new paradigm and may require careful tuning of the reward function to avoid reward hacking. The paper focuses primarily on language models; generalization to other modalities (e.g., vision) is mentioned in appendices but not the primary focus.
This paper has high potential impact on the field of interpretability. By providing a robust, learnable, and causally-grounded method for attribution, it enables more reliable identification of task-relevant circuits and weight changes. The application to removing refusals via weight restoration offers a new tool for auditing and potentially mitigating safety issues in LLMs. The unification of gradient-based and mask-learning perspectives may stimulate further theoretical work in attribution. The method's ability to handle non-differentiable objectives via RL opens up new avenues for interpreting complex, emergent behaviors. Matryoshka Attribution (MAttr) is a novel mask-learning method that uses a differentiable sigmoid top-$k$ operator to learn a single, nested ordering of internal components for attribution, achieving state-of-the-art performance on the Mechanistic Interpretability Benchmark and enabling the identification of specific weight changes responsible for LLM behaviors like refusal through reinforcement learning.
Most minimally invasive surgery is still performed with hand-held laparoscopic instruments, and the surgeon's instrument kinematics are lost when the operation ends; only the endoscope video is kept. This paper presents an end-to-end pipeline that captures this motion in the operating room and uses it to train a surgical robot policy, validated on live animals. We introduce a surgical instrument-state logger that mounts on the shaft of a standard laparoscopic instrument and recovers its pose and jaw state from an inertial sensor, a time-of-flight sensor and a Hall sensor, with no external camera or tracker. A data pipeline measures the latency of every sensor channel against a robot ground truth and aligns the channels before forming observation-action pairs. On these demonstrations we train a diffusion policy with a fine-tuned DINOv3 backbone, selecting its design by closed-loop rollouts in a physics simulator reconstructed from depth maps of an ex-vivo rabbit appendix. The policy is then retrained on 849 in-vivo demonstrations from four live rabbits and deployed on four additional live rabbits with electrosurgery armed. With the surgeon selecting the surgical phase, the policy completed the appendectomy in three of the four animals. The results show that demonstrations recorded from a surgeon's own instruments are sufficient to train, select and deploy a bimanual surgical policy in vivo. The robot serves only as the timing reference for sensor calibration and as the executor, and collects no demonstrations. Both demonstration corpora are released to support future surgical robot learning research.
Primary: Seoul National University
All Institutions: Eulji University, Seoul National University, Seoul National University College of Medicine, Seoul National University Hospital
The paper presents a robust end-to-end pipeline for learning bimanual surgical policies from hand-held instrument demonstrations, validated in vivo with electrosurgery. It makes a significant contribution to surgical robotics by introducing a novel, low-cost instrument-state logger and demonstrating that such data is sufficient to train safe, effective policies for live animal surgery, thereby bridging the gap between hand-held surgical practice and robotic automation.
The paper proposes a novel hardware-software pipeline for collecting surgical demonstrations without a robot. The core hardware contribution is a "surgical instrument-state logger" mounted on the shaft of standard laparoscopic instruments, using an IMU, Time-of-Flight (ToF) sensor, and Hall sensor to estimate pose and jaw state. The methodology rigorously addresses the critical issue of sensor latency by fitting a first-order lag model to each channel against a robot ground truth (FR3) and applying per-channel latency matching. The policy is a Diffusion Policy with a fine-tuned DINOv3 backbone, selected via closed-loop rollouts in a physics simulator (Isaac Sim) reconstructed from depth maps. This selection process is a strong methodological choice, prioritizing closed-loop safety metrics (trocar violations, stage completion) over offline validation error, which the authors show to be anti-correlated with safety in this context.
The experimental validation is the paper's strongest feature. It moves beyond simulation and ex-vivo testing to in-vivo execution on live rabbits. The authors trained on 849 in-vivo demonstrations and deployed the policy on four additional live rabbits with electrosurgery armed. The policy completed the appendectomy in 3 out of 4 animals under shared autonomy. The safety metrics are detailed, including RCM error monitoring and electrosurgery gating. The comparison against video-based tracking (showing 10.7mm error vs 1.36mm for the logger) provides strong evidence for the hardware approach. The ablation study on policy configuration (31 candidates) is thorough and statistically grounded (Fisher's exact test).
High. The authors release both demonstration corpora (ex-vivo and in-vivo) and the code repository. The hardware design is described in sufficient detail (sensor models, mounting, firmware logic) for replication. The simulator setup (Isaac Sim, FEM parameters) is specified. The latency matching procedure is clearly defined.
The primary limitation is the reliance on the surgeon to manually select the surgical phase during deployment, as the vision-based phase predictor failed (24.6% agreement). This limits the autonomy claim. The study is limited to laparoscopic appendectomy in rabbits, so generalization to other procedures or species is not demonstrated. The robot (FR3) is still required for calibration and execution, so it is not a fully robot-free pipeline. The sample size for in-vivo deployment is small (4 animals).
This work has significant potential impact on surgical robotics by demonstrating that high-quality demonstration data can be collected from standard hand-held instruments, removing the need for expensive robot-based teleoperation for data collection. This could democratize the collection of surgical datasets and accelerate the development of surgical robot policies. The hardware logger is a practical, low-cost solution that could be widely adopted. The paper presents a robust end-to-end pipeline for learning bimanual surgical policies from hand-held instrument demonstrations, validated in vivo with electrosurgery. It makes a significant contribution to surgical robotics by introducing a novel, low-cost instrument-state logger and demonstrating that such data is sufficient to train safe, effective policies for live animal surgery, thereby bridging the gap between hand-held surgical practice and robotic automation.