Machine Learning Papers

Last 7 Days (September 05 – September 11, 2026)

← Previous Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 86)
Ivan Moshkov, Stephen Ge, George Armstrong ... · NVIDIA · arXiv
We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised fine-tuning and reinforcement learning, and evaluate checkp...
#2 TOP PAPER (Score: 82)
Zili Wang, Zhaopeng Qiu, Yuekai Zhang ... · NVIDIA · MLSys 2025
Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) post-training. Online co-training can further increase the draft's accuracy, yielding greater speedups. However, scaling this approach to co-training on large models with ...
#3 TOP PAPER (Score: 81)
Aashiq Muhamed, Virginia Smith · Carnegie Mellon University · arXiv
Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work ...