Machine Learning Papers

Last 7 Days (August 25 – August 31, 2026)

← Previous Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 88)
Qinglin Ye, Zhiyuan Gu, Jingjie Xia ... · University of Chinese Academy of Sciences +8 · arXiv
Search-augmented reasoning remains difficult for small language models. On-policy distillation (OPD) from trained teachers offers a promising direction, but suffers from two issues: (1) high-quality multi-turn search trajectories depend on dynamic retriever responses, making SFT ...
#2 TOP PAPER (Score: 88)
Jiaming Zhou, Qihang Zhang, Gangwei Xu ... · Ant Group (Robby Ant Research) · arXiv
Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task can be performed simply by specifying it in the context, without any parameter upd...
#3 TOP PAPER (Score: 85)
Yixin Tao, Weiqiang Zheng · Shanghai University of Finance and Economics +1 · arXiv (preprint)
We settle the minimax-optimal alternating regret, a regret notion motivated by alternating learning dynamics in games, for both online linear optimization (OLO) and online convex optimization (OCO). For OLO over the probability simplex $Δ_d$, we give an algorithm with $O(\log d...