arxiv:2605.09959
Jiaxin Huang
teapot123
AI & ML interests
None yet
Recent Activity
upvoted a paper 1 day ago
Negative Self-Distillation: Learning to Reason by Avoiding Flaws upvoted a paper 25 days ago
EnvHarness: Awakening Static Worlds for Agent Learning upvoted a paper 4 months ago
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories