Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It Paper • 2610.03195 • Published 4 days ago • 22
Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models Paper • 2609.33355 • Published 9 days ago • 49
World Pilot: Steering Vision-Language-Action Models with World-Action Priors Paper • 2606.12403 • Published Jun 10 • 27
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Paper • 2606.12191 • Published Jun 10 • 73
Human Psychometric Questionnaires Mischaracterize LLM Behavior Paper • 2509.10078 • Published May 29 • 36
SWE-Explore: Benchmarking How Coding Agents Explore Repositories Paper • 2606.07297 • Published Jun 5 • 124
Human Psychometric Questionnaires Mischaracterize LLM Behavior Paper • 2509.10078 • Published May 29 • 36
Human Psychometric Questionnaires Mischaracterize LLM Behavior Paper • 2509.10078 • Published May 29 • 36
SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations Paper • 2606.05563 • Published Jun 4 • 56
ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time? Paper • 2606.05553 • Published Jun 4 • 50
RobotValues: Evaluating Household Robots When Human Values Conflict Paper • 2606.03312 • Published Jun 2 • 26
ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time? Paper • 2606.05553 • Published Jun 4 • 50