SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks Paper • 2602.12670 • Published Feb 13 • 66
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering? Paper • 2603.15401 • Published Mar 16 • 21
SkillSmith: Compiling Agent Skills into Boundary-Guided Runtime Interfaces Paper • 2605.15215 • Published May 12 • 1
SkVM: Compiling Skills for Efficient Execution Everywhere Paper • 2604.03088 • Published Apr 6 • 11
SWE-Pruner Pro: The Coder LLM Already Knows What to Prune Paper • 2607.18213 • Published Jul 20 • 47
SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents Paper • 2601.16746 • Published Jan 23 • 74
WebTestPilot: Agentic End-to-End Web Testing against Natural Language Specification by Inferring Oracles with Symbolized GUI Elements Paper • 2602.11724 • Published Feb 12 • 1
ARC: Compiling Hundreds of Requirement Scenarios into A Runnable Web System Paper • 2602.13723 • Published Mar 21 • 1
ARC: Compiling Hundreds of Requirement Scenarios into A Runnable Web System Paper • 2602.13723 • Published Mar 21 • 1
WebTestPilot: Agentic End-to-End Web Testing against Natural Language Specification by Inferring Oracles with Symbolized GUI Elements Paper • 2602.11724 • Published Feb 12 • 1