An Empirical Study of Agent Skills' Downstream Utility
Agent Skills package procedural guidance and resources for reuse, but a relevant Skill does not necessarily improve task performance.
ProofPaper ↗
Key points
- Existing studies characterize Skill content and evaluate downstream performance, yet provide limited explanations of how utility depends on content, execution configuration, and multi-Skill organization.
- We conduct an empirical study on 87 SkillsBench tasks, defining downstream utility as the pass-rate difference from No-Skill on the same tasks under the same model--harness configuration.
- We compare the same Skills across nine configurations, then examine alternative published Skills and organizations of fixed Skill sets under three selected configurations.
- We derive 17 authoring practices linking executable procedures to recovery, preservation of task requirements, and checks on final artifacts.
Sources (1)
- [1]An Empirical Study of Agent Skills' Downstream UtilityarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 06:05 AM
Agent Skills package procedural guidance and resources for reuse, but a relevant Skill does not necessarily improve task performance.
Existing studies characterize Skill content and evaluate downstream performance, yet provide limited explanations of how utility depends on content, execution configuration, and multi-Skill organization.
Extractive summary: sentences quoted from the sources.