ResearchPod Summary
As LLM agents increasingly rely on 'Agent Skills'—packaged as SKILL.md files—to extend their capabilities, the assumption that these skills are inherently reusable has come under scrutiny. This paper investigates the quality of the current ecosystem by analyzing 138,133 public SKILL.md files across 20,556 repositories. The authors developed a two-tier taxonomy of 31 distinct defects, ranging from official specification violations (e.g., missing routing metadata) to best-practice failures (e.g., hardcoded credentials or platform-specific dependencies). The study evaluates the prevalence of these defects and tests their functional impact on skill retrieval using a deterministic stress test.
The study finds that the vast majority of public skills are not optimized for reuse. Specifically, 91.8% of the analyzed files contain at least one detected defect, with 89.3% violating official specification requirements. The most common issues are 'ordinary' packaging problems rather than complex security attacks: weak routing metadata, bloated instruction bodies, and poor organization of supporting resources.
The functional impact of these defects is clear: skills with valid routing metadata are retrieved significantly more reliably than those with routing defects. Furthermore, the authors observe that skills marked as AI-generated often exhibit higher rates of safety and portability issues, suggesting that current automated generation workflows are not sufficiently quality-assured. The paper concludes by proposing a quality-assured generation workflow that integrates spec-aware prompting, automated linting, and safety gating.
This research demonstrates that the 'reusability' of agent skills is currently more of a promise than a reality. For developers and researchers, these findings highlight that simply sharing code is insufficient; without adherence to standardized packaging and metadata requirements, skills remain brittle and difficult to integrate into broader agent ecosystems. The provided taxonomy and evidence-based guidelines offer a roadmap for improving the reliability of the agent-skill supply chain.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.