mirror of
https://github.com/agentskills/agentskills.git
synced 2026-06-18 15:54:06 +08:00
863f7c28570200de824fe172ad4b7c9683b76796
A how-to guide for evaluating skill output quality using structured evals. Covers the full eval workflow: designing test cases, running with-skill vs. baseline comparisons, writing assertions, LLM-based grading, aggregating benchmarks, analyzing patterns, human review, and LLM-driven iterative improvement. Derived from the workflow implemented by the `skill-creator` Skill, but written as a standalone guide that readers can follow without using that tool. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Agent Skills
Agent Skills are a simple, open format for giving agents new capabilities and expertise.
Skills are folders of instructions, scripts, and resources that agents can discover and use to perform better at specific tasks. Write once, use everywhere.
Getting Started
- Documentation - Guides and tutorials
- Specification - Format details
- Example Skills - See what's possible
This repo contains the specification, documentation, and reference SDK. Also see a list of example skills here.
About
Agent Skills is an open format maintained by Anthropic and open to contributions from the community.
License
Code in this repository is licensed under Apache 2.0. Documentation is licensed under CC-BY-4.0. See individual directories for details.
Languages
Python
99.1%
Shell
0.9%