Commit Graph

2 Commits

  • fix: align report format across evaluate.py, agent spec, and template
    - evaluate.py: add CRITICAL ISSUES (axes ≤ 2) section, VERDICT line
    - agent-evaluator.md: match format_report output exactly (title, evidence markers, bar graphs)
    - templates/evaluation-report.md: match evaluate.py output format
    - All now produce identical AGENT SELF-EVALUATION REPORT structure
    
    Single authoritative format: evaluate.py's format_report() output.
  • feat(skills,agents): add agent-self-evaluation skill and agent-evaluator persona
    Add structured 5-axis self-evaluation framework for agent output quality:
    - Accuracy, Completeness, Clarity, Actionability, Conciseness
    - Evidence-based scoring with concrete improvement suggestions
    - Standalone Python evaluator script with keyword heuristics
    - Detailed scoring anchors reference guide
    - High-score and low-score annotated examples
    - Reusable evaluation report template
    - Optional hook integration for session-stop evaluation
    
    Agent persona (agent-evaluator) provides a dedicated subagent
    for applying the rubric to agent output with tool-backed verification.
    
    All files tested: Python script runs, examples score correctly
    (high 4.2, low 3.4), frontmatter parses clean, 183 lines (under 500).