evals-skills

evals-skills

Guides AI coding agents in building product-specific AI evaluations.

Agent SkillDevOpen source
Type
Agent Skill
Open source
Yes
GitHub Stars
★ 1.2k
Source
skill-github

Overview

evals-skills is a set of skills designed to help AI coding agents construct AI evaluations tailored to specific products. These skills assist in avoiding common evaluation pitfalls and provide support across various stages, from data review to generating synthetic data. After installation via npx, users can specify one or more skills to aid in evaluation tasks. Ideal for developers who need detailed testing and optimization of AI models.

Capabilities

  • ▪Route to the appropriate evaluation skill
  • ▪Audit evaluation pipelines and identify potential issues
  • ▪Intelligently select samples for error analysis
  • ▪Generate diverse synthetic test inputs
  • ▪Design LLM-based subjective quality evaluators
  • ▪Calibrate LLM evaluators with human labels

Use cases

AI model performance evaluationError pattern identificationSynthetic data generationEvaluating retrieval-augmented generation (RAG) quality

Setup

Requires: API KeyNode环境
Install using npx: `npx skills add https://github.com/ai-evals-course/evals-skills` or install a specific skill individually: `npx skills add https://github.com/ai-evals-course/evals-skills --skill error-discovery`

This information was compiled by AI from public sources and may contain inaccuracies — please refer to the source.

FAQ

How do I get started with these skills?

First run `npx skills add https://github.com/ai-evals-course/evals-skills` to install all skills, or specify a single skill to install.

Which skill is most important?

`error-discovery` is the most critical skill, used to detect error patterns in data.

Related skills