Create verifiable evaluation records through a draft, review, revise, submit workflow.
Copy the install command and let the AI configure it · recommended for beginners
Please install the "io.github.cyanheads/evals-mcp-server" MCP server from askskill: Run: claude mcp add 'io-github-cyanheads-evals-mcp-server' -- node @cyanheads/evals-mcp-server
Create a draft evaluation record for this LLM response quality test, including goals, sample description, scoring rubric, assigned graders, and submission checklist.
A structured evaluation draft with workflow-ready fields and assigned reviewers.
Revise this evaluation record based on the following review comments: add failure cases, standardize scoring criteria, clarify each grader's responsibilities, and provide revision notes.
An updated evaluation record with revision notes explaining what changed and why.
Check whether this evaluation record meets submission requirements: confirm draft, review, and revision steps are complete, verify required grader decisions are present, and list any missing items.
A pre-submission check result stating readiness, missing items, and recommended fixes.
Evaluate AI agent outputs for CI gates, regressions, and canary promotions.
Review MCP servers for quality, security, scores, and improvement plans.
Score, assess, and compare MCP servers before you decide to trust them.
Review code brutally honestly with scores, real issues, and actionable fixes.
Inspect Minecraft project evidence before writing development code.
Enable AI to read and write Obsidian vault content through Team Relay.