Audit LLM-as-judge results for drift, bias, and human agreement.
Copy the install command and let the AI configure it · recommended for beginners
No copy-paste install info for "judge-audit-mcp" yet — see the docs or source repo.
Audit these LLM-as-judge evaluation runs, compare score changes across batches, identify possible judge drift, and summarize the most obvious drift cases.
A cross-run score difference analysis highlighting samples or batches that may show judge drift.
Use controlled probes to audit bias in this judge model and analyze whether it shows systematic preference for certain prompt styles, answer lengths, or phrasing.
A bias measurement report describing possible systematic preferences and how they appear.
Compare these LLM judge results with human ratings, identify high-disagreement samples, and summarize the main differences between model and human judgments.
A judge-versus-human agreement analysis plus a list of high-disagreement cases.
Researchers or developers can use it to audit an LLM judge across repeated evaluation runs and check for score drift. This helps detect unstable evaluation behavior early.
When a team suspects the judge model may favor certain answer formats, they can use controlled probes to measure bias. It is useful for checking whether scores are influenced by irrelevant factors.
When building an automated evaluation workflow, teams can compare LLM judge outputs with human ratings. This helps determine whether the model judge is close enough to human standards.
This is an MCP server for auditing LLM-as-judge evaluations. It can detect judge drift across runs, measure bias with controlled probes, and compare model judge agreement with human raters.
Based on the given description, it focuses on three issues: judge drift, judge bias, and agreement between model judging and human ratings. If you need specific metrics or methods, see the source repository.
The provided material only says it is an MCP server and does not include installation steps, runtime requirements, or key requirements. See the source repository for details.
Audit financial data quality and evaluate AI inference, bias, and KPIs.
Audit app store metadata and refine keyword fields inside MCP clients.
Audit liability clauses against company standards with evidence-based verdicts or abstention.
Audit MCP servers for conformance and generate scores with Markdown reports.
Audit AI apps or model setups to find risks and improvements.
Scan configured MCP servers locally and generate inventory with risk scores.