Evaluate RAG outputs for faithfulness, hallucinations, and retrieval quality.
Copy the install command and let the AI configure it · recommended for beginners
No copy-paste install info for "groundcheck" yet — see the docs or source repo.
Use groundcheck to evaluate this RAG QA result, provide a faithfulness score, and identify any claims not supported by the retrieved context.
Returns a faithfulness score and highlights unsupported parts of the answer.
Use groundcheck to detect whether this RAG answer contains hallucinations, and briefly explain suspicious claims and why.
Outputs hallucination findings with suspicious claims and explanations.
Use groundcheck to compare retrieval quality metrics for two RAG retrieval result sets and decide which one better supports the final answer.
Provides retrieval metric comparisons and identifies the stronger setup.
Developers building a RAG-based knowledge assistant can use it to verify whether answers stay grounded in retrieved content and to detect hallucinations. This helps isolate whether issues come from generation or retrieval.
Researchers or data analysts comparing retrieval approaches can use it to measure retrieval quality metrics. It is useful for deciding which retrieval setup better supports strong answers.
When an AI agent needs to automatically review RAG output quality, this MCP server can perform the evaluation. It is designed to work through MCP sampling and requires no API keys.
It is an MCP server that lets AI agents evaluate RAG outputs. Its stated capabilities include faithfulness scoring, hallucination detection, and retrieval quality metrics.
Based on the provided description, it does not require API keys. It works using MCP sampling.
The provided material does not include installation or configuration steps. Please see the source repository for integration details.
Run read-only checks on MCP servers and return an A-F evaluation report.
Evaluate AI agent outputs for CI gates, regressions, and canary promotions.
Evaluate MCP retrieval servers for quality, coverage, and citation integrity.
Intelligent RAG tool that chooses between private knowledge and web search.
Evaluate AI safety classifier robustness against decomposition, obfuscation, and multi-agent attacks.
Verify factual claims with live sources, confidence scores, and cited verdicts.