retriEVAL
LLM evals as MCP tools: score outputs for faithfulness, relevancy, and hallucination.
The current version, v1.0.0, was published to the official MCP registry on 2026-08-24. It is distributed as the Docker image ghcr.io/hcarrillo001/retrieval-mcp:1.0.0, filed under AI & LLMs by this index, and declares 4 environment variables. Removals and unreachable sources are measured across every server this index tracks, contract drift across the servers that answer in consecutive snapshots; the index's current counts put this record in context.
Using retriEVAL in Claude, Cursor, Gemini CLI, Cline, or Zed?
MCP tool contracts can change remotely with no version bump. The mcpindex gate pins each contract and HOLDs the call when it drifts-before your agent acts. Zero credentials. This is not the package install for this server itself (use Install this server for that).
Rewrites your MCP host config so each server launches behind the gate. Inspect first: curl -fsSL https://mcpindex.ai/install.sh | less
uv tool install mcpindex-gate && mcpindex-config-wireVerdict not yet evaluated for this tool. The semantic screen takes adversarial cases first; coverage rolls out as the corpus expands (15/150 labels to graduation). The deterministic conformance probe is built but has not yet run on the public corpus, so a recorded verdict here is REVIEW or UNVERIFIED, never a clearing ALLOW. Until a verdict is recorded, an agent should treat this tool as not-yet-cleared and fall back to its own checks. Method: the eval, four-state verdict, honest limits.
Own this server? Screen its description →
That verdict was true at screening time (snapshot 2026-08-24).
Contracts can change after screening, with no version bump. The gate pins retriEVAL’s tool contracts on first sight and holds any silent change before your agent acts - the check that keeps being true on Tuesday.
See your first HOLD in 2 minutes →
Related: how to trust an MCP server · screen before install · silent contract drift
ANTHROPIC_API_KEYJudge model API key. Not needed if you set RETRIEVAL_JUDGE_BACKEND=ollama to run the judge locally.
RETRIEVAL_JUDGE_BACKENDWhich judge to score with: anthropic (default), openai, or ollama for a fully local judge.
RETRIEVAL_JUDGE_MODELJudge model id, e.g. claude-sonnet-4-6 or qwen2.5:14b for Ollama.
RETRIEVAL_BUDGET_USDSpend cap for judge calls. Unset means no cap.
BridgeNode — x402 pay-per-request AI inference. OpenAI-compatible API + MCP server, Solana USDC.
Persistent project context for Google Gemini. Python/FastMCP. IANA-registered .faf format.
Static worst-case token-budget analysis for LLM-agent workflows + signed budget certificates.