# llm-eval-mcp

> llm-eval-mcp — cdywolf-llm-eval-mcp. Use this tool when you need to evaluate the reliability of large language models (LLMs) and assess their performance under various conditions. It provides adversarial task generation, automated judging, and confidence statistics to help identify potential weaknesses and biases in LLMs. Ideal for use cases where LLM reliability is critical, such as high-stakes decision-making or sensitive information processing.

Canonical page: https://skillsregistry.net/skills/cdywolf-llm-eval-mcp  
JSON: https://api.skillsregistry.net/v1/skills/cdywolf-llm-eval-mcp

## Description

MCP server that provides tools for evaluating LLM agent reliability, including adversarial task generation, automated LLM-as-judge assessment, and confidence statistics.

## Trust

- **Trust score (0–1):** 0.70
- **Verification tier:** verified
- **Last scanned:** 2026-09-03

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** ai-ml
- **Updated:** 2026-09-03

## Source

- **Source listing:** [Glama](https://glama.ai/mcp/servers/fxhhmmlzgn)
- **Repository:** <https://github.com/cdywolf/llm-eval-mcp>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "cdywolf-llm-eval-mcp"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/cdywolf-llm-eval-mcp` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/cdywolf-llm-eval-mcp/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
