# agent-eval

> agent-eval — rudrendupaul-agent-eval. Use this tool when you need to statistically evaluate changes in LLM agent behavior, solving problems of uncertain performance improvements or regressions. It takes in agent performance data as input and outputs p-values, effect sizes, and confidence intervals to inform decision-making. Ideal for use cases where agent updates or modifications are made, and a rigorous assessment of their impact is required.

Canonical page: https://skillsregistry.net/skills/rudrendupaul-agent-eval  
JSON: https://api.skillsregistry.net/v1/skills/rudrendupaul-agent-eval

## Description

MCP server exposing statistical regression testing for LLM agents as a "run" tool: p-value, effect size, and confidence interval on whether agent behavior actually changed.

## Trust

- **Trust score (0–1):** 0.68
- **Verification tier:** scanned
- **Last scanned:** 2026-08-29

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** ai-ml
- **Updated:** 2026-08-29

## Source

- **Source listing:** [Glama](https://glama.ai/mcp/servers/xcuatrxlgl)
- **Repository:** <https://github.com/RudrenduPaul/agent-eval>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "rudrendupaul-agent-eval"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/rudrendupaul-agent-eval` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/rudrendupaul-agent-eval/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
