# AgentOps EvalBench MCP

> AgentOps EvalBench MCP — abhinavvarma02-agentops-evalbench-mcp. Use this tool when you need to evaluate and optimize the performance of large language models (LLMs) by assessing their answers for accuracy, reliability, and efficiency. It solves problems related to LLM testing, validation, and comparison by providing automated scoring and observability features. The tool accepts document uploads and test sets as inputs and outputs detailed evaluation metrics, making it ideal for use cases where LLM model performance and reliability are critical.

Canonical page: https://skillsregistry.net/skills/abhinavvarma02-agentops-evalbench-mcp  
JSON: https://api.skillsregistry.net/v1/skills/abhinavvarma02-agentops-evalbench-mcp

## Description

Enables LLM evaluation and observability by uploading documents, building test sets, running RAG pipelines, and automatically scoring answers for groundedness, hallucination risk, retrieval quality, latency, and cost, with tools exposed to MCP-compatible clients.

## Trust

- **Trust score (0–1):** 0.58
- **Verification tier:** scanned
- **Last scanned:** 2026-08-30

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** devops-ci
- **Updated:** 2026-08-30

## Source

- **Source listing:** [Glama](https://glama.ai/mcp/servers/c4jhmjnz3m)
- **Repository:** <https://github.com/AbhinavVarma02/Agentops-Evalbench-MCP>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "abhinavvarma02-agentops-evalbench-mcp"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/abhinavvarma02-agentops-evalbench-mcp` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/abhinavvarma02-agentops-evalbench-mcp/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
