# mcp-llm-eval

> Use this tool when you need to evaluate and validate the performance of large language models (LLMs) against specific datasets, scoring responses and enforcing quality thresholds to ensure reliable AI outputs. It solves problems of model quality control and validation, providing a reusable CI/CD primitive for AI model evaluation. The tool takes in datasets and model inputs, producing scored responses and quality assessments as output.

Canonical page: https://skillsregistry.net/skills/berkayildi-mcp-llm-eval  
JSON: https://api.skillsregistry.net/v1/skills/berkayildi-mcp-llm-eval

## Description

A local MCP server that packages LLM evaluation gates as reusable CI/CD primitives, enabling AI agents to run datasets against models, score responses, and enforce quality thresholds.

## Trust

- **Trust score (0–1):** 0.44
- **Verification tier:** scanned
- **Last scanned:** 2026-09-03

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** mcp-remote
- **Runtime environment:** api
- **Category:** devops-ci
- **Updated:** 2026-09-03

## Source

- **Source listing:** [Glama](https://glama.ai/mcp/servers/n4lk43z5hk)
- **Repository:** <https://github.com/berkayildi/mcp-llm-eval>

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "berkayildi-mcp-llm-eval"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/berkayildi-mcp-llm-eval` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/berkayildi-mcp-llm-eval/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
