# evalbench

> evalbench — cognis-digital-evalbench. Use this tool when you need to evaluate the performance of large language models (LLMs) or agents in an offline setting, with capabilities to track regressions and ensure model quality. It provides a harness for testing and validation, accepting model inputs and producing evaluation outputs. Ideal for use cases where model performance and reliability are critical, such as in machine learning development and deployment pipelines.

Canonical page: https://skillsregistry.net/skills/cognis-digital-evalbench  
JSON: https://api.skillsregistry.net/v1/skills/cognis-digital-evalbench

## Description

Offline LLM / agent eval harness with regression gates

## Trust

- **Trust score (0–1):** 0.50
- **Verification tier:** unverified

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** container
- **Runtime environment:** vm
- **Updated:** 2026-09-21

## Source

- **Source listing:** [GitHub](https://github.com/cognis-digital/evalbench)

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "cognis-digital-evalbench"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/cognis-digital-evalbench` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/cognis-digital-evalbench/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
