# genpark-ai-evaluation-custom-benchmark-runner-skill

> genpark-ai-evaluation-custom-benchmark-runner-skill — alphaparkinc-genpark-ai-evaluation-custom-benchmark-runner-skill. Use this tool when you need to evaluate and benchmark AI models on real-world tasks or custom datasets. It solves problems of assessing AI performance and comparing models on specific use cases, providing outputs such as evaluation metrics and benchmark results. It takes in custom benchmarks and AI models as inputs, and is ideal for use cases where standardized evaluation metrics are insufficient.

Canonical page: https://skillsregistry.net/skills/alphaparkinc-genpark-ai-evaluation-custom-benchmark-runner-skill  
JSON: https://api.skillsregistry.net/v1/skills/alphaparkinc-genpark-ai-evaluation-custom-benchmark-runner-skill

## Description

Real-world AI task evaluation & custom benchmark runner (oqoqo style)

## Trust

- **Trust score (0–1):** 1.00
- **Verification tier:** verified
- **Last scanned:** 2026-09-28

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** container
- **Runtime environment:** vm
- **Updated:** 2026-09-28

## Source

- **Source listing:** [GitHub](https://github.com/alphaparkinc/genpark-ai-evaluation-custom-benchmark-runner-skill)

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "alphaparkinc-genpark-ai-evaluation-custom-benchmark-runner-skill"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/alphaparkinc-genpark-ai-evaluation-custom-benchmark-runner-skill` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/alphaparkinc-genpark-ai-evaluation-custom-benchmark-runner-skill/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
