# agentbench

> agentbench — exe215-agentbench. Use this tool when you need to evaluate the performance of your OpenClaw agent across a variety of real-world tasks, solving problems such as assessing agent capabilities and identifying areas for improvement. The agentbench tool takes an OpenClaw agent as input and outputs benchmarking results across 40 tasks, providing a comprehensive evaluation of the agent's abilities. This tool is ideal for developers and researchers looking to optimize their agent's performance in diverse scenarios.

Canonical page: https://skillsregistry.net/skills/exe215-agentbench  
JSON: https://api.skillsregistry.net/v1/skills/exe215-agentbench

## Description

Benchmark your OpenClaw agent across 40 real-world tasks.

## Trust

- **Trust score (0–1):** 1.00
- **Verification tier:** scanned
- **Last scanned:** 2026-09-19

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** instructions
- **Runtime environment:** llm
- **Category:** productivity
- **Updated:** 2026-09-19

## Source

- **Source listing:** [ClawHub](https://clawskills.sh/skills/exe215-agentbench)

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "exe215-agentbench"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/exe215-agentbench` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/exe215-agentbench/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
