# agent-evaluation

> agent-evaluation — rustyorb-agent-evaluation. Use this tool when you need to assess the performance and reliability of Large Language Model (LLM) agents, including evaluating their behavioral responses and capabilities. It provides benchmarking and testing functionalities to identify areas of improvement and measure reliability metrics. Ideal for use cases where LLM agent evaluation and optimization are crucial, such as in AI model development and deployment.

Canonical page: https://skillsregistry.net/skills/rustyorb-agent-evaluation  
JSON: https://api.skillsregistry.net/v1/skills/rustyorb-agent-evaluation

## Description

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics.

## Trust

- **Trust score (0–1):** 1.00
- **Verification tier:** scanned
- **Last scanned:** 2026-09-19

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** instructions
- **Runtime environment:** llm
- **Category:** ai-ml
- **Updated:** 2026-09-19

## Source

- **Source listing:** [ClawHub](https://clawskills.sh/skills/rustyorb-agent-evaluation)

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "rustyorb-agent-evaluation"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/rustyorb-agent-evaluation` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/rustyorb-agent-evaluation/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
