# llm-d-benchmarking-agent

> llm-d-benchmarking-agent — talbenamii-llm-d-benchmarking-agent. Use this tool when you need to benchmark and evaluate the performance of llm-d models in a secure and controlled environment. It solves problems related to model optimization, capacity planning, and results analysis by providing a chat-based interface to deploy and test llm-d stacks, and generates shareable reports. The tool takes plain-English goals as input and outputs detailed benchmarking results, making it ideal for use cases that require efficient and accurate model evaluation.

Canonical page: https://skillsregistry.net/skills/talbenamii-llm-d-benchmarking-agent  
JSON: https://api.skillsregistry.net/v1/skills/talbenamii-llm-d-benchmarking-agent

## Description

Chat-based assistant that benchmarks llm-d from a plain-English goal: it interviews you, deploys an llm-d stack if needed, drives the real llm-d-benchmark CLI in a security sandbox, and explains the results. Also a Kubernetes run orchestrator, results analyzer, capacity pre-flight, shareable reports, and an MCP server.

## Trust

- **Trust score (0–1):** 0.59
- **Verification tier:** scanned
- **Last scanned:** 2026-09-28

## Facts

- **Version:** 1.0.0
- **Skill type:** atomic
- **Execution layer:** container
- **Runtime environment:** vm
- **License:** Apache-2.0
- **Updated:** 2026-09-28

## Source

- **Source listing:** [GitHub](https://github.com/TalBenAmii/llm-d-benchmarking-agent)

## Use it

Resolve this record through the SkillsRegistry MCP server (no auth, read-only):

```
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp
```

```json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_skill",
    "arguments": {
      "slug": "talbenamii-llm-d-benchmarking-agent"
    }
  }
}
```

REST: `GET https://api.skillsregistry.net/v1/skills/talbenamii-llm-d-benchmarking-agent` · pull for local use: `GET https://api.skillsregistry.net/v1/skills/talbenamii-llm-d-benchmarking-agent/pull`

---
SkillsRegistry indexes agent skills from public registries and GitHub. Skills we have analysed are scanned with Circle-IR and scored on six dimensions; each listing states its scan coverage. More: https://skillsregistry.net/llms.txt
