InferBench's MCP server lets coding agents run, serve and benchmark local LLMs (text + image, llama.cpp + Stable Diffusion) on your own hardware on demand — measuring real tokens/sec and picking the optimal quant for your GPU from a 124-model catalog. Local-first, no cloud required.
Cognium trust score
29%
Tier
Scanned
Composite of vulnerability cleanliness, spec conformance, provenance, stability, and usage signals — scanned and weighted by Cognium. Human and agent signals are tracked separately.
Last scanned 2026-08-31.
Returns 7 tools: search_skills, get_skill, list_leaderboard, get_trust_breakdown, resolve_composition, plus the ChatGPT-connector search and fetch. Every tool is annotated read-only.
Resolve this skill directly via MCP tools/call get_skill.