Runs comprehensive benchmarks against local LLM models via Ollama or LM Studio, measuring tokens per second, time to first token, memory usage, and quality across reasoning, coding, math, and multilingual categories. Produces a global score (0-100) combining hardware fit and quality metrics, with a verdict ranging from Excellent to Not Recommended. Results can be shared to a public leaderboard at metrillm.dev for cross-hardware comparison.
Cognium trust score
50%
Tier
Unverified
Composite of vulnerability cleanliness, spec conformance, provenance, stability, and usage signals — scanned and weighted by Cognium. Human and agent signals are tracked separately.
Returns 7 tools: search_skills, get_skill, list_leaderboard, get_trust_breakdown, resolve_composition, plus the ChatGPT-connector search and fetch. Every tool is annotated read-only.
Resolve this skill directly via MCP tools/call get_skill.