llm-eval-search — fvahedian-llm-eval-search. Use this tool when you need to evaluate the performance of large language models (LLMs) using a search-based approach, solving problems related to model benchmarking and comparison. It takes in model configurations and evaluation metrics as inputs and outputs performance scores and rankings. Ideal for use cases where accurate model assessment is crucial, such as in natural language processing and machine learning applications.
Cognium trust score
75%
Tier
Verified
Composite of vulnerability cleanliness, spec conformance, provenance, stability, and usage signals — scanned and weighted by Cognium. Human and agent signals are tracked separately.
Last scanned 2026-09-28.
Returns 7 tools: search_skills, get_skill, list_leaderboard, get_trust_breakdown, resolve_composition, plus the ChatGPT-connector search and fetch. Every tool is annotated read-only.
Resolve this skill directly via MCP tools/call get_skill.