An ephemeral diagnostic agent for Kubernetes clusters that provides real-time NVIDIA GPU hardware introspection via stdio transport, designed for AI-assisted troubleshooting by SREs debugging complex hardware failures. Uses NVIDIA NVML library to expose GPU inventory, health monitoring, and XID error analysis tools through kubectl debug sessions, eliminating the need for persistent DaemonSets or standing infrastructure. Operates in read-only mode by default with an optional operator mode for destructive operations, featuring graceful degradation when GPU hardware is unavailable.
Cognium trust score
88%
Tier
Verified
Composite of vulnerability cleanliness, spec conformance, provenance, stability, and usage signals — scanned and weighted by Cognium. Human and agent signals are tracked separately.
Last scanned 2026-09-28.
Returns 7 tools: search_skills, get_skill, list_leaderboard, get_trust_breakdown, resolve_composition, plus the ChatGPT-connector search and fetch. Every tool is annotated read-only.
Resolve this skill directly via MCP tools/call get_skill.