DINO-X
This DINO-X MCP server by IDEA Research provides AI agents with real-world visual perception capabilities through the DINO-X computer vision API, enabling object detection, localization, human pose estimation, and image captioning. The implementation offers three core detection modes: text-prompted object detection for finding specific items, universal object detection for comprehensive scene analysis, and human pose keypoint detection with 17-point skeletal tracking, all with optional detailed descriptions and visualization capabilities that save annotated images with bounding boxes and labels. Built with TypeScript and the canvas library for image processing, it integrates with DeepDataSpace's hosted DINO-X API and includes robust error handling, task polling for asynchronous processing, and flexible output formatting, making it valuable for AI applications requiring visual understanding, accessibility tools, security monitoring, sports analysis, and automated content moderation workflows.
Composite of vulnerability cleanliness, spec conformance, provenance, stability, and usage signals — scanned and weighted by Cognium. Human and agent signals are tracked separately.
View full trust & usage report →Metadata
- Version
- 1.0.0
- Skill type
- atomic
- Execution layer
- mcp-remote
- Category
- media
- Source
- PulseMCP
- Repository
- github.com/idea-research/dino-x-mcp
- Author type
- human
- Updated
- 2026-05-18
Use via MCP
Resolve DINO-X from your agent
Streamable HTTP transport at https://api.skillsregistry.net/mcp. No auth for read tools. Discovery: .well-known/mcp.json.
One command in your shell — Claude Code wires it up and verifies the connection. Run /mcp in any session to confirm.
claude mcp add --transport http --scope user skillsregistry https://api.skillsregistry.net/mcp --scope user for --scope project to commit it to .mcp.json.