Analyze screenshots and return descriptions or structured UI coordinates.
Copy the install command and let the AI configure it · recommended for beginners
No copy-paste install info for "visual-intelligence-mcp" yet — see the docs or source repo.
Please inspect this screenshot and describe the key UI elements and visible states.
A concise text description of the screenshot UI.
Analyze the current screenshot and output structured JSON coordinates for automated clicking.
Machine-readable coordinate data.
Based on this local screenshot, determine the page state and point out actionable areas.
Page state notes and actionable regions.
When an automation agent needs to “look at the screen,” it can send a local screenshot here. The tool helps determine the current UI state and support next actions.
Useful when you need to click, locate, or annotate UI elements. The tool can return structured JSON coordinates that programs can use directly.
When you only need to know what is on a screenshot or what the page shows, it can produce descriptive results and reduce manual inspection.
It is an MCP tool that routes “look at screen/screenshot” requests to a multimodal model via an API relay. It provides image recognition capabilities for Codex/Claude Code.
It can return a text description of the screenshot or structured JSON coordinates for UI automation.
It is suited for local screenshot analysis and UI automation agents that need to understand visible screen content. For more details, see the source repository.
Enable vision-less LLMs to understand screenshots and images through a vision proxy.
Analyze screenshots, extract text, and compare visuals with image understanding.
Analyze local images for coding agents with markdown and structured evidence.
Analyze screenshots, text, and UI mockups through one vision MCP tool.
Capture screens and webpages, then compare images for visual regression testing.
Analyze images with AI for OCR, scene description, detection, and comparison.