Let a local AI watch screens and automate native desktop actions.
Copy the install command and let the AI configure it · recommended for beginners
No copy-paste install info for "mcp-vision" yet — see the docs or source repo.
Observe the admin page currently open on my screen, click the "New User" button, then fill in the name "Zhang San", email "[email protected]", and role "Editor", and finally click Save. Use only local UI interactions.
The AI identifies buttons and fields from the visual layout, then completes the clicks, typing, and save action automatically.
Read the customer list from the spreadsheet on my screen and copy each row's name and phone number into the matching fields in the CRM app on the right; if a pop-up appears, close it before continuing.
The AI recognizes UI elements across desktop apps and performs the copy, switch, and paste workflow.
Open the settings area in the current test environment, click through the left-side menu, and verify each page loads correctly and shows its buttons; if you find errors or blank sections, record the page name and issue location.
The AI navigates the interface automatically and returns a report of broken pages and visibility issues.
Analyze screenshots, run OCR, and monitor vision workflows with local Ollama models.
Analyze screenshots, text, and UI mockups through one vision MCP tool.
Detect and analyze objects in images with zero-shot vision models.
Analyze images through one MCP tool using any OpenAI-compatible endpoint.
Turn screenshots and images into code, text, and diagnostic insights.
Let AI see and control a Linux desktop for visual task automation.