Add local photo understanding, OCR, comparison, and metadata extraction to coding assistants.
Copy the install command and let the AI configure it · recommended for beginners
No copy-paste install info for "photo-vlm-mcp" yet — see the docs or source repo.
Use photo-vlm-mcp to read this app screenshot, extract all visible text, and group it into headings, buttons, and body text.
Structured OCR output for organizing UI copy or troubleshooting issues.
Use photo-vlm-mcp to compare these two webpage screenshots, describe differences in layout, copy, button states, and visual elements, and list possible regression issues.
A clear difference list with potential issues, suitable for testing and review.
Use photo-vlm-mcp to analyze this photo’s scene, identify major objects and environment features, and extract any available image metadata.
Scene description, object recognition results, and a metadata summary for archiving or downstream processing.
Enable non-vision AI clients to analyze images with local Ollama vision models.
Analyze images with AI for OCR, scene description, detection, and comparison.
Enable text-only AI to analyze images offline with local Ollama models.
Analyze screenshots, run OCR, and monitor vision workflows with local Ollama models.
Detect and analyze objects in images with zero-shot vision models.
Analyze videos with frame extraction, scene detection, and metadata retrieval.