Let text-only LLMs understand images and videos through cloud vision models.
Copy the install command and let the AI configure it · recommended for beginners
No copy-paste install info for "vision-mcp" yet — see the docs or source repo.
Please read all text in this image and output it by paragraphs; if there is a table, preserve its structure as much as possible.
Returns OCR results from the image, including cleaned text or table content.
Please describe the main subjects in this image, what they are doing, and summarize the key information in the scene.
Returns a structured description and summary of the image content.
Please inspect this video, summarize its main content, and list recognizable key visual details and basic media information.
Returns a summary of the video content and available media information details.
Developers can use this MCP tool to connect text-only models with cloud vision models so they can understand and discuss images. It fits workflows that need multimodal understanding when the base model has no native vision support.
Researchers, product managers, or office workers can use it to run OCR on screenshots, scans, or photos and quickly obtain copyable text. It is especially useful when visual content needs to be turned into text for downstream work.
When a task involves video summarization, identifying visual elements, or grounding, this tool helps the model inspect videos and return relevant visual information. It is useful for content review, asset understanding, and multimedia-assisted analysis.
It is an MCP tool that gives text-only LLMs image and video understanding through cloud vision models. Its described capabilities include vision chat, OCR, grounding, and media information tools.
Based on the description, it supports images and videos and lets the model understand their content. For exact format support, see the source repository.
The description mentions a free GLM fallback, which indicates GLM can serve as a backup vision option in some cases. For exact behavior and limitations, see the source repository.
Give text-only LLMs vision support for analyzing local or online images.
Describe images, extract text, and run custom vision prompts on local files.
Enable text-only coding agents to analyze visual inputs as structured text.
Enable any LLM to describe images from paths, URLs, or base64.
Adds visual understanding to text-only LLMs with multi-provider fallback routing.
Use vision models via OpenRouter for OCR, image analysis, and object detection.