Analyze images and extract text with GLM-4.6V-Flash inputs.
Copy the install command and let the AI configure it · recommended for beginners
No copy-paste install info for "glm-vision-mcp" yet — see the docs or source repo.
Please read all text in this local screenshot and output it in the original paragraph order. If there are headings or tables, label them separately.
Returns OCR text from the screenshot while preserving the original structure as much as possible.
Analyze the main content of this document image, summarize the key information, and list important fields or conclusions shown in it.
Outputs a summary of the image content and a list of important information from the document.
Please read the content from this image URL, describe what is shown, and extract all visible text.
Returns a visual description and the extracted visible text.
Office workers or students can use it to recognize text from screenshots and document photos, reducing manual typing. It fits text extraction from screenshots, document images, and similar visuals.
Researchers or developers can ask the assistant to analyze image content and summarize key points, not just extract text. It can help interpret screenshots, document pages, or other images.
When images come from URLs, base64 data, or local files, this tool provides a unified way to analyze and extract text from them. It suits workflows that receive images from different sources.
It provides image understanding and OCR through GLM-4.6V-Flash. It can analyze image content and extract text from screenshots, documents, and similar images.
It supports URL, base64, and local file inputs. That means online images, local images, and encoded image data can all be used as sources.
The provided information does not include installation or configuration details. Please see the source repository for specifics.
Convert images into structured descriptions and OCR for text-only LLM understanding.
Analyze images in Claude Code with Zhipu AI's GLM-4.6V model.
Let text-only LLMs understand images and videos through cloud vision models.
Enable vision-less LLMs to understand screenshots and images through a vision proxy.
Give text-only LLMs vision support for analyzing local or online images.
Analyze images with vision APIs from URLs, local files, or base64 inputs.