tan-yong-sheng/ai-vision-mcp

tan-yong-sheng/ai-vision-mcp

от tan-yong-sheng
MCP-сервер для ИИ-анализа изображений и видео через Google Gemini и Vertex AI. Поддерживает детекцию объектов, дизайн-аудит с WCAG-проверкой, сравнение снимков и анализ YouTube-роликов. Пригодится ...

AI Vision MCP Server

A powerful Model Context Protocol (MCP) server that provides AI-powered image and video analysis using Google Gemini and Vertex AI models.

Features

  • Dual Provider Support: Choose between Google Gemini API and Vertex AI
  • Multimodal Analysis: Support for both image and video content analysis
  • Flexible File Handling: Upload via multiple methods (URLs, local files, base64)
  • Storage Integration: Built-in Google Cloud Storage support
  • Comprehensive Validation: Zod-based data validation throughout
  • Error Handling: Robust error handling with retry logic and circuit breakers
  • TypeScript: Full TypeScript support with strict type checking

Quick Start

Pre-requisites

You could choose either to use google provider or vertex_ai provider. For simplicity, google provider is recommended.

Below are the environment variables you need to set based on your selected provider. (Note: It’s recommended to set the timeout configuration to more than 5 minutes for your MCP client).

(i) Using Google AI Studio Provider

Инструменты были проиндексированы:
analyze_image

Анализирует изображение с помощью AI-моделей компьютерного зрения. Поддерживает URL, данные в base64 и локальные пути к файлам.

Анализировать изображение

Параметры
  • imageSourcestringобязательный

    Image source - can be a URL, base64 data (data:image/...), or local file path

  • modeenum

    Analysis mode: general (default), palette (extract design tokens), hierarchy (analyze visual hierarchy), components (catalog UI components)

  • optionsobject
  • promptstringобязательный

    The prompt describing how you want to compare the images. If the task is front-end or UI comparison, the prompt you provide must be: "Compare the given screenshots and describe differences in layout structure, component arrangement, color scheme, typography, and visual hierarchy. Pay attention to common sections such as the navbar, header, footer, and main content areas to identify style or layout inconsistencies." + your additional requirements. For other tasks, the prompt you provide must clearly describe what to compare, identify, or analyze between the images.

analyze_video

Анализирует видео с помощью AI-моделей компьютерного зрения. Поддерживает URL и локальные пути к файлам.

Анализировать видео

Параметры
  • optionsobject
  • promptstringобязательный

    The prompt describing what you want to know about the video.

  • videoSourcestringобязательный

    Video source - can be a URL or local file path

compare_images

Сравнивает несколько изображений с помощью AI-моделей компьютерного зрения. Поддерживает URL-адреса, данные в формате base64 и локальные пути к файлам.

Сравнить изображения

Параметры
  • imageSourcesstring[]обязательный

    Array of image sources (URLs, base64 data, or file paths) - minimum 2 images. Maximum determined by MAX_IMAGES_FOR_COMPARISON environment variable (default: 4)

  • optionsobject
  • promptstringобязательный

    The prompt describing how you want to compare the images. If the task is front-end or UI consistency, the prompt you provide must specify what to evaluate — such as layout alignment, component structure, spacing, typography, color consistency, and visual hierarchy. Pay special attention to shared sections like the navbar, header, footer, and main content areas to identify layout shifts or inconsistent styles between versions. For other tasks, the prompt you provide must clearly describe what aspects to compare or analyze — such as visual differences, content changes, design variations, or quality degradation.

detect_objects_in_image

Обнаруживает объекты на изображении с помощью AI-моделей компьютерного зрения и создаёт размеченные изображения с ограничивающими рамками. Поддерживает URL, данные в base64 и локальные пути к файлам. Обработка файлов: явно указанный filePath → точный путь, иначе → временная директория. Использует оптимизированные параметры по умолчанию для обнаружения объектов.

Обнаружить объекты на изображении

Параметры
  • imageSourcestringобязательный

    Image source - can be a URL, base64 data (data:image/...), or local file path

  • outputFilePathstring

    Optional explicit output path for the annotated image. If provided, the image is saved to this exact path. Relative paths are resolved against the MCP server's current working directory.

  • promptstringобязательный

    Text prompt describing what to detect or recognize in the image. Avoid including any instructions about output structure or formatting — these are automatically managed by the workflow.

  • viewportHeightnumber

    Optional logical viewport height (for web screenshots). Used to distinguish between actual image dimensions and logical viewport size.

  • viewportWidthnumber

    Optional logical viewport width (for web screenshots). Used to distinguish between actual image dimensions and logical viewport size.

Похожие MCP-сервера

botmonster/image2svg-mcp

botmonster/image2svg-mcp

MCP-сервер для конвертации растровых изображений в SVG. Работает с PNG, JPG, WEBP, принимает base64 или URL, даёт полный контроль над параметрами векторизации. Идеален для автоматизации дизайна и п...

Python5
MKirovBG/scribefy-mcp

MKirovBG/scribefy-mcp

Сервер Scribefy для MCP — извлекает транскрипты YouTube по ссылке прямо в Claude Desktop, Cursor, Windsurf и других AI-клиентах. Также бесплатно ищет видео, получает метаданные и похожие ролики.

JavaScript1
strato-space/media-gen-mcp

strato-space/media-gen-mcp

Строго типизированный MCP сервер для генерации изображений (OpenAI gpt-image) и видео (Sora, Veo). Поддерживает редактирование, загрузку с URL и диска, сжатие. Интегрируется с любыми MCP-клиентами.

TypeScript9
Adityaaery20/media-mcp

Adityaaery20/media-mcp

MCP-сервер для обработки изображений и видео без ключей и конфигурации. Позволяет изменять размер, конвертировать форматы и сжимать файлы. Также доступны обрезка, фильтры и анализ. Видео требует ff...

Python4
kenneives/design-token-bridge-mcp

kenneives/design-token-bridge-mcp

MCP сервер для трансляции дизайн-токенов: извлекает из Tailwind, CSS, Figma или W3C DTCG и генерирует темы под Kotlin, SwiftUI, Tailwind и CSS-переменные. Полезен разработчикам и дизайнерам в пайплайне v0 → Figma → Claude Code.

TypeScript5
albertnahas/icogenie-mcp

albertnahas/icogenie-mcp

MCP сервер @icogenie/mcp дает AI-агентам генерировать SVG-иконки по тексту. Помогает разработчикам быстро создавать иконки для интерфейсов в Claude. Поддерживает пакетную работу, библиотеку и ежедневные кредиты.

TypeScript6
© Каталог MCP, 2026. Все права защищены.
Проект не аффилирован с Anthropic и любыми упомянутыми продуктами.
Все названия и торговые марки принадлежат их владельцам.
Контакты для связи: hi@mcp-katalog.ru

Лука Никитин