OpenRouter MCP Multimodal
The MCP server for multimodal AI agents.
One install · 14 tools · 300+ OpenRouter models · text, vision, audio & video — analysis and generation.
Quick start · Tools · Examples · Security · Development · FAQ
What is this?
OpenRouter MCP Multimodal is a production-grade Model Context Protocol (MCP) server — listed on the official MCP Registry as io.github.stabgan/openrouter-multimodal. It connects AI coding agents (Cursor, Claude Desktop, VS Code, Windsurf, Cline, and others) to OpenRouter's unified LLM API over stdio.
Unlike text-only MCP servers, one install covers the full multimodal surface:
| Capability | Tools | Highlights |
|---|---|---|
| Chat | chat_completion |
300+ models, :nitro / :exacto suffixes, provider routing, web search, response caching, reasoning tokens |
| Vision | analyze_image, generate_image |
OCR, captioning, VQA, image generation with reference inputs |
| Audio | analyze_audio, generate_audio |
Transcription, speech/music generation |
| Video | analyze_video, generate_video, generate_video_from_image, get_video_status |
Clip understanding, Veo / Sora / Seedance / Wan generation with progress notifications |
| Catalog | search_models, get_model_info, validate_model, rerank_documents, health_check |
Model discovery, validation, reranking, ops health |






