trackmcp
Back to directory
WindoC

gemini-ocr-mcp

View on GitHub

A FastMCP-based OCR server powered by Google Gemini. Handles both image file paths and base64‑encoded images to return extracted text. Easy to integrate into MCP workflows — just set your Gemini API key and model, run the server, and call `ocr_image_file` or `ocr_image_base64`

7 stars PythonOthers Updated Aug 5, 2026

Documentation

Gemini OCR MCP Server

This project provides a simple yet powerful OCR (Optical Character Recognition) service through a FastMCP server, leveraging the capabilities of the Google Gemini API. It allows you to extract text from images either by providing a file path or a base64 encoded string.

Objective

Extract the text from the following image:

CAPTCHA

and convert it to plain text, e.g., fbVk

Features

- File-based OCR: Extract text directly from an image file on your local system.

- Base64 OCR: Extract text from a base64 encoded image string.

- Easy to Use: Exposes OCR functionality as simple tools in an MCP server.

- Powered by Gemini: Utilizes Google's advanced Gemini models for high-accuracy text recognition.

Prerequisites

- Python 3.8 or higher

- A Google Gemini API Key. You can obtain one from Google AI Studio.

Setup and Installation

1. Clone the repository:

bash
git clone https://github.com/WindoC/gemini-ocr-mcp
    cd gemini-ocr-mcp

2. Create and activate a virtual environment:

bash
# Install uv standalone if needed

    ## On macOS and Linux.
    curl -LsSf https://astral.sh/uv/install.sh | sh
    
    ## On Windows.
    powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

3. Install the required dependencies:

bash
uv sync

MCP Configuration Example

If you are running this as a server for a parent MCP application, you can configure it in your main MCP `config.json`.

Windows Example:

json
{
  "mcpServers": {
    "gemini-ocr-mcp": {
      "command": "uv",
      "args": [
        "--directory",
        "x:\\path\\to\\your\\project\\gemini-ocr-mcp",
        "run",
        "gemini-ocr-mcp.py"
      ],
      "env": {
        "GEMINI_MODEL": "gemini-2.5-flash-preview-05-20",
        "GEMINI_API_KEY": "YOUR_GEMINI_API_KEY"
      }
    }
  }
}

Linux/macOS Example:

json
{
  "mcpServers": {
    "gemini-ocr-mcp": {
      "command": "uv",
      "args": [
        "--directory",
        "/path/to/your/project/gemini-ocr-mcp",
        "run",
        "gemini-ocr-mcp.py"
      ],
      "env": {
        "GEMINI_MODEL": "gemini-2.5-flash-preview-05-20",
        "GEMINI_API_KEY": "YOUR_GEMINI_API_KEY"
      }
    }
  }
}

Note: Remember to replace the placeholder paths with the absolute path to your project directory.

Tools Provided

`ocr_image_file`

Performs OCR on a local image file.

- Parameter: `image_file` (string): The absolute or relative path to the image file.

- Returns: (string) The extracted text from the image.

`ocr_image_base64`

Performs OCR on a base64 encoded image.

- Parameter: `base64_image` (string): The base64 encoded string of the image.

- Returns: (string) The extracted text from the image.

Frequently asked questions

What is gemini-ocr-mcp?

gemini-ocr-mcp is A FastMCP-based OCR server powered by Google Gemini. Handles both image file paths and base64‑encoded images to return extracted text. Easy to integrate into MCP workflows — just set your Gemini API key and model, run the server, and call `ocr_image_file` or `ocr_image_base64`

How do I install gemini-ocr-mcp?

Open the GitHub repository and follow its README. Most MCP servers are added to your client's MCP config, then called by your agent.

Is gemini-ocr-mcp open source?

Yes — it is hosted on GitHub at https://github.com/WindoC/gemini-ocr-mcp and has 7 stars.

Related MCP tools

Run your own MCP server? See who uses it and what to fix.

Measure it with TrackMCP