← Back to Blog

How to Run an AI OCR Document Reader Agent with OE Runtime

Extract text from scanned documents, invoices, receipts, and forms automatically. OE Runtime uses Azure Vision to run OCR, classifies the document type, and structures the extracted data into ready-to-use fields — no manual data entry required.

👁️
Step 1
Extract Text
Run OCR on the image and extract all visible text preserving reading order and layout.
Step 2
Classify & Structure
Identify the document type and extract structured fields like vendor, date, line items, or totals.
Step 3
Report
Return document type, extracted fields, full raw text, low-confidence regions, and suggested next action.

What You Need


Create the Project Folder

Create a folder called ocr-vision/ and add these two files:

ocr-vision/
├── SKILL.md         # the portable skill (agentskills.io)
├── agent.yaml       # wires SKILL.md to your connector
└── oe-config.json  # LLM key + connector credentials

The Skill Files

Create SKILL.md inside an ocr-vision/ directory. This is the portable skill — it follows the agentskills.io format and runs unchanged on Claude, Cursor, Windsurf, or OE Runtime:

---
name: ocr-vision
description: Extract and analyze text from images using OCR
license: Apache-2.0
metadata:
  author: Open Enthrium
  version: "1.0"
---

You are a document analysis agent. Extract text from images using OCR,
then analyze, classify, and summarize the extracted content.
Complete all steps fully before writing your report.

## Step 1: Extract Text
Call the Azure Vision connector to run OCR on the provided image.
POST /read/analyze with body:
{
  "url": "<image URL to analyze>"
}
Save the Operation-Location URL from the response headers.
Then poll GET <operation-path> (the path portion of the Operation-Location URL)
every 2 seconds until status is "succeeded".
Extract all text lines from the analyzeResult.readResults.

## Step 2: Classify and Structure
Analyze the extracted text to determine the document type
(invoice, receipt, contract, ID, form, letter, or other).
Extract key structured fields based on type — e.g. for an invoice: vendor, date, line items, total.

## Step 3: Report
Produce a document analysis result:
- Document type identified
- Structured fields extracted with values
- Full raw text extracted
- Any low-confidence regions or unreadable sections
- Suggested next action (e.g. file, process payment, follow up)

Create agent.yaml in the same directory to wire the skill to your connector:

name: Document Reader
description: Extract and analyze text from images using OCR
connectors:
  - connection_name: Azure Vision
    connection_type: azure-vision
skills:
  - path: ./
    trigger_type: auto

The Config File

Create oe-config.json in the same directory:

{
  "llm": {
    "provider": "openai",
    "model": "gpt-4o",
    "apiKey": "YOUR_OPENAI_API_KEY"
  },
  "server": {
    "enabled": false,
    "port": 3333,
    "apiKey": "your-secret-api-key"
  },
  "connectors": [
    {
      "connection_name": "Azure Vision",
      "connection_type": "azure-vision",
      "apiKey": "YOUR_AZURE_VISION_KEY",
      "headerName": "Ocp-Apim-Subscription-Key",
      "baseUrl": "https://YOUR_RESOURCE.cognitiveservices.azure.com/vision/v3.2"
    }
  ]
}

Create an Azure Computer Vision resource at portal.azure.com and copy the endpoint and API key from the resource's Keys and Endpoint page. The free tier (F0) allows up to 5,000 OCR transactions per month.

Download OE Runtime

OE Runtime — Direct Downloads

Run the Agent

From the parent folder containing your skill directory:

MethodBest forDownload
1 npx recommended No install needed — always runs the latest version —
2 Windows .exe Download once, run offline on Windows ⊞ Windows (.exe)
3 macOS binary Download once, run offline on Mac  macOS
4 Linux binary Server deployments, cron jobs, Docker 🐧 Linux
5 API Server integration Call from any app, webhook, or automation pipeline 📮 Postman Collection

1 npx recommended

npx -y @openenthrium/oe-runtime@latest ./ocr-vision

2 Windows

oe-runtime-win.exe ./ocr-vision

3 macOS

chmod +x oe-runtime-macos
./oe-runtime-macos ./ocr-vision

First run blocked? System Settings → Privacy & Security → Allow Anyway.

4 Linux

chmod +x oe-runtime-linux
./oe-runtime-linux ./ocr-vision

5 API Server integration

Add a "server" block to oe-config.json, then start with --serve:

{
  "llm": { ... },
  "server": { "enabled": true, "port": 3333, "apiKey": "your-secret-key" },
  "connectors": [ ... ]
}
npx -y @openenthrium/oe-runtime@latest --serve --config oe-config.json

Run with inline YAML:

curl -X POST http://localhost:3333/run \
  -H "Content-Type: application/json" \
  -H "X-API-Key: your-secret-key" \
  -d '{"yaml": "...", "params": {}}'

Or run from a file on the server:

curl -X POST http://localhost:3333/run-file \
  -H "Content-Type: application/json" \
  -H "X-API-Key: your-secret-key" \
  -d '{"file": "/path/to/agent.yaml", "params": {}}'

Use Cases

Build your own agents with OE Runtime

Download OE Runtime and run any AI agent locally or as a server — no cloud required.

Get OE Runtime →
Series OE Runtime Agent Guides — 21 Connectors
Series overview →