Extract text from scanned documents, invoices, receipts, and forms automatically. OE Runtime uses Azure Vision to run OCR, classifies the document type, and structures the extracted data into ready-to-use fields — no manual data entry required.
Create a folder called ocr-vision/ and add these two files:
Create SKILL.md inside an ocr-vision/ directory. This is the portable skill — it follows the agentskills.io format and runs unchanged on Claude, Cursor, Windsurf, or OE Runtime:
--- name: ocr-vision description: Extract and analyze text from images using OCR license: Apache-2.0 metadata: author: Open Enthrium version: "1.0" --- You are a document analysis agent. Extract text from images using OCR, then analyze, classify, and summarize the extracted content. Complete all steps fully before writing your report. ## Step 1: Extract Text Call the Azure Vision connector to run OCR on the provided image. POST /read/analyze with body: { "url": "<image URL to analyze>" } Save the Operation-Location URL from the response headers. Then poll GET <operation-path> (the path portion of the Operation-Location URL) every 2 seconds until status is "succeeded". Extract all text lines from the analyzeResult.readResults. ## Step 2: Classify and Structure Analyze the extracted text to determine the document type (invoice, receipt, contract, ID, form, letter, or other). Extract key structured fields based on type — e.g. for an invoice: vendor, date, line items, total. ## Step 3: Report Produce a document analysis result: - Document type identified - Structured fields extracted with values - Full raw text extracted - Any low-confidence regions or unreadable sections - Suggested next action (e.g. file, process payment, follow up)
Create agent.yaml in the same directory to wire the skill to your connector:
name: Document Reader description: Extract and analyze text from images using OCR connectors: - connection_name: Azure Vision connection_type: azure-vision skills: - path: ./ trigger_type: auto
Create oe-config.json in the same directory:
{
"llm": {
"provider": "openai",
"model": "gpt-4o",
"apiKey": "YOUR_OPENAI_API_KEY"
},
"server": {
"enabled": false,
"port": 3333,
"apiKey": "your-secret-api-key"
},
"connectors": [
{
"connection_name": "Azure Vision",
"connection_type": "azure-vision",
"apiKey": "YOUR_AZURE_VISION_KEY",
"headerName": "Ocp-Apim-Subscription-Key",
"baseUrl": "https://YOUR_RESOURCE.cognitiveservices.azure.com/vision/v3.2"
}
]
}Create an Azure Computer Vision resource at portal.azure.com and copy the endpoint and API key from the resource's Keys and Endpoint page. The free tier (F0) allows up to 5,000 OCR transactions per month.
From the parent folder containing your skill directory:
| Method | Best for | Download | |
|---|---|---|---|
| 1 | npx recommended | No install needed — always runs the latest version | — |
| 2 | Windows .exe | Download once, run offline on Windows | ⊞ Windows (.exe) |
| 3 | macOS binary | Download once, run offline on Mac | macOS |
| 4 | Linux binary | Server deployments, cron jobs, Docker | 🐧 Linux |
| 5 | API Server integration | Call from any app, webhook, or automation pipeline | 📮 Postman Collection |
npx -y @openenthrium/oe-runtime@latest ./ocr-vision
oe-runtime-win.exe ./ocr-vision
chmod +x oe-runtime-macos
./oe-runtime-macos ./ocr-vision
First run blocked? System Settings → Privacy & Security → Allow Anyway.
chmod +x oe-runtime-linux
./oe-runtime-linux ./ocr-vision
Add a "server" block to oe-config.json, then start with --serve:
{
"llm": { ... },
"server": { "enabled": true, "port": 3333, "apiKey": "your-secret-key" },
"connectors": [ ... ]
}
npx -y @openenthrium/oe-runtime@latest --serve --config oe-config.json
Run with inline YAML:
curl -X POST http://localhost:3333/run \
-H "Content-Type: application/json" \
-H "X-API-Key: your-secret-key" \
-d '{"yaml": "...", "params": {}}'
Or run from a file on the server:
curl -X POST http://localhost:3333/run-file \
-H "Content-Type: application/json" \
-H "X-API-Key: your-secret-key" \
-d '{"file": "/path/to/agent.yaml", "params": {}}'
Download OE Runtime and run any AI agent locally or as a server — no cloud required.
Get OE Runtime →