No internet. No API key. No cost per token. Run a full end-to-end AI agent pipeline on your own machine using Ollama as the local LLM and OE Runtime as the agent executor — two files, one command, zero cloud dependency.
qwen2.5:7b or llama3.1:8b)Cloud LLMs are powerful but they come with trade-offs: every prompt you send is transmitted over the internet, billed per token, and subject to the provider's data retention policies. For many real-world workloads — internal documents, customer PII, regulated data, air-gapped infrastructure — that is simply not acceptable.
Running agents locally with Ollama solves all three problems at once:
✅ OE Runtime supports Ollama natively. Set "provider": "ollama" in your config and the runtime connects to the local Ollama server automatically — no extra setup, no code changes.
Download Ollama from ollama.com and install it. Then pull a model — qwen2.5:7b is a strong general-purpose choice that runs well on most developer machines with 8 GB RAM:
# Start the Ollama server (runs in the background)
ollama serve
# In a second terminal — pull the model (downloads once, cached locally)
ollama pull qwen2.5:7b
The model downloads once (~4.7 GB for qwen2.5:7b) and is cached on disk. Subsequent runs are instant — no re-download.
Any model available through ollama pull works with OE Runtime. These are tested and recommended:
| Model | Size | Best for |
|---|---|---|
| qwen2.5:7b | 4.7 GB | General agents, reasoning, structured output — best balance of speed and quality |
| llama3.2:3b | 2.0 GB | Low-RAM machines, fast responses, simple tasks |
| phi4-mini | 2.5 GB | Microsoft model, very capable for its size, edge deployments |
| mistral:7b | 4.1 GB | European data-residency requirements, strong instruction following |
| deepseek-r1:7b | 4.7 GB | Reasoning-heavy tasks, multi-step planning agents |
| llama3.1:8b | 4.9 GB | Robust tool calling, complex multi-step workflows |
Create a folder called hello-world-offline/ and add two files:
Create SKILL.md and agent.yaml inside hello-world-offline/:
---
name: hello-world-offline
description: Runs entirely on a local LLM via Ollama — no internet, no API cost
license: Apache-2.0
metadata:
author: Open Enthrium
version: "1.0"
---
You are a helpful AI assistant running fully offline on a local model.
No data leaves this machine.
## Step 1: Greet
Introduce yourself, state today's date, and share one interesting
fact about local AI. Keep it to three sentences.
agent.yaml
name: Hello World (Offline)
description: Runs entirely on a local LLM via Ollama — no internet, no API cost
skills:
- path: ./
trigger_type: auto
Create oe-config.json in the same folder. Set "provider" to "ollama" — OE Runtime connects to Ollama at http://localhost:11434 automatically. No API key required:
{
"llm": {
"provider": "ollama",
"model": "qwen2.5:7b",
"apiKey": "ollama"
},
"server": {
"enabled": false,
"port": 3333,
"apiKey": "your-secret-api-key"
}
}
The "apiKey": "ollama" field is a placeholder — Ollama does not require authentication. OE Runtime uses it internally to satisfy the OpenAI-compatible client.
💡 Want to use a different model? Just change "model" to any model you have pulled — for example "phi4-mini", "mistral:7b", or "llama3.2:3b". No other changes needed.
Make sure Ollama is running (ollama serve), then from the parent folder containing your skill directory:
| Method | Best for | Download | |
|---|---|---|---|
| 1 | npx recommended | No binary download — npx handles it. Ollama is the only local dependency. | — |
| 2 | Windows .exe | Fully offline once downloaded — no Node.js, no npx needed | ⊞ Windows (.exe) |
| 3 | macOS binary | Download once, run offline on Mac | macOS |
| 4 | Linux binary | Air-gapped servers, edge devices, secure infrastructure | 🐧 Linux |
| 5 | API Server integration | Expose agents as a local HTTP API — call from any app on the same network | 📮 Postman Collection |
npx -y @openenthrium/oe-runtime@latest ./hello-world-offline
oe-runtime-win.exe ./hello-world-offline
chmod +x oe-runtime-macos
./oe-runtime-macos ./hello-world-offline
First run blocked? System Settings → Privacy & Security → Allow Anyway.
chmod +x oe-runtime-linux
./oe-runtime-linux ./hello-world-offline
Add a "server" block to oe-config.json, then start with --serve:
{
"llm": { "provider": "ollama", "model": "qwen2.5:7b", "apiKey": "ollama" },
"server": { "enabled": true, "port": 3333, "apiKey": "your-secret-key" }
}
npx -y @openenthrium/oe-runtime@latest --serve --config oe-config.json
Run with inline YAML:
curl -X POST http://localhost:3333/run \
-H "Content-Type: application/json" \
-H "X-API-Key: your-secret-key" \
-d '{"yaml": "...", "params": {}}'
Or run from a file on the server:
curl -X POST http://localhost:3333/run-file \
-H "Content-Type: application/json" \
-H "X-API-Key: your-secret-key" \
-d '{"file": "/path/to/agent.yaml", "params": {}}'
The agent runs on qwen2.5:7b locally, generates a response, and prints it to the terminal:
────────────────────────────────────────────────────
🚀 OE Runtime Standalone v1.7.2
────────────────────────────────────────────────────
🤖 Hello World (Offline)
Runs entirely on a local LLM via Ollama
LLM ollama / qwen2.5:7b
────────────────────────────────────────────────────
Hello! I'm an OE Runtime agent running entirely on your local machine
using qwen2.5:7b via Ollama. Today is August 14, 2026. Local AI models
like qwen2.5 can now match cloud LLMs on most everyday tasks — at zero
cost and with complete data privacy.
✅ Done
No internet request was made. No token was billed. The entire pipeline — agent executor, LLM inference, and response — ran on your own hardware.
The same Ollama config works with any OE Runtime connector. Switch out the YAML, add connector credentials, and connect to real local data sources — the LLM stays on-device throughout.
Here is a complete example: an agent that queries a local PostgreSQL database and summarises the results — entirely offline.
---
name: local-db-summary
description: Queries a local PostgreSQL database and returns a plain-English summary
license: Apache-2.0
metadata:
author: Open Enthrium
version: "1.0"
---
You are a data analyst running fully offline. Query the database and
respond with a clear, concise summary. No data leaves this machine.
## Step 1: List Tables
Query the Local DB connector:
SELECT table_name, pg_size_pretty(pg_total_relation_size(table_name::text)) AS size
FROM information_schema.tables
WHERE table_schema = 'public'
ORDER BY pg_total_relation_size(table_name::text) DESC;
List each table with its size.
## Step 2: Summarise
Based on the tables you found, write a two-sentence summary of what
this database appears to store and which table holds the most data.
name: Local DB Summary (Offline)
description: Queries a local PostgreSQL database and returns a plain-English summary
connectors:
- connection_name: Local DB
connection_type: postgresql
skills:
- path: ./
trigger_type: auto
{
"llm": {
"provider": "ollama",
"model": "qwen2.5:7b",
"apiKey": "ollama"
},
"connectors": [
{
"connection_name": "Local DB",
"connection_type": "postgresql",
"host": "localhost",
"port": 5432,
"database": "mydb",
"user": "postgres",
"password": "your-password"
}
]
}
npx -y @openenthrium/oe-runtime@latest ./hello-world-offline
The local model reads the query results, reasons over them, and returns a plain-English summary — no internet required at any step. Swap the connector for SSH, local file storage, a private REST API, or any of OE Runtime's 45+ connector types and the pattern stays identical.
Download OE Runtime and run any AI agent locally or as a server — no cloud required.
Get OE Runtime →