← Back to Blog

How to Run AI Agents Completely Offline with Ollama

No internet. No API key. No cost per token. Run a full end-to-end AI agent pipeline on your own machine using Ollama as the local LLM and OE Runtime as the agent executor — two files, one command, zero cloud dependency.

🏠
Step 1
Install Ollama & Pull a Model
Download Ollama and pull qwen2.5:7b or any local model you prefer.
Step 2
Two Config Files
agent.yaml defines the task. oe-config.json points to Ollama — no API key needed.
Step 3
Run — Fully Offline
One npx command. The agent runs on your local model. No data leaves your machine.

What You Need


Why Run Agents Offline?

Cloud LLMs are powerful but they come with trade-offs: every prompt you send is transmitted over the internet, billed per token, and subject to the provider's data retention policies. For many real-world workloads — internal documents, customer PII, regulated data, air-gapped infrastructure — that is simply not acceptable.

Running agents locally with Ollama solves all three problems at once:

✅ OE Runtime supports Ollama natively. Set "provider": "ollama" in your config and the runtime connects to the local Ollama server automatically — no extra setup, no code changes.

Step 1 — Install Ollama and Pull a Model

Download Ollama from ollama.com and install it. Then pull a model — qwen2.5:7b is a strong general-purpose choice that runs well on most developer machines with 8 GB RAM:

# Start the Ollama server (runs in the background)
ollama serve

# In a second terminal — pull the model (downloads once, cached locally)
ollama pull qwen2.5:7b

The model downloads once (~4.7 GB for qwen2.5:7b) and is cached on disk. Subsequent runs are instant — no re-download.

Recommended Models

Any model available through ollama pull works with OE Runtime. These are tested and recommended:

ModelSizeBest for
qwen2.5:7b4.7 GBGeneral agents, reasoning, structured output — best balance of speed and quality
llama3.2:3b2.0 GBLow-RAM machines, fast responses, simple tasks
phi4-mini2.5 GBMicrosoft model, very capable for its size, edge deployments
mistral:7b4.1 GBEuropean data-residency requirements, strong instruction following
deepseek-r1:7b4.7 GBReasoning-heavy tasks, multi-step planning agents
llama3.1:8b4.9 GBRobust tool calling, complex multi-step workflows

Step 2 — Create the Project Folder

Create a folder called hello-world-offline/ and add two files:

hello-world-offline/
├── SKILL.md         # the portable skill (agentskills.io)
├── agent.yaml       # wires SKILL.md to Ollama — no cloud needed
└── oe-config.json  # points to local Ollama — no API key needed

The Skill Files

Create SKILL.md and agent.yaml inside hello-world-offline/:

SKILL.md
---
name: hello-world-offline
description: Runs entirely on a local LLM via Ollama — no internet, no API cost
license: Apache-2.0
metadata:
  author: Open Enthrium
  version: "1.0"
---

You are a helpful AI assistant running fully offline on a local model.
No data leaves this machine.

## Step 1: Greet
Introduce yourself, state today's date, and share one interesting
fact about local AI. Keep it to three sentences.
agent.yaml
name: Hello World (Offline)
description: Runs entirely on a local LLM via Ollama — no internet, no API cost
skills:
  - path: ./
    trigger_type: auto

The Config File

Create oe-config.json in the same folder. Set "provider" to "ollama" — OE Runtime connects to Ollama at http://localhost:11434 automatically. No API key required:

{
  "llm": {
    "provider": "ollama",
    "model": "qwen2.5:7b",
    "apiKey": "ollama"
  },
  "server": {
    "enabled": false,
    "port": 3333,
    "apiKey": "your-secret-api-key"
  }
}

The "apiKey": "ollama" field is a placeholder — Ollama does not require authentication. OE Runtime uses it internally to satisfy the OpenAI-compatible client.

💡 Want to use a different model? Just change "model" to any model you have pulled — for example "phi4-mini", "mistral:7b", or "llama3.2:3b". No other changes needed.

Download OE Runtime

OE Runtime — Direct Downloads

Run the Agent

Make sure Ollama is running (ollama serve), then from the parent folder containing your skill directory:

MethodBest forDownload
1 npx recommended No binary download — npx handles it. Ollama is the only local dependency. —
2 Windows .exe Fully offline once downloaded — no Node.js, no npx needed ⊞ Windows (.exe)
3 macOS binary Download once, run offline on Mac  macOS
4 Linux binary Air-gapped servers, edge devices, secure infrastructure 🐧 Linux
5 API Server integration Expose agents as a local HTTP API — call from any app on the same network 📮 Postman Collection

1 npx recommended

npx -y @openenthrium/oe-runtime@latest ./hello-world-offline

2 Windows

oe-runtime-win.exe ./hello-world-offline

3 macOS

chmod +x oe-runtime-macos
./oe-runtime-macos ./hello-world-offline

First run blocked? System Settings → Privacy & Security → Allow Anyway.

4 Linux

chmod +x oe-runtime-linux
./oe-runtime-linux ./hello-world-offline

5 API Server integration

Add a "server" block to oe-config.json, then start with --serve:

{
  "llm": { "provider": "ollama", "model": "qwen2.5:7b", "apiKey": "ollama" },
  "server": { "enabled": true, "port": 3333, "apiKey": "your-secret-key" }
}
npx -y @openenthrium/oe-runtime@latest --serve --config oe-config.json

Run with inline YAML:

curl -X POST http://localhost:3333/run \
  -H "Content-Type: application/json" \
  -H "X-API-Key: your-secret-key" \
  -d '{"yaml": "...", "params": {}}'

Or run from a file on the server:

curl -X POST http://localhost:3333/run-file \
  -H "Content-Type: application/json" \
  -H "X-API-Key: your-secret-key" \
  -d '{"file": "/path/to/agent.yaml", "params": {}}'

What You Get Back

The agent runs on qwen2.5:7b locally, generates a response, and prints it to the terminal:

────────────────────────────────────────────────────
  🚀   OE Runtime Standalone  v1.7.2
────────────────────────────────────────────────────

🤖  Hello World (Offline)
    Runs entirely on a local LLM via Ollama

    LLM        ollama / qwen2.5:7b

────────────────────────────────────────────────────

Hello! I'm an OE Runtime agent running entirely on your local machine
using qwen2.5:7b via Ollama. Today is August 14, 2026. Local AI models
like qwen2.5 can now match cloud LLMs on most everyday tasks — at zero
cost and with complete data privacy.

✅  Done

No internet request was made. No token was billed. The entire pipeline — agent executor, LLM inference, and response — ran on your own hardware.

Add Connectors: Go Beyond Hello World

The same Ollama config works with any OE Runtime connector. Switch out the YAML, add connector credentials, and connect to real local data sources — the LLM stays on-device throughout.

Here is a complete example: an agent that queries a local PostgreSQL database and summarises the results — entirely offline.

SKILL.md

---
name: local-db-summary
description: Queries a local PostgreSQL database and returns a plain-English summary
license: Apache-2.0
metadata:
  author: Open Enthrium
  version: "1.0"
---

You are a data analyst running fully offline. Query the database and
respond with a clear, concise summary. No data leaves this machine.

## Step 1: List Tables
Query the Local DB connector:
SELECT table_name, pg_size_pretty(pg_total_relation_size(table_name::text)) AS size
FROM information_schema.tables
WHERE table_schema = 'public'
ORDER BY pg_total_relation_size(table_name::text) DESC;
List each table with its size.

## Step 2: Summarise
Based on the tables you found, write a two-sentence summary of what
this database appears to store and which table holds the most data.

agent.yaml

name: Local DB Summary (Offline)
description: Queries a local PostgreSQL database and returns a plain-English summary
connectors:
  - connection_name: Local DB
    connection_type: postgresql
skills:
  - path: ./
    trigger_type: auto

oe-config.json

{
  "llm": {
    "provider": "ollama",
    "model": "qwen2.5:7b",
    "apiKey": "ollama"
  },
  "connectors": [
    {
      "connection_name": "Local DB",
      "connection_type": "postgresql",
      "host": "localhost",
      "port": 5432,
      "database": "mydb",
      "user": "postgres",
      "password": "your-password"
    }
  ]
}
npx -y @openenthrium/oe-runtime@latest ./hello-world-offline

The local model reads the query results, reasons over them, and returns a plain-English summary — no internet required at any step. Swap the connector for SSH, local file storage, a private REST API, or any of OE Runtime's 45+ connector types and the pattern stays identical.

Enterprise Use Cases

Build your own agents with OE Runtime

Download OE Runtime and run any AI agent locally or as a server — no cloud required.

Get OE Runtime →