← Back to Blog

How to Run an AI Text-to-Speech Agent with OE Runtime

Convert written content into natural-sounding speech audio automatically. OE Runtime connects to ElevenLabs, prepares your text for audio delivery, generates the audio with your chosen voice settings, and returns the file URL — all from a single YAML file.

🔊
Step 1
Prepare Content
Clean and optimise the text for audio delivery — expand abbreviations and add natural pause markers.
Step 2
Generate Audio
Call ElevenLabs with voice settings (stability 0.7, similarity boost 0.8) and retrieve the audio URL.
Step 3
Report
Return character count, voice settings used, audio URL, estimated duration, and generation confirmation.

What You Need


Create the Project Folder

Create a folder called speech-audio/ and add these two files:

speech-audio/
├── SKILL.md         # the portable skill (agentskills.io)
├── agent.yaml       # wires SKILL.md to your connector
└── oe-config.json  # LLM key + connector credentials

The Skill Files

Create SKILL.md inside a speech-audio/ directory. This is the portable skill — it follows the agentskills.io format and runs unchanged on Claude, Cursor, Windsurf, or OE Runtime:

---
name: speech-audio
description: Convert text to natural-sounding speech audio
license: Apache-2.0
metadata:
  author: Open Enthrium
  version: "1.0"
---

You are a speech synthesis agent. Convert provided text to audio using ElevenLabs.
Choose a clear, professional voice appropriate for the content type.
Complete all steps fully before writing your report.

## Step 1: Prepare Content
Prepare the following text for speech synthesis. Clean it for audio delivery —
expand abbreviations and confirm it reads well aloud:

"Welcome to Open Enthrium. Your AI automation platform is ready. You can now
build, deploy, and run intelligent agents that connect to any tool or service.
Let's get started."

## Step 2: Select Voice
GET /voices to retrieve the list of available ElevenLabs voices.
Select a professional, clear voice suitable for corporate narration.
Save the voice_id of the selected voice.

## Step 3: Generate Audio
POST /text-to-speech/<voice_id> with body:
{
  "text": "<prepared text from step 1>",
  "model_id": "eleven_multilingual_v2",
  "voice_settings": { "stability": 0.7, "similarity_boost": 0.8 }
}
The response is an audio file saved to a local temp path. Report that path.

## Step 4: Report
Summarize the result:
- Text converted (character count)
- Voice selected (name and voice_id)
- Voice settings used (stability, similarity_boost)
- Output audio file path

Create agent.yaml in the same directory to wire the skill to your connector:

name: Text to Speech Agent
description: Convert text to natural-sounding speech audio
connectors:
  - connection_name: ElevenLabs
    connection_type: elevenlabs
skills:
  - path: ./
    trigger_type: auto

The Config File

Create oe-config.json in the same directory:

{
  "llm": {
    "provider": "openai",
    "model": "gpt-4o",
    "apiKey": "YOUR_OPENAI_API_KEY"
  },
  "server": {
    "enabled": false,
    "port": 3333,
    "apiKey": "your-secret-api-key"
  },
  "connectors": [
    {
      "connection_name": "ElevenLabs",
      "connection_type": "elevenlabs",
      "apiKey": "YOUR_ELEVENLABS_API_KEY",
      "headerName": "xi-api-key",
      "baseUrl": "https://api.elevenlabs.io/v1"
    }
  ]
}

Get your ElevenLabs API key from elevenlabs.io/app/settings. The agent automatically fetches available voices via GET /voices and selects the best match for the content — no need to hardcode a voice ID. Browse the full voice library at elevenlabs.io/voice-library. The generated audio is saved to a local file and the path is returned in the report.

Download OE Runtime

OE Runtime — Direct Downloads

Run the Agent

From the parent folder containing your skill directory:

MethodBest forDownload
1 npx recommended No install needed — always runs the latest version —
2 Windows .exe Download once, run offline on Windows ⊞ Windows (.exe)
3 macOS binary Download once, run offline on Mac  macOS
4 Linux binary Server deployments, cron jobs, Docker 🐧 Linux
5 API Server integration Call from any app, webhook, or automation pipeline 📮 Postman Collection

1 npx recommended

npx -y @openenthrium/oe-runtime@latest ./speech-audio

2 Windows

oe-runtime-win.exe ./speech-audio

3 macOS

chmod +x oe-runtime-macos
./oe-runtime-macos ./speech-audio

First run blocked? System Settings → Privacy & Security → Allow Anyway.

4 Linux

chmod +x oe-runtime-linux
./oe-runtime-linux ./speech-audio

5 API Server integration

Add a "server" block to oe-config.json, then start with --serve:

{
  "llm": { ... },
  "server": { "enabled": true, "port": 3333, "apiKey": "your-secret-key" },
  "connectors": [ ... ]
}
npx -y @openenthrium/oe-runtime@latest --serve --config oe-config.json

Run with inline YAML:

curl -X POST http://localhost:3333/run \
  -H "Content-Type: application/json" \
  -H "X-API-Key: your-secret-key" \
  -d '{"yaml": "...", "params": {}}'

Or run from a file on the server:

curl -X POST http://localhost:3333/run-file \
  -H "Content-Type: application/json" \
  -H "X-API-Key: your-secret-key" \
  -d '{"file": "/path/to/agent.yaml", "params": {}}'

Use Cases

Build your own agents with OE Runtime

Download OE Runtime and run any AI agent locally or as a server — no cloud required.

Get OE Runtime →
Series OE Runtime Agent Guides — 21 Connectors
Series overview →