Skip to main content
Use this guide when you want to turn a document, changelog, or article into a short audio conversation between two speakers.

What you’re building

A four-sentence paragraph about tides, scripted by a chat model and voiced by two speakers in one request. 49 seconds, $0.0074.

Host: Welcome back! Today we’re talking tides. Most people know the Moon causes them, but why do we get two high tides a day instead of just one?Guest: It comes down to ocean bulges. The Moon’s gravity pulls hardest on the ocean facing it, but it also creates a second bulge on the exact opposite side.Host: Wait, on the far side too? How does the Moon pull water away from itself?Guest: It pulls the Earth’s center harder than the far water! Because the pull is weaker out there, the ocean stretches outward, forming that second bulge.Host: Ah, so the Earth basically spins right beneath both of these bulges every day?Guest: Exactly! Coastlines rotate through both, which is why we usually see two high and two low tides roughly every twenty-five hours.
The /api/v1/audio/speech endpoint accepts a list of turns as input, and each turn can name its own voice and delivery instructions. The model voices the whole conversation in one request and returns a single audio stream, so you never split the dialogue into per-line requests or stitch clips back together in order. By the end, you will have a script that:
  1. Asks a chat model for a host-and-guest dialogue as structured JSON.
  2. Sends every turn to a Gemini TTS model in one request, with a different voice for each speaker.
  3. Saves the result as a WAV file you can play.
Building with a coding agent? Paste the URL of this page into Claude Code, Codex, or Cursor and ask it to add a “listen to this post” button to your blog using this recipe.

Before you start

You need:
  • An OpenRouter API key available as OPENROUTER_API_KEY
  • Bun, or Node.js 24 or newer, both of which run TypeScript files directly
  • About one cent of credits per minute of audio
The three steps below form one script. Paste them in order into podcast.ts and run it with bun podcast.ts or node podcast.ts. This guide uses google/gemini-3.8-flash-lite-tts. Not every TTS provider supports multi-speaker input, and those that do not return a 400 rather than reading the whole dialogue in one voice. Text-to-Speech lists the supported models and the full parameter reference.

Step 1: Write the script

Ask a chat model for the dialogue and use structured outputs to get it back as a list of turns. A schema saves you from parsing speaker labels out of free text, and the direction field gives the voice model something to act on.
The model returns a script like this one:

Step 2: Voice the whole script in one request

Map each speaker to a voice, then send the turns as input, with each turn’s direction as its instructions. A top-level instructions applies only to turns that omit their own, so this request leaves it out.
Gemini TTS returns pcm only. Requesting mp3 returns 400 Gemini TTS only supports response_format="pcm". The 30 available voice names are listed under supported_voices in the models API.

Step 3: Save it as a WAV file

The response body is raw 16-bit PCM at 24 kHz, mono, as the Content-Type header says. Most players will not open raw PCM, so add a 44-byte WAV header in front of it:
To ship MP3 instead, convert the WAV with ffmpeg -i podcast.wav podcast.mp3.

What it costs

The 49-second clip above cost $0.0074 to voice. Look up any request’s cost with GET /api/v1/generation?id=<generation ID>, using the ID from the response header. Writing the script with google/gemini-3.8-flash adds a fraction of a cent. For higher-quality speech, google/gemini-3.8-flash-tts takes the same request at a higher output price.

Troubleshooting

Check your work

The script should print a dialogue of six to eight alternating turns, then save a podcast.wav that plays two clearly different voices taking turns in the order the script lists them.