Audio module

AI voice and transcription, every model, one subscription.

Omnibo gives you 29 AI audio models from 8 providers — MiniMax Speech, Seed TTS, Gemini TTS, GPT-4o mini TTS, Grok Voice and transcription — on a single $29.99/month subscription. Speech generation is metered per million characters of input and transcription per minute of audio, and both figures are shown before the run. A standalone voice subscription starts at about $22/month and covers voice alone.

What you can do in the Omnibo audio module

The Omnibo audio module covers both directions: text into speech, and recorded audio into text. Both are metered in the same balance as everything else you make.

Welcome back to the show — today we’re talking about|
52 / 9,999
Model & options
MiniMax Speech 2.8 HD
Audio models · 29 available
ACE-Step0.2 cr/s
CassetteAI Sound Effects12 cr/clip
Eleven Flash v2.560,000 cr/M chars
Eleven Multilingual v2120,000 cr/M chars
Eleven Music720 cr/min
Eleven Music v2.53 cr/s
ElevenLabs Sound Effects v22.4 cr/s
Eleven v3120,000 cr/M chars
ElevenLabs Sound Effects v2 (fal.ai)2.4 cr/s
MiniMax Music 2.6180 cr/clip
Scribe v24.4 cr/min
Stable Audio 2.5240 cr/clip
Kling Text to Audio42 cr/clip
Lyria 3 Clip48 cr/clip
Lyria 3.596 cr/clip
Seed Audio 1.03 cr/s
Gemini 3.5 Transcribe6 cr/min
MiniMax Speech 2.8 HD120,000 cr/M chars
MiniMax Speech 2.8 Turbo72,000 cr/M chars
Gemini 3.1 Flash TTS55,200 cr/M chars
GPT Transcribe5.4 cr/min
Gemini 2.5 Flash TTS27,600 cr/M chars
Gemini 2.5 Pro TTS55,200 cr/M chars
GPT-4o mini TTS22,800 cr/M chars
Grok Transcribe2 cr/min
Grok Voice18,000 cr/M chars
Seed TTS 1.018,000 cr/M chars
TTS-118,000 cr/M chars
TTS-1 HD36,000 cr/M chars
See all 29 audio models →
~6.2 cr
for 52 characters

Rebuilt in HTML from the live components — “MiniMax Speech 2.8 HD, 120,000 cr / M characters” is real text on this page, not an image.

Speech

Turn a script into narration on any voice model your plan reaches, up to 9,999 characters a run.

Voices

Pick a voice per run, so the same script can be auditioned in several deliveries before you commit to one.

Transcription

Turn recorded audio into text, metered per minute rather than per character.

Carry forward

Use a generated voice track as the reference audio on a video run without leaving Omnibo.

Every audio model in Omnibo, and what a minute of speech costs

Omnibo carries 29 audio models across 8 providers. Speech models are metered per million characters of input text; transcription models are metered per minute of audio. Both units are shown per row rather than averaged into one number. ACE-Step costs 0.2 cr / second; Eleven Multilingual v2 costs 120,000 cr / M characters. A Pro plan includes 9,000 credits a month.

Catalogue last updated 14 September 2026 · rates read from the live catalogue on every revalidation

Required plan for these audio models, cheapest first:Hobby · 17Pro · 12

Every audio model in Omnibo with its provider, credit rate, required plan and the date it was added
ModelProviderCostPlanAdded
ACE-StepNewfal.ai0.2 cr / secondHobby2026-09-14
CassetteAI Sound EffectsNewfal.ai12 cr / clipHobby2026-09-14
Eleven Flash v2.5NewElevenLabs60,000 cr / M charactersHobby2026-09-14
Eleven Multilingual v2NewElevenLabs120,000 cr / M charactersPro2026-09-14
Eleven MusicNewfal.ai720 cr / minutePro2026-09-14
Eleven Music v2.5NewElevenLabs3 cr / secondPro2026-09-14
ElevenLabs Sound Effects v2NewElevenLabs2.4 cr / secondHobby2026-09-14
Eleven v3NewElevenLabs120,000 cr / M charactersPro2026-09-14
ElevenLabs Sound Effects v2 (fal.ai)Newfal.ai2.4 cr / secondHobby2026-09-14
MiniMax Music 2.6Newfal.ai180 cr / clipPro2026-09-14
Scribe v2NewElevenLabs4.4 cr / minuteHobby2026-09-14
Stable Audio 2.5Newfal.ai240 cr / clipPro2026-09-14
Kling Text to AudioKling AI42 cr / clipHobby2026-09-13
Lyria 3 ClipGoogle48 cr / clipHobby2026-09-13
Lyria 3.5Google96 cr / clipPro2026-09-13
Seed Audio 1.0ByteDance3 cr / secondHobby2026-09-13
Gemini 3.5 TranscribeGoogle6 cr / minuteHobby2026-09-08
MiniMax Speech 2.8 HDMiniMax120,000 cr / M charactersPro2026-08-30
MiniMax Speech 2.8 TurboMiniMax72,000 cr / M charactersPro2026-08-30
Gemini 3.1 Flash TTSGoogle55,200 cr / M charactersPro2026-08-08
GPT TranscribeOpenAI5.4 cr / minuteHobby2026-08-08
Gemini 2.5 Flash TTSGoogle27,600 cr / M charactersHobby2026-07-05
Gemini 2.5 Pro TTSGoogle55,200 cr / M charactersPro2026-07-05
GPT-4o mini TTSOpenAI22,800 cr / M charactersHobby2026-07-05
Grok TranscribexAI2 cr / minuteHobby2026-07-05
Grok VoicexAI18,000 cr / M charactersHobby2026-07-05
Seed TTS 1.0ByteDance18,000 cr / M charactersHobby2026-07-05
TTS-1OpenAI18,000 cr / M charactersHobby2026-07-05
TTS-1 HDOpenAI36,000 cr / M charactersPro2026-07-05

Scroll the table sideways for cost, plan and date

How much does AI voice generation cost per month?

The voice tools sell character allowances that expire monthly, and none of them writes the script you are about to narrate.

How much does AI voice generation cost per month?
ServiceMonthlyWhat it coversEffective
ElevenLabs Creator$22100,000 charactersVoice only — no chat, image or video
Omnibo Pro$29.99~0 runs on MiniMax Speech 2.8 HDPlus 57 chat, 25 image, 20 video models

Scroll the table sideways for the full comparison

The Omnibo row is arithmetic you can check: 9,000 credits a month ÷ 120,000 credits per run on MiniMax Speech 2.8 HD = 0 runs on MiniMax Speech 2.8 HD. Credits do not roll over on any of these, Omnibo included. The difference is that the others each want their own subscription, and none of them will write your script or voice it. September 2026 list prices for the cheapest tier of each service that permits commercial, watermark-free output.

Work made in Omnibo

Every image on this site was generated in Omnibo, tagged with the model that made it and what that run cost. No stock, no other tools.

Close-up portrait of a woman on a rain-wet city street at night, generated from the shared brief
FLUX.2 [pro]36 cr
A camel wool coat photographed as a fashion editorial on a warm sand backdrop, generated in Omnibo
FLUX.2 [pro]36 cr
The same rain-wet street portrait brief, interpreted by a different image model
GPT Image 264 cr
A rally car mid-corner in golden-hour dust, generated in Omnibo
FLUX.2 [pro]36 cr
The same rain-wet street portrait brief again, with warmer shopfront light behind the subject
Nano Banana Pro161 cr
A saxophonist under amber stage light in a small jazz club, generated in Omnibo
FLUX.2 [pro]36 cr
The same rain-wet street portrait brief, rendered with harder contrast and a colder key
Grok Imagine 2.096 cr
A screenwriter’s desk at dawn, papers and a cold coffee lit by window light, generated in Omnibo
FLUX.2 [pro]36 cr

Questions about AI voice and transcription on Omnibo

How is Omnibo speech generation metered?

Per million characters of the text you submit, which is how the voice providers meter it themselves. A run is capped at 9,999 characters, and the composer shows the credit cost of the exact text in the box before you generate.

How is Omnibo transcription metered?

Per minute of audio, not per character — transcription models take sound in rather than text, so the table below quotes them in credits per minute.

Can I use Omnibo voice output commercially?

Yes on every paid Omnibo plan, subject to each provider’s own terms, which are linked per model on the catalogue page. Free-plan output is for evaluation.

Does the audio balance come out of the same credits as everything else?

Yes. Omnibo has one balance for chat, image, video and audio, so an unused voice allowance is not stranded in a separate subscription at the end of the month.

Write it, film it and voice it on one balance.

Start free. No card. All 29 audio models are on the same account as the 57 chat, 25 image, 20 video models.

Start free