Since GPT-4o Realtime, "real time" has become the magic word in boardrooms.

The demos are stunning: 200ms latency, natural interruptions, a conversation that finally feels like a conversation.

So teams demand realtime. Everywhere. For everything.

That's an architecture mistake dressed up as ambition. You're choosing the thrill of the demo, not the fit of the use case. And the invoice follows.

That's an architecture mistake dressed up as ambition.

What the Hype Doesn't Show

Realtime processes audio natively, with no intermediate transcription. It's real, it's impressive, and it comes with three trade-offs the demos gloss over.

Cost. Native speech-to-speech is structurally more expensive than a separate pipeline. At volume, support, lead qualification, the gap isn't marginal, it's budgetary.

Voice. In a pure realtime stack, synthesis is baked into the model. Less TTS customization, no cloned voice. Your brand's voice becomes the provider's voice.

Fit. A user exploring a catalog by voice deserves 200ms. A prospect recording an async message will never hear the difference. Paying for realtime there is burning budget on a thrill nobody feels.

Two Architectures, One Setting — Not a Doctrine

At Scenaro, each scenario chooses its architecture. It's not a platform decision, it's a per-context setting.

The realtime model: native audio, roughly 200ms, natural interruptions. Ideal when conversational fluency is the differentiator, a premium advisor in full duplex, voice-driven product exploration. Trade-off: less control over the voice.

The STT→LLM→TTS pipeline: 800ms to 1.5s, but full control over every layer. You pick your transcription engine, your model, your voice, cloned if needed. Ideal for volume, cost control, and sonic identity.

And in between, the hybrid variant: realtime inference, separate voice synthesis. Realtime speed on reasoning, freedom on voice.

None of the three is "the future." Each answers a different question.

The Right Question

The right question isn't "realtime or not?

" It's: what latency, what voice, and what cost fit this specific scenario?

A wine advisor in full duplex on a homepage justifies realtime. A qualification scenario in push-to-talk doesn't need 200ms, a separate pipeline does the job at a fraction of the cost, with the brand's exact voice.

That's why at Scenaro the choice happens scenario by scenario, in the Cockpit. The same conversational experience can run a realtime scenario and a pipeline scenario side by side. You compare, measure, adjust. Scenario versioning even lets you test one architecture against the other on the same journey before rolling it out.

The most expensive mistake never changes: deploying realtime everywhere because the demo impressed, then discovering on the invoice that most conversations never needed it. Voice architecture is decided per use case. Never per trend.