Voice Agents with no Awkward Pauses

GLM 5.1
On a specialized Wafer endpoint
550 ms
Client-observed p50 TTFT, down from the prior provider’s 800 ms
-30%
p50 TTFT at ~25% higher peak load
  • Dedicated Endpoint
  • TTFT-based SLA
  • Signed BAA · US-only

Overview

Neon Health runs HIPAA-compliant voice agents for healthcare. By moving GLM-5.1 inference to a dedicated Wafer endpoint, the team reduced client-observed p50 time-to-first-token by ~30%, with that target now written into the SLA.

Solution

Client-observed p50 TTFT cut from 800 ms to ~550 ms

Neon Health’s own measurement after moving GLM 5.1 to a dedicated Wafer endpoint — at ~25% higher peak load

What changed

Same model, faster serving — GLM 5.1 on a dedicated, TTFT-tuned Wafer stack

Headroom under the wall — with 250 ms of the 800 ms turn budget freed up

Inside the SLA — with previous 800 ms sat below the 600 ms target

Where the 800 ms goes — and the LLM’s slice

From end of user speech to first agent audio. Past 800 ms, the call stops feeling like a conversation

Latency that doesn’t fall off a cliff

«The lowest latency we’ve seen — and it doesn’t go off a cliff when you increase the requests per minute»

«[Wafer] has the lowest latency we’ve seen from any provider we’ve tried. And it doesn’t go off a cliff when you increase the requests per minute.»
Harry BleyanCo-founder & CTO, Neon Health

Your voice agents live or die on latency

Wafer measures your traffic, optimizes a dedicated endpoint for your TTFT SLA, and gets you to production in under two weeks — without crashing during call volume spikes

Signed BAA · US-only data residency · TTFT written into the SLA