Voice Agents with no Awkward Pauses
- GLM 5.1
- On a specialized Wafer endpoint
- 550 ms
- Client-observed p50 TTFT, down from the prior provider’s 800 ms
- -30%
- p50 TTFT at ~25% higher peak load
- Dedicated Endpoint
- TTFT-based SLA
- Signed BAA · US-only
Overview
Neon Health runs HIPAA-compliant voice agents for healthcare. By moving GLM-5.1 inference to a dedicated Wafer endpoint, the team reduced client-observed p50 time-to-first-token by ~30%, with that target now written into the SLA.
Solution
Client-observed p50 TTFT cut from 800 ms to ~550 ms
Neon Health’s own measurement after moving GLM 5.1 to a dedicated Wafer endpoint — at ~25% higher peak load


What changed
Same model, faster serving — GLM 5.1 on a dedicated, TTFT-tuned Wafer stack
Headroom under the wall — with 250 ms of the 800 ms turn budget freed up
Inside the SLA — with previous 800 ms sat below the 600 ms target
Where the 800 ms goes — and the LLM’s slice
From end of user speech to first agent audio. Past 800 ms, the call stops feeling like a conversation


Latency that doesn’t fall off a cliff
«The lowest latency we’ve seen — and it doesn’t go off a cliff when you increase the requests per minute»


«[Wafer] has the lowest latency we’ve seen from any provider we’ve tried. And it doesn’t go off a cliff when you increase the requests per minute.»
Harry BleyanCo-founder & CTO, Neon HealthYour voice agents live or die on latency
Wafer measures your traffic, optimizes a dedicated endpoint for your TTFT SLA, and gets you to production in under two weeks — without crashing during call volume spikes
Signed BAA · US-only data residency · TTFT written into the SLA
