Quantized TTS on iOS: Core ML & the latency ceiling
Meta description: Run quantized TTS on iPhone with Core ML. Learn phoneme buffer design, ANE vs GPU tradeoffs, and how to achieve sub-200ms first-audio latency on-device.
TL;DR
Running a distilled VITS or Kokoro-class TTS model on iPhone via Core ML is achievable — but the latency ceiling is unforgiving. Sub-200ms first audio requires chunked phoneme buffers, careful ANE scheduling, and a hard stop at INT8 quantization for prosody-sensitive layers. This is the architecture that gets you there.
Why on-device TTS, and why now
Cloud TTS is getting complicated. OpenAI announced in 2025 it would test sponsored content inside ChatGPT, and that’s unlikely to be the last such move. The on-device case was already strong: no API cost, no latency jitter from network round-trips, no data leaving the device. For TTS specifically, the privacy argument comes bundled with better tail latency.
The numbers are concrete. A typical cloud TTS round-trip runs 300–600ms on a good connection. In my testing, Core ML on a Neural Engine-capable iPhone hits 120–180ms to first audio for a quantized model — if you architect the pipeline correctly.
The phoneme-to-mel pipeline
Most distilled TTS architectures share a common spine: text → G2P → duration predictor → mel spectrogram → vocoder. The Core ML deployment splits this into two inference passes:
- Encoder + Duration Predictor — runs once per utterance chunk, produces aligned mel frames
- HiFi-GAN or MB-MelGAN Vocoder — converts mel frames to 22.05kHz PCM, streamed in chunks
The trick for sub-200ms is to not wait for the full utterance. Fire the vocoder on the first 50–80 mel frames while the encoder continues processing the remainder.
// Chunked phoneme buffer dispatch
func synthesizeChunked(phonemes: [Int32], chunkSize: Int = 64) async throws {
var offset = 0
while offset < phonemes.count {
let slice = Array(phonemes[offset..<min(offset + chunkSize, phonemes.count)])
let melFrames = try await encoderModel.predict(phonemes: slice)
let audio = try await vocoderModel.predict(mel: melFrames)
audioEngine.scheduleBuffer(audio)
offset += chunkSize
}
}
First audio hits the speaker before synthesis completes. That’s where the latency budget is won.
ANE vs GPU: the delegate tradeoff
Core ML’s MLComputeUnits enum gives you three paths. In my testing on iPhone 13 Pro (A15, VITS-small at 22kHz, 5-word utterances):
| Compute Unit | First-Audio Latency | Power Draw | Best For |
|---|---|---|---|
.cpuAndNeuralEngine | 130–180ms | Low | Attention-heavy encoder |
.cpuAndGPU | 200–280ms | High | Upsampling vocoder |
.all (auto) | 140–200ms | Medium | Baseline only |
The ANE excels at the encoder’s attention layers. The vocoder, with its upsampling convolutions, consistently runs faster on GPU. So I split the model: encoder on .cpuAndNeuralEngine, vocoder on .cpuAndGPU. The combined pipeline lands in the 150–175ms range on A15 and newer. Don’t trust .all — measure and override.
The INT8 quantization cliff
Most teams get this wrong: INT8 is not uniformly safe across TTS layers. Quantizing the duration predictor to INT8 degrades prosody significantly — irregular pauses, flattened intonation, clipped phoneme boundaries. The quality cliff is sharp, not gradual.
| Layer Group | Safe Quantization | Notes |
|---|---|---|
| Text encoder | INT8 | Minimal perceptible impact |
| Duration predictor | FP16 only | INT8 breaks prosody |
| Mel decoder | INT8 | Acceptable with calibration |
| Vocoder upsampling | FP16 only | Audible artifacts at INT8 |
Mixed-precision lands the model at 35–55MB — well within the 80MB threshold I treat as the on-device viability ceiling for non-game apps.
Scheduling the audio buffer
AVAudioPlayerNode.scheduleBuffer(_:completionHandler:) is the right primitive. Keep a double-buffer: one chunk playing, one synthesizing. The completion handler triggers the next synthesis dispatch, keeping the audio thread hot without overrunning memory.
One thing to plan for: sustained ANE load will thermal-throttle over extended inference sessions. Apple’s Core ML performance documentation covers the scheduling behavior and is the authoritative source on current thresholds, which vary by chip generation. In practice this means designing synthesis as burst-plus-pause rather than a continuous stream — something that matters especially for accessibility tooling and hands-free workflows, where thermal headroom is a real constraint, not an edge case.
3 actionable takeaways
-
Split your model across compute units. Encoder on ANE, vocoder on GPU. Measure both paths independently on your target device before trusting Core ML’s auto-scheduler —
.allis a starting point, not a final answer. -
Protect the duration predictor from INT8 quantization. It’s the highest-impact layer for prosody quality. Keep it at FP16 regardless of model size pressure; the perceptual cost of getting this wrong far outweighs the storage savings.
-
Stream mel chunks, not complete utterances. Chunked phoneme buffers with 50–80 frame slices are the architectural difference between 150ms and 400ms first-audio latency. Schedule the vocoder on the first chunk while the encoder processes the rest.
Related topics: ios swift mobile architecture