MVP Factory
ai startup development

Streaming On-Device LLMs to Compose with MediaPipe

KW
Krystian Wiewiór · · 5 min read

Meta: Wire MediaPipe’s LLM Inference API to Compose with Kotlin Flow: SharedFlow buffer sizing, JNI-to-coroutine dispatch, and context window compression for Pixel 8-class devices.


TL;DR

MediaPipe’s LLM Inference API gives you on-device Gemma and Phi-3 inference on Android, but the pipeline from JNI token callback to reactive Compose UI has several non-obvious failure modes: dropped tokens from under-buffered SharedFlows, UI jank from wrong dispatcher boundaries, and silent prompt truncation when you exceed the 1024–4096 token context window. This post walks through the production-grade wiring.


The architecture at a glance

When MediaPipe generates a token on-device, the callback originates from a JNI thread — not the main thread, not a coroutine dispatcher, but a raw native thread with no coroutine context. That token needs to travel through:

  1. JNI callback → Kotlin lambda
  2. Kotlin lambda → SharedFlow emission
  3. SharedFlow → StateFlow accumulation in ViewModel
  4. Compose recomposition on the main thread

Every boundary is a potential bottleneck or data loss point.


The JNI dispatcher boundary

Most teams get this wrong by emitting directly into a SharedFlow from the JNI callback thread without switching dispatchers first.

// ❌ Dangerous — emitting from unknown native thread
inferenceSession.generateAsync(prompt) { partialResult, _ ->
    _tokenFlow.tryEmit(partialResult) // silently drops tokens under load
}

// ✅ Correct — dispatch through a dedicated coroutine scope
inferenceSession.generateAsync(prompt) { partialResult, _ ->
    emissionScope.launch(Dispatchers.Default) {
        _tokenFlow.emit(partialResult) // suspends if buffer full
    }
}

The distinction between tryEmit and emit is load-bearing. tryEmit returns false and drops the token if the buffer is full — with no exception, no log, no signal. On a Pixel 8 running Gemma 2B at roughly 15–20 tokens per second, you will see drops under any meaningful UI load if you rely on tryEmit with a default buffer.


SharedFlow buffer sizing

Inference rate on device varies; the UI frame budget at 60fps does not. Your SharedFlow buffer must absorb burst emissions without stalling the native inference engine.

Device ClassApprox Tokens/secRecommended BufferOverflow Policy
Pixel 8 / Snapdragon 8 Gen 215–2564SUSPEND
Mid-range (Dimensity 700)5–1232SUSPEND
Low-end (< 4 GB RAM)2–616DROP_OLDEST
private val _tokenFlow = MutableSharedFlow<String>(
    replay = 0,
    extraBufferCapacity = 64,
    onBufferOverflow = BufferOverflow.SUSPEND
)

Use SUSPEND on capable hardware — backpressure should signal the emission scope to slow down, not silently lose data. On genuinely constrained devices where inference is already the bottleneck, DROP_OLDEST prevents unbounded coroutine queue growth at the cost of occasional visual glitching that is less harmful than an OOM.


Connecting to Compose state

Expose a StateFlow<String>, not the raw SharedFlow, to the UI layer. Raw SharedFlow collection in Compose can produce redundant recompositions when multiple collectors exist. Use runningFold to accumulate the token stream reactively:

val responseState: StateFlow<String> = _tokenFlow
    .runningFold("") { acc, token -> acc + token }
    .stateIn(viewModelScope, SharingStarted.Eagerly, "")

In your composable, a single collectAsState call on this StateFlow gives you stable, testable, lifecycle-aware streaming output with minimal boilerplate.


Context window: the ceiling you will hit

MediaPipe’s LLM Inference API enforces a hard context window configured at model initialization — typically 1024 to 4096 tokens depending on model variant and available device RAM. Exceeding this limit does not throw an exception. It silently truncates the prompt from the beginning.

In my experience building production systems with on-device inference, prompt compression is not optional for any multi-turn conversation feature. Strategies ranked by practical effectiveness:

  1. Rolling window — retain the last N tokens of history, discard oldest turns first
  2. Priority tagging — mark system instructions non-evictable, conversation turns evictable
  3. Extractive compression — summarize older turns with a heuristic before eviction
fun compressHistory(history: List<Message>, maxTokens: Int): List<Message> {
    var tokenCount = estimateTokens(systemPrompt)
    return history.reversed()
        .takeWhile { msg ->
            tokenCount += estimateTokens(msg.content)
            tokenCount <= maxTokens * 0.85 // 15% safety margin
        }
        .reversed()
}

The 15% safety margin is not cosmetic — character-based token estimation is approximate, and hitting the hard limit mid-generation produces corrupted partial output that is worse than a clean truncation.


Conclusion

On-device LLM inference on Android is production-ready, but the plumbing between MediaPipe’s native inference engine and a reactive Compose UI requires deliberate engineering at every layer. The JNI boundary, SharedFlow configuration, and context window ceiling are each capable of silently degrading your user experience in ways that won’t surface during early development.

Three things worth internalizing before shipping:

  • Use suspending emit over tryEmit from JNI callbacks, with at least 32–64 buffer capacity and SUSPEND overflow on capable hardware. Silent token loss is the hardest bug to diagnose in a streaming UI — there’s nothing to grep for.
  • Accumulate to StateFlow via runningFold in the ViewModel and never expose raw SharedFlow to Compose. A stable StateFlow eliminates redundant recompositions and makes your streaming logic unit-testable without a UI harness.
  • Build prompt compression before you need it. The 1024–4096 token ceiling arrives faster than expected in real conversations. The rolling window strategy takes an afternoon to implement; retrofitting it under user complaints does not.

Tags: android kotlin jetpackcompose mobile architecture


Share: Twitter LinkedIn