MVP Factory
ai startup development

Wiring Android's Jetpack Compose to a Quantized On-Device LLM for Real-Time Code Completion: Token Streaming, Editor State Management, and the Latency Budget That Fits on Pixel 9

KW
Krystian Wiewiór · · 5 min read

TL;DR

Shipping real-time code completion on-device is achievable on flagship Android hardware, but only if you treat latency as a first-class constraint from day one. CodeGemma 2B INT4 gives you the model. StateFlow gives you the streaming pipeline. A carefully tuned debounce gate keeps inference from running wild. Get all three right and the editor feels fluid. Miss any one of them and it feels broken.


Why on-device, and why now

Cloud-hosted completions solve latency by throwing bandwidth at it — until the user is on a plane, behind a corporate proxy, or simply unwilling to send their proprietary code to a remote endpoint. The Pixel 9 series, with its dedicated NPU and 12GB RAM headroom, finally makes the hardware argument for local inference credible.

A quantized 2B-parameter model at INT4 precision occupies roughly 1-1.2GB on disk and fits comfortably in GPU VRAM on current flagship hardware, leaving enough headroom for the host app. The target: p95 first-token latency under 200ms. Below that threshold, completion feels responsive. Above it, users abandon the suggestion before it finishes rendering.


Model selection: CodeGemma 2B INT4

Not all small models are created equal for code. CodeGemma 2B, fine-tuned on code corpora and exported via TFLite with INT4 weight quantization, lands in a useful range — small enough to fit the memory budget, fast enough to meet the latency target, and accurate enough to be worth showing:

ModelSize (INT4)p50 First Token (Pixel 9)Code MBPP Pass@1
CodeGemma 2B INT4~1.1 GB~120msCompetitive for 2B class
Generic 3B INT8~3.0 GB~280msLower code-specific accuracy
7B INT4~4.2 GB>500msOut-of-budget for real-time

The INT4 quantization is non-negotiable here. INT8 doubles your memory footprint and pushes you past the latency ceiling on anything but the most extreme silicon.


The inference pipeline: StateFlow as the streaming backbone

Token streaming from a TFLite inference loop maps naturally onto a Kotlin StateFlow. Each decoded token gets emitted as a partial buffer update, and your Compose TextField or custom editor widget observes the flow with collectAsState().

// ViewModel
private val _completionTokens = MutableStateFlow("")
val completionTokens: StateFlow<String> = _completionTokens.asStateFlow()

fun streamCompletion(prompt: String) {
    viewModelScope.launch(Dispatchers.Default) {
        _completionTokens.value = ""
        inferenceEngine.streamTokens(prompt).collect { token ->
            _completionTokens.update { it + token }
        }
    }
}

In Compose, you wire this directly to your editor overlay:

val completion by viewModel.completionTokens.collectAsState()
// Render `completion` as a ghost-text overlay anchored to cursor position

StateFlow gives you conflation for free, which matters more than it might seem. If your UI frame drops, you never process a stale intermediate state — you always get the latest accumulated buffer.


Debounced trigger logic: avoiding inference storms

Here is what most teams get wrong about on-device completion: they fire inference on every keystroke. On a desktop with a cloud endpoint this is merely wasteful. On a device where inference consumes 15-20% CPU, it is catastrophic.

The trigger model that works in production:

  1. Debounce the cursor idle event — 150-250ms window after the last keystroke
  2. Gate on syntactic signal — only trigger after a ., (, space following a keyword, or newline
  3. Cancel in-flight inference on any new keystroke via Job.cancel()
editorState
    .onEach { cancelCurrentInference() }
    .debounce(180)
    .filter { isTriggerContext(it.cursorContext) }
    .collectLatest { state ->
        viewModel.streamCompletion(state.buildPrompt())
    }

collectLatest handles the cancellation semantics for you — any new emission cancels the previous coroutine, which propagates cancellation down into the inference loop.


GPU delegate vs. NNAPI: the decision tree

The delegate selection deserves its own logic layer. In my experience building production systems with TFLite, the naive “always prefer GPU” strategy breaks on mid-range devices and causes silent fallback latency spikes.

Is device flagship-tier (Snapdragon 8 Gen 3 / Tensor G4+)?
├─ YES → Try GPU Delegate → if init < 2s, proceed
│         └─ FAIL → Fall back to NNAPI with INT8 cast
└─ NO  → Try NNAPI → check benchmark on first run
          └─ p50 > 350ms → fall back to CPU (disable feature)

Instrument delegate init time and first-inference latency at app startup, cache the result in SharedPreferences, and skip the expensive fallback detection on subsequent launches.


Compose editor state management

Maintain a separate EditorState data class that tracks cursor offset, visible line range, and the last accepted completion. This decouples inference triggers from the raw TextFieldValue and lets you build deterministic test cases for trigger logic without a running model.


Three things worth remembering

  1. Set the 200ms first-token budget as a hard constraint before writing a line of inference code. Profile your target device class early; if the model doesn’t fit the budget in a synthetic benchmark, no amount of Compose optimization will save it.

  2. Use collectLatest + debounce as your inference storm prevention layer. The cancellation semantics are built in; fighting keystroke latency at the Compose layer is the wrong battle.

  3. Instrument delegate selection and cache the result. Cold delegate initialization is a one-time penalty you can amortize; paying it on every inference is an architecture bug, not a hardware limitation.


Tags: android jetpackcompose mobile architecture kotlin


Share: Twitter LinkedIn