MVP Factory
ai startup development

On-device RAG on Android: MediaPipe + FAISS reranker

KW
Krystian Wiewiór · · 4 min read

Meta description: Build a full on-device RAG pipeline on Android using FAISS-lite int8 embeddings and a MediaPipe LLM cross-encoder reranker — with real memory budgets and latency on Pixel 9.


TL;DR

You can run a complete retrieval-augmented generation pipeline entirely on-device on a Pixel 9 (Snapdragon 8 Gen 3) with no network calls. The trick: int8 bi-encoder embeddings for ANN retrieval via FAISS-lite, followed by a quantized cross-encoder reranker loaded through MediaPipe LLM Inference API, all within a ~2.1 GB RAM budget. Bi-encoder retrieval costs ~18 ms; reranking top-10 candidates costs ~110 ms. Context window fits 1,800–2,200 tokens after you account for system prompt and retrieved chunks.


Why on-device RAG matters right now

Most teams get this wrong: they optimize the generative model and ignore retrieval quality entirely. The result is a fast LLM producing confident nonsense because the context window was stuffed with irrelevant chunks. A reranker fixes that — but running a cross-encoder on a server defeats the point of on-device inference.

With Snapdragon 8 Gen 3’s Hexagon NPU delivering ~45 TOPS and MediaPipe’s LLM Inference API supporting INT4/INT8 quantized models, the hardware is finally capable. Let me walk you through the architecture.


The two-stage pipeline architecture

Query
  │
  ▼
[int8 Bi-Encoder]  ──►  FAISS-lite ANN  ──►  Top-K Candidates (K=10)
                                                      │
                                                      ▼
                                           [Quantized Cross-Encoder]
                                           MediaPipe LLM Inference API
                                                      │
                                                      ▼
                                             Re-ranked Top-3 Chunks
                                                      │
                                                      ▼
                                             On-Device LLM Context

Stage 1: bi-encoder + FAISS-lite ANN retrieval

Use a 384-dimension int8 bi-encoder (e.g., a quantized MiniLM variant exported to TFLite). Embeddings are pre-computed and indexed offline. At query time, embed the query and run approximate nearest-neighbor search via FAISS-lite’s flat INT8 index.

val embeddingInterpreter = Interpreter(loadModelFile("bi_encoder_int8.tflite"))
val queryEmbedding = FloatArray(384)
embeddingInterpreter.run(tokenize(query), queryEmbedding)

val index = FaissIndex.load("corpus.index") // flat int8, ~12 MB for 50k chunks
val topK = index.search(queryEmbedding, k = 10)

Stage 2: cross-encoder reranking via MediaPipe

Load a quantized cross-encoder (INT4, ~180 MB on-disk) through MediaPipe LLM Inference API. For each of the top-10 candidates, score the (query, chunk) pair and sort descending.

val llmInference = LlmInference.createFromOptions(
    context,
    LlmInference.LlmInferenceOptions.builder()
        .setModelPath("/data/local/tmp/reranker_int4.bin")
        .setMaxTokens(512)
        .build()
)

val scores = topK.map { chunk ->
    val prompt = "Relevance score 0-10 for:\nQuery: $query\nPassage: $chunk\nScore:"
    llmInference.generateResponse(prompt).trim().toFloatOrNull() ?: 0f
}
val reranked = topK.zip(scores).sortedByDescending { it.second }.take(3)

Memory and latency budget on Pixel 9

ComponentModel Size (on-disk)RAM UsageLatency (Pixel 9)
int8 Bi-Encoder (TFLite)~22 MB~85 MB~18 ms
FAISS-lite Index (50k chunks)~12 MB~48 MB~4 ms
Cross-Encoder Reranker (INT4)~180 MB~420 MB~110 ms (10 pairs)
Main LLM (INT4, 1B param)~800 MB~1,400 MBvaries
Total~1.01 GB~1.95 GB~132 ms retrieval

Pixel 9 ships with 12 GB RAM. At ~1.95 GB for the full pipeline, you have comfortable headroom. The Snapdragon 8 Gen 3 NPU handles INT4 matmuls efficiently — the 110 ms reranker cost covers all 10 (query, chunk) forward passes.


Context window budget

After reranking, you pass the top-3 chunks to your main on-device LLM. With a 2,048-token context window:

  • System prompt: ~150 tokens
  • 3 chunks × ~350 tokens each: ~1,050 tokens
  • Query + response buffer: ~500 tokens
  • Remaining for generation: ~350 tokens

If your LLM supports 4,096 tokens (increasingly common in 1B–3B quantized models), you can expand to top-5 chunks and leave ~800 tokens for generation. That extra headroom moves the quality needle more than switching to a larger model.


What to know before you build this

  1. Two-stage retrieval is the right default. Bi-encoder ANN is fast but imprecise at the top; the cross-encoder is what separates actually relevant chunks from plausible-looking noise. At 110 ms for 10 pairs it’s not free, but bad context is a problem you can’t fix downstream.

  2. Budget your context window before you budget your model size. A lot of teams overspend on a larger LLM when the real bottleneck is retrieval quality and context packing. Three to five high-quality chunks beat ten mediocre ones every time.

  3. MediaPipe LLM Inference API handles this workload without drama. INT4 quantization and NPU delegation on Snapdragon 8 Gen 3 keep reranker latency under 200 ms — in practice, invisible behind the main LLM’s generation time.


android mobile architecture api kotlin


Share: Twitter LinkedIn