Real-time on-device OCR with CameraX: under 18ms
Meta description: Wire Android CameraX to a quantized CRAFT+CRNN OCR pipeline with NNAPI delegation and region batching — staying under 18ms on mid-range devices.
TL;DR
Combining CameraX with a two-stage quantized OCR pipeline (CRAFT for detection, CRNN for recognition) is doable on mid-range Android hardware. But it only works if you make deliberate choices at every layer: delegate selection, batching strategy, and frame scheduling. The naive implementation blows your latency budget in the first 50ms. This is how to build one that doesn’t.
The problem with naive OCR pipelines
Going on-device with a two-stage text detection and recognition pipeline carries real inference cost. On a Snapdragon 680-class device, an unoptimized pipeline runs at 60–120ms per frame. At 30fps you have ~33ms per frame delivery cycle; targeting 18ms inference leaves headroom for preprocessing and UI updates without dropping frames. The unoptimized pipeline blows that budget three times over.
To hit 18ms end-to-end on mid-range hardware, every architectural decision has to pull in the same direction.
Stage 1: NNAPI vs GPU delegate — know your hardware
The first decision point is delegate selection. This is where most teams go wrong — treating it as a single global configuration for the entire app.
CRAFT (text region detection) is a convolutional network with spatial operations that map well to GPU parallelism. CRNN (character sequence recognition) is recurrent-heavy and benefits more from NNAPI on devices where the DSP vendor driver is mature.
| Model | Preferred delegate | Reason |
|---|---|---|
| CRAFT (detection) | GPU delegate | Dense convolutions, spatial pooling |
| CRNN (recognition) | NNAPI (DSP path) | Sequential ops, lower memory bandwidth |
| CRNN fallback | CPU (XNNPACK) | NNAPI driver instability on some OEMs |
Benchmark both delegates on your target device tier. NNAPI performance varies significantly across OEM driver implementations — sometimes dramatically. Build a runtime delegate probe that runs a warm-up inference pass and selects based on measured latency, not assumptions.
val options = Interpreter.Options().apply {
val nnApiDelegate = NnApiDelegate()
addDelegate(nnApiDelegate)
setNumThreads(2)
}
// Probe: run 3 warm-up passes, measure median
Stage 2: Text region proposal batching
This is where naive implementations fall apart. If you run CRNN inference once per detected text region, you pay the interpreter initialization and memory transfer overhead on every single call. On a dense document that means 20–40 serial inference calls per frame.
The fix: batch all region proposals from CRAFT into a single CRNN inference pass. Pad or resize all candidate crops to a fixed input height (typically 32px), stack them into a batch tensor, run one forward pass.
// Batch crops from CRAFT output
val batchTensor = Array(regions.size) { i ->
preprocessCrop(regions[i], targetHeight = 32)
}
crnnInterpreter.runForMultipleInputsOutputs(
arrayOf(batchTensor), outputMap
)
On a Snapdragon 680 device with batches of 10–20 regions, batching reduces recognition latency by roughly 60–70% versus serial calls — primarily by eliminating repeated JNI overhead and keeping the accelerator warm across the batch. This isn’t a micro-optimization. It’s the difference between a usable demo and something you’d actually ship.
Stage 3: The CameraX frame pipeline and drop strategy
CameraX ImageAnalysis delivers frames at the camera’s native rate — typically 30fps. The default STRATEGY_KEEP_ONLY_LATEST drops all queued frames and always delivers the most recent one. That’s the right default. Don’t fight it.
val isProcessing = AtomicBoolean(false)
val imageAnalysis = ImageAnalysis.Builder()
.setBackpressureStrategy(ImageAnalysis.STRATEGY_KEEP_ONLY_LATEST)
.setOutputImageFormat(ImageAnalysis.OUTPUT_IMAGE_FORMAT_YUV_420_888)
.build()
.also { analysis ->
analysis.setAnalyzer(executor) { imageProxy ->
if (!isProcessing.compareAndSet(false, true)) {
imageProxy.close()
return@setAnalyzer
}
scope.launch(Dispatchers.Default) {
try {
runOcrPipeline(imageProxy)
} finally {
isProcessing.set(false)
imageProxy.close()
}
}
}
}
The AtomicBoolean gate is the non-obvious piece. STRATEGY_KEEP_ONLY_LATEST prevents frame queue buildup, but without an in-flight guard you can still dispatch a second coroutine before the first completes if the executor has available threads. The compareAndSet ensures only one inference pass runs at a time.
Use YUV_420_888 output and convert only the Y-plane (luminance) for CRAFT input. Color information isn’t needed for text region detection, and skipping the chroma planes cuts preprocessing cost by roughly 40% on the same Snapdragon 680 benchmark device.
Latency budget breakdown (mid-range target)
| Stage | Target time |
|---|---|
| YUV → grayscale crop | ~1ms |
| CRAFT detection (GPU) | ~8ms |
| Region proposal batching | ~1ms |
| CRNN recognition (NNAPI) | ~6ms |
| Result post-processing | ~1ms |
| Total | ~17ms |
Three things worth remembering
Benchmark delegate selection per model, not per app. CRAFT and CRNN have different computational profiles — profile them independently on representative mid-range hardware before committing to a strategy.
Batch your region proposals. Serial per-region inference is the single largest preventable latency source in a two-stage OCR pipeline. Batching isn’t an optimization; it’s a correctness requirement for real-time use.
Trust STRATEGY_KEEP_ONLY_LATEST and gate with an AtomicBoolean. Frame drop isn’t failure; it’s the mechanism that keeps your pipeline synchronized with real-world camera output. The backpressure strategy handles the queue. The atomic flag handles in-flight concurrency. You need both.
#android #mobile #architecture #kotlin