CameraX + Depth AI: Real-time spatial audio under 22ms
Meta description: Wire Android CameraX to a quantized MiDaS-small INT8 model and binaural audio rendering. Here’s the exact frame pipeline that stays under 22ms on mid-range Android devices.
TL;DR
Monocular depth + spatial audio is a useful pairing for real-world awareness apps. The pipeline: CameraX YUV frame → INT8 MiDaS-small via GPU delegate → depth-to-position mapping → AAudio binaural parameter update. The hard constraint is 22ms end-to-end on mid-range hardware. Miss it consistently and you get drift between visual and audio cues, perceptually jarring and functionally useless. This post shows exactly where the budget goes and how to defend it.
Why this problem is hard
Ocean surveillance startup Quartermaster, which raised $140M recently, calls their domain “the largest blind spot on Earth.” That framing transfers: real-time spatial awareness under tight compute, where latency isn’t a UX preference but a functional requirement.
On Android, 22ms maps to roughly 45fps, the minimum for spatial audio to feel coherent. The common failure mode: teams benchmark inference in isolation, ship to a mid-range device, and discover that capture + preprocessing ate 12ms before the model even started.
The 22ms budget, broken down
| Stage | Target Budget | Typical Overage Risk |
|---|---|---|
| CameraX YUV capture + callback | 3ms | Low |
| YUV → RGB + resize to 256×256 | 4ms | Medium (CPU path) |
| MiDaS-small INT8 inference (GPU) | 8ms | High on older GPUs |
| Depth map → stereo position params | 2ms | Low |
| AAudio parameter update | 2ms | Low |
| Total | 19ms | 3ms headroom |
You want 3ms of headroom, not zero. Thermal throttling on mid-range devices can push GPU inference from 8ms to 14ms under sustained load. Design for headroom, not the happy path.
CameraX frame extraction
Use ImageAnalysis with a non-blocking executor and STRATEGY_KEEP_ONLY_LATEST. This is non-negotiable — dropping frames is correct behavior when inference can’t keep up.
val analysis = ImageAnalysis.Builder()
.setTargetResolution(Size(640, 480))
.setBackpressureStrategy(ImageAnalysis.STRATEGY_KEEP_ONLY_LATEST)
.setOutputImageFormat(ImageAnalysis.OUTPUT_IMAGE_FORMAT_YUV_420_888)
.build()
analysis.setAnalyzer(inferenceExecutor) { imageProxy ->
processFrame(imageProxy) // owns close()
}
The YUV_420_888 format avoids an extra GPU copy compared to RGBA. On the preprocessing side, do the resize on the GPU using a RenderScript replacement (CameraX’s built-in effect pipeline or a Vulkan compute shader). A CPU bicubic resize at 640×480 → 256×256 will consistently blow your 4ms budget.
Quantized MiDaS-small via GPU delegate
INT8 quantization of MiDaS-small cuts the model from ~13MB to ~4MB and yields roughly 1.8× throughput on Adreno and Mali GPUs. Use the TFLite GPU delegate with experimentalBestEffort precision. The depth artifacts from FP16 fallback are imperceptible to the downstream audio stage.
val gpuDelegate = GpuDelegate(
GpuDelegate.Options().apply {
setPrecisionLossAllowed(true) // FP16 fallback on unsupported ops
setQuantizedModelsAllowed(true)
}
)
val options = Interpreter.Options().addDelegate(gpuDelegate)
val interpreter = Interpreter(loadModelFile(), options)
NNAPI is tempting but inconsistent across OEMs. In my experience building production systems, GPU delegate with setPrecisionLossAllowed(true) outperforms NNAPI on mid-range Snapdragon and MediaTek devices when you factor in driver variance. Benchmark both on your actual device matrix. The spread differs by chipset family and it’s wide enough to matter.
Depth-to-spatial audio mapping
The depth map from MiDaS is inverse relative depth, not metric distance. This trips people up. You’d expect closer objects to have higher values, and with MiDaS they do, but the scale is relative, not metric. For spatial audio, that’s actually fine — you only need relative positioning. Map the center-of-mass of the shallowest depth region to azimuth/elevation parameters for your HRTF renderer.
fun depthMapToAzimuth(depthMap: FloatArray, width: Int, height: Int): Float {
// Find centroid of pixels in top 10th percentile (closest objects)
val threshold = depthMap.max()!! * 0.9f
var weightedX = 0f; var totalWeight = 0f
depthMap.forEachIndexed { i, v ->
if (v >= threshold) { weightedX += (i % width) * v; totalWeight += v }
}
val normalizedX = (weightedX / totalWeight) / width // 0..1
return (normalizedX - 0.5f) * 180f // -90 to +90 degrees
}
Feed azimuth and a depth-derived distance parameter into AAudio with an HRTF convolution effect. OpenSL ES works, but AAudio’s lower-latency path is measurably better: typically 6-10ms vs. 12-20ms for the audio parameter propagation round-trip.
Before you ship
-
Drop frames deliberately.
STRATEGY_KEEP_ONLY_LATESTis your primary latency defense. An analyzer that processes every frame will cause ANRs on mid-range devices within minutes of sustained use. -
Benchmark on thermal-stressed hardware. Run your inference loop for 10 minutes before measuring. A model that hits 8ms cold will often hit 14ms hot. Design your budget around the steady-state number.
-
Prefer GPU delegate over NNAPI for cross-device consistency. NNAPI’s performance variance across OEM drivers makes it a poor default for shipping production builds. Profile both, but GPU delegate is the safer baseline.
#android #mobile #kotlin #architecture #jetpackcompose