MVP Factory
ai startup development

On-device semantic search: Core ML ANE + quantized embeddings

KW
Krystian Wiewiór · · 4 min read

Meta description: Deploy a quantized MiniLM on iPhone via Core ML, force ANE scheduling, profile token throughput, and keep your HNSW index under the on-device memory ceiling.

Tags: ios swift mobile architecture api


TL;DR

You can run real-time semantic search entirely on-device by combining a quantized sentence-transformer (MiniLM-L6 INT8) with an HNSW index, all wired to the Apple Neural Engine via MLComputeUnits.cpuAndNeuralEngine. On iPhone 15 Pro, INT8 delivers ~18ms/query at recall@10 of ~0.91. Dropping to INT4 cuts that to ~11ms but recall@10 falls to ~0.83 — a tradeoff you need to measure, not assume. The real ceiling isn’t compute: it’s keeping your HNSW index under ~150MB to avoid jettison under memory pressure.


Why on-device semantic search is worth the complexity

In my experience building production systems that process sensitive user data, the moment you eliminate a round-trip to a cloud embedding API, you unlock a qualitatively different product. No network dependency, no latency spike on cellular, no privacy surface. The engineering tradeoffs are real, though, and most teams underestimate the ANE scheduling nuances before they’ve profiled a single inference.

Here’s the architecture.


The model pipeline: MiniLM-L6 INT8 via Core ML

The standard starting point is all-MiniLM-L6-v2 converted to Core ML with INT8 weight quantization using coremltools:

import coremltools as ct

model = ct.convert(
    traced_model,
    inputs=[ct.TensorType(name="input_ids", shape=(1, 128))],
    compute_precision=ct.precision.FLOAT16,
    minimum_deployment_target=ct.target.iOS17
)

spec = ct.optimize.coreml.linear_quantize_weights(
    model, config=ct.optimize.coreml.OptimizationConfig(
        global_config=ct.optimize.coreml.OpLinearQuantizerConfig(
            mode="linear_symmetric", dtype="int8"
        )
    )
)

On the Swift side, force ANE execution explicitly — do not leave this to the runtime scheduler:

let config = MLModelConfiguration()
config.computeUnits = .cpuAndNeuralEngine  // NOT .all — avoids GPU fallback

let model = try MLModel(contentsOf: modelURL, configuration: config)

Using .all risks silent GPU fallback during thermal throttling. .cpuAndNeuralEngine is more predictable on A17 Pro and M-series chips, though extreme sustained thermals can still affect ANE availability.


Profiling token throughput with Instruments

Use Xcode Instruments with the Core ML template, not Time Profiler — it exposes per-layer compute unit attribution and lets you verify ANE utilization rather than guessing.

Typical results on iPhone 15 Pro (A17 Pro, sequence length 64):

QuantizationAvg LatencyANE UtilizationRecall@10
FP32 (baseline)47ms~20%0.93
FP1628ms~55%0.92
INT818ms~82%0.91
INT411ms~78%0.83

INT8 is the pragmatic sweet spot for most applications. INT4’s recall degradation matters if your search corpus is domain-specific or your user queries are short — anecdotally, queries under 8 tokens appear more sensitive to quantization error relative to embedding variance, though this warrants measurement against your specific query distribution before drawing hard conclusions.


The HNSW index memory ceiling

Most teams get this wrong: the embedding model isn’t your memory problem. A 384-dimensional INT8 MiniLM is roughly 22MB. Your HNSW index is.

At 384 float32 dimensions per vector:

Corpus SizeHNSW Memory (M=16, ef=200)Fit on iPhone?
10K docs~23MBYes
50K docs~115MBMarginal
100K docs~230MBNo — jettison risk

The practical ceiling for a foreground app sits around 150MB total for the index before iOS memory pressure events start terminating background processes (tighter on older devices; profile on your minimum-supported hardware). At 50K documents you’re marginal; consider sharding the index or using product quantization (PQ) on the vectors stored in the index (distinct from weight quantization) to compress stored embeddings by 4–8x.

Use hnswlib via a Swift/C++ bridge or integrate usearch, which ships a native Swift API with INT8 vector storage support.


Three things worth getting right

  1. Always specify MLComputeUnits.cpuAndNeuralEngine explicitly. Silent GPU fallback during thermal events will sabotage your latency SLA in production and you won’t see it in the simulator.

  2. Measure recall@10 before shipping INT4. The 7ms latency gain is real, but so is the ~8-point recall drop. Run your actual query distribution against your actual corpus — benchmark on your data, not synthetic embeddings.

  3. Budget 150MB for the HNSW index, not the model. If your corpus exceeds ~40K documents, prototype PQ compression or index sharding early. Retrofitting memory architecture after launch is expensive.

The ANE is fast enough for production semantic search on-device. The discipline is in measurement, not assumption.


Share: Twitter LinkedIn