Sub-30ms semantic search: Swift 6 + Core ML actors
Meta description: Sub-30ms p95 on-device semantic search using Swift 6 actor isolation for Core ML. Benchmarks compare async let vs TaskGroup for ANE inference on A17 Pro.
TL;DR
Swift 6’s strict concurrency model is a feature, not a burden, for ML inference pipelines. Model your Core ML prediction queue as a custom actor, keep embedding buffers actor-isolated, and use async let for ANE-bound tasks to free your UI thread from inference scheduling. The payoff: sub-30ms p95 semantic search latency on device, with zero data races at compile time.
The problem most teams hit first
The concrete cost of getting this wrong: p95 query latency climbs past 60ms and frame drops appear under active search. In my experience building production systems that do on-device inference, the root cause is almost always the same — treating Core ML prediction as just another async call.
It isn’t. The Apple Neural Engine (ANE) has its own scheduling queue, and if you invoke it from the wrong context — say, directly from a @MainActor-bound view model — you serialize ANE work with your UI update cycle and tank frame rate.
Swift 6 makes this mistake a compile error. That’s good. But it means you need to model your inference boundary correctly from day one, not patch it in later.
Modeling the inference actor
The Core ML prediction pipeline belongs behind a custom actor. You get a dedicated serial executor for all MLModel.prediction() calls, isolated MLMultiArray embedding buffers with no cross-thread copies, and a clean boundary the Swift 6 concurrency checker can statically verify.
actor EmbeddingEngine {
private let model: MLModel
init(modelURL: URL) throws {
let config = MLModelConfiguration()
config.computeUnits = .cpuAndNeuralEngine
self.model = try MLModel(contentsOf: modelURL, configuration: config)
}
func embed(_ text: String) async throws -> [Float] {
let input = try EmbeddingInput(text: text)
let output = try model.prediction(input: input)
return output.embeddingVector
}
}
The key detail is computeUnits: .cpuAndNeuralEngine. This lets Core ML schedule the quantized model across both CPU and ANE without your code managing the split manually.
async let vs TaskGroup: what the numbers actually show
Here’s what most teams get wrong about batching embeddings for real-time search: they reach for TaskGroup assuming it parallelizes ANE work. It doesn’t — the ANE is a serial resource. What it does do is add overhead from child task bookkeeping.
For a semantic search pipeline embedding a query against a local corpus of 500 documents on an A17 Pro:
| Strategy | p50 latency | p95 latency | Notes |
|---|---|---|---|
| Sequential await | 42ms | 61ms | Baseline — one call at a time |
| async let (2 concurrent) | 18ms | 27ms | Sweet spot for ANE scheduling |
| TaskGroup (8 children) | 22ms | 34ms | Child task overhead hurts here |
| TaskGroup (2 children) | 19ms | 28ms | Near parity with async let |
Benchmark methodology: iPhone 15 Pro (A17 Pro), iOS 17.4. Model: MiniLM-L6-v2 4-bit palettized via
coremltools 7.2. Corpus of 500 documents with embeddings pre-indexed in a flat vector store; query embedding computed live per run. Results averaged over 200 runs after 10 warm-up inferences. Cold model load excluded from all measurements.
async let wins for small-batch embedding because it suspends cheaply and frees the calling context — @MainActor in the search coordinator — while Core ML schedules inference on the ANE. To be precise: the benefit is not that Swift’s task suspension controls ANE pipelining; that is a Core ML internals detail. The gain is that @MainActor is no longer blocked waiting for a result, so UI work proceeds in parallel with Core ML’s own internal scheduling. TaskGroup is the right call for genuinely heterogeneous work — embedding plus BM25 scoring running concurrently, for instance — not when you are repeatedly hitting a single serial compute resource.
Structuring the async search pipeline
The search coordinator lives at @MainActor but delegates inference immediately:
struct EmbeddedDocument {
let id: String
let vector: [Float]
}
@MainActor
final class SearchCoordinator: ObservableObject {
private let engine: EmbeddingEngine
private let index: VectorIndex
func search(query: String) async throws {
async let queryEmbedding = engine.embed(query)
async let candidates = index.topK(k: 20) // returns [EmbeddedDocument]
let (qVec, docs) = try await (queryEmbedding, candidates)
results = docs
.map { RankedResult(id: $0.id, score: cosineSimilarity(qVec, $0.vector)) }
.sorted { $0.score > $1.score }
}
}
@MainActor suspends at the await and is free to process UI events during inference. The ANE runs on Core ML’s internal scheduler — Swift is not orchestrating it, just staying out of its way.
The quantization factor
A 4-bit quantized embedding model runs at roughly 3–4ms per inference on an A17 Pro via ANE. The same model at FP32 is approximately 14ms. The accuracy delta on semantic similarity tasks, measured by Spearman correlation against FP32 embeddings, is typically under 2% for English text. That tradeoff is worth it — a 3–4x speedup for negligible precision loss is not a close call for real-time UX.
Use coremltools to convert and validate quantized models before shipping. The ct.optimize.coreml.palettize_weights API handles 4-bit palettization with minimal accuracy regression for embedding workloads.
Production pitfalls
Cold model load is the one that bites teams most. First inference after MLModel(contentsOf:) on a cold ANE can spike 200–400ms as the model is compiled and staged to the neural engine — excluded from the benchmark table deliberately. Warm your model at app launch, not on first user query. A background Task in your app initializer is enough.
MLMultiArray memory pressure is subtler. Holding multiple instances in flight across concurrent embedding requests escalates memory pressure quickly. The actor boundary serializes these naturally for single-model setups, but if you pool multiple EmbeddingEngine instances for throughput, track live buffer count explicitly and apply back-pressure to the queue before you hit memory warnings.
Three things to take away
-
Model Core ML as an actor boundary from day one. Swift 6’s concurrency checker enforces correctness statically. Isolate
MLModeland allMLMultiArraybuffers inside a custom actor before you write your first prediction call. -
Prefer
async letoverTaskGroupfor ANE-bound workloads. The benefit is freeing@MainActorduring inference, not parallelizing the ANE itself — which is a serial resource. ReserveTaskGroupfor genuinely heterogeneous concurrent work. -
Ship quantized models and warm them at launch. 4-bit quantization delivers a 3–4x inference speedup with under 2% accuracy loss for embedding tasks. Pre-warm at app start to eliminate the cold-load spike before your first user query.
Tags: ios, swift, mobile, architecture, coreml