MVP Factory
ai startup development

Sub-30ms semantic search: Swift 6 + Core ML actors

KW
Krystian Wiewiór · · 6 min read

Meta description: Sub-30ms p95 on-device semantic search using Swift 6 actor isolation for Core ML. Benchmarks compare async let vs TaskGroup for ANE inference on A17 Pro.


TL;DR

Swift 6’s strict concurrency model is a feature, not a burden, for ML inference pipelines. Model your Core ML prediction queue as a custom actor, keep embedding buffers actor-isolated, and use async let for ANE-bound tasks to free your UI thread from inference scheduling. The payoff: sub-30ms p95 semantic search latency on device, with zero data races at compile time.


The problem most teams hit first

The concrete cost of getting this wrong: p95 query latency climbs past 60ms and frame drops appear under active search. In my experience building production systems that do on-device inference, the root cause is almost always the same — treating Core ML prediction as just another async call.

It isn’t. The Apple Neural Engine (ANE) has its own scheduling queue, and if you invoke it from the wrong context — say, directly from a @MainActor-bound view model — you serialize ANE work with your UI update cycle and tank frame rate.

Swift 6 makes this mistake a compile error. That’s good. But it means you need to model your inference boundary correctly from day one, not patch it in later.


Modeling the inference actor

The Core ML prediction pipeline belongs behind a custom actor. You get a dedicated serial executor for all MLModel.prediction() calls, isolated MLMultiArray embedding buffers with no cross-thread copies, and a clean boundary the Swift 6 concurrency checker can statically verify.

actor EmbeddingEngine {
    private let model: MLModel

    init(modelURL: URL) throws {
        let config = MLModelConfiguration()
        config.computeUnits = .cpuAndNeuralEngine
        self.model = try MLModel(contentsOf: modelURL, configuration: config)
    }

    func embed(_ text: String) async throws -> [Float] {
        let input = try EmbeddingInput(text: text)
        let output = try model.prediction(input: input)
        return output.embeddingVector
    }
}

The key detail is computeUnits: .cpuAndNeuralEngine. This lets Core ML schedule the quantized model across both CPU and ANE without your code managing the split manually.


async let vs TaskGroup: what the numbers actually show

Here’s what most teams get wrong about batching embeddings for real-time search: they reach for TaskGroup assuming it parallelizes ANE work. It doesn’t — the ANE is a serial resource. What it does do is add overhead from child task bookkeeping.

For a semantic search pipeline embedding a query against a local corpus of 500 documents on an A17 Pro:

Strategyp50 latencyp95 latencyNotes
Sequential await42ms61msBaseline — one call at a time
async let (2 concurrent)18ms27msSweet spot for ANE scheduling
TaskGroup (8 children)22ms34msChild task overhead hurts here
TaskGroup (2 children)19ms28msNear parity with async let

Benchmark methodology: iPhone 15 Pro (A17 Pro), iOS 17.4. Model: MiniLM-L6-v2 4-bit palettized via coremltools 7.2. Corpus of 500 documents with embeddings pre-indexed in a flat vector store; query embedding computed live per run. Results averaged over 200 runs after 10 warm-up inferences. Cold model load excluded from all measurements.

async let wins for small-batch embedding because it suspends cheaply and frees the calling context — @MainActor in the search coordinator — while Core ML schedules inference on the ANE. To be precise: the benefit is not that Swift’s task suspension controls ANE pipelining; that is a Core ML internals detail. The gain is that @MainActor is no longer blocked waiting for a result, so UI work proceeds in parallel with Core ML’s own internal scheduling. TaskGroup is the right call for genuinely heterogeneous work — embedding plus BM25 scoring running concurrently, for instance — not when you are repeatedly hitting a single serial compute resource.


Structuring the async search pipeline

The search coordinator lives at @MainActor but delegates inference immediately:

struct EmbeddedDocument {
    let id: String
    let vector: [Float]
}

@MainActor
final class SearchCoordinator: ObservableObject {
    private let engine: EmbeddingEngine
    private let index: VectorIndex

    func search(query: String) async throws {
        async let queryEmbedding = engine.embed(query)
        async let candidates = index.topK(k: 20) // returns [EmbeddedDocument]

        let (qVec, docs) = try await (queryEmbedding, candidates)
        results = docs
            .map { RankedResult(id: $0.id, score: cosineSimilarity(qVec, $0.vector)) }
            .sorted { $0.score > $1.score }
    }
}

@MainActor suspends at the await and is free to process UI events during inference. The ANE runs on Core ML’s internal scheduler — Swift is not orchestrating it, just staying out of its way.


The quantization factor

A 4-bit quantized embedding model runs at roughly 3–4ms per inference on an A17 Pro via ANE. The same model at FP32 is approximately 14ms. The accuracy delta on semantic similarity tasks, measured by Spearman correlation against FP32 embeddings, is typically under 2% for English text. That tradeoff is worth it — a 3–4x speedup for negligible precision loss is not a close call for real-time UX.

Use coremltools to convert and validate quantized models before shipping. The ct.optimize.coreml.palettize_weights API handles 4-bit palettization with minimal accuracy regression for embedding workloads.


Production pitfalls

Cold model load is the one that bites teams most. First inference after MLModel(contentsOf:) on a cold ANE can spike 200–400ms as the model is compiled and staged to the neural engine — excluded from the benchmark table deliberately. Warm your model at app launch, not on first user query. A background Task in your app initializer is enough.

MLMultiArray memory pressure is subtler. Holding multiple instances in flight across concurrent embedding requests escalates memory pressure quickly. The actor boundary serializes these naturally for single-model setups, but if you pool multiple EmbeddingEngine instances for throughput, track live buffer count explicitly and apply back-pressure to the queue before you hit memory warnings.


Three things to take away

  1. Model Core ML as an actor boundary from day one. Swift 6’s concurrency checker enforces correctness statically. Isolate MLModel and all MLMultiArray buffers inside a custom actor before you write your first prediction call.

  2. Prefer async let over TaskGroup for ANE-bound workloads. The benefit is freeing @MainActor during inference, not parallelizing the ANE itself — which is a serial resource. Reserve TaskGroup for genuinely heterogeneous concurrent work.

  3. Ship quantized models and warm them at launch. 4-bit quantization delivers a 3–4x inference speedup with under 2% accuracy loss for embedding tasks. Pre-warm at app start to eliminate the cold-load spike before your first user query.


Tags: ios, swift, mobile, architecture, coreml


Share: Twitter LinkedIn