Clef-flash: local decisions in Swift

Development preview. Local text support is under development and has not been released in speech-swift.

Clef-flash reads a state and a set of questions, then scores the allowed answers together. The Swift module runs the 9B model locally on Apple Silicon with MLX and 4-bit weights. No Python process or hosted API is used during inference.

Use from Swift

import Clef
let model = try await Clef.fromPretrained()
let result = try model.decide(
    state: "Please turn the kitchen lights on.",
    questions: [
        ClefQuestion(id: "action", instructions: "Which action?",
            kind: .choice(["on": "Lights on", "off": "Lights off", "other": "Other"])),
        ClefQuestion(id: "urgent", instructions: "Is this urgent?", kind: .noul())
    ])
print(result.decisions[0].probabilities)

Choice questions return probabilities for each option. A noul question returns the probability of true. A score question takes an ordered list such as .score(["normal", "urgent"]) and returns the expected zero-based index. The module does not execute commands.

Command line

speech clef decide request.json
speech clef decide request.json --model-dir /path/to/Clef-flash-9B-MLX-4bit
speech clef decide request.json --offline --max-tokens 4096

The default model is aufklarer/Clef-flash-9B-MLX-4bit. The first command downloads it if needed. Use a local directory or the offline flag to avoid a download.

Standalone command

swift build -c release --product clef-decide --disable-sandbox
scripts/build_mlx_metallib.sh release
.build/release/clef-decide /path/to/clef-flash-mlx-4bit request.json
{"state":"Turn the lights on.","questions":[
  {"id":"action","type":"choice","instructions":"Which action?",
   "choices":{"on":"Lights on","off":"Lights off","other":"Other"}}
]}

The output includes probabilities, input token count, and decision time excluding model load. The command accepts an ordered question array; it is not a SystemOne HTTP endpoint.

How it works

A Qwen3.5 backbone reads the text and schema. The joint decision head attends to that context, lets questions interact, and combines contextual scores with a prior from the model's output embeddings. It scores the options directly instead of generating a JSON answer token by token.

Requirements and limits

Local validation

October 3 update: Two fresh runs of twenty warm repeats each averaged about 276 ms for the same 303-token, three-field request (observed range 274–284 ms), excluding model load. The earlier 291 ms run below supplied the separate RSS and physical-footprint measurements. This is one repeated text request; speech recognition, speech output, and visual inference were not measured.

On an M5 Pro with 48 GB unified memory, twenty warmed runs of one 303-token, three-field request averaged 0.291 seconds (median 0.290; range 0.280–0.299), excluding model load. Native GPU operations and 512-token chunks reduced the original 0.921-second mean by about 68%, with unchanged weights and selected answers. Peak process RSS was 5.11 GiB; peak macOS memory footprint was 6.81 GiB. RSS alone does not capture all unified-memory use. Longer inputs can need more memory. A later correctness run measured 462–897 ms while other applications were active, so the 291 ms mean is not a fixed latency guarantee. These are development measurements on one repeated request, not a broad quality benchmark. All 303 reference token IDs match, and the maximum probability difference from the Python reference is about 0.00266.

Sources