Clef-flash: local decisions in Swift
Development preview. Local text support is under development and has not been released in speech-swift.
Clef-flash reads a state and a set of questions, then scores the allowed answers together. The Swift module runs the 9B model locally on Apple Silicon with MLX and 4-bit weights. No Python process or hosted API is used during inference.
Use from Swift
import Clef
let model = try await Clef.fromPretrained()
let result = try model.decide(
state: "Please turn the kitchen lights on.",
questions: [
ClefQuestion(id: "action", instructions: "Which action?",
kind: .choice(["on": "Lights on", "off": "Lights off", "other": "Other"])),
ClefQuestion(id: "urgent", instructions: "Is this urgent?", kind: .noul())
])
print(result.decisions[0].probabilities)
Choice questions return probabilities for each option. A noul question returns the probability of true. A score question takes an ordered list such as .score(["normal", "urgent"]) and returns the expected zero-based index. The module does not execute commands.
Command line
speech clef decide request.json
speech clef decide request.json --model-dir /path/to/Clef-flash-9B-MLX-4bit
speech clef decide request.json --offline --max-tokens 4096
The default model is aufklarer/Clef-flash-9B-MLX-4bit. The first command downloads it if needed. Use a local directory or the offline flag to avoid a download.
Standalone command
swift build -c release --product clef-decide --disable-sandbox
scripts/build_mlx_metallib.sh release
.build/release/clef-decide /path/to/clef-flash-mlx-4bit request.json
{"state":"Turn the lights on.","questions":[
{"id":"action","type":"choice","instructions":"Which action?",
"choices":{"on":"Lights on","off":"Lights off","other":"Other"}}
]}
The output includes probabilities, input token count, and decision time excluding model load. The command accepts an ordered question array; it is not a SystemOne HTTP endpoint.
How it works
A Qwen3.5 backbone reads the text and schema. The joint decision head attends to that context, lets questions interact, and combines contextual scores with a prior from the model's output embeddings. It scores the options directly instead of generating a JSON answer token by token.
Requirements and limits
- Apple Silicon, macOS 15+, and MLX. The supported affine 4-bit export is approximately 5.3 GB; runtime memory is additional.
- Text only. Images, video, the 27B Clef model, and other weight formats are not supported.
- Use each instance serially. The default limit is 4096 encoded tokens, configurable to 16384; oversized requests fail rather than silently truncate the state.
- Probabilities are not calibrated guarantees. Keep command execution and confirmation separate from model scoring.
- Published hosted latency is not a measurement of this local Swift implementation.
Local validation
October 3 update: Two fresh runs of twenty warm repeats each averaged about 276 ms for the same 303-token, three-field request (observed range 274–284 ms), excluding model load. The earlier 291 ms run below supplied the separate RSS and physical-footprint measurements. This is one repeated text request; speech recognition, speech output, and visual inference were not measured.
On an M5 Pro with 48 GB unified memory, twenty warmed runs of one 303-token, three-field request averaged 0.291 seconds (median 0.290; range 0.280–0.299), excluding model load. Native GPU operations and 512-token chunks reduced the original 0.921-second mean by about 68%, with unchanged weights and selected answers. Peak process RSS was 5.11 GiB; peak macOS memory footprint was 6.81 GiB. RSS alone does not capture all unified-memory use. Longer inputs can need more memory. A later correctness run measured 462–897 ms while other applications were active, so the 291 ms mean is not a fixed latency guarantee. These are development measurements on one repeated request, not a broad quality benchmark. All 303 reference token IDs match, and the maximum probability difference from the Python reference is about 0.00266.