API và giao thức

Module AudioCommon định nghĩa các giao thức không phụ thuộc mô hình và các kiểu dữ liệu dùng chung. Mọi mô hình tuân thủ các giao thức này có thể thay thế cho nhau qua các interface đó.

Tổng quan giao thức

┌─────────────────────────────────────────────────────────┐
│                    AudioCommon                          │
│                                                         │
│  AudioChunk          SpeechGenerationModel (TTS)        │
│  AlignedWord         SpeechRecognitionModel (STT)       │
│  SpeechSegment       ForcedAlignmentModel               │
│                      SpeechToSpeechModel                │
│                      VoiceActivityDetectionModel (VAD)   │
│                      TurnCompletionProvider (EOT)       │
│                      SpeakerEmbeddingModel              │
│                      SpeakerDiarizationModel            │
│                      SpeakerExtractionCapable           │
└─────────────────────────────────────────────────────────┘

SpeechRecognitionModel

Giao thức cho các mô hình chuyển giọng nói thành văn bản.

public protocol SpeechRecognitionModel: AnyObject {
    var inputSampleRate: Int { get }
    func transcribe(audio: [Float], sampleRate: Int, language: String?) -> String
    func transcribeWithLanguage(audio: [Float], sampleRate: Int, language: String?) -> TranscriptionResult
}

Các kiểu tuân thủ: Qwen3ASRModel, WhisperASRModel, ParakeetASRModel, ParakeetStreamingASRModel, OmnilingualASRModel (CoreML), OmnilingualASRMLXModel (MLX)

SpeechGenerationModel

Giao thức cho các mô hình chuyển văn bản thành giọng nói.

public protocol SpeechGenerationModel: AnyObject {
    var sampleRate: Int { get }
    func generate(text: String, language: String?) async throws -> [Float]
    func generateStream(text: String, language: String?) -> AsyncThrowingStream<AudioChunk, Error>  // has default impl
}

generateStream() có một bản triển khai mặc định bọc generate() thành một chunk duy nhất. Các mô hình có streaming thật sự (ví dụ Qwen3-TTS) sẽ ghi đè nó.

Các kiểu tuân thủ: Qwen3TTSModel, CosyVoiceTTSModel, VoxCPM2TTSModel, KokoroTTSModel, IndexTTS2TTSModel

IndexTTS2TTSModel thêm overload generate có âm thanh tham chiếu cho nhân bản zero-shot và dùng IndexTTS2SynthesisOptions để điều khiển tốc độ nói và khoảng nghỉ.

Nạp Qwen3-TTS cục bộ

fromLocal kiểm tra các tệp bắt buộc và tính đầy đủ của các shard checkpoint, sau đó suy ra kích thước, lượng tử hóa và kiến trúc từ config.json. configuration tùy chọn chỉ là giá trị dự phòng cho metadata bị thiếu; metadata sai hoặc mâu thuẫn sẽ phát sinh Qwen3TTSLoadingError.invalidConfiguration trước khi cấp phát mô hình.

public static func fromLocal(
    modelDirectory: URL,
    tokenizerDirectory: URL,
    configuration: Qwen3TTSConfig? = nil,
    wiredMemoryPolicy: Qwen3TTSWiredMemoryPolicy = .none,
    progressHandler: ((Double, String) -> Void)? = nil
) throws -> Qwen3TTSModel

public static func fromPretrained(
    modelId: String = Qwen3TTSModel.defaultModelId,
    tokenizerModelId: String = "Qwen/Qwen3-TTS-Tokenizer-12Hz",
    cacheDir: URL? = nil,
    tokenizerCacheDir: URL? = nil,
    offlineMode: Bool = false,
    wiredMemoryPolicy: Qwen3TTSWiredMemoryPolicy = .pin(fraction: 0.9),
    progressHandler: ((Double, String) -> Void)? = nil
) async throws -> Qwen3TTSModel
Chính sách bộ nhớ

fromLocal mặc định dùng .none nên không đổi giới hạn wired-memory Metal toàn tiến trình. fromPretrained giữ mặc định .pin(fraction: 0.9). Cả hai bộ nạp pretrained đều nhận tokenizerCacheDir riêng.

ForcedAlignmentModel

Giao thức cho căn chỉnh dấu thời gian ở cấp từ.

public protocol ForcedAlignmentModel: AnyObject {
    func align(audio: [Float], text: String, sampleRate: Int, language: String?) -> [AlignedWord]
}

SpeechToSpeechModel

Giao thức cho các mô hình hội thoại giọng nói tới giọng nói.

public protocol SpeechToSpeechModel: AnyObject {
    var sampleRate: Int { get }
    func respond(userAudio: [Float]) -> [Float]
    func respondStream(userAudio: [Float]) -> AsyncThrowingStream<AudioChunk, Error>
}

Các kiểu tuân thủ: PersonaPlexModel

VoiceActivityDetectionModel

Giao thức cho phát hiện hoạt động giọng nói.

public protocol VoiceActivityDetectionModel: AnyObject {
    var inputSampleRate: Int { get }
    func detectSpeech(audio: [Float], sampleRate: Int) -> [SpeechSegment]
}

TurnCompletionProvider

Giao thức cho các bộ phân loại kết thúc lượt nói mà StreamingVADProcessor hỏi ở mỗi khoảng dừng VAD đã xác nhận. VAD chỉ nghe được khoảng lặng; mô hình hoàn tất lượt nói lắng nghe ngữ điệu của cả phát ngôn, nên khoảng dừng giữa câu khiến tác nhân chờ, còn câu đã nói xong nhận được phản hồi ngay. Tương ứng với sc_turn_completion_vtable_t của speech-core. Xem hướng dẫn Smart Turn.

public protocol TurnCompletionProvider: AnyObject {
    /// Probability in [0, 1] that the turn is complete, given the audio of the
    /// turn so far (Smart Turn looks at the last 8 s).
    func turnCompleteProbability(audio: [Float], sampleRate: Int) throws -> Float
}

Các kiểu tuân thủ: SmartTurnModel (Smart Turn v3.2)

SpeakerEmbeddingModel

Giao thức cho trích xuất embedding người nói.

public protocol SpeakerEmbeddingModel: AnyObject {
    var inputSampleRate: Int { get }
    var embeddingDimension: Int { get }
    func embed(audio: [Float], sampleRate: Int) -> [Float]
}

Các kiểu tuân thủ: WeSpeakerModel

SpeakerDiarizationModel

Giao thức cho các mô hình phân tách người nói, gán nhãn người nói cho từng đoạn âm thanh.

public protocol SpeakerDiarizationModel: AnyObject {
    var inputSampleRate: Int { get }
    func diarize(audio: [Float], sampleRate: Int) -> [DiarizedSegment]
}

Các kiểu tuân thủ: DiarizationPipeline (Pyannote), SortformerDiarizer

SpeakerExtractionCapable

Giao thức mở rộng cho các engine phân tách có hỗ trợ trích xuất các đoạn của một người nói mục tiêu dựa trên embedding tham chiếu. Không phải engine nào cũng hỗ trợ (Sortformer chạy end-to-end và không tạo ra embedding người nói).

public protocol SpeakerExtractionCapable: SpeakerDiarizationModel {
    func extractSpeaker(audio: [Float], sampleRate: Int, targetEmbedding: [Float]) -> [SpeechSegment]
}

Các kiểu tuân thủ: DiarizationPipeline (chỉ Pyannote)

Các kiểu dùng chung

AudioChunk

public struct AudioChunk {
    public let samples: [Float]   // PCM samples
    public let sampleRate: Int    // Sample rate (e.g. 24000)
}

SpeechSegment

public struct SpeechSegment {
    public let startTime: Float   // Start time in seconds
    public let endTime: Float     // End time in seconds
}

AlignedWord

public struct AlignedWord {
    public let text: String       // The word
    public let startTime: Float   // Start time in seconds
    public let endTime: Float     // End time in seconds
}

DiarizedSegment

public struct DiarizedSegment {
    public let startTime: Float   // Start time in seconds
    public let endTime: Float     // End time in seconds
    public let speakerId: Int     // Speaker identifier (0-based)
}

DialogueSegment

Một đoạn đã được phân tích từ văn bản hội thoại nhiều người nói, có thẻ người nói và cảm xúc (tùy chọn). Dùng cùng DialogueParser và DialogueSynthesizer cho tổng hợp hội thoại của CosyVoice3.

public struct DialogueSegment: Sendable, Equatable {
    public let speaker: String?   // Speaker identifier ("S1", "S2"), nil for untagged
    public let emotion: String?   // Emotion tag ("happy", "whispers"), nil if none
    public let text: String       // Cleaned text to synthesize
}

DialogueParser

Phân tích văn bản hội thoại nhiều người nói với thẻ người nói inline ([S1]) và thẻ cảm xúc ((happy)).

public enum DialogueParser {
    static func parse(_ text: String) -> [DialogueSegment]
    static func emotionToInstruction(_ emotion: String) -> String
}

Các cảm xúc dựng sẵn: happy/excited, sad, angry, whispers/whispering, laughs/laughing, calm, surprised, serious. Các thẻ không xác định sẽ được truyền nguyên dạng như chỉ thị tự do.

DialogueSynthesizer

Điều phối tổng hợp hội thoại nhiều đoạn với nhân bản giọng theo từng người nói, khoảng lặng giữa các lượt và crossfade.

public enum DialogueSynthesizer {
    static func synthesize(
        segments: [DialogueSegment],
        speakerEmbeddings: [String: [Float]],
        model: CosyVoiceTTSModel,
        language: String,
        config: DialogueSynthesisConfig,
        verbose: Bool
    ) -> [Float]
}

DialogueSynthesisConfig

public struct DialogueSynthesisConfig: Sendable {
    public var turnGapSeconds: Float      // Default: 0.2
    public var crossfadeSeconds: Float    // Default: 0.0
    public var defaultInstruction: String // Default: "You are a helpful assistant."
    public var maxTokensPerSegment: Int   // Default: 500
}

PipelineLLM

Giao thức để tích hợp mô hình ngôn ngữ vào các pipeline giọng nói. Cầu nối một LLM tới luồng ASR → LLM → TTS của VoicePipeline.

public protocol PipelineLLM: AnyObject {
    func chat(messages: [(role: MessageRole, content: String)],
              onToken: @escaping (String, Bool) -> Void)
    func cancel()
}

Adapter dựng sẵn: Qwen3PipelineLLM kết nối Qwen35MLXChat tới giao thức này kèm dọn dẹp token, hủy thao tác và gom cụm chờ.

AudioIO

Trình quản lý I/O âm thanh có thể tái sử dụng, loại bỏ boilerplate AVAudioEngine. Xử lý thu micro, resampling, phát lại và đo mức âm thanh.

let audio = AudioIO()
try audio.startMicrophone(targetSampleRate: 16000) { samples in
    pipeline.pushAudio(samples)
}
audio.player.scheduleChunk(ttsOutput)
audio.stopMicrophone()

AudioIO bao gồm StreamingAudioPlayer cho đầu ra TTS và AudioRingBuffer để truyền âm thanh an toàn theo luồng giữa luồng thu và luồng suy luận.

SystemAudioTap

Thu mix đầu ra của hệ thống — những gì máy Mac đang phát — dưới dạng Float32 mono ở tần số lấy mẫu do bên gọi chọn (process taps của Core Audio, macOS 14.4+). Tap mono toàn cục mặc định loại trừ tiến trình hiện tại, nên phần phát của chính ứng dụng không bao giờ bị thu lại, và tap được bọc trong một thiết bị tổng hợp riêng tư chỉ chứa tap đó.

let tap = SystemAudioTap()
try tap.start(targetSampleRate: 16000) { samples in
    pipeline.pushAudio(samples)
}
tap.stop()

Việc thu vẫn tiếp tục khi đổi thiết bị đầu ra mặc định — bộ resample nội bộ được dựng lại khi tần số mix thay đổi. Ứng dụng chủ phải khai báo NSAudioCaptureUsageDescription. Nếu quyền ghi âm bị từ chối, tap vẫn có thể được tạo nhưng chỉ trả về im lặng: hãy theo dõi framesCaptured tăng trong khi nonSilentFrames vẫn là 0 và chủ động báo lỗi.

Các nguồn độc lập có dấu thời gian

Cả hai lớp thu đều có overload kèm dấu thời gian. CapturedAudioChunk chứa samples mono ở sampleRate cùng hostTime Mach tùy chọn của frame đầu vào đầu tiên. Vì hai nguồn dùng chung đồng hồ host, ứng dụng có thể giữ PCM riêng biệt và sắp xếp các sự kiện bản chép lời phát sinh mà không trộn âm thanh.

try tap.startTimestamped(targetSampleRate: 16000) { systemChunk in
    systemPipeline.pushAudio(systemChunk.samples)
    recordSystemTime(systemChunk.hostTime)
}

let microphone = AudioIO(enableAEC: true, enablePlayback: false)
try microphone.startMicrophoneTimestamped(targetSampleRate: 16000) { micChunk in
    microphonePipeline.pushAudio(micChunk.samples)
    recordMicrophoneTime(micChunk.hostTime)
}

Với chế độ thu chỉ để nghe, enableAEC: true bật Apple Voice Processing trước khi đọc định dạng micrô, còn enablePlayback: false bỏ qua trình phát không dùng đến. Trên macOS, âm thanh phát từ thiết bị được loại khỏi mẫu micrô; nếu không thể khởi động AEC, lỗi sẽ được báo thay vì quay về đầu vào thô.

Đầu vào tệp và nhận dạng tăng dần

CapturedAudioChunk còn có frameIndex tăng đơn điệu và dấu isFinal. AudioFileLoader.stream tạo các chunk này theo yêu cầu với bộ nhớ giới hạn, resample bằng trạng thái bộ chuyển đổi liên tục và mặc định lấy trung bình tất cả kênh đầu vào. Dùng .first hoặc .select([indices]) khi đã biết cách định tuyến kênh.

StreamingRecognitionModel và StreamingRecognitionSession cung cấp một API tăng dần duy nhất, giữ nguyên cache cho Parakeet Streaming và Nemotron Core ML/MLX:

let source = AudioFileLoader.stream(
    url: inputURL,
    options: AudioFileStreamOptions(targetSampleRate: 16_000,
                                    chunkDuration: 0.32,
                                    channelSelection: .mixAll))
let session = try model.makeStreamingSession(language: "en-US")

for try await chunk in source {
    for update in try session.push(chunk) {
        render(update.text, final: update.isFinal)
    }
}
for update in try session.finish() {
    render(update.text, final: true)
}

SentencePieceModel

Trình đọc protobuf dùng chung cho các tệp .model của SentencePiece, nằm trong AudioCommon. Mọi module cần giải mã các piece SentencePiece (PersonaPlex, OmnilingualASR, các bản port ASR / TTS sau này) đều dựng decoder riêng dựa trên trình đọc duy nhất này thay vì cài lại định dạng wire của protobuf.

public struct SentencePieceModel: Sendable {
    public struct Piece: Sendable, Equatable {
        public let text: String
        public let score: Float
        public let type: Int32
        public var pieceType: PieceType? { get }
        public var isControlOrUnknown: Bool { get }
    }
    public enum PieceType: Int32 {
        case normal = 1, unknown = 2, control = 3,
             userDefined = 4, unused = 5, byte = 6
    }
    public let pieces: [Piece]
    public var count: Int { get }
    public subscript(_ id: Int) -> Piece? { get }
    public init(contentsOf url: URL) throws
    public init(modelPath: String) throws
    public init(data: Data) throws
}

Được dùng bởi: OmnilingualASR.OmnilingualVocabulary, PersonaPlex.SentencePieceDecoder. Được kiểm thử bằng 7 unit test trong Tests/AudioCommonTests/SentencePieceModelTests.

MLXCommon.SDPA

Các hàm hỗ trợ scaled dot-product attention dùng chung trên mọi module attention MLX (Qwen3-ASR / Qwen3-TTS / Qwen3-Chat / CosyVoice / PersonaPlex / OmnilingualASR). Mỗi module tự giữ projection của riêng mình — SDPA chỉ lo phần boilerplate reshape → attention → merge.

public enum SDPA {
    // Flat [B, T, H*D] input: project/reshape happens inside
    public static func multiHead(
        q: MLXArray, k: MLXArray, v: MLXArray,
        numHeads: Int, headDim: Int, scale: Float,
        mask: MLXArray? = nil
    ) -> MLXArray

    // GQA / MQA variant with separate query and KV head counts
    public static func multiHead(
        q: MLXArray, k: MLXArray, v: MLXArray,
        numQueryHeads: Int, numKVHeads: Int, headDim: Int, scale: Float,
        mask: MLXArray? = nil
    ) -> MLXArray

    // Already-shaped [B, H, T, D] (RoPE / KV cache paths)
    public static func attendAndMerge(
        qHeads: MLXArray, kHeads: MLXArray, vHeads: MLXArray,
        scale: Float,
        mask: MLXArray? = nil
    ) -> MLXArray

    // Same, with ScaledDotProductAttentionMaskMode enum (newer API)
    public static func attendAndMerge(
        qHeads: MLXArray, kHeads: MLXArray, vHeads: MLXArray,
        scale: Float,
        mask: MLXFast.ScaledDotProductAttentionMaskMode
    ) -> MLXArray

    // Low-level head merge: [B, H, T, D] → [B, T, H*D]
    public static func mergeHeads(_ attn: MLXArray) -> MLXArray
}

Tất cả các lệnh reshape đều dùng -1 cho chiều batch, nhờ đó các hàm hỗ trợ này có thể kết hợp với các đồ thị MLX.compile(shapeless:) mà có batch thay đổi tại runtime (ví dụ Qwen3-TTS Talker giải mã tự hồi quy).

Máy chủ HTTP API

Tệp thực thi speech-server phơi bày mọi mô hình trong speech-swift dưới dạng endpoint HTTP REST và một endpoint WebSocket triển khai OpenAI Realtime API. Các mô hình được tải lười ở lần yêu cầu đầu tiên; truyền --preload để khởi động sẵn tất cả ngay khi bật.

swift build -c release
.build/release/speech-server --port 8080

# Tải sẵn mọi mô hình khi khởi động
.build/release/speech-server --port 8080 --preload

Các endpoint REST

EndpointMethodYêu cầuPhản hồi
/transcribePOSTBody audio/wavJSON { text } (Qwen3-ASR)
/v1/audio/transcriptionsPOSTForm multipart { file, model, response_format?, language?, prompt?, temperature? }JSON, văn bản, JSON chi tiết, SRT hoặc VTT
/v1/audio/speechPOSTJSON { model, input, voice, response_format?, speed? }Dữ liệu âm thanh audio/wav hoặc audio/pcm
/speakPOSTJSON { text, engine?, language?, voice? }Body audio/wav (Qwen3-TTS, CosyVoice, Kokoro)
/respondPOSTBody audio/wavBody audio/wav (PersonaPlex)
/enhancePOSTBody audio/wavBody audio/wav (DeepFilterNet3)
/vadPOSTBody audio/wavDanh sách JSON các đoạn
/diarizePOSTBody audio/wavDanh sách JSON DiarizedSegment
/embed-speakerPOSTBody audio/wavJSON [Float] (256 chiều)

/v1/audio/speech ánh xạ tên mô hình và giọng TTS tiêu chuẩn sang các engine cục bộ. Mặc định trả về WAV; response_format: "pcm" trả về âm thanh PCM16 mono 24 kHz, little-endian và không có header. Các định dạng nén không được hỗ trợ sẽ bị từ chối rõ ràng.

/v1/audio/transcriptions nhận một tệp WAV qua form multipart và mặc định trả về {"text":"..."}; response_format cũng có thể yêu cầu văn bản, JSON chi tiết, SRT hoặc VTT. Trên Apple, alias mô hình chọn engine ASR cục bộ. Máy chủ speech-core đóng gói cho Linux và Windows dùng Parakeet tải lười, nhận WAV PCM16/24/32 hoặc Float32 và giới hạn tải lên ở 25 MiB và 10 phút.

# Chuyển một tệp thành văn bản
curl -X POST http://localhost:8080/transcribe \
  --data-binary @recording.wav \
  -H "Content-Type: audio/wav"

# Tổng hợp giọng nói
curl -X POST http://localhost:8080/speak \
  -H "Content-Type: application/json" \
  -d '{"text": "Hello world", "engine": "cosyvoice"}' \
  -o output.wav

# OpenAI-compatible transcription
curl http://localhost:8080/v1/audio/transcriptions \
  -F "[email protected];type=audio/wav" \
  -F "model=whisper-1"

# OpenAI-compatible speech synthesis
curl http://localhost:8080/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{"model":"tts-1","voice":"alloy","input":"Hello world","response_format":"wav"}' \
  -o openai-output.wav

# Vòng giọng nói tới giọng nói đầy đủ
curl -X POST http://localhost:8080/respond \
  --data-binary @question.wav \
  -o response.wav

OpenAI Realtime API (/v1/realtime)

Endpoint WebSocket tại ws://host:port/v1/realtime triển khai giao thức OpenAI Realtime. Mọi tin nhắn đều là JSON với trường phân biệt type; payload âm thanh là PCM16 mã hóa base64 ở 24 kHz mono.

Trong khi tải mô hình lạnh hoặc sinh dài, máy chủ phát các sự kiện JSON nhẹ realtime.keepalive và frame điều khiển websocket pong khoảng mỗi 15 giây cho đến khi đầu ra mô hình sẵn sàng. Client có thể bỏ qua các sự kiện này hoặc dùng chúng làm chỉ báo hoạt động.

Sự kiện Client → Server

Sự kiệnMục đích
session.updateCấu hình engine, ngôn ngữ, giọng và định dạng âm thanh
input_audio_buffer.appendNối thêm một chunk PCM16 base64 vào buffer đầu vào
input_audio_buffer.commitCommit âm thanh đã đệm để chuyển thành văn bản
input_audio_buffer.clearLoại bỏ buffer đầu vào hiện tại
response.createYêu cầu tổng hợp TTS cho văn bản/chỉ thị được cung cấp

Sự kiện Server → Client

Sự kiệnÝ nghĩa
session.createdBắt tay xong, cấu hình mặc định đã được phát
session.updatedĐã xác nhận session.update gần nhất
input_audio_buffer.committedÂm thanh đã được chấp nhận và xếp hàng để chuyển thành văn bản
conversation.item.input_audio_transcription.completedKết quả ASR kèm văn bản chép cuối cùng
response.audio.deltaChunk PCM16 base64 của âm thanh đã tổng hợp
response.audio.doneKhông còn chunk âm thanh cho phản hồi này
response.donePhản hồi hoàn tất (metadata + thống kê độ trễ)
errorKhung lỗi với type và message

Lượt nói tự động bằng VAD phía máy chủ

input_audio_buffer.commit thủ công vẫn là mặc định. Đặt session.turn_detection thành server_vad để chạy Silero VAD trên máy chủ và tự động chép lời mỗi lượt nói đã kết thúc. Âm thanh tiền tố được giữ lại, bộ đệm khi rảnh được giới hạn, các sự kiện speech_started, speech_stopped và committed được gửi theo thứ tự, còn lượt nói liên tục sẽ bị buộc kết thúc ở giới hạn đã cấu hình.

ws.send(JSON.stringify({
  type: "session.update",
  session: {
    turn_detection: {
      type: "server_vad",
      threshold: 0.5,
      prefix_padding_ms: 300,
      silence_duration_ms: 500,
      max_turn_duration_ms: 120000
    }
  }
}));
const ws = new WebSocket('ws://localhost:8080/v1/realtime');

// ASR: push audio, request transcription
ws.send(JSON.stringify({ type: 'input_audio_buffer.append', audio: base64PCM16 }));
ws.send(JSON.stringify({ type: 'input_audio_buffer.commit' }));
// → conversation.item.input_audio_transcription.completed

// TTS: request synthesis and stream audio deltas
ws.send(JSON.stringify({
  type: 'response.create',
  response: { modalities: ['audio', 'text'], instructions: 'Hello world' }
}));
// → response.audio.delta (repeated), response.audio.done, response.done

Máy chủ nằm trong sản phẩm SPM AudioServer. Một client trình duyệt ví dụ được cung cấp tại Examples/websocket-client.html — mở nó song song với máy chủ đang chạy để vận hành toàn bộ vòng ASR + TTS.

Tải mô hình

Tất cả các mô hình được tải từ HuggingFace ở lần dùng đầu tiên và lưu trong ~/Library/Caches/qwen3-speech/. Module AudioCommon cung cấp HuggingFaceDownloader dùng chung để xử lý tải xuống, lưu cache và xác minh toàn vẹn.

Hủy hợp tác

Đối với phiên âm MLX, Qwen3ASRModel.transcribeCheckingCancellation(audio:sampleRate:options:) ném CancellationError khi phát hiện tác vụ bị hủy trước khi trích xuất đặc trưng, mã hóa, nạp trước bộ giải mã hoặc thực hiện một bước giải mã. Công việc đã gửi đến GPU có thể hoàn tất. Các API đồng bộ transcribe và transcribeBatch vẫn không ném lỗi và bỏ qua việc hủy tác vụ. speech-server sử dụng điểm vào hỗ trợ hủy cho Qwen3-ASR.