Skip to content

Commit 82a0e5e

Browse files
Adding Parakeet Export and Runtime (#136)
Parakeet export and runtime with unit tests and parity test infra. Small whisper fixes. Rename SpeechModel to SpeechRecognitionModel, among similar renames.
1 parent bfdd41c commit 82a0e5e

25 files changed

Lines changed: 5040 additions & 550 deletions

Package.swift

Lines changed: 6 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -176,13 +176,13 @@ let package = Package(
176176
]
177177
),
178178
.executableTarget(
179-
name: "speech-runner",
179+
name: "speech-recognizer",
180180
dependencies: [
181181
"CoreAISpeech",
182182
"CoreAIShared",
183183
.product(name: "ArgumentParser", package: "swift-argument-parser"),
184184
],
185-
path: "swift/Sources/Tools/speech-runner",
185+
path: "swift/Sources/Tools/speech-recognizer",
186186
swiftSettings: [
187187
.enableUpcomingFeature("MemberImportVisibility")
188188
]
@@ -268,7 +268,10 @@ let package = Package(
268268
),
269269
.testTarget(
270270
name: "SpeechTests",
271-
dependencies: ["CoreAISpeech"],
271+
dependencies: [
272+
"CoreAISpeech",
273+
"TestUtilities",
274+
],
272275
path: "swift/Tests/SpeechTests"
273276
),
274277
],

models/README.md

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -149,6 +149,7 @@ uv run models/<name>/export.py
149149
### Audio Models
150150

151151
- [CLAP](clap)
152+
- [Parakeet TDT](parakeet)
152153
- [Wav2Vec 2.0](wav2vec2)
153154
- [Whisper](whisper)
154155

models/parakeet/README.md

Lines changed: 84 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,84 @@
1+
# Parakeet TDT
2+
3+
Automatic speech recognition (ASR) model from NVIDIA. Pairs a FastConformer encoder with a Token-and-Duration Transducer (TDT) decoder that predicts `(token, duration)` pairs each step, letting greedy decoding skip blank-only frames for ~2-4x faster inference vs. standard RNN-T.[^1]
4+
5+
## Setup
6+
7+
If you haven't installed `uv`, install it by
8+
9+
```bash
10+
brew install uv
11+
```
12+
13+
## Export
14+
15+
```sh
16+
uv run export.py
17+
```
18+
19+
Saves a bundle directory at `<repo-root>/exports/<model>_<dtype>_<static|dynamic>/` containing three `.aimodel` assets (`encoder`, `decoder_step`, `joint`), the processor (feature extractor + tokenizer), and a `metadata.json` describing the bundle. Pass `--output-dir <path>` to override the destination.
20+
21+
```sh
22+
uv run export.py --help
23+
```
24+
25+
**Options:**
26+
27+
| Flag | Description | Default |
28+
| ------------------ | ---------------------------------------------- | ----------------------------- |
29+
| `--model` | Model variant | `nvidia/parakeet-tdt-0.6b-v3` |
30+
| `--output-dir` | Output directory for the bundle | `<repo-root>/exports/` |
31+
| `--dtype` | `float16`, `float32` | `float32` |
32+
| `--dynamic` | Encoder accepts variable audio length | static (5s default) |
33+
| `--audio-seconds` | Length of dummy audio for static encoder trace | `5.0` |
34+
| `--overwrite` | Overwrite existing bundle ||
35+
36+
**Supported models:**
37+
38+
| Model | Parameters |
39+
| ----------------------------- | ---------- |
40+
| nvidia/parakeet-tdt-0.6b-v3 | 0.6B |
41+
42+
## Running
43+
44+
### In your iOS and macOS applications
45+
46+
```swift
47+
import CoreAISpeech
48+
49+
// Load an exported bundle directory (metadata.json + encoder/decoder_step/joint .aimodel assets + processor/).
50+
let model = try await SpeechRecognitionModel(resourcesAt: "coreai-models/exports/parakeet-tdt-0.6b-v3_float32_static")
51+
52+
// Transcribe an audio file — decoded and resampled to the model's sample rate automatically:
53+
let (text, stats) = try await model.transcribe(audioURL: URL(fileURLWithPath: "audio.wav"))
54+
print(text)
55+
56+
// Or transcribe raw mono PCM you already hold at model.sampleRate:
57+
let (text2, _) = try await model.transcribe(pcm: pcmSamples)
58+
```
59+
60+
### On your Mac using built-in Command Line Tool
61+
62+
```bash
63+
swift run -c release speech-recognizer --model path/to/exported_bundle_dir --audio-path path/to/audio.wav
64+
```
65+
66+
Accepts any audio the system can decode (`wav`, `flac`, `m4a`, …). Add `--warmup` to run a full transcription pass (encode + decode) on silence before timing, or `--verbose` for debug output. Omit the audio file to run a silence latency benchmark.
67+
68+
## Why three graphs?
69+
70+
Parakeet TDT's runtime decoding is autoregressive with duration-aware time advancement: each step samples a `(token, duration)` pair from the joint network, then advances the encoder frame pointer by `duration` (and only runs the LSTM prediction net when the token is not blank). That control flow lives in `ParakeetTDTGenerationMixin.generate`, not in `forward`, so `torch.export` cannot capture it as a single graph. The bundle exposes the three building blocks the runtime needs:
71+
72+
| Graph | Inputs | Outputs |
73+
| -------------- | ------------------------------------------------------------------- | -------------------------------------------------- |
74+
| `encoder` | `input_features (B, T_audio, n_mels)` | `encoder_hidden_states (B, T_enc, decoder_hidden)` |
75+
| `decoder_step` | `input_ids (B, 1)`, `hidden_state`, `cell_state` | `decoder_output`, `new_hidden_state`, `new_cell_state` |
76+
| `joint` | `decoder_hidden_states (B, 1, H)`, `encoder_hidden_states (B, 1, H)` | `logits (B, 1, vocab + len(durations))` |
77+
78+
The encoder graph already includes `encoder_projector`, so the joint network's two addends share the same hidden size.
79+
80+
## Streaming
81+
82+
This recipe exports the full-utterance encoder; cache-aware / chunked-attention streaming is not yet implemented in `transformers` for Parakeet. The `decoder_step` and `joint` graphs are already streaming-shaped (single-step, explicit LSTM state in/out), so once a chunked encoder lands upstream the same bundle layout extends to streaming with only an encoder swap.
83+
84+
[^1]: [TDT paper](https://arxiv.org/abs/2304.06795) · [Parakeet TDT v3 paper](https://arxiv.org/abs/2509.14128) · [HuggingFace](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3)

0 commit comments

Comments
 (0)