Each clip is 1.50× faster than with MLX, with identical sampler settings, and a bulk job of 1,000 sentences 1.80× faster. It needs 73% less memory. Read from the same 8-bit file on both engines, it is 1.43× faster per request and 2.01× in bulk.
Mixed precision
Shanghai Jiao Tong University, with Cambridge and Geely · weights CC BY-NC 4.0, non-commercial only · model_1250000, shared reference WAV and transcript, seed 0 · One request: medians of 60 runs · Bulk: 1,000 LibriSpeech sentences, median of 5 passes
* PyTorch runs its own default sampler (Euler, 32 steps), not the RK4 settings MLX and NukeTorch share, so its column is not like-for-like. § NukeTorch’s batch API on an earlier build; on the current build it takes 1,771.53 s over two passes, 1.98× faster. MLX generating them one after another, as it ships. Median of five passes; the slowest NukeTorch pass was 1,963.20 s and the fastest MLX pass 3,454.79 s. PyTorch was not tested on the 1,000-sentence batch. † Everything each engine held at its peak, measured in runs separate from the timings. MLX keeps memory it has finished with in its own cache, so its bulk peak reflects its largest working set; its median over the second half of the bulk run was still 105,165 MB.
In detail
F5 produces audio by running a sampler that evaluates the model 28 times per request. A request takes 1,367 ms instead of 2,046 ms, with the sampler settings held identical on both sides. NukeTorch's timings come from paired runs in balanced order, alternating with a stock MLX worker in the same process; the MLX column is MLX measured on its own. In bulk the gain grows to 1.80×, NukeTorch’s batch API against MLX one request at a time.
What it means for you
F5 is slow on a Mac, so here the time is the point. On the test clip the audio comes back after 1.37 seconds instead of 2.05 under MLX: 1.5× faster on every generation. For voice cloning, narration or dubbing done in bulk, 1,000 sentences took 33 minutes instead of 59 under MLX, so work that takes an hour under MLX takes about 33 minutes. Transcription error rates were 0.05 points better: a median of 2.47% against 2.52% across 20,404 words in each of five passes.
On a single request F5 generates audio faster than it plays, at 1.19× real time against MLX's 0.80×, and in bulk at 3.25× real time. And its memory footprint drops from 3.2 GB to 0.9 GB, so it stops needing a well-specified machine. On the 1,000-sentence job the gap is starker: NukeTorch stays at 6.9 GB, while MLX grows to fill 106 GB of this 128 GB Mac.
Identical weights on both sides · F5-TTS 8-bit
Both engines were also run on MLX’s published 8-bit F5-TTS checkpoint, the same file read by both exactly as published, with the same sampler, texts, reference and seeds. NukeTorch is 1.43× faster than MLX on a single request (1,415 ms against 2,022 ms) and 2.01× faster across the 1,000 sentences (30 minutes against 61). Transcription error rates: 2.35% against 2.46%. People pick the 8-bit file to save memory, but on MLX it makes bulk jobs 1.04× slower than its FP32 model and still needs 2,167 MB for a single request. On NukeTorch the same file runs in 658 MB, and a 1,000-sentence job peaks at 1.6 GB against MLX’s 106 GB.
Both engines read MLX’s published 8-bit checkpoint, the same file, with the same sampler, texts, reference and seeds. One request: median of 60 calls. Bulk: median of 5 passes, NukeTorch’s batch API against MLX one request at a time.
Does the output still match
All passes returned 1,000 non-empty outputs. Both engines generate the same total audio length per pass, and transcription error rates are 0.05 points better: a median of 2.47% for NukeTorch against 2.52% for MLX, scored against the corpus text. Word error rate measures intelligibility only, not naturalness or speaker similarity.
Nuance
- NukeTorch and MLX use the same sampler settings.
- The bulk comparison is NukeTorch's default bulk mode against MLX one request at a time, as it ships. The median of five passes is shown.
- Timing boundaries differ slightly: MLX and PyTorch request timing includes reading the reference and preparing the text; NukeTorch starts at its speak call.
- Precision: NukeTorch runs FP16; MLX and PyTorch run the FP32 weights as shipped. With the same 8-bit checkpoint on both sides, NukeTorch is still 1.43× faster per request and 2.01× in bulk.
- Memory is everything each engine held at its peak, NukeTorch in FP16 against MLX and PyTorch in FP32. MLX’s 106 GB bulk peak partly reflects memory it keeps cached after use; on a smaller Mac it would likely use less and run slower.
- The whole clip returns at once, so the listener waits the full 1.4 seconds before hearing anything. We have not measured sentence-by-sentence streaming and make no claim about live conversation.
- Output lengths differ slightly between samplers: 39,181 samples for MLX and NukeTorch, 40,704 for PyTorch. Real-time factor is computed per request from actual output duration.
- Text outside English, or a reference that is not mono PCM16, falls back to an MLX worker rather than failing.
How to use it
macOS on Apple silicon · the same model files and checkpoint revision you use today · licensing undetermined