Tensorunfolded.

unfold.py
1def unfold(T: nuke.Tensor, mode: fp16) -> nuke.Tensor:
2    """U_Ω^(n)(T) = ∮ exp(i/ħ ∫ A_μ dx^μ) ⋆ (∇_ξ T ∧ J)"""

Neural Unfold Kernel Engine

NukeTorch starts from a different idea of the tensor, a new mathematical formulation we will publish, and builds its kernels from it. Same models, same weights, same output, less in the way of the hardware - the fastest AI inference engine. We started with Apple silicon, and we are measuring every step in public.

Decide, speak, listen. All three faster on a Mac.

NukeTorch runs open models of three kinds on Apple silicon. Each is measured against what people run today: the model maker’s own package for System 1, MLX for speech. Same model files, no changes to your code.

System 1 · Decider, Laya, Von
about 6×
Decider against PyTorch · 5.3× to 8.3× per request, 6.0× on 200 texts
  • Von 4.3× and Laya 2.2× faster on 1,000 texts
  • Decider 3.7× to 6.5× faster than llama.cpp per request
  • Up to 70% less memory for a request
  • The same answer as each maker’s package on at least 99.80% of 4,000 questions
Text to speech · Kokoro, F5-TTS, Fish Speech
about 3×
Kokoro against MLX · 2.5× per request, 3.4× in bulk
  • F5-TTS is 1.5× faster per clip
  • Fish Speech: 1.2× per clip and 4.4× in bulk, at the same precision as MLX
  • Kokoro needs 74% less memory and F5 73% less per request; on a 1,000-sentence F5 job, 6.9 GB against MLX’s 106 GB
Speech to text · Parakeet, Whisper
about 3×
Parakeet against MLX · 1.7× per clip, 3.0× on 1,000 clips
  • Whisper Tiny 2.5× faster on 1,000 clips
  • Whisper large-v3-turbo 1.8× faster on 1,000 clips, at the same precision as MLX
  • Parakeet needs 77% less memory than MLX on a bulk job
  • The same or a better word error rate than MLX on 1,000 real recordings

Measured on one Apple M4 Max with the same model files the authors publish. No results on NVIDIA hardware yet. Each model’s benchmark report lists what it does not show.

How it works

Where the speed comes from

Speed comes from Unfolding: NukeTorch builds its kernels from its own formulation of the tensor, so the hardware spends its time on arithmetic instead of waiting for the next operation.

PyTorch, eager the common default
MLX the fastest alternative on a Mac today
NukeTorch unfolded, less in between
Schematic, not measured. Coloured blocks are arithmetic; the gaps are the hardware waiting. Every number on this page is measured end to end.

Anticipating

Questions we expect

Including the ones that do not flatter us. What each set of results does and does not show is in each model’s benchmark report.

What is this not good for?

We're expanding to more models, but we're a small team with limited capacity, and right now every model has to be ported by us. As the team grows, we plan to make the framework much easier for others to use.

Is this just MLX with extra steps?

No. NukeTorch is a new engine, built from the ground up on our own mathematical principles. We pursue one question: what is a number? We used as few open-source libraries as we could; we don't even use NumPy in the engine, though some of our benchmark scripts and weight converters do. We learned from others, of course, so we won't claim every idea in our codebase is new. But we wrote every line ourselves, our intention was clear, and we exceeded our goal.

Why not contribute this upstream to MLX or PyTorch?

We'd love to, and we'd love to work closely with both teams. But for today, let us enjoy being the fastest engine on Apple silicon for every model we've measured.

How do I check any of this myself?

Every model has a benchmark report with the pinned checkpoints, the environment, the procedure, and the limits of each number. For now, we run private demos on request.

How is it licensed?

We're working on it. For now, it's undetermined.

What about NVIDIA, AMD, or TPU?

We'd love to work with them. Let's just say we have plans.

Do I have to change my model code?

Nope. But if you want every last bit of performance, we can work with you.

Will you publish the internals?

We'll share our position on this soon.

Decider runs up to 8.3× faster. Von up to 4.3×, Laya up to 2.2×. With up to 70% less memory.

System 1 models answer typed questions about a text with a probability for every answer, without generated text. Compared with each model maker’s own package on a Mac, NukeTorch picks the same answer on at least 99.80% of 4,000 questions.

Evidence, in one table

Three models, one Mac

Each model is compared with its maker’s own package on a Mac: on PyTorch for all three, and for Decider also through the package’s llama.cpp engine.

Model One request vs PyTorch, short / long text Batch vs PyTorch Memory, one request Evidence
Decider 2B 5.3× / 8.3× faster 6.0× faster 48% less Same precision
Laya 1.7× / 1.4× faster 2.2× faster 65% less Mixed precision
Von 1.3 4.1× / 1.7× faster 4.3× faster 70% less Mixed precision
One request

One text and 4 questions, from text in to a probability for every answer out. A short text and a long one, median of 30 calls each.

Batch

Laya and Von: 1,000 texts from AG News and IMDB, 4 questions each. Decider: the first 200 of them. NukeTorch’s batch call against the maker’s package as it runs.

Same answer

The share of the 4,000 questions on which an engine picks the same option as the maker’s package at full precision.

System 1 · model 1 of 3

Decider 2B

5.27× faster than the maker’s package on PyTorch for a short text and 8.29× for a long one, and 3.71× and 6.50× faster than its llama.cpp engine. A batch of 200 texts is 6.02× faster. It picks the same answer as the package at full precision on 99.975% of 4,000 questions.

Same precision

Mapika · decider-2b v11 · One request: one text and 4 questions, median of 30 calls · Batch: the first 200 texts, 800 questions

MetricNukeTorchllama.cppPyTorchvs llama.cppvs PyTorch
One request, short text (ms)88.72329.36467.943.71× faster5.27× faster
One request, long text (ms)421.452,737.913,494.406.50× faster8.29× faster
200 texts (s) §24.75109.98148.994.44× faster6.02× faster
Input tokens per second, 200 texts4,0999236814.44× faster6.02× faster
Same answer as the reference, 4,000 questions99.975%not scored99.950%n/a0.025 points better
Peak memory, one request, long text (MB) †4,1511,264 †8,029228% more48% less
Peak memory, 200 texts (MB) †4,1491,252 †8,369231% more50% less

§ The first 200 texts of the set, 4 questions each: NukeTorch’s batch call, median of 3 passes, against the package one request after another, one pass each. PyTorch is the maker’s package decider-ai on PyTorch (MPS, FP16, its Mac default); llama.cpp is the same package’s GGUF engine on the published BF16 file. The reference for the same answer is the package at full precision (PyTorch FP32). † Everything each engine held at its peak. llama.cpp’s figure counts only part of what it holds: it maps the model file into memory, and its footprint leaves those pages out.

Nuance
  • Precision: NukeTorch holds the published BF16 weights as FP16; the package runs FP16 on PyTorch, its Mac default, and BF16 on llama.cpp.
  • The batch is NukeTorch’s batch call against the package one request after another; the package has no in-process batch for these requests.
  • Label accuracy: topic 86.6% against 86.6%, sentiment 95.0% against 95.0%, text kind 99.1% against 99.2%, NukeTorch against the package at full precision.

How to use it

macOS on Apple silicon · the same published checkpoint as the maker’s package · licensing undetermined

Benchmark reportReproducing the Decider results
MachineApple M4 Max, 40 GPU cores, 128 GB, macOS 26.7
Reportv1.2. Author-run measurements; no independent reproduction yet.
One requestMedian of 30 timed calls per engine and text
BulkFirst 200 texts: NukeTorch median of 3 passes, the others one pass

Exactly what was run

MachineApple M4 Max, 40-core GPU, 128 GB unified memory, macOS 26.7
ModelMapika/decider-2b v11 (revision 533964da), BF16 safetensors; llama.cpp: Mapika/decider-2b-GGUF BF16 (revision ff2e5e68)
NukeTorchRelease build; the published model, its BF16 weights held as FP16
PyTorch MPSdecider-ai 1.8.1, torch 2.14.1, transformers 5.18.0, Decider(path): MPS, float16, patch_mps with its MLX Metal kernel (the package default on a Mac)
llama.cppdecider-ai GGUF engine, llama-cpp-python 0.3.36 built with Metal; one change to its bundled llama.cpp so that its BF16 kernels compile on this chip (Metal language 3.1): without it the BF16 file stops at its first matrix product
Questionstopic: choice; sentiment: score; positive: noul; review: noul

How each number is measured

Single requestModel resident, 3 untimed calls, then 30 timed calls per engine and text. The timer covers text in → tokens → model → probabilities out; model load is excluded.
BatchThe first 200 texts of the 1,000-text set (AG News test and IMDB test, interleaved), the same 4 questions each: 800 answers, 101,464 input tokens as the package counts them. NukeTorch receives the whole list through its batch call (median of 3 passes); the package has no in-process batch of such requests, so it receives them one after another (one pass each for PyTorch and llama.cpp).
Speed ratioThe other engine's time ÷ NukeTorch's.
MemoryProcess physical footprint sampled from outside the process (peak = the kernel's lifetime maximum; median over the second half of the run), the same probe for every engine, separate runs from the timing. llama.cpp maps the model file into memory, which its footprint does not count, so its figure is not comparable.
AgreementOn the 4,000 questions of the 1,000 texts, the option each engine chooses against the package at full precision (PyTorch FP32).

What this does not show

  • Thirty calls on two texts and a batch of the first 200 texts, with one set of 4 questions, on one machine.
  • The batch compares NukeTorch’s batch call with the package one request after another; PyTorch and llama.cpp ran it once.
  • llama.cpp’s memory figure counts only part of what it holds: its footprint leaves out the model file it maps into memory. Its answers were not scored.
  • No P95 or P99, no statistical significance claim.

System 1 · model 2 of 3

Laya

1.68× faster than the maker’s package on PyTorch for a short text and 1.42× for a long one, and 2.23× faster on 1,000 texts against the package’s fastest batch setting. 65% less memory for a request.

Mixed precision

Convai Innovations · Apache-2.0 · ModernBERT-large, 421M parameters · One request: one text and 4 questions, median of 30 calls · Batch: 1,000 texts, 4,000 questions, median of 5 passes

MetricNukeTorchPyTorchvs PyTorch
One request, short text (ms)27.1245.521.68× faster
One request, long text (ms)115.68164.041.42× faster
1,000 texts (s) §34.4176.602.23× faster
Texts per second, 1,000 texts29.0613.052.23× faster
Same answer as the reference, one request at a time99.95%FP32, the referencen/a
Same answer as the reference, 1,000 texts99.80%99.92%0.12 points worse
Peak memory, one request (MB) †1,1333,19665% less
Peak memory, 1,000 texts (MB) †3,2877,12754% less

§ NukeTorch’s batch API against the package’s predict_batch in its fastest setting, batches of 8 sorted by length; in input order the package takes 138.46 s and NukeTorch is 4.02× faster. The reference for the same answer is the package at full precision; PyTorch’s own figure is its Mac default. † Everything each engine held at its peak, measured in runs separate from the timings.

Nuance
  • Precision: NukeTorch runs the published FP16 weights natively. The package runs FP32 for fewer than 5 question rows in a forward pass and FP16 from 5 up, its Mac default, so its single requests run FP32.
  • The package pads each batch of 8 texts to its longest one, so its batch time depends on how the texts are grouped; both groupings it offers were measured, and the faster one is in the table.
  • Label accuracy is the same on every engine: topic 90.8%, sentiment 86.0%, text kind 95.9% at full precision.

How to use it

macOS on Apple silicon · the same published checkpoint as the maker’s package · licensing undetermined

Benchmark reportReproducing the Laya results
MachineApple M4 Max, 40 GPU cores, 128 GB, macOS 26.7
Reportv1.1. Author-run measurements; no independent reproduction yet.
One requestMedian of 30 timed calls per engine and text
Bulk1,000 texts: median of 5 passes per engine

Exactly what was run

ModelLaya, English checkpoint (Convai Innovations): ModernBERT-large encoder with a two-layer decision head, 421M parameters, Apache-2.0
Checkpointconvaiinnovations/laya@7b928d828b7b0e022f929d9bd2e44165aa270148, model.safetensors (842,609,210 bytes, SHA-256 891102d372688fc2a094dac56a384bc537b87c63f21f9f3dac0be2b7cbc8d86c), FP16 as published. Both engines read this one file
NukeTorchRelease build; the published FP16 weights, its native format
PyTorchThe maker's package laya 0.3.27 on PyTorch 2.14.1 and Transformers 5.18.0, Python 3.12.10; laya.load(directory, device="mps"), system_one for a single request, predict_batch for the batch; backend eager
PrecisionNukeTorch: FP16 weights; the encoder's matrix products read FP16 activations and sum in FP32; the residual stream, queries, keys, values and every norm's statistics stay FP32. PyTorch: the package's Mac default, FP32 for fewer than 5 question rows in a forward pass and FP16 autocast from 5 up; the full-precision reference is the same package with autocast off
Questionstopic (choice); sentiment (score); positive (noul); review (noul). Temperatures as shipped in the checkpoint's configuration
Texts1,000 texts: the first 500 of a seeded shuffle of AG News test (fancyzhx/ag_news@eb185aade064) and of IMDB test (stanfordnlp/imdb@e6281661ce1c), interleaved
Single-request textsThe first news item (400 input tokens over the 4 sequences) and the first review (2,048, every sequence at the model's 512-token limit)
MachineApple M4 Max, 40-core GPU, 128 GB unified memory, macOS 26.7, on AC power, no other user workload during timing
RunsRun 1: PyTorch (single requests, batches). Run 2: NukeTorch, same machine, same texts and order, same protocol
WarmupsSingle request: 3 untimed calls on both engines. Batch: one short untimed batch before the first pass

How each number is measured

Single requestMedian of 30 timed calls. No confidence interval or significance test is claimed.
TimersBoth engines: the text and the questions already in memory → the probabilities of every option. Model load is excluded.
BatchHeadline is the median of the pass times. Throughput = 1,000 ÷ wall time.
Batch executionNukeTorch's batch API receives the whole list and returns the answers in input order. PyTorch's predict_batch receives the whole list with batch_size=8, the value in the maker's own throughput benchmark; it pads each batch of 8 texts to its longest sequence, so the sort_by_length option, which groups similar lengths, is reported beside the input order. Other settings were tried on the first 200 texts before the series (best of two passes, sorted by length): batches of 8 took 15.48 s, batches of 4 texts 15.38 s, 16 texts 16.67 s, 32 texts 16.63 s; the package's full-precision setting took 17.68 s with batches of 8. None was more than 1% faster than the setting reported.
Speed ratioPyTorch's time ÷ NukeTorch's time. The percentage is (ratio − 1) × 100 and describes speed, not time saved. Ratios are computed from unrounded values.
AgreementA question agrees when the most probable option is the same as in the full-precision reference. Probability change is the largest absolute difference over the options of a question, after the checkpoint's temperature.
Warm-upNot timed. The passes of a series ran one after another in one process. Machine temperature was not recorded.
ScopeThirty calls on two texts and 5 passes over 1,000 English texts on one machine, with one set of 4 questions. They do not establish behaviour on other texts, question sets, languages or hardware.

What this does not show

  • Thirty calls on two texts and 5 passes over 1,000 English texts, with one set of 4 questions, on one machine.
  • The batch compares NukeTorch’s batch API with the package’s fastest batch setting; in input order the package is slower.
  • Label accuracy shows the engines agree; it is not a claim about the model’s quality.
  • No P95 or P99, no statistical significance claim.

System 1 · model 3 of 3

Von 1.3

4.15× faster than the maker’s package on PyTorch for a short text and 1.74× for a long one, and 4.28× faster on 1,000 texts. 70% less memory for a request.

Mixed precision

wfzyx · Apache-2.0 · ModernBERT-large, 395M parameters · One request: one text and 4 questions, median of 30 calls · Batch: 1,000 texts, 4,000 questions

MetricNukeTorchPyTorchvs PyTorch
One request, short text (ms)22.5193.394.15× faster
One request, long text (ms)326.07568.201.74× faster
1,000 texts (s) §33.18142.164.28× faster
Texts per second, 1,000 texts30.147.034.28× faster
Same answer as the reference, one request at a time99.92%99.80%0.12 points better
Same answer as the reference, 1,000 texts99.80%99.80%same
Peak memory, one request (MB) †1,0893,59170% less
Peak memory, 1,000 texts (MB) †3,2554,55529% less

§ NukeTorch’s batch API, median of 5 passes, against the package one request after another, as it runs (it has no batch call), median of 2 passes. The reference for the same answer is the package with its chain controller off; PyTorch’s own figure is the package as shipped, controller on. † Everything each engine held at its peak, measured in runs separate from the timings.

Nuance
  • Precision: the package runs the published FP32 weights in FP32; NukeTorch rounds them to FP16 when it loads them, its native format.
  • NukeTorch does not carry the package’s chain controller (extra passes for date and number facts); it changed the answers of 5 of the 1,000 texts.

How to use it

macOS on Apple silicon · the same published checkpoint as the maker’s package · licensing undetermined

Benchmark reportReproducing the Von results
MachineApple M4 Max, 40 GPU cores, 128 GB, macOS 26.7
Reportv1.1. Author-run measurements; no independent reproduction yet.
One requestMedian of 30 timed calls per engine and text
Bulk1,000 texts: NukeTorch median of 5 passes, PyTorch of 2

Exactly what was run

ModelVon 1.3 (weights 1.2): ModernBERT-large encoder with an option-marker scoring head, 395M parameters, Apache-2.0
Checkpointwfzyx/von@498ceba33390b32cfefaab6422ec380318ba9b99: model.safetensors (1,579,143,688 bytes, SHA-256 af57d5d2ab15715a753a1eb4add4271d1aecce7e76f3365f629c082df329297a, the encoder) and option_marker.pt (SHA-256 3faf27f88d30aaf9aa37860d4cdef99f1d05450cf40364f6236ac892d4d139ed: the same encoder tensors and the eight tensors of the head), FP32 as published
NukeTorchRelease build; reads the encoder from model.safetensors and the head's eight tensors, unchanged, from option_marker.pt; FP32 weights rounded to FP16 when loading
PyTorchThe maker's package von-sdk 1.3.7 on PyTorch 2.14.1 and Transformers 5.18.0, Python 3.12.10; OptionMarkerBackend(checkpoint_dir, device="mps").evaluate(state, questions); FP32; order-invariant options and the yes/no band rule as shipped; chain controller on (its default) for timing and memory
Questionstopic (choice); sentiment (score); positive (noul); review (noul). Temperature from the checkpoint's calibration map
Texts1,000 texts: the first 500 of a seeded shuffle of AG News test (fancyzhx/ag_news@eb185aade064) and of IMDB test (stanfordnlp/imdb@e6281661ce1c), interleaved; the same list as the Laya report. No text overflows the model's window
MachineApple M4 Max, 40-core GPU, 128 GB unified memory, macOS 26.7, on AC power, no other user workload during timing
RunsRun 1: PyTorch (single requests, the pass over the texts, memory). Run 2: NukeTorch on the current source, same machine, same texts and order
WarmupsSingle request: 3 untimed calls on both engines. Texts: a few untimed requests before the first pass

How each number is measured

Single requestMedian of 30 timed calls. No confidence interval or significance test is claimed.
TimersBoth engines: the text and the questions already in memory → the probabilities of every option. Model load is excluded.
TextsNukeTorch: median of 5 passes through its batch API. PyTorch: median of 2 passes, one evaluate call per text; the package has no batch call. This compares each system as it is used, not the same loop on both.
Speed ratioPyTorch's time ÷ NukeTorch's time. The percentage is (ratio − 1) × 100. Ratios are computed from unrounded values.
AgreementA question agrees when the most probable option is the same as in the reference. Probability change is the largest absolute difference over the options of a question, at the checkpoint's temperature.
Not carried overNukeTorch does not run the package's chain controller. A text that overflows the 8,192-token window loses its middle on both engines; no text of this set overflows.
ScopeThirty calls on two texts and passes over 1,000 English texts on one machine, with one set of 4 questions.

What this does not show

  • Thirty calls on two texts and passes over 1,000 English texts, with one set of 4 questions, on one machine.
  • The package has no batch call, so its 1,000 texts run one request after another; two passes were run.
  • NukeTorch does not carry the package’s chain controller; agreement is scored against the package with it off.
  • No P95 or P99, no statistical significance claim.

Kokoro runs about 3× faster. F5 runs 1.5× faster. Both need up to 74% less memory.

The fastest way to run these models on a Mac today. Same model files, no changes to your code, and output that passes the same quality checks.

Kokoro-82M · one request and in bulk
One request
18.5 ms
vs 46.7 ms · 2.5×
1,000 sentences
46 s
vs 156 s · 3.4×

Against MLX. A single sentence, and a bulk job of 1,000 real sentences.

F5-TTS v1 Base · time per clip
1.37 s
vs 2.05 s with MLX · 1.5×

Identical sampler settings on both sides. In bulk, 1.8× faster across five passes.

Peak memory footprint
Kokoro
0.54 GB
vs 2.05 GB
F5-TTS
0.88 GB
vs 3.22 GB

Peak physical footprint on a single request, NukeTorch against MLX, measured the same way for both.

Measured on one Apple M4 Max with the same model files the authors publish. Single requests: medians of 60 runs on a short sentence. Bulk: 1,000 real sentences from the LibriSpeech corpus. Memory measured with the same outside probe for every engine. Output passes the same automated quality checks, and our own ears; no formal listening study yet. No NVIDIA results. Why it is faster: how NukeTorch works.

Why it matters

The difference

The same three results land differently depending on what you build. These are the people we think they matter to most.

2.05 GB → 0.54 GB

You ship a Mac app with a local voice

74% less memory lowers the machine your app needs, so more of your users' Macs qualify. The audio never leaves the device, and there is no per-minute cloud bill.

Kokoro, peak memory footprint
47 ms → 18.5 ms

You run a local voice assistant

Every step between a person finishing a sentence and hearing a reply adds its own wait: transcription, the language model, then speech. The voice starts sooner, and the memory it gives back goes to the language model, which is usually what runs out first.

Kokoro, time to first audio
60 min → ~33 min

You generate speech in bulk

Voice cloning, narration and dubbing with F5: in our bulk test of 1,000 real sentences, work that takes an hour under MLX took about 33 minutes. With Kokoro, the same 1,000 sentences take 46 seconds instead of 2.6 minutes. Passages of several minutes are next on our test list.

F5-TTS and Kokoro, 1,000-sentence bulk job

Evidence, in one table

Three models, one Mac

MLX is the comparison we hold ourselves to, because it is the fastest alternative on a Mac. PyTorch appears in the Kokoro and F5 tables because it is the reference the model authors publish against.

Model One request vs MLX 1,000 sentences vs MLX Memory vs MLX Evidence
Kokoro-82M 2.5× faster 3.4× faster 74% less Mixed precision
F5-TTS v1 Base 1.5× faster 1.8× faster 73% less Mixed precision
Fish Speech S2 Pro 1.2× faster 4.4× faster 18% less Same precision
Time to first audio

How long the listener waits before the voice starts. In these tests the whole clip returns at once, so it is also the time to the finished clip.

Bulk job

1,000 different sentences from the LibriSpeech corpus, generated back to back. NukeTorch runs in its default bulk mode, which schedules the requests itself; MLX generates them one after another, as it ships.

Real-time factor

Seconds of audio produced per second of computing. Above 1.0, audio is generated faster than it plays.

Peak memory footprint

The most memory the engine held while it ran, everything included. It decides how many copies fit on a machine, and what else can run alongside.

Text to speech · model 1 of 3

Kokoro-82M

About 3× faster than MLX: 2.5× on a single request and 3.4× on a bulk job of 1,000 sentences. The model needs 74% less memory, and it stays well ahead even when run at MLX’s own full-precision settings.

Mixed precision

Maintained by hexgrad · Apache 2.0 · v1.0, voice af_heart, seed 0 · One request: 45,600 samples at 24 kHz (about 1.9 seconds), medians of 60 runs · Bulk: 1,000 LibriSpeech sentences, median of 5 passes

MetricNukeTorchMLXPyTorchvs MLXvs PyTorch
One request, text to finished audio (ms)18.4746.6561.682.53× faster3.34× faster
1,000 sentences in bulk (s) §46.36156.38not run3.37× fastern/a
Real-time factor, one request (×)102.8540.7330.802.53× faster3.34× faster
Real-time factor, bulk (×)150.4244.57not run3.37× fastern/a
Transcription error rate, bulk †2.25%2.12%not run0.13 points worsen/a
Peak memory, one request (MB)5362,0482,15074% less75% less
Peak memory, 1,000 sentences (MB) §2,8267,987not run65% lessn/a

§ NukeTorch in its default bulk mode, which schedules the requests itself; MLX generating them one after another, as it ships. Median of five passes; the slowest NukeTorch pass was 46.38 s and the fastest MLX pass 156.27 s. PyTorch was not tested on the 1,000-sentence batch. Memory is everything each engine held at its peak, measured in runs separate from the timings; MLX keeps memory it has finished with in its own cache, so its bulk peak reflects its largest working set. † The difference comes from NukeTorch’s own text-to-phoneme step, not the synthesis: fed the same Misaki phonemes as MLX, NukeTorch scores 2.14% against 2.12%.

In detail

On the test clip a single request finishes after 18.5 ms instead of 46.7 ms under MLX. In bulk the gap widens to 3.4×; NukeTorch schedules the requests itself, while MLX works through them one at a time. The five bulk passes landed within a third of a second of each other, 46.1 to 46.4 seconds, so this is a steady result rather than a lucky run.

What it means for you

On a single sentence on a fast Mac, the time saved is small in absolute terms. It adds up when Kokoro is one step in a chain. In a local voice assistant, transcription, the language model and speech each add a wait, and about 28 ms back from the last step is 28 ms nobody sits through. For bulk narration, 1,000 sentences take 46 seconds instead of 2.6 minutes.

The memory matters in every setup. About 0.54 GB instead of 2 GB is room for a larger language model on the same Mac, or a lower minimum spec for an app that ships Kokoro. If you maintain Kokoro, it is also a better answer for your readme's Mac section than PyTorch's Apple fallback.

Does the output still match

On the 1,000-sentence bulk test, NukeTorch’s output has 2.25% of words wrong against MLX’s 2.12%, 0.13 points worse. The cause is NukeTorch’s own text-to-phoneme step, which follows Misaki’s rules without its part-of-speech tagger and without its fallback for words outside the dictionary. Given the same Misaki phonemes, NukeTorch scores 2.14%, 0.02 points worse than MLX.

Same settings as MLX · Kokoro at full precision

To answer the obvious objection, that we beat slow settings, NukeTorch was also run the way MLX and PyTorch run: FP32 weights, Misaki’s phonemes (Misaki’s own time measured separately and added to NukeTorch’s), and one request at a time in bulk. It is still 2.35× faster than MLX on a single request (19.88 ms against 46.65 ms), 3.10× faster than PyTorch, and 2.58× faster than MLX across the 1,000 sentences (60.6 s against 156.4 s). Transcription error rates: 2.14% against 2.12%, 0.02 points worse. Peak memory on a single request is 636 MB against MLX’s 2,048 MB. Most of the speed is the runtime itself; half precision and NukeTorch’s own bulk scheduling add the rest.

Nuance
  • The single-request clip is short: about 1.9 seconds of audio from a three-word sentence. Short clips can flatter an engine with low fixed costs per request. The bulk test uses 1,000 real sentences instead; passages of several minutes are not yet measured.
  • The bulk comparison is NukeTorch's default bulk mode against MLX generating one request at a time, as it ships. It compares the two as you would run them, not identical scheduling.
  • All three engines were measured on the same machine. Single requests are medians of 60 runs, timed from raw text to finished audio on every engine; the slowest NukeTorch run was 18.80 ms and the fastest MLX run 46.03 ms. No P95 or P99 claim, here or anywhere on this site.
  • Precision: NukeTorch runs FP16; PyTorch and MLX run FP32, their default configuration, as their packages ship. We did not tune them. In NukeTorch, FP16 and FP32 give the same intelligibility: 2.25% of words wrong in both, 998 of 1,000 transcripts identical. Run at FP32 with MLX’s own phonemes, it is still 2.35× faster per request and 2.58× in bulk.
  • Bulk transcription error rate is 0.13 points worse than MLX’s because of NukeTorch’s own text-to-phoneme step; with Misaki’s phonemes it is 0.02 points worse.

How to use it

macOS on Apple silicon · the same model files and checkpoint revision you use today · licensing undetermined

Benchmark reportReproducing the Kokoro results
MachineApple M4 Max, 40 GPU cores, 128 GB, macOS 26.7
Reportv1.1 · FP32 report v1.0. Author-run measurements; no independent reproduction yet.
One requestMedian of 60 timed calls per engine (2 rounds of 30)
BulkMedian of 5 passes per engine; PyTorch not tested on the 1,000-sentence batch

Exactly what was run

ModelKokoro 82M v1.0 (hexgrad), voice af_heart, 24 kHz output
Checkpointshexgrad/Kokoro-82M@f3ff3571791e39611d31c381e3a41a3af07b4987 (PyTorch); prince-canuma/Kokoro-82M@e02c9eada7ce7416798af36b190a8a2dd2ecd566 (MLX); NukeTorch runs the same model
NukeTorchRelease build; batch API for the batch series
MLXmlx-audio 0.5.6 on MLX 0.32.2, run as shipped
PyTorchkokoro 0.9.4 (KModel/KPipeline) on PyTorch 2.10.0, model and inputs on the MPS device, torch.mps.synchronize() before each timer stop
Reference precisionMLX and PyTorch: FP32 weights as shipped.
Generation settingsVoice af_heart, speed 1.0, English. NukeTorch uses its own text-to-phoneme front end; MLX and PyTorch use the packaged Misaki G2P pipeline.
Single-request fixture"Watch out, driver!" (18 phonemes); every engine returns 45,600 samples (1.9 s)
Batch corpus1,000 utterances of LibriSpeech test-clean (the first 1,000 in ID order, utterances over 30 s skipped; OpenSLR SLR12, CC BY 4.0); sentence case with final punctuation
MachineApple M4 Max, 40-core GPU, 128 GB unified memory, macOS 26.7, on AC power, no other user workload during timing
RunsRun 1 measured MLX and PyTorch (single request and batch). Run 2 measured NukeTorch (single request from raw text, and batch) with its current text front end (the single-request sentence gives bit-identical audio to run 1's build). Each batch pass ran in its own process.
WarmupsNukeTorch 3, MLX and PyTorch 2 before each 30-call round; 2 utterances before each batch pass

How each number is measured

Single requestMedian of 60 timed calls (two rounds of 30, pooled). No confidence interval or significance test is claimed.
BatchHeadline is the median of the five pass times. Throughput = 1,000 ÷ wall time. RTFx = seconds of generated audio ÷ processing seconds (higher is faster). Input characters per second is a text-throughput figure: the same 110,187 characters divided by the wall time.
Batch executionNukeTorch's batch API receives the whole list; MLX's runner processes one request at a time. Both are reported as system throughput at each engine's default settings. The MLX wall time includes its WAV writes between requests; against its generation time alone (155.14 s), NukeTorch is 3.35× faster.
Warm-upSingle request: each 30-call round is preceded by warm-up calls (3 for NukeTorch, 2 for MLX and PyTorch) that are not timed. Batch: 2 warm-up utterances before each pass, not timed. Each pass ran in its own process, one after another. Machine temperature was not recorded.
Speed ratioFor times, other engine's time ÷ NukeTorch's time; for rates, NukeTorch's rate ÷ other engine's rate. The percentage is (ratio − 1) × 100 and describes speed, not time saved. Ratios are computed from unrounded values.
ScopeSixty calls on one sentence and five passes over one English corpus on one machine. They do not establish behaviour on other voices, languages, text domains or hardware.

Other configurations

Full precision, like MLXFrom the Kokoro FP32 report, v1.0. NukeTorch at FP32 with Misaki’s phonemes, one request at a time: 19.88 ms per request (2.35× faster than MLX, 3.10× faster than PyTorch); 60.63 s for 1,000 sentences (2.58× faster than MLX); 2.14% words wrong against 2.12%; 636 MB peak memory against 2,048 MB. NukeTorch’s own FP16 against FP32 with its own phonemes: 18.45 ms against 19.09 ms per request, 2.25% words wrong in both, 998 of 1,000 transcripts identical (one round of 30 calls and one bulk pass per precision).

What this does not show

  • One 1.9-second fixture for single requests. Bulk uses 1,000 real sentences, but nothing of several minutes.
  • The bulk comparison is NukeTorch’s default bulk mode against MLX one request at a time. MLX and PyTorch were measured in one run and NukeTorch in a second, not alternated.
  • Medians on one machine. No P95 or P99, no statistical significance claim. No test of several users at once.
  • Transcription error rate is 0.13 points worse than MLX’s on default settings, traced to NukeTorch’s own text-to-phoneme step.
  • A 128 GB M4 Max, where a memory result matters least. Not yet run on a base MacBook Air or an older M1 or M2 machine.
  • No human listening study, besides our own ears. Quality evidence is automatic transcription.

Text to speech · model 2 of 3

F5-TTS v1 Base

Each clip is 1.50× faster than with MLX, with identical sampler settings, and a bulk job of 1,000 sentences 1.80× faster. It needs 73% less memory. Read from the same 8-bit file on both engines, it is 1.43× faster per request and 2.01× in bulk.

Mixed precision

Shanghai Jiao Tong University, with Cambridge and Geely · weights CC BY-NC 4.0, non-commercial only · model_1250000, shared reference WAV and transcript, seed 0 · One request: medians of 60 runs · Bulk: 1,000 LibriSpeech sentences, median of 5 passes

MetricNukeTorchMLXPyTorch *vs MLXvs PyTorch *
One request, time to first audio (ms)1,366.532,045.653,229.831.50× faster2.36× faster
1,000 sentences in bulk (s) §1,955.623,515.26not run1.80× fastern/a
Real-time factor, one request (×)1.190.800.531.50× faster2.28× faster
Real-time factor, bulk (×)3.251.81not run1.80× fastern/a
Transcription error rate, bulk2.47%2.52%not run0.05 points bettern/a
Peak memory, one request (MB) †8783,2244,34473% less80% less
Peak memory, 1,000 sentences (MB) †6,851106,189not run94% lessn/a

* PyTorch runs its own default sampler (Euler, 32 steps), not the RK4 settings MLX and NukeTorch share, so its column is not like-for-like. § NukeTorch’s batch API on an earlier build; on the current build it takes 1,771.53 s over two passes, 1.98× faster. MLX generating them one after another, as it ships. Median of five passes; the slowest NukeTorch pass was 1,963.20 s and the fastest MLX pass 3,454.79 s. PyTorch was not tested on the 1,000-sentence batch. † Everything each engine held at its peak, measured in runs separate from the timings. MLX keeps memory it has finished with in its own cache, so its bulk peak reflects its largest working set; its median over the second half of the bulk run was still 105,165 MB.

In detail

F5 produces audio by running a sampler that evaluates the model 28 times per request. A request takes 1,367 ms instead of 2,046 ms, with the sampler settings held identical on both sides. NukeTorch's timings come from paired runs in balanced order, alternating with a stock MLX worker in the same process; the MLX column is MLX measured on its own. In bulk the gain grows to 1.80×, NukeTorch’s batch API against MLX one request at a time.

What it means for you

F5 is slow on a Mac, so here the time is the point. On the test clip the audio comes back after 1.37 seconds instead of 2.05 under MLX: 1.5× faster on every generation. For voice cloning, narration or dubbing done in bulk, 1,000 sentences took 33 minutes instead of 59 under MLX, so work that takes an hour under MLX takes about 33 minutes. Transcription error rates were 0.05 points better: a median of 2.47% against 2.52% across 20,404 words in each of five passes.

On a single request F5 generates audio faster than it plays, at 1.19× real time against MLX's 0.80×, and in bulk at 3.25× real time. And its memory footprint drops from 3.2 GB to 0.9 GB, so it stops needing a well-specified machine. On the 1,000-sentence job the gap is starker: NukeTorch stays at 6.9 GB, while MLX grows to fill 106 GB of this 128 GB Mac.

Identical weights on both sides · F5-TTS 8-bit

Both engines were also run on MLX’s published 8-bit F5-TTS checkpoint, the same file read by both exactly as published, with the same sampler, texts, reference and seeds. NukeTorch is 1.43× faster than MLX on a single request (1,415 ms against 2,022 ms) and 2.01× faster across the 1,000 sentences (30 minutes against 61). Transcription error rates: 2.35% against 2.46%. People pick the 8-bit file to save memory, but on MLX it makes bulk jobs 1.04× slower than its FP32 model and still needs 2,167 MB for a single request. On NukeTorch the same file runs in 658 MB, and a 1,000-sentence job peaks at 1.6 GB against MLX’s 106 GB.

Same 8-bit fileNukeTorchMLXvs MLX
One request (ms)1,415.172,021.681.43× faster
1,000 sentences in bulk (s)1,815.523,650.812.01× faster
Real-time factor, one request1.15×0.80×1.43× faster
Real-time factor, bulk3.50×1.74×2.01× faster
Transcription error rate, bulk2.35%2.46%0.10 points better
Peak memory, one request (MB)6582,16770% less
Peak memory, 1,000 sentences (MB)1,563106,18999% less

Both engines read MLX’s published 8-bit checkpoint, the same file, with the same sampler, texts, reference and seeds. One request: median of 60 calls. Bulk: median of 5 passes, NukeTorch’s batch API against MLX one request at a time.

Does the output still match

All passes returned 1,000 non-empty outputs. Both engines generate the same total audio length per pass, and transcription error rates are 0.05 points better: a median of 2.47% for NukeTorch against 2.52% for MLX, scored against the corpus text. Word error rate measures intelligibility only, not naturalness or speaker similarity.

Nuance
  • NukeTorch and MLX use the same sampler settings.
  • The bulk comparison is NukeTorch's default bulk mode against MLX one request at a time, as it ships. The median of five passes is shown.
  • Timing boundaries differ slightly: MLX and PyTorch request timing includes reading the reference and preparing the text; NukeTorch starts at its speak call.
  • Precision: NukeTorch runs FP16; MLX and PyTorch run the FP32 weights as shipped. With the same 8-bit checkpoint on both sides, NukeTorch is still 1.43× faster per request and 2.01× in bulk.
  • Memory is everything each engine held at its peak, NukeTorch in FP16 against MLX and PyTorch in FP32. MLX’s 106 GB bulk peak partly reflects memory it keeps cached after use; on a smaller Mac it would likely use less and run slower.
  • The whole clip returns at once, so the listener waits the full 1.4 seconds before hearing anything. We have not measured sentence-by-sentence streaming and make no claim about live conversation.
  • Output lengths differ slightly between samplers: 39,181 samples for MLX and NukeTorch, 40,704 for PyTorch. Real-time factor is computed per request from actual output duration.
  • Text outside English, or a reference that is not mono PCM16, falls back to an MLX worker rather than failing.

How to use it

macOS on Apple silicon · the same model files and checkpoint revision you use today · licensing undetermined

Benchmark reportReproducing the F5-TTS results
MachineApple M4 Max, 40 GPU cores, 128 GB, macOS 26.7
Reportv1.1 · 8-bit report v1.0. Author-run measurements; no independent reproduction yet.
One requestMedian of 60 timed calls per engine (2 rounds of 30)
BulkMedian of 5 passes per engine; PyTorch not tested on the 1,000-sentence batch

Exactly what was run

ModelF5-TTS v1 Base (SWivid), 24 kHz output
CheckpointsSWivid/F5-TTS@84e5a410d9cead4de2f847e7c9369a6440bdfaca (PyTorch); lucasnewman/f5-tts-mlx@2d719cadec8fd3887c8599475a5f65924d523654 (MLX and NukeTorch: the same file)
NukeTorchRelease build; batch API for the batch series
MLXf5-tts-mlx 0.2.6 on MLX 0.31.2, run as shipped
PyTorchthe F5-TTS reference implementation on PyTorch 2.10.0, model and inputs on the MPS device, torch.mps.synchronize() before each timer stop
Reference precisionMLX and PyTorch: FP32. Both checkpoints are stored in FP32; the F5-TTS reference loader casts to FP16 only on CUDA GPUs and loads FP32 on Apple silicon (MPS), and f5-tts-mlx keeps the stored FP32 weights.
Generation settingsFlow-matching sampler. NukeTorch and MLX: RK4 over 8 time points (7 steps × 4 evaluations = 28 denoiser evaluations), CFG strength 2 (conditional and unconditional forward per evaluation), sway coefficient −1. PyTorch: Euler, 32 steps (32 denoiser evaluations), CFG 2, as shipped. The solvers differ, so the PyTorch comparison is a comparison of the shipped configurations, not of equal solver work.
Single-request fixture"Watch out, driver!" with a fixed short English reference clip and transcript; NukeTorch and MLX return 39,181 samples, PyTorch 40,704
Batch corpus1,000 utterances of LibriSpeech test-clean (the first 1,000 in ID order, utterances over 30 s skipped; OpenSLR SLR12, CC BY 4.0); sentence case with final punctuation
MachineApple M4 Max, 40-core GPU, 128 GB unified memory, macOS 26.7, on AC power, no other user workload during timing
RunsOne series. Single-request rounds alternated engines. Each batch pass ran NukeTorch first, then MLX.
WarmupsNukeTorch 3, MLX and PyTorch 2 before each 30-call round; 2 utterances before each batch pass

How each number is measured

Single requestMedian of 60 timed calls (two rounds of 30, pooled). No confidence interval or significance test is claimed.
ControlThe NukeTorch single-request rounds were run paired with a stock MLX control in the same process (order ABBA, then BAAB). Had the control drifted more than 2% from its reference series, the runner would have idled 420 s and repeated the pair; it did not (+1.7% and −1.3%).
BatchHeadline is the median of the five pass times. Throughput = 1,000 ÷ wall time. RTFx = seconds of generated audio ÷ processing seconds (higher is faster). Input characters per second is a text-throughput figure: the same 110,187 characters divided by the wall time.
Batch executionNukeTorch's batch API receives the whole list; MLX's runner processes one request at a time. Both are reported as system throughput at each engine's default settings. The MLX wall time includes its WAV writes between requests; against its generation time alone (3,513.94 s), NukeTorch is 1.80× faster.
Warm-upSingle request: each 30-call round is preceded by warm-up calls (3 for NukeTorch, 2 for MLX and PyTorch) that are not timed. Batch: 2 warm-up utterances before each pass, not timed. Each pass ran in its own process, one after another. Machine temperature was not recorded.
Speed ratioFor times, other engine's time ÷ NukeTorch's time; for rates, NukeTorch's rate ÷ other engine's rate. The percentage is (ratio − 1) × 100 and describes speed, not time saved. Ratios are computed from unrounded values.
ScopeSixty calls on one sentence and five passes over one English corpus on one machine. They do not establish behaviour on other voices, languages, text domains or hardware.

8-bit against 8-bit

From the F5-TTS 8-bit report, v1.0.

Checkpointslucasnewman/f5-tts-mlx@2d719cadec8fd3887c8599475a5f65924d523654, file model_v1_8b, read by both engines as published; duration predictor duration_v2, quantized at load by both (identical codes)
MLXf5-tts-mlx 0.2.6 on MLX 0.31.2, --quantization-bits 8, run as shipped
WeightsMLX's published 8-bit checkpoint (group-64 affine uint8, scale and offset per 64 weights), loaded by both engines. MLX: its quantized matmuls as shipped

What this does not show

  • One timing fixture for single requests; bulk uses 1,000 sentences, five passes per engine. No statistical significance or general speed claim.
  • No streaming test. The whole clip returns at once.
  • The PyTorch column compares different samplers and should be discounted; the MLX comparison does not have this problem.
  • No P95 or P99. No test of several users at once.
  • Bulk memory was measured over one pass per engine; PyTorch was not tested on the 1,000-sentence batch.
  • The current-build figure is two passes; the MLX figure it is compared with is an earlier record, not run again.
  • No NVIDIA comparison, including against TensorRT-LLM.
  • No human listening study, besides our own ears. Quality evidence is automatic transcription.

Text to speech · model 3 of 3

Fish Speech S2 Pro

The heaviest model, and the cleanest comparison, with both engines at the same precision. 1.2× faster than MLX on a single request and 4.4× on a bulk job.

Same precision

Fish Audio · Fish Audio Research License · Built with Fish Audio · both engines on the checkpoint's BF16 weights, no quantization · One request: equal work at 33 frames, medians of 60 runs · Bulk: 1,000 LibriSpeech sentences, median of 5 passes

MetricNukeTorchMLXvs MLX
One request, full clip, 33 frames (ms)1,248.351,509.851.21× faster
Time per generated frame (ms)37.8345.751.21× faster
Real-time factor, one request (×)1.231.021.21× faster
1,000 sentences in bulk (s) §1,445.996,373.304.41× faster
Real-time factor, bulk (×)4.651.064.41× faster
Transcription error rate, bulk2.00%1.94%0.06 points worse
Peak memory, one request (MB)12,45815,20618% less
Peak memory, 1,000 sentences (MB)38,668108,47964% less

§ NukeTorch in its default bulk mode, which schedules the requests itself; MLX generating them one after another, as it ships. Median of five passes; the slowest NukeTorch pass was 1,544.39 s and the fastest MLX pass 6,346.40 s. Memory is everything each engine held at its peak, measured in runs separate from the timings.

In detail

S2 Pro is the heaviest model measured here, and on a single request it shows the smallest gain: 1.21×, with both engines reading the same BF16 weights and doing the same 33 frames of work.

In bulk, NukeTorch schedules the 1,000 requests itself, while MLX, as shipped, generates them one after another; the bulk job is 4.4× faster.

What it means for you

On a single request, 1.21× faster means about one hour of compute saved in every six. In bulk, 1,000 sentences took 24 minutes instead of 1 hour 46 minutes under MLX, with transcription error rates of 2.00% against 1.94%, 0.06 points worse. For a business whose margin is the gap between what a minute of audio costs to produce and what it sells for, that matters, and it does not require changing the model, the weights or the output.

Nuance
  • MLX is the only comparison here; PyTorch was not tested for this model.
  • Same precision on both sides: both engines load the same BF16 model files, byte for byte, and neither quantizes its weights.
  • Fish samples its audio, so the two engines produce different valid renderings. For the single-request test, NukeTorch was capped at the 33 frames MLX produced, so both do the same work.
  • The bulk comparison is NukeTorch's default bulk mode against MLX one request at a time, as it ships. NukeTorch's first pass took 6.7 to 8.9% longer than the other four, for a cause not established; a sixth pass on an idle machine (1,424 s) matched passes 2 to 5 and is not counted. In that first pass, two of the 1,000 requests reached the length limit rather than ending on their own; in the other four, every request ended on its own.

How to use it

macOS on Apple silicon · licensing undetermined

Benchmark reportReproducing the Fish Speech results
MachineApple M4 Max, 40 GPU cores, 128 GB, macOS 26.7
Reportv1.1. Author-run measurements; no independent reproduction yet.
One requestMedian of 60 timed calls per engine (2 rounds of 30), at a fixed 33 frames
BulkMedian of 5 passes per engine; PyTorch not tested for this model

Exactly what was run

ModelFish Speech S2 Pro (Fish Audio), 44.1 kHz output, 2,048 samples per generated frame
Checkpointsfishaudio/s2-pro@1de9996b6be38b745688de084d87a5633f714e4e (NukeTorch); mlx-community/fish-audio-s2-pro@eccd57bf5c1ebc13cb2f993df867f4e49931a36a (MLX); the two repositories' model shards are byte-identical (same SHA-256)
NukeTorchRelease build; batch API for the batch series
MLXmlx-audio 0.5.6 on MLX 0.32.2, run as shipped
PyTorchNot run for this model
Reference precisionMLX: the BF16 model as shipped. No weight quantization on either side.
Generation settingsFish's own sampler on both engines: temperature 0.7, top-p 0.7, top-k 30, repetition penalty, seed 0, no reference voice. Sampled tokens differ between the two engines' random streams, so each engine produces its own rendering of the text.
Single-request fixture"Watch out, driver!"; MLX stops at 33 frames on every call for this fixture, and NukeTorch was capped at 33 frames so both decode the same number of frames
Batch corpus1,000 utterances of LibriSpeech test-clean (the first 1,000 in ID order, utterances over 30 s skipped; OpenSLR SLR12, CC BY 4.0); sentence case with final punctuation
MachineApple M4 Max, 40-core GPU, 128 GB unified memory, macOS 26.7, on AC power, no other user workload during timing
RunsOne series. Single-request rounds alternated engines. Each batch pass ran NukeTorch first, then MLX.
Warmups2 per engine before each 30-call round; 1 utterance before each batch pass

How each number is measured

Single requestMedian of 60 timed calls (two rounds of 30, pooled). No confidence interval or significance test is claimed.
BatchHeadline is the median of the five pass times. Throughput = 1,000 ÷ wall time. RTFx = seconds of generated audio ÷ processing seconds (higher is faster). Input characters per second is a text-throughput figure: the same 110,187 characters divided by the wall time.
Batch executionNukeTorch's batch API receives the whole list; MLX's runner processes one request at a time. Both are reported as system throughput at each engine's default settings. The MLX wall time includes its WAV writes between requests; against its generation time alone (6,371.35 s), NukeTorch is 4.41× faster.
Warm-upSingle request: each 30-call round is preceded by warm-up calls (2 per engine) that are not timed. Batch: 1 warm-up utterance before each pass, not timed. Each pass ran in its own process, one after another. Machine temperature was not recorded.
Speed ratioFor times, other engine's time ÷ NukeTorch's time; for rates, NukeTorch's rate ÷ other engine's rate. The percentage is (ratio − 1) × 100 and describes speed, not time saved. Ratios are computed from unrounded values.
ScopeSixty calls on one sentence and five passes over one English corpus on one machine. They do not establish behaviour on other voices, languages, text domains or hardware.

What this does not show

  • One sentence for single requests, at a fixed 33 frames on both engines; whether the capped NukeTorch rendering ends the phrase as naturally as MLX’s was not evaluated.
  • Fish samples its tokens, so each engine produces its own rendering, and the generated length varies slightly between passes.
  • NukeTorch’s first bulk pass took 6.7 to 8.9% longer than the other four; the cause is not established, and the pass stays in the record.
  • The bulk comparison is NukeTorch’s default bulk mode against MLX one request at a time, not matched concurrency.
  • PyTorch was not tested for this model.
  • Medians on one machine. No P95 or P99, no statistical significance claim.
  • No human listening study, besides our own ears. Quality evidence is automatic transcription.

Parakeet transcribes about 3× faster. Whisper up to 2.5× faster. With up to 77% less memory.

Compared with MLX, the fastest alternative on a Mac. Same model files, no changes to your code, the same or a better word error rate on 1,000 real recordings, and a fraction of the memory.

Parakeet TDT 0.6B v3 · one clip and in bulk
One clip
20 ms
vs 34 ms · 1.7×
1,000 clips
12.9 s
vs 38.4 s · 3.0×

Against MLX, both in FP16. A 6.4-second utterance, and two hours of real recordings.

Whisper · 1,000 clips
Tiny
12.6 s
vs 32.0 s · 2.5×
large-v3-turbo
135 s
vs 239 s · 1.8×

Against MLX, with FP16 weights on both sides.

Peak memory · 1,000 clips
Parakeet
5.7 GB
vs 24.3 GB
Whisper turbo
3.7 GB
vs 5.1 GB

Everything each engine held at its peak, NukeTorch against MLX.

Measured on one Apple M4 Max with the published checkpoints each benchmark report lists; for Parakeet, MLX and PyTorch ran FP16 copies of their published FP32 checkpoints. One clip: 6.4 seconds of audio holding 5.6 seconds of English speech, medians of 60 calls. Bulk: 1,000 LibriSpeech recordings, medians of 5 passes, with NukeTorch in its default bulk mode and MLX one clip at a time. Parakeet is also compared with NVIDIA’s NeMo-Speech.cpp and with parakeet.cpp; no comparison with WhisperKit, and no results on NVIDIA hardware yet. Why it is faster: how NukeTorch works.

Why it matters

The difference

The same results land differently depending on what you build.

38 s → 13 s

You transcribe in bulk

Meeting notes, call recordings, podcast archives. Two hours of recorded speech transcribed by Parakeet in 13 seconds instead of 38 under MLX, with the same word error rate, in 5.7 GB of memory instead of 24.3 GB.

Parakeet, 1,000 real recordings
34 ms → 20 ms

You run a local voice assistant

Listening is the first wait in every turn. Parakeet turns a 6.4-second utterance into text in 20 ms instead of 34 under MLX, before the language model and the voice even start. Pair it with the text to speech results for both ends of the conversation.

Parakeet, one clip
239 s → 135 s

You ship Whisper in a Mac app

The same Whisper weights at the same precision, no retraining, 1.8× faster on large-v3-turbo across 1,000 recordings and with 28% less memory. Whisper Tiny is 2.5× faster.

Whisper large-v3-turbo, 1,000 real recordings

Evidence, in one table

Three models, one Mac

MLX is the comparison we hold ourselves to, because it is the fastest alternative on a Mac. PyTorch appears in each model's table because it is the reference the model authors publish against. For Parakeet we also ran NVIDIA’s NeMo-Speech.cpp and parakeet.cpp.

Model One clip vs MLX 1,000 clips vs MLX Memory in bulk vs MLX Words wrong, NukeTorch vs MLX Evidence
Parakeet TDT 0.6B v3 1.7× faster 3.0× faster 77% less 1.78% vs 1.78% Same precision
Whisper large-v3-turbo 1.2× faster 1.8× faster 28% less 3.18% vs 3.20% Same precision
Whisper Tiny 2.0× faster 2.5× faster 69% less 7.04% vs 7.04% Same precision
One clip

A 6.4-second English utterance, timed from audio in memory to the final transcript. Median of 60 calls.

1,000 clips

Real recordings from the LibriSpeech corpus, 7,389 seconds of speech. NukeTorch runs in its default bulk mode, which schedules the clips itself; MLX transcribes them one after another, as it ships.

Word error rate

The share of words transcribed wrong against the human transcript. Lower is better; equal numbers mean the speed came without an accuracy cost.

Speech to text · model 1 of 3

Parakeet TDT 0.6B v3

In FP16, 1.69× faster than MLX on a single clip and 2.97× across 1,000 recordings, with the same word error rate. In 8-bit, against NVIDIA’s native runtime NeMo-Speech.cpp, 2.61× on a single clip and 4.73× in bulk. In bulk it needs 77% less memory than MLX.

Same precision

NVIDIA · CC BY 4.0 · 25 languages · One clip: 6.4 seconds of audio holding 5.6 seconds of English speech, median of 60 calls · Bulk: 1,000 LibriSpeech recordings, 7,389 seconds of speech, median of 5 passes

One clipPrecisionTime (ms)Real-time factorPeak memory (MB) †NukeTorch speedNukeTorch memory
NukeTorchFP1619.98320.3×1,735——
MLXFP1633.69189.9×1,8821.69× faster8% less
parakeet.cppFP1655.62115.1×2,7692.78× faster37% less
PyTorchFP1672.8887.8×2,8913.65× faster40% less
NukeTorch8-bit17.83358.9×2,203——
NeMo-Speech.cpp8-bit46.61137.3×9732.61× faster126% more
1,000 clipsPrecisionTime (s)Real-time factorWord error ratePeak memory (MB) †NukeTorch speedNukeTorch memoryNukeTorch accuracy
NukeTorch §FP1612.91572.4×1.78%5,666———
MLXFP1638.40192.4×1.78%24,3022.97× faster77% lesssame
NukeTorch §8-bit11.40648.0×1.78%3,919———
NeMo-Speech.cpp8-bit53.96136.9×1.77%1,5064.73× faster160% more0.01 points worse

Speed, memory and accuracy columns compare NukeTorch with the engine in that row: FP16 with FP16, 8-bit with 8-bit. § 1,000 LibriSpeech test-clean recordings, 7,389 seconds of speech, median of five passes per engine; NukeTorch’s batch API against MLX and NeMo-Speech.cpp one clip after another, as they ship. PyTorch and parakeet.cpp were not tested on the 1,000-clip batch. † Everything each engine held at its peak, measured in runs separate from the timings.

In detail

Parakeet’s decoder works through the audio step by step, emitting up to ten symbols per slice of sound. On a single clip NukeTorch is 1.69× faster than MLX and 2.78× faster than parakeet.cpp in FP16, and 2.61× faster than NVIDIA’s NeMo-Speech.cpp in 8-bit.

In bulk NukeTorch also schedules the clips itself: 2.97× faster than MLX in FP16 and 4.73× faster than NeMo-Speech.cpp in 8-bit. Over the bulk job MLX’s memory climbs to 24.3 GB, against 5.7 GB for NukeTorch in FP16; in 8-bit NukeTorch needs 3.9 GB and NeMo-Speech.cpp 1.5 GB.

What it means for you

Two hours of recorded speech — meeting notes, call recordings, podcast archives — transcribed in 13 seconds instead of 38 under MLX, at the same word error rate; in 8-bit, 11 seconds instead of 54 under NeMo-Speech.cpp.

In a local voice assistant, listening is the first wait in every turn. A 6.4-second utterance becomes text in 20 ms instead of 34 ms under MLX, before the language model and the voice even start.

Nuance
  • Precision: FP16 is compared with FP16 and 8-bit with 8-bit. NukeTorch runs FP16 natively; MLX and PyTorch run FP16 copies of the published FP32 checkpoints, parakeet.cpp its published FP16 file. NeMo-Speech.cpp runs the 8-bit file NVIDIA publishes for it (no FP16 file is published for that runtime), against NukeTorch’s 8-bit models: int8 weights with one scale per output channel, made from the same checkpoint when the model first loads.
  • The bulk comparison is NukeTorch’s batch API against MLX and NeMo-Speech.cpp one clip at a time, as they ship.
  • NeMo, NVIDIA’s Python toolkit, has no Apple GPU path, so it is a note rather than a column: on the CPU in FP32 it takes 1,043.73 ms per clip, 52× slower than NukeTorch, and 694.3 s for the 1,000 clips with its own batching, at 1.78% of words wrong.
  • The one-clip sample was generated with Kokoro, not recorded. The 1,000-clip test uses real recordings.
  • English only, although Parakeet v3 supports 25 languages.

How to use it

macOS on Apple silicon · the same model files and checkpoint revision you use today · licensing undetermined

Benchmark reportReproducing the Parakeet results
MachineApple M4 Max, 40 GPU cores, 128 GB, macOS 26.7
Reportv1.3. Author-run measurements; no independent reproduction yet.
One clipMedian of 60 timed calls per engine (2 rounds of 30)
BulkMedian of 5 passes per engine; PyTorch and parakeet.cpp not tested on the 1,000-clip batch

Exactly what was run

ModelNVIDIA Parakeet TDT 0.6B v3 (600M parameters, 25 languages)
CheckpointsNVIDIA nvidia/parakeet-tdt-0.6b-v3@541d1f99c6b0c3cd0b11a95167540bb8edefd82b (PyTorch); mlx-community/parakeet-tdt-0.6b-v3@ed2b7e8c15f9aaa0b5772e2efb986255eaef7e15 (MLX); NukeTorch runs the same model; NeMo-Speech.cpp: NVIDIA's parakeet-tdt-0.6b-v3.q8_0.gguf from the same repository; parakeet.cpp: tdt-0.6b-v3-f16.gguf from mudler/parakeet-cpp-gguf; NeMo: the .nemo checkpoint from_pretrained downloads
NukeTorchRelease build; FP16 weights, its native format; 8-bit models with int8 weights, made from the same checkpoint when the model first loads.
MLXmlx-audio 0.5.6 on MLX 0.32.2, public loader and decode, FP16 weights, audio features in float16
NeMo-Speech.cppNVIDIA NeMo-Speech.cpp 0.2.0, source build, metal-asr preset, Metal device; bench asr with its defaults (offline TDT, full context)
parakeet.cppLocalAI parakeet.cpp (ggml), source build with PARAKEET_GGML_METAL=ON, Metal device; bench --decoder tdt with its defaults
PyTorchTransformers 5.17.0 on PyTorch 2.10.0, ParakeetForTDT.from_pretrained(dtype=float16), model and inputs on the MPS device, torch.mps.synchronize() before each timer stop
NeMo (note)nemo_toolkit 3.0.0 on PyTorch 2.10.0, ASRModel.from_pretrained and transcribe with defaults: CPU, FP32, greedy_batch TDT
PrecisionNukeTorch FP16, and its 8-bit models for the 8-bit comparisons; MLX and PyTorch: the published FP32 checkpoints cast to FP16 (every tensor, NumPy cast, configuration unchanged); NeMo-Speech.cpp q8_0 as published; parakeet.cpp f16 as published; NeMo FP32 as published
DecodingGreedy TDT decoding, English, up to 10 symbols per frame, on NukeTorch, MLX and PyTorch; NeMo-Speech.cpp and parakeet.cpp at their TDT defaults.
Single-request clipSynthetic English sentence generated with Kokoro (voice af_heart) at 24 kHz, resampled to 16 kHz mono, zero-padded to 6.4 s
Batch corpus1,000 utterances of LibriSpeech test-clean (the first 1,000 in ID order, utterances over 30 s skipped; OpenSLR SLR12, CC BY 4.0). FLAC converted losslessly to 16-bit WAV; reference transcripts from the corpus
MachineApple M4 Max, 40-core GPU, 128 GB unified memory, macOS 26.7, on AC power, no other user workload during timing
RunsRun 1: MLX, PyTorch, the native runtimes and NeMo (single-request rounds and batch passes alternated engines). Run 2: a NukeTorch-only check, not shown on this page. Run 3: NukeTorch in FP16 and in 8-bit, same machine, clips, order and protocol.
WarmupsNukeTorch 3, MLX, PyTorch and NeMo-Speech.cpp 2, parakeet.cpp 2 (its bench's own warmup plus the first timed call dropped) before each 30-call round; 1–2 utterances before each batch pass

How each number is measured

Single requestMedian of 60 timed calls (two rounds of 30, pooled). NeMo-Speech.cpp's bench reports the distribution of each round (median = the mean of the two round medians, min and max over both rounds, no SD). No confidence interval or significance test is claimed.
TimersNukeTorch, MLX and PyTorch: audio samples already in memory → text. parakeet.cpp: its bench's per-file processing time, audio decoded beforehand. NeMo-Speech.cpp: its bench's client-side latency per item, which includes reading the 200 KB WAV. All exclude model load.
BatchHeadline is the median of the five pass times. Throughput = 1,000 ÷ wall time. RTFx = seconds of input audio ÷ processing seconds (higher is faster).
Batch executionNukeTorch's batch API receives the whole list; MLX's runner and NeMo-Speech.cpp's bench process one utterance at a time. All are reported as system throughput at each engine's default settings.
Speed ratioFor times, other engine's time ÷ NukeTorch's time; for rates, NukeTorch's rate ÷ other engine's rate. The percentage is (ratio − 1) × 100 and describes speed, not time saved. Ratios are computed from unrounded values.
Warm-upNot timed. Each pass ran in its own process, one after another. Machine temperature was not recorded.
WERSubstitutions + deletions + insertions ÷ reference words, after lowercasing and removing punctuation.
ScopeSixty calls on one clip and five passes over one English read-speech corpus on one machine. They do not establish behaviour on other speakers, languages, domains or hardware.

What this does not show

  • FP16 is compared with FP16 and 8-bit with 8-bit; NukeTorch’s 8-bit models and the 8-bit file NVIDIA publishes for NeMo-Speech.cpp are different 8-bit formats.
  • One machine. No P95 or P99, no statistical significance claim.
  • Clips of up to 30 seconds; no hour-long files. English read speech only.
  • No streaming test: every figure is the full transcript, so there is no time-to-first-word figure.
  • The bulk comparison is NukeTorch’s own scheduling against the other runtimes one clip at a time, not matched concurrency.
  • Sessions were used one at a time; concurrent calls on one model were not tested.
  • Everything ran on one Mac; no results on NVIDIA hardware.

Speech to text · model 2 of 3

Whisper large-v3-turbo

The heaviest speech-to-text model here and the cleanest comparison, with FP16 weights on every engine: 1.2× faster than MLX on a single clip, 1.8× across 1,000 recordings, with 28% less memory.

Same precision

OpenAI · MIT · multilingual · One clip: 6.4 seconds of audio holding 5.6 seconds of English speech, median of 60 calls · Bulk: 1,000 LibriSpeech recordings, median of 5 passes

MetricNukeTorchMLXPyTorchvs MLXvs PyTorch
One clip, request time (ms)193.21228.57368.991.18× faster1.91× faster
1,000 clips (s) §134.81239.18not run1.77× fastern/a
Clips per second, bulk7.424.18not run1.77× fastern/a
Real-time factor, one clip (×)33.128.017.31.18× faster1.91× faster
Real-time factor, bulk (×)54.830.9not run1.77× fastern/a
Word error rate, 1,000 clips3.18%3.20%not run0.02 points bettern/a
Peak memory, one clip (MB)1,8432,7653,27733% less44% less
Peak memory, 1,000 clips (MB) §3,7085,120not run28% lessn/a

§ 1,000 LibriSpeech test-clean recordings, 7,389 seconds of speech, median of five passes per engine. NukeTorch runs in its default bulk mode, which schedules the clips itself; MLX transcribes them one after another, as it ships. PyTorch was not tested on the 1,000-clip batch. Memory is everything each engine held at its peak, measured in runs separate from the timings. MLX keeps memory it has finished with in its own cache, so its bulk peak reflects its largest working set.

In detail

large-v3-turbo is the heaviest speech-to-text model on this page, and on a single clip it shows the smallest gain, 1.18×. All three engines run FP16 weights, which makes this the most like-for-like result on the page. Across 1,000 recordings, with NukeTorch scheduling the clips itself, the gain is 1.77×.

What it means for you

If you ship Whisper in a Mac app, this is the result that transfers directly: the same weights, the same precision, no retraining, and 1,000 recordings transcribed in 2¼ minutes instead of 4, in 3.7 GB of memory instead of 5.1 GB. Accuracy is 0.02 points better: 3.18% of words wrong against MLX’s 3.20%.

Nuance
  • FP16 weights on every engine.
  • Decoding settings differ slightly: all three decode greedily in English, but MLX also produces its default timestamp tokens, PyTorch caps output at 64 new tokens, and NukeTorch produces no timestamps.
  • The bulk comparison is NukeTorch’s default bulk mode against MLX one clip at a time, as it ships.
  • The one-clip sample was generated with Kokoro, not recorded. The 1,000-clip test uses real recordings.
  • The comparison buyers will ask for, WhisperKit on Mac and faster-whisper on NVIDIA, has not been run yet.

How to use it

macOS on Apple silicon · the same model files and checkpoint revision you use today · licensing undetermined

Benchmark reportReproducing the Whisper large-v3-turbo results
MachineApple M4 Max, 40 GPU cores, 128 GB, macOS 26.7
Reportv1.1. Author-run measurements; no independent reproduction yet.
One clipMedian of 60 timed calls per engine (2 rounds of 30)
BulkMedian of 5 passes per engine; PyTorch not tested on the 1,000-clip batch

Exactly what was run

ModelOpenAI Whisper large-v3-turbo (809M parameters, multilingual)
CheckpointsOpenAI openai/whisper-large-v3-turbo@41f01f3fe87f28c78e2fbf8b568835947dd65ed9 (PyTorch); mlx-community/whisper-large-v3-turbo@a4aaeec0636e6fef84abdcbe3544cb2bf7e9f6fb (MLX); NukeTorch runs the same model
NukeTorchRelease build; batch API for the batch series
MLXmlx-whisper 0.4.3 on MLX 0.32.2, run as shipped
PyTorchTransformers 5.3.0 on PyTorch 2.10.0, model and inputs on the MPS device, torch.mps.synchronize() before each timer stop
Reference precisionMLX and PyTorch: FP16 weights as shipped
DecodingGreedy decoding, English. NukeTorch and PyTorch decode without timestamps; MLX runs as shipped, which emits its default timestamp tokens (these are stripped from the text but are decoded). MLX: temperature 0, previous-text conditioning off. PyTorch: eager attention, at most 64 new tokens.
Single-request clipSynthetic English sentence generated with Kokoro (voice af_heart) at 24 kHz, resampled to 16 kHz mono, zero-padded to 6.4 s
Batch corpus1,000 utterances of LibriSpeech test-clean (the first 1,000 in ID order, utterances over 30 s skipped; OpenSLR SLR12, CC BY 4.0). FLAC converted losslessly to 16-bit WAV; reference transcripts from the corpus
MachineApple M4 Max, 40-core GPU, 128 GB unified memory, macOS 26.7, on AC power, no other user workload during timing
RunsOne run. Single-request rounds alternated engines; the batch passes ran all NukeTorch models, then all MLX models.
WarmupsNukeTorch 3, MLX and PyTorch 2 before each 30-call round; 2 utterances before each batch pass

How each number is measured

Single requestMedian of 60 timed calls (two rounds of 30, pooled). No confidence interval or significance test is claimed.
BatchHeadline is the median of the five pass times. Throughput = 1,000 ÷ wall time. RTFx = seconds of input audio ÷ processing seconds (higher is faster). Output tokens/s = a pass's transcript tokens ÷ that pass's wall time, median of the five.
Batch executionNukeTorch's batch API receives the whole list; MLX's runner processes one utterance at a time. Both are reported as system throughput at each engine's default settings.
Speed ratioFor times, other engine's time ÷ NukeTorch's time; for rates, NukeTorch's rate ÷ other engine's rate. The percentage is (ratio − 1) × 100 and describes speed, not time saved. Ratios are computed from unrounded values.
Warm-upSingle request: each 30-call round is preceded by warm-up calls (3 for NukeTorch, 2 for MLX and PyTorch) that are not timed. Batch: 2 warm-up utterances before each pass, not timed. Each pass ran in its own process, one after another. Machine temperature was not recorded.
WERSubstitutions + deletions + insertions ÷ reference words, after lowercasing and removing punctuation.
ScopeSixty calls on one clip and five passes over one English read-speech corpus on one machine. They do not establish behaviour on other speakers, languages, domains or hardware.

What this does not show

  • One machine. No P95 or P99, no statistical significance claim.
  • Clips of up to 30 seconds; no hour-long files. English read speech only.
  • No streaming test: every figure is the full transcript, so there is no time-to-first-word figure.
  • The bulk comparison is NukeTorch’s own scheduling against MLX one clip at a time, not matched concurrency.
  • No comparison yet with other Mac runtimes such as WhisperKit or whisper.cpp, or with NVIDIA.

Speech to text · model 3 of 3

Whisper Tiny

The smallest model: 2.0× faster than MLX on a single clip, 2.5× across 1,000 recordings, and 11× faster than PyTorch.

Same precision

OpenAI · Apache 2.0 · multilingual · One clip: 6.4 seconds of audio holding 5.6 seconds of English speech, median of 60 calls · Bulk: 1,000 LibriSpeech recordings, median of 5 passes

MetricNukeTorchMLXPyTorchvs MLXvs PyTorch
One clip, request time (ms)13.1926.62151.582.02× faster11.49× faster
1,000 clips (s) §12.6332.03not run2.54× fastern/a
Clips per second, bulk79.1631.22not run2.54× fastern/a
Real-time factor, one clip (×)485.2240.442.22.02× faster11.49× faster
Real-time factor, bulk (×)584.9230.7not run2.54× fastern/a
Word error rate, 1,000 clips7.04%7.04%not runsamen/a
Peak memory, one clip (MB)3566941,63849% less78% less
Peak memory, 1,000 clips (MB) §9383,072not run69% lessn/a

§ 1,000 LibriSpeech test-clean recordings, 7,389 seconds of speech, median of five passes per engine. NukeTorch runs in its default bulk mode, which schedules the clips itself; MLX transcribes them one after another, as it ships. PyTorch was not tested on the 1,000-clip batch. Memory is everything each engine held at its peak, measured in runs separate from the timings. MLX keeps memory it has finished with in its own cache, so its bulk peak reflects its largest working set.

In detail

On a single clip Tiny is 11.5× faster than PyTorch and 2.0× faster than MLX; across 1,000 recordings it is 2.5× faster than MLX, with FP16 weights on every engine.

What it means for you

Tiny is the model people pick when speed and footprint matter more than accuracy: captions, voice commands, on-device search. Here 1,000 recordings take under 13 seconds instead of 32 under MLX, in 938 MB of memory instead of 3.1 GB, at exactly the same word error rate.

Nuance
  • FP16 weights on every engine.
  • The bulk comparison is NukeTorch’s default bulk mode against MLX one clip at a time, as it ships.
  • The one-clip sample was generated with Kokoro, not recorded. The 1,000-clip test uses real recordings.
  • Tiny's accuracy is modest on any engine: about 7% of words wrong on this corpus.

How to use it

macOS on Apple silicon · the same model files and checkpoint revision you use today · licensing undetermined

Benchmark reportReproducing the Whisper Tiny results
MachineApple M4 Max, 40 GPU cores, 128 GB, macOS 26.7
Reportv1.1. Author-run measurements; no independent reproduction yet.
One clipMedian of 60 timed calls per engine (2 rounds of 30)
BulkMedian of 5 passes per engine; PyTorch not tested on the 1,000-clip batch

Exactly what was run

ModelOpenAI Whisper Tiny (39M parameters, multilingual)
CheckpointsOpenAI openai/whisper-tiny@169d4a4341b33bc18d8881c4b69c2e104e1cc0af (PyTorch); mlx-community/whisper-tiny@78c52ab98ca87f570bc57ad852e15ef7060f9f76 (MLX); NukeTorch runs the same model
NukeTorchRelease build; batch API for the batch series
MLXmlx-whisper 0.4.3 on MLX 0.32.2, run as shipped
PyTorchTransformers 5.3.0 on PyTorch 2.10.0, model and inputs on the MPS device, torch.mps.synchronize() before each timer stop
Reference precisionMLX and PyTorch: FP16 weights as shipped
DecodingGreedy decoding, English, no timestamps, on all three engines. MLX: temperature 0, previous-text conditioning off. PyTorch: eager attention, at most 64 new tokens.
Single-request clipSynthetic English sentence generated with Kokoro (voice af_heart) at 24 kHz, resampled to 16 kHz mono, zero-padded to 6.4 s
Batch corpus1,000 utterances of LibriSpeech test-clean (the first 1,000 in ID order, utterances over 30 s skipped; OpenSLR SLR12, CC BY 4.0). FLAC converted losslessly to 16-bit WAV; reference transcripts from the corpus
MachineApple M4 Max, 40-core GPU, 128 GB unified memory, macOS 26.7, on AC power, no other user workload during timing
RunsOne session. Single-request rounds per engine; batch passes per engine, MLX then NukeTorch.
WarmupsNukeTorch 3, MLX and PyTorch 2 before each 30-call round; 2 utterances before each batch pass

How each number is measured

Single requestMedian of 60 timed calls (two rounds of 30, pooled). No confidence interval or significance test is claimed.
BatchHeadline is the median of the five pass times. Throughput = 1,000 ÷ wall time. RTFx = seconds of input audio ÷ processing seconds (higher is faster). Output tokens/s = a pass's transcript tokens ÷ that pass's wall time, median of the five.
Batch executionNukeTorch's batch API receives the whole list; MLX's runner processes one utterance at a time. Both are reported as system throughput at each engine's default settings.
Speed ratioFor times, other engine's time ÷ NukeTorch's time; for rates, NukeTorch's rate ÷ other engine's rate. The percentage is (ratio − 1) × 100 and describes speed, not time saved. Ratios are computed from unrounded values.
Warm-upSingle request: each 30-call round is preceded by warm-up calls (3 for NukeTorch, 2 for MLX and PyTorch) that are not timed. Batch: 2 warm-up utterances before each pass, not timed. Each pass ran in its own process, one after another. Machine temperature was not recorded.
WERSubstitutions + deletions + insertions ÷ reference words, after lowercasing and removing punctuation.
ScopeSixty calls on one clip and five passes over one English read-speech corpus on one machine. They do not establish behaviour on other speakers, languages, domains or hardware.

What this does not show

  • One machine. No P95 or P99, no statistical significance claim.
  • Clips of up to 30 seconds; no hour-long files. English read speech only.
  • No streaming test: every figure is the full transcript, so there is no time-to-first-word figure.
  • The bulk comparison is NukeTorch’s own scheduling against MLX one clip at a time, not matched concurrency.
  • No comparison yet with other Mac runtimes such as WhisperKit or whisper.cpp, or with NVIDIA.

Get in touch

A demo, a pilot, or just a question — write to us.