Menu
HomeDriversSkillsSolutionsUpdatesGitHub

PhanthyMotus

Back to updates
Releaseperception2026-09-11

English transcription holds up in noise

A new parakeet-en model trained on ~1.7M hours of diverse audio — deliberately including non-speech material to suppress hallucination — replaces an audiobook-trained option that was a poor match for a robot microphone. Punctuation and casing included, 104 MB, RTF 0.039 on CPU.

English transcription holds up in noise

Speech recognition gains an English model, parakeet-en, chosen for what a robot's microphone actually hears.

Why

The only English-only model here was trained on 960 hours of clean audiobook narration — precisely the distribution an embodied robot is least like. There is always cooling-fan rumble in a robot's microphone, which left the option effectively unusable.

parakeet-en is trained on roughly 1.7 million hours of diverse audio, and deliberately includes non-speech material so it does not hallucinate words out of silence.

What you get

Punctuation and casingIncluded in the transcript; the old model had neither
Size104 MB int8 — the smallest offline English model here
SpeedReal-time factor 0.039, CPU only, measured on an Orin 6

A 7.4-second clip transcribes in 0.29 seconds, punctuated, with proper nouns correctly capitalized.

Pick parakeet-en from the speech-recognition model dropdown on the card. CPU only for now — the GPU combination has not been benchmarked and its transcripts have not been reviewed, and we do not list combinations we have not checked.