Comparison
NUMU PRO vs running Whisper yourself
Where each tool wins, without pretending the other one is bad.
Both are built on the same class of speech recognition. Run well, self-hosted Whisper produces excellent raw transcripts.
Choose OpenAI Whisper if…
Open weights, no per-minute cost if you own the hardware, and full control over where the data goes.
you have GPU capacity, engineering time, and a strict requirement to keep everything in-house.
Choose NUMU PRO if…
you want speaker diarization, summaries, subtitles, and an editor without building and maintaining that pipeline yourself.
Side by side
| NUMU PRO | OpenAI Whisper | |
|---|---|---|
| Setup | Upload a file | Install, configure, maintain a GPU pipeline |
| Diarization | Built in | Separate model you wire up yourself |
| Summaries | Built in | A second model and prompt layer |
| Subtitles | SRT, VTT, ASS, burn-in | Basic output, formatting is on you |
| Cost model | From $0.05 / min at volume | Hardware and engineering time |
OpenAI Whisper: common questions
If Whisper is free, why pay anything?
Because raw Whisper gives you text and nothing else. Diarization, summaries, subtitle timing, an editor, and export formats are all separate work you build and maintain. The model is free; the pipeline is not.
When should I self-host instead?
When data cannot leave your infrastructure at all, you have GPU capacity, and someone owns the pipeline as part of their job. Those conditions are real and we will not argue you out of them.
Can I get on-premise from you instead?
Yes — the Enterprise plan runs the engine inside your infrastructure, including air-gapped. You get the surrounding product without sending recordings out.
Try it on your own audio.
Sixty free minutes a month is enough to compare accuracy yourself instead of trusting a table.