Skip to content
ToolProof

How to Transcribe Audio Offline for Free (8 Models Tested)

ChatGPT's audio upload is for paid plans only. We ran 8 free speech-to-text models offline on a 22-minute call to compare accuracy and speed.

Published
Reading time
7 min
Category
AI Tools

You can transcribe audio for free, without uploading it anywhere, by installing a desktop app such as Vibe (opens in a new tab) or Buzz (opens in a new tab) that runs an open speech-recognition model on your own computer. In our test on a real 22-minute conference call, OpenAI's Whisper large-v3-turbo and NVIDIA's Parakeet TDT 0.6B v2 were the most accurate, with word error rates of about 10 to 11 percent, and Parakeet finished about five times faster on the same CPU. The smallest models got about one word in five wrong.

ChatGPT can now transcribe audio files you upload, but OpenAI's help center says audio uploads are not available on the Free plan (opens in a new tab), are capped at 512 MB, and are processed on a best-effort basis, so very long recordings may time out. An offline app costs nothing, has no length limit beyond your patience, and keeps the recording on your computer.

Which free offline app to use

All three apps below are free and open source, and they run the speech model on your own computer. We didn't install them for this guide (our test machine has no desktop), so these details come from each project's official page as of October 2026.

AppBest forPlatformsModels
Vibe (opens in a new tab)Transcribing audio and video filesWindows, Mac, LinuxWhisper, Parakeet TDT v3, Nemotron
Buzz (opens in a new tab)Transcribing files, plus live microphone captionsWindows, Mac (Apple silicon), LinuxWhisper
Handy (opens in a new tab)Dictating into any text box with a shortcutWindows, Mac, LinuxWhisper, Parakeet V3
  • Vibe says it transcribes files "fully offline," handles batches of files, and exports to TXT, SRT, VTT, DOCX, PDF and other formats. It also lists speaker identification.
  • Buzz exports TXT, SRT and VTT. Its README warns that the Windows installer isn't signed, so Windows will show a warning you have to click through, and that the last version for Intel Macs is 1.4.5.
  • Handy is for speaking into apps, not for transcribing files you already have.

Each app lets you choose which model to download. That choice matters more than the app, so the rest of this guide is about the models.

Which model to choose

Based on our test, here is what we'd pick:

  • Most accurate: Whisper large-v3-turbo (often shown as "turbo" or "Large v3 Turbo"). It had the lowest error rate, but without a graphics card it took almost 10 minutes for 22 minutes of audio.
  • Best balance, especially without a graphics card: Parakeet. Version 2 (English only) was nearly as accurate as Whisper turbo and about five times faster. The apps above offer version 3, which also handles 25 European languages; in our test it was just as fast but less accurate, mainly because it returned no text for some short clips (details below).
  • If your app only offers smaller Whisper models: use base.en at minimum. small.en was clearly more accurate, but it was the slowest model in our test.
  • Avoid for anything important: the tiny models. They got roughly one word in five wrong on our phone-quality call, and Whisper tiny.en wasn't even faster than Parakeet.

If a model name ends in .en, it's the English-only version. OpenAI's Whisper documentation (opens in a new tab) says the English-only versions tend to do better in English, especially at the tiny and base sizes, so pick them when all your audio is in English.

Our results

We transcribed the same recording with eight models and compared each transcript with a human-made reference transcript.

ModelErrors on the call (WER)Time for the 22-minute callErrors on clean speech (of 99 words)Model filesPeak memory
Whisper large-v3-turbo10.5%9 min 37 s2988 MB2.3 GB
Parakeet TDT 0.6B v211.1%1 min 45 s1631 MB1.3 GB
Whisper small.en12.8%10 min 3 s0358 MB1.4 GB
Parakeet TDT 0.6B v314.0%1 min 46 s1640 MB1.3 GB
Whisper base.en15.9%2 min 53 s2153 MB0.8 GB
Moonshine base16.0%1 min 20 s3274 MB0.7 GB
Moonshine tiny20.7%54 s3118 MB0.5 GB
Whisper tiny.en21.4%1 min 45 s399 MB0.6 GB

WER (word error rate) counts every wrong, missing or extra word as a share of the words actually spoken, so 10.5% means roughly one word in ten needs fixing. Times are for 4 CPU cores with no graphics card; a newer laptop, a Mac with Apple silicon, or an app that uses your graphics card can be much faster, but the order of the models should stay about the same. Model sizes are for the compressed (int8) files we used; the same model can be a different size in each app.

What the numbers show:

  • Phone-quality speech is much harder than clean speech. On 41 seconds of clear, single-speaker speech (99 words), every model made three mistakes or fewer. On the call, with several speakers on phone lines, the best model got about one word in ten wrong.
  • Bigger isn't always slower. Whisper turbo was more accurate than Whisper small and still finished a little sooner (9 min 37 s versus 10 min 3 s). OpenAI's Whisper documentation (opens in a new tab) describes turbo as "an optimized version of large-v3 that offers faster transcription speed with a minimal degradation in accuracy".
  • The smallest model isn't the fastest. Whisper tiny.en took 1 min 45 s, the same as Parakeet v2, while making about twice as many errors. Model size on disk is only a rough guide to speed, so it is worth testing a short file before committing to a long one.
  • Parakeet v3 skipped some short clips. It returned no text for 17 of the 210 speech segments, all under five seconds long, such as "115 million." Parakeet v2 transcribed the same segments. That accounts for most of the gap between the two versions in our test. Apps that feed the model longer chunks of audio may not show the same problem.

What the mistakes looked like

Here is one sentence from the start of the call, as each model heard it. The reference transcript says: "Later we will conduct a question and answer session and instructions will follow at that time."

ModelTranscript
Whisper turboLater, we will conduct a class student answer session, and instructions will follow at the time.
Parakeet v2Later we will conduct a question and answer session and instructions will follow at the time.
Parakeet v3Later we will conduct a question and answer session and instructions will follow at the time.
Whisper base.enLater we look in Zaka Class to the Nancer session and instructions will follow at the time.
Moonshine baseLater, we look at ZACA class in an answer session, and instructions will follow at the time.

The company name, ZAGG, comes up 10 times in the call, and none of the eight models spelled it correctly even once. Names, product terms and numbers are where you should proofread first.

What we tested

  • Recording: a 21.8-minute earnings call (ZAGG Inc., third quarter 2020) with six speakers, from Rev's Earnings-21 dataset (opens in a new tab) (CC BY-SA 4.0). It comes with a human-made reference transcript, which is what makes an error rate possible. The audio is real conference-call quality, not studio sound.
  • Clean speech: four short clips, 41 seconds in total: three audiobook clips that ship with the models as test files, and the "ask not what your country can do for you" clip from the whisper.cpp samples (opens in a new tab).
  • Engine: sherpa-onnx (opens in a new tab) 1.13.8 with the int8 models from its model releases (opens in a new tab). The apps above use their own engines, so their speed will differ.
  • Splitting: we cut the call into 210 speech segments of up to 20 seconds with the Silero voice-activity model, and gave every model the same segments.
  • Scoring: both transcripts were normalized with Whisper's English text normalizer (so "3rd" and "third", or "um" and nothing, aren't counted as errors) and scored with the jiwer library.
  • Machine: a Linux cloud server with 4 Intel Xeon cores at 2.8 GHz and no graphics card, October 11, 2026. Each model ran alone, using 4 threads.

FAQ

Is ChatGPT's audio transcription free?

No. As of October 2026, OpenAI's help article (opens in a new tab) says audio uploads are available on paid ChatGPT plans and not on the Free plan. Files can be up to 512 MB.

What is the most accurate free transcription model?

In our test, Whisper large-v3-turbo, with a 10.5% word error rate on a phone-quality call, followed closely by Parakeet TDT 0.6B v2 at 11.1%. On clean, single-speaker audio, most models we tested were nearly perfect.

Do I need a graphics card to transcribe audio offline?

No. Our test machine had none. Parakeet v2 transcribed the 22-minute call in 1 minute 45 seconds on 4 CPU cores. A graphics card mainly helps the larger Whisper models, and Vibe and Buzz both list GPU support.

Can offline models transcribe languages other than English?

Yes, but not all of them. Whisper's models without ".en" in the name are multilingual, and NVIDIA's model card (opens in a new tab) says Parakeet v3 covers 25 European languages. The ".en" models and Parakeet v2 are English only. We only tested English.

Is offline transcription private?

The audio is processed on your computer instead of being uploaded, which is the main privacy advantage. Vibe, for example, says "no data ever leaves your device." Apps still need the internet to download models and updates, and optional extras such as AI summaries can use online services, so check what you turn on.

  • #transcription
  • #speech to text
  • #Whisper
  • #Parakeet
  • #privacy