Crek
Blog

Speech-to-text todos in under a second with Groq Whisper and an LLM

How we made voice-to-todo feel instant: Groq Whisper writes down what was said, the user checks it, and the LLM drafts in the background while they read.

Crek

7 min read ·

You press a button, say “call mom, pay rent Friday and gym tomorrow at 7”, and three todos appear. Getting that to feel instant turned out to be less about a faster model and more about where the waiting happens.

This is how we rebuilt voice capture in Crek, a free habit tracker and to-do list, around Groq's Whisper API. The same shape works for any app that turns speech into structured data: notes, calendar events, expenses, form filling.

YOUCREKaudiotextready firstAddSpeakHEARDCall mom, pay rent Friday andgym tomorrow at 7 a.m.|Continue3 DRAFTSCall momPay rentFri, Oct 2GymOct 1 · 7 AMWrites it downWhisper on Groq · 0.9 sDrafts your todosGemini · 2.6 sstarts the moment the text is backSavedsame checks as typing
The new flow. The draft request goes out the moment the words are back, so by the time you press Continue it has usually answered. Change a word first and it drafts again.

The slow version: one call that did everything

The first version sent the recording to Gemini in a single request. Gemini accepts audio, so one call could transcribe the sentence and pull the todos out of it: a title, the details, a due date resolved against the user's timezone.

It worked, and it was simple. It also had three problems:

  • The user waited for all of it. Listening, writing down the transcript and extracting the todos all happened before anything appeared on screen.
  • The model typed the transcript back. We showed the transcript so people could spot a misheard word, which meant the model had to write the whole sentence out before it started on the todos. Output tokens are the slow part of any LLM reply.
  • A misheard word meant starting over. If it heard “call Tom” instead of “call mom”, the only fix was to record the whole thing again, which was another full call.

So we split the job in two, with the user in the middle.

BEFOREBrowserrecords WAVDraft route/api/todos/draftGeminiaudio inaudioaudiotranscript + todosdraftsAFTERserver · API keys live hereBrowserHeard screenTranscribe route/api/transcribe · newGroqwhisper-large-v3-turboDraft route/api/todos/draftGeminitext in, todos out1 · audioaudio, entexttranscript2 · text, at oncetexttodos onlydrafts
Before, one Gemini call heard the audio and wrote back the transcript and the todos. After, Whisper on Groq does the listening through a new route, and Gemini only ever reads text. Outlined boxes are new; both routes run on the server, which is the only place the API keys live.

Step 1: transcribe with Groq Whisper

Speech-to-text is a solved problem with fast, cheap models. Groq's speech-to-text API runs OpenAI's Whisper on its own hardware, and it is quick: in our tests a 3 to 5 second clip came back as text in 0.3 to 0.9 seconds, including the upload.

Which Whisper model

Groq offers two:

ModelSpeedWord error ratePrice per hour
whisper-large-v3-turbo216× real time12%$0.04
whisper-large-v3189× real time10.3%$0.111

We use turbo. It costs about a third as much, and the one thing large-v3 adds that matters is translation, which we don't need. The accuracy gap matters less than it looks, because in the new flow the user reads every transcript before anything is made from it.

Two details worth knowing before you pick either:

  • Every request is billed as at least 10 seconds. A two-second todo costs the same as a ten-second one, about $0.0001 on turbo.
  • Set the language if you know it. We pass language: "en". Detection is least reliable on exactly the short clips a voice input produces.

The request

It is a multipart upload with the file and a few fields. We call it with plain fetch rather than an SDK, because the app runs on Cloudflare Workers and a hand-written request has nothing Node-specific to trip over.

const body = new FormData();
body.append("file", new File([audio], "recording.wav", { type: "audio/wav" }));
body.append("model", "whisper-large-v3-turbo");
body.append("language", "en");
body.append("temperature", "0");

const response = await fetch("https://api.groq.com/openai/v1/audio/transcriptions", {
  method: "POST",
  headers: { Authorization: `Bearer ${process.env.GROQ_API_KEY}` },
  body,
});
const { text } = await response.json();

The browser never calls Groq itself. It posts the recording to our own route, and the server adds the key. A key shipped to the browser is a key anyone can read in the network tab.

On the recording side, browsers disagree about formats: Chrome records WebM, Safari MP4. We decode whatever the browser produced and re-encode it as 16 kHz mono WAV before uploading. That is one format from every browser, and 16 kHz mono is what Whisper converts everything to anyway. Groq accepts WebM and Ogg directly too, so sending the compressed original is a smaller upload if you don't need one format everywhere.

Whisper says “Thank you.” when nobody spoke

We sent three seconds of silence to see what came back, expecting an empty string. We got:

Thank you.

Room noise came back as ".". This is a known Whisper habit. It learned from audio paired with subtitles, and the quiet parts of that data are full of captions like “Thank you.” and “Thanks for watching!”. Hand it silence and it writes a caption.

The usual advice is to check each segment's no_speech_prob and drop the ones that are probably silence. On Groq that didn't help us: no_speech_prob was 0 for pure silence and for pure noise alike. So we reject transcripts that are only one of those phrases:

const SILENCE = new Set([
  "", "you", "bye", "thanks", "thank you",
  "thank you very much", "thanks for watching", "thank you for watching",
]);

function isSilence(transcript: string) {
  const words = transcript.toLowerCase().replace(/[^a-z' ]/g, "").trim();
  return SILENCE.has(words);
}

It matches the whole transcript, never a part of it: “thank you” inside a real sentence is speech. A silent recording now fails in under a second with “Didn't catch anything”, before an LLM ever sees it.

Step 2: show the words before anything is made of them

The transcript goes on screen as editable text under the word Heard. If Whisper wrote “Tom” where you said “mom”, you fix one word by typing it. You don't record the sentence again.

That screen also changes what accuracy means. With one call, a misheard word quietly became a wrong todo. With a confirm step, it's a typo you can see and fix before anything is created.

The obvious objection is that it adds a step, and the step adds a wait. It adds a step, but it doesn't have to add a wait.

Step 3: start the LLM call before the user confirms

The drafting takes Gemini about 2.6 seconds. Reading a sentence on the Heard screen takes a person a few seconds too. So the draft request goes out the moment the transcript arrives, not when the user presses Continue, and the two happen at the same time.

0 s24681012KEPT AS HEARDYourecordingreading “Heard”GroqGeminidone while you readContinuedrafts in 4 msEDITEDYoureading + editingGroqGeminithrown awaysent again
One run on our dev server. When the words are kept as heard, Gemini finishes while you are still reading and Continue shows the drafts in 4 ms. When you edit them, the early answer is thrown away and a fresh call runs after Continue.

When Continue is pressed, there are two cases:

// When the transcript arrives: start drafting, don't wait for it.
let pending = requestDraft(transcript);
showHeard(transcript);

// When the user presses Continue:
if (pending.transcript !== edited) {
  pending.controller.abort();   // they changed it: this answer is for other words
  pending = requestDraft(edited);
}
const drafts = await pending.outcome;

If the text is untouched, Continue collects the answer that's already there. In our test run that took 4 ms from the key press to three drafts on screen. If the text was edited, the early request is dropped and a new one is sent with the corrected words, and the user waits the normal couple of seconds.

A few things make this safe rather than just fast:

  • A late answer can't land on the wrong screen. Every exit (Escape, closing the overlay, recording again) aborts the request in flight. After waiting, the code also checks that the answer is still for the current request before showing it.
  • A failed early answer isn't reused. If Gemini is busy, “Try again” goes back to the Heard screen with the words kept and sends a fresh request. It doesn't hand back the same failure.
  • Aborting doesn't save money. Abort stops the browser waiting. The provider has already started work, and it still bills the call.

What it costs

The trade is two LLM calls instead of one whenever someone edits the transcript. The early one is wasted. With a text-only prompt that is a small cost for removing the whole wait from the common case. We log whether each transcript was edited, which tells us how often we pay it and whether the confirm step is worth its tap.

Step 4: send the LLM text, and don't make it type the transcript back

Gemini now gets the confirmed transcript as a plain text part after the prompt, and nothing else. Two small changes came with that:

  • No audio tokens in. Text is a far smaller input than a recording.
  • No transcript out. The old schema asked for the transcript alongside the todos, so the model spent output tokens retyping its own input. The new schema asks only for the todos, and the server puts the transcript back on each draft from the request.

The reply is still treated as untrusted input. Each draft goes through the same validator as a todo typed by hand before it can be saved, so whatever the model invents, or whatever a tampered client sends, gets the same checks.

Keeping older app versions working

Installed copies of our mobile app still post the recording straight to the draft endpoint and expect todos back. Rather than break them, the endpoint checks what it was sent. JSON means a confirmed transcript from the new flow. Multipart means a recording from an older build, so it transcribes it with Groq first and then drafts. Those builds keep working and get faster, just without the confirm step, and the audio branch comes out once nobody runs them.

What we measured, and what's left

On our dev server, with a sentence of a few seconds:

  • Transcription: 0.3 to 0.9 s with whisper-large-v3-turbo on Groq.
  • Drafting: about 2.6 s with Gemini on text, hidden behind the reading time.
  • Continue to drafts: 4 ms when the transcript was kept as heard.

What we haven't done yet, written down so it doesn't get forgotten:

  • No timeout on the Groq or Gemini calls. A hung upstream holds the request open.
  • No retry when Gemini answers 503 “high demand”, which happened often enough in testing to notice.
  • English only. The language is fixed to en. Supporting more means passing the user's language, or letting Whisper detect it and accepting the misses on short clips.

The part that surprised us most wasn't the model at all. Most of the wait was never removed; it was moved to where the user was already busy reading.

Questions

How fast is Groq's Whisper API?

In our tests, whisper-large-v3-turbo on Groq returned the transcript of a 3 to 5 second clip in 0.3 to 0.9 seconds, upload included. Groq lists the turbo model at about 216 times real time and $0.04 per hour of audio, and bills every request as at least 10 seconds.

Why does Whisper return “Thank you” when nobody spoke?

Whisper learned from audio paired with subtitles, and quiet stretches in that data often carry captions like “Thank you.” or “Thanks for watching!”. Given silence, it writes one. On Groq the no_speech_prob value was 0 even for pure silence in our tests, so we reject any transcript made up only of those phrases.

Should I send audio straight to a multimodal LLM or transcribe it first?

One multimodal call is simpler, but the user waits for all of it at once and can only fix a misheard word by recording again. Transcribing first with a fast speech-to-text model shows the words within a second, lets the user correct them, and hands the LLM text, which is faster and cheaper than audio.

Does starting the LLM request before the user confirms waste tokens?

Only when the user edits the text. The early answer is then thrown away and a second call runs on the corrected words. Aborting the first request in the browser does not stop the provider billing it. With a text-only prompt that second call is small, and it is worth logging how often transcripts get edited.