Transcription is live again: 60 free minutes a month, Pro is $19/mo. Other tools are coming back. Students: lecture notes still live at VocalScribe.
A clean transcript starts with a clean recording. Put the microphone close to each speaker, record in a quiet room, have people speak one at a time, and do a short test before you start. Then upload the audio track, set the language, and tick Label different speakers if more than one person talks.
Transcription tools turn speech into text. They work best when the speech is the loudest, clearest thing in the file. Every other sound, from an air conditioner to a second conversation, is something the model has to guess around. Most of the mistakes people blame on transcription software are decided before the file is ever uploaded, at the moment of recording. The good news is that the fixes are cheap. Almost all of them are about where you put the microphone and how people speak, not about buying equipment.
This guide applies to every kind of recording people bring to Kai The Scribe: meetings, interviews, podcasts, lectures, voice memos and video narration.
Distance is the biggest factor you control. The further a microphone is from a mouth, the more it records the room instead of the voice: echo off walls, the hum of electronics, chairs moving. A phone or laptop microphone across a table is the most common reason a transcript comes back patchy.
Background noise competes with speech. Music, traffic, a busy café and the hum of a fan all make it harder to pick out words.
Crosstalk, where two people talk at once, is the hardest thing for any transcription tool to separate. It also confuses speaker labelling, because the words from two voices land in the same moment.
Record ten seconds of every voice, then play it back on headphones. Listen for a muted or wrong microphone, a hum, one voice much quieter than the others, and clipping, the harsh crackle when a level is set too high. A test takes less than a minute and catches the problems that would otherwise cost you the whole session. Do one each time you use new equipment or a new room.
Kai The Scribe accepts MP3, WAV, M4A, MP4, MOV and WEBM files. A few choices make the upload smoother:
Even a clean recording produces a draft, not a final record. Names, numbers and technical terms are where any transcript is most likely to slip. Search for them, fix the spelling once with find and replace, and play back any line you plan to quote or publish. For step-by-step guides by recording type, see transcribing an interview, transcribing a lecture, and podcast transcription.
Get the microphone closer to the person speaking. Distance lets room echo and background noise into the recording, and no setting in a transcription tool can fully take them back out.
No. A basic external or clip-on microphone near the speaker usually beats an expensive one across the room. Position matters more than price.
Any of MP3, WAV, M4A, MP4, MOV or WEBM works. For long recordings, upload the audio rather than the video: it is much smaller and keeps you under the 100 MB upload cap.
Kai The Scribe transcribes the audio you give it; it does not edit or denoise the file. If a recording is very noisy, an audio editor’s noise reduction can help, but it is no substitute for recording cleanly.
Yes, when you know it. Auto Detect works for most files, but picking the language yourself removes one thing that can go wrong, especially on a recording that opens with music or silence.