Transcription

Speech to text, captions and subtitles.

Nothing launched in this category yet.

Be the first

About Transcription

Transcription became a solved problem faster than most people noticed, and the practical consequence is that the interesting tools in this subcategory are no longer competing on accuracy. They compete on what happens around the transcript: speaker identification, searchability, editing workflows, translation and the ability to run without sending your audio anywhere.

Accuracy differences that remain are concentrated in specific places. Names, technical jargon, acronyms and heavily accented or overlapping speech are where errors cluster, and tools that let you supply a vocabulary list of expected terms will noticeably outperform ones that do not on domain-specific material. If you transcribe the same kind of content repeatedly, this feature is worth more than a general accuracy claim.

Speaker attribution — knowing who said what — is harder than transcription and varies much more between tools. It degrades sharply with overlapping speech and with participants who sound similar. If attribution matters for your use, test on a genuinely messy recording rather than a clean one, because clean recordings hide the difference entirely.

Local transcription is now practical on ordinary hardware and is the right default for anything sensitive. Interviews, medical or legal conversations, therapy sessions and internal meetings are all material you may have obligations about, and processing locally removes the question of what a vendor retains. The speed penalty on a modern laptop is modest.

Timestamps and formats are the difference between a transcript you can use and a wall of text. Word-level timestamps enable transcript-driven editing, jump-to-moment search and caption generation. Standard caption formats matter if the output is going onto video. Check what you get before assuming.

Translation and subtitling sit here too, and the honest caveat is that machine translation of speech compounds two error sources. For anything consequential, treat the output as a draft for a human who speaks the language. For making a recording roughly accessible to a wider audience, it is genuinely useful as-is.

From the blog

Reading on launching, ranking and transcription.

All posts