Speech is Cheap / Measured results

A complete test of the transcription pipeline.

Explore every eligible LibriSpeech-Long test recording, its reference transcript, and the words returned by our pipeline. This evaluation ran locally on an RTX 3060 on 2026-09-18 UTC.

476 recordings 73 speakers Frozen before test inference
Measured local pipeline accuracy. The public API was not part of this run. These recordings contain English read audiobooks, up to four minutes each. Cloud latency, API reliability, punctuation, speaker accuracy, and competitor comparisons were not evaluated.

Word error rate
Reference words / audio
Failed first attempts
Median local request

WER and speaker uncertainty

10,000 seeded bootstrap draws, resampling speakers within each split. Whiskers show marginal 95% intervals. Pooled WER is weighted by reference words.

Every recording

CleanOtherClick a point to inspect its transcript

Each point is one recording. The same speaker may contribute several recordings. Points are not independent speakers.

Inspect the recordings

Search and diagnostic filters affect this table only. The headline results always describe the complete selected split. CSV exports contain primary scoring fields and a separate lexical_wer column.

RecordingSplitAudioWordsWERS / D / IRequestStatus

Read the result correctly

What was measured

One first attempt for every eligible recording in the two held-out test splits. The twelve development pilot recordings and their timing repeat are excluded. There was no WER-based stopping and no configuration tuning after the protocol was frozen.

How words and errors are counted

WER = substitutions + deletions + insertions, divided by normalized reference words. The aggregate adds the edit counts and reference-word counts across recordings before dividing. It is not the average of recording percentages, and 100% minus WER is not a general accuracy guarantee.

The primary score uses the pinned Whisper English normalizer. It normalizes case, punctuation, numbers, contractions and some spelling variants; it also removes some fillers and bracketed labels. The sensitivity score uses Unicode NFKC, case folding, punctuation/symbol removal, and whitespace collapse only. Both reference and hypothesis receive the same selected normalization. Raw and normalized text are available for every recording.

Failed or timed-out first attempts are empty hypotheses in the primary all-input score. Retries never replace them. The completed-only view is secondary. No test retries were run.

What the confidence interval means

The score is exact for these fixed inputs, outputs, and scoring rules. The speaker bootstrap estimates how the ratio changes with the speaker mix. It retains all recordings of each resampled speaker and resamples separately within clean and other splits. The intervals do not cover shared-book dependence, reference errors, benchmark exposure in training, or performance on other kinds of audio.

Local timing, costs, and automatic language selection

Automatic language selection remained enabled. Non-English labels in this English corpus are diagnostic observations, not verified language annotations. No transcript was corrected or removed because of its language label.

No cloud GPU compute was purchased. The existing home workstation was used; electricity was not metered. These timing measurements do not establish public API latency or a cloud-provider speed comparison.

Sources, versions, and exclusions

Pinned LibriSpeech-Long dataset. CC BY 4.0 attribution: Park et al., Long-Form Speech Generation with Spoken Language Models, 2024; Panayotov et al., LibriSpeech: an ASR corpus based on public domain audio books, 2015.

Whisper English normalizer revision c0d2f624c09dc18e709e37c2ad90c039a4eb72a2; jiwer 3.1.0; RapidFuzz 3.14.1. Full input hashes and normalized transcripts are in the download. This scoring configuration is not an assertion of parity with any external leaderboard.

References and model training exposure have not been independently audited. The evaluated build is pinned; current live production equivalence has not been independently reattested. There is no matched competitor run.

Reference

Pipeline output

Word errors under the selected normalization
Segment language diagnostics and input hashes