# LibriSpeech-Long: measured Speech is Cheap pipeline WER

The complete eligible test corpus was evaluated locally on an NVIDIA GeForce RTX 3060 on 2026-09-18 UTC. This measures the pinned Speech is Cheap inference pipeline. Requests did not pass through the public Jobs API or RunPod cloud infrastructure.

| Split | Recordings | Speakers | Reference words | S / D / I | WER | 95% speaker-bootstrap interval |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| Test clean | 269 | 40 | 146,398 | 2488 / 649 / 4342 | 5.11% | 4.44% to 5.87% |
| Test other | 207 | 33 | 106,488 | 4442 / 2239 / 3159 | 9.24% | 7.82% to 10.78% |
| Pooled | 476 | 73 | 252,886 | 6930 / 2888 / 7501 | 6.85% | 6.11% to 7.64% |

The all-input primary result is 17,319 word errors over 252,886 normalized reference words. The 476 selected recordings contain 25.5075 audio hours. Failed first attempts: 0 / 476. No test retries were run. The twelve development pilot recordings and one pilot timing repeat are excluded.

WER is the ratio of summed edit counts to summed reference words, not the mean of per-recording percentages. It does not measure punctuation, timestamp alignment, speaker attribution, or semantic correctness. Do not convert it to an unqualified “accuracy percentage.”

The secondary lexical-only normalization produces 7.17% pooled WER. Its per-split results and intervals are in `results.json`. This sensitivity result uses the same recorded predictions. It is not another model run.

The sum of local HTTP request times was 1352.675 seconds; the full collection interval was 23.507 minutes. The sum of handler intervals from Docker log timestamps was 1350.144 seconds. The warm-request median was 2.365 seconds and p95 was 5.970 seconds. First requests on workers and warm requests are separated in the JSON. Collection elapsed includes worker replacements and runner overhead. Local HTTP aggregate throughput was 67.89 times real time; collection throughput was 65.11 times real time.

These timing statistics describe this local run. Report preparation and occasional browser checks ran concurrently on the same host, and no CPU or memory cgroup limits were imposed. This was not an isolated performance experiment. Handler times are logging-based proxies, not RunPod billed execution times. They do not establish cloud latency or a provider comparison. No cloud compute was purchased. Electricity was not metered, so total electricity cost is unknown.

The inputs and scoring protocol were frozen before test inference. Dataset revision: `a5bb3ef1a678743ef6c6d1014f8f731d8c56c5e7`. Selection seed: `sic106-20260918-v1`. The selected set is 269 clean and 207 other test recordings across 40 and 33 speakers respectively. The predeclared exclusion is `test_clean/6930-81414-0002`, a 5.34-second recording below the supported six-second minimum. It was not padded or merged. Describe this as the eligible 476-recording corpus, not the unmodified 477-recording corpus.

Requests used automatic language selection, 30-second segments, and confidence threshold 0.5. Audio labels, speaker parsing, word timestamps, streaming, and split channels were off. The local request deadline was 60 seconds. The worker was replaced after every 100 jobs. First failures or timeouts remain empty hypotheses in the all-input score; completed-only metrics are secondary. No model settings changed after the pilot. Operational checks at 25%, 50%, and 75% did not inspect accuracy.

The primary normalizer is the unchanged OpenAI Whisper English normalizer at revision `c0d2f624c09dc18e709e37c2ad90c039a4eb72a2`. Edit alignment uses jiwer 3.1.0 and RapidFuzz 3.14.1. `normalization-examples.json` documents number, contraction, filler, punctuation, label, and spelling behavior. The complete scorer and dependency lock are included. This is not a claim of exact scoring parity with an external leaderboard.

Confidence intervals use 10,000 bootstrap draws with seed 10620260918. Each draw resamples speakers within each split and retains every recording for a selected speaker. The word-weighted ratio is recomputed on each draw; the 2.5th and 97.5th percentiles form marginal intervals. The score is exact for the fixed corpus and outputs. The intervals describe sensitivity to speaker mix and do not correct reference errors, shared-book dependence, corpus bias, or benchmark exposure during model training.

Automatic detection assigned non-English labels to 469 segments across 173 recordings. All returned text was retained. These are model labels, not verified language annotations. Inspect the segment-level data in the explorer; no references or outputs were manually corrected.

The evidence applies to English read audiobooks, with reconstructed recordings up to four minutes. It does not establish conversational, overlapping-speaker, multilingual, diarization, or hour-long accuracy. References and model training exposure have not been independently audited. There is no matched competitor run. Current live production equivalence has not been independently reattested.

Reproduce the metrics offline after extracting the bundle:

```sh
docker build -t sic106-results .
docker run --rm --network none -v "$PWD:/data:ro" sic106-results
```

`reproduce.py` verifies reference hashes, recomputes every normalized transcript and edit count, and reproduces all primary/secondary, all-input/completed-only aggregate metrics and confidence intervals. Audio is not required to rescore saved predictions. Audio-byte and decoded-PCM hashes are provided for a later inference reproduction; the original source is linked below. Full internal runtime logs and deployment pins are preserved in the private engineering handoff, rather than embedded in this public bundle.

Files: `results.json` is the complete machine-readable record; `results.csv` gives one row per recording; `predictions.jsonl` retains raw/normalized text, alignments, and returned segments; `manifest.json` lists identifiers and input hashes; `scoring.py` and `whisper_normalizers/` contain the scorer. The HTML is self-contained and makes no external data requests. CSV `wer` is the primary normalized ratio, while `lexical_wer` is the secondary ratio. Percent displays multiply either by 100.

[Pinned dataset](https://huggingface.co/datasets/ilyakam/librispeech-long/tree/a5bb3ef1a678743ef6c6d1014f8f731d8c56c5e7). CC BY 4.0 attribution: Park et al., *Long-Form Speech Generation with Spoken Language Models*, 2024; Panayotov et al., *LibriSpeech: an ASR corpus based on public domain audio books*, 2015. The Whisper normalizer is MIT-licensed; see `WHISPER_LICENSE`.

Scorer SHA-256: `5ab68463de1c123f05b8c92ec045f4845e1316c2279fdcec675ccd0b613bece5`. Frozen private input-manifest SHA-256: `d81ec18d11934a3a6f9c17120317816faffe925f04d17b9d79bf05cbbb3194c5`.
