# Podcast consecutive same-speaker audio: 20,000-manifest estimate Population snapshot: 2026-10-08 15:31:04.617032+00:00. **Exactly 20,000 randomly selected manifests read successfully**, all from processed podcasts with positive good-audio duration. All languages; no IPTV or YouTube. ## Capacity In the sample: **12,142 maximal runs longer than 60 seconds**, totaling **321.88 hours**, from **3,910 episodes**. | Output | Estimated corpus total (approx. 95% interval) | |---|---:| | Maximal consecutive runs >60 s | 14.53 (14.03–15.03) million | | Audio in those complete runs | 385,186.87 (369,615.48–400,946.43) hours | | Whole-segment chunks >60 s from those runs | 17.46 (16.77–18.16) million | | Audio in those chunks, excluding residual tails | 327,694.09 (314,731.60–340,966.69) hours | **Stricter gap rule:** if adjacent segments must be no more than two seconds apart on the original timeline, estimated maximal runs: **14.33 (13.83–14.83) million**; audio: **378,025.81 (362,776.59–393,460.24) hours**. ## Maximum concatenated length distribution Durations below are **seconds of retained audio**, not wall-clock spans. A maximal run is extended fully until a speaker change or an ineligible listed segment; it is not stopped at one minute. Percentiles are empirical sample percentiles, weighted equally per run or episode as indicated. The observed maximum is not an estimate of the corpus maximum. | Percentile | All maximal runs | Maximal runs >60 s | Longest run per sampled episode¹ | |---|---:|---:|---:| | P10 | 4.39 | 63.29 | 6.04 | | P25 | 6.91 | 68.11 | 14.03 | | P50 | 13.60 | 80.05 | 25.48 | | P75 | 19.66 | 103.53 | 51.32 | | P90 | 35.70 | 142.37 | 89.54 | | P95 | 50.82 | 179.01 | 124.64 | | P99 | 97.23 | 304.15 | 234.93 | | Observed max | 1,142.21 | 1,142.21 | 1,142.21 | ¹ Among all 20,000 sampled positive-good episodes; zero when no transcript-backed eligible run exists. Additional per-episode-speaker maximum and qualifying-episode distributions are in `results.json`. ## Rules and estimation - Each segment must pass literal `good_audio.passed=true`, remain downstream eligible, have a nonempty transcript and a known dominant speaker. Source identity, chronology and audio routing metadata are validated. - Only adjacent listed segments with the same episode-local speaker ID are joined. A failed-quality, downstream-excluded, untranscribed or unknown-speaker segment breaks the run. A speaker change breaks it even if the earlier speaker returns later. - Primary scenario imposes no extra limit on gaps between adjacent listed segments; gaps are removed from concatenated audio. The <=2-second sensitivity restricts those gaps. This measures adjacency of listed segments, not a guarantee of uninterrupted original-waveform speech. - Threshold is strictly >60 seconds on the 48 kHz crop grid. Every qualifying maximal run is one long sample; no audio reuse or overlapping-window combinatorics. The separate chunk count greedily packs whole segments until >60 seconds and starts another sample; short residual tails are excluded only from that chunk-duration total. - Sampling: a seeded independent Bernoulli row sample, then uniform subsampling to exactly 20,000, with no episode-length cap or row-order LIMIT. Failed reads are not replaced. Population: 24,212,736 positive-good podcast jobs, 2,523,517.61 good-audio hours. - Counts and hours use sampled yield per good-audio second calibrated to the exact database total. Approximate 95% intervals use 2,000 episode-level bootstrap resamples. Uncalibrated estimates are retained as a cross-check. - Existing manifest speaker and quality predictions are trusted; this does not independently verify ASR accuracy, voice consistency, semantic coherence or corpus-wide duplicate content. No audio downloads or production mutations. - Sample inspected: 641,148 good segments / 2,108.80 good hours. 14.12% of good duration has no transcript text. 0.00 eligible seconds fail an additional >=98% dominant-speaker-share criterion. - Seven run-boundary tests and 14 data-integrity checks passed; no unresolved manifest failures.