◒   PODCAST / STITCH LAB

CONSECUTIVE SPEECH · AN AUDIO FIELD STUDY

Listen between
the cuts.

What happens when good segments from the same speaker become one long sample? Hear the joins. Watch the spectrum. Decide for yourself.

Explore 15 examples ↓
15English listening examples
5retained-duration buckets
20,000manifests in the report
0fades or level adjustments

01 / LISTENING ROOM

Find a length. Follow the joins.

45.4 minutes, 175 joins. Three complete maximal runs per bucket. Videos load on demand; each white line marks a cut.

02 / THE CAPACITY REPORT

How much long-form audio?

Podcast data only · All languages
Snapshot: 8 October 2026

Estimated complete runs >60s14.53 million95% interval: 14.03–15.03 million
Estimated retained audio385,187 hours95% interval: 369,615–400,946 hours
Observed in the sample12,142 runs321.88 hours across 3,910 episodes

Most runs are just over a minute

Qualifying maximal runs in the 20,000-manifest sample. Select a bar to hear that bucket.

How long can a run get?

Empirical percentiles, seconds of retained audio.

PercentileAll runsRuns >60sEpisode max¹

¹ Longest run for each of all 20,000 sampled episodes; zero if no eligible run. The observed maximum is 19:02, not an estimate of the corpus maximum.

What counts as a consecutive run?

Every constituent must pass the manifest’s good-audio check, remain downstream eligible, contain transcript text, and have a known dominant speaker. We extend the run through adjacent listed segments with the same episode-local speaker ID. A different speaker or any ineligible segment breaks it, even when the first speaker returns later.

Only the retained segments are concatenated. Gaps between them are removed, with no crossfade, added silence or volume normalization. There is no maximum gap in the primary analysis. “Consecutive” describes listed-segment adjacency; it does not promise uninterrupted original speech. Each example shows the original times and the gap removed at each join.

Each complete maximal run strictly longer than 60 seconds is one sample; no overlapping combinations are counted. Source selection respects the manifest: enhanced segments use cleaned audio, and other segments use their correct original audio.

Estimation, sensitivity & limitations

Exactly 20,000 manifests were sampled from 24,212,736 processed podcast jobs with positive good-audio duration. Independent Bernoulli sampling followed by uniform subsampling used seed 208202610, with no duration cap or row-order limit. All manifests were read successfully. The sample contains 641,148 good segments and 2,108.80 good hours.

Counts and hours use sampled yield per good-audio second, calibrated to 2,523,517.61 corpus good-audio hours. Approximate 95% intervals use 2,000 episode-level bootstrap resamples. Quality, transcript and speaker labels are model outputs; this is not a human audit of training suitability, semantic coherence, voice identity or duplicate content. 14.12% of sampled good duration has no transcript.

A stricter ≤2-second gap rule estimates 14.33 million runs and 378,026 retained hours. Splitting primary runs greedily into whole-segment chunks >60s produces an estimated 17.46 million chunks and 327,694 hours, excluding short residual tails. These are alternative constructions, not extra audio to add to the complete-run estimate.

The 15 listening examples are curated English examples, not a random sample. Full-run manifest text passed language screening; Whisper-small language detection checked 30-second windows covering every example, requiring English as the top language with probability ≥0.8. These probabilities are not guarantees. Transcripts shown are manifest text, not newly corrected captions.

Videos encode exact 48 kHz mono float crops as AAC at 128 kbps; listening delivery is lossy. The underlying concatenations use exact sample counts. Spectrograms use 96 mel bands, 50–12,000 Hz, a 1,024-sample Hann window at 24 kHz, a 20 ms hop, and a fixed −100 to 0 dB power scale. Pixel columns average power over time; inspect joins by ear as well as visually.