How few tokens can an audio codec decode at once?

Descript DAC · Fish 400M MorePitchLoss1 · Fish DAC S2-Pro · Qwen3-TTS 12Hz
MR-STFT, log-mel and log1p-mel deterioration through a 50% error budget

Listen to the measured error bands

Open the interactive listening report →: original, full-decode control and independently windowed reconstructions for every codec, dataset, metric and positive-error band. The listening ladder holds one source per dataset fixed across every comparison. Its displayed percentages are individual clip increases, not the pooled confidence-bound thresholds below.

Answer at a 50% added-error budget

The table gives the smallest measured window whose upper 95% confidence bound is at most 50%, and for which all larger measured windows also pass. A temporal frame carries one ID from every codebook. Scalar IDs are the number of codebook indices delivered in that window.

CodecMetricFramesScalar IDsmsMean increase %95% CI %Clips
Descript DACMR-STFT76381.345.0[41.0, 49.3]60
Descript DACLog-mel L114126162.546.3[43.4, 49.6]60
Descript DACLog1p-mel L117153197.445.6[42.4, 49.3]60
Fish 400M MorePitchLoss1MR-STFT330139.342.5[38.9, 46.3]60
Fish 400M MorePitchLoss1Log-mel L1440185.839.4[36.9, 42.2]60
Fish 400M MorePitchLoss1Log1p-mel L1660278.642.3[38.2, 46.8]60
Fish DAC S2-ProMR-STFT313101439.643.1[37.8, 48.6]60
Fish DAC S2-ProLog-mel L1313101439.643.0[38.7, 47.7]60
Fish DAC S2-ProLog1p-mel L1767603529.439.4[32.0, 47.9]55
Qwen3-TTS 12HzMR-STFT464320.042.9[38.8, 47.3]60
Qwen3-TTS 12HzLog-mel L1696480.045.7[41.7, 49.9]60
Qwen3-TTS 12HzLog1p-mel L1152401200.039.8[32.5, 48.1]60

A 50% increase refers to the numeric reconstruction error. It does not mean a 50% loss in perceived quality. Each codec is compared with its own full-decode baseline; absolute errors also appear below so that a weak baseline cannot be mistaken for a better codec.

One window that satisfies all three metrics

For the pooled sample, take the largest of the three conservative 50% thresholds. This gives one measured window size that stays within the budget across all three metrics under the stated rule.

CodecMinimum temporal framesScalar codebook IDsWindow duration (ms)
Descript DAC17153197.4
Fish 400M MorePitchLoss1660278.6
Fish DAC S2-Pro767603529.4
Qwen3-TTS 12Hz152401200.0

How to compare these results

Compare both temporal frame counts and durations. Each Fish frame spans four times as much audio as a Descript frame. The joint thresholds above therefore measure different durations even when one codec uses fewer frames.

The full-decode baselines matter. Pooled MR-STFT baseline errors are Descript DAC: 0.987, Fish 400M MorePitchLoss1: 1.162, Fish DAC S2-Pro: 0.980, Qwen3-TTS 12Hz: 1.108. Relative tolerance to windowing describes context sensitivity rather than establishing the best absolute reconstruction quality.

The strict 5% upper-confidence-bound criterion is not established for S2-Pro in the tested sweep. The longest partial-window conditions include only the longest eligible clips. These unresolved results are retained rather than inferring an unmeasured cutoff.

Joint thresholds across the full error-budget range

Each cell satisfies the conservative threshold rule for MR-STFT, log-mel and log1p-mel together.

Budget %Descript DACFish 400M MorePitchLoss1Fish DAC S2-ProQwen3-TTS 12Hz
5160 frames / 1857.6 ms49 frames / 2275.6 msnot established189 frames / 15120.0 ms
1084 frames / 975.2 ms25 frames / 1161.0 ms384 frames / 17832.9 ms123 frames / 9840.0 ms
2043 frames / 499.2 ms15 frames / 696.6 ms251 frames / 11656.4 ms48 frames / 3840.0 ms
3029 frames / 336.7 ms10 frames / 464.4 ms138 frames / 6408.7 ms31 frames / 2480.0 ms
4022 frames / 255.4 ms8 frames / 371.5 ms90 frames / 4179.6 ms23 frames / 1840.0 ms
5017 frames / 197.4 ms6 frames / 278.6 ms76 frames / 3529.4 ms15 frames / 1200.0 ms

Shared sample and model geometry

CodecSample rateFrames/sCodebooksms/frame
Descript DAC4410086.133911.610
Fish 400M MorePitchLoss14410021.5331046.440
Fish DAC S2-Pro4410021.5331046.440
Qwen3-TTS 12Hz2400012.5001680.000

All models use the same 60 clips across Cartoon Voices, LibriSpeech Other 500, and Podcast 800. Selection uses seed 20260928; the default 20-clips-per-dataset sample matches the preceding Fish window experiment. All codebooks are retained. Clips are downmixed to mono and resampled to the codec rate without normalization.

Qwen tokenizer: native geometry and decoder controls

Qwen/Qwen3-TTS-Tokenizer-12Hz is the audio tokenizer, not a text-to-speech language model. Despite its name, its native temporal rate is 12.5 frames/s: 1,920 samples per frame at 24 kHz, or 80 ms. Each frame contains 16 discrete codebook IDs. This study retains every book and uses the pinned checkpoint and implementation.

The public Qwen decode helper normally splits long code sequences into 300-frame chunks with 25 frames of left context. That policy is not an independent-window experiment. Here the underlying model.decoder(codes) receives the complete tensor for the control and independent tensors for the treatments, with no shared context. Its eight-layer transformer has a per-layer 72-frame sliding window (5.76 s); this does not itself determine a safe measured cutoff.

Native durations, not nominal model names or scalar index counts, must be compared: a Qwen frame spans 80 ms, a Fish frame 46.44 ms, and a Descript frame 11.61 ms. Codebook sizes in the measured registry: [2048, 2048, 2048, 2048, 2048, 2048, 2048, 2048, 2048, 2048, 2048, 2048, 2048, 2048, 2048, 2048].

Sources: official HF model and the pinned Qwen implementation recorded in the protocol.

Qwen decoder size and measured CPU/GPU throughput

The deployed Qwen decoder contains 114.323 million parameters; the full encoder-plus-decoder contains 153.715 million. Counts include frozen codebook lookup parameters, not only trainable weights.

DeviceBatchFramesDecoded secondsMedian decode ms× realtimeObserved min–max ×
cuda11098.7251.97167.80165.42–168.76
cpu11098.722944.682.962.88–3.10

Batch-one decoding of one real Podcast clip; three warmup calls and ten measured calls per device. CUDA timing synchronizes before and after each call. This measures discrete lookup plus direct waveform decoding, excluding encoding, resampling, disk IO, model loading and metrics. GPU: NVIDIA L40S; CPU: 4 Torch threads. Both use FP32 and eager attention; TF32 is disabled. Observed min–max ranges describe repeat timing variability, not confidence intervals or performance guarantees on other clips/hardware.

All 16 codebooks have 2,048 entries: 176 fixed-width bits per frame, 2.2 kb/s, or 0.99 MB per hour of audio as an ideal bit-packed token payload (decimal MB, no headers, padding or entropy coding). Cached int64 tensors in this experiment instead use 5.76 MB/hour before serialization overhead; those files are not a packed codec format.

Actual package versions: torch 2.12.0+cu130, torchaudio 2.11.0a0+4e3e282, transformers 4.57.6, accelerate 1.12.0, qwen-tts 0.1.1, onnxruntime 1.30.0. The existing Torch installation was preserved. GPU checks verified finite reconstruction, all-book integer lookup geometry, and independent one-frame decode. The maximum difference between three distinct eight-frame windows decoded together and separately was 2.1e-06.

Full-roundtrip quality in the benchmark workbench

A separate normal workbench run evaluates all four codecs on exactly the same 60 checked source IDs: 240 successful full roundtrips, with no failed or missing metrics. Run ID: 0b4452295400. Reconstructions are saved for playback and comparison in the workbench. The following values are unweighted means of 20 clips per dataset, not decoder-window deterioration percentages.

Full roundtrip means; error bars show clip-to-clip sample standard deviation, not confidence intervals.
CodecDatasetNLog-mel ↓Mel L1 ↓Mel SC ↓MR-STFT ↓Speaker cosine ↑SI-SDR dB ↑SNR dB ↑
Descript DACcartoon-voices200.3580.00260.1400.9550.9219.2829.706
Descript DAClibrispeech-other-500200.3130.00300.1271.0510.93311.69512.122
Descript DACpodcast-800200.3020.00650.1110.9340.95211.66811.945
Fish 400M MorePitchLoss1cartoon-voices200.4620.00390.2381.1160.8514.3585.331
Fish 400M MorePitchLoss1librispeech-other-500200.4450.00550.2501.2520.7864.6395.831
Fish 400M MorePitchLoss1podcast-800200.4250.01150.2161.1100.8606.2816.665
Fish DAC S2-Procartoon-voices200.3400.00270.1550.9360.9326.3227.204
Fish DAC S2-Prolibrispeech-other-500200.3310.00410.1871.0400.9155.5706.978
Fish DAC S2-Propodcast-800200.3300.00890.1570.9730.9366.5037.306
Qwen3-TTS 12Hzcartoon-voices200.3770.00340.2181.0860.931-1.1412.717
Qwen3-TTS 12Hzlibrispeech-other-500200.3660.00450.2061.2000.9302.4354.662
Qwen3-TTS 12Hzpodcast-800200.3460.00970.1751.0430.9544.3305.625

Speaker similarity uses SpeechBrain ECAPA-TDNN at its native 16 kHz; all other metrics use the workbench’s 24 kHz sample-zero alignment, no loudness normalization. Waveform SNR and SI-SDR are phase/timing sensitive, not MOS. Log1p-mel is reported separately throughout the controlled window study. This run contains no new ASR transcriptions or WER claims. Source hashes, runtime fixes and complete per-clip metrics are stored under data/qwen-paired-benchmark/ and the run database.

The workbench uses CPU resampling and requests deterministic Torch algorithms; the controlled window experiment uses GPU resampling and feature extraction. Their definitions and source files match, but numerical paths can produce slightly different quantized codes and absolute errors. Each window-treatment ratio uses its own measured matched full control, never the separate workbench baseline.

Relationship to the preceding Fish experiment

The earlier 62-frame result used an absolute added log-mel margin of 0.01 nat. Against that sample baseline, this is about a 2.3% increase. The present report asks a different question: how small a window can satisfy budgets from 5% to 50% in each of three errors? A much smaller window at a 50% budget is therefore consistent with the earlier result. The sample, full-encode control and independent hard-join policy are shared.

Relative deterioration and uncertainty

Solid lines: ratio of mean errors. Shaded bands: paired bootstrap 95% intervals. Horizontal line: 50% budget.
The same measurements with the vertical range focused on 0–50% decisions.

The same comparison in token counts

Full vertical range: ratio of mean errors, with paired bootstrap 95% bands and the 50% budget.
Each temporal frame is a bundle of nine IDs for Descript, ten IDs for Fish, or sixteen IDs for Qwen. Descript frames cover 11.61 ms; Fish frames cover 46.44 ms; Qwen frames cover 80 ms.

Individual codebook tokens per decoder call

Scalar token count = temporal frames × retained codebooks. All books remain active; this is not an experiment dropping books. Each model therefore has a different minimum token bundle. Equal token counts do not imply equal duration or equal packed bit cost.

Absolute error by token count

Source-relative error against temporal frames per call. Dotted lines use the matched full-decode control, retaining the same eligible clips as the solid treatment curve.

Dataset breakdown by token count

The same subgroup measurements and 95% paired bootstrap bands with temporal token frames on the x axis. The vertical range is focused on the 50% decision boundary.

Token-count distributions across clips

Median and 10th–90th percentile of individual clip increases against temporal token frames per call. Shading shows clip variation, not a confidence interval; the CSV retains values outside the displayed range.

Absolute reconstruction error

Solid lines: source-versus-windowed error. Dotted lines: source-versus-full error on the same eligible clips.
CodecDatasetMR-STFT baselineLog-mel baselineLog1p-mel baseline
Descript DACoverall0.986550.324740.003440
Descript DACcartoon-voices0.955090.358480.002372
Descript DAClibrispeech-other-5001.068970.313780.002606
Descript DACpodcast-8000.935580.301970.005343
Fish 400M MorePitchLoss1overall1.161610.444790.005817
Fish 400M MorePitchLoss1cartoon-voices1.116620.461590.003575
Fish 400M MorePitchLoss1librispeech-other-5001.256840.447640.004732
Fish 400M MorePitchLoss1podcast-8001.111360.425120.009144
Fish DAC S2-Prooverall0.979640.333960.004422
Fish DAC S2-Procartoon-voices0.937830.340000.002466
Fish DAC S2-Prolibrispeech-other-5001.028240.331820.003617
Fish DAC S2-Propodcast-8000.972840.330070.007185
Qwen3-TTS 12Hzoverall1.108450.363420.004916
Qwen3-TTS 12Hzcartoon-voices1.085030.377350.003124
Qwen3-TTS 12Hzlibrispeech-other-5001.199380.367380.003906
Qwen3-TTS 12Hzpodcast-8001.040950.345520.007719

Decoder divergence at fixed codes

Full decoder output serves as the reference. These errors isolate the numerical effect of windowing; they are distinct from deterioration relative to the original source.

Source-relative error can decrease for an individual clip even while the windowed waveform differs from full decoding. Divergence from the control and increased error against the source answer different questions.

Error budgets from 5% through 50%

Minimum measured windows under the conservative upper-confidence-bound rule; duration normalizes differing token rates.

Descript DAC — threshold ranges

Budget %MetricMean-rule framesCI-rule framesCI-rule msMeasured fail/pass bracketPass CI %Clips
5MR-STFT6274859.1(73, 74] frames[3.7, 4.6]60
5Log-mel L11151231428.0(122, 123] frames[4.1, 5.0]60
5Log1p-mel L11411601857.6(159, 160] frames[3.8, 5.0]60
10MR-STFT3334394.7(33, 34] frames[8.3, 9.9]60
10Log-mel L16063731.4(62, 63] frames[8.6, 10.0]60
10Log1p-mel L18084975.2(83, 84] frames[7.9, 10.0]60
20MR-STFT1717197.4(16, 17] frames[16.7, 19.9]60
20Log-mel L13233383.1(32, 33] frames[17.3, 19.4]60
20Log1p-mel L14043499.2(42, 43] frames[16.6, 19.7]60
30MR-STFT1112139.3(11, 12] frames[23.4, 28.0]60
30Log-mel L12224278.6(23, 24] frames[24.8, 28.2]60
30Log1p-mel L12729336.7(28, 29] frames[24.5, 28.9]60
40MR-STFT89104.5(8, 9] frames[31.4, 37.7]60
40Log-mel L11718209.0(17, 18] frames[33.7, 38.5]60
40Log1p-mel L12022255.4(21, 22] frames[32.5, 38.4]60
50MR-STFT7781.3(6, 7] frames[41.0, 49.3]60
50Log-mel L11414162.5(13, 14] frames[43.4, 49.6]60
50Log1p-mel L11617197.4(16, 17] frames[42.4, 49.3]60

Fish 400M MorePitchLoss1 — threshold ranges

Budget %MetricMean-rule framesCI-rule framesCI-rule msMeasured fail/pass bracketPass CI %Clips
5MR-STFT32391811.2(38, 39] frames[2.5, 4.2]60
5Log-mel L130321486.1(31, 32] frames[3.6, 5.0]60
5Log1p-mel L140492275.6(48, 49] frames[2.5, 4.8]59
10MR-STFT1720928.8(19, 20] frames[6.3, 8.6]60
10Log-mel L11718835.9(17, 18] frames[7.7, 9.5]60
10Log1p-mel L121251161.0(24, 25] frames[6.3, 9.0]60
20MR-STFT99418.0(8, 9] frames[15.4, 19.3]60
20Log-mel L199418.0(8, 9] frames[16.7, 19.4]60
20Log1p-mel L11315696.6(14, 15] frames[13.3, 18.2]60
30MR-STFT56278.6(5, 6] frames[23.1, 28.5]60
30Log-mel L166278.6(5, 6] frames[25.6, 29.9]60
30Log1p-mel L1910464.4(9, 10] frames[20.8, 27.5]60
40MR-STFT44185.8(3, 4] frames[31.6, 38.3]60
40Log-mel L145232.2(4, 5] frames[29.4, 33.9]60
40Log1p-mel L178371.5(7, 8] frames[29.6, 37.0]60
50MR-STFT33139.3(2, 3] frames[38.9, 46.3]60
50Log-mel L134185.8(3, 4] frames[36.9, 42.2]60
50Log1p-mel L156278.6(5, 6] frames[38.2, 46.8]60

Fish DAC S2-Pro — threshold ranges

Budget %MetricMean-rule framesCI-rule framesCI-rule msMeasured fail/pass bracketPass CI %Clips
5MR-STFT384——not established in tested sweep[—, —]—
5Log-mel L1384——not established in tested sweep[—, —]—
5Log1p-mel L1384——not established in tested sweep[—, —]—
10MR-STFT24138417832.9(256, 384] frames[1.8, 7.3]4
10Log-mel L120025511842.2(254, 255] frames[6.2, 9.8]15
10Log1p-mel L138438417832.9(256, 384] frames[3.0, 8.9]4
20MR-STFT881265851.4(125, 126] frames[9.8, 19.2]23
20Log-mel L1701346222.9(133, 134] frames[12.7, 18.9]21
20Log1p-mel L115725111656.4(250, 251] frames[9.9, 18.3]15
30MR-STFT57632925.7(62, 63] frames[20.1, 29.2]58
30Log-mel L149612832.8(60, 61] frames[20.4, 29.2]58
30Log1p-mel L11051386408.7(137, 138] frames[19.9, 30.0]21
40MR-STFT35391811.2(38, 39] frames[29.2, 38.9]60
40Log-mel L135391811.2(38, 39] frames[29.4, 37.7]60
40Log1p-mel L180904179.6(89, 90] frames[25.2, 37.8]45
50MR-STFT24311439.6(30, 31] frames[37.8, 48.6]60
50Log-mel L125311439.6(30, 31] frames[38.7, 47.7]60
50Log1p-mel L161763529.4(75, 76] frames[32.0, 47.9]55

Qwen3-TTS 12Hz — threshold ranges

Budget %MetricMean-rule framesCI-rule framesCI-rule msMeasured fail/pass bracketPass CI %Clips
5MR-STFT12115312240.0(152, 153] frames[1.7, 3.6]15
5Log-mel L112215512400.0(154, 155] frames[2.5, 4.9]15
5Log1p-mel L116818915120.0(188, 189] frames[2.0, 4.0]12
10MR-STFT44624960.0(61, 62] frames[5.7, 9.4]32
10Log-mel L1481048320.0(103, 104] frames[5.3, 10.0]19
10Log1p-mel L1951239840.0(122, 123] frames[3.6, 9.6]16
20MR-STFT16211680.0(20, 21] frames[14.5, 19.7]60
20Log-mel L121231840.0(22, 23] frames[15.0, 19.4]60
20Log1p-mel L141483840.0(47, 48] frames[12.5, 19.8]50
30MR-STFT912960.0(11, 12] frames[20.7, 27.2]60
30Log-mel L112141120.0(13, 14] frames[23.4, 29.0]60
30Log1p-mel L124312480.0(30, 31] frames[17.5, 27.5]58
40MR-STFT56480.0(5, 6] frames[30.7, 38.9]60
40Log-mel L189720.0(8, 9] frames[30.3, 38.2]60
40Log1p-mel L115231840.0(22, 23] frames[24.6, 36.5]60
50MR-STFT34320.0(3, 4] frames[38.8, 47.3]60
50Log-mel L166480.0(5, 6] frames[41.7, 49.9]60
50Log1p-mel L112151200.0(14, 15] frames[32.5, 48.1]60

Differences between datasets

Dataset means and confidence bands expose subgroup differences. Each panel uses its own eligible clips.

Subgroup thresholds use the observed sweep; integer refinements target the pooled primary endpoint. Subgroup brackets can therefore be wider.

CodecDatasetMetric50% fail/pass framesmsPass CI %Clips
Descript DACcartoon-voicesLog1p-mel L1(22, 23]267.0[37.1, 48.4]20
Descript DACcartoon-voicesLog-mel L1(12, 13]150.9[42.6, 49.5]20
Descript DACcartoon-voicesMR-STFT(6, 7]81.3[37.3, 44.5]20
Descript DAClibrispeech-other-500Log1p-mel L1(18, 19]220.6[37.6, 48.0]20
Descript DAClibrispeech-other-500Log-mel L1(17, 18]209.0[40.7, 49.0]20
Descript DAClibrispeech-other-500MR-STFT(9, 10]116.1[33.2, 43.6]20
Descript DACpodcast-800Log1p-mel L1(14, 15]174.1[41.5, 50.0]20
Descript DACpodcast-800Log-mel L1(11, 12]139.3[42.7, 49.8]20
Descript DACpodcast-800MR-STFT(5, 6]69.7[35.0, 45.5]20
Fish 400M MorePitchLoss1cartoon-voicesLog1p-mel L1(6, 7]325.1[36.8, 47.3]20
Fish 400M MorePitchLoss1cartoon-voicesLog-mel L1(2, 3]139.3[41.1, 48.2]20
Fish 400M MorePitchLoss1cartoon-voicesMR-STFT(2, 3]139.3[34.3, 42.1]20
Fish 400M MorePitchLoss1librispeech-other-500Log1p-mel L1(4, 5]232.2[36.3, 43.9]20
Fish 400M MorePitchLoss1librispeech-other-500Log-mel L1(3, 4]185.8[41.7, 49.1]20
Fish 400M MorePitchLoss1librispeech-other-500MR-STFT(3, 4]185.8[37.2, 47.1]20
Fish 400M MorePitchLoss1podcast-800Log1p-mel L1(6, 7]325.1[32.1, 47.5]20
Fish 400M MorePitchLoss1podcast-800Log-mel L1(3, 4]185.8[32.6, 44.0]20
Fish 400M MorePitchLoss1podcast-800MR-STFT(2, 3]139.3[30.6, 46.3]20
Fish DAC S2-Procartoon-voicesLog1p-mel L1(112, 113]5247.7[9.4, 36.5]4
Fish DAC S2-Procartoon-voicesLog-mel L1(38, 39]1811.2[25.0, 41.4]20
Fish DAC S2-Procartoon-voicesMR-STFT(40, 48]2229.1[25.4, 47.1]20
Fish DAC S2-Prolibrispeech-other-500Log1p-mel L1(239, 240]11145.6[10.4, 47.7]6
Fish DAC S2-Prolibrispeech-other-500Log-mel L1(35, 36]1671.8[35.6, 50.0]20
Fish DAC S2-Prolibrispeech-other-500MR-STFT(57, 58]2693.5[26.9, 46.2]18
Fish DAC S2-Propodcast-800Log1p-mel L1(61, 62]2879.3[30.3, 47.7]20
Fish DAC S2-Propodcast-800Log-mel L1(24, 25]1161.0[35.5, 48.9]20
Fish DAC S2-Propodcast-800MR-STFT(14, 15]696.6[34.0, 48.2]20
Qwen3-TTS 12Hzcartoon-voicesLog1p-mel L1(28, 29]2320.0[23.2, 47.7]20
Qwen3-TTS 12Hzcartoon-voicesLog-mel L1(10, 11]880.0[32.9, 45.9]20
Qwen3-TTS 12Hzcartoon-voicesMR-STFT(6, 7]560.0[33.6, 48.8]20
Qwen3-TTS 12Hzlibrispeech-other-500Log1p-mel L1(8, 9]720.0[28.3, 46.3]20
Qwen3-TTS 12Hzlibrispeech-other-500Log-mel L1(4, 5]400.0[37.1, 45.0]20
Qwen3-TTS 12Hzlibrispeech-other-500MR-STFT(2, 3]240.0[34.2, 44.2]20
Qwen3-TTS 12Hzpodcast-800Log1p-mel L1(16, 18]1440.0[23.5, 39.1]20
Qwen3-TTS 12Hzpodcast-800Log-mel L1(5, 6]480.0[37.1, 48.4]20
Qwen3-TTS 12Hzpodcast-800MR-STFT(2, 3]240.0[36.6, 50.0]20

Clip-to-clip ranges

Shading is the 10th–90th percentile of individual clip error increases, not a confidence interval. The vertical view is limited to −5% through 110%; the CSV retains all extrema.

Mean error ratios and averages of per-clip percentages are different statistics. The primary endpoint is the ratio of arithmetic means. The distribution plot uses individual clip ratios to show the spread; the raw aggregate file additionally includes minima and maxima.

Numerical ranges at each metric’s conservative 50% threshold

CodecMetricFramesMedian %P10 %P90 %Min %Max %
Descript DACLog1p-mel L11748.334.680.126.689.3
Descript DACLog-mel L11446.034.159.725.886.9
Descript DACMR-STFT742.028.765.016.198.7
Fish 400M MorePitchLoss1Log1p-mel L1639.528.461.712.280.9
Fish 400M MorePitchLoss1Log-mel L1439.628.350.221.572.3
Fish 400M MorePitchLoss1MR-STFT343.023.260.319.585.1
Fish DAC S2-ProLog1p-mel L17642.37.892.8-0.0233.2
Fish DAC S2-ProLog-mel L13138.926.465.618.0106.7
Fish DAC S2-ProMR-STFT3140.220.774.39.799.6
Qwen3-TTS 12HzLog1p-mel L11529.612.394.33.0197.0
Qwen3-TTS 12HzLog-mel L1640.430.669.518.5104.8
Qwen3-TTS 12HzMR-STFT439.325.962.918.197.9

These are individual clip percentages. They need not fall inside the confidence interval of the pooled ratio, and the maximum can exceed 50% even when the pooled conservative criterion passes.

Stricter observed per-clip caps at 50%

The primary confidence-bound rule limits the pooled average. The following alternatives constrain the measured individual clip distribution: its 90th percentile, or its maximum. Each rule also requires every larger measured window to pass. These are empirical sample caps, without a population confidence guarantee; eligible sample size changes with window length.

CodecMetricP90 cap: frames / msClipsMaximum cap: frames / msClips
Descript DACMR-STFT10 / 116.1 ms6014 / 162.5 ms60
Descript DACLog-mel L117 / 197.4 ms6024 / 278.6 ms60
Descript DACLog1p-mel L124 / 278.6 ms6041 / 476.0 ms60
Fish 400M MorePitchLoss1MR-STFT5 / 232.2 ms608 / 371.5 ms60
Fish 400M MorePitchLoss1Log-mel L15 / 232.2 ms608 / 371.5 ms60
Fish 400M MorePitchLoss1Log1p-mel L18 / 371.5 ms6015 / 696.6 ms60
Fish DAC S2-ProMR-STFT63 / 2925.7 ms58240 / 11145.6 ms15
Fish DAC S2-ProLog-mel L154 / 2507.8 ms5884 / 3901.0 ms49
Fish DAC S2-ProLog1p-mel L1108 / 5015.5 ms32252 / 11702.9 ms15
Qwen3-TTS 12HzMR-STFT9 / 720.0 ms6031 / 2480.0 ms58
Qwen3-TTS 12HzLog-mel L111 / 880.0 ms6023 / 1840.0 ms60
Qwen3-TTS 12HzLog1p-mel L142 / 3360.0 ms5766 / 5280.0 ms28

An actual spectrogram example

The following figures use the same Podcast 800 clip for all models: 04c87851-859c-4a3d-be14-b501dadb468b_seg00187_0001366305_0001375009. Each codec uses its joint conservative 50% window size. White marks in the difference panel indicate the hard joins. Display color scales match within each model and metric; difference color scales saturate at their 99th percentile. This is one example, so its individual error increase can exceed the pooled budget.

CodecFramesWindow msMR-STFT increase %Log-mel increase %Log1p-mel increase %
Descript DAC17197.418.534.939.9
Fish 400M MorePitchLoss16278.630.632.355.1
Fish DAC S2-Pro763529.413.319.424.3
Qwen3-TTS 12Hz151200.015.923.824.7

Descript DAC — mel spectrograms

Left: log-mel; right: log1p-mel. The last row isolates the decoder-window effect.

Fish 400M MorePitchLoss1 — mel spectrograms

Left: log-mel; right: log1p-mel. The last row isolates the decoder-window effect.

Fish DAC S2-Pro — mel spectrograms

Left: log-mel; right: log1p-mel. The last row isolates the decoder-window effect.

Qwen3-TTS 12Hz — mel spectrograms

Left: log-mel; right: log1p-mel. The last row isolates the decoder-window effect.

Coverage and interpretation

A clip contributes only if its sequence contains a join. Short clips are excluded at larger windows.

There is no universal monotonic threshold: changing window length also moves the joins onto different audio content, and decoder context is reset at each call. Reported pass/fail brackets are measured ranges, not interpolated guarantees. A passing pooled mean does not guarantee that every clip or dataset passes.

Confidence intervals are pointwise exploratory intervals for this fixed sample; they are not simultaneous coverage guarantees across all adaptive comparisons. The metrics describe spectral error, and no listener test or ASR metric was used to interpret these window thresholds.

What the S2-Pro architecture tells us

The pinned S2-Pro configuration includes an eight-layer causal transformer after code lookup, with a 128-frame attention window per layer, followed by upsampling and a waveform decoder with four transformer layers. At 21.53 code frames per second, 128 code frames represent about 5.94 seconds; stacking layers can extend the effective history beyond that per-layer window. An independent decoder call starts a new sequence, so it loses the preceding codes as attention context.

This is a plausible mechanism for the measured S2-Pro sensitivity to short windows. The experiment compares complete inference paths and does not isolate architecture from training. Its thresholds apply to decoding independent code windows using the documented hard-join policy.

Local source evidence: vendor/fish-speech/fish_speech/configs/modded_dac_vq.yaml, models/dac/rvq.py:DownsampleResidualVectorQuantize.decode, and models/dac/modded_dac.py:DAC.from_indices, under the pinned Fish source revision in the protocol.

Exact experiment and metric definitions

Each source is encoded once as a complete clip. The same integer code tensor is decoded in one call for the control and then in contiguous independent windows. Equal-length windows may be batched, while each batch row remains an independent sequence. The final short window is retained. Waveforms are concatenated at hard joins and trimmed to source duration; no overlap, crossfade, decoder state, delay correction, or loudness adjustment is used.

All metrics use sample-zero aligned 24 kHz mono audio. The mel transform uses 80 Slaney magnitude bands, Slaney area normalization, a 1024-sample periodic Hann window, 256-sample hop, 0–12 kHz range, centered constant-padded STFT, and epsilon 10⁻⁵.

Relative increase: 100 × (mean treatment error / mean matched baseline error − 1). Negative values are retained. Ten thousand paired clip resamples form the percentile 95% interval. The initial broad sweep is followed by integer refinement of pooled threshold brackets; ranges that remain unsampled are reported explicitly.

Reproducibility and artifacts

The measured output contains 37,200 clip/window conditions. Experiment execution took 42.7 minutes of cumulative measured session wall time, including the inherited three-codec experiment but excluding idle time between extensions, on NVIDIA L40S with Torch 2.12.0+cu130. CUDA timings here describe total experiment execution, including metric work; they are not decoder throughput. Source tensors and full-control metrics are checked for consistency when resuming; cached discrete codes are retained under codes/.

Reproduce the experiment and both report formats from the repository root. The experiment resumes recorded conditions and reuses cached codes.

.venv/bin/python scripts/benchmark_decoder_degradation.py
.venv/bin/python scripts/validate_decoder_degradation.py
.venv/bin/python scripts/decoder_degradation_spectrograms.py
.venv/bin/python scripts/report_decoder_degradation.py
node scripts/export_decoder_degradation_pdf.mjs

Files include selected_clips.csv, protocol.json, per_clip.jsonl, per_clip.csv, aggregate.csv, thresholds.csv, and report.json.

CodecCheckpoint revisionMeasured SHA-256
Descript DAC0.0.1a88eed82a7024ccc1facdb1e605c4c2f99281c8118c22c9895ffa846d8fb61aa
Fish 400M MorePitchLoss1step3802000-generator-20260915a97c5ebe1cbfcd0a341e268de5a31b42b1daac4d64741a25687ae271c3c94321
Fish DAC S2-Pro1de9996b6be38b745688de084d87a5633f714e4e74fc41c5a7151c6f350af8bd7e5d6e3accfcc7f3dfbfac23afd35af07052bb2f
Qwen3-TTS 12Hz7dd38ad4e9bad454aae9cd937d0cd577604fe229836b7b357f5ea43e889936a3709af68dfe3751881acefe4ecf0dbd30ba571258