Actual is the noise currently in the released dataset (downloaded from Hugging Face) for this exact clip. New is the fixed pipeline, which measures the active region and targets 15 dB so the noise stays in the background and the constraints stay audible. Measured SNR (speech vs the noise floor in silent gaps) is shown per row: the released noise sits around 8 dB and masks the speech, the fix moves it to about 15 dB.
| Environment | Actual (released, ~8 dB) | New (fixed, 15.0 dB) |
|---|---|---|
| coffee shop | measured 8.4 dB | measured 16.1 dB |
| convention hall | measured 8.4 dB | measured 15.4 dB |
| call center | measured 8.6 dB | measured 15.6 dB |
| summer outdoor | measured 8.0 dB | measured 14.9 dB |
| mountain outdoor | measured 7.8 dB | measured 15.3 dB |
| static noise | measured 7.6 dB | measured 15.1 dB |
Interruption is not a level fix, so it is not in the table above. The released version overlays a competing speaker on top of the user, who never stops talking, so it just sounds like babble. The new design is a real interruption as a side sequence: the user is interrupted, answers, then returns to the task, with all constraints preserved across the break.
5 old vs 5 new interruptions →
Redesigned bank: 5 interruption types (speech / sound / self) →