Speech-variation samples on our actual benchmark instructions. Listen and tell us what sounds natural and usable.
Genuine per-accent English voices (American, British, Indian, Nigerian, Singaporean, and 9 more) via free Microsoft neural voices. These are the accents we would actually add to the benchmark.
What the instruct-driven TTS can control — speaking rate, gender, age, pitch, and emotion all work. It also shows 41 accent attempts: notice they all sound like one neutral voice. That is why real accents come from per-accent voices instead.
Straight from the dataset: v1 is the single concatenated clip the evaluation used, v2 is the same instruction as separate per-turn clips (a real conversation). Both files already ship in the release, so the difference is purely how the turns are delivered to the model.
The actual released noise (downloaded from the dataset) measures about 8 dB and masks the speech. The fix measures the active region and targets 15 dB, so noise stays in the background. Actual vs new per environment, same clip, with measured SNR.