OmniAgentBench — accent & voice demo 🗣️

Speech-variation samples on our actual benchmark instructions. Listen and tell us what sounds natural and usable.

Real accents 14 accents · free

Genuine per-accent English voices (American, British, Indian, Nigerian, Singaporean, and 9 more) via free Microsoft neural voices. These are the accents we would actually add to the benchmark.

Qwen3-TTS voice controls rate · gender · age · pitch · emotion

What the instruct-driven TTS can control — speaking rate, gender, age, pitch, and emotion all work. It also shows 41 accent attempts: notice they all sound like one neutral voice. That is why real accents come from per-accent voices instead.

Multi-turn: v1 vs v2 actual released audio

Straight from the dataset: v1 is the single concatenated clip the evaluation used, v2 is the same instruction as separate per-turn clips (a real conversation). Both files already ship in the release, so the difference is purely how the turns are delivered to the model.

Noise: actual vs new fix comparison

The actual released noise (downloaded from the dataset) measures about 8 dB and masks the speech. The fix measures the active region and targets 15 dB, so noise stays in the background. Actual vs new per environment, same clip, with measured SNR.

Feedback wanted: which accents sound natural and clear enough to include as a speech-variation dimension? Any that sound wrong or unintelligible? And for multi-turn, does v1 or v2 feel like the fairer test?