Measured release listening

Scylla's Band
Voice Benchmark

Long-form intelligibility, voice, language, and affect results. WER is an ASR signal—not a substitute for listening—so each result presents the source script, generated audio, ASR transcript, and word-level diff together.

Release findings

Micro WER weights each spoken word equally. P90 shows the difficult tail. Listen alongside the displayed script and ASR comparison when interpreting these measurements.

Audio and transcripts