Orpheus Q8 Tara is included below with six new recordings using the same emotional and speech passages, including native laughter, chuckle, and sigh cues. Its listening copies use the same −26 LUFS target as Azelma/Eponine.
Every saved voice/reference for Chatterbox Nano and Chatterbox Turbo in this archive: Azelma, Eponine, Alba, and Tara-reference. Each has a conversation clip and a vocal-events prompt containing [chuckle] and [sigh]. Listen to judge the actual laughter, voice quality, and consistency; no quality ranking is implied.
The Azelma and Eponine references lead because publisher metadata identifies them as female. Tara-reference is conditioned on a short synthetic Orpheus Tara clip; it is not Orpheus running inside Chatterbox. These are all saved voices here, not an exhaustive list of voices either model can condition on.
The Chatterbox timings below were measured on orbital, MacBook Air M1, 8 GB, CPU-only. They are full waveform generation times, excluding model download/load and voice conditioning. They do not measure streaming or network latency. RTF = generation seconds / audio seconds; below 1 means faster than real time. Batch/offline listening is supported regardless of RTF.
Listening copies retain their original gain-only treatment: Azelma/Eponine approximately −26 LUFS; Alba/Tara-reference approximately −23 LUFS (some peak-constrained clips are quieter). These two audition sets are not level-matched against each other. Original audio is linked for every clip.
Orpheus Q8 · Tara · emotion and laughter
Six new recordings using the same passages as the other demos: expression, laughter, chuckle and sigh, ordinary speech, conversation, and technical pronunciation. Orpheus uses native <laugh>, <chuckle>, and <sigh> cues; the remaining text is unchanged.
Listening copies use static gain toward −26 LUFS, matching the Azelma/Eponine samples. Older Alba/Tara-reference Chatterbox clips use −23 LUFS. No pitch or speed changes; original WAVs are linked below.
Expression · surprise and delight
12.03 s audio · 23.31 s full request
Passage and original WAV
Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.
Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?
The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.
Generated September 6 on eimu through the existing Orpheus-3b-FT-Q8_0 Tara service. Full request time includes queueing, model generation, SNAC decoding, server WAV writing and HTTP transfer. It is not isolated generation time or measured first-audio latency. The service's existing settings and GPU pacing were retained.
Original Orpheus-3b-FT speech from Tabi, with Q8 and Q4 recordings. These are saved benchmark clips from September 5 PDT / September 6 UTC, using an older passage without laughter tags. They are separate from the Chatterbox auditions.
The original WAVs have no level matching. Their request timings include streaming SNAC decoding and exclude initialization. The Q4 runs use an improved runtime; these recordings do not isolate quantization or provide a matched timing comparison with the other models.
Spoken passage and recording details
Hey Aku, the MacBook is handling my voice now. Your gaming GPU can take the night off.
This is the saved benchmark prompt; the recordings were generated on Tabi. Orpheus used Intel Vulkan for the model and CPU SNAC for audio decoding. The MacBook sentence describes the old fixture, not current voice routing.
Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?
Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?
Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?
Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?
Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?
Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?
Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?
Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?