New September 6 audition using the requested WD disk for model storage. Chatterbox Nano and Turbo, each with Azelma, Eponine, Tara-reference, and Alba: ordinary speech and laughter. Quality and convincing laughter matter; generation need not be real time.
16 of 16 new clips are available locally. 16 ready. Six new Orpheus Q8 Tara clips below cover the same emotional and speech passages, outside that Tabi run count.
Tabi hardware: Intel Core Ultra 7 258V, 8 CPU threads; Arc 140V / Xe2; 32 GB shared RAM; Gentoo. Those are previously verified hardware facts. The run notes retain actual execution settings and disk paths when recorded.
New-run times measure full waveform generation, excluding loading and voice conditioning. RTF is generation time divided by audio duration. Streaming and network/playback latency are not measured for these new clips. Orpheus request and first-audio timings are labelled in its section. Clips are original WAVs without new level matching; listen to judge quality and the requested vocal event.
Azelma and Eponine lead, followed by the other saved references. Tara-reference uses synthetic Orpheus reference audio to condition Chatterbox. Model names are explicit; ordering is not a quality ranking.
Orpheus Q8 · Tara · emotion and laughter
Six new recordings using the same passages as the other demos: expression, laughter, chuckle and sigh, ordinary speech, conversation, and technical pronunciation. Orpheus uses native <laugh>, <chuckle>, and <sigh> cues; the remaining text is unchanged.
Original WAVs are used here, as for the Nano/Turbo recordings on this page, without level, pitch, or speed changes.
Expression · surprise and delight
12.03 s audio · 23.31 s full request
Passage and original WAV
Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.
Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?
The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.
Generated September 6 on eimu through the existing Orpheus-3b-FT-Q8_0 Tara service. Full request time includes queueing, model generation, SNAC decoding, server WAV writing and HTTP transfer. It is not isolated generation time or measured first-audio latency. The service's existing settings and GPU pacing were retained.
Original Orpheus-3b-FT speech from Tabi, with Q8 and Q4 recordings. These are saved benchmark clips from September 5 PDT / September 6 UTC, using an older passage without laughter tags. They are separate from the Chatterbox auditions.
The original WAVs have no level matching. Their request timings include streaming SNAC decoding and exclude initialization. The Q4 runs use an improved runtime; these recordings do not isolate quantization or provide a matched timing comparison with the other models.
Spoken passage and recording details
Hey Aku, the MacBook is handling my voice now. Your gaming GPU can take the night off.
This is the saved benchmark prompt; the recordings were generated on Tabi. Orpheus used Intel Vulkan for the model and CPU SNAC for audio decoding. The MacBook sentence describes the old fixture, not current voice routing.