A voice worth keeping.

Listen first; reveal the model afterwards. The same three passages test conversation, emotional phrasing, and technical pronunciation. Some models also have a separate vocal-events demonstration.

The A–H clips were generated locally on the MacBook Air CPU. The Orpheus section adds six new Q8 Tara recordings from eimu using the same emotional and speech passages, with level-matched listening copies. Static level matching keeps louder samples from winning by volume alone. No denoising, time stretching, or compression. The original WAVs are available under each clip.

These are candidate configurations, not an exhaustive ranking. Reference choice affects cloning quality; the Tara reference is an existing 6.144-second synthetic Orpheus clip. Your listening judgment decides the winner.

Orpheus Q8 · Tara · emotion and laughter

Six new recordings using the same passages as the other demos: expression, laughter, chuckle and sigh, ordinary speech, conversation, and technical pronunciation. Orpheus uses native <laugh>, <chuckle>, and <sigh> cues; the remaining text is unchanged.

Listening copies use static gain toward −23 LUFS, matching the original demo target, with peak headroom. No pitch or speed changes; original WAVs are linked below.

Expression · surprise and delight

12.03 s audio · 23.31 s full request

Passage and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Original WAV

Laughter · laugh and chuckle

7.08 s audio · 14.46 s full request

Passage and original WAV

Hey Aku. <laugh> The laptop found its voice. <chuckle> That little victory made my day.

Original WAV

Vocal events · chuckle and sigh

10.75 s audio · 21.11 s full request

Passage and original WAV

Well, that was subtle. <chuckle> The desktop survived, and nobody had to buy a new graphics card. <sigh> I could get used to this.

Original WAV

Ordinary speech

5.97 s audio · 11.88 s full request

Passage and original WAV

Hey Aku. The laptop found its voice. That little victory made my day.

Original WAV

Conversation

14.85 s audio · 28.36 s full request

Passage and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Original WAV

Technical pronunciation

14.59 s audio · 28.62 s full request

Passage and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Original WAV

Recording and timing details

Generated September 6 on eimu through the existing Orpheus-3b-FT-Q8_0 Tara service. Full request time includes queueing, model generation, SNAC decoding, server WAV writing and HTTP transfer. It is not isolated generation time or measured first-audio latency. The service's existing settings and GPU pacing were retained.

Prompts, timing records, levels and audio checksums

Earlier Tabi Q8/Q4 benchmark recordings

Orpheus · Tara

Original Orpheus-3b-FT speech from Tabi, with Q8 and Q4 recordings. These are saved benchmark clips from September 5 PDT / September 6 UTC, using an older passage without laughter tags. They are separate from the Chatterbox auditions.

The original WAVs have no level matching. Their request timings include streaming SNAC decoding and exclude initialization. The Q4 runs use an improved runtime; these recordings do not isolate quantization or provide a matched timing comparison with the other models.

Spoken passage and recording details

Hey Aku, the MacBook is handling my voice now. Your gaming GPU can take the night off.

This is the saved benchmark prompt; the recordings were generated on Tabi. Orpheus used Intel Vulkan for the model and CPU SNAC for audio decoding. The MacBook sentence describes the old fixture, not current voice routing.

Audio provenance and checksums

Orpheus Q8 · Tara

Original runtime · CPU SNAC, one thread.

Run 1

27.77 s request · 6.83 s audio · RTF 4.068 · 1.71 s first audio

Original WAV

Run 2

42.54 s request · 6.83 s audio · RTF 6.232 · 1.98 s first audio

Original WAV · Saved measurements

Voice A

Conversation

Text and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Untouched original

Gain-only adjustment: -4.83 dB; approximately -23.0 LUFS.

Expression

Text and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Untouched original

Gain-only adjustment: -3.09 dB; approximately -23.0 LUFS.

Technical

Text and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Untouched original

Gain-only adjustment: -3.83 dB; approximately -23.0 LUFS.

Reveal model and measurements

Pocket TTS 3.1.0 · alba

Generation time / audio duration: 0.23–0.24. Timing is descriptive, not a quality score.

Peak process RSS: 0.94 GiB.

Voice B

Conversation

Text and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Untouched original

Gain-only adjustment: +2.47 dB; approximately -23.0 LUFS.

Expression

Text and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Untouched original

Gain-only adjustment: +2.54 dB; approximately -23.0 LUFS.

Technical

Text and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Untouched original

Gain-only adjustment: +3.38 dB; approximately -23.0 LUFS.

Vocal Events

Text and original WAV

Well, that was subtle. [chuckle] The desktop survived, and nobody had to buy a new graphics card. [sigh] I could get used to this.

Untouched original

Gain-only adjustment: +3.70 dB; approximately -23.0 LUFS.

Reveal model and measurements

Chatterbox Nano · alba

Generation time / audio duration: 0.73–0.81. Timing is descriptive, not a quality score.

Peak process RSS: 2.17 GiB.

Voice C

Conversation

Text and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Untouched original

Gain-only adjustment: +2.38 dB; approximately -23.0 LUFS.

Expression

Text and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Untouched original

Gain-only adjustment: +2.50 dB; approximately -23.0 LUFS.

Technical

Text and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Untouched original

Gain-only adjustment: +2.18 dB; approximately -23.0 LUFS.

Reveal model and measurements

Kokoro 0.9.4 / 82M v1.0 · af_heart

Generation time / audio duration: 0.22–0.23. Timing is descriptive, not a quality score.

Peak process RSS: 1.64 GiB.

Voice D

Conversation

Text and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Untouched original

Gain-only adjustment: +2.30 dB; approximately -23.0 LUFS.

Expression

Text and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Untouched original

Gain-only adjustment: +3.75 dB; approximately -23.0 LUFS.

Technical

Text and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Untouched original

Gain-only adjustment: +2.50 dB; approximately -23.0 LUFS.

Vocal Events

Text and original WAV

Well, that was subtle. [chuckle] The desktop survived, and nobody had to buy a new graphics card. [sigh] I could get used to this.

Untouched original

Gain-only adjustment: +2.80 dB; approximately -23.0 LUFS.

Reveal model and measurements

Chatterbox Turbo · alba

Generation time / audio duration: 1.64–1.75. Timing is descriptive, not a quality score.

Peak process RSS: 2.66 GiB.

Voice E

Conversation

Text and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Untouched original

Gain-only adjustment: -2.60 dB; approximately -23.0 LUFS.

Expression

Text and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Untouched original

Gain-only adjustment: -2.07 dB; approximately -23.0 LUFS.

Technical

Text and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Untouched original

Gain-only adjustment: -2.09 dB; approximately -23.0 LUFS.

Reveal model and measurements

Pocket TTS 3.1.0 · azelma

Generation time / audio duration: 0.23–0.24. Timing is descriptive, not a quality score.

Peak process RSS: 0.95 GiB.

Voice F

Conversation

Text and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Untouched original

Gain-only adjustment: +2.71 dB; approximately -23.0 LUFS.

Expression

Text and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Untouched original

Gain-only adjustment: +2.01 dB; approximately -23.0 LUFS.

Technical

Text and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Untouched original

Gain-only adjustment: +3.95 dB; approximately -23.3 LUFS.

Vocal Events

Text and original WAV

Well, that was subtle. [chuckle] The desktop survived, and nobody had to buy a new graphics card. [sigh] I could get used to this.

Untouched original

Gain-only adjustment: +3.35 dB; approximately -23.0 LUFS.

Reveal model and measurements

Chatterbox Nano · tara-reference

Generation time / audio duration: 0.71–0.74. Timing is descriptive, not a quality score.

Peak process RSS: 2.33 GiB.

Voice G

Conversation

Text and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Untouched original

Gain-only adjustment: +1.29 dB; approximately -23.0 LUFS.

Expression

Text and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Untouched original

Gain-only adjustment: +1.24 dB; approximately -23.0 LUFS.

Technical

Text and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Untouched original

Gain-only adjustment: +1.20 dB; approximately -23.0 LUFS.

Reveal model and measurements

Kokoro 0.9.4 / 82M v1.0 · af_bella

Generation time / audio duration: 0.21–0.24. Timing is descriptive, not a quality score.

Peak process RSS: 2.04 GiB.

Voice H

Conversation

Text and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Untouched original

Gain-only adjustment: +2.92 dB; approximately -23.0 LUFS.

Expression

Text and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Untouched original

Gain-only adjustment: +3.41 dB; approximately -23.0 LUFS.

Technical

Text and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Untouched original

Gain-only adjustment: +4.13 dB; approximately -23.0 LUFS.

Vocal Events

Text and original WAV

Well, that was subtle. [chuckle] The desktop survived, and nobody had to buy a new graphics card. [sigh] I could get used to this.

Untouched original

Gain-only adjustment: +4.95 dB; approximately -23.0 LUFS.

Reveal model and measurements

Chatterbox Turbo · tara-reference

Generation time / audio duration: 1.51–1.68. Timing is descriptive, not a quality score.

Peak process RSS: 3.15 GiB.