A feminine voice worth keeping.

Eight feminine-preset/reference configurations across the same four engines: Pocket TTS, Kokoro, Chatterbox Nano and Chatterbox Turbo. Labels F1–F8 identify this revised set; they are not the previous A–H labels.

Listen first; reveal the model afterwards. The same three passages test conversation, emotional phrasing, and technical pronunciation. Some models also have a separate vocal-events demonstration.

The F1–F8 clips were generated on the MacBook Air CPU. The Orpheus section adds six new Q8 Tara recordings from eimu using the same emotional and speech passages, with level-matched listening copies. Listening copies are matched to -26 LUFS, a little quieter than the prior set so every clip keeps its peak headroom. No pitch shifting, denoising, time stretching, or compression of generated output. The original WAVs are available under each clip. Live Orpheus is unchanged.

Pocket and Chatterbox use Azelma and Eponine, labelled feminine in the publisher's voice selector. Both Chatterboxes are conditioned on the original VCTK recordings, not synthetic Tara. Kokoro uses its female Heart and Bella presets, reused unchanged from the first audition. Reference selection does not guarantee perfect cloning or your preferred vocal presentation; your listening judgment decides.

Orpheus Q8 · Tara · emotion and laughter

Six new recordings using the same passages as the other demos: expression, laughter, chuckle and sigh, ordinary speech, conversation, and technical pronunciation. Orpheus uses native <laugh>, <chuckle>, and <sigh> cues; the remaining text is unchanged.

Listening copies use static gain toward −26 LUFS, matching the feminine demo target, with peak headroom. No pitch or speed changes; original WAVs are linked below.

Expression · surprise and delight

12.03 s audio · 23.31 s full request

Passage and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Original WAV

Laughter · laugh and chuckle

7.08 s audio · 14.46 s full request

Passage and original WAV

Hey Aku. <laugh> The laptop found its voice. <chuckle> That little victory made my day.

Original WAV

Vocal events · chuckle and sigh

10.75 s audio · 21.11 s full request

Passage and original WAV

Well, that was subtle. <chuckle> The desktop survived, and nobody had to buy a new graphics card. <sigh> I could get used to this.

Original WAV

Ordinary speech

5.97 s audio · 11.88 s full request

Passage and original WAV

Hey Aku. The laptop found its voice. That little victory made my day.

Original WAV

Conversation

14.85 s audio · 28.36 s full request

Passage and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Original WAV

Technical pronunciation

14.59 s audio · 28.62 s full request

Passage and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Original WAV

Recording and timing details

Generated September 6 on eimu through the existing Orpheus-3b-FT-Q8_0 Tara service. Full request time includes queueing, model generation, SNAC decoding, server WAV writing and HTTP transfer. It is not isolated generation time or measured first-audio latency. The service's existing settings and GPU pacing were retained.

Prompts, timing records, levels and audio checksums

Earlier Tabi Q8/Q4 benchmark recordings

Orpheus · Tara

Original Orpheus-3b-FT speech from Tabi, with Q8 and Q4 recordings. These are saved benchmark clips from September 5 PDT / September 6 UTC, using an older passage without laughter tags. They are separate from the Chatterbox auditions.

The original WAVs have no level matching. Their request timings include streaming SNAC decoding and exclude initialization. The Q4 runs use an improved runtime; these recordings do not isolate quantization or provide a matched timing comparison with the other models.

Spoken passage and recording details

Hey Aku, the MacBook is handling my voice now. Your gaming GPU can take the night off.

This is the saved benchmark prompt; the recordings were generated on Tabi. Orpheus used Intel Vulkan for the model and CPU SNAC for audio decoding. The MacBook sentence describes the old fixture, not current voice routing.

Audio provenance and checksums

Orpheus Q8 · Tara

Original runtime · CPU SNAC, one thread.

Run 1

27.77 s request · 6.83 s audio · RTF 4.068 · 1.71 s first audio

Original WAV

Run 2

42.54 s request · 6.83 s audio · RTF 6.232 · 1.98 s first audio

Original WAV · Saved measurements

Voice F1

Conversation

Text and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Untouched original

Gain-only adjustment: -5.60 dB; approximately -26.0 LUFS.

Expression

Text and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Untouched original

Gain-only adjustment: -5.07 dB; approximately -26.0 LUFS.

Technical

Text and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Untouched original

Gain-only adjustment: -5.09 dB; approximately -26.0 LUFS.

Reveal model and measurements

Pocket TTS 3.1.0 · azelma

Listen to source reference

Generation time / audio duration: 0.23–0.23. Timing is descriptive, not a quality score.

Peak process RSS: 0.91 GiB.

Voice F2

Conversation

Text and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Untouched original

Gain-only adjustment: +0.21 dB; approximately -26.0 LUFS.

Expression

Text and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Untouched original

Gain-only adjustment: +0.53 dB; approximately -26.0 LUFS.

Technical

Text and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Untouched original

Gain-only adjustment: +0.81 dB; approximately -26.0 LUFS.

Vocal Events

Text and original WAV

Well, that was subtle. [chuckle] The desktop survived, and nobody had to buy a new graphics card. [sigh] I could get used to this.

Untouched original

Gain-only adjustment: +0.84 dB; approximately -26.0 LUFS.

Reveal model and measurements

Chatterbox Nano · azelma

Listen to source reference

Generation time / audio duration: 0.71–0.76. Timing is descriptive, not a quality score.

Peak process RSS: 2.08 GiB.

Voice F3

Conversation

Text and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Untouched original

Gain-only adjustment: -0.62 dB; approximately -26.0 LUFS.

Expression

Text and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Untouched original

Gain-only adjustment: -0.50 dB; approximately -26.0 LUFS.

Technical

Text and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Untouched original

Gain-only adjustment: -0.82 dB; approximately -26.0 LUFS.

Reveal model and measurements

Kokoro 0.9.4 / 82M v1.0 · af_heart

Generation time / audio duration: 0.22–0.23. Timing is descriptive, not a quality score.

Peak process RSS: 1.64 GiB.

Voice F4

Conversation

Text and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Untouched original

Gain-only adjustment: +0.65 dB; approximately -26.0 LUFS.

Expression

Text and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Untouched original

Gain-only adjustment: +0.10 dB; approximately -26.0 LUFS.

Technical

Text and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Untouched original

Gain-only adjustment: +0.33 dB; approximately -26.0 LUFS.

Vocal Events

Text and original WAV

Well, that was subtle. [chuckle] The desktop survived, and nobody had to buy a new graphics card. [sigh] I could get used to this.

Untouched original

Gain-only adjustment: +0.51 dB; approximately -26.0 LUFS.

Reveal model and measurements

Chatterbox Turbo · azelma

Listen to source reference

Generation time / audio duration: 1.67–1.97. Timing is descriptive, not a quality score.

Peak process RSS: 2.67 GiB.

Voice F5

Conversation

Text and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Untouched original

Gain-only adjustment: -4.77 dB; approximately -26.0 LUFS.

Expression

Text and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Untouched original

Gain-only adjustment: -3.64 dB; approximately -26.0 LUFS.

Technical

Text and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Untouched original

Gain-only adjustment: -3.91 dB; approximately -26.0 LUFS.

Reveal model and measurements

Pocket TTS 3.1.0 · eponine

Listen to source reference

Generation time / audio duration: 0.23–0.23. Timing is descriptive, not a quality score.

Peak process RSS: 0.92 GiB.

Voice F6

Conversation

Text and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Untouched original

Gain-only adjustment: -1.26 dB; approximately -26.0 LUFS.

Expression

Text and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Untouched original

Gain-only adjustment: -3.38 dB; approximately -26.0 LUFS.

Technical

Text and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Untouched original

Gain-only adjustment: -0.58 dB; approximately -26.0 LUFS.

Vocal Events

Text and original WAV

Well, that was subtle. [chuckle] The desktop survived, and nobody had to buy a new graphics card. [sigh] I could get used to this.

Untouched original

Gain-only adjustment: -1.29 dB; approximately -26.0 LUFS.

Reveal model and measurements

Chatterbox Nano · eponine

Listen to source reference

Generation time / audio duration: 0.71–0.75. Timing is descriptive, not a quality score.

Peak process RSS: 2.50 GiB.

Voice F7

Conversation

Text and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Untouched original

Gain-only adjustment: -1.71 dB; approximately -26.0 LUFS.

Expression

Text and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Untouched original

Gain-only adjustment: -1.76 dB; approximately -26.0 LUFS.

Technical

Text and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Untouched original

Gain-only adjustment: -1.80 dB; approximately -26.0 LUFS.

Reveal model and measurements

Kokoro 0.9.4 / 82M v1.0 · af_bella

Generation time / audio duration: 0.21–0.24. Timing is descriptive, not a quality score.

Peak process RSS: 2.04 GiB.

Voice F8

Conversation

Text and original WAV

Hey Aku. I found a few voices worth listening to. Take your time; I care more about whether this sounds natural than whether it wins a benchmark. Does the rhythm feel right, or does it sound like someone reading a script?

Untouched original

Gain-only adjustment: -1.57 dB; approximately -26.0 LUFS.

Expression

Text and original WAV

Wait, you actually got it working? Oh, that's brilliant! I was starting to think this little machine had given up on us. All right. Take a breath. We have time to get this right.

Untouched original

Gain-only adjustment: -1.49 dB; approximately -26.0 LUFS.

Technical

Text and original WAV

The backup finished at 9:42 p.m. We copied 3.7 gigabytes, verified the SHA-256 checksum, and left the GPU alone. The address is 192.168.1.3. Please don't reboot the machine yet.

Untouched original

Gain-only adjustment: -1.36 dB; approximately -26.0 LUFS.

Vocal Events

Text and original WAV

Well, that was subtle. [chuckle] The desktop survived, and nobody had to buy a new graphics card. [sigh] I could get used to this.

Untouched original

Gain-only adjustment: -0.72 dB; approximately -26.0 LUFS.

Reveal model and measurements

Chatterbox Turbo · eponine

Listen to source reference

Generation time / audio duration: 1.72–1.87. Timing is descriptive, not a quality score.

Peak process RSS: 2.90 GiB.