Give your AI a voice: pick one of 96, or clone your own
Voices show up in three places in Foreverse: audiobook listening, roleplay read-aloud, and a companion's voice notes and calls. This walkthrough covers the catalog (49 Chinese voices on an Alibaba key including dialect groups, 26 on xAI, 13 on OpenAI, 8 on CosyVoice), how inline preview bills (a real synthesis call, cached so replays are free), and the five-step voice clone: record or upload, name it, bind it to a companion. Cloning itself is free — and strictly your own voice only.

A question we get often enough to answer in writing: “why does the voice list keep changing size?” Nothing mysterious — the catalog follows the channel. Synthesis voices are provider assets, so whichever key you bring decides which shelf you browse. An Alibaba DashScope key puts 49 Chinese voices on it, xAI puts 26, OpenAI 13, SiliconFlow's CosyVoice 8. Bring no key and the official channel offers its built-in voice models: perfectly usable, smaller shelf. Once that clicks, choosing a voice stops being folklore and becomes shopping.
The shelf feeds three places: whole-book audiobook listening, read-aloud in roleplay chat, and a companion's voice notes and calls. Each has its own picker; the pickers look identical and remember independently, so the audiobook can run a deep narrator while the companion speaks in the voice you cloned, and neither setting tramples the other.
Browsing the shelf
Take the largest shelf as the example. DashScope's 49 voices come grouped — standard first, then dialects, then character voices. The dialect group is the fun one if your story has regional color: a Shanghainese voice, a Cantonese one. Each row shows a name, a gender badge, and a one-line temperament description copied verbatim from the provider's official catalog; we didn't editorialize. There's a search field that matches names, IDs, descriptions, and gender, and a “switch model” chip that stacks a model picker on top — choose a different channel and the catalog hot-swaps in place without closing the sheet.
Practical browsing advice, from doing this too many times ourselves: filter by gender first, preview the two or three whose descriptions read right, and stop. Auditioning all 49 in order is a way to spend an evening.
What a preview costs
The inline preview button is live ammunition: one tap fires one real TTS synthesis call, with no confirm dialog in the way. We allow ourselves that because the bill is engineered small. The sample line caps at 40 characters, so a single tap costs a fraction of a cent on most providers, and the audio caches on-device — same voice, same line, second tap plays instantly with zero calls. Our measurements: first preview speaks in about two seconds, replays in under one, and the request log shows no second entry. When a preview fails, the row says why in words — invalid key, rate limit, network — rather than offering a bare retry button and a mystery.
Cloning: record a line, or upload one
However large the shelf, one voice is never on it. The clone studio lives in Settings → AI models & services → My Voices, with a shortcut at the bottom of any companion's voice picker (“+ Clone a voice”).
Five steps, and the first two are rules. Consent first: you confirm the voice is your own and that you know it can be deleted at any time. Channel second: xAI Grok or SiliconFlow CosyVoice, listed according to the keys you've configured — the xAI lane needs Custom Voices enabled on your team account, and if it isn't, the error message says exactly that instead of failing vaguely; CosyVoice has no such gate. Then the sample: record on the spot, where the first line of the guided script doubles as the spoken consent statement, or upload an existing file — the CosyVoice lane asks you to type the file's exact transcript alongside it. Name it, save. The cloning step itself is free, on your own key.
The save moment is the good part: a “bind to companion” prompt appears immediately, and one tap on a companion's avatar wires the voice in — including switching the speech model to the clone's home channel, so you don't finish the flow in Settings hunting for a dropdown. From then on the clone sits pinned in a “My Voices” group across all three pickers. On the companion's picker it stays visible under any model; select it while a different provider's model is active and the app switches models for you, with a note saying it did.
Deletion is symmetrical. Each clone's row offers preview, rename, delete — and delete cascades to the provider's side, not just the local list. A companion currently speaking in that voice falls back to the default. No half-deleted voice lingers on someone's server.
Three places, three ways to choose
The listening page (tap the center of the page while reading to raise the control bar, then “Listen”) exposes four entries along the bottom: engine, voice, cache, original text. Engine is the fork that matters. Android's system TTS synthesizes locally — offline, free, fine for burning through a webnovel on a commute. Online voice models sound far better and bill per synthesis; segment caching means a chapter you've already listened to never bills twice. Both lanes list their voices under the same “voice” entry, clones included.
Roleplay read-aloud is the “Read aloud” action in the chat's “+” panel, and it reads the latest reply; the “Auto narration” toggle in the ⋮ menu speaks each new reply as it lands. Voice choice here persists per model, so different cards can keep different voices without renegotiating every session. For narration-heavy cards, the “Read dialogue only” toggle in the same menu section skips the stage directions — cheaper, and easier on the ears.
The companion's voice is set on their profile page. Voice notes cache after synthesis, so replaying one doesn't bill again. And one honest little behavior worth knowing: if you switch the speech model and the old voice doesn't exist in the new model's catalog, the setting clears back to default on its own — rather than keeping a stale value that would fail on every synthesis until you figured out why.
Edges, stated plainly
The catalog numbers are what the providers published as of July 2026; shelves change when they do. Cloning is restricted to your own voice, full stop — giving a companion some specific other person's voice is not a supported path, and the consent script exists to make that boundary audible. System TTS is free but fixed; online voices are better but every synthesis is a real call, so long books belong on segment caching. And every synthesis, like every text call, lands in the app's API request log where you can audit it line by line.
FAQ
Is voice cloning free? Is it safe?
The cloning step is free and runs on your own key; synthesizing speech with the cloned voice afterwards bills at that provider's TTS rates. Three safety rules: you may only clone your own voice, and the first line of the recording script is a consent statement you read aloud; a cloned voice can be deleted at any time; deletion cascades to the provider's side too, and any companion using that voice falls back to the default instead of keeping an orphaned clone.
Why do I only see a handful of voices when others have dozens?
The catalog follows your providers. Voices are provider assets: an Alibaba DashScope key surfaces 49 Chinese voices including dialect and character groups, xAI lists 26, OpenAI 13, SiliconFlow's CosyVoice 8. With no key at all, the official channel offers its built-in voice models — usable, just a smaller shelf. The picker has a 'switch model' chip at the top; change the channel and the catalog hot-swaps in place.
Does previewing a voice cost money?
Yes, and deliberately little. Tapping the inline preview runs one real TTS synthesis on a sample line capped at 40 characters, so a single tap costs a fraction of a cent on most providers. The result is cached on-device: the same voice with the same line plays instantly on the second tap with zero calls. In our tests, first preview spoke in about two seconds, replays in under one with no new request logged.
Can I use a cloned voice for audiobook listening?
Yes. Once saved, the clone appears pinned in a 'My Voices' group across all three pickers — audiobook listening, roleplay read-aloud, and the companion's voice setting. The listening page also has a free lane: choose the Android system TTS engine for local, offline, zero-cost narration, with the trade-off that its voices are fixed and can't be cloned.
Questions or ideas? Join our Discord →