Voice replies
Voice is optional and can be enabled per verb. Start with text only, then add voice once the personality feels right. If you want voice on every eligible reply, set voice response frequency to100. Lower values make voice replies intermittent on purpose.
Voice output is a sidecar to the normal AI response. Verba generates the
in-character text reply first, then may send the same response as audio when
Voice Engine is enabled and the configured frequency selects it. A clear
natural-language request for a voice message can also trigger audio, but Voice
Engine must still be enabled and the Verb must have a usable voice.
Slack and Twitch do not install a standalone voice command. Voice is not
handled as an isolated command response on those integrations.
Delivery by surface
Calling is unavailable on providers that do not expose native bot calling.
Voice attachment visual
The Voice Engine settings also include a Voice Attachment Visual section. You can upload a custom still image for voice message videos instead of using the default Verba image.- The live preview uses the same
20:7crop as the final voice attachment - The generated video output is framed to
400x140 - You can replace or remove the image at any time
- If no custom image is uploaded, Verba falls back to the default voice banner
Voice cloning
Upload a short, clean sample to create a custom voice. The best results come from 6 to 12 seconds of clear speech with minimal background noise. Reference text is optional. If you can, paste the exact transcript of the sample. It improves similarity and keeps the voice more stable across replies.Supported languages
The voice engine supports a focused language set:- Auto (recommended default)
- English
- Chinese
- Japanese
- Korean
- German
- French
- Russian
- Portuguese
- Spanish
- Italian
Premium voice model access
Voice model availability is plan-based in the same way as the AI and Image engines. When a selected premium voice model is outside your current tier, Verba shows an upgrade prompt with the number of additional premium voice models available on a higher plan. That count is dynamic and can change as the voice catalog changes.Discord voice chat
On Discord, Ultra verbs can join voice chat with/vc-join, and they can also
join from a normal mention request such as asking the bot to join VC/call in
server chat.
Lower-tier verbs can still keep those commands enabled, but they respond with
an in-character upgrade message instead of joining live VC. Normal generated
voice messages stay free.
When the live voice path is healthy, the bot can:
- Listen in the connected voice channel
- Transcribe incoming speech
- Generate a reply
- Speak the reply back into VC
- Voice Engine is enabled for the verb
- The selected voice model is available to that plan
- The bot can access and speak in the target voice channel
- The active speech provider is healthy and has available quota
If the selected live speech provider is unavailable or out of quota, Discord
VC can join successfully but still fail to transcribe or speak until that
provider becomes available again.
Voice notes on messaging platforms
WhatsApp voice notes can be transcribed and passed to the selected verb. When voice replies are enabled, supported WhatsApp conversations can receive a voice response as well. This is different from Discord live voice-channel calling, which has its own Ultra plan requirement. Telegram can receive voice/audio media through its bot integration. Slack receives an uploaded audio file, Twitch receives a playable link because chat has no general audio upload API, and Email receives an attachment. These are message attachments, not calls.Discord is currently the only integration with a native live-call path.
WhatsApp calling is intentionally not exposed, and Slack, Twitch, Telegram
bots, and Email have no Verba call action.
Keep it natural
- Short replies sound better.
- Avoid long paragraphs in voice mode.
- Set a frequency that feels human.
Safety and permissions
Only upload audio you own or have permission to use.AI engine
Lower temperature for clearer voice output.

