Clear, practical technology insights BSOD Code Lookup · Windows Error Code Lookup · Wi-Fi Troubleshooting · PC Troubleshooting Checklist

Grok Custom Voices: What xAI’s Voice Cloning Feature Supports

xAI has launched Custom Voices for its speech APIs. Understand voice cloning, account restrictions, and why API support does not confirm public voice sharing in the Grok app.

Table of Contents

xAI announced Custom Voices on April 30, 2026. The feature creates a synthetic voice from a recording for use with Grok’s Text to Speech and Voice Agent APIs.

This is voice cloning, not transcription. Transcription turns spoken words into text; voice cloning produces a voice that can speak new text.

Grok will have a voice copying feature. Picture 1

What is confirmed

The launch announcement describes recording your own voice in the xAI console, completing a verification process, and using the resulting voice in supported speech services. It presents narration and branded conversational agents as potential uses.

The current Custom Voices documentation explains that a generated voice_id can be used with the speech APIs. It lists a maximum reference clip length of 120 seconds and recommends a quiet recording with a single speaker.

That documentation currently lists availability in the United States except Illinois, and says creation through the API is restricted to Enterprise teams. Check the console and documentation for your account’s current options.

The earlier description of a consumer-app voice library with public links should not be treated as an established feature. The current API documentation says custom voices are scoped to a team and are not available to other users.

API support therefore does not confirm that anyone can save another person’s voice through a Grok app link. Look for an explicit product announcement before relying on that workflow.

Grok will have a voice copying feature. Picture 2

Prepare and review a voice sample

If the feature is available to you, use your own voice and follow the service’s verification instructions. Record without background music, speak in the style you intend to generate, and test a short script before preparing longer narration.

Listen for pronunciation, pacing, and unintended sounds. Label synthetic recordings where listeners might mistake them for an actual recorded statement, and do not use a voice model to impersonate someone.

For other narration workflows, see text-to-speech with stock voices in Speechma and the OmniVoice voice creation guide. These are different products with their own controls and terms.

Discussion

Reader Comments 0

Sign in with email or Google to join the discussion.