AI Voice Production: 5 Best Practices

This article presents practical tips and important considerations for professionals responsible for planning AI-generated voice content, based on experience from actual development projects.
Creating an AI-generated narration demo can be surprisingly simple. Once you hear an AI voice reading a script in a reasonably natural way, you will probably think, “This could work.”
And in fact, it can.
However, once the project moves into full production, more delicate quality control and detailed direction are required to make the content suitable for real-world use.
For example, the production process may involve fine-tuning subtle elements such as making the overall impression calmer or optimizing the pace. It may also require generating, comparing, and selecting from several different versions.
These adjustments are an essential part of improving the quality of AI-generated voice content. Designing the process in advance makes it possible to manage production efficiently and reach results that stakeholders can agree on.
Based on knowledge gained from development projects, this article explains how to guide an AI voice content project toward a successful outcome.
AI Voice Content Uses AI Twice
When adding AI voice generation to a product, the process may appear to consist of only one step: passing text to a TTS, or text-to-speech, system.
In practice, however, there is another necessary step.
AI also needs to prepare the dialogue or narration that will be read aloud.
If the original written material is sent directly to a TTS system, the sentences may be too long and difficult to understand when heard because the source was written for reading rather than listening. The system may also mispronounce kanji and other words.
For example:
- Should 方 be pronounced kata or hō?
- How should a proper noun be pronounced?
- Where should a sentence be divided?
- Which words require emphasis?
These issues cannot always be prevented through TTS settings alone. A step should therefore be added beforehand to transform the original material into a script suitable for narration.
A text-generation AI, or LLM, handles this conversion.
The production pipeline therefore looks like this:
Original content
→ 1. Script-generation AI
→ 2. Text-to-Speech
→ Voice content

Understanding this structure makes it easier to identify where adjustments should be made.
Pronunciation errors are mainly related to Step 1, the script, while voice tone and pacing are mainly related to Step 2, the speech synthesis.
Instead of sharing only subjective feedback, the planning and development teams can discuss whether the issue comes from the script structure in Step 1 or the speech-generation settings in Step 2.
This separation can significantly reduce the cost of trial and error.
LLM-based TTS systems such as Gemini-TTS also allow users to provide natural-language instructions about delivery, such as:
- “Use a calm, explanatory tone.”
- “Whisper this section.”
- “Read this part at a slightly faster pace.”
In other words, the production direction written in ordinary language by the planning team can be used directly as material for controlling the voice generation.
As a general cost reference, standard Gemini-TTS pricing can be approximately $0.03 for one minute of generated audio, while script generation may cost approximately $0.002 per script. These are estimates as of July 2026. Always check the latest official pricing before implementation.
Important
Do not change Steps 1 and 2 at the same time.
If both the script and the voice-generation settings are changed simultaneously, it becomes difficult to determine which change produced the result.
As a rule, testing should be separated by stage.
Separating the stages also allows some tests to be conducted in parallel, which can shorten the overall production schedule.
Mindset: Treat Voice Quality as a Moving Target

Understanding one characteristic unique to audio production can make the process easier to manage.
Voice quality is subjective, and requirements tend to continue changing.
During the improvement phase, stakeholders may provide a wide range of opinions regarding:
- Tone
- Pace
- Pauses
- Energy
- Pronunciation
- Emotional expression
- Overall impression
The impression created by a voice is closely connected to the listener’s preferences and the situation in which the content is used.
For that reason, the goal should not be to identify one objectively correct answer. Instead, the process should aim to create the intended experience.
It is also natural and positive for new improvement ideas to emerge as stakeholders listen to more versions. This is part of the process of improving voice content.
A successful production plan should therefore not assume that the content will be completed in a single generation.
Instead, the standard process should be to structure feedback, compare multiple versions, test improvements, and refine the output step by step.
The following five practices make this possible.
Five Practices for Successful AI Voice Production

1. Share a List of Adjustable Parameters Before Development Begins
Before starting production, the planning and development teams should share a list of the parameters that can be changed.
These may include:
- Model
- Voice
- Speaking speed
- Degree of output variation
- Temperature or similar settings
- Prompt
- Audio format
- Post-processing settings
When the available controls are clearly documented, a subjective request such as “Make it sound a little calmer” can be translated into a concrete discussion:
“We will change this setting and generate another version for comparison.”
This list also becomes the foundation for managing differences between versions, as explained later.
2. Begin Voice Testing with Predefined Voices
There are two main approaches to preparing a voice.
The first is to select from voices already provided by the system, as with Gemini.
The second is to use Voice Clone technology offered by services such as ElevenLabs or Fish Audio to reproduce a particular person’s voice.
Starting with Voice Clone from the beginning can cause several types of testing to become mixed together:
- Creating and refining the cloned voice
- Improving the script
- Adjusting delivery
- Evaluating pronunciation
- Testing tone and pacing
This can cause the project scope to expand in too many directions.
During the initial PoC phase, it is better to start with predefined voices.
Once the team has confirmed that the requirements for the script and delivery can be met, testing can move to Voice Clone.
The recommended approach is to divide the process into clear stages and verify one element at a time.
3. Turn Subjective Feedback into Concrete Requirements
A qualitative impression such as:
“The tone feels slightly too high.”
should be converted into a more specific requirement, such as:
“Tone: calm and explanatory. The opening greeting may be slightly brighter.”
Once subjective impressions are expressed as concrete requirements, they become shared evaluation criteria for the entire team.
Recording each requirement together with the speaker and date also clarifies the history of the discussion and supports efficient agreement among stakeholders.
After every improvement, the team should listen again using the entire requirements list—not only the latest request.
This acts as a regression test for voice production. It helps confirm that a new change has not damaged a quality that was previously achieved.
It also makes it possible to detect deterioration in another part of the content at an early stage.
4. The Planning Team Must Prioritize Trade-Offs
Some requirements may conflict with each other.
For example:
- A calm, composed tone
- Fast-paced energy and momentum
Different viewpoints and ideas are valuable when the team is trying to create better content.
However, when requirements conflict, their priority must be clarified and recorded as part of the decision criteria.
This makes it possible to make fast decisions without losing the project’s central direction.
Only someone who understands the purpose and intended experience of the content can make these trade-off decisions.
For that reason, they should not be left entirely to the development team.
5. Include a Group Listening Session in the Production Plan
People perceive voices differently. Therefore, the final decision should not depend on a review by only one person.
Stakeholders should gather to listen to the generated audio in real time and discuss it together.
This helps eliminate differences in interpretation and provides a faster path toward improving quality.
The key is to listen while referring to the requirements list and then record the discussion results as updates to that list.
Important
Saving only each generated audio version is not enough.
For every version, the team should also record and share:
- What was changed
- The differences in settings
- The prompt used
- The script version
- The model and voice used
- The remaining issues to be considered
This version management makes it possible to return to or reproduce a previously successful result at any time.
From a technical perspective, managing setting differences in a format such as JSON can be effective and can be implemented relatively easily in collaboration with the development team.
Post-Processing Is Another Important Option
Another concept that can speed up discussions is post-processing.
Some requirements cannot always be handled completely through AI voice-generation settings alone.
Examples include:
- Fine adjustments to playback speed
- Preventing audio clipping
- Normalizing volume levels
- Inserting silence between chapters
- Cutting unnecessary pauses
- Combining multiple generated segments
- Applying fades
These issues can be addressed through a post-processing stage in which the generated audio is modified programmatically.
Keeping both options in mind—
- adjusting the AI settings, or
- handling the issue through a post-processing program
—expands the range of ways the team can respond to production requests.
Production Converges When the Team Aims for Agreement

If the team seeks a state of permanent and universal satisfaction with AI-generated voice content, the production process may never end.
The practical goal should instead be agreement:
“We have satisfied this requirements list, in this order of priority, to this defined level.”
Record the settings, express the requirements in words, listen together, and make a decision.
By incorporating this process from the beginning, the team can continuously improve the quality without becoming lost in unstructured trial and error every time new feedback is received.
The five preparations described in this article are:
- Creating a list of adjustable parameters
- Starting voice testing with predefined voices
- Consolidating feedback into a requirements list
- Agreeing on priorities
- Holding a group listening session
None of these require a specialized tool. They are practical preparations that can be introduced immediately.
A useful first step is to add these five items to the agenda for the project kickoff meeting.
This article is based on our development experience.
The performance and specifications of AI speech-synthesis technologies continue to change, and the most appropriate production process may also evolve.
For a more detailed technical explanation, please refer to our upcoming white paper, “How to Stabilize Quality in Voice Content Production Using Generative AI.”
We also provide consultation and development support for AI-generated voice content.

Struggling to turn ideas into reality? With a proven track record of over 1,000 clients, our agile and flexible team will accelerate your business growth.
Book a Free ConsultationMore on "Generative AI & ML"

Turning Google Colab into an API Server to Run Speech-to-Speech Voice AI
This PoC turns Google Colab into a temporary WebSocket server for speech-to-speech AI. It combines faster-whisper large-v3, Gemini 2.5 Flash-Lite, VOICEVOX and Silero VAD, then improves latency through sentence buffering, streaming responses and parallel TTS generation.

What We Learned Building Voice AI with Gemini: The Major Difference Between a PoC and a Commercial Service
A Gemini Live API PoC can deliver real-time voice conversations quickly, but production introduces concurrency, quotas, observability, model lifecycle and device-specific audio issues. This article explains why the architecture moved from direct browser access to LiveKit and Vertex AI.

A Concept for Developing AI Through Artificial Languages
Can AI learn logic more efficiently through an artificial language than through natural language? This article compares Esperanto, Lojban, and Ithkuil, then presents a custom GPT-2 model trained from scratch on Lojban-based data that achieved 100% accuracy on prepared three-valued logic tests.