Write for a performer, not a settings panel
Tell the voice who is speaking, who is listening, what the passage is trying to do, how fast it should move, and where restraint, warmth, urgency, or pauses matter. Keep the instruction compatible with the Persian text. Correct short vowels, Ezafe, names, and ambiguous readings separately, because direction changes delivery but does not reliably repair a wrong word.
A strong voice direction answers six questions
You do not need all six in every prompt. The point is to replace vague praise words with decisions a listener could actually hear. “Make it amazing and professional” offers almost no direction. “Read as a calm museum guide, at a measured pace, with a short pause before the final date” does.
Four director notes you can adapt
Clear reading for a learner
Read as a patient Persian tutor addressing an adult learner. Keep a natural pace, but leave a small pause at sentence boundaries. Articulate complete words clearly without turning the passage into a word list. Keep the tone respectful and matter-of-fact.
Storytelling
Read as a warm storyteller speaking to one attentive listener. Begin quietly, let the tension rise through the middle of the scene, and slow slightly before the final sentence. Use purposeful pauses, not theatrical breaks after every phrase.
News or documentary narration
Use a composed, credible broadcast delivery. Keep the pace steady, names and figures distinct, and emotion restrained. Pause briefly between the main claim and the supporting detail. Do not sound promotional.
Reflective personal essay
Read in a close, thoughtful voice, as if recalling an event rather than performing it. Allow gentle hesitation around the memory, but keep the sentences connected. Let the last line settle instead of pushing it toward a dramatic conclusion.
Direct one contrast at a time
If a first attempt is too fast, change pace and keep the rest of the direction stable. If the pace is right but the voice sounds melodramatic, reduce emotional intensity. Changing the voice, script, punctuation, pace, and five adjectives at once makes it impossible to tell which edit helped.
Weak
“Make it natural, premium, emotional, polished, cinematic, and perfect.”
Stronger
“Use a warm, restrained documentary voice. Keep a steady pace and pause once before the quoted line.”
Persian reading decisions come before voice direction
Direction describes how a chosen reading should be performed. Persian text often leaves short vowels and Ezafe unwritten, and a homograph can support several readings. Those decisions require vocabulary, grammar, and sentence meaning. A note that says “sound more confident” does not tell the system whether گل is gol, flower, orgel, mud.
VowelMarks separates the two jobs while keeping them together. ThePersian speech workspace works out contextual vowel marks and Ezafe before playback. Choose a Premium Voice preset for clear reading, storytelling, or news, or open Custom Voice Studio to describe the performance yourself.
Punctuation should support the script, not fight it
Use periods, question marks, commas, quotation marks, and paragraph breaks where the Persian meaning calls for them. A clean paragraph gives the voice real structure. Repeating punctuation or inserting a comma after every few words may force pauses, but it can also break Ezafe chains, separate verbs from their complements, and make the result harder to understand.
A five-minute review loop
- Generate one representative paragraph.Do not begin with the entire article or story.
- Verify the Persian reading.Check names, short vowels, Ezafe, and ambiguous words before judging performance.
- Choose one delivery change.Pace, warmth, restraint, pause placement, or emphasis is enough for one comparison.
- Listen without reading the prompt.Judge what is actually audible, not what the instruction says should be audible.
- Save the winning voice recipe.Reuse the same voice and direction across the rest of the project, then make local exceptions only where needed.
What not to put in a voice instruction
- Private information that is not meant to become part of a provider request.
- A second version of the entire script when the text already exists in the text field.
- Conflicting roles such as “whispered, booming, intimate, and stadium loud.”
- Claims about an exact real person whose voice you do not have permission to imitate.
- Pronunciation guesses that have not been checked against the intended word or name.
Related guides
- Persian voiceover workflow from script to finished narration
- How to fix Persian speech pronunciation before changing style
- Why Farsi TTS chooses the wrong vowel or Ezafe
- Persian word stress and sentence prominence
Frequently asked questions
What should I write in an AI voice prompt for Persian?
Name the speaker's role, listener, purpose, emotional temperature, pace, and one or two important delivery choices. Keep the direction concrete and compatible with the script instead of stacking many vague adjectives.
How do I make an AI Persian voice sound like a storyteller?
Describe a calm storyteller speaking to a specific audience, use complete paragraphs with clean punctuation, request measured pacing and purposeful pauses, and identify where tension or warmth should rise. Test one paragraph before generating the full piece.
Can a director note fix wrong Persian pronunciation?
Not reliably. Direction shapes performance. Wrong short vowels, Ezafe, homographs, or names are reading problems and should be corrected with contextual pronunciation evidence first.
How much detail should a custom voice instruction contain?
Usually two to five precise sentences are enough for a short passage. Add detail only when it changes an audible decision such as pace, emotional restraint, pause placement, speaker role, or intended audience.
Why does the same AI voice prompt produce different results?
Generative speech can vary between attempts, and the script itself affects how a voice interprets direction. Voice identity, prompt wording, punctuation, passage length, and provider behavior can all change the result.
Sources and review notes
Sources are listed for the claims they support. Original practice examples are identified in the article. Product capabilities and access terms were checked on the review date and can change.
- Google AI for Developers, Text-to-Speech Generation. Used for current guidance on natural-language speech direction, speaker profiles, and the limits of prompt adherence.
- Google Cloud, Speech Synthesis Markup Language. Used for the established distinction between text content and explicit speech controls such as pauses, speaking rate, and emphasis.
- University of Texas at Austin, Persian Stress. Used for the distinction between word-level stress patterns and broader sentence emphasis.
- Mousavi et al., Grapheme-to-Phoneme Conversion in Persian. Used to separate Persian reading selection, including short vowels and Ezafe, from vocal performance direction.