Prepare the text for voiceover and describe the presentation. Return JSON {"script":"...","voice_style":"...","emphasis_notes":"..."}: short text “for the ear” (up to 80 words), voice style (tone, tempo, emotion) and notes on pauses/emphasis and complex words.
Speech synthesis reads the text with your voice, and it sounds exactly as you wrote it - the model reads it literally. This means that you need to write “for the ear” and not for the eye. Mechanics: the engine takes text, punctuation marks and style marks and uses them to build intonation, pauses and tempo. Levers. The first is that punctuation controls the rhythm: a period gives a pause, an ellipsis gives thoughtfulness, short sentences sound confident, long ones sound heavy. Second, set not only the text, but also the presentation: voice (calm announcer, friendly, energetic), emotion and tempo. “Read this” and “read it warmly, slowly, as if for a friend” give different results. Pro trick: write complex words, names and abbreviations the way they sound, or break them up - otherwise the engine will read it in its own way. The second technique: place pauses intentionally (with commas, periods, breaking into lines) where the meaning requires a breath - the naturalness of speech is based on pauses, and not on the voice. A typical mistake is to submit for dubbing a “text to be read with the eyes” with long periods and footnotes; it sounds like it's falling apart. Write briefly, speak the draft out loud in advance.
Unlock access to submit solutions for instant AI review.