Take ONE of your real tasks (work or household) where your voice could help. Describe it and select the appropriate voice AI skill. Return JSON {"task":"...","capability":"...","why":"..."}, where capability is one of: conversation (live conversation), dictation (dictation), transcription (transcription), synthesis (speech synthesis).
Voice AI can do four different things, and confusing them is the first mistake of a newbie. The first is a live conversation: you speak, the assistant hears and answers in a voice almost without a pause, you can interrupt and clarify, as with a person on the phone. The second is dictation: your speech turns into text, you speak instead of typing. Third, transcription: a long recording (meeting, lecture) is decomposed into text, divided by speaker. Fourth - speech synthesis: the text is voiced in a natural voice with the desired intonation. The mechanics are simple: everywhere in the middle there is a model that translates sound into meaning and back. Pro trick: don’t ask yourself “how to use voice AI,” but first name the task and then choose a skill for it. Dictate a letter - dictation; understand a two-hour call - transcription plus summary. Insight: the voice wins where the hands are busy or the thought is longer than it is convenient to type. A common mistake is to try to dictate perfectly the first time; the voice is as strong as a draft, which is then cleaned by AI.