Describe one short video. Return JSON {"idea":"...","prompt":"...","duration":"..."}: idea - idea in one phrase, prompt - description of the scene using the formula “subject + one action + environment + mood", duration - duration in seconds (number + “sec”).
AI video works like a very fast animator: the neural network has seen millions of videos and learned to predict how pixels change from frame to frame. You give the text - she completes the movement. It is important to understand: this is not gluing together ready-made clips, but generating each frame anew, so the clearer the description, the more stable the picture. The working formula of the first prompt: subject + one action + environment + mood. Not “a beautiful city”, but “an elderly fisherman mending a net on a wooden pier, morning fog.” Insider: a short video of 3-5 seconds almost always comes out cleaner than a long one - the model manages to maintain the logic of movement, hands and faces do not “float”. Another trick: one action per video. Two teams in a row (“walks, then sits down”) confuse the model, and it mixes them into mush. A typical beginner mistake is to dump ten details at once: the weather, a bunch of items, three characters. The model won't hold everything. Start with one clear scene and one character, then complicate it.