Take the scene apart. Return JSON {"subject":"...","action":"...","setting":"...","shot":"...","style":"..."}: subject - who/what, action - ONE action, setting - where and when, shot - plan and angle (close/medium/general, camera level), style—visual style.
A video prompt has a skeleton, like a movie still. Break it down into parts: subject (who/what), action (what he is doing), environment (where and when), shot (close-up, medium, general), perspective (eye level, above, below), light and mood (warm sunset, cold morning), style (realism, watercolor, 3D). The neural network reads the description like a director’s request: the more accurately each part is named, the less it thinks out for you. The key technique is ONE action per scene. “A child launches a paper airplane” will hold on, but “runs, jumps, catches, laughs” will fall apart into artifacts. The second technique: first describe a dense static frame, like a photograph, and only then add movement - the model first builds a stable composition, and then animates it. A typical mistake is to write emotions instead of actions: “atmospheric”, “epic” are poorly understood by the model. Replace it with something visible: “the camera zooms in slowly, specks of dust in a beam of light.” The rule is simple: if an action cannot be shown on the screen, the model will ignore it.
Unlock access to submit solutions for instant AI review.