You do not have a camera.
You have a machine that dreams a short motion out of a single still image, and it dreams badly the moment you ask it for something the still does not already contain.
I learned this across a 10-episode series, and every rule below was paid for in failed generations.
None of it is theory.
The medium's real physics A real camera moves through a space that exists whether or not you point at it.
The model has no space.
It has one flat image and a statistical guess about what "zoom out" tends to look like in its training data.
When the frame widens, the model is not revealing more of a room that was always there.
It is inventing pixels to fill the new area, drawn from everything it has ever seen.
That single fact reorganizes everything you know about directing: There is no coverage.
Every "angle" is a separate generation from a separate still.
Continuity is not captured; it is engineered, frame by frame.
Nothing survives the cut for free.
The model does not know that shot 12 and shot 13 are the same character in the same room.
Anything you want to persist (damage state, light, color) must be re-declared or re-anchored every single time.
The model abhors an empty frame.
Its deepest reflex is to resolve ambiguity: a silhouette becomes a face, fog becomes a mountain range, a clean retro interior grows drips and cobwebs because "analog" reads as "abandoned".
Spawn pressure is constant.
Background figures flicker into existence in any populated-looking scene.
Every motion prompt in my pipeline ends with an anti-spawn guard: "Do not add extra characters.
Keep everything as pictured." Drop that guard and the figures come back.
A widening or traveling frame is an invitation for the model to hallucinate.
Direct this camera and you are not choosing what to show.
You are choosing what to withhold from its imagination.
The classical grammar, re-pointed If you carry film vocabulary, it all still applies.
The mechanism just changes completely.
Classical tool Here Lens choice There is no lens.
The "look" is a prompt suffix asserted in words on every clip.
Depth of field is a keyword, not an aperture.
Blocking You cannot choreograph.
The reliable unit is micro-motion: one head turn, one hand, environmental drift. "A does X while B does Y while camera does Z" produces morphing garbage.
The 180° rule The model has no memory of the line.
You hold it in the writing, naming screen directions explicitly, shot after shot.
Coverage Multi-clip + frame chaining: the last frame of one clip becomes the start frame of the next (max 3 in a chain, because error compounds).
Camera move A semantic suggestion the model interprets loosely.
Moves that reveal new area (tilt, pan, zoom out, crane) are the highest-risk category.
One hard rule sits on top: one move per clip.
Two simultaneous move-instructions are two conflicting statistical pulls on the same frame; the result is a smeared average, or the model picks one, or it morphs.
One clean move plus one or two atmospheric elements.
That is the whole motion budget of a clip.
The shot that failed four times The widest interior reveal of episode 9: the fully-mended android, gold in its seams, self-luminous, revealed in its workshop.
Plan: single start frame, slow zoom out.
On a real set, a trivial dolly-back.
Four reshoots failed the same way.
As the frame widened past the borders of the source still, the model filled the new area with furniture and clutter that do not exist in the universe.
The anti-spawn guard did not save it, because that guard forbids extra characters, and the model was inventing environment.
No adjective fixed it.
I stopped asking the model to imagine the destination and handed it one that already existed: a wide environmental shot from elsewhere in the same episode.
Start frame plus end frame, and a prompt rewritten for two-frame continuity.
It worked on the first try.
With the destination pinned to a real frame, the model had nothing left to invent.
The whole failure-an