The expensive mistake is not picking the wrong model. It is picking one, getting a mediocre result, and concluding that the source photograph was the problem. I have watched people throw away a perfectly good still because the first attempt came back wrong, and then shoot a replacement that had the same flaw for a different reason.
So I spent an afternoon doing the boring version: one photograph, the same short brief, run through every model sitting in the same workspace. Not a benchmark. A log.
The still was a matte black desk lamp on a grey paper backdrop, shot flat, no shadow to speak of. The brief each time: the lamp head tilts down about fifteen degrees as the camera eases closer. Nothing clever.
Run 1–3: The Text-Only Baseline
First three passes were text-to-video with no reference image at all, which sounds pointless until you do it. You are asking the model to invent a lamp. It invents one. Then it invents a different one.
That is the point of the exercise. It tells you which parts of your description are load-bearing and which are decoration. “Matte black” survived every run. “Fifteen degrees” did not survive any of them — no model treats a number in a prompt as a measurement, and if you have ever wondered why your precise instructions get approximated, that is why.
Useful only as a diagnostic. Skip it if you are on a deadline.
Worth saying plainly, because people new to image to video get this backwards: the text-only mode is not a cheaper version of the image mode. It is a different job. One invents a subject, the other preserves one. If you came here to animate a photograph you already have, the text pass is a tuning fork, not a shortcut.
Run 4–6: Same Still, Reference-Led
Now the actual work. Upload the lamp as the opening frame, same sentence, and the difference is not subtle — the object is no longer something the model interprets, it is something the model has to keep.
This is where a browser workspace that turns stills into clips earns the subscription, because switching models is a dropdown rather than four accounts and four billing relationships. The roster in one place covers Seedance, MiniMax H3 and the Hailuo family, Veo, Kling and Wan. Same still, same prompt, six answers, one afternoon.
I am not going to rank them. Rankings from one photograph are worthless, and the model that handles a rigid product object well is not necessarily the one you want for a face. What I will say is that the spread between models on the same input was wider than the spread between my good prompt and my sloppy prompt on the same model. That surprised me. I had assumed prompt craft dominated.
The Two Settings That Moved the Needle
Everything else was noise. These two were not.
Resolution, chosen before generating. The ladder goes 480p through 4K on the Seedance side, 768P and 2K on MiniMax H3. Generating a draft at 4K to “see it properly” is how an afternoon becomes a day. Draft small. Commit large. I now run every first pass at the bottom of the ladder and it has cost me nothing.
Reference discipline. What MiniMax H3 does with an audio reference is a good illustration of the general rule: each file you attach has to have one job, and the model reads them together. Attach a still for composition, a clip for motion style, a sound bed for rhythm, and it works. Attach three stills that disagree about what the subject looks like and you get an average of three lamps.
One hard limit worth knowing before you plan around it: audio cannot be the only input. It rides along with something visual or not at all.
The shape menu deserves a line too. Six fixed ratios are on offer plus an adaptive setting, ultrawide at one end and full vertical at the other, and choosing up front costs nothing while cropping afterwards costs you the subject. I set shape in the same breath as resolution now, and any image to video job that skips that step pays for it later.
Where the Comparison Broke Down
Two honest failures from the log.
The flat grey backdrop defeated the push-in on more than one run. With no depth cues, “camera moves closer” and “object gets bigger” are the same pixels, and you cannot tell which one you got. A product still with a floor line or a visible edge would have made the test cleaner. My fault, not the models’.
And the lamp is a rigid object with no articulated parts, which is the easy case. Anything with hands, hair, or text on it would have produced a different ranking and probably a different conclusion. One photograph is one data point. Treat this log as a method, not a result.
Four Things Worth Copying
Small wins, in the order they paid off.
Run the throwaway text-only pass first when you have time. It shows you which words in your brief are doing nothing.
Draft at the lowest resolution the workspace offers, every time, without exception.
Switch models before you rewrite the prompt. If the output is wrong in the same way across three models, the problem is your still or your sentence. If it is wrong differently each time, it is the model, and switching is a two-second fix rather than a two-hour rewrite.
Keep the failures. The run where the camera move vanished into the background taught me more about what image to video needs from a source photograph than any of the successful passes did, and it cost me one draft at 480p.
That last point is the whole afternoon compressed. Image to video is cheap to iterate and expensive to argue about in the abstract. Six runs and a log beat a week of reading comparisons, including this one. Pick a still off your own drive tonight and burn six drafts on it. You will learn more about image to video from your own failures than from mine.

