Current AI video models copy how a product looks far better than how it behaves. A matte ceramic mug on a table is an easy shot. A clear glass bottle poured by a hand, with a readable ingredient label, stacks four of the hardest problems into four seconds.
So ask a narrower question. Which parts of your product, and which moments in your ad, will a model get right? Below, products are sorted by material and action, the research shows where models fail, and a six-question scorer tells you how to split a brief between generation and real footage.
Why does AI video get the look right and the physics wrong?
Video models learn from how things appear in footage. They carry no built-in sense of weight, friction or flow, so surfaces come out convincing while interactions come out uneven.
The evidence points the same way across labs. In its Sora technical report, OpenAI wrote that the model does not accurately simulate many basic interactions, such as glass shattering, and that food does not always change on screen when someone eats it. Google DeepMind and INSAIT then tested Sora, Runway, Pika, Lumiere, Stable Video Diffusion and VideoPoet on fluid dynamics, optics, solid mechanics, magnetism and thermodynamics in the Physics-IQ benchmark. Their conclusion: physical understanding is "severely limited, and unrelated to visual realism".
That second half matters for your review. A clip can look expensive and still show a pump that bends when pressed.
Models improve fast, and newer releases handle cases these studies missed. The pattern has held through each release so far: surfaces first, interactions last.
Which product materials are easy for AI video, and which are hard?
Sort your range by what the camera has to show. The table rates common materials from the easiest to the hardest to keep accurate in motion.
| Material | Example products | What tends to go wrong | Rating |
|---|---|---|---|
| Matte and rigid | Ceramics, sneakers, furniture, boxed goods | Proportions drift slightly between shots | Easy |
| Soft goods and fabric | Bedding, apparel on a hanger, plush toys | Folds and drape change shape from frame to frame | Moderate |
| Glossy surfaces and metal | Watches, cookware, metallic caps | Reflections jump or show a room that is not there | Moderate |
| Printed packaging | Skincare jars, supplements, food packs | Small text is redrawn, misspelled or smeared | Hard |
| Fine repeating detail | Knitwear, chains, faceted stones, woven baskets | Patterns swim or merge when the camera moves | Hard |
| Clear glass and plastic | Perfume, spirits, water bottles | Edges and refraction shift; contents float | Hard |
| Liquids and gels | Drinks, serums, sauces, cleaning products | Pours break physics and fill levels change on their own | Hardest |
Rate the shot as well as the product. A perfume bottle standing on stone is manageable, while the same bottle sprayed onto a wrist puts glass, liquid and hand contact into one moment.

Why do hands and moving parts fail more often than the product itself?
An interaction is harder than an object. When a hand grips a pump, two solid things have to meet, press and move together without passing through each other.
The VideoPhy benchmark measured this directly. Its authors wrote 688 prompts describing physical interactions, generated videos with models including CogVideoX, Pika, Runway Gen-2 and Luma Dream Machine, and had people judge whether each video followed the prompt and obeyed physical common sense. The best model tested, CogVideoX-5B, managed both in 39.6% of cases. For solid on solid interactions, the category that covers a hand twisting a cap or pressing a button, it managed 24.4%.

