All insights

Can AI video show hands using your product? What works and what to avoid

Which hand shots AI video handles and which break, from a still hold to pouring and applying, with a seven-question scorer and shot designs that avoid the failures.

Marianne Larsen
AI Producer, AI Vidia

Yes for a still hold or brief contact with a rigid product, and rarely for a hand that must twist, pump, pour or apply a product for several seconds. Video models draw how hands look better than how contact behaves, and benchmark results show solid on solid contact as the weakest case. Score your shot below, then cut around the failure or film the real hand.

An unlabelled amber dropper bottle and a white ceramic bowl on a light marble counter, with a hand entering the frame from the left, its fingers softly out of focus at the edge of the shot
Image generated with AI: a fictional, unlabelled dropper bottle on a counter with a hand entering the frame.
On this page7 sections
  1. 01Can AI video do hands now?
  2. 02Why do contact shots break when a still hand looks fine?
  3. 03Which hand shots are low risk and which are high risk?
  4. 04Do pours, squeezes and textures fail the same way?
  5. 05How do you design the shot so the failure never shows?
  6. 06When does a hand shot turn into a claim?
  7. 07How do you review hand shots before sign-off?

A hand opening a jar, twisting a dropper or pouring a sauce sells a product faster than a pack standing on a shelf. It is also the moment where generated video shows its seams.

The answer depends on what the hand does and for how long. A hand resting on a lid for half a second is a different request from a hand that unscrews it while the camera pushes in. Score your own shot in the block further down, then use the sections that match the band you land in.

Can AI video do hands now?

Better than it used to, and still uneven wherever fingers meet the product. Hands in easy poses come out clean more often than they did. Hands that grip, twist, press or pour still produce the failures a media buyer spots in the first second: a sixth finger, a thumb that slides through the pack, a lid that changes shape as it turns.

The research explains the pattern. The authors of HandRefiner write that diffusion models often produce malformed hands, with wrong finger counts and irregular shapes, because the structure of a hand involves heavy deformation and occlusion. A hand around a product adds a second object that hides part of the hand and must keep its own shape.

Treat every hand shot as a small production decision. Some are safe to generate. Some need a different camera angle. A few need the real product and a real hand on camera.

Why do contact shots break when a still hand looks fine?

A video model draws each frame to look like the footage it has seen. It has no rule that a finger cannot pass through a bottle. Kang and colleagues tested this in a controlled physics setting (ICML 2025) and found that models generalize case by case, copying the closest example they have seen, and fail outside the range of their training data. Your particular pump, a cap with a quarter turn or an unusual clasp sits outside that range more often than a generic action does.

The VideoPhy benchmark measured contact directly. People judged whether each generated video followed its prompt and obeyed physical common sense. The best model tested, CogVideoX-5B, managed both in 39.6% of cases overall. For solid on solid interactions it managed 24.4%. The paper names a bottle toppling off a table as an example of that category.

Contact between solid objects was the hardest interactionShare of generated videos that followed the prompt and obeyed physics, best model tested (CogVideoX-5B), by interaction type
Contact between solid objects was the hardest interaction. Share of generated videos that followed the prompt and obeyed physics, best model tested (CogVideoX-5B), by interaction type
ItemValue
All prompts40%
Solid on solid24%
Solid and fluid53%
Fluid on fluid44%

VideoPhy, Bansal et al., arXiv 2406.03520v2, Table 3: human evaluation, October 2024. Open models of 2024, not current commercial tools.

These were open models from 2024, so read the chart for its order and not its levels. The follow-up study VideoPhy-2 (March 2025, 200 real-world actions) reported that the best model reached 22% joint adherence on its hard action subset and named conservation of mass and momentum as a weak spot.

A hand closing around a rigid product is a solid on solid case. That reading is our inference: neither paper tests product ads. That is why contact moments deserve the most attention in a brief.

Which hand shots are low risk and which are high risk?

Sort the shots in your brief by what the hand does. The table runs from the safest request to the riskiest.

Hand actions from lowest to highest risk in generated video. Ratings are AI Vidia production guidance based on the sources listed below, not a measured benchmark.
Hand actionExampleWhat tends to go wrongRisk
No contactProduct on a counter, hand out of frameNothing hand relatedLow
Still holdA hand rests on a lid or holds a tube to cameraFinger count and thumb position drift between framesLow to medium
Grip and liftA hand picks up a bottle and raises itFingers merge into the pack and the weight looks wrongMedium
PourA sauce, oil or drink leaves the containerThe stream thins or breaks and the fill level changesHigh
Fine manipulationTwisting a cap, pumping, using a dropper, closing a claspThe part changes shape and a finger passes through itHigh
Apply to skinSmoothing a cream or serum onto a hand or faceTexture and skin look wrong, and the result reads as a claimHigh

Contact type is only the first input. Longer contact gives the model more frames to get wrong. Two visible hands double the fingers to count. A product that hides the fingers forces the model to invent what is behind it. Clear glass, liquids and fast camera moves each add risk on top.

Two hands lifting the dropper cap out of a small amber glass bottle with a blank label on a light grey counter, shown in a close-up with the fingers and the dropper clearly visible
Image generated with AI: a fictional, unbranded dropper bottle at the moment a hand lifts out the dropper, the kind of fine handling the scorer rates as high risk.
ScorecardScore a hand shot before you generate itPick the option that matches the shot in your brief. The total is a risk band, not a prediction for a specific model.
  1. What does the hand do?

    Choose the hardest thing the hand does in the shot.

  2. How long does the contact last?
  3. How many hands are visible?
  4. Does the product hide the fingers?

    For example a hand wrapped around a bottle or a jar.

  5. What is the product made of?
  6. How does the camera move?
  7. Does the action prove a claim?

    Texture, absorption, a result or how well it works.

Score: 0 of 21 · 0 of 7 answered

Answer every question to see the verdict.

How to read your score
ScoreVerdictWhat it means
0 to 4Low risk: generate it and review itGenerate the shot, check finger count and product outline at quarter speed, and keep the take that passes. Keep the contact short.
5 to 9Medium risk: reframe before you generateUse one of the alternatives: cut on contact, a first-person angle with fewer visible fingers, or the product at rest after the hand leaves. Expect several rejected takes if you keep the full action. If the shot proves a claim, film it with the real product.
10 to 21High risk: film it or design around itFilm the real hand and the real product, or composite a real hand plate into a generated scene. If the shot proves a claim, the demonstration must reflect how the product really performs, so generating it is not an option.

Points are AI Vidia production guidance built on the sources in this article. They are not a measured benchmark.

A medium band means the idea can work with a changed shot. A high band means the action itself is the problem, and no prompt fixes that reliably.

Need new ads?
Describe your product and what you want to test.
Send your brief

Do pours, squeezes and textures fail the same way?

Liquids fail differently. In the VideoPhy data, solid and fluid interactions scored 53.1% for the best model, the strongest of the contact categories, which still leaves almost half of the prompts failing. A stream that looks right for two seconds can thin out, change thickness or fill a glass that was already full.

Check the stream, the fill level and the container in every pour. The stream keeps one thickness from the container to the surface. The level rises once and stays where it lands. The container keeps its shape while it tilts.

Squeezes and dollops follow the same logic. A gel that holds a peak on the fingertip is a texture claim, and so is a cream that melts into skin. That leads to the second question about any hand shot: is it a scene or a proof?

How do you design the shot so the failure never shows?

Most of the fixes happen in the shot plan and not in the prompt. These five alternatives cover the medium and high bands, listed from the cheapest to the most expensive.

  1. Cut on contact. The hand enters at 0:01 and the edit cuts at 0:02, before the fingers close. The next shot already shows the product opened or in use.
  2. First-person framing. Put the camera at the eye line of the person using the product, so only a thumb and forefinger reach into frame. Fewer visible fingers means fewer places to fail.
  3. Product at rest after the hand leaves. Show the hand set the product down, then hold on the product alone for the last seconds of the beat.
  4. A real hand plate composited in. Film a real hand doing the action on a plain background and composite it into the generated scene.
  5. Real footage for the demonstration. Film the real product and a real hand for any shot that proves what the product does.
A three-panel storyboard of a shot plan on a desk: a hand entering the frame toward a bottle, a plain cut frame in the middle, and the bottle standing alone on a counter
Image generated with AI: the cut-on-contact plan as three panels, hand enters, cut, product at rest.

Write the plan into the brief with timestamps. A line such as "hand enters 0:01, cut 0:02, product at rest 0:02 to 0:04" tells the producer where the contact ends and gives your editor the cut point.

When does a hand shot turn into a claim?

A hand pouring a sauce into a bowl is a scene. A hand smoothing a cream that visibly vanishes into skin is a statement about what the cream does. Advertising rules treat the second one as evidence.

The classic case is FTC v. Colgate-Palmolive, decided by the US Supreme Court in 1965. A shaving cream commercial showed sand on a sheet of plexiglass standing in for sandpaper. The Court held that presenting a mock-up as a real test is deceptive, even when the underlying claim about the product is true. By analogy, a generated clip of a serum absorbing or a spray clearing a stain is a mock-up.

The UK regulator says the same about visuals. Its guidance on production techniques (updated 18 August 2025) states that visual claims should not misleadingly exaggerate the effect a product can achieve, and that before and after images must reflect typical results. On 11 June 2026 the ASA added in its note on AI and deepfakes that the CAP Code is media-neutral, so the rules apply however the ad was made, and that the advertiser stays responsible when automated tools produce the content.

This is guidance from one US case and one UK regulator, and not legal advice for your market. It supports a simple working rule. Mark every hand shot in the brief as scene or proof. Scene shots can be generated. Proof shots use the real product and a real hand.

How do you review hand shots before sign-off?

Review is where a clean hand gets selected. It does not come from the first take. The steps below are a short pass for every clip.

  1. Play the clip at quarter speed, once for each visible hand.
  2. Count fingers on the first frame, the contact frame and the last frame.
  3. Watch the product outline through the contact moment. The cap, pump and label must keep their shape.
  4. Check every visible word against your flat label artwork, as described in our guide to hard-to-generate products.
  5. Write a timestamped note for each defect, such as "cap warps at 0:06" or "thumb passes through the pump at 0:03".
  6. Decide per clip: keep, regenerate, reframe with one of the alternatives above, or film.

Set a limit. After three rejected takes of the same contact moment, change the shot and leave the prompt alone. The same defect list belongs in your creative QA gate and your brand safety review, so approvers see the same notes.

Try this in 20 minutes. Open your current brief and list every second where a hand touches the product. Score each one with the block above, mark which ones prove a claim, and pick an alternative for every medium or high result. If you want a producer to run the same pass, the process page shows where review sits in a project, and the AI video ads service describes what gets delivered.

Frequently asked questions

01Can AI video generate a realistic hand holding my product?
Yes for a still or slow hold that lasts about a second, with the fingers in view and a rigid product. Reliability drops when the hand grips hard, wraps around the pack or holds it for several seconds. Review the take at quarter speed, count fingers at the first, contact and last frames, and regenerate or reframe when the thumb or the pack changes shape.
02Why does AI struggle with hands?
The authors of HandRefiner attribute malformed hands to the way hand structure involves heavy deformation and occlusion, which makes it hard for diffusion models to learn. A hand has many joints, changes shape with every grip and is often partly hidden. Adding a product adds a second object that hides fingers and must keep its own shape, so contact shots fail more often than hands alone.
03Can I show a hand applying my cream or serum in AI video?
Use AI video for the scene and film the real product for the result. A generated clip of cream absorbing into skin is a mock-up, and the ASA and the US Supreme Court in FTC v. Colgate-Palmolive both treat misleading demonstrations as a problem. Film the application with the real product, or cut before the texture is visible.
04How do you fix a bad hand in a generated clip?
Change the shot before you change the prompt. Cut earlier so the failing frames never appear, reframe from a first-person angle with fewer fingers, or composite a real hand plate into the scene. Tools such as HandRefiner correct hands in still images, and a video clip needs a reframe or a filmed hand instead. After three rejected takes, change the shot.
05How many hand shots should one video ad contain?
Plan one or two contact moments in a 15 second ad and keep each under two seconds. This is AI Vidia production guidance and not a measured rule. Fewer contact moments mean fewer clips to reject, and you can film the ones that carry proof instead of trying to generate them.

How we made this article

Sources were read on 30 September 2026. The percentages come from Table 3 of VideoPhy (human evaluation of 2024 open models) and from the VideoPhy-2 abstract. The risk ratings, the scorer and the review steps are AI Vidia production guidance built on those sources, not a measured benchmark. The reading that hand contact counts as a solid on solid interaction is our inference. No video tools were tested for this article, and the images are AI-generated illustrations of a fictional product.

We research and draft with AI assistance and check figures against the sources below.

Sources

  1. 01Bansal et al.: VideoPhy, Evaluating Physical Commonsense for Video Generation, arXiv 2406.03520v2, October 2024
  2. 02Bansal et al.: VideoPhy-2, A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation, arXiv 2503.06800, March 2025
  3. 03Lu et al.: HandRefiner, Refining Malformed Hands in Generated Images by Diffusion-based Conditional Inpainting, arXiv 2311.17957
  4. 04Kang et al.: How Far is Video Generation from World Model, A Physical Law Perspective, ICML 2025, arXiv 2411.02385
  5. 05US Supreme Court: FTC v. Colgate-Palmolive Co., 380 U.S. 374 (1965), Cornell Legal Information Institute
  6. 06ASA and CAP: Beauty and cosmetics, the use of production techniques (updated 18 August 2025)
  7. 07ASA: AI and deepfakes, four things advertisers need to know before they hit run (11 June 2026)

Next step

Get a proposal for your next ads.

Tell us about your product, audience and ad production needs. We review your brief and reply with a proposed next step.

Send your brief

Read next