Alibaba's 14B image-to-video model, running without a local GPU.
Wan 2.2 I2V-A14B is one of the strongest open image-to-video models available. It takes a still image and a description of the motion and produces 720p video from it — the subject, framing and colour of your source frame are preserved while the scene moves.
Running it locally means a 14-billion-parameter model and a lot of VRAM. Here it runs server-side, so the requirement is a browser and an image. Clip length is 5 or 8 seconds, and because it is image-to-video rather than text-to-video the composition is decided by the frame you supply rather than re-rolled by the model — which is what makes it repeatable.
Render time is genuinely long — around two minutes for a 5-second clip, and almost all of that is the model itself rather than queueing. Length is limited to 5 or 8 seconds; the provider rejects anything else. For text-to-video with no source image, use Kling instead.
| Price | 494 credits per second — about 45c per second |
|---|---|
| Type | Video |
| Provider | fal |
| Aspect ratios (3) | 16:9, 9:16, 1:1 |
| Durations | 5s, 8s |
Credit prices are the live studio rates. Cash equivalents are approximate at starter-pack rates and improve on the larger packs — see pricing.
Every AI model on comfyarts — video models, prices and measured capabilities side by side.