Every AI Video Model in VIBE
39 AI video models and 10 AI image models, in one app, on your phone.
VIBE is an AI video generator app for iOS and Android that bundles 39 AI video models and 10 AI image models behind a single interface. That includes Google Veo 3.1, OpenAI Sora 2 and Sora 2 Pro, Kling 3, Seedance 2.5, WAN 2.7, Grok Imagine, Luma Ray 3.2, LTX 2.5 Fast, Flux 3 and PixVerse 6. Eleven video models are free to use without an account, and the rest are available with a subscription. The full specification of every model is listed below.
- 39 AI video models from Google, OpenAI, Kuaishou, ByteDance, Alibaba, xAI and more
- 10 AI image models, including one that is free every day
- 11 video models free with no account and no payment
- Resolutions from 360p to true 4K, clips from 2 to 30 seconds
- Text-to-video, image-to-video, audio-synced video, and talking avatars
- Run the same prompt on several models and compare the results
AI video models
| Model | Developer | Resolution | Duration | Input | Audio | Tier |
|---|---|---|---|---|---|---|
| Seedance Pro FastCamera-fixed (static camera) mode | ByteDance | 480p | 2, 3, 4, 6, 8, 10, 12s | Text + Image | No | Free |
| WAN 2.2 | Alibaba | 720p | 5s | Text + Image | No | Free |
| LTX 2 Fast | Lightricks | 720p | 6β20s (2s steps) | Text only | Optional | Free |
| LTX 2 DistilledPrompt enhancement | Lightricks | 720p | 5s or 10s | Text + Image | No | Free |
| Seedance | ByteDance | 720p | 5s or 10s | Text + Image | No | Free |
| Luma Ray Flash 2Start-image anchoring | Luma AI | 720p | 5s or 9s | Text + Image | No | Free |
| Motion 2.0 | Leonardo.Ai | 720p | 5s or 10s | Text + Image | No | Free |
| Seedance 1.5 ProLast-frame image, camera-fixed mode, seed control | ByteDance | 1080p | 5, 10, 12s | Text + Image | Optional | Premium |
| PixVerse 6Negative prompt; audio toggle defaults OFF; 360p forced | PixVerse | 360p | 5, 8, 10, 15s | Text + Image | Optional | Free |
| Sora | OpenAI | 1080p | 5, 10, 15, 20s | Text only | No | Premium |
| Sora 2World simulation / physics; synchronized dialogue + sound effects; image reference input | OpenAI | 720p | 4, 8, 12s | Text + Image | Yes | Premium |
| Sora 2 ProHighest-quality Sora tier; image reference input | OpenAI | 720p Β· 1080p | 4, 8, 12s | Text + Image | Yes | Premium |
| Google Veo 2 | Google DeepMind | 1080p | 5s or 8s | Text only | No | Premium |
| WAN 2.6Multi-shot scene transitions for longer storytelling | Alibaba | 720p Β· 1080p | 5, 10, 15s | Text + Image | Yes | Premium |
| WAN 2.7Upload a voice or music track (wav/mp3, 3β30s, β€15 MB) to sync the video to it; 720p forced server-side | Alibaba | 720p | 2β15s (any whole second) | Text + Image | Yes | Premium |
| Happy Horse 1.1Audio always on (no toggle) | Alibaba | 720p Β· 1080p | 3β15s (any whole second) | Text + Image | Yes | Premium |
| Kling ProCinematic look | Kuaishou | 1080p | 5s or 10s | Text + Image | No | Premium |
| Kling v2.5 Turbo ProFast 1080p turbo tier | Kuaishou | 1080p | 5s or 10s | Text + Image | No | Premium |
| Kling v2.6 | Kuaishou | 1080p | 5s or 10s | Text + Image | Optional | Premium |
| Kling O3 ProVoice binding for character/voice continuity across generations | Kuaishou | 1080p | 3β15s (any whole second) | Text + Image | Optional | Premium |
| Kling O3 Standard | Kuaishou | 1080p | 3β15s (any whole second) | Text + Image | Optional | Premium |
| Kling 3 ProMulti-language audio; structured stories of up to 6 shots, each with its own prompt | Kuaishou | 1080p | 3β15s (any whole second) | Text + Image | Optional | Premium |
| Kling 3 Standard | Kuaishou | 1080p | 3β15s (any whole second) | Text + Image | Optional | Premium |
| Grok Imagine | xAI | 480p Β· 720p | 6, 8, 10s | Text + Image | No | Free |
| Vidu Q3 | Vidu | 360p Β· 540p Β· 720p Β· 1080p | 4, 5, 6, 8s | Text + Image | Yes | Premium |
| Google Veo 3 FastNegative prompt, seed control | Google DeepMind | 4K | 5s or 8s | Text only | Yes | Premium |
| Google Veo 3.1Up to 3 reference images for subject consistency (reference-to-video); last-frame input | Google DeepMind | 720p Β· 1080p | 4, 6, 8s | Text + Image | Optional | Premium |
| Google Veo 3.1 FastLast-frame input for frame-to-frame transitions | Google DeepMind | 720p Β· 1080p | 4, 6, 8s | Text + Image | Optional | Premium |
| Google Veo 3.1 LiteCheapest Veo tier; reference-image input; 720p only | Google DeepMind | 720p | 4, 6, 8s | Text + Image | Yes | Premium |
| PRUNA VUpload your own audio (mp3/wav/flac) to sync video to music or speech; talking-avatar style lip sync from a single photo | Pruna AI | 720p | 2β10s (any whole second) | Text + Image | Yes | Free |
| LTX 2.3 FastTrue 4K (2160p) output | Lightricks | 1080p Β· 1440p Β· 2160p | 6β20s (2s steps; over 10s requires 1080p) | Text + Image | Optional | Premium |
| Seedance 2.0Synchronized dialogue, sound effects and music | ByteDance | 480p | 4β15s (any whole second) | Text + Image | Optional | Premium |
| Talking AvatarTalking avatar: a photo + a voice recording become a lip-synced talking video. NO text prompt at all. Output frame matches the input photo. | β | 480p Β· 720p | Matches the uploaded voice recording (no duration picker) | Image only | Yes | Premium |
| Seedance 2.0 MiniBudget tier of Seedance 2.0 with synchronized audio | ByteDance | 480p Β· 720p | 4β15s (any whole second) | Text + Image | Optional | Premium |
| Grok Imagine 1.5Image-to-video only (prompt optional, guides motion); synchronized audio that Grok Imagine 1.0 lacks | xAI | 480p Β· 720p | 4, 6, 8, 10, 12, 15s | Image only | Yes | Premium |
| Luma Ray 3.2Reasoning video model for cinematic results; start-image anchoring (5s clips only) | Luma AI | 720p Β· 1080p | 5s or 10s (10s is text-to-video only) | Text + Image | No | Premium |
| LTX 2.5 FastLast-frame interpolation (first frame + last frame); selectable frame rate; scales from 720p to 4K | Lightricks | 720p Β· 1080p Β· 2k Β· 4k | 2β20s (2,3,4,5,6,8,10,12,14,16,18,20; over 10s requires 24/25 fps) | Text + Image | Yes | Free |
| Flux 3Storyboard mode: up to 6 images the clip moves through (3+ images require an explicit duration). Draft mode for a fast cheap 720p preview. | Black Forest Labs | 720p Β· 1080p | 5, 6, 8, 10, 12, 15, 20s, or auto | Text + Image | Yes | Premium |
| Seedance 2.5Up to 4 reference images to keep a character consistent across clips, OR a first + last frame (mutually exclusive; first/last-frame mode forces the adaptive aspect ratio). Synchronized audio including dialogue. | ByteDance | 480p Β· 720p | 4β30s (any whole second) | Text + Image | Yes | Premium |
AI image models
Generate an image from text, then animate it with image-to-video without leaving the app.
| Model | Developer | Tokens / image | Image input | Tier |
|---|---|---|---|---|
| Nano Banana Pro | 4 (1K) / 5 (2K) / 10 (4K) | Yes | Premium | |
| Flux 2 Pro | Black Forest Labs | 4 | Yes | Premium |
| Seedream 4.5 | ByteDance | 5 (2K) / 8 (4K) | Yes | Premium |
| Qwen Image | Alibaba | 3 | Yes | Premium |
| Kling V3 | Kuaishou | 3 | Yes | Premium |
| Grok Imagine | xAI | 2 | Yes | Premium |
| Nano Banana | 2 | Yes | Premium | |
| Ideogram 3 | Ideogram | 5 | Yes | Premium |
| Imagen 4 | 3 | No | Premium | |
| Flux Schnell | Black Forest Labs | 1 | No | Free |
Specifications reflect what the VIBE app offers today. New models are usually added within days of release, and this page is updated with them.
Which AI video model should you use?
There is no single best model. These are the picks that hold up in practice, by goal.
Most realistic footage
Google Veo 3.1 for photorealistic camera work and lighting, and OpenAI Sora 2 for physical realism β how objects move, collide, and settle. Both generate synchronized audio in the same pass as the picture.
Longest clips
Seedance 2.5 reaches 30 seconds, which is the longest in the app. LTX 2.5 Fast, LTX 2.3 Fast and Flux 3 reach 20 seconds. Kling 3, Kling O3, WAN 2.7 and Happy Horse 1.1 reach 15.
Highest resolution
LTX 2.5 Fast and LTX 2.3 Fast output true 4K. Most flagship models β Veo 3.1, Kling 3, Sora 2 Pro, Luma Ray 3.2 β top out at 1080p, which is what social platforms and ad placements actually need.
Free to use
Seedance Pro Fast is the fastest free model, LTX 2.5 Fast the highest quality, and PRUNA V adds free audio and lip sync. Grok Imagine and PixVerse 6 are free too. No account, no payment.
Consistent characters
Google Veo 3.1 accepts up to three reference images and Seedance 2.5 accepts up to four, so a character, product, or set stays recognisable across separate clips instead of changing every generation.
Sound and voice
WAN 2.7 and PRUNA V let you upload your own music or voice recording and sync the video to it. The Talking Avatar model turns a photo plus a voice recording into a lip-synced talking video.
Tokens, free models, and what a subscription changes
VIBE uses tokens rather than a per-model subscription. Each generation costs a number of tokens based on the model you picked, the resolution, and the length of the clip β a cheap free model costs a couple of tokens per second, while a flagship model at 1080p costs considerably more. You see the exact cost before you generate.
Eleven of the video models are free. You can use them without creating an account, entering payment details, or starting anything that later charges you. Among the image models, Flux Schnell is free, with a daily allowance for users who are not subscribed.
A subscription unlocks the premium models β Veo 3.1, Sora 2 and Sora 2 Pro, Kling 3, Seedance 2.5, WAN 2.6 and 2.7, Flux 3, the Talking Avatar model and the rest β plus higher resolutions, longer clips, and synchronized audio. Token packs can also be bought separately if you would rather not subscribe.
AI video models β frequently asked questions
Which app has the most AI video models?
VIBE includes 39 AI video models and 10 AI image models in a single iOS and Android app, including Google Veo 3.1, OpenAI Sora 2, Kling 3, Seedance 2.5, WAN 2.7, Grok Imagine, Luma Ray 3.2, LTX 2.5 Fast and Flux 3. Most competing tools offer between one and three models.
Can I use Sora 2, Veo 3.1 and Kling in the same app?
Yes. All three are available inside VIBE on iOS and Android, so you can run the same prompt through each one and compare the results without separate subscriptions or separate accounts.
Which AI video model is the most realistic?
Google Veo 3.1 is generally strongest for photorealistic lighting and camera movement, while OpenAI Sora 2 is strongest for physical realism β how objects move and interact. Which one wins depends on the shot, so testing the same prompt on both is the practical approach.
Which AI video generator makes the longest videos?
Inside VIBE, Seedance 2.5 generates the longest single clip at up to 30 seconds. LTX 2.5 Fast, LTX 2.3 Fast and Flux 3 reach 20 seconds, and Kling 3, Kling O3, WAN 2.7 and Happy Horse 1.1 reach 15 seconds.
Which AI video models are free?
Eleven video models are free in VIBE: Seedance Pro Fast, WAN 2.2, LTX 2 Fast, LTX 2 Distilled, Seedance, Luma Ray Flash 2, Motion 2.0, PixVerse 6, PRUNA V, Grok Imagine and LTX 2.5 Fast. No account or payment is required to use them.
Can any of these models generate 4K video?
Yes. LTX 2.5 Fast and LTX 2.3 Fast output true 4K. Most other flagship models generate 1080p, which is what YouTube, TikTok, Instagram and ad platforms actually require.
Which AI video models generate sound?
Sora 2 and Sora 2 Pro always generate synchronized audio, as do WAN 2.6 and 2.7, Happy Horse 1.1, Vidu Q3, Google Veo 3.1 Lite, LTX 2.5 Fast, Flux 3, Seedance 2.5 and PRUNA V. Veo 3.1, Veo 3.1 Fast, Kling 3, Kling O3, Kling v2.6, LTX 2.3 Fast and Seedance 2.0 let you switch audio on or off, which also changes the cost.
How often are new AI models added?
New models are usually added within days of public release, and this page is updated when they are. Recent additions include Seedance 2.5, LTX 2.5 Fast, Flux 3, Kling 3, WAN 2.7 and Grok Imagine 1.5.
Model guides
Deeper detail on the models people ask about most.