Gemini Omni
Google Gemini Omni Flash - Text, image & reference video with native audio and multilingual lip-sync. On Influverse a 5s clip at 720p costs 25 credits, charged from the same wallet as every other model. There is no separate subscription for this model and no training step to use your own character with it.
By Influverse AI Team Last updated: 8 October 2026 Editorial policy
Generate with Gemini OmniCapabilities
| Vendor | |
|---|---|
| Duration | 3 to 10 seconds |
| Resolutions | 720p |
| Audio | Native audio on every generation |
| Text to video | Yes |
| Image to video | Yes |
| First and last frame | No |
| Motion control | No |
| Extend an existing clip past its last frame | No |
| Edit an existing clip from a prompt | Yes |
| Reference video | No |
| Reference images | Up to 10 |
What it costs in credits
Derived from the same pricing source the app charges against, so these figures and the cost preview on the generation screen can never disagree.
| Duration | Resolution | Credits |
|---|---|---|
| 3s | 720p | 15 |
| 5s | 720p | 25 |
| 10s | 720p | 50 |
Credits come from your plan and from credit packs. See the pricing page for the plan ladder.
When to use Gemini Omni
Gemini Omni Flash is Google's fast multimodal video model and the other place to go for talking content. Like HappyHorse it generates native audio and multilingual lip-sync on every clip, and it accepts text, an image and a large set of reference images. Unlike HappyHorse it will also take a finished clip back in and rework it: you hand it a video and a prompt, and it re-renders that footage the way you describe. Be precise about what that gives you, because the obvious assumption is that the clip gets longer, and that is the one thing this mode does not do. Nothing is added after the last frame of the video you supply. If the line you wrote runs longer than the take, the answer is a fresh generation or a different model, not this mode. What the mode is genuinely good for is a variation on a take you already like: change the setting behind the speaker, the wardrobe, the grade or the mood, and start from the framing and the delivery you already approved instead of rolling the dice again. Its duration menu is the shorter of the two talking models and it offers a single resolution, so the trade is control and speed against range. Reach for it when you want a spoken clip quickly, when you are working in a language other than English, or when you have a take that works and want a second version of it. It sits in the middle of the catalogue on price with audio included, which makes it reasonable value for talking content and poor value for anything silent, since you cannot switch the audio off. If the clip is going under a music bed or a separately recorded voiceover, generate it on Veo 3.1 Lite or Kling 2.6 Pro instead and keep the difference. If the piece needs to run past this model's duration ceiling, move to HappyHorse or Seedance, or build it out of several clips and cut them together. As with HappyHorse, remember that the dedicated talking-shot pipeline exists when you need the character to say exact words in the character's own voice.
Strengths
- Native audio and multilingual lip-sync on every clip
- Takes a finished clip back in and reworks it from a new prompt
- Large reference image set
- Fast turnaround for spoken content
Limits
- Single resolution and the shortest duration menu of the models that always generate audio
- Audio cannot be switched off, so it is poor value for silent clips
- No first and last frame control
- The edit mode reworks the clip you supply and adds nothing after its last frame
Best for
- Short spoken beats and quick talking clips
- Multilingual social content
- A second version of a take whose delivery already works
Other models to consider
Back to the full model catalogue
More from Influverse
Generate with Gemini Omni
Every model is open on every plan, the free one included. Start your free 7-day trial with 120 credits. $0 today, cancel anytime.
Start creating free