Google video model

Gemini Omni

Google Gemini Omni Flash - Text, image & reference video with native audio and multilingual lip-sync. On Influverse a 5s clip at 720p costs 25 credits, charged from the same wallet as every other model. There is no separate subscription for this model and no training step to use your own character with it.

By Influverse AI Team Last updated: 8 October 2026 Editorial policy

Generate with Gemini Omni

Capabilities

VendorGoogle
Duration3 to 10 seconds
Resolutions720p
AudioNative audio on every generation
Text to videoYes
Image to videoYes
First and last frameNo
Motion controlNo
Extend an existing clip past its last frameNo
Edit an existing clip from a promptYes
Reference videoNo
Reference imagesUp to 10

What it costs in credits

Derived from the same pricing source the app charges against, so these figures and the cost preview on the generation screen can never disagree.

DurationResolutionCredits
3s720p15
5s720p25
10s720p50

Credits come from your plan and from credit packs. See the pricing page for the plan ladder.

When to use Gemini Omni

Gemini Omni Flash is Google's fast multimodal video model and the other place to go for talking content. Like HappyHorse it generates native audio and multilingual lip-sync on every clip, and it accepts text, an image and a large set of reference images. Unlike HappyHorse it will also take a finished clip back in and rework it: you hand it a video and a prompt, and it re-renders that footage the way you describe. Be precise about what that gives you, because the obvious assumption is that the clip gets longer, and that is the one thing this mode does not do. Nothing is added after the last frame of the video you supply. If the line you wrote runs longer than the take, the answer is a fresh generation or a different model, not this mode. What the mode is genuinely good for is a variation on a take you already like: change the setting behind the speaker, the wardrobe, the grade or the mood, and start from the framing and the delivery you already approved instead of rolling the dice again. Its duration menu is the shorter of the two talking models and it offers a single resolution, so the trade is control and speed against range. Reach for it when you want a spoken clip quickly, when you are working in a language other than English, or when you have a take that works and want a second version of it. It sits in the middle of the catalogue on price with audio included, which makes it reasonable value for talking content and poor value for anything silent, since you cannot switch the audio off. If the clip is going under a music bed or a separately recorded voiceover, generate it on Veo 3.1 Lite or Kling 2.6 Pro instead and keep the difference. If the piece needs to run past this model's duration ceiling, move to HappyHorse or Seedance, or build it out of several clips and cut them together. As with HappyHorse, remember that the dedicated talking-shot pipeline exists when you need the character to say exact words in the character's own voice.

Strengths

  • Native audio and multilingual lip-sync on every clip
  • Takes a finished clip back in and reworks it from a new prompt
  • Large reference image set
  • Fast turnaround for spoken content

Limits

  • Single resolution and the shortest duration menu of the models that always generate audio
  • Audio cannot be switched off, so it is poor value for silent clips
  • No first and last frame control
  • The edit mode reworks the clip you supply and adds nothing after its last frame

Best for

  • Short spoken beats and quick talking clips
  • Multilingual social content
  • A second version of a take whose delivery already works

Other models to consider

Back to the full model catalogue

More from Influverse

Generate with Gemini Omni

Every model is open on every plan, the free one included. Start your free 7-day trial with 120 credits. $0 today, cancel anytime.

Start creating free