Ir para o conteúdo principal
5 minLesson 6/6

A practical way to decide between image, video and audio models instead of guessing.

OnZero gives you several models in every medium, and the right choice depends on what you are optimizing for - identity accuracy, speed, photorealism, or sharp text and layout. Picking with intent beats picking at random and hoping.

Image models

The default image model prioritizes identity and face accuracy - the safest choice whenever a locked character is attached. Faster variants trade a little of that precision for speed. Photoreal-leaning models push toward camera-like realism, while text-strong models handle on-canvas typography and layout best. If a face keeps drifting, the model is often the first thing to check.

Video and audio models

Video models trade off differently across cost, clip length and how well they follow precise motion direction - a quick social clip and a carefully choreographed keyframed shot often call for different engines. Audio has its own split too: character voice and voiceover prioritize clarity and consistency, while music and sound-effect generators prioritize mood matching over precision.

  1. 1Ask what you are optimizing for: identity accuracy, speed, realism, or text and layout.
  2. 2Default to the identity-first model whenever a locked character is attached.
  3. 3Switch models if results consistently miss on the same dimension.
  4. 4Check the model list any time a new job type comes up - defaults change less often than your needs do.

When in doubt, start with the default model for the medium - it is set as default because it is the safest general-purpose choice - and only switch once you know specifically what you are trading away.

Fundamentals, patterns, motion, audio, iteration and model choice - together these six lessons are the craft underneath every other track in this academy. Apply them anywhere you prompt in OnZero and every other skill gets easier.