Browse documentation
Reference

Models

The model catalog available through the studio, MCP, and CLI.

Image

IDStrengthBase credits
soulFashion-grade realism, 89 aesthetic presets4/image
flux-proPrecise prompt following, crisp detail8/image
gpt-image-2Photorealism, typography, and complex instructions4/image
seedream-5-proFlagship visual reasoning and multilingual text4/image
seedream-5-liteFast high-resolution creative generation3/image
nano-banana-2Fast reasoning, text, and reference consistency5/image
nano-banana-proMaximum compositional reasoning up to 4K8/image
nano-banana-2-liteSub-two-second generation at fixed 1K3/image
grok-imagine-2Versatile photoreal and stylized generation4/image
flux-2-proProduction-ready detail and consistency3/image
flux-2-flexFine-grained control over precision and speed6/image
recraft-v4Design-led imagery and brand-ready composition3/image
recraft-v4-proPremium design and photoreal production assets13/image
z-imageFast, economical generation for high-volume work2/image
wan-2.2-imageHigh-fidelity cinematic still generation1/image
kling-o1-imagePrecise multi-reference composition2/image
flux-2-maxHighest-quality reference editing15/image
flux-kontext-maxPremium instruction-led image editing5/image

Video

Video pricing depends on duration, resolution, quality, and engine. Call estimate_generation_cost for an account-independent calculation.

  • foxcut-dop — camera-directed motion
  • minimax-h3-max — ultra-fast 480p/768p native-audio video with first/last frames
  • seedance-2.0 and seedance-2.5 — multi-shot continuity
  • kling-3.0 and kling-2.5 — photoreal motion and frame control
  • sora-2 and veo-3.1 — high-end world simulation
  • minimax-hailuo and wan-2.5 — fast or value-oriented motion

Talking avatar

  • foxcut-speak — native speech and lip sync
  • latentsync — fast audio-driven synchronization
  • kling-avatar — expressive full-scene performance