Freepik is now Magnific

    Wan 3.0 generates scenes, not shots.

    Thirty seconds of native footage with sound, in a single generation. Image, video and audio references in one call, and the whole Magnific stack around it.

    Inside the model

    One model. The whole shoot.

    No separate models for text, image, video or audio input. Alibaba Cloud's Wan 3.0 takes every reference through one call, sets up the whole take, and holds them all in place.

    2–30s

    Native footage from 2 to 30 seconds in a single generation. Set the exact length, or let intelligent duration read your material and decide.

    Universal referencing

    Up to 10 images, 5 clips and 5 audio tracks in one call, with pixel-level consistency across characters, props, sound and spatial layout.

    Text that stays text

    Interfaces, animations and on-screen type render as structured content, not as texture.

    Adaptive ratio

    Five fixed frames from 16:9 to 9:16, plus adaptive: the model reads your material and picks the framing itself.

    Sound, native by default.

    Audio renders in the same pass as the picture, in sync from frame one.

    4K

    Master in 4K with Magnific

    The model renders up to 1080p. When the take is right, Magnific's video upscaler masters it in 4K without leaving the suite.

    before image
    after image
    Before
    After

    From first reference to 4K master

    Wan 3.0 generates the scene. Magnific covers everything around it, before and after.

    Universal referencing

    Image, video, audio. One call.

    Feed Wan 3.0 up to 10 images, 5 clips and 5 audio tracks in a single generation, combined however the scene needs. Pixel-level consistency holds it together: the character keeps the face, the prop keeps its place, shot after shot.

    Duration

    Thirty seconds. One pass.

    Thirty seconds of native footage in a single generation, picture and sound together. Set the exact second you need, or let intelligent duration read the story and decide. The shot becomes a scene.

    On-screen text

    Interfaces, titles and type that hold.

    Software interfaces, animations and on-screen type render as structured content, readable in every frame.

    Native sound

    The take arrives with its soundtrack

    Sound here isn't a post step. Effects and music render in the same pass as the image, synced to what happens on screen. It's on by default. Switch it off when silence is the point.

    Magnific upscaler

    Finish in 4K, inside Magnific.

    Wan 3.0 renders native takes up to 1080p. Pick the winner, run it through the Magnific upscaler, and deliver in 4K without leaving the suite. Nothing gets re-shot.

    before image
    after image
    Before
    After

    From one prompt to a finished scene

    • Advertising

      A 30-second spot in one generation, with product, soundtrack and pacing locked from the first frame.

    • E-commerce

      Product heroes with the exact product. Reference the shots you already have and keep every detail in place.

    • Social

      Reels that arrive with their soundtrack. Effects and music included, no audio pass.

    • Film & Originals

      Micro-dramas and action scenes with room to breathe: a whole arc in a single generation.

    • Studios & agencies

      One reproducible Spaces workflow, packaged as an App, for output your whole team can repeat.

    New

    Block the camera in 3D

    Stage the move in 3D Scenes, then hand it to Wan 3.0 as one of its video references. You direct the camera before you spend a single generation.

    Frequently asked questions

    • Wan 3.0 is Alibaba Cloud's new all-in-one video model. No separate text-to-video, image-to-video or editing variants: a single call takes every input. On top of that, native 30-second generation, sound rendered with the picture, and universal referencing with pixel-level consistency.
    • From 2 to 30 seconds, natively. Intelligent duration can read your prompt and material and set the length for you.
    • Yes. Audio generates in the same pass as the video and is on by default. Switch it off when you want silence.
    • Up to 10 images, 5 video clips and 5 audio tracks in one call, combined however the scene needs, with pixel-level consistency across characters, props, sound and space.
    • 16:9, 4:3, 1:1, 3:4 and 9:16, plus adaptive: the model picks the framing from your intent and material.
    • Wan 3.0 generates up to 1080p. Inside Magnific you finish in 4K with the video upscaler, in the same place you built the references and picked the take.
    • Yes. Run Wan 3.0 from Claude, ChatGPT, Cursor or your own apps through Magnific's MCP and API, with your regular Magnific credits.

    Made with Wan 3.0

    Regional restrictions apply. Subject to Acceptable Use Policy.