Nº 090 · AI ·6 min read · August 02, 2026

MiniMax H3: when sound stopped being a separate department

Fig. 01 MiniMax H3: when sound stopped being a separate department

What the model does

On July 31, 2026, MiniMax released H3. Not a video model with audio layered on afterward. A model that generates both in the same forward pass, from the same prompt, at the same time. Fifteen seconds of 2K footage. Native stereo audio. One generation. No separate sound step.

The model accepts text, still images, video clips, and audio as input, and it handles image-to-video, text-to-video, first-and-last-frame animation, subject reference, and motion reference through a unified pretraining approach rather than a stack of separate tools. The API is live now under the model ID MiniMax-H3. Consumer access runs through the Hailuo AI app at around $0.13 per second. Open weights are promised for release within days.

The announcement covers capability claims and comparison tables, as these announcements do. What it does not address is the stranger question sitting underneath the technical specs. For the entire history of cinema as a professional practice, image and sound have been two departments. Two budgets. Two unions. A director of photography on one side, a sound designer on the other, with a director somewhere in between trying to make those two timelines mean the same thing.

H3 collapses that gap in a single generation pass. The image and the audio arrive together, as one object. And now you have to decide what you think about that.

What sound actually is

When synchronized sound arrived in cinema with The Jazz Singer in 1927, the industry called it talking pictures. The name reveals the misunderstanding. Nobody was waiting for pictures that talked. They were waiting to find out what pictures would do with the full apparatus of music, silence, texture, presence, and off-screen space that theater and radio had spent decades learning to control.

The first wave of sound films mostly wasted it. Studios built glass-walled soundproof booths and locked the cameras inside them, killing the visual grammar that silent cinema had taken thirty years to develop. Directors who had trained as composers of pure image found themselves hostage to a microphone. The screen became a stage, and not in a good way.

It took roughly a decade to work out what sound was actually for. Rouben Mamoulian's Applause (1929) was among the first films to separate the microphone from the camera: two tracks recorded simultaneously, cut against each other to create something neither could carry alone. That discovery, that sound and image could be edited independently and brought back into productive tension, is the foundation of everything that came after. Every scene where music contradicts what the camera is showing, where silence is louder than action, where you hear what a character cannot see: all of it flows from there.

Sound in cinema is not accompaniment. It is argument. It has its own point of view about what a scene is about, sometimes the same as the image, sometimes deliberately different, and the collision between them is where meaning lives.

H3 generates audio and image as one object. The audio is not making an argument about the image. It is statistically coherent with the image. Those are not the same thing.

What changes for working filmmakers

The gap H3 closes is real, and it is worth naming clearly before getting to the harder part.

  • A rough cut or proof of concept can now arrive with functional audio already in place. This changes how you communicate intent to collaborators, clients, and investors who need to hear the film, not just see it.
  • The distance between a compelling prototype and an actual production has narrowed. Expect the bar for what qualifies as a "persuasive demo" to shift accordingly, because the bar for everyone shifts at the same time.
  • At around $0.13 per second, fifteen seconds of 2K video with co-generated audio costs less than two dollars. The economics of early-stage experimentation are genuinely different from what they were last week.

These are real benefits. Use them. None of them have anything to do with what a sound designer does when they are doing their best work.

What the anxiety gets wrong

The common concern about H3 is that sound designers are next on the list. Some of them are, specifically the part of the job that consists of selecting ambient beds, matching room tone, and confirming that footsteps correspond to the floor material. That category of work was already compressible. The observation is not wrong. It is just pointing at something that was already moving before July 31.

The less common question is what H3 does to the pipeline that used to produce experienced sound designers in the first place. Everyone who does this work at a high level started somewhere that was not creative. They were the person who organized the session, labeled the takes, recorded the foley, did the cleanup pass at two in the morning. That is how you learn what audio is doing in a scene: not by studying it in the abstract, but by touching every element of it across hundreds of hours, for other people's projects, until the decisions stop being conscious.

When that pipeline compresses, the entry point changes. Not because the craft disappears, but because the path through execution to vision gets shorter, and the space that used to buy a young professional time to develop a point of view gets smaller. That is a different problem than "jobs are being automated." It is a structural problem about how craft knowledge actually transfers from one generation to the next.

H3 is not a threat to the sound designer who knows what a scene should hear. It is a problem for the pipeline that used to train that person.

Where this lands

Integrated generation, one model producing image and audio from a single prompt in one pass, is now the baseline. H3 is live. The capability exists. This is not a prediction about next year.

The useful question for filmmakers is not whether the audio H3 produces is convincing. In most cases, it will be. The useful question is what "convincing" actually costs when it arrives automatically, and what "intentional" costs when you have to reach for it separately, against a model that never stops generating plausible sound.

A sound designer working at full capacity is not making the audio convincing. They are making it mean something that the image alone does not. That requires a position. A point of view about what the scene is about at a level the director has not always named yet. The model does not have that. The model has patterns derived from the image in front of it.

This work did not belong to the model before H3. It does not belong to the model after it. But the margin for doing execution-level audio without a position, and calling it sound design, just got a lot thinner.

Know the difference between those two things well enough that you do not confuse them. That clarity is the work.

About the author

Read the manifesto Write in