The Prompt Bible
MiniMax H3 Video + Music 3.0 — The Complete Production Reference
Write it like a producer. Get it like a film.
What this is
The Prompt Bible is a single, complete reference for writing prompts that get what you want out of the two engines behind Solligence Café:
- MiniMax H3 — an open, general-purpose multimodal model that generates video with native stereo sound from text, image, video, and audio.
- MiniMax Music 3.0 — an open-weights music model that composes, arranges, performs, and produces a complete song in a single generation from a creative concept and optional lyrics.
It is the reference we use inside the Café, written down. If you've ever stared at a prompt box and thought "it looks almost right, but not quite," this is the book for you.
It is not:
- A beginner's "what is AI video" tutorial.
- A list of prompts to copy-paste without understanding. (There are 100+ working prompts in the appendices — but the value is in knowing why each one is built the way it is, so you can build the next one yourself.)
- A hype piece. Every claim about what these models do and don't do is traceable to MiniMax's official documentation.
Who it's for
| You… | Start with |
|---|---|
| Are a creator (Reels / Shorts / YouTube) who wants cinematic b-roll, beat-synced edits, and product shots without a crew | §1 → §4 → §10 → §11 |
| Run a brand or D2C store that wants ads and product films in hours, not weeks | §1 → §4 (use cases 1 & 4) → §11 |
| Make music and want to describe the exact record you hear in your head | §1 → §6 → §7 |
| Build products on top of H3 / Music 3.0 (APIs, workflows, apps) | §2 → §3 → §5 → §6 → §8 |
| Are a director/DP translating a shot list into a model | §1 → §3 → §4 → App C |
| Want the one-page version you keep pinned | Appendix A, right now |
How to use it
The 5-minute path. Read §1 (The Five Portable Laws). Do one worked example in §4. Render it in the Café. That's the whole loop the Café is built around — every render is an ad.
The 45-minute path. §1 → §3 (video syntax) or §6 (music syntax), depending on your output. Keep the relevant appendix table open beside you.
The full path. All of it, in order. The sections are cumulative: §1 gives you the mental model, §2–§4 make you fluent in H3, §5–§8 make you fluent in Music 3.0, and §9–§11 turn fluency into a system you can run every week.
Two rules for using the book:
- Render while you read. Every section ends with a prompt to paste. The model teaches you more in one render than ten pages of theory.
- The prompt is a wish list, not a contract. You will re-roll. That's the job. §1.5 explains the control dial that tells you when to add detail and when to let the model surprise you.
The promise (and the honest version of it)
The promise: After this book you will be able to sit down with any idea — a brand film, a product shot, a 15-second stylized short, a full song — and write a prompt that gets you 80% of the way there in the first render, and to the rest in two or three iterations.
The honest version: As of 2026, every major video model still struggles with the same four things: physics, multi-character consistency, hands, and on-screen text. H3 is at or near the front of the pack on text rendering, and its reference system is the strongest general-purpose tool we know for character consistency — but "at the front of the pack" is not "solved." §9 teaches you to diagnose which of the four you're hitting, and what to change. That diagnosis skill is worth more than any single prompt template, because models change and templates rot.
What you'll be able to do
By the end, you will be able to write prompts that:
- Control camera (move, speed, angle, shot size) with a working vocabulary — §3.6
- Direct lighting and color like a DP, including palette anchors for cross-shot consistency — §3.8
- Script motion that reads as physical, not floaty — §3.5, §9.2
- Cast a recurring character across many shots without them changing faces — §4.4
- Put correct, legible text on screen (H3's genuine strength) — §4.5
- Score a clip with native, diegetic, stereo sound that behaves like a sound designer's work — §3.11, §4.3
- Move a 768p render up to 2K with H3's own regeneration — §4.7
- Describe a complete song — genre, tempo, key, arc, vocal character, arrangement, mix — in the language Music 3.0 actually parses — §6
- Control section-by-section arrangement and vocal delivery across a five-minute song — §7.3
- Write lyrics that sing cleanly (syllable discipline, one-metaphor rule, narrative arc) — §7.9
- Combine both engines into a finished, platform-ready video + score package — §8
- Run all of the above on a repeatable weekly production system for the Café — §11
The one idea the whole book is built on
You are not typing words at a machine. You are briefing a very fast, very literal, very under-specified production team — a director, a DP, a sound designer, an arranger, and a mix engineer — in a single pass, in a language they have only half learned from training data.
- When you're vague, they improvise. (Sometimes beautifully. Often not what you wanted.)
- When you're precise, they execute.
- When you're precise in the right places and open in the others, they execute your vision while giving you their surprise.
Every rule in this book is a specific instance of that one idea. Learn the idea; the rules become obvious.
How this book relates to the Café
The Café is built on these two models. H3 runs at 720p-class output with native audio; Music 3.0 is the scoring engine. Everything in this book maps onto what you can do in the Café's studio today:
- Every worked example in §4 and §6 is a prompt you can paste in the Café and render.
- §11 is the Café's production system, written down: the 5-slot weekly cadence, the Founder Pass engine, and the compliance rules that keep every render an ad (UPI, ₹99/hr, 10 min free — said once, at the end, never in a price-war table).
If you're reading this as a customer, the book teaches you to get more out of your hours. If you're reading it as a maker, it's the operational manual for the thing you're selling. Either way: render something while you read.
Edition & licensing
- Edition: v1.0 — model-verified against MiniMax H3 (v1.1 prompt guides, 2026-08-08) and MiniMax Music 3.0 (launch, 2026-08).
- Model drift note: These are fast-moving models. The rules in this book are built on how diffusion/flow models and LLM-conditioned audio models work, which is durable. Specific numeric specs (durations, resolutions, price points) are accurate as of this edition's date and will be the first thing to change — check the Café studio's current options before betting a deadline on a number.
- Usage: Free on the web as the Café's pillar page. The full collection (all 100+ starter prompts, extended style playbooks) is available as a standalone PDF.
- Credit: This book distills MiniMax's official prompt-writing guides, the official API documentation, the Music 3.0 announcement, and the practitioner field guides of the Sora and Suno communities. MiniMax and OpenAI are trademarks of their respective owners; this book is an independent reference, not an endorsement. See Appendix H for the full source list.
Next: §1 — The Five Portable Laws.
§1 — The Five Portable Laws
Before we touch a single H3 or Music 3.0 prompt, we install the operating system. These five laws are not MiniMax-specific. We extracted them by reading how every serious prompting guide in this space — OpenAI's Sora 2 guide, the Suno community field guide, Runway's Gen-4 guide, Kling's formula, and MiniMax's own H3 and Music guides — actually tells you to think. Where all of them agree, you have a law. Learn the five laws and every specific rule in this book becomes a corollary.
1.1 Law One — The Briefing Law
You are not typing words at a machine. You are briefing a production team that is fast, literal, and under-specified — in a single pass.
When you submit a prompt, a large model does what a production crew would: it takes your description as direction, fills every gap with the statistically most likely completion of your training distribution, and executes. Three consequences follow, and they drive everything else:
- Gaps become improvisation. Anything you don't specify — time of day, outfit, lens, weather, the singer's age, the mix width — the model decides for you, by drawing from what's most common in its training data. OpenAI puts it plainly: "Unless you describe these details, the model will make them up." This is why two people can write the same prompt and get different results, and why your result may be beautiful and wrong.
- Literal beats implied. The crew has never seen your storyboard. If you want them to infer that "neon signs" implies "it's night and the street is wet," you're making them improvise the wetness. State what you want to see and hear. Sora's weak/strong table is the cleanest statement of this:
| You wrote (weak) | The model may infer | You should write (strong) |
|---|---|---|
| "a beautiful street at night" | …a street. Maybe. | "wet asphalt, zebra crosswalk, neon signs reflecting in puddles" |
| "moves quickly" | …movement of some kind | "cyclist pedals three times, brakes, and stops at the crosswalk" |
| "cinematic" | …whatever "cinematic" averaged to in training | "anamorphic 2.0× lens, shallow depth of field, volumetric light" |
- Repetition is the crew, not a bug. The same prompt generates different videos and different songs every time. Sora calls it a feature: "Treat your prompt as a creative wish list, not a contract." Iteration — re-rolling, nudging, pinning — is the normal production loop, not a failure of the model. Plan for it (§9.4).
The briefing law in one sentence: Specify what must be true; leave open what may surprise you; and expect to re-roll.
1.2 Law Two — The Container Law
The container is owned by parameters. The content is owned by the prompt. Never ask the prompt to do the container's job — and know the one big exception.
Every generation has two kinds of decisions:
- Container decisions — how long, what resolution, what aspect, which model, which reference assets. These are the shot's frame. In the Sora API they are explicit fields:
size: "1280x720",seconds: "8",characters: [id]. The Sora guide is emphatic: "These parameters are the video's container… your prompt controls everything else." Writing "make it longer" or "in 4K" in prose for a model whosesecondsfield is the only length control is talking past the crew. - Content decisions — subject, action, camera, light, style, sound, score. These live in prose.
The H3 exception — and why it matters for you. H3 is different from Sora on one point, and it's a feature of H3's design: the official H3 prompt formats take duration and aspect ratio in prose — the base format opens with "10 seconds, 16:9," and the reference format opens with "15 seconds, 9:16." H3's omni-representation reads your whole prompt (including those tokens) as semantic context for the generation, so in H3 you do brief the container in words. This is the single most common thing people copying Sora habits get wrong on H3.
So the working rule in the Café is:
| Decision | H3 video | Music 3.0 |
|---|---|---|
| Duration | In prose (10 seconds / 15s, 9:16) — must be an integer 4–15s | Not in prose; the model produces up to ~5 min, shape via sections |
| Resolution / aspect | Aspect in prose; resolution is a model/output setting | Output is 44.1kHz/256kbps MP3 via API settings |
| References (images/video/audio) | Attached as inputs; named in prose (Image 1 = …) | Lyrics + prompt fields; cover uses audio_url |
| Model / quality | Model choice (768P vs 2K output, regeneration) | model: music-3.0, is_instrumental, lyrics_optimizer |
Memory hook: On H3, the container gets a sentence. On Sora, the container gets a field. On Music 3.0, the container is mostly fixed — put your words where they bend the music.
1.3 Law Three — The One-Move Law
Motion is the hardest part of any generation. Give each subject one primary action and at most two secondary actions, each with a single direction — and one camera move per shot.
This is the rule with the most independent confirmations across every guide we read, because it's the one that most separates "beautiful accident" from "usable clip":
- Sora (OpenAI): "Movement is often the hardest part to get right, so keep it simple. Each shot should have one clear camera move and one clear subject action. Actions work best when described in beats or counts." Their canonical fix: "Actor walks across the room" → "Actor takes four steps to the window, pauses, and pulls the curtain in the final second."
- H3 (MiniMax official): formalizes it as a vector rule — one primary action + ≤ 2 secondary actions per subject; one vector per action — and a frame-delta rule: describe the change between frames, not an instantaneous state. "Slowly rises and hovers" is one vector (up, then rest). "Rises while trembling" is two competing vectors — the model must average a direction it doesn't have.
- The physics corollary (why the rule works): video models generate frames by predicting change. Objects in training have weight, momentum, and permanence. When you specify a state ("a ball in the air") the model can hold it or let it melt; when you specify a change ("the ball drops two feet and bounces once, losing height") you're giving it a physical trajectory to interpolate, which is exactly what it's good at.
Practical consequences you'll use constantly:
- Count beats. Steps, taps, blinks, sips. A 5-second clip fits roughly 3–5 beats of action. (§3.5 has the full method.)
- Never stack opposing vectors in one clause: not "rises and falls," not "zooms in and zooms out" (camera: cut between them), not "spins left and right."
- Hands are the frontier. Two simultaneous hand actions = two vectors on the model's least reliable region. One hand, one job, or rest the other hand still. (§9.2 diagnostic.)
- Stillness is a vector too. "The camera is static" and "she holds still, breathing" are instructions, and they work — static camera + one small action is the most reliable motion recipe in the entire space.
1.4 Law Four — The Early-Style Law
The medium and the style are the strongest lever you own. Set them first, in the first sentence, and let everything downstream inherit them.
Every vendor converges on this. Sora: "Style is one of the most powerful levers for guiding the model… Establish this style early so the model can carry it through consistently. The same details will read very differently depending on whether you call for a polished Hollywood drama, a handheld smartphone clip, or a grainy vintage commercial."
Why it works mechanistically: style tokens change the model's entire completion distribution. "16mm black-and-white documentary" doesn't add a filter at the end — it re-weights everything that follows: the grain, the focus, the camera grammar, even the kind of action the model finds plausible, because in its training data that style comes with a grammar of movement. H3's feature pages exploit this directly: every official example opens with the medium — "Live-action footage… fused with hand-drawn glowing animation" — before a single action appears.
The opening-sentence formula (video):
[Duration, aspect]. [Medium / style, in concrete production terms]. [Scene in one sentence].e.g. — 15 seconds, 9:16. Live-action kitchen at dusk, fused with hand-drawn glowing animation. A small robot in a black hoodie sings while the kitchen hums.
For music, the same law takes the form of the opening clause: genre + mood first, always, because the model builds the whole arrangement's identity from that clause and then elaborates. (A melancholic yet defiant pop-house song, 74 BPM… — see §6.2.)
Style in concrete terms is the skill to practice, and it's the same skill in both media:
| Abstract (weak) | Concrete (strong) |
|---|---|
| "cinematic" | "shot on 35mm, anamorphic, 2.39:1, warm tungsten practicals" |
| "cool" | "vaporwave, magenta-cyan grade, VHS tracking lines" |
| "Indian wedding" | "Hindi Bollywood drama, golden-hour open lawn, marigold tones, dhol-pipe score" |
| "nice music" | "lo-fi R&B, 72 BPM, vinyl crackle, brushed drums, close-miked breathy vocal" |
| "anime" | "2D cel anime, 2000s TV-broadcast look, flat-shaded light, film grain" |
Appendix C is the full dictionary: 40+ styles, each with a concrete production-language recipe and a working H3 opener and a Music 3.0 opener.
1.5 Law Five — The Dial Law
Detail is a dial, not a switch. Every element of a shot has an importance; set each one's detail level by how much it matters to the result.
This is the law that turns the other four from dogma into judgment. Sora states the tradeoff directly: "Detailed prompts give you control and consistency, while lighter prompts open space for creative outcomes. The right balance depends on your goals." And the Suno field guide quantifies the music version: tag count 3–5 for a simple song, 8–15 for detailed control, ~20 maximum — beyond this the model gets confused.
The practical instrument is a three-stop dial you apply element-by-element:
| Stop | When | What it looks like |
|---|---|---|
| Lock (full detail) | Must be true — brand text, the product, the character's face, the one action the idea depends on, the UPI-able hook | 1–3 precise sentences; the strongest, most concrete language you have |
| Steer (light detail) | Should lean a direction — mood, palette, energy | One clause each ("warm tungsten," "builds to a lift") |
| Open (omit) | Anything where surprise is desired — background life, incidental props, the model's own staging | Not mentioned at all. (Saying "the model's choice" out loud is worse than silence; it spends attention on a non-decision.) |
A locked example, annotated (a product shot — the Café's bread and butter):
10 seconds, 9:16. [STEER: medium] Live-action commercial film, shallow depth of field,
matte black studio. [LOCK: scene] A matte-black wireless earbud case stands on a
dark stone plinth, lid closed, single logo tile in brushed gold at front center.
[LOCK: action] The camera slowly pushes in; at the halfway point the lid opens
and lifts one centimeter, revealing the two earbuds. [STEER: light] One soft key
light from camera left, warm edge light from the right. [OPEN: nothing else]
Sound: near-silent studio; a soft mechanical click as the lid opens; a low hum
rises under the reveal.Count the stops: 2 locks, 2 steers, 1 open. The model has no freedom to change the product, the lid, or the logo — and total freedom over how the background breathes. That's the dial working.
Where beginners lose: they max-detail everything, then blame the model when the result is "rigid" — or they max-open everything, then blame it for being "random." Neither is the model failing. You set the dial. Diagnose before you re-roll: which element was supposed to be locked that drifted? Lock only that one, keep the rest, re-roll. That single habit — §9.4's pin-and-nudge protocol — is the difference between someone who "can't prompt" and someone who ships.
1.6 How the models actually read your words (the mechanism under the laws)
The laws are practical; this is why they're true. You don't need the math, but you need the shape of it, because the shape predicts behavior when a prompt misfires.
Video models (H3, and the diffusion/flow family). Modern video models are latent models: they don't generate pixels, they generate compressed "latent" frames, and a decoder (H3's H3-VAE) turns latents into 768p/2K pixels. Your text prompt is encoded by a language-model encoder into a semantic conditioning signal that steers the generation step by step. Three behaviors fall out of this architecture, and you'll feel all three while prompting:
- Conditioning is suggestive, not contractual. The text signal biases each denoising step; it doesn't force it. That's why the same prompt re-rolls to different results (Law One) and why an over-specific 900-character prompt can reduce coherence — you've crowded the signal with details the model can't jointly satisfy, and it averages. (This is the music side of the same fact: Suno's field guide measures prompt adherence decaying within a single generation — tightest in the first 30–60 seconds, then the model drifts back to its statistical defaults. Front-load what must be true. §7.2.)
- Salience follows concreteness and position. Concrete nouns out-pull abstract adjectives (the model's conditioning vectors for "wooden table" are more defined than for "furniture"); and in H3's format, the opening of the description and the structured fields (soundscape, music) are where attention is most reliably allocated — which is exactly why H3's official format is a form, not a free-text box.
- Relationships are the new frontier — and H3's answer. The hard thing in training data is not objects but relationships between elements (who's holding what, which sound belongs to which action, how two clips relate across a cut). H3's headline architectural contribution — the Contextual Omni Representation — is a dedicated system for exactly this: it represents not just "a robot" and "a kitchen" but the robot in the kitchen, the hum belonging to the fridge, the hand-print on the glass. That's why H3's native audio is stereo and diegetic-correct rather than a glued-on soundtrack, and why its reference system can keep a character consistent across shots. When you use H3's structured format, you're feeding that system in its native tongue. (§2.4 goes deeper.)
Music models (Music 3.0, and the LLM-conditioned audio family). Music 3.0 is an 8B language model over musical tokens. Your prompt is read by a language model in the literal sense: it parses your description into a structured caption — genre, tempo, key, emotional contour, instrument entry/exit, vocal delivery — and then samples a song trajectory token by token under that conditioning, with a flow-matching + neural audio decoder (the Flow-VAE) rendering the waveform. Consequences:
- It parses like a musician's brief, not like a keyword cloud. The official guide's first rule — "write prompts as vivid English sentences, not comma-separated tags" — is a direct consequence of the model being an LLM. Sentences have a syntax the LLM understands; tag clouds are exactly the disorganized input LLMs handle worst.
- Its "statistical gravity" is musical. Music LLMs are trained on the distribution of existing music, which means they have gravity wells: strong statistical attraction toward common structures (the Suno community names the pop gravity well — nearly every genre request drifts toward pop structure unless you counter it) and genre clouds (tags that co-occur so heavily they travel together:
trappulls808pullsbass;orchestralpullscinematicpullsepic). Music 3.0's Structured Caption training is MiniMax's answer to gravity: by conditioning on an explicit arrangement framework (what enters, what exits, what the vocal does in each section), the model is taught to hold your brief across a full five minutes instead of drifting home to the gravity well. (§7.1 is the full treatment, with the countermeasures.) - Continuity is the bottleneck, and it's a language-model problem. Keeping a song's identity coherent over 5 minutes is a long-range language modeling task — which is why Music 3.0's architecture splits an 8B global model (song structure) from a 0.6B local model (acoustic detail per frame) and fuses their hidden states into the audio decoder. Practical meaning for you: your section tags and section-level directions are doing real structural work.
[Chorus – with lift]is not decoration; it's a conditioning event. (§6.9, §7.3.)
The two mechanisms, one diagram in words:
VIDEO (H3): your prose ──► LLM encoder ──► semantic conditioning ──► latent video
│ diffusion/flow
└──► audio (stereo, native, jointly generated)
──► decoder (H3-VAE) ──► pixels+sound
MUSIC (M3.0): your sentence-brief ──► 8B Global LM (structure) ─┐
lyrics + tags ──► 0.6B Local LM (acoustics) ──┼─► fused hidden states
└─► flow matching ──► Flow-VAE ──► waveformNotice the shared root: in both engines, a language model is reading your words and steering a neural renderer. Everything you learn for one transfers to the other. That's why one book can cover both.
1.7 The five laws, condensed (pin this)
- Briefing — you're briefing a fast, literal, under-specified crew; gaps become their improvisation; re-rolling is the job.
- Container — parameters own the container; prose owns the content. H3 exception: duration + aspect go in prose.
- One-Move — one primary action, ≤2 secondary, one vector each; one camera move per shot; count beats; state change, not state.
- Early-Style — medium/style first, in concrete production language; it re-weights everything downstream.
- Dial — every element is Lock / Steer / Open; set by importance; diagnose drift by element, then pin-and-nudge.
Under all five: a language model reads your words; concreteness and position win salience; structure fights gravity.
Exercise — close this section by doing the book's core move
Take one idea you actually have (a product, a place, a character, a song you want to make). Write it three ways:
- Vague (5 words). Render it. Note what the model improvised.
- Maxed (300+ words, every detail). Render it. Note where it feels rigid or incoherent.
- Dialed (Lock/Steer/Open per element, ~80 words). Render it. Note what was true and what surprised you in a good way.
The gap between render 1 and render 3 is the whole book, in ten minutes. You now know where to put your attention for the next four sections.
Next: §2 — H3: what it is, and what that means for your prompts.
§2 — H3: What It Is, and What That Means for Your Prompts
Before syntax, geography. This section gives you the map of MiniMax H3 — what the model is, the three ways you can feed it, and the spec limits that shape every prompt you'll write in this book. Nothing here is speculation; it's all from MiniMax's official documentation and announcement.
2.1 The model in one paragraph
H3 is MiniMax's general-purpose, open-weights multimodal video model — the successor in the Hailuo line (Hailuo 01 → Hailuo 02 → H3). "General-purpose multimodal" is doing a lot of work in that sentence, so let's unpack it, because each adjective is a feature you can use in a prompt:
- General-purpose. One model, not a family of specialist forks. The same H3 does a 15-second brand film, a physics-y product drop, a stylized anime beat, a talking character, and a motion-graphics UI reveal. That's rare — most vendors split "photoreal model" and "animation model" into separate products. Practical meaning: you can change the style in prose without changing anything else, and you can plan a whole channel's visual language around one engine.
- Multimodal in, video+sound out. H3 accepts text, images, video, and audio as inputs in a unified context, and it outputs video with native stereo sound — the audio is generated jointly with the video, not added afterwards by a separate music tool. When a robot sings in an H3 clip, the model composed the singing as part of the scene, with the spatial behavior (panning, room, sync) that real production sound has. Practical meaning: one render = picture and a usable soundscape. Your edit gets a head start; your "Can AI do this?" tests get sound-on proof.
- Open weights. The model is published (GitHub + Hugging Face) — the same architecture that runs behind MiniMax's API and the Hailuo consumer app can be self-hosted. Practical meaning for the Café: our product is not hostage to a vendor's per-call API terms; for business customers, "open weights" also means data control and price independence — a genuine differentiator we can say out loud (and we do, in §11).
- Production posture. MiniMax markets H3 explicitly for production: film, advertising, branding, e-commerce, gaming — with a per-second price for 2K output claimed at under a third of mainstream closed models. The model is aimed at people shipping, and its documentation is written for them.
What H3 is not (to keep expectations honest):
- It is not a long-form engine. Clips run 4–15 seconds. Long pieces are edits of many H3 shots — the same way a film is cuts, not one take. (§8.4, §11.)
- It does not "follow a screenplay." It executes shot-level direction. You are the editor; H3 is the (very good) shot generator.
- Its weaknesses are the industry's weaknesses: physics edge cases, multi-character scenes, hands, and — its relative strength — on-screen text. H3 is currently at or near the front of the pack for text rendering; "front of the pack" is a real edge you can build product shots on, but §9 tells you how to verify it clip by clip.
2.2 The three generation modes
H3 has exactly three ways in. Know which one you're reaching for before you write a word — the mode determines the prompt format (§2.3), and choosing the wrong mode is the most expensive mistake in this book.
Mode 1 — Text-to-Video (T2VA)
Input: prompt only. Use when: you have an idea and no assets; you want the model to stage everything.
This is the "from scratch" mode. H3 will invent the subject, the scene, the camera, and the sound from your prose alone. It's the mode for:
- Fast ideation and "can AI do X?" tests (your cheapest creative partner)
- Scenes where nothing you describe is pre-existing — abstract visuals, environments, mood pieces
- B-roll you'll cut over other content
The cost: every element of the shot is a gap in your brief, and gaps become improvisation (Law One). T2VA renders are the most surprising — in both directions. Treat your first T2VA render as a mood board, lock what you like (re-roll, or capture the good frames as references), and drive from there.
Mode 2 — Image-to-Video (I2VA / First-Last-Frame)
Input: prompt + 0, 1, or 2 images (first frame and/or last frame). Use when: you have frames — a design, a still render, a photo, an approved look — and want motion through them.
Two sub-flavors:
- First-frame (I2VA): one image = the opening frame; the model animates from it. The classic "bring my still to life" move. The image is the composition, the lighting, the character design — your prose only has to say what happens.
- First-last (FL2VA / L2VA): two images = the bookends; the model must invent the middle. This is the most underused, highest-leverage mode in H3. You control the state change of the entire clip — the before and after are locked, the how is the model's choreography. Perfect for: product reveals (closed → open), morphs (sketch → finished), a character's two poses with a motion between, transitions between two designed frames.
Specs that matter: images 256–5,760 px on each side, aspect ratio between 2:5 and 5:2, JPG/PNG/WEBP/HEIC/HEIF, ≤ 30 MB each. (Zero images = you're just in T2VA again.)
The framing insight: in first-last work, write the middle explicitly — the official FL2VA examples do exactly this: "…switching smoothly from a bright orange knit hat to a black wide-brim hat, with a bright yellow background throughout." The words that matter most in FL2VA describe what changes, and what stays put.
Mode 3 — Reference Generation (Ref2VA)
Input: prompt + up to 9 reference images, 3 reference video clips, 3 reference audio clips (≤ 12 files total). Use when: you have assets with roles — a character design, a brand's motion language, a voice, a song — and want the model to keep them consistent while doing new things.
This is H3's headline capability and the reason a one-person studio can keep a recurring character (a mascot, a host, a product) consistent across a whole series. Each reference asset is assigned a role in your prompt — subject, environment, costume, prop, style, voice, performance, motion, camera, or editing rhythm — and the model's job is to carry that asset's defining features through the generated video.
Specs that matter:
- Images: ≤ 9, each 256–5,760 px, ≤ 30 MB
- Video clips: ≤ 3, each 2–15 s, total ≤ 15 s, H.264/H.265 (in-video audio AAC/MP3)
- Audio clips: ≤ 3, each 2–15 s, total ≤ 15 s, WAV/MP3
- Mixed input: ≤ 12 files total
Practical meaning: your reference budget is roughly nine images, fifteen seconds of video, fifteen seconds of audio per render. Plan your asset sheet around that: a character (2–3 angles), the environment (1–2), the costume (1–2), a style frame (1), a voice clip (1), a music stem (1) = a full production kit that fits.
The three modes, one decision table:
| You have… | You want… | Mode | Format |
|---|---|---|---|
| An idea only | Everything generated | T2VA | base |
| A still / a frame | Motion from it | I2VA | base |
| Two frames | The journey between them | FL2VA/L2VA | base |
| A character / brand asset / voice / motion language | New scenes, same identity | Ref2VA | ref (6 sections) |
2.3 The two prompt families
H3's official documentation defines two prompt formats, and which one you use is decided by your mode:
| Family | Used by | Required sections |
|---|---|---|
| Base prompt | T2VA, I2VA, FL2VA, L2VA | integrated_multimodal_description + overall_soundscape + non_diegetic_music |
| Reference prompt (ref) | Ref2VA | subject_definitions + summary + retention_analysis + detailed_description + overall_soundscape + non_diegetic_music |
The ref format is the base format plus three reference-management sections at the front: a definitions inventory (what each asset is, what its role is), a summary (what the shot is, driven by the references), and a retention analysis (explicitly: which features of each reference to preserve, and how to avoid conflicts between references). That last section is a formal acknowledgment of the hardest problem in reference-based generation — two references can contradict each other — and giving the model an explicit conflict-resolution brief is a genuinely sophisticated tool.
Do not mix the families. Base format must never mention reference assets ("the robot from the image" is a ref-format sentence; in base format the robot is simply described). Ref format must define every asset you attach. §3 is the full syntax of both; §4.4 is the ref format in deep practice.
2.4 The architecture, in the parts that change how you prompt
You don't need the papers. You need the four components, translated into prompting consequences.
(a) Contextual Omni Representation — the "relationship engine"
H3's core representational move. The model builds a unified semantic representation of everything in context — text, images, video, audio — and, critically, the relationships among them: which sound belongs to which on-screen event, how a reference character relates to a new scene, how two clips relate across a cut. (The launch details: the training pipeline distills up to ~100K tokens of source material into ~4K working context — the model is built to compress context into instruction.)
Prompting consequence: H3 rewards relational prompting. Instead of listing objects and hoping, state the relations: "the hum belongs to the fridge, under the robot's singing"; "the hand-print on the glass is the man's, from a second before." Sound-to-source binding, character-to-scene binding, and cut-to-cut binding are H3's home turf. This is also why H3's native audio is spatially correct — the sound is generated inside the relationship graph, not mixed on afterwards.
(b) H3-VAE — the rebuilt tokenizer
A high-compression latent tokenizer that gives the model 4× the effective sequence length — which is what makes native 2K possible without a bolt-on upscaler.
Prompting consequence: detail density is your friend within a shot. H3 can hold more fine-grained texture (small text, fine pattern, multiple small props) than a short-context model could — but Law Five still applies; detail must be earned by importance, not sprayed.
(c) H3-Omni Transformer — one backbone, trained for both jobs
A single transformer trained with separate workloads for understanding (reading your inputs) and generation (writing video+audio), which is the efficiency engine behind the unified API.
Prompting consequence: the same model that reads your references also directs the shot — so how clearly you name roles (Image 1 = the character; Image 2 = the environment) directly affects execution. Ambiguous role-assignment is understood less well than explicit role-assignment, even if the model "sees" the assets.
(d) In-Context Regeneration — the 2K path
H3's 2K output is produced by the base model regenerating its own 768p output in-context — the model re-renders its own result at higher resolution, recovering fine detail (including small text) that a pixel-domain upscaler would only guess at. Via the API this is the Video Regeneration task: you resubmit the same content plus your 768p clip as role: base_video.
Prompting consequence (a full workflow, §4.7): the cheap-and-fast path is 768p for iteration, 2K only for the keeper. You can afford 5–10 fast 768p rolls to find the shot, then pay the 2K cost once, on the winner — with the original prompt intact. This is the H3-native version of "dailies, then the final."
The Context-IR exception: letting the model write the prompt
There's one API task — H3-Context-IR — that inverts the whole book: you hand it your raw materials (text/images/video/audio) and it returns an enhanced, structured prompt (the six-section form), which you then feed to generation. It "deeply interprets multimodal context… and produces a structured representation with richer semantic detail while preserving the user's original intent."
How we use it in the Café: as the first draft, not the author. Run Context-IR on your references → read its output → edit it with the rules in §3 (it's a competent junior writer; it will get relationships right and dial-levels wrong) → generate. You get the model's own comprehension of your assets, corrected by a human eye. (§4.8.)
2.5 The spec sheet (your hard constraints)
Memorize the shape; look up the numbers when it matters. These are the API's hard limits as of this edition:
| Parameter | H3 value | Prompting note |
|---|---|---|
| Output resolution | 768P / 2K | 2K via Regeneration (§4.7) |
| Duration | 4–15 seconds, integers only | In prose on H3: 10 seconds. No 3s, no 6.5s. |
| Aspect ratio | common ratios + adaptive | In prose: 9:16 / 16:9 / 1:1… |
| Prompt length | ≤ 7,000 characters | You will never hit this; a disciplined shot is 300–900 |
| First/last-frame images | 0–2; 256–5,760 px; AR 2:5–5:2; ≤ 30 MB | JPG/PNG/WEBP/HEIC |
| Reference images | ≤ 9; ≤ 30 MB each | |
| Reference video | ≤ 3 clips; 2–15 s each; ≤ 15 s total; ≤ 50 MB | H.264/HEVC |
| Reference audio | ≤ 3 clips; 2–15 s each; ≤ 15 s total; ≤ 15 MB | WAV/MP3 |
| Mixed total | ≤ 12 files | |
| API flow | async: create → poll → download | The Café handles polling; you get files |
| Priced access | pay-as-you-go API | The Café is where you pay ₹99/hr instead |
Three constraints that shape how you write (not just what):
- The 15-second reference ceiling means your reference video material must be a 15-second summary of your intent — the best 15 seconds of motion/style/pace you have, not a whole scene. Trim references like an edit.
- The 12-file ceiling means a kit, not a wardrobe. Pick the 6–9 references that carry the shot; the rest are in your archive for the next render.
- Integer seconds, 4–15 means your idea must fit a 4–15s beat, or it's a sequence of H3 shots cut together — the only way H3 makes "long." Design the sequence in §8.4; the cuts are the length.
2.6 Where H3 sits (so you can answer the obvious question)
The honest landscape in 2026: several frontier video models (Sora, Veo, Runway, Kling, Seedance, and H3) each hold an edge somewhere. The pattern: photoreal physics and native audio are converging across all of them; the differentiators are (a) reference systems for consistency, (b) text rendering, (c) open weights / self-hosting, and (d) price per second of usable output. H3's claimed edges are exactly (a)–(d) — its open multimodal context, its reference grammar, its text rendering, its open weights, and its sub-third 2K pricing.
We teach H3 in this book because it's the engine the Café runs and because its documented prompting grammar (the base/ref formats) is the most explicit of any major model — explicit enough that "prompting H3 well" is a learnable craft with a defined shape, rather than a vibes exercise. The Five Laws (§1) transfer to any model; everything from §3 on is H3-native. If you also use Sora or others, §1 and the appendices will keep you fluent there too.
Next: §3 — H3 Video Prompting: the complete syntax of the base and reference formats.
§3 — H3 Video Prompting: The Complete Syntax
This is the reference section. Every rule here comes from MiniMax's official H3 prompt-writing guides (v1.1) and API documentation; where a rule has a reason, the reason is stated, because reasons survive model updates and rules rot.
Read it twice the first time: once for the shape of the two formats, once for the individual rules.
3.1 The base format (T2VA / I2VA / FL2VA / L2VA)
A base prompt is three named sections, in this order:
integrated_multimodal_description
[the shot: subject, scene, action, camera, light — plain prose]
overall_soundscape
Sound: …
Dialogue: …
non_diegetic_music
Music: …Three standing rules about the format itself:
- The section names are load-bearing. Use exactly
integrated_multimodal_description,overall_soundscape,non_diegetic_music. These names are how H3's omni-representation routes your text into the right part of the generation (visual description vs sound vs score). integrated_multimodal_descriptionis the only free-prose section. It's where Law Four (style first) and Law Three (one move) live.- The two sound sections use key-value lines, one fact per line:
Sound: …/Dialogue: …/Music: …. Not paragraphs. The model parses lines. - Omit
non_diegetic_musicentirely if you don't want background music. Its presence is a request; its absence is the (better) default of "natural sound only." - Never mention reference assets in the base format. If you attached an image, the base format doesn't know that — describe what the image shows. (Mentioning "the image" belongs to the ref format, §3.2.) If you're using references, switch formats — don't improvise a hybrid.
The skeleton, annotated
integrated_multimodal_description
10 seconds, 9:16. ← LAW 2 (H3 exception): duration + aspect, first
Cinematic live-action, warm tungsten, ← LAW 4: medium/style in production terms
shallow depth of field.
A young woman in a mustard cardigan ← subject: identity features only (§3.4)
stands at a rain-streaked bus window, ← scene: 1–2 concrete nouns
looking out. Rain runs down the glass. ← LAW 3: one action, one vector (the rain)
She lifts her hand and rests it on the ← LAW 3: her one primary action
glass, palm flat, fingers spread.
The camera is static. ← LAW 3: camera = one move (none)
overall_soundscape
Sound: city traffic murmur, rain on glass in the foreground, a distant bus horn.
Sound: a soft exhale as her hand settles.
(no Dialogue line — she speaks nothing)
(omit non_diegetic_music)That's the entire base format: one prose block, then sound lines. 90% of H3 work you'll do is this shape with different words.
3.2 The ref format (Ref2VA)
When you attach reference assets, the prompt gains three sections at the front:
subject_definitions
Image 1: A small round robot, black hoodie with a red-lettered chest pocket, large glowing blue eyes. Role: subject.
Image 2: A cozy kitchen at dusk, warm lighting, wooden shelves. Role: environment.
Voice 1: A cheerful synthetic voice, mid-pitch, slight robot cadence.
summary
A 15-second 9:16 clip. The robot (Image 1) sings in the kitchen (Image 2) while
bobbing its head. Camera static at medium shot. Audio is the robot's voice (Voice 1)
over a light electronic beat.
retention_analysis
Image 1: preserve the hoodie, the red chest lettering, and the eye glow exactly; the
robot's body proportions must stay identical to the reference in every frame.
Image 2: preserve the warm dusk lighting and shelf layout; new props may appear but
must match the warm wooden palette.
Voice 1: preserve the mid-pitch synthetic timbre; the singing must stay in that voice,
not drift to a human-sounding one.
detailed_description
15 seconds, 9:16. The robot stands center-frame in the kitchen, facing camera,
singing while it bobs its head and taps its feet. A light electronic beat plays
underneath, entering at the second line of the chorus.
The camera is static; medium shot.
overall_soundscape
Sound: kitchen ambience, low fridge hum, the beat.
Dialogue: (robot, singing) "Loading loading, find your mind and your voice"
non_diegetic_music
Music: light electronic beat, cheerful, steady, enters under the chorus and continues.The three reference sections, one by one
subject_definitions — the inventory. One line per asset: Asset N: [what it is]. Role: [role].
- The description must be accurate to the asset — the model cross-checks your words against the actual image/video/audio, and a description that drifts from the asset creates a conflict the model has to average away.
- Roles must match usage. The role vocabulary is:
subject · environment · costume · prop · style · voice · performance · motion · camera · editing rhythm · sequence. A video reference used for its choreography is amotionorperformancereference — naming itsubjectwhen it's really a movement reference mis-routes the model's attention. - One role per asset. If an image should be both the character and the style, that's two references (or a decision: pick the one that matters more and say so in
retention_analysis). - Cap: ≤ 5 definitions per media type (images / videos / audio) — consistent with the 12-file input ceiling, it forces kit thinking (§2.5).
summary — the shot, in 2–4 sentences. Subject + scene + camera + audio, only insofar as the references drive it. No embellishment — this section is the model's "what am I making" statement. Omit it only if, bizarrely, you attached references you won't actually use.
retention_analysis — the conflict contract. One line per reference, each stating what to preserve, how it appears in the output, and how it coexists with the other references. This is the section that makes or breaks multi-reference shots. The official guides require relational phrasing: consistent with / different from / independent of. And a hard rule: never promise behavior the reference doesn't contain. If your character reference shows a standing robot, do not write "preserve the robot's running gait" — there is no gait in the asset to preserve, and the model will improvise one.
3.3 The one-sentence-one-fact rule (and the full syntax code)
H3's most important micro rule, because it's the one whose violation you can feel immediately:
One sentence = one key fact. Avoid paragraph dumps. Avoid more than two dependent clauses per sentence. Avoid stacked keywords (more than ~5 in a row). Avoid chained time adverbs ("first… then… then… finally").
Why: recall §1.6 — your prose becomes a conditioning signal, and signal-to-noise matters. A 60-word sentence with four facts is four faint signals; four 15-word sentences are four clear ones. The official guides are unusually explicit about this because H3's context compression rewards discrete, attributable facts.
The syntax code (post this next to your screen):
| Rule | Do | Don't |
|---|---|---|
| One fact per sentence | "The lamp is off. A match flares." | "The lamp, which is off, flares as he strikes a match that lights it" |
| Concrete nouns | "wooden table", "yellow taxi" | "furniture", "a vehicle" |
| Minimal adjectives | "a slow walk" | "a beautiful, graceful, elegant, serene walk" |
| Distance | one of near / mid / far | "deeply, tightly, in the far distance" |
| Identity | face, hair, clothing (visible, essential) | personality words, backstory, mood |
| Clips | 300–900 chars, 15–30 sentences | 3,000-char essays |
The 300–900 character working band. H3's hard limit is 7,000 characters; the quality band is far smaller. Everything you need for a 10-second shot fits in ~800. When your base description passes 1,500 characters, you're not prompting — you're writing a treatment, and the model is averaging your anxieties. Cut to the facts; move the rest to your storyboard notes.
3.4 Describing subjects (and the anti-mood rule)
For every character or object in the shot, you have a short identity card:
- People: face shape/features, hair, and the 1–3 wardrobe items that define them. That's the whole card — 4–5 items maximum.
- Objects: material + one distinguishing feature. ("a matte-black earbud case with a gold logo tile")
- The anti-mood rule: never describe a face with mood words. "A sad face" is not a visual feature the model can render consistently — and worse, it steals the expression's job from the action. Mood belongs in: the action ("she stares at the letter, shoulders dropping"), the sound (a sigh), or the score (the
non_diegetic_musicline). The face gets features; the scene gets mood. - Lighting consistency across all characters in one shot: if two people share a frame, they share a light. Name the light once, at shot level, and don't give character A warm light and character B cool light in the same sentence — that's a conflict the model resolves by blending (and the blend reads as "AI").
3.5 Motion: vectors, beats, and the frame-delta rule
The complete motion doctrine, consolidated from the official rules + §1.3:
- Count your actions per subject: 1 primary + at most 2 secondary. ("The robot bobs its head" is primary; "taps its feet, eyes flickering" are the two secondaries. A third — "and waves" — is where clips start to smear.)
- One vector per action. An action has a direction and a profile: up-then-rest, in-then-stop, open-then-hold. "Rises and hovers" ✓. "Rises while trembling" ✗ (two vectors: rise + oscillate). If two motions genuinely co-occur (a character walking and talking), that's one vector of locomotion plus an audio event, not two visual vectors.
- Describe change between frames, not instantaneous states. The model interpolates change. "The glass is full" is a state; "water rises to the brim over three seconds" is a change. Same for blur, focus, occlusion: what is changing, and how fast.
- Count beats to fit the clock. A 5-second clip ≈ 3–5 beats; 10s ≈ 5–8; 15s ≈ 8–12. A beat = one perceptible action unit (a step, a sip, a blink). More beats than the clock allows = the model compresses them into a rush, and the rush is where physics goes to die.
- Stillness is a legal vector. "The camera is static" / "he holds still, breathing" — specify stillness when stillness is the point; don't assume the model knows you want it.
- Hands: one hand, one job (§9.2 for the diagnostic). When both hands must act, make one hold (static, easy) and one do (the vector).
3.6 Camera language (the full working dictionary)
The combination rule first (it's a hard rule in the official guide): one compound move ("slowly pushes in") or two simple moves ("push in, pan left") — never opposing moves in one sentence. "Zooms in and zooms out" is impossible; make it a cut between two shots.
Moves
| Move | What it does | Typical phrase |
|---|---|---|
| push in / dolly in | closer over time; intimacy, emphasis | "the camera slowly pushes in" |
| pull back / dolly out | wider over time; reveal, exit | "pulls back to reveal the room" |
| pan | horizontal sweep | "pans left across the street" |
| tilt | vertical sweep | "tilts up from the shoes to the face" |
| track / follow | moves with the subject | "tracks her from the door to the window" |
| orbit / arc | circles the subject | "arcs around the car, 90°" |
| crane / pedestal | vertical travel of the whole frame | "cranes up over the crowd" |
| static | no move — a choice, say it | "the camera is static" |
| handheld | micro-shake; documentary feel | "handheld, slight shake" |
Speed words (modify any move)
slowly · gently · swiftly · smoothly · abruptly — one speed word per move. ("slowly and quickly" is the motion version of a conflicting vector.)
Short-form tags (API convenience for T2V)
The video-generation API documents bracket tags appended to key descriptions: [pan] [zoom] [static]. Use them when you want the minimum possible camera instruction; use full phrases when camera is a protagonist of the shot.
Framing (shot size) — one explicit tag, placed at the cut
| Tag | Use for |
|---|---|
| close-up | face/hands; emotion, detail, text-on-skin |
| medium shot | waist-up; dialogue, gesture, most "character" shots |
| wide / establishing | place, scale, context; opening shots |
Placement rule: state shot size once at the start for a single-shot clip; with multiple shots, state it at each new camera setup (with the cut, §3.9). Don't restate it mid-shot — restating reads as a change.
3.7 Composition and depth
- Depth of field is a steering word: "shallow depth of field, sharp on her, background soft" is one clause that does more for a shot's look than five adjectives.
- Three-layer framing (foreground / midground / background) is the Sora-confirmed way to get a composed frame rather than a flat one — and it's exactly the relationship H3's omni-rep is built for (§2.4a): "Foreground: a coffee cup on the bench. Midground: the traveler, mid-frame. Background: the train, soft focus."
- Negative composition works, sparingly: "no text, no logos in the background" is one clean instruction; long "avoid" lists are attention spent on absences (and H3, unlike some models, has no dedicated negative-prompt field — omission and one short negation are your tools, §9.3).
3.8 Lighting: source + direction + temperature
The complete lighting sentence has at most three clauses: source, direction, color temperature — plus optionally one atmosphere sentence (mist, smoke, rain) as its own line.
| Slot | Options (mix one each) |
|---|---|
| Source | window / key light / practicals (lamps, screens) / sun / moon / candles / neon |
| Direction | from camera left / right / above / behind (rim) / low angle |
| Temperature | warm (tungsten, ~3200K feel) / cool (daylight, ~5600K) / mixed |
| Quality (optional) | soft (diffuse) / hard (direct) |
Worked pairings (each is a legal 3-clause lighting line):
Soft window light from camera left, warm tungsten fill from a desk lamp.
Hard noon sun from above, cool shadows, a thin rim from the open door.
Neon signage from the right — magenta key, teal bounce off wet asphalt.The palette-anchor upgrade (from Sora, works on H3): name 3–5 colors that the shot is allowed to use — "palette: amber, cream, walnut brown" — and cross-shot consistency stops being a gamble. For a series (same character, many clips), the palette anchor is part of your character card: same face, same wardrobe, same light, same palette = the model has four reasons to render the same world.
Cross-shot lighting rule: a location keeps its light logic across every shot that returns to it. The "same room, different light per cut" problem is the #1 cause of edits that feel wrong without being able to say why.
3.9 Scene transitions and multi-shot structure
H3 can hold multiple shots in one prompt — and its official guidance is specific about how.
The in-prose transition vocabulary
(cut) (cut to) (cut back to) (dissolve) / (fade in) (fade out) (montage) (scene shift: from A to B)
Selection logic:
(dissolve)— same world, time passing, or a soft beat change. The official rule: sequential events in one continuous space are described as a crossfade/soft dissolve.(cut)— a hard beat; distinct actions, or a new camera setup within the same space.(cut to)/(scene shift: from A to B)— a location change. The official rule: hard cuts are for distinct locations; one location-change per transition, maximum.(montage)— a quick sequence of related beats (a "day in the life" block); use for 3+ micro-shots.
Two legal ordering strategies
- Spatial-first (preferred for most work): each shot block = location → subject → action. The model reads each block as a self-contained setup. This is the shape of the FL2VA and multi-shot official examples.
- Chronological (for stories with time passing): each shot needs a time anchor ("at dawn", "by noon") and character-state carry-over ("she is now in the grey coat, hair wet"). This is the harder strategy — use it when time itself is the subject.
The budget: a 15-second clip with 3 shots = ~4–5 seconds each; that's 1 action + 1 camera move + 1 light per shot, minimums per §3.5–3.8. Four or more shots in one 15s clip is where H3 starts averaging. For more than 3 shots, generate the shots separately and cut them (the §8.4 sequence workflow) — this is also where "long content" actually comes from.
3.10 Dialogue: the counterintuitive rules
H3 generates voiced dialogue natively, and the official rules for it are strict — because speech is the highest-bandwidth element in the shot and the easiest to get wrong:
- Only lines actually spoken inside the clip's time window. A 10s clip fits a handful of short lines; a monologue that needs 40 seconds doesn't fit a 10-second clip, no matter how you ask.
- Every line is explicitly attributed:
Speaker says: "line"— includingoff-screen:for voices without a visible speaker. Unattributed lines are a coin flip about who speaks. - The on-screen (nearest-camera) speaker says ≤ 2 lines total. This is the rule nobody expects and everybody needs: the model handles off-screen speech better than heavy on-screen speech. Structure for it — a presenter getting two lines while a narrator (off-screen) carries the rest is both better-executed and better-made.
- No quotation-mark nesting, no
x3shorthand. If a line repeats, write the repetition out. (Music 3.0's lyric side does use(typ-typ-typical)ad-lib repetition — video dialogue does not; the two engines parse repetition differently.) - Original language is preserved. H3 speaks the language of the line you wrote; it translates only if you ask it to.
- Speech goes in
Dialogue:lines underoverall_soundscape— never buried in the visual description. And any non-dialogue vocal sound (singing, humming, sighs, laughs) is aSound:line, not dialogue — the routing matters for how the voice is rendered.
Worked dialogue block (2 on-screen + 1 off-screen, the sweet spot):
overall_soundscape
Sound: café ambience, cups on saucers, low murmur.
Dialogue: Barista says: "Your flat white, right here."
Dialogue: Customer says off-screen: "Give me a second, I'm finishing this."3.11 `overall_soundscape`: writing like a sound designer
The soundscape section is where H3's native-audio design really shows, and the official rules read like a sound-designer's checklist:
- Open neutral. The first
Sound:line states the environment factually: "Sound: a quiet sushi stall at dusk, low ambient chatter." No emotion words, no music words — the first line is the room, not the feeling. - Environment: 1–2 contextual sounds, max 3 total, each with one spatial modifier. "quiet traffic in the distance" — the modifier is the where/how-far, not an intensity ("loud traffic" is a banned construction; distance is how sound-designers set level).
- Time-synced SFX: one event per line, source object explicit, cause→action→sound order. "Sound: a wooden chair scrapes on tile as the character leans back." One modifier, no unseen sources, and foreground SFX must belong to the main on-screen action — if the sound's source isn't on screen, the model may attach it to the wrong visible event (the omni-rep's relationship engine works for you if the source is stated).
- Foreground vs background is a real mix concept — the official food-demo example does it explicitly: "ambient chatter in the background; a soft bump from the chef's movements; the singing bowl in the foreground." You can mix in prose:
in the foreground/in the background/underare working words. - Background music does NOT go here. It goes in
non_diegetic_music(§3.12). Putting score in the soundscape blurs diegetic/non-diegetic and the mix gets muddy.
A full soundscape, layered the official way:
overall_soundscape
Sound: a quiet workshop at night; rain on a skylight in the background; a clock ticking in the distance.
Sound: the robot's small footsteps on concrete — foreground, with each step.
Sound: a low mechanical whir as the robot turns its head.
Dialogue: (robot) "Almost lost it… but I got it!"3.12 `non_diegetic_music`: the one-sentence score
The official rules for the score section are deliberately small — this is the section where less is more:
- Only include it when the shot calls for score (or when you explicitly want music). Otherwise: omit the section. Natural sound only is the default, and it's the correct default for most product and documentary-style work.
- One sentence: genre/era + emotion + (optionally) one behavior detail.
Music: lo-fi hip-hop, warm and understated, plays softly under the dialogue.Music: 80s synth, tense, swells at the reveal.Music: Indian classical — soft sarod and tabla, contemplative.
- What the one sentence may NOT contain: instrument lists, BPM numbers, or arrangement detail. Those words belong to a music production (see §6 — where they are exactly what you do write, for the Music 3.0 engine). In a video prompt, the score is dressed, not produced: two or three words of genre, one word of emotion, one clause of behavior.
- Multiple sections of music (a scene shift that changes the score): one sentence per section, with the follow-up tied to the shift:
Music: (new, after the cut to the station) synth pulse, urgent.
The division of labor to internalize:
| Want… | Write it in… | With… |
|---|---|---|
| Diegetic sound (room, SFX, voices) | overall_soundscape | source + sync + foreground/background |
| Non-diegetic score | non_diegetic_music | genre + emotion + one behavior, one sentence |
| A real composed track, produced | Music 3.0 (§6) | full structured caption |
This is the hinge of the whole book: H3 dresses a scene with sound; Music 3.0 is the production. When your shot needs a score with arrangement and structure, you compose it in Music 3.0 and either (a) use it as a reference audio (≤ 15 s of it, role: the music bed) or (b) lay it over the finished clip in your edit. (§8.2.)
3.13 Character consistency across shots (the character-slot system)
For any recurring character (your mascot, your host, the protagonist of a series), the official continuity method is the character slot:
- First appearance = full definition. The first shot that shows the character gets the complete identity card (§3.4): face, hair, the 2–3 defining wardrobe items. This is the anchor.
- Every later mention = the same slot, verbatim. Later shots reference the character with the identical feature set — same words, same order. Do not "vary it up." The model's consistency comes from repetition of the definition, and in reference work, from the same reference image every time.
- Changes must be caused. If the character gets a scar, loses a glove, or swaps coats, the narrative must contain the change ("after the fight, his left sleeve is torn") — an uncaused difference is read as inconsistency, and inconsistency is exactly what you're trying to prevent.
- Carry state at transitions: pose, posture, and gaze continue across a dissolve; a
(cut)to a new setup is your license to re-block the character. - The reference upgrade: in Ref2VA work, the character is the reference image (role: subject) plus the slot text. The image carries the face; the text carries the wardrobe and the rules. That pair — stable image + stable slot text — is the closest thing to a "character asset" in text-to-video, and it's the backbone of §4.4.
3.14 Audio hand-off across shots
Sound, like character, has a continuity rule the official guide states explicitly — first mention = full description; carry-over in the same scene; re-explain on scene change.
Concretely:
- A sound bed that persists across cuts inside one scene: introduce it fully on first appearance, then a short carry line —
Sound: (continued) the station PA, distant. - A new scene: re-establish its environment sound in full. Don't assume the model carries the world's acoustics across a location change — it won't, and the edit will feel like a dub.
- Persisting background elements across a cut get the official marker:
Sound: (continued from previous shot) …
Why this matters commercially: your Reels will be cut from many H3 shots. If each shot's soundscape is re-estably briefed, your edit is a mix — rooms change, beds continue, SFX sync to action — and the whole thing sounds produced. If you don't brief sound per shot, your edit is ten clips with ten random ambiences, and no amount of DAW work fixes a dub that was wrong at the source.
3.15 The complete base prompt — master template
Everything above, assembled. Copy this; fill the brackets; it's the shape of every H3 base prompt in this book:
integrated_multimodal_description
[Duration integer 4–15] seconds, [aspect]. [Medium/style — concrete production terms].
[Scene: 1–2 sentences, concrete nouns; three layers if composed].
[Subject identity card: features, hair, wardrobe — 4–5 items max].
[Lighting: source + direction + temperature (≤3 clauses); atmosphere line if any].
[Action: 1 primary + ≤2 secondary, each one vector; beats counted to the clock].
[Camera: one move + speed, or static; shot-size tag].
[Transitions, only if multi-shot: (cut)/(dissolve)/(scene shift) with per-shot setup].
overall_soundscape
Sound: [environment, 1–2 sounds, neutral, spatial modifiers].
Sound: [time-synced SFX — source explicit, cause→action→sound; fg/bg marked].
Dialogue: [Speaker says: "…"] ← only lines that fit the window; ≤2 for on-screen speaker
non_diegetic_music ← omit entirely if no score wanted
Music: [genre/era + emotion + one behavior detail].Pre-flight checklist (30 seconds, before every render):
- [ ] Duration is an integer 4–15, stated first?
- [ ] Style/medium in sentence one?
- [ ] Every subject has a feature card (no mood words on faces)?
- [ ] Light: one source, one direction, one temperature?
- [ ] Actions ≤ 3 per subject, one vector each; beats fit the seconds?
- [ ] One camera move (or "static")?
- [ ] Sound section: neutral first line; SFX synced to visible sources; music out of it?
- [ ] Dialogue attributed; on-screen speaker ≤ 2 lines?
- [ ] No paragraph sentences; no keyword stacks; no "first…then…then…"?
- [ ] Every element you care about is locked; the rest is genuinely open?
If all ten are yes: render. Then §9 tells you how to read the result.
Next: §4 — H3 in practice: the six use cases, with full worked prompts.
§4 — H3 in Practice: Six Use Cases, Fully Worked
MiniMax organizes H3's world into six production use cases, and so do we. For each: the job it does, the prompt shape that wins it, a complete, paste-ready prompt, a why-it-works teardown, and the variations that make it yours. All prompts follow the §3 syntax exactly — read one teardown and the rest become fast.
Use Case 1 — Brand Films & Cinematic Content
The job: the 10–15-second brand film. Trailers, product hero moments, "who we are" statements. The shot that makes a brand look like a film, not a slide.
The winning shape: one location, one subject (the product or a human relationship with it), one reveal, static or slow-push camera, hard-to-soft light, diegetic sound doing the heavy lifting (the brand should sound designed, not scored over).
Worked example — "The Earbud Reveal" (10s, 9:16)
integrated_multimodal_description
10 seconds, 9:16. Cinematic live-action product film, matte black studio,
shallow depth of field, one soft key light from camera left with a warm rim
from camera right. Palette: matte black, brushed gold, warm amber.
A matte-black wireless earbud case stands centered on a dark stone plinth,
lid closed, a brushed-gold logo tile at front center. The camera is static,
medium close-up, framing the case from chest height of the plinth.
At the halfway point the lid opens and lifts one centimeter, revealing
the two earbuds inside, then holds.
Dust motes drift slowly through the key light.
overall_soundscape
Sound: near-silent studio room tone in the background.
Sound: a soft mechanical click as the lid opens — foreground.
Sound: a low, warm hum rises under the reveal and settles.
non_diegetic_music
Music: minimal electronic, warm and understated, a slow swell under the reveal,
ending with the hum.Why it works (teardown):
- Law 4 in line 1: "Cinematic live-action product film" sets the entire distribution before a single object appears.
- Lock/Steer/Open (§1.5): the product, the lid, the logo = locked (3 precise sentences). Light = steered (one 3-clause line + palette anchors). Dust motes = steered, one vector. Everything else = open — the studio's background life is the model's to breathe.
- One move: the lid's open-and-hold is one vector with a rest — the single most reliable motion shape in this book.
- Sound does the reveal: the click is time-synced to the visible cause (lid opens → click), foreground-marked; the hum is the emotional layer. The one-sentence score (§3.12) does not fight the diegetic design — it arrives under it.
- Palette anchors (black/gold/amber) make this clip cut next to any other clip in the same series without a color pass.
Variations that make it yours:
- 16:9 version: same prompt,
16:9, and add a foreground layer ("a marble counter edge in the lower frame, soft focus") for the horizontal composition. - The anti-hero: swap the studio for a real place — "a monsoon-wet metro station platform, 07:10 AM, sodium practicals" — and the same one-reveal structure becomes a brand world instead of a product shot.
- Two beats instead of one: a 15s version = lid opens (beat 1) → earbud lifts out of the case on nothing, hover (beat 2 — note: floating product shots are a beloved format; the vector is "rises two centimeters, hovers").
Use Case 2 — Visual Creative & Content Packaging
The job: the effect — the stylized, patterned, texture-led short that stops a scroll. Trend content, aesthetic b-roll, motion-graphics-adjacent visuals, the "wait, what is this" clip.
The winning shape: the medium is the subject. The idea is a transformation or a pattern, the style is extreme and specific, and the sound is textural (rhythmic, synthetic, toy-like) — this is where H3's stereo design shines, because the whole clip is the texture.
Worked example — "Neon Market, Hand-Drawn Ghosts" (15s, 9:16)
integrated_multimodal_description
15 seconds, 9:16. Live-action night market footage fused with hand-drawn
glowing animation, 2D line-work over live footage, deep blue base grade
with neon magenta and cyan accents.
A crowded covered market at night, string lights overhead, stall lights
in the midground. The camera tracks slowly left at shoulder height through
the crowd. (dissolve) Hand-drawn glowing figures — simple white line
drawings of people — drift between the real shoppers, bobbing gently,
leaving short glowing trails that fade. The trails weave in the foreground.
Rain begins to fall on the open edge of the market; real rain, drawn
figures stay dry.
overall_soundscape
Sound: market ambience in the background — vendors, distant traffic,
crowd murmur.
Sound: each drawn figure makes a soft electronic hum as it drifts —
the hums pan left to right with the figures, in the foreground.
Sound: rain enters at the open edge, spreading across the frame.
non_diegetic_music
Music: dark ambient with a slow pulse, synth and sub-bass, curious and
slightly uncanny, building gently with the rain.Teardown highlights:
- The fusion line is the whole concept in one clause: "live-action… fused with hand-drawn glowing animation" — this is the H3 Feature-Highlights pattern verbatim (style-first, medium-as-subject).
- Relationship prompting (§2.4a): "the hums pan with the figures" — the sound-to-source binding is explicit, which is exactly what the omni-rep executes.
- The one-rule world: "real rain, drawn figures stay dry" — a single contrast rule that makes the magic legible. (The audience's eye needs one consistent law of the world; you gave it one.)
- Dissolve, not cut: the transition into the drawn layer is a
(dissolve)— same location, a change in reality, which is the dissolve's job (§3.9).
Variations: the "documentary world + one impossible thing" structure is a template. Swap the market for: a library (glowing book-ghosts), a gym (drawn sweat-drops), an office (line-drawn coffee steam that writes). Each render is a different "Can AI do this?" post, same engine, same 15s clock. (This is literally the IG "hand-drawn × live" channel theme in marketing/strategy/CHANNEL-THEMES.md — the prompt above is its production spec.)
Use Case 3 — AI Narrative Content (short drama / vertical series)
The job: the 15-second scene with a story inside it — the vertical short-drama beat (the ReelShort/DramaBox format). Character-driven, dialogue-forward, serial.
The winning shape: two characters max, one location per clip, the §3.10 dialogue rules (on-screen ≤ 2 lines, narrator carries the rest), the §3.13 character slot for continuity, and one emotional beat per 15 seconds — not a plot point, a feeling (tension, a secret, a decision). The plot accumulates across episodes, which are clips you cut together (§8.4).
Worked example — "The Verandah" (15s, 9:16)
subject_definitions
Image 1: A woman in her late 20s, dark hair in a loose braid, maroon kurta. Role: subject.
Image 2: A woman in her 50s, silver-streaked hair, teal sari with a gold border. Role: subject.
Image 3: A brick verandah at dusk, brass kettle on a small table, warm window light. Role: environment.
summary
A 15-second 9:16 vertical drama scene. The younger woman (Image 1) sits facing
her mother (Image 2) on the verandah (Image 3). The younger woman delivers one
line, then her mother answers. Medium two-shot, static camera. Warm dusk light.
retention_analysis
Image 1: preserve the face, the braid, and the maroon kurta exactly in every frame.
Image 2: preserve the silver-streaked hair and the teal sari; her posture stays
upright and still in every frame.
Image 3: preserve the dusk warmth and the brass kettle's position; new props may
appear only if they match the warm palette.
detailed_description
15 seconds, 9:16. The two women sit across the small table on the verandah,
the brass kettle between them. Warm window light from camera left, soft shadows.
The camera is static, medium two-shot, both women from the waist up, the younger
in the left of frame, the mother on the right.
The younger woman looks down at the kettle, then up at her mother.
(dissolve) The mother's hand rests on the younger woman's hand on the table.
Palette: maroon, teal, brass, dusk orange.
overall_soundscape
Sound: dusk on a street in the background — distant crickets, one radio voice.
Sound: the kettle ticks softly as it cools — foreground, near the table.
Dialogue: Daughter says: "I'm not asking for permission. I'm telling you."
Dialogue: Mother says off-screen: "Then sit, and tell me while it's still warm."Teardown highlights:
- Reference discipline: three references, each with a role and a retention line that names the exact features to keep — the two faces are the series' identity, and the retention contract says so in writing.
- The dialogue structure is the scene: one on-screen line (the daughter — she's nearest camera, one line ≤ 2 ✓) plus the mother's answer as off-screen — the official "off-screen speaks more" rule, which also feels more natural (we hear the mother's voice while seeing the daughter's face react — that reaction is the shot).
- One emotional beat: the hand-on-hand
(dissolve)is the whole scene's turn. Nothing else moves. A 15-second drama scene is a feeling with a trigger, and this prompt is two sentences of feeling. - Kettle as time-keeper: "ticks as it cools" = a diegetic clock the audience subconsciously trusts.
The series engine: every episode = same two subject_definitions lines (verbatim) + same retention_analysis + a new 15-second beat. The model gets the same identity contract every time — that is what makes a one-person studio able to run a character series. (The "Can AI [X]?" YT channel and the vertical-drama ICP both run on exactly this.)
Use Case 4 — Product & E-Commerce
The job: the product film that sells in 5–15 seconds without a voiceover script — performance-marketing creative, D2C launches, ad creative that can be re-rolled per audience.
The winning shape: the product is the only actor. One action (the use), one environment that implies the life the product enters, camera that serves the product (not a personality), and sound built from the product's own physics — the click, the pour, the slide. H3's text-rendering strength makes this the use case where on-screen copy is part of the ad (§4.5).
Worked example — "The Pour" (5s, 9:16 — the shortest legal H3 clip)
integrated_multimodal_description
5 seconds, 9:16. Live-action beverage commercial, warm morning kitchen,
soft window light from camera left, shallow depth of field.
Palette: cream, tea-amber, warm wood.
A glass of amber tea on a wooden counter, steam rising from it. A ceramic
mug enters frame from the right and the tea is poured from the kettle held
just above it — one steady pour, the stream unbroken, filling to two-thirds.
The camera slowly pushes in on the glass during the pour.
overall_soundscape
Sound: a quiet kitchen in the background, one window with birdsong.
Sound: the pour — a soft steady stream in the foreground, from the first
second to the fourth.
Sound: the kettle sets down with a soft ceramic clack on the counter.
non_diegetic_music
(omitted — natural sound only)Teardown: 5 seconds = H3's minimum = the ad shape. One action (the pour) with a physical detail that makes it real ("the stream unbroken" — the model must keep the stream continuous; that constraint is what stops the classic melting-liquid failure). The clack is the product's own sound design — no score needed, and omitting non_diegetic_music is the confident choice: the product sounds like itself.
- The D2C multiplier: this one prompt is your week of ad creative — re-roll 5 times, cut the best 3, caption each differently (the §11 compliance tail on each). H3's cost-per-second makes 10 rolls a rounding error against a stock-footage license.
Use Case 5 — Digital Experience & Game Creative
The job: UI motion, interaction demos, feature reveals, game trailers — screens and systems as film.
The winning shape: the interface is set in a physical world (the model renders UI-in-space far better than abstract motion graphics), camera moves with the user's hand (a hand or a pointer is your character), and the sound is the UI's own — taps, whooshes, confirmations — which H3 renders as design-led SFX rather than film sound.
Worked example — "The Checkout, Zero Friction" (10s, 16:9)
integrated_multimodal_description
10 seconds, 16:9. Premium product film, a smartphone floating in a soft
grey studio void, screen fully lit, the interface crisp and legible.
One soft overhead key light, gentle reflection under the phone.
The phone fills the frame in a medium shot, camera static. On screen: a
one-step checkout card. A hand enters from the bottom of frame and taps
the card — the card expands with a smooth 0.3s ease, showing a UPI symbol
and a "Pay" button. The thumb taps Pay. The screen fills with a soft
green confirmation check that draws itself in.
The hand exits; the phone drifts up one centimeter and holds.
overall_soundscape
Sound: silent studio in the background.
Sound: a soft tactile tap as the finger lands — foreground.
Sound: a light UI whoosh as the card expands.
Sound: a two-note confirmation chime as the check draws.Teardown: the UI is the subject with a feature card (§3.4: material + distinguishing feature = "a one-step checkout card"); the hand is the character (one hand, one job — §3.5); and the sound section is a sound-design spec for the interface itself — each SFX line is time-synced to its visible cause, which is the only way UI demos stop feeling like foley guesswork. (The UPI symbol on screen is also your §4.5 text test — verify it renders cleanly; H3 is your best bet in the market, but verify per roll.)
Use Case 6 — Animation & Stylized Visuals
The job: the non-live-action channel — 2D, 3D-stylized, papercraft, stop-motion-feel, anime, "the brand as a drawing." The Café's own visual language lives here, and it's where a small studio out-looks any live-action competitor, because style is a moat live-action can't cross.
The winning shape: name the medium with production specificity (Law 4, maximized), keep the world simple (stylized worlds survive the model's physics weakness — a papercraft diorama has fewer physical promises than a live kitchen), and let the style's grammar supply the motion (stop-motion moves in ticks; cel anime moves in held frames; 3D-stylized moves in smooth arcs).
Worked example — "The Paper City" (15s, 9:16)
integrated_multimodal_description
15 seconds, 9:16. Cut-paper papercraft stop-motion aesthetic — layered
paper diorama, visible paper edges and tiny punched details, matte pastel
palette (dusty blue, cream, coral), soft even studio light from above with
gentle shadows between the layers.
A tiny paper city in three depth layers: rooftops in the background,
streets with a paper train in the midground, a paper balcony in the
foreground. The train travels left to right along the midground track
in slow, slightly jerky stop-motion ticks. Steam puffs rise from its
chimney as small folded-paper puffs that unfold and fall.
The camera is static, framed so all three layers are visible.
overall_soundscape
Sound: a quiet studio in the background.
Sound: the paper train's wheels tick — tick tick tick — in the foreground,
in time with its motion.
Sound: each steam puff makes a soft paper-creak as it unfolds.
Sound: one small paper door in the foreground layer opens and closes once,
mid-clip, with a paper flap sound.
non_diegetic_music
Music: music-box and soft glockenspiel, slow and wistful, underneath.Teardown: the style supplies the motion grammar — "slow, slightly jerky stop-motion ticks" is both a style statement and the motion vector (Law 3 + Law 4 fused; the ticks are one vector: advance + jitter). The three-layer diorama is the §3.7 composition doing double duty as style (papercraft = layers) and safety (simple geometry = fewer physics promises). And the sound design is the medium's own physics: paper creaks, wheels tick — the audience's ear confirms the eye's material. This is the exact prompt shape behind the Café's papercraft-explainer skill: change the city to your topic's metaphor, keep the grammar, new render.
4.5 (interlude) — On-screen text & brand: H3's genuine edge
Among the four hard problems of 2026 video models (physics, multi-character, hands, text), H3 is at or near the front of the pack on text — its launch materials cite accurate text and brand rendering as a headline capability, and its in-context 2K regeneration specifically recovers small text that upscalers guess at.
What that means practically:
- Put the words where the brand needs them, and say so.
A white sign reads "SOLLIGENCE CAFÉ" in a clean sans serif— H3 will render the actual letters. Brief the text like a prop: what it says, where it sits, what it's made of. - Keep on-screen text short. One word to a short phrase, large, high-contrast. Multi-line paragraphs are still a coin toss for everyone in 2026 — including the leaders. (Long-form text belongs in your edit, as a graphics layer — always a fallback.)
- Verify per roll. This is a statistical strength, not a guarantee. The diagnostic: zoom to 2K (the regeneration pass, §4.7) and check the letters. If the logo came out as "S0LL1GENCE," that roll is B-roll, not hero. §9.2 has the full text-check.
- The 2K path is your text-quality lever. Small brand text is exactly what in-context regeneration recovers — another reason the 768p-iterate → 2K-commit workflow (§4.7) is the workflow for any shot with type.
4.6 Physics & motion realism: what to promise, what not to
The industry-wide hard list (physics / multi-character / hands / text) has promises you can reliably make on H3 and traps you should route around:
Reliable (promise these in a brief):
- One subject, one primary action, gravity-consistent (a cup falls down; the official frame-delta rule is exactly the tool for it)
- Rigid bodies with one vector (doors, lids, drawers — the §4.1 earbud lid is the archetype)
- Fluid with a continuity constraint stated ("the stream unbroken", "fills to two-thirds" — the constraint is what the model enforces)
- Cloth/flags/hair as secondary motion (one primary action + "her scarf lifts" = the legal 1+1)
- Camera moves (the model's best region — camera is learned motion, not physical)
Trap list (route around, don't fight):
- Two simultaneous hand actions with fine detail → one hand acts, one holds (§3.5)
- Two characters interacting (handing, touching) → prefer the near-miss (hands approach; the cut lands the contact) or one character + an object
- Liquid + cloth + glass all in one frame → one material hero per shot
- Sustained running/walking across a changing background → prefer in-place or short-distance; long tracking shots are where legs drift
- Mirrors, reflections of moving subjects, glass with pass-through → the model must double-simulate; budget a re-roll or route around
The universal fix when physics breaks: reduce the state change the model must invent. Smaller distance, slower speed, one material, a rest at the end of the vector ("…and holds"). §9.2 turns this into a diagnostic table you apply in 30 seconds.
4.7 The 2K workflow: dailies, then the final
H3's production pipeline, from the API: 768P for iteration → 2K Regeneration for the keeper.
- Roll at 768P. Fast, cheap, same model. Run 5–10 rolls of the shot. You are choosing performance (the action, the timing, the sound sync) — at 768p you can judge all of it.
- Pick one keeper. Judge on: did the locked elements stay true (§1.5)? Did the one move land? Does the sound sync to the visible cause?
- Regenerate to 2K. Via the API's Video Regeneration task: the same
content(your prompt + inputs) + the 768p clip asrole: base_video. The base model re-renders its own output in-context at 2K — recovering fine detail (texture, small text, edges) that a pixel upscaler would only hallucinate. - Verify at 2K, then stop. The 2K pass is your final quality gate — check text, check the one move, deliver. Do not re-iterate at 2K; that's where budget goes to die. If the 2K reveals a problem, fix the prompt and re-roll at 768p — the iteration layer is always 768p.
Why this is a strategy, not a detail: your cost model is now "ten cheap dailies + one expensive final" — the same economics as a real shoot (dailies are cheap; the final grade is one pass). It's also why H3's pricing posture (2K at under a third of mainstream per-second) is a real argument the Café can make: a 2K finished shot for less than the mainstream models charge for a 720p one. (§11 says it exactly this way — once, at the close.)
4.8 H3-Context-IR: the model drafts, you edit
The last H3 tool inverts authorship. The H3-Context-IR API task takes your raw multimodal inputs (your references, a rough idea) and returns an enhanced structured prompt — the six-section form, written by the model's own comprehension of your assets.
The Café's usage pattern (junior-writer protocol):
- Feed it the kit: your character image, environment frame, voice clip, and a one-sentence intent ("the robot sings the chorus in the kitchen").
- Read its draft like an editor. It will be strong on relationships (it just interpreted your assets — that's its job) and weak on dial-levels (it tends to over-describe; it doesn't yet know which of your facts matter most).
- Apply the §3 code: cut to one-fact sentences, enforce 1+2 actions, check the one-move rule, lock/steer/open each element, enforce the ≤2-line on-screen dialogue rule, split diegetic/non-diegetic.
- Generate. You get the model's faithful reading of your assets plus a human's editorial judgment — the best of both authorship models.
Use Context-IR for: character-heavy Ref2VA shots (where the asset-relationship reading is the whole value), new style explorations (let it propose the medium's vocabulary, then correct), and any shot where you'd otherwise stare at a blank ref-format template.
Section drill — the five prompts, rendered in your head
Before §5 (Music 3.0), do the section's exercise: pick one use case above that matches something you actually want to make, paste its prompt into the Café, render it, and apply the §3.15 pre-flight checklist to your own version. One render. Then read §5 — you'll arrive with dirt under your fingernails, which is exactly the state the music sections assume.
Next: §5 — Music 3.0: what it is, and the one concept that runs the whole engine.
§5 — Music 3.0: What It Is, and the One Concept That Runs the Whole Engine
The video half of this book took six sections because H3 has two prompt families and three input modes. The music half is leaner: Music 3.0 has one input shape (a sentence-brief + optional lyrics) and one output (a complete song). But it has one central concept that the rest of the video book doesn't — and once you see it, "how do I prompt a song" stops being a mystery.
5.1 The model in one paragraph
Music 3.0 is MiniMax's next-generation, open-weights music model. Given a creative concept and optional lyrics, it composes, arranges, performs, and produces a complete song in a single generation — up to five minutes long, as a finished 44.1kHz/256kbps mix.
Four adjectives that are features you use in prompts:
- Complete-song, single generation. Not a 30-second loop you extend six times; the model plans the whole song — intro through outro — in one pass. This is the direct answer to the industry's oldest pain (the "great 30 seconds, unusable rest" problem), and it changes your workflow from assembly to direction.
- Open weights. Like H3, Music 3.0 is published (Hugging Face: MiniMaxAI/MiniMax-Music3). Same business meaning as §2.1: self-hostable, data-control-friendly, price-independent. (§11 says this out loud for business customers.)
- Performing, not synthesizing. MiniMax's own framing: the goal is vocals that sound performed rather than synthesized and instruments with physical realism — the model is being steered toward the sound of people playing, which is the lever §7.6 pulls ("physical recording reality").
- Intent-first. The headline claim: more accurate interpretation of creative intent — the model is trained to hold your brief across the full song instead of drifting to what music usually sounds like (§7.1's gravity wells, and the counter).
The three documented upgrades, and what each one unlocks for you:
| Upgrade (MiniMax's words) | What it means for prompting |
|---|---|
| More accurate interpretation of creative intent (Structured Captions) | Your description can specify what changes and where — emotion, instrumentation, vocal delivery per section — and the model executes the arc instead of averaging it. → §6.6, §7.3 |
| More complete, more varied arrangements (Hybrid-LM, global+local) | Section tags in your lyrics are real structural events the model respects — bridges, instrumental breaks, a final-chorus lift. → §6.9, §7.3 |
| Clearer, more natural sound (fused hidden states → flow matching → Flow-VAE) | Vocal performance is directable: breath, falsetto, phrasing, harmonies; instrumental technique is directable: glissando, legato, bowing. → §6.4, §7.4 |
What it is not:
- It is not a stems engine. You get one finished mix per generation. (Stems/mixing happen in your DAW, if you want them — or you re-generate with a different arrangement brief.)
- It is not a remix engine in the sample-crate sense. The closest tools are covers (re-perform an existing song's lyrics in a new style, §6.8) and the H3 reference-audio path (§8.2).
- Its ceiling is ~5 minutes per generation. Longer sets = multiple songs, cut together (a tracklist, not a track — §8.4).
5.2 The architecture, in the parts that change how you prompt
Three components, each with a prompting consequence. (Full technical detail: MiniMax's launch post; the practical translation is here.)
(a) The tokenizer — 8 layers of RVQ, structure first
An 8-layer residual vector quantizer represents the music: layer 1 = core semantics and structure; layers 2–8 = progressively finer acoustic detail. Training is staged: the structural layer is taught first, then all layers jointly.
Prompting consequence: structure is the backbone of the model's attention. What you write about form (sections, the arc, who enters when) lands on the layer the model learned first and most stably; acoustic flavor words land on the residual layers, which do more with less. This is why a 20-word structural brief plus a 40-word flavor sentence outperforms a 60-word undifferentiated paragraph. (§7.1 makes this into a writing rule.)
(b) The Hybrid-LM — an 8B global mind + a 0.6B local ear
An 8B Global LLM (initialized from Qwen3.5-8B) predicts the semantic music tokens frame by frame with full-song context; a 0.6B Local LLM predicts the acoustic tokens within each frame. They train in two stages — global alignment first, then joint — so the global model holds the song together while the local model dresses each moment.
Prompting consequence: long-range instructions are real. "The chorus should feel twice as big as the verse" is a global instruction the 8B model is built to carry; "the vocal breathes more in the bridge" is a local one. You can therefore write a song brief the way an arranger works: global statement first (what the song is, what it does overall), then section events (what changes, where). This is the direct ancestor of the §6.6 canonical pattern.
(c) The synthesis stack — fused hidden states → flow matching → Flow-VAE
At inference, Music 3.0 does not decode discrete tokens into audio. It takes the continuous hidden states of both language models, fuses them, and feeds a 2.4B flow-matching module that maps into a 123M Flow-VAE decoder (inherited from MiniMax's speech architecture, retrained for music) to render the waveform.
Prompting consequence: the performance layer is richer than a token decoder would allow — and it's where vocal technique and instrumental technique words get their purchase: "expressive vibrato," "slightly aged timbre," "glissando," "legato," "brushed drums" are not decoration on this stack; they condition the continuous representation that becomes the sound. Say the technique you want performed, not just the instrument you want present. (§6.4, §7.4.)
The one concept: the Structured Caption
Here is the idea the whole music half of this book is built on — the reason Music 3.0 can hold a five-minute brief:
MiniMax trained Music 3.0 against Structured Captions: a fine-grained, temporal description framework in which a song is specified as (1) a global identity — genre, tempo, time signature, key, use case, production character — and (2) a timeline of changes — emotional contour, entry/exit of instruments, groove and low-end energy, and section-level vocal delivery, harmony, and effects.
Two things follow:
- Your prompt is parsed toward that framework. When you write a sentence-brief, the model maps it onto the caption structure it learned: your genre word becomes the global identity; your "builds into an uplifting chorus" becomes a temporal instruction; your vocal description becomes a section-level delivery spec. The model's "native language" for music is the structured caption — so the more your prose mirrors that shape (global → arc → sections → performance), the less translation loss.
- There's a built-in amplifier for non-experts. MiniMax also built a template-based Prompt Enhancement System: curated caption templates that expand a simple description into a detailed, musically coherent one, using proper arrangement terminology. In practice you don't call the enhancer directly through the Café — you imitate it. The canonical demo prompts in §6.3 are what the enhancer writes. Learn to write in that shape and you're using the system by hand.
The two engines, one root — and the one difference: H3 and Music 3.0 are the same architecture in spirit (an LLM reading your brief, steering a neural renderer — §1.6). The difference is what the brief is for: H3's brief directs a scene (spatial, seconds-long, one shot's worth of events); Music 3.0's brief directs a work (temporal, minutes-long, an arc of events). A video prompt is a still life with motion; a music prompt is a journey. That single difference explains every music-specific rule that follows.
5.3 The API surface (what the Café actually calls)
The practical surface, from the official docs — the parameters, and which one is your lever:
`POST /v1/music_generation` — the main tool
| Parameter | Values / shape | Your lever? |
|---|---|---|
model | music-3.0, music-2.6, free variants, music-cover | — (the Café picks) |
prompt | the style brief — free text, sentence form | ✅ this is the prompt |
lyrics | your lyrics with section tags (10–1,000 chars in cover flow) | ✅ half the creative control |
lyrics_optimizer | true = model writes the lyrics from your brief | when you have no words yet |
is_instrumental | true = no vocals, pure arrangement | the instrumental switch |
audio_url / audio_base64 | the source for covers (mutually exclusive with cover_feature_id) | the cover path |
cover_feature_id | from the preprocess step (below); pair with modified lyrics | the two-step cover |
audio_setting | { sample_rate: 44100, bitrate: 256000, format: "mp3" } | container, fixed |
output_format | url (or file) | container |
The prompt + the lyrics are the two creative levers. Everything else is container. (Law Two, music version: on Music 3.0, the container is mostly fixed — put your words where they bend the music.)
The two supporting tools
POST /v1/lyrics_generation—mode: "write_full_song",prompt: <your style brief>→ returns complete lyrics with proper section structure (Verse/Chorus/Bridge). This is the "I have the sound, write me the words" half of the workflow — the same brief you'll use for the music, aimed at the lyricist. (If you write your own lyrics, skip it.)POST /v1/music_cover_preprocess— feed it an existing song'saudio_url(free step): returns acover_feature_id(valid 24h) +formatted_lyrics(ASR-extracted, section-tagged, editable) +structure_result(segment types + timestamps) + duration. Then generate withcover_feature_id+ your modified lyrics + a new style prompt = a cover with new words and/or new style, same song skeleton.
The two-step cover, end to end (the workflow in §6.8)
your song.mp3 ──► preprocess ──► { cover_feature_id, formatted_lyrics, structure }
│
▼ (you edit the words / keep them)
generate { model: music-cover, cover_feature_id,
lyrics: edited, prompt: new style brief }
│
▼
new performance of the (edited) song5.4 The business note (read this once, then it changes how you think about the product)
From the official Music docs, verbatim:
*Starting August 20, 2026, the paid APIs (Music Generation and Lyrics Generation) will no longer be available to new users… The free music generation APIs (Music-3.0-free, Music-2.6-free, music-cover-free) will be discontinued. To experience or use music generation capabilities, please visit MiniMax Audio, or use the open-source MiniMax Music 3 model on Hugging Face.*
Translation for the Café: the product path for Music 3.0 is open weights, self-hosted — the same pattern as H3 (which already runs self-hosted in the Café's Cloud Run ComfyUI setup). This is not a problem; it's the same story as H3, and it's a selling point the Café can say out loud:
"Both engines behind the Café — video and music — are open-weight models we run ourselves. Your briefs, your renders, your data. No per-call API dependence, no vendor terms changing under you. You pay by the hour, in rupees, by UPI."
(That sentence lives in §11, where it earns its one allowed appearance. It is simplicity language — "your studio, your terms" — not a superiority claim, which is why it's legal under our positioning rules.)
For the bible's readers: if you build on Music 3.0, know that the durable path is the open-weights model (GPU + inference serving) or the MiniMax Audio app — plan your product's music dependency on open weights, not on a paid API's continued existence.
5.5 The spec sheet (your hard constraints)
| Parameter | Music 3.0 value | Prompting note |
|---|---|---|
| Max duration | ~5 minutes per generation | Long beyond that = a set of songs (§8.4) |
| Output | 44.1kHz / 256kbps MP3 (API default) | Container; fixed |
| Vocals | full singing, spoken, duets, choirs — all directable | is_instrumental: true to switch off |
| Lyrics | optional; section-tagged; lyrics_optimizer if none | the other creative lever |
| Covers | one-step (audio in) or two-step (preprocess + edit) | the re-performance path |
| Language | English prompts work best; other-language scene phrases OK for flavor | official guide, §6.1 |
| Structure control | section tags + modifiers are real events | §6.9, §7.3 |
Three constraints that shape how you write:
- One generation = one mix. You don't get stems; you get a record. If you need a stem, that's your DAW — or a second generation with a brief that excludes things (§7.7).
- Five minutes is a structure budget, not a length tax. The model holds identity across five minutes because of the structured-caption training (§5.2c) — so a well-briefed 4:30 song is more coherent than a poorly-briefed one. Spend the length on arc, not on padding.
- The brief is a sentence, the lyrics are the script. Two levers, different jobs: the
promptfield sets how it sounds; thelyrics(with tags) set what happens and when. A song that "sounds right but happens wrong" needs a lyrics fix; one that "happens right but sounds wrong" needs a prompt fix. (§7.2's diagnostic starts here.)
Next: §6 — the complete syntax of the Music 3.0 prompt: the sentence skeleton, the canonical pattern, and the section-tag grammar.
§6 — Music 3.0 Prompting: The Complete Syntax
The video book had two formats; the music book has two levers — the prompt field (the style brief) and the lyrics field (the script) — and this section gives you the complete syntax of both. Every rule here is from MiniMax's official prompt-writing guide and the Music 3.0 announcement's own demo prompts, which are the best teaching material MiniMax has published.
6.1 The first rule: sentences, not tags
Write prompts as vivid English sentences — not comma-separated tags.
This is the official guide's opening line, and it's the most counter-cultural rule in the music prompting world, because the community tradition (especially from the Suno years) was the opposite: tag clouds, comma lists, keyword soup. Here's why the official rule wins on Music 3.0 specifically: the model is an 8B language model (§5.2b). Sentences have syntax an LLM parses natively; tag clouds are exactly the disorganized input language models handle worst. The Suno community converged on the same discovery independently (labeled genre: "…" vocal: "…" fields) — MiniMax formalized it.
The working interpretation: the brief is a creative brief for a musician — the way an A&R person or a session arranger would describe a record in one or two spoken sentences. If you can't say it to a musician and have them nod, it isn't a brief yet.
6.2 The sentence skeleton (the official shape)
A complete Music 3.0 prompt follows this pattern — five slots, in order, each one sentence or clause:
A [mood/emotion] [BPM, optional] [genre + sub-genre] [song/track/piece].
[Vocal description — or, for instrumentals, "Instrumental with …"].
[Narrative/theme — what the song is about].
[Atmosphere/scene — the world it lives in].
[Key instruments + production elements].Slot-by-slot, with the official's own examples:
| Slot | Job | Official example |
|---|---|---|
| 1 — Mood + genre | required — the global identity | "A melancholic yet defiant Pop-House song" / "A smoky 74 BPM Neo-Soul fusion" |
| 2 — Vocal (or instrumental) | the voice as a character | "featuring emotional vocals" / "Vocals: Sultry, sophisticated male baritone with smooth jazz inflections and breathy delivery" |
| 3 — Narrative/theme | the subject | "about lighting a torch in the cold dark night as a form of romantic rebellion" |
| 4 — Atmosphere/scene | the world | "evoking a sunrise drive along a coastal highway" (the official's instrumental-scene move) |
| 5 — Instruments + production | the band + the room | "mellow beats with lo-fi elements" / "a warm fretless bassline, shimmering Rhodes piano, and brushed jazz drums" |
The two worked skeletons, straight from the official guide:
Vocal track:
A melancholic yet defiant Pop-House song, featuring emotional vocals, about
lighting a torch in the cold dark night as a form of romantic rebellion,
energetic rhythm with synth elements.Instrumental:
A warm and uplifting 100 BPM indie folk instrumental piece, evoking a sunny
afternoon stroll through a small town market, featuring bright acoustic guitar
fingerpicking, gentle ukulele strums, light hand claps, and a whistled melody
that feels like pure contentment.Notice what the instrumental version does: slot 2 is replaced by a scene ("evoking a sunny afternoon stroll…"). For instrumentals, the scene is the soul — with no voice to carry it, the world is what the listener falls into.
6.3 The canonical demo pattern (what the enhancer writes)
The Music 3.0 launch post demos every song with the same high-resolution pattern — this is what MiniMax's Prompt Enhancement System produces, and it's the most reliable shape for a serious brief:
[Genre / sub-genre], [BPM], [key]. [Emotional arc across sections],
with [vocal: register + texture + phrasing],
[instrument list: primary + supporting],
and [mix/space: the final 3D picture].One canonical example, annotated (from the official demos):
Futuristic melodic EDM / progressive house, ← genre + sub-genre
126 BPM, B-flat major. ← tempo + key (both optional, both powerful)
Reflective verses about memory and digital ← the ARC: verse state → chorus state
identity build into an uplifting,
hook-driven chorus, ← (the arc is a *change*, like H3's vectors)
with an emotive lead vocal, ← vocal character (one noun phrase)
bright layered synths, pulsing bass, ← instruments: 2 primary + 1 support
crisp four-on-the-floor drums, subtle glitch ← the *texture* layer
textures,
and a wide polished festival mix. ← the ROOM: stereo image + finishRead it as a producer would: what kind of record, at what speed, in what key, doing what over time, sung by whom, played by whom, and where you're hearing it. That's a complete production brief in 45 words — and the demos show it holds across everything: Shanghai jazz with "a vintage room-like mix," a "wide polished festival mix," a "natural live-room sound," a "warm spacious club mix." The final slot — the mix/space description — is the most underused lever in the whole music space. It's one sentence, and it changes the dimensionality of the result ("wide," "narrow," "close-miked," "vintage room," "live," "festival") more than three adjectives of mood ever will.
A second canonical, for the other end of the range (ultra-specific vocal persona):
Warm healing Chinese Mandopop ballad, deep emotional with light rock elements,
middle-aged mature male and female duet, mature parental voices aged around
50-70, mother sings gentle tender lead parts with warm mature soft female
vocals, slightly aged timbre, wise and tender tone, father provides deep low
steady supportive harmonies, rich mature male vocals, slightly weathered
gravelly low register, starts with piano and lush strings opening, chorus
builds gradually with drums and electric guitar layers, ends fading back to
pure acoustic gentle close, … soft whispered intimate outro like a gentle
murmur calling home.This is the maximum version — age, timbre, role-split, accent, and a full arrangement timeline ("starts with… chorus builds… ends fading…") in one brief. Two lessons: (1) vocal personas can be cast in fine detail — the more specific, the more character you get; (2) you can script the arrangement's time in prose — "starts with X, builds with Y, ends in Z" is a legitimate and powerful arc description, the sentence-level twin of the section tags (§6.9).
6.4 The vocal: cast it like a character
The official rule: describe vocals as a character — register + texture + technique + phrasing. "female vocal" is the named anti-pattern. The working stack, four attributes:
| Attribute | The question it answers | Range of answers |
|---|---|---|
| Register/gender | who | "mature male baritone-tenor" / "clear female mezzo-soprano" / "aged around 50-70" |
| Texture | how it sounds up close | "breathy" · "gravelly, slightly weathered" · "velvety light rasp" · "airy, buoyant" · "crystal-clear" |
| Technique | what it does | "restrained phrasing" · "soaring octave hooks" · "playful staccato phrasing, falsetto flips" · "expressive vibrato" · "rhythmic sing-rap flow" |
| Effects/space | where it sits | "with lush reverb" · "close-miked, intimate" · "layered harmonies" · "heavy autotune" |
Official phrase bank (steal these verbatim):
- smooth emotional vocals · raw, unpolished vocals shifting between whispers and screams
- breathy delivery with intimate phrasing · powerful soulful vocals with gospel inflections
- sultry, sophisticated baritone with jazz inflections · ethereal, crystal-clear vocals with lush reverb
- relaxed, soul-flavored vocals with ad-libs and melodic scats · aggressive vocal delivery with rhythmic intensity
The duet rule: for two voices, split the roles explicitly — who leads, who supports, what each one's texture is (the 50–70-year-old parental-duet demo above is the template: mother = tender lead, father = deep supportive harmony). Undivided "duet" = a coin flip about who sings what.
Music 3.0's vocal engine is the most technique-receptive in the market (the Flow-VAE stack, §5.2c — it explicitly models breath, pronunciation, falsetto, harmonies, and removes the classic high-frequency "digital hiss"). Say the performance you want, and it performs it.
6.5 Narrative, scene, and the one-image rule
Two slots, one discipline: the song needs a world, and the world needs one image.
- Narrative (what it's about): one sentence of subject, ideally with a tension — "about letting go of perfectionism and embracing your true self like flowing water" (official). The "like X" simile is doing real work: it hands the model a concrete to ground the abstract.
- Scene (where it lives): one place or moment — "a high-end rooftop lounge at night" (the official's example of a scene giving the model "a coherent world"). For instrumentals this slot replaces the vocal slot entirely (§6.2).
- The one-image rule (our addition, from the same logic as §7.9's one-metaphor rule): one scene, one image, developed. "A city" is a direction; "a rain-wet metro platform at 07:10, sodium light on wet tiles" is a world. The model composes toward a coherent image; it cannot compose toward an idea. (The official demos never break this: every one has exactly one scene-image.)
6.6 Instruments, production, and the mix
The official discipline for slot 5 — and the one with the most counterintuitive rule:
Specify 2–3 key instruments precisely; leave the rest to the model.
This is Law Five (the Dial) in music clothing. The model is an arranger, not a toaster — it will fill a band around your 2–3 anchors, and the filling is usually good. The failure mode is the opposite of the video book's: in music, over-specifying the arrangement is where briefs die (you crowd the arranger's model the same way you crowd a shot — §1.6).
The working layers (each optional; 1–2 per layer, max):
| Layer | Examples (official instrument/production reference) |
|---|---|
| Strings/guitar | acoustic guitar fingerpicking · electric guitar riffs · fretless bass · violin · cello · erhu · guzheng · pipa |
| Keys/synth | piano · Rhodes piano · synth pad · synth lead · arpeggiator · music box · organ |
| Drums/perc | brushed jazz drums · electronic drums · 808 hi-hats · trap percussion · cajon · bongos |
| Wind/brass | saxophone · trumpet · flute · harmonica · bamboo flute · xiao |
| Texture/effects | vinyl crackle · tape hiss · ambient pads · glitch elements · rain sounds |
| The mix | "wide polished festival mix" · "vintage room-like mix" · "natural live-room sound" · "warm spacious club mix" · "clean modern mixing" · "narrow mono, close-miked" |
BPM: the clock, in numbers. The official BPM reference, and how to use it:
| Feel | BPM | Prompt phrasing |
|---|---|---|
| Very slow, meditative | 40–60 | "a meditative 50 BPM…" |
| Slow ballad | 60–80 | "a slow 70 BPM ballad…" |
| Mid-tempo groove | 80–110 | "a groovy 95 BPM…" |
| Upbeat, energetic | 110–130 | "an upbeat 120 BPM…" |
| Fast, driving | 130–160 | "a driving 140 BPM…" |
(BPM in the brief is advisory — the model treats it as a strong steer, not a metronome lock. If a dead BPM matters — a track that must hit 128 for a beat-synced edit — that's a post step, not a prompt. §8.3.)
Keys and time signatures work the same way: "B-flat major," "E major," "7/8" are all legal brief language. Use a key when the emotional color of the key matters (minor for ache, major for lift, the specific key for a match to an existing track — §8.3).
The genre map (official reference — 9 families):
| Family | Genres |
|---|---|
| Pop & Dance | Pop · Dance Pop · Electropop · Synth-pop · Dream Pop · K-pop · J-pop · C-pop · City Pop · House · Future Bass · EDM |
| Rock & Alt | Rock · Indie Rock · Pop Rock · Post-Rock · Shoegaze · Punk · Metal · Alternative |
| R&B / Soul / Funk | R&B · Neo-Soul · Contemporary R&B · Funk · Gospel · Soul |
| Hip-Hop | Hip-Hop · Trap · Boom Bap · Lo-fi Hip-Hop · Cloud Rap · Drill · Afrobeats |
| Electronic | Ambient · Techno · Drum and Bass · Chillwave · Vaporwave · Amapiano |
| Folk / Acoustic | Folk · Indie Folk · Country · Chinese Traditional · Celtic Folk |
| Jazz / Blues | Jazz · Smooth Jazz · Jazz Fusion · Bossa Nova · Blues · Avant-Garde Jazz |
| Classical | Classical · Orchestral · Cinematic · Film Score · Epic · Neoclassical · Piano Solo |
| World | Reggae · Latin · Waltz · Tango · Flamenco · (and any named regional style — "light bossa nova / Brazilian pop" is an official demo) |
The blend rule: genre blends are first-class ("Avant-Garde Jazz and Neo-Soul fusion" is an official opener). Order matters — the first genre is the spine, later ones are the seasoning (the same "first tag has most influence" law the Suno community measured — §7.1).
Language note (official): English prompts work best; other-language scene phrases are fine for flavor (the demos mix Mandarin scene texture into English briefs freely). Lyrics, of course, can be in any language the singer is briefed to use.
6.7 The two-lever diagram (your whole music workflow in one picture)
┌─────────────────────────────────────────────┐
prompt ───► │ HOW IT SOUNDS (the style brief, §6.2-6.6) │
(sentence) │ genre·BPM·key·arc·vocal·instruments·mix │
└─────────────────────────────────────────────┘
║
lyrics ───► ┌─────────────────────────────────────────────┐ ║
(tags + words) │ WHAT HAPPENS, WHEN (the script, §6.9) │──┼──► MUSIC 3.0 ──► one finished song
[Verse]… │ section tags, modifiers, ad-libs, dynamics │ ║ (≤ 5 min,
└─────────────────────────────────────────────┘ ║ 44.1k/256k MP3)
║
is_instrumental: true ──► (kills the vocal lever; the brief carries everything)
lyrics_optimizer: true ──► (model writes the script from the brief)
audio_url / cover_feature_id ──► (cover path, §6.8)The diagnostic (use it before you re-roll):
- Sounds right, happens wrong (wrong sections, missing bridge, no lift) → fix the lyrics/tags.
- Happens right, sounds wrong (wrong genre drift, muddy mix, wrong voice) → fix the prompt.
- Both wrong → the brief/lyrics conflict (a "gentle acoustic" prompt over a 160-BPM rap script) → reconcile the two levers first; the model averages conflicts (§7.1).
6.8 The cover path (re-performance, with or without new words)
Music 3.0's cover capability is the "I love this song, I want my version" tool — two workflows:
One-step (quick): model: music-cover + audio_url + a style prompt. Lyrics are auto-extracted by ASR and re-performed in the new style. (Official example: prompt: "Jazz, smooth, late night lounge, saxophone" over any song.)
Two-step (control): music_cover_preprocess on the song → you get the extracted, section-tagged lyrics + structure → you edit the words → generate with cover_feature_id + edited lyrics + style prompt. This is the adaptation path: new words for a familiar melody-skeleton, a language swap, a re-interpretation. (The feature_id is valid 24h; lyrics in this flow is 10–1,000 chars.)
Where the Café uses it: the Founder Pass "your song, our sound" offer — and the music-video pipeline (§8.2) where a client's existing track gets a new performance bed under their video.
6.9 The lyrics: the section-tag grammar (the *other* half of the prompt)
Your lyrics field is a script with production notes. The model reads section tags as structural events — this is the Hybrid-LM's global model doing its job (§5.2b) — and the official demos confirm the working vocabulary:
The core tags
[Intro] [Verse] [Pre-Chorus] [Chorus] [Bridge] [Instrumental] [Solo] [Outro]
+ numbered variants: [Verse 1] [Chorus 1] …
+ role variants seen in official demos: [Hook] [Guitar Break] [Breakdown to drums-Chorus] [All in Chorus]
+ parenthetical form: ( intro ) ( outro )Tag *modifiers* — the section-level dial
The official demos use dynamic modifiers on the tag itself — this is the §5.2b "section events are real" principle, in practice:
[Chorus – with lift] ← the *dynamics* of the section
[Final Chorus – softer / layered] ← the *arrangement* of the section
[Breakdown to drums-Chorus] ← the *transition* into the next sectionRule: a tag is where; a modifier is how. [Chorus] says "chorus happens"; [Chorus – with lift] says "chorus happens and gets bigger." One modifier per tag; put the most important dynamic on the tag and the texture in the prompt field (the two levers stay separated, §6.7).
Ad-libs and vocal events in the script
The demos use parenthetical ad-libs freely — (oh), (typ-typ-typical), ( intro ) — and spelled-out instrumental sections with sung syllables:
[Instrumental dance bridge]
Uuuuuuuuun Just Beguuuuuuu Uuuuuuuuun(Contrast with video: H3's dialogue rules forbid shorthand repetition (§3.10) — music's lyrics embrace it. Same family of models, different parsing. Don't port one engine's habits to the other.)
Layout discipline
- Blank line between sections — always. Tags start their own lines.
- The bridge is "new, optional but recommended" (MiniMax's own words, in the final-chorus-lift demo): a bridge before the final chorus is what makes the final chorus mean something — it's the rest before the last push.
- Close the song. An
[Outro](or a marked final chorus) gives the model a place to land; an open-ended lyric invites a run-on ending. (The community's[End]-tag habit from the Suno years carries over in spirit — give the model a defined finish.) - Syllable & beat discipline (from the community field guide, §7.9 has the full craft): 6–10 syllables/line, 4 lines/section, ±1–2 for lines in the same position, read aloud and clap. Lines that can't be sung to a beat will be mangled — the model fits syllables to the grid, and ragged lines are where "word salad" is born.
6.10 Instrumentals: the full-brief song
With is_instrumental: true, the vocal lever is gone — so the brief carries 100%, and the scene becomes the soul (§6.2's instrumental rule, formalized):
A meditative 60 BPM cinematic piano piece, intimate and unhurried, evoking a
snowfall on an empty city rooftop at night, featuring solo felt piano with
long pedal resonances, a single sustained string pad entering at the turn,
and a quiet close-miked mix with a small-room space.Instrumental checklist: scene = a place, not a mood · one hero instrument + one enterer · the arc in time ("a pad entering at the turn") · the mix as a room. That's a complete instrumental brief in four sentences.
6.11 Master template — the Music 3.0 brief, assembled
A [mood (1–2 words, contrast allowed)] [BPM] [genre / sub-genre] [song],
[emotional arc: verse-state → chorus-state, one clause],
with [vocal: register + texture + technique + space — or "a [scene] instrumental"],
[2–3 named instruments, primary first] + [one texture layer],
and [the mix: wide/narrow, room/live/festival, finish].
Lyrics:
[Intro]
(4 lines, 6–10 syllables each)
[Verse 1]
(8 lines, the story's first beat)
[Pre-Chorus]
(4 lines, the tension)
[Chorus]
(8 lines, the thesis — the words you'd hum)
[Verse 2]
(8 lines, the complication)
[Pre-Chorus]
(4 lines)
[Chorus – with lift]
(8 lines, the same hook, bigger)
[Bridge]
(4–6 lines, the turn — a new angle, new space)
[Final Chorus – softer / layered] (or: [Final Chorus – with lift])
(8 lines, the same hook, transformed)
[Outro]
(2–4 lines, the landing)Pre-flight checklist (30 seconds, before you generate):
- [ ] Brief is 1–3 sentences (not a tag cloud)? Genre first?
- [ ] BPM stated if tempo matters? Key if color/match matters?
- [ ] Vocal cast with ≥2 of {register, texture, technique, space}? (Or scene for instrumental?)
- [ ] Exactly 2–3 instruments named — not a band list?
- [ ] The mix slot filled (wide/room/live)?
- [ ] Arc stated as a change (verse → chorus), not a static adjective?
- [ ] Lyrics: every section tagged, blank-line separated, 6–10 syllables/line, bridge present, a defined outro?
- [ ] Do the two levers agree (gentle brief + gentle script)?
All yes → generate. Then §7 is what you do with the result — including the parts of the result that fight back.
Next: §7 — Advanced control: gravity wells, front-loading, dynamics, realism, and the craft of the lyric.
§7 — Advanced Music Control: Gravity, Front-Loading, Dynamics, Realism, and the Lyric
§6 gave you the shape. This section gives you the physics — how the model actually behaves when the shape meets a real request — and the advanced techniques that separate "it made a song" from "it made my song." The community field guides (especially the Suno one, which has the deepest published instrumentation of how music models actually behave) contribute most of this; where a behavior is model-specific, it's flagged.
7.1 How the model actually hears your words (the style-mesh model)
The single most valuable mental model from the entire music-prompting ecosystem:
A music model does not "read" your brief like a human. It maps your text into a probabilistic style-mesh — it finds the concepts in your words, pulls on the co-occurrence statistics of its training data (which styles, instruments, and structures have historically appeared together), and blends the result. Left unconstrained, it falls toward the statistical defaults: the gravity wells.
Three consequences, each with a countermeasure:
(a) Gravity wells — the model drifts home
Music training data is not evenly distributed. Some styles are statistical attractors: the community has measured (on Suno, and the pattern is architectural — it applies to any LLM-conditioned music model) that nearly every genre request drifts toward pop structure — rock↔pop, funk↔pop, and the whole mainstream stack are so heavily linked in the data that absence of counterweight = pop. You can feel it as the "it came out radio-ready when I asked for desert blues" experience.
Countermeasures (in order of strength):
- Specificity is ballast. "desert blues" is specific; "blues" is a door into the whole cloud. The more precisely you name the genre + era + region + texture, the less surface area there is to drift from. (This is why the §6.3 canonical pattern works: it is dense.)
- Force the weird combination. The community's "escape" move: "orchestral phonk," "emo industrial," "math-rock gospel" — a combination the model has few defaults for, so it must actually build from your words instead of grabbing the nearest attractor.
- Strategic contrast / exclusion. Name what you do not want where the model is likely to add it — but see §7.7: in Music 3.0's sentence format, exclusion is one careful clause ("with no drums, no autotune"), not a list; over-exclusion destabilizes the arrangement (the model needs something to build with — max ~2–3 exclusions, and place them last in the brief).
- The structured caption is the deep counter. Remember §5.2a: Music 3.0 was trained against structured captions precisely so that a held brief beats the gravity. The more your brief mirrors the caption shape (global identity → arc → sections → performance), the more of the model's capacity goes to your arrangement instead of its defaults. In short: on Music 3.0, "write like the Structured Caption" is the anti-gravity system.
(b) Genre clouds — tags travel in packs
Closely related: concepts that co-occur in training data move together. Asking for trap pulls 808 bass pulls sub-heavy low end whether you asked or not; orchestral pulls cinematic pulls epic pulls strings. Two uses:
- Exploit: when you want the cloud, name the center tag and let the orbit come free ("Amen-break hip-hop" arrives with its whole sonic neighborhood).
- Defuse: when the orbit is unwanted, name the non-cloud partner — "trap hi-hats with an upright bass" replaces the cloud's low end with something from a different cloud. You don't fight the gravity; you redirect it with a heavier anchor.
(c) Adherence decay — front-load what must be true
Measured community behavior (Suno v5.5, but architecturally general): prompt adherence is tightest in the first 30–60 seconds of a generation and decays from there — the model repeats established motifs, simplifies, loops progressions, and drifts toward its defaults as the song proceeds. Music 3.0's global Hybrid-LM is MiniMax's direct answer to this failure mode (a full-song context model holding identity to 5 minutes) — but the writer's countermeasure still stands:
Put your most distinctive, must-be-true elements in the opening of the brief, and reinforce the song's identity again in the arc clause.
Concretely: the signature texture, the signature vocal character, the one instrument the song is known by — all in sentence one. The late-song sections inherit from what was strongest, and you choose what that is.
The word-level findings (from the community's 200+ track mass-testing — treat as priors, verify on your model):
- Emotion words beat technique words. "desperate" shaped the output more than "minor key with reverb." (The model's semantic layer responds to affect; the acoustic layer responds to physical description — use both, in their respective slots: emotion in the mood/arc slots, physics in the mix/instrument slots.)
- Some words are load-bearing, some are noise. Community-verified working: "anthemic chorus," "building intensity," named emotion words. Community-verified weak/ignored: "ethereal" (alone), "crescendo" (say "building intensity" instead), pure mixing-technique words in the wrong slot.
- The practical version of Law Five for music: mood = emotion words · arrangement = structure words · sound = physics words. Each slot has its own vocabulary, and mixing vocabularies across slots is where briefs go to get average.
7.2 The two-layer diagnostic (fix the right lever)
From §6.7, expanded into a field procedure. When a render is wrong, classify the error before touching anything:
| Error type | Example | Fix lever | The move |
|---|---|---|---|
| Wrong sound | came out pop when you asked for shoegaze; muddy low end; wrong voice age | prompt | Sharpen genre/era/texture; add the mix slot; re-cast the vocal; add 1–2 exclusions at the end |
| Wrong happenings | no bridge; chorus doesn't lift; intro too long; section order shuffled | lyrics/tags | Add/rename the section tag; add the dynamic modifier ([Chorus – with lift]); tighten the arc clause in the brief |
| Wrong words | mangled syllables, word salad, broken pronunciation | lyrics (craft) | §7.9: fix line lengths, remove the un-singable line, simplify the rhyme |
| Both | gentle brief over a 160-BPM rap script | reconcile | Make the two levers tell one story; then re-generate |
The re-roll protocol (the music version of pin-and-nudge): change one lever per iteration. Generate 3–4 candidates per setting (music models, like video, re-roll to different results — Law One). Keep the best of each in a listening log: which candidate nailed the sound? Which nailed the structure? The assembly move (when the best-sound and best-structure come from different candidates): take the structure of candidate A and the sound character of candidate B, write a brief that describes both, and re-generate — the model can often synthesize across your own history better than any single roll did. (The community calls this "assemble from parts"; on Music 3.0 it works because the structured caption can express your past outcomes as new conditions.)
7.3 Section-level control: directing the arrangement *in time*
The §6.9 grammar, with the full advanced toolkit — this is where "the model holds a 5-minute song" becomes "the model plays my 5-minute song":
The three tools, by scope:
- Global (the brief's arc clause): "verses stay intimate and sparse; the choruses open wide; the bridge drops to near-silence before the final lift." — the shape of the whole song in one sentence. Write this first; it's the arranger's program note.
- Section tags + dynamic modifiers (the lyrics):
[Chorus – with lift][Stripped Back][Building Intensity][Sudden Break]— events, localized. One modifier per tag. - In-section prose (the brief's texture/arrangement clauses): "the drums enter at the second verse and never fully drop out; the strings arrive only in the final chorus" — instrument choreography, which is the one thing the model plans globally (the Hybrid-LM's §5.2b job) and executes locally.
The dynamics vocabulary (the working set):
| Intent | Phrase (brief) | Tag modifier (lyrics) |
|---|---|---|
| get bigger | "builds in density toward the chorus" | [Chorus – with lift] |
| get smaller | "the arrangement strips back to voice and one instrument" | [Stripped Back] |
| sudden change | "a sudden drop to near-silence, then the full band" | [Sudden Break] |
| rise to climax | "a slow build from solo vocal to full orchestral wall" | [Building Intensity] → [Massive Finale] |
| breath | "an instrumental passage for the song to breathe" | [Instrumental] / [Interlude] |
The chorus rule of thumb: the chorus is where listeners and models both concentrate. If your song's identity is the chorus (most songs are), the chorus gets the densest brief language: the vocal's chorus delivery ("the chorus vocals double an octave higher"), the arrangement's chorus weight ("the chorus is a wall of sound against the sparse verses"), and the tag's chorus dynamics. Sparse verses + a specified chorus is the single highest-leverage arrangement move in the book.
The final-chorus move (the official's own recommendation): the bridge is optional but recommended — its job is to empty the room so the final chorus can mean something. [Bridge] (quiet, new angle) → [Final Chorus – with lift] (or, for ballads, – softer / layered) is the two-line climax system behind most of MiniMax's own demos. Steal it; it works every time.
7.4 Directing vocal performance (the Flow-VAE layer)
Music 3.0's vocal engine is explicitly built for performance direction (§5.2c: it models breath, pronunciation, falsetto, harmonies, and removes the "digital hiss" that gave away older models). The addressable vocabulary, in four tiers:
- Timbre: "slightly aged, warm" · "velvety with a light rasp" · "crystal-clear" · "gravelly, weathered" — who (see §6.4's stack).
- Delivery: "restrained, intimate phrasing" · "belted with open vowels" · "spoken-sung, rhythmic" · "ragged, on the edge of breaking" — how the words ride the melody.
- Technique: "expressive vibrato on held notes" · "falsetto flips on the hook" · "ad-libs and melodic scats" · "stacked harmonies a third below" — what the voice does (the demos use all of these; they are executed, not decorative).
- Space: "close-miked, dry and intimate" · "drenched in hall reverb" · "with a delay echo on the last word of each line" · "light autotune sheen" — where the voice sits in the mix.
The duet/choir system: assign roles (lead / harmony / chant / backing), registers (who is high, who is low), and textures (who is clean, who is gritty) to each voice — the parental-duet demo (§6.3) is the full template. A choir brief works the same way: "a lone soprano, then a backing choir enters a step behind her, then a massive final-chorus choir" — the choir as a moving object with entry and growth.
What not to do: don't stack all four tiers on every section — that's the music version of max-detail-everything (§1.5). Pick the one vocal event that defines each section; the rest stays to the model's performance instinct, which on Music 3.0 is genuinely good.
7.5 Making it sound *real*: the physical-recording stack
When the goal is "this should sound like a recorded performance, not a generated one" — bedroom folk, singer-songwriter, one-take, field recording — the word "realistic" is the weakest tool in the language (models under-respond to it; it's too close to a judgment to be a specification). The community's solution, and the one that works: describe the physical recording, not the idea of realism.
The acoustic-realism stack (steal the lines):
small-room acoustics · room tone (air, faint hiss) · close-mic presence
off-axis mic placement · proximity effect · single-mic capture
one-take performance · natural timing drift (human micro-rubato)
natural dynamics (no brickwall feel) · breath detail (inhales, exhales)
mouth noise · pick noise, fret squeak · chair creak, body shift
tape saturation · analog warmth, slight wow-and-flutter · gentle preamp drive
narrow stereo image · consistent background noise floor · imperfections keptThese are physical facts a microphone would capture — and physical facts are exactly what the Flow-VAE's continuous representation was retrained for (§5.2c: "string attack and bowing… the sense of space around the voice"). The official "acoustic bossa nova" demo sits at one end of this spectrum; the full stack above sits at the other ("one person, one guitar, one take, a real room").
The electronic counter-stack (the opposite problem: models default to a generic electronic sound — the community's "sawtooth trap"): you can't reliably forbid the default, but you can replace it with a named synthesis family:
FM synthesis bass · wavetable movement · formant-driven bass
granular textures · evolving modulation · non-repeating bass cycles
rounded harmonic profile · controlled high harmonics · mono-stable low endThe principle is the general one, stated for electronics: give the model a specific to latch onto instead of the default it would pick. The same logic runs the whole book — §4.5's "name the physical material of the UI," §4.6's "state the fluid's continuity," §7.1's "name the non-cloud partner."
7.6 Exclusion: the careful kind
Music 3.0 takes prose, so "negative prompting" takes the form of careful clauses — with the community's hard-won rules:
no Xworks better than "don't add X" — the parser treats the former as a clean constraint.- Exclusions go at the end of the brief — positives are processed first; a trailing "no drums, no autotune" reads as a final veto.
- Two to three exclusions, maximum. The model needs material to build with; a brief that excludes half the instrument world arrives with nothing to arrange. (If you need many exclusions, you're describing a different genre — name that genre instead.)
- Belt-and-braces for the stubborn ones (community practice): state the exclusion and name the replacement — "no drums; the groove is carried by a fingerpicked bass" — because absence invites the model to improvise a filler, while replacement gives it a job.
Common exclusion targets (the field guide's table): only-female-vocals → exclude male vocal · pure acoustic → exclude electronic/synth/drum machine · pure rock → exclude electronic/hip-hop/pop · kill the wordless intro → exclude humming/oohs/wordless intro vocals (then start on the first lyric line, §7.8).
7.7 The one-sentence rhythm: matching the brief to the model's parsing
A meta-rule the whole §7 section rests on: the model parses sentences with roles. The same word does different work in different slots (§7.1c: emotion in mood slots, physics in mix slots, structure in arc slots). The writing discipline that follows:
- One idea per sentence. (The H3 one-fact rule, §3.3, is the same rule — the two engines share the constraint even when their formats differ.)
- Role words anchor the sentence: "the vocal is…" · "the arrangement…" · "the mix…" · "the groove…" — the anchor tells the model which layer of the structured caption your sentence is writing to.
- Order = priority. First sentence = global identity; arc clause = the shape; texture sentences = the details; exclusions = last. (Position is salience — §1.6, both engines.)
7.8 Structure surgery: intros, starts, and endings
The three most common structural failures, and their specific fixes:
- The phantom intro (the model opens with a hum, a wordless "ooh," or a long runway before the first line — the most-reported music-model annoyance of the last two years): fix with both levers — an exclusion clause in the brief ("no wordless intro; the vocal enters on the first bar") and a script that starts on the lyric (first line of the verse as the first content of the lyrics field; the community's
[START_ON: TRUE]+ first-words trick generalizes to any model that reads tags). - The runaway ending (the song overruns: a mid-verse fade, an instrumental ramble, a hard cut at the length limit): close the script with an explicit
[Outro]containing the landing line — a defined finish is what the model converges to. (Community effectiveness on Suno: ~85–90%; on Music 3.0's 5-minute horizon, a defined ending is even more important — you have more runway to overshoot on.) - The flat middle (verses 1 and 2 feel identical; no arc): this is the arc clause's job (§7.3) — state in the brief what differs between the two verses ("the second verse adds a counter-melody and the vocal gets more urgent"). Two verses that are briefed as different come out different.
7.9 Lyric writing for AI voices (the craft layer)
The model sings syllables against a beat grid. Ragged script = mangled delivery. The community's complete craft rules, all of them portable to Music 3.0:
1. Syllable discipline.
- Target 6–10 syllables per line (most genres); 4 lines per standard section, 8 for longer ones.
- Lines in the same structural position (e.g., all four lines of a chorus) should vary by ±1–2 syllables — the grid needs near-regularity, not regularity.
- The clap test: read the line aloud while clapping the beat. If you can't clap it cleanly, rewrite it before the model gets a chance to mangle it.
- One-syllable-short fix: repeat the last word ("stay with me" → "stay with me, stay") or add a natural filler ("stay with me tonight, love"). Never pad with a weird word — the model will sing the weird word.
2. The one-metaphor rule. One song, one central metaphor, explored in all its facets — water (or fire, or light, or architecture — one). A song that switches river → fire → clouds → earth gives the listener and the model no visual anchor, and the arrangement wanders with it. (The official demos obey this religiously: "cloud copy of me" is a single digital-ghost metaphor, sustained across every section.)
3. The narrative arc (verses tell; chorus says; bridge turns).
- Verses: story beats — concrete, progressing, each verse a new scene of the same world (Verse 1: the glance · Verse 2: the crack in the facade).
- Pre-chorus: the tension — 4 lines of rising pressure toward the release.
- Chorus: the thesis — the sentence the listener hums. Short, reusable, emotionally named.
- Bridge: the turn — a new angle, a new space, the same hook seen from behind (this is the §7.3 final-chorus move's setup).
- Final chorus: the same hook, transformed in meaning by the journey — the words can be identical; the brief says what changed ("softer, layered, the harmonies underneath").
4. The hummability test (the final gate): before you generate, say the chorus to yourself out loud, three times, at tempo. If the third time feels like chewing gravel, the line lengths or the rhyme are wrong — fix the script, not the model. The model will sing what you give it exactly; your job is to give it something singable.
5. Language and pronunciation: the model sings the language of the lyric (the demos run Mandarin, English, and mixed freely). For hard-to-pronounce names and words, spell the intended pronunciation in a brief note ("the word 'Solligence' is sung 'SOL-li-gens'") — the vocal engine's pronunciation modeling (§5.2c) gives you a hook to pull on. (For on-screen text, that's H3's §4.5 job — one model writes the letters, one sings the word; §8.2 joins them.)
7.10 The advanced checklist (before any non-trivial generation)
- [ ] Brief mirrors the structured caption: global identity → arc → performance?
- [ ] Must-be-true elements front-loaded (sentence one)?
- [ ] Emotion words in mood slots; physics words in mix slots; structure words in arc slots?
- [ ] 2–3 instruments, primary first; the mix slot filled with a space word?
- [ ] Exclusions ≤ 3, at the end, each with a replacement?
- [ ] Arc clause names what changes (verse → chorus; v1 → v2; bridge → final)?
- [ ] Chorus has the densest brief language of any section?
- [ ] Lyrics: 6–10 syllables/line, clapped out loud, one metaphor, bridge present, defined outro?
- [ ] Two levers agree (no brief/script conflict)?
- [ ] Generation plan: 3–4 candidates, judge sound and structure separately, keep best-of-each?
Next: §8 — Convergence: scoring video with Music 3.0, and the Café's one-render-two-engines loop.
§8 — Convergence: Video × Music, and the One-Render-Two-Engines Loop
Half this book is H3, half is Music 3.0. The product — what you actually ship, and what the Café sells — is the combination: a finished, platform-ready short that is picture and sound from one production act. This section is the join: how the two engines meet, the diegetic/non-diegetic hinge, beat-syncing, and the production loop that turns both into a weekly output machine.
8.1 The diegetic / non-diegetic hinge (the one distinction that runs everything)
Every sound in a finished piece is one of two kinds, and the whole book's audio rules exist to keep the two routed correctly:
| Diegetic (in-world) | Non-diegetic (over-world) | |
|---|---|---|
| What it is | sound that exists in the scene — room tone, SFX, voices, a character's instrument | score that the audience hears but the characters don't |
| Where H3 owns it | overall_soundscape (§3.11) — generated with the video, native stereo, source-synced | non_diegetic_music (§3.12) — one-sentence dressing only |
| Where Music 3.0 owns it | — (it's a music engine; its output is score or source music) | the actual production — composed, arranged, performed, mixed |
The hinge rule: H3 dresses a scene with sound; Music 3.0 is the production. A scene's world sound is H3's job; a piece's score is Music 3.0's job. The moment you find yourself trying to make H3 produce a real arrangement (instrument lists, BPM, mix detail in a video prompt) — you're using the wrong engine. That detail belongs in a Music 3.0 brief, full-stop (§3.12's division-of-labor table, now load-bearing).
Why this matters more than it looks: the two engines' same words mean different things. "strings, 120 BPM, wide mix" in a Music 3.0 brief = produce that arrangement. The same words in an H3 non_diegetic_music line = violate H3's one-sentence rule and muddy the mix. One vocabulary, two grammars. The hinge is the firewall.
8.2 The two ways to marry a H3 picture with a Music 3.0 score
Path A — In-world (diegetic) marriage: use the score as a reference. If the music exists in the scene (a band playing, a radio, a character singing, a music video), feed a ≤ 15-second excerpt of the Music 3.0 track to H3 as a reference audio (role = the music/voice the scene contains), and describe the source in the scene ("the band on the stage is playing this; the crowd sways to it"). H3's omni-rep then syncs the picture to the sound — motion, pace, and spatial audio all referencing the actual track. This is how you get a music video where the performance is the brief.
- Constraint: the H3 reference-audio budget is ≤ 15s total (§2.5) — so you sync on the most characteristic 15 seconds of the track (the hook), not the whole song.
Path B — Over-world (non-diegetic) marriage: lay the score in the edit. If the music is score (the audience's bed, the characters don't hear it), generate the picture in H3 with a light non_diegetic_music note (or none), generate the score in Music 3.0, and combine them in your edit — the track under the cut. This is the default for brand films, product ads, and most social content.
- The lock: when you'll lay a specific track over the cut, brief the H3 picture's pace and edit points to match that track's sections (§8.3), so the cut and the score land together in the edit.
Which path? One question: Do the characters in the scene hear the music? Yes → Path A (it's in the world). No → Path B (it's over the world).
8.3 Beat-syncing: making the picture and the score *land together*
The reason a Reel hits is that visual cuts and musical events share a timeline. You can't make H3 "cut to the beat" directly (it doesn't know your score), but you can brief the picture's structure to the score's structure, then cut to match. The workflow:
- Get the score first. Generate the Music 3.0 track before the picture. Its sections are the edit's spine. (This is the key workflow inversion: music leads, picture follows.)
- Map the score's sections to picture beats. A 15s H3 clip has room for ~3–5 beats (§3.5). Line the score's section boundaries (verse→pre-chorus→chorus, the drop, the lift) against those beats. The chorus should land on the visual reveal; the drop on the cut.
- Brief each H3 shot to hold until its musical event. Give each shot a duration that ends at a section boundary in the score, and write the shot's one move so its climax (the reveal, the turn, the hit) lands on the boundary. (E.g., "the lid opens on the drop" = the 1-vector motion timed to the score.)
- Cut in the edit, on the sections. The DAW's waveform is the conductor; your H3 shots are the orchestra, already rehearsed to the score's phrasing.
The practical minimum (for a first cut): match the three biggest events — the intro→verse, the chorus/drop, and the final-hit/outro — to the three most important visual beats. You don't need 50 on-beat cuts; you need the three that matter to land. Everything else can breathe.
The honest limit: sub-frame beat-quantization (every single cut grid-locked) is a post skill (the DAW's snap-to-grid), not a prompt skill. The prompt's job is to get the big events aligned; the edit's job is to quantize the small ones. Don't spend prompt budget on what the DAW does for free.
8.4 "Long" content = a *sequence* of H3 shots (the only way H3 makes long)
H3's hard ceiling is 15 seconds (§2.5). So "a 60-second brand film" or "a 3-minute music video" is not one generation — it's a cut of many H3 shots to a spine, and the spine is usually the score. This is the single most important production idea in the book, because it's how a one-person studio makes anything long:
The sequence workflow (music leads):
- Score first (Music 3.0) — the full track is your timeline and your structure. Its sections are your scene list.
- One H3 shot per musical section — a 3-minute song with ~8 sections = ~8 H3 shots, each 8–15s, each briefed to that section's mood and that beat's event (§8.3). Same character slot, same palette, same light logic across all of them (§3.13, §3.8) — that's what makes 8 separate generations feel like one film.
- Cut to the score's sections in the edit; the picture changes with the music, so the cut is motivated (it never looks like a slideshow, because the music is already telling the audience why it's cutting).
- The music video case (Path A): for a performance MV, the shots are the band/character playing, synced to the track as reference audio (§8.2A) — the same 8-shot sequence, but the sync is in-world.
The continuity contract that makes the sequence feel like one piece (the three invariants, repeated across every shot):
- Same character slot (identical feature text + same reference image) — §3.13
- Same palette anchors + light logic (per location) — §3.8
- Same "one world-law" (the single rule of the piece's reality, §4.2) — the audience trusts a world with one consistent law
Break any one of the three and the 8-shot sequence reads as 8 unrelated clips. Keep all three and it reads as a film a crew made — which is the whole illusion, and it's yours to build.
The "long video" honesty note (positioning rule #3): we build long things as edits of short, great shots — never by claiming a model "makes full-length video." That's both what the tech actually does and what our positioning allows us to say. (The Café's tagline is 5-seconds-is-a-feature; the edit is where length lives.)
8.5 The Café's one-render-two-engines loop (the product, in one diagram)
This is the loop the Café is, and the loop §11 industrializes. One brief, two engines, one finished short:
┌──────────────────────────────────────────┐
YOUR IDEA ───────► │ 1. SCORE IT (Music 3.0) │
(concept + brief) │ the track = the timeline + structure │
└──────────────────────────────────────────┘
│ sections = scene list
▼
┌──────────────────────────────────────────┐
│ 2. PICTURE IT (H3, one shot per section) │
│ same slot · same palette · one world-law│
│ diegetic sound natively in each shot │
└──────────────────────────────────────────┘
│ ~N shots
▼
┌──────────────────────────────────────────┐
│ 3. CUT TO THE SCORE (the edit) │
│ 3 big events on-beat · continuity held │
└──────────────────────────────────────────┘
│
▼
FINISHED, PLATFORM-READY SHORT
(9:16 Reel · 16:9 brand film · 1:1 · 4–15s shots)Why two engines, one loop (the value story):
- The picture has sound from the first render (H3's native audio) — so even a first draft of the picture is watchable, not silent. You're evaluating the piece, not the footage.
- The score is produced, not faked (Music 3.0) — a real arrangement you can keep, reuse, and re-cut to. It's an asset, not a stock track.
- The same brief drives both — the mood/arc you wrote for the score is the same arc you brief the picture against. One creative decision, two renderings. That's the "simplicity" the Café sells: you describe the feeling; the studio returns the piece.
The 5-second truth (positioning rule #3, applied): each H3 shot is 4–15s — short by design, reliable, iterable. The piece is long; the shot is not. Every 5-second beat is a feature, because it's a beat you can re-roll in minutes when the one that matters underdelivers. That's the entire production philosophy in one line.
8.6 The section drill
Take one idea you have. Run the full loop on paper before touching a render box:
- Write the Music 3.0 brief for its score (§6.11 template) — 1–3 sentences.
- Let the score's sections become your scene list (one H3 shot each).
- Write one H3 prompt for the most important section (the chorus / the reveal) using the §3.15 template.
- Note the three invariants (character slot / palette / world-law) you'll hold across the other shots.
You have now, on paper, a complete production plan for a finished short — two engines, one loop, three invariants. That plan is the Café. Render step 3 and you've started shipping.
8.7 A complete worked run (one product ad, start to finish)
The loop in full, with real briefs at each step — the closest thing in this book to a recipe. (The product: a matte-black pour-over kettle, for a D2C coffee brand.)
Step 1 — Score first (Music 3.0). prompt: A warm minimalist acoustic cue at 84 BPM, a sparse fingerpicked guitar, an unhurried build from a single chord to a soft two-chord lift, with a wide, open, dry mix, no vocal. (is_instrumental) → 3 candidates; keep the one whose second chord-lift lands at a usable timestamp (~10.8s in) — that timestamp is now an edit target. (Its section map: intro [0–4s] → solo guitar [4–11s] → the lift [11–16s].)
Step 2 — The scene list from the score's sections (one H3 shot each, all 9:16, all one world-law: matte kitchen, morning light, steam is the only magic):
| Shot | Score section | H3 brief (the full version is in the style: §10.3) |
|---|---|---|
| 1 (4s) | intro | kettle on a matte counter, cold still; 5 seconds… the kettle sits still; a thin thread of steam. Camera static. Sound: near-silent room. (Music omitted.) |
| 2 (7s) | solo guitar | the kettle pours: 7 seconds… the kettle pours a single unbroken stream into the glass carafe; the water rises, one soft surface ripple. Camera static, close. Sound: the pour, foreground. (Music omitted.) |
| 3 (5s) | the lift | the reveal: 5 seconds… the carafe lifts off the counter, steam rising, the morning light through the glass. Camera slow push-in. Sound: a soft set-down, foreground. (Music omitted.) |
Step 3 — Cut to the score. Shot 1 starts at 0:00 · Shot 2 starts at the guitar entry (~4s) · Shot 3 starts at the lift (10.8s) — the one on-beat moment that matters (§8.3: three events, not fifty). Total: 16s of a 16s track. The picture ends where the music ends.
Step 4 — The three invariants held (that's what makes 3 generations feel like one film): same matte palette anchors in all three (matte black, cream, pale wood, warm white) · same light recipe (one soft key from the left, single soft shadow) · one world-law (steam is the only magic; everything else is real and still). The music leads; the picture is its portrait.
Step 5 — The tail (in the edit, once, at the end): Made in Solligence Café · UPI · ₹99/hr · 10 min free → solligence.in
The entire cost of this ad: one Music 3.0 brief + three H3 prompts + one cut. No crew, no studio, no shoot — and every roll along the way was a piece of content (§11.1). That is what ₹99/hr buys.
Next: §9 — when it fights back: failure modes, and the 30-second diagnostic.
§9 — When It Fights Back: Failure Modes, and the 30-Second Diagnostic
Models fail in patterns, not in random noise. The same failures recur across every generation, across every model in the market, because they trace to the architecture (latent prediction, probabilistic blending, statistical gravity — §1.6). Learn the patterns once and every bad render becomes legible: you can name the failure, find its row in the table, and apply the fix — in about 30 seconds, before you waste a re-roll guessing.
This section is the field manual. Everything here is a compression of the rules already taught; the value is in the diagnosis → fix table format, built to be used while staring at a bad render.
9.1 The four hard problems of 2026 video models (and where H3 stands)
The industry's open frontier is the same four things for every major model. Knowing which one you're hitting is half the fix:
| Hard problem | What it looks like | H3's position (2026) | Your posture |
|---|---|---|---|
| Physics | objects melt, liquids behave like jelly, gravity direction flips, momentum vanishes | strong on simple rigid-body + stated-change; edge cases (complex fluid+cloth+glass) still slip | Constrain the state change (§9.2) |
| Multi-character consistency | two people: one drifts, faces swap, identities merge | strong single character (its reference system is the market's best general-purpose); two interacting characters is the open case | One hero per shot (§9.2); near-miss instead of contact |
| Hands | extra fingers, fused knuckles, hands that pass through objects | improving; still the least reliable region on every model | One hand, one job (§3.5) |
| On-screen text | misspelled logos, scrambled small type, "S0LL1GENCE" | at/near the front of the pack — a headline H3 capability, with 2K regeneration recovering small text | Use it, verify it (§4.5) — H3's real edge |
The honest frame for customers (and for you, briefly, before a deadline): three of the four are industry limits with a shrinking gap year to year; the fourth (text) is where H3 currently leads. So: build on H3's strength (text + single-character + simple physics), route around its shared-weakness zone (two-character contact, hands, complex materials) with composition, and verify the strength per render. That's not a limitation list — it's a targeting strategy.
9.2 The diagnostic tables (the 30-second fix)
VIDEO — name the symptom, take the fix:
| You see | It is | The fix (in order of reach) |
|---|---|---|
| Object/melt/shape drift mid-clip | physics over-invention | 1) shorten the state change ("falls two feet" not "tumbles around") 2) add a rest at the end ("…and holds") 3) one material hero per shot 4) re-roll (sometimes it's just a bad roll) |
| Liquid looks like jelly / stream breaks | fluid without constraint | state the continuity explicitly: "the stream unbroken," "fills to two-thirds, steady" (§4.4) |
| Character's face/clothes change frame to frame (single char) | identity under-spec | tighten the §3.4 feature card; in ref work, check the retention_analysis line names the exact features; re-roll |
| Two characters' identities bleed/swap | multi-character frontier | prefer one hero + object; for genuine two-shots: give both full distinct cards + the near-miss (hands approach; the cut lands the contact, §4.6) |
| Hands: extra/missing/fused fingers | hands frontier | one hand acts, one holds still (§3.5); move the fine-detail action to a close-up of a different body part; re-roll |
| Camera does two things at once / wobbles oddly | vector conflict | one move per shot, one speed word (§3.6); if you need two moves, cut between two shots |
| Action happens but too fast / rushed | beat overload | count your beats vs the seconds (5s ≈ 3–5 beats, §3.5); delete beats or extend the clip |
| On-screen text scrambled | text (verify!) | zoom the 2K pass and read the letters (§4.5); if wrong: shorter string, higher contrast, bigger size, re-roll — or move the type to your edit as a graphics layer (always available) |
| Sound doesn't match the action / SFX on wrong event | source unbound | state the source object + sync explicitly in overall_soundscape ("the click as the lid opens — foreground", §3.11) |
| Music in the soundscape sounds like a wall of noise | diegetic/non-diegetic bleed | move score OUT of overall_soundscape into non_diegetic_music (one sentence) — or cut the score and let the diegetic design stand (§3.12) |
| Whole clip is "almost right, slightly wrong" (the common one) | dial mis-set | don't re-roll yet — find the one element that was supposed to be locked that drifted (§1.5); lock only it; keep the rest; re-roll |
MUSIC — name the symptom, take the fix (from §7.2, field form):
| You hear | It is | The fix |
|---|---|---|
| Sounds like generic pop / radio-ready | gravity well (§7.1a) | sharpen genre+era+region; force a weird combination; name the non-cloud partner; max 2–3 trailing exclusions |
| Muddy / congested low end | no mix slot / default mix | fill the mix slot with a space word ("clean, open, controlled low end", §6.6); or the electronic counter-stack (§7.5) |
| Genre right, but arrangement flat / no build | arc unstated | write the arc clause (what changes verse→chorus, §7.3); give the chorus the densest language |
| Chorus doesn't lift / song is static | no section dynamics | tag modifiers: [Chorus – with lift], bridge before final chorus (§7.3's climax system) |
| Wrong section order / no bridge / runaway ending | structure unstated | lyrics fix: explicit tags, blank-line separation, defined [Outro] (§6.9, §7.8) |
| Phantom intro (hum/ooh before the words) | model default | brief exclusion ("no wordless intro") and script that starts on the first lyric line (§7.8) |
| Word salad / mangled syllables | un-singable script | §7.9: fix line lengths (6–10 syllables, clap test), one metaphor, simplify the rhyme — fix the script, not the model |
| Vocals sound "synthetic" / hiss / rigid | performance under-brief | direct the performance (breath, phrasing, vibrato, space — §7.4); for "recorded" feel, the physical-recording stack (§7.5) |
| Drifts to pop mid-song | long-form adherence decay | front-load identity (sentence one) + reaffirm in the arc clause (§7.1c); 3–4 candidates, keep the best |
| Two levers fighting (gentle brief, 160-BPM rap script) | brief/script conflict | make the two levers agree on one story, then re-generate (§6.7) |
The meta-diagnostic (when the tables don't obviously fit): ask the three §1.5 questions in order — Which element was supposed to be locked that drifted? Which element was supposed to be open that I now want? Is this a sound problem or a happening problem (music)? The answer picks your lever and your fix. If you can't name the specific element, you can't fix it — so the first move is always naming, not re-rolling.
9.3 The strip-back protocol (when a shot keeps misfiring)
Sometimes the problem isn't one element — it's that the whole shot is over-briefed into incoherence (the model averaging a crowd of demands, §1.6). The recovery is a deliberate simplification ladder, in this order:
- Freeze the camera. ("The camera is static.") Camera motion is the most independent variable; killing it removes an entire class of motion failures and lets you judge the action cleanly.
- One action, one subject. Delete every secondary action; keep the single primary vector (§3.5).
- Clear the background. "Plain / uncluttered background" — remove competing subjects the model must also simulate.
- One material, one light. Reduce to a single material hero and a single light source (§3.8, §4.6).
- Shorten the clock. If it's a 15s shot, make it a 5s of the same core — does the idea work when small? (Often the 15s version failed because the 5s core was never stable.)
Then rebuild upward, one layer per successful render: once the stripped shot works, add one complexity back (a second action, or a camera move, or a background element) and re-render. You're re-earning complexity the model can actually hold — and you end with a shot that's both right and robust to re-rolls. This is the exact "strip back, then layer" recovery Sora's official guide describes for its edit tool; on H3 you do it with the prompt itself.
When to stop stripping: if the idea needs the complexity (a two-character contact is the shot), the fix is route around — reframe so the hard thing happens across a cut (the near-miss, §4.6) or in two shots — instead of asking one 15s render to do it all. Composition is a stronger tool than raw model capability, and it's always available.
9.4 Pin-and-nudge: the iteration protocol (the habit that makes you "good at prompting")
The single most useful behavior in this book — the difference between someone who "can't prompt" and someone who ships:
- Render N candidates (3–5 for video; 3–4 for music). Law One: they'll differ.
- Pick the one closest to the brief. Not the "best-looking" — the closest to what you asked for. (A gorgeous wrong answer teaches you nothing; a close answer has your answer in it.)
- Name the one gap. The single element that drifted and must be locked (§9.2's meta-diagnostic).
- Change one thing. Lock that one element with stronger language; keep every other sentence identical. Re-render.
- Repeat 2–4 until the locked element holds across re-rolls (test it: render the unchanged prompt 2 more times; if the element holds 2/2, it's pinned).
- Only then add the next complexity (§9.3's rebuild order).
Why "one thing" is non-negotiable: change two variables and you can't tell which one worked — you've turned iteration into search. One variable per render keeps every re-roll a legible experiment. (This is also why the §6.7 two-lever split matters: it gives you named variables to change one at a time, instead of one undifferentiated blob of words.)
The pin log (keep it — it's your style system): a running note of pinned phrasings per character / product / palette / world-law ("the robot: black hoodie, red chest lettering, glowing blue eyes — always these three, in this order"). This log is your brand's production grammar. It's how the §8.4 continuity contract gets enforced across weeks — and it's the asset a solo studio has that a freelance contract never builds. (It doubles as the §11 batch engine's source of truth.)
9.5 What you will *not* fix by prompting (honest limits)
Some things are post jobs, and pretending they're prompt jobs burns hours:
- Sub-frame beat quantization (every cut grid-locked to the DAW) → the DAW's snap-to-grid, not the prompt (§8.3).
- Long-form text / multi-line paragraphs on screen → still a coin toss for every model in 2026; long copy belongs in your edit as graphics (§4.5).
- A dead BPM / exact sample-accurate tempo → the model treats BPM as a steer; exact tempo is a post (time-stretch) or a different tool's job.
- Stems / mix separation → Music 3.0 returns one finished mix; stems are a DAW (or a stem-separation model's) job (§5.1).
- Pixel-perfect logo replication → H3 renders your logo well (verify per roll) but "pixel-perfect" is a vector-over-video post step for hero frames.
The principle: prompt for what the model generates; post what the model approximates. Every hour you spend prompting a post-job is an hour stolen from the next creative decision. The model is a performer, brilliant and improvisionary — the edit suite is the control room. Know which room you're in.
9.6 The 30-second pre-delivery checklist (run it before you call a render "done")
- [ ] The locked element held across 2 re-rolls (pinned, not lucky)?
- [ ] Hands / physics / text checked at 2K (the three verify-or-route items)?
- [ ] Sound synced to visible sources; diegetic/non-diegetic correctly split?
- [ ] Music: sound and structure both right (the two-layer diagnostic, §7.2)?
- [ ] Continuity (for sequences): same slot, same palette, one world-law across shots?
- [ ] The three big beats land on the score's three big events?
- [ ] Compliance tail present, said once, at the end (UPI · ₹99/hr · 10 min free — §11)?
If all seven are yes: it ships. If one is no: you now know exactly which one, and you have the fix for it — which is the entire point of this section.
9.7 The failure catalog (a card to keep beside the tables)
The 18 most common failures, each in one breath — for the week you're shipping under pressure and can't re-derive:
Video — 10: ① object melts mid-clip → shrink the state change + add a rest · ② liquid jellies → state the continuity ("unbroken, fills to two-thirds") · ③ single-character face/clothes drift → tighten the feature card / retention line, re-roll · ④ two characters swap → one hero + near-miss · ⑤ hand horror → one hand acts, one holds · ⑥ camera does two things → one move, or cut · ⑦ action rushed → fewer beats (5s ≈ 3–5) or longer clip · ⑧ text scrambled → 2K verify, or move type to edit · ⑨ SFX on the wrong event → name the source + sync · ⑩ soundscape is a wall → score to non_diegetic_music.
Music — 8: ⑪ generic pop → sharpen genre+era, weird combo, non-cloud partner · ⑫ muddy low end → the mix slot ("controlled low end") / the counter-stack · ⑬ no build → the arc clause · ⑭ chorus flat → the densest language + [Chorus – with lift] · ⑮ wrong section order → fix the tags · ⑯ phantom intro → exclusion + script starts on the lyric · ⑰ word salad → fix the script (clap test) · ⑱ synthetic vocal → the performance tier + the recording stack.
When in doubt, the strip-back ladder (§9.3) first — it resolves most of the ten video cases in one move.
Next: §10 — the style playbook: seven looks, each a complete production recipe.
§10 — The Style Playbook: Seven Looks, Each a Complete Production Recipe
Style is Law Four — the strongest lever — and a style is more than an opening sentence. A working style is a bundle: the medium line, the motion grammar, the light recipe, the sound design, the palette, and the failure profile (where this look breaks, and how to route around it). This section gives you seven complete bundles — the seven looks the Café actually produces — each ready to fill in with your subject.
How to use: pick the look that fits the brief → take its template → swap in your subject/content → apply its failure routing. Everything else in this book is why; this section is what you copy.
10.1 CINEMATIC LIVE-ACTION (the brand-film default)
The look in production language (the Law-4 opener): Cinematic live-action, shot on 35mm, anamorphic 2.0×, 2.39:1, shallow depth of field, natural light with warm practicals, filmic grain, fine halation on highlights.
The grammar of the look:
- Motion: slow, motivated, human-scale — pushes, tracks, statics. Never fast whip-moves (those are the other style's job).
- Light: one key source named with direction + temperature (§3.8); windows, practicals, golden hour. The 3–5 palette anchors are the grade.
- Sound: diegetic-first — room tone, one synced SFX, ambience in the background (§3.11). Score under, one sentence.
- The tell: the audience should feel filmed, not rendered — the grain and the halation and the imperfect light do that work.
Template:
[Duration], [aspect]. Cinematic live-action, [lens/format], [DOF], [light recipe].
Palette: [5 anchor colors].
[Subject card]. [Scene: fg/mid/bg layers]. [One primary action, one vector].
[Camera: one slow move or static].
Sound: [room tone] in the background; [one synced SFX] in the foreground.
Music: [one genre word + one emotion + one behavior].Worked (the §4.1 earbud reveal is the worked version; here's the human variant):
12 seconds, 16:9. Cinematic live-action, 35mm anamorphic, shallow depth of field,
warm morning window light from camera left, soft tungsten fill.
Palette: amber, cream, walnut brown, slate blue, soft grey.
A woman in her 30s, dark hair in a low bun, olive shirt, sits at a window desk,
a paper coffee cup in both hands, looking out. Rain streaks the glass.
She lifts the cup, sips once, and sets it down. The camera is static, medium close-up.
Sound: quiet room in the background; rain on glass; the soft clack of the cup on wood.
Music: warm lo-fi, understated, a gentle lift as she looks up.Failure profile & routing: this look demands physical plausibility — it has the least stylistic cover. Human faces + hands + cloth are all in frame. Route: one human, one action, stillness as the default motion; if hands fail, reframe to hands-off (the cup is already in her hands — hold, don't take). The 35mm/grain language is also a consistency anchor across a sequence (§8.4) — reuse it verbatim in every shot of the piece.
10.2 HAND-DRAWN × LIVE (glowing lines over reality — the Café's signature)
The look in production language: Live-action footage fused with hand-drawn glowing animation — simple white line-work over live footage, lines that glow and leave short fading trails, deep base grade with neon accents.
The grammar of the look:
- The world-law (the whole concept in one sentence): real things behave; drawn things drift. Real rain falls; drawn figures bob. Real objects obey physics; line-creatures float, trail, and fade. That one law is what makes the style read as magic instead of glitch (§4.2's lesson, maximized).
- Motion: the drawn layer moves in slow drifts with trails — the trail is the motion signature (a line-creature that moves without a trail looks like a UI, not a ghost).
- Light: the live layer gets a deep, low-contrast grade (the lines need darkness to glow against); the drawn layer owns the neon.
- Sound: textural and synthetic — each drawn element has its own small sound (hums, chimes, paper-creaks), panned with its motion (§4.2's teardown). This look is made for H3's native stereo.
Template:
[15] seconds, [9:16]. Live-action [place] fused with hand-drawn glowing animation,
simple [color] line-work, [deep grade] base, [neon accent] trails.
[Live scene: crowd/street/room, one real action].
(dissolve) [Drawn figures — what they are] drift [direction], bobbing,
leaving short glowing trails that fade; the trails weave [fg/mid].
[One world-law: what stays real, what becomes drawn].
Sound: [live ambience] in the background; each drawn figure makes [its sound],
panning with the figures, in the foreground.
Music: [dark ambient / pulse], [curious/uncanny], building with [the real event].Failure profile & routing: the drawn layer's style is robust (it's 2D — few physics promises); the live layer underneath is the risk (crowds = many faces). Route: keep the live layer slightly defocused or at distance (shallow DOF on the crowd = the faces become texture, and the lines carry the frame); the (dissolve) into the drawn layer is the reality switch — put it early so the magic is established before the audience decides if it's a glitch. This is the IG channel theme (CHANNEL-THEMES.md) — same grammar, new place + new drawn-creature = a new render, every week, forever.
10.3 MINIMALIST PRODUCT (the D2C ad, no voiceover)
The look in production language: Minimalist product film, matte studio void, one soft key light, single-source shadow, high-key negative space, the product as the only object in frame.
The grammar of the look:
- The void is a feature: one product, one plinth/counter, nothing else. Negative space is the composition; the product's silhouette is the frame.
- Motion: the product is the actor — one physical action (open, pour, lift, hover) with the §4.4 continuity constraint ("the stream unbroken"). Camera: static or one slow push-in; the product moves, the camera mostly doesn't (the opposite of cinematic — here the object is the subject).
- Sound: the product's own physics — the click, the pour, the slide, the clack. Often no score at all (omitting
non_diegetic_musicis the confident move, §4.4). - Text: this is where H3's text edge earns its keep — the brand mark on the product, or a one-word end-card, rendered in the shot (§4.5). Verify per roll.
Template:
[5–10] seconds, [9:16]. Minimalist product film, matte [color] studio void,
one soft key from [direction], single soft shadow, generous negative space.
[The product, feature-card: material + one distinguishing detail + brand mark].
[One action: open/pour/lift/hover — with its continuity constraint].
The camera [is static / slowly pushes in].
Sound: near-silent room; [the product's own sound, foreground, time-synced].
(Music: omitted — or one 2-word minimal cue under the action.)Failure profile & routing: the risk is material rendering (glass, liquid, metal are the hard materials, §4.6). Route: one material hero per shot (the glass pour OR the metal lid, not both); state the fluid's continuity; if a material keeps melting, change the action (a lid opening survives where a bottle pouring doesn't — the rigid body is the safe shape). The 5-second minimum clip is the native ad shape here — short is the style.
10.4 PAPERCRAFT / STOP-MOTION DIORAMA (the tactile explainer)
The look in production language: Cut-paper papercraft stop-motion — layered paper diorama, visible paper edges, tiny punched details, matte pastel palette, soft even light from above with gentle shadows between the layers.
The grammar of the look:
- Layers are the grammar: 3 depth layers (fg / mid / bg) are both the composition (§3.7) and the style (papercraft = physical depth). Simple geometry = few physics promises = the safest look in this book (§4.6's routing logic, applied to a style).
- Motion: ticks. "Slow, slightly jerky stop-motion ticks" — the jitter is the tell; smooth 24fps motion would read as 3D, not papercraft. One vector with a stutter profile.
- Sound: the medium's physics — paper creaks, wheels tick, flaps slap (§4.6 worked example). The ear confirms the eye.
- Light: flat, even, from above — the opposite of cinematic; shadows exist only between layers (and that's where they create the depth).
Template:
[10–15] seconds, [9:16]. Cut-paper papercraft stop-motion, [N] depth layers,
matte [palette] pastels, soft top light, shadows between layers, visible paper edges.
[The metaphor-as-diorama: what the topic is, built in paper].
[One moving element in the midground layer, in stop-motion ticks].
[Its paper-physics sound: creak/tick/flap]. The camera is static, all layers visible.
Music: [music-box / glockenspiel / soft], [wistful / playful], underneath.Failure profile & routing: extremely robust — the failure is boredom, not breakage. Route: the motion must carry meaning (the paper train = the idea moving; the paper heart filling = the idea growing) — a moving element that illustrates beats a moving element that merely animates. (This is the Café's papercraft-explainer skill: topic → metaphor → diorama → ticks → new render.)
10.5 3D-STYLIZED (game CG / toy-grade characters)
The look in production language: 3D stylized render, toy-grade soft shapes, subsurface-scattered skin, exaggerated proportions, clean studio light with a single rim, high saturation, smooth arcs of motion.
The grammar of the look:
- Proportions are the style: big eyes, small features, rounded edges — the design does the work; the render must stay soft (subsurface scattering, no hard micro-detail — a 3D-stylized character with photoreal skin pores is a style war).
- Motion: smooth arcs, no ticks — the opposite of papercraft; this look lives on fluid 24fps arcs and squash-and-stretch. One vector, eased.
- Sound: toy-grade — bouncy, plucky, soft-impact SFX (the "boing"/"pop" register); characters' moves have cartoon sounds even in a "real" 3D render.
- Consistency: this is where Ref2VA shines hardest — a 3D character is defined by its model sheet; one reference image (role: subject) + the retention contract (§3.13) keeps the same character across a whole series.
Template:
[10] seconds, [9:16]. 3D stylized, toy-grade soft shapes, subsurface skin,
exaggerated [proportions], clean studio light + one [color] rim, saturated palette.
[Character: full feature card — or "Image 1 = the character" in ref format].
[One physical action with squash-and-stretch].
[Camera: one smooth arc or static].
Sound: [soft studio room]; [toy-grade SFX for the action, foreground].
Music: [playful, bouncy], [one behavior].Failure profile & routing: the risk is style war (the model drifting to photoreal or to flat-cartoon within one clip). Route: name the anti-style once at the front ("soft toy-grade, not photoreal") and hold the same style sentence in every shot of a series; if a frame goes photoreal, that's a re-roll, not a rewrite — the style is the lock.
10.6 MUSIC VIDEO / LYRIC TYPOGRAPHY (the MV)
The look in production language: Music-video aesthetic — [the song's world as a place], [the performance as the subject], [lyric typography as a designed element of the frame].
The grammar of the look (this is the §8.2 Path-A + §8.3 loop, styled):
- Performance is the subject: the singer/band is the character slot; the score is the reference audio (≤15s of the hook, Path A) — the picture syncs to the real track.
- The three-beat sync (§8.3): the chorus/drop lands on the visual reveal; the verse = the quiet shots; the final hit = the biggest shot. Brief each H3 shot to hold until its musical event.
- Lyric typography (the MV's signature): on H3, short lyric lines can render in-world (a sign, a screen, a wall — §4.5's text strength). The long lyric sheet belongs in the edit (graphics layer) — in-world type is for the 3–5 words that matter, placed as props in the scene.
- Light/sound: follow the song's world (a dark-ambient song = the §10.2 deep grade; a funk song = the §10.1 warm practicals) — the score dictates the look (music leads, §8.4).
Template (one shot of the MV — the chorus/reveal):
[12] seconds, [9:16]. [The song's world as a place], [grade matching the track],
[the performer: character slot], performing to the track.
Reference: [Voice/Song 1 = the track's hook, ≤15s]. The performance [one move]
lands on the [chorus/drop].
[3–5 words of the hook on a [sign/screen/wall] in-world].
Sound: the track (reference) drives the mix; [one world SFX] under it.Failure profile & routing: the sync is the product — if the reveal doesn't land on the drop, fix the cut (the DAW), not the render. In-world lyric text: keep it ≤ a short phrase, large, high-contrast, and verify the 2K pass per roll (§4.5) — if the type scrambles, the line moves to the edit's graphics layer (the style survives; only the delivery changes).
10.7 CHARACTER-LED MENU / CO-OP INTRO (two characters, one world)
The look in production language: Stylized [2D / 3D-toy] title sequence — two distinct characters, a shared designed world, title typography as the frame's hero, menu-grade polish.
The grammar of the look (the two-character problem, solved by design):
- Two characters, zero contact. The whole look is built around the near-miss (§4.6): the two characters share a composed frame (left/right thirds, distinct silhouettes, distinct palettes — contrast is the consistency here), but the interaction happens across a cut (they face each other; the "(cut)" lands the handshake). Two characters in contact in one render is the industry's hardest case; two characters in conversation across a cut is a style.
- The title is a character: the wordmark is briefed like a subject (§3.4: material + one feature — "the title 'SOLLIGENCE' in chunky hand-drawn lettering, cream on deep blue, holding the top third").
- Sound: a title has a leitmotif — one 5–10s music-box/synth cue (Music 3.0, instrumental) is the sound design; the characters' actions get toy-grade SFX (§10.5) under it.
Template:
[10–15] seconds, [9:16]. [Stylized medium] title sequence, [the shared world],
[palette: char-A color vs char-B color — complementary].
[Character A card, left third] faces [Character B card, right third];
[one small action each — both "hold" shapes, no contact].
[The title as a designed element, top third, one material].
The camera is static (or one slow push on the title).
Sound: [quiet room]; [small SFX per action];
Music: [5–10s instrumental cue] under the title.Failure profile & routing: if the two characters bleed (identities swap — the §9.1 multi-character frontier), sharpen the contrast (bigger palette/silhouette difference + full distinct cards each) and/or separate them across two shots (A alone → B alone → the title over both as still figures). Still figures + a title is the fallback that always works — a menu is, in the end, mostly still.
10.8 Choosing a style (the decision, in 20 seconds)
| Your content is… | Start with the style | Why |
|---|---|---|
| A brand / a person / a place, "premium" | 10.1 Cinematic | the default trust-signal |
| Something magical / weird / trend-y | 10.2 Hand-drawn × live | the style is the hook; safest "wow" |
| A product / a D2C ad | 10.3 Minimalist product | object-as-actor; H3's text edge on the mark |
| An idea / an explainer / a kid's topic | 10.4 Papercraft | safest physics; motion must mean something |
| A mascot / a game character / a series | 10.5 3D-stylized (+Ref2VA) | the consistency workhorse |
| A song / an artist / a music video | 10.6 MV (Path A) | the score leads; §8's loop |
| Two characters / a title / a menu | 10.7 Character-led | near-miss design; the title is the hero |
The style is a decision made once — at the top of the brief — and held (the same opening sentence in every shot of the piece, §8.4). Change the style mid-piece and you've made two videos. Pick one, lock it, and let the content vary inside it. That's the whole craft of looking like a studio.
10.9 Mixing styles (and the one rule that keeps it honest)
Occasionally a brief genuinely needs two looks — a product film whose hero dissolves into a papercraft metaphor; a brand film that crosses into hand-drawn magic. The rule for doing it without it becoming a war:
- One style per shot; the crossing is a transition. The
(dissolve)/(scene shift)is the only place a style may change inside a piece — exactly like a scene change, it's a declared move (§3.9). Two styles within one continuous shot = the model averaging two worlds = mush. - The arrival style is the brand style; the departure style is the idea style. If the piece is a brand film that uses paper-magic as its metaphor, cinematic is the home base and papercraft is the excursion (and it returns: the final shot is cinematic). The brand is always where the piece begins and ends.
- The crossing gets its own sound rule. The style switch is the one place
overall_soundscapemay change mid-piece — the ambience tail with the dissolve (the §4.2 rule) — and the non-diegetic cue bridges it (one cue whose behavior spans the switch: "…shifting in texture with the dissolve"). The sound is what sells the visual switch to the brain.
That's the whole of it: cross on the transition, return to the brand style, and let the sound do the convincing.
Next: §11 — the Café production system: the cadence, the Founder Pass engine, and the compliance rules.
§11 — The Café Production System: From Skills to a Weekly Machine
The previous ten sections made you fluent. This section is industrial: how fluency becomes a repeatable weekly output that behaves like a brand, like a funnel, and like a business. It's the Café's operating system, written down — the same system that runs marketing/strategy/CHANNEL-THEMES.md and LAUNCH-PLAN.md, now generalized so anyone with this book can run it.
11.1 The flywheel: every render is an ad (house rule #1)
The Café's first house rule, and the one that makes the whole system self-funding:
Every render is an ad. Nothing gets generated that doesn't also ship as content.
The logic: in a normal studio, R&D renders (testing, iterating, the §9.4 pin-and-nudge re-rolls) are cost — burned hours that produce nothing but learning. In the Café, the learning is the product's proof. The pin-and-nudge iterations, the style experiments, the "can AI do X?" tests — all of them are content, because they show the studio working, which is exactly what sells a studio.
The flywheel, concretely:
render (any purpose: test, client, experiment)
│
├──► ships as content (the process IS the ad: "we just did X in 5 minutes")
│ │
│ ▼
│ audience sees *capability in action* (not a claim — a receipt)
│ │
│ ▼
│ interest → the 10-min free trial → the hour (UPI, ₹99)
│ │
└────────────┘ (every render also *costs* the studio ~minutes,
not weeks — the cost of the ad is the cost of the work)Why this beats the alternatives:
- Claim-based marketing ("we make amazing AI video") costs nothing to make and trust to work. The render-receipt costs minutes to make and works on its own — the audience watches the thing actually happen.
- Stock/competitor comparison marketing violates the positioning rule ("sell simplicity, never superiority"). The flywheel never compares — it just shows.
- Consistency falls out for free: if every render ships, the channel never goes quiet. There is no "content day" and "work day" — the work is the content calendar.
The one discipline it requires: label the process. A raw test-render ships with its story — "prompt → result," "3 rolls, kept the second," "we were trying to make the hands work" — because the process is the product's proof. (This is exactly the X "Prompt→Render" channel theme: the prompt is the before, the render is the after, the gap between them is the ad.)
11.2 Platform-native output: one render, the right shape
H3's container (Law Two) meets each platform's native shape. The map — one brief, rendered in the platform's aspect + the beat length that platform rewards:
| Platform | Native shape | H3 container | The content role |
|---|---|---|---|
| X (Twitter) | text-first, audio-off, 16:9 or 1:1 | 5–15s, 1:1 or 16:9 | the Prompt→Render proof: the text of the prompt is the post; the render is the attachment. (X is where the words are the product.) |
| Instagram Reels | sound-on, beat-driven, 9:16 | 5–15s, 9:16 | the style show: hand-drawn×live, cars×music, the look is the stop-scroll. (IG is where the sound+look is the product.) |
| YouTube Shorts | sound-on, 9:16, searchable | 15s (the max), 9:16 | the "Can AI [X]?" test: one question, one proof, evergreen (search compounds — the §8.4 sequence makes 60s+ from 15s shots). |
| YouTube (long) | 16:9, search-driven | a sequence of H3 shots cut to the score (§8.4) | the deep proof: a full "we made X" build, the pinned-comment funnel. |
| 16:9 or document, trust-first | 5–15s 16:9, or a document (prompt + result as a post) | the business case: the founder's process post (the pitch is "hours, not weeks; UPI, ₹99/hr" — the reliability story, not the wow). |
The rule that keeps this honest: one fixed format per platform, never cross-posted (the CHANNEL-THEMES.md law). The same idea can render for three platforms — but each gets its own shape, aspect, and own story frame, because each platform's algorithm rewards a different kind of attention. Cross-posting a 9:16 Reel to X is how brands look like spam; the native version of the same idea looks like a studio.
The container quick-reference (H3, §2.5): duration 4–15s integers · aspect in prose (9:16 / 16:9 / 1:1) · 768P iterate → 2K the keeper (§4.7) · the music is ≤15s reference-audio (Path A) or laid in the edit (Path B).
11.3 The compliance tail (house rules 2–4, applied to *every* post)
Three positioning rules, turned into a mechanical habit so no post ever needs a legal thought:
- Sell simplicity, never superiority (rule #2). The post shows what got made and how fast — it never says "better than [competitor]" or "movie-grade / physics-accurate" (that's the quality war, and it ages badly). The vocabulary is time and ease: "in five minutes," "one prompt," "no crew, no studio, no timeline." (If a draft reaches for "unlike other tools…" — that's the rule. Cut it.)
- 5 seconds is a feature (rule #3). Long pieces are edits of short shots (§8.4) — and we say that out loud as the strength: "every beat is 5 seconds, so every beat is re-rollable." Never "full-length video." The edit is where length lives; the shot is proudly small.
- UPI always (rule #4) — and the whole offer, once, at the end. Every commercial post ends with the standard tail, said once, never in a price-comparison table:
Made in Solligence Café · UPI · ₹99/hr · 10 min free → solligence.in
The tail is one line, one screen, one tap. It keeps the "fun" themes (the magic of §10) as business assets — the wonder is the hook; the UPI line is the door. (It also future-proofs: when the price changes, one line changes everywhere.)
The open-weights sentence (the business customer's hook, from §5.4): for B2B posts (LinkedIn) only, the one allowed "how it's built" line — "Both engines behind the Café — video and music — are open-weight models we run ourselves. Your briefs, your renders, your data." — said as simplicity ("your studio, your terms"), never as a spec-sheet. It's the whole differentiator in one breath, and it's true (H3 and Music 3.0 are both open-weights — §2.1, §5.1).
11.4 The 50 Founder Passes as a *production engine*
The 50-person Founder Pass program (marketing/programs/FOUNDER-PASS.md) is usually read as a sales program. Read it as a production system and it's the Café's content factory with a built-in audience:
- Each Founder Pass = a named human with a real brief. Their first render is a real request (their product, their song, their character) — which means the "every render is an ad" flywheel now has 50 different real stories instead of one studio's experiments.
- The weekly slot: one Founder Pass render per week, shipped with their story and their permission (the "Founder Pass Fridays" IG slot, per CHANNEL-THEMES.md). 50 passes = 50 weeks of a real-people series, pre-built. (That's the "real-people testimonial = weekly slot, not platform theme" decision from the channel plan — the pass is the scalable testimonial.)
- The pass as the pin log's source: each founder's first render becomes a row in the §9.4 pin log — their product, their palette, their one world-law. Fifty founder kits = a library of reusable, real-world style bundles.
- The compounding: a founder who gets one great render comes back. The pass isn't a discount — it's the first frame of a relationship, and the relationship is the channel.
The honest constraint (the one that keeps it real): a Founder Pass render is a collaboration, not a free service — the founder supplies the brief and the judgment (the §9.4 pin-and-nudge is theirs, with the studio's help), which is exactly why it's a pass and not a voucher, and exactly why the render that ships is theirs to tell the story of.
11.5 The weekly SOP (the 5-slot cadence, operationalized)
From CHANNEL-THEMES.md's week-1 calendar (~25 posts/week across platforms), reduced to the production rhythm — what actually happens, in order, in a week:
Monday — Plan (30 min). Pick the week's 5 hero ideas (one per platform-slot: X proof ×2, IG style ×2, YT question ×1, +1 Founder Pass). For each: write the one-sentence brief + choose the §10 style + note the container (aspect/length). No rendering yet. (The decision is the Monday; the pixels are the week.)
Tuesday–Thursday — Render (the pin-and-nudge days). For each hero: 3–5 rolls at 768p → pick closest → pin-and-nudge the one gap (§9.4) → 2K the keeper (§4.7). Every roll that ships is content (§11.1) — the process posts (prompt→render, the failed-then-fixed story) are queued as they happen, not batched later. The Founder Pass render runs this same loop, with the founder in the loop on the pin decision.
Friday — Cut + ship (the edit day). The §8.5 loop where music leads: score first (Music 3.0), picture shots to its sections, cut to the three big beats (§8.3), the standard tail on each (§11.3), platform-native shapes (§11.2). Everything ships in the week it was made — the flywheel's one non-negotiable.
Saturday — Measure (20 min). Not vanity metrics — the three that feed next Monday's plan: which render stopped the most thumbs (the style that works), which prompt got the most "how?" replies (the demand), which founder came back (the relationship). Those three answers are next week's five ideas.
The through-line: decide on Monday, render in the middle, ship by Friday, measure on Saturday. The §1 flywheel only works if the ship is weekly and unconditional — a render that doesn't ship is a cost, not an ad. The whole system is just §11.1, scheduled.
11.6 The cost truth (why ₹99/hr is the *point*)
The system's economics, stated plainly (this is the §11 "say it once" number, for the B2B/LinkedIn slot):
- H3's per-second pricing is the floor of the model market (2K at under a third of mainstream closed models, §2.1) — and the Café's self-hosted, open-weights setup means no per-call API markup sits on top of it (§5.4).
- The Café charges by the hour, not by the render. A render is minutes of model time + minutes of your judgment (§9.4's pin-and-nudge is faster than it looks once you're fluent — 3–5 rolls, one gap, done). The hour buys the studio's attention, not the GPU.
- So the offer is a rate, not a list: "₹99/hour, by UPI, 10 minutes free" is all of it — no tier table, no per-render SKU, no "contact us for pricing." That simplicity is the product — the positioning rule made manifest in the price. (A price this simple is itself the proof of the brand: a studio that can explain itself in one line.)
The 10-minute free trial is the flywheel's ignition: it's exactly one pin-and-nudge cycle — enough to render, enough to see the method work, not enough to build a habit you'd have to charge for. Ten minutes of "oh, that's how" is the conversion.
11.7 The first 30 days (a week-by-week of the system above)
If you start the Café from zero, this is the shape of month one — the §11.5 SOP applied across four weeks, with the goal of each:
| Week | Render focus | Content focus | The one number to watch |
|---|---|---|---|
| 1 — The grammar | the §10 styles, one each, for the brand itself (the S-tile, the café counter, the UPI tap — C.12) | establish the 5 fixed platform slots; every slot ships; the tail on every post | volume: 20+ renders shipped, zero silence days |
| 2 — The proof | the "Can AI [X]?" set (C.7) + the product set (C.1) | YT searches the questions; X the prompt→render; IG the styles | "how?" replies per render (the demand signal) |
| 3 — The people | the first 5 Founder Pass renders | Founder Pass Fridays begin (IG) + founder personal-profile posts tagging the Café | pass activation: 5/5 founders got a render they kept |
| 4 — The system | the full §8 loop on one hero piece (music leads → shots → cut) | the 60s+ YouTube build (the §8.4 sequence) + the pinned-comment funnel live | one completed hero piece, end to end, on time |
The 30-day test of the whole system: if, by day 30, you have 20+ shipped renders, a live 5-slot weekly cadence, 5 active Founder Passes, and one completed hero piece — the machine is built; now (and only now) does the question of paid amplification enter, pointed at the one render that stopped the most thumbs (§11.5's Saturday metric). Everything else in the business is downstream of that first honest signal.
11.8 The one-page operating summary (pin this)
- Every render is an ad — ship it, label the process, never let a render die in a folder.
- One fixed format per platform — same idea, native shape each; never cross-post.
- Music leads, picture follows — the score is the spine; shots are 4–15s beats; long = an edit.
- Three invariants across every shot of a piece — same character slot · same palette · one world-law.
- Pin-and-nudge, one variable at a time — keep the pin log; it's your style system.
- The tail, once, at the end — Made in Solligence Café · UPI · ₹99/hr · 10 min free → solligence.in.
- Decide Monday, ship Friday, measure Saturday — the flywheel is a schedule.
You now have the whole machine: the model truth (§1–§8), the failure science (§9), the looks (§10), and the system that turns all of it into a brand and a business (§11). The appendices are the tools you keep beside you.
Next: the appendices — the cheat sheet, the rule tables, and the prompt libraries.
Appendix A — The One-Page Master Cheat Sheet
The whole book, one page. Print it. Pin it above the render box.
The five laws (§1)
- Briefing — you're briefing a fast, literal, under-specified crew; gaps become their improvisation; re-rolling is the job.
- Container — parameters own the container; prose owns the content. H3 exception: duration + aspect go in prose.
- One-Move — 1 primary + ≤2 secondary actions per subject, one vector each; one camera move per shot; count beats (5s ≈ 3–5); state change, not state.
- Early-Style — medium/style first, in concrete production language; it re-weights everything.
- Dial — every element is Lock / Steer / Open; set by importance; diagnose drift by element, then pin-and-nudge.
H3 video — the shape (§3)
integrated_multimodal_description
[integer 4–15] seconds, [aspect]. [STYLE/medium]. [scene: fg/mid/bg].
[subject: features only — no mood on faces]. [light: source+dir+temp, ≤3 clauses].
[1 primary + ≤2 secondary actions]. [one camera move + speed, or static].
[transitions: (cut)/(dissolve)/(scene shift) — per-shot setup].
overall_soundscape
Sound: [env, 1–2, neutral, spatial] · Sound: [synced SFX, source named, fg/bg]
Dialogue: [Speaker says: "…"] (on-screen speaker ≤ 2 lines)
non_diegetic_music
Music: [genre + emotion + one behavior] ← omit if no scoreRef format adds: subject_definitions (asset + role) · summary · retention_analysis (what to keep + how it coexists). Specs: 4–15s integers · refs ≤9 img / ≤15s video / ≤15s audio / ≤12 files · prompt ≤7000 chars · 768p→2K via Regeneration.
Music 3.0 — the shape (§6)
prompt: A [mood] [BPM] [genre/sub-genre] song, [arc: verse→chorus],
with [vocal: register+texture+technique+space], [2–3 instruments],
and [the mix: wide/room/live]. (2–3 exclusions, at the end, max)
lyrics: [Intro] [Verse 1] [Pre-Chorus] [Chorus] [Verse 2] [Chorus – with lift]
[Bridge] [Final Chorus – softer/layered] [Outro]
6–10 syllables/line · 4 lines/section · blank line between · one metaphorSwitches: is_instrumental · lyrics_optimizer · covers (one-step audio_url / two-step cover_feature_id) · ≤5 min · 44.1k/256k MP3.
The join (§8)
- Diegetic (in-world) = H3's
overall_soundscape· Non-diegetic (score) = Music 3.0's production. - Music in the scene? → H3 reference audio (≤15s of the hook). Score over the cut? → lay it in the edit.
- Music leads, picture follows: score's sections = scene list; one H3 shot per section; cut to the 3 big beats.
- Sequence invariants: same character slot · same palette · one world-law.
Diagnose in 30 seconds (§9)
- Video: melt/break → shrink the state change + add a rest · face/clothes drift → tighten the feature card / retention line · two chars swap → one hero + near-miss · hands → one hand acts, one holds · text scrambled → 2K verify or move to edit · almost-right → find the one unlocked element.
- Music: generic pop → sharpen genre/era + weird combo + non-cloud partner · flat → write the arc clause, densest language on the chorus · structure wrong → fix the tags · word salad → fix the script (clap test).
- Recovery: freeze camera → one action → clear bg → one material/light → shorter clock; then re-add one layer per success.
- Iterate: N candidates → pick closest to brief → change one thing → pin (holds 2/2 re-rolls) → next. Keep the pin log.
Ship it (§11)
- Every render is an ad — label the process; ship weekly.
- Native shapes: X = prompt→render (text-first, 1:1/16:9) · IG = style, sound-on (9:16) · YT = "Can AI [X]?" (15s 9:16, search) · LinkedIn = the business case (16:9).
- The tail, once, at the end: Made in Solligence Café · UPI · ₹99/hr · 10 min free → solligence.in
- The week: decide Mon · render Tue–Thu · cut+ship Fri · measure Sat (which stopped thumbs / which got "how?" / which founder came back).
Appendix B — Rule Tables (the look-up dictionaries)
Reference tables for the whole book. Copy the row you need; it is pre-vetted phrasing.
B.1 Camera moves (H3) — one per shot (§3.6)
| Intent | Phrase | Speed words |
|---|---|---|
| static (the default that never fails) | "The camera is static, [wide / medium / close-up]." | — |
| slow approach | "The camera pushes in slowly on [subject]." | slow / gradual / gentle |
| reveal | "The camera rises and tilts up to reveal [scene]." | slow / smooth |
| follow | "The camera tracks [subject] from [side] as [action]." | steady |
| orbit | "The camera orbits [subject] in a slow half-circle." | slow |
| wide-out (for scale/ending) | "The camera pulls back to a wide, revealing [context]." | slow |
| locked-off two-shot | "The camera is static, framing both [A] and [B] in [left/right thirds]." | — |
Rules: one move per shot (two = a cut of two shots) · pair every move with one speed word · for sequences, vary the moves shot-to-shot (static → push → static) to create rhythm without chaos.
B.2 Lighting recipes (H3) — ≤3 clauses each (§3.8)
| Look | Recipe (source + direction + temperature) |
|---|---|
| morning window | "soft daylight through a window from camera left, cool fill, warm practical lamp" |
| golden hour | "low golden-hour sun from camera right, long shadows, warm tungsten fill" |
| night city | "cool neon practicals, mixed color temperature, rim light from [source]" |
| studio product | "one soft key from [direction], single soft shadow, matte background" |
| candle/ember | "warm flickering practicals, deep shadows, low contrast around the light" |
| studio (3D-stylized) | "clean soft key + one [color] rim, even fill, no hard shadows" |
| hand-drawn×live base | "deep low-contrast grade, [neon] practicals, dark negative space for the lines" |
| papercraft | "soft even top light, shadows only between layers, matte" |
Palette anchors (§3.8): name 3–5 specific colors ("amber, cream, walnut, slate, soft grey") — one sentence; the model grades to them.
B.3 Action & beat patterns (H3) — §3.5, §4.6
| Pattern | Phrase |
|---|---|
| single vector | "[subject] [one verb phrase], in [one direction/speed]." |
| state change (NOT state) | "the lid opens" ✓ / "the lid is open" ✗ · "she sits down" ✓ / "she is seated" ✗ |
| the rest | "…then [subject] holds [state]." (stops the improvisation) |
| two beats | "[action 1]; then [action 2]." (a semicolon, a cut worth) |
| the dissolve (reality switch) | "(dissolve) [the second world/layer] takes over; [what continues, what changes]." |
| the near-miss (two-character contact) | "hands approach, the [contact] happens on the (cut)." |
| fluid continuity | "the [liquid] stream unbroken; fills to [level]; no splash." |
| stop-motion ticks | "the [element] moves in slow, slightly jerky stop-motion ticks." |
| squash-and-stretch | "the [character] [action] with a soft squash-and-stretch." |
B.4 Transition operators (H3, per shot) (§3.9)
| Operator | Use for |
|---|---|
(cut) | a hard jump; also the landing of a near-miss |
(dissolve) | a reality switch / time pass / style crossfade |
(scene shift) | a location change within one logical moment |
| (none) | continuous action inside one shot — the default |
B.5 Sound design (H3) — §3.11 quick rows
| Slot | Row |
|---|---|
| ambience | "Sound: [room/street/space] ambience in the background — [one concrete source]." |
| synced SFX | "Sound: [SFX] as [visible source event], in the foreground." |
| stereo move | "[SFX] moves from left to right as [subject crosses]." |
| voice | "Dialogue: [Speaker] says: '[line]'." (on-screen speaker ≤ 2 lines) |
| score | → non_diegetic_music only (one sentence) — never a wall in the soundscape |
B.6 Non-diegetic one-liners (H3) — §3.12
Music: [genre word], [emotion word], [one behavior — "understated, lifting at the end"]. Examples: Music: warm lo-fi, understated, a gentle lift. · Music: dark pulse, tense, building. · Music: omitted. (the confident move for product films)
B.7 Genre families (Music 3.0) — §6.8
| Family | Entry phrases |
|---|---|
| pop / indie | "an indie pop song" · "a synth-pop anthem" · "a bedroom pop ballad" |
| rock / metal | "a post-punk song" · "a desert blues track" · "a math-rock instrumental" |
| hip-hop / rap | "a boom-bap hip-hop track" · "an amen-break hip-hop beat" · "a UK garage-influenced rap" |
| electronic / dance | "a melodic techno track" · "a future bass drop" · "a UK garage record" |
| film / score | "a cinematic string score" · "a wistful orchestral cue" · "a pulsing electronic underscore" |
| acoustic / roots | "an acoustic folk ballad" · "a Texas outlaw country track" · "a bossa nova in the Tom Jobim style" |
| vocal-specific | "a jazz vocal in the Billie Holiday style" · "a bluegrass with tenor lead" |
| instrumental | "a [BPM] [style] instrumental, driven by [instrument]" (set is_instrumental) |
Blend rule: base genre + at most one fusion ("orchestral phonk" ✓ · "jazz-funk-metal-gospel-techno" ✗) + a texture/era word ("1980s, tape-warm").
B.8 BPM bands (Music 3.0) — §6.7
| Feel | BPM | Example phrase |
|---|---|---|
| ballad / downbeat | 60–80 | "a 72 BPM slow-burn" |
| groove / mid | 90–110 | "a 100 BPM mid-tempo groove" |
| energetic / pop | 110–130 | "a 124 BPM lift" |
| intense / club | 130+ | "a 140 BPM drive" |
| rap specific | 70–95 | "a 85 BPM boom-bap" |
(BPM is a steer, not a dial — treat it as "about this tempo," §9.5.)
B.9 Vocal direction (Music 3.0) — §6.4, §7.4
| Tier | Phrases |
|---|---|
| timbre | "slightly aged, warm" · "velvety with a light rasp" · "gravelly, weathered" · "crystal-clear, youthful" |
| delivery | "restrained, intimate phrasing" · "belted with open vowels" · "spoken-sung, rhythmic" · "ragged, on the edge of breaking" |
| technique | "expressive vibrato on held notes" · "falsetto flips on the hook" · "ad-libs and melodic scats" · "stacked harmonies a third below" |
| space | "close-miked, dry and intimate" · "drenched in hall reverb" · "a delay echo on the last word of each line" · "light autotune sheen" |
| gender/age | "a female vocal" · "a boyish vocal" · "a duet: [lead] + [harmony role]" · "a lone soprano, then a backing choir" |
B.10 Dynamics / structure tags (Music 3.0) — §6.9, §7.3
| Tag | Modifier | Means |
|---|---|---|
[Intro] | — | the runway (or omit → starts on the lyric, §7.8) |
[Verse 1] [Verse 2] | — | the story (brief what differs between them) |
[Pre-Chorus] | — | the tension |
[Chorus] | – with lift | the thesis, bigger |
[Bridge] | – stripped | the room emptied, before the final |
[Final Chorus] | – with lift / – softer, layered | the climax system |
[Outro] | — | the defined landing (put the last line here) |
| free events | [Sudden Break] · [Building Intensity] · [Stripped Back] · [Massive Finale] · [Instrumental] · [Ad-lib] · [Humming] | one modifier per tag |
B.11 The exclusion set (Music 3.0) — §7.6 (≤3, at the end)
no drums · no autotune · no electronic elements · no male vocals · no wordless intro; the vocal enters on the first bar · no reverb — always pair each with a replacement ("no drums; the groove is a fingerpicked bass").
B.12 Character / subject slots (H3) — §3.4, §3.13
| Slot | Content |
|---|---|
| face | "a man in his 40s, short salt-and-pepper beard, tired eyes" |
| build / style | "a tall woman, black hoodie, red chest lettering, silver hoop earring" |
| prop / vehicle / creature | material + one distinguishing feature + scale cue |
| the retention line (ref work) | "From Image 1, retain: [the exact features]; [how it coexists with the new scene]." |
Feature-first, mood-second, never emotion-on-face in a video brief.
B.13 Mix vocabulary (Music 3.0) — the `and [the mix…]` slot, §6.6
| Space / feel | Phrase |
|---|---|
| big / modern | "a wide, open mix" · "a big, wide production" |
| small / intimate | "a close, dry, intimate mix" · "a narrow mono mix" |
| real / recorded | "a natural, live room mix" · "small-room acoustics with room tone" |
| controlled (electronic) | "a clean, controlled low end" · "mono-stable low end, rounded harmonics" |
| warm / vintage | "a warm, tape-saturated mix" · "analog warmth, gentle preamp drive" |
| textured / haunted | "a spacious, reverberant mix" · "drenched in hall reverb" |
| finished / radio | "a polished, balanced mix" (the default — name it only when you want the default) |
B.14 Texture / era words (Music 3.0) — the anti-gravity ballast, §7.1a
1980s, tape-warm · 70s, analog · 90s, vinyl-crackle · lo-fi, dusty · clean, digital · one-take · small-room · live-to-two-track · broadcast · bedroom, DIY · honky-tonk, dry · cathedral, wide
(Pair one era word + one texture word with the genre — that is the specificity-is-ballast move, two words long.)
B.15 Mood / emotion words (Music 3.0) — the affect slot, §7.1c
wistful · hopeful · defiant · tender · eerie · triumphant · nostalgic · restless · serene · bold · playful · aching · euphoric · resigned
*(Emotion words go in the mood/arc slots; physics words (B.13) go in the mix slot; structure words (B.10) go in the arc clause. One vocabulary per slot — §7.7.)*
B.16 The slot-grammar master (both engines, one table, §7.7)
| Slot | H3 (video) | Music 3.0 |
|---|---|---|
| container | duration + aspect, sentence one (prose) | — (the engine owns length) |
| identity | style/medium, sentence two | genre + BPM + era, sentence one |
| structure | beats + transitions | arc clause + section tags |
| subject | feature card (no emotion) | vocal character (4 tiers) |
| sound | overall_soundscape (diegetic) | mix slot + performance tier |
| score | non_diegetic_music (1 sentence) | the whole engine |
| constraints | transitions + one world-law | exclusions (≤3, last) |
The universal row: one sentence = one instruction, in the right slot, at the right position. That single line is the book.
Appendix C — Prompt Libraries (paste-ready, by category)
~80 production-grade starting briefs. Each H3 prompt is a complete one-paragraph brief (container → style → subject → action → camera → sound → music); each music prompt is the §6.11 two-lever shape. Swap the bracketed parts for your content. These are starting briefs — pin-and-nudge (§9.4) from here.
C.1 Product / e-commerce (H3)
5 seconds, 9:16. Minimalist product film, matte charcoal void, one soft key from the left, single soft shadow. A brushed-steel thermos, brushed grain visible, a small cream logo on the lid. The lid opens slowly, steam rises. Camera static, product centered, generous negative space. Sound: near-silent room; the soft click of the lid, foreground. (Music omitted.)6 seconds, 9:16. Minimalist product film, pale grey seamless, soft top light. A glass water bottle, fine embossed lettering, condensation on the glass. Water is poured in a single unbroken stream, the surface rises and settles. Camera slowly pushes in. Sound: the pour, foreground; a faint room. Music: a two-note minimal cue underneath.5 seconds, 1:1. Minimalist product film, warm ivory background, soft window light. A linen tote bag, raw hem, woven texture. The bag lifts off the counter and settles, fabric settling once. Camera static. Sound: a soft fabric whisper, foreground. (Music omitted.)6 seconds, 9:16. Product film, matte black plinth on dark slate, one cool rim light. A wireless earbud case, matte graphite, the seam catching the rim. The case opens, the earbuds visible, a tiny status light blinks once. Camera static, close. Sound: a precise mechanical click, foreground. Music: a single low synth pulse under the open.5 seconds, 9:16. Minimalist product film, pastel pink seamless, soft even light. A ceramic mug, glaze pooling at the lip, a hand-lettered word on the side. The mug rotates slowly on the counter, one full turn. Camera static. Sound: near-silent; a faint ceramic tick at the end of the turn. (Music omitted.)6 seconds, 9:16. Product film, warm wood counter, golden-hour side light. A leather wallet, grain visible, single brass snap. The wallet opens to a spread, then closes. Camera low angle, static. Sound: the leather creak, foreground. Music: a warm two-chord guitar cue.5 seconds, 1:1. Minimalist product film, pale blue void, soft key. A folded knit beanie, soft yarn, a small woven label. The beanie unfolds and rises like a slow wave, settling into shape. Camera static. Sound: a soft fabric hush, foreground. (Music omitted.)6 seconds, 9:16. Product film, matte white seamless, one hard small shadow. A wristwatch, brushed case, dark dial. The second hand sweeps one full rotation, the dial face catching a moving glint. Camera macro, static. Sound: a soft mechanical tick each rotation; quiet room. (Music omitted.)5 seconds, 9:16. Minimalist product film, deep green background, soft window light. A glass perfume bottle, faceted, amber liquid. The bottle tilts slightly, the liquid sloshing and settling. Camera static, close. Sound: a soft liquid slosh, foreground. Music: one sustained high note under.6 seconds, 9:16. Product film, pale concrete, soft top light. A running shoe, mesh and foam, a single color accent. The shoe lifts off the floor and hovers, one slow rotation. Camera static, low. Sound: a soft air whoosh, foreground. Music: a light beat, one bar of it.5 seconds, 1:1. Minimalist product film, cream seamless, warm soft light. A candle, cream wax, a small paper label. The flame lights and sways once, the label glowing. Camera static, close. Sound: a match strike, foreground; faint room. (Music omitted.)6 seconds, 9:16. Product film, slate background, cool key + warm practical. A ceramic coffee grinder, brass collar. A hand cranks the handle twice, grounds fall, the hopper emptying. Camera static. Sound: the crank ratchet, foreground. Music: a soft two-chord cue.
C.2 Brand / founder / place (H3, §10.1)
12 seconds, 16:9. Cinematic live-action, 35mm anamorphic, shallow depth of field, warm morning window light from the left. A small one-room studio, laptops, plants, a mug with steam. A founder sits at the desk, looks at the camera, smiles once, and begins to work. Camera static, medium. Sound: quiet room, soft keyboard, one distant dog bark. Music: warm lo-fi, understated, lifting at the end.15 seconds, 16:9. Cinematic live-action, 35mm, golden hour, a street that becomes a café. A man walks down the lane; (dissolve) the same man steps through the café door, the bell rings. Camera tracks him the whole way. Sound: street ambience to room tone, the door bell on the dissolve. Music: a wistful guitar line.12 seconds, 9:16. Cinematic live-action, 35mm, night city, neon practicals, rim light. A woman in a rain jacket crosses a lit crosswalk, umbrella, reflections on wet glass. She stops, looks up at the lights. Camera slow push-in. Sound: rain on umbrella, distant traffic. Music: dark lo-fi, understated.15 seconds, 16:9. Cinematic live-action, 35mm, a sunlit market, warm tungsten practicals, mixed temperatures. A vendor arranges fruit in a single deliberate line; a customer hands over cash; a small exchange of words. Camera one slow track. Sound: market ambience, the crinkle of a bag. Music: a warm acoustic cue.10 seconds, 16:9. Cinematic live-action, 35mm, an empty studio at dawn, one window, cool light. The room is still; a single chair faces the camera. A figure sits in it, once. Camera static, wide. Sound: room tone, one floor creak. Music: a single sustained string note.12 seconds, 1:1. Cinematic live-action, 35mm, a kitchen, warm morning light. A parent and child stand at the counter; the child hands over a drawing; the parent pins it to the fridge. Camera static, medium two-shot (no contact beyond the handoff). Sound: kitchen ambience, the pin thud. Music: a warm two-chord cue.15 seconds, 16:9. Cinematic live-action, 35mm, a rooftop at dusk, one sodium street lamp, mixed practicals. Two silhouettes lean on the railing, one lifts a phone, the screen light on their face. Camera slow pull-back to the skyline. Sound: wind, distant city. Music: a slow ambient pad.10 seconds, 9:16. Cinematic live-action, 35mm, a small office, afternoon window. A designer moves a rendered still from a tablet to a wall frame — the frame on the wall matches the tablet exactly. Camera static. Sound: the soft tap of the frame. Music: a light cue.
C.3 Cars × Music (the IG beat-sync theme, H3 + Music 3.0)
15 seconds, 9:16. Cinematic live-action, 35mm, a coastal road at golden hour, wet asphalt, long shadows. A classic red sports car drives toward the camera, one continuous run, wheels turning, road unspooling. Camera static, low, the car filling the frame. Sound: a deep engine note and tire wash, foreground; wind. (Music: Path B — lay the 3.0 track; brief the run to hit the drop at frame-center.)15 seconds, 9:16. Cinematic live-action, 35mm, a desert highway at night, one sodium headlight pair, dust. The headlight pair approaches, a matte black SUV emerges from the dust. Camera static, wide. Sound: a low engine growl, foreground; desert hush. (Music: Path B — the track's verse under the run, chorus as it emerges.)15 seconds, 9:16. Cinematic live-action, 35mm, a rain-soaked city street, neon reflections, mixed practicals. A white sports car takes a corner, water spraying, tires gripping. Camera one slow track. Sound: tires on wet asphalt, foreground; rain. (Music: Path B — the drop lands on the corner.)15 seconds, 9:16. Cinematic live-action, 35mm, an empty parking structure, cool fluorescent, one warm car interior. The car's door opens, a silhouette steps out, the interior light spills. Camera static, medium. Sound: the door, the footfall, the engine idle. (Music: Path B — a single sustained note under.)15 seconds, 9:16. Cinematic live-action, 35mm, a mountain pass at dawn, alpenglow, thin fog. A vintage sedan crests the pass and stops; the engine ticks as it cools. Camera slow pull-back. Sound: the engine ticking, foreground; mountain wind. (Music: Path B — a wistful string cue.)15 seconds, 9:16. Cinematic live-action, 35mm, a tunnel, one sodium light pair, the walls repeating. The car runs the tunnel, lights strobing off the wet wall. Camera static, head-on. Sound: tire slap in the tunnel, reverbed; engine. (Music: Path B — a pulse track, the strobe on the beat.)
C.4 Hand-drawn × Live (the Café's signature, H3, §10.2)
15 seconds, 9:16. Live-action night street fused with hand-drawn glowing animation, simple white line-work, deep blue base grade, cyan trails. A busy crowd walks, real and sharp. (dissolve) glowing line-figures drift above the crowd, bobbing, leaving short fading trails that weave between heads. What stays real: the rain, the wet ground; what becomes drawn: the people. Camera slow push-in. Sound: city rain in the background; each line-figure makes a small chime, panning with the figure, foreground. Music: dark ambient pulse, curious, building with the trail density.15 seconds, 9:16. Live-action café interior fused with hand-drawn glowing animation, warm white lines, deep warm base, amber trails. A barista pours a latte, real; (dissolve) the steam becomes a drawn swirl that rises and loops, a small drawn creature living in the steam. Real: the cup, the counter. Drawn: the steam-creature. Camera static, medium. Sound: the pour, foreground; the steam-creature a soft whistle. Music: warm ambient, understated.15 seconds, 9:16. Live-action crosswalk at night fused with hand-drawn glowing lines, white line-work, deep teal base, green trails. The lights turn; the crowd is real; (dissolve) each crossing person leaves a glowing line-trail that lingers and fades. Real: the rain, the asphalt. Drawn: the trail-people. Camera high, static. Sound: wet feet, foreground; trails a faint hum. Music: a dark pulse, tense-release with the light change.15 seconds, 9:16. Live-action bedroom at night fused with hand-drawn glowing lines, pale blue lines, deep base, soft trails. A girl sleeps, real; (dissolve) a drawn owl orbits her desk lamp, bobbing, its trail a faint glow on the wall. Real: the room, the lamp. Drawn: the owl. Camera static, slow push. Sound: room tone; the owl a soft wing-flick. Music: a music-box ambient, wistful.
C.5 Papercraft / stop-motion explainers (H3, §10.4)
12 seconds, 9:16. Cut-paper papercraft stop-motion, three depth layers, matte pastel palette (butter, mint, coral, slate, cream), soft top light, shadows between layers, visible paper edges. A paper city on the mid layer; a paper train runs the mid layer in slow, slightly jerky stop-motion ticks, its paper wheels ticking. A paper sun rises in the back layer. Camera static, all layers. Sound: paper creaks and wheel ticks, foreground. Music: a music-box cue, wistful.12 seconds, 9:16. Cut-paper papercraft, three layers, matte pastels, soft top light. A paper heart on the mid layer fills slowly from the bottom with a paper wave; the wave's crest ticks upward. The back layer is a paper sky. Camera static. Sound: a soft paper hush as the wave rises. Music: a gentle two-note loop, lifting as the heart fills.12 seconds, 9:16. Cut-paper papercraft, three layers, matte pastels, soft top light. A paper book opens in the mid layer; paper birds fly out of its pages into the front layer, in jerky ticks. Camera static. Sound: the page turn, foreground; soft wing flaps. Music: a glockenspiel cue, playful.12 seconds, 9:16. Cut-paper papercraft, three layers, matte pastels, soft top light. A paper plant on the mid layer grows: stem, leaf, bloom, each in slow ticks; the bloom opens last. Camera slow push-in. Sound: a soft paper unfurl on the bloom. Music: a warm two-chord cue, resolving on the bloom.
C.6 3D-stylized / mascot (H3 + Ref2VA, §10.5)
10 seconds, 9:16. 3D stylized, toy-grade soft shapes, subsurface skin, exaggerated big eyes, clean studio light with one cyan rim, saturated palette. A small round robot, matte white, one big single eye, stubby arms. It bounces twice with a soft squash-and-stretch, then waves one stubby arm. Camera static, centered. Sound: a soft studio room; two plucky boings on the bounce, foreground. Music: a playful bouncy cue, one bar.10 seconds, 9:16. 3D stylized, toy-grade, subsurface, exaggerated, clean studio + warm rim. A chubby round cat-robot, cream fur, one big blue eye, a small backpack. It tilts its head, then hops once. Camera static. Sound: a soft thud on the hop, foreground. Music: a plucky cue. (Ref: Image 1 = the character sheet; retain: cream fur, one blue eye, the backpack.)12 seconds, 9:16. 3D stylized, toy-grade, subsurface, saturated, clean studio. A small robot barista, matte white, one eye, a tiny apron. It pours a paper cup with a steady stream, the foam rising. Camera static, close. Sound: the pour, foreground; a soft "ding" at the fill line. Music: a light bouncy cue. (Ref: Image 1 = the character; retain: the apron, the one eye, the matte white.)
C.7 "Can AI [X]?" (YouTube, 15s 9:16 — one question, one proof)
15 seconds, 9:16. Cinematic live-action, 35mm, a window seat, morning light, rain. A person sips tea, looks out; (dissolve) the same person, the same cup, now in a slightly different room, same gesture. Camera static. Sound: rain, the cup on the saucer. Music: a warm cue. (Question on screen: "Can AI hold a character across a change of place?")15 seconds, 9:16. 3D stylized, toy-grade, subsurface, clean studio. The same round robot (Ref Image 1) in three outfits, each a (cut), same build, same eye. Camera static. Sound: three soft thuds, one per cut. Music: a plucky cue. (Question: "Can AI keep *the same* character?")15 seconds, 9:16. Minimalist product film, matte void, soft key. A glass; (cut) the glass becomes a metal one; (cut) the metal becomes a ceramic one — same shape, same one soft key, same single shadow. Camera static. Sound: a soft set-down each cut, foreground. (Question: "Can AI hold *one* object across materials?")15 seconds, 9:16. Cinematic live-action, 35mm, a piano, warm practicals. A hand plays a chord; the notes rise as small glowing paper notes into the room. Camera slow push. Sound: the chord, foreground. (Question: "Can AI make sound you can *see*?")15 seconds, 9:16. Cut-paper papercraft, three layers, matte pastels. A paper moon sets in the back layer; (dissolve) a paper sun rises; the mid-layer clock ticks once. Camera static. Sound: the clock tick, foreground. (Question: "Can AI *animate a metaphor*?")15 seconds, 9:16. Live-action city at night fused with hand-drawn lines, white lines, deep base. (dissolve) the whole crowd becomes line-figures at once. Camera slow pull-back. Sound: the city hush; a soft collective hum. (Question: "Can AI turn *reality* into a drawing?")15 seconds, 9:16. Cinematic live-action, 35mm, a single long take, a hand writes a word on fogged glass, the word stays crisp. Camera static, close. Sound: the marker on glass, foreground. (Question: "Can AI write *legible* text?")
C.8 Music 3.0 — tracks (prompt + lyrics shape; §6, §7)
- brand anthem —
prompt: A warm, confident indie pop song at 118 BPM, verses intimate and sparse, a wide anthemic chorus with a lift, with a clear warm female lead vocal, fingerpicked guitar, and a wide open mix. lyrics: [Intro] / [Verse 1] (4 lines) / [Pre-Chorus] / [Chorus – with lift] / [Verse 2] / [Final Chorus – with lift] / [Outro] - lo-fi study —
prompt: A mellow lo-fi hip-hop instrumental at 82 BPM, steady and loop-friendly, with a soft upright bass and a clean, dry mix, no lead vocal. (is_instrumental) - desert blues —
prompt: A gritty 70s-influenced desert blues track at 96 BPM, verses sparse and dry, the chorus opening wide, with a slightly aged male vocal, a Twangy Telecaster, and a dry honky-tonk mix, no autotune, no electronic drums. lyrics: [Verse 1] / [Chorus] / [Verse 2] / [Bridge – stripped] / [Final Chorus – with lift] - funk drop —
prompt: A funky groove at 108 BPM, a slap bass and a wah-tele lead, verses tight, the drop wide and loud, with a rhythmic male vocal, and a punchy room mix. lyrics: [Verse] / [Drop – with lift] / [Verse] / [Drop – massive] - synth-pop anthem —
prompt: A 1985-influenced synth-pop anthem at 122 BPM, an upsolving chorus, with a bright female vocal, a drum machine, and a big gated-reverb mix. lyrics: [Verse] / [Pre] / [Chorus – with lift] / [Bridge] / [Final Chorus – with lift] - bossa —
prompt: A sunny bossa nova at 110 BPM, relaxed and unhurried, with a warm female vocal, a nylon guitar, and a close intimate room mix. lyrics: [Verse] / [Chorus] / [Instrumental] / [Chorus] - post-punk —
prompt: A 1980s post-punk track at 128 BPM, verses tense, the chorus a cold wide hook, with a detached male vocal, a fuzz bass, and a dry mix, no autotune. lyrics: [Verse] / [Chorus – with lift] / [Verse] / [Bridge] / [Final Chorus] - folk one-take —
prompt: A one-take folk ballad at 72 BPM, a slightly aged warm vocal close-miked, a fingerpicked acoustic, small-room acoustics with room tone, and a narrow mono mix, no autotune. lyrics: [Verse] / [Verse] / [Chorus – softer] / [Outro] - ukg record —
prompt: A UK garage record at 132 BPM, a shuffled groove, the drop wide, with a clear male vocal, a drum machine and a sub bass, and a wide club mix. lyrics: [Verse] / [Drop] / [Verse] / [Drop – massive] - cinematic cue (score) —
prompt: A wistful cinematic string score, a slow build from solo cello to full strings, wide and open, no vocal. (is_instrumental)(the Café's default brand underscore — the §8.4 spine) - music-box —
prompt: A wistful music-box instrumental, slow, a single bell melody over a soft pad, dry and close, no drums. (is_instrumental)(papercraft default — §10.4) - pluck (mascot) —
prompt: A playful bouncy instrumental at 120 BPM, a plucky marimba over a light beat, dry and close, no vocal. (is_instrumental)(3D-stylized default — §10.5) - dark pulse (hand-drawn) —
prompt: A dark ambient electronic pulse at 100 BPM, a sub bass and a slow tremolo pad, tense and curious, wide, no vocal. (is_instrumental)(hand-drawn×live default — §10.2) - spoken-word bed —
prompt: A sparse cinematic underscore at 70 BPM, a low pad and a single high string, understated, leaving space for a voice, no vocal. (is_instrumental) - title leitmotif —
prompt: A 5-second-style music-box title cue, a three-note motif, dry, no drums, no vocal. (is_instrumental)(character-led menu — §10.7) - bossa-parents —
prompt: A warm bossa-nova duet at 108 BPM, a parent and child alternating, the child's vocal boyish, the parent's warm, a nylon guitar, a close intimate mix. lyrics: [Verse – parent] / [Chorus – child] / [Verse – parent] / [Chorus – both, layered](the official demo, adapted)
C.9 Café social proofs (the flywheel content, H3)
5 seconds, 1:1. Minimalist product film, matte void. The Solligence wordmark assembles from light, the yellow S tile landing last. Camera static. Sound: a soft tick per letter, foreground. (Music: one note.) (X prompt→render post: the prompt is the text.)10 seconds, 9:16. [any §10 style]. [the idea]. [one action]. (The on-screen text overlay "5 minutes, one prompt" is added in the edit; the render is the receipt.)15 seconds, 9:16. [the §11.5 week's hero]. [style]. [the idea]. [action]. Camera [move]. Sound [diegetic]. Music [one cue]. (The Founder Pass render — the founder's brief, their story, the standard tail in the edit.)5 seconds, 9:16. A hand taps a phone; a tiny 3D cup renders above it in a second. Camera static, close. Sound: the tap, foreground. (Music: one pluck.) ("a coffee, rendered" — the tagline as a render.)10 seconds, 1:1. A prompt (as on-screen text) fills a frame; (dissolve) the rendered result it describes appears. Camera static. Sound: a soft "materialize" whoosh. (Music: one cue.) (the literal Prompt→Render post — the whole product in 10 seconds)
C.10 More music — the full genre × mood grid (Music 3.0, prompt-only)
A dreamy ambient pop song at 90 BPM, verses hushed, the chorus airy and wide, with a soft airy female vocal, a piano and a wide open mix, no autotune. lyrics: [Verse]/[Chorus – with lift]/[Bridge – stripped]/[Final Chorus – softer, layered]A 2000s-influenced rock anthem at 140 BPM, a massive final chorus, with a raspy male vocal, a driven electric guitar, and a big wide mix, no autotune. lyrics: [Verse]/[Pre]/[Chorus – with lift]/[Bridge]/[Final Chorus – with lift]A smooth R&B groove at 92 BPM, verses laid-back, the hook silky, with a sultry male vocal, a clean electric piano, and a warm close mix, no autotune. lyrics: [Verse]/[Hook]/[Verse]/[Hook]A punk-pop burst at 160 BPM, loud and fast, the hook shouted, with an energetic male vocal, a fuzzy guitar, and a dry punchy mix. lyrics: [Verse]/[Chorus – with lift]/[Verse]/[Chorus – massive]A dark trap at 140 BPM, a hard 808 groove, the hook moody, with an autotuned male vocal, a sub bass, and a wide club mix. lyrics: [Hook]/[Verse]/[Hook]/[Bridge]/[Final Hook – with lift]A gospel-inflected soul song at 100 BPM, a big singalong chorus, with a powerful female vocal, a church-organ-adjacent electric piano, and a wide choral mix. lyrics: [Verse]/[Chorus – with lift]/[Bridge – stripped]/[Final Chorus – with lift]A folksy singer-songwriter ballad at 78 BPM, a small intimate room, with a warm female vocal close-miked, a nylon guitar, and a dry close mix, no autotune. lyrics: [Verse]/[Chorus – softer]/[Verse]/[Outro]A desert reggae at 80 BPM, a laid-back offbeat groove, the hook easy, with a relaxed male vocal, a skank guitar, and a warm room mix. lyrics: [Verse]/[Hook]/[Instrumental]/[Hook]A symphonic metal instrumental at 150 BPM, a soaring main riff, aggressive and epic, wide and big, no vocal. (is_instrumental)A lo-fi study beat at 80 BPM, tape-warm, a dusty vinyl crackle, a soft upright bass, dry and close, no vocal. (is_instrumental)A cinematic trailer hit, a 30-second-style build to one massive brass-and-drum hit, no vocal. (is_instrumental)A waltz at 150 BPM (3/4 feel), a music-box and a solo cello, wistful and small, dry, no vocal. (is_instrumental)(music-box 3/4 — the papercraft variant)A 1980s new-wave track at 118 BPM, an icy hook, with a detached male vocal, a gated drum machine, and a cold wide mix. lyrics: [Verse]/[Chorus]/[Bridge]/[Final Chorus]A K-drama-style ballad at 88 BPM, a tearjerker chorus, with a clear female vocal, a piano and a string swell, and a wide mix. lyrics: [Verse]/[Pre]/[Chorus – with lift]/[Bridge – stripped]/[Final Chorus – with lift]A phonk at 130 BPM, a cowbell lead, dark and driving, no vocal, wide and punchy. (is_instrumental)(the cars×music drop — the Path-B track for §C.3)A melodic house at 124 BPM, a euphoric drop, a pad and a pluck, no vocal, wide. (is_instrumental)(the second cars×music option)A strumming acoustic funk at 112 BPM, a bright positive groove, with an upbeat female vocal, a Telecaster, and a dry room mix. lyrics: [Verse]/[Chorus – with lift]/[Bridge]/[Final Chorus]A rain-soaked lo-fi folk at 84 BPM, mellow and grey, with a soft male vocal, a felted piano, and a close dry mix, no autotune. lyrics: [Verse]/[Chorus – softer]/[Outro]A big-band-adjacent swing at 132 BPM, a horn section and a walking bass, a brass hook, no vocal, wide and roomy. (is_instrumental)A lullaby at 66 BPM, a music-box and a soft harp, a single high melody, dry, no vocal. (is_instrumental)(the sleep/child-content default)
C.11 More H3 — one-line scene generators (each: container + style + subject + action + camera + sound)
5 seconds, 9:16. Cinematic live-action, 35mm, a rainy window, cool light. A single droplet races down the glass, then stops. Camera static, macro. Sound: one tap, foreground; rain hush.5 seconds, 9:16. Cinematic live-action, 35mm, a dim kitchen, one warm bulb. Steam from a kettle rises and vanishes. Camera static, close. Sound: the kettle, foreground.5 seconds, 1:1. Cinematic live-action, 35mm, a bookshelf, warm side light. A hand pulls one book out, turns its cover to the camera. Camera static. Sound: the spine, foreground. (Music omitted.)5 seconds, 9:16. Cinematic live-action, 35mm, a stage, one spotlight, haze. A microphone stands alone; (the light) flickers once. Camera static. Sound: room ring, foreground.5 seconds, 9:16. Minimalist product film, matte void. A paper plane rises off a desk and glides. Camera static. Sound: a soft paper whisper, foreground. (Music omitted.)5 seconds, 1:1. 3D stylized, toy-grade, soft studio. A rubber duck bobs once on a matte pool. Camera static. Sound: a soft plop, foreground. (Music omitted.)5 seconds, 9:16. Cut-paper papercraft, two layers, pastels. A paper envelope opens; a paper letter floats up. Camera static. Sound: the paper, foreground. (Music: music-box, one note.)5 seconds, 9:16. Live-action bedroom fused with hand-drawn lines, pale lines, deep base. (dissolve) a drawn star floats up from the desk lamp. Camera static. Sound: a soft chime, foreground.5 seconds, 9:16. Cinematic live-action, 35mm, a coffee cup, warm key. The crema surface ripples once, then stills. Camera macro, static. Sound: a soft pour-finish, foreground.5 seconds, 1:1. Cinematic live-action, 35mm, a vinyl turntable, warm practicals. The needle drops; the dust on the record glints. Camera static, top. Sound: the drop, the low rumble, foreground.5 seconds, 9:16. Cinematic live-action, 35mm, a bicycle wheel, golden hour. The wheel spins to a stop, spokes blurring to still. Camera low, static. Sound: the hub, foreground; wind.5 seconds, 9:16. 3D stylized, toy-grade, clean studio. A paper cup jiggles on a counter, steam a small loop. Camera static. Sound: a soft thud, foreground. (Music: one pluck.)5 seconds, 1:1. Minimalist product film, pale seamless. A lipstick cap twists up, the product rising. Camera static, macro. Sound: the twist, foreground. (Music omitted.)5 seconds, 9:16. Cinematic live-action, 35mm, a train window, passing fields, soft light. A child's hand prints once on the glass. Camera static, close. Sound: the rail rhythm, foreground.5 seconds, 9:16. Cut-paper papercraft, three layers. A paper clock's hands move one tick forward, slowly. Camera static, close. Sound: one tick, foreground. (Music: one note.)5 seconds, 9:16. Live-action city crossing fused with hand-drawn lines, white lines, deep base. (dissolve) the zebra stripes glow. Camera high, static. Sound: a soft hum, foreground.
C.12 The Café's own recurring renders (the brand's standing assets)
- The S-tile mark (5s, 1:1) —
A matte yellow square on deep charcoal; a black 'S' assembles from a single continuous line, drawing itself in one unbroken stroke; the square settles. Camera static. Sound: one soft tick at the line's end, foreground. (Music omitted.)(the brand's 5-second signature — §11.1's "every render is an ad" made literal) - The UPI tap (5s, 9:16) —
Close-up, warm light: a thumb taps a phone screen; a small 3D coffee cup renders above the screen in under a second, steam included. Camera static, close. Sound: the tap, foreground; one soft pluck as the cup appears. (Music: one note.)("₹99/hr, UPI" as a visual — the offer in five seconds) - The 10-minute clock (10s, 1:1) —
A minimal clock face, matte, one warm key; the minute hand sweeps exactly one full turn, steady, while a small scene (any hero) plays behind it. Camera static. Sound: one soft tick per quarter, foreground. (Music: a slow two-note loop.)("10 minutes free" as a render — the trial, visualized) - The café counter (15s, 9:16) —
Cinematic live-action, 35mm, a warm small counter, morning light, a single cup of coffee, steam. A hand places a phone on the counter; the phone screen shows a render assembling in real time. Camera slow push-in. Sound: the phone set-down, foreground; quiet room. Music: warm lo-fi, understated, lifting at the end. (End-card in the edit: "Made in Solligence Café · UPI · ₹99/hr · 10 min free")(the one that is the standard tail — the brand's home video)
*That is the 100. Every entry is a starting brief — run §9.4's pin-and-nudge on the one element that matters most to the brand. The C.1–C.9 groups are the production core; C.10–C.12 are the standing assets and the one-line generators. Swap the brackets, keep the grammar.*
Appendix D — Sources, License & Version Notes
D.1 The source inventory (what the book is built on)
Tier 1 — official, model-verified (every model claim in the book traces here):
| Source | What it supplies |
|---|---|
MiniMax H3 open-source repository — github.com/MiniMax-AI/MiniMax-H3 (Apache 2.0) + its two official prompting guides (h3-video-base-prompt-guide.md, h3-video-ref-prompt-guide.md, Chinese + English) | §1 laws' video counterparts; §2 identity/modes/architecture; §3 all syntax (integrated_multimodal_description / overall_soundscape / non_diegetic_music, ref format, subject/motion/light/camera rules, 768P→2K, Context-IR) |
| MiniMax blog: "MiniMax H3" (launch) | §2 market position (quality-per-dollar framing) |
| MiniMax blog: "H3 Feature Highlights" | §4 all six use cases (Brand Film, Visual Creative, AI Narrative, Product/E-commerce, Digital Experience, Animation) |
MiniMax H3 API documentation (platform.minimax.io/docs) | §2.5 spec sheet (model h3, parameters, 4–15s, 768P/1080P/2K, reference budgets, pricing position) |
MiniMax Music 3.0 open-source repository — github.com/MiniMax-AI/MiniMax-Music3 (Apache 2.0) | §5 identity, architecture (Structured Caption concept, Hybrid-LM global context, Flow-VAE continuous RVQ + AudioX2-style vocoder lineage), specs, open-weights posture |
| MiniMax blog: "MiniMax Music 3.0" (launch) | §5 demos (Mandayn folk, bilingual parental duet, acoustic bossa nova, electronic Pop), the Structured Caption narrative, the vocal engine's capabilities (breath/pronunciation/falsetto/harmonies, "no digital hiss") |
MiniMax Music API documentation (platform.minimax.io/docs) | §5.3 API surface (music_generation, lyrics_generation, cover, audio_preprocess; fields incl. is_instrumental, lyrics_optimizer, stream, prompt ≤1000 chars, lyrics ≤5000 chars, outputs 44.1kHz/256kbps), the deprecation timeline |
Tier 2 — the broader ecosystem (informs technique; flagged as such where used):
| Source | What it supplies |
|---|---|
| OpenAI's Sora 2 Prompting Guide (official cookbook) | the "one sentence = one instruction" and "multi-subject = multi-sentence" syntax; the strip-back recovery ladder (§9.3); the physical-recording realism stack (§7.5) |
| Suno Field Guide (community, ~92k chars) | the style-mesh model, gravity wells, genre clouds, adherence decay/front-loading, the two-lever split, section-tag + dynamic-modifier grammar, the section-structure / exclusion / mid-song / prompt-length findings, the syllable and one-metaphor lyric rules (§§6.6–6.11, 7.1–7.9) |
| Runway / Kling / Luma / Veo public prompting guidance (search summaries) | the universal core (subject first, action+camera, lighting+shot type, style, audio, negative keywords) that Law Three / Law Four generalize across vendors |
The honesty line (how to read Tier 2): community findings were mass-tested mostly on Suno, not on Music 3.0. They are used in this book as architectural priors — behaviors explained by how LLM-conditioned music models work — and each such section says "verify on your model." When MiniMax's own official examples were available (the four launch demos), they anchor the section. Nothing in the book presents a Tier-2 guess as a Tier-1 fact.
D.2 A note on the models' moving targets (read this yearly)
Both models are fast-moving, young models (H3 and Music 3.0 are both 2025–2026 releases, both open-weights, both iterating). This book is written to survive version drift by design:
- The Laws (§1) are the stable layer. They describe why these models behave as they do (how text-conditioned generation works) — that layer does not version.
- The syntax (§§3, 6) is the stable-most current layer — it is pinned to the official format of the version in Appendix D.1. If a new version ships a changed format, re-verify §§2–3 and §5–6 against the new official guides; everything else (Laws, failure science, style playbook, system) transfers unchanged.
- The failure table (§9) is a living document. As the four hard problems (§9.1) close one by one, remove their rows — a failure table that still lists solved problems is a lie. The table's structure (symptom → it-is → fix) is the permanent asset.
D.3 Version record
| Version | Date | Change |
|---|---|---|
| v1.0 | 2026-06-06 | First complete edition. H3 pinned to the open-source release + its two official prompting guides; Music 3.0 pinned to the open-source release + its launch demos + the current API surface. Tier-2 technique flagged per §D.1. |
(Each subsequent edition: one row. Re-verify the syntax chapters against the then-current official guides; bump the pin.)
D.4 License & use
- The models (MiniMax H3, MiniMax Music 3.0) are open weights under the Apache 2.0 license (see their respective repositories). You may self-host, modify, and ship products on them, subject to each model's license terms — read them before commercial deployment; the license is the contract, this book is not it.
- This book (the Solligence Café Prompt Bible) is published by Solligence Digital Solutions LLP for the Solligence Café (solligence.in). It may be: freely read (the pillar page at
solligence.in/prompt-bible), used in production, and cited with attribution ("the Solligence Café Prompt Bible, v1.0"). - Redistribution: the free pillar page and the paid PDF share one text. The PDF is a product (curated layout, print formats, the extended libraries) — it is sold, not freely redistributed; the pillar page is the free edition.
- The prompt libraries (Appendix C) are provided as starting briefs: you own everything you render; the examples are templates, not locked assets. Swap the brackets, keep the grammar.
D.5 One last word (the whole book in a paragraph)
These two engines are honest tools. They will do what you ask, well, and they will improvise what you didn't ask — that is not a flaw to be fought, it is the contract to be understood. Ask precisely (the Laws), shape the container (the syntax), steer by element (the dials), expect the four hard problems (the failure table), and iterate like a crew (pin-and-nudge). Do that, and a one-person studio with a phone and a UPI id ships finished, sounding shorts on a weekly rhythm — and every single render is an ad.
That is the book. Now go render.
Solligence Café · solligence.in · UPI · ₹99/hr · 10 min free
The Solligence Café Prompt Bible v1.0 — June 2026. Made in Solligence Café.
The book is free. The studio is the point.
Every prompt in this Bible runs in Solligence Café — the UPI-powered studio built on these two open engines. Ten minutes free, ₹99/hr after.
Open the StudioMade in Solligence Café · UPI · ₹99/hr · 10 min free