
Introducing References: sound control for Music v2
- Category
- Product
- Published

Music v2 has generated millions of tracks since launch. Watching what people ask for and what they keep, one pattern holds: the users who get the most out of the model make a small number of concrete decisions and state them plainly. Clever prompts lose to specific ones, over and over.
This guide covers what those decisions are. Every example below was generated with the exact prompt shown, unedited. Press play to hear it, and press Try this prompt to drop it into the composer and change it.
A prompt answers five questions, whether you intend it to or not: genre, mood, instrumentation, tempo, and production era. Any question you leave open, the model answers with the most statistically likely choice — which is to say, the most average one.
Here's a prompt that leaves all five open:
upbeat electronic track
And the same request with all five answered:
French house, 122 BPM, filtered disco sample chops, sidechained bass, tight four-on-the-floor kick, rooftop-at-sunset mood, warm analog glue
The second prompt contains no special technique. Someone knew what they wanted and wrote it down.
The rule: decide genre, mood, instruments, tempo, and era before the model decides for you.
The model understands music the way it's discussed — in studios, in credits, in arguments about the snare. Studio language moves real levers. Sidechained, close-mic'd, bone-dry, tape saturation, plate reverb: each of these audibly changes the mix.
Here's one ballad in three rooms. The song stays the same; only the production words change:
slow soul ballad, 68 BPM, female vocal — bone-dry drums, close-mic'd vocal, dead room
slow soul ballad, 68 BPM, female vocal — cavernous plate reverb, tape echo throws, gospel room
slow soul ballad, 68 BPM, female vocal — tape saturation, wow and flutter, dusty vinyl crackle
If you don't have the vocabulary, describe the space. "Sounds like it was recorded in a stairwell" will get you a stairwell, with the drips.
The rule: say it the way a producer would say it.
Give numbers. The model holds a stated BPM and key precisely enough to write against a click. The same prompt at two tempos, everything else untouched:
synthwave arpeggio anthem, F minor, 88 BPM, neon nostalgia, analog polysynth pads
synthwave arpeggio anthem, F minor, 128 BPM, neon nostalgia, analog polysynth pads
The model will take "slow" and "fast" if that's all you have. A number removes the guesswork, and the moment you plan to layer a generation with anything else — a vocal take, a sample, a second generation — you need the guesswork gone.
The rule: BPM and key are numbers, and the model respects numbers.
The model follows instructions about time. Tell it what enters when, the way you'd brief a band:
UK garage, 132 BPM — start with just a shuffled drum loop, add a warm sub bassline after four bars, then bring in chopped vocal stabs for the drop
The load-bearing words are small: start with, just, then, bring in. Notice what "just a drum loop" is doing — without the just, the model fills the silence, and it has a great deal to fill it with.
The rule: narrate the arrangement in order, and mark the silences.
The prompt decides how a lyric gets delivered. The same two lines under opposite instructions:
whispered indie folk, close and intimate, fingerpicked acoustic guitar
stadium rock, belted powerhouse vocal, huge drums, wall of guitars
Whispered, belted, conversational, deadpan, stacked harmonies — the model takes delivery notes like a session singer who never pushes back. When you want no voice at all, the word is instrumental. It's the most reliable word in the system.
The rule: tell the singer how to sing.
A loop leaves no room for surprises, so a loop prompt is an exercise in exclusion. State the bars, the BPM, the key — and state what's banned:
boom bap drum break, 90 BPM, dusty and swung, four bars, no melody — just drums
dreamy guitar loop, 140 BPM, F sharp minor, emo trap, four bars, instrumental
Both repeat seamlessly — leave them running. "No melody, just drums" is the entire reason the drum break stays a drum break. There's a library of loops made exactly this way in Sounds.
The rule: for loops, the negative space is the prompt.
Everything above treats the model as a band. It's also a sound machine — it renders sound itself, and it holds no opinion on whether the sound you're describing could physically occur.
Weather does not keep time. Nobody has told the model:
a thunderstorm in 6/8 — thunder lands on the downbeat, rain fills the offbeats, distant lightning as crash cymbals
You can tell it a sound was secretly another sound all along:
one enormous cathedral bell that, as it decays, slowly turns out to have been a choir all along
You can introduce two sounds that have never met:
a dial-up modem handshake harmonised by a gospel choir, in E flat
And you can hand it chaos with instructions to organise:
an orchestra tuning up that accidentally becomes techno — the A gets a kick drum, chaos becomes a groove
The method: take a sound that exists, give it one property it has no business having, and describe the result with a straight face. The model doesn't know these sounds are impossible, and nothing requires you to tell it. (We also asked for the inside of a grandfather clock playing swung jazz. It's very good. We're leaving that one for you.)
The rule: description outranks physics.
The sounds in Exhibit A are raw material, and every technique in this guide works on them. To prove it, we took one impossible sound and walked it to a finished track.
Step one: the raw material. A swarm, tuned. The exclusions carry this prompt — no drums, no synths, no melody. Just the swarm, given a key:
a dense swarm of bees that is, on closer listening, an F minor chord
Step two: production vocabulary, applied to livestock. One producer word — sidechained — tells the swarm to duck out of the way of a kick drum four times a bar. Watch the waveform: you can see it pump.
the same swarm of bees, sidechained — the bees duck out of the way of a four-on-the-floor kick drum, 126 BPM
Step three: the track. Every rule in this guide in a single prompt — a stated key and BPM, a narrated three-act arrangement, and an instruction for the swarm to become the lead. One generation. The same bees throughout:
dark techno built from a swarm of bees, F minor, 126 BPM. Start with just the tuned swarm. After four bars the kick arrives and sidechains it. Then the swarm sharpens into the lead synth line — rolling bassline, sparse industrial percussion, the swarm always audible, always the lead
Nothing in this build was new. Five decisions, producer vocabulary, numbers, narrated time, exclusions — the same tools as every section above, pointed at an absurd input. The techniques compose. That's the finding.
If you'd rather steer with sound than sentences: upload step one as a Reference and let the swarm itself drive the generation.
The rule: an impossible sound is a starting point.
Every era of recorded music sounds the way it does because of what was in the room — the shellac, the ribbon mics, the tape machines, the samplers, the loudness war. The model knows all of it. "Production era" was the fifth decision back in the first section of this guide, and it turns out to be a road you can drive down mid-song.
So we asked for the whole century. One melody, played forward through time, and the tape never stops:
one melody carried through a century of recording, without stopping — it starts on a crackling 1920s gramophone as hot jazz, turns into 1950s rock and roll with slapback echo, blooms into 1970s disco, gets swallowed by a 1990s jungle rave, and lands in the present as glossy hyperpop. The same tune the whole way.
Listen for the handoffs. A muted trumpet carries the tune out of the crackle; a twanging guitar takes it at the border of 1956; disco strings lift it; a rave stab chews it up; and it arrives in 2026 wearing supersaws. Ninety years of engineering, crossed in forty-six seconds, by a prompt that names five dates and holds one melody.
There's no hidden machinery here either. Naming an era, narrating time — the same two techniques from the top of this guide, at the largest scale they'll stretch to.
The rule: production era is a dial like tempo, and the model can turn it mid-song.
Some ideas resist description. The melody you hummed into your phone, the swing of one specific bassline, a texture you can only point at. References covers this: upload a track — a voice memo qualifies — and Music v2 matches its feel, groove, and palette without copying the notes.
The prompt still steers everything the reference leaves open. In practice the pairing that works is a reference that carries the feel and a prompt that says what should be different: "same energy, half-time drums, female vocal." Start from something you made yourself, and the output stops sounding like the model.
The rule: use the prompt to say what the reference should become.
One prompt using the full toolkit — five decisions, production words, a number, an arrangement, a delivery per section:
Bedroom pop about a city at 5am, 96 BPM, A minor — muted guitar chops, soft breakbeat, tape-warm bass. Verse close and conversational, chorus opens wide with stacked harmonies.
Twenty seconds to write. Every decision in the output is one a person made.
Everything in this guide reduces to making the decisions the model would otherwise make for you. In practice that means three habits:
Every prompt on this page is a starting point. Press Try this prompt on whichever is closest to the thing in your head, and make it yours from there.