Music tech · AI music · How it works

How an AI song is actually made, start to finish

Not the marketing version. The actual file structure, the actual prompts, and the four hours of unglamorous work that sit between a good idea and something you would play to somebody.

ZK

Zaib Khan

· 3 min read

People assume this is one box you type a sentence into. It is closer to four separate jobs, and the interesting parts are the joins between them. This is the process I use for every song that goes out, written down properly.

One: the brief, which is the whole thing

Before any tool is opened, I need facts. Not adjectives. The difference between a song that works and one that does not is almost entirely decided here.

Useless: they are very much in love, she is beautiful, they have been together a long time.

Useful: they met at a friend's birthday in Leeds in 2019, he spent the entire night pretending he could dance, she let him carry on, and he still calls her Doctor Sahiba because she corrects everybody, including him, especially him.

The second one contains a chorus. The first one contains nothing. Half of what our order form does is drag people from the first kind of answer to the second, which is why it asks for the story rather than for a list of qualities.

Two: writing the lyric

I draft with a language model, then rewrite by hand. The draft is useful for structure and useless for meaning. Its instinct is to reach for the nearest available image, which is always the one every other song has used.

The rule I work to: every verse must contain one thing that could not be said about any other couple. If a line would survive being moved into a different song, it comes out.

Sections get labelled explicitly, because the music model reads them:

  • [Intro], kept short, usually instrumental
  • [Verse 1], the situation, concrete and specific
  • [Pre-Chorus], the turn, where the feeling arrives
  • [Chorus], the title made literal, repeated exactly
  • [Verse 2], the complication or the passage of time
  • [Bridge], genuinely different, or the final chorus lands flat
  • [Final Chorus], often with one word changed

Three: the style prompt

This is the part with the most folklore around it and the least documentation. A style prompt is one comma-separated line describing the record you want, and it behaves less like an instruction and more like a search query across everything the model has heard.

What works, in order of how much it matters:

  1. 1Instruments, named specifically. Harmonium and tabla, not traditional Indian instruments.
  2. 2Vocal character. Warm male baritone, restrained, close-mic'd.
  3. 3Era and production. Nineties playback, wide reverb, strings slightly forward.
  4. 4Tempo and feel. Slow six-eight, unhurried.
  5. 5Mood, last and briefly. Everything before this has already done the work.

And an exclusion list, which matters more than people expect. Without one you get autotune, trap hats and a modern pop mix bolted onto something that wanted none of them.

Four: the video, and the thing that breaks it

Video generation has one failure that swallows everything else: the same person does not stay the same person between shots. Generate twenty-eight scenes independently and you get twenty-eight cousins.

The fix is unglamorous. You write one locked description per character and paste it verbatim into every prompt, changing only the wardrobe. Mine run to about sixty words and specify jaw, hairline, build, skin texture and eyes. Skin texture matters more than it sounds: without an explicit note about visible pores you get a plastic sheen that reads as fake instantly.

Then each scene gets a shot, an action and a single camera move. One move. Ask for a push-in and a pan and a rack focus and you get something that lurches.

The maths of it

For a three-minute song with a video, honestly:

  • Reading the brief and finding the angle: 30 minutes
  • Lyric, draft plus rewrites: 90 minutes
  • Generating and choosing the recording, usually 20 to 40 attempts: 45 minutes
  • Character locks and scene prompts: 60 minutes
  • Generating and rejecting clips: 90 minutes
  • Edit, cut to the beat, colour: 60 minutes

Call it six hours of attention. Which is why anything advertising an instant personalised song for nine dollars is selling you the ninety seconds and skipping the six hours, and you can hear it.

The tools removed the cost of production. They did not remove the work of deciding what is good.

ZK

Zaib Khan

Founder, Sureela

Has been building software since school. Writes and produces every Sureela song, and acts when he gets the chance. Moved from London to the San Francisco Bay Area to build this.

Have one written for you.

Two photographs and a few lines about the people involved. An original song and a music video, back in 24 hours.

Make ours

From $49. Remade once, free, if it is not right.

Written by Zaib Khan, 2026.

Read next

Music tech

Writing in Hinglish, and why machines struggle with it

Hindi, Urdu and English in the same verse is normal for millions of people and unusually hard for AI. What actually goes wrong, and how to work around it.

AI music

What AI music gets right, and where it falls apart

An honest assessment of AI music generation in 2026: the things it genuinely does better than a session musician, and the four places it still reliably fails.