This post was machine-translated from Korean with AI.

[music]

How do you even make money from "AI music generation"?

updated

The boss set the direction — make and sell store background music (BGM) with AI — and I spent nearly half a year clinging to it.

I spent far more time "clinging to what wouldn't work" than finding good tools.

What I was building, and why

The boss's direction was this: mass-generate BGM with AI for spaces like cafés, bakeries, hotels, and offices, then supply and distribute it. It started as a personal project.

For store music, "variety" is everything — play it for an hour and you can't have similar tracks repeating.

I started with Google's Lyria 3 Pro. Eight cents a track, API automation, commercial use allowed — the terms were good.

But once I listened back, there was a problem. Two tracks, "Afternoon Breeze" and "Spring Garden," had different instruments but nearly identical melody, structure, and rhythm. Different titles, effectively the same song.

Here were the entire prompts I put into the two tracks:

TrackPrompt
Afternoon Breezebossa nova, nylon guitar, light shaker, 85 bpm, instrumental
Spring Gardenlight jazz, flute + piano, 85 bpm, instrumental

Different genres (bossa nova vs light jazz), different instruments — and yet, listening, you go, "wait, isn't this basically the same song?"

Afternoon Breeze — Lyria 3 Pro (bossa nova, 85 BPM)
Spring Garden — Lyria 3 Pro (light jazz, 85 BPM)

That's fatal for store use, so I went digging through other tools to see how they keep things varied.

I actually benchmarked nine music-generation tools up front — Lyria 3 Pro, Suno, Udio, MiniMax, plus the likes of Stable Audio, Soundraw, and Beatoven.

Each had its own strengths. Suno has Weirdness and Style Influence sliders on its Style·Lyrics fields; Udio lets you pick song length in steps starting at 0:32.

But under the boss's condition — "bulk generation for stores" — what mattered was automation, unit cost, and variety, and the field narrowed on those.

Here are the nine I looked over. Quantitative values like per-track cost and max length only really show up once you build with each tool, so the table sticks to access and features. (The clear unit costs: Lyria at $0.08/track, and MiniMax, which I ended up adopting, around $0.003.)

ToolAccessNotable feature
SunoFree 50 credits/dayWeirdness·Style Influence sliders, structure tags, Persona
UdioFree 600 credits/month4 length steps (0:32–2:10), Clarity slider
SoundrawFree trial (paid to download)Mood/Genre filters without prompts, Energy Curve editor
Beatoven.ai15 min freeUse-case categories (café, workout…), per-section emotion
Stable Audio 2.0Free creditsTime-coded structure tags, Negative Prompt
AIVA3 free tracks/monthEmotion presets, Key/Tempo/Duration control
Google MusicFXFreeLyria's sibling model (Lyria 2 based)
MubertFree trialTag-combination (genre·mood·activity) generation
Meta MusicGenFree (open source)Melody upload reference, no sign-up

Moving to MiniMax, a fresh string of traps

I ended up switching to MiniMax (Music 2.6), and it wasn't easy here either. The very first sample track came out as a Chinese-language country song.

Curious? Go try it. It's fresh, I'll give it that.

On top of that, I'd written mood directions in parentheses in the lyrics box. Under [Intro], things like (guitar arpeggio, subtle piano touch).

And the AI literally sang "guitar arpeggio." Even English stage directions like [Verse 1 — low register, introspective] came out as lyrics.

Only after listening did I learn that what you put in parentheses isn't a performance note — it gets synthesized as lyrics. Special characters like em dashes (—) and ellipses (…) can get sung too, so plain English and plain punctuation are the safe bet.

Here's what it sounded like. I played a track where I'd only jotted the mood into the intro, and the vocal was clearly singing:

"guitar arpeggio… subtle piano touch…" (the stage direction laid right onto the melody)

The notes I'd written about the mood became the lyrics wholesale. I threw away half of that day's tracks.

And hearing it set to a melody and sung made it tragicomic. It was only because the lyrics were Korean that I caught it — if they'd been English, I might not have noticed so easily.

The field and filter limits are rough, too. Styles caps at 2,000 characters, Lyrics at 3,500, and song length at 5 minutes.

Section tags aren't freeform either — only a fixed set of 14 ([Intro] [Verse] [Chorus] [Hook] [Drop] [Bridge] [Outro] and the like) are valid. Use a nonstandard tag without knowing this, and it gets sung as lyrics too.

What's maddening is that even when you do use the right tags, it sometimes reads them aloud anyway.

To make fitness EDM, I put three favorite famous DJs' real names in the prompt, and got blocked three times in a row with "Under Review."

Drop the real names for phrases like "festival-ready big room house," swap the word "sample" for "chopped vocal synth texture," and cut things like "top-40," and it finally went through. That's how I learned you can't put living artists' real names in a prompt.

What I don't get, though: MiniMax is random-seed based, so the same prompt yields a different track each time. So why block it?

Tracks with zero similarity pop out anyway, which also meant there was little point in saving the prompt for a song I happened to like.

Four passes on a single hum

The thing I clung to longest was a quiet "hum" — a wordless, humming vocal that suits a spa or café.

Something funny (well, painful) happened here. I'd handed the quality scoring to Gemini. Feed it the audio and it scores it.

Gemini gave my first hum version a perfect 12. But to the boss's ear it was off — rejected.

For the third version, made by pinning down the melody and reference tracks (like La La Land's "City of Stars"), Gemini gave a perfect 16 and even a comment that "the mood improved." And yet, to the boss's ear, the vocal was too loud and nothing like the reference. Rejected again.

Along the way I tried turning the vocal down, but MiniMax has no volume tag at all.

So I used three things at once: keywords like "instrumental as primary" and "vocal quieter than instruments" in Styles, trimming the lyric phonemes under 150 characters, and when that still failed, splitting just the vocal out with Demucs and remixing it back at -8 to -12dB.

Even after all that, the "quiet hum" I was aiming for never came out.

The AI score stayed perfect while the boss's ear kept saying "NO." There was no human texture like breath, just the same pattern repeating — honestly, an eerie sound.

"This is supposed to play in a spa?" I thought. For the fourth and fifth attempts I didn't even hit generate. That's how completely I gave up on the hum.

Giving up opened a path

But once I gave up, an unexpected path opened. I dropped the hum, just flipped on the Instrumental switch, and ran that café-jazz prompt (942 characters) I'd been about to toss. Out came a single 8-minute-49-second instrumental piece.

The same prompt in vocal mode gave me 1.5 to 3 minutes. For store BGM, this long instrumental was far more natural than an awkward hum.

The pure instrumentals also averaged 19.6/20 on Gemini's scoring, with BPM landing within ±4 of target. In the end, just removing the vocal satisfied both the AI score and the boss's ear.

It was cost-efficient, too. MiniMax charges the same credits per track whether it's 5 minutes or 8. So to fill a one-hour compilation you need 12 five-minute tracks but only 7–8 eight-minute ones.

Longer prompts were also better. Comparing via Gemini's scoring, a loose 280-character listing averaged around 13, while 1,400 characters spelling out references, chord progressions, and mixing direction scored in the 19s.

Explaining at length clearly beat tossing something short.

That said, this might be a MiniMax-specific trait, so take it with a grain of salt. There's no guarantee Suno scores as well under the same conditions. And I'm still testing Suno, so I don't have full hands-on numbers there yet — just so you know.

For the license, I'm using a $100/year starter plan the boss happened to subscribe to early. (It's no longer sold.)

It used to be $100/year with a 100-tracks-per-day cap; now they meter token usage instead. It resets every 5 hours, and per 5-hour window I can make around 20 tracks.

Generation used to be billed separately and now it's folded in, so that part is a little disappointing.

The scorer cost money too

Choosing the evaluation tool was a cost as well. Scoring the same 22 tracks ran about $0.37 on Gemini and about $13.2 on GPT-4o audio. A 35× gap, so I went with Gemini.

You're probably curious what the scoring criteria were. It rated each item 1 to 4 points. Instrumentals got five items — sound quality · structural clarity · instrument expression · mood consistency · BPM accuracy — at 4 points each, for 20 total.

Vocal tracks added pronunciation clarity and emotional expression, and Korean vocals added language naturalness, reaching 28. Hums were judged separately on naturalness, mood, human feel, and commercial fit.

A 4 on each item was defined as "ready for commercial use," a 1 as "noise / off."

For how different the tracks were from each other (variety), instead of human ears I also used an audio-embedding model called CLAP to score similarity objectively. Six pairs averaged 0.79 — a "varied" verdict (0.95+ would mean nearly the same song).

Switching generation engines saved money too. Lyria was about $528 a year, while MiniMax's starter plan was $100 a year, roughly $0.003 per track — saving about $428 (81%) and giving 5.5× the monthly generation quota.

Oh, and the model that claimed "free with just an API key" spat out an insufficient-balance error when I actually called it.

There was one more variable in tool choice beyond cost. Suno has the best quality, but it only runs on the web, so bulk automation wasn't possible.

Store BGM means cranking out hundreds of tracks, and you can't pull them one by one by hand, so the answer was MiniMax with its API automation. (I did separately apply for the Suno Enterprise API beta. I'm not holding my breath — it'll probably fall through.)

What stuck with me

Just like with LLM LoRA, I stopped trusting AI scores. (I keep ending up writing only about not trusting AI's grading.. heh.)

At least for music you judge by ear, even a perfect score from Gemini doesn't matter if it sounds bad to the boss. AI tends to overrate things where subtle naturalness matters — hums, ambient — so we made it a rule: when the AI score and the boss's ear diverge by 2 or more points, we follow the boss's ear, no exceptions.

And we decided not to cling to what won't work. It took burning four passes on the hum to learn it, but the moment we gave up, the better path — instrumental — appeared right away.

As an aside, I ran the math that stacking a catalog of around 1,000 tracks could reach a few thousand dollars a month. But that's a simulation built on a pile of assumptions, hard to call real revenue.

Making the tracks and selling them turned out to be entirely different problems.

I'm still tinkering with music these days — lately experimenting with writing lyrics more precisely on a structural basis (lines per section, density per genre). I'll write that one up once it's more cooked.

My lyricist is the boss's 7-year-old daughter. She must be busy — she hasn't written me anything.