Blog

Finagle My Song · Blog

Can ChatGPT master your song? I measured what it actually does

There's a tip going around the AI music communities right now: upload your WAV to ChatGPT, tell it what kind of master you want, and a few minutes later it hands one back. People are saying the results beat the free mastering sites.

I was sceptical, for one specific reason: ChatGPT can't hear.

So I ran the experiment. One track, same prompt, two separate chats. Then I measured everything. Some of what I found surprised me — just not in the direction I expected.

The setup

One rock track of mine, mixed but not mastered — sitting at −15.3 LUFS with the true peak right at the digital ceiling. The prompt, verbatim, both times:

hey can you master this track? the vocals arent loud enough and also its muddy in places - id like a real crunchy rocky guitar sound

Two complaints it could measure, one ask it couldn't. Hold that thought.

The tip I saw promised about 5 minutes. Run 1 took 21 minutes 33 seconds. Run 2 took 18 minutes 22 seconds. More on where that time went in a second.

What it's actually doing

ChatGPT doesn't listen to your song. It writes Python scripts.

The scripts measure the file — loudness, peaks, spectral balance — then apply EQ, compression and limiting based on those numbers. Mastering by spreadsheet. And to be fair, measuring is the part it's genuinely good at: both runs opened with the identical diagnosis, down to the decimal — peaks near 0 dBFS, about −15.3 LUFS. My own analyzer says −15.30 LUFS, −0.03 dBTP. Correct, twice.

Then it got strange. Both runs decided the job would be easier with stems, and tried to download Demucs — an open-source AI stem separator. Both runs failed to get the model. Both then spent several minutes searching the web for mirror links, trying alternative download routes, poking at DNS settings — before quietly giving up and falling back to mid/side processing instead. That's where most of those 20 minutes went. Not mastering. Flailing at a download.

I only know this because I watched the work log. The final summary doesn't mention it — it just notes the vocal changes "use centre/side processing rather than isolated-track mixing." Which is true. It's also the polished version of events.

The numbers

Original and both masters, through the same analyzer:

                    Original   Run 1    Run 2
Loudness (LUFS)       -15.3    -13.0    -12.5
True peak (dBTP)      -0.03     -1.1     -1.0
Dynamic range (LU)     15.3     11.8     11.5
High-mid (2-8kHz)      3.5%     4.1%     4.5%
Treble (8-20kHz)       0.6%     0.6%     0.6%

Finding 1: it told the truth

I expected it to slam a limiter, hand me a loudness-war sausage, and round the numbers in its favour. It did none of that.

ChatGPT claimed −13.0 LUFS / −1.1 dBTP for run 1 and −12.5 / −1.0 for run 2. My analyzer, independently, measured its files at −13.0 and −12.5. Exactly. Every number it reported checked out, at both ends of the pipeline.

And the loudness decisions were sensible: a modest 2.5 dB lift with a proper −1 dBTP ceiling. It bought that the honest way — the dynamic range went from 15.3 LU to about 11.7, which is real limiting, but nowhere near sausage territory.

Credit where due: the measurement layer doesn't lie. If you ask it what it did — and you should — the numbers are real.

Finding 2: same file, same prompt, two different masters

Not wildly different. Half a dB apart in loudness, similar shape. Cousins, not twins.

But look at the decisions. Run 1 cleaned up mud between 220–520 Hz; run 2 chose 170–480 Hz. Run 2 added oversampled soft clipping for density; run 1 didn't. Run 1 lifted the vocal with centre-channel presence; run 2 built it a completely different way, with parallel compression.

There is no fixed mastering chain in there. It improvises a new one every run — new scripts, new bands, new settings — and this time the dice landed close together. That's the same measurements steering it to the same neighbourhood, not a repeatable process. An engineer can recall a session. This can't even recall itself.

Finding 3: the sound

Look at the last two rows of that table. My mix went in with 3.5% of its energy in the high-mids and 0.6% in the treble. The masters came back at 4.1–4.5% and… 0.6%. The tonal balance barely moved. Now re-read the prompt: "a real crunchy rocky guitar sound."

Then I did the part almost nobody does — turned the masters down to match the original's level and A/B'd at equal volume.

The vocals: genuinely nice. They came forward like I asked, and the treatment put a bit of character on them that I liked. Credit where due, again — "the vocals aren't loud enough" got properly fixed.

The guitars: not crunchy. Both runs claimed "parallel saturation for a crunchier, denser guitar sound," and at matched volume I got a slightly denser guitar. Not a crunchy one. Not close to what I was hearing in my head when I typed the prompt.

And that's the mechanism showing through. "Vocals aren't loud enough" is a number — centre-channel level — and the number moved. "Muddy in places" is a number — low-mid energy — and the number moved. "Crunchy" isn't a number. It's a sound. Nothing in that pipeline ever listened, so nothing ever noticed the crunch didn't arrive. It did everything the spreadsheet could see, and missed the one thing you could only hear.

The kicker: when the crunch isn't there, there's no knob to turn. You can't nudge the saturation up. Your only move is to re-roll — and Finding 2 says a re-roll isn't this master with more crunch, it's a different master.

About the stems trick

Some people go further: split the track into stems first, have ChatGPT master each one, then recombine. Funny thing — ChatGPT agrees. Both of my runs tried to do exactly that on their own, and only skipped it because the download failed.

But be clear about what stem-by-stem processing is. The moment you treat parts of the mix separately and re-sum them, you've changed the balance between the parts of your song. That's not mastering — that's remixing. It might genuinely sound better! But it sounds better because you changed the mix, not because the polish got deeper. Different knob.

The verdict

Better than I expected, and it failed exactly where the mechanism says it must.

The measurements are honest. The loudness decisions are sensible. The complaints you can express as numbers get real fixes — as a free option for someone with no mastering software, that's genuinely usable. But it's not repeatable, there are no dials on the result, and the part of your request that lives in your ears — the part that's the actual point — goes unheard, because nothing ever listens.

If you use it, two rules:

  1. Loudness-match before you judge. Even a 2.5 dB lift is enough to win a first listen on volume alone. Turn the ChatGPT master down to your original's level, then A/B. If it still sounds better at the same volume, it's actually better.
  2. Make it show its work. Ask what it changed and by how much. In my runs the reported numbers checked out — use them. Sanity-check the moves, or redo them in your DAW with proper control.

That first rule is the reason I built Finagle My Song. Upload a track and it analyses your LUFS, dynamic range and tonal balance free — and if you master with it, every A/B against your original is loudness-matched automatically, so you're judging the sound, not the volume. And when the master isn't quite there yet, you get the dials — loudness target, compression, EQ — instead of a re-roll. No account needed.

But even if you never touch my tool: never judge a master at two different volumes. Not ChatGPT's, not a website's, not an engineer's. Match the levels first. It's the single habit that makes everything else in mastering honest.