Lunar Boom Learning

5.1 · What Makes Generated Music Good?

Section 5.1 of 5.6

What Makes Generated Music Good?

Separating quality into measurable dimensions

Learn how to evaluate generated music across audio fidelity, musical coherence, condition adherence, diversity, originality, editability, reliability, and usefulness for a specific application.

About 16 minutes

Guiding question

Can a high-fidelity clip still be musically weak?

By the end, you’ll be able to
  • Explain why generated-music quality cannot be represented adequately by one universal score.
  • Distinguish audio fidelity from musical coherence and structural development.
  • Evaluate whether generated music follows text, melody, harmony, timing, or reference conditions.
  • Distinguish diversity from randomness and originality from simple variation.
  • Explain why originality assessment requires comparison with training and reference material.
  • Define editability, reliability, and usefulness as task-dependent quality dimensions.
  • Recognise trade-offs among fidelity, adherence, diversity, structure, and practical usability.
  • Create a task-specific evaluation rubric with dimensions, weights, gates, rating anchors, and sampling rules.
  • Combine automated metrics with controlled human listening rather than treating either as sufficient alone.

Chapter 4 asked whether a model could be trained, adapted, reproduced, and debugged correctly. Chapter 5 asks a more difficult question:

Is the resulting system actually good enough for its intended use?

There is no single property called music quality that captures every answer.

A generated clip may have clean audio but weak composition. It may follow the prompt but repeat one phrase. It may sound impressive in isolation but fail when edited into a video. It may produce one excellent sample while failing most other prompts.

Evaluation therefore begins by dividing quality into separate dimensions.

Core dimensions of generated-music quality

Swipe sideways to view the full comparison

DimensionCentral questionExample failure
Audio fidelityDoes the audio sound technically clean and plausible?Metallic artifacts, clipping, hiss, unstable stereo image, or codec distortion
Local musical coherenceDo neighbouring musical events make sense together?Abrupt pitch jumps, broken rhythm, or conflicting harmony
Global musical coherenceDoes the piece develop convincingly across its full duration?Endless loops, forgotten motifs, weak transitions, or no ending
Condition adherenceDoes the output follow the requested prompt and controls?A solo-guitar prompt produces dense electronic percussion
DiversityDoes the system produce meaningfully different valid outputs?Every seed produces almost the same arrangement
OriginalityDoes the output avoid unacceptable reproduction of known material?A generated phrase closely matches a training example
EditabilityCan the output be modified and integrated into a workflow?No usable edit points, stems, continuation, or stable timing
ReliabilityHow often does the system meet its requirements?One excellent sample appears among many unusable results
UsefulnessDoes the result solve the intended user task?A pleasant track cannot meet the required duration or format

Audio fidelity

Fidelity concerns how the sound is rendered rather than whether the composition is interesting.

Possible defects include:

    1. Clipping or harsh saturation
    2. Metallic codec artifacts
    3. Hiss, hum, or unstable background noise
    4. Warbling pitch
    5. Smearing of attacks and percussion
    6. Unnatural vocal consonants
    7. Stereo instability
    8. Sudden changes in loudness or ambience
    9. Silence or corrupted endings

A technically clean loop can still be harmonically dull, structurally repetitive, or unrelated to the prompt.

In practice

Clean sound is not strong music

A perfectly rendered four-chord loop may receive a high fidelity rating and a low structural or originality rating. Fidelity measures signal quality, not the complete musical result.

Local and global coherence

A five-second segment mainly reveals local relationships. A three-minute track also reveals whether the model can organise time.

Global evaluation can examine:

    1. Recognisable section boundaries
    2. Motif return and variation
    3. Contrast between sections
    4. Harmonic direction
    5. Changes in density and instrumentation
    6. Transition quality
    7. Build-up and release
    8. Ending quality

Short excerpts should not be used to claim long-form structural quality.

Break a prompt into testable attributes

Swipe sideways to view the full comparison

Prompt componentPossible evaluation questionExample
InstrumentationAre the requested instruments audible?Solo nylon-string guitar
Genre or styleDoes the musical vocabulary match the requested category?Traditional bossa nova
TempoIs the perceived or measured pace appropriate?Slow at approximately 70 BPM
MoodDo listeners perceive the requested character?Calm and reflective
TextureDoes the arrangement have the requested density?Sparse accompaniment
ProductionDoes the sound follow the requested recording treatment?Dry and intimate
StructureDoes the piece contain the requested sections?Intro, verse, chorus, and outro
ExclusionsAre prohibited elements absent?No drums or vocals

Examples in practice

Real-world example: MusicLM

MusicLM evaluated text-to-music generation using separate measures of audio quality and adherence to text descriptions. The work also introduced MusicCaps, a benchmark of approximately 5,500 music-text pairs with detailed descriptions written by human experts.

Real-world example: MusicGen

MusicGen combined automated evaluation with human listening and reported audio quality separately from relevance to the text condition. This avoids treating acoustic quality and condition adherence as one interchangeable property.

Automated prompt-alignment proxies

Models such as CLAP and MuLan place text and audio in related embedding spaces. Similarity between a prompt embedding and an audio embedding can provide a scalable proxy for semantic alignment.

The score remains dependent on the embedding model. It may respond strongly to broad genre or instrument cues while missing timing, arrangement, negation, subtle production instructions, or long-form structure.

Embedding similarity should therefore complement attribute checks and listening tests.

Diversity is not the same as randomness

A diverse model can produce several appropriate musical solutions. A random model may produce outputs that differ greatly but fail the task.

Diversity can be examined at several levels:

    1. Different melodies for the same prompt
    2. Different arrangements and instrumentation
    3. Different rhythmic interpretations
    4. Different outputs for genuinely different prompts
    5. Coverage across the intended genre or cultural domain
    6. Variation across seeds without loss of quality

Diversity should be measured among outputs that first meet basic validity and adherence requirements.

In practice

Diversity can trade off with consistency

More exploratory sampling may produce a wider range of ideas while increasing the number of unstable or off-prompt samples. Evaluate diversity together with quality and success rate.

Originality and non-replication

Originality is not established by comparing two generated samples only with each other.

A system can generate a diverse batch while every output resembles different training examples. Conversely, a style-consistent collection can share common instrumentation and rhythm without copying a particular work.

An originality evaluation may include:

    1. Exact token or fingerprint matches
    2. Near-neighbour retrieval
    3. Melody and interval comparisons
    4. Chroma and harmonic similarity
    5. Spectrogram or embedding similarity
    6. Prompt-targeted extraction tests
    7. Human review of flagged pairs

The metric and threshold must match the kind of reproduction being investigated.

Editability and workflow fit

Creators may need more than a finished stereo file.

Useful evaluation questions include:

    1. Can the track be cut on beat?
    2. Is the tempo stable enough for synchronisation?
    3. Does it contain a clean intro and ending?
    4. Can sections be extended without obvious seams?
    5. Can unwanted elements be removed or regenerated?
    6. Are stems or symbolic controls available?
    7. Does the output preserve quality after normal editing?
    8. Does the file meet the required technical format?

A less impressive track may be more valuable if it is easier to adapt to the project.

Usefulness depends on the application

The same output can be excellent for one task and unsuitable for another.

A loose ambient texture may work well beneath narration but fail as a listening-focused single. A complex cinematic cue may sound impressive but leave no space for dialogue. A melody may inspire a songwriter while being too rough for release.

Evaluation should begin with the intended user, setting, constraints, and decision that the score will support.

Different applications require different rubrics

Swipe sideways to view the full comparison

ApplicationHigh-priority dimensionsPossible secondary dimensions
Background music for videoMood adherence, duration, editability, dialogue compatibility, reliabilityComplex structural development
Listening-focused songFidelity, global coherence, identity, originality, emotional effectStem availability
Sample-pack loopClean audio, timing stability, loop quality, originality, format complianceLong-form development
Songwriting assistantIdea diversity, controllability, editability, inspirational usefulnessFinished mastering quality
Adaptive game musicTransition quality, loopability, state control, latency, reliabilityFixed linear song form
Music restoration or enhancementFidelity, artifact reduction, source preservationPrompt diversity

Quality dimensions can conflict

Improving one dimension can weaken another.

Examples include:

    1. Stronger guidance improves prompt adherence but may reduce variation.
    2. Higher sampling temperature increases variety but may reduce stability.
    3. Aggressive denoising removes noise but may smear musical detail.
    4. Strict style adaptation improves domain fit but narrows broader diversity.
    5. Longer duration increases usefulness for songs but exposes more structural failures.
    6. Heavy post-processing improves loudness consistency but may damage dynamics.

Report trade-offs rather than hiding them inside one average.

Human listening remains necessary

Listeners can assess qualities that are difficult to capture fully with automated metrics, including musical development, emotional effect, production plausibility, and task usefulness.

Controlled listening should consider:

    1. Consistent playback equipment and loudness
    2. Randomised sample order
    3. Hidden model identities
    4. Clear rating instructions
    5. Defined listening windows
    6. Suitable reference and anchor samples
    7. Listener screening or expertise where relevant
    8. Repeated items for reliability checks
    9. Reporting uncertainty and disagreement

The correct listening design depends on the purpose of the evaluation.

Example five-point prompt-adherence scale

Swipe sideways to view the full comparison

ScoreAnchor
1The output contradicts or omits most requested attributes
2One broad attribute is present, but several important requirements are absent
3The main request is recognisable, with clear omissions or conflicts
4Most important attributes are present, with minor deviations
5All material prompt attributes are clearly satisfied
Automated and human evaluation

Swipe sideways to view the full comparison

MethodStrengthLimitation
Automated metricsFast, repeatable, and scalable across large sample setsDepend on representation, reference data, and metric assumptions
Expert listeningCan assess technical and domain-specific musical detailsExpensive, slow, and influenced by expertise and taste
General listener studyRepresents broader user reactionsMay be less sensitive to specialised defects
Creator task studyMeasures workflow usefulness directlyRequires realistic tools and task design
Combined evaluationTriangulates statistical, perceptual, and operational evidenceRequires more careful planning and reporting

Build a task-specific quality rubric

  1. Define the intended use

    State the user, task, operating environment, and decision the evaluation will support.

  2. Define mandatory gates

    List failures that make an output unusable regardless of other scores.

  3. Select quality dimensions

    Choose only dimensions relevant to the task and define each one precisely.

  4. Write rating anchors

    Explain what low, medium, and high scores mean using observable evidence.

  5. Select prompts and conditions

    Cover common, difficult, rare, compositional, and exclusion-based requests.

  6. Define the sampling policy

    Specify seeds, candidates per prompt, duration, decoding settings, and whether selection is allowed.

  7. Choose automated measures

    Use metrics that correspond to defined dimensions and document their assumptions.

  8. Design human evaluation

    Control playback, randomisation, blinding, anchors, listener groups, and instructions.

  9. Report distributions

    Show averages, variation, failure rates, prompt-level results, and disagreement.

  10. Define the decision rule

    State what evidence permits release, revision, limited deployment, or rejection.

What an evaluation prompt set should cover

Swipe sideways to view the full comparison

Prompt groupPurpose
Common requestsMeasure ordinary user performance
Rare genres and instrumentsReveal coverage gaps and representation weaknesses
Contradictory combinationsTest how the model resolves conflicting instructions
Negative instructionsTest whether excluded elements remain absent
Precise tempo and durationMeasure temporal control
Long-form structureMeasure section planning and global coherence
Minimal promptsMeasure defaults and ambiguity handling
Paraphrased promptsTest robustness to equivalent wording
Adversarial or malformed inputsTest safe and predictable failure behaviour

Report more than the mean

An average can hide unstable behaviour.

A quality report should consider:

    1. Mean and median ratings
    2. Variation and confidence intervals
    3. Prompt-level performance
    4. Listener disagreement
    5. Failure and regeneration rates
    6. Performance by genre, language, culture, instrument, and duration where relevant
    7. Best, typical, and worst examples
    8. Automated-human metric correlation
    9. Known blind spots and unmeasured properties

The purpose is to show where the system works, where it fails, and how certain the evaluation is.

Interactive lesson

Build and Apply a Music Quality Rubric

Try it: Begin with the listening-focused song preset and rate each clip without using the overall score. Compare the dimension profiles, then switch to background video music and observe how the weights and gates change. Finish by designing a custom rubric and documenting the decision rule.

Generated-music quality is multidimensional. Evaluate fidelity, local and global coherence, condition adherence, diversity, originality, editability, reliability, and usefulness separately before applying task-specific weights and mandatory gates.

Turning dimensions into scalable measurements

This section defined what should be evaluated. The next section, Automated Evaluation, examines how distributional audio metrics, classifiers, embeddings, symbolic statistics, control-alignment scores, and replication searches can measure parts of those dimensions at scale.

Check your understanding

Ready for a quick check?

Test whether you can separate music-quality dimensions and design a rubric suited to a defined use case.

Ready to continue?

Save this section to your account and continue from any device.

Checking your account…
Sources for this lesson (14)
  1. MusicLM: Generating Music From Text — Andrea Agostinelli et al. (2023)
  2. MusicLM: Generating Music From Text — Andrea Agostinelli et al. (2023)
  3. Simple and Controllable Music Generation — Jade Copet et al. (2023)
  4. Fréchet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms — Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi (2019)
  5. Adapting Frechet Audio Distance for Generative Music Evaluation — Azalea Gui, Hannes Gamper, Sebastian Braun, and Dimitra Emmanouilidou (2024)
  6. CLAP: Learning Audio Concepts From Natural Language Supervision — Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang (2023)
  7. MuLan: A Joint Embedding of Music Audio and Natural Language — Qingqing Huang et al. (2022)
  8. Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio — Roser Batlle-Roca, Wei-Hsiang Liao, Xavier Serra, Yuki Mitsufuji, and Emilia Gómez (2024)
  9. ITU-R BS.1534: Method for the Subjective Assessment of Intermediate Quality Level of Audio Systems — ITU Radiocommunication Sector (2015)
  10. ITU-R BS.1116: Methods for the Subjective Assessment of Small Impairments in Audio Systems — ITU Radiocommunication Sector (2015)
  11. ITU-R BS.1283: Subjective Assessment of Sound Quality — ITU Radiocommunication Sector (1997)
  12. AI Risk Management Framework Playbook: Map — NIST
  13. AI Risk Management Framework Playbook: Measure — NIST
  14. AI Risk Management Framework Playbook: Manage — NIST
Browse all contextual sources →