Lunar Boom Learning

5.3 · Human Listening Tests

Section 5.3 of 5.6

Human Listening Tests

Measuring perception, adherence, and preference

Learn how to design controlled listening studies using blind presentation, task-specific questions, rating anchors, paired comparisons, representative listener groups, loudness control, randomisation, quality checks, and appropriate statistical reporting.

About 16 minutes

Guiding question

What exact judgement should the listener make?

By the end, you’ll be able to
  • Explain why human listening tests must begin with a precise evaluation question.
  • Distinguish audio quality, prompt adherence, musical coherence, usefulness, and overall preference.
  • Choose among single-stimulus ratings, paired comparisons, ABX discrimination, and MUSHRA-style tests.
  • Design a blind or double-blind presentation that hides irrelevant system information.
  • Control loudness, playback level, sample order, clip duration, and listening environment.
  • Use rating anchors and instructions that reduce ambiguity among listeners.
  • Select listener groups according to the intended user population and evaluation purpose.
  • Recognise when expert and general-listener results should be analysed separately.
  • Use repeated items, reference samples, attention checks, and completion checks to identify unreliable responses.
  • Report uncertainty, listener disagreement, exclusions, and study limitations alongside average ratings.
  • Apply informed-consent, privacy, compensation, and data-retention safeguards to participant studies.

Automated metrics can compare embeddings, distributions, tempo estimates, and repetition patterns. They cannot fully determine whether listeners find a generated piece natural, coherent, relevant, useful, or preferable.

Human listening tests address those questions by presenting controlled audio examples to people and recording defined judgements.

The difficult part is not asking people which clip they like. The difficult part is designing a study where the response can be interpreted correctly.

Begin with one judgement

A listening question should identify the property listeners are meant to evaluate.

Useful questions include:

    1. Which clip sounds more technically natural?
    2. Which clip better matches the text prompt?
    3. Which continuation better preserves the supplied melody?
    4. How coherent is the piece across its full duration?
    5. Which result would you use in the stated creative task?
    6. Can you hear a difference between the reference and processed clip?

Which one is better? combines several possible criteria and makes the result difficult to explain.

Questions that should not be silently combined

Swipe sideways to view the full comparison

CriterionQuestionPossible disagreement
Audio qualityHow clean and natural does the audio sound?A polished clip may be musically dull
Prompt adherenceHow well does the clip match the requested description?An accurate clip may sound unpleasant
Musical coherenceHow well do the musical ideas connect and develop?A coherent piece may not match the prompt
OriginalityDoes the piece appear distinct from the comparison material?An unusual piece may still be low quality
UsefulnessWould the output work for the stated task?A simple track may be more usable than a complex favourite
PreferenceWhich clip do you personally prefer?Preference can reflect taste rather than system accuracy

Select a test method that matches the claim

ITU guidance distinguishes several subjective audio-assessment methods because different questions require different designs.

A study comparing subtle codec artifacts should not automatically use the same procedure as a study comparing creative preference among generated songs.

The method should be selected according to the expected difference, reference availability, number of systems, listener burden, and intended conclusion.

Common listening-test formats

Swipe sideways to view the full comparison

FormatBest suited toMain limitation
Single-stimulus ratingScoring many independent clipsListeners may use the scale differently
Paired comparisonChoosing between two systems on one criterionMany systems require many pairings
ABX discriminationTesting whether two conditions are perceptibly differentDoes not directly measure liking or quality
MUSHRA-style comparisonComparing several intermediate-quality audio conditions against a reference and anchorsRequires a meaningful common reference and careful adaptation
RankingOrdering several candidatesBecomes cognitively demanding as the set grows
Task-based evaluationMeasuring usefulness in a real workflowMore expensive and difficult to standardise

Hide information that should not affect the judgement

Model names, company names, prices, popularity, human or AI labels, and expected rankings can alter listener expectations.

Use neutral condition identifiers. Avoid interface styling that reveals the system. If a human facilitator interacts with participants, conceal the assignment when practical.

Blinding does not mean hiding the evaluation task or participant rights. Listeners should still receive clear instructions and consent information.

Randomise and counterbalance presentation order

The first clip may become an implicit reference. The second may benefit from recency. A strong previous clip can make the next clip seem worse. Repeated exposure can also create familiarity or fatigue.

Randomising or counterbalancing order distributes these effects across systems rather than allowing one condition to benefit systematically.

Store the actual order shown to each participant so order effects can be audited.

Order controls for different tests

Swipe sideways to view the full comparison

TestRecommended controlRecorded information
A versus BBalance which system appears as A and which plays firstLabel assignment and playback order
Single-stimulus ratingsRandomise clip sequence per listenerPosition and preceding stimulus
Multi-stimulus testRandomise visible stimulus positionsCondition-to-interface mapping
Long listening sessionDistribute systems across early and late positionsSession block and elapsed time
Repeated-item reliability checkSeparate repeated copies within the sessionDistance between repetitions

Control loudness without erasing the property being tested

A louder sample can attract attention and influence perceived impact or quality. Uncontrolled level differences therefore confound many comparisons.

For ordinary model comparisons, use a documented loudness-measurement and matching procedure, verify true peaks, and inspect whether normalisation introduced clipping or changed dynamics.

When loudness or dynamics are themselves part of the evaluation, do not normalise away the difference. Instead, define a controlled playback-level policy and state that level is part of the tested output.

Control the listening environment

Laboratory tests can standardise the room, transducers, playback level, and supervision. Remote tests provide access to more participants but introduce unknown headphones, speakers, rooms, distractions, and system processing.

The appropriate level of control depends on the claim. Subtle artifact detection needs stricter equipment and listener controls than broad preference testing.

At minimum, record the playback route where possible and require participants to confirm that they can hear the examples clearly.

Laboratory and remote listening

Swipe sideways to view the full comparison

AreaControlled laboratoryRemote or crowdsourced study
Playback equipmentSelected and calibratedParticipant-dependent
EnvironmentControlled room and noiseUnknown room and interruptions
SupervisionDirectly monitoredLimited or automated
Participant reachSmaller and more localPotentially broad and scalable
Subtle artifact sensitivityGenerally stronger when properly designedMay be reduced by hardware and environment
Quality controlsTraining and observationReferences, attention checks, completion checks, and response filtering

Match clip duration to the evaluation claim

Short clips reduce fatigue and allow more comparisons. They may hide structural drift, repetition, weak transitions, or poor endings.

Use enough duration to expose the property being judged:

    1. Immediate artifacts may need only a few seconds.
    2. Prompt instrumentation and mood may need a representative phrase.
    3. Melody adherence requires the relevant conditioned region.
    4. Song structure requires complete or strategically selected sections.
    5. Ending quality requires the final passage.

Do not evaluate a three-minute coherence claim using only the strongest ten-second excerpt.

Use one interpretable scale per question

A numerical scale is useful only when listeners understand what its points represent.

Provide observable anchors. For example:

    1. 1: the prompt is mostly contradicted or absent
    2. 3: the main request is recognisable but several material attributes are missing
    3. 5: all material prompt attributes are clearly present

Avoid changing the meaning of the scale between trials. Do not label one endpoint bad when the real question is adherence rather than quality.

Choosing a response format

Swipe sideways to view the full comparison

ResponseUseful forCaution
Five-point category ratingSimple quality or adherence judgementsLimited resolution and listener-specific scale use
Continuous scaleFine distinctions across several stimuliPrecision of the interface may exceed perceptual certainty
Pairwise preferenceDirect model comparisonDoes not reveal absolute adequacy
Strong or weak A/B preference with tiePreference intensity and uncertaintyResponse categories need clear anchors
Binary detectionPresence of a defined event or artifactComplex qualities cannot be reduced safely to yes or no
Free-text explanationDiscovering failure themesHarder to aggregate and code consistently

Select listeners according to the intended decision

Audio engineers may detect codec artifacts that general listeners overlook. Musicians may describe harmony and structure more precisely. Target creators may judge whether the result fits a workflow. General listeners may better represent a consumer audience.

No group is universally correct. Define whose judgement matters for the use case and recruit accordingly.

Expert and general listeners

Swipe sideways to view the full comparison

GroupPotential contributionPotential limitation
Audio engineersSensitive to distortion, dynamics, imaging, and artifactsMay not represent ordinary users
MusiciansDetailed judgement of rhythm, melody, harmony, and formTraining and genre background can shape preferences
Target creatorsDirect assessment of workflow usefulnessMay prioritise utility over listening enjoyment
General listenersBroad audience preference and acceptabilityMay not identify subtle technical causes
Internal model developersDeep knowledge of intended behaviourHigh risk of expectation and familiarity bias

Train listeners on the task, not on the expected winner

Before the measured trials, explain the question, interface, scale, replay controls, and listening environment. Provide practice examples that demonstrate the range of relevant defects or adherence levels.

Do not tell participants which model is expected to perform better. Training should reduce misunderstanding without steering preferences.

Check response quality without deleting disagreement

Remote studies need safeguards against inaudible playback, random clicking, incomplete listening, bots, duplicate participation, or misunderstanding.

Possible checks include:

    1. Requiring playback before submission
    2. Repeating selected trials
    3. Including obvious reference or anchor cases
    4. Recording completion time
    5. Asking equipment and environment questions
    6. Detecting impossible or internally inconsistent responses
    7. Limiting duplicate participation
    8. Reviewing free-text explanations where collected

Define exclusion rules before examining model rankings.

Limit fatigue and adaptation

Long sessions reduce attention and can change how listeners use the scale.

Reduce burden by:

    1. Dividing trials into blocks
    2. Providing breaks
    3. Limiting repeated long-form clips
    4. Using balanced incomplete designs when every listener cannot hear every system pair
    5. Randomising which prompt-system combinations each listener receives
    6. Recording session position so fatigue effects can be analysed

A shorter high-quality study may provide better evidence than a larger exhausting one.

Plan the number of listeners and ratings

There is no universal number of listeners for every music test.

Required evidence depends on:

    1. Expected difference between systems
    2. Variation among listeners and prompts
    3. Number of models and comparisons
    4. Repeated-measures or independent design
    5. Desired uncertainty
    6. Planned subgroup analysis
    7. Number of ratings per sample
    8. Missing or excluded responses

Use a pilot study to estimate variability, then plan the full design around the smallest difference that matters for the release decision.

Real-world example: MusicLM

MusicLM used an A-versus-B listening task focused specifically on text adherence.

Listeners received a text caption and two ten-second music clips. They chose among strong or weak preference for either clip and no preference. The instructions asked listeners to focus on which music better matched the caption rather than general audio quality.

The study collected 1,200 ratings and aggregated pairwise wins across the compared systems. This design demonstrates how a listening task can isolate one dimension instead of asking for vague overall quality.

Real-world example: MusicGen

MusicGen evaluated human judgements of two separate properties:

    1. Overall perceptual quality
    2. Relevance to the text input

The study used crowdsourced raters, normalised samples to a shared loudness target, required each selected sample to receive several ratings, and applied CrowdMOS-style procedures to remove responses that did not meet predefined listening-quality checks.

Separating overall quality from relevance made it possible for one model to perform differently across the two dimensions.

In practice

Real-world example: MusicRL

MusicRL collected 300,000 pairwise user preferences and used them to align a music generator with human feedback. Its analysis found that text adherence and audio quality explained only part of musical preference, illustrating why overall liking should not be treated as a simple combination of two technical scores.

Aggregate pairwise choices carefully

Simple summaries include win rate, loss rate, tie rate, and preference margin for every pair.

When many systems are compared through an incomplete set of pairs, a Bradley-Terry-style model can estimate relative strengths from the pairwise outcomes.

The analysis should preserve listener and prompt structure where relevant. Thousands of ratings from a small number of prompts do not provide the same evidence as ratings spread across many independent prompts.

Report the complete study, not only the winning score

A listening-test report should include:

    1. Evaluation question and hypotheses
    2. Systems and checkpoints
    3. Prompt and clip selection
    4. Candidate-selection policy
    5. Number and type of listeners
    6. Recruitment and compensation
    7. Listener instructions and training
    8. Playback environment and equipment policy
    9. Loudness and peak treatment
    10. Clip duration and replay rules
    11. Randomisation and counterbalancing
    12. Rating scales and anchors
    13. Quality-control and exclusion rules
    14. Statistical method and uncertainty
    15. Group-level and prompt-level results
    16. Listener disagreement
    17. Failed or missing trials
    18. Known limitations

Without this information, the result is difficult to reproduce or interpret.

Protect participants and their data

Listening studies can collect personal information, demographic data, equipment details, work history, musical training, open-text responses, and behavioural records.

Before collection:

    1. Determine whether institutional or legal human-subject review applies
    2. Explain the purpose, tasks, duration, risks, compensation, and withdrawal process
    3. Obtain valid informed consent where required
    4. Collect only information necessary for the study
    5. Define retention and deletion periods
    6. Restrict access to identifiable data
    7. Explain how anonymised or aggregated results may be shared
    8. Avoid deceptive recruitment or coercive payment conditions

Legal requirements depend on jurisdiction, organisation, funding, and the nature of the research.

Design a complete listening study

  1. State the decision

    Define what product, research, or release decision the study will support.

  2. Select one primary criterion

    Choose quality, adherence, coherence, preference, usefulness, or another defined property.

  3. Select the test format

    Use single-stimulus, paired comparison, ABX, MUSHRA-style, ranking, or task-based evaluation as appropriate.

  4. Choose the listener population

    Recruit experts, target users, general listeners, or separate strata according to the intended use.

  5. Freeze the stimuli

    Record prompts, seeds, checkpoints, excerpts, candidate-selection rules, and durations.

  6. Standardise presentation

    Control loudness, peaks, playback, interface, replay, and condition labels.

  7. Randomise and counterbalance

    Distribute order, interface position, and session placement across conditions.

  8. Write anchors and practice trials

    Teach the task without revealing the expected outcome.

  9. Define response-quality rules

    Predefine repeated items, reference checks, completion checks, and exclusions.

  10. Pilot the study

    Test the interface, instructions, variability, duration, and technical delivery.

  11. Collect and analyse

    Preserve listener, prompt, order, and condition structure in the analysis.

  12. Report uncertainty and limitations

    Show distributions, disagreement, exclusions, subgroup results, and practical significance.

Interactive lesson

Design and Run a Blind Music Listening Test

Try it: Begin with the prompt-adherence comparison. Complete the study-design panel before hearing the clips. Compare what happens when model labels are revealed, order is fixed, loudness is unmatched, or expert and general listeners are pooled. Finish by exporting a complete protocol.

A reliable listening test defines one judgement, hides irrelevant system information, controls loudness and playback, randomises order, uses anchored responses, recruits the relevant listener population, applies predefined quality checks, and reports uncertainty and disagreement.

From perceived quality to source relationships

Listening tests can reveal that a generation sounds familiar, imitative, or unusually close to another work. They cannot establish where the relationship came from or whether the system memorised a training example.

The next section, Similarity, Memorisation, and Attribution, combines perceptual review with retrieval, provenance, training-data audits, and attribution evidence.

Check your understanding

Ready for a quick check?

Test whether you can design, control, analyse, and report a human listening study for generated music.

Ready to continue?

Save this section to your account and continue from any device.

Checking your account…
Sources for this lesson (15)
  1. Methods for the Subjective Assessment of Small Impairments in Audio Systems — ITU Radiocommunication Sector (2015)
  2. Guidance for the Selection of the Most Appropriate ITU-R Recommendation for Subjective Assessment of Sound Quality — ITU Radiocommunication Sector
  3. General Methods for the Subjective Assessment of Sound Quality — ITU Radiocommunication Sector (2019)
  4. Method for the Subjective Assessment of Intermediate Quality Level of Audio Systems — ITU Radiocommunication Sector (2023)
  5. Algorithms to Measure Audio Programme Loudness and True-Peak Audio Level — ITU Radiocommunication Sector (2023)
  6. MusicLM: Generating Music From Text — Andrea Agostinelli et al. (2023)
  7. Simple and Controllable Music Generation — Jade Copet et al. (2023)
  8. MusicRL: Aligning Music Generation to Human Preferences — Geoffrey Cideron et al. (2024)
  9. CROWDMOS: An Approach for Crowdsourcing Mean Opinion Score Studies — Flavio Ribeiro, Dinei Florêncio, Cha Zhang, and Mike Seltzer (2011)
  10. Rank Analysis of Incomplete Block Designs: The Method of Paired Comparisons — Ralph Allan Bradley and Milton E. Terry (1952)
  11. The Belmont Report — National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research (1979)
  12. 45 CFR 46: Protection of Human Subjects — Office for Human Research Protections
  13. Informed Consent FAQs — Office for Human Research Protections
  14. The Research Provisions — Information Commissioner's Office
  15. AI Risk Management Framework Playbook: Measure — NIST
Browse all contextual sources →