Skip to main content
CrossCoach
Sign inRequest access
A lone studio microphone standing in near-total darkness, the air around it alive with a glowing violet-to-magenta waveform that radiates outward like a voice made visible, the only light in the frame.
Reading · Speaker comparison

Voice on Trial: Forensic Speaker Comparison

Voice comparison now rests on likelihood ratios, calibration and case-relevant validation rather than visual “voiceprints.” Six sections explain the evidence and its limits.

15 min readBased on the the forensic voice-comparison literature
I

The voiceprint was never a print

In 1966, Lawrence Kersta became the first voice-identification expert to testify in court. He said his "voiceprints" were as reliable as fingerprints.

Durand Begault and Fausto Poza later traced the claim to a Nature study in which novice high-school girls compared isolated words from twelve speakers. Reported error rates were 0% to 3%. The task was closed-set: the corresponding sample was always present, participants knew that, and they could compare freely among exemplars. Those conditions differ substantially from casework in which the true speaker may not be represented. Kersta also manufactured spectrographs and profited from examiner certification. The term voiceprint encouraged an analogy with fingerprints, although voices vary with health, age, emotion, effort and transmission channel.

Oscar Tosi's Michigan State research, funded by the US Department of Justice in 1968, reported false-identification rates of 2% to 6% and false-elimination rates of 5% to 12% under laboratory conditions. Tosi argued that casework would perform better because experienced examiners would decline difficult comparisons. Poza and Begault respond that the decision to proceed relied on subjective data-quality judgements without objective criteria. Casework also introduces imperfect channels and greater within-speaker variation.

When the National Academy of Sciences reviewed the field in the late 1970s, it concluded that available error-rate estimates "do not constitute a generally adequate basis" for courts to assess aural-visual voice identification. Poza and Begault wrote in 2005 that the conclusion remained applicable.

A data plate: the 1966 sales claim that voiceprints were 'as reliable as fingerprints', set against the 1978 court's own words, 'we avoid the term voiceprint'.
Fig. 1 · The voiceprint was never a print. Sold in 1966 as being as reliable as a fingerprint, the analogy was refused even by the court that admitted the evidence in 1978.

In 1978, the Second Circuit admitted spectrographic evidence in United States v Williams, relying on Tosi's 6.3% false-identification rate and analogies with handwriting and gun-barrel comparison. It described the central task as "the simple step of visual pattern-matching, a step easily comprehended and evaluated by a jury." Begault and Poza criticise that approach as subjective and impressionistic, noting the absence of evidence that trained examiners identify voices by ear more accurately than lay people. The expert, Lundgren, reported thirteen correspondences against a ten-point standard set by his trade association.

The court itself avoided the term voiceprint, writing that a spectrogram "has been often called a voiceprint. We avoid the term as potentially leading to an unwarranted association with fingerprint evidence." Current evidence should not be framed as visual voiceprint identification.

A spectrogram has been often called a "voiceprint." We avoid the term as potentially leading to an unwarranted association with fingerprint evidence.
United States v Williams, 583 F.2d 1194 (2d Cir. 1978), n.5
Challenge 01 · Put it to the test

Show me the study

Counsel holds up the witness’s report, where the spectrograms are said to "match," and asks for the evidence behind the method.

The question

"You testified that the spectrograms ‘match.’ Can you point me to a single published study, on case-realistic recordings, that establishes the error rate of your eye-matching method, or did the National Academy of Sciences conclude in the 1970s that no such adequate basis exists?"

Your answer
II

A likelihood ratio evaluates the evidence, not guilt

In 2009, Geoffrey Stewart Morrison argued in Science & Justice that forensic voice comparison was moving from categorical identification towards likelihood ratios, as DNA evidence had done in the 1990s. A likelihood ratio compares the probability of the observed voice evidence under a same-speaker proposition with its probability under a different-speaker proposition. An LR of 100 means the evidence is 100 times more probable under the same-speaker proposition. An LR of 1/1000 means it is 1000 times more probable under the different-speaker proposition. The numerator captures similarity and the denominator typicality; similarity in a common feature provides little support because it would also occur often among different speakers.

An LR is the probability of the evidence given each proposition, not the probability of either proposition given the evidence. Saying that an LR of 100 means the suspect is 100 times more likely to be the speaker reverses the conditional and commits the prosecutor's fallacy. A conclusion about guilt would also require prior information that belongs to the court.

Morrison's 2013 calibration tutorial explains that a raw system produces a score, not an interpretable LR. Calibration shifts and scales scores using known same-speaker and different-speaker data. In one worked example, calibration reduced log-likelihood-ratio cost from 1.802 to 0.750, where lower is better. Morrison distinguishes auditory, acoustic-phonetic and fully automatic approaches. In Cambier-Langeveld's 2004–2005 exercise, only four of twelve analysts reported an LR. The report should identify which type of system produced the value and what was calibrated.

A likelihood ratio is an expression of the probability of obtaining the evidence given same- versus different-origin hypotheses. There are logical and legal reasons why the forensic scientist must present a strength-of-evidence statement in this form and must not present the probability of the hypotheses given the evidence.
Morrison (2009), on the likelihood-ratio paradigm
A data plate showing a likelihood ratio as a fraction, similarity over typicality, equal to how probable the evidence is under each hypothesis — not the probability of guilt.
Fig. 2 · The evidence, not the verdict. A likelihood ratio weighs similarity against typicality — how probable the recording is if the speakers are the same versus different. Reading it as the probability the suspect is guilty is the prosecutor's fallacy.
Challenge 02 · Put it to the test

Does 100 mean it’s him?

Counsel reads the likelihood ratio back from the report and offers a tidy, wrong restatement of it.

The question

"You testified to a likelihood ratio of 100, so it is 100 times more likely my client is the speaker on the recording, correct?"

Your answer
III

Cllr must be measured under conditions like the case

In a 2012 New South Wales fraud trial, a practitioner said there was "much to support a hypothesis that the defendant is the speaker," based on auditory assessment and spectrograms of the word "no." No performance figure for comparable conditions was given. Enzinger and Morrison built a statistical system using the same features, vowel-resonance movement and mean pitch, then tested it using a compressed mobile call and a reverberant police interview. Its Cllr was 0.834. A system that always returns "no information" scores 1, so those features provided little useful information under the simulated case conditions.

Morrison's 2011 paper presents Cllr, log-likelihood-ratio cost, as a performance measure for LR systems. Lower values are better; 1 is the value of an uninformative system. The measure penalises a very large LR supporting a false proposition more heavily than a modest LR in the wrong direction. Measured validity "depends both on the system and the test set." If the questioned recording is a one-minute mobile call and the known recording is five minutes of interview speech, the validation pairs should reproduce those durations and channels. Studio-recording results do not describe performance on a noisy telephone comparison.

On the same test data, Enzinger and Morrison's fully automatic system achieved Cllr 0.401, compared with 0.834 for the manual acoustic analysis. Performance belongs to a particular method under particular conditions, not to voice comparison generally or to an examiner's experience.

A data plate: a Cllr scale from 0 (perfect) to 1 (no information), with the automatic system at 0.401 and the hand-run acoustic-phonetic method at 0.834, near the useless end, on the same case-like data.
Fig. 3 · Ask for the Cllr. On the same case-like recordings, an automatic system scored 0.401 while the hand-run acoustic-phonetic method scored 0.834 — where 1 is a system that gives no information at all.

Kelly et al. tested the commercial VOCALISE system in 2019 using forensic_eval_01, a public case-like dataset containing 223 recordings from 61 speakers, with 111 same-speaker and 9720 different-speaker comparisons. Across six configurations, pooled Cllr ranged from 0.246 for a configuration adapted to case-like conditions to 0.462 for built-in configurations. Case-relevant adaptation nearly halved the cost.

The relevant question is the Cllr measured on validation data comparable in duration, channel, language and speaking style to the recordings in the case. A Tippett plot shows cumulative same-speaker and different-speaker comparisons against log LR and should include credible-interval bands. If performance is supported only by experience or training, the error rate may remain, in the words of the Angleton ruling quoted by Enzinger and Morrison, "unknown and may vary considerably, depending on the conditions of the particular application." Daubert asks whether the technique "can be (and has been) tested."

A system which always responded with a likelihood ratio of 1, and therefore gave no information to assist the trier of fact in making their decision, would result in a Cllr-pooled value of 1.
Enzinger & Morrison 2017, on the acoustic-phonetic system’s 0.834
Challenge 03 · Put it to the test

What’s your Cllr?

The witness has concluded the recordings offer "substantial support" for the prosecution. Counsel asks what that conclusion is worth.

The question

"You concluded the recordings offer ‘substantial support’ that my client is the speaker. What is the Cllr of your method, measured on test recordings that match the duration, channel, and speaking style of the two recordings in this case? And if you have never measured it, how is the jury to know your method performs any better than a system that simply answers ‘same speaker’ every time?"

Your answer
IV

Three places the number falls apart

A likelihood ratio depends on several modelling decisions. The first is the relevant population. Its denominator asks how typical the questioned voice is among other speakers, so the analyst must define those speakers.

Hughes and Foulkes (2015) used the same New Zealand English vowel data with reference populations matched, mismatched or mixed by social class and age. For social class, the matched system had Cllr 0.513 and the mismatched system 0.659. For age, the corresponding values were 0.561 and 0.712. Changing between matched and mismatched populations shifted some individual LRs by more than a factor of ten. The authors note that "one cannot know for certain the sociolinguistic community to which the offender belongs." The chosen population must therefore be justified and its uncertainty acknowledged.

The second set of decisions concerns calibration and recording conditions. Morrison's 2017 paper examined a New South Wales case in which an automatic system used a reverberant jail landline call saved as MP3 for same-speaker training, pristine head-mounted-microphone studio recordings for different-speaker training, and a questioned mobile call recorded from beneath a mattress. No calibration step was used. Morrison's replication produced Cllr 0.957, close to the uninformative value of 1, and values averaging 18% higher than an unbiased version. Better reference conditions and calibration reduced Cllr to 0.674. The reported utterance-level values of 1.6 to 5.5 therefore carried little demonstrated information.

Third, experts may diverge. Cambier-Langeveld's 2007 exercise sent one constructed case to twelve experts in ten countries. For Q10, an 18-second different-speaker clip, three participants reported same-speaker conclusions. Automatic systems produced LRs for true same-speaker comparisons from about 40 to 40,000, and for true different-speaker comparisons from 4.36 × 10⁻²³ to about 6. The range shows that method and implementation can materially change the answer.

A data plate listing three judgements that move the likelihood ratio: the relevant population (swings one LR more than tenfold), calibration (none used, Cllr 0.957), and recording conditions (studio is not a phone under a mattress).
Fig. 4 · Three places the number falls apart. The relevant population, the calibration and the recording conditions each move the likelihood ratio — and a weakness in any one can carry it close to meaningless.

Be ready to identify the relevant population and explain its selection, describe calibration, and show that the validation data reproduce the case channel, codec, noise level and duration. Weakness in any of those components affects the meaning of the LR.

A system which gives no useful information, a system that always outputs a likelihood ratio of 1 irrespective of the input, will have a Cllr of 1. The first variant of the system had a Cllr very close to 1.
Morrison 2018, on the replicated NSW practitioner’s system (Cllr 0.957)
Challenge 04 · Put it to the test

How did you pick the population?

The witness reported a likelihood ratio in the hundreds. Counsel goes straight at the typicality term it depends on.

The question

"You reported a likelihood ratio in the hundreds. That number depends on comparing my client's voice against some group of ‘other speakers.’ You had to choose that group. I'm instructed that changing it can move a result like yours more than tenfold. So how did you decide which population to compare against, when the whole question in this trial is who the speaker was, and what does your number become if you chose the wrong group?"

Your answer
V

Bias, and the witness who never heard the suspect

Two distinct human-factors issues arise in voice evidence: contextual bias affecting experts and memory limitations affecting earwitnesses.

Kukucka, Kassin, Zapf and Dror surveyed 403 forensic examiners in 21 countries in 2017. Participants estimated their own accuracy at a mean of 96%; 148, or 37%, claimed 100% accuracy. While 71% regarded cognitive bias as a concern in forensic science, 52% saw it as a concern in their own discipline and 26% thought their own judgements were affected. This is the bias blind spot. Seventy-one per cent believed they could reduce bias by trying to set expectations aside. The authors say effort is not a sufficient remedy because bias can operate automatically and without awareness. Procedural controls, including Linear Sequential Unmasking, regulate what information is received and when. Laboratories adopting them reported that they were neither onerous nor expensive.

If a voice examiner knew the police theory, suspect history or existence of a confession before forming an opinion, the relevant question is what procedure limited that information. Saying it was put aside does not remove the exposure.

Earwitness evidence presents a different problem: memory for an unfamiliar voice is weak. Yarmey and colleagues compared six-voice line-ups with a single-voice showup. Correct identification was poor in both, and showups produced more false identifications of innocent people. Clifford and Denot found accuracy near 50% shortly after exposure but 9% after one to three weeks. In Orchard and Yarmey's study, whispered speech reduced identification accuracy, and confidence in a distinctive voice did not improve performance. Across Yarmey's studies, confidence and accuracy correlated at about .25. People able to recognise a schoolmate's face years later could not reliably distinguish familiar from unfamiliar voices even after seven speech samples.

A data plate: the giant figure '37%' — the share of forensic examiners who rated themselves 100% accurate, while 71% believed willpower alone reduces bias.
Fig. 5 · The bias blind spot. In a survey of 403 forensic examiners, 37% rated their own judgements 100% accurate, and 71% believed bias can be reduced simply by trying — the very move the research says does not work.

Do not treat a calibrated system comparison and an earwitness identification as equivalent. They involve different mechanisms, performance evidence and safeguards. Expertise may improve technical comparison, but it does not establish immunity from contextual bias.

148 examiners (36.72% of the total sample) reported a belief that their own judgments are 100% accurate.
Kukucka, Kassin, Zapf & Dror (2017)
Challenge 05 · Put it to the test

What did you know before you listened?

Counsel turns from the method to the mind running it, and asks what the witness knew before forming a view.

The question

"Doctor, you told the jury your method is rigorous, yet a survey of 403 forensic examiners found most thought willpower alone could neutralise bias and 37 percent rated themselves 100 percent accurate. Before you formed your opinion in this case, did you know the police theory, the suspect’s record, or that there was a confession, and what specific procedure walled that information off from your analysis?"

Your answer
VI

Is the recording even of a real person?

A voice case may now require authentication before speaker comparison. Liu and colleagues reported the ASVspoof 2021 challenge, which received submissions from 54 teams across three tasks, including fabricated speech "in the voice of a target speaker." On the deepfake progress data, 23 of 33 systems achieved equal error rates below 10%, with the best below 1%. On the less familiar evaluation data, every system exceeded 15%. Performance did not generalise from the development material.

Some detectors had learned the duration of silence at the beginning and end of clips, a feature of the training corpus rather than synthetic speech. Removing non-speech with a voice-activity detector meant performance "is substantially degraded." The paper warns that a countermeasure using a database artifact is unlikely to "lead to reliable detection in the wild." A detector validated on known attacks may therefore miss speech produced by an unfamiliar cloning method.

Language analysis for the determination of origin (LADO) raises a separate issue. Patrick and Fraser note that language varieties do not follow national borders, speakers accommodate and change during displacement, and some analysts are native-speaker informants without linguistic training. A categorical opinion that a person is not from a country may affect an asylum or deportation decision despite limited evidential support. Any opinion should be confined to observed linguistic features and the limits of the available population data; in some cases no origin opinion can be supported.

A data plate: deepfake detectors scored below 1% error on known attacks but every system exceeded 15% on unfamiliar attacks — performance did not generalise.
Fig. 6 · Is the recording even real? The best deepfake detectors scored under 1% error on known attacks, but every system exceeded 15% on unfamiliar ones — a detector tuned to known attacks may miss a new cloning method.

A witness should not guarantee that a recording is a human utterance rather than a clone. A spoofing detector tested on known attacks may not generalise to new ones, while speaker comparison may assume that the questioned sample is authentic. Partial spoofing is especially difficult: the ASVspoof authors note that replacing "I won the election" with "I lost the election" changes the meaning while leaving most of the recording unaltered. The report should state which authentication assumptions were tested and which remain unresolved.

While 23 (out of 33) systems have EERs of less than 10% for the progress subset, and while the best performing system even has an EER of less than 1%, all have EERs exceeding 15% for the evaluation set.
Liu et al. (2023), ASVspoof 2021, Sec. III-C
How far does each voice-evidence move actually stand?
Calibrated LR with a Cllr on case-matched recordings
Automatic system, but validated on mismatched audio
Acoustic-phonetic features without calibration
Aural-spectrographic "voiceprint" eye-matching
Lay earwitness identification of an unfamiliar voice
Inferring nationality from speech (LADO)
ValidatedInstrument-basedSubjective comparison

Relative scientific footing across the six moves in this reading, from a calibrated likelihood ratio on case-matched recordings down to claims the literature cannot support. Bar length is illustrative, not a measured metric.

What to carry into the witness box
  • 01"I matched the voiceprints" is the sentence that ends a career. Spectrogram eye-matching was never validated, and even the court that admitted it refused the fingerprint analogy the word smuggles in.
  • 02A likelihood ratio answers the evidence, not the verdict. It is not the probability your suspect is guilty, and saying so is the prosecutor’s fallacy. It means nothing until the system is calibrated.
  • 03A method’s accuracy is a measured number, the Cllr, on test recordings that match this case’s channel, duration, and speaking style. A useless system scores 1, and "off the shelf" is not "tuned to the case." Ask for the number and the Tippett plot.
  • 04The number rests on three judgements: the relevant population, calibration, and condition match. Change the population alone and a single comparison can swing more than an order of magnitude. Be ready to defend each choice, including the one you made without knowing who the offender was.
  • 05Keep the expert and the earwitness apart. Your instrumented, calibrated comparison is not a frightened bystander’s memory of a masked voice, and your expertise does not exempt you from bias. Willpower is not a debiasing procedure; controlling what you were told, and when, is.
  • 06You cannot guarantee the recording is a living human and not a clone. Your spoofing detector was validated on known attacks and may not catch a new one. And you do not infer nationality from speech.
Challenge 06 · Put it to the test

Is it even a human?

The recording in this case is of unknown provenance. Counsel asks the question that now comes before "who is speaking."

The question

"Your countermeasure was tuned against the spoofing attacks in the challenge data. The recording in this case was, you concede, of unknown provenance. Can you tell this jury, to any stated degree of confidence, that it was produced by a living human and not by a voice-cloning tool released after your detector was validated?"

Your answer
Ask the tutor

Still have questions about the research?

Ask anything about the forensic voice-comparison literature. The tutor answers from the document itself — and keeps one eye on how it might come up under cross-examination.

Your question
References
Next reading

Gait, Body and Clothing Comparison

Keep going

Counsel is briefed on this literature. Take it into the witness box and practise voice comparison.