
Firearm & Toolmark Identification: What the Marks Can Support
Firearm and toolmark examiners compare the marks a gun or tool leaves and decide whether two came from the same source. The comparison looks mechanical, but the decision is a judgement with no defined threshold. Seven sections examine what the discipline can support: the uniqueness premise, the black-box error rates and how inconclusives are counted, contextual bias and non-blind verification, whether objective methods change the picture, toolmarks on bone, and the conclusion language that has been condemned in court.
The uniqueness premise and the missing threshold
Firearm and toolmark identification rests on a rule the discipline calls the theory of identification. When the marks on two items agree well enough, the examiner concludes they came from the same source. The governing standard, from the Association of Firearm and Tool Mark Examiners, is that agreement counts as an identification when it exceeds the best agreement seen between marks known to have been made by different tools, so that another tool making the mark is a practical impossibility. The theory states plainly that this judgement is subjective.
The difficulty is in the phrase "well enough." There is no number behind it. Ronald Nichols, who wrote the discipline's leading defence, concedes the point directly: there is no universal agreement as to how much correspondence exceeds the best known non-matching situation. So the threshold at which two marks become a match is not written down anywhere. It sits in the individual examiner's head, built from training and experience.
Under that judgement lies an assumption. Saks and Koehler, writing in Science in 2005, named it the assumption of discernible uniqueness: that marks made by different sources are observably different, and that an examiner can see the difference. They called it an assumption lacking theoretical or empirical foundations. The point is not that two guns are identical. It is that "unique" has never been measured, so it cannot be sworn to as a finding, and even a true uniqueness would not tell you an examiner can reliably perceive it.
William Tobin and Peter Blau put the circularity in the open in Jurimetrics: an examiner may call an identification when there is sufficient agreement, and sufficient agreement is defined as enough agreement for an identification. There is no rule that a single genuine difference forces an exclusion. The examiner is free to count the marks that agree and explain away the ones that do not.
“is inherently vague and tautological, with circular reasoning. An examiner”
There is a second problem the examiner has to manage, and it cuts toward false identifications rather than missed ones. Tools that come off the same production line can leave the same marks. These are called subclass characteristics, and they are shared by many items made on the same machine. A single button used to rifle barrels can serve thousands of guns; a broach, hundreds. A 2024 study funded by the US National Institute of Standards and Technology drove the point home with numbers: on unfinished broached breech faces, an objective algorithm put 89% of comparisons between different guns above the threshold it uses to declare a match. The signature the examiner attributes to one gun can be shared by every gun made on the same tooling, none of which the examiner has ever seen.
None of this means the work is guesswork. A trained examiner distinguishes marks that untrained people cannot, and subclass is a known hazard that a careful examiner watches for. The honest reading of the foundation is narrower. The discipline can compare marks and describe how well they correspond. What it cannot do is point to a measured threshold, or a measured rate of shared marks across the world's guns, that would justify naming one source to the exclusion of the rest.

Where is the threshold?
Counsel accepts your training, then asks what number a match rests on.
"I put it to you that there is no fixed number, no minimum count of marks and no percentage, at which the law or your own profession says two marks match. It comes down to when you, personally, feel you have seen enough. That is the truth of it, isn't it?"
The black-box studies and what they measure
The discipline's answer to the validity question is a handful of black-box studies, in which examiners judge comparisons whose true source is known and their answers are scored. The headline numbers are low. The 2014 Ames study, which the President's Council of Advisors on Science and Technology relied on, reported a false-positive rate of about 1% on cartridge cases. The 2022 Ames II study reported false-positive rates under 1% for both bullets and cartridge cases. Taken alone, those numbers sound reassuring.
Two things sit underneath them. First, the errors were not spread evenly. In the 2014 study, most of the false identifications came from a small number of examiners. Second, and more important, the studies count only conclusive answers. When examiners in Ames II were shown cartridge cases fired by different guns, they reached the correct answer, elimination, only a third to a half of the time. The rest they called inconclusive, and inconclusives are left out of the error rate. A discipline that correctly clears a different gun half the time at best is not the same as one that is right 99% of the time.
The design of the studies matters as much as the scoring. Several are closed-set proficiency tests: the examiner knows the answer is somewhere in the set, which is not the casework question. Cliff Spiegelman and William Tobin, reviewing the studies courts rely on, noted that if being kept from the answers counts as a blind study, then most university exams are blind too. The examiners who take part are volunteers, and in Ames II roughly a third of those who enrolled dropped out, which biases the sample toward the confident. Nicholas Scurich and Hal Stern found that in the Ames II data, about a third of the reported correct eliminations rested on class differences the examiner could see at a glance, not on the fine individual marks the method is meant to test.
And one widely cited figure needs its history told. A 2020 study of consecutively manufactured Beretta barrels reported a false-identification rate of about one in 1,250. That number appears only after the author reclassified six of the seven false identifications in the raw data as administrative or transcription errors. On the answers as submitted, the rate was one in 179.
“If we are to accept that definition of what constitutes a blind study, then most college tests are also blind, since most test administrators do not provide answers before students take tests.”
There is a further result that speaks to reliability rather than accuracy. When the Ames II examiners re-examined the same comparisons months later, they disagreed with their own earlier conclusions about 35% of the time on non-matching cartridge cases. When two examiners judged the same non-matching comparison, they agreed only about 36% of the time. An opinion that the same examiner will not reproduce, and that a second examiner will often not share, is not a stable measurement.
The fair statement of the black-box work is this. These are the best data the discipline has, and on the reported basis the false-positive numbers are genuinely low. But the numbers describe specific examiners judging specific guns in a test they knew was a test, with every difficult non-answer set aside. They are, as the Ames II authors themselves say, an upper bound on accuracy, not a measurement of casework error.

A third of the time
Counsel accepts the studies exist, then asks what they measured.
"I put it to you that when examiners in these studies were shown cartridges fired by different guns, they reached the correct answer, elimination, only a third to a half of the time. The rest they called inconclusive, and every one of those was left out of the error rate. A discipline that cannot reliably tell two different guns apart cannot honestly claim near-certainty when it says two are the same, can it?"
The inconclusive-accounting war
The single most contested question in firearms is how to count an inconclusive answer. It sounds like bookkeeping. It changes the error rate by more than an order of magnitude.
A 2023 study in the Proceedings of the National Academy of Sciences by Guyll and colleagues called firearm examination highly valid, reporting accuracy near 99%. That figure counts only conclusive decisions. The same study's own numbers tell the other story. Put the inconclusives back in and the ability to correctly identify a same-source pair falls from 99.9% to 93.4%, and the ability to correctly clear a different-source pair, the specificity, falls from 99.1% to 63.5%. In the same data, examiners answered inconclusive six times more often on different-gun comparisons than on same-gun comparisons.
That asymmetry is the heart of the problem. Inconclusive answers cluster on the comparisons of guns that do not match. So when the field scores an inconclusive as correct, or leaves it out, it is quietly propping up its ability to clear the innocent while hiding how often it fails to. Alan Dorfman and Richard Valliant worked the arithmetic on one dataset: the false-positive rate for different-source bullets is 0.70% if inconclusives count as correct, and 66% if they count as potential errors. Same data. Same answers. A scoring choice the field made in its own favour.
The reformers' sharpest demonstration is also the simplest. If an examiner never has to be counted wrong for saying inconclusive, then an examiner can reach a perfect score by saying inconclusive to everything. Itiel Dror and Nicholas Scurich pointed to a study in which 98.3% of decisions were inconclusive, leaving almost no room for a measured error. A low published error rate can mean an examiner is accurate. It can equally mean an examiner is willing to abstain.
“The number of inconclusive responses observed in firearm error rate studies is staggering.”
The other side of this argument is worth stating fairly, because it is coherent. An identification or an elimination asserts something about the world that is either right or wrong. An inconclusive asserts nothing about the source. Hal Arkes and Jonathan Koehler put it simply: an inconclusive is not an error, it is a pass. Alex Biedermann and Kyriakos Kotsoglou add that calling a non-assertion an error forces the field to abandon ground truth and substitute a majority vote for what the answer should have been, which brings its own problems. On this view an inconclusive on genuinely poor marks is the honest and cautious call, and forcing a conclusion would manufacture error.
Both things can be true at once. An inconclusive need not be scored as wrong, and counting inconclusives as correct still deflates the error rate and hides a one-directional failure to exclude. The conclusion for the witness box is not that the discipline is worthless. It is that the honest error rate is unknown, bounded well above the advertised figure, and that a low number on a proficiency test measures willingness to commit as much as it measures accuracy.

Scored in your favour
Counsel holds the low error rate up and asks how it was calculated.
"I put it to you that the low error rate you rely on is calculated only after every inconclusive answer is set aside, so that on the very comparisons where examiners could not decide, the figure pretends those cases never happened. Count them honestly and your discipline's false-positive rate rises from under one per cent to something in the region of a third or more, doesn't it?"
Bias, verification, and the jury
A firearms examiner rarely works on the marks alone. The case comes with a theory, a suspect, and often the conclusion of a first examiner. Cognitive science shows that context shapes a judgement without the person feeling it move, and the firearms-specific evidence is direct.
The clearest result is a field study inside real casework. Erwin Mattijssen and colleagues at the Netherlands Forensic Institute compared blind and non-blind peer review of firearm conclusions. When the second examiner could not see the first examiner's conclusion, they disagreed about the strength of the comparison 42.3% of the time. When they could see it, disagreement fell to 12.5%. The reviewer who already knows the answer drifts toward it. So the second signature on the report, the one presented in court as independent verification, was for the most part not an independent check at all.
The discipline's own view of bias makes the problem worse. In a survey of 403 examiners, most regarded their judgements as nearly infallible; 37% claimed a personal accuracy rate of 100%, and 71% believed an examiner can defeat bias simply by trying to ignore expectations. That belief is exactly the one the science rejects. Bias is not a failure of effort or honesty; it operates outside awareness, and trying to suppress a thought is not a reliable way to remove it. An examiner who tells the court "I set the context aside" is describing the very move the research says does not work.
There is a structural reason firearms cannot use the standard fix. Sequential unmasking, which works for DNA and fingerprints, asks the examiner to record the evidence trace before seeing the reference. In firearms the examiner must view the questioned and the reference marks together, side by side, to compare them at all. So the examiner necessarily sees the answer they are aiming at while forming the opinion, which is the condition under which an honest analyst can find agreement a blind analyst would not.
“71% believed that examiners can reduce bias by simply trying to ignore their expectations.”
The last link is the jury. Brandon Garrett and colleagues ran mock jurors through the same case with different conclusion language. Declaring a match raised the odds of conviction roughly five to six times over an inconclusive, and softening the wording made almost no difference: jurors treated "consistent with" much as they treated a flat identification. In the same study, conviction rates were 65% when the examiner said identification and 57% when the examiner said consistent with, a gap too small to matter. And cross-examination, the safeguard courts rely on to expose a weak opinion, did not significantly move jurors' verdicts.
The honest position on the stand is not a claim of immunity. It is a description of what was done: whether the analysis was recorded before the case theory was known, and whether the verifier was blind to the first conclusion. Those are procedures a lab either followed or did not. "I kept an open mind" is neither, and the evidence says the jury will over-weight the conclusion regardless.

Was the check blind?
Counsel asks about the second signature on the report.
"I put it to you that the independent verification on your report was nothing of the kind. The second examiner was handed your conclusion before forming their own; and when a reviewer already knows the answer you reached, a silent agreement is not confirmation, it is deference, particularly to a senior colleague. That is what happened here, isn't it?"
Can objective methods rescue it?
If the trouble is a subjective threshold, the obvious answer is to make the comparison objective. Researchers have built methods that do exactly that: congruent matching cells, which divide a 3D surface scan into small regions and count how many agree; algorithmic likelihood ratios that score similarity numerically; and automated correlation systems that search image databases. This is real progress, and the numbers on the best datasets are strong. One 3D likelihood-ratio method reported values exceeding a billion for same-source pairs.
Two cautions travel with those numbers, and they are the ones a cross-examiner will use. The first concerns databases. The National Integrated Ballistic Information Network, NIBIN, searches images of cartridge cases and returns a ranked list of candidates. A candidate is not a match. As the network's own evaluation states, a high-confidence candidate must be manually confirmed by an examiner on a comparison microscope before it becomes a hit. The database narrows the field and generates an investigative lead; the identification, if there is one, is still the human examiner's, with all the limits the earlier sections describe.
The second caution concerns the numbers themselves. A high similarity score is not the same as a valid likelihood ratio. Geoffrey Morrison and Ewald Enzinger showed that a score measuring only how alike two items are cannot yield an interpretable weight of evidence, because it ignores how common that degree of similarity is in the relevant population. A system can separate matches from non-matches almost perfectly and still report a number that is far too strong. And the reference data behind these methods are narrow. The billion-to-one method was built on a small set of one make of pistol firing one type of ammunition, and its authors warn it can barely be generalised even to the same model of gun. A fresh reference set must be built for each firearm and each cartridge type, or the error rate rises.
There is also an honesty point the authors themselves make. When a validation study reports zero errors, that means no errors were seen in a small sample, not that the true rate is zero. And the error that matters most in a real case, a specimen mix-up or a labelling mistake in the lab, can be many times more likely than any coincidental-match probability the algorithm produces.
“High confidence candidates must be manually confirmed in order to constitute a hit.”
So the objective methods change the picture without settling it. They replace a private judgement with a reproducible number, a shared data format, and a stated error rate, which is genuine progress the field should be credited for. But the strongest results come from small, mostly consecutively manufactured datasets that cannot themselves test the uniqueness assumption; the scores need calibration against the right population before they become weights of evidence; and the tool in daily use, the database, produces leads for a human to confirm rather than identifications. A courtroom-ready, validated, calibrated objective method is not yet here.

The database matched it
You told the jury a database matched the cartridge to the suspect's gun. Counsel asks what the database actually did.
"You told the jury the ballistic database matched this cartridge to my client's gun. But a database of that kind only searches images and returns a ranked list of candidates for a human examiner to check under a microscope. It produces an investigative lead, not a match, doesn't it? The identification, if there is one, is still your own judgement, with all its limits."
Toolmarks beyond firearms
The same comparison logic is applied well beyond firearms, to the marks left by screwdrivers, bolt cutters, saws and knives, and often with far less validation behind it. The most sobering results come from marks on the body.
Consider cut marks in cartilage. Christian Crowder and colleagues, and separately a group led by Jennifer Love, tested whether trained analysts could read a knife's class, its type, from cuts in costal cartilage using the accepted method. The result was a misclassification rate above 65%. Three trained analysts, asked the comparatively easy question of what kind of blade made the cut, got it wrong roughly two-thirds of the time. And that is a coarser claim than naming one specific knife. On bone, the same pattern appears in a narrower form: the direction of the bevel on a cut was wrong more than half the time, against under 17% in wax.
Two mechanisms explain why these marks mislead. The first is the angle of the cut. Puentes and Cardoso found that pushing a knife through cartilage with a forward stroke rather than straight down shrinks the spacing of the striations by 60 to 70%, so that one knife can leave marks that mimic a differently spaced blade, and the angle cannot be recovered from the wound. The second is the assumption the examiner brings. A 2025 study of 472 cuts from 34 saws, the largest of its kind, confirmed some long-standing rules and overturned others. It found that the flare of a saw kerf, long used to tell a jury which side the sawyer stood on, has no reliable relationship to the handle side at all.
The subclass problem returns here too, and it is the engine of false identifications between different tools. When different tools are made one after another, their marks can be alike enough to be mistaken for a match. A 2025 breech-face study found that on some manufacturing methods, up to 40% of comparisons between different tools crossed the threshold the algorithm calls an identification.
“The application of current accepted method for tool mark analysis of cut costal cartilage resulted in >65% misclassification rate”
The honest use of toolmark evidence on the body is real but narrow. The literature supports class-level statements and exclusions: this cut is consistent with a serrated blade, or a particular saw is excluded by kerf width. The founding text of saw-mark analysis says as much, calling a positive identification of one saw a rare occurrence, and the discipline's own guidance describes naming a specific tool from a mark as an unacceptable practice. What the marks cannot support is the confident individualisation, this saw and no other, that has still been offered in court. Where the discipline stays at the level of class and exclusion, and states the conditions it is assuming about angle and substrate, it is on defensible ground. Beyond that it is not.

Which tool, really?
You told the jury the defendant's tool made the marks. Counsel asks what the marks can actually establish.
"You told the jury my client's saw made these marks on the bone. But in a controlled test, when trained analysts were asked to identify only the type of blade that cut through cartilage, a far easier question than naming one tool, they got it wrong roughly two-thirds of the time. You are asking these twelve people to go further than that and name the very tool, aren't you?"
Conclusion language, documented failures, and admissibility
How a firearms conclusion is worded is where cases are won and lost. The signature overstatement is source attribution to the exclusion of all other firearms in the world, and its relatives: practical impossibility, a hundred per cent certainty, a match to an exact statistical certainty, it is like a fingerprint. Since United States v. Green in 2005, a growing line of courts has ruled these phrases inadmissible, and the discipline's own governing language now bars them. The United States Department of Justice's uniform language tells its examiners they shall not claim a source identification to the exclusion of all other sources, shall not use the word individualise, and shall not assert that their examinations are infallible or have a zero error rate.
The gap between what the examiner means and what the jury hears is measurable. When Thomas Busey and colleagues surveyed jurors, 71% took the word identification to mean to the exclusion of all others. And a reanalysis of the error-rate studies by Aggadi and colleagues found that the strength implied by the word identification runs ahead of the measured strength of the evidence by up to five orders of magnitude. The word does work in the courtroom that the data cannot support.
The discipline also has a documented failure to reckon with, and it is the closest parallel it owns. Comparative bullet-lead analysis compared the trace-element composition of bullets and testified that two came from the same batch of lead, often the same box. It ran for close to 40 years, admitted almost without challenge, because the FBI laboratory was the sole provider and the analysis required a nuclear reactor, so no defence expert could test it. Then in 2002 a metallurgical study showed that the whole method rested on three assumptions, representativeness, homogeneity and uniqueness of the molten source, and that industry production data falsified all three: many different batches of lead were chemically indistinguishable. The FBI discontinued the practice in 2005. Bullet-lead analysis was precise, instrument-based, and wrong. Precision of measurement never guaranteed validity of the inference, and four decades of admissibility never made it science.
“If this technique had not been used for years to send people to prison, no reasonable scholars of forensic evidence would consider it ready for court.”
Firearm and toolmark conclusions fail most often in the wording. Each phrase below has been condemned by courts or by the discipline's own language standards; the alternative keeps the sentence inside what the comparison can support.
Courts are moving, unevenly. Since 2005 a line of decisions has limited firearm testimony, some restricting the examiner to describing similarities, some barring the certainty language, and the Maryland high court in Abruquah v. State in 2023 pushed back on unqualified identification language. But most courts still admit the evidence, largely on precedent and long use rather than fresh findings of reliability, which is the very ground the bullet-lead story warns against. The examiner who wants to be believed, and to stay believed on appeal, states what the marks can support: that they correspond, how strongly, under what assumptions, and with what error rate, and not that one gun fired the round to the exclusion of the world.
- 01There is no numeric threshold for "sufficient agreement." The match point is a subjective judgement, and the uniqueness under it has never been measured. Say that plainly rather than defend experience as a measured standard.
- 02The black-box error rates exclude inconclusives, and examiners correctly excluded different guns only a third to a half of the time. The studies are proficiency tests their own authors call an upper bound, not a casework error rate.
- 03How inconclusives are counted swings the error rate by more than an order of magnitude, and inconclusives cluster on different-source comparisons. A low published rate reflects willingness to abstain as much as accuracy.
- 04Non-blind verification is not an independent check: reviewers who see the first conclusion disagree far less often. Point to whether the verification was blind and whether the analysis was recorded before the case theory was known.
- 05Objective methods and databases are progress, not a rescue. A database hit is a lead a human must confirm, and a high similarity score is not a valid likelihood ratio unless calibrated on the right population.
- 06On the body, toolmark comparison supports class and exclusion, not individualisation. Trained analysts misread even the blade type in cartilage about two-thirds of the time.
- 07"To the exclusion of all other firearms," "zero error rate," and "it is like a fingerprint" are condemned in court and barred by the discipline's own language. Comparative bullet-lead analysis ran 40 years on precision before it collapsed. Precision is not validity.
To the exclusion of the world
You have testified to an identification. Counsel asks what it excludes.
"You told this jury that the bullet that killed the victim came from my client's gun to the exclusion of every other firearm on earth. There is no database of the world's guns, no count, no measurement behind that claim. You cannot actually support it, can you?"
Still have questions about the research?
Ask anything about the firearm and toolmark identification literature. The tutor answers from the document itself — and keeps one eye on how it might come up under cross-examination.
- Saks, M. J., & Koehler, J. J. (2005). The coming paradigm shift in forensic identification science. Science, 309(5736), 892–895.
- Nichols, R. G. (2007). Defending the scientific foundations of the firearms and tool mark identification discipline. Journal of Forensic Sciences, 52(3), 586–594.
- Schwartz, A. (2007). Commentary on the scientific foundations of firearms and tool mark identification. Journal of Forensic Sciences, 52(6), 1414–1415.
- Page, M., Taylor, J., & Blenkin, M. (2011). Uniqueness in the forensic identification sciences — Fact or fiction? Forensic Science International, 206(1–3), 12–18.
- Tobin, W. A., & Blau, P. (2013). Hypothesis testing of the critical underlying premise of discernible uniqueness in firearms-toolmarks forensic practice. Jurimetrics, 53(2), 121–142.
- Franklin, R., & Morris, K. (2024). Subclass characteristics on consecutively manufactured breech faces. Journal of Forensic Sciences, 69, 2041–2053.
- Baldwin, D. P., Bajic, S. J., Morris, M., & Zamzow, D. (2014). A study of false-positive and false-negative error rates in cartridge case comparisons. Ames Laboratory Technical Report IS-5207. US Department of Energy.
- Monson, K. L., Smith, E. D., & Peters, E. M. (2023). Accuracy of comparison decisions by forensic firearms examiners. Journal of Forensic Sciences, 68(1), 86–100.
- Monson, K. L., Smith, E. D., & Peters, E. M. (2023). Repeatability and reproducibility of comparison decisions by forensic firearms examiners. Journal of Forensic Sciences, 68(5), 1721–1740.
- Spiegelman, C., & Tobin, W. A. (2013). Analysis of experiments in forensic firearms/toolmarks practice offered as support for low rates of practice error and claims of inferential certainty. Law, Probability and Risk, 12(2), 115–133.
- Scurich, N., & Stern, H. S. (2023). Letter: undisclosed different-class comparisons in the Ames II study. Journal of Forensic Sciences, 68(3), 1093–1094.
- Guyll, M., Madon, S., Yang, Y., Burd, K. A., & Wells, G. (2023). Validity of forensic cartridge-case comparisons. Proceedings of the National Academy of Sciences, 120(20), e2210428120.
- Dorfman, A. H., & Valliant, R. (2022). Inconclusives, errors, and error rates in forensic firearms analysis: Three statistical perspectives. Forensic Science International: Synergy, 5, 100273.
- Arkes, H. R., & Koehler, J. J. (2022). Inconclusives and error rates in forensic science: A signal detection theory approach. Law, Probability and Risk, 20(4), 153–168.
- Scurich, N., & John, R. S. (2023). Three-way ROCs for forensic decision making. Statistics and Public Policy, 10(1), 2239306.
- Dror, I. E., & Scurich, N. (2020). (Mis)use of scientific measurements in forensic science. Forensic Science International: Synergy, 2, 333–338.
- Biedermann, A., & Kotsoglou, K. N. (2021). Forensic science and the principle of excluded middle: Inconclusive decisions and the structure of error rate studies. Forensic Science International: Synergy, 3, 100147.
- Mattijssen, E. J. A. T., Witteman, C. L. M., Berger, C. E. H., & Stoel, R. D. (2020). Cognitive biases in the peer review of bullet and cartridge case comparison casework: A field study. Science & Justice, 60(4), 337–346.
- Kukucka, J., Kassin, S. M., Zapf, P. A., & Dror, I. E. (2017). Cognitive bias and blindspots: A survey of forensic science examiners. Journal of Applied Research in Memory and Cognition, 6(4), 452–459.
- Garrett, B. L., Scurich, N., & Crozier, W. E. (2020). Mock jurors’ evaluation of firearm examiner testimony. Law and Human Behavior, 44(5), 412–423.
- King, W. R., Wells, W., Katz, C. M., Maguire, E. R., & Frank, J. (2013). Opening the black box of NIBIN. US National Institute of Justice, Report No. 243875.
- Morrison, G. S., & Enzinger, E. (2018). Score-based procedures for the calculation of forensic likelihood ratios. Science & Justice, 58(1), 47–58.
- Riva, F., & Champod, C. (2014). Automatic comparison and evaluation of impressions left by a firearm on fired cartridge cases. Journal of Forensic Sciences, 59(3), 637–647.
- Love, J. C., Derrick, S. M., Wiersema, J. M., & Peters, C. (2012). Microscopic analysis of sharp force trauma in bone and cartilage: A validation study. Journal of Forensic Sciences, 57(2), 355–360.
- Crowder, C., Rainwater, C. W., & Fridie, J. S. (2013). Microscopic analysis of sharp force trauma in bone and cartilage: A validation study of class characteristics. Journal of Forensic Sciences, 58(5), 1119–1126.
- Puentes, K., & Cardoso, H. F. V. (2013). Reliability of cut mark analysis in human costal cartilage. Forensic Science International, 231(1–3), 244–248.
- Busey, T., Klutzke, B., Nuzzi, L., & Vanderkolk, J. (2022). Validating strength-of-support conclusion scales for fingerprint, footwear, and toolmark impressions. Journal of Forensic Sciences, 67(3), 936–950.
- Aggadi, N., Zeller, D., & Busey, T. (2025). Quantifying the strength of firearms comparisons based on error rate studies. Journal of Forensic Sciences, 70(1), 84–96.
- Randich, E., Duerfeldt, W., McLendon, W., & Tobin, W. (2002). A metallurgical review of the interpretation of bullet lead compositional analysis. Forensic Science International, 127(3), 174–191.
- Tobin, W. A. (2004). Comparative bullet lead analysis: A case study in flawed forensics. The Champion (NACDL), July 2004, 12–21.
- Kaye, D. H. (2018). Firearm-mark evidence: Looking back and looking ahead. Case Western Reserve Law Review, 68(3), 723–757.
- Abruquah v. State, 483 Md. 637 (2023) (Maryland Court of Appeals).
Ballistics & Gunshot Residue: What the Physics and the Particle Can Support
Counsel is briefed on this literature. Take it into the witness box and practise firearms, ballistics & GSR.