Data vs. Instinct: What Empirical Jury Research Says About Selection Accuracy
By David R. Drwencke · · 9 min read
Introduction
If you've tried more than a handful of cases, you've had this moment: you strike a juror because "something feels off," and you can't fully articulate why. You're not alone, and you're not necessarily wrong. But the empirical research on jury selection accuracy asks a harder question than "does instinct work?" It asks: does instinct work better than data? The answer, backed by decades of jury research and newer machine learning studies, is a qualified no. Systematic, data-driven approaches to jury selection consistently outperform pure instinct, but the margin is modest, highly case-dependent, and dwarfed by the strength of the underlying evidence. For trial attorneys, that's not a reason to throw out your gut — it's a reason to anchor it in validated data and structured methods.
This post walks through what the research actually shows, broken down by what it means for plaintiff's counsel, criminal defense, corporate defense, and prosecutors alike, because the data-vs-instinct question doesn't care which side of the "v." you're on.
1. How Accurate Are Instinctive Attorney Picks, Really?
Traditional jury selection — the kind built on years in the well, read-the-room experience, and pattern recognition — has been tested directly against more systematic methods, and the results are humbling for even the most confident trial lawyers.
- Reviews of empirical work find that attorneys using traditional, intuition-based selection perform about as well as early scientific jury selection (SJS) methods overall. In other words, decades of formal jury consulting methodology barely moved the needle over gut instinct in many studies.
- Across studies, demographic composition changes typically shift verdicts by only about 5–15% — a real effect, but a modest one, and far smaller than most attorneys assume when they're mentally sorting a panel by age, race, or occupation.
- One systematic selection method taught to lawyers improved prediction accuracy in only two of four mock cases — and only where the relationship between demographics and attitudes was strong and stable across the case type.
2. What Does Modern Data Say About Jury Selection Accuracy?
This is where the research gets genuinely interesting. Recent studies applying machine learning to juror questionnaires offer a controlled, apples-to-apples comparison between human jury consultants and predictive models — using the exact same data.
- In one controlled study, professional jury consultants and ML models were given identical pretrial questionnaire data and asked to predict juror verdict leanings.
- The consultants' majority judgment served as the human baseline. Two model types were tested: Random Forest and k-nearest neighbors (KNN).
- Consultant accuracy landed in the mid-60% range. Random Forest reached 0.818 accuracy, and KNN reached 0.796 — improvements of 12.3 and 10.1 percentage points, respectively, over the human consultants.
- The confidence intervals for that improvement (roughly 6 to 20 percentage points) indicate the models weren't marginally better by chance — they were consistently and meaningfully more accurate under identical informational constraints.
3. What Kind of Data Actually Predicts Juror Behavior?
Here's the finding that should change how attorneys build voir dire, regardless of practice area: attitudes and case-specific biases beat raw demographics, every time.
- Studies of scientific jury selection show that demographic traits — gender, race, age — have highly inconsistent relationships to verdicts, often flipping direction depending on case type. Men convict more often than women in some criminal trials and less often in others. There is no stable, generalizable demographic formula.
- Attitudinal measures tied specifically to the case type — for example, attitudes toward rape in a sexual assault trial, or attitudes toward corporate responsibility in a products liability case — are far stronger predictors of verdicts than demographics.
- Validated instruments matter here. The Pretrial Juror Attitude Questionnaire (PJAQ) predicts about 21% of pre-deliberation verdicts, 15.1% of post-deliberation verdicts, and accounts for 7.6% of the variance in verdict change during deliberation.
Why This Matters Across Practice Areas
- Plaintiff's attorneys often rely on demographic shorthand (assuming certain jurors are more "sympathetic"). The data suggests reframing voir dire around attitudes toward personal responsibility, damages caps, and corporate accountability instead.
- Criminal defense counsel frequently focus on race and age as proxies for leniency. The research says attitudes toward police credibility and the presumption of innocence are far more predictive.
- Corporate defense teams often overweight juror occupation or education level. Attitudes toward "big company" liability and regulatory skepticism carry more predictive weight.
- Prosecutors, similarly, benefit more from probing attitudes toward victim credibility and burden of proof than from demographic assumptions about who "believes the state."
4. The Limits of Selection: Evidence Still Dominates
Even the most sophisticated selection strategy has a ceiling — and the research is unambiguous about where that ceiling sits.
- Comprehensive reviews spanning 45 years of jury decision-making research confirm that trial evidence quality is the primary determinant of verdicts, far outweighing juror composition.
- Empirical studies of scientific jury selection conclude that jury composition usually changes verdicts only within a 5–15% band, and that strong evidence overwhelms juror predispositions.
- As jury researchers Kressel and Kressel summarized it: when the evidence is strong, other factors — including selection strategy — matter relatively little.
5. Practical Implications for Trial Attorneys
So what should you actually do differently on Monday morning? The research points to a few concrete rules of thumb.
Use Structured Data, Not Just Gut
Incorporate validated bias scales or case-tailored attitude questionnaires into pretrial research wherever feasible. Just as importantly, track your own voir dire decisions against actual outcomes over time. Build a firm-specific dataset. Test your instincts the way you'd test any other hypothesis — with follow-up data, not just confidence.
Prioritize Attitudes Over Demographics
Design voir dire questions to surface core beliefs relevant to your specific case theme — liability standards, punishment philosophy, corporate responsibility, police credibility — rather than relying on crude demographic proxies that the research shows are unreliable across case types.
Recognize Realistic Accuracy Ceilings
Even skilled, experienced jury consultants concede they can only estimate leanings, not predict verdicts with certainty. Accuracy depends heavily on data quality, voir dire time constraints, and juror candor. Set expectations with clients accordingly — jury selection reduces risk, it doesn't eliminate it.
Consider Technology and Analytics Where Feasible
Emerging machine learning tools show that algorithms can outperform human consultants under equal information constraints. Separately, NLP-based analysis of voir dire transcripts is beginning to flag disparate questioning patterns and potential bias in real time, with implications for both trial fairness and strategic decision-making — a development every trial team should be tracking.
Frequently Asked Questions
Q: Does data-driven jury selection guarantee a better outcome than instinct-based selection? No. The research shows a consistent but modest advantage for structured, data-driven methods — typically outperforming pure instinct by meaningful margins in controlled studies, but the underlying strength of the evidence at trial remains the dominant factor in verdicts. Q: Are demographics like race, gender, and age useless in jury selection? Not useless, but unreliable as standalone predictors. Studies show demographic-verdict relationships flip direction depending on case type, meaning a pattern that holds in one trial may reverse in another. Case-specific attitudes are consistently stronger predictors than demographics alone. Q: How much does jury composition actually change trial outcomes? Empirical studies of scientific jury selection estimate that jury composition shifts verdicts within roughly a 5–15% band. This is a real effect worth pursuing but far from decisive when evidence is strong on either side. Q: Can machine learning models really outperform experienced jury consultants? Under controlled conditions with identical questionnaire data, yes. One study found Random Forest and KNN models outperformed consultant judgments by roughly 10–12 percentage points, with confidence intervals confirming the improvement was statistically meaningful rather than incidental. Q: What should attorneys prioritize if they can't afford full jury consulting services? Focus voir dire questions on case-specific attitudes rather than demographics. Validated instruments like the Pretrial Juror Attitude Questionnaire (PJAQ) demonstrate that attitudinal data explains meaningfully more verdict variance than demographic profiling, and this approach can be applied without expensive technology. Q: Does this research apply equally to criminal and civil trials? Yes, with case-specific calibration. The underlying principle — attitudes tied to case themes outperform demographics, and evidence strength dominates both — holds across criminal defense, prosecution, plaintiff, and corporate defense contexts, though the specific attitudinal predictors differ by case type.Closing Thoughts
The data-versus-instinct debate in jury selection isn't really a debate once you look at the numbers. It's a calibration problem. Trial attorneys who rely purely on instinct are leaving measurable accuracy on the table; attorneys who rely purely on data without contextual judgment miss the nuance that keeps algorithms from replacing experienced counsel altogether. The winning approach blends both: structured, attitude-focused questioning, honest tracking of your own selection outcomes, and a clear-eyed acknowledgment that no jury selection strategy — however sophisticated — will out-muscle a strong evidentiary record.
Tools like StrikeList AI are part of this broader shift toward structured jury analytics — helping trial teams organize voir dire data, track juror attitudes systematically, and bring a more evidence-based lens to a process that has historically run almost entirely on instinct. Whether or not you use a platform like it, the research is clear: the future of effective jury selection isn't data instead of instinct, or instinct instead of data. It's both, working together, with a realistic understanding of what selection can and can't do for your case.
Frequently Asked Questions
- Q: Does data-driven jury selection guarantee a better outcome than instinct-based selection?
- No. The research shows a consistent but modest advantage for structured, data-driven methods — typically outperforming pure instinct by meaningful margins in controlled studies, but the underlying strength of the evidence at trial remains the dominant factor in verdicts. Q: Are demographics like race, gender, and age useless in jury selection? Not useless, but unreliable as standalone predictors. Studies show demographic-verdict relationships flip direction depending on case type, meaning a pattern that holds in one trial may reverse in another. Case-specific attitudes are consistently stronger predictors than demographics alone. Q: How much does jury composition actually change trial outcomes? Empirical studies of scientific jury selection estimate that jury composition shifts verdicts within roughly a 5–15% band. This is a real effect worth pursuing but far from decisive when evidence is strong on either side. Q: Can machine learning models really outperform experienced jury consultants? Under controlled conditions with identical questionnaire data, yes. One study found Random Forest and KNN models outperformed consultant judgments by roughly 10–12 percentage points, with confidence intervals confirming the improvement was statistically meaningful rather than incidental. Q: What should attorneys prioriti