Most job interviews are conversations with a verdict at the end. The hiring manager skims the resume five minutes before the call, asks whatever comes to mind, and walks away with a feeling about the person. That feeling drives the hire.
The loose process feels natural because nearly everyone has sat through it. I spent 17 years in agency recruiting sending candidates into these interviews and hearing about them afterward: distracted managers, resumes skimmed as the call connected. I even screened candidates this way myself. And one pattern I only later understood as the real problem: even the well-prepared interviewers asked no standard questions. Each candidate for the same role got a different interview. The format was consistent, but the structure wasn't.
The research on this way of working is grim, and in 2022 it got sharper. A major re-analysis of selection research put the validity of the unstructured interview at 0.19 (Sackett et al., 2022), less than half the 0.42 validity of a structured one.
Validity here is the correlation between interview scores and how the person actually performs on the job, on a scale from -1 to 1. Zero means no relationship; 1 means perfect positive prediction. Few hiring methods have validity much above 0.4. That is why 0.42 puts structured interviews at the top of the field, and why 0.19 is a serious indictment of the default, traditional interview format.
What did the 2022 re-analysis change?
It flipped a ranking the field had trusted since 1998 and demoted the method most companies use to near the bottom. When Sackett's team corrected a statistical error that had inflated validity estimates for decades, structured interviews moved to #1 and unstructured interviews fell to the bottom.
The old ranking came from Schmidt and Hunter's 1998 meta-analysis, the most-cited work in personnel selection, with more than 6,500 citations by 2023. That paper ranked cognitive ability tests as the strongest predictor of job performance, and a generation of hiring practice was built on it.
Then Sackett, Zhang, Berry, and Lievens re-ran the numbers and found a flaw hiding in plain sight. A hiring study can only track the people who got hired. Nobody knows how the rejected candidates would have performed. So researchers add a statistical boost to the results to account for that blind spot. The boost was designed for studies that test real job applicants. But about 80% of studies test people already in the job instead (Sackett et al., 2022). So for decades, we had the wrong boost, wrong studies, and inflated numbers.
When Sackett's team removed the unsupported corrections, the table reordered itself:
| Method | Predictive validity, Sackett et al. 2022 | Predictive validity, Schmidt & Hunter 1998 |
|---|---|---|
| Structured interviews | 0.42 | 0.51 |
| Job knowledge tests | 0.40 | 0.48 |
| Empirically keyed biodata | 0.38 | 0.35 |
| Work sample tests | 0.33 | 0.54 |
| Cognitive ability tests | 0.31 | 0.51 |
| Unstructured interviews | 0.19 | 0.38 |
Values are estimated operational validities (ρ) for predicting overall job performance.
Structured interviews moved to #1. Cognitive ability dropped to the middle of the pack, and a follow-up using only 21st-century studies put it lower still, at 0.22 (Sackett, Demeke, et al., 2024). Unstructured interviews, the method most companies use for their most expensive decisions, landed near the bottom.
The gap between the top and bottom of that table is the headline: structured interviews are 2.2x more predictive of job performance than unstructured interviews (0.42 ÷ 0.19; Sackett et al., 2022).
A ratio of two correlations is hard to picture. Dunlap (1994) described a common-language conversion for correlations under a bivariate-normal assumption: take two randomly selected people and estimate how often the person with the higher interview score also has the higher job-performance score. Applied illustratively, chance gets that ordering right 50 times out of 100. An unstructured interview at 0.19 raises the estimate to about 56, a 6-call edge. A structured interview at 0.42 raises it to about 64, a 14-call edge. These are illustrative pairwise odds rather than a literal hiring-accuracy rate. The practical point is the same: structure more than doubles the edge over chance. The tool behind most hiring decisions in most companies barely improves on chance.
Another pattern in the table matters. The strongest predictors sample the job itself through structured interviews, job knowledge tests, and work samples. Methods that infer performance from psychological traits cluster lower. The researchers call this "sample over sign." Asking someone to show or walk through the real work beats guessing from a trait score.
What counts as a structured interview?
Structure is a set of design choices, not a vibe. The definitive review counted 15 components across content and evaluation (Campion, Palmer, & Campion, 1997). A team building its first structured process should prioritize four: same job-based questions, answer-by-answer scoring, anchored rating scales, and trained interviewers.
Content choices cover the job analysis, consistent questions, and question types. Evaluation choices cover answer-level ratings, anchored scales, notes, training, and independent judgments. Later studies found that teams adopt these components unevenly, keeping the comfortable parts while skipping the ones that discipline judgment (Chapman & Zweig, 2005; Roulin, Bourdage, & Wingate, 2019).
The four priorities fit together: common questions create comparable evidence, while anchored answer-level scores give interviewers a shared way to judge it.
Same job-based questions. Every candidate gets the same questions, in the same order, drawn from an analysis of what the job requires. The job determines the questions instead of each interviewer's personal favorites.
Score each answer. Rate every response as it happens, against set criteria. A single gut rating at the end leaves room for the halo effect: one strong answer or charming opening can color everything that follows.
Anchored rating scales. Each score point gets a plain description of what an answer at that level sounds like. An anchor for a 4 might read: names the tradeoff, explains the call they made. Anchors keep one interviewer's 3 from being another interviewer's 5.
Trained interviewers. Nobody is born knowing how to follow up consistently or score against anchors. In a study of 131 professional interviewers, formal training was linked to more consistent questioning, better notes, and much more standardized scoring (Roulin, Bourdage, & Wingate, 2019).
For a TA leader, these components buy something beyond prediction. A structured process produces a written evidence trail behind every hire and every rejection: what was asked, what was said, how it was scored, and against what standard. When a hiring decision gets challenged, by a rejected candidate, a regulator, or your own leadership, that trail is the difference between a documented process and a shrug. Campion's team said as much back in 1997: fully unstructured interviews are hard to defend.
What problem does structure solve?
Noise. Bias gets the attention; noise does the quiet damage. Noise is the scatter in judgment: two interviewers score the same answer differently, or the same interviewer scores differently at 9 a.m. than at 4 p.m. Structure attacks noise by replacing one vague overall judgment with small, scorable ones applied the same way every time.
Bias is error that leans one way, while noise is error that scatters. Highhouse and Brooks (2023) argue that a century of selection research points to one fix: structure the process, set standards, apply them consistently, and combine judgments independently.
Structure works because it breaks one vague question about whether the team likes a candidate into small scorable ones: did they name the constraint, did they explain their method. Small questions leave less room for mood, halo, and drift.
The independence piece matters more than most teams expect. In assessment-center research, panels that discussed candidates and settled on a consensus rating did worse than a simple average of independent scores (Highhouse & Brooks, 2023). Discussion can feel rigorous while erasing the value of independent judgments. Have each interviewer score first, then compare the evidence.
Do structured interviews reduce bias?
They shrink the gap, but they don't erase it. Every assessment method produces some difference in average scores between demographic groups, and the question is how large. Structured interviews produced one of the smallest gaps of any high-validity method in the 2022 data, and no method gets to zero (Sackett et al., 2022).
Researchers measure the gap with a statistic called d. The rule of thumb: 0.2 is a small gap, 0.8 is a large one. Structured interviews came in at d = 0.23 for the Black-White score difference, near the small end of the scale. Cognitive ability tests came in at 0.79, near the large end, more than three times the size of the structured-interview gap.
Those figures drive a second finding. For decades the field believed in a hard tradeoff: the most accurate method, cognitive testing, also produced the biggest group gaps, so companies had to choose between prediction and fairness. A 2024 follow-up broke that assumption. Hiring systems that drop cognitive tests and lean on structured interviews and related methods lose little to no validity while cutting adverse impact substantially (Berry, Lievens, Zhang, & Sackett, 2024).
The fairness-versus-prediction tradeoff was overstated. A hiring process can be built for both, and structured interviews can carry a lot of that weight.
If structure works, why is it so rare?
Because the resistance is human, and the economics have made it worse. Interviewers like the interview they already run, they trust their own intuition more than the evidence warrants, and they hear "scorecard" as a threat to their professional standing. On top of that, building real structure has been slow, expensive work.
People like the interview they run. Lievens and De Paepe (2004) asked interviewers what they actually want: discretion over questions and scoring, informal personal contact with candidates, and interviews that are easy to prepare. Structure cuts against all three. In their sample, only 3% of interviewers reported the highest level of structure, with every question set in advance and every answer rated against example answers. HR managers show the same pull. Van der Zee, Bakker, and Bakker (2002) found they held stronger intentions toward unstructured interviews than structured ones, and their real practice matched. People run the interview they like, and the one they think their peers approve of. Notably, concern about assessing "fit" showed no relationship with structure use, so it does not explain the gap.
The intuition trap. Highhouse (2008) argues that HR professionals can know the research and still believe skilled interviewers read between the lines. Two beliefs drive it: that near-perfect prediction is possible if the assessor is good enough, and that experience sharpens intuitive judgment of people. The second belief fails on the evidence. Chapman and Zweig (2005) found years of interviewing experience were negatively related to using better question types (r = -.30). Experience can harden habits instead of improving them. The objection I have heard most from interviewers is some version of a built-in BS meter, and I understand it because I believed in mine. My questions really did get sharper with years in the seat. What I missed was that I tailored those questions to each candidate, which meant my evidence was never comparable from one candidate to the next. Experience had improved my interviewing and done nothing for my structure. I did not fully see the difference until I was deep into building a structured interviewing product and the research forced it on me.
The status threat. Nolan, Dalal, and Carter (2020) found that practitioners believed others in their company would see them as having less control when they ran a structured interview. They also rated the structured format as a greater threat to their professional worth. The implied threat is that a scorecard writes off the interviewer's BS meter and hard-earned experience.
The build cost. Structure is not free, and cost has been a dealbreaker for most organizations. In agency recruiting the pressure is time-to-submit: get the candidate on the phone, ask a few questions, write it up, send it out. Structure looks like a delay, though in 17 years I never saw anyone test that assumption. It was simply assumed. And the upfront work is real. A structured process requires a job analysis, questions written from it, anchored rating guides for every question, and interviewers trained to use them, per role. I was once part of a team that built this by hand for an enterprise client hiring dozens of product managers: job analysis, hiring-manager intakes, an interview with one of the company's top-performing PMs to map the skill set, a weighted skill model built in a spreadsheet, and a paid subject-matter expert to write the questions and rubrics. It worked. That engagement ran at roughly three candidate submissions for every two hires (a single client's result, one I led and tracked, not a study), and the process stayed fast. It also took a level of effort most companies will never fund for most roles, which is why enterprises hire consultancies at several hundred dollars an hour and nearly everyone else builds by hand or skips it. When the interviewers in Lievens and De Paepe's study said they wanted interviews that are easy to prepare, they were naming a real preparation burden. The method's evidence base has been public for decades. Its practice has been gated by budget and by the clock.
Underneath it all sits a training vacuum. Fewer than 34% of interviewers in Chapman and Zweig's sample had received any formal interview training. The typical interview is run from memory by someone using a process the company does not measure.
The most useful finding in this literature is about how to sell the change. When Nolan's team presented validity evidence as what employers lose by using unstructured interviews, the share of practitioners who preferred the structured format rose from about 36% to 60%. The study measured stated preference rather than later behavior, so loss framing is a rollout tactic, not a forecast. A structured-interview rollout is change management: it needs knowledge, tools, training, and sustained attention to the decisions that shape your people, culture, and work output.
What the research doesn't say
First, 0.42 is the mean across studies, not a guarantee for your team. The 80% credibility interval runs from 0.18 to 0.66 (Sackett et al., 2022). Where your team lands depends on how well you build and run the process, and the "structured" label covers interviews with very different levels of rigor.
Even at the low end of that range, structured interviews still show a positive relationship with job performance. The lower credibility value for unstructured interviews is -0.01, meaning that some settings may show no useful predictive relationship at all.
Second, a meta-analysis reports correlations across hundreds of studies. It cannot guarantee any single hire. A well-structured process will still miss sometimes, because people and jobs are messy. Across the available studies, structure shifts the odds in your favor more than any other method measured. No one can promise the exact validity your team will see.
How do you start?
Start with one open role. Define the job, ask every candidate the same questions, score each answer against prewritten anchors, and keep each interviewer's judgment independent until the evidence is ready to compare. Five steps are enough to build the core:
- Write down the skills the job requires on day one.
- Turn those requirements into questions, and ask every candidate the same ones.
- Write down what a good answer sounds like before the first interview.
- Score each answer as it happens, against those anchors.
- Have each interviewer score alone before the group compares evidence.
Those five steps create the core of a structured interview. The process takes more effort than scheduling a conversation. It also scored a higher mean validity than every other common method in Sackett et al.'s 2022 re-analysis. Most companies still have not made the switch.
The rest of this series covers how: what makes a scoring framework work (edition 2, on Ratio's RATE framework), why interviewers disagree and what fixes it, and what happens to interviews now that candidates apply with AI.
About Ratio
Ratio extracts and weights the skills a job needs for success from whatever you have on it (a job description, intake notes, a meeting transcript), then builds the complete interview kit: the questions, the follow-ups, and the scoring guides. The AI does the prep. A person runs the interview, judges the answers, and makes the call. Ratio never auto-scores candidates.
References
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068. https://doi.org/10.1037/apl0000994
- Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274. https://doi.org/10.1037/0033-2909.124.2.262
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2023). Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors. Industrial and Organizational Psychology, 16(3), 283–300. https://doi.org/10.1017/iop.2023.24
- Sackett, P. R., Demeke, S., Bazian, I. M., Griebie, A. M., Priest, R., & Kuncel, N. R. (2024). A contemporary look at the relationship between general cognitive ability and job performance. Journal of Applied Psychology, 109(5), 687–713. https://doi.org/10.1037/apl0001159
- Berry, C. M., Lievens, F., Zhang, C., & Sackett, P. R. (2024). Insights from an updated personnel selection meta-analytic matrix: Revisiting general mental ability tests' role in the validity-diversity trade-off. Journal of Applied Psychology, 109(10), 1611–1634. https://doi.org/10.1037/apl0001203
- Campion, M. A., Palmer, D. K., & Campion, J. E. (1997). A review of structure in the selection interview. Personnel Psychology, 50(3), 655–702. https://doi.org/10.1111/j.1744-6570.1997.tb00709.x
- Chapman, D. S., & Zweig, D. I. (2005). Developing a nomological network for interview structure: Antecedents and consequences of the structured selection interview. Personnel Psychology, 58(3), 673–702. https://doi.org/10.1111/j.1744-6570.2005.00516.x
- Roulin, N., Bourdage, J. S., & Wingate, T. G. (2019). Who is conducting "better" employment interviews? Antecedents of structured interview components use. Personnel Assessment and Decisions, 5(1), 37–48. https://doi.org/10.25035/pad.2019.01.002
- Highhouse, S. (2008). Stubborn reliance on intuition and subjectivity in employee selection. Industrial and Organizational Psychology, 1(3), 333–342. https://doi.org/10.1111/j.1754-9434.2008.00058.x
- Highhouse, S., & Brooks, M. E. (2023). Improving workplace judgments by reducing noise: Lessons learned from a century of selection research. Annual Review of Organizational Psychology and Organizational Behavior, 10, 519–533. https://doi.org/10.1146/annurev-orgpsych-120920-050708
- Van der Zee, K. I., Bakker, A. B., & Bakker, P. (2002). Why are structured interviews so rarely used in personnel selection? Journal of Applied Psychology, 87(1), 176–184. https://doi.org/10.1037/0021-9010.87.1.176
- Lievens, F., & De Paepe, A. (2004). An empirical investigation of interviewer-related factors that discourage the use of high structure interviews. Journal of Organizational Behavior, 25(1), 29–46. https://doi.org/10.1002/job.246
- Nolan, K. P., Dalal, D. K., & Carter, N. (2020). Threat of technological unemployment, use intentions, and the promotion of structured interviews in personnel selection. Personnel Assessment and Decisions, 6(2), 38–53. https://doi.org/10.25035/pad.2020.02.006
- Dunlap, W. P. (1994). Generalizing the common language effect size indicator to bivariate normal correlations. Psychological Bulletin, 116(3), 509–511. https://doi.org/10.1037/0033-2909.116.3.509
Frequently asked questions
Do structured interviews actually work?
Yes, and they were the strongest predictor of overall job performance in the largest recent re-analysis of selection research. Structured interviews showed an estimated operational validity of 0.42 against 0.19 for unstructured interviews, which makes them 2.2x more predictive as a method (Sackett et al., 2022). Validity is a correlation on a scale from -1 to 1, not an accuracy percentage. The 80% credibility interval runs from 0.18 to 0.66, so what a team gets depends on how well it builds and runs the process.
What does a validity of 0.42 actually mean?
It is the correlation between interview scores and later job performance, on a scale from -1 to 1. Zero means no relationship; 1 means perfect positive prediction. It does not mean 42% accurate. Under Dunlap's (1994) bivariate-normal common-language conversion, a structured interview raises the estimated odds of correctly ordering two randomly selected people by score and performance from 50% to about 64%. An unstructured interview raises them to about 56%. These are illustrative pairwise odds, not a literal hiring-accuracy rate (Sackett et al., 2022).
Do structured interviews work for junior candidates with no work history?
Yes, with the right question type. Behavioral questions need past experience to draw on; situational questions do not, because they ask what the candidate would do in a described situation. The 2023 follow-up notes structured interviews flex both ways, which some top predictors, like work samples, cannot (Sackett et al., 2023).
Are structured interviews cold or robotic?
A well-run structured interview can still feel human. Structure governs the scored questions and the scale. Rapport and the candidate's own questions belong around the scored section, so that every candidate provides comparable evidence inside a human conversation.
Do structured interviews eliminate bias?
No method eliminates bias. Structured interviews show smaller Black-White score differences than most alternatives (d = 0.23 vs. 0.79 for cognitive ability tests, less than a third the size), which reduces one part of the problem (Sackett et al., 2022).
Does cognitive testing predict job performance better?
Current estimates say no. The corrected 2022 estimate puts cognitive ability at 0.31 against 0.42 for structured interviews, and the 21st-century-only estimate is 0.22 (Sackett et al., 2022; Sackett, Demeke, et al., 2024). The 2024 analysis found that dropping cognitive tests from a selection system costs little validity and cuts adverse impact substantially (Berry et al., 2024).
Was this useful?