Why do two interviewers disagree about the same answer?
Put two experienced interviewers in the same room with the same candidate. Ask them afterward to rate the candidate's answer to one question. They will often land a full point apart, and each will be confident.
Nothing went wrong with either interviewer. That is what unstandardized judgment looks like. Highhouse and Brooks (2023) call it noise: unwanted variation in judgment that has nothing to do with the candidate. It comes in two forms. Disagreement noise is two judges reading the same evidence differently. Occasion noise is one judge reading the same evidence differently at 9 a.m. and 4 p.m., after a good candidate or a bad one, on a Monday or a Friday.
Bias gets more attention because it has a direction and a name. Noise is quieter. It shows up as a remark about the tough grader on the panel, or the afternoon group that liked everyone, and it gets absorbed as the cost of doing business.
Noise is the largest source of error a hiring team can control, and the research on how to reduce it is about a hundred years old. Highhouse and Brooks summarize it in one line: set a standard, apply the standard consistently, and aggregate independent judgments.
Everything in this paper is a practical version of that sentence.
Which part of structure do interviewers skip?
They skip the scoring. Interviewers adopt the visible, conversational parts of structure at high rates and leave the part that actually reduces noise, which is the evaluation half: written anchors, answer-level scores, and a fixed rule for combining them. The gap is not knowledge. Every interviewer in the 2019 sample knew what a rating scale was.
Roulin, Bourdage, and Wingate (2019) surveyed 131 professional interviewers on their use of seven structure components. Reported use, on a 5-point scale:
| Component | Mean use |
|---|---|
| Follow-up questioning | 4.03 |
| Question consistency | 3.99 |
| Rapport building | 3.89 |
| Note taking | 3.85 |
| Standardized evaluation | 2.66 |
Interviewers ask the same questions, take notes, build rapport, and follow up. Then they score however they score. It is a habit gap. Scoring against a written standard feels like grading, and grading feels like it is not the interviewer's job.
Chapman and Zweig (2005) found the same pattern earlier and added the number that explains it: fewer than 34% of their interviewers had any formal interview training. Nobody had shown them what a 3 means.
Standard questions with unstandardized scoring is the most common shape of structured interviewing in practice. It gets most of the work and a fraction of the benefit.
What does a scoring standard consist of?
Five parts, four of them evaluation components that Campion, Palmer, and Campion (1997) rank among the ones that matter most: rate each answer, use anchored rating scales, do not discuss candidates between interviews, and combine scores by a fixed rule. Add interviewer training, which they also list, and a team has the whole scoring standard. The rest of this section takes them one at a time.
Score the answer, not the interview
One global rating at the end of an interview is where halo effects live. A strong first answer colors the next five. A likeable candidate gets the benefit of the doubt on a thin answer. A nervous one gets docked for a strong answer delivered badly.
Scoring each answer on its own breaks one vague judgment into several narrow ones. Instead of asking whether the team likes a candidate, an interviewer asks whether the answer named real specifics, whether the candidate described their own approach or the team's, whether they said what made it hard, and whether they showed it worked. Narrow judgments are easier to make, easier to defend, and much harder for a first impression to override.
This is also why the skill, not the question, should be the unit of the debrief. A panel that discusses a candidate question by question can let a strong answer on one skill leak into the rating on another. Reviewing one skill at a time keeps the evidence where it belongs.
Write the anchors before the interview
An anchor is a description of what a score sounds like. Not a definition of the skill; a description of the answer.
The useful shape is two ends and a floor. What does an empty answer sound like? What does a strong one sound like? And, for the questions that carry the most weight, what specific phrase would tell you the candidate is reciting rather than recalling?
Anchors work because they move the standard from the interviewer's head to the page. Two interviewers who share a written anchor for a strong answer can still disagree, but they disagree about the evidence, which is a productive argument. Two interviewers with no anchor disagree about their own private standards, which is not.
A small design choice matters here. Anchors that list every acceptable strong answer turn into an answer key, and an answer key can be coached. Anchors that describe the shape of a strong answer and name the empty one are harder to game and easier to apply. An anchor might tell the interviewer to look for specific detail rather than the name of a tool, which says what to reject without telling the candidate what to say.
Score alone before anyone talks
This is the step with the most evidence behind it and the least adoption.
Highhouse and Brooks cite assessment-center research on a large managerial sample. The consensus overall rating, the number the assessors agreed on after discussion, added no validity over cognitive ability and personality measures. Zero. A simple average of the dimension scores assessors had given independently added .09. Optimally weighted independent scores added .12.
The discussion did not just fail to help. The number it produced was worse than the average of what people thought before the discussion started.
The mechanism is familiar to anyone who has sat on a panel. The most senior or most confident person speaks first. Everyone else's read quietly reshapes itself around that one. What comes out looks like agreement and is one loud opinion with three echoes.
The fix is almost free. Everyone writes their scores before anyone says a word. Then the scores go on the table, the gaps get discussed, and the panel records a final number with a written reason. The discussion still happens. It just happens against four independent readings instead of one.
Keep the original scores. The originals plus the final are the audit log, and they are also how a team learns which panelists drift and which anchors are unclear.
Combine by a rule, not by feel
The last evaluation component is the one that sounds most clinical and is the least controversial in the research. Once each answer has a score, the scores should combine into a candidate result the same way every time.
Kuncel, Klieger, Connelly, and Ones (2013) showed that combining the same information by a mechanical rule outperforms combining it by clinical judgment. That does not mean a machine should decide. It means the interviewer's judgment belongs in the scores, not in a second, private reweighting of the scores after the fact. If the rule says the role-defining skill carries the most weight, it carries the most weight for every candidate, not just the ones the panel wanted to advance.
A fixed rule also makes the decision explainable. A hiring manager can defend a candidate who scored 4 on both must-have skills and 2 on a preferred one, under a rule that weights must-haves higher. Nobody can defend feeling stronger about her.
What does interviewer calibration look like in practice?
Calibration is a short, repeated exercise in which interviewers score the same real answers alone and then argue out where they disagreed. Anchors, independent scoring, and a fixed rule are the standard; calibration is what makes a group apply the standard the same way. It is a working session, not a training deck, and it takes about 30 minutes.
The research is consistent about what calibration is not: a lecture on validity coefficients. Lievens and De Paepe (2004) recommend vivid, concrete cases over statistics. Nolan, Dalal, and Carter (2020) found that the framing of the evidence matters more than the evidence itself. When the research was framed as what a team loses by interviewing unstructured, the share of practitioners preferring the structured interview rose from about 36% to 60%. Same facts, loss framing. That study measured stated preference rather than later behavior, so treat loss framing as a rollout tactic rather than a forecast of what the team will do.
The calibration exercise that works is short and specific:
- Take three real answers to one question, ideally from past interviews with names removed.
- Every interviewer scores all three alone, in writing, against the anchors.
- Put the spread on the screen.
- Argue out the two answers with the widest gap.
Most groups discover that they disagree with the person next to them by a full point on at least one answer. That moment does more than any slide. It makes the noise visible, and it ties the argument to the anchor wording rather than to personalities.
Run it once at rollout and once a quarter. Roulin and colleagues found that formal training correlated with standardized evaluation at r = .39, the strongest of any relationship in their study, and stronger than its link to question consistency or note taking. Training is how the scoring habit gets built, and it decays.
One more finding is worth carrying into the room. Years of interviewing experience were negatively related to question sophistication in that same study (r = -.30). The veteran on the panel is not exempt from calibration. Experience hardens whatever habits an interviewer started with, and if nobody ever asked them to write down what a 3 means, the habit they hardened is improvisation.
What about panels?
A panel is one interview scored by several people, so everything above applies with more force. A panel multiplies the sources of noise and adds one the solo interview does not have: people watching each other. The practices that hold it together are independent scores first, revealed before any reasons, and one written rating per skill at the end.
The process used by industrial-organizational psychologists who run panels for a living has four steps. Rate alone first, on every skill, in the moment. Reveal scores, not reasons; see where the panel agrees and disagrees before anyone explains. Then the reasons, with scores allowed to move if the discussion earns it. Then record one rating per skill, in writing, with a brief explanation linking the evidence to the criteria.
Two design questions have to be answered before the first panel runs, and most teams answer them by accident.
What counts as consensus? Full agreement, or within a point? The honest answer is that a panel can decide a one-point gap does not matter, settle on a final score, and move on, as long as the gap itself is written down. What it cannot do is pretend the gap never existed.
Do panelists carry equal weight? In the math, yes. You invite people to a panel because their judgment counts. Seniority is handled in the debrief, where the senior person or a neutral moderator leads the discussion, and in the design of the process, where different stages carry different people. It is not handled by weighting one person's score higher.
The pull toward the hiring manager's number is the panel's biggest known failure. Everyone drifts toward the most senior score once it is visible. Independent scoring first is the only reliable guard, and it costs nothing.
The open problem: what do you write down when there is no evidence?
Every scoring rubric handles a good answer and a bad one. Almost none handle the case where the candidate simply did not give the interviewer anything to judge. They misread the question. The interviewer ran out of time. The candidate talked for two minutes and never landed.
The default is to split the difference and give a middle score. That is the one thing not to do, because a middle score reads later as a real observation, and it silently pulls the candidate's result toward the mean on a question that produced no information.
The better default is to record no score and a note. That treats the question as unanswered rather than answered badly, and it tells the debrief what it needs to know: this skill still needs evidence. It also protects the candidate from being penalized for the interviewer's clock.
This is a live problem in the field, not a solved one. Interviewer guidance routinely tells the interviewer to follow the scoring guidance for insufficient evidence, and most teams discover at that moment that they have none. Write the rule before the first interview, because the interviewers will invent one otherwise, and each will invent a different one.
Where to start on Monday
Five changes cover most of the scoring gap, and a team can put all five in place before its next hiring loop without buying anything. They work in order: anchors give interviewers something to score against, independent scoring keeps the readings separate, and the fixed rule stops the debrief from quietly rewriting them.
- Write two anchors per question: what an empty answer sounds like, and what a strong one sounds like.
- Score every answer on its own, in writing, before the debrief.
- Open the debrief by putting the independent scores side by side. Discuss the gaps. Record a final score per skill with one sentence of reason.
- Combine scores by the same rule for every candidate.
- Run a 30-minute calibration on three real answers before the first live interview, and again in three months.
None of this removes the interviewer's judgment. It moves the judgment into the scores, where it can be compared, and out of the debrief, where it gets overwritten.
Edition 1 of this series makes the evidence case for structure overall, and edition 2 covers the scoring framework these anchors sit inside.
What this paper does not claim
Structured scoring reduces noise. It does not remove bias, and no method does. The validity research reports meaningful variability across settings: Sackett et al. (2022) give structured interviews a mean validity of 0.42 with an 80% credibility interval of 0.18 to 0.66, and how a team implements the scoring half of structure is a large part of where it lands in that range.
The practices here are the ones with the most consistent evidence behind them. They are not a guarantee, and this paper is not legal advice.
About Ratio
Ratio builds the scoring standard so teams do not have to, working from whatever you have on the job: a job description, intake notes, a hiring manager's meeting transcript. It produces the skills that matter for the role, one question per skill, a planned follow-up for each part of the answer, and a written scoring guide for every question. Interviewers score alone on the same rubric, and panels reveal scores together before recording a final one. Ratio never auto-scores candidates.
References
- Campion, M. A., Palmer, D. K., & Campion, J. E. (1997). A review of structure in the selection interview. Personnel Psychology, 50(3), 655–702. https://doi.org/10.1111/j.1744-6570.1997.tb00709.x
- Chapman, D. S., & Zweig, D. I. (2005). Developing a nomological network for interview structure: Antecedents and consequences of the structured selection interview. Personnel Psychology, 58(3), 673–702. https://doi.org/10.1111/j.1744-6570.2005.00516.x
- Highhouse, S., & Brooks, M. E. (2023). Improving workplace judgments by reducing noise: Lessons learned from a century of selection research. Annual Review of Organizational Psychology and Organizational Behavior, 10, 519–533. https://doi.org/10.1146/annurev-orgpsych-120920-050708
- Kuncel, N. R., Klieger, D. M., Connelly, B. S., & Ones, D. S. (2013). Mechanical versus clinical data combination in selection and admissions decisions: A meta-analysis. Journal of Applied Psychology, 98(6), 1060–1072.
- Lievens, F., & De Paepe, A. (2004). An empirical investigation of interviewer-related factors that discourage the use of high structure interviews. Journal of Organizational Behavior, 25(1), 29–46. https://doi.org/10.1002/job.246
- Nolan, K. P., Dalal, D. K., & Carter, N. (2020). Threat of technological unemployment, use intentions, and the promotion of structured interviews in personnel selection. Personnel Assessment and Decisions, 6(2), 38–53. https://doi.org/10.25035/pad.2020.02.006
- Roulin, N., Bourdage, J. S., & Wingate, T. G. (2019). Who is conducting "better" employment interviews? Antecedents of structured interview components use. Personnel Assessment and Decisions, 5(1), 37–48. https://doi.org/10.25035/pad.2019.01.002
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068. https://doi.org/10.1037/apl0000994
Origin: 2_wiki/Judgment-Noise.md; 2_wiki/Interview-Structure-Components.md; 2_wiki/Adoption-Challenges.md; 2_wiki/Panel-Interviews.md; 2_wiki/Scoring-System.md (public-safe parts only); 2_wiki/Selection-Method-Stage-Placement.md (Kuncel); 2_wiki/Sackett-Stat-Compliance.md; 6_product/prompts/analyses/Campion-1997-Structured-Interview-Review-LLM-Brief.md; 8_gtm/11_marketing/white_papers/White-Paper-Series-Topics-and-First-Edition-2026-07-18.md (topic 13); content library CL-005, CL-006, CL-029. Revised 2026-09-29 against Ratio_Writing_Style_Guide.md (September 29, 2026) and the /ratio-whitepaper validator; references 1, 2, 3, 5, 6, 7 and 8 reconciled against the verified entries in WP01. Reference 4 (Kuncel et al., 2013) still needs its DOI and page range checked against the PDF in 9_archives/research_papers/.
Frequently asked questions
Why do interviewers disagree about the same candidate?
Most disagreement is noise rather than insight: unwanted variation in judgment that has nothing to do with the candidate (Highhouse & Brooks, 2023). It takes two forms. Disagreement noise is two interviewers reading the same evidence differently. Occasion noise is one interviewer reading the same evidence differently at 9 a.m. than at 4 p.m. Written anchors and independent scoring before any discussion are the two practices with the most evidence behind reducing it.
How do you score interview answers consistently?
Score each answer on its own rather than giving one rating at the end, against anchors written before the interview, in writing, before anyone discusses the candidate. Then combine the scores by the same fixed rule for every candidate. Campion, Palmer, and Campion (1997) list these as the evaluation components that carry the most weight, and adding interviewer training completes the standard.
What is an anchored rating scale in an interview?
An anchored rating scale pairs each score point with a plain description of what an answer at that level sounds like, written before the interview. The useful shape is two ends and a floor: what an empty answer sounds like, what a strong one sounds like, and what reciting rather than recalling sounds like. Anchors that describe the shape of a strong answer are harder for a candidate to game than anchors that list every acceptable answer.
Should interviewers score before or after the debrief?
Before, always, and in writing. In the assessment-center research Highhouse and Brooks (2023) cite, the consensus rating assessors agreed on after discussion added no validity over cognitive ability and personality measures, while a simple average of their independent dimension scores added .09 and optimally weighted independent scores added .12. The discussion still has value. It has to happen against independent readings rather than produce them.
What does interviewer calibration training involve?
The exercise with the most support is short and concrete: every interviewer scores the same three real answers alone and in writing against the anchors, then the group argues out the two answers with the widest spread. Roulin, Bourdage, and Wingate (2019) found formal training correlated with standardized evaluation at r = .39, the strongest relationship in their study. Run it at rollout and again each quarter, because the habit decays.
Was this useful?