The early version tried to cover too much.
Ratio began as a single-round interview product. We pulled a long list of skills from the job and turned them into questions. Then we tried to fit the whole thing into one conversation. User feedback and our own testing showed the cost quickly: rushed interviews, questions left unscored, and no room to follow up when an answer needed another pass.
Cutting the scope changed the product. The skill model became tighter, must-haves became scarce, and the interview split across rounds. Fewer planned questions left room for structured follow-ups and for the rest of a real interview: rapport, candidate questions, role context, and a candid discussion about fit.
Then another problem surfaced. A good question doesn't tell an interviewer how to judge the answer.
That is the problem RATE was built to solve.
Why do better interview questions still produce weak decisions?
A well-written question improves what the candidate is asked. The decision can still drift when the interviewer has no shared way to read the answer. One person rewards detail, another rewards confidence, and a third trusts familiarity. Structure has to govern information gathering and evaluation, not the question list alone.
Edition one established the larger evidence case. Structured interviews predict job performance better than unstructured interviews (Sackett et al., 2022). That result describes structured interviews as a method without validating a particular commercial framework or guaranteeing a result for one employer.
Campion, Palmer, and Campion (1997) mapped 15 ways an interview can become more structured. Asking consistent questions is one component. Taking notes, using anchored scales, rating each answer, and applying the same evaluation process also matter. A team can standardize the script and still lose the method at the moment of judgment.
That failure is easy to miss because the questions look professional and the scorecard exists, while interviewers still translate the candidate's words through private standards they can't explain to one another.
| Where structure can leak | What happens |
|---|---|
| Too many questions | The round gets rushed and answers go unscored |
| Dense scoring guides | Interviewers ignore the guide or fall back on instinct |
| Different criteria by question | Raters have to learn a new standard every few minutes |
| Improvised follow-ups | Candidates receive different chances to clarify the same gap |
| One overall score at the end | Early impressions can color everything that follows |
RATE addresses the middle of that table: how to listen, follow up, and score while the interview is happening.
What does the RATE framework measure?
RATE is a four-part reading method for interview answers. Relevance asks whether the answer is grounded. Approach looks for the candidate's method and ownership. The Trade-offs dimension surfaces the constraint or hard choice. Evidence asks what happened, how the person checked it, or what they learned. Together, the four dimensions turn an impression into evidence another interviewer can inspect.
The model has two layers:
| Layer | RATE dimensions | What the layer establishes |
|---|---|---|
| Foundation | Relevance + Approach | The answer addresses the question with specific detail and shows what the candidate did or would do |
| Depth | Trade-offs + Evidence | The answer exposes judgment and gives a result, check, consequence, or learning |
R: Relevance
Relevance grounds the answer in the question. In a behavioral interview, that means a real episode rather than a general account of how the candidate tends to work. In a technical or situational question, the candidate should engage with the actual facts, inputs, and constraints. Broad principles alone leave the answer ungrounded.
Specificity alone isn't enough. A tool name, date, or company can decorate a vague answer. The detail has to connect to a distinct decision or piece of work.
A: Approach
Approach shows the method and the candidate's part in it. Interviewers should be able to identify what this person chose, checked, changed, or owned. Team language is fine when the individual's contribution remains clear.
This dimension catches a common gap between knowing the vocabulary and doing the work. A list of sensible steps may sound polished while leaving sequence, ownership, and judgment unresolved.
T: Trade-offs
The Trade-offs dimension reveals what made the work hard. The signal may be a deadline, budget limit, competing goal, threshold, cost, or choice that closed off another option. Naming a generic tension such as speed versus quality carries little weight until the candidate explains what moved, what stayed protected, and why.
The quality of the choice comes later. The first scoring question is whether the answer exposes a decision that another trained interviewer can inspect.
E: Evidence
Evidence covers the result and more. It can be an outcome, a test, a consequence, a lesson applied later, or a signal that the plan failed. In a hypothetical scenario, it is the check that would tell the candidate whether the approach worked.
An answer can earn Evidence without a dramatic number. The useful distinction is between a claim that things improved and an observable sign that supports it.
Why do Trade-offs and Evidence reveal more depth?
Relevance and Approach show that a candidate can engage with the work. Trade-offs and Evidence show how far that understanding goes. They force the answer past smooth procedure and into consequence. The interviewer can inspect what the person gave up, which risk they accepted, how they checked the result, and what they would carry into the next decision.
I saw the difference while screening a candidate whose resume appeared strong. The initial screen would have been enough for me to submit the person in my agency years. Tougher assessment questions exposed that the candidate couldn't demonstrate the depth of experience and knowledge the resume seemed to promise. The reason for that gap was never established, but the gap itself became visible.
Trade-offs is useful because experienced work rarely offers every benefit at once. A senior product manager choosing between a retention fix and an enterprise feature has to commit scarce engineering capacity. A finance leader facing quarter-end pressure has to protect the accounting treatment while accepting a business consequence. A recruiter deciding whether to advance a thin candidate has to weigh speed against the cost of a weak submission.
The Dreyfus skill-acquisition model describes a broad movement from reliance on abstract rules toward judgment grounded in concrete experience (Dreyfus & Dreyfus, 1980). RATE borrows that lens cautiously. It asks for reasoning that can be challenged with another question, rather than treating any mention of a trade-off as proof of expertise.
Evidence completes the account. A candidate who names a difficult choice but can't say what followed may be describing a plan rather than work they understood. A candidate who names the result but can't connect it to an action may be claiming the team's outcome. RATE keeps those gaps separate so one polished strength doesn't hide another missing part.
How should interviewers score and follow up with RATE?
Use the same short listening prompts for every answer of the same question type. Mark each RATE dimension as clear, unclear, or absent. Follow up only where the answer leaves a real gap. This keeps the scorecard usable during a live conversation while giving candidates a consistent chance to supply evidence they didn't offer at first.
Earlier RATE anchors carried much more detail. During testing, interviewers either fell back on gut feel or used only the short lead-in before the added explanation. We followed their behavior and cut the extra text. The lead-ins became fixed by question type, while deeper clarification moved into optional follow-ups.
The questions should vary naturally. The scoring language should repeat.
| RATE dimension | Behavioral question | Technical or situational question |
|---|---|---|
| Relevance | Provide a real example | Give specific detail |
| Approach | Explain what they did | Explain their approach |
| Trade-offs | Discuss the tradeoff | Discuss the tradeoff |
| Evidence | Share the outcome | Describe what good looks like |
Read each cell as a yes-or-no check. Did they provide a real example? Did they explain what they did? Did they discuss the tradeoff? Did they share the outcome? When the interviewer can't answer, the matching follow-up should pull for that dimension without feeding the candidate an answer.
Campion et al. (1997) placed follow-up questioning on a structure continuum. Prohibiting all follow-ups was the highest-control option, while limited or pre-planned follow-ups were the next level. No guidance was the lowest level because interviewers may follow up differently. The authors warned that prompting can clarify missing information or contaminate the interview by coaching an answer. They described the net validity effect as unclear. Their practical suggestion was narrower: standardized prompts tied to job requirements may reduce missing information without changing what the question measures.
Here is a sample RATE worksheet. The answer is fictional and exists only to show the method.
Sample situational question: Engineering has room for one more feature this quarter. Sales wants an enterprise add-on, while usage data points to an onboarding fix that may retain more users. Which one do you support?
Sample answer summary: The candidate chooses the onboarding fix, reviews where users leave the product, gives Sales a date for reconsidering the add-on, and says success means fewer first-week exits. The answer never addresses the revenue delayed by that choice.
| Dimension | Initial read | Reason | Prepared follow-up if needed |
|---|---|---|---|
| Relevance | Clear | Uses the capacity limit and usage data | None |
| Approach | Clear | Commits to a sequence and names the review step | None |
| Trade-offs | Unclear | Chooses retention but leaves the revenue cost unstated | What revenue are you willing to delay for that fix? |
| Evidence | Clear | Names first-week exits as the success check | None |
This scorecard makes the reasoning visible without prescribing the business answer. Once the follow-up is answered, the team can map the RATE pattern to its own rating scale and job standard.
What does RATE's research lineage actually support?
RATE is Ratio's synthesis of ideas from structured interviewing, behavioral description, problem solving, cognitive demand, skill development, and behaviorally anchored scales. Those sources explain why its dimensions and anchors are plausible design choices. They do not amount to an empirical validation study of RATE, its four-part combination, or its ability to predict job performance.
| RATE design choice | Research tradition it draws from | Boundary |
|---|---|---|
| Ground the answer in a real case | Patterned behavior-description interviewing (Janz, 1982) | One early study does not validate RATE's Relevance dimension |
| Listen for method, choices, and checks | IDEAL problem solving (Bransford & Stein, 1984) | A general problem-solving model is not an interview-scoring validation |
| Scale question difficulty with work depth | Depth of Knowledge (Webb, 1997) and skill acquisition (Dreyfus & Dreyfus, 1980) | Both originated outside employment-interview scoring |
| Use concrete, shared rating language | Behaviorally anchored rating scales (Smith & Kendall, 1963) | BARS supports behavioral anchoring as an approach, not these specific RATE anchors |
| Standardize questions, scoring, and follow-ups | Structured-interview review (Campion et al., 1997) | The review supports components of structure, not RATE as a branded system |
The claim discipline is simple: RATE is research-informed and operationally specified. It has not yet been shown, in a defined outcome study, to predict job performance or make different raters agree. Those are empirical questions. They require a study with a stated sample, jobs, raters, outcomes, and comparison condition.
There is another limit. RATE scores the evidence contained in one answer. It cannot determine a person's overall worth, diagnose honesty, or replace the job analysis that determines which skills deserve questions in the first place. A clean rubric attached to the wrong question still produces a clean score for the wrong thing.
Why should the interviewer still make the judgment?
Human scoring keeps the evaluator accountable for the decision while software handles preparation, consistency, and recordkeeping. That boundary keeps an algorithm from silently turning interview language, voice, or behavior into a candidate judgment the interviewer cannot inspect. Human scoring still cannot remove human bias or guarantee fair outcomes.
Candidate reactions have been tested directly. One experiment involved 123 participants in highly automated or videoconference interviews under high or low stakes. In the high-stakes condition, automation produced lower perceived control and reduced acceptance through lower social presence and fairness (Langer, König, & Papathanasiou, 2019). This was one experimental setting, not a universal measure of candidate sentiment.
Automation may also change the answers it is meant to assess. In a separate mock interview with 124 participants, mainly German students, the automatic-evaluation group reported less deceptive impression management, felt they had fewer opportunities to perform, and gave shorter answers (Langer, König, & Hemsing, 2020). Together, the studies show that changing the evaluator can change the candidate's experience and response behavior.
Regulators also treat automated employment decisions as a distinct risk area. As of September 2026, New York City's Local Law 144 requires covered automated employment decision tools to have had a bias audit within the past year, to make information about that audit publicly available, and to give candidates notice before use; the city began enforcing the law on July 5, 2023 (New York City Department of Consumer and Worker Protection, 2026). In the EU, the AI Act classifies certain employment systems as high risk and requires human oversight (European Parliament and Council of the European Union, 2024). The AI Omnibus amending the Act entered into force on July 27, 2026 and moved the application date for Annex III high-risk systems, the category that covers employment, to December 2, 2027 (European Commission, 2026).
These rules are specific, evolving, and dependent on how a system is used. Keeping the final rating human is a product boundary, not a blanket legal exemption.
How can a team start using RATE?
Start with one role and a short interview round. Give every interviewer the same job-linked questions, the same four RATE checks, and the same neutral follow-ups. Score each answer before discussing the candidate. The first goal is a traceable decision: another person should be able to see what evidence supported the rating.
- Choose the few skills the round can assess well. Leave time for follow-ups, candidate questions, and human conversation.
- Write one clear question per skill. Make sure each question gives the candidate a fair chance to show all four RATE dimensions.
- Use the fixed RATE checks. Relevance, Approach, Trade-offs, and Evidence should mean the same thing from one candidate to the next.
- Prepare neutral follow-ups. Pull for the missing dimension without suggesting the preferred answer.
- Score independently before discussion. Record the evidence behind the rating, then compare judgments with the panel.
Edition 3 of this series covers the calibration exercise that gets a panel applying these checks the same way. The paper after it will test consistency directly, once the benchmark's remaining parity check is complete.
About Ratio
Ratio extracts and weights the skills a job needs for success from whatever you have on it (a job description, intake notes, a meeting transcript), then builds the complete interview kit: the questions, the follow-ups, and the scoring guides. AI prepares the materials; a person runs the interview, judges the answers, and makes the call. Ratio never auto-scores candidates.
References
- Bransford, J. D., & Stein, B. S. (1984). The IDEAL problem solver: A guide for improving thinking, learning, and creativity. W. H. Freeman.
- Campion, M. A., Palmer, D. K., & Campion, J. E. (1997). A review of structure in the selection interview. Personnel Psychology, 50(3), 655–702. https://doi.org/10.1111/j.1744-6570.1997.tb00709.x
- Dreyfus, S. E., & Dreyfus, H. L. (1980). A five-stage model of the mental activities involved in directed skill acquisition (ORC 80-2). University of California, Berkeley, Operations Research Center. https://doi.org/10.21236/ADA084551
- European Commission. (2026). Regulatory framework for artificial intelligence. Retrieved September 29, 2026, from https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai
- European Parliament and Council of the European Union. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence. Official Journal of the European Union. https://eur-lex.europa.eu/eli/reg/2024/1689/oj
- Janz, T. (1982). Initial comparisons of patterned behavior description interviews versus unstructured interviews. Journal of Applied Psychology, 67(5), 577–580. https://doi.org/10.1037/0021-9010.67.5.577
- Langer, M., König, C. J., & Hemsing, V. (2020). Is anybody listening? The impact of automatically evaluated job interviews on impression management and applicant reactions. Journal of Managerial Psychology, 35(4), 271–284. https://doi.org/10.1108/JMP-03-2019-0156
- Langer, M., König, C. J., & Papathanasiou, M. (2019). Highly automated job interviews: Acceptance under the influence of stakes. International Journal of Selection and Assessment, 27(3), 217–234. https://doi.org/10.1111/ijsa.12246
- New York City Department of Consumer and Worker Protection. (2026). Automated employment decision tools. Retrieved September 29, 2026, from https://www.nyc.gov/site/dca/about/automated-employment-decision-tools.page
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068. https://doi.org/10.1037/apl0000994
- Smith, P. C., & Kendall, L. M. (1963). Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales. Journal of Applied Psychology, 47(2), 149–155. https://doi.org/10.1037/h0047060
- Webb, N. L. (1997). Criteria for alignment of expectations and assessments in mathematics and science education (Research Monograph No. 6). National Institute for Science Education, University of Wisconsin-Madison, and Council of Chief State School Officers. https://wcer.wisc.edu/publications/other-publications
Frequently asked questions
What does RATE stand for in interviewing?
RATE stands for Relevance, Approach, Trade-offs, and Evidence. Relevance grounds the answer in the question. Approach shows the candidate's method and ownership. The Trade-offs dimension exposes a constraint or hard choice. Evidence supplies an outcome, check, consequence, or learning.
How do you score interview answers with RATE?
Check whether each RATE dimension is clear, unclear, or absent. Relevance and Approach establish the foundation of the answer; Trade-offs and Evidence add depth. Use a prepared follow-up when a dimension is unclear, then map the completed evidence pattern to the job's rating scale. Score each answer before discussing the candidate with anyone else.
Is RATE empirically validated?
No. RATE draws from established work on structured interviews, behavioral interviewing, problem solving, skill development, cognitive demand, and behaviorally anchored scales, but it has not been independently validated. No independent study has yet established the predictive validity or interrater reliability of RATE as a combined framework. It is research-informed and operationally specified, which is a different claim.
Can a structured interview include follow-up questions?
Yes. A structured interview can include limited or pre-planned follow-ups, one of the options Campion et al. (1997) described on their structure continuum. The key is consistency: follow-ups should clarify the same dimensions, stay tied to the job, and avoid coaching candidates toward a preferred answer. The review found the net validity effect of limiting follow-ups unclear.
Should a machine score the candidate's answers?
Ratio's position is no, and the product boundary reflects it. Human scoring keeps the evaluator accountable for the decision while software handles preparation, consistency, and recordkeeping. Regulators treat automated employment decisions as a distinct risk area: NYC Local Law 144 requires a bias audit and candidate notice for covered automated employment decision tools, and the EU AI Act classifies certain employment systems as high risk and requires human oversight, with those obligations applying from December 2, 2027 (as of September 2026).
Was this useful?