Interviewer gut feel is confident and frequently wrong. It's not that experienced interviewers can't spot good candidates — it's that unstructured impressions are inconsistent between interviewers, and inconsistent with themselves over time.
We looked at outcome data across roughly 1,000 hires we've been involved in placing, comparing each hire's interview scores against their performance six and twelve months in, wherever a client was willing to share that assessment. The gap between structured and unstructured evaluation was larger than we expected going in.
What structured scoring actually changes
A scorecard forces every interviewer to rate the same defined competencies against the same scale, before comparing notes. That single change — scoring before discussing — is what most improves accuracy. Discussion before scoring lets the most confident voice in the room anchor everyone else's rating.
The mechanism is well documented outside hiring too: groups asked to state an independent judgment before discussing converge on more accurate answers than groups that discuss first and then vote. Interview panels are a textbook case of the failure mode — the first person to speak, especially if they speak with conviction, measurably shifts what everyone else reports, even when they were leaning the other way five minutes earlier.
Where unstructured impressions go wrong
The specific failure isn't that interviewers are bad judges of people — most are perfectly good at forming an impression. It's that the impression forms too early and too globally. A strong opening answer creates a halo that colors how every subsequent answer gets interpreted, and a single awkward moment can do the same in reverse, regardless of whether either one actually reflects the competency being assessed.
Unstructured interviews also make it much harder to compare candidates against each other fairly. “I really liked candidate A” and “I really liked candidate B” aren't comparable data points — they're two different people's global impressions, formed under different circumstances, on different days, possibly assessing different things without either interviewer realizing it.
What the outcome data actually showed
We split the 1,000 hires into those assessed with a structured, pre-agreed scorecard and those assessed primarily on interviewer discussion and consensus, then compared each group against later performance signal — retention, manager-reported performance where available, and whether the hire was promoted or exited within 18 months.
It's not about removing judgement
The scorecard doesn't replace judgement, it structures where judgement gets applied — to specific, pre-agreed competencies, scored independently, rather than to a vague overall impression formed in the first five minutes.
This is the most common misconception clients raise when we introduce scorecards: that structure means less trust in the interviewer's expertise. It's the opposite. A good scorecard is built by the people who best understand the role, deciding in advance what actually matters — and then trusting each interviewer's independent judgment on those specific things, instead of one dominant voice's overall gut read carrying the whole panel.
Where scorecards get adopted badly
The most common way scorecard adoption fails isn't resistance, it's ceremony without discipline — teams introduce a scorecard template, interviewers fill it in, and then the debrief still opens with “so, what did everyone think?” before anyone's independent score has actually been reviewed. The form exists, but the mechanism that makes it work — scoring before discussing — never actually gets enforced.
The second common failure is competencies that are too vague to score consistently. “Communication” or “culture fit” sound like real dimensions but mean something different to every interviewer who rates them, which reintroduces exactly the inconsistency a scorecard is supposed to remove. The fix isn't dropping those dimensions, it's defining them concretely enough that two interviewers watching the same answer would plausibly give it the same score.
How we run our own panels
Every search we run gets a scorecard built before the first interview is scheduled, agreed with the hiring manager against the specific scope of the role rather than pulled from a generic template. Interviewers submit scores independently through the client portal before any debrief conversation happens, and the debrief itself starts from the scores already on record, not from an open “so, thoughts?” that invites the first strong opinion to set the tone.
It's a small process change, and it took longer to get client hiring managers comfortable with than we expected — the instinct to talk it through first is strong, even among people who agree with the reasoning. What's made it stick is showing clients their own data after a handful of hires: the searches that scored independently before discussing consistently produced hires whose six-month performance matched the interview scores more closely than the searches where the debrief happened first.
Building a scorecard that actually works
A scorecard is only as good as the competencies it's built around, and most of the value comes from getting that design right before the first interview, not from the scoring mechanics themselves.
Define three to five competencies, no more
A scorecard with ten dimensions produces noise, not signal — interviewers rush through it and revert to a gut overall score anyway. Three to five forces real prioritization of what matters.
Score independently before any group discussion
Every interviewer submits their scores before the debrief starts. This single rule change does more for accuracy than any other part of the process.
Tie each competency to a specific question or exercise
A competency with no dedicated question attached to it in the loop won't get assessed consistently, whatever the scorecard says on paper.
Get hiring notes like this by email
One note a fortnight, real numbers from live searches — no filler, unsubscribe anytime.
On this page
Get these by email
One hiring note a fortnight. Real numbers, no filler.
Keep reading
All articlesWhy your best candidates drop out at week three
Sep 8, 2026
When to run a retained search instead of posting the role
Aug 19, 2026
What Series B engineering salaries actually look like in 2026
Sep 2, 2026
