Most hiring mistakes are not made in the interview. They are made afterward, when three interviewers compare vague impressions and the loudest voice wins. A structured interview scorecard fixes that by making everyone rate the same candidate against the same criteria, on the same scale, with written evidence. It is one of the better-evidenced ways to reduce bias and make hiring decisions you can actually defend later.
This guide gives you a scorecard template you can copy today, the rules that keep it honest, and the part most teams get wrong: how to capture the interview evidence behind each rating without creating a GDPR or EU AI Act problem for yourself.
What is a structured interview scorecard?
A structured interview scorecard is a fixed template that every interviewer fills in for every candidate, rating them against predefined competencies rather than a general gut feel. The competencies are chosen before anyone interviews, the rating scale is the same for everyone, and each rating is backed by a note about what the candidate actually said or did.
A predefined template used by every interviewer to rate a candidate against the same competencies, on the same scale, with written evidence for each rating and a final recommendation. It replaces free-form impressions with consistent, comparable criteria agreed before the interview.
The value is comparability. When two candidates are scored against identical criteria with evidence attached, the debrief becomes a conversation about the evidence, not a contest of confidence. That is also what makes the decision easier to justify if it is ever challenged.
The four parts of a scorecard that works
A usable scorecard is short and specific. Four elements do the work.
- Competencies defined up front. Pick the handful of competencies that genuinely predict success in this role, and write them down before interviewing. Resist the generic twenty-item list; four to six real competencies beat a long checklist nobody reads.
- A clear rating scale. Use a 1 to 4 scale. An even number of points removes the safe middle option and forces a lean. Define what each number means in one line, so a 3 means the same thing to every interviewer.
- A written-evidence field per competency. This is the part that separates a scorecard from a rating. Every score must point to something concrete the candidate said or did. No evidence, no score.
- An overall recommendation with a rationale. A single hire or no-hire call with two or three sentences explaining why, tied back to the competency ratings.
A template you can copy
Adapt the competencies to your role, keep the structure:
| Competency | Rating (1-4) | Evidence from the interview |
|---|---|---|
| Role-specific skill | What the candidate said or demonstrated | |
| Problem-solving | A concrete example from the conversation | |
| Collaboration | A specific behaviour or story | |
| Communication | Observed, with a quoted or paraphrased example |
Scale: 1 = clear concern, 2 = below bar, 3 = meets bar, 4 = strong. Finish with one overall recommendation and a short rationale. When you run the debrief, review evidence before revealing ratings, so scores are not anchored by whoever speaks first. Our guide to running a defensible interview debrief covers that sequence in detail.
The hard part: capturing evidence without breaking GDPR
The evidence field is where scorecards succeed or fail. If interviewers reconstruct quotes from memory an hour later, the evidence is thin and often wrong. The obvious fix is to record and transcribe the interview so each rating can point to what was actually said. That is the right instinct, and it introduces obligations you have to handle deliberately.
Interview recordings are sensitive personal data tied to a named individual. To use them as scorecard evidence lawfully, you need a lawful basis, appropriate consent for the recording, a defined retention period, a way for candidates to exercise their rights, and a clear answer to where the audio physically lives. For the consent side specifically, see our guide on recording job interviews in Germany and the EU.
Residency is the part teams underestimate. Keeping interview audio and transcripts on EU infrastructure removes the cross-border transfer question entirely, which is one fewer thing to defend when a works council or a candidate asks. Where a tool hosts the data is worth confirming in writing before you route real interviews through it; our comparison of EU vs UK data residency for interview notetakers walks through how to check.
Let the tool capture, keep the judgment human
There is a strong temptation to let software score the candidate for you. Resist it. Under the EU AI Act, recruitment is a high-risk use case, and a tool that scores, ranks, or filters candidates automatically pulls you into heavier obligations and real fairness risk. A scorecard filled in by humans, supported by a tool that only records and summarizes, sits in a far lighter-touch position.
That is the line Numi is built on. Numi is an EU-hosted AI meeting assistant used for interviews: it records, transcribes, and summarizes each conversation on European infrastructure and extracts action items and decisions, so every interviewer has an accurate record to ground their rating. It does not score or rank candidates for you, and it does not train on your data. The scorecard, and the decision, stay with your panel, which is exactly where a defensible process keeps them.
Putting it together
Write down four to six real competencies, rate each on a 1 to 4 scale, attach evidence to every score, and finish with one recommendation and a rationale. Capture the evidence from an accurate record rather than memory, keep that record on EU infrastructure with consent and a retention schedule, and let humans do the scoring. That is a process that produces better hires and holds up when someone asks how the decision was made. For a shortlist of EU-hosted tools that capture the record without grading people, see our roundup of the best GDPR-compliant AI notetakers for recruiters.