A Technical Interview Scorecard Template That Interviewers Actually Fill In
By Pooja Goenka ยท 2026-10-02
Disclosure first: we build an interview platform with scorecards in it. Everything below works in a spreadsheet, and the template is written so you can copy it into one today.
Most interview scorecards fail the same way. There are eight criteria, each scored 1 to 5, and no one has written down what a 3 means. So every interviewer invents their own scale, writes "good communication, solid fundamentals" in the notes box, and gives a 4. The debrief then compares four private scales as if they were one.
A scorecard is only useful if two interviewers who saw the same answer would give it the same score. That needs fewer criteria and much more specific wording than most teams use.
Why structure is worth the effort
Structured interviews, where every candidate gets the same questions and is scored against the same written criteria, predict job performance better than unstructured conversations. That finding has held for decades. Schmidt and Hunter's 1998 meta-analysis put structured interviews well ahead of unstructured ones, and a 2022 re-analysis by Sackett and colleagues found structured interviews were the strongest single predictor in their review.
The scorecard is where the structure lives. Without one, a "structured" interview is just a list of questions.
The template
Four competencies for a mid-level software engineer. Change them for the role, but keep the count low. Four is enough, and six is the most an interviewer will fill in properly.
| Competency | 1: Concern | 2: Below bar | 3: Meets bar | 4: Strong |
|---|---|---|---|---|
| Problem solving | Could not get to a working approach, even with hints | Reached an approach only after the interviewer gave the key idea | Found a working approach, with small hints at most, and explained why it works | Compared two approaches and chose one for a stated reason, such as time, memory or simplicity |
| Code quality | Code would not run, or could not be followed | Ran, but with bugs the candidate did not find when asked to test | Correct, readable, and the candidate tested it with at least one edge case | Also named the cases it would fail on in production and how to guard against them |
| Technical judgement | Accepted a flawed suggestion without question | Spotted a problem only when pointed at it | Questioned a suggestion and explained what was wrong with it | Explained the trade-off and when the flawed version would actually be fine |
| Communication | Interviewer could not follow the reasoning | Explained what they did, not why | Thought out loud, and checked assumptions before coding | Adjusted the explanation when the interviewer looked lost |
Below the table, every scorecard gets three more fields:
- Evidence. For each score, one sentence describing what the candidate said or did. "Wrote a brute-force solution, then reduced it to O(n) using a hash map when asked about scale" is evidence. "Strong problem solver" is not.
- Overall recommendation. Strong no, no, yes, strong yes. No "maybe". A 2.5 average means you have not decided, and the debrief should hear that.
- What would change my mind. One line. It is the most useful field in the debrief, because it tells the panel what to ask the other interviewers.
Why there are only four levels
A five-point scale gives everyone a safe middle. Interviewers who are unsure give a 3, and the 3s say nothing. With four levels you have to choose a side of the bar. The two middle levels are labelled "below bar" and "meets bar" on purpose, because that is the decision you are actually making.
The "technical judgement" row
This is the row most templates do not have, and the one I would least want to drop. A lot of engineering work now starts from code or an answer someone else wrote, often an AI assistant. The useful skill is noticing what is wrong with it.
You can test that directly. Give the candidate a short, confident answer or code snippet with one real problem in it, and ask them to review it. Score what they catch, and score what they wrongly flag too, because a candidate who calls everything a bug is no more useful than one who misses everything.
Three rules that make the scorecard work
Fill it in before you talk to anyone. The scorecard is submitted before the debrief and before reading anyone else's. Once someone hears a colleague say "I thought they were great", their own memory of the interview shifts. We wrote about how to run the debrief itself separately.
Write the anchors before the first candidate. If you calibrate the anchors after seeing candidates, you calibrate them to the candidates you happened to like.
One interviewer, one or two competencies. An interviewer who owns problem solving and code quality goes deep on those. An interviewer asked to score all four in 45 minutes guesses at two of them.
Running it in a spreadsheet
One tab per role with the anchors at the top. One row per candidate per interviewer: competency scores, evidence, recommendation and the "change my mind" line. Hide each interviewer's row from the others until everyone has submitted. Google Sheets protected ranges or a simple form can do that, and a form that writes to a sheet is enough.
Before the debrief, sort by competency and look for the rows where two interviewers are two or more levels apart. Start the discussion there.
Where software helps
The spreadsheet version breaks down in two places: keeping scores hidden until everyone has submitted, and finding disagreements quickly when you have ten candidates in a week. That is what LogicWiz Interviews does for hiring teams. Interviewers score privately against your rubric, nothing is revealed until all scores are in, and the debrief view puts the disagreements side by side. The template above works the same way in a sheet.
Common questions
How many competencies should an interview scorecard have?
Four is a good default for one role, and six is the practical maximum. With more, interviewers stop reading the anchors and score from a general impression.
Should a scorecard use a 1 to 5 scale?
A four-point scale usually works better for hiring decisions, because it removes the neutral middle and forces each interviewer to say which side of the bar the candidate is on.
What should go in the evidence field?
One sentence describing what the candidate said or did, specific enough that another interviewer could picture it. Adjectives like "strong" or "solid" are not evidence.
When should interviewers fill in the scorecard?
Immediately after the interview, and before speaking to anyone else on the panel. Scores written after the debrief starts are shaped by the first opinion spoken out loud.
Can I use this template for non-engineering roles?
Yes. Keep the structure of four competencies, written anchors, an evidence field, a forced recommendation and the "what would change my mind" line, and replace the competencies with ones that matter for the role.