A rubric with four dimensions scored one to five is not automatically useful. I have seen plenty of them collect dust in shared drives because the dimensions were too vague to apply consistently, or too disconnected from what actually mattered in the role to produce useful signal. The format was right. The content was wrong.
The difference between a rubric that genuinely improves hiring decisions and one that creates an illusion of structure is almost entirely in how the dimensions are defined. This post covers what we have learned about that design problem, both from building Intervieux and from the broader selection research literature.
The Dimension Problem: Claims vs. Behaviors
Most rubrics I see when TA teams first show them to us measure traits rather than behaviors. Dimension labels like "communication skills," "problem-solving ability," and "cultural fit" are common. The problem with trait-based dimensions is that they require the evaluator to infer a trait from a response, which introduces exactly the kind of inconsistency and bias that structured rubrics are supposed to prevent.
Behavior-based dimensions describe what a person does in specific situations. "Identifies and communicates trade-offs when competing priorities conflict" is a behavioral dimension. "Communication skills" is not. The behavioral version gives the evaluator something to look for in the candidate's answer rather than asking them to make a holistic judgment about a personal quality.
This connects to a distinction that experienced interviewers understand intuitively but rarely make explicit: the difference between a claim and an example. "I handle pressure well" is a claim. "When our warehouse system went down the day before a major shipment, I coordinated manually with three teams over six hours while simultaneously pulling backup data from our second system" is an example. A rubric that scores on behavioral dimensions naturally incentivizes examples, because claims have nothing to be evaluated against the behavioral description.
Behavioral Anchors at Each Score Level
A dimension label alone doesn't tell an evaluator how to distinguish a 4 from a 2. Score-level anchors describe what a response at each level looks like, concretely. The more specific the anchor, the more consistent the scoring across different evaluators or across the same evaluator on different days.
For a dimension like "manages competing priorities under time pressure" on an operations coordinator role, anchors might look like this:
Score 5: The candidate describes a specific situation with multiple simultaneous demands. They explain the decision-making process they used to triage. The example shows a deliberate outcome, not just luck or survival.
Score 3: The candidate describes a plausible situation but at a generic level. They say they prioritized but don't explain the logic. The outcome is implied rather than stated.
Score 1: The candidate offers a principle statement ("I make a list and tackle the most important things first") without any specific example, or the example does not address competing demands at all.
With anchors like this, two evaluators reviewing the same response will arrive at similar scores more often than not. Without them, two evaluators reviewing the same response may disagree by two full points, which makes the rubric effectively useless for ranking candidates.
How Many Dimensions and Which Ones
More dimensions are not better. A rubric with eight dimensions sounds thorough, but it tends to produce scores that are highly correlated with each other (because evaluators find it hard to hold eight independent criteria in their heads simultaneously) and exhausting to apply. Four to six dimensions is the range that works in practice.
Selecting the right dimensions requires starting from a different question than "what qualities does a good employee have?" The useful question is: "What specifically determines success in the first 90 days of this role?" The answer to that question is typically narrower and more concrete than a general competency profile.
For an operations coordinator role at a growing logistics company, the 90-day success question might produce dimensions like: prioritization under competing demands, cross-functional communication (specifically: can they translate between ops and non-ops stakeholders), process adherence vs. exception judgment (when to follow the protocol and when to escalate), and attention to detail in data-facing tasks. Those four dimensions will produce a better shortlist than a generic profile that includes "leadership potential" and "growth mindset."
The Dimension-Role Fit Check
A useful test before finalizing any rubric dimension is to ask: if a candidate scored a 5 on this dimension, would their hiring manager notice it in the first 30 days? If the answer is no, the dimension may be measuring something real but not something relevant to this specific role at this specific stage.
We are not saying generic competencies are worthless. Plenty of them predict long-term performance reasonably well. But first-round screening at volume is a filter problem, not a long-range performance prediction problem. You are trying to distinguish the top 5 to 10 percent of a batch from the rest in a way that reduces second-round conversations wasted on candidates who clearly don't fit. Role-specific, stage-relevant dimensions do that job better than a general competency profile.
The Limits of Rubric Design
Even a well-designed rubric has a ceiling. It can't compensate for questions that don't give candidates the opportunity to demonstrate the dimension you're trying to measure. A rubric scoring for "situational judgment under ambiguity" applied to a question like "tell me about your previous job" will produce noise, not signal, because the question doesn't invite the kind of response the dimension requires.
Question design and rubric design are coupled problems. The questions need to be specific enough that a strong candidate for this role has a natural opportunity to demonstrate the behavioral dimensions. Behavioral interview questions ("tell me about a time when...") work well for this. Hypothetical questions ("what would you do if...") work somewhat less well because they measure stated preference rather than actual behavior. Competency-based follow-ups ("what specifically made you choose that approach over the alternatives?") are useful for probing depth after a candidate has given a surface answer.
At Intervieux, we build rubric-question pairing into the setup process specifically because the two need to be coherent. A rubric that doesn't match the questions it's applied to is as problematic as a rubric with vague dimensions. The goal is that a candidate who genuinely fits the role has a clear path to demonstrate that fit, and a candidate who doesn't fit has no obvious way to fake it.
See also: Structured Interviews Beat Gut Feel: What the Research Shows and Building Interview Logic That a Good Recruiter Would Recognize.