One of the easiest mistakes when building tools for hiring is deciding what "good" looks like before you understand what a good recruiter actually does. Plenty of products that touch the hiring workflow end up doing something no experienced recruiter would recognize as sound: they score on vocabulary density, reward response length, or flag candidates based on proxy signals that have nothing to do with whether the person can do the job.
They look rigorous. They are not doing the right thing.
Building Intervieux's scoring logic required us to be specific about what a skilled recruiter actually does during a first-round screen, and then ask whether we were doing that or something else. This is a post about that inquiry.
What a Skilled Recruiter Is Actually Doing
Before writing a line of scoring logic, we spent time breaking down the specific cognitive acts that happen during a good first-round screen. Not "assessing fit" as an abstract goal. The specific things.
First, an experienced recruiter is checking whether the candidate understood the question. A candidate who answers a different question than the one asked is revealing something. Not necessarily that they are a bad candidate, but that they missed the specificity of what was asked, which is relevant information about their listening and communication style. A naive scoring system might give a long, fluent, well-organized answer a high score even if that answer addressed a different topic than the one the question raised.
Second, a good recruiter is distinguishing between a claim and an example. "I am very organized" is a claim. "When our fulfillment workflow changed mid-quarter, I rebuilt the tracking sheet so that both the warehouse and the account team were using the same status codes, which cut our daily reconciliation time from 40 minutes to about 10" is an example. The difference matters enormously for first-round screening. Claims are easy. Examples are verifiable and specific. A strong recruiter will probe the former and credit the latter. The scoring logic needs to make the same distinction.
Third, a good recruiter applies role-specific criteria, not general "quality of answer" heuristics. A strong response for an operations coordinator role sounds different from a strong response for a customer success manager role. Both might be well-organized and fluent, but the behavioral content that matters is different. Generic scoring optimizes for presentation quality. Role-specific scoring optimizes for fit with the actual requirements of the job.
What We Explicitly Do Not Score On
Being specific about what we ignore is as important as being specific about what we measure. Several signals that correlate with perceived quality have weak or no connection to actual job performance.
Vocabulary level is one. Complex sentence structure and elevated word choice can make an answer feel more sophisticated without the underlying thinking being better. We do not reward vocabulary. We look at whether the answer addresses the question, whether it includes specific behavioral evidence, and whether it matches the criteria defined in the rubric for this role.
Response length is another. Some of the best-scoring responses we have processed are concise. A candidate who answers clearly and stops has often understood the question better than a candidate who takes three minutes to say the same thing twice before wrapping up. Length is neutral in our scoring unless the question requires a minimum level of specificity to be answerable at all.
Tone and warmth are the third category. This is the uncomfortable one to name, but it matters: an enthusiastic, conversationally warm response reliably sounds better to human evaluators than a flat, precise one, even when the informational content is equivalent. We don't score on tone. We score on what was actually said.
The Rubric as the Architecture
The entire scoring architecture in Intervieux is organized around one principle: the rubric the recruiter defines is the source of truth for this batch, for this role. The scoring logic does not layer on top of that rubric with general "good interview" heuristics. It maps each response against the specific criteria and behavioral anchors that the recruiter wrote before the batch began.
This has a consequence that we are explicit about: the system's output is only as good as the rubric. A vague rubric with poorly written score-level descriptions will produce consistent scores that are not useful. A well-designed rubric with specific behavioral anchors at each level will produce scores that an experienced recruiter reviews and agrees with. We can't compensate for bad rubric design with smarter scoring logic. The two are coupled.
This is why we put substantial effort into the rubric-building step of setup. A lot of tools in this space treat the "input your job requirements" step as a formality. We treat it as the most important step. If the criteria are wrong, everything downstream is wrong.
The Test We Use Internally
The internal test that guides every change we make to the scoring logic is simple: if an experienced recruiter reviewed our score alongside the candidate's response, would they agree with the decision we made?
Not always. Sometimes a recruiter would weigh things differently than we do. But "agree more often than not, disagree for reasons that are explicable rather than arbitrary" is the bar we set for ourselves. When we find a case where a recruiter consistently disagrees with a scoring decision on a class of responses, that's a signal that our logic is doing something different from what an expert would do. That's the category of failure we care most about catching.
We are also explicit about what the system is not claiming to do. It is not claiming to make the hiring decision. It is claiming to make the first-filter decision: which candidates in this batch of 200 or 400 have demonstrated sufficient competency on the rubric dimensions to warrant a recruiter's time in the next stage. That is a narrower claim. It is also the specific claim that matters for high-volume teams who don't have 50 recruiter-hours to spend on first-round screening for every role.
The Human Judgment Layer
Automated structured screening does not replace recruiter judgment. It is the first layer of a process that still requires a human decision at every subsequent stage. The shortlist Intervieux produces is an input to the recruiter's judgment, not a replacement for it. The recruiter reviews the scored shortlist, reads the evidence excerpts behind each score, and decides who to call. If a score doesn't look right given what they know about the role, they override it. That's the right architecture.
What the automated layer provides is consistency across the batch. Every candidate in a pool of 400 gets the same questions in the same order. The scoring criteria are applied identically to candidate 1 and candidate 400. The recruiter's judgment is still present, in the rubric design, in the review of the shortlist, and in every subsequent stage. It is not present in the evaluation of 400 individual responses, which is where the consistency problem lives and where automated structured scoring actually helps.
See also: Scoring Rubrics That Actually Predict Performance and Structured Interviews Beat Gut Feel: What the Research Shows.