Sortinghat

Hiring Scorecards: How to Turn a Job Description Into Weighted Evaluation Criteria

A job description tells candidates about the role. A hiring scorecard tells your team who to reject, and why. Here is one built from a real job description, with the actual weights.

By , Founder9 min read

A Director of Engineering at a 40-person AI startup and a Director of Engineering at a 4,000-person services company are not the same job. Same title, same seniority, often the same salary band. Entirely different bar. Any hiring scorecard that gives one candidate one score across both roles is measuring the title, not the job.

That is the practical failure behind most wasted submissions. Not bad sourcing. A definition of the role that was never specific enough to reject anyone against.

This post walks through one real role, from the job description to the eight weighted criteria the system generated, to the candidate scored against them.

Key takeaways

  • A job description is not a scorecard. One is written for candidates and describes the role. The other is written for your team and defines the decision rule.
  • Weights are the actual scorecard. A list of criteria with no priority reads as eight equal must-haves, which means nothing gets shortlisted.
  • Context has to be scoreable. "Tier one college, and either a listed company or an early-stage startup" is a real hiring rule. It should be a criterion, not a note in someone's head.
  • Test the scorecard before you run it. Score two people you already know, one you would hire and one you rejected. If it cannot separate them, it is not measuring the job.
8
Weighted criteria generated from one job description
Worked example below
2.5x
Submission to first interview, technical roles
Sortinghat team, directional
2x
Submission to first interview, non-technical roles
Sortinghat team, directional

Why a job description is not a hiring scorecard

A JD is a sales document. It describes the role attractively enough that someone applies. It is written for the candidate, and it succeeds when people want the job.

A scorecard is the opposite. It is written for the people making decisions, and its job is to make rejection consistent. Two recruiters reading the same JD will reject different people. Two recruiters working from the same weighted scorecard will not.

The gap between them is where the cost sits. A vague brief produces a search against nothing in particular, a screening call that asks generic questions, and a submission that gets rejected three weeks later for a requirement nobody wrote down. The same failure shows up in req qualification, where the questions that would have surfaced it were never asked.

One real role, eight weighted criteria

Take a live role: Director of Engineering, India (Site Lead), for an AI company standing up a new India site. Full-time, Bangalore, one position, ₹1.5 to 2 crore.

The job description covers the usual ground. Build and lead the engineering team. Own delivery and outcomes. Maintain technical standards. Stay hands-on. Collaborate with US-based researchers. Establish and run the site. Shape the culture from the start.

Sortinghat job description review: a parsed Director of Engineering role beside structured fields for client, salary and location
Fig 1The job description and the structured role data come from the same input, so nothing is retyped between the client call and the search.

From that, the system generated eight evaluation criteria, each with a scoring guideline and a priority weight from 1 to 5. Candidates are rated out of 5 on each one, and the final score is the weighted average using those importance weights.

Sortinghat evaluation criteria: eight generated criteria with editable priority weights from 5 down to 2
Fig 2Eight criteria, generated from the job description, with the weights editable before anyone is scored.
CriterionWhat it testsWeight
Leadership ExperienceDemonstrated leadership in software engineering, at least 5 years managing teams and other managers5
Team BuildingProven ability to build or scale engineering teams from initial stages in startups or fast-paced environments5
Technical DepthHands-on full-stack product engineering experience, enough to set and hold a technical bar4
Operational SkillsManaging hiring, budgeting, site setup and operational processes4
Cross-Cultural CommunicationCommunicating across cultures and time zones, particularly with US-based teams3
India Site LeadershipEstablishing or leading an India site for a US or global company3
Reinforcement Learning KnowledgeFamiliarity with reinforcement learning, agent evaluation, reward models or ML infrastructure2
Academic PedigreeEducational background from recognised institutions2

Why criteria weights decide the shortlist

Look at what the weighting does to this role.

Reinforcement Learning Knowledge sits at 2. On a role at an AI company, the instinct is to make that a must-have. But this is a site lead hire. The person has to build a team and stand up an office. A brilliant RL engineer who has never hired anyone fails this role, and the weights say so before anyone wastes a call on it.

Meanwhile India Site Leadership sits at 3, above the RL knowledge and above academic pedigree. That is the client's actual risk, and a list of unweighted criteria would have buried it among seven other things that all looked equally mandatory.

A criteria list with no weights is eight equal must-haves. Nobody passes eight must-haves, so the recruiter quietly ignores the list and goes back to gut feel.

This is also the part clients are best at correcting. Most hiring managers cannot write a good brief from a blank page. Almost all of them can look at eight criteria with weights and say "swap those two." Send the scorecard, not the questionnaire.

How to get client sign-off on the weights

Send it before the search starts, not after the first rejection. Three practical rules:

  • Send it as a decision, not a draft. "Here is how we will evaluate. Tell us what to change" gets a response. "What are your requirements?" does not.
  • Force a trade-off. Ask which two criteria they would drop to fill the role a month sooner. The answer reorders the weights faster than any intake call.
  • Log the answer against the role. Feedback recorded on the criteria improves the next scorecard for that client. Feedback recorded in an email thread improves nothing.

Adding your own criteria in plain English

The generated criteria cover what the JD says. They cannot cover what the client believes but did not write down, and that unwritten part is usually where the rejections come from.

So the criteria are not a fixed list. A recruiter can add their own, in ordinary language, and give it a weight. Something like:

Candidates from tier one colleges should be from a listed company or an early-stage startup.

That is a real hiring rule, it is genuinely how experienced recruiters think, and it is the kind of thing that normally lives in one person's head and dies when they leave. Written as a criterion, it becomes something the system reasons about: it works out the size and stage of the companies on a CV and scores accordingly.

Which produces the effect from the top of this post. The same candidate, evaluated for the same title at two different companies, gets two different scores, because the criteria and the weights are different. That is not a quirk. It is the entire point.

A second example, from a non-technical desk

The same mechanism applies away from engineering, and the weights shift in ways that are worth seeing.

Criterion, enterprise sales roleWeightWhy
Quota attainment history5The only criterion that predicts the next year with any reliability
Deal size band5Someone selling ₹5 lakh deals does not transfer to ₹2 crore deals
Sales cycle length experience4Short-cycle sellers churn out of long-cycle roles inside a year
Domain familiarity3Helps, learnable, not worth rejecting over
Tooling and CRM discipline2Trainable in weeks

Note that the CV-impressive criteria sit at the bottom on both scorecards. That is usually the sign that the weights are honest.

Testing a scorecard before you run it

A scorecard is a hypothesis about who succeeds in the role. Running it across 200 applicants before testing it means finding out you were wrong at 200 times the cost.

So test it against people whose outcome you already know. Drop in two or three benchmark CVs, one you would definitely hire and one you rejected, and watch how they score. Adjust the criteria descriptions, the weights and the scoring guidelines until the separation matches your judgement.

If the scorecard cannot tell your best hire from your clearest rejection, it is not measuring the job yet. Better to find that out with three CVs than with three weeks of submissions.

Sortinghat evaluation panel: a candidate scored 96 against the eight criteria, each with its own justification
Fig 3The same eight criteria, applied to one candidate. Every percentage opens to show the reasoning behind it.

What a scorecard changes downstream

A scorecard is not documentation. It is the input to everything after it.

  • The search runs against the criteria rather than a job title, which is what makes natural language search useful instead of approximate.
  • Screening questions come from the criteria, so the call tests the things that actually decide the hire.
  • The submission summary evidences against the criteria, which is why clients read it.
  • Client feedback gets logged against the criteria, so the next role for that client starts better defined.

When our own team works this way, candidates move from submission to first interview at roughly 2.5x the previous rate on technical roles and 2x on non-technical ones. Those are our observed figures across our own delivery, not a controlled study, and they vary by client and by how well the role was defined to begin with. Treat them as directional.

The mechanism behind them is unglamorous. Fewer people get submitted. The ones who do are evidenced against criteria the client agreed to before the search started.

Frequently asked questions

What is a hiring scorecard?

A structured definition of who qualifies for a role, made up of evaluation criteria, a scoring guideline for each, and a weight reflecting how much each one matters. Candidates are rated against every criterion and the final score is the weighted average, which lets a team evaluate consistently rather than relying on each recruiter's reading of a job description.

How is a hiring scorecard different from a job description?

A job description is written for candidates and describes the role attractively enough that someone applies. A scorecard is written for the hiring team and defines the decision rule, including what disqualifies someone. One is a sales document, the other is criteria.

How do you weight hiring criteria?

By asking what the role actually risks failing on, not by what sounds most impressive. On a site-lead hire, team building and operational skills should outweigh specialist technical knowledge, because building the team is the job. Weights of 1 to 5 across six to ten criteria is a workable range.

Can the same candidate score differently for the same role at two companies?

They should. Same job title does not mean same job. If the criteria and weights reflect what each company actually needs, including company stage and size, the same profile will score differently against each, and that difference is the useful part.

How long does it take to build a hiring scorecard?

The criteria generate from the job description in one step. The part worth spending time on is adjusting the weights and adding the rules the client believes but did not write down, which is usually a fifteen to thirty minute conversation.

How do you test a hiring scorecard before using it?

Score two people whose outcome you already know: one you would definitely hire and one you clearly rejected. If the scorecard cannot separate them, it is not measuring the job yet. Adjusting it against three known CVs is far cheaper than discovering the problem after three weeks of submissions.

How to build your first scorecard

Take your last three roles that closed slowly. For each one, find the requirement that only surfaced after a rejection.

That requirement should have been a weighted criterion before the search began. The fact that it was not is the entire cost of skipping the scorecard, and it is a cost most firms pay on every role without ever putting a number on it.

Then build one scorecard for a role you are working right now, test it against two candidates you already have an opinion about, and send it to the client before you send a single CV.

Build a scorecard from one of your live briefs

Bring a job description you are working right now. We will generate the criteria, set the weights with you, and score two candidates you already have an opinion about.

Book a demo

Founder of Sortinghat, an AI-native ATS and CRM for staffing, search and RPO firms. Writes about recruiter capacity, sourcing economics and what actually changes when AI reaches a delivery desk. More about the author