Sortinghat

Diversity Hiring: What Works, What Is Theatre, and What Agencies Can Change

Most diversity effort is concentrated at the shortlist, which is the stage where the least leverage exists.

By , Founder5 min read

A diverse shortlist assembled from a non-diverse pipeline is arithmetic, not progress. Most of the effort in diversity hiring goes into the final stage, where the candidate set is already determined, and comparatively little goes into the earlier stages where the composition is actually set. For a staffing firm that matters commercially, because clients increasingly ask and most agencies answer with a policy rather than a method.

Last reviewed August 2026. Assessment drawn from delivery observations; figures deliberately omitted where unverified.

Key takeaways

  • Composition is decided upstream of the shortlist. By where you source, how the role is written and who your database contains.
  • Quotas at shortlist stage produce theatre. A required ratio drawn from a homogeneous pool changes the list, not the outcome.
  • Requirement inflation is the most common exclusion mechanism. And it is usually unintentional, which is why nobody catches it.
  • Structured evaluation reduces variance. Consistent criteria applied consistently is the least glamorous and most defensible intervention.
Upstream
Where composition is actually decided
Assessment
Requirements
The most common unintentional filter
Assessment
Structure
The intervention with the clearest rationale
Assessment
Evaluation criteria screen showing generated criteria with descriptions and editable priority weights
Fig 1Criteria generated from a job description, weighted before anyone is scored.

Why the shortlist is the wrong place to intervene

By the time a shortlist is assembled, the candidate set has already been filtered several times: by where the role was advertised, by how it was written, by who was sourced, and by who responded.

Applying a diversity requirement at that final stage changes which of the surviving candidates are presented. It does not change who survived. If the sourced pool was homogeneous, the shortlist is a rearrangement.

This is why clients who mandate shortlist ratios frequently see the ratio met and the hiring outcome unchanged, then conclude the effort does not work.

Requirement inflation, the mechanism nobody notices

The most consistent exclusion mechanism in hiring is a requirement that is not actually required.

A degree from a specific tier when the work does not need it. Continuous employment when a career break is irrelevant to capability. A number of years that is a proxy for something else. Each one is written without intent and each one filters unevenly.

The corrective is unglamorous: for every requirement, ask what happens if a candidate does not have it. If the honest answer is that they could still do the job, it is a preference rather than a requirement and it should be weighted accordingly rather than used as a gate. That is exactly what weighted criteria make visible.

Where sourcing actually decides composition

Three upstream levers, in order of effect.

Where you search. A database built from one channel over five years reflects that channel. Searching it produces the same composition indefinitely, which is a data problem rather than a market one.

How the role is described. Language, seniority framing and the stated requirements all affect who responds. This is testable: run two versions and compare.

Who you contact. Outbound gives you control that inbound does not, since you choose the list rather than receiving one.

What structured evaluation does and does not fix

Applying the same criteria consistently to every candidate reduces variance between assessors. That is worth having and it is the clearest-rationale intervention available.

What it does not do is correct for a biased input. Structured evaluation applied to a homogeneous pool produces consistent decisions about a homogeneous pool.

It also introduces its own risk. Any scoring system trained or configured on past outcomes can reproduce the pattern in those outcomes, which is precisely why the EU AI Act classifies recruitment AI as high-risk and why NYC Local Law 144 requires bias auditing for automated employment decision tools. A score you cannot interrogate is a score you should not defend.

What an agency can genuinely offer a client

Not a promise about outcomes, which is not yours to make. Four things that are.

A pipeline composition report at sourcing stage, not at shortlist. That is where the client can still act.

A requirements review flagging which criteria are gates and which are preferences.

Consistent criteria applied to every candidate, with the reasoning visible.

Honesty about what the market contains. If a talent pool is genuinely constrained, saying so is more useful than delivering a ratio that misrepresents it.

Candidate evaluation panel showing an overall score broken into criteria with written justification for each
Fig 2Every score opens to show the reasoning behind it.

Frequently asked questions

Why do diverse shortlists not change hiring outcomes?

Because composition is decided upstream. By the time a shortlist is assembled the pool has already been filtered by where the role was advertised, how it was written and who was sourced. A ratio applied at that point rearranges survivors.

What is requirement inflation in hiring?

Criteria that are stated as requirements but are not actually required: a specific institution tier, continuous employment, or a years figure standing in as a proxy. They are usually written without intent and they filter unevenly.

How do you test whether a requirement is genuine?

Ask what happens if a candidate does not have it. If the honest answer is that they could still do the job, it is a preference rather than a gate and should be weighted rather than used to exclude.

Does structured evaluation reduce bias?

It reduces variance between assessors, which is worth having. It does not correct a biased input, and any scoring configured on past outcomes can reproduce the pattern in them, which is why bias auditing is a regulatory requirement in several jurisdictions.

What should a staffing agency promise a client on diversity?

Not an outcome, which is not the agency's to promise. A pipeline composition report at sourcing stage, a review flagging which requirements are genuine gates, consistent criteria with visible reasoning, and honesty about what the market actually contains.

Where does sourcing affect diversity most?

In which database you search, since a pool built from one channel reflects that channel indefinitely; in how the role is described, which is testable by running two versions; and in outbound, where you choose the list rather than receiving one.

The audit worth running on one open role

Take a live role and list every stated requirement. For each, write what would happen if a candidate lacked it. Then count how many survive that test as genuine gates.

On most job descriptions the number is considerably smaller than the list, and the difference is doing more filtering than anyone intended.

See which criteria are gates

Bring a job description and we will show you the requirements weighted rather than applied as a single filter.

Book a demo

Founder of Sortinghat, an AI-native ATS and CRM for staffing, search and RPO firms. Writes about recruiter capacity, sourcing economics and what actually changes when AI reaches a delivery desk. More about the author