Sortinghat

Talent Pool Software: How to Rank Your Own Database Before You Pay a Job Portal

We audited 3 million candidate records sitting in one firm's Google Drive folders. A tenth were duplicates, a tenth were junk, and 80% were unique profiles nobody had touched in years.

By , Founder9 min read

One staffing firm we work with had 3 million candidate records. Not in an ATS. In Google Drive folders belonging to individual recruiters, most of which nobody else could open. They had no talent pool software, so every time a new role came in the team went to Naukri and paid to source candidates the firm already owned, several times over.

This is not unusual. It is the normal state of a staffing firm that has been operating for more than a few years. The database exists, it is large, and it is functionally invisible.

So we audited it. Here is what 3 million records actually contained, what we did with them, and what happened to their sourcing costs once the pool became searchable.

Key takeaways

  • Only 20% of a large candidate database is genuinely dead. In our audit, 10% were duplicates and 10% held no usable company information. The other 80% were real, unique profiles.
  • Out of date is not the same as irrelevant. Most of that 80% had changed jobs since the record was created. That makes the record stale, not the person.
  • The barrier is access, not volume. Records scattered across individual recruiters' drives are not a talent pool. They are storage with a headcount problem attached.
  • The saving shows up per role. That firm's job portal spend fell from roughly $100 a role to $60, and it is still falling as the pool gets worked.
3M
Candidate records audited across one firm's drives
Sortinghat, 2026
80%
Unique, still-relevant profiles after dedup and junk removal
Sortinghat, 2026
$100 → $60
Job portal spend per role, before and after
Same firm, same period

What a 3 million record candidate database audit found

The records were spread across Google Drive folders owned by individual recruiters. Some of those recruiters had left the business. Nobody had a complete view, and no single search could touch more than a fraction of it.

When we consolidated and audited the lot, it broke into three groups.

SegmentShareWhat it was
Duplicates~10%The same person entered more than once, from different sources at different times, with the history split across the copies
Junk~10%No usable company or role information. Not enough to identify what the person does, let alone place them
Unique and viable~80%Real, distinct profiles. Mostly out of date, because these people had changed roles since the record was created

The 80% is the finding that matters, and it is the opposite of what most firm owners assume. Ask a staffing founder how much of their old database is worth anything and the answer is usually a shrug and a low number.

Out of date is not the same as irrelevant. A backend engineer whose record says they were at a fintech in 2022 has almost certainly moved. They are still a backend engineer. What is stale is the record, not the person, and the fix is enrichment rather than deletion.

What to do with the 10% that is genuinely junk

Delete it, and do so deliberately rather than letting it sit.

A record with no company, no role and no reachable contact is not a lead you have not got round to. It is noise in every search result you will ever run, and it is a compliance exposure. Under India's DPDP Act and the GDPR, holding personal data with no lawful basis and no defensible retention position is a liability, not an asset. Junk records fail that test twice over: you cannot justify keeping them and you cannot use them.

What makes a record placeable in your talent pool

Three tests. All three have to pass.

Reachable

At least one channel that works right now. In India that usually means a live WhatsApp number before it means an email address. A work address from a job the person left in 2023 is not a channel, and a personal email captured at the time of first contact is worth more than either.

Relevant

Their current skills and level map onto the open role, not onto the role they held when the record was created. This is where the 80% sits: relevant people behind stale records. It is also the test most databases fail silently, because the record still looks complete.

Receptive

Some signal that they are open to a move. A recent application, a reply to earlier outreach, a stated notice period, a contract nearing its end. This is the hardest of the three to maintain and the one that decays fastest.

A record passing all three is a candidate. Passing two, it is a lead. Passing one, it is storage. What we did with those 3 million records was move as many as possible from the third category into the first, by refreshing what each person is doing now and how active they currently are.

Why candidate scoring has to show its evidence

Scoring in most systems is a black box, which is why recruiters ignore it. A score that says 96 and nothing else is an opinion. Recruiters do not act on opinions from software they did not write.

So the score has to be answerable. Open a candidate on a role and you should see which criteria they met, at what level, and the reasoning behind each one, in plain language.

Sortinghat evaluation panel: a candidate scored 96, split into weighted criteria with a percentage and written justification for each
Fig 1Every score opens. The candidate scored 96 overall, with the reasoning behind each criterion readable underneath it.

That is the difference between a ranking a recruiter uses and one they scroll past on the way to opening a job portal. A recruiter who can see why one candidate outranked another will work the list. One who cannot will fall back on gut feel and a fresh portal search.

Sortinghat pipeline for a Director of Engineering role, with candidates in stage columns showing match scores of 96, 93 and 92
Fig 2Scores carry through the pipeline, so a recruiter triaging a column is reading ranked candidates rather than a list.

Database-first sourcing on one real role

The sequence matters more than the tooling.

  1. Search the existing pool in plain language before doing anything else.
  2. Apply availability, location and rate as conditions in the query rather than filtering afterwards.
  3. Refresh the top matches whose information is older than your threshold, so you know what these people are doing now.
  4. Run outreach across WhatsApp, email and call in parallel, not one after the other.
  5. Only then decide whether the role genuinely needs a fresh portal pull.

Most roles do not get past step four. That is the entire mechanism behind the cost reduction. Nothing was cancelled. The reason to reach for the portal was removed on most roles.

Sortinghat Advanced People Search returning 71 ranked candidates for a plain-English query about engineers in Bangalore
Fig 3Searching the pool in plain language. 71 candidates, ranked, with the career timeline visible before anyone opens a profile.
DimensionPortal-firstDatabase-first
First candidate contactedAfter the posting goes liveWithin the hour
Marginal cost per roleA fresh portal pullNear zero
Candidate familiarityColdHas dealt with you before
Effect on the databaseNoneEvery search refreshes the pool
Where it still winsNew skills, new geographies, roles you have never workedNothing in your history to draw on

What database-first sourcing changed on the P&L

That firm was spending roughly $100 a role on job portals. It is now around $60, and the number is still moving, because the pool improves every time a recruiter works it.

Portal spend is usually the second or third largest controllable cost in a staffing firm after salaries, and it is the one nobody audits, because it arrives as an annual subscription rather than a per-role charge. Nothing on that invoice tells you how many of last month's roles could have been filled without it.

The second effect is capacity. Recruiters who start inside the pool instead of a portal carry roughly 2x the roles they carried before, because the slowest part of the process was never the screening. It was starting from zero on every requisition. If you want to model this for your own desks, our recruiter capacity guide walks through the arithmetic.

There is a third effect that takes longer to show up. A pool that gets worked gets better. Every search refreshes records, every conversation confirms a notice period, every rejection logs a reason. A portal pull compounds nothing, because next month you start again.

Frequently asked questions

What is talent pool software?

Talent pool software consolidates, refreshes and ranks the candidates a staffing firm has already sourced, so recruiters can search their own database before paying to source externally. It differs from an ATS record store in that the pool is actively kept current and scored against live roles, rather than passively archived.

How much of an old staffing database is still usable?

In our audit of 3 million candidate records at one firm, about 10% were duplicates and about 10% held no usable company information. The remaining 80% were unique, real profiles that were out of date rather than worthless. Your own split will differ, and a sample audit of 500 records is the only reliable way to know it.

Does database-first sourcing replace job portals like Naukri?

No. It reduces how often you need them. Genuinely new skill requirements, new geographies and roles you have never worked still justify a fresh portal pull. The change is that these become the exception rather than the default on every role. In our customer data, job portal spend fell from roughly $100 per role to $60.

How is candidate scoring different from keyword matching?

Keyword matching checks whether terms appear in a CV. Scoring weighs the whole profile against the role's criteria and shows the reasoning behind each part of the result, so a recruiter can verify it rather than trust it blindly.

What should you do with duplicate candidate records?

Merge rather than delete, and only with a tool that preserves notes, call history, attachments and submission history from every source record. A merge that loses the 2023 note explaining why a client rejected someone will cost you that client's trust when you submit them again.

How long can a staffing firm keep candidate data?

Retention has to be defensible rather than indefinite. India's DPDP Act and the GDPR both require a lawful basis for holding personal data and a stated retention position. Records with no reachable contact, no placement history and no recent activity are usually a compliance exposure rather than an asset.

What talent pool software does not solve

Three limits worth stating, because a pool that is oversold gets abandoned in month three.

It cannot create depth you do not have. A firm entering a new vertical has no history in it, so the pool returns almost nothing and every role starts on a portal. That is correct behaviour, not a failure, and it resolves after two or three placements build the segment.

It cannot fix contact data that was never captured. If your team has historically recorded only a work email, the pool inherits that gap. Enrichment helps, and it is not a substitute for capturing a personal channel at first contact.

It will not make anyone search it. A pool nobody opens is an expensive archive. Adoption is the constraint, and it is the reason we treat implementation as delivery work rather than a login handover.

How to audit your own candidate database

Before you buy anything, find out where your database actually lives. If the honest answer includes individual recruiters' drives, folders and inboxes, you do not have a data quality problem yet. You have an access problem, and it is the cheaper of the two to fix.

Then take 500 records at random and check two things: does the contact information work, and is the stated role still current. That percentage is the real size of your talent pool, and it is the only number that will settle the internal argument.

If you are moving off an existing system, do the audit before the migration rather than after. Our guide on what actually breaks in an ATS migration covers the order of operations, and the migration page sets out how we handle it.

See your own database scored

We will consolidate your existing records, show you the duplicate and junk split, and rank what is placeable against a live role.

Book a demo

Founder of Sortinghat, an AI-native ATS and CRM for staffing, search and RPO firms. Writes about recruiter capacity, sourcing economics and what actually changes when AI reaches a delivery desk. More about the author