Sortinghat

India's DPDP Act and Candidate Data: What Agencies Must Change

Every staffing firm in India holds a large database of personal data collected over years, mostly without a documented basis.

By , Founder5 min read

A candidate database is personal data at scale. Under India's data protection framework, that brings obligations around what you collected it for, how long you keep it, and what happens when somebody asks you to delete it. Most Indian staffing firms hold records going back a decade, gathered through channels nobody documented, and have never had to answer any of those questions.

Last reviewed August 2026. Summary of published framework. Take legal advice before acting.

Key takeaways

  • A CV is personal data and a database of them is a liability as well as an asset. Both statements are true at once, which is the part firms find uncomfortable.
  • Purpose limitation is the hardest requirement for agencies. Data collected for one role, reused for another, over years.
  • Retention needs a defensible position. Indefinite storage of records with no reachable contact is hard to justify.
  • Deletion requests will arrive. And the firm needs to be able to find every copy of a person, which duplicates make impossible.
~10%
Duplicates in one audited 3M record database
Sortinghat
~10%
Records with no usable information
Sortinghat
Purpose
The requirement agencies find hardest
Regulatory framework
Advanced people search returning ranked candidates for a plain-English query, with career timelines and fit badges
Fig 1Searching an existing database in plain language, with the career timeline visible before anyone opens a profile.

Why a candidate database is the exposure

India's India's DPDP Act framework treats personal data as something held for a stated purpose, with a lawful basis, for a defensible period.

A staffing database fails all three tests by default. It was collected across years through portals, referrals, uploads and imports. The purpose was whatever role was open at the time. The retention period is forever, because nobody ever deleted anything.

That is not negligence, it is how the industry worked. It is also now a documented position you may be asked to produce.

Purpose limitation, and why it bites agencies hardest

The principle is that data collected for one purpose should not be used indefinitely for unrelated ones.

For an internal talent team that is manageable: they collected applications for their own roles. For a staffing firm, the entire business model is reusing a candidate record across many clients and many years, which is precisely the pattern the principle constrains.

The workable position is that the purpose was recruitment services generally rather than one specific role, stated clearly at collection. That requires the statement to exist at collection, which for most historical records it does not. New records are fixable. The back catalogue is the harder question.

Retention, and the case for deleting junk

Indefinite retention of everything is the weakest position available and it is what most firms currently have.

The useful reframe is that a large part of what you hold has no value anyway. An audit of one firm's 3 million records found roughly 10 percent duplicates and 10 percent with no usable company or role information at all. A record with no reachable contact, no placement history and no recent activity is noise in every search and a liability with no offsetting asset.

Deleting it improves search quality and reduces exposure at the same time, which makes it the rare compliance action with an immediate operational payoff. The mechanics are in duplicate candidate records.

Deletion requests and why duplicates make them dangerous

A candidate asking you to delete their data creates a specific operational problem: you have to find every record of them.

If one person exists three times under different spellings and two email addresses, a firm that deletes one record has not complied and has told the regulator it did. That converts a data hygiene problem into a compliance failure.

This is the practical argument for matching at the point of entry rather than periodic cleanup. It is also why the identity problem in blue collar databases carries more risk than the white collar equivalent, since phone-based identity produces more duplicates.

Where the data lives, and who else has it

Two questions most firms cannot answer quickly.

Where is it hosted? Enterprise and GCC clients ask this before they ask about delivery capability. Data hosted in the client's own country is a straightforward answer; anything else needs explaining.

Which third parties have it? Every integration is a place candidate data leaves your system. Job portals, messaging platforms, assessment tools, background verification providers. You remain responsible for it, which is covered in what integrations actually move.

A practical sequence

Fix collection first. New records get a clear purpose statement and a documented basis. This stops the problem growing while you deal with the rest.

Delete the junk. No contact, no history, no activity. This reduces exposure and improves search in one action.

Deduplicate. So that a deletion request can actually be honoured.

Document retention. A stated period with a rationale beats no position, even if the period is long.

Map your integrations. Know which third parties hold candidate data and on what basis.

Candidate activity timeline showing an automatically logged call written to the record with the stage move attached
Fig 2Calls, meetings and messages written to the record without anyone typing.

Frequently asked questions

Does the DPDP Act apply to recruitment databases?

A candidate database is personal data at scale, so the framework's requirements around purpose, lawful basis and retention apply to it. Staffing firms are more exposed than internal talent teams because reuse across clients and years is the business model.

How long can a staffing firm keep candidate data?

There is no single number, but indefinite retention of everything is the weakest position. A stated retention period with a rationale is defensible; holding records with no reachable contact and no placement history is difficult to justify.

What happens when a candidate asks for deletion?

You must find and remove every record of that person. Duplicates make this dangerous, because deleting one record while two remain means you have failed to comply and reported that you did.

Should staffing firms delete old candidate records?

Records with no reachable contact, no placement history and no recent activity cost more in search noise than they provide in value, and they carry exposure with no offsetting benefit. Deleting them improves search and reduces risk simultaneously.

Where should candidate data be hosted?

Enterprise and GCC clients typically ask this before they ask about delivery capability. Hosting in the client's own country is a straightforward answer; anything else requires explanation during procurement.

Do integrations create data protection risk?

Every integration is a point where candidate data leaves your system, and you remain responsible for it. Knowing which third parties hold candidate data, and on what basis, is part of the position you need to be able to state.

The question to answer before a client asks it

If a client's legal team asked today where your candidate data sits, how long you keep it, and which third parties have copies, could somebody answer in the meeting?

On most desks the answer is no, and the gap is documentation rather than practice. That is a week of work and it removes a recurring obstacle in enterprise procurement.

See your duplicate and junk split

We will audit your existing records and show you what is duplicated, what is unusable, and what is placeable today.

Book a demo

Founder of Sortinghat, an AI-native ATS and CRM for staffing, search and RPO firms. Writes about recruiter capacity, sourcing economics and what actually changes when AI reaches a delivery desk. More about the author