Sortinghat

AI Interviews: How the Questions Are Built, What Gets Scored, and Where the Human Sits

An AI interview is only as good as its questions, and most are drawn from a generic bank. Here is what changes when they come from the role's own criteria instead.

By , Founder10 min read

Most AI interview tools ask the same questions of everyone with the same job title. That produces a score which correlates with how articulate someone is, not with whether they can do the job. Question design, not voice quality, is what separates a useful AI interview from an expensive filter, and it is almost never what buyers evaluate in a demo.

Key takeaways

  • Questions must trace to the role, not a bank. Ours are generated from three inputs: the job description, the weighted evaluation criteria, and the company context.
  • A score without evidence is an opinion. The rubric has to show which answer produced which judgement, or recruiters will ignore the ranking.
  • Human review before rejection is a legal obligation. The EU AI Act classifies recruitment AI as high-risk and requires oversight of automated decisions.
  • Completion rate is a design metric. An interview candidates abandon filters for patience, and the people with the most options abandon first.

Where AI interview questions come from

Three inputs, and every question traces back to one of them.

The job description

The role, the level, the stated requirements and the responsibilities. This is the floor. A question set built from the JD alone is better than a generic bank and still misses most of what decides the hire.

The evaluation criteria

This is the part that matters. If the role's criteria weight team building at 5 and reinforcement learning knowledge at 2, then the interview should spend its time accordingly. Questions inherit the weighting, so the conversation concentrates on what actually decides the submission rather than distributing attention evenly across a checklist.

The company context

The same title at a 40-person startup and a 4,000-person services firm is not the same job. Company stage and size change what a good answer sounds like, and a question set that ignores this produces a score that transfers badly between clients.

Evaluation criteria screen showing generated criteria with descriptions and editable priority weights, plus a field for adding custom criteria
Fig 1Questions inherit these weights. A criterion at 5 gets interview time, a criterion at 2 gets a single check.

Behind the generation sits a question framework built by HR practitioners with two decades in staffing. That is not a credential for its own sake. The difference between a question that surfaces a real answer and one that invites a rehearsed answer is craft, and it does not emerge from a template.

What the AI interview actually scores

Three categories, and the distinction between them is where most tools quietly fail.

CategoryExampleConfidenceWeight it should carry
Verifiable factsNotice period, location, years in a technologyHighBinary or numeric, treat as a gate
Evidenced claimsWhat they built, owned, and what happenedMediumThe core of the assessment
Inferred signalsStructure, communication, domain fluencyLowLeast weight of the three

The failure mode is weighting the third category like the first. An articulate candidate outscores a capable one, the firm submits them, and nobody finds out for two months. Fluency is easy to measure and weakly predictive, which is a dangerous combination in any scoring system.

Screening call versus interview: two different jobs

These get conflated constantly, and running one when you needed the other wastes both the candidate's time and yours.

Screening callAI interview
QuestionAre they eligible and available?Can they do this job well?
Length3 to 7 minutesLonger, and depth varies by role
TestsInterest, notice, compensation, CV claimsCapability against weighted criteria
Runs onA broad contacted listA qualified shortlist
OutputAdvance or return to poolA ranked assessment with evidence

The sequence that works is qualification first, assessment second. Running a full interview on an unqualified list burns candidate goodwill on people who were never going to pass the notice period check.

Why generic question banks produce useless scores

A bank sorted by job title asks the same twelve questions of every backend engineer. Three problems follow.

It measures preparation, not capability. Common questions have common answers, and candidates who have interviewed recently outperform candidates who have not, independent of skill.

It cannot weight. Twelve equally-asked questions imply twelve equally-important requirements, which is never true of a real role.

It transfers badly between clients. The same score means different things at two companies hiring the same title, which makes the number unusable for exactly the comparison you wanted it for.

Why a score has to show its evidence

Scoring in most systems is a black box, which is why recruiters ignore it. A number with no reasoning attached is an opinion from software the recruiter did not write and has no reason to trust.

The score has to be answerable. Open a candidate and you should see which criteria they met, at what level, and the reasoning underneath each one in plain language. Leadership experience, technical depth, operational skills, each scored separately with the justification readable.

Candidate evaluation panel showing an overall score of 96 broken into criteria, each with a percentage and written justification
Fig 2Every percentage opens. A recruiter who can read the reasoning will use the ranking.

That readability is not a nicety. It is what makes the ranking usable, and it is also what makes the process defensible when a client or a candidate asks how a decision was reached.

Bias controls, and what to ask any vendor

Vague assurances here are worthless. Four things a buyer should be able to verify before signing anything.

  1. Which attributes the system is explicitly prevented from using
  2. Whether accent, fluency or speech pace influence scoring, and how that is tested
  3. What the audit process is, how often it runs, and who runs it
  4. Whether audit results are available to customers

A vendor who cannot answer these has not thought about it, or has and would rather you did not ask. Under NYC's Local Law 144, bias auditing is a requirement rather than a courtesy for automated employment decision tools used on New York roles.

Insert required Sortinghat's specific answer to all four questions above, plus completion rate and average duration for AI interviews with the sample size. If any of the four cannot be answered yet, say so and give a date. A stated gap is more credible than a claim you cannot evidence, and a buyer's legal team will ask.

What clients ask before they approve an AI interview

If you place into enterprises, this conversation is coming. Five questions their legal or talent team will put to you.

  1. Are candidates told it is automated, and when? Before the first question, or it is not disclosure.
  2. Who makes the rejection decision? The only safe answer is a named human.
  3. Has the tool been bias audited, by whom, and how recently?
  4. What happens to the recording and the transcript? Retention period, and whether it trains anything.
  5. What is the accommodation path? For candidates with disabilities, and for anyone who declines.

Have written answers to all five before you need them. A supplier who has to go away and find out has already told the client something about how carefully this was set up.

Where the human checkpoint sits

The AI interview does not reject anybody.

It assesses, evidences, flags and ranks. A recruiter reviews before any candidate is dropped and before anything reaches a client. That is partly a quality decision and mostly a legal one.

The EU AI Act classifies AI used in recruitment and candidate selection as high-risk, with obligations in force since 2 August 2026 that include human oversight of automated decisions. Illinois regulates AI-analysed video interviews specifically. Colorado has its own AI Act. A workflow with no human checkpoint is not deployable in several of the markets your clients hire into, regardless of how well it performs.

The practical division most firms reach: the AI handles assessment at volume, the recruiter handles the judgement call, the sell and the client conversation.

Disclosure and the candidate's side

Candidates should be told, before the interview begins, that it is automated, on whose behalf it is being run, and whether it is recorded. One sentence at the start, not a line in a follow-up email.

There should also be a human path available and visibly offered. A process with no opt-out creates a candidate experience problem and a compliance problem simultaneously, and the candidates most likely to want the opt-out are the ones with options.

Completion rate is the number that tells you whether you got this right. An interview people abandon is not assessing anyone. It is selecting for tolerance, which is not a hiring criterion anyone would write down.

Candidate experience, and the completion problem

Three things move completion rate, and none of them is voice quality.

Length

Every extra question costs completion. The discipline is to cut anything that does not trace to a weighted criterion, which usually removes a third of a first draft.

Clarity about what happens next

A candidate who does not know whether a person will see this, or when they will hear back, disengages faster. Saying both at the start costs fifteen seconds and buys more completion than any interface change.

A visible human path

Knowing you can speak to a person makes people more likely to complete the automated version, not less. The opt-out is used rarely and its visibility does most of the work.

The uncomfortable part is who abandons. It is disproportionately the candidates with options, which means a poorly designed interview filters out exactly the people you were trying to reach and reports it as a completion statistic.

Where AI interviews do not belong

Three situations where running one costs more than it saves.

Senior and executive roles. The first conversation is persuasion. A candidate at that level who receives an automated assessment concludes the role is not serious, and they are not wrong.

Roles with no clear criteria. If you cannot write down what disqualifies someone, the interview has nothing to assess against. That is a briefing problem, and automating on top of it produces confident scores about the wrong things.

Where the client will not accept it. Some clients have policies. Find out before you run it on their candidates, not after.

Frequently asked questions

How are AI interview questions generated?

From three inputs: the job description, the weighted evaluation criteria for that role, and the company context. Each question traces back to something being assessed, which is what separates a role-specific interview from a generic question bank sorted by job title.

Can an AI interview reject a candidate?

It should not. The EU AI Act classifies recruitment AI as high-risk and requires human oversight of automated decisions. The workflow that holds up is the AI ranking and flagging, with a person making every rejection and every submission decision.

Is AI interviewing legal?

It is regulated rather than prohibited, and the requirements differ by jurisdiction. The EU AI Act, NYC Local Law 144, Illinois AIVIA and the Colorado AI Act all impose obligations around disclosure, bias auditing and human oversight. The rules follow the candidate's location, not your office's.

What should an AI interview actually score?

Verifiable facts such as notice period and compensation, evidenced claims about what someone built or owned, and inferred signals such as structure and domain fluency. The third group carries the least confidence and should carry the least weight, which is where most tools go wrong.

How is an AI interview different from an AI screening call?

A screening call qualifies: interest, availability, compensation and the CV claims that decide the submission. An interview assesses capability against the role's criteria in more depth. The first is about eligibility, the second is about fit.

What is a good completion rate for an AI interview?

There is no cross-industry standard worth relying on. Measure your own, segment it by role type and seniority, and treat a falling rate as a design problem rather than evidence of candidate quality. Candidates with the most options abandon first.

Interview scoring and the scorecard are the same object

A common mistake is treating the interview as a separate stage with its own criteria. It is not, and separating them creates the inconsistency that clients notice.

The scorecard defines what matters and how much. The search runs against it. The screening call verifies the eligibility parts of it. The interview assesses the capability parts of it. The submission summary evidences against it. One definition, four applications.

When these drift apart, a candidate scores 90 on the interview and gets rejected by the client for something that was never assessed, because it lived only in the brief. Keeping them tied together is what makes the score predictive rather than decorative.

How to evaluate an AI interview tool

Run it on a role you have already filled. Feed it one candidate you hired and one you clearly rejected.

If the ranking does not separate them, the questions are not measuring the job. That test takes an afternoon, uses outcomes you already know, and is worth more than any demo.

See an AI interview built from your own criteria

Bring a role you have filled. We will generate the questions from its criteria and score two candidates you already have an opinion about.

Book a demo

Founder of Sortinghat, an AI-native ATS and CRM for staffing, search and RPO firms. Writes about recruiter capacity, sourcing economics and what actually changes when AI reaches a delivery desk. More about the author