Recruitment advertising for frontline employers

We advertise frontline jobs on social media. This page is about judging the hires that came out the other end.

Book a demo

Hiring guideChecked against 29 CFR 825 at source, 15 September 2026

Quality of Hire: How to Measure It Without Fooling Yourself

Quality of hire is a score you put on the people you hired, computed the same way every time. For hourly frontline work the version that works has three parts you can pull from a timeclock and a schedule: was the hire still working here on day 90 (50 points), did they reach the role’s normal output inside the window your own last ten hires set (30 points), and did they work the shifts they were scheduled for (20 points). The formula you will find everywhere else averages a performance rating, a ramp estimate, an engagement score and a culture-fit score, most of them out of one person’s head — an elaborate way of asking whether the hiring manager liked them. One warning before the arithmetic: if you hire six people a year, quality of hire is a story about six people, not a statistic. Below: the components, the scoring, a worked example, the sample-size problem, and how to attribute the score to a source.

The formula everyone quotes, and what it actually measures

Open any article on quality of hire and you will find a version of this:

Quality of hire = (performance rating + ramp-up score + engagement score + culture-fit score + hiring manager satisfaction) ÷ number of indicators

There is no canonical source for it. It appears in vendor documentation and HR articles in different forms with different weights, which is the first problem: two employers using “the standard formula” are not computing the same thing. The deeper problem is the inputs themselves.

One opinion, counted four times

The performance rating is the supervisor’s. The engagement score, in most small operations, is the supervisor’s impression. Culture fit is that impression with a nicer name. Hiring manager satisfaction is, explicitly, the supervisor being asked whether they are satisfied. Four of the five inputs trace to one person, and averaging them does not cancel that person’s bias — it disguises it as arithmetic. A hire who is liked scores 4.5 on all four; a competent, abrasive hire scores 3 on all four. The composite has told you nothing you did not know in the break room, and it is the kind of judgement that causes trouble if challenged: unfalsifiable, about the person rather than the work. Standards that hold up are hiring criteria.

It arrives too late, and compares to nothing

The composite is scored at six or twelve months, because that is when a performance rating exists. A number that tells you in October that January’s campaign produced weak hires is history, not a metric — you are advertising again next month. And because every employer picks its own inputs, a published quality-of-hire benchmark compares nothing to nothing, which is why it is one of the four metrics we refuse to put a number against on hiring metrics benchmarks. The score below is useful against your own last quarter and your own other sources, and against nobody else’s.

A metric you cannot recompute by hand, from records a machine wrote, is not a metric. It is a rating with a decimal point.

A definition you can compute from a timeclock and a schedule

Three components, all recorded by systems rather than opinion, all available on day 90.

1. Ninety-day survival — 50 points

Was the person still employed on their ninetieth day? Yes or no. This carries half the score because in frontline work it carries half the meaning: a hire gone by day 90 returned nothing on the advertising, the training or the supervisor’s time. Our first-90-days page argues that most of those exits were decided before the ad was written — the shift was nights when the ad said flexible — which is why this belongs in a hiring metric, not a management one.

Two rules keep it honest. Count day-one no-shows and accepted offers who never started as zeros; excluding them is how a source with a bad show-up rate looks respectable. And do not net out the ones you dismissed — an involuntary exit inside 90 days is a hiring miss too.

2. Time to full productivity — 30 points

Full productivity is not a feeling. Almost every hourly role has a countable thing per unit of time and the crew already runs at a known rate: picks per hour, routes per day, clients per shift, units off the line. Pick one and define full productivity as sustaining that rate across five consecutive shifts with no trainer standing there.

The target window should be yours: the median days your last ten hires in that role took to hit the rate. If you have never tracked it, use what the trainer says it should take and correct it after ten more hires. A benchmark window measures a stranger’s training program.

3. Reliability in the first ninety days — 20 points

Scheduled shifts against shifts worked. Every supervisor tracks this informally and almost nobody writes it down. It is the earliest signal of whether the job as lived matches the job as advertised, and whether the commute and the childcare work at that start time.

Count only unexcused absences and no-call/no-shows. Approved time off, agreed swaps and anything granted as an accommodation come out before you score anybody. The discipline is worth building in: 29 CFR 825.220(c) states that employers cannot use the taking of FMLA leave as a negative factor in employment actions and that it cannot be counted under no-fault attendance policies. Most brand-new hires are not yet eligible — 29 CFR 825.110 sets 12 months of employment and 1,250 hours of service among the conditions — but a habit that quietly counts protected leave against people will eventually reach someone it applies to.

Scoring it out of 100

One rule before the bands: a hire gone before day 90 scores zero, full stop. You cannot score the ramp of somebody who did not finish it, and partial credit for a fast-ramping three-week quit is the arithmetic that makes a bad quarter look middling.

ComponentPointsHow it scores
Ninety-day survival50 or 0Employed on day 90 — 50. Gone for any reason, or never started — 0, and the whole score is 0.
Time to full productivity30 / 20 / 10 / 0Inside your target window — 30. Up to 1.5× — 20. Up to 2× — 10. Beyond that, or not there at day 90 — 0.
Reliability, first 90 days20 / 12 / 6 / 0No no-call/no-shows and two or fewer unexcused absences — 20. None and three or four — 12. One no-call/no-show — 6. Two or more — 0.

A cohort’s figure is the average across every hire including the zeros, reported beside the survival rate: the average alone cannot say whether a 58 means most hires were mediocre or most were good and a third left.

A worked example: twelve order pickers

A distribution operation hired twelve pickers over a year, from paid social advertising and from an employee referral bonus. Full productivity is the crew rate over five consecutive shifts; the target window, from the ten hires before this cohort, is 30 days. Every figure is invented to show the arithmetic, bar one noted below.

HireSourceDay 90Days to full rateUnexcused / NCNSScore
1AdsYes261 / 0100
2AdsYes413 / 082
3AdsYes380 / 090
4AdsYes552 / 166
5AdsLeft day 220
6AdsLeft day 610
7AdsLeft day 90
8ReferralYes240 / 0100
9ReferralYes292 / 0100
10ReferralYes341 / 090
11ReferralYes474 / 072
12ReferralLeft day 440

The cohort scores 700 ÷ 12 = 58.3 on 66.7% survival. By source: advertising 338 ÷ 7 = 48.3 on 57.1%, referrals 362 ÷ 5 = 72.4 on 80%. Among hires who reached day 90 the two are close — 84.5 against 90.5 — so the referral advantage is almost entirely survival: a referred picker arrives knowing what the shift is really like, and an advertised one knows what the ad said.

With six hires a year, this is a story, not a statistic

The example produced a clean-looking finding — referrals beat advertising by 24 points — and you should not act on it, because seven hires and five hires cannot support it. The exact binomial interval on 4 survivors out of 7 puts the honest range at 18% to 90%; on 4 out of 5 it is 28% to 99%. Test the gap and the p-value is 0.58: if the sources were identical, a gap this big would turn up more than half the time by chance.

Hires measuredSurvivorsObserved survivalHonest range (95%)
3266.7%9% – 99%
6466.7%22% – 96%
12866.7%35% – 90%
241666.7%45% – 84%
402870.0%53% – 83%
1006767.0%57% – 76%

Exact binomial (Clopper–Pearson) intervals, reproducible in a spreadsheet: lower bound BETA.INV(0.025, k, n−k+1), upper bound BETA.INV(0.975, k+1, n−k), k survivors and n hires. Just the arithmetic of small numbers.

To tell a source with 50% ninety-day survival apart from one with 75%, at conventional confidence and power, you need roughly 29 hires from each. Most frontline employers will never have that in a year.

What to do with six hires. Three things, none a dashboard.

  • Read them as cases, not as a rate. Why did each of the ones who left, leave? Those answers are specific, usually fixable, and worth more than an average of six numbers.
  • Pool until it means something. Keep the definition fixed and accumulate across quarters and similar roles. Twenty-four hires scored identically is a measurement; twelve definitions is not.
  • Act on the big gaps only. If a source produced eight hires and seven are gone by day 90, you do not need a statistic to stop using it.

The same applies to scoring by supervisor: two supervisors on four hires each measures noise. Set your early exits against the published rates for your sector — turnover rates by industry — to see whether your problem is unusual.

Attributing the score to a source, so it changes what you spend

A quality-of-hire score not broken out by source is a report. Broken out by source it is a budget decision, and the mechanics are simpler than the software makes them sound. You need one field: where this person came from, written on the hire record on day one and carried into payroll, so at day 90 you can sort survivors by source. That is the join between advertising and HR that almost nobody has; recruiting analytics covers the tagging.

Then compute one figure per source: cost per surviving hire — what a hire from that source cost, divided by that source’s ninety-day survival rate. Carrying the example forward, at 14 applicants per hire:

AdvertisingReferral bonus
Cost to produce one hire$9.83 × 14 applicants = $137.62 media$250 bonus
Ninety-day survival57.1% (4 of 7)80% (4 of 5)
Cost per surviving hire$240.84$312.50
Early quits per surviving hire0.750.25

The $9.83 is the only real number there: the median cost per applicant for warehouse and production roles in our 2026 benchmark — 891 Boostpoint-managed Meta campaigns, 1,334 campaign-months, management fee inside. The 14 applicants per hire is your ratio; the $250 is invented. And media cost per hire is not cost per hire.

On raw cost advertising wins, $240.84 against $312.50. But it leaves half an extra early quit per surviving hire, and that has a price. Set the totals equal: they break even when an early quit costs $143. Above that, the dearer source is the cheaper one. Whether an early quit costs you more than $143 is not ours to answer — build it from your own wages, training hours and lost output on the turnover cost calculator, and cost per hire by industry sets out the published surveys.

The point is not the $143. A source’s quality and its price are one number, not two, and ranking channels on cost per applicant alone is ranking them on half the equation.

Applicants per hire is what makes this computable, and it is yours — it depends on how well your form converts, which is what candidate conversion rate is about. Attribute by hire, not applicant volume: a source delivering 200 applications and two survivors is worse than one delivering 20 and three. Sizing the flow itself is on candidate pipeline.

What to stop measuring

  • Hiring manager satisfaction surveys. The manager picks, trains, schedules and then rates the outcome. Not an independent measurement.
  • Culture-fit scores. Unfalsifiable, correlated with liking, and the phrase most likely to turn up in a complaint.
  • Engagement scores on people with under 90 days of service. A new hire says it is going fine because they want it to be. Ask the day-seven question — is this the job you applied for — and treat the answer as feedback on the ad.
  • Composite indexes nobody can reproduce. If a supervisor cannot rebuild the number on the back of a schedule, it will not be used.
  • Twelve-month performance ratings used as a recruiting metric. By month twelve you are grading a year of management, not a hiring decision.

What this score cannot see

  • Excellence. The picker who trains everyone else and the one who hits the rate and clocks out both score 100. For the top of the distribution, look at promotion and retention at twelve months.
  • The slow burner. Someone who takes 55 days to ramp and then stays six years scores 66 and deserves better. Re-score the cohort at twelve months on survival alone.
  • Roles with no countable output. Where the work is judgement rather than volume there is no honest ramp measure. Score those on survival and reliability out of 70 and say so.
  • The line between hiring and onboarding. Two sources scoring differently under one supervisor is a sourcing signal. Every source scoring badly under that supervisor is not.

Where quality of hire is actually decided

The score is a thermometer. Almost everything that moves it happens before the person starts, which is why the metric belongs to hiring.

  1. The ad said what the job is. Wage, shift, site, in the first line. A week-one quit is usually leaving a job that was described differently, and no onboarding fixes a mismatch written into the ad.
  2. The criteria were fixed before you met anybody — what the shift and the licence require, weighted and written down: hiring criteria.
  3. The interview produced a record. An interview scorecard gives you something to correlate the ninety-day score against; without one you never learn which signals predicted survival.
  4. You checked something outside the room. A reference check aimed at attendance speaks to two of the three components, and a paid working interview tests the third before either side commits.
  5. The first ninety days were dated. The cadence on onboarding and the dated touches on first-90-days turnover are what the survival score measures.

Score the cohort, read the zeros one by one, and change one of those five things. That beats an index with four decimals.

Frequently asked questions

What is quality of hire?

It is a score an employer puts on the people it hired, computed the same way every time, to judge whether the hiring produced people worth hiring. There is no standard definition and no meaningful external benchmark, because every employer picks its own inputs. For hourly roles the computable version is ninety-day survival, time to full productivity, and reliability in the first ninety days.

How do you calculate quality of hire?

Score each hire out of 100: 50 points for still being employed on day 90, 30 for reaching the role's normal output inside your own target window, 20 for working the shifts they were scheduled for. A hire who left before day 90 scores zero. The cohort figure is the average across every hire including the zeros, reported beside the survival rate.

What is the standard quality of hire formula, and is it any good?

The formula in circulation averages a performance rating, a ramp-up score, an engagement score, a culture-fit score and hiring manager satisfaction. It is weak for two reasons: most of those inputs come from the same supervisor, so averaging disguises one opinion as arithmetic, and it cannot be computed until six or twelve months, by which point it changes nothing.

What is a good quality of hire score?

There is no published figure to compare against and we do not offer one. Use it against yourself: score this quarter's hires the same way as last quarter's and watch the direction. A score under 50 means the hire did not reach day 90, so a cohort average well under 50 means hires are leaving rather than underperforming.

How many hires do you need before quality of hire is a real number?

Roughly 25 to 30 per source before a comparison means anything. With 6 hires, an observed survival rate of 67% has a true range of about 22% to 96% — consistent with almost any conclusion. Below about 25 hires, read the cases and act only on differences so large you would not need a statistic to describe them.

How do you measure time to full productivity for an hourly role?

Pick the one countable thing the role produces per unit of time — picks per hour, routes per day, clients per shift — and define full productivity as sustaining the crew's rate across five consecutive shifts with no trainer present. Set the target window from the median of your own last ten hires, because what you are measuring is your training program.

Can you count absences in a quality-of-hire score?

Count unexcused absences and no-call/no-shows only. Approved time off, agreed swaps and anything granted as an accommodation come out before you score anybody. 29 CFR 825.220(c) states that employers cannot use the taking of FMLA leave as a negative factor in employment actions and that it cannot be counted under no-fault attendance policies; most new hires are not yet eligible under 29 CFR 825.110, but the habit matters.

Is quality of hire just 90-day retention?

Mostly, and that is the honest answer for frontline work. Survival carries half the score because a hire who leaves inside ninety days returns nothing regardless of how they performed. Ramp and reliability add what retention misses: two hires who both stayed differ if one reached the crew rate in three weeks and the other in eight.

How do you measure quality of hire by source?

Write the source on the hire record on day one and carry it into payroll, so at day 90 you can sort survivors by source. Then compute cost per surviving hire: what one hire from that source cost, divided by that source's ninety-day survival rate. That figure combines price and quality, which is what makes it capable of changing a budget.

Can you benchmark quality of hire against other employers?

No, and be suspicious of anyone selling a figure for it. Because each employer picks the inputs and the weights, a published quality-of-hire benchmark compares nothing to nothing. Cost per applicant is measured the same way everywhere; everything downstream of the application should be compared only with your own last quarter.

Does Boostpoint publish quality-of-hire data?

No. Our data ends at the application — we see applicants, not starts, and not who is still employed at day 90 — so we publish no quality-of-hire figure, no cost per hire and no retention rate. The one real number from our data here is the $9.83 median cost per applicant for warehouse and production roles, from 891 managed Meta campaigns.

Quality of hire only means something if you know where the hire came from.

We tag every campaign so applicants can be traced to a source, and report cost per applicant with the management fee inside it. Median across 891 managed campaigns: $13.88. Day 90 is yours to measure; the tag that makes it possible is ours to set up.

Book a Demo