โ† BlogBlog

ICP scoring for B2B lead qualification: the criteria that predict fit

Most ideal customer profile (ICP) scoring rubrics are a list of things the team could measure, weighted by how strongly someone argued in the meeting where the rubric was written. That is why reps quietly stop reading the score. A rubric earns attention when every criterion on it has been checked against your own closed-won list, and when the ones that fail the check get deleted rather than demoted to one point. Here is how we would build an ICP scoring rubric for B2B lead qualification from scratch: the arithmetic, the narrow place where a model helps, and the much larger place where a lookup table is the honest answer.

1,000Leads a model needs first
120Conversions it needs too
0.989Best AUC in the cited study
SecondsMatch latency at millions/day

The score is a budget for attention

The constraint in B2B sales is almost never lead quality. It is hours. Your team can hold some fixed number of first conversations this week, more leads than that will arrive, and something has to decide the order. That is the whole job of a score, and keeping it that narrow makes every later decision easier.

So, a working definition of ICP scoring criteria: the small set of company attributes you can observe before anyone talks to the account, which separate the deals you won from the deals you lost. Anything you only learn on the call is context, not a criterion, because by then the routing has already happened.

Two things follow. Only precision at the top matters, so if you work forty leads a week out of three hundred, nobody should be arguing about whether a middling account scores 58 or 61. And fit does not belong in the same number as behaviour: behaviour rises with your own ad spend, so a combined score quietly rewards you for advertising to people who were never going to buy.

The one test that tells you a criterion is real

Scoring a prospective customer against ICP criteria only pays off when the criteria separate winners from losers. You can check that, and the check takes an afternoon with a CSV export rather than a workshop.

Pull every closed-won and closed-lost deal from the last twelve to twenty-four months. For each criterion, compute the win rate among deals that carried it and the win rate among deals that did not, then divide. A ratio near 1.0 means the criterion knows nothing and is adding noise to every score it touches. Below 1.0 it is inverted, and if it is already in your rubric you have been sorting the queue backwards. Our working line is 1.3.

One caveat before you act on the numbers. If you closed forty deals last year, a criterion that splits them 12 against 28 tells you very little, and a percentage to one decimal place does not fix that. Write the sample size next to the ratio, treat anything under about thirty deals per side as a hint, and re-run the test when the data catches up.

What common ICP scoring criteria are really measuring
CriterionWhat it proxies forWhere it lies to youVerdict
Employee headcountBudget and internal process complexityA 200-person agency and a 200-person manufacturer buy nothing alikeKeep, but as bands and never as a raw number
Industry code (SIC / NAICS)Almost nothing on its ownSelf-reported, often the parent company's, and one code covers wildly different operationsKeep only if your wins cluster into two or three codes
Detected tech stackIntegration effort and buying maturityInferred from page tags, stale within a quarter, blind to anything server-sideKeep for one or two decisive tools, drop the rest
Recent funding roundWillingness to spend on something soonIt says they will spend, not that they will spend with youTiebreaker weight at most
Contact's job titleWhether this person can start a purchaseTitles are not standardised; "Head of Ops" is five different jobsMap to a role band, never match on the string
Site visits and email opensAttention, which is not intentRises with your own ad spend, so it scores your campaign rather than the accountSeparate axis, keep it out of the fit score
Geography and time zoneDelivery cost and contract frictionRarely predictive on its own, occasionally an absolute blockerUsually belongs in the knockout gate instead
Hiring for the function you replaceThe problem is funded and someone owns itLives in job-post text, so most enrichment vendors simply do not carry itOften the strongest signal available, and the hardest to buy

The pattern in that table is worth sitting with. Criteria that predict well tend to be the ones nobody sells you as a tidy data field, and criteria that dominate real rubrics tend to be whatever already arrives in the enrichment payload. That is not a conspiracy, it is gravity: free fields get added, each takes a slice of the weight, and the one signal that meant something gets outvoted by six that came bundled.

Some criteria are a door, not points

A knockout criterion is a binary condition that ends the evaluation regardless of everything else. The company sits in a regulated segment you cannot serve. It is a current customer, or a competitor. It contracts in a jurisdiction where the paperwork takes four months and your deal cycle is six weeks.

These cannot be written as negative points, and the reason is arithmetic. Take the rubric in the next section, which tops out at 28. Score a hard blocker as minus 10 and a well-matched but unservable account still lands at 14, comfortably above your median. So the order is fixed: the gate runs first and returns yes or no, the weighted rubric runs second and returns a number, behaviour is read third and adjusts urgency rather than fit.

Keep the gate short. Three to five knockouts is plenty, because every rule you add is a rule you will defend the week a rep brings an exception with a real logo attached. And log what the gate rejects, with the reason. That rejected pile is the cheapest ICP research you will run: when the same excluded shape keeps reappearing and occasionally closes through the side door, your ICP is wrong and the gate is what told you.

Three weights are all your data can carry

Use three weights and nothing finer. Decisive criteria get 3, supporting criteria get 2, tiebreakers get 1. The reason is that your evidence cannot tell 12 from 15. If the lift test gave 1.9 on one criterion and 2.1 on another across sixty deals, those are the same number wearing different hats. A rubric that encodes a distinction its data cannot support looks more rigorous and performs worse, because it moves accounts across the threshold on noise.

Cap each category as well. If firmographics can contribute at most a third of the maximum, no single enrichment outage silently rewrites your queue. Coverage gaps correlate with company type, so an uncapped field that goes missing for one segment removes that segment from your pipeline and nobody notices.

Do not force the weights to sum to 100. That habit comes from slide decks and it taxes you every time you add a criterion. And watch the temptation to keep adding criteria until the score agrees with your gut on the ten accounts everyone remembers: those ten are memorable because they were unusual. Five decisive criteria almost always means one is restating another, headcount band with revenue band being the pair we see doubled up most.

A worked example, with the arithmetic shown

Here is a complete B2B lead qualification framework built on ICP criteria, small enough to hold in your head. Seven criteria, each scored 0, 1 or 2 for no match, partial match, clean match. Contribution is weight times match, so the maximum is 28.

A seven-criterion rubric, applied to two leads
CriterionWeightLead ALead B
Hiring for the function you replace3match 2 โ†’ 6match 0 โ†’ 0
Headcount in the 50-500 band3match 2 โ†’ 6match 1 โ†’ 3
Contact sits in a role band that can start a purchase2match 2 โ†’ 4match 1 โ†’ 2
Industry inside your top-three cluster2match 2 โ†’ 4match 0 โ†’ 0
One of two decisive tools detected2match 1 โ†’ 2match 2 โ†’ 4
Funded in the last 18 months1match 0 โ†’ 0match 2 โ†’ 2
Working-hours overlap with your team1match 2 โ†’ 2match 2 โ†’ 2
Total (max 28)2413
Where Lead A's 24 points come fromWorked example - the rubric and the arithmetic are in the table above
Hiring for the functionweight 3 ร— match 2
6
Headcount bandweight 3 ร— match 2
6
Contact role bandweight 2 ร— match 2
4
Industry clusterweight 2 ร— match 2
4
Decisive tool detectedweight 2 ร— match 1
2
Working-hours overlapweight 1 ร— match 2
2
Recent fundingweight 1 ร— match 0
0

Normalised, Lead A scores 24 of 28, or 86%. Lead B scores 13 of 28, or 46%. Look at where Lead B's points came from: a detected tool and a funding round, the two criteria the table above called weak. Its score is built out of things that are easy to observe, which is exactly the profile of a lead that reads well in the CRM and dies on the call.

Set the thresholds by capacity rather than by taste. If forty first conversations fit in a week and three hundred leads arrive, your A band is the top thirteen percent of the distribution, whatever raw score that turns out to be. Round thresholds like 80 mean nothing outside the rubric that produced them, and they drift the moment you add a criterion.

Do the payoff in arithmetic once, with your own figures. Say you close 5% of everything entering the pipeline and your top band closes at 15%. The lift is 3, so forty calls into the top band produce roughly six wins where forty undifferentiated calls produce two. Nothing about the leads changed. And Lead B is not a bad lead here, it is a lead you call in week three.

AI-powered fit scoring beats a lookup table in exactly two cases

AI-powered fit scoring for an ICP lead list is sold as one product and is really two jobs stapled together: reading the world, and doing the sum. Only the first benefits from a model, and separating them saves most teams a year of expensive disappointment.

The first case where a model genuinely wins is labelled history plus criteria that interact in ways a linear rubric cannot express. Headcount matters more in one industry than another, and nobody can write that down. A model finds it. The entry ticket is data, and the vendors are unusually specific about how much: Salesforce's documented requirements for Sales Cloud Einstein are at least 1,000 leads created in the last 200 days with at least 120 converted, and for opportunity scoring at least 200 closed-won plus 200 closed-lost inside 24 months (retrieved Aug 4, 2026). Below that it falls back to a global model built from anonymised data across other Salesforce customers, so your leads get scored on other companies' patterns. If your ICP is unusual, that is precisely the wrong prior.

With enough data the results are not in doubt. A 2025 study in Frontiers in Artificial Intelligence built a B2B lead scoring model on a CRM export of 16,600 records and 22 fields after cleaning, compared fifteen classification algorithms, and reported gradient boosting at AUC 0.9891 with recall 0.9586 and precision 0.9106 (retrieved Aug 4, 2026). Read the feature list before you believe a number like that, though. The reported top predictors include lead source and lead classification, which are fine, and "reason for status", a CRM field somebody fills in after deciding what happened to the lead. Features of that shape are the classic way an offline evaluation comes back brilliant and the live model does nothing, because on a new lead the field is empty. For every feature, ask whether it could have been written after the outcome was known. If yes, it leaves the training set.

The second case is unstructured input, and this is where AI genuinely belongs in the pipeline. The strongest signal in the criteria table was whether the company is hiring for the function your product replaces, and that signal exists as prose: a careers page, a job post, sometimes a line in a changelog. No vendor sells it as a clean field because extracting it requires reading. A model reads twenty pages and returns three structured values, and the score is still a lookup table. That leaves you an auditable rubric with a reader in front of it, which is far easier to defend than a black box emitting a number. It is the same line we draw between agents and fixed workflows.

Three ways to score the same list
Hand rubric in a sheetRubric plus enrichmentModel-scored
What it needs from youYour closed-won export and an afternoonA vendor subscription and a field mapping somebody maintainsLabelled history at the volumes above, plus an owner
Where the cost sitsRep minutes spent looking things upPer-record fees, plus the coverage gaps you fill by hand anywaySetup, then the standing job of watching for drift
What it gets wrongWhatever the person who wrote it believesConfidently wrong values nobody re-checks, because they arrive pre-filledIt learns your past routing, so it repeats last year's bias with better manners
How you debug a bad scoreRead the rowRead the row and the vendor's field historyFeature attributions, and a careful look for leakage
When it is the right callUnder a few hundred deals a year, or criteria that fit on one screenVolume high enough that manual lookup is the bottleneckEnough history, interacting criteria, and someone accountable for the model

Our position, from building both: most fit scoring is not a model and does not want to be one. The rubric is legible, a rep can argue with it, and when it is wrong you see why in four seconds. If you are weighing an enrichment subscription against building the reader in-house, the build versus buy questions transfer without modification.

Calibrate against closed-won, then watch for drift

An ICP scoring methodology that has never met your history is a formatting exercise. Calibration fixes that, and it is mechanical. Export every deal from the last twelve months with its outcome. Reconstruct the criteria as they were when the deal entered rather than as they look now, because a company with eighty employees then has two hundred today. Score every deal with the new rubric, then compare the score distribution among wins against the distribution among losses.

That comparison is the verdict. If wins have a median of 71 and losses 66, with the bulk of both piles sitting on top of each other, the rubric does not work no matter how sensible each criterion sounded. You want visible separation, and you want to know where the overlap sits, because the overlap is where the score should hand over to a human instead of pretending.

The trap is survivorship. Your CRM only holds deals somebody worked, so calibrating on it tells you how well the rubric ranks leads you already liked. Score a sample of the leads nobody called, and if you have no record of those, start keeping one: without it you can measure false positives forever and never see a false negative.

Then drift, and the signals are specific. Reps start overriding the A band, which is the earliest warning and arrives as grumbling rather than as a metric. The top band's close rate slides toward your overall average, meaning the rubric has stopped separating anything. Or a criterion's hit rate jumps overnight, which almost always means a vendor changed a taxonomy rather than that your market moved in a day.

Cadence that has held up for us: re-run the lift test quarterly, rewrite the rubric annually. More often and you fit to noise, less often and you qualify this year's leads against a market that has moved. The same failure pattern runs through automation projects generally, which we wrote about in why AI deployments fail: the rules are rarely the problem, the unowned input data is.

The week-one checklist

Export the deals. Twelve to twenty-four months, won and lost, with the attributes you had at the time. Usually a two-hour job that has been waiting a year.

Run the lift test on every candidate criterion. Win rate with, win rate without, divide. Write the ratio and the sample size side by side. Anything near 1.0 stays out, whoever suggested it.

Write the knockout gate before the rubric. Three to five absolute conditions, each one something you would genuinely refuse to work, plus somewhere to log what it rejects.

Assign weights of 3, 2 and 1 only. Cap any single category at roughly a third of the maximum, and resist the eighth criterion you want to add because seven feels incomplete.

Re-score the history and look at the overlap. Wins against losses on one axis. If they sit on top of each other, go back to the lift test rather than adding criteria.

Set thresholds from capacity, not from round numbers. Count how many first conversations really fit in a week and take that share off the top.

Pick one criterion that needs reading and automate that specifically. Job posts are the usual first choice. One extractor, one field, measured against thirty hand-labelled accounts before you trust it.

Put a date in the calendar one quarter out. A rubric nobody has an appointment with will still be running unchanged in 2028.

What breaks when the matcher runs at volume

One place we learned this expensively was not a sales system at all. Indigo is a real-time feed tracker: it watches Upwork's RSS feeds, matches every new posting against the personal filter sets of thousands of subscribers, and pushes the matches to Telegram within seconds. One posting can match hundreds of people at once, which is how a modest stream of source items becomes millions of delivered messages a day. It is the same problem as scoring an inbound list against per-recipient criteria, about a thousand times faster, and that speed makes three things obvious.

Matching has to be decoupled from ingestion. Indigo polls the feeds on a schedule and everything downstream reads from a queue, so a slow or failing source deepens the queue instead of stalling delivery. The lead-scoring version of that failure is the enrichment API that times out and takes your routing with it. Same fix: score what you have, mark the field unknown, re-score when the data lands.

Duplicates have to be impossible, not unlikely. Feeds return overlapping windows on every poll, and a second delivery is indistinguishable from a second genuine match, so every item is deduplicated against recently seen IDs and processed exactly once. The scoring equivalent is one account entering through three forms in a week: without a dedup key you score it three times and three reps call the same person.

And a false positive is paid for by a human. A wrong match at fan-out scale does not cost one person one second, it costs everyone it reached, and the trust it burns does not come back on the next message. A lead queue behaves the same way: a rep who works two bad A-band accounts in a row stops believing the score, and after that you own a rubric nobody reads. That is the argument for holding the top band tight rather than generous.

If you want a second pair of eyes on a rubric you already have, that is the sort of thing we look at in a free audit - bring the criteria list and last year's closed-won export. But the test itself is yours, it takes an afternoon, and you do not need us to run it.

FAQ

How many criteria should an ICP scoring rubric have?

Five to eight. Fewer and you cannot separate the middle of the distribution; more and each criterion's weight gets so thin that the strong signals stop deciding anything. If a ninth criterion feels necessary, check whether it is restating one you already have.

What is the difference between ICP scoring and lead scoring?

ICP scoring measures fit: is this the kind of company you win with, which is a slow-moving property of the account. Lead scoring in the usual sense measures behaviour: did this person visit pricing, open the sequence, book time. Keep them as two numbers and route on both, because a combined score hides which half moved.

Do we need AI to score fit?

Usually not for the scoring itself. Salesforce documents 1,000 leads and 120 conversions before Einstein builds a custom model for you, and most teams sit under that, where a weighted rubric is not a compromise but the correct tool. Where a model does earn its place is reading the unstructured sources, mostly job posts and careers pages, and turning them into fields your rubric can use.

How often should we recalibrate the rubric?

Re-run the lift test quarterly on a rolling four quarters, and rewrite the rubric once a year or whenever you change segment or pricing. More often than that and you are fitting to noise; less often and you are qualifying new leads against a market that has already moved.

Changelog
  • 4 August 2026Published.
Free process audit

See what this would look like in your operations.

Get in touch

30 minutes ยท we map your 3 best automation opportunities ยท no obligation