Setting up lead scoring: a model that works without mountains of data
Points models fail in small companies at one of two extremes: either they are so fine-grained that nobody understands them, or so coarse that they separate nothing. A model with eight criteria on two axes works without mountains of data.
In short
- Lead scoring pays off from around 30 enquiries a month. Below that it costs more upkeep than it saves time.
- Two axes instead of one number: fit (who is this) and behaviour (what is the person doing). A single score blurs both.
- Eight criteria are enough — four per axis. Every additional one makes the model harder to explain without improving it.
- Negative points matter more than positive ones: they keep the wrong contacts out of sales.
The usual mistake is the one big number: 0 to 100, everything flows in, and above 70 the contact goes to sales. The problem is not the arithmetic but that two completely different things are being added together.
A person who fits the audience perfectly and does nothing, and one who does not fit and clicks everything, can end up with the same score — and need entirely different handling.
Two axes instead of one number
Axis 1 – Fit
Who is this person, regardless of their behaviour? Industry, company size, role, region. These details rarely change and answer the question: do we want this customer at all?
Axis 2 – Behaviour
What has the person done, and how recently? Filled in a form, visited the pricing page, replied to an email. These values change daily and answer the question: is now the right moment?
| Behaviour low | Behaviour high | |
|---|---|---|
| Fit high | watch, approach deliberately | straight to sales |
| Fit low | general communication | check — often applicants, competitors or students |
The bottom-right cell is the reason for the second axis. Form a single number and exactly those contacts go to sales — who lose faith in the model after the third such case.
The eight criteria
Fit (max. 100)
Industry fits: +30
Company size in target range: +25
Role close to the decision: +30
Region we serve: +15
All four come from the form or a quick look-up — no behaviour, no guesswork.
Behaviour (max. 100)
Enquiry form completed: +40
Pricing or offering page visited: +25
Replied to an email: +25
Second visit within 14 days: +10
Behaviour ages: after 30 days without activity the value halves, after 90 days it drops to zero.
Negative points
Disposable address: −60
Competitor by domain: −100
Job application intent visible: −100
Unsubscribed or objected: −100
The most important group. It stops attention going to contacts where no business is possible.
The threshold
Handover to sales: fit ≥ 60 and behaviour ≥ 40.
Two conditions, not one sum. That systematically removes the bottom-right cell.
Worth knowing
Ageing the behaviour score is the part most often missing — and the one without which a model becomes useless after a year.
Without ageing, active contacts accumulate points over months until nearly all sit above the threshold. The model then separates nothing; it merely confirms that someone was once interested. A behaviour score with no expiry is not a signal but an archive.
When scoring is worth it
Three conditions. Miss one and the effort is not justified.
- At least around 30 enquiries a month. Below that, one person can look at every contact individually — and does it better than any model.
- Visibly varying quality. If nearly all enquiries fit, there is nothing to sort.
- The data exists at all. Industry and role have to be asked for in the form or findable. A model computing on empty fields produces random values.
Introducing it in four weeks
| Week | What happens |
|---|---|
| 1 | Sort the last 50 enquiries by hand — good, medium, poor. No model, just judgement. |
| 2 | Set criteria and points so they reproduce that sorting as closely as possible. |
| 3 | Run the model in the background without it triggering anything. Note the discrepancies. |
| 4 | Adjust the points, then switch it on — with one person reading along for the first two weeks. |
Week one is the decisive step and gets skipped readily. Without a human-made reference there is no measure of whether the model is any good.
After going live, the same thing happens regularly: sales receives three contacts that obviously do not fit and declares the model useless. Usually it comes down to a single criterion weighted too heavily — often "visited the pricing page", which competitors and applicants also do.
So build a feedback route into the first weeks: on every handover, sales can flag "does not fit" with one click and a reason. After twenty such reports you know exactly which criterion weighs too much — and that is worth more than any theoretical tuning.
Help me set up a simple lead scoring model. Our situation: - Enquiries per month: [number] - What we offer: [offering] - Ideal customer: [industry, size, role, region] - Who definitely does not fit: [exclusions] - Details we collect in the form: [fields] - Behaviour we can measure: [e.g. page visits, form, email reply] - Typical time to close: [details] Tasks: 1. First tell me whether scoring is worth it at all at our enquiry volume. If not, say so clearly and name the simpler alternative. 2. Propose four criteria each for fit and behaviour, with points totalling 100 per axis. Use only details we actually have, according to my list. 3. Name negative criteria and their points. 4. Set an ageing rule for the behaviour score, matched to our cycle length. 5. Name two thresholds (fit and behaviour) for the handover and justify them. 6. Describe how we check in week 3 whether the model agrees with our own judgement. Do not invent industry benchmarks.
In closing
A points model is not an arithmetic exercise but a written-down opinion about who your customer is. That is exactly why it works with eight criteria and without mountains of data.
What it needs: two separate axes, negative points, an ageing behaviour score — and one week in which humans create the reference the model has to be measured against.
Common questions
When is lead scoring worth it?
From around 30 enquiries a month, if their quality visibly varies and the necessary details — industry, size, role — are collected at all. Below that, one person can look at every contact individually and does it better than any model.
Why two scores instead of one?
Because fit and behaviour answer different questions: "do we want this customer?" and "is now the right moment?". A single total blurs both — a competitor clicking a lot reaches the same number as a perfectly fitting quiet prospect. Two thresholds instead of one sum resolve that.
How many criteria should a model have?
Eight — four per axis — plus three or four negative ones. Every additional criterion makes the model harder to explain without measurably improving how well it separates. A model nobody in sales can explain does not get used.
Why do behaviour points have to age?
Because without ageing, active contacts accumulate points over months until nearly all sit above the threshold. The model then separates nothing and merely shows that someone was once interested. Common practice is halving after 30 days without activity and resetting to zero after 90.
What if sales does not trust the model?
Build in a feedback route: on every handover, "does not fit" can be flagged with one click and a reason. After about twenty reports it usually emerges that a single criterion weighs too heavily — often "visited the pricing page", which competitors and applicants also do.
Marketing that sets itself up
The Studio Engine beta is live. Claim your spot and help shape it from the start.
Join the beta →