Scoring B2B leads with an LLM, and where it fell over

I built this for my own pipeline before I sold it to anyone. That means I know exactly where it breaks.

Anadi8 min read

Answers“how to score b2b leads with an llm”

I built a lead scoring system for my own outreach because the scores I was getting out of the box were not usable. It is called the ANTA Lead Intelligence Agent, it runs in production, and I used it every week on my own pipeline before I ever sold it to anyone.

That last part is the reason this post is worth reading. I am not describing a client engagement from a safe distance, where the failures are diplomatically described as learnings. I am describing the thing I depend on for my own pipeline, which means I know exactly where it falls over — and I have to keep using it anyway.

The problem with off-the-shelf lead scores

Generic scoring products weight what they know about everyone: headcount, industry code, revenue band, technologies detected, funding events. Those are real signals. They are also the same signals every other vendor's customer is scoring on, which means acting on them puts you in the same queue as everyone else, contacting the same companies in the same week, with the same trigger in your opening line.

Worse, for my business they were weighting the wrong things. A generic scorer loves a recent funding round. But a company that just raised is often in the middle of hiring a team to do internally exactly what I would be proposing to do for them — the money is a signal of competition, not of demand. Meanwhile a bootstrapped 20-person company that just posted a job listing for an accounts payable admin whose whole day is processing documents is, for me, a far better prospect. No generic scorer will ever rank it that way, because that judgement is specific to what I sell.

Fig. 01Two ways to split the same account list
What a database answers

Firmographics

  • Headcount band, industry code, revenue, geography
  • Detected technologies, funding events, growth rate
  • Cheap, structured, sortable, available to every competitor
  • Identical answers for two companies with opposite problems
What decides the send

Situation

  • What their open roles imply they are struggling to do manually
  • What their site claims they do vs. what they are hiring for
  • Whether they just funded an internal team to do my job
  • Whether there is a person whose week I can visibly give back

VerdictUse firmographics to build the list. Never use them to rank it.

Neither column is wrong. The left is what a database can answer and the right is what actually decides whether I should send an email, and no amount of tuning the left produces the right.

That gap is the entire thesis of the tool. The firmographics are table stakes and a data vendor can supply them. The judgement about what those facts mean for my offer is the part that has to be encoded, and it turns out to be a language problem rather than a numeric one. Once you accept that, the choice of mechanism follows: you do not want a weighted sum, you want something that can read.

What actually replaced it

Fig. 02Pipeline, one run
  1. 01

    Source

    Apollo, coarse firmographic filters only

  2. 02

    Enrich

    Job postings, site copy, hiring signals

    thin context → flagged, not guessed

  3. 03

    Score

    Written ICP rubric + anti-profile, structured out

  4. 04

    Draft

    First touch, from the scoring evidence only

  5. 05

    Sequence

    Lemlist send, replies synced to HubSpot

    nothing sends unreviewed

A representative run is 312 accounts in about 41 seconds. The expensive stage is 03; everything before it exists to make sure stage 03 is reading something worth reading, and everything after it exists to make sure a human still owns the send.

The critical design decision is stage three, and specifically that the model must return why.

A bare number is unauditable. You cannot correct it, you cannot learn from it, and you certainly cannot defend it to a client who asks why their competitor scored higher. A number attached to "this company posted a role whose described duties are document intake and manual data entry" is a claim I can check in ten seconds and reject when it is wrong. That is the whole difference between a score I act on and a score I ignore.

So the output is not a number. It is a record:

Fig. 03Anatomy of one scored account
scoreinteger, 0–100
reasonone paragraph
Written against the ICP definitionMust cite the evidence below
evidencearray of quoted spans
Sourced from the enrichment payloadNothing here that was not passed in
disqualifyingrequired field, may be empty
The score is the least important field. The two that changed my behaviour are the evidence list — which makes the score checkable — and the disqualifying field, which forces the model to have looked for reasons against.

The rubric itself is a written document, not a formula — a description of who is a good fit, in prose, including an explicit anti-profile of who to reject and why. That document is the actual asset. The code around it has been rewritten twice; the rubric has been edited about thirty times, and every edit came from reading a score I disagreed with and working out what the written definition failed to say.

Where it fell over

Three failure modes, all of which showed up on real accounts, all still partially present.

It was agreeable

Early on, the scores clustered high. Ask a model to evaluate fit and it will find a way to argue for fit, because that is the shape of the request — you have asked "is this a good prospect?" and the cooperative answer is yes. On a run of 312 accounts, a scorer that says yes to most of them has sorted nothing.

Two things helped. Writing an explicit anti-profile — an actual list of who is not a customer and why, not a negation of the good-fit list — and forcing the output into a fixed structure with a required disqualifying-evidence field, so the model has to go looking for reasons against rather than only reasons for.

The scores still skew generous. I read the reasoning, not the number, and I treat the ranking as an ordering rather than as an absolute. That is a real limitation, not a solved problem.

It hallucinated plausible context

Given a thin company page, it would generate a confident description of what the company does that was mostly inference. On a scoring pass that is tolerable, because a wrong inference produces a wrong rank and a wrong rank costs me one skipped account.

On an outreach draft it is lethal. A cold email containing a specific and wrong claim about someone's business is worse than no email — it is an active negative signal about the sender, and it is the exact thing that makes people hate this category of tooling.

The fix was structural rather than a prompt tweak, and this is the part I would transplant into any similar system: the draft may only reference evidence that was actually passed into it. Anything not present in the source context does not get to appear in the email. The generation stage does not have access to the model's general knowledge of the company; it has access to an array of quoted spans, and if that array is thin the email is generic — or the account does not get one.

The unstructured signal is often just missing

The job-posting angle works beautifully when there is a job posting. Plenty of good-fit companies are not hiring this month, and for those the tool has little more to work with than the firmographics I already said were not enough.

I have not solved this. What I do instead is mark those accounts as low-confidence rather than low-fit, which are genuinely different things and which a single score collapses into one. Low-fit means I looked and it is not a match; low-confidence means I could not see enough to say. They belong in different queues — the second one is a research task, not a rejection.

| Failure mode | What it looked like | What actually fixed it | What is still true | | --- | --- | --- | --- | | Agreeableness | Scores clustered 70–90 | Anti-profile + required disqualifying field | Still skews generous; read the reason, not the number | | Hallucinated context | Confident, invented company descriptions | Drafts restricted to passed-in evidence | Thin accounts produce bland drafts | | Missing signal | Nothing to read but firmographics | Split low-confidence from low-fit | Unsolved; those accounts need a human |

What it is worth

The published numbers for the current version: run #2,417, 312 accounts scored in 41 seconds, reply rate 11.4%.

The reply rate is the only one of those that means anything, and I want to be careful about what it means. It is my reply rate, on my list, in my category, written in my voice, on a list I curated the rubric for over months. It is not a benchmark and it will not transfer. Anyone quoting a reply-rate number at you as a product claim is quoting a number produced under conditions you do not share.

What does transfer is the shape of the thing:

  • Use the structured data vendor for what it is good at — building the list — and nothing else.
  • Score against a written definition, including an explicit definition of who to reject.
  • Require the model to state its evidence, and make that evidence the only thing downstream generation may use.
  • Separate "not a fit" from "could not tell".
  • Keep a human on the send. Nothing in the ANTA pipeline sends unreviewed, and the day I automate that away is the day this stops being something I would defend in public.

Would I build it again

Yes, and I would build it in the same order: rubric first, scoring second, drafting third, sending last. The temptation is to build the pipeline first because the pipeline is the fun part, and then discover that the rubric — the part that is just writing — is what decides whether any of it works.

If you are considering this for your own pipeline, the honest test is whether you can write two pages describing who your good-fit customer is and who your anti-customer is, with reasons, without using the words "growing", "innovative" or "forward-thinking". If you can, this system will work for you. If you cannot, no scoring tool will help, because there is nothing to score against yet.

02 / Contact

You already know what to build.

Send the two-paragraph version. You'll get a real technical response, not a calendar link and a discovery questionnaire.