Because most scoring models reward visible activity, and your best buyers are often quiet. Someone who read three pages, never filled a form, and arrived from an AI answer looks cold to a scoring model. They are not cold. They are late in their research and nearly ready to talk.
This is the failure we see most often when a team turns on AI qualification. The model is not wrong exactly. It is measuring the wrong era of buyer behaviour.
So this is the framework we use when we build qualification into a site and a CRM. It is deliberately conservative, because the cost of a false negative in B2B is much higher than the cost of a false positive.
It reads the information attached to a lead and sorts it, faster and more consistently than a person would. It can enrich a company from a domain, summarise what the lead did on your site, classify the enquiry, and rank it against your ideal customer profile. That is genuinely useful work.
What it does not do is know things. It cannot tell that the person filling in the form is the champion who will drive the deal internally, or that the small company on the form is a subsidiary of a large one. It infers, and inference at speed is exactly what you want for sorting and exactly what you do not want for rejecting.
Keeping that line clear is most of the job. Automate the sorting. Keep a human on anything that ends a conversation.
Because much of the research happens before they reach you. In the study Webflow published alongside its Conf 2026 announcements in September 2026, the median company across 2,000 analyzed websites appeared in only 16% of the AI answers it would want to be part of, and was cited in those answers just 6% of the time. Buyers are getting shortlists from somewhere, and increasingly it is not your site.
The click data points the same way. The Pew Research Center reported on July 22, 2025 that when a Google AI summary appeared, users clicked a traditional search result in 8% of visits, compared with 15% when no summary appeared. Clicks on links inside the summary happened in only 1% of visits.
The practical consequence is that you see less of the journey than you used to. A lead who visits twice and books a call may have spent a fortnight comparing you against two competitors in a chat window you will never see. Scoring them on session count punishes them for research you cannot observe.
Firmographic fit first, because it is stable and checkable. Company size, industry, geography, and whether they run the technology your product depends on. These do not change between Tuesday and Thursday, and they map directly to whether you can serve this customer at all.
Next, intent that costs the buyer something. Requesting pricing, reading a comparison page, opening the security documentation, or bringing a second person into the conversation. Each of these takes real effort, which makes them harder to fake and more meaningful than a newsletter signup.
Third, what they said. The free text in an enquiry form is the richest signal you have and the least used, because it is hard to score by rule. This is the part where a language model genuinely helps. Summarising and classifying a hundred enquiries a week into problem type and urgency is work a model does well and a person does resentfully.
Raw page views, time on site, and email opens. Each is easy to inflate and easy to miss. A researcher building a comparison spreadsheet racks up page views. A ready buyer who already read about you elsewhere might view two pages. Scoring on volume rewards the wrong one.
We also throw out anything that penalises a free email domain by itself. Plenty of real founders enquire from a personal address, and plenty of noise arrives from corporate ones. Use it as one input, never as a filter.
And ignore lead source purity. In our work the source attributed to a lead is usually the last thing that happened, not the thing that mattered. Making qualification depend on a field you know is unreliable just moves the error somewhere less visible. Our piece on stopping form spam covers the separate problem of filtering junk, which is not the same as qualifying.
Through a defined interface, with the narrowest permissions that let it do its job. The Model Context Protocol has become the common way to do this. Its own documentation describes it as an open source standard for connecting AI applications to external systems, and compares it to a USB-C port for AI applications.
The safety comes from scoping, not from the protocol. Give the agent read access to the records it needs and write access only to the fields it owns, such as a score field and a summary field. It should not be able to change deal stages, delete records, or email anyone.
Log everything it writes, with the reasoning attached. When a salesperson disagrees with a score, you want to be able to read why the model gave it. A score with no explanation is a number nobody trusts, and untrusted automation gets ignored, which is the worst outcome of all.
Recommend. We feel strongly about this and it is the main place we push back on clients. Let the model rank, summarise, route and draft. Do not let it close a lead as unqualified without a person seeing it, at least until you have months of evidence.
The asymmetry is the whole argument. If the model wrongly marks a lead as hot, a salesperson wastes twenty minutes and learns something. If it wrongly marks your best fit customer as junk, that deal disappears and no report will ever show you it existed. One error is visible and cheap. The other is invisible and expensive.
A good middle path is a review queue. Everything the model would reject goes into a list somebody skims once a day. It takes a few minutes, and it is the only way to measure how often the model is wrong in the direction that hurts.
Score the past first. Run the model against last year's closed deals and see whether it would have ranked your actual customers highly. If your best five customers score in the bottom half, the model is measuring something other than fit, and no amount of tuning the threshold will save it.
Then run it in shadow mode. Let it score live leads without routing anything, and compare its judgement against what your team did. Two or three weeks of that will tell you more than any vendor benchmark, because it is measured on your pipeline rather than someone else's.
Only after that should it touch routing. And when it does, keep the shadow comparison running. Models drift as your market changes, and the thing that catches drift is a comparison you never turned off.
It supplies the signal. A form that asks one vague question gives a model nothing to work with. A form that asks what the person is trying to do, in their own words, gives it the most valuable field in the record. That is a design decision, not an automation one.
Platforms are moving in this direction too. Webflow announced Campaigns at its Conf 2026 event on September 2, 2026, a product for performance marketers that generates landing pages and variants, tracks conversions, and connects to HubSpot or Salesforce so teams can measure pipeline. The pattern of site and CRM being one system rather than two is becoming the default.
We would still start with the form. Getting one honest open question in front of a buyer improves qualification more than any scoring model layered on top of a bad form. Our notes on form design cover how to ask without adding friction.
Take your last fifty closed won deals and your last fifty rejected leads. Ask whether your current rules would sort them correctly. That exercise takes an afternoon and it usually settles the argument about whether you need AI here at all, or just better fields.
If the sorting is genuinely hard, start with summarising and ranking, keep a human on rejection, and run it in shadow mode for a month. If you want help wiring the site, the form and the CRM into something that actually feeds a model good data, we are happy to walk through it with you at phoenix.studio.
Tell us where you want to go. We'll tell you how we'd get you there.