Every data vendor has a green tick. It is the most reassuring pixel in the category and the least examined. What sits behind it is a question asked of a stranger’s mail server — and three common configurations make that question unanswerable.
What the check actually does
A verifier connects to the receiving mail server and walks an SMTP conversation as far as RCPT TO without sending a message, then reads the response code. Accepted is read as valid, rejected as invalid.
That is a reasonable inference about a cooperative server answering honestly. Three situations break it, and B2B lives inside all three.
The three that break it
| Configuration | What the server does | What the tick means |
|---|---|---|
| Catch-all | Accepts mail for every address at the domain, existing or not | Nothing about this person — every address returns success |
| Greylisting | Temporarily rejects unknown senders with a 4xx and expects a retry | A valid address can be recorded as unknown or undeliverable |
| Enterprise gateway | Detects probing and rate-limits, generalises or refuses to answer | The check is being actively defeated by design |
The catch-all case is the one that matters most, because it is common in exactly the companies worth selling to and because it fails silently. The verifier gets a real 250 and reports it correctly. The address is then discarded or rejected internally, and the bounce arrives hours or days later — after the send, against your domain reputation, which is the currency you can least afford to spend.
Why a better verifier does not fix it
Accuracy comparisons between verifiers show meaningful spread, and the leaders are well short of certainty. But the interesting part is not the headline number — it is where the residual error sits. It is not scattered evenly. It is concentrated in catch-alls, greylisting and gateways, which is to say concentrated in mid-market and enterprise B2B.
So the error is not merely present, it is correlated with the accounts you care about most. Buying a more accurate verifier moves the number without moving the shape.
The uncertainty is not spread evenly. It is thickest exactly where the deals are.
What a rep actually needs
A single word forces every case into one of two buckets, so it has to round. Two addresses that are not remotely the same object end up wearing the same tick:
- An address printed on the company's own contact page, read on a date you can check.
- An address assembled from first.last@ because four other people at that employer use that pattern.
The first is an observation. The second is an inference with a sample size. Both may well deliver, and a rep should treat them differently — which they can only do if the difference survives to the screen.
What we do instead
We publish this as a limit rather than a feature, on the evidence standard: we do not verify email addresses. Every address carries a score from 0 to 100 and the reason for that score. Published on a linkable page scores high. Built from an employer’s observed format scores lower and shows the sample count it was inferred from. Nothing is guessed and then labelled verified.
That is a worse-sounding claim and a more useful one. It also composes with the rest of the work: the reason a low score exists at all is that the alternative is spending domain reputation on a guess, and reputation is the constraint everything else runs on. At a spam-complaint budget of one in a thousand, “probably fine” is not a category you can afford to send to blind.