Ask a data vendor how they know a company's revenue and the honest ones will tell you they don't, not really. They know how many employees show up on LinkedIn, they know the square footage of the building if it's public record, they run that through a model, and the model spits out a bucket. Under $1 million. One to five million. Five to ten. It's less a measurement than an astrology chart built out of proxies, and everyone in GTM has quietly agreed to treat it like a fact because the alternative, for a long time, didn't exist.
The tell is the match rate. Push a file of small businesses through a typical firmographic vendor and you'll get a modeled revenue band back for something like 20 to 30 percent of the list. The other 70 to 80 percent just doesn't have one. Not "we don't know precisely," but nothing. A blank cell where a business's size used to be inferred.
That gap is the whole reason SMB qualification has stayed primitive for so long. You can't route, score, or prioritize an account you can't size, so teams either burn reps' time chasing accounts with no signal attached, or they quietly drop the SMB segment entirely and go build a mid-market motion instead, where at least somebody filed a 10-K.
What we actually tested
We took a file of 18,943 ecommerce merchants and matched it against card transaction data instead of a modeled revenue estimate. Not employee counts run through a regression. Actual money that moved through actual point-of-sale and payment systems: revenue, transaction counts, refunds, and growth, all observed rather than inferred.
18,943 ecommerce accounts sourced through LeadGenius, matched against Enigma's card transaction dataset. Fields tested: 12-month card revenue, 1/3/12-month growth, transaction counts, and card refunds. All figures below are drawn directly from that file.
The first number is the one that matters most, because it's the number every revenue range vendor has trained the market to expect will be low.
61.5 percent of accounts resolved to an actual 12-month revenue figure. Not a range. A number, the kind you can subtract, divide, and rank: "$61,361." That alone is roughly double what the industry treats as a normal match rate on a modeled band, and it's before you get to the part that a revenue range was never built to hold in the first place.
The bucket was hiding the interesting part
Here's the thing a range obscures that transaction data can't: two businesses can sit in the exact same bucket and be having completely different years. Take two merchants both parked in "$1M to $5M." One of them is up 40 percent and hiring. The other is down 30 percent and quietly winding toward a fire sale on Facebook Marketplace. A revenue-range file can't tell them apart. It was never designed to. It answers a size question, and growth is a different question entirely.
Notice what the granularity buys you immediately. Almost half of these accounts fall under $100,000 in trailing revenue, which is a real, meaningful distinction for anyone qualifying leads, and it's a distinction that gets flattened into a single undifferentiated "small" tag by most vendors. The businesses doing $8,000 a year and the ones doing $95,000 a year look identical to a range file. They should not look identical to you.
Growth is the part a range was never designed to see
Now the number that actually changes how you'd prioritize a book of accounts. Among the accounts where we could observe a trailing 12-month trend, growth splits almost exactly down the middle, and not in the reassuring way.
"A revenue range answers a size question. Transaction data answers a health question. Those have never been the same question, we just priced them as if they were."
Refunds: the signal nobody was collecting
If growth tells you direction, refunds tell you something closer to friction, and it's a field that essentially doesn't exist in a traditional firmographic file. There's no "returns" column on a modeled revenue estimate, because a model built off employee count and NAICS code has no way to observe it. Card data does, because refunds are just negative transactions running through the same rails as the sales.
Read on its own, refund volume is mostly noise: returns are a normal cost of doing business in ecommerce, especially in apparel, where a chunk of this file sits. Read next to the growth field, it starts doing real work. An account with rising refunds and declining revenue is a very different account than one with rising refunds and 30 percent growth, and a range file can't distinguish either pairing because it only ever had the one dimension to begin with.
The sample itself argues against a single "SMB" bucket
One more thing the file makes hard to ignore: "SMB ecommerce" isn't a category, it's a label stretched over five hundred and forty distinct NAICS descriptions. Apparel retailers, sporting goods, cosmetics, manufacturers, jewelers, snack bars, novelty shops. A qualification model built for one of those behaves badly applied blindly to the rest, and a coarse sector code can't tell you which one you're looking at with any precision.
What this replaces, concretely
| Revenue-range approach | Transaction-data approach |
|---|---|
| 20–30% of file returns a modeled band | 61.5% returns actual 12-month revenue |
| "$1M–$5M" for every account in the bucket | Exact figure, rankable and filterable to the dollar |
| No growth signal; size is treated as static | 1/3/12-month growth trend on 32.1% of the file |
| No visibility into refunds or churn risk | Refund activity observed directly from transactions |
| Modeled from employee count, web scrape, or NAICS proxy | Observed from actual card revenue moving through the business |
What this means if you own SMB qualification
None of this changes the first step of an SMB motion, which is still finding the account and confirming it exists. What it changes is everything that happens after that: whether you can tell a growing account from a dying one before a rep spends a cycle finding out the hard way, whether refund activity gets factored into a health score at all, and whether "SMB" gets treated as one segment or the five hundred and forty different ones it actually is.
The practical move isn't complicated. Where your current file returns a revenue band, ask what it would take to get a real transaction-level figure instead, and treat that gap the way GTM teams eventually treated missing email addresses: not a permanent limitation of the data, but a specific problem with a specific fix.
Qualification built on what a business actually earns, not the bucket it was guessed into.
LeadGenius pairs live-sourced SMB and ecommerce contact data with transaction-level revenue, growth, and risk signal, so your qualification model is scoring real accounts instead of modeled ranges.

.avif)

