ReplyLabs
FeaturesPricingCompareFAQUse casesBlogHelpSetup
Sign inGet started free
Get started

Product

  • Install
  • Features
  • Pricing
  • Compare
  • Roadmap

Resources

  • Use cases
  • Blog
  • Glossary
  • Cost calculator

Support

  • Setup Guide
  • Help Center
  • Contact Support
  • Report an Issue
  • Feature Requests

Company

  • Opt Out of Testing

Legal

  • Privacy Policy
  • Terms of Service
  • Cookie list
  • Subprocessors

Empra Consultancy LTD
hello@replylabs.io

ReplyLabs|PrivacyTermsCookiesSubprocessors

© 2026 Empra Consultancy LTD. All rights reserved.

All articles
Blog

Why Enrichment Coverage Is Never 100%, and What Is Normal

A fifth to a third of company URLs will not return data, and no vendor fixes that. Why coverage fails, what a realistic rate looks like, how to measure yours.

By ReplyLabs · 5 min read · 15 September 2026

On this page
  • The five reasons a row comes back empty
  • What realistic looks like
  • Why advertised match rates are not comparable
  • Measuring your own, in about ten minutes
  • What to do about the rows you cannot fill
  • Related reading

The short answer.

  • Expect roughly two thirds to four fifths of company URLs to return usable data on a first pass.
  • Most of the shortfall is not fixable by any tool: dead domains, parked pages, and facts that are simply not published.
  • Advertised match rates above 90% usually have a different denominator than the list you submitted.
  • Coverage is set by how the list was built far more than by which tool reads it.

Everyone selling enrichment publishes a match rate. Almost nobody publishes what happens to the rows that do not match, which is the number that decides whether your campaign has holes in it.

This is an attempt at the honest version: why coverage fails, what is realistic, and how to measure your own rather than trusting anybody's marketing.

The five reasons a row comes back empty

They are not equally fixable, and treating them as one number is what makes the topic confusing.

1. The page needs JavaScript to exist

A plain HTTP fetch gets a near-empty shell and the content arrives later from a script. Fixable, by rendering the page before reading it, at meaningfully higher cost per row. This is a large share of first-pass failures on modern marketing sites.

2. The site refuses automated requests

A CDN or WAF challenges anything that does not look like a browser. Partly fixable. Different request methods and fallbacks recover some of it. A site that has decided to block automation will win, and should.

3. The domain is dead, parked, or redirected

The company folded, rebranded, was acquired, or the domain in your list was never their real one. Not fixable. No amount of retrying finds data behind a domain-parking page, and this is pure list quality.

4. The page loads fine and does not say the thing

The most under-appreciated cause. You asked for employee count, or the tech stack, or which market they serve, and the website simply does not mention it. Not fixable by scraping, because the input does not contain the answer. It needs a different source, or a decision that this field is optional.

5. Transient failure

A timeout, a rate limit, a bad minute. Fixable by retrying, and it is the cheapest recovery available. A tool that does not retry is quietly converting temporary failures into permanent gaps in your list, which is the single most avoidable form of coverage loss.

What realistic looks like

For a list of company websites read for firmographic detail, a reasonable expectation on a first pass is somewhere between two thirds and four fifths of rows returning usable data. Adding rendering and fallbacks for the rows that failed recovers a meaningful part of the remainder, though never all of it.

Two things shift that range far more than tooling choice:

How the list was built. A list exported from a maintained CRM performs very differently from a list scraped off a conference attendee page two years ago. Most of the variance people attribute to their enrichment tool is actually variance in the age and provenance of their list.

What you asked for. "What does this company do" is answerable from almost any homepage. "How many engineers do they employ" is answerable from a minority of them. Coverage is a property of the question as much as of the source, and comparing two tools on different questions tells you nothing.

Why advertised match rates are not comparable

When a vendor publishes 95% coverage, ask what the denominator is. Three common constructions, all defensible and all measuring something different from what you want to know:

  • Over records the vendor could identify. Rows it could not resolve at all never enter the calculation. This is the most common construction and it can turn a 70% real rate into a 95% published one.
  • Over a curated benchmark list. Usually large, well-known, well-documented companies. Your list is not that.
  • Over any field returned. A row that came back with a country and nothing else counts as matched, even though the field you needed is empty.

None of these are lies. They are just not the number you are trying to buy.

Measuring your own, in about ten minutes

Take one real run and count four things:

BucketWhat it tells you
Rows with usable outputYour actual coverage. The only rate that matters.
Rows that failed to fetchBlocking and rendering. Partly recoverable with better tooling.
Rows that fetched but had no answerThe field is unanswerable from this source. Change the source or drop the field.
Rows with a dead or parked domainList quality. No tool will help.

If bucket four is large, stop evaluating enrichment tools and go and fix the list, because you are trying to solve a data-sourcing problem with a data-reading tool.

The awkward part of this exercise is that most setups cannot produce the split at all. If a failed row is a blank cell, all four buckets look identical, which is why per-row failure reasons are worth more than a slightly higher headline rate. Being able to separate "blocked" from "not published" is the difference between a fixable problem and a permanent one.

What to do about the rows you cannot fill

  • Retry the transient ones. Cheapest recovery there is.
  • Fall back for the blocked ones. A second, more expensive method on the subset that failed costs far less than running the expensive method on everything. That is the whole idea behind waterfall enrichment.
  • Drop the unanswerable field. If a website does not publish headcount, no tool reading websites will find it.
  • Never let a blank become a personalised sentence. A row with no data must not reach the generation step. This is the failure that costs reputation rather than money.

That last point is the one worth building around. Coverage below 100% is normal and survivable. Coverage below 100% that nobody can see is what puts Hi , I noticed that into somebody's inbox.

Related reading

  • Waterfall enrichment, and when to fall back
  • What 1,000 enriched leads actually cost
  • Why a page loads and still returns nothing
  • Firmographic enrichment, explained

Frequently asked questions

For a list of company websites scraped for firmographic detail, expect somewhere between two thirds and four fifths of rows to return usable data on a first pass. The exact figure depends far more on how the list was built than on which tool reads it.
On this page
  • The five reasons a row comes back empty
  • What realistic looks like
  • Why advertised match rates are not comparable
  • Measuring your own, in about ten minutes
  • What to do about the rows you cannot fill
  • Related reading

Keep reading

All articles
Blog

What 1,000 Enriched Leads Actually Cost, Line by Line

Blog

Scrape, Summarise, Score: Multi-Step Enrichment in Sheets

Blog

How Many Rows Can You Run AI On in Google Sheets?

Try it on your own list

ReplyLabs runs from a sidebar inside Google Sheets. Start free with $20 credit, no card needed.

Get started free