contact@designmemarketing.com Deer Park, Long Island
★★★★★ 5.0 on Google · 200+ Clients Free SEO Audit
Behind the Numbers

How SEO Tools Like Ahrefs and Semrush Actually Get Their Data

That search volume number looks so precise. That Domain Rating feels so official. We took the whole machine apart to show you where every figure really comes from, and how much of it is careful, educated guessing.

0data sources blended to estimate every search volume number
0wide volume buckets Google actually sorts keywords into
0domains studied showing authority correlates with, not causes, rankings
0times Google uses any of these scores to rank your site

You open Ahrefs and it tells you a keyword gets 2,600 searches a month. Not "about 2,000." Not "a few thousand." Two thousand six hundred. That kind of precision feels like a fact pulled straight from Google's own servers. It is not. No third-party tool has a live feed into Google's real search data. Not Ahrefs, not Semrush, not Moz, not any of them. Every number they show you is an estimate, built from outside data and dressed up to look exact.

This is not a scandal, and it does not make these tools useless. We use them every single day, and we would not run campaigns without them. But most people treat these numbers as gospel, and that quietly leads to bad decisions, wasted budget, and panic over changes that never really happened. So we pulled the whole thing apart, piece by piece, to show you exactly how each number is made. By the end, you will read every SEO tool differently, and better.

First, a short history of why we're all guessing

To understand why an entire industry exists just to estimate search volume, you have to know one thing: Google used to give it away for free. A decade ago, the old AdWords Keyword Tool handed back real, specific monthly search counts for any keyword you asked about. No guessing required. Then, around 2016, Google changed the rules. Unless you were actively spending money on ads, Keyword Planner stopped showing exact numbers and started showing broad ranges instead.

Overnight, the one public window into real search demand went dark for most of the world. And that single decision created a market. Ahrefs, Semrush, Moz, and the dozens of tools behind them exist, in large part, to rebuild the number Google took away. They are not tapping a secret pipe into Google. They are reverse-engineering a figure Google used to publish, using whatever outside signals they can buy, crawl, or model. Once you see it that way, everything else makes sense.

Where search volume numbers really come from

Every tool builds its search volume estimates from three ingredients, mixed in slightly different proportions.

Google Keyword Planner
The baseline. Built for advertisers, and it reports broad ranges, not exact counts.
Clickstream data
Anonymized browsing behavior bought from browser extensions, plugins, and user panels.
Their own models
Machine learning that fills gaps, un-groups similar terms, and guesses the final figure.

Each ingredient has real strengths and real blind spots. To understand how much to trust the output, you have to understand each input. Let us take them one at a time.

Ingredient one: Google's buckets, not Google's numbers

Keyword Planner is the closest thing to a public window into real search demand, but it was made for people buying ads, not for SEOs. Two things about it matter enormously. First, it groups similar keywords together, so "keyword tool," "keyword checker," and "keyword finder" can all get folded into one combined number. Second, and this is the big one, it does not report exact volumes at all. It sorts every keyword into one of a small set of wide buckets.

Google's actual monthly volume buckets

1 – 10
11 – 100
101 – 1K
1.1K – 10K
10.1K – 100K
← your "24,000" really lives here
100K – 1M
1M+

Look at what that means. When a tool proudly shows you "24,000 searches," that keyword really lives somewhere in the ten thousand to one hundred thousand bucket, a band that is ten times wide from floor to ceiling. The tool then uses clickstream data and modeling to plant a specific-looking flag inside that enormous range. The precision on your screen is manufactured. The honest underlying truth is a bucket, and the exact number is a guess about where in the bucket you sit.

Ingredient two: the clickstream data supply chain

To sharpen those buckets into specific numbers, the tools buy clickstream data. This is where it gets genuinely interesting, and a little uncomfortable. Clickstream is anonymized browsing behavior, a record of what real people search, click, and visit, harvested from software running quietly on millions of devices.

Where does it come from? Free browser extensions. Toolbars. Mobile apps. Some VPNs. Certain plugins you installed years ago and forgot about. When you clicked "accept" on the terms, many of them reserved the right to collect and sell your anonymized activity. That data flows to brokers, who package and sell it wholesale to anyone building a search model, including the SEO tools. You have almost certainly fed one of these panels without ever knowing its name.

Here is the catch that matters most. Clickstream is a sample, not a census. The tools observe what a few million people do, then extrapolate it to represent billions of searches. That works reasonably well for popular keywords, where the sample sees plenty of activity. It falls apart on niche, long-tail, or brand-new terms, where the panel might have seen the keyword a handful of times, or never. A small sample error on a small keyword gets multiplied up into a big, confident, wrong number. There is also panel bias: the kind of person who installs a free tracking extension is not a perfect mirror of every searcher, which skews the sample in ways no model fully corrects.

Same keyword, four different answers

Here is the proof that these are estimates and not facts. Ask four different sources for the volume of the exact same keyword, and you get four different numbers. If any of them had Google's real data, they would all agree. They do not. Tap a keyword and watch the spread.

Google Keyword Planner says 1K – 10K / mo (a 10x range)
Ahrefs2,600
Semrush3,600
Moz1,900

Illustrative of the typical spread. Same keyword, same month, three different tools, three different answers, all sitting inside one very wide Google range.

None of those numbers is lying. They are all honest estimates from different samples run through different models. But not one of them is "the truth," because outside of Google, the truth is not for sale.

Why the tools disagree so much

The gap between tools is not random noise or one company being worse than another. It comes from specific, understandable differences in how each one is built. Once you see the five reasons, the disagreement stops being confusing and starts being expected.

01Different clickstream panels. Each tool buys behavior data from different brokers, which means different samples of different people. Different inputs produce different outputs.
02Different blend ratios. One tool leans harder on Keyword Planner, another leans harder on clickstream. The recipe is proprietary, and the mix changes the number.
03Different models. Each company's machine learning makes its own assumptions about how a sample maps to the whole population. Those assumptions rarely match.
04Different indexes. For traffic and links, each tool crawls its own slice of the web, so they are literally looking at different data before they even start estimating.
05Different refresh schedules. One updates monthly, another quarterly. You may be comparing a fresh snapshot against a stale one without realizing it.

This is the same reason your rank tracker and your own eyes so often disagree, which we broke down in why your rankings fluctuate. Different vantage points, different answers, none of them the single truth.

How traffic estimates are built

Search volume is one guess. Traffic estimates stack several guesses on top of each other. Here is the key thing to understand: no tool can see your Google Analytics. They have no access to your real traffic. So when a tool claims a competitor gets 40,000 visits a month, it did not measure that. It estimated it from the outside, like this.

EstimateSearch volume for each keyword they rank for
×
EstimateClick-through rate for that ranking position
×
EstimateThe position they think you actually hold
=
ResultAn estimated traffic number

Look at what that means. The final traffic figure is an estimate, multiplied by an estimate, multiplied by an estimate. Every one of those inputs carries its own margin of error, and multiplying them together stacks the errors up. That is exactly why a tool can show you a scary traffic drop that your own analytics never recorded. The tool did not watch your visitors leave. It just recalculated its guess.

The weakest link: the click-through rate assumption

Of those three inputs, the middle one is the shakiest, and almost nobody talks about it. To turn a ranking position into a traffic number, every tool assumes a click curve: position one gets some set percentage of clicks, position two gets less, and so on down the page. That curve looks something like this.

A typical modeled click-through rate by position

#1#2#3#4#5#6#7#8#9#10
Illustrative of a typical modeled curve. Real click-through rates vary enormously by search and are shifting as AI results grow.

The problem is that this curve is a fantasy of a clean, simple results page that barely exists anymore. Real click-through rate depends entirely on what the actual page looks like. Ads at the top push every organic result down. A map pack, a featured snippet, a shopping carousel, or a video block can swallow most of the clicks before anyone scrolls. And now AI Overviews answer the question right on the page, so a growing share of searches end with no click at all. A position-based curve cannot see any of that. It assumes clicks that may never happen, then multiplies that assumption straight into the traffic estimate.

The one thing they actually measure: backlinks

Everything so far has been estimated. Backlinks are the exception, because the tools crawl the web themselves rather than buying a sample. This is where they do real measurement, and it is the strongest data any of them own. AhrefsBot is the second most active crawler on the entire internet, behind only Google's own, which is why Ahrefs tends to find new links quickly and keep a large, fresh index.

But even here, complete is the wrong word. Each tool's link index is a huge sample of the web, not all of it. They catch most links, miss some, and lag on others, especially lost links, which can sit in a report long after they are gone. Different crawlers find different links, which is why two tools rarely agree on the exact backlink count for the same site. They also handle link types differently: Domain Rating, for example, counts only followed links and ignores nofollow, sponsored, and user-generated ones entirely. Useful, thorough, genuinely measured, and still not the whole web.

While we're here: keyword difficulty is a guess too

Every tool also shows a keyword difficulty score, usually zero to one hundred, that promises to tell you how hard a keyword is to rank for. Treat it with the same caution as everything else. Each tool calculates difficulty its own way, mostly by looking at the backlink strength of the pages currently ranking on page one, then compressing that into a single number. Different tools, different formulas, different scores for the same keyword. It is a useful rough gauge of competitiveness, and a terrible thing to treat as precise.

How Domain Rating and Domain Authority are calculated

Now the metric everyone quotes and almost nobody understands. Domain Rating, Domain Authority, and Authority Score are three different companies' attempts to score a website's strength on a scale of 0 to 100. They are built almost entirely from backlink data, each from that tool's own index, and they are not interchangeable. Tap through them.

MeasuresThe strength of a backlink profile. It counts unique referring domains, weights each by its own Domain Rating, and divides by how many sites that domain links out to, so a link gets diluted when the source links to everyone.
IgnoresTraffic, spam signals, domain age, content quality, on-page factors. It is purely a link-graph score, and only followed links count.
Scale0 to 100, logarithmic. Ahrefs is refreshingly transparent that this is all it measures.
MeasuresA prediction, not a measurement. Moz uses a machine learning model trained on more than 40 signals to predict how likely a domain is to rank, then scores that prediction.
IgnoresIt is a modeled guess at ranking ability, built from Moz's own link index plus spam signals. It does not read your content or judge search intent.
Scale1 to 100, logarithmic. Moz states plainly that Google does not use it.
MeasuresA blended score. Semrush mixes backlink signals with estimated organic traffic and natural-versus-spam link patterns, all through machine learning.
IgnoresLike the others, it is a proxy. The added traffic ingredient is itself an estimate, so it blends measured links with modeled traffic.
Scale0 to 100. Not comparable to DR or DA, because it is built from a different index and formula.

Because each tool uses its own crawler and its own math, the same website can score DR 55 on Ahrefs and DA 48 on Moz without either being wrong. Comparing one site's DR to another site's DA is meaningless. Pick one metric and track only that one over time.

There is one more trap in the number itself. That 0 to 100 scale is logarithmic, not linear, so the distance between scores is wildly uneven.

20 30 70 80
20 to 30: a handful of good links, done in months. 70 to 80: years of work and hundreds of new referring domains.

Does Google use any of this?

This is the part that matters most, and the answer is blunt. Google does not use Domain Rating, Domain Authority, or Authority Score. At all. Not for crawling, not for indexing, not for ranking. Google's own John Mueller has said it about as plainly as a person can.

"We don't use domain authority. That's a metric from an SEO company."

John Mueller, Google Search Advocate

Now the honest nuance, because we would rather give you the whole picture than a tidy half-truth. In 2024, a large leak of internal Google documents surfaced a feature called siteAuthority, which suggests Google does have some domain-level authority concept of its own inside its systems. That does not validate the vendor scores. It means Google may run its own private version, computed from its own data, that has nothing to do with Moz's number or Ahrefs' number. Your DR is still not being fed into Google. Both things are true at once.

So why do high-DR sites so often rank well? Correlation, not causation. A strong backlink profile helps you rank, and a strong backlink profile also raises your DR. Both flow from the same source, which makes them move together, but one does not cause the other. A page with DR 80 and thin, off-target content will still get buried. The score is a symptom of doing the work, not a substitute for it. We went deeper on that in why your competitor ranks higher.

The blind spot nobody's numbers cover

There is one more limitation worth naming, and it is growing fast. Every figure in every one of these tools is built from traditional search engine data, which really means Google. None of them count the searches people now make inside ChatGPT, Perplexity, Gemini, or Claude, and they do not fully capture what happens inside AI Overviews. That means the real demand for almost any topic is higher than the volume number shows, and the gap widens every month as more people ask an AI instead of typing into a search bar. The tools are estimating a slice of search behavior that is shrinking as a share of the whole. Worth keeping in mind before you write off a topic because a tool said the volume was low.

Want to see it for yourself?

You do not have to take our word for any of this. The disagreement is easy to reproduce in about ten minutes, and once you have seen it with your own keywords, you will never read these numbers the same way again.

  1. Pick one keyword and check its search volume in Ahrefs, Semrush, and Moz. Write down the three numbers and notice how far apart they are.
  2. Open Google Keyword Planner and look up the same keyword. See the wide range it actually sits in, and notice how all three tool numbers fall somewhere inside it.
  3. Compare all of that to your own Google Search Console data for a keyword you already rank for. That impressions and clicks figure is the only real number in the whole exercise.
  4. Track one metric in one tool for three months instead of chasing absolute figures. The trend is trustworthy even when the exact number is not.

So what are these tools actually good for?

A fair question after all of that, and the answer is: quite a lot, as long as you use them for what they are. The trick is knowing which jobs they are built for and which ones they quietly fail at.

Trust them for
Don't trust them for
  • Comparing keywords against each other
  • Spotting trends over time in one tool
  • Competitive backlink analysis
  • Vetting and prospecting link opportunities
  • Discovering keywords you never knew existed
  • Rough, directional sizing of a market
  • Exact search volume for a keyword
  • Exact traffic, yours or a competitor's
  • Precise month-to-month change
  • Treating DR or DA as a Google score
  • Low-volume and long-tail precision
  • Measuring real AI-search demand

Used correctly, they are genuinely useful, and we rely on them constantly. Treat the numbers as relative and directional, not absolute. A keyword showing higher volume than another in the same tool is a fair comparison. Watching your own estimated trend climb over months is meaningful. Sizing up a competitor's backlink profile is exactly what these tools are built for, which is why they anchor real work like keyword research.

What you should never do is treat an estimate as gospel, jump between tools because one flatters you this month, or chase a Domain Rating like it is the goal. The real scoreboard is not a third-party estimate at all. It is your Google Search Console data, your actual clicks, impressions, and queries straight from Google, paired with honest rank tracking and, above everything, the leads and calls your business receives. That is why our SEO reporting leans on what is real and uses the tools for context, not as the source of truth. If you want the bigger picture on where your effort goes, we laid it out in where your SEO budget goes.

The tools are a map, drawn from the best outside information money can buy. A good map is worth having, and we would never work without one. Just never forget it is a drawing of the territory, not the territory itself. The businesses that win are the ones who read the map with clear eyes, and keep them fixed on the real ground: the customers actually finding them, calling them, and choosing them.

Sources

  • Ahrefs, Website Authority Checker (how Domain Rating is calculated, DR correlation study of 218,713 domains)
  • Google Search Central and John Mueller public statements (Google does not use domain authority metrics)
  • Moz and Semrush documentation (Domain Authority and Authority Score methodology)
  • Google Ads Keyword Planner (bucketed volume ranges and close-variant grouping)
  • Industry reporting on clickstream data sourcing, position-based CTR curves, and the 2024 Google documentation leak (siteAuthority)

Frequently asked questions

How do Ahrefs and Semrush know search volume?

They estimate it. They start with Google Keyword Planner, which returns broad ranges built for advertisers, then layer on clickstream data, anonymized browsing behavior bought from browser extensions and panels of real users, and run both through their own models. Because clickstream is a sample and Keyword Planner groups similar terms together, the final number is a modeled estimate, not a real count. That is why the same keyword shows different volumes in every tool.

Why do Ahrefs and Semrush show different numbers for the same keyword?

Because they use different clickstream panels, blend Google Keyword Planner data in different proportions, run different models, crawl different slices of the web, and refresh on different schedules. You are comparing separate estimates built from separate inputs, so they rarely match. If any tool had Google's real numbers, they would all agree. They do not, which is the clearest proof that these are estimates.

Is Domain Rating or Domain Authority a Google ranking factor?

No. Domain Rating is from Ahrefs, Domain Authority is from Moz, and Authority Score is from Semrush. Google does not use any of them. Google's John Mueller has said plainly, we don't use domain authority, that's a metric from an SEO company. These scores are each built from backlink data in that tool's own index, and they correlate with rankings only because strong backlinks help both the score and the ranking.

Can SEO tools see my real website traffic?

No. They cannot see your Google Analytics or your real numbers. They estimate your traffic from the outside by taking every keyword they think you rank for, multiplying an estimated search volume by an estimated click-through rate for your estimated position, and adding it up. It is an estimate built on estimates, which is why a tool can show a traffic drop your own analytics never recorded.

SEO built on what's real, not just what's estimated

We use the same tools everyone else does, but we know exactly where the estimates end and the real data begins. Our reporting is built on your actual Google numbers and the leads you can count, with the tools used for context, not as the whole story. To see what that looks like for your business, tell us about it below or get a quote.