Google Leak Explained: What the 2024 Docs Show
The Google leak is the internal Search documentation that became public in May 2024: more than 2,500 pages describing Google’s Content Warehouse API, listing 14,014 attributes that Google stores about pages, sites, links and user clicks. It shows what Google stores. It does not show how any of it is weighted in ranking. That distinction is the whole story, and most coverage blurred it.
A note on the keyword: most “Google leak” searches are about password dumps and Gmail breaches. This page is about the Search documentation leak. It is for marketers and owners who keep hearing the leak cited as proof of something and want to know what it actually says.
On this page
- What was the Google leak?
- How the documents surfaced
- What the leak is, and what it is not
- The attributes that matter for content
- What did the leak change in SEO thinking?
- Reading the leak without overreading it
- Where the Google leak fits in an SEO audit
What was the Google leak?
It was a public copy of documentation for Google’s internal Content Warehouse API, the system that stores what Google knows about documents on the web. Mike King’s iPullRank analysis explains that the documentation was “accidentally published” to a public code repository as part of a client library, under an Apache 2.0 license, and describes it as the “current, active architecture of Google Search Content Storage as of March of 2024.”
The scale is what made it matter. iPullRank counted 2,596 modules with 14,014 attributes. Each attribute has a name, a type and, often, a one-line description written by a Google engineer. Those descriptions are the valuable part. They say, in Google’s internal shorthand, what a field is for.
We work from the same material. Ambeltek’s internal audit tooling is built on a parsed copy of the package, version 0.4.0, which contains 2,593 module files; we have not reconciled the small gap with iPullRank’s count. Every attribute quoted on this page was checked against those files.
How the documents surfaced
The documents sat in a public repository for about six weeks, from March 27 to May 7, 2024, and were already gone when the first analyses appeared. The timeline below comes from the two original write-ups.
Scroll the diagram sideways to see all of it.
Rand Fishkin’s SparkToro post names the source as Erfan Azimi, an SEO practitioner, and describes checking the material with former Google employees, one of whom said it had “all the hallmarks of an internal Google API.” Google did not dispute it. Spokesperson Davis Thompson told reporters, as Entrepreneur reported: “We would caution against making inaccurate assumptions about Search based on out-of-context, outdated, or incomplete information.”
Read that statement carefully. It does not say the documents are fake. It says they are easy to misread. On that point Google is right.
What the leak is, and what it is not
The leak is a data dictionary, not an algorithm. It tells you which fields exist and, often, what each one holds. It does not tell you which fields feed ranking, how they are combined, or how much each one counts. iPullRank, the most thorough analysis, is explicit: there is “no detail about Google’s scoring functions,” and “we do not know how features are weighted in the various downstream scoring functions.”
Scroll the diagram sideways to see all of it.
Two terms recur in the descriptions and explain where stored values would be used. Q* (Qstar) is where siteAuthority is “applied,” and the notes of a January 2025 call with Google engineer Pandu Nayak, filed in the DOJ case, describe Q* as “Google’s measure of quality of a document.” The same notes define twiddlers as functions that “re-rank a set of already selected results,” which is where hostAge is said to sandbox fresh spam. Both terms tell you roughly where a field sits in the pipeline. Neither tells you its weight.
So every attribute below comes with the same caveat, stated once here and meant everywhere: Google stores this attribute; the leak does not show how much it is weighted. Some fields are marked deprecated or legacy in their own descriptions. Some may never have influenced a ranking.
The attributes that matter for content
Out of 14,014 attributes, a handful describe things a content team can actually influence. These are quoted verbatim from the module files, with the module and attribute name exactly as they appear.
| Attribute | Verbatim description (abridged only where marked) | Why a content team should care |
|---|---|---|
QualityNsrNsrData.titlematchScore | “Titlematch score of the site, a signal that tells how well titles are matching user queries.” | Site-level, not page-level: titles across the whole site |
QualityNsrPQData.contentEffort | “LLM-based effort estimation for article pages (see landspeeder/4311817).” | Google stores a model’s estimate of effort |
PerDocData.OriginalContentScore | ”…Only pages with little content have this field…” | Originality is tracked for thin pages specifically |
QualityAuthorityTopicEmbeddingsVersionedItem.siteFocusScore | “Number denoting how much a site is focused on one topic.” | Topical focus is measured at site level |
QualityAuthorityTopicEmbeddingsVersionedItem.siteRadius | “The measure of how far page_embeddings deviate from the site_embedding.” | Off-topic pages are measurably off-topic |
CompressedQualitySignals.siteAuthority | “site_authority: converted from quality_nsr.SiteAuthority, applied in Qstar.” | A site-wide authority value exists |
PerDocData.hostAge | “The earliest firstseen date of all pages in this host/domain. These data are used in twiddler to sandbox fresh spam in serving time…” | New hosts get extra scrutiny for spam |
QualityNsrNsrData.smallPersonalSite | “Score of small personal site promotion go/promoting-personal-blogs-v1” | A promotion for small personal sites exists |
The diagram groups them by what they describe.
Scroll the diagram sideways to see all of it.
Two more fields are worth knowing because they describe the page experience rather than the text. QualityNsrNsrData.clutterScore is described as a “Delta site-level signal in Q* penalizing sites with a large number of distracting/annoying resources loaded by the site.” And CompressedQualitySignals.exactMatchDomainDemotion is “converted from QualityBoost.emd.boost,” which is the old exact-match-domain demotion, still stored in 2024.
What did the leak change in SEO thinking?
Mostly, it moved several long-running arguments from speculation to documentation, while leaving their weights unknown. Four in particular:
-
Clicks. The leak includes click fields such as
goodClicks,badClicksandlastLongestClicksinside modules namedQualityNavboostCraps.... SparkToro tied them to NavBoost, Google’s click-based system. The antitrust case had already described NavBoost as memorizing clicks for all queries from the prior 13 months, according to the plaintiffs’ post-trial brief in US v. Google. Together, the two sources make “Google uses clicks” hard to deny. Our NavBoost explainer goes through both. -
Site-level signals.
siteAuthority,siteFocusScoreand a site-widetitlematchScoreshow Google stores judgments about whole sites, not just pages. That fits what our guide to topical authority argues: a focused site is easier to understand. -
Chrome data.
chromeInTotalis “Site-level Chrome views.” SparkToro also pointed to Chrome clickstream use elsewhere in the docs. -
New-site scrutiny.
hostAgeis used “to sandbox fresh spam in serving time.” That is narrower than the folk idea of a sandbox for every new site: it is written about fresh spam. -
Topic-specific trust flags. SparkToro highlighted fields that look like allowlists for sensitive topics. One is
QualityNsrNsrData.isElectionAuthority, described as a “Bit to determine whether the site has the election authority signal.” For most businesses this is irrelevant. For publishers in news, health or elections, it is a reminder that some queries are held to a different standard, and that no amount of on-page work substitutes for being a recognized authority on the topic.
It also surfaced fields few people had written about, such as smallPersonalSite, a stored promotion score for small personal sites, and contentEffort, an LLM-based estimate of effort on article pages.
Reading the leak without overreading it
Use the leak to decide what to measure, not to claim what Google rewards. Four rules keep it honest:
- Quote the attribute, then the caveat. “Google stores
contentEffort” is true. “Google ranks by effort” is not something the leak shows. - Prefer attributes with descriptions. Many fields have no description at all. Their names invite guesses; do not build strategy on a guess.
- Check the date. The documents reflect March 2024. Google has run eight core updates since then, and any of them may have changed what is used. See our note on Google core updates.
- Look for corroboration. The strongest leak-based conclusions are the ones confirmed elsewhere, such as clicks through trial testimony, or quality through Google’s own E-E-A-T guidance.
In practice, the leak points to the same habits that most lists of SEO ranking factors already recommend: titles that match what people search and what the page delivers, content that shows real effort and originality, a site that stays on topic, and pages that satisfy the click. The difference is that you can now name the stored fields those habits relate to.
Where the Google leak fits in an SEO audit
We use the leak as a checklist of things Google bothers to store, then test each one with evidence we can see. Titles are checked against the queries in Search Console, because titlematchScore is described as how well titles match user queries. Off-topic sections are flagged, because siteRadius measures how far pages drift from the site. Thin pages are reviewed for originality, because OriginalContentScore exists specifically for pages with little content. Each check is useful on its own merits; the leak only tells us Google thought the same property was worth recording.
That is how our AI visibility audit uses the material: as a map of what Google measures, paired with the caveat that it is stored, not weighted. If you want to see the method applied to your own pages, the audit is the place to start, and the SEO audit guide shows the parts you can run yourself.
Frequently asked questions
Was there a Google data leak?
There have been several unrelated incidents, but in SEO 'the Google leak' means the Content Warehouse API documentation that was published to a public code repository on March 27, 2024 and removed on May 7, 2024. It exposed internal attribute names and descriptions from Google Search's storage systems. It did not expose user data or passwords.
Did Google confirm the leak was real?
Google did not dispute the documents. Spokesperson Davis Thompson said: 'We would caution against making inaccurate assumptions about Search based on out-of-context, outdated, or incomplete information.' Reports at the time treated that as confirmation that the documents were genuine.
Does the Google leak show ranking factor weights?
No. The documents list modules and attributes with short descriptions, but no scoring functions or weights. iPullRank's analysis states plainly that we do not know how features are weighted downstream. An attribute being stored does not prove it is used in ranking, or how much.
How big was the Google Search leak?
iPullRank counted 2,596 modules with 14,014 attributes, and SparkToro described it as more than 2,500 pages of API documentation. The package version we parsed, google_api_content_warehouse 0.4.0, contains 2,593 module files.
Does the leak prove Google uses clicks for ranking?
It shows Google stores click attributes such as goodClicks, badClicks and lastLongestClicks in modules named after NavBoost. Separately, the US v. Google antitrust case described NavBoost as a ranking component that uses 13 months of click data. The leak alone shows storage; the trial record shows use.
Who leaked the Google documents?
The documents were shared with SparkToro's Rand Fishkin by Erfan Azimi, an SEO practitioner. They had been published accidentally to a public repository as part of Google's Content Warehouse API client library, then removed.