SEOStrategy

Google Leak Explained: What the 2024 Docs Show

The Google leak is the internal Search documentation that became public in May 2024: more than 2,500 pages describing Google’s Content Warehouse API, listing 14,014 attributes that Google stores about pages, sites, links and user clicks. It shows what Google stores. It does not show how any of it is weighted in ranking. That distinction is the whole story, and most coverage blurred it.

A note on the keyword: most “Google leak” searches are about password dumps and Gmail breaches. This page is about the Search documentation leak. It is for marketers and owners who keep hearing the leak cited as proof of something and want to know what it actually says.

On this page

What was the Google leak?

It was a public copy of documentation for Google’s internal Content Warehouse API, the system that stores what Google knows about documents on the web. Mike King’s iPullRank analysis explains that the documentation was “accidentally published” to a public code repository as part of a client library, under an Apache 2.0 license, and describes it as the “current, active architecture of Google Search Content Storage as of March of 2024.”

The scale is what made it matter. iPullRank counted 2,596 modules with 14,014 attributes. Each attribute has a name, a type and, often, a one-line description written by a Google engineer. Those descriptions are the valuable part. They say, in Google’s internal shorthand, what a field is for.

We work from the same material. Ambeltek’s internal audit tooling is built on a parsed copy of the package, version 0.4.0, which contains 2,593 module files; we have not reconciled the small gap with iPullRank’s count. Every attribute quoted on this page was checked against those files.

How the documents surfaced

The documents sat in a public repository for about six weeks, from March 27 to May 7, 2024, and were already gone when the first analyses appeared. The timeline below comes from the two original write-ups.

Timeline of the 2024 Google Search documentation leak Four points on a line. March 27, 2024: documentation pushed to a public repository as part of the Content Warehouse API client library. May 7, 2024: removed from the repository. May 27 and 28, 2024: iPullRank and SparkToro publish their analyses after Erfan Azimi shares the documents with Rand Fishkin. Days later: Google spokesperson Davis Thompson cautions against inaccurate assumptions based on out-of-context, outdated or incomplete information, without disputing authenticity. From accidental commit to public story Dates from iPullRank and SparkToro's May 2024 write-ups 27 Mar 2024 7 May 2024 27 to 28 May Days later Docs pushed to a public repository (client library) Removed from the repository iPullRank and SparkToro publish Source: Erfan Azimi Google cautions against "inaccurate assumptions" About six weeks in a public repository (27 Mar to 7 May). No user data or passwords were involved.

Scroll the diagram sideways to see all of it.

Source: Ambeltek diagram based on iPullRank's and SparkToro's May 2024 write-ups and Google's statement as reported by Entrepreneur.

Rand Fishkin’s SparkToro post names the source as Erfan Azimi, an SEO practitioner, and describes checking the material with former Google employees, one of whom said it had “all the hallmarks of an internal Google API.” Google did not dispute it. Spokesperson Davis Thompson told reporters, as Entrepreneur reported: “We would caution against making inaccurate assumptions about Search based on out-of-context, outdated, or incomplete information.”

Read that statement carefully. It does not say the documents are fake. It says they are easy to misread. On that point Google is right.

What the leak is, and what it is not

The leak is a data dictionary, not an algorithm. It tells you which fields exist and, often, what each one holds. It does not tell you which fields feed ranking, how they are combined, or how much each one counts. iPullRank, the most thorough analysis, is explicit: there is “no detail about Google’s scoring functions,” and “we do not know how features are weighted in the various downstream scoring functions.”

What the leaked documentation shows versus what it cannot show Two columns. What the leak shows: module names such as QualityNsrNsrData; attribute names such as titlematchScore; data types and encodings; short engineer-written descriptions; which systems a field is converted for, such as Qstar or twiddlers. What the leak cannot show: weights or scoring functions; whether a stored field is used in ranking today; how attributes are combined; what changed after March 2024. Bottom line: stored, not weighted. Stored is not weighted What a data dictionary can and cannot tell you The leak shows Module names (QualityNsrNsrData) Attribute names (titlematchScore) Data types and encodings One-line engineer descriptions Where a field is applied (Qstar, twiddlers) The leak cannot show Weights or scoring functions Whether a field ranks pages today How attributes are combined What changed after March 2024 Anything about AI Overviews ranking Rule for this page: Google stores the attribute; the leak does not show how much it is weighted.

Scroll the diagram sideways to see all of it.

Source: Ambeltek diagram based on the leaked google_api_content_warehouse 0.4.0 module files and iPullRank's May 2024 analysis.

Two terms recur in the descriptions and explain where stored values would be used. Q* (Qstar) is where siteAuthority is “applied,” and the notes of a January 2025 call with Google engineer Pandu Nayak, filed in the DOJ case, describe Q* as “Google’s measure of quality of a document.” The same notes define twiddlers as functions that “re-rank a set of already selected results,” which is where hostAge is said to sandbox fresh spam. Both terms tell you roughly where a field sits in the pipeline. Neither tells you its weight.

So every attribute below comes with the same caveat, stated once here and meant everywhere: Google stores this attribute; the leak does not show how much it is weighted. Some fields are marked deprecated or legacy in their own descriptions. Some may never have influenced a ranking.

The attributes that matter for content

Out of 14,014 attributes, a handful describe things a content team can actually influence. These are quoted verbatim from the module files, with the module and attribute name exactly as they appear.

AttributeVerbatim description (abridged only where marked)Why a content team should care
QualityNsrNsrData.titlematchScore“Titlematch score of the site, a signal that tells how well titles are matching user queries.”Site-level, not page-level: titles across the whole site
QualityNsrPQData.contentEffort“LLM-based effort estimation for article pages (see landspeeder/4311817).”Google stores a model’s estimate of effort
PerDocData.OriginalContentScore”…Only pages with little content have this field…”Originality is tracked for thin pages specifically
QualityAuthorityTopicEmbeddingsVersionedItem.siteFocusScore“Number denoting how much a site is focused on one topic.”Topical focus is measured at site level
QualityAuthorityTopicEmbeddingsVersionedItem.siteRadius“The measure of how far page_embeddings deviate from the site_embedding.”Off-topic pages are measurably off-topic
CompressedQualitySignals.siteAuthority“site_authority: converted from quality_nsr.SiteAuthority, applied in Qstar.”A site-wide authority value exists
PerDocData.hostAge“The earliest firstseen date of all pages in this host/domain. These data are used in twiddler to sandbox fresh spam in serving time…”New hosts get extra scrutiny for spam
QualityNsrNsrData.smallPersonalSite“Score of small personal site promotion go/promoting-personal-blogs-v1”A promotion for small personal sites exists

The diagram groups them by what they describe.

Leaked attributes relevant to content, grouped by what they describe Four groups of stored attributes. Site-level quality and focus: CompressedQualitySignals.siteAuthority, QualityAuthorityTopicEmbeddingsVersionedItem.siteFocusScore and siteRadius, QualityNsrNsrData.titlematchScore. Page content: QualityNsrPQData.contentEffort, PerDocData.OriginalContentScore. Users and clicks: QualityNavboostCrapsCrapsData.lastLongestClicks, QualityNsrNsrData.chromeInTotal. Age and demotions: PerDocData.hostAge, QualityNsrNsrData.clutterScore, CompressedQualitySignals.exactMatchDomainDemotion. Every attribute is stored; none has a published weight. Content-relevant attributes, by theme All stored; none has a known weight Site-level quality and focus CompressedQualitySignals.siteAuthority ...TopicEmbeddingsVersionedItem.siteFocusScore ...TopicEmbeddingsVersionedItem.siteRadius QualityNsrNsrData.titlematchScore Page content QualityNsrPQData.contentEffort (LLM-based effort estimation) PerDocData.OriginalContentScore (pages with little content) Users and clicks QualityNavboostCrapsCrapsData.lastLongestClicks QualityNsrNsrData.chromeInTotal ("Site-level Chrome views.") Age and demotions PerDocData.hostAge QualityNsrNsrData.clutterScore CompressedQualitySignals.exactMatchDomainDemotion Names shortened with "..." where the module is QualityAuthorityTopicEmbeddingsVersionedItem.

Scroll the diagram sideways to see all of it.

Source: Ambeltek diagram from the leaked google_api_content_warehouse 0.4.0 module files (attribute names verbatim).

Two more fields are worth knowing because they describe the page experience rather than the text. QualityNsrNsrData.clutterScore is described as a “Delta site-level signal in Q* penalizing sites with a large number of distracting/annoying resources loaded by the site.” And CompressedQualitySignals.exactMatchDomainDemotion is “converted from QualityBoost.emd.boost,” which is the old exact-match-domain demotion, still stored in 2024.

What did the leak change in SEO thinking?

Mostly, it moved several long-running arguments from speculation to documentation, while leaving their weights unknown. Four in particular:

  • Clicks. The leak includes click fields such as goodClicks, badClicks and lastLongestClicks inside modules named QualityNavboostCraps.... SparkToro tied them to NavBoost, Google’s click-based system. The antitrust case had already described NavBoost as memorizing clicks for all queries from the prior 13 months, according to the plaintiffs’ post-trial brief in US v. Google. Together, the two sources make “Google uses clicks” hard to deny. Our NavBoost explainer goes through both.

  • Site-level signals. siteAuthority, siteFocusScore and a site-wide titlematchScore show Google stores judgments about whole sites, not just pages. That fits what our guide to topical authority argues: a focused site is easier to understand.

  • Chrome data. chromeInTotal is “Site-level Chrome views.” SparkToro also pointed to Chrome clickstream use elsewhere in the docs.

  • New-site scrutiny. hostAge is used “to sandbox fresh spam in serving time.” That is narrower than the folk idea of a sandbox for every new site: it is written about fresh spam.

  • Topic-specific trust flags. SparkToro highlighted fields that look like allowlists for sensitive topics. One is QualityNsrNsrData.isElectionAuthority, described as a “Bit to determine whether the site has the election authority signal.” For most businesses this is irrelevant. For publishers in news, health or elections, it is a reminder that some queries are held to a different standard, and that no amount of on-page work substitutes for being a recognized authority on the topic.

It also surfaced fields few people had written about, such as smallPersonalSite, a stored promotion score for small personal sites, and contentEffort, an LLM-based estimate of effort on article pages.

Reading the leak without overreading it

Use the leak to decide what to measure, not to claim what Google rewards. Four rules keep it honest:

  1. Quote the attribute, then the caveat. “Google stores contentEffort” is true. “Google ranks by effort” is not something the leak shows.
  2. Prefer attributes with descriptions. Many fields have no description at all. Their names invite guesses; do not build strategy on a guess.
  3. Check the date. The documents reflect March 2024. Google has run eight core updates since then, and any of them may have changed what is used. See our note on Google core updates.
  4. Look for corroboration. The strongest leak-based conclusions are the ones confirmed elsewhere, such as clicks through trial testimony, or quality through Google’s own E-E-A-T guidance.

In practice, the leak points to the same habits that most lists of SEO ranking factors already recommend: titles that match what people search and what the page delivers, content that shows real effort and originality, a site that stays on topic, and pages that satisfy the click. The difference is that you can now name the stored fields those habits relate to.

Where the Google leak fits in an SEO audit

We use the leak as a checklist of things Google bothers to store, then test each one with evidence we can see. Titles are checked against the queries in Search Console, because titlematchScore is described as how well titles match user queries. Off-topic sections are flagged, because siteRadius measures how far pages drift from the site. Thin pages are reviewed for originality, because OriginalContentScore exists specifically for pages with little content. Each check is useful on its own merits; the leak only tells us Google thought the same property was worth recording.

That is how our AI visibility audit uses the material: as a map of what Google measures, paired with the caveat that it is stored, not weighted. If you want to see the method applied to your own pages, the audit is the place to start, and the SEO audit guide shows the parts you can run yourself.

Frequently asked questions

Was there a Google data leak?

There have been several unrelated incidents, but in SEO 'the Google leak' means the Content Warehouse API documentation that was published to a public code repository on March 27, 2024 and removed on May 7, 2024. It exposed internal attribute names and descriptions from Google Search's storage systems. It did not expose user data or passwords.

Did Google confirm the leak was real?

Google did not dispute the documents. Spokesperson Davis Thompson said: 'We would caution against making inaccurate assumptions about Search based on out-of-context, outdated, or incomplete information.' Reports at the time treated that as confirmation that the documents were genuine.

Does the Google leak show ranking factor weights?

No. The documents list modules and attributes with short descriptions, but no scoring functions or weights. iPullRank's analysis states plainly that we do not know how features are weighted downstream. An attribute being stored does not prove it is used in ranking, or how much.

How big was the Google Search leak?

iPullRank counted 2,596 modules with 14,014 attributes, and SparkToro described it as more than 2,500 pages of API documentation. The package version we parsed, google_api_content_warehouse 0.4.0, contains 2,593 module files.

Does the leak prove Google uses clicks for ranking?

It shows Google stores click attributes such as goodClicks, badClicks and lastLongestClicks in modules named after NavBoost. Separately, the US v. Google antitrust case described NavBoost as a ranking component that uses 13 months of click data. The leak alone shows storage; the trial record shows use.

Who leaked the Google documents?

The documents were shared with SparkToro's Rand Fishkin by Erfan Azimi, an SEO practitioner. They had been published accidentally to a public repository as part of Google's Content Warehouse API client library, then removed.

OK

Olga Kunger

Founder & Lead Strategist, Ambeltek

Olga leads Ambeltek's web development, AI SEO, and GEO work — helping brands rank on Google and get cited by AI engines. More about Olga →

Want this done for your site?

AI SEO programs that rank on Google and earn AI citations. Growth tier from $9k, fixed scope before work starts.