Skip to content
Issue docs

Sitemap file over the 50,000-entry or 50 MB limit

Importantsitemap_file_over_limitIssue 149

What is this issue?

The sitemaps.org protocol puts two hard ceilings on a single sitemap file:

  • at most 50,000 entries, and
  • at most 50 MB uncompressed.

Both apply to a <urlset> counting <url> elements and to a <sitemapindex> counting <sitemap> children.

This check reports a file past either ceiling.

For a sitemap to pass this check:

  • Every file it publishes is inside both limits.

The remedy is the same for both, which is why they are one check: split the file, and list the parts from a sitemap index. An index can itself list 50,000 children, so two levels cover 2.5 billion URLs.

Example: a marketplace generates one sitemap.xml containing every listing. It passes 80,000 entries. Search engines stop reading the file, and the 30,000 listings past the ceiling are not the only ones affected — the file is rejected whole.

Why it matters

An oversized file is rejected, not truncated. This is the part that surprises people. Search Console reports the whole file as an error; it does not process the first 50,000 and discard the rest. Every page in the file loses its sitemap-based discovery at once.

Discovery falls back to internal links alone. Pages that are well linked survive. Pages that are deep in a catalogue, newly published, or reachable only from a paginated listing are the ones that quietly stop being found — which is exactly the set a sitemap exists to help.

The 50 MB limit is measured uncompressed. Serving sitemap.xml.gz does not raise it. A compressed file that unpacks to 60 MB is over the limit.

Nothing else reports it. This is why it is graded IMPORTANT rather than STANDARD: there is no other check whose weight already covers the harm, and the failure is mechanical rather than a matter of degree.

Fixing it restores sitemap-based discovery for every page in the file.

How to fix it

  1. Split the file. Break it into several <urlset> files, each comfortably inside both ceilings — 10,000 to 20,000 entries per file is a common working size and leaves room to grow between deploys.

  2. Publish a sitemap index that lists them.

    <?xml version="1.0" encoding="UTF-8"?>
    <sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
      <sitemap><loc>https://example.com/sitemap-products-1.xml</loc></sitemap>
      <sitemap><loc>https://example.com/sitemap-products-2.xml</loc></sitemap>
      <sitemap><loc>https://example.com/sitemap-pages.xml</loc></sitemap>
    </sitemapindex>
  3. Split by section, not by arbitrary batches, where you can. One file per content type makes Search Console's per-file coverage numbers answer a question you actually have.

  4. Point robots.txt at the index, and submit the index in Search Console. Individual child files do not need submitting separately.

  5. If the byte limit is what you hit, look at what is in each entry. Image and video extensions, long <lastmod> values and generous whitespace all add up; a file can pass the 50,000 count and still be over 50 MB.

  6. Remember that the index has the same ceilings. If you need more than 50,000 child files, nest a second level of indexes.

Examples

Example 1 — one file over the entry limit

Problematic:

<!-- https://example.com/sitemap.xml — 82,431 <url> elements -->
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url><loc>https://example.com/listing/1</loc></url>
  ...
  <url><loc>https://example.com/listing/82431</loc></url>
</urlset>

Corrected: five files under an index.

<!-- https://example.com/sitemap.xml -->
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <sitemap><loc>https://example.com/sitemap-listings-1.xml</loc></sitemap>
  <sitemap><loc>https://example.com/sitemap-listings-2.xml</loc></sitemap>
  <sitemap><loc>https://example.com/sitemap-listings-3.xml</loc></sitemap>
  <sitemap><loc>https://example.com/sitemap-listings-4.xml</loc></sitemap>
  <sitemap><loc>https://example.com/sitemap-listings-5.xml</loc></sitemap>
</sitemapindex>

Example 2 — inside the entry limit, over the byte limit

A file with 30,000 entries, each carrying several <image:image> blocks with titles and captions, reaches 68 MB.

Problematic: the entry count looks fine, so nothing obvious is wrong.

Corrected: split it the same way. Serving it gzipped does not help — the 50 MB ceiling is measured on the uncompressed document.


Example 3 — the boundary

  • 50,000 entries: passes.
  • 50,001 entries: reported.
  • exactly 52,428,800 bytes (50 MB): passes.
  • one byte more: reported.

How PixyScan detects this

This is answered from the sitemap alone — no page has to be fetched — so it runs during the site-level pass, as soon as the sitemap has been walked.

  1. Fetch each sitemap file. PixyScan follows the sitemap from robots.txt and from the conventional locations, through nested indexes.

  2. Count what each file declares. For a <urlset> this is the number of <url> elements in the document; for a <sitemapindex> it is the number of <sitemap> children. The count is of what the file declares, not of what PixyScan read — the walk stops reading entries part-way through a very large file, and reporting that truncated number would show a 60,000-URL sitemap as a small one.

  3. Measure each file's size in bytes, uncompressed, as the protocol specifies. Bytes, not characters: a path in Cyrillic, Japanese or accented Latin is two or three bytes per character, so counting characters would under-measure exactly the large international sitemaps this check is aimed at.

  4. Report every file past 50,000 entries or past 50 MB, naming which ceiling each one crossed and by how much. Both limits are inclusive: exactly 50,000 entries and exactly 50 MB pass.

What is deliberately never reported:

  • A file that could not be read. A sitemap that 404s or times out was not measured, and "we could not fetch it" is neither under the limit nor over it. That is a separate finding.

What we store

Storage Level

Site Level — the finding is about the sitemap, not about any one page in it, so it is raised once per scan with the offending entries listed on it.


Database Table / Prisma Model

AuditIssue (url_id null), read from the sitemap walk, and the file tree stored on SiteCrawlBehaviourData.


Fields Used

Field Type Description
SiteCrawlBehaviourData.sitemapFileTree Json Every sitemap file the walk found, and which index listed it
(walk) file.entryCount Number <url> or <sitemap> elements the file declares
(walk) file.byteSize Number The file's uncompressed size in bytes
(walk) file.fetchedOk Boolean False when the file could not be read, so it is not judged

Detection Dependencies

  • XML Sitemap — every file, including nested index files
  • HTTP Response — the body of each sitemap file, whose byte length is the measurement

Further reading