Sitemap lastmod dated in the future
What is this issue?
<lastmod> says when a page last changed. This check reports entries whose <lastmod> is
dated after the crawl — a date on which the page cannot yet have been edited.
For a sitemap to pass this check:
- No entry's
<lastmod>is more than a day ahead of the present.
The day of slack is deliberate. <lastmod> may be a date with no time —2026-09-05 —
which is read as midnight UTC, so a site publishing "today" from a timezone well ahead of
UTC legitimately writes tomorrow's date for part of its working day. A value inside 24
hours is a clock; a value beyond it is a wrong date.
Example: a CMS stores a "scheduled publish" timestamp alongside the real modification time, and the sitemap generator reads the wrong one. Every article queued for next month is dated next month.
Why it matters
<lastmod> is the one sitemap hint Google says it uses — and it uses it only while it
believes it. Google's own guidance is that it ignores the value when a site's dates are
consistently inaccurate.
The penalty is site-wide, not per entry. Once the values stop being trusted, they stop being trusted for the whole sitemap. One generator bug costs every page the recrawl prioritisation the field exists to provide — including the pages whose dates are perfectly correct.
Recrawl scheduling is what you lose. On a large site the practical effect is that genuinely updated pages take longer to be re-fetched, because nothing distinguishes them from the rest of the file.
Why this is graded STANDARD. The effect is real and site-wide, but indirect: no page is deindexed, no page becomes uncrawlable, and the fix is a one-line change in the generator.
Fixing it makes the field worth having.
How to fix it
Read the list on the finding — each entry names the URL and the
<lastmod>exactly as the sitemap declares it.Find which timestamp the generator is reading. A future date almost always means it has picked up a scheduled-publish field, an "expires" field, or a placeholder default rather than the page's actual modification time.
Emit the real modification time, and only when you have one.
<lastmod>is optional; omitting it is far better than declaring a value you cannot stand behind.Write a full W3C datetime with a timezone —
2026-09-04T11:00:00+02:00— or a plain date. A time with no timezone is technically invalid and is read differently by different consumers.Do not update
<lastmod>on every build. Regenerating it as "now" for every page on every deploy is the other way sites lose Google's trust in the field: it stops distinguishing anything.Check the server clock if the dates are only slightly ahead. A drifting clock on the machine generating the sitemap produces exactly this.
Examples
Example 1 — a scheduled-publish date leaking into lastmod
Problematic: the crawl runs on 4 September 2026.
<url>
<loc>https://example.com/blog/october-preview</loc>
<lastmod>2026-10-01</lastmod>
</url>Corrected: emit the real modification time, or nothing at all.
<url>
<loc>https://example.com/blog/october-preview</loc>
<lastmod>2026-08-22T09:14:00+00:00</lastmod>
</url>Example 2 — a timezone, not a fault
The crawl runs on 4 September 2026 at 12:00 UTC. A site in Auckland publishes on 5 September local time and writes the date it sees.
<lastmod>2026-09-05</lastmod>Not reported. Midnight UTC on the 5th is twelve hours ahead of the crawl, well inside the day of tolerance.
Example 3 — the boundary
With a crawl at 2026-09-04T12:00:00Z:
2026-09-05T12:00:00Z— exactly 24 hours ahead: passes.2026-09-05T13:00:00Z— 25 hours ahead: reported.2026-09-05T16:00:00+05:30— 10:30 UTC, 22.5 hours ahead: passes.2026-09-05T09:00:00-06:00— 15:00 UTC, 27 hours ahead: reported.
How PixyScan detects this
Answered from the sitemap alone, during the site-level pass.
Read the
<lastmod>of every entry the walk parsed.Interpret it as an instant. A date-only value is midnight UTC. A value with an offset —
2026-09-05T16:00:00+05:30— is the instant it names, not the wall clock it shows. A value carrying a time but no timezone is invalid under the protocol's own datetime profile and is read as UTC, so the answer does not depend on where the crawler happens to be running.Compare it with the time of the crawl, plus 24 hours of tolerance. The tolerance exists because a date-only value from a site in a far-eastern timezone is legitimately "tomorrow" in UTC for part of the day; the largest real offset is +14, so a day covers every timezone with room to spare.
Report every entry past that, once per page however many files list it, naming the URL and the declared value.
What is deliberately never reported:
- An entry with no
<lastmod>. The field is optional and omitting it is not a fault. - A value that is not a date at all.
yesterday, an empty element, or a malformed timestamp is not evidence of a future date.
A limit worth knowing: PixyScan reads up to 5,000 entries from each sitemap file, so on a very large file the list is a floor rather than a complete census.
What we store
Storage Level
Site Level — the finding is about the sitemap, not about any one page in it, so it is raised once per scan with the offending entries listed on it.
Database Table / Prisma Model
AuditIssue (url_id null), read from the sitemap walk; the per-page value is also kept on Url.sitemapLastmod.
Fields Used
| Field | Type | Description |
|---|---|---|
| (walk) entry.loc | String | The <loc> as declared |
| (walk) entry.lastmod | String? | The <lastmod> verbatim, kept as text so the site's own wording survives |
| Url.sitemapLastmod | String? | The same value stored against the page, for the sitemap explorer |
Detection Dependencies
- XML Sitemap — the
<lastmod>element of each entry - The crawl timestamp — the moment the comparison is made against