Lorem ipsum placeholder text left in the page
What is this issue?
This issue checks whether your page contains placeholder text like "Lorem Ipsum" or other dummy content that indicates unfinished or draft content has been published to a live website.
A page passes this check if:
- It contains no placeholder text patterns
- All visible text is real, meaningful content written for users
Example of passing implementation
<p>Welcome to our company blog where we share industry insights and tips.</p>Example of failing implementation
<p>
Lorem ipsum dolor sit amet, consectetur adipiscing elit. Integer nec odio.
</p>Common placeholder text patterns checked
- "Lorem ipsum dolor sit amet"
- "Consectetur adipiscing elit"
- "Integer nec odio"
- "Praesent libero"
- "Sed cursus ante"
Why it matters
Having placeholder text on a live website is critical because:
- Rankings: Search engines will index your page for irrelevant Latin words instead of your actual content, hurting your SEO
- Indexability: Pages with placeholder text may be flagged as "thin content" or "auto-generated content" and de-indexed
- User experience: Users see meaningless text instead of helpful information, damaging trust and credibility
- Duplicate content: Lorem Ipsum text appears on thousands of websites, creating duplicate content issues
Impact on SEO health score
Resolving this issue is critical for your SEO health score. Publishing placeholder text to production is a serious issue that can result in search engines penalizing your entire website.
How to fix it
Step 1: Search your entire project
Use search tools to find all instances of "lorem ipsum", "dolor sit amet", or other placeholder text in:
- Your codebase files
- Your database content
- Your CMS templates
- Your design mockups
Step 2: Replace with real content
Write actual, meaningful content that:
- Accurately describes your product, service, or topic
- Provides value to your users
- Is written in the correct language for your audience
Step 3: Prevent future occurrences
- Add automated checks in your build process to detect placeholder text
- Train your content team to never publish placeholder text
- Use staging environments to review content before going live
- Add
<meta name="robots" content="noindex">to pages that are still in draft
Step 4: Check content that is in the markup but not on screen
Placeholder copy parked in a collapsed panel, an unused tab or a JS-driven clone is still in the HTML that a crawler downloads. PixyScan does not report it — a finding you cannot locate by looking at the page is one you cannot act on — but it is worth removing anyway, and it becomes reportable the moment the block it sits in is shown.
Example fix
<!-- Before: Placeholder text -->
<div class="hero">
<h1>Lorem ipsum dolor sit amet</h1>
<p>Consectetur adipiscing elit. Integer nec odio.</p>
</div>
<!-- After: Real content -->
<div class="hero">
<h1>Welcome to Our SEO Platform</h1>
<p>Boost your search rankings with our powerful SEO tools and analytics.</p>
</div>Examples
Example 1: Basic Lorem Ipsum detection
Scenario: A website launched with placeholder text in the hero section.
Problematic state (fails):
<section class="hero">
<h1>Lorem ipsum dolor sit amet</h1>
<p>Consectetur adipiscing elit. Integer nec odio.</p>
</section>Corrected state (passes):
<section class="hero">
<h1>Welcome to Our Company</h1>
<p>We provide innovative solutions for your business needs.</p>
</section>Example 2: Placeholder text a reader cannot see
Scenario: Placeholder copy parked in a collapsed panel and hidden in the markup itself.
Not reported (passes) — the check reads visible text, and this is not visible:
<div style="display: none;">
<p>Lorem ipsum dolor sit amet, consectetur adipiscing elit.</p>
</div>
<p>Real content here.</p>Reported (fails) — the placeholder is not hidden, only the block above it is:
<div style="display: none;">
<p>Draft notes, not shown.</p>
</div>
<p>Lorem ipsum dolor sit amet, consectetur adipiscing elit.</p>Only what the element's own markup declares — a hidden attribute, an inline
display: none / visibility: hidden — can be judged this way. A rule that arrives from
a stylesheet or a class needs a layout engine to resolve, so that text is still read as
visible and still reported.
Example 3: Partial replacement
Scenario: Some sections have real content, but others still have Lorem Ipsum.
Problematic state (fails):
<article>
<h2>Our Services</h2>
<p>We offer web development and SEO services.</p>
<p>Lorem ipsum dolor sit amet, consectetur adipiscing elit.</p>
</article>Corrected state (passes):
<article>
<h2>Our Services</h2>
<p>We offer web development and SEO services.</p>
<p>Contact us today to learn more about our solutions.</p>
</article>How PixyScan detects this
PixyScan scans your page's HTML to detect placeholder text through these steps:
Detection process
Fetch the page: PixyScan downloads the raw HTML of your page (no JavaScript execution)
Extract the visible text: The crawler reads the text of the whole
<body>, whatever markup it sits in — paragraphs, spans, divs, sections, articles, list items, headings, links, table cells and any depth of nesting. No tag is checked in preference to another, so how the page happens to be structured cannot hide the placeholder.Nav, header and footer are kept: a stub footer is exactly where placeholder copy tends to be left, and a reader sees it there. (The readability grade for the same page ignores those regions — a site-wide nav is not that page's prose.)
What is dropped is everything a reader never sees:
<script>,<style>,<noscript>and<template>— code and inert markup, not content- Anything the markup declares hidden: a
hiddenattribute,aria-hidden="true", or an inlinedisplay: none/visibility: hidden <iframe>fallback text and<svg><title>accessible names- HTML comments, and text that only exists in attributes such as
altortitle
Normalize the text: Whitespace is collapsed to single spaces and zero-width and soft-hyphen characters are removed, so line breaks, source indentation and invisible characters cannot sit between the two words and change the result.
Scan for the placeholder: The crawler matches "lorem" followed by "ipsum" with at most a few non-alphanumeric characters between them. The tolerance is what makes markup-split copy detectable —
<span>Lorem</span><span>ipsum</span>andLorem<br>ipsumboth extract asLoremipsum, with no separator surviving the element boundary — while still requiring the two halves to be adjacent, so a page that merely mentions both words far apart is not flagged.Flag the issue: The page fails this check if the placeholder is found anywhere in the visible text → STANDARD issue, raised once per page however many times it occurs, and recorded against that page.
Important limitations
- PixyScan only checks static HTML text—not content added by JavaScript
- If placeholder text is loaded dynamically via JavaScript, it won't be detected
- Only inline styles can be judged as hidden. A
display: nonethat comes from a stylesheet or a CSS class needs a layout engine to resolve, so that text is still read as visible
What we store
Storage Level
Page Level — This issue is evaluated for each individual page crawled.
Database Table / Prisma Model
ContentReadabilityData
Stored Fields
| Field | Type | Description |
|---|---|---|
| containsLoremIpsum | Boolean | Whether the page contains Lorem Ipsum placeholder text |
The finding itself
The reported issue is a separate row in AuditIssue, and one row is what a failing page gets — however many times the placeholder occurs on it. Repeating a placeholder block twenty times in a template is still one page to fix.
| Column | Value |
|---|---|
scanId |
The scan that read the page |
urlId |
The urls row for the page the text was read from |
issueCode |
no_lorem_ipsum |
details |
{ message, url } — the short message and that page own URL |
The urlId is resolved by the API from the URL the page was posted under, and it is what
the report joins on to name the affected page. details.url is the same page carried on
the finding itself, so a row listed against the wrong page is visible in the data rather
than something to be inferred from a report pointing at the site root.
Detection Dependencies
- The following data sources are required to evaluate this issue:
- HTML Document — The crawled page content is analyzed for Lorem Ipsum text patterns
- Page Content — Text content extracted from the page body