Skip to content
Issue docs

Lorem ipsum placeholder text left in the page

Standardno_lorem_ipsumIssue 3

What is this issue?

This issue checks whether your page contains placeholder text like "Lorem Ipsum" or other dummy content that indicates unfinished or draft content has been published to a live website.

A page passes this check if:

  • It contains no placeholder text patterns
  • All visible text is real, meaningful content written for users

Example of passing implementation

<p>Welcome to our company blog where we share industry insights and tips.</p>

Example of failing implementation

<p>
  Lorem ipsum dolor sit amet, consectetur adipiscing elit. Integer nec odio.
</p>

Common placeholder text patterns checked

  • "Lorem ipsum dolor sit amet"
  • "Consectetur adipiscing elit"
  • "Integer nec odio"
  • "Praesent libero"
  • "Sed cursus ante"

Why it matters

Having placeholder text on a live website is critical because:

  • Rankings: Search engines will index your page for irrelevant Latin words instead of your actual content, hurting your SEO
  • Indexability: Pages with placeholder text may be flagged as "thin content" or "auto-generated content" and de-indexed
  • User experience: Users see meaningless text instead of helpful information, damaging trust and credibility
  • Duplicate content: Lorem Ipsum text appears on thousands of websites, creating duplicate content issues

Impact on SEO health score

Resolving this issue is critical for your SEO health score. Publishing placeholder text to production is a serious issue that can result in search engines penalizing your entire website.

How to fix it

Step 1: Search your entire project

Use search tools to find all instances of "lorem ipsum", "dolor sit amet", or other placeholder text in:

  • Your codebase files
  • Your database content
  • Your CMS templates
  • Your design mockups

Step 2: Replace with real content

Write actual, meaningful content that:

  • Accurately describes your product, service, or topic
  • Provides value to your users
  • Is written in the correct language for your audience

Step 3: Prevent future occurrences

  • Add automated checks in your build process to detect placeholder text
  • Train your content team to never publish placeholder text
  • Use staging environments to review content before going live
  • Add <meta name="robots" content="noindex"> to pages that are still in draft

Step 4: Check content that is in the markup but not on screen

Placeholder copy parked in a collapsed panel, an unused tab or a JS-driven clone is still in the HTML that a crawler downloads. PixyScan does not report it — a finding you cannot locate by looking at the page is one you cannot act on — but it is worth removing anyway, and it becomes reportable the moment the block it sits in is shown.

Example fix

<!-- Before: Placeholder text -->
<div class="hero">
  <h1>Lorem ipsum dolor sit amet</h1>
  <p>Consectetur adipiscing elit. Integer nec odio.</p>
</div>

<!-- After: Real content -->
<div class="hero">
  <h1>Welcome to Our SEO Platform</h1>
  <p>Boost your search rankings with our powerful SEO tools and analytics.</p>
</div>

Examples

Example 1: Basic Lorem Ipsum detection

Scenario: A website launched with placeholder text in the hero section.

Problematic state (fails):

<section class="hero">
  <h1>Lorem ipsum dolor sit amet</h1>
  <p>Consectetur adipiscing elit. Integer nec odio.</p>
</section>

Corrected state (passes):

<section class="hero">
  <h1>Welcome to Our Company</h1>
  <p>We provide innovative solutions for your business needs.</p>
</section>

Example 2: Placeholder text a reader cannot see

Scenario: Placeholder copy parked in a collapsed panel and hidden in the markup itself.

Not reported (passes) — the check reads visible text, and this is not visible:

<div style="display: none;">
  <p>Lorem ipsum dolor sit amet, consectetur adipiscing elit.</p>
</div>
<p>Real content here.</p>

Reported (fails) — the placeholder is not hidden, only the block above it is:

<div style="display: none;">
  <p>Draft notes, not shown.</p>
</div>
<p>Lorem ipsum dolor sit amet, consectetur adipiscing elit.</p>

Only what the element's own markup declares — a hidden attribute, an inline display: none / visibility: hidden — can be judged this way. A rule that arrives from a stylesheet or a class needs a layout engine to resolve, so that text is still read as visible and still reported.

Example 3: Partial replacement

Scenario: Some sections have real content, but others still have Lorem Ipsum.

Problematic state (fails):

<article>
  <h2>Our Services</h2>
  <p>We offer web development and SEO services.</p>
  <p>Lorem ipsum dolor sit amet, consectetur adipiscing elit.</p>
</article>

Corrected state (passes):

<article>
  <h2>Our Services</h2>
  <p>We offer web development and SEO services.</p>
  <p>Contact us today to learn more about our solutions.</p>
</article>

How PixyScan detects this

PixyScan scans your page's HTML to detect placeholder text through these steps:

Detection process

  1. Fetch the page: PixyScan downloads the raw HTML of your page (no JavaScript execution)

  2. Extract the visible text: The crawler reads the text of the whole <body>, whatever markup it sits in — paragraphs, spans, divs, sections, articles, list items, headings, links, table cells and any depth of nesting. No tag is checked in preference to another, so how the page happens to be structured cannot hide the placeholder.

    Nav, header and footer are kept: a stub footer is exactly where placeholder copy tends to be left, and a reader sees it there. (The readability grade for the same page ignores those regions — a site-wide nav is not that page's prose.)

    What is dropped is everything a reader never sees:

    • <script>, <style>, <noscript> and <template> — code and inert markup, not content
    • Anything the markup declares hidden: a hidden attribute, aria-hidden="true", or an inline display: none / visibility: hidden
    • <iframe> fallback text and <svg><title> accessible names
    • HTML comments, and text that only exists in attributes such as alt or title
  3. Normalize the text: Whitespace is collapsed to single spaces and zero-width and soft-hyphen characters are removed, so line breaks, source indentation and invisible characters cannot sit between the two words and change the result.

  4. Scan for the placeholder: The crawler matches "lorem" followed by "ipsum" with at most a few non-alphanumeric characters between them. The tolerance is what makes markup-split copy detectable — <span>Lorem</span><span>ipsum</span> and Lorem<br>ipsum both extract as Loremipsum, with no separator surviving the element boundary — while still requiring the two halves to be adjacent, so a page that merely mentions both words far apart is not flagged.

  5. Flag the issue: The page fails this check if the placeholder is found anywhere in the visible text → STANDARD issue, raised once per page however many times it occurs, and recorded against that page.

Important limitations

  • PixyScan only checks static HTML text—not content added by JavaScript
  • If placeholder text is loaded dynamically via JavaScript, it won't be detected
  • Only inline styles can be judged as hidden. A display: none that comes from a stylesheet or a CSS class needs a layout engine to resolve, so that text is still read as visible

What we store

Storage Level

Page Level — This issue is evaluated for each individual page crawled.


Database Table / Prisma Model

ContentReadabilityData


Stored Fields

Field Type Description
containsLoremIpsum Boolean Whether the page contains Lorem Ipsum placeholder text

The finding itself

The reported issue is a separate row in AuditIssue, and one row is what a failing page gets — however many times the placeholder occurs on it. Repeating a placeholder block twenty times in a template is still one page to fix.

Column Value
scanId The scan that read the page
urlId The urls row for the page the text was read from
issueCode no_lorem_ipsum
details { message, url } — the short message and that page own URL

The urlId is resolved by the API from the URL the page was posted under, and it is what the report joins on to name the affected page. details.url is the same page carried on the finding itself, so a row listed against the wrong page is visible in the data rather than something to be inferred from a report pointing at the site root.


Detection Dependencies

  • The following data sources are required to evaluate this issue:
  • HTML Document — The crawled page content is analyzed for Lorem Ipsum text patterns
  • Page Content — Text content extracted from the page body

Further reading