Skip to content
Issue docs

Session ID exposed in the URL

Importanturl_session_id_parameterIssue 120

What is this issue?

This check reports a page whose address carries a session identifier.

It looks for a fixed list of parameter names that mean nothing else -- PHPSESSID, JSESSIONID, ASPSESSIONID, sid, sessid, sessionid, session_id, CFID, CFTOKEN, zenid, osCsid -- in two places:

  • The query string: ?PHPSESSID=8f3c1a...
  • The path parameter form that Java servlet containers still emit when a client refuses cookies: /cart;jsessionid=8f3c1a...

Names that could legitimately mean something else are deliberately not on the list. A conference site's ?session=morning is a real parameter and is not reported.

Four of the names are judged on their VALUE as well, because they are also ordinary English: sid (store id, site id, series id) and sessionid, session_id, sessid -- which a conference agenda, a webinar platform or an LMS uses for a THING rather than a token. A session token is long and opaque, so ?sessionid=12 passes and ?sessionid=8f3c1ab29d4e7f60a1b2c3d4e5f60718 does not.

A parameter with an empty value is never reported, whatever its name: ?PHPSESSID= is a logged-out page that emitted the parameter with no session in it.

For a URL to pass:

  • It carries no session identifier; session state lives in a cookie.

Why it matters

A session ID in the URL gives every visitor, and every crawl, a different address for the same page. There is no upper bound on how many.

  • Unbounded duplicates: one page becomes an unlimited set of URLs, each of which a crawler may fetch and try to index.
  • Split signals: any link that is shared carries somebody's session ID with it, so link equity accumulates on an address that will never be visited again.
  • Crawl budget: a crawler can spend an entire allowance walking session variants of the same handful of pages.
  • Security: a shared link hands over a live session. Search results, referrer headers and browser history have all leaked session IDs this way.

This is graded important rather than standard, unlike the other four URL hygiene checks: the ranking cost is real and measurable, and it is the same cost that the parameterised-URL checks are graded for.

How to fix it

  1. Keep the session in a cookie. Every server platform supports cookie-based sessions; this is a configuration setting, not a rewrite. In PHP set session.use_only_cookies and disable session.use_trans_sid; in a Java container disable URL rewriting for sessions.

  2. Turn off the cookie-less fallback. The ;jsessionid= form exists only to support clients that refuse cookies. Almost nothing does any more, and crawlers are the main thing that triggers it.

  3. Redirect session URLs to the clean address. A 301 from any URL carrying a session parameter to the same URL without it stops the ones already indexed from persisting.

  4. Make the canonical clean. Pages served with a session parameter must carry a canonical pointing at the parameterless address.

  5. Check what is already indexed. Search for your domain with the parameter name; anything that appears needs the redirect above.

  6. Never put a session ID in a sitemap or an internal link.

Examples

Scenario: A cart page on a site with cookie-based sessions.

Passes because: the address carries no session identifier.

https://example.com/cart

Example 2: A session ID in the query string

Scenario: PHP with session.use_trans_sid left on.

Fails because: PHPSESSID is a session identifier, and every visitor gets a different address for this page.

https://example.com/cart?PHPSESSID=8f3c1ab29d4e7f60

Corrected version:

https://example.com/cart

Example 3: The path-parameter form

Scenario: A Java application falling back to URL rewriting because the crawler did not accept a cookie.

Fails because: the session identifier is in the path, where it is easy to miss.

https://example.com/cart;jsessionid=0A1B2C3D4E5F6789

Corrected version: disable session URL rewriting, and 301-redirect any address carrying ;jsessionid= to the clean path.

https://example.com/cart

How PixyScan detects this

  1. Parses the page's final address.

  2. Reads the query-string parameter names and compares each against a fixed list of names that only ever mean "session identifier". The comparison ignores case, so SessionId and sessionid are both found.

  3. Checks the value for the overloaded names. sid, sessionid, session_id and sessid are also ordinary English, so they count only when the value looks like a session token: sixteen or more characters of letters, digits, hyphen or underscore. The rest of the list is unambiguous and is matched on the name alone. Every name requires a non-empty value.

  4. Reads the path parameters too. Everything after a ; inside a path segment is a path parameter, and /cart;jsessionid=... never appears in the query string at all. A check that read only the query string would report nothing on exactly the sites where this is worst.

  5. Reports the parameter names it found, spelled as the URL spells them. The session VALUE is deliberately not recorded: it is a live credential.

Parameters whose names merely contain a session name as a substring -- ?consider=1, ?residency=uk -- are not matched.

What we store

Storage Level

Page Level


Database Table / Prisma Model

audit_issues.details


Stored Fields

Field Type Description
url String The page URL that was examined
sessionParams String[] The session parameter names found, as spelled

Detection Dependencies

  • HTTP Response
  • The page URL as the crawler fetched it

Note

The evidence is kept on the finding rather than in a per-page table. The URL is stored with every session value replaced by REDACTED, and only the parameter NAME is kept as written: the value is a working credential for somebody else’s session, and this row is both written to the database and rendered on screen.

Further reading