Article
6 minute readHallucinated Links in AI Answers: Diagnose and Fix Broken URLs
Hallucinated links in AI answers can point to pages that never existed. Find the origin, triage each URL and redirect only to verified equivalents.
Hallucinated links are a common side effect of AI answers. An answer may cite a real page, an outdated URL or a plausible address that never existed on your site. Before you add a redirect, you need to know which case you face and whether the destination really answers the cited question.
This guide shows how to capture the link, find where it came from, triage it with a simple sheet and repair only what deserves repair. As a result, readers who click a broken citation reach an accurate answer instead of a confusing page.
Preserve the complete link first
Start by saving the evidence exactly as you found it. Record the full URL, the answer context, the timestamp and the platform. Keep query strings and fragments in the original record, because they often explain the failure.
Then check the response and the redirect chain with a suitable HTTP tool or a browser. Distinguish a truly missing page from a temporary failure, an access restriction or a client-specific issue.
One failed fetch does not prove a URL is permanently broken. Therefore, test it more than once before you decide anything.
Find where hallucinated links come from
A link in an AI answer that returns a 404 has only a few possible origins. Each leaves different traces, and working out which one applies decides whether the fix is yours at all:
- Moved or removed page: the page existed and changed address, which your redirect history, old sitemaps or a web archive will show.
- Assembled from patterns: the system learned that integrations often live at
/integrations/{name}/, so the path looks like yours but never existed. - Different site's structure: your domain attached to a competitor's path, which is rare and usually a one-off.
- Refused by your server: a bot rule, geo-restriction or rate limit blocked the fetching client while browsers see the page fine.
- Junk in the URL: tracking parameters, trailing punctuation or an encoded fragment broke an otherwise correct link.
Only the first and fourth cases are defects on your side. A pattern guess deserves a redirect only when it maps to a real page and recurs. Meanwhile, tolerant URL parsing can absorb junk, and a different site's structure is nobody's task.
Classify each broken link
Once you know the origin, choose the matching action. The table below covers the most common cases. Above all, never redirect every unknown URL to the homepage, because a successful status that fails the reader's task is not a fix.
| Case | Action to investigate |
|---|---|
| Former page with a clear successor | Consider a relevant redirect |
| Typo or fabricated path | Check recurrence and whether a genuine equivalent exists |
| Valid page blocked for the client | Review access rules and bot handling |
| Working link with unrelated content | Audit whether the citation supports the answer |
Worked example: a guessed integration URL
Suppose a hypothetical answer links to /integrations/calendar-sync/, but your real documentation lives at /help/calendar-connection/. First, verify that both names refer to the same integration. Then check that the documentation covers the claim the answer made.
If the mismatch recurs and the destination is a genuine equivalent, a deliberate redirect may help affected users. However, if the answer invented an integration you do not offer, a redirect to a vaguely related page would hide the misinformation instead of fixing it.
In that case, the honest response is a clear public statement of what the product does. A reader who lands on a page that never mentions the promised feature will simply conclude your site is confusing.
Triage reported links in one sheet
Keep one row per distinct broken URL, not per report, and count the reports. The columns below give you enough to decide the action without investigating the same URL twice:
url | first_seen | reports | status_served | archive_shows_page | closest_real_page | equivalence | action | owner
/integrations/calendar-sync/ | 2026-09-02 | 4 | 404 | never existed | /help/calendar-connection/ | same feature, confirmed | redirect | docs
/pricing/enterprise-2024 | 2026-09-05 | 1 | 404 | existed until 2025-03 | /pricing | superseded | redirect | web
/help/offline-mode | 2026-09-07 | 3 | 404 | never existed | (none) | feature does not exist | no redirect; correct the claim | product
/docs/api?ref=ai | 2026-09-08 | 2 | 403 | current page | /docs/api | same page | fix the block rule | infraThe equivalence column is where judgment lives. "Same feature, confirmed" means someone read both the answer's claim and the candidate page and agreed the page answers what the answer promised.
Without that check, a redirect can complete the request while sending the reader to the wrong thing. So never fill this column automatically.
Look for patterns in your logs
Your server logs show which missing paths are requested again and again. Review repeated requests, referrer information where available and the affected topic. Then separate internal tests and obvious automated noise from plausible user demand.
A missing path can reveal a real documentation need. Still, it does not prove your product supports the feature the URL suggests. Confirm both the need and the facts before you create a new page.
Reduce new hallucinated links
You cannot stop external systems from assembling URLs. However, you can make your real URLs easier to find than the guessed ones. The measures are ordinary site hygiene, applied consistently:
- One canonical URL per topic: make it discoverable from navigation, the sitemap and internal links with descriptive anchor text.
- Stable, predictable paths: keep documentation paths when you reorganize, and redirect old addresses to new ones.
- A useful 404 page: list the closest real pages, search the site for the path's words and log the request with its referrer.
- A monthly 404 review: check recurring paths from verified external fetchers, because they are the earliest sign of a new guess pattern.
Then retest the original answers on their own schedule. A repaired site and a corrected answer are two separate outcomes, so report each one at its own strength.
Verify every repair
After a fix, check the final response, the relevance of the destination and the navigation. Also record why each redirect exists, so a future cleanup does not remove it by mistake.
Remember that fixing your site does not force an external system to stop citing the old address. Therefore, count a corrected answer only when it appears across several runs. The real goal is a reader who reaches an accurate answer with minimal confusion, not just a lower 404 count.
Track hallucinated links in a worksheet
Copy the worksheet columns below into a spreadsheet to track hallucinated links, and keep one row per captured URL. The filled row is an illustrative example, not a customer result, so replace it with your own records.
| Captured URL | HTTP result | Answer claim | Equivalent destination | Recurrence | Decision | Verification |
|---|---|---|---|---|---|---|
| /integrations/calendar-sync/ | Check required | Calendar connection | Verify actual guide | Unknown | Investigate | Pending |
Use AI to classify captured URLs
A model can sort a long list of captured links quickly, as long as you supply the response checks and page content. Use the prompt below only after you add those records.
Classify these captured AI citation URLs using supplied response checks and page content. Return missing, redirected, blocked, temporary failure or working. Recommend a redirect only when a verified equivalent destination exists.Then confirm every proposed redirect yourself. The model can suggest candidates, but only a person should approve an equivalence.
Conclusion
Hallucinated links need diagnosis before repair. Capture the full URL, find its origin, classify it, and redirect only when a verified equivalent page exists.
In short, the aim is an accurate answer for the reader, not a smaller error count. Open your 404 log this week, pick the most repeated path from an external fetcher and run it through the triage sheet.
Sources
These sources informed the research for this guide. The checklist, examples and workflow are independently written and are not results of a SEOVision experiment.
- New Study: How Often Do AI Assistants Hallucinate Links? (16 Million URLs Studied): research starting point, not an endorsement of this workflow
- 67% of ChatGPT’s Top 1,000 Citations Are Off-Limits to Marketers (+ More Findings): research starting point, not an endorsement of this workflow
Frequently asked questions
Quick answers to the questions readers ask most about this topic.
Why do AI answers link to pages that don't exist?
Some links point to moved or removed pages. Others are assembled from common URL patterns, so they look plausible but never existed on your site.
Should I redirect every broken AI citation?
No. Redirect only when a verified equivalent page exists and the path recurs. Never send unknown URLs to the homepage.
What if the AI answer invents a feature in the URL?
Do not redirect it to a vaguely related page. Publish a clear, accurate statement of what your product does instead.
Will fixing my site change the AI answer?
Not necessarily. Repairing your site and changing an external answer are separate outcomes, so retest the answer across several runs.
Sources
These references support the platform guidance discussed above. Worked examples are illustrative unless identified as measured results.
Keep reading