Technical SEO
Indexation Management And Crawl Control
Get the right pages indexed and keep the wrong ones out.
Index bloat is the quiet killer on large sites. When Google is spending its crawl allowance on filter permutations, session URLs and tag archives, the pages you actually care about get crawled less often and reindexed more slowly.
You probably need this if
- Far more pages are indexed than you have real content
- Important pages sit in discovered but not indexed for weeks
- New content takes a long time to appear in the index
- Search Console shows large numbers of duplicate or alternate pages
What you get
- Index inventory comparing what should be indexed against what is
- Crawl budget analysis from server log files
- Robots directives, meta robots and canonical strategy documented per template
- Index bloat remediation plan with the correct removal method for each case
- XML sitemap architecture split by template for cleaner diagnostics
How we approach it
-
Inventory the index
Establish what is actually indexed versus what should be, because the gap in both directions is usually larger than expected.
-
Read the logs
Server logs show where crawl budget genuinely goes, which is frequently URLs nobody knew existed.
-
Apply the right control
Noindex, robots disallow, canonical and 410 all do different things, and using the wrong one is the most common mistake in this area.
-
Segment sitemaps
Split sitemaps by template so indexation problems can be traced to a specific page type rather than guessed at.
Indexation Management: common questions
Should I use noindex or robots.txt disallow?
Noindex keeps the page crawlable but out of the index, which is right when a page must not rank. Robots disallow prevents crawling entirely and is right for saving crawl budget. Never combine them, because a disallowed page cannot be crawled and so its noindex is never seen.
What causes discovered but not indexed?
Google knows the URL but has not prioritised crawling it. On large sites this usually indicates crawl budget pressure or a quality assessment that the page is unlikely to be worth indexing. Improving internal linking and page quality is the fix, not resubmitting the URL.
How do I fix index bloat?
Identify the low-value pattern, then apply the right control. Noindex pages that must stay accessible to users, robots-disallow crawl traps, 410 pages that are genuinely gone, and canonical near-duplicates to the version you want ranked.
Related services
Find out what is actually holding you back
A short conversation and a look at your site is usually enough to tell you where the real problem is. No obligation, and no pressure to sign anything.