SEO checklist
Robots.txt, sitemap, and indexing checklist
Use this checklist to review robots.txt rules, sitemap URL lists, robots meta tags, and obvious indexing blockers before publishing.
Start with what should be indexable
Before editing robots.txt or a sitemap, decide which pages should be available to search engines. A simple list of important pages makes the rest of the check much easier.
Robots.txt can guide crawlers, but it is not a privacy tool. If a page should not be public, fix access and publishing controls first.
- Review robots.txt rules.
- Check sitemap URLs for duplicates or staging domains.
- Confirm page-level robots meta tags do not conflict with the sitemap.
Quick workflow
Generate or review a simple robots.txt file, paste sitemap URLs into a checker, then inspect page-level robots tags when a page behaves differently than expected.
Do
- Keep sitemap URLs canonical and indexable.
- Use noindex deliberately for pages that should stay out of search.
- Test small changes before a large migration.
Don't
- Do not block important pages by accident.
- Do not put localhost or staging URLs in the sitemap.
- Do not rely on robots.txt to hide sensitive content.
FAQs
Should every page be in the sitemap?
No. Include important canonical pages that you want search engines to discover, not test pages, duplicates, or private URLs.
Can robots.txt remove a page from search results?
Not reliably. Use page-level noindex for public pages you do not want indexed, and use access controls for private content.