What Is robots.txt?
A plain text file that tells search engine crawlers which pages to visit and which to skip.
Every time Google sends a crawler to your site, it checks one file before it does anything else. That file is robots.txt. It sits in the root of your domain, it takes about two seconds to read, and it can have a significant effect on what search engines see and what they ignore.
Most site owners never look at it. A surprising number have rules in theirs they didn’t write and don’t know about. This guide explains what robots.txt does, how to read it, and what mistakes to avoid.
What robots.txt Actually Does
robots.txt is a plain text file that tells web crawlers which parts of your site they’re allowed to visit. That’s the whole job. It’s part of a standard called the Robots Exclusion Protocol, which dates back to 1994 and is supported by every major search engine including Google, Bing, and DuckDuckGo.
When Googlebot arrives at your domain, the first request it makes is for yourdomain.com/robots.txt. It reads the file, checks whether any rules apply to it, and then behaves accordingly before it crawls anything else.
The key word there is “behaves accordingly.” robots.txt is a request, not a lock. Reputable crawlers like Googlebot follow it. Malicious scrapers, spam bots, and badly written scripts often don’t. So robots.txt is not a security measure. It won’t stop someone from accessing a page, and it won’t protect sensitive information. For that, you need proper authentication or access controls.
How to Read a robots.txt File
The file uses a simple syntax. Here’s a typical example:
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
User-agent: Googlebot
Disallow: /staging/
Sitemap: https://yourdomain.com/sitemap.xml
Breaking that down line by line:
User-agent: * means “this rule applies to all crawlers.” The asterisk is a wildcard. If you want a rule to apply to a specific crawler only, you’d write its name instead, like User-agent: Googlebot.
Disallow: /wp-admin/ tells crawlers not to visit anything under the /wp-admin/ path. That includes /wp-admin/edit.php, /wp-admin/users.php, and every other page in that folder. This is standard on WordPress sites because the admin area has no business appearing in search results.
Allow: /wp-admin/admin-ajax.php creates an exception to the rule above. The admin-ajax.php file handles dynamic requests on the frontend of many WordPress sites. Blocking it can break things, so the Allow line overrides the broader Disallow.
Sitemap: is not a directive about what to block. It’s a pointer telling crawlers where your XML sitemap lives. Google doesn’t require it here since you can submit your sitemap via Search Console, but including it doesn’t hurt.
What You Can and Cannot Control With robots.txt
robots.txt controls crawling, not indexing. That distinction matters more than most guides make clear.
If you block a page with Disallow, you stop crawlers from visiting it. But you don’t stop it from appearing in search results. Google can still index a page it has never crawled if another site links to it. In that case, Google might show the page in results with a message like “No information is available for this page.” The URL exists in Google’s index. The content doesn’t.
To prevent a page from appearing in search results entirely, you need a noindex tag in the page’s HTML, not a robots.txt rule. The two tools do different jobs and work at different levels.
There’s also a common mistake that combines both and produces the opposite of what’s intended: blocking a page with robots.txt and adding a noindex tag. If Googlebot can’t crawl the page, it can’t read the noindex tag. The block prevents the instruction from being seen. The page ends up in an awkward middle state where Google knows it exists but can’t confirm it should be excluded.
What Should Actually Be in Your robots.txt
For most WordPress sites, the defaults are fine. WordPress generates a virtual robots.txt automatically if you don’t have a physical file, and it blocks /wp-admin/ while allowing admin-ajax.php. If you’re using a plugin like Yoast SEO or The SEO Framework, one of them is likely managing the file for you.
Things that are reasonable to block:
- Admin and login pages (
/wp-admin/,/wp-login.php): not for security, but because they add no SEO value and waste crawl budget - Staging subdirectories or subdomains if they’re publicly accessible
- Internal search result pages (
/search/,/?s=): these create near-duplicate content and give crawlers nothing useful - Tag and date archive pages on WordPress if they’re generating thin content
- Cart and checkout pages on WooCommerce stores
- Duplicate parameter URLs if your site generates them (
?sort=,?ref=)
Things you should not block:
- CSS and JavaScript files: Google needs to render your pages properly, and blocking these can hurt how your site is understood and ranked
- Images, unless they’re irrelevant to search (blocking images removes them from Google Image Search)
- Your main content pages, obviously, but it’s more common than you’d think for a misconfigured rule to catch pages it wasn’t meant to
Crawl Budget and Why It Matters for Larger Sites
Crawl budget is the number of URLs Googlebot is willing to crawl on your site within a given timeframe. For small sites with a few dozen pages, it’s not something you need to think about. Google will crawl everything.
For larger sites with thousands of pages, it starts to matter. If Googlebot is spending time on low-value URLs like filtered search results, session IDs, or thin archive pages, it has less capacity left for your important content. robots.txt is one tool for steering crawlers away from pages that don’t need to be visited, which frees up budget for pages you want indexed.
This is less relevant for most site owners reading this. But if you’re running a large WordPress site with lots of automatically generated pages, it’s worth reviewing what crawlers are spending their time on via Google Search Console’s crawl stats report.
Where Your robots.txt File Is and How to Check It
Your robots.txt file lives at the root of your domain. You can check it right now by visiting yourdomain.com/robots.txt in a browser. It’ll either show you a plain text file or a 404 error, which means one doesn’t exist.
On WordPress, if no physical file is present, WordPress generates a virtual one on the fly. Most SEO plugins let you edit it from the plugin settings without touching files directly. If you’re on shared hosting with cPanel, you can create or edit the file via the File Manager. On managed WordPress hosts, the process varies, but most have an interface for it.
Google also has a robots.txt tester in Google Search Console. It lets you check whether a specific URL would be blocked or allowed based on your current rules. It’s useful for verifying that a new rule does what you intended before you deploy it.
Common Mistakes to Avoid
The most dangerous mistake is accidentally blocking your entire site. It looks like this:
User-agent: *
Disallow: /
That single rule tells every crawler not to visit any page on your domain. It’s often added to staging environments to prevent premature indexing, which is correct in that context. The problem is when it gets copied to the live site during a migration or launch and nobody notices until rankings drop. It’s one of the more painful mistakes in web publishing because it can take weeks to recover from once you fix it.
The second common mistake is treating robots.txt as a privacy tool. Blocking a page with robots.txt while leaving it publicly accessible does not make it private. If anything, listing a path in robots.txt can draw attention to it. Some security researchers and scrapers actively read robots.txt to find paths worth investigating.
The third is forgetting to update it after a site restructure. If you block /old-section/ and then rename that section to /new-section/, the old rule does nothing and the new section is crawled freely. Rules need maintenance like everything else.
Frequently Asked Questions
Does robots.txt affect my Google rankings?
Indirectly, yes. Blocking pages that waste crawl budget can help Google focus on your important content. Accidentally blocking pages you want indexed will remove them from search results. The file itself is not a ranking signal, but what it allows or prevents crawlers to see has a direct effect on indexing.
What happens if I don’t have a robots.txt file?
Crawlers will proceed as if no restrictions exist and visit whatever they can reach on your site. On WordPress, a virtual robots.txt is generated automatically. On other platforms, the absence of the file is usually fine for small sites where you want everything crawled and indexed.
Can I block specific bots but allow Googlebot?
Yes. You can have multiple User-agent blocks in one file. For example, you can block AhrefsBot or SemrushBot while leaving a general allow for other crawlers. Each User-agent block applies only to the crawler named in it.
Is robots.txt the same as a meta robots tag?
No. robots.txt controls which URLs crawlers visit. A meta robots tag (like <meta name="robots" content="noindex">) controls how a page is treated in search results once it has been crawled. They work at different levels. Use robots.txt to manage crawling, and meta robots tags to manage indexing.
Should my sitemap be in my robots.txt?
It’s good practice but not required. Including a Sitemap: directive in robots.txt helps crawlers find your sitemap without needing to be told about it separately. You can also submit your sitemap directly via Google Search Console, which is more reliable.
If you want to check your site’s response time or run a quick DNS lookup after adjusting how your domain is configured, the Server Response Tester and DNS Lookup tool are both free to use. For more on how search engines interact with your hosting setup, see the guide on does web hosting affect SEO.