Configure WordPress robots.txt correctly

robots.txt is a small text file with enormous power over your site's SEO. A single misplaced line — or worse, a forgotten setting — can make Google stop crawling your entire site. Site owners have accidentally hidden their homepage from Google and watched traffic plummet, not realizing a checkbox in WordPress Settings was the culprit. This guide explains how robots.txt actually works, how WordPress handles it, and exactly what to allow and disallow so your site stays visible in search results.

1. What is robots.txt and Why Does It Matter?

robots.txt is a plain text file placed at the root of your domain (e.g., yourdomain.com/robots.txt) that communicates instructions to web crawlers — primarily Googlebot, Bingbot, and other search engines. The file tells bots which sections of your website they're permitted to crawl and which they should skip. Think of it as a digital "do not disturb" sign for your server.

The syntax is straightforward and human-readable, using simple directives like User-agent (which bot), Disallow (forbidden paths), and Allow (exceptions). Each rule is one line, and the file is interpreted from top to bottom.

Why does this matter for SEO? Because search engines have finite crawl budgets. Google's crawler has limited resources and time to crawl your site each day. If you waste that budget on crawling duplicate content, old cache files, or temporary pages, Googlebot has less time to crawl your important pages — your product pages, blog posts, and revenue-generating content. A well-configured robots.txt acts as a traffic director, pointing bots toward content that matters and away from the noise. This improves crawl efficiency and indirectly helps rankings.

2. How WordPress Handles robots.txt (Virtual vs. Physical Files)

Since WordPress 5.3, the platform has generated a virtual robots.txt file automatically. This is fundamentally different from most other applications. When you visit yoursite.com/robots.txt, there is often no actual file on your server — instead, WordPress generates the content dynamically at request time. This approach gives WordPress developers flexibility without requiring site owners to manually manage a file.

The default WordPress virtual robots.txt typically looks like this:

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://yourdomain.com/sitemap.xml

This default behavior is intentional and protective: WordPress blocks crawler access to the admin panel (since there's no public content there) but allows the specific AJAX endpoint that the front-end may use for form submissions and dynamic content. The Sitemap directive tells bots where to find your content map.

However, if you need to customize this beyond the defaults — for example, to block search results pages, cache directories, or plugin assets — you have two options: install an SEO plugin that manages robots.txt for you (creating a physical file in the process), or upload a physical robots.txt file via FTP. Once a physical file exists on the server, WordPress stops generating the virtual version and uses your file instead.

3. The #1 WordPress robots.txt Mistake

Every SEO specialist has seen this happen: a site owner accidentally checks "Discourage search engines from indexing this site" in WordPress Settings → Reading. This single checkbox changes the entire robots.txt output to:

User-agent: *
Disallow: /

The line Disallow: / is the nuclear option. It tells every search engine bot: "You are not welcome anywhere on this site." Googlebot, Bingbot, and all others will respect this and won't crawl any page. Your site will vanish from Google search results within days.

This setting is intended for staging sites or development environments — places you don't want search engines to discover yet. Unfortunately, it's sometimes enabled during development, the site goes live, and the developer forgets to disable it. Many sites have lost months of search traffic this way.

Critical Check: Go to WordPress Admin → Settings → Reading and verify that "Discourage search engines from indexing this site" is unchecked. If it's checked and your site is live, disable it immediately, then resubmit your site to Google Search Console to request re-indexing.

4. Three Ways to Edit or Create robots.txt

Method 1: Using Yoast SEO (Recommended for Most Users)

  1. Install and activate the Yoast SEO plugin (Free version works)
  2. Navigate to Yoast SEO → Tools → File editor
  3. Select the robots.txt tab
  4. Edit the robots.txt content directly in the web interface
  5. Click Save Changes

When you save, Yoast creates a physical robots.txt file on your server at the WordPress root. This file now takes precedence over WordPress's virtual version. You can edit it whenever you need without touching FTP, and Yoast provides helpful UI hints about best practices.

Method 2: Rank Math SEO

Rank Math is another popular WordPress SEO plugin with a similar workflow. Go to Rank Math → Tools → File Manager → robots.txt. Edit and save. It works the same way as Yoast — creates a physical file that overrides the virtual one.

Method 3: Upload robots.txt Manually via FTP (For Direct Control)

  1. Create a plain text file named robots.txt (no .txt.txt or extensions)
  2. Add your robots.txt rules to it (see the examples below)
  3. Connect to your server via FTP or File Manager
  4. Navigate to your WordPress root directory (the same folder containing wp-config.php)
  5. Upload the robots.txt file
  6. WordPress will now use your physical file instead of generating a virtual one

This method gives you the most control. You own the file and can edit it any time by re-uploading a new version. Note: when you upload via FTP to AsiaGB hosting, use the root path (e.g., public_html/robots.txt) — do not nest it in a subdirectory.

5. Recommended robots.txt Rules for WordPress

Standard WordPress Setup

User-agent: *
Disallow: /wp-admin/
Disallow: /wp-login.php
Disallow: /wp-includes/
Disallow: /wp-content/cache/
Disallow: /wp-content/plugins/
Disallow: /?s=
Disallow: /search/
Disallow: /tag/
Allow: /wp-admin/admin-ajax.php
Allow: /wp-content/uploads/
Allow: /wp-content/themes/

Sitemap: https://yourdomain.com/sitemap.xml

This configuration balances security and crawl efficiency. It blocks the admin panel, login page, and WordPress core files — none of which are useful for users or search engines to find in results. It also blocks search result pages and tag archives to prevent duplicate content issues. The Allow directives ensure that images in uploads and theme CSS/JS are still crawlable (important for image indexing and rendering).

For WooCommerce Sites

E-commerce sites have additional pages that shouldn't be indexed, particularly checkout and cart pages (they change constantly and create duplicate URLs with query parameters). Add these Disallow rules to the standard set:

Disallow: /cart/
Disallow: /checkout/
Disallow: /my-account/
Disallow: /order-pay/
Disallow: /order-received/
Disallow: /orders/
Disallow: /?add-to-cart=
Disallow: /?product_id=
Disallow: /?page_id=

WooCommerce generates many temporary pages during checkout. By blocking them in robots.txt, you preserve crawl budget for your actual product pages and category pages, which matter for SEO.

6. Should You Block wp-content/uploads or wp-content/themes?

This is where many WordPress sites get it wrong. Some site owners think "I'll block /wp-content/ entirely to save crawl budget." This is a mistake. Here's why:

Allow /wp-content/uploads/: Your images live in this folder. If bots can't crawl images, Google can't index them, and you miss out on image search traffic. Additionally, when Google tries to render your site (which it does to understand CSS), it needs to load CSS files. If the CSS is blocked, Google sees a broken layout and may score your page lower for mobile-friendliness. Always allow uploads and themes.

/wp-content/plugins/: Plugin files are generally safe to block since they're not user-facing content. Most plugins that matter for search (like Yoast) have critical files in uploads/ anyway. Blocking /wp-content/plugins/ is optional but low-risk.

7. How to Verify robots.txt is Working Correctly

Using Google Search Console

Google Search Console (GSC) provides a dedicated robots.txt tester. Here's how to use it:

  1. Log in to Google Search Console
  2. Select your property
  3. Go to Settings → Crawling → robots.txt Tester (if you don't see this, try the "Coverage" report and look for a robots.txt link)
  4. Type in a URL you want to check (e.g., /wp-admin/ or /blog/my-post/)
  5. Click Test
  6. Green = Allowed ✓ | Red = Blocked ✗

This tool is invaluable because it shows you exactly what Google sees your robots.txt file as. If you upload a new robots.txt, it may take 24-48 hours for Google to re-fetch and recognize it in GSC, but the test tool usually reflects it faster.

Using URL Inspection Tool

For individual pages, use GSC's URL Inspection tool (top search bar in GSC). After entering a URL, the report shows an "Indexing" section that will flag if a page is "Blocked by robots.txt." This tells you exactly which robots.txt rule is causing the block and gives you a chance to fix it immediately.

8. Common robots.txt Mistakes to Avoid

Mistake #1: Blocking /wp-content/ entirely

Disallow: /wp-content/ blocks images, CSS, and fonts. Google can't render your site properly, and image search traffic vanishes. Never block the entire /wp-content/ folder.

Mistake #2: Typos in paths

robots.txt is case-sensitive on Linux servers. Disallow: /Blog/ will not block /blog/. Always use lowercase paths unless your URLs actually use uppercase (rare). Test with GSC to be sure.

Mistake #3: Using Disallow without understanding priorities

If you write:

Disallow: /search/
Disallow: /search/advanced/

The first rule already blocks /search/advanced/. The second rule is redundant. While it doesn't hurt, it wastes space. robots.txt supports wildcard patterns with asterisks:

Disallow: /search*

This blocks /search/, /search-results/, /searchdata/, and any URL containing "search." Be careful with wildcards — test before deploying.

Mistake #4: Confusing robots.txt with noindex

These are not the same:

If a page is blocked by robots.txt, bots never reach it, so they can't read the noindex tag. Use robots.txt for admin/system pages; use noindex for public pages you don't want in results.

9. The Impact of Crawl Budget and How robots.txt Saves It

Large sites (100,000+ pages) have crawl budget constraints. Google allocates a certain number of crawl requests per day to each domain. If your site has infinite duplicate pages (pagination, session IDs, filters), Googlebot may waste the entire budget crawling variations and never reach new content you publish.

A strategic robots.txt reduces waste:

Every URL blocked is one less resource spent, and every resource saved is one more crawl of your valuable pages. For small sites (under 10,000 pages), crawl budget is usually not an issue, but good robots.txt practices scale with your site.

10. How to Fix robots.txt After Making a Mistake

Accidentally blocked your site? It's fixable. Here's the recovery process:

  1. Fix the robots.txt file immediately. Remove the bad Disallow rule or change Disallow: / to Disallow: /wp-admin/.
  2. If using Yoast SEO: Save the changes in the Yoast interface; the file updates on the server instantly.
  3. If using FTP: Re-upload the corrected robots.txt file.
  4. Go to Google Search Console → Settings → Crawling → robots.txt Tester. Verify your important URLs are now "Allowed" (green).
  5. Submit affected pages for reindexing. Use the "Request Indexing" feature in GSC to ask Google to recrawl your pages. Google usually re-crawls within 24-48 hours.
  6. Monitor your site in GSC for the next 2 weeks. Watch the Coverage report — your blocked pages should reappear as "Indexed" over time.

It's not instant, but robots.txt changes are one of the quickest things to recover from. Stay alert, test frequently, and you'll catch problems before they tank your traffic.

Bottom Line: robots.txt is a powerful SEO tool when configured correctly, but it's also a loaded gun. Understand what Disallow actually does, avoid the two fatal mistakes (Disallow: / and blocking /wp-content/), and use Google Search Console's robots.txt tester to verify every rule before deploying. A well-tuned robots.txt works silently in the background, directing crawlers efficiently and protecting your crawl budget — one of the invisible foundations of good SEO.

Frequently Asked Questions

Can I use robots.txt to completely hide my site from Google?

Technically yes — Disallow: / will block Google from crawling. However, this won't guarantee your site is hidden. If other sites link to you, Google may still index your URLs based on those links, even though Googlebot never visited your site. To truly hide your site, use one of these more reliable methods:

  • Use meta robots noindex tag on all pages
  • Use password protection (.htaccess authentication)
  • Remove your site from Google Search Console and request removal
How often does Google re-read my robots.txt file?

Google crawls robots.txt on every Googlebot visit — typically several times per day on active sites. However, there's no guarantee of when it will fetch your robots.txt. For critical changes, test in Google Search Console immediately to confirm the new version is recognized. If Google is still using an old version after 48 hours, it's worth submitting a reindexing request for your homepage to force a fresh crawl.

Is robots.txt a security measure? Can it protect /wp-admin/?

No. robots.txt is a courtesy for legitimate bots, not a barrier against hackers. Malicious bots and hacking scripts ignore robots.txt entirely and will still attempt to access /wp-admin/. For real security, use:

  • htaccess password protection on /wp-admin/
  • WAF (Web Application Firewall) rules
  • Change login URL with a plugin
  • Fail2ban to block repeated failed login attempts

robots.txt is for SEO organization, not security.

Should I block /wp-content/plugins/ in robots.txt?

It's optional. Plugin files (.js, .css, .php) are not typically indexed or needed in search results, so blocking /wp-content/plugins/ won't hurt your SEO. However, some plugins may need certain files to be crawlable for JavaScript rendering. The safest approach: allow plugins but block the cache subfolder (Disallow: /wp-content/plugins/cache/) if you have one. Test critical pages in GSC to ensure they render and score well on mobile.

What's the difference between * and / in Disallow rules?

Disallow: / blocks the entire site — nothing can be crawled.

Disallow: * blocks everything in the root. The asterisk is a wildcard meaning "any character sequence." So Disallow: /search* blocks /search/, /search-results, /searchdata/, etc. Use wildcards carefully and test in GSC first to confirm you're blocking what you intend.

How do I verify that Google Search Console is reading my robots.txt correctly?

Use the robots.txt Tester in Google Search Console (Settings → Crawling → robots.txt Tester). Type test URLs and verify the result. Also check the Coverage report in GSC — if pages you thought were blocked are actually indexed, there's likely a conflict or the robots.txt is different than expected. Remember: GSC can take 24-48 hours to recognize updates to your physical robots.txt file, so be patient after making changes.

Need WordPress Hosting with SEO in Mind?

AsiaGB WordPress Hosting comes pre-configured with proper robots.txt, DirectAdmin control panel for full file access, and SEO-friendly server settings. Scale your site with 99% uptime SSD storage and daily backups.

Explore WordPress Hosting Plans →