Have you noticed unusual server connections recently, with requests masquerading as OpenAI’s official web crawler “GPTBot” exhibiting aggressive or anomalous behavior?
While monitoring traffic on Mountos infrastructure, we captured an intriguing IP address: 20.171.207.116. Its HTTP headers claimed to originate from OpenAI’s GPTBot, yet its traffic pattern resembled aggressive automated directory fuzzing rather than standard search indexing.
Upon in-depth network verification, we confirmed that this IP indeed falls within Microsoft/OpenAI’s published subnet blocks. However, its request volume prompted a deeper investigation into whether spoofed crawlers or compromised cloud nodes are proliferating across the web.
This guide provides OpenAI’s official crawler IP CIDR ranges and outlines actionable strategies to distinguish legitimate AI scrapers from malicious threats.
1. Real vs. Fake GPTBot: The Anomalous Case
During our server audit, the flagged IP was 20.171.207.116 exhibiting the following User-Agent string:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.2; +https://openai.com/gptbot)
Reverse DNS and WHOIS lookups showed that the IP is registered to Microsoft Corporation (AS8075 MICROSOFT-CORP-MSN-AS-BLOCK) and falls strictly within OpenAI’s publicly declared IP range (20.171.207.0/24).
However, the bot initiated over 100,000 rapid requests targeting the root directory in a single day, ignoring gentle crawl rates. Whether due to misconfigured AI model training or unauthorized endpoint fuzzing, such aggressive behavior necessitated an immediate firewall block.
2. OpenAI Official IP Ranges: Verified CIDR Lists
To help system administrators authenticate legitimate crawlers, OpenAI publishes verified JSON feeds of their network ranges:
GPTBot Published IP Ranges
{
"creationTime": "2023-11-30T11:51:00.000000",
"prefixes": [
{"ipv4Prefix": "52.230.152.0/24"},
{"ipv4Prefix": "52.233.106.0/24"},
{"ipv4Prefix": "20.171.206.0/24"},
{"ipv4Prefix": "20.171.207.0/24"},
{"ipv4Prefix": "4.227.36.0/25"},
{"ipv4Prefix": "20.125.66.80/28"},
{"ipv4Prefix": "172.182.193.160/28"}
]
}
ChatGPT-User Published IP Ranges
Used when end-users prompt ChatGPT to browse the web in real time:
{
"creationTime": "2025-02-20T20:15:50.707457",
"prefixes": [
{"ipv4Prefix": "23.98.179.16/28"},
{"ipv4Prefix": "172.183.222.128/28"},
{"ipv4Prefix": "51.8.155.64/28"},
{"ipv4Prefix": "51.8.155.48/28"},
{"ipv4Prefix": "135.237.131.208/28"},
{"ipv4Prefix": "51.8.155.112/28"},
{"ipv4Prefix": "52.159.249.96/28"},
{"ipv4Prefix": "172.178.141.112/28"},
{"ipv4Prefix": "172.178.140.144/28"},
{"ipv4Prefix": "172.178.141.128/28"},
{"ipv4Prefix": "4.196.118.112/28"},
{"ipv4Prefix": "20.215.188.192/28"},
{"ipv4Prefix": "4.197.22.112/28"},
{"ipv4Prefix": "57.154.175.0/28"},
{"ipv4Prefix": "52.236.94.144/28"},
{"ipv4Prefix": "23.98.186.192/28"},
{"ipv4Prefix": "23.98.186.176/28"},
{"ipv4Prefix": "13.83.167.128/28"},
{"ipv4Prefix": "20.97.189.96/28"},
{"ipv4Prefix": "20.161.75.208/28"},
{"ipv4Prefix": "52.225.75.208/28"},
{"ipv4Prefix": "52.156.77.144/28"},
{"ipv4Prefix": "40.84.221.208/28"},
{"ipv4Prefix": "40.84.221.224/28"}
]
}
OAI-SearchBot Published IP Ranges
Used by SearchGPT and prototype search experiences:
{
"creationTime": "2025-02-10T21:00:00.000000",
"prefixes": [
{"ipv4Prefix": "20.42.10.176/28"},
{"ipv4Prefix": "172.203.190.128/28"},
{"ipv4Prefix": "51.8.102.0/24"},
{"ipv4Prefix": "135.234.64.0/24"}
]
}
Any web request claiming to be GPTBot or ChatGPT whose source IP falls outside these official CIDRs is a forged request and should be filtered immediately.
3. Anatomy of a Legitimate GPTBot Request
To verify genuine OpenAI scrapers, inspect three criteria:
- User-Agent String: Must clearly state
GPTBot/1.xalong with the linkhttps://openai.com/gptbot. - IP Whitelist: The origin IP must match OpenAI’s published subnets.
- Behavioral Compliance: Legitimate crawlers respect
robots.txtdisallow rules and maintain reasonable request concurrency.
4. Why Threat Actors Spoof AI Crawlers
Malicious bots spoof well-known User-Agents for several malicious purposes:
- Bypassing Web Application Firewalls (WAF): Many basic firewalls automatically whitelist common crawlers by User-Agent alone.
- Vulnerability Scanning: Probing for exposed
.envfiles, git repositories, and unauthenticated API endpoints under the guise of an AI bot. - Unauthorized Scraping: Evading website scraping limits to train unvetted private models.
5. Defense Strategies: How to Protect Your Servers
To safeguard your infrastructure, implement the following defense layers:
1. WAF Spoofing Filter (Cloudflare / AWS WAF)
Create a firewall rule that blocks fake bots:
- Condition:
(http.user_agent contains "GPTBot") and not (ip.src in {OpenAI_CIDR_List}) - Action: Block (Drop connection)
2. Configure robots.txt
If you choose to block OpenAI from indexing your website entirely, add:
User-agent: GPTBot
Disallow: /
3. Implement Strict Rate Limiting
Deploy IP-based rate limiting rules to throttle any single IP sending more than 60 requests per minute, regardless of its declared identity.
4. Continuous Log Auditing
Regularly inspect Nginx / Apache access logs to detect anomaly spikes in request frequency and monitor Cloudflare Bot Management scores.
Summary
As artificial intelligence crawlers become omnipresent, validating source authenticity is vital for server health and intellectual property protection. Rely on verified IP ranges and robust WAF filtering to ensure only legitimate search and AI agents interact with your digital platforms.
Support Independent Perspectives & In-Depth Insights
Every thoughtful analysis and candid critique comes from our dedication to truth and quality. We choose not to follow sensational algorithms or clickbait headlines.
Sustaining independent research requires reader support. Make a one-time or monthly contribution, securely processed by Google.
Payments secured by Google · Manage or cancel anytime in your Google Account



Comments