As a website webmaster or system administrator, have you ever experienced this nightmare: on a quiet, ordinary night, your server monitoring triggers piercing alerts, CPU utilization unexpectedly skyrockets to 100%, RAM is overwhelmed, and legitimate visitors trying to load your pages are greeted by cold “502 Bad Gateway” or “504 Gateway Timeout” errors?

When you scramble to SSH into the terminal and inspect your web server’s access logs, what meets your eyes is rarely enthusiastic browsing by real readers, but tens of thousands of high-concurrency requests originating from specific IP subnets.

These automated utilities are not legitimate search engine crawlers like Googlebot or Bingbot that abide by web standards; rather, they are “bad bots and scrapers” that maliciously devour server bandwidth, ruthlessly scrape original articles to populate spam content farms, unauthorizedly harvest corpora to train AI models, and aggressively probe your applications for unpatched security vulnerabilities.

Among malicious web traffic, unauthorized scrapers emerging from specific China Telecom subnets have become a persistent scourge for webmasters worldwide due to their hyper-frequency, massive concurrency, and complete disregard for internet conventions. This article discloses recent high-risk IP subnets and presents an end-to-end multi-layer defense guide spanning CDN edges, web application servers, and Linux kernel firewalls.


1. Why You Must Decisively Block These Bad Bots

Novice webmasters sometimes harbor wishful thinking, wondering: “Could more traffic coming in help our website’s SEO?” The reality is precisely the opposite!

These malicious crawlers deliver purely destructive blows to server infrastructure and site stability:

  1. Paralyzing Server Performance and Exhausting Connections: Modern dynamic web stacks (such as WordPress, Drupal, or custom backends) invoke PHP/Node.js worker threads and query database pools on every dynamic request. Malicious crawlers firing dozens or hundreds of requests per second instantly exhaust database connection pools and saturate CPU threads, rendering the site completely unresponsive to authentic visitors.
  2. Flagrant Disregard for the robots.txt Standard: Ethical search engine spiders proactively inspect robots.txt at the site root and honor specified Crawl-delay directives; grey-market scrapers blatantly disregard all exclusion rules, often prioritizing paths explicitly marked Disallow.
  3. Spoofed User-Agents and Vulnerability Probing: They routinely disguise themselves as standard desktop Chrome browsers or spoof Googlebot identities while secretly scanning for sensitive paths such as /wp-login.php, /.env, /phpmyadmin, and /.git/config, searching for forgotten backup archives or weak credential configurations.
  4. Intellectual Property Theft and Bandwidth Cost Explosions: Carefully crafted original writing is ripped and published on scraper sites within minutes of release; if your servers run on pay-as-you-go bandwidth metering (such as AWS, GCP, or Linode), you will also be blindsided by astronomical bandwidth bills at the end of the month.

Server monitoring dashboard analyzing high-frequency bot traffic and abnormal CPU load curves


2. High-Risk IP and CIDR Blocklist

Based on cross-verified server logs and webmaster incident reports, the following China Telecom (Guangdong and adjacent regions) IP addresses and /24 subnets have exhibited repeated brute-force scraping and directory probing behaviors. It is strongly advised to add them to your global drop list:

1. Isolated High-Risk Individual IPs

  • 14.153.206.108
  • 14.155.182.76
  • 14.155.204.129

2. Heavily Active /24 Subnets (Block Full CIDR)

  • 14.155.183.0/24 (Spanning 14.155.183.1 to 14.155.183.254)
  • 14.155.184.0/24 (Spanning 14.155.184.1 to 14.155.184.254)
  • 14.155.185.0/24 (Spanning 14.155.185.1 to 14.155.185.254)
  • 14.155.230.0/24 (Spanning 14.155.230.1 to 14.155.230.254)

3. Implementing Three-Layer Defense in Depth

To permanently rid your servers of scraper harassment, a single line of defense rarely suffices. We recommend establishing a three-tier defensive perimeter: “CDN Edge Layer $\rightarrow$ Web Server Layer $\rightarrow$ Linux Kernel Firewall Layer.”

Systems engineer configuring Linux firewall rules and IP blocklists in terminal


First Line of Defense: Cloudflare CDN / WAF Edge Filtering (Top Recommendation)

Routing site traffic through Cloudflare (enabling the orange cloud proxy) represents the most cost-effective defensive posture with zero computational overhead on your origin server. Malicious requests are intercepted at edge nodes closest to the requester without ever reaching your host.

1. Create Custom WAF Rules

Log in to the Cloudflare dashboard, navigate to “Security” $\rightarrow$ “WAF” $\rightarrow$ “Custom rules”:

  • Rule name: Block Known Bad Crawler Subnets
  • Filter criteria: Select “IP Source Address” $\rightarrow$ “is in” or use the raw expression editor:
    (ip.src in {14.153.206.108 14.155.182.76 14.155.204.129 14.155.183.0/24 14.155.184.0/24 14.155.185.0/24 14.155.230.0/24})
  • Action: Select “Block”.

2. Enable Bot Fight Mode and Geo-Challenges

  • Under “Security” $\rightarrow$ “Bots”, toggle on “Bot Fight Mode”.
  • If your core target audience is situated in Taiwan, Japan, Europe, or the Americas, and your business conducts zero commercial activity in Mainland China, you can configure a geographic rule: whenever ip.geoip.country eq "CN" and the requester is not a verified search engine, enforce a “Managed Challenge.” This single step weeds out 99% of headless scraper scripts automatically.

Second Line of Defense: Severing Connections at the Nginx Reverse Proxy

If your setup bypasses third-party CDNs, or if you require an internal secondary security perimeter at the web server layer, block troublesome IPs directly inside Nginx.

Crucially, when mitigating aggressive bots, utilize Nginx’s specialized return 444; rather than standard deny directives (which transmit a 403 Forbidden payload). Non-standard status code 444 instructs Nginx to immediately shut down the TCP connection without transmitting any HTTP response headers or error bodies, saving maximum socket and network resources.

Create /etc/nginx/conf.d/block_bad_bots.conf:

# Map high-risk scraper IPs and subnets
geo $bad_client {
    default         0;
    14.153.206.108  1;
    14.155.182.76   1;
    14.155.204.129  1;
    14.155.183.0/24 1;
    14.155.184.0/24 1;
    14.155.185.0/24 1;
    14.155.230.0/24 1;
}

server {
    listen 80;
    listen 443 ssl http2;
    server_name example.com;

    # Immediately sever connection if client matches blocklist
    if ($bad_client) {
        return 444;
    }

    # Remaining server block configuration...
}

Reload Nginx to activate changes:

sudo nginx -t && sudo systemctl reload nginx

Third Line of Defense: Linux Kernel Firewall (iptables + ipset)

When crawler volume escalates to the point of starving Nginx worker connections, the ultimate low-level remedy is dropping network packets outright at the Linux kernel space (DROP).

Many administrators write dozens of repetitive iptables -A INPUT -s ... -j DROP commands. However, lengthy rule chains force Linux to perform linear packet inspections, which degrades networking throughput. The proper, high-performance approach is leveraging “ipset” hash tables:

1. Install Packages and Initialize the Hash Table

# Ubuntu / Debian
sudo apt-get install ipset iptables-persistent -y

# Create an ipset hash table named bad_crawlers for network subnets
sudo ipset create bad_crawlers hash:net

2. Populate the Table with Target Subnets

sudo ipset add bad_crawlers 14.153.206.108
sudo ipset add bad_crawlers 14.155.182.76
sudo ipset add bad_crawlers 14.155.204.129
sudo ipset add bad_crawlers 14.155.183.0/24
sudo ipset add bad_crawlers 14.155.184.0/24
sudo ipset add bad_crawlers 14.155.185.0/24
sudo ipset add bad_crawlers 14.155.230.0/24

3. Bind the ipset Table to iptables

# Drop incoming packets that match the bad_crawlers set without mercy
sudo iptables -I INPUT -m set --match-set bad_crawlers src -j DROP

# Persist rules across system reboots
sudo ipset save > /etc/ipset.conf
sudo netfilter-persistent save

Using $O(1)$ ipset hash lookups, even if your blocklist swells to tens of thousands of subnets, the Linux kernel determines matches and discards malicious packets within microseconds, consuming virtually zero CPU and memory overhead.


Defensive Strategy Comparison Matrix

Here is how each defensive approach stacks up in terms of resource usage and effectiveness:

Defense TechniqueProtection LayerServer OverheadDeployment ComplexityEffectiveness Rating
robots.txtProtocol AdvisoryNegligibleExtremely SimpleIneffective (Ignored by bad bots)
Cloudflare WAFCDN / Edge LayerZero Overhead (Optimal)LowExceptional; traffic never hits origin
Nginx return 444Web Application LayerMinimalModerateGood; drops TCP connection without payload
Linux ipset + iptablesOS Kernel LayerVery Low (Hash Lookup)Moderate-HighOutstanding; vital for massive concurrency

Conclusion: Architectural Static Generation Is the Ultimate Cure

Mitigating malicious crawlers often feels like an endless game of cat-and-mouse; threat actors switch IP subnets today only to return behind fresh proxies next week.

Beyond reviewing server logs and updating blocklists regularly, making fundamental architectural upgrades provides true peace of mind—such as adopting modern static site generators (SSG) like Astro, which pre-render dynamic database operations into pure HTML, CSS, and JavaScript at build time.

When your server no longer queries a relational database on every incoming request, handling millions of scraping requests is as lightweight as serving flat static files. You eliminate database crash anxieties entirely and ensure your servers remain reliably healthy and rock solid.

✦ Independent Journalism · Reader Support ✦

Support Independent Perspectives & In-Depth Insights

Every thoughtful analysis and candid critique comes from our dedication to truth and quality. We choose not to follow sensational algorithms or clickbait headlines.

Sustaining independent research requires reader support. Make a one-time or monthly contribution, securely processed by Google.

Payments secured by Google · Manage or cancel anytime in your Google Account

TagsServer ProtectionBad BotsIP BlockingCloudflareNginxNetwork Security