top of page

How websites block scrapers, and what developers should take away

11 minutes ago
4 min read

Large marketplaces are among the most scraped sites on the web, and they respond with layered defences rather than one single wall. If you run a site and want to limit automated extraction, their playbook is worth studying. If you build data pipelines, it is worth studying for the opposite reason: it tells you exactly which behaviours get flagged. This post covers both sides, plainly.

Why one defence is never enough

Any single check can be worked around. Header checks catch naive scripts but are trivial to spoof. Rate limits slow things down but punish heavy legitimate users. The sites that hold up combine signals, so that a request has to look normal across many dimensions at once.

The main layers site owners use

1. Traffic monitoring and throttling

Start with logs. Watch for bursts of requests from one IP, subnet, or network block, and for navigation paths no human follows. Cap requests per IP, session, or account over a time window, and set tighter caps on sensitive endpoints. Throttling or a challenge is usually a better first response than a hard ban.

2. Header and client validation

Reject requests with missing client identifiers or ones that announce command-line tools. Check that language, referrer, and encoding headers are consistent with the browser the request claims to be. Inconsistency is often a stronger signal than any single bad value.

3. Fingerprinting and behaviour

Canvas, WebGL, fonts, timezone, TLS handshake details, and header ordering combine into a device profile. Headless browsers often leave detectable traces. On top of that, behaviour models look at mouse paths, scroll rhythm, and click timing. Perfectly regular movement or instant page-to-page jumps stand out.

4. Challenge tests

Image puzzles, invisible checks, and proof-of-work challenges can be shown only to suspicious traffic. Treat a triggered challenge as a signal in itself, and feed it back into throttling decisions.

5. Sessions and cookies

Require cookies, assign unique identifiers, and score sessions by history and consistency. A session that switches IP, locale, or client string mid-visit deserves a closer look.

6. Honeypots

Hidden links or form fields that no real visitor would touch are effective traps. Rotate them regularly, and disallow them in robots.txt so well-behaved crawlers never trip them.

7. Dynamic loading and markup churn

Loading key content through JavaScript blocks simple HTML parsers. Randomising class names and element IDs breaks scrapers built on fixed selectors. Rendering text as images also deters basic tools, but it hurts accessibility and search visibility, so most teams avoid it.

8. Protected endpoints

Internal APIs need token authentication, signed requests, short expiry windows, per-user limits, and caps on query depth to stop over-fetching.

9. Policy and legal measures

Terms of service that require written consent for automated collection give you a basis for cease-and-desist notices. Enforcement tends to be slow, so treat this as a backstop.

10. Managed bot services

Commercial bot management platforms share threat data across many sites and score traffic in real time. They cost money and take integration work, but they cover the most ground.

Newer measures aimed at AI crawlers

Many sites now disallow named AI crawlers in robots.txt, add explicit clauses against AI-driven collection in their terms, and watch for the large, predictable fetch patterns typical of model training. Robots.txt is a guideline rather than a security control, but ignoring it can strengthen a site owner's legal position.

Quick comparison

Technique

Strength

Trade-off

Rate limiting

Cuts load, slows bulk collection

Can affect power users

Header checks

Stops basic scripts

Easy to spoof

Fingerprinting

Catches advanced bots

Needs detection infrastructure

Challenges

Strong against automation

Friction if overused

Honeypots

Low cost, clear signal

Needs rotation, false positive risk

Markup churn

Breaks selector-based tools

SEO and maintenance cost

Bot services

Broadest coverage

Cost and integration effort

Keep the site usable

Whitelist search engine crawlers, show generic error pages without explaining why a request was blocked, prefer throttling over bans, and test that content stays accessible. Aggressive blocking that drives away real visitors defeats the purpose.

What this means for developers collecting public data

The legal picture is fairly consistent: public product data such as titles, prices, and ratings is generally fair to collect, while login-protected content, personal data, and copyrighted material are off limits. A site's terms may still prohibit automation, so read them before you build.

Technically, two defences hit pipelines hardest: JavaScript-rendered content and changing markup. The Minexa.ai API handles JavaScript rendering and location-dependent content without extra configuration. Extraction is tied to page structure, so each column maps to a fixed position instead of a guess. When a site redesigns and the page no longer matches, the scraper returns an empty result rather than wrong values, and retraining takes a few minutes using the same setup flow. That makes markup churn a visible event in your data instead of a silent quality problem.

Scrapers are trained once in the Chrome extension and then called from your own code, with your own cron jobs deciding when to run. You can set up your first scraper here and test it against a public list page you already know well.

Whichever side you are on, the same rule applies: understand the signals, respect the boundaries, and design for change instead of assuming pages stay the same.

Recent Posts

See All

Comments


Heading 2

bottom of page