How websites block scrapers, and what developers should take away
Large marketplaces are among the most scraped sites on the web, and they respond with layered defences rather than one single wall. If you run a site and want to limit automated extraction, their playbook is worth studying. If you build data pipelines, it is worth studying for the opposite reason: it tells you exactly which behaviours get flagged. This post covers both sides, plainly.
Why one defence is never enough
Any single check can be worked around. Header checks catch naive scripts but are trivial to spoof. Rate limits slow things down but punish heavy legitimate users. The sites that hold up combine signals, so that a request has to look normal across many dimensions at once.
The main layers site owners use
1. Traffic monitoring and throttling
Start with logs. Watch for bursts of requests from one IP, subnet, or network block, and for navigation paths no human follows. Cap requests per IP, session, or account over a time window, and set tighter caps on sensitive endpoints. Throttling or a challenge is usually a better first response than a hard ban.
2. Header and client validation
Reject requests with missing client identifiers or ones that announce command-line tools. Check that language, referrer, and encoding headers are consistent with the browser the request claims to be. Inconsistency is often a stronger signal than any single bad value.
3. Fingerprinting and behaviour
Canvas, WebGL, fonts, timezone, TLS handshake details, and header ordering combine into a device profile. Headless browsers often leave detectable traces. On top of that, behaviour models look at mouse paths, scroll rhythm, and click timing. Perfectly regular movement or instant page-to-page jumps stand out.
4. Challenge tests
Image puzzles, invisible checks, and proof-of-work challenges can be shown only to suspicious traffic. Treat a triggered challenge as a signal in itself, and feed it back into throttling decisions.
5. Sessions and cookies
Require cookies, assign unique identifiers, and score sessions by history and consistency. A session that switches IP, locale, or client string mid-visit deserves a closer look.
6. Honeypots
Hidden links or form fields that no real visitor would touch are effective traps. Rotate them regularly, and disallow them in robots.txt so well-behaved crawlers never trip them.
7. Dynamic loading and markup churn
Loading key content through JavaScript blocks simple HTML parsers. Randomising class names and element IDs breaks scrapers built on fixed selectors. Rendering text as images also deters basic tools, but it hurts accessibility and search visibility, so most teams avoid it.
8. Protected endpoints
Internal APIs need token authentication, signed requests, short expiry windows, per-user limits, and caps on query depth to stop over-fetching.
9. Policy and legal measures
Terms of service that require written consent for automated collection give you a basis for cease-and-desist notices. Enforcement tends to be slow, so treat this as a backstop.
10. Managed bot services
Commercial bot management platforms share threat data across many sites and score traffic in real time. They cost money and take integration work, but they cover the most ground.
Newer measures aimed at AI crawlers
Many sites now disallow named AI crawlers in robots.txt, add explicit clauses against AI-driven collection in their terms, and watch for the large, predictable fetch patterns typical of model training. Robots.txt is a guideline rather than a security control, but ignoring it can strengthen a site owner's legal position.
Quick comparison
Technique | Strength | Trade-off |
Rate limiting | Cuts load, slows bulk collection | Can affect power users |
Header checks | Stops basic scripts | Easy to spoof |
Fingerprinting | Catches advanced bots | Needs detection infrastructure |
Challenges | Strong against automation | Friction if overused |
Honeypots | Low cost, clear signal | Needs rotation, false positive risk |
Markup churn | Breaks selector-based tools | SEO and maintenance cost |
Bot services | Broadest coverage | Cost and integration effort |
Keep the site usable
Whitelist search engine crawlers, show generic error pages without explaining why a request was blocked, prefer throttling over bans, and test that content stays accessible. Aggressive blocking that drives away real visitors defeats the purpose.
What this means for developers collecting public data
The legal picture is fairly consistent: public product data such as titles, prices, and ratings is generally fair to collect, while login-protected content, personal data, and copyrighted material are off limits. A site's terms may still prohibit automation, so read them before you build.
Technically, two defences hit pipelines hardest: JavaScript-rendered content and changing markup. The Minexa.ai API handles JavaScript rendering and location-dependent content without extra configuration. Extraction is tied to page structure, so each column maps to a fixed position instead of a guess. When a site redesigns and the page no longer matches, the scraper returns an empty result rather than wrong values, and retraining takes a few minutes using the same setup flow. That makes markup churn a visible event in your data instead of a silent quality problem.
Scrapers are trained once in the Chrome extension and then called from your own code, with your own cron jobs deciding when to run. You can set up your first scraper here and test it against a public list page you already know well.
Whichever side you are on, the same rule applies: understand the signals, respect the boundaries, and design for change instead of assuming pages stay the same.


Comments