How to do web scraping without rebuilding it every month
Most guides on how to do web scraping end once the first script prints clean rows. That is usually the easy part. The harder part comes later, when the site changes its layout, starts loading data through JavaScript, or begins rate limiting your requests. For anything running in production, maintenance is the main cost, not the first build.
This guide walks through the full process with that in mind.
The four stages of any scraper
Fetch the page. A plain GET request is enough for static HTML. Pages that build their content in the browser need a headless browser to run the JavaScript first. Pages behind a login also need you to manage sessions and cookies.
Parse the DOM. The raw HTML is turned into a tree you can search, with tools like Beautiful Soup or the live DOM APIs in Playwright.
Extract and structure. CSS selectors or XPath find the elements you want. You pull out text and attributes such as links, then map them to a schema.
Validate. Check that the data is complete, remove duplicates, and only then send it on to dashboards, analytics or AI pipelines.
Stage three is where scrapers break. Each selector assumes something about the page markup, and markup changes without warning.
Crawler, scraper or API?
System | Job | Use it when |
Crawler | Finds URLs | You need to map a site |
Scraper | Pulls fields from known pages | There is no official API |
Official API | Returns structured data directly | The provider offers one |
If an official API covers the fields you need, use it. Scraping makes sense when it does not.
The tool landscape
Python libraries: Requests and HTTPX to fetch pages, Beautiful Soup and PyQuery to parse them, Scrapy as a full framework for crawling and export.
Browser automation: Playwright, Selenium and Puppeteer for pages that render with JavaScript.
No-code tools: visual point-and-click scrapers for people who do not write code.
Managed APIs: services that take care of proxies, rendering and CAPTCHAs for you.
Where the real work goes
Dynamic content. A static request to a JavaScript-heavy site often returns a nearly empty shell, so you need a rendering layer.
Anti-bot measures. Rate limits, IP blocks, CAPTCHAs and browser fingerprinting are common. Pace your requests and treat robots.txt as an ethical baseline.
Legal limits. Public factual data carries less risk. Personal data falls under laws such as GDPR and CCPA, and paywalled content can breach a site's terms.
Pagination. You have to follow next-page links, keep track of visited URLs, and avoid looping over the same pages.
A note on LLM extraction
Sending pages to an LLM works well when layouts vary or for one-off edge cases. At volume, two problems show up. When a page has two similar values, such as a sale price and an original price, the model can pick the wrong one without any sign of error. And because you pay per token, full HTML pages get expensive across thousands of URLs. Many teams use a hybrid setup: deterministic extraction for the bulk of the data and a model only for the exceptions.
Training once instead of writing selectors
Minexa.ai takes a different route for stage three. You train a scraper once in its Chrome extension. It detects the repeating list on the page and its data points, including image links and attributes hidden in the markup, and you confirm what it found. After that, the Minexa.ai API reuses the scraper on any page with the same structure, and extraction takes milliseconds per page.
Each column is tied to a fixed position in the page structure. If a value is missing, the field comes back empty rather than filled with a guess. After a major redesign, you retrain the scraper in a few minutes. Check your column names afterwards, because labels can change slightly when a scraper is retrained.
Two things stay on your side when using the API. You write a short JS scenario that tells it what to click for pagination, and you run your own cron jobs to handle recurring URL batches.
To try this on your own target site, install the Chrome extension and train your first scraper.
A practical checklist
Look for an official API before you scrape.
Use browser developer tools to check whether the content is static or rendered with JavaScript.
Pace your requests and respect robots.txt.
Validate and deduplicate data before anything downstream uses it.
Plan from the start for how you will handle layout changes.
The get started guide walks through the move from extension to API step by step.
Further reading: How to do web scraping when maintenance is the job


Comments