top of page

How to scrape investment and VC data from Inc42 using the Minexa API

Jul 14
3 min read

If you track Indian startup funding rounds, monitor sector trends, or build datasets for VC research, Inc42 is one of the most consistently updated sources available. The challenge is that the data sits inside article cards spread across paginated listing pages, and there is no official API to pull it programmatically.

This guide shows how to extract structured startup and investment data from Inc42 using the Minexa API. The workflow has two phases: train a scraper once using the Minexa Chrome extension, then call the API to extract data at scale across as many pages as you need.

Watch the full tutorial first

The video below walks through the complete extraction workflow on Inc42, from opening the extension to getting structured JSON back from the API.

What the extracted data looks like

Each row in the output corresponds to one article card on the Inc42 startups listing page. Here is a sample of two records with the most useful fields:

[
  {
    "article_link": "https://inc42.com/startups/how-nudge-is-reinventing-ecommerce-for-the-agentic-ai-era/",
    "article_title_2": "How Nudge Is Reinventing Ecommerce For The Agentic AI Era",
    "author_name": "Ankush Das",
    "event_date": "3rd July, 2026",
    "data_card_id": "562648",
    "industry_links": "https://inc42.com/startups/"
  },
  {
    "article_link": "https://inc42.com/startups/30-startups-to-watch-startups-that-caught-our-eye-in-june-2026/",
    "article_title_2": "30 Startups To Watch: Startups That Caught Our Eye In June 2026",
    "author_name": "Akshit P.",
    "event_date": "1st July, 2026",
    "data_card_id": "563124",
    "industry_links": "https://inc42.com/startups/"
  }
]

The data_card_id field is particularly useful for deduplication. Because the same article can appear across multiple paginated runs, this numeric identifier lets you filter duplicates reliably without parsing titles or URLs.

The article_titles field returns a structured array of typed objects. Each object encodes the section label, full article title, author name, and publication date within a single traversable structure per card. This means you can extract all four values from one field rather than joining separate columns after the fact.

The card_classes field encodes layout metadata per card, including aspect ratio identifiers. This lets you distinguish featured cards from standard listings programmatically, which is useful if you want to weight or filter by editorial prominence.

Step 1: Train the scraper in the extension

Navigate to the Inc42 startups page in Chrome with the Minexa extension installed. Open the extension popup and confirm you are on the correct page.

Click the 'I'm on the right page' button in the extension popup. Minexa will begin detecting the list structure, pagination method, and all data points on the page automatically.

The extension will show the pagination options it detected. Review and click Continue to proceed to the next configuration step.

After validating pagination, choose whether to extract list-level data only or also follow each article link into its detail page. For a broad investment monitoring dataset, list-level extraction covers the fields most analysts need.

Start the scraping job. Minexa highlights the detected container covering the full article card list so you can confirm it has the right scope before committing.

Once the scraper is created, you will see all extracted data points listed. Click 'API Request' to get the scraper ID and ready-to-run code samples.

Step 2: Call the Minexa API from Python

With the scraper trained, you can now call the Minexa API to extract data at scale. Replace YOUR_API_KEY and use the scraper ID assigned during training (for example, 6183).

import requests

url = "https://api.minexa.ai/data"

payload = {
    "scraper_id": 6183,
    "columns": "top_40",
    "urls": [
        "https://inc42.com/startups/?utm_medium=referral&utm_source=menu"
    ]
}

headers = {
    "Content-Type": "application/json",
    "Authorization": "Bearer YOUR_API_KEY"
}

response = requests.post(url, json=payload, headers=headers)
print(response.json())

The columns parameter accepts either a named list of fields or a shorthand like top_40 to return the forty highest-ranked data points the scraper identified. For recurring monitoring jobs across many Inc42 pages, set up your own cron job and pass updated URL lists to the API on whatever schedule fits your workflow.

Export the results as JSON or Excel directly from the Minexa dashboard, or consume the API response directly in your pipeline.

For a related guide on building programmatic data pipelines from financial listing pages, see: Scraping tax and accounting data from IDX using the Minexa API.

Recent Posts

See All

Comments


Heading 2

bottom of page