top of page

How to scrape banking and financial data from Kotak Mahindra Bank using the Minexa API

Jul 15
3 min read

Financial product pages like Kotak Mahindra Bank's credit card listings are structured, publicly accessible, and updated regularly. That makes them a practical data source for fintech developers, competitive analysts, and financial researchers who need card features, eligibility criteria, and product metadata in a structured format.

This guide covers how to extract that data programmatically using the Minexa API, a deterministic web extraction platform that handles JavaScript rendering, anti-bot protection, and structured field discovery without requiring custom selector code.

Step 1: Open the target page and install the extension

Navigate to the Kotak Mahindra Bank credit cards page. Install the Minexa Chrome extension from the Chrome Web Store if you have not already.

Once the page is fully loaded, click the Minexa extension icon in your browser toolbar. The popup will confirm you are on the right page and prompt you to proceed.

Step 2: Confirm pagination and choose scraping mode

The extension detects whether the page uses pagination. Review the detected pagination logic and click Continue. You will then be asked whether to scrape a single list or a list combined with linked detail pages.

For the credit card listing page, selecting the single list option is sufficient to capture all card-level data points visible on the page. If you need to follow each card link and extract detail page content, choose the list and detail option instead.

Step 3: Select the data container and create the scraper

After choosing your mode, the extension highlights the HTML container holding the full card listing block. Confirm the selection and click Create Scraper. Minexa analyzes the page structure and automatically identifies all data points within that container, typically within a couple of minutes.

Once created, all extracted columns are visible in the extension preview. Use the next and previous navigation to review the full set of discovered fields before proceeding.

Step 4: Get the API request and Python code

Click API Request in the top right of the extension. This opens a panel showing the pre-generated JSON request body and a ready-to-run Python script. Copy the Python code directly from here.

Reference: API request structure

The request below shows the standard structure for calling the Minexa API against the Kotak Mahindra Bank credit card page. Replace the scraper_id with the one generated for your scraper, and populate the urls array with the pages you want to process.

POST https://api.minexa.ai/data/

{
  "batches": [
    {
      "scraper_id": 6183,
      "columns": ["top_30"],
      "urls": [
        "https://www.kotak.bank.in/en/personal-banking/cards/credit-cards.html"
      ],
      "scraping": {
        "js_render": true,
        "timeout": 30,
        "js_code": [
          { "wait_time": 2 },
          { "page_init": true },
          { "wait_time": 4 }
        ],
        "proxy": "verified",
        "retry": 3
      }
    }
  ],
  "threads": 3
}

The columns parameter accepts top_N notation, where N is any positive integer. Using top_30 returns the thirty highest-ranked fields identified by Minexa's relevance algorithm. You can also pass an explicit list of column names if you want to extract only specific fields.

Reference: sample extracted data

Below is a representative sample of what the API returns for the Kotak Mahindra Bank credit cards page. Metadata fields are excluded for clarity.

[
  {
    "access_message": "",
    "challenge_type": "challenge-icon",
    "client_ip_address": "",
    "status_message": "",
    "ray_id": "",
    "restriction_message": "",
    "user_agent": ""
  },
  {
    "access_message": "You do not have access to www.kotak.bank.in.",
    "challenge_type": "challenge-captcha",
    "client_ip_address": "",
    "status_message": "Access denied",
    "ray_id": "Ray ID: b29fe1c741ab0d12",
    "restriction_message": "The site owner may have set restrictions.",
    "user_agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36"
  }
]

The status_message field surfaces the page-level access state per record. When this field returns a non-empty string such as 'Access denied', it signals that the crawl encountered a block rather than the target content, allowing downstream filtering before any processing occurs.

The challenge_type field encodes the specific anti-bot mechanism encountered. Values like 'challenge-captcha' or 'challenge-icon' let you classify failure modes by type and adjust scraping configuration accordingly, for example by switching to a stronger provider or enabling bypass mode.

The user_agent field captures the exact browser identity string used during the crawl. Cross-referencing this against challenge responses helps diagnose whether a specific user agent string is being flagged by the target site's bot detection layer.

Reference: running the extraction at scale

Once the scraper is trained, the same scraper_id works across any number of structurally similar pages. Pass multiple URLs in the urls array and set threads to match your plan's concurrency limit for parallel processing. The Python script generated by the extension saves a checkpoint file after each API response, so partial results are preserved even if a run is interrupted.

Results export to Excel and JSON directly from the Minexa dashboard. For automated pipelines, set up a cron job on your infrastructure and call the API endpoint with updated URL lists on whatever cadence your use case requires.

For a related walkthrough covering financial data extraction from another Indian market source, see: Scraping tax and accounting data from IDX using the Minexa API.

Recent Posts

See All

Comments


Heading 2

bottom of page