Skip to main content

Authenticated Scraping

Scrape pages behind login walls using Fetchyr

2 min read


This guide covers login flows, MFA handling, and session reuse for scraping authenticated pages.

Overview

Many sites require authentication to access their data. KLOAKD's Fetchyr module handles login flows, MFA, and form automation — then passes a session artifact to kolektr.page() or webgrph.crawl().

Step 1 — Create a session

from kloakd import Kloakd

client = Kloakd(api_key="sk-live-...", organization_id="your-org-id")

session = client.fetchyr.login(
    url="https://example.com/login",
    username_selector="#email",
    password_selector="#password",
    username="user@example.com",
    password="secret",
    submit_selector="button[type=submit]",
)
print(f"Session artifact: {session.artifact_id}")

Step 2 — Handle MFA (if required)

mfa = client.fetchyr.detect_mfa(
    url="https://example.com/mfa",
    session_artifact_id=session.artifact_id,
)
if mfa.mfa_detected:
    result = client.fetchyr.submit_mfa(
        challenge_id=mfa.challenge_id,
        code="123456",
    )
    print(f"MFA submitted: {result.success}")

MFA types

| Type | Description | |------|-------------| | totp | Time-based OTP (Google Authenticator, Authy) | | sms | SMS code | | email | Email code |

Step 3 — Extract authenticated pages

data = client.kolektr.page(
    "https://example.com/dashboard",
    schema={"title": "css:h1", "content": "css:article"},
    session_artifact_id=session.artifact_id,
)
for record in data.records:
    print(record)

Step 4 — Crawl authenticated pages

Use the client.crawl() orchestrator with session_artifact_id for a one-call solution:

result = client.crawl(
    "https://example.com/dashboard",
    max_depth=2,
    session_artifact_id=session.artifact_id,
    extract_schema={"title": "css:h1", "content": "css:article"},
)
print(f"Crawled {result.total_pages} authenticated pages")
for page in result.pages:
    print(f"  {page.url}: {page.extracted_data}")

Or use the low-level webgrph.crawl() with polling:

import time

crawl = client.webgrph.crawl(
    "https://example.com/dashboard",
    max_depth=2,
    session_artifact_id=session.artifact_id,
)
print(f"Crawl started: {crawl.crawl_id}")

# Poll for completion
while True:
    status = client.webgrph.get_crawl_status(crawl.crawl_id)
    if status.get("status") == "completed":
        break
    time.sleep(2)

pages = client.webgrph.get_crawl_pages(crawl.crawl_id)
print(f"Crawled {len(pages.get('pages', []))} authenticated pages")

Session reuse

Sessions expire after 30 minutes of inactivity. Reuse the same session_artifact_id across multiple requests:

# One login, many extractions
pages = ["/dashboard", "/profile", "/settings"]
for path in pages:
    data = client.kolektr.page(
        f"https://example.com{path}",
        schema={"title": "css:h1"},
        session_artifact_id=session.artifact_id,
    )
    print(f"{path}: {len(data.records)} records")

Tips

  • Use client.crawl() with session_artifact_id — one call handles authenticated discovery, fetching, and extraction
  • Create sessions before crawling — login once, crawl many pages
  • Check mfa_detected before extracting — the session won't work until MFA is complete
  • Store artifact_id for reuse — avoids re-login for subsequent jobs
  • Sessions auto-expire — 30 minutes of inactivity, no cleanup needed

Next steps

Was this page helpful?