Authenticated Scraping
Scrape pages behind login walls using Fetchyr
2 min read
This guide covers login flows, MFA handling, and session reuse for scraping authenticated pages.
Overview
Many sites require authentication to access their data. KLOAKD's Fetchyr module handles login flows, MFA, and form automation — then passes a session artifact to kolektr.page() or webgrph.crawl().
Step 1 — Create a session
from kloakd import Kloakd
client = Kloakd(api_key="sk-live-...", organization_id="your-org-id")
session = client.fetchyr.login(
url="https://example.com/login",
username_selector="#email",
password_selector="#password",
username="user@example.com",
password="secret",
submit_selector="button[type=submit]",
)
print(f"Session artifact: {session.artifact_id}")
Step 2 — Handle MFA (if required)
mfa = client.fetchyr.detect_mfa(
url="https://example.com/mfa",
session_artifact_id=session.artifact_id,
)
if mfa.mfa_detected:
result = client.fetchyr.submit_mfa(
challenge_id=mfa.challenge_id,
code="123456",
)
print(f"MFA submitted: {result.success}")
MFA types
| Type | Description |
|------|-------------|
| totp | Time-based OTP (Google Authenticator, Authy) |
| sms | SMS code |
| email | Email code |
Step 3 — Extract authenticated pages
data = client.kolektr.page(
"https://example.com/dashboard",
schema={"title": "css:h1", "content": "css:article"},
session_artifact_id=session.artifact_id,
)
for record in data.records:
print(record)
Step 4 — Crawl authenticated pages
Use the client.crawl() orchestrator with session_artifact_id for a one-call solution:
result = client.crawl(
"https://example.com/dashboard",
max_depth=2,
session_artifact_id=session.artifact_id,
extract_schema={"title": "css:h1", "content": "css:article"},
)
print(f"Crawled {result.total_pages} authenticated pages")
for page in result.pages:
print(f" {page.url}: {page.extracted_data}")
Or use the low-level webgrph.crawl() with polling:
import time
crawl = client.webgrph.crawl(
"https://example.com/dashboard",
max_depth=2,
session_artifact_id=session.artifact_id,
)
print(f"Crawl started: {crawl.crawl_id}")
# Poll for completion
while True:
status = client.webgrph.get_crawl_status(crawl.crawl_id)
if status.get("status") == "completed":
break
time.sleep(2)
pages = client.webgrph.get_crawl_pages(crawl.crawl_id)
print(f"Crawled {len(pages.get('pages', []))} authenticated pages")
Session reuse
Sessions expire after 30 minutes of inactivity. Reuse the same session_artifact_id across multiple requests:
# One login, many extractions
pages = ["/dashboard", "/profile", "/settings"]
for path in pages:
data = client.kolektr.page(
f"https://example.com{path}",
schema={"title": "css:h1"},
session_artifact_id=session.artifact_id,
)
print(f"{path}: {len(data.records)} records")
Tips
- Use
client.crawl()withsession_artifact_id— one call handles authenticated discovery, fetching, and extraction - Create sessions before crawling — login once, crawl many pages
- Check
mfa_detectedbefore extracting — the session won't work until MFA is complete - Store
artifact_idfor reuse — avoids re-login for subsequent jobs - Sessions auto-expire — 30 minutes of inactivity, no cleanup needed
Next steps
- crawl() orchestrator — one-call API for crawling and extracting
- Fetchyr module — full API reference for login, MFA, workflows
- Choosing an approach — when to use authenticated scraping
- Batch extraction guide — crawl then extract pattern
Was this page helpful?
