Skip to main content

Batch Extraction

Extract data from many pages using crawl then extract

3 min read


The client.crawl() orchestrator with extract_schema is the most efficient way to get structured data from an entire site — one call, no polling.

Overview

When you need data from many pages, you have two options:

  1. One-call approach (recommended): client.crawl() with extract_schema — handles discovery, fetching, and extraction automatically
  2. Manual approach: webgrph.crawl() then kolektr.page() each page — for fine-grained control over pagination and rate limits
from kloakd import Kloakd

client = Kloakd(api_key="sk-live-...", organization_id="your-org-id")

result = client.crawl(
    "https://example.com",
    max_depth=2,
    extract_schema={"title": "css:h1", "content": "css:article"},
)

print(f"Crawled {result.total_pages} pages")
for page in result.pages:
    if page.extracted_data:
        print(f"  {page.url}: {page.extracted_data.get('title')}")

Streaming version

for event in client.crawl_stream(
    "https://example.com",
    max_depth=2,
    extract_schema={"title": "css:h1", "content": "css:article"},
):
    if event.type == "page_fetched" and event.extracted_data:
        print(f"  {event.url}: {event.extracted_data.get('title')}")
    elif event.type == "crawl_complete":
        print(f"Done! {event.total_pages} pages")

Manual approach

For cases where you need per-page pagination control or custom error handling:

Basic pattern

import time
from kloakd import Kloakd

client = Kloakd(api_key="sk-live-...", organization_id="your-org-id")

# Step 1 — Crawl the site (async, returns 202)
crawl = client.webgrph.crawl("https://example.com", max_depth=2)
print(f"Crawl started: {crawl.crawl_id}")

# Wait for crawl to complete
while True:
    status = client.webgrph.get_crawl_status(crawl.crawl_id)
    if status.get("status") == "completed":
        break
    time.sleep(2)

pages = client.webgrph.get_crawl_pages(crawl.crawl_id)
print(f"Found {len(pages.get('pages', []))} pages")

# Step 2 — Extract from each page
all_records = []
for page in pages.get("pages", []):
    data = client.kolektr.page(
        page["url"],
        schema={"title": "css:h1", "content": "css:article"},
        fetch_artifact_id=crawl.artifact_id,
    )
    all_records.extend(data.records)

print(f"Total records: {len(all_records)}")

Pattern with pagination

For pages with many records, use limit and offset:

crawl = client.webgrph.crawl("https://example.com/products", max_depth=1)

# Wait for crawl to complete
import time
while True:
    status = client.webgrph.get_crawl_status(crawl.crawl_id)
    if status.get("status") == "completed":
        break
    time.sleep(2)

pages = client.webgrph.get_crawl_pages(crawl.crawl_id)

for page in pages.get("pages", []):
    offset = 0
    while True:
        data = client.kolektr.page(
            page["url"],
            schema={"name": "css:.product-name", "price": "css:.price"},
            fetch_artifact_id=crawl.artifact_id,
            limit=100,
            offset=offset,
        )
        all_records.extend(data.records)
        if not data.has_next:
            break
        offset += 100

Auto-paginate helper

Python SDK provides page_all() for automatic pagination:

all_records = client.kolektr.page_all(
    "https://example.com/products",
    schema={"name": "css:.product-name", "price": "css:.price"},
)
print(f"Total: {len(all_records)} records")

Rate limit awareness

Batch extraction makes many API calls. Watch for rate limits:

from kloakd.errors import RateLimitError
import time

pages = client.webgrph.get_crawl_pages(crawl.crawl_id)
for page in pages.get("pages", []):
    try:
        data = client.kolektr.page(
            page["url"],
            schema={"title": "css:h1"},
            fetch_artifact_id=crawl.artifact_id,
        )
    except RateLimitError as e:
        print(f"Rate limited. Waiting {e.retry_after}s...")
        time.sleep(e.retry_after)
        continue

Tips

  • Use client.crawl() with extract_schema — one call replaces the entire crawl-poll-extract loop
  • Use artifact chaining — pass fetch_artifact_id from crawl to extract in the manual approach
  • Limit crawl depth — max_depth=2 is usually enough for most sites
  • Use auto-paginate — page_all() handles pagination automatically for single-page extraction
  • Handle rate limits — catch RateLimitError and retry

Next steps

Was this page helpful?