Batch Extraction
Extract data from many pages using crawl then extract
3 min read
The client.crawl() orchestrator with extract_schema is the most efficient way to get structured data from an entire site — one call, no polling.
Overview
When you need data from many pages, you have two options:
- One-call approach (recommended):
client.crawl()withextract_schema— handles discovery, fetching, and extraction automatically - Manual approach:
webgrph.crawl()thenkolektr.page()each page — for fine-grained control over pagination and rate limits
One-call approach (recommended)
from kloakd import Kloakd
client = Kloakd(api_key="sk-live-...", organization_id="your-org-id")
result = client.crawl(
"https://example.com",
max_depth=2,
extract_schema={"title": "css:h1", "content": "css:article"},
)
print(f"Crawled {result.total_pages} pages")
for page in result.pages:
if page.extracted_data:
print(f" {page.url}: {page.extracted_data.get('title')}")
Streaming version
for event in client.crawl_stream(
"https://example.com",
max_depth=2,
extract_schema={"title": "css:h1", "content": "css:article"},
):
if event.type == "page_fetched" and event.extracted_data:
print(f" {event.url}: {event.extracted_data.get('title')}")
elif event.type == "crawl_complete":
print(f"Done! {event.total_pages} pages")
Manual approach
For cases where you need per-page pagination control or custom error handling:
Basic pattern
import time
from kloakd import Kloakd
client = Kloakd(api_key="sk-live-...", organization_id="your-org-id")
# Step 1 — Crawl the site (async, returns 202)
crawl = client.webgrph.crawl("https://example.com", max_depth=2)
print(f"Crawl started: {crawl.crawl_id}")
# Wait for crawl to complete
while True:
status = client.webgrph.get_crawl_status(crawl.crawl_id)
if status.get("status") == "completed":
break
time.sleep(2)
pages = client.webgrph.get_crawl_pages(crawl.crawl_id)
print(f"Found {len(pages.get('pages', []))} pages")
# Step 2 — Extract from each page
all_records = []
for page in pages.get("pages", []):
data = client.kolektr.page(
page["url"],
schema={"title": "css:h1", "content": "css:article"},
fetch_artifact_id=crawl.artifact_id,
)
all_records.extend(data.records)
print(f"Total records: {len(all_records)}")
Pattern with pagination
For pages with many records, use limit and offset:
crawl = client.webgrph.crawl("https://example.com/products", max_depth=1)
# Wait for crawl to complete
import time
while True:
status = client.webgrph.get_crawl_status(crawl.crawl_id)
if status.get("status") == "completed":
break
time.sleep(2)
pages = client.webgrph.get_crawl_pages(crawl.crawl_id)
for page in pages.get("pages", []):
offset = 0
while True:
data = client.kolektr.page(
page["url"],
schema={"name": "css:.product-name", "price": "css:.price"},
fetch_artifact_id=crawl.artifact_id,
limit=100,
offset=offset,
)
all_records.extend(data.records)
if not data.has_next:
break
offset += 100
Auto-paginate helper
Python SDK provides page_all() for automatic pagination:
all_records = client.kolektr.page_all(
"https://example.com/products",
schema={"name": "css:.product-name", "price": "css:.price"},
)
print(f"Total: {len(all_records)} records")
Rate limit awareness
Batch extraction makes many API calls. Watch for rate limits:
from kloakd.errors import RateLimitError
import time
pages = client.webgrph.get_crawl_pages(crawl.crawl_id)
for page in pages.get("pages", []):
try:
data = client.kolektr.page(
page["url"],
schema={"title": "css:h1"},
fetch_artifact_id=crawl.artifact_id,
)
except RateLimitError as e:
print(f"Rate limited. Waiting {e.retry_after}s...")
time.sleep(e.retry_after)
continue
Tips
- Use
client.crawl()withextract_schema— one call replaces the entire crawl-poll-extract loop - Use artifact chaining — pass
fetch_artifact_idfrom crawl to extract in the manual approach - Limit crawl depth —
max_depth=2is usually enough for most sites - Use auto-paginate —
page_all()handles pagination automatically for single-page extraction - Handle rate limits — catch
RateLimitErrorand retry
Next steps
- crawl() reference — full crawl orchestrator API
- Pagination guide — deep dive into limit/offset
- extract() reference — full extract API
Was this page helpful?
