Pagination
Handle large result sets with limit, offset, and auto-paginate
2 min read
Overview
Both kolektr.page() and webgrph.crawl() support pagination via limit and offset parameters. The Python SDK also provides page_all() and crawl_all() helpers for automatic pagination.
limit and offset
| Parameter | Default | Description |
|-----------|---------|-------------|
| limit | 100 | Max records to return per request |
| offset | 0 | Number of records to skip |
Example: manual pagination
from kloakd import Kloakd
client = Kloakd(api_key="sk-live-...", organization_id="your-org-id")
all_records = []
offset = 0
while True:
data = client.kolektr.page(
"https://example.com/products",
schema={"name": "css:.product-name", "price": "css:.price"},
limit=100,
offset=offset,
)
all_records.extend(data.records)
if not data.has_next:
break
offset += 100
print(f"Total: {len(all_records)} records")
TypeScript
let allRecords: Record[] = [];
let offset = 0;
while (true) {
const data = await client.kolektr.page('https://example.com/products', {
schema: { name: 'css:.product-name', price: 'css:.price' },
limit: 100,
offset,
});
allRecords = [...allRecords, ...data.records];
if (!data.hasNext) break;
offset += 100;
}
Auto-paginate helpers
Python
# Extract all records from a single page
all_records = client.kolektr.page_all(
"https://example.com/products",
schema={"name": "css:.product-name", "price": "css:.price"},
)
# Crawl all pages from a site
for page in client.webgrph.crawl_all("https://example.com", max_depth=3):
print(page.url, page.depth)
Response fields
| Field | Type | Description |
|-------|------|-------------|
| records | array | Current page of records |
| total | integer | Total records available |
| has_next | boolean | More pages available |
Rate limit awareness
Each paginated request counts against your rate limit. Use larger limit values to reduce API calls:
# Fewer requests, more records per request
data = client.kolektr.page(url, schema=schema, limit=500, offset=0)
Maximum limit is 1000 for Developer tier and above. Playground tier is capped at 100.
Next steps
- kolektr.page() reference — full API details
- webgrph.crawl() reference — full API details
- Batch extraction guide — crawl then extract pattern
