Kolektr
Structured data extraction — the engine behind extract()
2 min read
Kolektr is the extraction module that powers kolektr.page(). It extracts structured data from any web page using CSS selectors, XPath, or AI.
The high-level client.crawl() orchestrator can extract structured data from every page on a site in a single call — pass an extract_schema and Kolektr runs automatically on each fetched page.
For the full API reference, see kolektr.page().
Always pass fetch_artifact_id from a previous evadr.fetch() call to avoid re-fetching pages behind anti-bot protection.
page
result = client.kolektr.page(
url="https://example.com/products",
schema={
"name": "css:.product-name",
"price": "css:.product-price",
"rating": "css:.product-rating",
},
fetch_artifact_id=page.artifact_id,
limit=50,
offset=0,
)
print(result.total, result.has_more)
for record in result.records:
print(record)
page_all (auto-pagination)
all_records = client.kolektr.page_all(
url="https://example.com/products",
schema={"name": "css:.product-name", "price": "css:.product-price"},
)
print(f"Total: {len(all_records)} records")
extract_html
result = client.kolektr.extract_html(
html="<html>...</html>",
schema={"title": "css:h1", "body": "css:article"},
)
print(result.records)
Schema syntax
| Syntax | Example | Description | |--------|---------|-------------| | css: | css:.price | CSS selector | | xpath: | xpath://span | XPath expression | | ai: | ai:product price | AI-powered extraction | | attr: | attr:img@src | Element attribute |
Next steps
- crawl() orchestrator — extract from an entire site in one call with
extract_schema - kolektr.page() reference — full API endpoint docs
- Batch extraction guide — crawl then extract pattern
- Pagination guide — handling large result sets
