extract()
Extract structured data from any web page
3 min read
Extracting from many pages? Pass an extract_schema to client.crawl() and it will extract structured data from every page on the site in a single call — no manual iteration needed.
Endpoint
POST /api/v1/organizations/{org_id}/kolektr/page
Description
Extract structured data from a web page using CSS selectors, XPath, or AI-powered extraction. Pass a fetch_artifact_id from a previous fetch() call to skip re-fetching the page.
Request body
| Field | Type | Required | Default | Description |
|-------|------|----------|---------|-------------|
| url | string | Yes | — | Target page URL |
| schema | object | No | null | Field-to-selector mapping |
| fetch_artifact_id | string | No | null | Reuse HTML from a prior fetch |
| session_artifact_id | string | No | null | Authenticated session from Fetchyr |
| api_map_artifact_id | string | No | null | API map from Skanyr discovery |
| limit | integer | No | 100 | Max records to return |
| offset | integer | No | 0 | Pagination offset |
Schema syntax
| Syntax | Example | Description |
|--------|---------|-------------|
| css: | css:.price | CSS selector |
| xpath: | xpath://span | XPath expression |
| ai: | ai:product price | AI-powered extraction |
| attr: | attr:img@src | Element attribute |
Response
{
"records": [
{ "title": "Book A", "price": "$12.99" },
{ "title": "Book B", "price": "$8.50" }
],
"total": 2,
"has_next": false,
"artifact_id": "art-abc123"
}
Response fields
| Field | Type | Description |
|-------|------|-------------|
| records | array | Extracted data records |
| total | integer | Total records found |
| has_next | boolean | More pages available |
| artifact_id | string | Reusable artifact ID |
Examples
Python
from kloakd import Kloakd
client = Kloakd(api_key="sk-live-...", organization_id="your-org-id")
result = client.kolektr.page(
"https://books.toscrape.com",
schema={"title": "css:h3 a", "price": "css:p.price_color"},
)
for record in result.records:
print(record)
TypeScript
import { Kloakd } from 'kloakd-sdk';
const client = new Kloakd({
apiKey: 'sk-live-...',
organizationId: 'your-org-id',
});
const result = await client.kolektr.page(
'https://books.toscrape.com',
{ schema: { title: 'css:h3 a', price: 'css:p.price_color' } }
);
console.log(result.records);
Go
client := kloakd.MustNew(kloakd.Config{
APIKey: "sk-live-...",
OrganizationID: "your-org-id",
})
result, _ := client.Kolektr.Page(ctx, "https://books.toscrape.com",
&kloakd.PageOptions{
Schema: map[string]string{"title": "css:h3 a", "price": "css:p.price_color"},
})
for _, r := range result.Records {
fmt.Println(r)
}
cURL
curl -X POST https://api.kloakd.dev/api/v1/organizations/$ORG_ID/kolektr/page \
-H "Authorization: Bearer sk-live-..." \
-H "Content-Type: application/json" \
-d '{"url": "https://books.toscrape.com", "schema": {"title": "css:h3 a", "price": "css:p.price_color"}}'
Artifact chaining
Pass a fetch_artifact_id to skip re-fetching:
page = client.evadr.fetch("https://books.toscrape.com")
result = client.kolektr.page(
"https://books.toscrape.com",
schema={"title": "css:h3 a", "price": "css:p.price_color"},
fetch_artifact_id=page.artifact_id,
)
Errors
| Status | Error | When |
|--------|-------|------|
| 400 | ValidationError | Invalid URL or schema |
| 401 | AuthenticationError | Missing or invalid API key |
| 403 | ForbiddenError | Org ID mismatch |
| 429 | RateLimitError | Quota exceeded |
| 502 | UpstreamError | Target site unreachable |
Learn more
- crawl() orchestrator — extract from an entire site in one call with
extract_schema - Kolektr module — advanced extraction options
- Pagination guide — handling large result sets
- Batch extraction guide — crawl then extract pattern
