Skip to main content

extract()

Extract structured data from any web page

3 min read


Extracting from many pages? Pass an extract_schema to client.crawl() and it will extract structured data from every page on the site in a single call — no manual iteration needed.

Endpoint

POST /api/v1/organizations/{org_id}/kolektr/page

Description

Extract structured data from a web page using CSS selectors, XPath, or AI-powered extraction. Pass a fetch_artifact_id from a previous fetch() call to skip re-fetching the page.

Request body

| Field | Type | Required | Default | Description | |-------|------|----------|---------|-------------| | url | string | Yes | — | Target page URL | | schema | object | No | null | Field-to-selector mapping | | fetch_artifact_id | string | No | null | Reuse HTML from a prior fetch | | session_artifact_id | string | No | null | Authenticated session from Fetchyr | | api_map_artifact_id | string | No | null | API map from Skanyr discovery | | limit | integer | No | 100 | Max records to return | | offset | integer | No | 0 | Pagination offset |

Schema syntax

| Syntax | Example | Description | |--------|---------|-------------| | css: | css:.price | CSS selector | | xpath: | xpath://span | XPath expression | | ai: | ai:product price | AI-powered extraction | | attr: | attr:img@src | Element attribute |

Response

{
  "records": [
    { "title": "Book A", "price": "$12.99" },
    { "title": "Book B", "price": "$8.50" }
  ],
  "total": 2,
  "has_next": false,
  "artifact_id": "art-abc123"
}

Response fields

| Field | Type | Description | |-------|------|-------------| | records | array | Extracted data records | | total | integer | Total records found | | has_next | boolean | More pages available | | artifact_id | string | Reusable artifact ID |

Examples

Python

from kloakd import Kloakd

client = Kloakd(api_key="sk-live-...", organization_id="your-org-id")

result = client.kolektr.page(
    "https://books.toscrape.com",
    schema={"title": "css:h3 a", "price": "css:p.price_color"},
)
for record in result.records:
    print(record)

TypeScript

import { Kloakd } from 'kloakd-sdk';

const client = new Kloakd({
  apiKey: 'sk-live-...',
  organizationId: 'your-org-id',
});

const result = await client.kolektr.page(
  'https://books.toscrape.com',
  { schema: { title: 'css:h3 a', price: 'css:p.price_color' } }
);
console.log(result.records);

Go

client := kloakd.MustNew(kloakd.Config{
    APIKey:         "sk-live-...",
    OrganizationID: "your-org-id",
})

result, _ := client.Kolektr.Page(ctx, "https://books.toscrape.com",
    &kloakd.PageOptions{
        Schema: map[string]string{"title": "css:h3 a", "price": "css:p.price_color"},
    })
for _, r := range result.Records {
    fmt.Println(r)
}

cURL

curl -X POST https://api.kloakd.dev/api/v1/organizations/$ORG_ID/kolektr/page \
  -H "Authorization: Bearer sk-live-..." \
  -H "Content-Type: application/json" \
  -d '{"url": "https://books.toscrape.com", "schema": {"title": "css:h3 a", "price": "css:p.price_color"}}'

Artifact chaining

Pass a fetch_artifact_id to skip re-fetching:

page = client.evadr.fetch("https://books.toscrape.com")
result = client.kolektr.page(
    "https://books.toscrape.com",
    schema={"title": "css:h3 a", "price": "css:p.price_color"},
    fetch_artifact_id=page.artifact_id,
)

Errors

| Status | Error | When | |--------|-------|------| | 400 | ValidationError | Invalid URL or schema | | 401 | AuthenticationError | Missing or invalid API key | | 403 | ForbiddenError | Org ID mismatch | | 429 | RateLimitError | Quota exceeded | | 502 | UpstreamError | Target site unreachable |

Learn more

Was this page helpful?