Skip to main content

Quickstart

First crawl in 5 minutes

3 min read


KLOAKD's crawl() method is a single call that discovers pages, fetches them through the anti-bot engine, and extracts structured data. This guide covers that plus low-level methods for fine-grained control.

1. Get your API key

Try instantly — visit the Playground for free discoveries with zero setup. No API key required.

For programmatic access — sign in to the dashboard, create an organization, and copy your API key from Settings → API Keys. API access requires a Developer plan or above.

2. Install the SDK

Python

pip install kloakd-sdk

TypeScript / Node.js

npm install kloakd-sdk

Go

go get github.com/kloakd/kloakd-go

Java (Maven)

<dependency>
  <groupId>dev.kloakd</groupId>
  <artifactId>kloakd-sdk</artifactId>
  <version>0.1.0</version>
</dependency>

3. Crawl a site — one call does everything

import os
from kloakd import Kloakd

client = Kloakd(
    api_key=os.environ["KLOAKD_API_KEY"],
    organization_id=os.environ["KLOAKD_ORG_ID"],
)

result = client.crawl(
    "https://books.toscrape.com",
    max_pages=50,
    extract_schema={"title": "css:h3 a", "price": "css:p.price_color"},
)

print(f"Discovered: {result.total_pages_discovered} pages")
print(f"Fetched:    {result.pages_fetched}")
print(f"Failed:     {result.pages_failed}")

for page in result.pages[:5]:
    if page.ok:
        print(f"  {page.url} → {page.structured_data}")

TypeScript equivalent

import { Kloakd } from 'kloakd-sdk';

const client = new Kloakd({
  apiKey: process.env.KLOAKD_API_KEY!,
  organizationId: process.env.KLOAKD_ORG_ID!,
});

const result = await client.crawl('https://books.toscrape.com', {
  maxPages: 50,
  extractSchema: { title: 'css:h3 a', price: 'css:p.price_color' },
});

console.log(`Discovered: ${result.totalPagesDiscovered} pages`);
console.log(`Fetched:    ${result.pagesFetched}`);
console.log(`Failed:     ${result.pagesFailed}`);

for (const page of result.pages.slice(0, 5)) {
  if (page.ok) {
    console.log(`  ${page.url} →`, page.structuredData);
  }
}

What just happened?

crawl() is a fully managed pipeline:

  1. Discover — BFS traversal finds all pages on the site (Webgrph)
  2. Fetch — Each page is fetched through the 4-tier anti-bot engine (Evadr). Auto-escalates from HTTP → headless browser → residential proxy as needed. No configuration required.
  3. Extract — If extract_schema is provided, structured data is extracted from each page (Kolektr), reusing the cached fetch artifact — no second HTTP round-trip

Per-page failures are caught and marked success=False — the crawl never aborts on a single page error.

4. Stream crawl progress (optional)

For long-running crawls, stream real-time progress events:

import asyncio
from kloakd import AsyncKloakd

client = AsyncKloakd(
    api_key=os.environ["KLOAKD_API_KEY"],
    organization_id=os.environ["KLOAKD_ORG_ID"],
)

async def main():
    async for event in client.crawl_stream(
        "https://example.com",
        max_pages=500,
        extract_schema={"title": "css:h1"},
    ):
        if event.type == "page_fetched":
            print(f"[{event.page}/{event.total}] {event.url} OK")
        elif event.type == "page_failed":
            print(f"[{event.page}/{event.total}] {event.url} FAIL: {event.error}")
        elif event.type == "crawl_complete":
            print("Done!")

asyncio.run(main())

Or use the sync client's generator:

from kloakd import Kloakd

client = Kloakd(
    api_key=os.environ["KLOAKD_API_KEY"],
    organization_id=os.environ["KLOAKD_ORG_ID"],
)

for event in client.crawl_stream("https://example.com", max_pages=500):
    if event.type == "page_fetched":
        print(f"[{event.page}/{event.total}] {event.url} OK")
    elif event.type == "crawl_complete":
        print("Done!")

5. Low-level methods (advanced)

Need fine-grained control? Use the individual modules directly:

Fetch a single page with anti-bot bypass

page = client.evadr.fetch("https://example.com")
print(f"Status: {page.status_code}, Tier: {page.tier_used}")
print(f"Anti-bot bypassed: {page.anti_bot_bypassed}")
print(f"HTML length: {len(page.html)}")

Extract structured data from a single page

result = client.kolektr.page(
    "https://example.com",
    schema={"title": "css:h1", "content": "css:article"},
)
for record in result.records:
    print(record)

Artifact chaining — skip redundant work

Pass artifact IDs between methods to reuse fetched HTML:

# Fetch once (with anti-bot bypass)
page = client.evadr.fetch("https://protected-site.com")

# Extract data — reuses the fetched HTML, no second request
data = client.kolektr.page(
    "https://protected-site.com",
    schema={"title": "css:h1", "content": "css:article"},
    fetch_artifact_id=page.artifact_id,
)

Next steps

Was this page helpful?