Quickstart
First crawl in 5 minutes
3 min read
KLOAKD's crawl() method is a single call that discovers pages, fetches them through the anti-bot engine, and extracts structured data. This guide covers that plus low-level methods for fine-grained control.
1. Get your API key
Try instantly — visit the Playground for free discoveries with zero setup. No API key required.
For programmatic access — sign in to the dashboard, create an organization, and copy your API key from Settings → API Keys. API access requires a Developer plan or above.
2. Install the SDK
Python
pip install kloakd-sdk
TypeScript / Node.js
npm install kloakd-sdk
Go
go get github.com/kloakd/kloakd-go
Java (Maven)
<dependency>
<groupId>dev.kloakd</groupId>
<artifactId>kloakd-sdk</artifactId>
<version>0.1.0</version>
</dependency>
3. Crawl a site — one call does everything
import os
from kloakd import Kloakd
client = Kloakd(
api_key=os.environ["KLOAKD_API_KEY"],
organization_id=os.environ["KLOAKD_ORG_ID"],
)
result = client.crawl(
"https://books.toscrape.com",
max_pages=50,
extract_schema={"title": "css:h3 a", "price": "css:p.price_color"},
)
print(f"Discovered: {result.total_pages_discovered} pages")
print(f"Fetched: {result.pages_fetched}")
print(f"Failed: {result.pages_failed}")
for page in result.pages[:5]:
if page.ok:
print(f" {page.url} → {page.structured_data}")
TypeScript equivalent
import { Kloakd } from 'kloakd-sdk';
const client = new Kloakd({
apiKey: process.env.KLOAKD_API_KEY!,
organizationId: process.env.KLOAKD_ORG_ID!,
});
const result = await client.crawl('https://books.toscrape.com', {
maxPages: 50,
extractSchema: { title: 'css:h3 a', price: 'css:p.price_color' },
});
console.log(`Discovered: ${result.totalPagesDiscovered} pages`);
console.log(`Fetched: ${result.pagesFetched}`);
console.log(`Failed: ${result.pagesFailed}`);
for (const page of result.pages.slice(0, 5)) {
if (page.ok) {
console.log(` ${page.url} →`, page.structuredData);
}
}
What just happened?
crawl() is a fully managed pipeline:
- Discover — BFS traversal finds all pages on the site (Webgrph)
- Fetch — Each page is fetched through the 4-tier anti-bot engine (Evadr). Auto-escalates from HTTP → headless browser → residential proxy as needed. No configuration required.
- Extract — If
extract_schemais provided, structured data is extracted from each page (Kolektr), reusing the cached fetch artifact — no second HTTP round-trip
Per-page failures are caught and marked success=False — the crawl never aborts on a single page error.
4. Stream crawl progress (optional)
For long-running crawls, stream real-time progress events:
import asyncio
from kloakd import AsyncKloakd
client = AsyncKloakd(
api_key=os.environ["KLOAKD_API_KEY"],
organization_id=os.environ["KLOAKD_ORG_ID"],
)
async def main():
async for event in client.crawl_stream(
"https://example.com",
max_pages=500,
extract_schema={"title": "css:h1"},
):
if event.type == "page_fetched":
print(f"[{event.page}/{event.total}] {event.url} OK")
elif event.type == "page_failed":
print(f"[{event.page}/{event.total}] {event.url} FAIL: {event.error}")
elif event.type == "crawl_complete":
print("Done!")
asyncio.run(main())
Or use the sync client's generator:
from kloakd import Kloakd
client = Kloakd(
api_key=os.environ["KLOAKD_API_KEY"],
organization_id=os.environ["KLOAKD_ORG_ID"],
)
for event in client.crawl_stream("https://example.com", max_pages=500):
if event.type == "page_fetched":
print(f"[{event.page}/{event.total}] {event.url} OK")
elif event.type == "crawl_complete":
print("Done!")
5. Low-level methods (advanced)
Need fine-grained control? Use the individual modules directly:
Fetch a single page with anti-bot bypass
page = client.evadr.fetch("https://example.com")
print(f"Status: {page.status_code}, Tier: {page.tier_used}")
print(f"Anti-bot bypassed: {page.anti_bot_bypassed}")
print(f"HTML length: {len(page.html)}")
Extract structured data from a single page
result = client.kolektr.page(
"https://example.com",
schema={"title": "css:h1", "content": "css:article"},
)
for record in result.records:
print(record)
Artifact chaining — skip redundant work
Pass artifact IDs between methods to reuse fetched HTML:
# Fetch once (with anti-bot bypass)
page = client.evadr.fetch("https://protected-site.com")
# Extract data — reuses the fetched HTML, no second request
data = client.kolektr.page(
"https://protected-site.com",
schema={"title": "css:h1", "content": "css:article"},
fetch_artifact_id=page.artifact_id,
)
Next steps
- SDK docs — full SDK reference for Python, TypeScript, Go, Java
- Authentication — API keys and org context
- Error reference — full error taxonomy
