Python SDK
kloakd-sdk for Python — one API for everything, full module coverage with async support
4 min read
Install
pip install kloakd-sdk
Python 3.10+ | 88/88 tests | 87% coverage | httpx transport
Client
from kloakd import Kloakd
client = Kloakd(
api_key="sk-live-...",
organization_id="your-org-uuid",
timeout=30.0,
max_retries=3,
)
Quickstart — one call does everything
result = client.crawl(
"https://example.com",
max_pages=50,
extract_schema={"title": "css:h1", "content": "css:article"},
)
for page in result.pages:
if page.ok:
print(f"{page.url} → {page.structured_data}")
That's it. crawl() internally:
- Discovers all pages on the site (Webgrph BFS)
- Fetches each page through the 4-tier anti-bot engine (auto-escalates to headless browser for Cloudflare-protected sites — no
force_browserneeded) - Extracts structured data from each page if
extract_schemais provided (Kolektr, using cached fetch artifacts — no second HTTP round-trip)
Per-page failures are caught and marked success=False — the crawl never aborts on a single page error.
Primary method
client.crawl()
result = client.crawl(
"https://example.com",
max_depth=3,
max_pages=100,
extract_schema={"title": "css:h1", "price": "css:.price"},
include_external_links=False,
session_artifact_id=None, # reuse a Fetchyr session for login-protected sites
)
print(f"Discovered: {result.total_pages_discovered} pages")
print(f"Fetched: {result.pages_fetched}")
print(f"Failed: {result.pages_failed}")
for page in result.pages:
if page.ok:
print(f" {page.url} [{page.tier_used}] {page.structured_data}")
else:
print(f" {page.url} FAILED: {page.error}")
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
| url | str | required | Seed URL to start crawling from |
| max_depth | int | 3 | Maximum BFS depth |
| max_pages | int | 100 | Maximum pages to crawl |
| extract_schema | dict[str, str] | None | CSS selector schema for structured extraction |
| include_external_links | bool | False | Follow off-domain links |
| session_artifact_id | str | None | Reuse an authenticated session artifact (from Fetchyr) |
Returns: SiteCrawlResult
| Field | Type | Description |
|---|---|---|
| success | bool | Whether the crawl completed |
| url | str | Seed URL |
| total_pages_discovered | int | Pages found during BFS |
| pages_fetched | int | Pages successfully fetched |
| pages_failed | int | Pages that failed (included with success=False) |
| pages | list[CrawlPage] | All pages with HTML and optional structured data |
| crawl_artifact_id | str | Artifact ID for the site hierarchy |
| error | str | Error message if crawl failed entirely |
client.crawl_stream() — async streaming
For long-running crawls, use the async streaming version to receive real-time progress events:
import asyncio
from kloakd import AsyncKloakd
client = AsyncKloakd(api_key="sk-live-...", organization_id="...")
async def main():
async for event in client.crawl_stream(
"https://example.com",
max_pages=500,
extract_schema={"title": "css:h1"},
):
if event.type == "page_fetched":
print(f"[{event.page}/{event.total}] {event.url} OK")
elif event.type == "page_failed":
print(f"[{event.page}/{event.total}] {event.url} FAIL: {event.error}")
elif event.type == "crawl_complete":
result = event.metadata["result"]
print(f"Done: {result.pages_fetched} fetched, {result.pages_failed} failed")
asyncio.run(main())
Or use the sync client's generator:
from kloakd import Kloakd
client = Kloakd(api_key="sk-live-...", organization_id="...")
for event in client.crawl_stream("https://example.com", max_pages=500):
if event.type == "page_fetched":
print(f"[{event.page}/{event.total}] {event.url} OK")
elif event.type == "crawl_complete":
print("Done!")
Event types
| Type | Description |
|---|---|
| discovery_started | Crawl discovery has begun |
| discovery_progress | Pages found during BFS (pages_found field) |
| discovery_complete | Discovery finished, fetch phase starting |
| page_fetching | About to fetch page N of total |
| page_fetched | Page fetched successfully |
| page_failed | Page fetch failed (crawl continues) |
| crawl_complete | All pages processed, final summary in metadata["result"] |
Low-level methods
Need fine-grained control? The individual modules are still available:
evadr.fetch()
page = client.evadr.fetch("https://example.com")
print(f"Status: {page.status_code}, Tier: {page.tier_used}")
print(f"Artifact ID: {page.artifact_id}")
webgrph.crawl()
crawl = client.webgrph.crawl(
"https://example.com",
max_depth=3,
max_pages=100,
)
print(f"Crawl started: {crawl.crawl_id}")
kolektr.page()
result = client.kolektr.page(
"https://example.com",
schema={"title": "css:h1", "price": "css:.price"},
)
for record in result.records:
print(record)
Artifact chaining
Pass artifact IDs between low-level methods to skip redundant work:
page = client.evadr.fetch("https://example.com")
data = client.kolektr.page(
"https://example.com",
schema={"title": "css:h1"},
fetch_artifact_id=page.artifact_id,
)
crawl = client.webgrph.crawl(
"https://example.com",
max_depth=2,
session_artifact_id=page.artifact_id,
)
Error handling
from kloakd.errors import AuthenticationError, RateLimitError, KloakdError
try:
result = client.crawl("https://example.com")
except AuthenticationError:
print("Invalid API key")
except RateLimitError as e:
print(f"Rate limited. Retry after {e.retry_after}s")
except KloakdError as e:
print(f"Error: {e}")
Advanced namespaces
client.skanyr # API discovery
client.nexus # AI strategy engine
client.parlyr # Natural language queries
client.fetchyr # RPA & authenticated scraping
Next steps
- Quickstart — get started in 5 minutes
- API reference — detailed endpoint docs
- Error reference — full error taxonomy
