Skip to main content

Python SDK

kloakd-sdk for Python — one API for everything, full module coverage with async support

4 min read


Install

pip install kloakd-sdk

Python 3.10+ | 88/88 tests | 87% coverage | httpx transport

Client

from kloakd import Kloakd

client = Kloakd(
    api_key="sk-live-...",
    organization_id="your-org-uuid",
    timeout=30.0,
    max_retries=3,
)

Quickstart — one call does everything

result = client.crawl(
    "https://example.com",
    max_pages=50,
    extract_schema={"title": "css:h1", "content": "css:article"},
)

for page in result.pages:
    if page.ok:
        print(f"{page.url} → {page.structured_data}")

That's it. crawl() internally:

  1. Discovers all pages on the site (Webgrph BFS)
  2. Fetches each page through the 4-tier anti-bot engine (auto-escalates to headless browser for Cloudflare-protected sites — no force_browser needed)
  3. Extracts structured data from each page if extract_schema is provided (Kolektr, using cached fetch artifacts — no second HTTP round-trip)

Per-page failures are caught and marked success=False — the crawl never aborts on a single page error.

Primary method

client.crawl()

result = client.crawl(
    "https://example.com",
    max_depth=3,
    max_pages=100,
    extract_schema={"title": "css:h1", "price": "css:.price"},
    include_external_links=False,
    session_artifact_id=None,  # reuse a Fetchyr session for login-protected sites
)

print(f"Discovered: {result.total_pages_discovered} pages")
print(f"Fetched:    {result.pages_fetched}")
print(f"Failed:     {result.pages_failed}")

for page in result.pages:
    if page.ok:
        print(f"  {page.url} [{page.tier_used}] {page.structured_data}")
    else:
        print(f"  {page.url} FAILED: {page.error}")

Parameters

| Parameter | Type | Default | Description | |---|---|---|---| | url | str | required | Seed URL to start crawling from | | max_depth | int | 3 | Maximum BFS depth | | max_pages | int | 100 | Maximum pages to crawl | | extract_schema | dict[str, str] | None | CSS selector schema for structured extraction | | include_external_links | bool | False | Follow off-domain links | | session_artifact_id | str | None | Reuse an authenticated session artifact (from Fetchyr) |

Returns: SiteCrawlResult

| Field | Type | Description | |---|---|---| | success | bool | Whether the crawl completed | | url | str | Seed URL | | total_pages_discovered | int | Pages found during BFS | | pages_fetched | int | Pages successfully fetched | | pages_failed | int | Pages that failed (included with success=False) | | pages | list[CrawlPage] | All pages with HTML and optional structured data | | crawl_artifact_id | str | Artifact ID for the site hierarchy | | error | str | Error message if crawl failed entirely |

client.crawl_stream() — async streaming

For long-running crawls, use the async streaming version to receive real-time progress events:

import asyncio
from kloakd import AsyncKloakd

client = AsyncKloakd(api_key="sk-live-...", organization_id="...")

async def main():
    async for event in client.crawl_stream(
        "https://example.com",
        max_pages=500,
        extract_schema={"title": "css:h1"},
    ):
        if event.type == "page_fetched":
            print(f"[{event.page}/{event.total}] {event.url} OK")
        elif event.type == "page_failed":
            print(f"[{event.page}/{event.total}] {event.url} FAIL: {event.error}")
        elif event.type == "crawl_complete":
            result = event.metadata["result"]
            print(f"Done: {result.pages_fetched} fetched, {result.pages_failed} failed")

asyncio.run(main())

Or use the sync client's generator:

from kloakd import Kloakd

client = Kloakd(api_key="sk-live-...", organization_id="...")

for event in client.crawl_stream("https://example.com", max_pages=500):
    if event.type == "page_fetched":
        print(f"[{event.page}/{event.total}] {event.url} OK")
    elif event.type == "crawl_complete":
        print("Done!")

Event types

| Type | Description | |---|---| | discovery_started | Crawl discovery has begun | | discovery_progress | Pages found during BFS (pages_found field) | | discovery_complete | Discovery finished, fetch phase starting | | page_fetching | About to fetch page N of total | | page_fetched | Page fetched successfully | | page_failed | Page fetch failed (crawl continues) | | crawl_complete | All pages processed, final summary in metadata["result"] |

Low-level methods

Need fine-grained control? The individual modules are still available:

evadr.fetch()

page = client.evadr.fetch("https://example.com")
print(f"Status: {page.status_code}, Tier: {page.tier_used}")
print(f"Artifact ID: {page.artifact_id}")

webgrph.crawl()

crawl = client.webgrph.crawl(
    "https://example.com",
    max_depth=3,
    max_pages=100,
)
print(f"Crawl started: {crawl.crawl_id}")

kolektr.page()

result = client.kolektr.page(
    "https://example.com",
    schema={"title": "css:h1", "price": "css:.price"},
)
for record in result.records:
    print(record)

Artifact chaining

Pass artifact IDs between low-level methods to skip redundant work:

page = client.evadr.fetch("https://example.com")

data = client.kolektr.page(
    "https://example.com",
    schema={"title": "css:h1"},
    fetch_artifact_id=page.artifact_id,
)

crawl = client.webgrph.crawl(
    "https://example.com",
    max_depth=2,
    session_artifact_id=page.artifact_id,
)

Error handling

from kloakd.errors import AuthenticationError, RateLimitError, KloakdError

try:
    result = client.crawl("https://example.com")
except AuthenticationError:
    print("Invalid API key")
except RateLimitError as e:
    print(f"Rate limited. Retry after {e.retry_after}s")
except KloakdError as e:
    print(f"Error: {e}")

Advanced namespaces

client.skanyr     # API discovery
client.nexus      # AI strategy engine
client.parlyr     # Natural language queries
client.fetchyr    # RPA & authenticated scraping

Next steps

Was this page helpful?