Skip to main content

Advanced Modules

Platform overview, module architecture, and how modules compose

3 min read


The orchestrator

client.crawl() is the high-level orchestrator that chains discovery, fetching, and optional extraction into a single call:

result = client.crawl(
    "https://example.com",
    max_depth=2,
    extract_schema={"title": "css:h1", "content": "css:article"},
)

It calls webgrph.crawl() for discovery, evadr.fetch() for each page, and kolektr.page() for extraction — all in one synchronous call. Use client.crawlStream() for real-time progress events.

The 3 low-level methods

Behind the orchestrator, KLOAKD has three low-level methods for fine-grained control:

| Method | Module | Returns | Use case | |--------|--------|---------|----------| | evadr.fetch() | Evadr | Raw HTML + artifact | Get a single page with anti-bot bypass | | webgrph.crawl() | Webgrph | Page hierarchy (async) | Map an entire site (requires polling) | | kolektr.page() | Kolektr | Structured records | Extract data from a single page |

See Choosing an Approach for guidance.

Advanced modules

Beyond the 3 core methods, KLOAKD has 4 specialized modules for advanced use cases:

| Module | Role | SDK namespace | |--------|------|---------------| | Skanyr | Discover hidden API endpoints from JS bundles and network activity | client.skanyr | | Nexus | AI strategy engine — analyze a site and decide the optimal extraction approach | client.nexus | | Parlyr | Conversational NLP — turn natural language prompts into extraction plans | client.parlyr | | Fetchyr | RPA & authentication — handle login flows, MFA, form automation | client.fetchyr |

Artifact chaining

The key to KLOAKD's efficiency is artifact reuse. When evadr.fetch() retrieves a page, it stores the result as an artifact. Every subsequent module can reference that artifact instead of re-fetching.

evadr.fetch(url)                  → artifact_id: "art-abc123"
kolektr.page(url, fetch_artifact_id=artifact_id)   → reuses fetch HTML
webgrph.crawl(url, session_artifact_id=artifact_id) → reuses fetch for root page

Zero redundant HTTP requests. Zero redundant anti-bot bypass attempts.

Composition patterns

Fetch → Extract

page = client.evadr.fetch("https://example.com")
data = client.kolektr.page("https://example.com", fetch_artifact_id=page.artifact_id)

Crawl → Extract (batch)

Use client.crawl() with extract_schema for a one-call solution:

result = client.crawl(
    "https://example.com",
    max_depth=2,
    extract_schema={"title": "css:h1", "content": "css:article"},
)
for page in result.pages:
    print(f"{page.url}: {page.structured_data}")

Or use the low-level webgrph.crawl() + kolektr.page() manually:

import time

crawl = client.webgrph.crawl("https://example.com", max_depth=2)

# Poll for crawl completion
while True:
    status = client.webgrph.get_crawl_status(crawl.crawl_id)
    if status.get("status") == "completed":
        break
    time.sleep(2)

pages = client.webgrph.get_crawl_pages(crawl.crawl_id)
for page in pages.get("pages", []):
    data = client.kolektr.page(page["url"], fetch_artifact_id=crawl.artifact_id)

Fetchyr → Extract (authenticated)

session = client.fetchyr.login(
    url="https://example.com/login",
    username_selector="#email",
    password_selector="#password",
    username="user@example.com",
    password="secret",
)
data = client.kolektr.page(
    "https://example.com/dashboard",
    session_artifact_id=session.artifact_id,
)

Skanyr → Kolektr (API data)

discovery = client.skanyr.discover("https://example.com")
data = client.kolektr.get_api_data("https://example.com/api/v1/products")

Error model

All errors inherit from a base KloakdError class:

| Error | HTTP | When | |-------|------|------| | AuthenticationError | 401 | Invalid API key | | ForbiddenError | 403 | Organization ID mismatch (IDOR protection) | | NotEntitledError | 403 | Plan doesn't include this module | | RateLimitError | 429 | Quota exceeded (includes retry_after) | | UpstreamError | 502 | Target site is unreachable | | ApiError | 4xx/5xx | Other server errors |

See Error Reference for full details.

Retry policy

All SDKs implement exponential backoff with a 1-hour cap:

  • Retryable: 429, 500, 502, 503, 504
  • Backoff: base_delay × 2^attempt, capped at 60s
  • 429 responses: respects Retry-After header and retry_after body field
  • Default: 3 retries
Was this page helpful?