Advanced Modules
Platform overview, module architecture, and how modules compose
3 min read
The orchestrator
client.crawl() is the high-level orchestrator that chains discovery, fetching, and optional extraction into a single call:
result = client.crawl(
"https://example.com",
max_depth=2,
extract_schema={"title": "css:h1", "content": "css:article"},
)
It calls webgrph.crawl() for discovery, evadr.fetch() for each page, and kolektr.page() for extraction — all in one synchronous call. Use client.crawlStream() for real-time progress events.
The 3 low-level methods
Behind the orchestrator, KLOAKD has three low-level methods for fine-grained control:
| Method | Module | Returns | Use case |
|--------|--------|---------|----------|
| evadr.fetch() | Evadr | Raw HTML + artifact | Get a single page with anti-bot bypass |
| webgrph.crawl() | Webgrph | Page hierarchy (async) | Map an entire site (requires polling) |
| kolektr.page() | Kolektr | Structured records | Extract data from a single page |
See Choosing an Approach for guidance.
Advanced modules
Beyond the 3 core methods, KLOAKD has 4 specialized modules for advanced use cases:
| Module | Role | SDK namespace |
|--------|------|---------------|
| Skanyr | Discover hidden API endpoints from JS bundles and network activity | client.skanyr |
| Nexus | AI strategy engine — analyze a site and decide the optimal extraction approach | client.nexus |
| Parlyr | Conversational NLP — turn natural language prompts into extraction plans | client.parlyr |
| Fetchyr | RPA & authentication — handle login flows, MFA, form automation | client.fetchyr |
Artifact chaining
The key to KLOAKD's efficiency is artifact reuse. When evadr.fetch() retrieves a page, it stores the result as an artifact. Every subsequent module can reference that artifact instead of re-fetching.
evadr.fetch(url) → artifact_id: "art-abc123"
kolektr.page(url, fetch_artifact_id=artifact_id) → reuses fetch HTML
webgrph.crawl(url, session_artifact_id=artifact_id) → reuses fetch for root page
Zero redundant HTTP requests. Zero redundant anti-bot bypass attempts.
Composition patterns
Fetch → Extract
page = client.evadr.fetch("https://example.com")
data = client.kolektr.page("https://example.com", fetch_artifact_id=page.artifact_id)
Crawl → Extract (batch)
Use client.crawl() with extract_schema for a one-call solution:
result = client.crawl(
"https://example.com",
max_depth=2,
extract_schema={"title": "css:h1", "content": "css:article"},
)
for page in result.pages:
print(f"{page.url}: {page.structured_data}")
Or use the low-level webgrph.crawl() + kolektr.page() manually:
import time
crawl = client.webgrph.crawl("https://example.com", max_depth=2)
# Poll for crawl completion
while True:
status = client.webgrph.get_crawl_status(crawl.crawl_id)
if status.get("status") == "completed":
break
time.sleep(2)
pages = client.webgrph.get_crawl_pages(crawl.crawl_id)
for page in pages.get("pages", []):
data = client.kolektr.page(page["url"], fetch_artifact_id=crawl.artifact_id)
Fetchyr → Extract (authenticated)
session = client.fetchyr.login(
url="https://example.com/login",
username_selector="#email",
password_selector="#password",
username="user@example.com",
password="secret",
)
data = client.kolektr.page(
"https://example.com/dashboard",
session_artifact_id=session.artifact_id,
)
Skanyr → Kolektr (API data)
discovery = client.skanyr.discover("https://example.com")
data = client.kolektr.get_api_data("https://example.com/api/v1/products")
Error model
All errors inherit from a base KloakdError class:
| Error | HTTP | When |
|-------|------|------|
| AuthenticationError | 401 | Invalid API key |
| ForbiddenError | 403 | Organization ID mismatch (IDOR protection) |
| NotEntitledError | 403 | Plan doesn't include this module |
| RateLimitError | 429 | Quota exceeded (includes retry_after) |
| UpstreamError | 502 | Target site is unreachable |
| ApiError | 4xx/5xx | Other server errors |
See Error Reference for full details.
Retry policy
All SDKs implement exponential backoff with a 1-hour cap:
- Retryable:
429,500,502,503,504 - Backoff:
base_delay × 2^attempt, capped at 60s 429responses: respectsRetry-Afterheader andretry_afterbody field- Default: 3 retries
