Choosing an Approach
When to use crawl, fetch, or extract
3 min read
KLOAKD provides one orchestrator and three low-level methods. Here's how to choose.
Decision tree
- Need data from an entire site? →
client.crawl()(one call, handles everything) - Need data from a single page? →
client.crawl(url, max_depth=0)orkolektr.page() - Need raw HTML from a single page? →
evadr.fetch() - Need to map a site without fetching? →
webgrph.crawl()(low-level, async) - Need to bypass anti-bot? → All methods handle this automatically
- Need authenticated access? → Pass
session_artifact_idtoclient.crawl()
Comparison
| | client.crawl() | evadr.fetch() | webgrph.crawl() | kolektr.page() |
|--|-----------|-----------|-------------|------------|
| Returns | Pages + optional extracted data | Raw HTML | Page hierarchy (async) | Structured records |
| Pages | Many (BFS) | 1 | Many (BFS) | 1 |
| Anti-bot | Per-page 4-tier escalation | 4-tier escalation | Per-page escalation | Reuses fetch artifact |
| Extraction | Optional via extract_schema | No | No | Yes (schema) |
| Polling | No (synchronous) | No (synchronous) | Yes (async) | No (synchronous) |
| Streaming | crawlStream() | fetchStream() | crawlStream() | No |
| Use case | Discover + fetch + extract a site | Get HTML from one page | Map site, find pages | Get structured data from one page |
Common patterns
Pattern 1: Crawl and extract an entire site (recommended)
result = client.crawl(
"https://example.com",
max_depth=2,
extract_schema={"title": "css:h1", "content": "css:article"},
)
for page in result.pages:
print(f"{page.url}: {page.structured_data}")
Pattern 2: Crawl without extraction
result = client.crawl("https://example.com", max_depth=3)
for page in result.pages:
print(f" {page.url} (depth {page.depth}) — {page.status_code}")
Pattern 3: Fetch a single page
page = client.evadr.fetch("https://protected-site.com")
print(page.html, page.artifact_id)
Pattern 4: Extract from a single page
result = client.kolektr.page(
"https://example.com/products",
schema={"name": "css:.product-name", "price": "css:.product-price"},
)
Pattern 5: Authenticated crawl
session = client.fetchyr.login(
url="https://example.com/login",
username_selector="#email",
password_selector="#password",
username="user@example.com",
password="secret",
)
result = client.crawl(
"https://example.com/dashboard",
max_depth=2,
session_artifact_id=session.artifact_id,
extract_schema={"title": "css:h1"},
)
Pattern 6: Fetch then extract (manual chaining)
page = client.evadr.fetch("https://protected-site.com")
data = client.kolektr.page(
"https://protected-site.com",
schema={"title": "css:h1"},
fetch_artifact_id=page.artifact_id,
)
When to use advanced modules
- Skanyr — when you suspect the site has hidden APIs (faster than scraping HTML)
- Nexus — when you don't know which approach to use (AI decides for you)
- Parlyr — when you want to describe what you want in natural language
- Fetchyr — when the page requires login, MFA, or form automation
See Advanced Modules for details.
