Skip to main content

Choosing an Approach

When to use crawl, fetch, or extract

3 min read


KLOAKD provides one orchestrator and three low-level methods. Here's how to choose.

Decision tree

  1. Need data from an entire site? → client.crawl() (one call, handles everything)
  2. Need data from a single page? → client.crawl(url, max_depth=0) or kolektr.page()
  3. Need raw HTML from a single page? → evadr.fetch()
  4. Need to map a site without fetching? → webgrph.crawl() (low-level, async)
  5. Need to bypass anti-bot? → All methods handle this automatically
  6. Need authenticated access? → Pass session_artifact_id to client.crawl()

Comparison

| | client.crawl() | evadr.fetch() | webgrph.crawl() | kolektr.page() | |--|-----------|-----------|-------------|------------| | Returns | Pages + optional extracted data | Raw HTML | Page hierarchy (async) | Structured records | | Pages | Many (BFS) | 1 | Many (BFS) | 1 | | Anti-bot | Per-page 4-tier escalation | 4-tier escalation | Per-page escalation | Reuses fetch artifact | | Extraction | Optional via extract_schema | No | No | Yes (schema) | | Polling | No (synchronous) | No (synchronous) | Yes (async) | No (synchronous) | | Streaming | crawlStream() | fetchStream() | crawlStream() | No | | Use case | Discover + fetch + extract a site | Get HTML from one page | Map site, find pages | Get structured data from one page |

Common patterns

result = client.crawl(
    "https://example.com",
    max_depth=2,
    extract_schema={"title": "css:h1", "content": "css:article"},
)
for page in result.pages:
    print(f"{page.url}: {page.structured_data}")

Pattern 2: Crawl without extraction

result = client.crawl("https://example.com", max_depth=3)
for page in result.pages:
    print(f"  {page.url} (depth {page.depth}) — {page.status_code}")

Pattern 3: Fetch a single page

page = client.evadr.fetch("https://protected-site.com")
print(page.html, page.artifact_id)

Pattern 4: Extract from a single page

result = client.kolektr.page(
    "https://example.com/products",
    schema={"name": "css:.product-name", "price": "css:.product-price"},
)

Pattern 5: Authenticated crawl

session = client.fetchyr.login(
    url="https://example.com/login",
    username_selector="#email",
    password_selector="#password",
    username="user@example.com",
    password="secret",
)

result = client.crawl(
    "https://example.com/dashboard",
    max_depth=2,
    session_artifact_id=session.artifact_id,
    extract_schema={"title": "css:h1"},
)

Pattern 6: Fetch then extract (manual chaining)

page = client.evadr.fetch("https://protected-site.com")
data = client.kolektr.page(
    "https://protected-site.com",
    schema={"title": "css:h1"},
    fetch_artifact_id=page.artifact_id,
)

When to use advanced modules

  • Skanyr — when you suspect the site has hidden APIs (faster than scraping HTML)
  • Nexus — when you don't know which approach to use (AI decides for you)
  • Parlyr — when you want to describe what you want in natural language
  • Fetchyr — when the page requires login, MFA, or form automation

See Advanced Modules for details.

Was this page helpful?