Skip to main content

crawl()

Map and crawl an entire website

6 min read


New to KLOAKD? Use the high-level client.crawl() orchestrator instead of the raw webgrph.crawl() endpoint. It handles discovery, fetching, and optional extraction in a single call — no polling required. See the Quickstart for a 30-second example.

High-level orchestrator: client.crawl()

The crawl() orchestrator chains discovery (webgrph.crawl()), fetching (evadr.fetch()), and optional extraction (kolektr.page()) into a single call. It returns a SiteCrawlResult with every page fetched and ready to use.

Parameters

| Field | Type | Required | Default | Description | |-------|------|----------|---------|-------------| | url | string | Yes | — | Root URL to crawl | | max_depth | integer | No | 3 | Maximum crawl depth | | max_pages | integer | No | 100 | Maximum pages to crawl | | extract_schema | object | No | null | Schema for structured extraction on every page | | session_artifact_id | string | No | null | Authenticated session from Fetchyr | | include_external_links | boolean | No | false | Follow links to other domains |

Examples

from kloakd import Kloakd

client = Kloakd(api_key="sk-live-...", organization_id="your-org-id")

# Crawl + extract in one call
result = client.crawl(
    "https://example.com",
    max_depth=2,
    extract_schema={"title": "css:h1", "content": "css:article"},
)

print(f"Crawled {result.total_pages_discovered} pages")
for page in result.pages:
    print(f"  {page.url} — {page.status_code}")
    if page.structured_data:
        print(f"    Title: {page.structured_data.get('title')}")
import { Kloakd } from 'kloakd-sdk';

const client = new Kloakd({
  apiKey: 'sk-live-...',
  organizationId: 'your-org-id',
});

const result = await client.crawl('https://example.com', {
  maxDepth: 2,
  extractSchema: { title: 'css:h1', content: 'css:article' },
});

console.log(`Crawled ${result.totalPagesDiscovered} pages`);
for (const page of result.pages) {
  console.log(`  ${page.url} — ${page.statusCode}`);
}
result, _ := client.Crawl(ctx, "https://example.com",
    &kloakd.CrawlOrchestratorOptions{
        MaxDepth:      2,
        ExtractSchema: map[string]string{"title": "css:h1", "content": "css:article"},
    })
fmt.Printf("Crawled %d pages\n", result.TotalPagesDiscovered)
for _, page := range result.Pages {
    fmt.Printf("  %s — %d\n", page.URL, page.StatusCode)
}
SiteCrawlResult result = client.crawl("https://example.com",
    CrawlOptions.builder()
        .maxDepth(2)
        .extractSchema(Map.of("title", "css:h1", "content", "css:article"))
        .build());
System.out.printf("Crawled %d pages%n", result.totalPagesDiscovered());
for (CrawlPage page : result.pages()) {
    System.out.printf("  %s — %d%n", page.url(), page.statusCode());
}

Streaming: client.crawlStream()

Get real-time progress events as the crawl discovers and fetches pages:

for event in client.crawl_stream("https://example.com", max_depth=2):
    print(f"[{event.type}] {event.url or ''}")
for await (const event of client.crawlStream('https://example.com', { maxDepth: 2 })) {
  console.log(`[${event.type}] ${event.url ?? ''}`);
}

CrawlProgressEvent types

| Event type | Description | |------------|-------------| | discovery_start | Crawl discovery started | | discovery_complete | All pages discovered | | page_fetching | Started fetching a page | | page_fetched | Page fetched successfully | | page_failed | Page fetch failed | | crawl_complete | Crawl finished, results ready |


Low-level API: webgrph.crawl()

The raw endpoint for direct API calls or when you need fine-grained control over polling and pagination.

Endpoint

POST /api/v1/organizations/{org_id}/webgrph/crawl

Description

Crawl a website using BFS traversal, building a complete page hierarchy. Anti-bot bypass is built in — crawl uses the same 4-tier escalation as fetch() for each page.

Request body

| Field | Type | Required | Default | Description | |-------|------|----------|---------|-------------| | url | string | Yes | — | Root URL to crawl | | max_depth | integer | No | 3 | Maximum crawl depth | | max_pages | integer | No | 100 | Maximum pages to crawl | | include_external_links | boolean | No | false | Follow links to other domains | | session_artifact_id | string | No | null | Authenticated session from Fetchyr | | limit | integer | No | 100 | Max results per page | | offset | integer | No | 0 | Pagination offset |

Response

Returns 202 Accepted with a crawl_id — the crawl runs asynchronously. Poll GET /api/v1/organizations/{org_id}/webgrph/crawl/{crawl_id}/status for progress:

{
  "crawl_id": "crawl-abc123",
  "status": "pending",
  "completed_pages": 0,
  "total_pages": 0,
  "artifact_id": "art-xyz789"
}

When status is "completed", retrieve pages via GET /api/v1/organizations/{org_id}/webgrph/crawl/{crawl_id}/pages:

{
  "pages": [
    {
      "url": "https://example.com",
      "depth": 0,
      "status_code": 200,
      "title": "Example Site"
    }
  ],
  "total": 42
}

Page object

| Field | Type | Description | |-------|------|-------------| | url | string | Page URL | | depth | integer | Depth from root (0 = root) | | status_code | integer | HTTP status code | | title | string or null | Page title | | children | array | Child page objects |

Examples

Python

import time
from kloakd import Kloakd

client = Kloakd(api_key="sk-live-...", organization_id="your-org-id")

# Start crawl (returns 202 immediately)
crawl = client.webgrph.crawl(
    "https://example.com",
    max_depth=3,
    max_pages=100,
)
print(f"Crawl started: {crawl.crawl_id}")

# Poll for completion
while True:
    status = client.webgrph.get_crawl_status(crawl.crawl_id)
    print(f"  {status.get('completed_pages', 0)}/{status.get('total_pages', 0)} pages")
    if status.get("status") == "completed":
        break
    time.sleep(2)

# Retrieve discovered pages
pages = client.webgrph.get_crawl_pages(crawl.crawl_id)
for page in pages.get("pages", []):
    print(f"  {page['url']} (depth {page['depth']})")

TypeScript

import { Kloakd } from 'kloakd-sdk';

const client = new Kloakd({
  apiKey: 'sk-live-...',
  organizationId: 'your-org-id',
});

const crawl = await client.webgrph.crawl('https://example.com', {
  maxDepth: 3,
  maxPages: 100,
});
console.log(`Crawl started: ${crawl.crawlId}`);

// Poll for completion
while (true) {
  const status = await client.webgrph.getCrawlStatus(crawl.crawlId);
  console.log(`  ${status.completedPages}/${status.totalPages} pages`);
  if (status.status === 'completed') break;
  await new Promise(r => setTimeout(r, 2000));
}

Go

crawl, _ := client.Webgrph.Crawl(ctx, "https://example.com",
    &kloakd.CrawlOptions{
        MaxDepth: 3,
        MaxPages: 100,
    })
fmt.Printf("Crawl started: %s\n", crawl.CrawlID)

// Poll for completion
for {
    status, _ := client.Webgrph.GetCrawlStatus(ctx, crawl.CrawlID)
    fmt.Printf("  %d/%d pages\n", status.CompletedPages, status.TotalPages)
    if status.Status == "completed" {
        break
    }
    time.Sleep(2 * time.Second)
}

cURL

curl -X POST https://api.kloakd.dev/api/v1/organizations/$ORG_ID/webgrph/crawl \
  -H "Authorization: Bearer sk-live-..." \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com", "max_depth": 3, "max_pages": 100}'

Artifact chaining

Chain webgrph.crawl() with kolektr.page() to get structured data from every page:

import time

crawl = client.webgrph.crawl("https://example.com", max_depth=2)

# Wait for crawl to complete
while True:
    status = client.webgrph.get_crawl_status(crawl.crawl_id)
    if status.get("status") == "completed":
        break
    time.sleep(2)

pages = client.webgrph.get_crawl_pages(crawl.crawl_id)
for page in pages.get("pages", []):
    data = client.kolektr.page(
        page["url"],
        schema={"title": "css:h1", "content": "css:article"},
        fetch_artifact_id=crawl.artifact_id,
    )
    print(f"{page['url']}: {len(data.records)} records")

Errors

| Status | Error | When | |--------|-------|------| | 400 | ValidationError | Invalid URL | | 401 | AuthenticationError | Missing or invalid API key | | 403 | ForbiddenError | Org ID mismatch | | 429 | RateLimitError | Quota exceeded | | 502 | UpstreamError | Target site unreachable |

SSE streaming

For real-time updates instead of polling, use the async client's crawl_stream():

import asyncio
from kloakd import AsyncKloakd

async def main():
    client = AsyncKloakd(api_key="sk-live-...", organization_id="your-org-id")
    crawl = await client.webgrph.crawl("https://example.com", max_depth=2)

    async for event in client.webgrph.crawl_stream(crawl.crawl_id):
        print(f"  [{event.type}] {event.url}")

asyncio.run(main())

Learn more

Was this page helpful?