crawl()
Map and crawl an entire website
6 min read
New to KLOAKD? Use the high-level client.crawl() orchestrator instead of the raw webgrph.crawl() endpoint. It handles discovery, fetching, and optional extraction in a single call — no polling required. See the Quickstart for a 30-second example.
High-level orchestrator: client.crawl()
The crawl() orchestrator chains discovery (webgrph.crawl()), fetching (evadr.fetch()), and optional extraction (kolektr.page()) into a single call. It returns a SiteCrawlResult with every page fetched and ready to use.
Parameters
| Field | Type | Required | Default | Description |
|-------|------|----------|---------|-------------|
| url | string | Yes | — | Root URL to crawl |
| max_depth | integer | No | 3 | Maximum crawl depth |
| max_pages | integer | No | 100 | Maximum pages to crawl |
| extract_schema | object | No | null | Schema for structured extraction on every page |
| session_artifact_id | string | No | null | Authenticated session from Fetchyr |
| include_external_links | boolean | No | false | Follow links to other domains |
Examples
from kloakd import Kloakd
client = Kloakd(api_key="sk-live-...", organization_id="your-org-id")
# Crawl + extract in one call
result = client.crawl(
"https://example.com",
max_depth=2,
extract_schema={"title": "css:h1", "content": "css:article"},
)
print(f"Crawled {result.total_pages_discovered} pages")
for page in result.pages:
print(f" {page.url} — {page.status_code}")
if page.structured_data:
print(f" Title: {page.structured_data.get('title')}")
import { Kloakd } from 'kloakd-sdk';
const client = new Kloakd({
apiKey: 'sk-live-...',
organizationId: 'your-org-id',
});
const result = await client.crawl('https://example.com', {
maxDepth: 2,
extractSchema: { title: 'css:h1', content: 'css:article' },
});
console.log(`Crawled ${result.totalPagesDiscovered} pages`);
for (const page of result.pages) {
console.log(` ${page.url} — ${page.statusCode}`);
}
result, _ := client.Crawl(ctx, "https://example.com",
&kloakd.CrawlOrchestratorOptions{
MaxDepth: 2,
ExtractSchema: map[string]string{"title": "css:h1", "content": "css:article"},
})
fmt.Printf("Crawled %d pages\n", result.TotalPagesDiscovered)
for _, page := range result.Pages {
fmt.Printf(" %s — %d\n", page.URL, page.StatusCode)
}
SiteCrawlResult result = client.crawl("https://example.com",
CrawlOptions.builder()
.maxDepth(2)
.extractSchema(Map.of("title", "css:h1", "content", "css:article"))
.build());
System.out.printf("Crawled %d pages%n", result.totalPagesDiscovered());
for (CrawlPage page : result.pages()) {
System.out.printf(" %s — %d%n", page.url(), page.statusCode());
}
Streaming: client.crawlStream()
Get real-time progress events as the crawl discovers and fetches pages:
for event in client.crawl_stream("https://example.com", max_depth=2):
print(f"[{event.type}] {event.url or ''}")
for await (const event of client.crawlStream('https://example.com', { maxDepth: 2 })) {
console.log(`[${event.type}] ${event.url ?? ''}`);
}
CrawlProgressEvent types
| Event type | Description |
|------------|-------------|
| discovery_start | Crawl discovery started |
| discovery_complete | All pages discovered |
| page_fetching | Started fetching a page |
| page_fetched | Page fetched successfully |
| page_failed | Page fetch failed |
| crawl_complete | Crawl finished, results ready |
Low-level API: webgrph.crawl()
The raw endpoint for direct API calls or when you need fine-grained control over polling and pagination.
Endpoint
POST /api/v1/organizations/{org_id}/webgrph/crawl
Description
Crawl a website using BFS traversal, building a complete page hierarchy. Anti-bot bypass is built in — crawl uses the same 4-tier escalation as fetch() for each page.
Request body
| Field | Type | Required | Default | Description |
|-------|------|----------|---------|-------------|
| url | string | Yes | — | Root URL to crawl |
| max_depth | integer | No | 3 | Maximum crawl depth |
| max_pages | integer | No | 100 | Maximum pages to crawl |
| include_external_links | boolean | No | false | Follow links to other domains |
| session_artifact_id | string | No | null | Authenticated session from Fetchyr |
| limit | integer | No | 100 | Max results per page |
| offset | integer | No | 0 | Pagination offset |
Response
Returns 202 Accepted with a crawl_id — the crawl runs asynchronously. Poll GET /api/v1/organizations/{org_id}/webgrph/crawl/{crawl_id}/status for progress:
{
"crawl_id": "crawl-abc123",
"status": "pending",
"completed_pages": 0,
"total_pages": 0,
"artifact_id": "art-xyz789"
}
When status is "completed", retrieve pages via GET /api/v1/organizations/{org_id}/webgrph/crawl/{crawl_id}/pages:
{
"pages": [
{
"url": "https://example.com",
"depth": 0,
"status_code": 200,
"title": "Example Site"
}
],
"total": 42
}
Page object
| Field | Type | Description |
|-------|------|-------------|
| url | string | Page URL |
| depth | integer | Depth from root (0 = root) |
| status_code | integer | HTTP status code |
| title | string or null | Page title |
| children | array | Child page objects |
Examples
Python
import time
from kloakd import Kloakd
client = Kloakd(api_key="sk-live-...", organization_id="your-org-id")
# Start crawl (returns 202 immediately)
crawl = client.webgrph.crawl(
"https://example.com",
max_depth=3,
max_pages=100,
)
print(f"Crawl started: {crawl.crawl_id}")
# Poll for completion
while True:
status = client.webgrph.get_crawl_status(crawl.crawl_id)
print(f" {status.get('completed_pages', 0)}/{status.get('total_pages', 0)} pages")
if status.get("status") == "completed":
break
time.sleep(2)
# Retrieve discovered pages
pages = client.webgrph.get_crawl_pages(crawl.crawl_id)
for page in pages.get("pages", []):
print(f" {page['url']} (depth {page['depth']})")
TypeScript
import { Kloakd } from 'kloakd-sdk';
const client = new Kloakd({
apiKey: 'sk-live-...',
organizationId: 'your-org-id',
});
const crawl = await client.webgrph.crawl('https://example.com', {
maxDepth: 3,
maxPages: 100,
});
console.log(`Crawl started: ${crawl.crawlId}`);
// Poll for completion
while (true) {
const status = await client.webgrph.getCrawlStatus(crawl.crawlId);
console.log(` ${status.completedPages}/${status.totalPages} pages`);
if (status.status === 'completed') break;
await new Promise(r => setTimeout(r, 2000));
}
Go
crawl, _ := client.Webgrph.Crawl(ctx, "https://example.com",
&kloakd.CrawlOptions{
MaxDepth: 3,
MaxPages: 100,
})
fmt.Printf("Crawl started: %s\n", crawl.CrawlID)
// Poll for completion
for {
status, _ := client.Webgrph.GetCrawlStatus(ctx, crawl.CrawlID)
fmt.Printf(" %d/%d pages\n", status.CompletedPages, status.TotalPages)
if status.Status == "completed" {
break
}
time.Sleep(2 * time.Second)
}
cURL
curl -X POST https://api.kloakd.dev/api/v1/organizations/$ORG_ID/webgrph/crawl \
-H "Authorization: Bearer sk-live-..." \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com", "max_depth": 3, "max_pages": 100}'
Artifact chaining
Chain webgrph.crawl() with kolektr.page() to get structured data from every page:
import time
crawl = client.webgrph.crawl("https://example.com", max_depth=2)
# Wait for crawl to complete
while True:
status = client.webgrph.get_crawl_status(crawl.crawl_id)
if status.get("status") == "completed":
break
time.sleep(2)
pages = client.webgrph.get_crawl_pages(crawl.crawl_id)
for page in pages.get("pages", []):
data = client.kolektr.page(
page["url"],
schema={"title": "css:h1", "content": "css:article"},
fetch_artifact_id=crawl.artifact_id,
)
print(f"{page['url']}: {len(data.records)} records")
Errors
| Status | Error | When |
|--------|-------|------|
| 400 | ValidationError | Invalid URL |
| 401 | AuthenticationError | Missing or invalid API key |
| 403 | ForbiddenError | Org ID mismatch |
| 429 | RateLimitError | Quota exceeded |
| 502 | UpstreamError | Target site unreachable |
SSE streaming
For real-time updates instead of polling, use the async client's crawl_stream():
import asyncio
from kloakd import AsyncKloakd
async def main():
client = AsyncKloakd(api_key="sk-live-...", organization_id="your-org-id")
crawl = await client.webgrph.crawl("https://example.com", max_depth=2)
async for event in client.webgrph.crawl_stream(crawl.crawl_id):
print(f" [{event.type}] {event.url}")
asyncio.run(main())
Learn more
- Quickstart —
client.crawl()in 30 seconds - Webgrph module — hierarchy, reader view, SSE streaming
- Batch extraction guide — crawl then extract pattern
- Pagination guide — handling large crawls
