Skip to main content

Webgrph

Site structure mapping, crawl, and hierarchy — the engine behind crawl()

2 min read


Webgrph is the site mapping module that powers the discovery phase of client.crawl(). It maps the full structure of a website using BFS traversal and artifact chaining.

The high-level client.crawl() orchestrator calls webgrph.crawl() internally — you don't need to poll or manage crawl IDs yourself. See crawl() for the one-call API.

For the low-level API reference, see webgrph.crawl().

crawl

import time

crawl = client.webgrph.crawl(
    "https://example.com",
    max_depth=3,
    max_pages=500,
    session_artifact_id=page.artifact_id,
)
print(f"Crawl started: {crawl.crawl_id}")

# Poll for completion
while True:
    status = client.webgrph.get_crawl_status(crawl.crawl_id)
    print(f"  {status.get('completed_pages', 0)}/{status.get('total_pages', 0)} pages")
    if status.get("status") == "completed":
        break
    time.sleep(2)
const crawl = await client.webgrph.crawl('https://example.com', {
  maxDepth: 3,
  sessionArtifactId: page.artifactId,
});
console.log(`Crawl started: ${crawl.crawlId}`);

crawl_all (auto-pagination)

for p in client.webgrph.crawl_all("https://example.com", max_depth=3):
    print(p.url, p.depth)

crawlStream (SSE)

for await (const event of client.webgrph.crawlStream('https://example.com')) {
  console.log(event.type, event.url, event.depth);
}

get_hierarchy

tree = client.webgrph.get_hierarchy("https://example.com")
print(tree.root.url, len(tree.root.children))

PageNode fields

| Field | Type | Description | |-------|------|-------------| | url | string | Page URL | | depth | integer | Depth from root | | status_code | integer | HTTP status | | title | string or null | Page title | | children | PageNode list | Child pages |

Next steps

Was this page helpful?