Webgrph
Site structure mapping, crawl, and hierarchy — the engine behind crawl()
2 min read
Webgrph is the site mapping module that powers the discovery phase of client.crawl(). It maps the full structure of a website using BFS traversal and artifact chaining.
The high-level client.crawl() orchestrator calls webgrph.crawl() internally — you don't need to poll or manage crawl IDs yourself. See crawl() for the one-call API.
For the low-level API reference, see webgrph.crawl().
crawl
import time
crawl = client.webgrph.crawl(
"https://example.com",
max_depth=3,
max_pages=500,
session_artifact_id=page.artifact_id,
)
print(f"Crawl started: {crawl.crawl_id}")
# Poll for completion
while True:
status = client.webgrph.get_crawl_status(crawl.crawl_id)
print(f" {status.get('completed_pages', 0)}/{status.get('total_pages', 0)} pages")
if status.get("status") == "completed":
break
time.sleep(2)
const crawl = await client.webgrph.crawl('https://example.com', {
maxDepth: 3,
sessionArtifactId: page.artifactId,
});
console.log(`Crawl started: ${crawl.crawlId}`);
crawl_all (auto-pagination)
for p in client.webgrph.crawl_all("https://example.com", max_depth=3):
print(p.url, p.depth)
crawlStream (SSE)
for await (const event of client.webgrph.crawlStream('https://example.com')) {
console.log(event.type, event.url, event.depth);
}
get_hierarchy
tree = client.webgrph.get_hierarchy("https://example.com")
print(tree.root.url, len(tree.root.children))
PageNode fields
| Field | Type | Description | |-------|------|-------------| | url | string | Page URL | | depth | integer | Depth from root | | status_code | integer | HTTP status | | title | string or null | Page title | | children | PageNode list | Child pages |
Next steps
- webgrph.crawl() reference — full API endpoint docs
- Batch extraction guide — crawl then extract pattern
- Pagination guide — handling large crawls
Was this page helpful?
