TypeScript SDK
kloakd-sdk for Node.js and browser — one API for everything, full module coverage
4 min read
Install
npm install kloakd-sdk
Node.js 18+ | 76/76 tests | 92.76% statement coverage | native fetch
Client
import { Kloakd } from 'kloakd-sdk';
const client = new Kloakd({
apiKey: process.env.KLOAKD_API_KEY!,
organizationId: process.env.KLOAKD_ORG_ID!,
timeout: 30_000,
maxRetries: 3,
});
Quickstart — one call does everything
const result = await client.crawl('https://example.com', {
maxPages: 50,
extractSchema: { title: 'css:h1', content: 'css:article' },
});
for (const page of result.pages) {
if (page.ok) {
console.log(`${page.url} →`, page.structuredData);
}
}
That's it. crawl() internally:
- Discovers all pages on the site (Webgrph BFS)
- Fetches each page through the 4-tier anti-bot engine (auto-escalates to headless browser for Cloudflare-protected sites — no
forceBrowserneeded) - Extracts structured data from each page if
extractSchemais provided (Kolektr, using cached fetch artifacts — no second HTTP round-trip)
Per-page failures are caught and marked success: false — the crawl never aborts on a single page error.
Primary method
client.crawl()
const result = await client.crawl('https://example.com', {
maxDepth: 3,
maxPages: 100,
extractSchema: { title: 'css:h1', price: 'css:.price' },
includeExternalLinks: false,
sessionArtifactId: undefined, // reuse a Fetchyr session for login-protected sites
});
console.log(`Discovered: ${result.totalPagesDiscovered} pages`);
console.log(`Fetched: ${result.pagesFetched}`);
console.log(`Failed: ${result.pagesFailed}`);
for (const page of result.pages) {
if (page.ok) {
console.log(` ${page.url} [tier ${page.tierUsed}]`, page.structuredData);
} else {
console.log(` ${page.url} FAILED: ${page.error}`);
}
}
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
| url | string | required | Seed URL to start crawling from |
| maxDepth | number | 3 | Maximum BFS depth |
| maxPages | number | 100 | Maximum pages to crawl |
| extractSchema | Record<string, string> | undefined | CSS selector schema for structured extraction |
| includeExternalLinks | boolean | false | Follow off-domain links |
| sessionArtifactId | string | undefined | Reuse an authenticated session artifact (from Fetchyr) |
Returns: SiteCrawlResult
| Field | Type | Description |
|---|---|---|
| success | boolean | Whether the crawl completed |
| url | string | Seed URL |
| totalPagesDiscovered | number | Pages found during BFS |
| pagesFetched | number | Pages successfully fetched |
| pagesFailed | number | Pages that failed (included with success: false) |
| pages | CrawlPage[] | All pages with HTML and optional structured data |
| crawlArtifactId | string \| null | Artifact ID for the site hierarchy |
| error | string \| null | Error message if crawl failed entirely |
client.crawlStream() — async streaming
For long-running crawls, use the async streaming version to receive real-time progress events:
for await (const event of client.crawlStream('https://example.com', {
maxPages: 500,
extractSchema: { title: 'css:h1' },
})) {
if (event.type === 'page_fetched') {
console.log(`[${event.page}/${event.total}] ${event.url} OK`);
} else if (event.type === 'page_failed') {
console.log(`[${event.page}/${event.total}] ${event.url} FAIL: ${event.error}`);
} else if (event.type === 'crawl_complete') {
const result = event.metadata.result;
console.log(`Done: ${result.pagesFetched} fetched, ${result.pagesFailed} failed`);
}
}
Event types
| Type | Description |
|---|---|
| discovery_started | Crawl discovery has begun |
| discovery_progress | Pages found during BFS (pagesFound field) |
| discovery_complete | Discovery finished, fetch phase starting |
| page_fetching | About to fetch page N of total |
| page_fetched | Page fetched successfully |
| page_failed | Page fetch failed (crawl continues) |
| crawl_complete | All pages processed, final summary in metadata.result |
Low-level methods
Need fine-grained control? The individual modules are still available:
evadr.fetch()
const page = await client.evadr.fetch('https://example.com');
console.log(`Status: ${page.statusCode}, Tier: ${page.tierUsed}`);
console.log(`Artifact ID: ${page.artifactId}`);
webgrph.crawl()
const crawl = await client.webgrph.crawl('https://example.com', {
maxDepth: 3,
maxPages: 100,
});
console.log(`Crawl started: ${crawl.crawlId}`);
kolektr.page()
const result = await client.kolektr.page(
'https://example.com',
{ schema: { title: 'css:h1', price: 'css:.price' } }
);
console.log(result.records);
Artifact chaining
Pass artifact IDs between low-level methods to skip redundant work:
const page = await client.evadr.fetch('https://example.com');
const data = await client.kolektr.page('https://example.com', {
schema: { title: 'css:h1' },
fetchArtifactId: page.artifactId ?? undefined,
});
const crawl = await client.webgrph.crawl('https://example.com', {
maxDepth: 2,
sessionArtifactId: page.artifactId ?? undefined,
});
Error handling
import { AuthenticationError, RateLimitError, KloakdError } from 'kloakd-sdk';
try {
const result = await client.crawl('https://example.com');
} catch (e) {
if (e instanceof AuthenticationError) console.log('Invalid API key');
else if (e instanceof RateLimitError) console.log(`Retry after ${e.retryAfter}s`);
else if (e instanceof KloakdError) console.log(e.message);
}
Advanced namespaces
client.skanyr // API discovery
client.nexus // AI strategy engine
client.parlyr // Natural language queries
client.fetchyr // RPA & authenticated scraping
Next steps
- Quickstart — get started in 5 minutes
- API reference — detailed endpoint docs
- Error reference — full error taxonomy
