Skip to main content

TypeScript SDK

kloakd-sdk for Node.js and browser — one API for everything, full module coverage

4 min read


Install

npm install kloakd-sdk

Node.js 18+ | 76/76 tests | 92.76% statement coverage | native fetch

Client

import { Kloakd } from 'kloakd-sdk';

const client = new Kloakd({
  apiKey: process.env.KLOAKD_API_KEY!,
  organizationId: process.env.KLOAKD_ORG_ID!,
  timeout: 30_000,
  maxRetries: 3,
});

Quickstart — one call does everything

const result = await client.crawl('https://example.com', {
  maxPages: 50,
  extractSchema: { title: 'css:h1', content: 'css:article' },
});

for (const page of result.pages) {
  if (page.ok) {
    console.log(`${page.url} →`, page.structuredData);
  }
}

That's it. crawl() internally:

  1. Discovers all pages on the site (Webgrph BFS)
  2. Fetches each page through the 4-tier anti-bot engine (auto-escalates to headless browser for Cloudflare-protected sites — no forceBrowser needed)
  3. Extracts structured data from each page if extractSchema is provided (Kolektr, using cached fetch artifacts — no second HTTP round-trip)

Per-page failures are caught and marked success: false — the crawl never aborts on a single page error.

Primary method

client.crawl()

const result = await client.crawl('https://example.com', {
  maxDepth: 3,
  maxPages: 100,
  extractSchema: { title: 'css:h1', price: 'css:.price' },
  includeExternalLinks: false,
  sessionArtifactId: undefined,  // reuse a Fetchyr session for login-protected sites
});

console.log(`Discovered: ${result.totalPagesDiscovered} pages`);
console.log(`Fetched:    ${result.pagesFetched}`);
console.log(`Failed:     ${result.pagesFailed}`);

for (const page of result.pages) {
  if (page.ok) {
    console.log(`  ${page.url} [tier ${page.tierUsed}]`, page.structuredData);
  } else {
    console.log(`  ${page.url} FAILED: ${page.error}`);
  }
}

Parameters

| Parameter | Type | Default | Description | |---|---|---|---| | url | string | required | Seed URL to start crawling from | | maxDepth | number | 3 | Maximum BFS depth | | maxPages | number | 100 | Maximum pages to crawl | | extractSchema | Record<string, string> | undefined | CSS selector schema for structured extraction | | includeExternalLinks | boolean | false | Follow off-domain links | | sessionArtifactId | string | undefined | Reuse an authenticated session artifact (from Fetchyr) |

Returns: SiteCrawlResult

| Field | Type | Description | |---|---|---| | success | boolean | Whether the crawl completed | | url | string | Seed URL | | totalPagesDiscovered | number | Pages found during BFS | | pagesFetched | number | Pages successfully fetched | | pagesFailed | number | Pages that failed (included with success: false) | | pages | CrawlPage[] | All pages with HTML and optional structured data | | crawlArtifactId | string \| null | Artifact ID for the site hierarchy | | error | string \| null | Error message if crawl failed entirely |

client.crawlStream() — async streaming

For long-running crawls, use the async streaming version to receive real-time progress events:

for await (const event of client.crawlStream('https://example.com', {
  maxPages: 500,
  extractSchema: { title: 'css:h1' },
})) {
  if (event.type === 'page_fetched') {
    console.log(`[${event.page}/${event.total}] ${event.url} OK`);
  } else if (event.type === 'page_failed') {
    console.log(`[${event.page}/${event.total}] ${event.url} FAIL: ${event.error}`);
  } else if (event.type === 'crawl_complete') {
    const result = event.metadata.result;
    console.log(`Done: ${result.pagesFetched} fetched, ${result.pagesFailed} failed`);
  }
}

Event types

| Type | Description | |---|---| | discovery_started | Crawl discovery has begun | | discovery_progress | Pages found during BFS (pagesFound field) | | discovery_complete | Discovery finished, fetch phase starting | | page_fetching | About to fetch page N of total | | page_fetched | Page fetched successfully | | page_failed | Page fetch failed (crawl continues) | | crawl_complete | All pages processed, final summary in metadata.result |

Low-level methods

Need fine-grained control? The individual modules are still available:

evadr.fetch()

const page = await client.evadr.fetch('https://example.com');
console.log(`Status: ${page.statusCode}, Tier: ${page.tierUsed}`);
console.log(`Artifact ID: ${page.artifactId}`);

webgrph.crawl()

const crawl = await client.webgrph.crawl('https://example.com', {
  maxDepth: 3,
  maxPages: 100,
});
console.log(`Crawl started: ${crawl.crawlId}`);

kolektr.page()

const result = await client.kolektr.page(
  'https://example.com',
  { schema: { title: 'css:h1', price: 'css:.price' } }
);
console.log(result.records);

Artifact chaining

Pass artifact IDs between low-level methods to skip redundant work:

const page = await client.evadr.fetch('https://example.com');

const data = await client.kolektr.page('https://example.com', {
  schema: { title: 'css:h1' },
  fetchArtifactId: page.artifactId ?? undefined,
});

const crawl = await client.webgrph.crawl('https://example.com', {
  maxDepth: 2,
  sessionArtifactId: page.artifactId ?? undefined,
});

Error handling

import { AuthenticationError, RateLimitError, KloakdError } from 'kloakd-sdk';

try {
  const result = await client.crawl('https://example.com');
} catch (e) {
  if (e instanceof AuthenticationError) console.log('Invalid API key');
  else if (e instanceof RateLimitError) console.log(`Retry after ${e.retryAfter}s`);
  else if (e instanceof KloakdError) console.log(e.message);
}

Advanced namespaces

client.skanyr     // API discovery
client.nexus      // AI strategy engine
client.parlyr     // Natural language queries
client.fetchyr    // RPA & authenticated scraping

Next steps

Was this page helpful?