Skip to main content

Java SDK

kloakd-sdk for Java 21+ — one API for everything, full module coverage with async support

4 min read


Install

<dependency>
  <groupId>dev.kloakd</groupId>
  <artifactId>kloakd-sdk</artifactId>
  <version>0.1.0</version>
</dependency>

Java 21+ | 67/67 tests | 88.7% line | 70.7% branch | 96.2% method | stdlib HttpClient

Client

Kloakd client = Kloakd.builder()
    .apiKey(System.getenv("KLOAKD_API_KEY"))
    .organizationId(System.getenv("KLOAKD_ORG_ID"))
    .timeout(Duration.ofSeconds(30))
    .maxRetries(3)
    .build();

Quickstart — one call does everything

SiteCrawlResult result = client.crawl("https://example.com",
    CrawlOptions.builder()
        .maxPages(50)
        .extractSchema(Map.of("title", "css:h1", "content", "css:article"))
        .build());

for (CrawlPage page : result.pages()) {
    if (page.success()) {
        System.out.println(page.url() + " → " + page.structuredData());
    }
}

That's it. crawl() internally:

  1. Discovers all pages on the site (Webgrph BFS)
  2. Fetches each page through the 4-tier anti-bot engine (auto-escalates to headless browser for Cloudflare-protected sites — no forceBrowser needed)
  3. Extracts structured data from each page if extractSchema is provided (Kolektr, using cached fetch artifacts — no second HTTP round-trip)

Per-page failures are caught and marked success: false — the crawl never aborts on a single page error.

Primary method

client.crawl()

SiteCrawlResult result = client.crawl("https://example.com",
    CrawlOptions.builder()
        .maxDepth(3)
        .maxPages(100)
        .extractSchema(Map.of("title", "css:h1", "price", "css:.price"))
        .build());

System.out.printf("Discovered: %d pages%n", result.totalPagesDiscovered());
System.out.printf("Fetched:    %d%n", result.pagesFetched());
System.out.printf("Failed:     %d%n", result.pagesFailed());

for (CrawlPage page : result.pages()) {
    if (page.success()) {
        System.out.printf("  %s [tier %d] %s%n", page.url(), page.tierUsed(), page.structuredData());
    } else {
        System.out.printf("  %s FAILED: %s%n", page.url(), page.error());
    }
}

Parameters

| Parameter | Type | Default | Description | |---|---|---|---| | url | String | required | Seed URL to start crawling from | | maxDepth | int | 3 | Maximum BFS depth | | maxPages | int | 100 | Maximum pages to crawl | | extractSchema | Map<String, String> | null | CSS selector schema for structured extraction | | includeExternalLinks | boolean | false | Follow off-domain links | | sessionArtifactId | String | null | Reuse an authenticated session artifact (from Fetchyr) |

Returns: SiteCrawlResult

| Field | Type | Description | |---|---|---| | success | boolean | Whether the crawl completed | | url | String | Seed URL | | totalPagesDiscovered | int | Pages found during BFS | | pagesFetched | int | Pages successfully fetched | | pagesFailed | int | Pages that failed (included with success: false) | | pages | List<CrawlPage> | All pages with HTML and optional structured data | | crawlArtifactId | String | Artifact ID for the site hierarchy | | error | String | Error message if crawl failed entirely |

client.crawlStream() — async streaming

For long-running crawls, use the async streaming version to receive real-time progress events:

client.crawlStream("https://example.com",
    CrawlOptions.builder()
        .maxPages(500)
        .extractSchema(Map.of("title", "css:h1"))
        .build())
    .forEach(event -> {
        switch (event.type()) {
            case "page_fetched" -> System.out.printf("[%d/%d] %s OK%n", event.page(), event.total(), event.url());
            case "page_failed" -> System.out.printf("[%d/%d] %s FAIL: %s%n", event.page(), event.total(), event.url(), event.error());
            case "crawl_complete" -> System.out.println("Done");
        }
    });

Event types

| Type | Description | |---|---| | discovery_started | Crawl discovery has begun | | discovery_progress | Pages found during BFS (pagesFound field) | | discovery_complete | Discovery finished, fetch phase starting | | page_fetching | About to fetch page N of total | | page_fetched | Page fetched successfully | | page_failed | Page fetch failed (crawl continues) | | crawl_complete | All pages processed |

Low-level methods

Need fine-grained control? The individual modules are still available:

evadr().fetch()

FetchResult page = client.evadr().fetch("https://example.com");
System.out.printf("Status: %d, Tier: %d%n", page.statusCode(), page.tierUsed());
System.out.printf("Artifact ID: %s%n", page.artifactId());

webgrph().crawl()

CrawlResult crawl = client.webgrph().crawl("https://example.com",
    CrawlOptions.builder().maxDepth(3).maxPages(100).build());
System.out.printf("Crawl started: %s%n", crawl.crawlId());

kolektr().page()

ExtractionResult result = client.kolektr().page("https://example.com",
    PageOptions.builder()
        .schema(Map.of("title", "css:h1", "price", "css:.price"))
        .build());
for (var record : result.records()) {
    System.out.println(record);
}

Artifact chaining

Pass artifact IDs between low-level methods to skip redundant work:

FetchResult page = client.evadr().fetch("https://example.com");

ExtractionResult data = client.kolektr().page("https://example.com",
    PageOptions.builder()
        .schema(Map.of("title", "css:h1"))
        .fetchArtifactId(page.artifactId())
        .build());

CrawlResult crawl = client.webgrph().crawl("https://example.com",
    CrawlOptions.builder()
        .maxDepth(2)
        .sessionArtifactId(page.artifactId())
        .build());

Error handling

try {
    SiteCrawlResult result = client.crawl("https://example.com", CrawlOptions.builder().build());
} catch (RateLimitException e) {
    System.out.printf("Retry after %ds%n", e.getRetryAfter());
} catch (AuthenticationException e) {
    System.out.println("Invalid API key");
} catch (KloakdException e) {
    System.out.println(e.getMessage());
}

Advanced namespaces

client.skanyr()    // API discovery
client.nexus()     // AI strategy engine
client.parlyr()    // Natural language queries
client.fetchyr()   // RPA & authenticated scraping

Next steps

Was this page helpful?