Java SDK
kloakd-sdk for Java 21+ — one API for everything, full module coverage with async support
4 min read
Install
<dependency>
<groupId>dev.kloakd</groupId>
<artifactId>kloakd-sdk</artifactId>
<version>0.1.0</version>
</dependency>
Java 21+ | 67/67 tests | 88.7% line | 70.7% branch | 96.2% method | stdlib HttpClient
Client
Kloakd client = Kloakd.builder()
.apiKey(System.getenv("KLOAKD_API_KEY"))
.organizationId(System.getenv("KLOAKD_ORG_ID"))
.timeout(Duration.ofSeconds(30))
.maxRetries(3)
.build();
Quickstart — one call does everything
SiteCrawlResult result = client.crawl("https://example.com",
CrawlOptions.builder()
.maxPages(50)
.extractSchema(Map.of("title", "css:h1", "content", "css:article"))
.build());
for (CrawlPage page : result.pages()) {
if (page.success()) {
System.out.println(page.url() + " → " + page.structuredData());
}
}
That's it. crawl() internally:
- Discovers all pages on the site (Webgrph BFS)
- Fetches each page through the 4-tier anti-bot engine (auto-escalates to headless browser for Cloudflare-protected sites — no
forceBrowserneeded) - Extracts structured data from each page if
extractSchemais provided (Kolektr, using cached fetch artifacts — no second HTTP round-trip)
Per-page failures are caught and marked success: false — the crawl never aborts on a single page error.
Primary method
client.crawl()
SiteCrawlResult result = client.crawl("https://example.com",
CrawlOptions.builder()
.maxDepth(3)
.maxPages(100)
.extractSchema(Map.of("title", "css:h1", "price", "css:.price"))
.build());
System.out.printf("Discovered: %d pages%n", result.totalPagesDiscovered());
System.out.printf("Fetched: %d%n", result.pagesFetched());
System.out.printf("Failed: %d%n", result.pagesFailed());
for (CrawlPage page : result.pages()) {
if (page.success()) {
System.out.printf(" %s [tier %d] %s%n", page.url(), page.tierUsed(), page.structuredData());
} else {
System.out.printf(" %s FAILED: %s%n", page.url(), page.error());
}
}
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
| url | String | required | Seed URL to start crawling from |
| maxDepth | int | 3 | Maximum BFS depth |
| maxPages | int | 100 | Maximum pages to crawl |
| extractSchema | Map<String, String> | null | CSS selector schema for structured extraction |
| includeExternalLinks | boolean | false | Follow off-domain links |
| sessionArtifactId | String | null | Reuse an authenticated session artifact (from Fetchyr) |
Returns: SiteCrawlResult
| Field | Type | Description |
|---|---|---|
| success | boolean | Whether the crawl completed |
| url | String | Seed URL |
| totalPagesDiscovered | int | Pages found during BFS |
| pagesFetched | int | Pages successfully fetched |
| pagesFailed | int | Pages that failed (included with success: false) |
| pages | List<CrawlPage> | All pages with HTML and optional structured data |
| crawlArtifactId | String | Artifact ID for the site hierarchy |
| error | String | Error message if crawl failed entirely |
client.crawlStream() — async streaming
For long-running crawls, use the async streaming version to receive real-time progress events:
client.crawlStream("https://example.com",
CrawlOptions.builder()
.maxPages(500)
.extractSchema(Map.of("title", "css:h1"))
.build())
.forEach(event -> {
switch (event.type()) {
case "page_fetched" -> System.out.printf("[%d/%d] %s OK%n", event.page(), event.total(), event.url());
case "page_failed" -> System.out.printf("[%d/%d] %s FAIL: %s%n", event.page(), event.total(), event.url(), event.error());
case "crawl_complete" -> System.out.println("Done");
}
});
Event types
| Type | Description |
|---|---|
| discovery_started | Crawl discovery has begun |
| discovery_progress | Pages found during BFS (pagesFound field) |
| discovery_complete | Discovery finished, fetch phase starting |
| page_fetching | About to fetch page N of total |
| page_fetched | Page fetched successfully |
| page_failed | Page fetch failed (crawl continues) |
| crawl_complete | All pages processed |
Low-level methods
Need fine-grained control? The individual modules are still available:
evadr().fetch()
FetchResult page = client.evadr().fetch("https://example.com");
System.out.printf("Status: %d, Tier: %d%n", page.statusCode(), page.tierUsed());
System.out.printf("Artifact ID: %s%n", page.artifactId());
webgrph().crawl()
CrawlResult crawl = client.webgrph().crawl("https://example.com",
CrawlOptions.builder().maxDepth(3).maxPages(100).build());
System.out.printf("Crawl started: %s%n", crawl.crawlId());
kolektr().page()
ExtractionResult result = client.kolektr().page("https://example.com",
PageOptions.builder()
.schema(Map.of("title", "css:h1", "price", "css:.price"))
.build());
for (var record : result.records()) {
System.out.println(record);
}
Artifact chaining
Pass artifact IDs between low-level methods to skip redundant work:
FetchResult page = client.evadr().fetch("https://example.com");
ExtractionResult data = client.kolektr().page("https://example.com",
PageOptions.builder()
.schema(Map.of("title", "css:h1"))
.fetchArtifactId(page.artifactId())
.build());
CrawlResult crawl = client.webgrph().crawl("https://example.com",
CrawlOptions.builder()
.maxDepth(2)
.sessionArtifactId(page.artifactId())
.build());
Error handling
try {
SiteCrawlResult result = client.crawl("https://example.com", CrawlOptions.builder().build());
} catch (RateLimitException e) {
System.out.printf("Retry after %ds%n", e.getRetryAfter());
} catch (AuthenticationException e) {
System.out.println("Invalid API key");
} catch (KloakdException e) {
System.out.println(e.getMessage());
}
Advanced namespaces
client.skanyr() // API discovery
client.nexus() // AI strategy engine
client.parlyr() // Natural language queries
client.fetchyr() // RPA & authenticated scraping
Next steps
- Quickstart — get started in 5 minutes
- API reference — detailed endpoint docs
- Error reference — full error taxonomy
