How it works

A high-level conceptual overview of how KLOAKD turns a single URL into structured, warehouse-ready data — without configuration, templates, or maintenance.

The Pipeline

From a URL to warehouse-ready data

Every URL flows through the same pipeline. Each step is handled by a specialized module, orchestrated automatically — you never configure which parts to use.

1. Input

Paste a URL. That's it. No selectors, no templates, no configuration. The platform takes it from here.

2. Access

The platform reaches the site even when it is protected by enterprise bot-management systems — no proxy, CAPTCHA, or browser service to configure separately.

3. Adaptive Fetch

The engine automatically chooses the most efficient way to retrieve each page and escalates only when a site requires it — fast by default, reliable when it matters.

4. Data-Layer Discovery

The platform locates the structured sources behind the page and extracts directly from them — capturing far more complete data than reading the HTML alone.

5. Adaptive Extraction

Extraction adapts on its own when a site changes, so collection continues without anyone stopping to repair code.

6. Structured Output

Kolektr auto-generates schemas, extracts multimodal data (text, images, OCR), and exports warehouse-ready JSON, CSV, or Parquet.

Adaptive Fetch

Efficient by default, reliable when it matters

At the core of the platform is an engine that decides, per request, how much effort a page actually needs. Simple pages are retrieved quickly and cheaply; only the pages that require more get it. That keeps costs low without sacrificing the 98% success rate on protected sites.

Fast by default

Everyday pages are retrieved with the lightest, quickest approach — keeping runs fast and costs low.

Escalates when needed

When a site pushes back, the engine automatically applies more capability until the page opens — no manual tuning.

Always returns data

The pipeline is designed to deliver a usable result rather than fail silently, even under difficult conditions.

Orchestration

Seven modules, one pipeline

Each module specializes in one capability. The platform orchestrates them automatically — you never need to decide which to use or in what order.

Evadr

Webgrph

Nexus

Parlyr

Fetchyr

Kolektr

Kloakd

At Scale

Consistent across large, protected sites

Large crawls are where most tools fall down. The platform is built to collect consistently across thousands of pages and to resume interrupted work rather than start over — so a single failure never costs you an entire run, and large jobs finish in a fraction of the time.

Consistent

Access is established once per site and reused, so protection is handled smoothly across the whole crawl.

Resumable

Interrupted work picks up where it stopped instead of restarting from the beginning.

Fast

Large jobs that would take a traditional crawler most of a day complete in a couple of hours.

Observability

Follow every run in real time

Each run streams events as it progresses — so your systems always know the current state, from start through to structured results.

GET /v1/pipeline/{id}/stream

event: started

event: progress

event: data_ready

event: complete

See it live

Create a free account and run the pipeline on any URL of your choosing.