How it works
A high-level conceptual overview of how KLOAKD turns a single URL into structured, warehouse-ready data — without configuration, templates, or maintenance.
The Pipeline
From a URL to warehouse-ready data
Every URL flows through the same pipeline. Each step is handled by a specialized module, orchestrated automatically — you never configure which parts to use.
1. Input
Paste a URL. That's it. No selectors, no templates, no configuration. The platform takes it from here.
2. Access
The platform reaches the site even when it is protected by enterprise bot-management systems — no proxy, CAPTCHA, or browser service to configure separately.
3. Adaptive Fetch
The engine automatically chooses the most efficient way to retrieve each page and escalates only when a site requires it — fast by default, reliable when it matters.
4. Data-Layer Discovery
The platform locates the structured sources behind the page and extracts directly from them — capturing far more complete data than reading the HTML alone.
5. Adaptive Extraction
Extraction adapts on its own when a site changes, so collection continues without anyone stopping to repair code.
6. Structured Output
Kolektr auto-generates schemas, extracts multimodal data (text, images, OCR), and exports warehouse-ready JSON, CSV, or Parquet.
Adaptive Fetch
Efficient by default, reliable when it matters
At the core of the platform is an engine that decides, per request, how much effort a page actually needs. Simple pages are retrieved quickly and cheaply; only the pages that require more get it. That keeps costs low without sacrificing the 98% success rate on protected sites.
Fast by default
Everyday pages are retrieved with the lightest, quickest approach — keeping runs fast and costs low.
Escalates when needed
When a site pushes back, the engine automatically applies more capability until the page opens — no manual tuning.
Always returns data
The pipeline is designed to deliver a usable result rather than fail silently, even under difficult conditions.
Orchestration
Seven modules, one pipeline
Each module specializes in one capability. The platform orchestrates them automatically — you never need to decide which to use or in what order.
Evadr
Webgrph
Nexus
Parlyr
Fetchyr
Kolektr
Kloakd
At Scale
Consistent across large, protected sites
Large crawls are where most tools fall down. The platform is built to collect consistently across thousands of pages and to resume interrupted work rather than start over — so a single failure never costs you an entire run, and large jobs finish in a fraction of the time.
Consistent
Access is established once per site and reused, so protection is handled smoothly across the whole crawl.
Resumable
Interrupted work picks up where it stopped instead of restarting from the beginning.
Fast
Large jobs that would take a traditional crawler most of a day complete in a couple of hours.
Observability
Follow every run in real time
Each run streams events as it progresses — so your systems always know the current state, from start through to structured results.
GET /v1/pipeline/{id}/stream
event: started
event: progress
event: data_ready
event: complete
See it live
Create a free account and run the pipeline on any URL of your choosing.
