🎉 Unlock the Power of AI for Everyday Efficiency with ChatGPT for just $29 - limited time only! Go to the course page, enrol and use code for discount!
Menu

Write For Us

We Are Constantly Looking For Writers And Contributors To Help Us Create Great Content For Our Blog Visitors.

Contribute
Kadoa Review: The Web Scraper That Repairs Itself Before Anyone Notices
General, AI Tools Review

Kadoa Review: The Web Scraper That Repairs Itself Before Anyone Notices


Sep 11, 2026    |    0

Every Monday, somewhere in your company, a person opens sixty browser tabs.

They copy a price. They paste it into a spreadsheet. Then a delivery date, a permit number, a job posting, a competitor's new listing. Three hours later there's a file called master_final_v4.xlsx and a mild headache, and the numbers in it are already a day old.

The worse version happens later. One of those sixty sites gets a redesign — nothing dramatic, a button moves, a table becomes cards. The spreadsheet doesn't announce this. It just quietly starts collecting garbage, and nobody notices for eleven days.

Here's the reframe worth sitting with: the problem was never getting the data. It's public. It's sitting right there on pages anyone can open. The problem is that the web was built to be looked at, not read by machines. So the moment you automate the looking, you've adopted a pet that breaks every time somebody else redecorates their living room.

Kadoa is a bet that an AI should own that pet.

So what is Kadoa?

Kadoa is a Swiss platform that turns public web pages into structured, scheduled, monitored data feeds. Based in Zurich, founded by Adrian Krebs (CEO), Johannes Engler (CTO) and Tavis Lochhead (CRO). The product dates to 2023 — its first blog post, that November, was about using GPT-4 Vision to read web pages the way a person would.

Investors have shown up twice: an early round with AIX Ventures, then a round led by VI Partners with Fyrfly announced in March 2025. How much? Nobody says. Crunchbase, Dealroom, CB Insights and the investors' own press release all leave the number out, so anyone quoting a total raised is guessing.

Now the part that shapes everything else, and that most reviews skip.

Kadoa changed who it sells to. From 2023 to 2025 it was a general-purpose, no-code scraper for basically anybody. Today the homepage reads "The web data layer for finance," and the line underneath promises pipelines for investment teams. Every testimonial is an anonymized job title — "Head of Data Science, US Hedge Fund," "Director of Research, Global Quant Firm." There are no named customers anywhere in public.

The technical documentation, meanwhile, is still perfectly general-purpose. It'll happily scrape a shoe store.

That gap is the story. The product still does what it always did; the sales motion moved upmarket and left the small team behind. Keep it in mind — it explains almost everything below, including the price you can't see.

How it actually works

You type a sentence, not a URL. Since March 2026 that's the default way in: describe what you want, and the assistant works out where to get it. From there it explores the site, proposes a data structure you can edit, builds and tests a pipeline, shows you a sample to approve, and goes live with scheduling and alerts.

Step two is the clever bit, and it's easy to skim past. Verbatim:

"Kadoa explores the site and finds the most reliable source for the data (e.g. API endpoint, embedded JSON, CSV file)"

Sit with what that means. Most scrapers photograph the plated dish — they read the visible page, the part designers keep rearranging. Kadoa walks into the kitchen. Behind almost every modern web page is a plainer, more stable feed the page itself draws from. It's ugly, it isn't meant for humans, and it barely ever changes. Going one layer down is why this survives redesigns that flatten ordinary scrapers.

And when something does break, Kadoa claims to fix it: "Kadoa repairs, tests, and redeploys broken pipelines automatically." Vendors say things like that constantly, so the mechanism matters more than the promise — when a site changes, it regenerates the extraction code and then **validates the new output against data it collected previously.**That's a proofreader who remembers what the sentence used to say, not one who merely checks that a sentence exists.

Two more things worth knowing before the money section. Previewing is free — you build a workflow, inspect a sample of what it caught, and approve or reject before anything counts against the bill. And the plumbing is grown-up: output as JSON, CSV, Parquet or Delta, delivered to Snowflake, Databricks, Amazon S3, Google Cloud Storage, SFTP, webhooks, Google Sheets, Slack or email. There's also an MCP server, which in plain English means ChatGPT or Claude can query your data directly instead of you exporting a file first.

Honest gaps: no Zapier, no Airtable. And the Chrome extension — the old one-click way to try it — is gone; the page that used to advertise it now returns a 404.

Kick the tires without talking to a salesperson

Kadoa runs six public datasets, live and free, which beats any sales deck:

  • Congress Trading Monitor — 68,012 disclosed trades from 447 filers, parsed straight from House, Senate and ethics-office filings
  • US Layoffs Tracker — 48,824 WARN Act notices covering 5,413,105 workers since 1988
  • Quant Job Market — 4,104 live roles across 74 funds and trading firms
  • US Datacenter Tracker — 2,241 sites, deduplicated
  • Plus a POTUS Tracker and a World Mining Monitor

Look at that datacenter tracker for a second, because it's the entire pitch in one object. Nobody publishes a master list of American datacenters. Kadoa built one out of local building permits — public, tedious, scattered across hundreds of county websites nobody would visit twice. Boring, until somebody stacks it up. That's the business.

Where it falls down

The strangest thing about Kadoa is that almost nobody has reviewed it. Across G2, Trustpilot, SourceForge, Slashdot, Product Hunt and TheresAnAIForThat there are two identifiable human reviewers in total — and the negative one is the same person posting on all six sites. So when a listicle reports that Kadoa scores 1.0 out of 5, that isn't sentiment. That's one furious ex-customer, one happy trial user from three years ago, and silence.

Read the complaint anyway, because it's specific. A CEO who used it for one to two years, writing in August 2025: "they change their entire API endpoints on a whim on the SAME VERSION, breaking all of the work you've done to integrate with them," and "the UI fails to save changes every other try." Single-source and unrebutted — but entirely consistent with a small company moving fast.

There is no published accuracy figure — and the number people mistake for one is measuring something else. The homepage advertises a "99+% run success rate." That says the job ran. It says nothing about whether the numbers it brought back were right. On extraction accuracy, the actual question, Kadoa publishes nothing.

What it does ship is quietly revealing: a completeness check comparing actual row count against expected, and "suspicious value detection" for values that don't fit the rest of the row. Those guardrails exist for a reason.

The one reproducible outside test comes from Parsera — a direct competitor, so weigh accordingly — which ran its own test against a single jewelry site in February 2026 and reported roughly 75% accuracy with product URLs missing on one catalog, and 29 of 95 products captured on another.

Coverage has edges. An independent scraping newsletter reported in March 2026 that Kadoa offers only five proxy locations — which matters if you need region-specific pricing — and published a screenshot of an "Unsupported scraping URL" error. Some targets get refused outright.

The documentation is silent on the hard cases. No page on sites behind a login, none on CAPTCHAs, none on rate limits. Supported content types are HTML, XML, RSS, JSON and PDF. Silence isn't proof of failure, but it's the first thing to raise on a sales call.

And the door for small teams is closing. Extension gone, pricing behind a conversation, data residency reserved for Enterprise.

What it costs

There are no prices. That isn't a summary — the pricing page shows two plans and not one dollar figure. Starter offers a free trial, "consumption-based pricing" and basic support behind a Try it free button. Enterprise offers custom limits, SAML SSO, an SLA and an account manager behind Contact us.

Ignore every number you'll find elsewhere. SourceForge lists $300/month. A competitor's blog remembers a retired $39/month tier. Kadoa's own January 2026 blog post still says plans start at $99/month. Three sources, three answers, none matching the live page.

What you can pin down is the unit, and that's where the budget lives. Kadoa bills in slots. From the docs, verbatim: "A slot holds one active workflow in production."

You are not buying data by the row. You're renting a parking space per live scraper. Three rules follow, and they're the ones to model before signing anything:

  1. Building and previewing is free and never uses a slot. Experimenting costs nothing.
  2. Every enabled workflow carries a 30-day minimum. The docs are blunt: "One ad-hoc run holds a slot for 30 days." Run something once, pay for a month.
  3. Kadoa classifies complexity, not you. Simple counts 0.1×, standard 1×, complex "2x or more" — though Kadoa puts complex at under 1% of workflows on its platform.

There's also a soft ceiling: rows and unblocking are "covered by fair use," with a multiplier if you exceed it. Your volume limit is a negotiation, not a contract term.

The fine print, in plain English

The security half is solid. SOC 2 Type 1 and Type 2 attestations. TLS 1.3 or higher in transit, encryption at rest, 30-day backups, annual third-party penetration testing, isolated execution environments — and US or EU data residency, which is Enterprise only.

On ethics, watch the verb. The docs say Kadoa "automatically checks robots.txt directives for all web interactions, promoting ethical data sourcing from websites." Checks and promotes. Not obeys, not blocks. It also detects personal information and can be configured to redact, flag or block it.

Then the awkward half: the legal documents are old. The Privacy Policy reads "Version 1.0 – February 2024." The Terms of Service read "Version 1.0 – Effective as of January 1, 2023." Both predate the finance repositioning — and the terms are dated before Kadoa's first blog post even existed. The privacy policy still runs under the old corporate name, reviewr.ai AG, lists no sub-processors, sets no retention periods, and says nothing about AI or model training — so a hard no-training commitment has to be asked for in writing.

Then there's the clause that should stop a buyer mid-sentence. Section 8.2:

"The Intellectual Property Rights of the Deliverables remain with Kadoa."

Followed by: the customer "may not commercially exploit such Deliverables in the form as provided by Kadoa."

In plain English: you rent your own scraped data. Internal use is fine. Reselling it as delivered is not, unless you've transformed it enough to count as genuinely derivative work. For anyone whose business model involves packaging data for someone else, that's not boilerplate — that's the negotiation.

Two more. Liability is excluded, not capped — no familiar "twelve months of fees" formula, just ordinary negligence excluded as far as the law allows. And it's Swiss law with exclusive jurisdiction in Zurich, so a London or New York buyer with a dispute is booking a flight.

Also, for the record, the terms forbid "systematic retrieval (e.g. scraping) of content from the Services." A scraping company that prohibits scraping it. Somewhere a lawyer wrote that with a completely straight face.

The thing nobody is pricing in

On 15 September 2026 — five days from now — Cloudflare starts blocking "mixed-use" crawlers by default on ad-supported pages, for new customers, new sites of existing customers, and every existing free-tier customer. Cloudflare sits in front of about a quarter of all websites. Matthew Prince's framing: "Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge."

European regulators, separately, adopted draft web-scraping guidelines on 8 July 2026 — still open for consultation until the end of October — in which robots.txt, ai.txt and CAPTCHA become factors weighed when judging whether collecting someone's data was reasonable. The same document repeats, without ambiguity, that "publicly available" is not an exemption.

Neither is a Kadoa problem specifically. Both are a supply problem for everyone. The easily-scrapeable web is shrinking on a published schedule — which raises the value of a maintained pipeline and, at the same time, the risk of paying for one that quietly stops reaching half its sources.

The verdict

Kadoa gets the hard engineering right. Going one layer beneath the visible page instead of photographing it is the correct instinct. Self-healing that validates against historical data is a real idea, not a slogan. Previews are free, the delivery into Snowflake and S3 is grown-up plumbing, and those six public datasets are the most honest sales material on the site.

The other column is what it costs in trust: no published price, no named customer, two reviewers on the entire internet, no accuracy figure, and legal documents untouched since before the company decided who it was for.

It doesn't make web data free. It makes web data somebody else's maintenance problem — and the terms of service are the price of that trade.

So: an investment or research team with real budget and a permanent data appetite should book the call, and should negotiate that IP clause before anything else. A small team that just needs competitor prices in a spreadsheet by Tuesday should look at Browse AI, Octoparse or Thunderbit — they publish their prices, they start free or near it, and they are still built for you.

Either way, stop opening sixty tabs.