Key takeaways
  • AWS has signed a definitive agreement to acquire DuckLabs, the company behind DuckDB. DuckDB stays open source under the MIT license, governed by the independent DuckDB Foundation.
  • Wekalp made a deliberately contrarian call: build its lakehouse execution layer on DuckDB running against Parquet on S3, not a distributed "big iron" engine, and cleanly separate context, storage and compute.
  • DuckDB has run at the core of Wekalp's processing engine in production for eighteen months, across the query engine, the tabular validation pipeline and the AI-driven analysis features.
  • AWS frames DuckDB as the engine for everyday queries of a terabyte or less, with its own services covering the exabyte end. That is the same split Wekalp arrived at by watching real customer workloads.
  • The acquisition changes the calculus, not the architecture: lower longevity risk, deeper S3 integration, and an easier answer for regulated customers who ask about open-source risk.

AWS has signed a definitive agreement to acquire DuckLabs, the Amsterdam-based company behind DuckDB. Hannes Mühleisen and Mark Raasveldt — the researchers who built DuckDB out of CWI, the Dutch institute that also gave the world Python — will continue to lead the team and the project's technical direction from inside AWS. DuckDB itself stays open source under the MIT license, governed by the independent DuckDB Foundation. The deal is expected to close in early September.

For most of the data world, this is a headline about consolidation. For us at Wekalp, it reads like a validation letter.

The bet we made a while ago

Wekalp is a B2B AI-native data lakehouse purpose-built for business, regulatory, and financial reporting and analytics. Our users work in a world where a one-dollar mismatch is an error, every number has to be traceable, every column has to be auditable, and every response has to land in time to meet a deadline. Whether the workload is a traditional SQL query or a Gen-AI conversation over the same data, DuckDB has proven to be a solid asset.

When we architected the platform, we made a slightly contrarian call: skip the "big iron" distributed engines by default and build our execution layer around DuckDB, running against Parquet on S3. The architecture was designed for composability, immutability, interoperability, security, and governance — deliberately preparing for the AI era. DuckDB let us cleanly separate the three foundational lakehouse layers of context, storage, and compute, which is what makes our AI initiatives not only closer to business needs but also secure, governed, and explainable.

At the time, that looked more like a cost and simplicity decision than a strategic one. DuckDB let us run SQL directly against Parquet in S3 with no separate cluster to provision, patch, or babysit — one less system standing between the data and the answers.

And it kept holding its own as data volumes grew. We designed our storage and partitioning strategy specifically to play to DuckDB's strengths: pruning partitions before a query ever touches disk, and keeping row groups sized for vectorized scans. Together, those push the practical ceiling of a single-node engine well beyond the naive case. (We wrote about the early part of that journey in how DuckDB and Pandas transformed our data platform.)

AWS's own announcement makes a related point about the industry at large: DuckDB's superpower is everyday queries of a terabyte or less, while S3, Redshift, Athena, EMR, Glue ETL, and SageMaker cover the exabyte-scale end. We arrived at that conclusion from the other direction — by watching our customers' actual workloads rather than reading a market thesis.

Eighteen months in production, not eighteen months of pitching

DuckDB has been running in production at the core of Wekalp's processing engine for eighteen months. Not as a proof of concept, and not as a component retrofitted to sound relevant to today's news. It has been handling sales volumes, regulatory lineage, and financial accuracy for our customers that entire time — through the same iterations that hardened the rest of the platform: schema evolution, PII encryption rollout, validation pipeline rebuilds. That work is what separates a demo from a system people trust.

Not a one-trick engine: where it has carried weight

This isn't a decision we made once and forgot, and it isn't a bet we're still waiting to see pay off. Across eighteen months of production load, DuckDB has kept showing up as the right tool as the platform matured:

  • Core query engine for the lakehouse. SQL running directly against Parquet in S3; no separate warehouse hop required.
  • Tabular validation pipeline. When we evaluated DuckDB against Polars for our data-quality and validation layer, DuckDB won on Parquet/S3 integration and SQL expressiveness.
  • AI-driven analysis features. Our conversational BI, visualization, and data exploration capabilities lean heavily on DuckDB underneath, running multiple agents to surface the right decision-making insights. DuckDB's ability to handle chained queries against Parquet without the latency of a heavier engine is what makes those features feel interactive rather than batch.

None of that required a separate cluster or a warehouse we'd pay for by the terabyte scanned. It runs in-process, close to the application — exactly the way DuckDB is designed to.

Why an AWS acquisition changes the calculus, not the architecture

We didn't choose DuckDB because we expected AWS to buy it. We chose it because it fit the shape of our customers' needs. But the acquisition still matters for a platform like ours, for four concrete reasons:

  • Longevity risk just went down. Betting a production platform on an open-source project always carries a question: who maintains this in five years? AWS acquiring DuckLabs while keeping the project MIT-licensed and Foundation-governed is about as good an answer as an infrastructure team can hope for — commercial backing without the rug-pull risk of a license change.
  • The sub-terabyte query thesis is now infrastructure strategy, not just a good idea. AWS is explicitly framing DuckDB as the layer that handles everyday sub-terabyte queries while its own analytics stack handles the exabyte end. That is the same split we already run internally: DuckDB-on-Parquet for day-to-day queries, S3 as the durable append-only store underneath. It's reassuring to see a hyperscaler's roadmap converge on the same layering.
  • Deeper S3 integration is coming, and we're already positioned for it. With Iceberg support live in S3 Tables and AWS signaling DuckDB integration across its services, the gap between "storage layer" and "query engine" will keep shrinking. Because our lakehouse already treats S3/Parquet as the source of truth and DuckDB as the engine sitting on top, we inherit that improvement curve for free instead of re-architecting to catch it.
  • It's reassuring for regulated customers. Our customers and prospects in regulated industries — banks and insurance — ask questions about open-source risk and platform stability. "The query engine at the core of our lakehouse is now backed by AWS, open source, and still MIT-licensed" is a much easier sentence to put in a security questionnaire than it was yesterday.

Agents behave a lot like people when they interact with data — they poke, they experiment, they run exploratory analysis on small slices before committing to anything.

That line from the AWS announcement is worth sitting with. It is precisely the access pattern DuckDB is fast at, and precisely why our agentic BI layer works the way it does.

What we're watching

A few threads worth following as this integration plays out:

  1. How AWS handles DuckDB inside Lambda and other serverless surfaces. We already lean on in-process execution, so serverless DuckDB is directly relevant to our ingestion path.
  2. Whether S3 Tables and Iceberg integration deepens enough to simplify our current partitioning approach.
  3. Whether AWS-native DuckDB tooling emerges that could sit alongside — not necessarily replace — the encryption and validation layers we've built ourselves.

For now, the practical takeaway is simple: this isn't an architecture we're pitching off the back of someone else's acquisition news. It has been carrying live production workloads, quietly, for a year and a half before this was a headline. And it's the same direction the biggest cloud vendor in the world just spent real money to bet on too.

If you want to see how that architecture shows up in practice, start with our technology stack or why AI-native data architecture is question-first.