LakehouseBox did not start as a plan for a company. It started in my garden.
Over a few months I kept running into the same gap from different directions. Each time it looked like a side problem, and each time it turned out to be the same one: there was no simple place to keep analytical data that an agent can set up, use and change by itself, on open standards, run in Europe. These are the five things that pushed me to build it, and the thought that ties them together.
1. A garden that needed a lakehouse
This summer I automated my garden. A weather station and soil probes report every minute, a camera takes a photo every fifteen minutes, and every morning a Claude Code routine reads the data and sends me a short report. The project is open on GitHub.
It is 2026, so the format was never in doubt: Apache Iceberg tables that DuckDB, Python or Spark can read. Finding somewhere to keep them was harder. Few services offered a modern catalog with the latest version of the format, and the closest one, Cloudflare's, needed more pieces just to get data in: a pipeline per table, a small program for the weather station, and a virtual machine on Google Cloud only to relay the camera's uploads. All of it ran on American infrastructure.
What I wanted was one cheap place where an agent could set up the whole project from the command line: the tables, the way the sensors and the camera send their data, and the agents that read it. The garden has run on LakehouseBox since 24 September, with nothing in between for me to run: no worker, no virtual machine, no queue.
2. Can analytics be sovereign?
The second reason was a question: can analytics in Europe run with nothing American underneath? I have argued before that sovereignty in the age of AI has to mean choice, not isolation. So I tried to make the choice.
It fell apart on measurement. A modern Iceberg service hands each client short-lived credentials scoped to one table, and that needs the object store to answer the STS requests that issue them. The European object stores we tested do not: OVHcloud returns 400 and Scaleway 412 to every STS action, and Hetzner's object storage has no API for it at all. Today a managed Iceberg catalog runs on one of the big American clouds or on Cloudflare. None of them is European.
So we run the object store and the catalog ourselves, as one system, on servers in Germany. That makes us responsible for durability, which is real work, and it makes the promise possible. Analytics is one of the few parts of a company's stack where full sovereignty is within reach, because every layer exists as open source or as a European commodity. The facts, including what we still depend on, are on the sovereignty page.

3. Agents that act, and see what happened
The third reason is the one I spend most of my time on.
For the last couple of years we have been giving agents access to data so they can answer questions. The more interesting step is when they act: change something in the world, see the effect in the data, and adjust. A garden is a good place to try that. In mine an agent now controls the watering. It decides when to water, and the soil probes tell it afterwards whether it was right. Its journal is a table beside the sensor readings, so every decision sits next to the data that judges it.
A loop like that needs plumbing: sensor readings, photos, the forecast, the irrigation controller. Each of those used to be an integration project. With AI it is something an agent writes in an afternoon. The infrastructure underneath has not caught up: most data platforms still assume a person in a console doing the setup. I wanted one that an agent like Claude Code can drive from end to end.
4. A European home for public data
The fourth reason is the second one seen from the publisher's side. Europe produces an enormous amount of public data, and the modern way to share it is as tables that people and agents can query directly, not files to download. But if no European object store can carry an Iceberg catalog, there is no European infrastructure to publish open data as Iceberg tables either. A publisher who wants to keep its data under European control has had no option.
On LakehouseBox a catalog can be made public. Anyone can read its tables without credentials, any account can attach it by name and join it with its own data, and public catalogs have their own free allowance: 5 GB, and up to 50 GB for datasets we review. Our own public catalog has weather, holidays, places and more, each with its source, its licence and a data card written for an agent. It is the idea I have been working on with Portolan: public data an agent can find, understand and query.

5. A place where the standards are current
The last reason is personal. I have spent years on open geospatial standards, from GeoParquet, which we introduced at CARTO in 2022, to native geometry in Parquet and Iceberg. Standards work is slow, and approval is only the first half of the story. The second half is implementation. This year I measured how much of Iceberg v3's geometry support had reached the engines people use every day. Some of it had.
Managed services wait for a standard to settle, and then for their own roadmap. I wanted a platform that does the opposite: Iceberg v3, its geometry and geography types and better clustering as soon as they work, with the best practices built in so nobody has to know them to benefit. LakehouseBox writes Iceberg v3 with native geometry today. This week we rewrote the places and rail tables in our public catalog that way, with rows ordered along a Hilbert curve, and a query for a box around central Madrid now reads 1.1 to 1.8 MB instead of about 26 MB. Large country polygons did not gain, and we wrote that down too.
When intelligence gets cheap
There is one more thought behind all five, and it is the one that excites me most.
For about twenty years analytics has lived by two slogans: you cannot improve what you cannot measure, and data is the new oil. The second came true in a way nobody intended. Like crude, most data was never refined. It was extracted, stored at a cost and left in the tank. The first slogan was missing its second half. You cannot improve what you do not analyse, and analysis needed specialists there were never enough of. So before any new data project the honest question was: why collect all this, if nobody will have the time to look at it?
AI is changing the answer. It is making intelligence cheap, and with it the analysis and the decisions that follow. If that holds, the cost of an analytics project starts to look like the cost of its sensors. Think of a stretch of coast where you want to understand pollution and currents. A project like that used to be expensive twice: once to gather the data, and again to pay the people who could make sense of it. The second cost is the one that is falling.
But cheap intelligence moves the bottleneck. When the analyst is an agent that looks at the data every hour and tries ten queries to keep one, everything around it has to be cheap too, in two ways.
Cheap in money: that is the lakehouse. Open Iceberg tables on object storage, the cheapest durable place to keep data, with the compute kept separate. The agent runs DuckDB wherever it already runs, so the analysis costs what that machine costs, and nobody bills it per query for reading its own data.
Cheap in effort, which matters more. If a person still has to build the pipelines, create the tables and wire up every sensor, the data engineer becomes the new bottleneck and the project is as expensive as before. So LakehouseBox is agentic from the bottom up. An agent can do all of it through the CLI and the API: the catalogs and tables, the intake a sensor posts to, the uploader a camera writes with, the keys scoped to each job and, in preview, agents that run on a schedule next to the data. A person approves the agent's connection once in the browser. After that, a new analytics project is a conversation, not a project plan.
That is why I see AI as a reason to be more ambitious, not less. For the first time we can analyse everything we measure, and act on it, including things like a garden or a beach that were never worth a data team. How many projects never started because nobody would have had time to look at the data?
Where we want to go
The goal is simple to say. Starting an analytics loop should be as easy as asking for one: "I want to use less water in my garden without losing a plant", or "I want to understand the pollution on this beach and act before it gets worse". The agent, with the platform underneath it, works out the rest: what to measure, which sensors and public data to connect, how to store it, the analysis, the automation and the report. Every decision and its result stay in the tables, so the next decision is a little better. Loops that get things done, and get better over time.
We are not there yet. An agent can already set up most of the data side by itself; knowing what to measure, and closing the loop well, is what we are working on now. It is early in other ways too: one region, no high availability yet, no paid plans yet, and a capabilities page that lists what is missing as plainly as what is there. The free plan is 5 GB with no card.
If you have an agent and some data it keeps losing, or a dataset you want to publish in Europe, the setup instructions are written for your agent to follow. And if you care about open standards, sovereignty or agents that close the loop, I would love to hear from you at hello@lakehousebox.com. Come build this with us.
