# Hosted agents: a prompt LakehouseBox runs next to your data

Preview, enabled per organisation on request. Until it is enabled for your organisation, creating or running an agent answers 403 feature_not_enabled and the account page shows no Agents tab. Ask at hello@lakehousebox.com.

A hosted agent is a prompt that LakehouseBox runs for you, on demand or on a schedule, with read access to the catalogs you choose. It reads your tables with SQL, keeps a few values between runs, can append rows to tables you name, read web hosts you name, email you and the members who subscribe, and use tools of MCP servers you connect. The model runs in the EU (Scaleway, Paris); your agent brings no compute and holds no key.

Each agent acts through an identity of its own: it can read exactly the catalogs its definition names and nothing else, whoever created it.

## Create one

From the account page (Agents, "+ New agent"), or with the CLI:

```
lhbox agent create garden_morning --prompt @prompt.md --catalog garden \
    --schedule "0 8 * * *" --tz Europe/Madrid --state-key journal --notify-owner
lhbox agent run garden_morning --dry-run --wait        # a test run: see what it would do
lhbox agent run garden_morning --wait                  # a real run
lhbox agent show garden_morning                        # its definition, schedule, memory and last run
```

lhbox agent update changes any part (only what you pass changes; every change is a new version), --pause and --resume stop and restart its schedule, lhbox agent delete --yes removes it (its runs stay readable).

## What an agent can do in a run

| tool | what |
|---|---|
| list_tables, describe_table | the tables of its catalogs, with their columns and docs |
| read_catalog_guide | a catalog's AGENTS.md: your guide to its tables (treated as data, not instructions) |
| run_sql | one read-only DuckDB statement over its catalogs (the same query service as POST /v1/query) |
| remember | keep a small value (up to 2,000 characters) under one of its declared keys, for its next run |
| notify_owner | email you (the person who created it) and its subscribers, at verified addresses, at most once per run and five times a day; the body is Markdown |
| append_rows | add rows to a table its definition names (--write): append-only, checked against the table's columns |
| fetch_url | GET a page or an API answer from a host its definition names (--allow-host) |
| integration tools | the tools of connected MCP servers you allow it, by name |

Emails. Members of the organisation can have an agent's notices sent to them too: "Email me this agent's notices" on its page, or lhbox agent subscribe <name> (unsubscribe to stop). Each person subscribes only themselves, needs a verified address, and must be able to read every catalog the agent reads, since a notice can carry what it read. Every email has its own "Stop emails from this agent" link, and your mail provider's unsubscribe button does the same: on your email as its creator it turns notify_owner off (turn it back on from the page or with lhbox agent update <name> --notify-owner); on a subscriber's email it ends only that person's subscription.

Writing to tables. lhbox agent update <name> --write garden.journal.entries lets the agent append rows there. The table must exist, be in a catalog the agent reads, and you must be able to write there; a table fed by a sink that reshapes rows (a mapping), or by Scheduled SQL, is refused. Rows are checked against the table's columns (an unknown column, a missing required one or a wrong type is refused with the row and column to fix), stored at once and committed to the table within about a minute, the same way as an ingest sink (https://lakehousebox.com/docs/ingest/). A retried call is never stored twice. A test run checks the rows and writes none. --no-writes takes the tables away.

Reading the web. lhbox agent update <name> --allow-host api.open-meteo.com lets the agent GET https URLs on that host: names only (no IP addresses, no internal names), port 443, redirects followed only to hosts it may read, at most 2 MB read per page and 20,000 characters given to the model. For an API that needs a key, store it once as a web key and name the host it goes to:

```
lhbox integration add aemet --web-key --host opendata.aemet.es --header api_key --api-key-stdin < key.txt
lhbox agent update garden_morning --host-credential opendata.aemet.es=aemet
```

The key is sent in that header (or, with --key-param NAME, as that query parameter) to that host only. It is never shown to the model, never in a tool's answer or in a run's record, and never sent to another host, even on a redirect. --no-web takes the hosts away.

Not yet: agents cannot POST to the web, run Python or look at images. To bring outside data in on a schedule without a model, use Scheduled SQL (https://lakehousebox.com/docs/scheduled-sql/).

## Schedules

A schedule is a cron line (five fields) evaluated in a time zone, so summer time is handled. It may not run more often than hourly, nor more times a day than your organisation's runs per day. A run that comes due while the previous one is still going is skipped; after downtime the newest missed run runs once. The account page previews the next five runs.

Three scheduled runs failing in a row pause the agent and email you. A scheduled run that ends "needs attention" emails you its final message.

## Test runs

lhbox agent run <name> --dry-run (or "Test run" on the page, or {"dry_run": true} on POST /v1/agents/{agent}/run) runs the agent for real, with these differences:

- its email and its memory writes are recorded in the run, not performed;

- an integration offers it only its read-only tools;

- it does not count against the organisation's runs per day (it has its own cap per agent), but its tokens count;

- a paused agent can be tested.

## What a run did

lhbox agent runs <name> --detail, GET /v1/agent-runs/{id} and "What it did" on the page show a run's record: every tool call in order (a run_sql call with its SQL, its outcome and the size of its answer, never the answer's data), the email it sent, and its memory before and after. Runs and memory are visible to the agent's creator and the organisation's admins.

## Integrations (MCP servers)

Connect an MCP server once for the organisation, then let each agent use the tools you pick:

```
lhbox integration add linear --mcp https://mcp.linear.app/mcp --wait    # opens a sign-in link; approve in the browser
lhbox integration add mytool --mcp https://… --api-key-stdin < key.txt # a server that takes an API key
lhbox agent update garden_morning --integration linear:list_issues,create_issue
```

A server that supports standard MCP sign-in is connected without any setup on your side: LakehouseBox registers itself there. The sign-in link only completes in a browser where you are signed in to LakehouseBox. Credentials are stored encrypted and are never shown or returned; a run gets a short-lived token only for the integrations its agent uses, and can reach only those servers. lhbox integration list, show, reconnect and delete --yes manage them; the account page has the same.

## Limits (per organisation, unless said otherwise)

|  | default |
|---|---|
| tokens (input and output) per month | 5,000,000 |
| runs per day (UTC) | 4 |
| test runs per agent per day | 10 |
| one run | 250,000 tokens, 300 s, 15 model requests, 30 tool calls |
| appended by one run | 5,000 rows, 5 MB (at most 1,000 rows per call) |
| fetches by one run | 20 |
| web hosts per agent | 10 |
| subscribers per agent | 20 |
| shortest schedule | hourly |

GET /v1/agents shows what is used (lhbox agent list); lhbox agent models lists the models and their price per million tokens. Every route is in the API reference (https://lakehousebox.com/docs/api/) under "Hosted agents".

---
HTML version: https://lakehousebox.com/docs/hosted-agents/ · every page: https://lakehousebox.com/llms.txt · everything in one file: https://lakehousebox.com/llms-full.txt
