Agent instructions for One copy of your data, read by every tool.; the page for people is https://lakehousebox.com/docs/tools/

# Engines

LakehouseBox is an Apache Iceberg REST catalog plus S3-compatible storage. Your engine talks to both directly. Each product has its own page with what it can do here, how to connect it, and its peculiarities; lhbox connect --engine … prints the recipe for the ones that have one (--catalog <name> with several catalogs), with the bucket filled in and the credential's secret shown as ******** unless --show-secrets. Catalog: https://catalog.lakehousebox.com. Storage: https://s3.lakehousebox.com.

| Product | Status | Reads | Writes | Iceberg v3 | Geometry, geography | Last verified |
|---|---|---|---|---|---|---|
| DuckDB (https://lakehousebox.com/docs/engines/duckdb/) | verified | yes | yes | reads and writes (1.5.5+) | geometry, read and written, pruned with &&; no geography (cannot open such a table) | 2026-10-03 (DuckDB 1.5.6 (and DuckDB-WASM 1.5.6 in the browser)) |
| PyIceberg (https://lakehousebox.com/docs/engines/pyiceberg/) | verified | yes | yes (format version 2) | reads; cannot write | no (reads schema and counts only) | 2026-10-03 (PyIceberg 0.12) |
| Spark (https://lakehousebox.com/docs/engines/spark/) | verified | yes | yes | reads and writes | reads and writes values; spatial functions need Apache Sedona | 2026-10-03 (Spark 3.5 + Iceberg 1.11; Spark 4.1 + Iceberg 1.12) |
| Trino (https://lakehousebox.com/docs/engines/trino/) | verified | yes | yes (format version 2 tested) | reads | geometry read; spherical distance for points only | 2026-10-03 (Trino 483) |
| Polars (https://lakehousebox.com/docs/engines/polars/) | verified | yes | yes | not tested | not tested | 2026-09-28 (Polars 1.44 + PyIceberg 0.12) |
| ClickHouse (https://lakehousebox.com/docs/engines/clickhouse/) | partial | with icebergS3, one table at a time; not through the catalog | no | not tested | not tested | 2026-09-28 (ClickHouse 25.8 LTS, 26.9) |
| Snowflake (https://lakehousebox.com/docs/engines/snowflake/) | verified | yes, on an allowed account | yes, on an allowed account | reads and writes | geometry and geography, read and written | 2026-10-03 (Snowflake 10.35) |
| Databricks (https://lakehousebox.com/docs/engines/databricks/) | untested | untested | untested | untested | untested | never run against the service |
| BigQuery (https://lakehousebox.com/docs/engines/bigquery/) | not possible today | not possible today | not possible today | not possible today | not possible today | never run against the service |

"Verified" means a recorded run against the service, on the date and versions shown; the product pages say exactly what was run. Something missing, or a product you use with LakehouseBox that is not here? Write to hello@lakehousebox.com.

## Gotchas, all of them measured

- DuckDB: ATTACH the bare bucket name. ATTACH 'acme--demo-data' is read-write. ATTACH 's3://acme--demo-data/' attaches READ-ONLY and every INSERT fails with "attached in read-only mode". PyIceberg and Spark want the s3://acme--demo-data/ form for warehouse, so copying one recipe's identifier into the other engine produces a false "DuckDB cannot write Iceberg".

- DuckDB 1.5.5 or newer. 1.5.0 writes manifest lists the catalog's maintenance cannot read, so tables it wrote were neither compacted nor expired by the scheduler. 1.5.5 writes the spec-compliant schema.

- Format version 2 by default; DuckDB 1.5.5+ also writes version 3. Every table has its own format version; lhbox table create without --format-version takes the catalog's default_format_version (2 unless you set it: lhbox catalog create <name> --default-format-version 3 or catalog update <name> --default-format-version 3; shown in catalog list and in the recipe's format_version_hint). An engine that creates a table itself (DuckDB CREATE TABLE … AS) chooses its own version, v2 today, and is never upgraded silently. Version 3 is needed for geometry/geography; it is written by DuckDB 1.5.5 or newer and by Spark 3.5 + Iceberg 1.11 (verified 2026-09-20 with a geometry(EPSG:28992) column loaded by a plain INSERT, CRS kept in the Parquet logical type). PyIceberg 0.12 reads v3 but cannot write it, and cannot load a geometry column with a non-default CRS. Who writes and who only reads each version: GET /v1/config/formats (no auth), echoed as writers/readers by table create and table get.

- The catalog only accepts its own tokens. Engines mint them from the catalog credential (catalog_credential in the recipe) at https://catalog.lakehousebox.com/v1/oauth/tokens (client_credentials; client_id = access key, client_secret = secret key) and refresh them themselves; they expire after an hour (the store's ICEBERG_OAUTH_TOKEN_EXPIRY, 3600 s: a vended credential never outlives the token that asked for it). LakehouseBox API keys (al_live_…) and tokens from POST /v1/tokens are 401 at the catalog by design.

- PyIceberg asks for vended credentials on every request. That is the intended path: the catalog returns per-table storage credentials which override any static s3.* key you pass. If a cross-bucket copy fails with ACCESS_DENIED, pass "header.X-Iceberg-Access-Delegation": "none" to use your own keys instead.

- Table buckets accept only Iceberg files. Data files must be .parquet/.orc/.avro/.lance and metadata must look like Iceberg metadata, under <namespace>/<table>/(data|metadata)/; anything else is refused on write. Put photos, documents and other blobs in the catalog's blob bucket <handle>--<catalog>--blobs (same credential, S3 path-style, region us-east-1; lhbox connect prints a boto3 example) and index them in a table.

- CTAS makes every column optional. CREATE TABLE … AS SELECT from DuckDB cannot express identifier fields, required columns or partitioning. When you need them, lhbox table create --column name:type[:required][:identifier] first, then INSERT.

- Maintenance is on and not disableable. Snapshot expiry (20 snapshots, 7 days) and orphan cleanup run on every catalog, and compaction (128 MB target files) on its format-version 2 tables; format-version 3 tables are not compacted yet (roadmap). A client that hard-codes vN.metadata.json file names will 404 after a maintenance commit; list the prefix or go through the catalog.

- Commits are optimistic: retry on 409. Iceberg writers commit against the table's current snapshot; when another writer or the platform's maintenance committed in between, the catalog answers 409 and DuckDB raises CommitFailedException … branch "main" has changed. DuckDB does not retry. An unattended job should catch that error, wait a few seconds and re-run the statement (measured 2026-09-20: two concurrent appenders with a client-side retry landed every row, no duplicates). Maintenance can commit into a table within seconds of its creation; a short blackout of the catalog after such a commit was also observed and is being fixed.

- CREATE OR REPLACE TABLE is not supported. DuckDB's iceberg extension refuses CREATE OR REPLACE TABLE on an attached Iceberg catalog. To redo a table, DROP TABLE demo_data.demo.cities; then CREATE TABLE … again (two commits, the second into a fresh table).

- TIMESTAMP WITH TIME ZONE values need pytz in Python. now() and every TIMESTAMP WITH TIME ZONE column come back through the Python client (duckdb.sql(…).fetchall(), .df()) only when the pytz module is installed, which a bare venv lacks: pip install pytz. Or print inside DuckDB with .show(), cast (now()::VARCHAR), or leave the column out of the SELECT. The duckdb binary and lhbox duckdb are not affected.

- Rotation. lhbox catalog rotate <name> replaces both catalog credentials: the old read/write key and every catalog token from it are refused at once; the old read-only key is deleted in the background right after the answer (about 10 s; warehouse.rotate_complete in the audit log records it); storage sessions already vended run out within 900 s. Fetch a fresh recipe afterwards.

## Direct S3 access from a script

lhbox credentials --catalog demo_data --namespace demo --table cities (POST /v1/credentials) returns storage credentials scoped to that table's object prefix, with their expiry, for a boto3 or aws-cli session that needs the files themselves rather than the table.

---
HTML version of these instructions: https://lakehousebox.com/docs/engines/ · every page: https://lakehousebox.com/llms.txt · everything in one file: https://lakehousebox.com/llms-full.txt
