Host a Portolan catalog on LakehouseBox, and query it as Iceberg tables
A guide for the Portolan community. Every command and output below comes from a run against the service on 2026-10-04.
You publish a Portolan catalog: STAC JSON, GeoParquet, Cloud-Optimized GeoTIFFs, README and AGENTS.md files. This guide puts that catalog on LakehouseBox, in the EU, without changing how you build it: portolan push uploads it to your catalog's file bucket, one command publishes the folder read-only (range requests and CORS included), and rashid checks the result against the public URLs. Then, optionally, each vector collection becomes an Iceberg table that DuckDB queries by name, and that anyone can read with no account.
Every command and output below comes from one run on 2026-10-04 against the live service. The organisation handle is shown as <handle> and keys are never shown.
What you need
- portolan-cli 0.8.0 to build and push the catalog, and rashid 0.1.8 to check it. Install portolan-cli in a virtual environment of its own: it pins
duckdb<1.5.5, and in a shared environment it downgraded DuckDB to 1.5.4, older than LakehouseBox needs for writing. - DuckDB 1.5.6 (the binary, for
lhbox duckdb, or the Python module) for the queries, with thespatialextension. - A LakehouseBox account and the
lhboxCLI, logged in (Get started). jqand the AWS CLI, used once each below.
The run used a small catalog named portland with two collections: stops, TriMet's 6,316 stops as GeoParquet (from the TriMet Portolan mirror on source.coop), and elevation, a 4.3 MB Cloud-Optimized GeoTIFF cut from the Copernicus DEM GLO-30.
1. Create a catalog
A LakehouseBox catalog has two places for data: Iceberg tables, and a file bucket for everything else. The STAC tree goes in the file bucket.
lhbox catalog create portland
name portland
bucket <handle>--portland
blob_bucket <handle>--portland--blobs
default_format_version 2
status ready
2. Build the Portolan catalog as usual
Nothing here is specific to LakehouseBox:
portolan init --auto --id portland --title "Portland transit and terrain" \
--description "TriMet stops and a Copernicus 30 m elevation model for Portland, Oregon" --license CC-BY-4.0
portolan add stops/stops.parquet
portolan add elevation/dem-2021/portland-dem.tif --datetime 2021-04-22
portolan readme
Fill in providers, contact and source_url in each collection's .portolan/metadata.yaml before portolan add: add writes the providers into collection.json, and a later portolan check --fix did not. portolan-cli wants a raster in a folder of its own (elevation/dem-2021/), as above.
rashid on the local tree, before anything is uploaded, gave 1 error: PTL-COL-001 on elevation, because portolan-cli 0.8.0 wraps a single raster in an item. It is portolan-cli's layout and has nothing to do with hosting. That is the baseline the hosted copy is compared against in step 5.
3. Push it to the file bucket
portolan push writes to any S3-compatible store. Give it the catalog's read/write key pair and the LakehouseBox endpoint:
eval "$(lhbox connect --catalog portland --write --show-secrets --json | jq -r '"export AWS_ACCESS_KEY_ID=\(.catalog_credential.client_id) AWS_SECRET_ACCESS_KEY=\(.catalog_credential.client_secret)"')"
export AWS_REGION=us-east-1 PORTOLAN_S3_ENDPOINT=https://s3.lakehousebox.com
portolan push s3://<handle>--portland--blobs/portland
→ No remote versions.json found (first push)
✓ [1/2] elevation: 1 version(s), 7 file(s)
✓ [2/2] stops: 1 version(s), 5 file(s)
✓ Uploaded catalog.json
✓ Uploaded versions.json
✓ Pushed 2 collection(s), 2 version(s), 16 file(s) (4.4 MB, avg 3.9 MiB/s)
The pair can write to the whole catalog, so keep it in your shell or an uncommitted .env, never in the catalog's repository. An uploader key is not enough here: it can only write, and portolan push reads the remote versions.json first.
4. Publish the folder read-only
lhbox catalog publish portland --confirm portland --files portland/
Catalog portland is PUBLIC since 2026-10-03T23:05:04.946Z.
public_url: https://s3.lakehousebox.com/<handle>--portland/
public_files_url: https://s3.lakehousebox.com/<handle>--portland--blobs/portland/ (the folder portland/ of the file bucket; the rest stays private)
Your STAC root is now https://s3.lakehousebox.com/<handle>--portland--blobs/portland/catalog.json. Anyone can read the folder's data files and list them; nobody outside your organisation can write or delete there. Publishing makes the catalog's tables public too, and its bytes count against the public allowance (5 GB; 50 GB once we review the dataset): see Public catalogs.
What a reader without an account got, file by file:
catalog.json 200 application/json
AGENTS.md 200 text/plain; charset=utf-8
stops/collection.json 200 application/json
stops/stops.parquet 200 application/vnd.apache.parquet
stops/stops.thumb.jpg 403
elevation/dem-2021/portland-dem.tif 200 application/octet-stream
elevation/dem-2021/portland-dem.thumb.jpg 403
The thumbnails answer 403: see Limits. A range request with the Portolan browser's origin:
curl -s -o /dev/null -D - -H "Range: bytes=0-16383" -H "Origin: https://browser.portolan-sdi.org" \
https://s3.lakehousebox.com/<handle>--portland--blobs/portland/elevation/dem-2021/portland-dem.tif
HTTP/2 206
accept-ranges: bytes
access-control-allow-origin: https://browser.portolan-sdi.org
access-control-expose-headers: ETag, Content-Length, Content-Type, Last-Modified, x-amz-request-id, x-amz-version-id, Content-Range, Accept-Ranges
content-range: bytes 0-16383/4267106
etag: "c1687c8f5a30273dc8b28d1e6b24abdf"
content-length: 16384
5. Check it with rashid against the public URLs
rashid reads the local tree and, with --live, probes every asset at the published URL:
rashid check . --live --live-base-url https://s3.lakehousebox.com/<handle>--portland--blobs/portland/ --summary
4 error(s), 0 warning(s), 2 info(s) across 4 files.
error PTL-LIV-002 3x HEAD returns a Content-Length matching the declared file:size
elevation/collection.json: asset 'thumbnail': HEAD Content-Length 218 does not match the declared file:size 12186
elevation/dem-2021/dem-2021.json: asset 'thumbnail': HEAD Content-Length 218 does not match the declared file:size 12186
stops/collection.json: asset 'thumbnail': HEAD Content-Length 218 does not match the declared file:size 18384
error PTL-COL-001 1x a single-file collection must expose its data as a collection-level asset
Against the baseline of step 2, hosting added three errors, all from the thumbnails, which are not served publicly yet. Nothing else: the GeoParquet and the COG passed the size, checksum and format checks, and the range and CORS probes passed. A copy of the tree downloaded from the public URLs, checked with --schema and --live, gave the same four errors.
6. Each vector collection as an Iceberg table (optional)
lhbox table import reads the published GeoParquet over HTTPS and writes it into an Iceberg table, with geometry as a typed column (format version 3) and the rows in Hilbert order (the output is JSON; an excerpt):
lhbox table import portland.transit.stops https://s3.lakehousebox.com/<handle>--portland--blobs/portland/stops/stops.parquet --format-version 3
importing 1 file(s) into transit.stops as ONE commit (api_create_then_insert, rows ordered by hilbert(geometry))…
"rows": 6316,
"format_version": 3,
"geometry_columns": ["geometry"],
"order_by": "hilbert(geometry)",
The STAC data asset stays the GeoParquet file you published; the table is a second copy for querying. Query it by name with lhbox duckdb --catalog portland (the stops within 1,000 feet of stop 2; the data is in EPSG:2913, in feet):
INSTALL spatial; LOAD spatial;
WITH s AS (SELECT geometry AS g FROM portland.transit.stops WHERE stop_id = 2)
SELECT t.stop_id, t.stop_name, round(ST_Distance(t.geometry, s.g)) AS feet
FROM portland.transit.stops t, s
WHERE t.geometry && ST_Buffer(s.g, 1000) AND ST_DWithin(t.geometry, s.g, 1000)
ORDER BY feet;
┌─────────┬──────────────────┬────────┐
│ stop_id │ stop_name │ feet │
│ int32 │ varchar │ double │
├─────────┼──────────────────┼────────┤
│ 2 │ A Ave & Chandler │ 0.0 │
│ 4 │ A Ave & 10th St │ 138.0 │
│ 6 │ A Ave & 8th St │ 683.0 │
│ 7 │ A Ave & 8th St │ 771.0 │
└─────────┴──────────────────┴────────┘
Write spatial filters as geometry && <shape> AND <exact test>: DuckDB skips files and row groups for && only (performance).
With no account, anyone reads the table by its root, and the GeoParquet by its URL:
INSTALL iceberg; LOAD iceberg; INSTALL httpfs; LOAD httpfs;
CREATE SECRET lhbox_public (TYPE S3, ENDPOINT 's3.lakehousebox.com', URL_STYLE 'path', USE_SSL true);
SELECT count(*) AS stops FROM iceberg_scan('s3://<handle>--portland/transit/stops');
SELECT count(*) AS stops FROM read_parquet('https://s3.lakehousebox.com/<handle>--portland--blobs/portland/stops/stops.parquet');
SELECT jurisdic, count(*) AS stops FROM iceberg_scan('s3://<handle>--portland/transit/stops') GROUP BY ALL ORDER BY stops DESC LIMIT 5;
┌───────┐
│ stops │
│ int64 │
├───────┤
│ 6316 │
└───────┘
┌───────┐
│ stops │
│ int64 │
├───────┤
│ 6316 │
└───────┘
┌────────────────┬───────┐
│ jurisdic │ stops │
│ varchar │ int64 │
├────────────────┼───────┤
│ Portland │ 3394 │
│ Beaverton │ 431 │
│ Clackamas Co. │ 403 │
│ Gresham │ 385 │
│ Washington Co. │ 320 │
└────────────────┴───────┘
The keyless secret is needed for iceberg_scan, even given the full https:// metadata URL: the table's metadata names its files as s3:// paths, and without the secret DuckDB looked for them at Amazon (NoSuchBucket).
Point the STAC collection at the table
Portolan's STAC Iceberg extension adds iceberg: fields to a collection. No tool writes them for you yet; the run added these to stops/collection.json by hand, with https://portolan-sdi.github.io/stac-iceberg-extension/v1.0.0/schema.json in stac_extensions:
"iceberg:catalog_type": "rest",
"iceberg:catalog_uri": "https://catalog.lakehousebox.com",
"iceberg:rest_prefix": "<handle>--portland",
"iceberg:authorization_type": "oauth2",
"iceberg:table_id": "transit.stops",
"iceberg:metadata_location": "https://s3.lakehousebox.com/<handle>--portland/transit/stops/metadata/v4.metadata.json",
"iceberg:format_version": 3
lhbox catalog public-url portland --table transit.stops prints the current metadata_location. A second portolan push did not upload the edited collection.json (no new version), so upload it yourself:
aws s3 cp stops/collection.json s3://<handle>--portland--blobs/portland/stops/collection.json \
--endpoint-url https://s3.lakehousebox.com --content-type application/json
rashid (--schema) reported nothing new for the extra fields. An agent with no account then went from the collection to the rows: it read iceberg:metadata_location from collection.json and passed it to iceberg_scan with the keyless secret above, and got 6,316 rows.
Cloud-Optimized GeoTIFFs
A COG goes in the file bucket like any other file and is served with range requests, so GDAL reads only the parts it needs. Reading the published DEM with rasterio (GDAL /vsicurl/), no account:
import rasterio, time
url = "https://s3.lakehousebox.com/<handle>--portland--blobs/portland/elevation/dem-2021/portland-dem.tif"
t = time.time()
with rasterio.Env(GDAL_DISABLE_READDIR_ON_OPEN="EMPTY_DIR", CPL_VSIL_CURL_ALLOWED_EXTENSIONS=".tif"):
with rasterio.open(url) as d:
print("size", d.width, "x", d.height, "crs", d.crs, "overviews", d.overviews(1), "blocks", d.block_shapes)
for name, lon, lat in [("Pioneer Courthouse Square", -122.6794, 45.5189), ("Council Crest", -122.7083, 45.4985)]:
print(name, round(next(d.sample([(lon, lat)]))[0], 1), "m")
print("seconds", round(time.time() - t, 2))
size 1440 x 720 crs EPSG:4326 overviews [2] blocks [(512, 512)]
Pioneer Courthouse Square 14.2 m
Council Crest 325.0 m
seconds 0.73
Rasters stay files: there is no Iceberg table for a COG.
AGENTS.md
Portolan wants an AGENTS.md next to catalog.json and next to each collection.json. They are ordinary files in the published folder, so portolan push uploads them and they are served like the rest (AGENTS.md answered 200 above). Replace portolan-cli's template with real content: where the files are, the CRS and its units, how to read the Iceberg table, the && pattern.
LakehouseBox also shows one AGENTS.md per catalog on its public page and gives it to agents that open the catalog. Set it to the same text:
lhbox catalog agents-md portland --set AGENTS.md
Wrote AGENTS.md to catalog portland: 1394 bytes, sha256 e2faa423d0c6593589232fc1d0dbba905a158e7bbc82140a20cb442d0db8dcd6.
Good practice
- Publish one folder that holds only the catalog (
--files portland/): everything later written there is public. - Keep the STAC
dataasset on the GeoParquet file, not on the table's data files. Files that DuckDB writes into an Iceberg table carry native geometry but no GeoParquetgeokey, and rashid skips them as plain Parquet. - Import vector collections with
--format-version 3, so geometry is a typed column with its CRS, and filter with&&. - When the GeoParquet changes, the table does not follow by itself: it is a copy made at import. Load the new version into the table too, and update
iceberg:metadata_locationif you link it.
Limits and gotchas
- Thumbnails are not served publicly yet. Photos (
.jpg,.png,.webp) in a public folder answer 403 to a reader without an account, unless your organisation is approved to serve them. Portolan requires a thumbnail per collection, so rashid--livereports onePTL-LIV-002per thumbnail and the Portolan browser shows none. To be approved, write to hello@lakehousebox.com with the catalog's name. The approved path was not part of this run. - SLD styles (
.sld.xml) are not served publicly (403 in the run); JSON styles are (200). PMTiles are served (200). - SVG files cannot be uploaded (415
UnsupportedFileType), so aportolan init --logowith an SVG logo cannot be pushed; use a PNG, with the thumbnail caveat above. The list of accepted files is on Files. - portolan-cli and LakehouseBox want different DuckDB versions: keep portolan-cli in its own environment.
- The table is a copy. A collection imported as a table is stored twice, and both copies count against the public allowance.
- No tool writes the
iceberg:block yet, andmetadata_locationchanges with every commit to the table. The table root (s3://<handle>--portland/transit/stops) stays the same. - Format version 3 tables are not compacted yet (snapshot expiry and orphan cleanup do run). A table you import once and append to rarely does not need it.
- PyIceberg 0.12 cannot load a geometry column with a non-default CRS such as EPSG:2913: read those tables with DuckDB.
Next
- Public catalogs: what publishing makes readable, and the allowances.
- Files: the file types a catalog's file bucket takes and serves.
- Importing files:
lhbox table importin detail. - DuckDB and performance: geometry,
&&and how tables are laid out.