BigQuery
| Status | not possible today |
|---|---|
| Reads | not possible today |
| Writes | not possible today |
| Iceberg v3 | not possible today |
| Geometry, geography | not possible today |
| Last verified | never run against the service |
Google BigQuery cannot query a LakehouseBox catalog directly today. To do so it would need either to talk to our Iceberg REST catalog or to read the table files from our S3-compatible storage. Google's documentation offers neither for a provider outside its own list. We have not run BigQuery against the service. What follows is what Google's documentation says (read on 2026-10-04), not a run of ours.
What blocks it
BigQuery has four ways to use Iceberg tables that live outside its own storage. Each one is limited to storage or catalogs Google names:
- Iceberg managed tables (BigQuery-managed, formerly "BigLake tables for Apache Iceberg in BigQuery") keep their data in a Cloud Storage bucket: the storage URI is "a fully qualified Cloud Storage URI. For example,
gs://mybucket/table" (Apache Iceberg managed tables). Google describes these tables as BigQuery-managed, with their files in Cloud Storage. - Iceberg external tables (pointed at one
metadata.jsonfile) take ags://URI, or an Amazon S3 or Azure Blob Storage URI (Create Apache Iceberg external tables). Google says these tables "are no longer recommended for most use cases", and they are read-only. s3://data goes through BigQuery Omni. Omni supports Amazon S3 and Azure Blob Storage, in six AWS regions and one Azure region (Introduction to BigQuery Omni). An Omni connection to S3 is an AWS IAM role, which BigQuery assumes withsts:AssumeRoleWithWebIdentity(Connect to Amazon S3). The documentation has no field for a storage endpoint or for an access key pair, so there is nowhere to puthttps://s3.lakehousebox.comor a LakehouseBox key.- Catalog federation (Preview) reads Iceberg tables from remote catalogs, but only from a fixed list of providers: "Databricks Unity Catalog, AWS Glue, Snowflake Horizon Catalog, Workday Data Lake, SAP Business Data Cloud" (About cross-cloud data access). The catalog options in Google's API have one entry per provider (Glue, Unity, Snowflake), and none for a general Iceberg REST endpoint (FederatedCatalogOptions). A federated catalog is also read-only: "Resource manipulation (such as creating, updating, or deleting resources) are not supported" (Set up cross-cloud Lakehouse for AWS Glue).
Google's own Iceberg REST catalog (the Lakehouse runtime catalog, formerly BigLake metastore) is a catalog other engines connect to. It does not connect BigQuery to someone else's REST catalog.
What would change this
Any of these, from Google: catalog federation accepting a general Iceberg REST catalog (an endpoint and OAuth client credentials), or external tables and Omni accepting an S3-compatible endpoint with a key pair. We will re-check this page against Google's documentation and run it the day one appears.
Getting the data into BigQuery today
None of these paths has been run against LakehouseBox:
- Read with another engine, load into BigQuery. DuckDB, Spark, Trino and PyIceberg read LakehouseBox tables (their pages say which versions were run). An engine can write the result to Parquet in a Cloud Storage bucket, and BigQuery loads Parquet from Cloud Storage. That gives BigQuery a copy, not the live table.
- Spark on Google Cloud. The Spark recipe uses Apache Iceberg's own REST catalog client, and nothing in it names where Spark runs, but we have not tried it on Google's managed Spark (Dataproc).
- Copying the files with Storage Transfer Service is not a way in. Storage Transfer Service can copy from S3-compatible storage into Cloud Storage (it needs a transfer agent, an access key pair and the endpoint; Transfer from S3-compatible sources). But an Iceberg table's metadata records the full
s3://path of every file, so the copied metadata in Cloud Storage would still name the files at LakehouseBox, not the copies.