Apache Spark

Statusverified
Readsyes
Writesyes
Iceberg v3reads and writes
Geometry, geographyreads and writes values; spatial functions need Apache Sedona
Last verified2026-10-03 (Spark 3.5 + Iceberg 1.11; Spark 4.1 + Iceberg 1.12)

Connect

spark.sql.catalog.<catalog_name>=org.apache.iceberg.spark.SparkCatalog
spark.sql.catalog.<catalog_name>.type=rest
spark.sql.catalog.<catalog_name>.uri=https://catalog.lakehousebox.com
spark.sql.catalog.<catalog_name>.warehouse=s3://<handle>--<catalog>/
spark.sql.catalog.<catalog_name>.credential=<client_id>:<client_secret>
spark.sql.catalog.<catalog_name>.io-impl=org.apache.iceberg.aws.s3.S3FileIO
spark.sql.catalog.<catalog_name>.s3.endpoint=https://s3.lakehousebox.com
spark.sql.catalog.<catalog_name>.s3.path-style-access=true
spark.sql.catalog.<catalog_name>.client.region=us-east-1
spark.sql.catalog.<catalog_name>.header.X-Iceberg-Access-Delegation=vended-credentials

Pass these as --conf flags or in spark-defaults.conf; the catalog is then addressed by its name in Spark SQL (SELECT * FROM demo_data.demo.cities for the catalog demo_data). Jars: iceberg-spark-runtime-3.5_2.12 and iceberg-aws-bundle, both 1.11. Java 17 or newer: the Iceberg 1.11 runtime is compiled for Java 17 (class file 61.0) and dies with UnsupportedClassVersionError on the Java 11 that the apache/spark:3.5 images ship; 1.9.2 was the last Java 11 line. client.region=us-east-1 is the SigV4 region the store accepts, not a placement (the data stays in Nuremberg); without it the AWS SDK looks for a region on the JVM and fails before the first request.

Using Spark with LakehouseBox and found something missing, wrong or out of date here? Write to hello@lakehousebox.com: what you ran, the version, and what happened. Product names and logos are trademarks of their owners; their use here does not imply endorsement.