Skip to content

Core: Support client-configured encryption in RESTCatalog - #13225

Open
smaheshwar-pltr wants to merge 20 commits into
apache:mainfrom
smaheshwar-pltr:rest-encrypt-on-main-cherry-pick
Open

smaheshwar-pltr wants to merge 20 commits into
apache:mainfrom
smaheshwar-pltr:rest-encrypt-on-main-cherry-pick

Conversation

@smaheshwar-pltr

@smaheshwar-pltr smaheshwar-pltr commented Jun 3, 2025 •

Copy link
Copy Markdown
Contributor

This PR implements client-side support for REST catalog encryption. With it, clients interacting with a REST catalog can read and write encrypted data.

It is similar to #13066, that integrates encryption with the Hive catalog.

cc @rdblue @RussellSpitzer @ggershinsky

@smaheshwar-pltr
smaheshwar-pltr marked this pull request as ready for review June 3, 2025 21:24
Comment thread core/src/main/java/org/apache/iceberg/rest/RESTTableOperations.java Outdated
@github-actions

github-actions Bot commented Jul 7, 2025

Copy link
Copy Markdown

This pull request has been marked as stale due to 30 days of inactivity. It will be closed in 1 week if no further activity occurs. If you think that’s incorrect or this pull request requires a review, please simply write any comment. If closed, you can revive the PR at any time and @mention a reviewer or discuss it on the dev@iceberg.apache.org list. Thank you for your contributions.

@github-actions github-actions Bot added the stale label Jul 7, 2025
@smaheshwar-pltr

Copy link
Copy Markdown
Contributor Author

(Bump to remove staleness)

@github-actions github-actions Bot removed the stale label Jul 8, 2025
@github-actions

github-actions Bot commented Aug 7, 2025

Copy link
Copy Markdown

This pull request has been marked as stale due to 30 days of inactivity. It will be closed in 1 week if no further activity occurs. If you think that’s incorrect or this pull request requires a review, please simply write any comment. If closed, you can revive the PR at any time and @mention a reviewer or discuss it on the dev@iceberg.apache.org list. Thank you for your contributions.

@github-actions github-actions Bot added the stale label Aug 7, 2025
@smaheshwar-pltr

Copy link
Copy Markdown
Contributor Author

(Bump to remove staleness)

@github-actions github-actions Bot removed the stale label Aug 8, 2025
@github-actions

github-actions Bot commented Sep 8, 2025

Copy link
Copy Markdown

This pull request has been marked as stale due to 30 days of inactivity. It will be closed in 1 week if no further activity occurs. If you think that’s incorrect or this pull request requires a review, please simply write any comment. If closed, you can revive the PR at any time and @mention a reviewer or discuss it on the dev@iceberg.apache.org list. Thank you for your contributions.

@github-actions github-actions Bot added the stale label Sep 8, 2025
@smaheshwar-pltr

Copy link
Copy Markdown
Contributor Author

(Bump to remove staleness)

@smaheshwar-pltr

Copy link
Copy Markdown
Contributor Author

Given #7770 is merged, curious for thoughts on this PR.
REST integration sounds on the cards. Also happy for this PR to be superseded if other folks have worked on it.

cc @huaxingao @RussellSpitzer @ggershinsky @rdblue

@RussellSpitzer

Copy link
Copy Markdown
Member

Could you elaborate on the api additions? I think it would help to have some more description on the general direction of this or

@huaxingao

Copy link
Copy Markdown
Contributor

@smaheshwar-pltr Could you please resolve the conflicts?

@ggershinsky

Copy link
Copy Markdown
Contributor

@huaxingao @smaheshwar-pltr Our team has a person who works on encryption with the REST catalog. If @smaheshwar-pltr does not object, we can follow up on this patch.

return encryptionManager;
}

private void encryptionPropsFromMetadata(TableMetadata metadata) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this method applied on the TableMetadata that is fetched directly from the REST catalog, and not from the metadata.json file? Both are possible, but the former must override (and check) the latter, to protect against the key removal and other attacks.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes this method will always be applied on a metadata field of a LoadTableResponse object received directly from the REST catalog (its only usage within this class is as such, and you can check the constructor usages within RESTSessionCatalog to confirm that the metadata coming in from there is as such too.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Question (not sure if there have been discussions here or if your team have thoughts): we want the key ID to come from the REST catalog service directly for security reasons.

It's typical for REST catalogs to provide metadata that corresponds to the metadata file in storage and not modify it apart from that. Given this, would it be preferable to have this field returned within the LoadTableResponse itself, to encourage catalogs to track it explicitly?

The concrete proposal here might be: ENCRYPTION_TABLE_KEY and ENCRYPTION_DEK_LENGTH become properties on the LoadTableResponse's config (mentioned in the REST spec here).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Well I see two scenarios when thinking about this:

  1. metadata.json is something that both the server and the clients can read (although clients wouldn't need to, given they get the metadata with the LoadTableResponse )
  2. metadata.json can only be accessed on the server side and clients are not given FS credentials (either vended or not) to reach it

For case (1) I totally agree, we can't rely on just metadata.json to store these encryption properties, and the catalog should store it separately too, and eventually populating (i.e. doing the override logic referred by @ggershinsky) the properties in the LoadTableResponse to be created.
For case (2) I'm not 100% sure, but still leaning toward the catalog taking on this responsibility.

Either way, for the client side there's not much we can do other than recommending clients to consider the metadata from LoadTableResponse only. The rest (no pun intended) is on the server side to be decided and will be implementation-specific. For this code snippet above, irrelevant IMHO.

Let me know your thoughts.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Afaik, HMS is not optimized for JSON storage. But maybe someone in the community will take on storing the full metadata object there, to improve table security. Having only the table properties is barely sufficient. I think we should recommend REST catalogs for encrypted tables.

It's not. But would storing a hash suffice? If so we could generate the hash of the whole JSON content and store it via an additional (Hive) table property. Then during table loading we can verify that the TableMetadata we just read in from a potentially untrusted storage (and yet metadata.json is not encrypted) is original or has been tampered with.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One aspect of having an encrypted metadata.json is when the table schema is also considered a sensitive piece of information. I haven't found this in the discussions but do you know if this has ever been considered @ggershinsky ?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would storing a hash suffice? If so we could generate the hash of the whole JSON content and store it via an additional (Hive) table property. Then during table loading we can verify that the TableMetadata we just read in from a potentially untrusted storage (and yet metadata.json is not encrypted) is original or has been tampered with.

I think it's a good idea

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One aspect of having an encrypted metadata.json is when the table schema is also considered a sensitive piece of information. I haven't found this in the discussions but do you know if this has ever been considered

Not sure. Though, it should be possible to have a REST implementation that hides the metadata.json file from the storage.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's a good idea

Sounds good, I can take this on and will produce a PR shortly.

Not sure. Though, it should be possible to have a REST implementation that hides the metadata.json file from the storage.

Yes, with REST that's true, I just meant it in a general sense, e.g. it's not currently possible to hide the schema of an encrypted table with Hive catalog. It may just be one more thing to note/document as a limitation of encryption wrt. Hive catalog - just merely wanted to highlight this though.

@ggershinsky

Copy link
Copy Markdown
Contributor

Also, it would be good to refactor (if possible) a code common to this PR and to #13066 , so that other catalogs will be able to re-use it.

@github-actions github-actions Bot added the build label Oct 27, 2025
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
…ption

Adds read and write support for Iceberg tables whose data files are encrypted at rest with Parquet
Modular Encryption (AES-GCM), under Iceberg's Standard/PME key management. Encryption is transparent
to the SQL user: writing to an encryption-enabled table produces encrypted Parquet, and reading one
decrypts it.

JVM engines get the data-file half of this for free -- OutputFileFactory hands an EncryptedOutputFile
to parquet-java, which applies PME itself. StarRocks cannot, because its Parquet reader and writer are
its own C++ implementation, so the data key has to be generated in the BE, applied through
parquet-cpp, and carried back to the FE for key_metadata. The metadata half -- manifest encryption and
the KEK chain -- StarRocks inherits from Iceberg exactly as Spark does.

Write. The FE decides from the table's EncryptionManager and signals the BE with the algorithm and DEK
length, never key material. The BE draws a fresh per-file DEK from RAND_bytes, attaches
FileEncryptionProperties, and calls disable_aad_prefix_storage() so the AAD prefix is withheld from
the file and a reader must supply it from key_metadata -- that is what binds a file to its identity in
the table and makes whole-file substitution detectable. The DEK returns to the FE at file completion
and is zeroized (OPENSSL_cleanse) when the writer is destroyed.

Read. The FE recovers the DEK by parsing StandardKeyMetadata directly from file.keyMetadata(), not via
EncryptionManager.decrypt(): in 1.11.0 that returns a StandardDecryptedInputFile, which implements
NativeEncryptionInputFile rather than NativelyEncryptedFile and exposes no nativeCryptoParameters(),
so there is no file key to take from it. The BE decrypts the footer, caches an immutable
EncryptionContext, and builds a per-reader page decryptor, since parquet's Decryptor mutates AAD per
page and cannot be shared.

Delete files each carry their own key. Position deletes get key material on
TIcebergDeleteFile.parquet_encryption_info; equality deletes already have their own scan range and
pick it up from the shared buildScanRange path.

Fails closed. A key that cannot be recovered fails the query. An encrypted footer with no DEK is
refused rather than parsed as plaintext. AES_CTR has no PME equivalent and is rejected at planning
instead of being silently written as GCM. An encrypted ORC position-delete file is refused, because
only the Parquet reader can decrypt. The DEK length is never defaulted by the BE: the FE resolves it
from encryption.data-key-length using Iceberg's own constant and default, and encryption enabled with
no length reaching the BE is an error. At commit the returned key is checked against that length, so a
file encrypted below the table's declared strength is refused rather than recorded.

A table that declares encryption while the catalog hands back a plaintext EncryptionManager is refused
on both paths (IcebergEncryption). encryption.key-id is a table property and reaches every catalog,
but the manager comes from the catalog's TableOperations, so the two can disagree -- and a plaintext
manager would otherwise read as "unencrypted table" and commit cleartext into a table configured to
protect it. Both gates test the manager the library returned rather than a version or catalog name, so
they stop firing on their own when a catalog gains encryption support.

Requires Iceberg 1.11.0: key_metadata holds the DEK in the clear, which is safe only because Iceberg
encrypts the manifest carrying it, and that chain -- manifest and manifest-list encryption plus the
encryption-keys list in table metadata -- arrives in 1.11.0. Also builds arrow/parquet-cpp with
-DPARQUET_REQUIRE_ENCRYPTION=ON and installs the PME headers; the symbols were already in
libparquet.a, only the headers were missing.

Scope, set by Iceberg rather than by this change: table format v3 (EncryptionUtil.checkCompatibility
rejects the encryption properties below v3), and Hive Metastore catalogs only -- in 1.11.0 only
HiveTableOperations builds a real EncryptionManager, so REST catalogs report themselves unencrypted
whatever their properties say. apache/iceberg#13225 is the pending fix; nothing here changes when it
lands beyond the dependency bump.

Tests: encrypted write round trip, AAD-prefix withholding, and rejection of an unset and an
out-of-range DEK length (parquet_file_writer_test); encrypted position-delete round trip plus a
negative case asserting a read without the key fails (iceberg_delete_builder_test); key_metadata
round-tripped against Iceberg's own StandardKeyMetadata parser in both directions
(StarRocksKeyMetadataTest); the write signal, the length default, and that the FE never sends key
material (IcebergTableSinkTest); read-path key recovery and fail-closed behaviour
(IcebergConnectorScanRangeSourceTest); refusal at commit of a key shorter than the table's policy
(IcebergMetadataTest); and both declared-vs-enforced gates (IcebergEncryptionTest).

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
…ption

StarRocks cannot read or write Iceberg tables whose data files are encrypted at
rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks.
Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME
key management: a per-file data encryption key encrypts the Parquet file, and the
DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata,
protected by the encrypted manifest that carries it.

JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory
hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO
on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to
be built:

- BE generates the per-file DEK and applies PME through parquet-cpp on write
- the DEK travels back to FE, which serializes it into key_metadata at commit
- on read FE recovers the DEK from key_metadata and ships it per scan range; BE
  decrypts the footer and pages itself

The metadata half -- manifest encryption and the KEK chain -- is inherited from the
Iceberg library exactly as Spark does, since FE planning and commit already go
through it. key_metadata serialization delegates to Iceberg's own (package-private)
StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is
byte-identical to what any other Iceberg engine produces.

All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink
(DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE
signal from one place so they cannot drift. Position-delete files carry their own
per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile
and records its own keyMetadata. An unencrypted delete file beside encrypted data
would disclose which rows were removed, which is what table encryption exists to
prevent.

Fails closed throughout:

- a key that cannot be recovered fails the query; never falls back to plaintext
- an encrypted footer with no DEK is refused rather than parsed as plaintext
- AES_CTR has no PME equivalent and is rejected at planning, not written as GCM
- BE never defaults the DEK length. Only FE can see encryption.data-key-length, so
  FE resolves it using Iceberg's own constant and default; encryption enabled with
  no length reaching BE is an error. At commit the key BE returns is checked against
  that length, so a file encrypted below the table's declared strength is refused
- a table that declares encryption while the catalog supplies a plaintext
  EncryptionManager is refused on both read and write. Those signals come from
  different places -- the property travels with the metadata for every catalog, the
  manager comes from the catalog's TableOperations -- so they can disagree, and a
  plaintext manager would otherwise read as "unencrypted table"
- whether a table is encrypted is decided from encryption.key-id on both the write
  signal and at commit. Deciding it from the manager's type on one side and the
  property on the other would let BE encrypt a file the commit records no key for,
  leaving data nothing can decrypt
- the encrypted page-header length prefix sits outside the AEAD, so it is bounded
  (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB
  cap the plaintext path uses) before it is used to size any buffer

Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds
the DEK in the clear, which is safe only because Iceberg encrypts the manifest
carrying it, and that chain arrives in 1.11.0. Table format v3 is required --
EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore
catalogs only for now: in 1.11.0 only HiveTableOperations builds a real
EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its
properties say (apache/iceberg#13225). Because the decision reads
table.encryption() rather than a version or catalog name, no StarRocks change is
needed when that lands. arrow/parquet-cpp is built with
-DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols
were already in libparquet.a.

Observability: no signal was removed or renamed. Existing Parquet reader/writer
scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and
refusals surface as query errors naming the table and the reason, so no new metric
was added. Key material is never logged or put in a profile.

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
Records the design behind Iceberg table encryption: the trust boundaries and where
key material is allowed to exist, the write and read flows across FE and BE, and
the key-management chain from the table master key down to the per-file DEK.

Written down because the constraints are not visible from the code. Why Iceberg
1.11.0 is a correctness floor rather than a version preference, why table format v3
is required, why the feature is limited to Hive Metastore catalogs today and what
changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check
exists at all -- a table can declare encryption the catalog will not apply, and the
two signals come from different places. Also records the gaps that are deliberately
out of scope, so the next person does not have to rediscover which are intentional.

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
…ption

StarRocks cannot read or write Iceberg tables whose data files are encrypted at
rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks.
Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME
key management: a per-file data encryption key encrypts the Parquet file, and the
DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata,
protected by the encrypted manifest that carries it.

JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory
hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO
on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to
be built:

- BE generates the per-file DEK and applies PME through parquet-cpp on write
- the DEK travels back to FE, which serializes it into key_metadata at commit
- on read FE recovers the DEK from key_metadata and ships it per scan range; BE
  decrypts the footer and pages itself

The metadata half -- manifest encryption and the KEK chain -- is inherited from the
Iceberg library exactly as Spark does, since FE planning and commit already go
through it. key_metadata serialization delegates to Iceberg's own (package-private)
StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is
byte-identical to what any other Iceberg engine produces.

All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink
(DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE
signal from one place so they cannot drift. Position-delete files carry their own
per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile
and records its own keyMetadata. An unencrypted delete file beside encrypted data
would disclose which rows were removed, which is what table encryption exists to
prevent.

Fails closed throughout:

- a key that cannot be recovered fails the query; never falls back to plaintext
- an encrypted footer with no DEK is refused rather than parsed as plaintext
- AES_CTR has no PME equivalent and is rejected at planning, not written as GCM
- BE never defaults the DEK length. Only FE can see encryption.data-key-length, so
  FE resolves it using Iceberg's own constant and default; encryption enabled with
  no length reaching BE is an error. At commit the key BE returns is checked against
  that length, so a file encrypted below the table's declared strength is refused
- a table that declares encryption while the catalog supplies a plaintext
  EncryptionManager is refused on both read and write. Those signals come from
  different places -- the property travels with the metadata for every catalog, the
  manager comes from the catalog's TableOperations -- so they can disagree, and a
  plaintext manager would otherwise read as "unencrypted table"
- whether a table is encrypted is decided from encryption.key-id on both the write
  signal and at commit. Deciding it from the manager's type on one side and the
  property on the other would let BE encrypt a file the commit records no key for,
  leaving data nothing can decrypt
- the encrypted page-header length prefix sits outside the AEAD, so it is bounded
  (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB
  cap the plaintext path uses) before it is used to size any buffer

Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds
the DEK in the clear, which is safe only because Iceberg encrypts the manifest
carrying it, and that chain arrives in 1.11.0. Table format v3 is required --
EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore
catalogs only for now: in 1.11.0 only HiveTableOperations builds a real
EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its
properties say (apache/iceberg#13225). Because the decision reads
table.encryption() rather than a version or catalog name, no StarRocks change is
needed when that lands. arrow/parquet-cpp is built with
-DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols
were already in libparquet.a.

Observability: no signal was removed or renamed. Existing Parquet reader/writer
scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and
refusals surface as query errors naming the table and the reason, so no new metric
was added. Key material is never logged or put in a profile.

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
Records the design behind Iceberg table encryption: the trust boundaries and where
key material is allowed to exist, the write and read flows across FE and BE, and
the key-management chain from the table master key down to the per-file DEK.

Written down because the constraints are not visible from the code. Why Iceberg
1.11.0 is a correctness floor rather than a version preference, why table format v3
is required, why the feature is limited to Hive Metastore catalogs today and what
changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check
exists at all -- a table can declare encryption the catalog will not apply, and the
two signals come from different places. Also records the gaps that are deliberately
out of scope, so the next person does not have to rediscover which are intentional.

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
Records the design behind Iceberg table encryption: the trust boundaries and where
key material is allowed to exist, the write and read flows across FE and BE, and
the key-management chain from the table master key down to the per-file DEK.

Written down because the constraints are not visible from the code. Why Iceberg
1.11.0 is a correctness floor rather than a version preference, why table format v3
is required, why the feature is limited to Hive Metastore catalogs today and what
changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check
exists at all -- a table can declare encryption the catalog will not apply, and the
two signals come from different places. Also records the gaps that are deliberately
out of scope, so the next person does not have to rediscover which are intentional.

The diagrams ship with their PlantUML sources and a render script, because they had
already gone stale once: a PNG kept asserting behaviour the code no longer had, and
with no source in tree there was nothing to check it against. Re-rendering the two
untouched diagrams reproduces their committed PNGs byte for byte, so the sources and
the images are known to agree.

Corrects four claims that no longer matched the code:

- the write path decides from the table's declaration (encryption.key-id), not from
  the EncryptionManager's type, and all three sinks derive that one signal
- the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table
  property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core,
  so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data.
  The BE rejects an unknown cipher; it is not a planning-time check
- the read gate runs in buildScanRange ahead of any per-file test, not inside
  buildEncryptionInfo, because such a table's files carry no key metadata and a
  per-file check would skip them
- REST catalogs: upstream fails silently, but StarRocks refuses rather than
  inheriting that silence

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
…ption

StarRocks cannot read or write Iceberg tables whose data files are encrypted at
rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks.
Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME
key management: a per-file data encryption key encrypts the Parquet file, and the
DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata,
protected by the encrypted manifest that carries it.

JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory
hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO
on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to
be built:

- BE generates the per-file DEK and applies PME through parquet-cpp on write
- the DEK travels back to FE, which serializes it into key_metadata at commit
- on read FE recovers the DEK from key_metadata and ships it per scan range; BE
  decrypts the footer and pages itself

The metadata half -- manifest encryption and the KEK chain -- is inherited from the
Iceberg library exactly as Spark does, since FE planning and commit already go
through it. key_metadata serialization delegates to Iceberg's own (package-private)
StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is
byte-identical to what any other Iceberg engine produces.

All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink
(DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE
signal from one place so they cannot drift. Position-delete files carry their own
per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile
and records its own keyMetadata. An unencrypted delete file beside encrypted data
would disclose which rows were removed, which is what table encryption exists to
prevent.

Fails closed throughout:

- a key that cannot be recovered fails the query; never falls back to plaintext
- an encrypted footer with no DEK is refused rather than parsed as plaintext
- AES_CTR has no PME equivalent and is rejected at planning, not written as GCM
- BE never defaults the DEK length. Only FE can see encryption.data-key-length, so
  FE resolves it using Iceberg's own constant and default; encryption enabled with
  no length reaching BE is an error. At commit the key BE returns is checked against
  that length, so a file encrypted below the table's declared strength is refused
- a table that declares encryption while the catalog supplies a plaintext
  EncryptionManager is refused on both read and write. Those signals come from
  different places -- the property travels with the metadata for every catalog, the
  manager comes from the catalog's TableOperations -- so they can disagree, and a
  plaintext manager would otherwise read as "unencrypted table"
- whether a table is encrypted is decided from encryption.key-id on both the write
  signal and at commit. Deciding it from the manager's type on one side and the
  property on the other would let BE encrypt a file the commit records no key for,
  leaving data nothing can decrypt
- the encrypted page-header length prefix sits outside the AEAD, so it is bounded
  (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB
  cap the plaintext path uses) before it is used to size any buffer

Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds
the DEK in the clear, which is safe only because Iceberg encrypts the manifest
carrying it, and that chain arrives in 1.11.0. Table format v3 is required --
EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore
catalogs only for now: in 1.11.0 only HiveTableOperations builds a real
EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its
properties say (apache/iceberg#13225). Because the decision reads
table.encryption() rather than a version or catalog name, no StarRocks change is
needed when that lands. arrow/parquet-cpp is built with
-DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols
were already in libparquet.a.

Observability: no signal was removed or renamed. Existing Parquet reader/writer
scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and
refusals surface as query errors naming the table and the reason, so no new metric
was added. Key material is never logged or put in a profile.

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
Records the design behind Iceberg table encryption: the trust boundaries and where
key material is allowed to exist, the write and read flows across FE and BE, and
the key-management chain from the table master key down to the per-file DEK.

Written down because the constraints are not visible from the code. Why Iceberg
1.11.0 is a correctness floor rather than a version preference, why table format v3
is required, why the feature is limited to Hive Metastore catalogs today and what
changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check
exists at all -- a table can declare encryption the catalog will not apply, and the
two signals come from different places. Also records the gaps that are deliberately
out of scope, so the next person does not have to rediscover which are intentional.

The diagrams ship with their PlantUML sources and a render script, because they had
already gone stale once: a PNG kept asserting behaviour the code no longer had, and
with no source in tree there was nothing to check it against. Re-rendering the two
untouched diagrams reproduces their committed PNGs byte for byte, so the sources and
the images are known to agree.

Corrects four claims that no longer matched the code:

- the write path decides from the table's declaration (encryption.key-id), not from
  the EncryptionManager's type, and all three sinks derive that one signal
- the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table
  property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core,
  so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data.
  The BE rejects an unknown cipher; it is not a planning-time check
- the read gate runs in buildScanRange ahead of any per-file test, not inside
  buildEncryptionInfo, because such a table's files carry no key metadata and a
  per-file check would skip them
- REST catalogs: upstream fails silently, but StarRocks refuses rather than
  inheriting that silence

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
…ption

StarRocks cannot read or write Iceberg tables whose data files are encrypted at
rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks.
Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME
key management: a per-file data encryption key encrypts the Parquet file, and the
DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata,
protected by the encrypted manifest that carries it.

JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory
hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO
on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to
be built:

- BE generates the per-file DEK and applies PME through parquet-cpp on write
- the DEK travels back to FE, which serializes it into key_metadata at commit
- on read FE recovers the DEK from key_metadata and ships it per scan range; BE
  decrypts the footer and pages itself

The metadata half -- manifest encryption and the KEK chain -- is inherited from the
Iceberg library exactly as Spark does, since FE planning and commit already go
through it. key_metadata serialization delegates to Iceberg's own (package-private)
StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is
byte-identical to what any other Iceberg engine produces.

All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink
(DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE
signal from one place so they cannot drift. Position-delete files carry their own
per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile
and records its own keyMetadata. An unencrypted delete file beside encrypted data
would disclose which rows were removed, which is what table encryption exists to
prevent.

Fails closed throughout:

- a key that cannot be recovered fails the query; never falls back to plaintext
- an encrypted footer with no DEK is refused rather than parsed as plaintext
- AES_CTR has no PME equivalent and is rejected at planning, not written as GCM
- BE never defaults the DEK length. Only FE can see encryption.data-key-length, so
  FE resolves it using Iceberg's own constant and default; encryption enabled with
  no length reaching BE is an error. At commit the key BE returns is checked against
  that length, so a file encrypted below the table's declared strength is refused
- a table that declares encryption while the catalog supplies a plaintext
  EncryptionManager is refused on both read and write. Those signals come from
  different places -- the property travels with the metadata for every catalog, the
  manager comes from the catalog's TableOperations -- so they can disagree, and a
  plaintext manager would otherwise read as "unencrypted table"
- whether a table is encrypted is decided from encryption.key-id on both the write
  signal and at commit. Deciding it from the manager's type on one side and the
  property on the other would let BE encrypt a file the commit records no key for,
  leaving data nothing can decrypt
- the encrypted page-header length prefix sits outside the AEAD, so it is bounded
  (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB
  cap the plaintext path uses) before it is used to size any buffer

Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds
the DEK in the clear, which is safe only because Iceberg encrypts the manifest
carrying it, and that chain arrives in 1.11.0. Table format v3 is required --
EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore
catalogs only for now: in 1.11.0 only HiveTableOperations builds a real
EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its
properties say (apache/iceberg#13225). Because the decision reads
table.encryption() rather than a version or catalog name, no StarRocks change is
needed when that lands. arrow/parquet-cpp is built with
-DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols
were already in libparquet.a.

Observability: no signal was removed or renamed. Existing Parquet reader/writer
scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and
refusals surface as query errors naming the table and the reason, so no new metric
was added. Key material is never logged or put in a profile.

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
Records the design behind Iceberg table encryption: the trust boundaries and where
key material is allowed to exist, the write and read flows across FE and BE, and
the key-management chain from the table master key down to the per-file DEK.

Written down because the constraints are not visible from the code. Why Iceberg
1.11.0 is a correctness floor rather than a version preference, why table format v3
is required, why the feature is limited to Hive Metastore catalogs today and what
changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check
exists at all -- a table can declare encryption the catalog will not apply, and the
two signals come from different places. Also records the gaps that are deliberately
out of scope, so the next person does not have to rediscover which are intentional.

The diagrams ship with their PlantUML sources and a render script, because they had
already gone stale once: a PNG kept asserting behaviour the code no longer had, and
with no source in tree there was nothing to check it against. Re-rendering the two
untouched diagrams reproduces their committed PNGs byte for byte, so the sources and
the images are known to agree.

Corrects four claims that no longer matched the code:

- the write path decides from the table's declaration (encryption.key-id), not from
  the EncryptionManager's type, and all three sinks derive that one signal
- the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table
  property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core,
  so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data.
  The BE rejects an unknown cipher; it is not a planning-time check
- the read gate runs in buildScanRange ahead of any per-file test, not inside
  buildEncryptionInfo, because such a table's files carry no key metadata and a
  per-file check would skip them
- REST catalogs: upstream fails silently, but StarRocks refuses rather than
  inheriting that silence

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
…ption

StarRocks cannot read or write Iceberg tables whose data files are encrypted at
rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks.
Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME
key management: a per-file data encryption key encrypts the Parquet file, and the
DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata,
protected by the encrypted manifest that carries it.

JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory
hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO
on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to
be built:

- BE generates the per-file DEK and applies PME through parquet-cpp on write
- the DEK travels back to FE, which serializes it into key_metadata at commit
- on read FE recovers the DEK from key_metadata and ships it per scan range; BE
  decrypts the footer and pages itself

The metadata half -- manifest encryption and the KEK chain -- is inherited from the
Iceberg library exactly as Spark does, since FE planning and commit already go
through it. key_metadata serialization delegates to Iceberg's own (package-private)
StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is
byte-identical to what any other Iceberg engine produces.

All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink
(DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE
signal from one place so they cannot drift. Position-delete files carry their own
per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile
and records its own keyMetadata. An unencrypted delete file beside encrypted data
would disclose which rows were removed, which is what table encryption exists to
prevent.

Fails closed throughout:

- a key that cannot be recovered fails the query; never falls back to plaintext
- an encrypted footer with no DEK is refused rather than parsed as plaintext
- AES_CTR has no PME equivalent and is rejected at planning, not written as GCM
- BE never defaults the DEK length. Only FE can see encryption.data-key-length, so
  FE resolves it using Iceberg's own constant and default; encryption enabled with
  no length reaching BE is an error. At commit the key BE returns is checked against
  that length, so a file encrypted below the table's declared strength is refused
- a table that declares encryption while the catalog supplies a plaintext
  EncryptionManager is refused on both read and write. Those signals come from
  different places -- the property travels with the metadata for every catalog, the
  manager comes from the catalog's TableOperations -- so they can disagree, and a
  plaintext manager would otherwise read as "unencrypted table"
- whether a table is encrypted is decided from encryption.key-id on both the write
  signal and at commit. Deciding it from the manager's type on one side and the
  property on the other would let BE encrypt a file the commit records no key for,
  leaving data nothing can decrypt
- the encrypted page-header length prefix sits outside the AEAD, so it is bounded
  (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB
  cap the plaintext path uses) before it is used to size any buffer

Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds
the DEK in the clear, which is safe only because Iceberg encrypts the manifest
carrying it, and that chain arrives in 1.11.0. Table format v3 is required --
EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore
catalogs only for now: in 1.11.0 only HiveTableOperations builds a real
EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its
properties say (apache/iceberg#13225). Because the decision reads
table.encryption() rather than a version or catalog name, no StarRocks change is
needed when that lands. arrow/parquet-cpp is built with
-DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols
were already in libparquet.a.

Observability: no signal was removed or renamed. Existing Parquet reader/writer
scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and
refusals surface as query errors naming the table and the reason, so no new metric
was added. Key material is never logged or put in a profile.

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
Records the design behind Iceberg table encryption: the trust boundaries and where
key material is allowed to exist, the write and read flows across FE and BE, and
the key-management chain from the table master key down to the per-file DEK.

Written down because the constraints are not visible from the code. Why Iceberg
1.11.0 is a correctness floor rather than a version preference, why table format v3
is required, why the feature is limited to Hive Metastore catalogs today and what
changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check
exists at all -- a table can declare encryption the catalog will not apply, and the
two signals come from different places. Also records the gaps that are deliberately
out of scope, so the next person does not have to rediscover which are intentional.

The diagrams ship with their PlantUML sources and a render script, because they had
already gone stale once: a PNG kept asserting behaviour the code no longer had, and
with no source in tree there was nothing to check it against. Re-rendering the two
untouched diagrams reproduces their committed PNGs byte for byte, so the sources and
the images are known to agree.

Corrects four claims that no longer matched the code:

- the write path decides from the table's declaration (encryption.key-id), not from
  the EncryptionManager's type, and all three sinks derive that one signal
- the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table
  property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core,
  so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data.
  The BE rejects an unknown cipher; it is not a planning-time check
- the read gate runs in buildScanRange ahead of any per-file test, not inside
  buildEncryptionInfo, because such a table's files carry no key metadata and a
  per-file check would skip them
- REST catalogs: upstream fails silently, but StarRocks refuses rather than
  inheriting that silence

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
…ption

StarRocks cannot read or write Iceberg tables whose data files are encrypted at
rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks.
Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME
key management: a per-file data encryption key encrypts the Parquet file, and the
DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata,
protected by the encrypted manifest that carries it.

JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory
hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO
on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to
be built:

- BE generates the per-file DEK and applies PME through parquet-cpp on write
- the DEK travels back to FE, which serializes it into key_metadata at commit
- on read FE recovers the DEK from key_metadata and ships it per scan range; BE
  decrypts the footer and pages itself

The metadata half -- manifest encryption and the KEK chain -- is inherited from the
Iceberg library exactly as Spark does, since FE planning and commit already go
through it. key_metadata serialization delegates to Iceberg's own (package-private)
StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is
byte-identical to what any other Iceberg engine produces.

All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink
(DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE
signal from one place so they cannot drift. Position-delete files carry their own
per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile
and records its own keyMetadata. An unencrypted delete file beside encrypted data
would disclose which rows were removed, which is what table encryption exists to
prevent.

Fails closed throughout:

- a key that cannot be recovered fails the query; never falls back to plaintext
- an encrypted footer with no DEK is refused rather than parsed as plaintext
- AES_CTR has no PME equivalent and is rejected at planning, not written as GCM
- BE never defaults the DEK length. Only FE can see encryption.data-key-length, so
  FE resolves it using Iceberg's own constant and default; encryption enabled with
  no length reaching BE is an error. At commit the key BE returns is checked against
  that length, so a file encrypted below the table's declared strength is refused
- a table that declares encryption while the catalog supplies a plaintext
  EncryptionManager is refused on both read and write. Those signals come from
  different places -- the property travels with the metadata for every catalog, the
  manager comes from the catalog's TableOperations -- so they can disagree, and a
  plaintext manager would otherwise read as "unencrypted table"
- whether a table is encrypted is decided from encryption.key-id on both the write
  signal and at commit. Deciding it from the manager's type on one side and the
  property on the other would let BE encrypt a file the commit records no key for,
  leaving data nothing can decrypt
- the encrypted page-header length prefix sits outside the AEAD, so it is bounded
  (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB
  cap the plaintext path uses) before it is used to size any buffer

Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds
the DEK in the clear, which is safe only because Iceberg encrypts the manifest
carrying it, and that chain arrives in 1.11.0. Table format v3 is required --
EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore
catalogs only for now: in 1.11.0 only HiveTableOperations builds a real
EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its
properties say (apache/iceberg#13225). Because the decision reads
table.encryption() rather than a version or catalog name, no StarRocks change is
needed when that lands. arrow/parquet-cpp is built with
-DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols
were already in libparquet.a.

Observability: no signal was removed or renamed. Existing Parquet reader/writer
scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and
refusals surface as query errors naming the table and the reason, so no new metric
was added. Key material is never logged or put in a profile.

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
Records the design behind Iceberg table encryption: the trust boundaries and where
key material is allowed to exist, the write and read flows across FE and BE, and
the key-management chain from the table master key down to the per-file DEK.

Written down because the constraints are not visible from the code. Why Iceberg
1.11.0 is a correctness floor rather than a version preference, why table format v3
is required, why the feature is limited to Hive Metastore catalogs today and what
changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check
exists at all -- a table can declare encryption the catalog will not apply, and the
two signals come from different places. Also records the gaps that are deliberately
out of scope, so the next person does not have to rediscover which are intentional.

The diagrams ship with their PlantUML sources and a render script, because they had
already gone stale once: a PNG kept asserting behaviour the code no longer had, and
with no source in tree there was nothing to check it against. Re-rendering the two
untouched diagrams reproduces their committed PNGs byte for byte, so the sources and
the images are known to agree.

Corrects four claims that no longer matched the code:

- the write path decides from the table's declaration (encryption.key-id), not from
  the EncryptionManager's type, and all three sinks derive that one signal
- the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table
  property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core,
  so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data.
  The BE rejects an unknown cipher; it is not a planning-time check
- the read gate runs in buildScanRange ahead of any per-file test, not inside
  buildEncryptionInfo, because such a table's files carry no key metadata and a
  per-file check would skip them
- REST catalogs: upstream fails silently, but StarRocks refuses rather than
  inheriting that silence

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
…ption

StarRocks cannot read or write Iceberg tables whose data files are encrypted at
rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks.
Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME
key management: a per-file data encryption key encrypts the Parquet file, and the
DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata,
protected by the encrypted manifest that carries it.

JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory
hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO
on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to
be built:

- BE generates the per-file DEK and applies PME through parquet-cpp on write
- the DEK travels back to FE, which serializes it into key_metadata at commit
- on read FE recovers the DEK from key_metadata and ships it per scan range; BE
  decrypts the footer and pages itself

The metadata half -- manifest encryption and the KEK chain -- is inherited from the
Iceberg library exactly as Spark does, since FE planning and commit already go
through it. key_metadata serialization delegates to Iceberg's own (package-private)
StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is
byte-identical to what any other Iceberg engine produces.

All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink
(DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE
signal from one place so they cannot drift. Position-delete files carry their own
per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile
and records its own keyMetadata. An unencrypted delete file beside encrypted data
would disclose which rows were removed, which is what table encryption exists to
prevent.

Fails closed throughout:

- a key that cannot be recovered fails the query; never falls back to plaintext
- an encrypted footer with no DEK is refused rather than parsed as plaintext
- AES_CTR has no PME equivalent and is rejected at planning, not written as GCM
- BE never defaults the DEK length. Only FE can see encryption.data-key-length, so
  FE resolves it using Iceberg's own constant and default; encryption enabled with
  no length reaching BE is an error. At commit the key BE returns is checked against
  that length, so a file encrypted below the table's declared strength is refused
- a table that declares encryption while the catalog supplies a plaintext
  EncryptionManager is refused on both read and write. Those signals come from
  different places -- the property travels with the metadata for every catalog, the
  manager comes from the catalog's TableOperations -- so they can disagree, and a
  plaintext manager would otherwise read as "unencrypted table"
- whether a table is encrypted is decided from encryption.key-id on both the write
  signal and at commit. Deciding it from the manager's type on one side and the
  property on the other would let BE encrypt a file the commit records no key for,
  leaving data nothing can decrypt
- the encrypted page-header length prefix sits outside the AEAD, so it is bounded
  (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB
  cap the plaintext path uses) before it is used to size any buffer

Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds
the DEK in the clear, which is safe only because Iceberg encrypts the manifest
carrying it, and that chain arrives in 1.11.0. Table format v3 is required --
EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore
catalogs only for now: in 1.11.0 only HiveTableOperations builds a real
EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its
properties say (apache/iceberg#13225). Because the decision reads
table.encryption() rather than a version or catalog name, no StarRocks change is
needed when that lands. arrow/parquet-cpp is built with
-DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols
were already in libparquet.a.

Observability: no signal was removed or renamed. Existing Parquet reader/writer
scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and
refusals surface as query errors naming the table and the reason, so no new metric
was added. Key material is never logged or put in a profile.

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
Records the design behind Iceberg table encryption: the trust boundaries and where
key material is allowed to exist, the write and read flows across FE and BE, and
the key-management chain from the table master key down to the per-file DEK.

Written down because the constraints are not visible from the code. Why Iceberg
1.11.0 is a correctness floor rather than a version preference, why table format v3
is required, why the feature is limited to Hive Metastore catalogs today and what
changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check
exists at all -- a table can declare encryption the catalog will not apply, and the
two signals come from different places. Also records the gaps that are deliberately
out of scope, so the next person does not have to rediscover which are intentional.

The diagrams ship with their PlantUML sources and a render script, because they had
already gone stale once: a PNG kept asserting behaviour the code no longer had, and
with no source in tree there was nothing to check it against. Re-rendering the two
untouched diagrams reproduces their committed PNGs byte for byte, so the sources and
the images are known to agree.

Corrects four claims that no longer matched the code:

- the write path decides from the table's declaration (encryption.key-id), not from
  the EncryptionManager's type, and all three sinks derive that one signal
- the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table
  property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core,
  so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data.
  The BE rejects an unknown cipher; it is not a planning-time check
- the read gate runs in buildScanRange ahead of any per-file test, not inside
  buildEncryptionInfo, because such a table's files carry no key metadata and a
  per-file check would skip them
- REST catalogs: upstream fails silently, but StarRocks refuses rather than
  inheriting that silence

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
…ption

StarRocks cannot read or write Iceberg tables whose data files are encrypted at
rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks.
Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME
key management: a per-file data encryption key encrypts the Parquet file, and the
DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata,
protected by the encrypted manifest that carries it.

JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory
hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO
on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to
be built:

- BE generates the per-file DEK and applies PME through parquet-cpp on write
- the DEK travels back to FE, which serializes it into key_metadata at commit
- on read FE recovers the DEK from key_metadata and ships it per scan range; BE
  decrypts the footer and pages itself

The metadata half -- manifest encryption and the KEK chain -- is inherited from the
Iceberg library exactly as Spark does, since FE planning and commit already go
through it. key_metadata serialization delegates to Iceberg's own (package-private)
StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is
byte-identical to what any other Iceberg engine produces.

All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink
(DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE
signal from one place so they cannot drift. Position-delete files carry their own
per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile
and records its own keyMetadata. An unencrypted delete file beside encrypted data
would disclose which rows were removed, which is what table encryption exists to
prevent.

Fails closed throughout:

- a key that cannot be recovered fails the query; never falls back to plaintext
- an encrypted footer with no DEK is refused rather than parsed as plaintext
- AES_CTR has no PME equivalent and is rejected at planning, not written as GCM
- BE never defaults the DEK length. Only FE can see encryption.data-key-length, so
  FE resolves it using Iceberg's own constant and default; encryption enabled with
  no length reaching BE is an error. At commit the key BE returns is checked against
  that length, so a file encrypted below the table's declared strength is refused
- a table that declares encryption while the catalog supplies a plaintext
  EncryptionManager is refused on both read and write. Those signals come from
  different places -- the property travels with the metadata for every catalog, the
  manager comes from the catalog's TableOperations -- so they can disagree, and a
  plaintext manager would otherwise read as "unencrypted table"
- whether a table is encrypted is decided from encryption.key-id on both the write
  signal and at commit. Deciding it from the manager's type on one side and the
  property on the other would let BE encrypt a file the commit records no key for,
  leaving data nothing can decrypt
- the encrypted page-header length prefix sits outside the AEAD, so it is bounded
  (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB
  cap the plaintext path uses) before it is used to size any buffer

Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds
the DEK in the clear, which is safe only because Iceberg encrypts the manifest
carrying it, and that chain arrives in 1.11.0. Table format v3 is required --
EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore
catalogs only for now: in 1.11.0 only HiveTableOperations builds a real
EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its
properties say (apache/iceberg#13225). Because the decision reads
table.encryption() rather than a version or catalog name, no StarRocks change is
needed when that lands. arrow/parquet-cpp is built with
-DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols
were already in libparquet.a.

Observability: no signal was removed or renamed. Existing Parquet reader/writer
scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and
refusals surface as query errors naming the table and the reason, so no new metric
was added. Key material is never logged or put in a profile.

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
Records the design behind Iceberg table encryption: the trust boundaries and where
key material is allowed to exist, the write and read flows across FE and BE, and
the key-management chain from the table master key down to the per-file DEK.

Written down because the constraints are not visible from the code. Why Iceberg
1.11.0 is a correctness floor rather than a version preference, why table format v3
is required, why the feature is limited to Hive Metastore catalogs today and what
changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check
exists at all -- a table can declare encryption the catalog will not apply, and the
two signals come from different places. Also records the gaps that are deliberately
out of scope, so the next person does not have to rediscover which are intentional.

The diagrams ship with their PlantUML sources and a render script, because they had
already gone stale once: a PNG kept asserting behaviour the code no longer had, and
with no source in tree there was nothing to check it against. Re-rendering the two
untouched diagrams reproduces their committed PNGs byte for byte, so the sources and
the images are known to agree.

Corrects four claims that no longer matched the code:

- the write path decides from the table's declaration (encryption.key-id), not from
  the EncryptionManager's type, and all three sinks derive that one signal
- the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table
  property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core,
  so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data.
  The BE rejects an unknown cipher; it is not a planning-time check
- the read gate runs in buildScanRange ahead of any per-file test, not inside
  buildEncryptionInfo, because such a table's files carry no key metadata and a
  per-file check would skip them
- REST catalogs: upstream fails silently, but StarRocks refuses rather than
  inheriting that silence

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
…ption

StarRocks cannot read or write Iceberg tables whose data files are encrypted at
rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks.
Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME
key management: a per-file data encryption key encrypts the Parquet file, and the
DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata,
protected by the encrypted manifest that carries it.

JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory
hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO
on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to
be built:

- BE generates the per-file DEK and applies PME through parquet-cpp on write
- the DEK travels back to FE, which serializes it into key_metadata at commit
- on read FE recovers the DEK from key_metadata and ships it per scan range; BE
  decrypts the footer and pages itself

The metadata half -- manifest encryption and the KEK chain -- is inherited from the
Iceberg library exactly as Spark does, since FE planning and commit already go
through it. key_metadata serialization delegates to Iceberg's own (package-private)
StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is
byte-identical to what any other Iceberg engine produces.

All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink
(DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE
signal from one place so they cannot drift. Position-delete files carry their own
per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile
and records its own keyMetadata. An unencrypted delete file beside encrypted data
would disclose which rows were removed, which is what table encryption exists to
prevent.

Fails closed throughout:

- a key that cannot be recovered fails the query; never falls back to plaintext
- an encrypted footer with no DEK is refused rather than parsed as plaintext
- AES_CTR has no PME equivalent and is rejected at planning, not written as GCM
- BE never defaults the DEK length. Only FE can see encryption.data-key-length, so
  FE resolves it using Iceberg's own constant and default; encryption enabled with
  no length reaching BE is an error. At commit the key BE returns is checked against
  that length, so a file encrypted below the table's declared strength is refused
- a table that declares encryption while the catalog supplies a plaintext
  EncryptionManager is refused on both read and write. Those signals come from
  different places -- the property travels with the metadata for every catalog, the
  manager comes from the catalog's TableOperations -- so they can disagree, and a
  plaintext manager would otherwise read as "unencrypted table"
- whether a table is encrypted is decided from encryption.key-id on both the write
  signal and at commit. Deciding it from the manager's type on one side and the
  property on the other would let BE encrypt a file the commit records no key for,
  leaving data nothing can decrypt
- the encrypted page-header length prefix sits outside the AEAD, so it is bounded
  (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB
  cap the plaintext path uses) before it is used to size any buffer

Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds
the DEK in the clear, which is safe only because Iceberg encrypts the manifest
carrying it, and that chain arrives in 1.11.0. Table format v3 is required --
EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore
catalogs only for now: in 1.11.0 only HiveTableOperations builds a real
EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its
properties say (apache/iceberg#13225). Because the decision reads
table.encryption() rather than a version or catalog name, no StarRocks change is
needed when that lands. arrow/parquet-cpp is built with
-DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols
were already in libparquet.a.

Observability: no signal was removed or renamed. Existing Parquet reader/writer
scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and
refusals surface as query errors naming the table and the reason, so no new metric
was added. Key material is never logged or put in a profile.

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 11, 2026
Records the design behind Iceberg table encryption: the trust boundaries and where
key material is allowed to exist, the write and read flows across FE and BE, and
the key-management chain from the table master key down to the per-file DEK.

Written down because the constraints are not visible from the code. Why Iceberg
1.11.0 is a correctness floor rather than a version preference, why table format v3
is required, why the feature is limited to Hive Metastore catalogs today and what
changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check
exists at all -- a table can declare encryption the catalog will not apply, and the
two signals come from different places. Also records the gaps that are deliberately
out of scope, so the next person does not have to rediscover which are intentional.

The diagrams ship with their PlantUML sources and a render script, because they had
already gone stale once: a PNG kept asserting behaviour the code no longer had, and
with no source in tree there was nothing to check it against. Re-rendering the two
untouched diagrams reproduces their committed PNGs byte for byte, so the sources and
the images are known to agree.

Corrects four claims that no longer matched the code:

- the write path decides from the table's declaration (encryption.key-id), not from
  the EncryptionManager's type, and all three sinks derive that one signal
- the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table
  property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core,
  so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data.
  The BE rejects an unknown cipher; it is not a planning-time check
- the read gate runs in buildScanRange ahead of any per-file test, not inside
  buildEncryptionInfo, because such a table's files carry no key metadata and a
  per-file check would skip them
- REST catalogs: upstream fails silently, but StarRocks refuses rather than
  inheriting that silence

Signed-off-by: Annie Yan <anniey@outlook.com>
@smaheshwar-pltr

smaheshwar-pltr commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor Author

@singhpk234 @szlta apologies if I’ve been unclear. I definitely agree it would be good to get something into this release. But I don't understand what's being proposed, it seems very non-trivial to me and not quite acceptable to just decide this without any spec change in.

Let me be concrete:

  • What triggers rejection: advertising storage vending, receiving credentials, or detecting that the delegated credential set is incomplete? Are we supposed to throw, or discard returned credentials and use configured credentials?
  • Recall that storage credentials can arrive through both storage-credentials and LoadTableResponse.config. That config also contains general REST stuff, planning and ordinary FileIO settings, and custom implementations can define their own properties. Do we prohibit any configuration entirely (this is extreme), maintain lists of ignorised credential properties, provider-specific validation? How do catalog-level defaults and overrides fit into this?
  • The restriction also needs to account for remote signing and credentials returned by scan planning. Again, this is non-trivial.
  • What about failure timing? For example, credentials returned by a create/register operation are inspected after the server may already have completed that operation.

I’m struggling to understand the value of this PR with this restrictions and no KMS vending being available to lift them (there's no KMS vending spec in). But I'd argue that hybrid configuration is valuable when KMS authorization is managed externally to the catalog. (I'd also argue that external KMS authorization is not problematic enough to introduce such intrusive enforcements^ but I'm not going to block on this).

I think I'm misunderstanding something because I can't imagine that just deciding all this now in this PR for the next release is good - if there's a straightforward implementation, please feel free to implement it on top of this PR, or supersede this PR if that’s easier! Very happy for that. (My point here is that I'm just not understanding unfortunately)

@singhpk234

singhpk234 commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

But I don't understand what's being proposed

Here is what i was proposing : c8116f3
essentially only enable this hybrid mode when client explicitly sets this as catalog prop, we just need an assertion wired through ... Let me know what do you think of the change ..

I can't imagine that just deciding all this now in this PR for the next release is good

To recap, folks advocated we need this in prev release 1.11 #13225 (comment), but we were not able to make it, it been ~4 months where some of us in the working group, talked through scenarios, came up with a rest spec proposal. I also raised this in community syncs, it was said there can be use cases we should support this in the sdk the hybrid mode. I was searching for use cases.

A couple of folks also reached out to me to see if i can help in getting this in 1.12, because presently engines such as Apache Spark can't work at all with RESTCatalog, I have been trying to follow up in this pr Jul 27 .... I understand totally it takes time !

With that being said, I don't think there is anything wrong in being ambitious to get this in a upcoming release, though we should not be making implementation decision in haste .... lets do the right thing !

Thank you for all the work so far on this !

super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 12, 2026
…ption

StarRocks cannot read or write Iceberg tables whose data files are encrypted at
rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks.
Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME
key management: a per-file data encryption key encrypts the Parquet file, and the
DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata,
protected by the encrypted manifest that carries it.

JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory
hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO
on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to
be built:

- BE generates the per-file DEK and applies PME through parquet-cpp on write
- the DEK travels back to FE, which serializes it into key_metadata at commit
- on read FE recovers the DEK from key_metadata and ships it per scan range; BE
  decrypts the footer and pages itself

The metadata half -- manifest encryption and the KEK chain -- is inherited from the
Iceberg library exactly as Spark does, since FE planning and commit already go
through it. key_metadata serialization delegates to Iceberg's own (package-private)
StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is
byte-identical to what any other Iceberg engine produces.

All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink
(DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE
signal from one place so they cannot drift. Position-delete files carry their own
per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile
and records its own keyMetadata. An unencrypted delete file beside encrypted data
would disclose which rows were removed, which is what table encryption exists to
prevent.

Fails closed throughout:

- a key that cannot be recovered fails the query; never falls back to plaintext
- an encrypted footer with no DEK is refused rather than parsed as plaintext
- AES_CTR has no PME equivalent and is rejected at planning, not written as GCM
- BE never defaults the DEK length. Only FE can see encryption.data-key-length, so
  FE resolves it using Iceberg's own constant and default; encryption enabled with
  no length reaching BE is an error. At commit the key BE returns is checked against
  that length, so a file encrypted below the table's declared strength is refused
- a table that declares encryption while the catalog supplies a plaintext
  EncryptionManager is refused on both read and write. Those signals come from
  different places -- the property travels with the metadata for every catalog, the
  manager comes from the catalog's TableOperations -- so they can disagree, and a
  plaintext manager would otherwise read as "unencrypted table"
- whether a table is encrypted is decided from encryption.key-id on both the write
  signal and at commit. Deciding it from the manager's type on one side and the
  property on the other would let BE encrypt a file the commit records no key for,
  leaving data nothing can decrypt
- the encrypted page-header length prefix sits outside the AEAD, so it is bounded
  (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB
  cap the plaintext path uses) before it is used to size any buffer

Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds
the DEK in the clear, which is safe only because Iceberg encrypts the manifest
carrying it, and that chain arrives in 1.11.0. Table format v3 is required --
EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore
catalogs only for now: in 1.11.0 only HiveTableOperations builds a real
EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its
properties say (apache/iceberg#13225). Because the decision reads
table.encryption() rather than a version or catalog name, no StarRocks change is
needed when that lands. arrow/parquet-cpp is built with
-DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols
were already in libparquet.a.

Observability: no signal was removed or renamed. Existing Parquet reader/writer
scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and
refusals surface as query errors naming the table and the reason, so no new metric
was added. Key material is never logged or put in a profile.

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 12, 2026
Records the design behind Iceberg table encryption: the trust boundaries and where
key material is allowed to exist, the write and read flows across FE and BE, and
the key-management chain from the table master key down to the per-file DEK.

Written down because the constraints are not visible from the code. Why Iceberg
1.11.0 is a correctness floor rather than a version preference, why table format v3
is required, why the feature is limited to Hive Metastore catalogs today and what
changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check
exists at all -- a table can declare encryption the catalog will not apply, and the
two signals come from different places. Also records the gaps that are deliberately
out of scope, so the next person does not have to rediscover which are intentional.

The diagrams ship with their PlantUML sources and a render script, because they had
already gone stale once: a PNG kept asserting behaviour the code no longer had, and
with no source in tree there was nothing to check it against. Re-rendering the two
untouched diagrams reproduces their committed PNGs byte for byte, so the sources and
the images are known to agree.

Corrects four claims that no longer matched the code:

- the write path decides from the table's declaration (encryption.key-id), not from
  the EncryptionManager's type, and all three sinks derive that one signal
- the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table
  property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core,
  so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data.
  The BE rejects an unknown cipher; it is not a planning-time check
- the read gate runs in buildScanRange ahead of any per-file test, not inside
  buildEncryptionInfo, because such a table's files carry no key metadata and a
  per-file check would skip them
- REST catalogs: upstream fails silently, but StarRocks refuses rather than
  inheriting that silence

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 12, 2026
…ption

StarRocks cannot read or write Iceberg tables whose data files are encrypted at
rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks.
Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME
key management: a per-file data encryption key encrypts the Parquet file, and the
DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata,
protected by the encrypted manifest that carries it.

JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory
hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO
on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to
be built:

- BE generates the per-file DEK and applies PME through parquet-cpp on write
- the DEK travels back to FE, which serializes it into key_metadata at commit
- on read FE recovers the DEK from key_metadata and ships it per scan range; BE
  decrypts the footer and pages itself

The metadata half -- manifest encryption and the KEK chain -- is inherited from the
Iceberg library exactly as Spark does, since FE planning and commit already go
through it. key_metadata serialization delegates to Iceberg's own (package-private)
StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is
byte-identical to what any other Iceberg engine produces.

All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink
(DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE
signal from one place so they cannot drift. Position-delete files carry their own
per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile
and records its own keyMetadata. An unencrypted delete file beside encrypted data
would disclose which rows were removed, which is what table encryption exists to
prevent.

Fails closed throughout:

- a key that cannot be recovered fails the query; never falls back to plaintext
- an encrypted footer with no DEK is refused rather than parsed as plaintext
- AES_CTR has no PME equivalent and is rejected at planning, not written as GCM
- BE never defaults the DEK length. Only FE can see encryption.data-key-length, so
  FE resolves it using Iceberg's own constant and default; encryption enabled with
  no length reaching BE is an error. At commit the key BE returns is checked against
  that length, so a file encrypted below the table's declared strength is refused
- a table that declares encryption while the catalog supplies a plaintext
  EncryptionManager is refused on both read and write. Those signals come from
  different places -- the property travels with the metadata for every catalog, the
  manager comes from the catalog's TableOperations -- so they can disagree, and a
  plaintext manager would otherwise read as "unencrypted table"
- whether a table is encrypted is decided from encryption.key-id on both the write
  signal and at commit. Deciding it from the manager's type on one side and the
  property on the other would let BE encrypt a file the commit records no key for,
  leaving data nothing can decrypt
- the encrypted page-header length prefix sits outside the AEAD, so it is bounded
  (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB
  cap the plaintext path uses) before it is used to size any buffer

Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds
the DEK in the clear, which is safe only because Iceberg encrypts the manifest
carrying it, and that chain arrives in 1.11.0. Table format v3 is required --
EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore
catalogs only for now: in 1.11.0 only HiveTableOperations builds a real
EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its
properties say (apache/iceberg#13225). Because the decision reads
table.encryption() rather than a version or catalog name, no StarRocks change is
needed when that lands. arrow/parquet-cpp is built with
-DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols
were already in libparquet.a.

Observability: no signal was removed or renamed. Existing Parquet reader/writer
scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and
refusals surface as query errors naming the table and the reason, so no new metric
was added. Key material is never logged or put in a profile.

Signed-off-by: Annie Yan <anniey@outlook.com>
super-annie added a commit to super-annie/starrocks that referenced this pull request Sep 12, 2026
Records the design behind Iceberg table encryption: the trust boundaries and where
key material is allowed to exist, the write and read flows across FE and BE, and
the key-management chain from the table master key down to the per-file DEK.

Written down because the constraints are not visible from the code. Why Iceberg
1.11.0 is a correctness floor rather than a version preference, why table format v3
is required, why the feature is limited to Hive Metastore catalogs today and what
changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check
exists at all -- a table can declare encryption the catalog will not apply, and the
two signals come from different places. Also records the gaps that are deliberately
out of scope, so the next person does not have to rediscover which are intentional.

The diagrams ship with their PlantUML sources and a render script, because they had
already gone stale once: a PNG kept asserting behaviour the code no longer had, and
with no source in tree there was nothing to check it against. Re-rendering the two
untouched diagrams reproduces their committed PNGs byte for byte, so the sources and
the images are known to agree.

Corrects four claims that no longer matched the code:

- the write path decides from the table's declaration (encryption.key-id), not from
  the EncryptionManager's type, and all three sinks derive that one signal
- the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table
  property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core,
  so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data.
  The BE rejects an unknown cipher; it is not a planning-time check
- the read gate runs in buildScanRange ahead of any per-file test, not inside
  buildEncryptionInfo, because such a table's files carry no key metadata and a
  per-file check would skip them
- REST catalogs: upstream fails silently, but StarRocks refuses rather than
  inheriting that silence

Signed-off-by: Annie Yan <anniey@outlook.com>
@smaheshwar-pltr

Copy link
Copy Markdown
Contributor Author

@singhpk234 thank you for the implementation! I've re-iterated some high-level thoughts in #18081 (review), happy to work on integrating that implementation to this PR (or whatever works) once we're aligned on the approach there - it generally makes sense to me with some discussions I think worth having as I mentioned above - and +1 on not rushing the implementation here. LMKWYT!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.