Core: Support client-configured encryption in RESTCatalog - #13225
smaheshwar-pltr wants to merge 20 commits into
Conversation
|
This pull request has been marked as stale due to 30 days of inactivity. It will be closed in 1 week if no further activity occurs. If you think that’s incorrect or this pull request requires a review, please simply write any comment. If closed, you can revive the PR at any time and @mention a reviewer or discuss it on the dev@iceberg.apache.org list. Thank you for your contributions. |
|
(Bump to remove staleness) |
|
This pull request has been marked as stale due to 30 days of inactivity. It will be closed in 1 week if no further activity occurs. If you think that’s incorrect or this pull request requires a review, please simply write any comment. If closed, you can revive the PR at any time and @mention a reviewer or discuss it on the dev@iceberg.apache.org list. Thank you for your contributions. |
|
(Bump to remove staleness) |
|
This pull request has been marked as stale due to 30 days of inactivity. It will be closed in 1 week if no further activity occurs. If you think that’s incorrect or this pull request requires a review, please simply write any comment. If closed, you can revive the PR at any time and @mention a reviewer or discuss it on the dev@iceberg.apache.org list. Thank you for your contributions. |
|
(Bump to remove staleness) |
d007aa2 to
889cbbf
Compare
|
Given #7770 is merged, curious for thoughts on this PR. |
|
Could you elaborate on the api additions? I think it would help to have some more description on the general direction of this or |
|
@smaheshwar-pltr Could you please resolve the conflicts? |
|
@huaxingao @smaheshwar-pltr Our team has a person who works on encryption with the REST catalog. If @smaheshwar-pltr does not object, we can follow up on this patch. |
| return encryptionManager; | ||
| } | ||
|
|
||
| private void encryptionPropsFromMetadata(TableMetadata metadata) { |
There was a problem hiding this comment.
Is this method applied on the TableMetadata that is fetched directly from the REST catalog, and not from the metadata.json file? Both are possible, but the former must override (and check) the latter, to protect against the key removal and other attacks.
There was a problem hiding this comment.
Yes this method will always be applied on a metadata field of a LoadTableResponse object received directly from the REST catalog (its only usage within this class is as such, and you can check the constructor usages within RESTSessionCatalog to confirm that the metadata coming in from there is as such too.
There was a problem hiding this comment.
Question (not sure if there have been discussions here or if your team have thoughts): we want the key ID to come from the REST catalog service directly for security reasons.
It's typical for REST catalogs to provide metadata that corresponds to the metadata file in storage and not modify it apart from that. Given this, would it be preferable to have this field returned within the LoadTableResponse itself, to encourage catalogs to track it explicitly?
The concrete proposal here might be: ENCRYPTION_TABLE_KEY and ENCRYPTION_DEK_LENGTH become properties on the LoadTableResponse's config (mentioned in the REST spec here).
There was a problem hiding this comment.
Well I see two scenarios when thinking about this:
- metadata.json is something that both the server and the clients can read (although clients wouldn't need to, given they get the metadata with the
LoadTableResponse) - metadata.json can only be accessed on the server side and clients are not given FS credentials (either vended or not) to reach it
For case (1) I totally agree, we can't rely on just metadata.json to store these encryption properties, and the catalog should store it separately too, and eventually populating (i.e. doing the override logic referred by @ggershinsky) the properties in the LoadTableResponse to be created.
For case (2) I'm not 100% sure, but still leaning toward the catalog taking on this responsibility.
Either way, for the client side there's not much we can do other than recommending clients to consider the metadata from LoadTableResponse only. The rest (no pun intended) is on the server side to be decided and will be implementation-specific. For this code snippet above, irrelevant IMHO.
Let me know your thoughts.
There was a problem hiding this comment.
Afaik, HMS is not optimized for JSON storage. But maybe someone in the community will take on storing the full metadata object there, to improve table security. Having only the table properties is barely sufficient. I think we should recommend REST catalogs for encrypted tables.
It's not. But would storing a hash suffice? If so we could generate the hash of the whole JSON content and store it via an additional (Hive) table property. Then during table loading we can verify that the TableMetadata we just read in from a potentially untrusted storage (and yet metadata.json is not encrypted) is original or has been tampered with.
There was a problem hiding this comment.
One aspect of having an encrypted metadata.json is when the table schema is also considered a sensitive piece of information. I haven't found this in the discussions but do you know if this has ever been considered @ggershinsky ?
There was a problem hiding this comment.
would storing a hash suffice? If so we could generate the hash of the whole JSON content and store it via an additional (Hive) table property. Then during table loading we can verify that the TableMetadata we just read in from a potentially untrusted storage (and yet metadata.json is not encrypted) is original or has been tampered with.
I think it's a good idea
There was a problem hiding this comment.
One aspect of having an encrypted metadata.json is when the table schema is also considered a sensitive piece of information. I haven't found this in the discussions but do you know if this has ever been considered
Not sure. Though, it should be possible to have a REST implementation that hides the metadata.json file from the storage.
There was a problem hiding this comment.
I think it's a good idea
Sounds good, I can take this on and will produce a PR shortly.
Not sure. Though, it should be possible to have a REST implementation that hides the metadata.json file from the storage.
Yes, with REST that's true, I just meant it in a general sense, e.g. it's not currently possible to hide the schema of an encrypted table with Hive catalog. It may just be one more thing to note/document as a limitation of encryption wrt. Hive catalog - just merely wanted to highlight this though.
|
Also, it would be good to refactor (if possible) a code common to this PR and to #13066 , so that other catalogs will be able to re-use it. |
…ption Adds read and write support for Iceberg tables whose data files are encrypted at rest with Parquet Modular Encryption (AES-GCM), under Iceberg's Standard/PME key management. Encryption is transparent to the SQL user: writing to an encryption-enabled table produces encrypted Parquet, and reading one decrypts it. JVM engines get the data-file half of this for free -- OutputFileFactory hands an EncryptedOutputFile to parquet-java, which applies PME itself. StarRocks cannot, because its Parquet reader and writer are its own C++ implementation, so the data key has to be generated in the BE, applied through parquet-cpp, and carried back to the FE for key_metadata. The metadata half -- manifest encryption and the KEK chain -- StarRocks inherits from Iceberg exactly as Spark does. Write. The FE decides from the table's EncryptionManager and signals the BE with the algorithm and DEK length, never key material. The BE draws a fresh per-file DEK from RAND_bytes, attaches FileEncryptionProperties, and calls disable_aad_prefix_storage() so the AAD prefix is withheld from the file and a reader must supply it from key_metadata -- that is what binds a file to its identity in the table and makes whole-file substitution detectable. The DEK returns to the FE at file completion and is zeroized (OPENSSL_cleanse) when the writer is destroyed. Read. The FE recovers the DEK by parsing StandardKeyMetadata directly from file.keyMetadata(), not via EncryptionManager.decrypt(): in 1.11.0 that returns a StandardDecryptedInputFile, which implements NativeEncryptionInputFile rather than NativelyEncryptedFile and exposes no nativeCryptoParameters(), so there is no file key to take from it. The BE decrypts the footer, caches an immutable EncryptionContext, and builds a per-reader page decryptor, since parquet's Decryptor mutates AAD per page and cannot be shared. Delete files each carry their own key. Position deletes get key material on TIcebergDeleteFile.parquet_encryption_info; equality deletes already have their own scan range and pick it up from the shared buildScanRange path. Fails closed. A key that cannot be recovered fails the query. An encrypted footer with no DEK is refused rather than parsed as plaintext. AES_CTR has no PME equivalent and is rejected at planning instead of being silently written as GCM. An encrypted ORC position-delete file is refused, because only the Parquet reader can decrypt. The DEK length is never defaulted by the BE: the FE resolves it from encryption.data-key-length using Iceberg's own constant and default, and encryption enabled with no length reaching the BE is an error. At commit the returned key is checked against that length, so a file encrypted below the table's declared strength is refused rather than recorded. A table that declares encryption while the catalog hands back a plaintext EncryptionManager is refused on both paths (IcebergEncryption). encryption.key-id is a table property and reaches every catalog, but the manager comes from the catalog's TableOperations, so the two can disagree -- and a plaintext manager would otherwise read as "unencrypted table" and commit cleartext into a table configured to protect it. Both gates test the manager the library returned rather than a version or catalog name, so they stop firing on their own when a catalog gains encryption support. Requires Iceberg 1.11.0: key_metadata holds the DEK in the clear, which is safe only because Iceberg encrypts the manifest carrying it, and that chain -- manifest and manifest-list encryption plus the encryption-keys list in table metadata -- arrives in 1.11.0. Also builds arrow/parquet-cpp with -DPARQUET_REQUIRE_ENCRYPTION=ON and installs the PME headers; the symbols were already in libparquet.a, only the headers were missing. Scope, set by Iceberg rather than by this change: table format v3 (EncryptionUtil.checkCompatibility rejects the encryption properties below v3), and Hive Metastore catalogs only -- in 1.11.0 only HiveTableOperations builds a real EncryptionManager, so REST catalogs report themselves unencrypted whatever their properties say. apache/iceberg#13225 is the pending fix; nothing here changes when it lands beyond the dependency bump. Tests: encrypted write round trip, AAD-prefix withholding, and rejection of an unset and an out-of-range DEK length (parquet_file_writer_test); encrypted position-delete round trip plus a negative case asserting a read without the key fails (iceberg_delete_builder_test); key_metadata round-tripped against Iceberg's own StandardKeyMetadata parser in both directions (StarRocksKeyMetadataTest); the write signal, the length default, and that the FE never sends key material (IcebergTableSinkTest); read-path key recovery and fail-closed behaviour (IcebergConnectorScanRangeSourceTest); refusal at commit of a key shorter than the table's policy (IcebergMetadataTest); and both declared-vs-enforced gates (IcebergEncryptionTest). Signed-off-by: Annie Yan <anniey@outlook.com>
…ption StarRocks cannot read or write Iceberg tables whose data files are encrypted at rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks. Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME key management: a per-file data encryption key encrypts the Parquet file, and the DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata, protected by the encrypted manifest that carries it. JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to be built: - BE generates the per-file DEK and applies PME through parquet-cpp on write - the DEK travels back to FE, which serializes it into key_metadata at commit - on read FE recovers the DEK from key_metadata and ships it per scan range; BE decrypts the footer and pages itself The metadata half -- manifest encryption and the KEK chain -- is inherited from the Iceberg library exactly as Spark does, since FE planning and commit already go through it. key_metadata serialization delegates to Iceberg's own (package-private) StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is byte-identical to what any other Iceberg engine produces. All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink (DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE signal from one place so they cannot drift. Position-delete files carry their own per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile and records its own keyMetadata. An unencrypted delete file beside encrypted data would disclose which rows were removed, which is what table encryption exists to prevent. Fails closed throughout: - a key that cannot be recovered fails the query; never falls back to plaintext - an encrypted footer with no DEK is refused rather than parsed as plaintext - AES_CTR has no PME equivalent and is rejected at planning, not written as GCM - BE never defaults the DEK length. Only FE can see encryption.data-key-length, so FE resolves it using Iceberg's own constant and default; encryption enabled with no length reaching BE is an error. At commit the key BE returns is checked against that length, so a file encrypted below the table's declared strength is refused - a table that declares encryption while the catalog supplies a plaintext EncryptionManager is refused on both read and write. Those signals come from different places -- the property travels with the metadata for every catalog, the manager comes from the catalog's TableOperations -- so they can disagree, and a plaintext manager would otherwise read as "unencrypted table" - whether a table is encrypted is decided from encryption.key-id on both the write signal and at commit. Deciding it from the manager's type on one side and the property on the other would let BE encrypt a file the commit records no key for, leaving data nothing can decrypt - the encrypted page-header length prefix sits outside the AEAD, so it is bounded (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB cap the plaintext path uses) before it is used to size any buffer Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds the DEK in the clear, which is safe only because Iceberg encrypts the manifest carrying it, and that chain arrives in 1.11.0. Table format v3 is required -- EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore catalogs only for now: in 1.11.0 only HiveTableOperations builds a real EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its properties say (apache/iceberg#13225). Because the decision reads table.encryption() rather than a version or catalog name, no StarRocks change is needed when that lands. arrow/parquet-cpp is built with -DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols were already in libparquet.a. Observability: no signal was removed or renamed. Existing Parquet reader/writer scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and refusals surface as query errors naming the table and the reason, so no new metric was added. Key material is never logged or put in a profile. Signed-off-by: Annie Yan <anniey@outlook.com>
Records the design behind Iceberg table encryption: the trust boundaries and where key material is allowed to exist, the write and read flows across FE and BE, and the key-management chain from the table master key down to the per-file DEK. Written down because the constraints are not visible from the code. Why Iceberg 1.11.0 is a correctness floor rather than a version preference, why table format v3 is required, why the feature is limited to Hive Metastore catalogs today and what changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check exists at all -- a table can declare encryption the catalog will not apply, and the two signals come from different places. Also records the gaps that are deliberately out of scope, so the next person does not have to rediscover which are intentional. Signed-off-by: Annie Yan <anniey@outlook.com>
…ption StarRocks cannot read or write Iceberg tables whose data files are encrypted at rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks. Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME key management: a per-file data encryption key encrypts the Parquet file, and the DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata, protected by the encrypted manifest that carries it. JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to be built: - BE generates the per-file DEK and applies PME through parquet-cpp on write - the DEK travels back to FE, which serializes it into key_metadata at commit - on read FE recovers the DEK from key_metadata and ships it per scan range; BE decrypts the footer and pages itself The metadata half -- manifest encryption and the KEK chain -- is inherited from the Iceberg library exactly as Spark does, since FE planning and commit already go through it. key_metadata serialization delegates to Iceberg's own (package-private) StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is byte-identical to what any other Iceberg engine produces. All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink (DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE signal from one place so they cannot drift. Position-delete files carry their own per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile and records its own keyMetadata. An unencrypted delete file beside encrypted data would disclose which rows were removed, which is what table encryption exists to prevent. Fails closed throughout: - a key that cannot be recovered fails the query; never falls back to plaintext - an encrypted footer with no DEK is refused rather than parsed as plaintext - AES_CTR has no PME equivalent and is rejected at planning, not written as GCM - BE never defaults the DEK length. Only FE can see encryption.data-key-length, so FE resolves it using Iceberg's own constant and default; encryption enabled with no length reaching BE is an error. At commit the key BE returns is checked against that length, so a file encrypted below the table's declared strength is refused - a table that declares encryption while the catalog supplies a plaintext EncryptionManager is refused on both read and write. Those signals come from different places -- the property travels with the metadata for every catalog, the manager comes from the catalog's TableOperations -- so they can disagree, and a plaintext manager would otherwise read as "unencrypted table" - whether a table is encrypted is decided from encryption.key-id on both the write signal and at commit. Deciding it from the manager's type on one side and the property on the other would let BE encrypt a file the commit records no key for, leaving data nothing can decrypt - the encrypted page-header length prefix sits outside the AEAD, so it is bounded (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB cap the plaintext path uses) before it is used to size any buffer Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds the DEK in the clear, which is safe only because Iceberg encrypts the manifest carrying it, and that chain arrives in 1.11.0. Table format v3 is required -- EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore catalogs only for now: in 1.11.0 only HiveTableOperations builds a real EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its properties say (apache/iceberg#13225). Because the decision reads table.encryption() rather than a version or catalog name, no StarRocks change is needed when that lands. arrow/parquet-cpp is built with -DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols were already in libparquet.a. Observability: no signal was removed or renamed. Existing Parquet reader/writer scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and refusals surface as query errors naming the table and the reason, so no new metric was added. Key material is never logged or put in a profile. Signed-off-by: Annie Yan <anniey@outlook.com>
Records the design behind Iceberg table encryption: the trust boundaries and where key material is allowed to exist, the write and read flows across FE and BE, and the key-management chain from the table master key down to the per-file DEK. Written down because the constraints are not visible from the code. Why Iceberg 1.11.0 is a correctness floor rather than a version preference, why table format v3 is required, why the feature is limited to Hive Metastore catalogs today and what changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check exists at all -- a table can declare encryption the catalog will not apply, and the two signals come from different places. Also records the gaps that are deliberately out of scope, so the next person does not have to rediscover which are intentional. Signed-off-by: Annie Yan <anniey@outlook.com>
Records the design behind Iceberg table encryption: the trust boundaries and where key material is allowed to exist, the write and read flows across FE and BE, and the key-management chain from the table master key down to the per-file DEK. Written down because the constraints are not visible from the code. Why Iceberg 1.11.0 is a correctness floor rather than a version preference, why table format v3 is required, why the feature is limited to Hive Metastore catalogs today and what changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check exists at all -- a table can declare encryption the catalog will not apply, and the two signals come from different places. Also records the gaps that are deliberately out of scope, so the next person does not have to rediscover which are intentional. The diagrams ship with their PlantUML sources and a render script, because they had already gone stale once: a PNG kept asserting behaviour the code no longer had, and with no source in tree there was nothing to check it against. Re-rendering the two untouched diagrams reproduces their committed PNGs byte for byte, so the sources and the images are known to agree. Corrects four claims that no longer matched the code: - the write path decides from the table's declaration (encryption.key-id), not from the EncryptionManager's type, and all three sinks derive that one signal - the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core, so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data. The BE rejects an unknown cipher; it is not a planning-time check - the read gate runs in buildScanRange ahead of any per-file test, not inside buildEncryptionInfo, because such a table's files carry no key metadata and a per-file check would skip them - REST catalogs: upstream fails silently, but StarRocks refuses rather than inheriting that silence Signed-off-by: Annie Yan <anniey@outlook.com>
…ption StarRocks cannot read or write Iceberg tables whose data files are encrypted at rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks. Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME key management: a per-file data encryption key encrypts the Parquet file, and the DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata, protected by the encrypted manifest that carries it. JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to be built: - BE generates the per-file DEK and applies PME through parquet-cpp on write - the DEK travels back to FE, which serializes it into key_metadata at commit - on read FE recovers the DEK from key_metadata and ships it per scan range; BE decrypts the footer and pages itself The metadata half -- manifest encryption and the KEK chain -- is inherited from the Iceberg library exactly as Spark does, since FE planning and commit already go through it. key_metadata serialization delegates to Iceberg's own (package-private) StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is byte-identical to what any other Iceberg engine produces. All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink (DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE signal from one place so they cannot drift. Position-delete files carry their own per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile and records its own keyMetadata. An unencrypted delete file beside encrypted data would disclose which rows were removed, which is what table encryption exists to prevent. Fails closed throughout: - a key that cannot be recovered fails the query; never falls back to plaintext - an encrypted footer with no DEK is refused rather than parsed as plaintext - AES_CTR has no PME equivalent and is rejected at planning, not written as GCM - BE never defaults the DEK length. Only FE can see encryption.data-key-length, so FE resolves it using Iceberg's own constant and default; encryption enabled with no length reaching BE is an error. At commit the key BE returns is checked against that length, so a file encrypted below the table's declared strength is refused - a table that declares encryption while the catalog supplies a plaintext EncryptionManager is refused on both read and write. Those signals come from different places -- the property travels with the metadata for every catalog, the manager comes from the catalog's TableOperations -- so they can disagree, and a plaintext manager would otherwise read as "unencrypted table" - whether a table is encrypted is decided from encryption.key-id on both the write signal and at commit. Deciding it from the manager's type on one side and the property on the other would let BE encrypt a file the commit records no key for, leaving data nothing can decrypt - the encrypted page-header length prefix sits outside the AEAD, so it is bounded (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB cap the plaintext path uses) before it is used to size any buffer Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds the DEK in the clear, which is safe only because Iceberg encrypts the manifest carrying it, and that chain arrives in 1.11.0. Table format v3 is required -- EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore catalogs only for now: in 1.11.0 only HiveTableOperations builds a real EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its properties say (apache/iceberg#13225). Because the decision reads table.encryption() rather than a version or catalog name, no StarRocks change is needed when that lands. arrow/parquet-cpp is built with -DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols were already in libparquet.a. Observability: no signal was removed or renamed. Existing Parquet reader/writer scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and refusals surface as query errors naming the table and the reason, so no new metric was added. Key material is never logged or put in a profile. Signed-off-by: Annie Yan <anniey@outlook.com>
Records the design behind Iceberg table encryption: the trust boundaries and where key material is allowed to exist, the write and read flows across FE and BE, and the key-management chain from the table master key down to the per-file DEK. Written down because the constraints are not visible from the code. Why Iceberg 1.11.0 is a correctness floor rather than a version preference, why table format v3 is required, why the feature is limited to Hive Metastore catalogs today and what changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check exists at all -- a table can declare encryption the catalog will not apply, and the two signals come from different places. Also records the gaps that are deliberately out of scope, so the next person does not have to rediscover which are intentional. The diagrams ship with their PlantUML sources and a render script, because they had already gone stale once: a PNG kept asserting behaviour the code no longer had, and with no source in tree there was nothing to check it against. Re-rendering the two untouched diagrams reproduces their committed PNGs byte for byte, so the sources and the images are known to agree. Corrects four claims that no longer matched the code: - the write path decides from the table's declaration (encryption.key-id), not from the EncryptionManager's type, and all three sinks derive that one signal - the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core, so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data. The BE rejects an unknown cipher; it is not a planning-time check - the read gate runs in buildScanRange ahead of any per-file test, not inside buildEncryptionInfo, because such a table's files carry no key metadata and a per-file check would skip them - REST catalogs: upstream fails silently, but StarRocks refuses rather than inheriting that silence Signed-off-by: Annie Yan <anniey@outlook.com>
…ption StarRocks cannot read or write Iceberg tables whose data files are encrypted at rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks. Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME key management: a per-file data encryption key encrypts the Parquet file, and the DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata, protected by the encrypted manifest that carries it. JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to be built: - BE generates the per-file DEK and applies PME through parquet-cpp on write - the DEK travels back to FE, which serializes it into key_metadata at commit - on read FE recovers the DEK from key_metadata and ships it per scan range; BE decrypts the footer and pages itself The metadata half -- manifest encryption and the KEK chain -- is inherited from the Iceberg library exactly as Spark does, since FE planning and commit already go through it. key_metadata serialization delegates to Iceberg's own (package-private) StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is byte-identical to what any other Iceberg engine produces. All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink (DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE signal from one place so they cannot drift. Position-delete files carry their own per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile and records its own keyMetadata. An unencrypted delete file beside encrypted data would disclose which rows were removed, which is what table encryption exists to prevent. Fails closed throughout: - a key that cannot be recovered fails the query; never falls back to plaintext - an encrypted footer with no DEK is refused rather than parsed as plaintext - AES_CTR has no PME equivalent and is rejected at planning, not written as GCM - BE never defaults the DEK length. Only FE can see encryption.data-key-length, so FE resolves it using Iceberg's own constant and default; encryption enabled with no length reaching BE is an error. At commit the key BE returns is checked against that length, so a file encrypted below the table's declared strength is refused - a table that declares encryption while the catalog supplies a plaintext EncryptionManager is refused on both read and write. Those signals come from different places -- the property travels with the metadata for every catalog, the manager comes from the catalog's TableOperations -- so they can disagree, and a plaintext manager would otherwise read as "unencrypted table" - whether a table is encrypted is decided from encryption.key-id on both the write signal and at commit. Deciding it from the manager's type on one side and the property on the other would let BE encrypt a file the commit records no key for, leaving data nothing can decrypt - the encrypted page-header length prefix sits outside the AEAD, so it is bounded (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB cap the plaintext path uses) before it is used to size any buffer Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds the DEK in the clear, which is safe only because Iceberg encrypts the manifest carrying it, and that chain arrives in 1.11.0. Table format v3 is required -- EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore catalogs only for now: in 1.11.0 only HiveTableOperations builds a real EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its properties say (apache/iceberg#13225). Because the decision reads table.encryption() rather than a version or catalog name, no StarRocks change is needed when that lands. arrow/parquet-cpp is built with -DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols were already in libparquet.a. Observability: no signal was removed or renamed. Existing Parquet reader/writer scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and refusals surface as query errors naming the table and the reason, so no new metric was added. Key material is never logged or put in a profile. Signed-off-by: Annie Yan <anniey@outlook.com>
Records the design behind Iceberg table encryption: the trust boundaries and where key material is allowed to exist, the write and read flows across FE and BE, and the key-management chain from the table master key down to the per-file DEK. Written down because the constraints are not visible from the code. Why Iceberg 1.11.0 is a correctness floor rather than a version preference, why table format v3 is required, why the feature is limited to Hive Metastore catalogs today and what changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check exists at all -- a table can declare encryption the catalog will not apply, and the two signals come from different places. Also records the gaps that are deliberately out of scope, so the next person does not have to rediscover which are intentional. The diagrams ship with their PlantUML sources and a render script, because they had already gone stale once: a PNG kept asserting behaviour the code no longer had, and with no source in tree there was nothing to check it against. Re-rendering the two untouched diagrams reproduces their committed PNGs byte for byte, so the sources and the images are known to agree. Corrects four claims that no longer matched the code: - the write path decides from the table's declaration (encryption.key-id), not from the EncryptionManager's type, and all three sinks derive that one signal - the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core, so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data. The BE rejects an unknown cipher; it is not a planning-time check - the read gate runs in buildScanRange ahead of any per-file test, not inside buildEncryptionInfo, because such a table's files carry no key metadata and a per-file check would skip them - REST catalogs: upstream fails silently, but StarRocks refuses rather than inheriting that silence Signed-off-by: Annie Yan <anniey@outlook.com>
…ption StarRocks cannot read or write Iceberg tables whose data files are encrypted at rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks. Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME key management: a per-file data encryption key encrypts the Parquet file, and the DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata, protected by the encrypted manifest that carries it. JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to be built: - BE generates the per-file DEK and applies PME through parquet-cpp on write - the DEK travels back to FE, which serializes it into key_metadata at commit - on read FE recovers the DEK from key_metadata and ships it per scan range; BE decrypts the footer and pages itself The metadata half -- manifest encryption and the KEK chain -- is inherited from the Iceberg library exactly as Spark does, since FE planning and commit already go through it. key_metadata serialization delegates to Iceberg's own (package-private) StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is byte-identical to what any other Iceberg engine produces. All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink (DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE signal from one place so they cannot drift. Position-delete files carry their own per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile and records its own keyMetadata. An unencrypted delete file beside encrypted data would disclose which rows were removed, which is what table encryption exists to prevent. Fails closed throughout: - a key that cannot be recovered fails the query; never falls back to plaintext - an encrypted footer with no DEK is refused rather than parsed as plaintext - AES_CTR has no PME equivalent and is rejected at planning, not written as GCM - BE never defaults the DEK length. Only FE can see encryption.data-key-length, so FE resolves it using Iceberg's own constant and default; encryption enabled with no length reaching BE is an error. At commit the key BE returns is checked against that length, so a file encrypted below the table's declared strength is refused - a table that declares encryption while the catalog supplies a plaintext EncryptionManager is refused on both read and write. Those signals come from different places -- the property travels with the metadata for every catalog, the manager comes from the catalog's TableOperations -- so they can disagree, and a plaintext manager would otherwise read as "unencrypted table" - whether a table is encrypted is decided from encryption.key-id on both the write signal and at commit. Deciding it from the manager's type on one side and the property on the other would let BE encrypt a file the commit records no key for, leaving data nothing can decrypt - the encrypted page-header length prefix sits outside the AEAD, so it is bounded (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB cap the plaintext path uses) before it is used to size any buffer Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds the DEK in the clear, which is safe only because Iceberg encrypts the manifest carrying it, and that chain arrives in 1.11.0. Table format v3 is required -- EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore catalogs only for now: in 1.11.0 only HiveTableOperations builds a real EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its properties say (apache/iceberg#13225). Because the decision reads table.encryption() rather than a version or catalog name, no StarRocks change is needed when that lands. arrow/parquet-cpp is built with -DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols were already in libparquet.a. Observability: no signal was removed or renamed. Existing Parquet reader/writer scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and refusals surface as query errors naming the table and the reason, so no new metric was added. Key material is never logged or put in a profile. Signed-off-by: Annie Yan <anniey@outlook.com>
Records the design behind Iceberg table encryption: the trust boundaries and where key material is allowed to exist, the write and read flows across FE and BE, and the key-management chain from the table master key down to the per-file DEK. Written down because the constraints are not visible from the code. Why Iceberg 1.11.0 is a correctness floor rather than a version preference, why table format v3 is required, why the feature is limited to Hive Metastore catalogs today and what changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check exists at all -- a table can declare encryption the catalog will not apply, and the two signals come from different places. Also records the gaps that are deliberately out of scope, so the next person does not have to rediscover which are intentional. The diagrams ship with their PlantUML sources and a render script, because they had already gone stale once: a PNG kept asserting behaviour the code no longer had, and with no source in tree there was nothing to check it against. Re-rendering the two untouched diagrams reproduces their committed PNGs byte for byte, so the sources and the images are known to agree. Corrects four claims that no longer matched the code: - the write path decides from the table's declaration (encryption.key-id), not from the EncryptionManager's type, and all three sinks derive that one signal - the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core, so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data. The BE rejects an unknown cipher; it is not a planning-time check - the read gate runs in buildScanRange ahead of any per-file test, not inside buildEncryptionInfo, because such a table's files carry no key metadata and a per-file check would skip them - REST catalogs: upstream fails silently, but StarRocks refuses rather than inheriting that silence Signed-off-by: Annie Yan <anniey@outlook.com>
…ption StarRocks cannot read or write Iceberg tables whose data files are encrypted at rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks. Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME key management: a per-file data encryption key encrypts the Parquet file, and the DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata, protected by the encrypted manifest that carries it. JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to be built: - BE generates the per-file DEK and applies PME through parquet-cpp on write - the DEK travels back to FE, which serializes it into key_metadata at commit - on read FE recovers the DEK from key_metadata and ships it per scan range; BE decrypts the footer and pages itself The metadata half -- manifest encryption and the KEK chain -- is inherited from the Iceberg library exactly as Spark does, since FE planning and commit already go through it. key_metadata serialization delegates to Iceberg's own (package-private) StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is byte-identical to what any other Iceberg engine produces. All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink (DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE signal from one place so they cannot drift. Position-delete files carry their own per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile and records its own keyMetadata. An unencrypted delete file beside encrypted data would disclose which rows were removed, which is what table encryption exists to prevent. Fails closed throughout: - a key that cannot be recovered fails the query; never falls back to plaintext - an encrypted footer with no DEK is refused rather than parsed as plaintext - AES_CTR has no PME equivalent and is rejected at planning, not written as GCM - BE never defaults the DEK length. Only FE can see encryption.data-key-length, so FE resolves it using Iceberg's own constant and default; encryption enabled with no length reaching BE is an error. At commit the key BE returns is checked against that length, so a file encrypted below the table's declared strength is refused - a table that declares encryption while the catalog supplies a plaintext EncryptionManager is refused on both read and write. Those signals come from different places -- the property travels with the metadata for every catalog, the manager comes from the catalog's TableOperations -- so they can disagree, and a plaintext manager would otherwise read as "unencrypted table" - whether a table is encrypted is decided from encryption.key-id on both the write signal and at commit. Deciding it from the manager's type on one side and the property on the other would let BE encrypt a file the commit records no key for, leaving data nothing can decrypt - the encrypted page-header length prefix sits outside the AEAD, so it is bounded (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB cap the plaintext path uses) before it is used to size any buffer Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds the DEK in the clear, which is safe only because Iceberg encrypts the manifest carrying it, and that chain arrives in 1.11.0. Table format v3 is required -- EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore catalogs only for now: in 1.11.0 only HiveTableOperations builds a real EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its properties say (apache/iceberg#13225). Because the decision reads table.encryption() rather than a version or catalog name, no StarRocks change is needed when that lands. arrow/parquet-cpp is built with -DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols were already in libparquet.a. Observability: no signal was removed or renamed. Existing Parquet reader/writer scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and refusals surface as query errors naming the table and the reason, so no new metric was added. Key material is never logged or put in a profile. Signed-off-by: Annie Yan <anniey@outlook.com>
Records the design behind Iceberg table encryption: the trust boundaries and where key material is allowed to exist, the write and read flows across FE and BE, and the key-management chain from the table master key down to the per-file DEK. Written down because the constraints are not visible from the code. Why Iceberg 1.11.0 is a correctness floor rather than a version preference, why table format v3 is required, why the feature is limited to Hive Metastore catalogs today and what changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check exists at all -- a table can declare encryption the catalog will not apply, and the two signals come from different places. Also records the gaps that are deliberately out of scope, so the next person does not have to rediscover which are intentional. The diagrams ship with their PlantUML sources and a render script, because they had already gone stale once: a PNG kept asserting behaviour the code no longer had, and with no source in tree there was nothing to check it against. Re-rendering the two untouched diagrams reproduces their committed PNGs byte for byte, so the sources and the images are known to agree. Corrects four claims that no longer matched the code: - the write path decides from the table's declaration (encryption.key-id), not from the EncryptionManager's type, and all three sinks derive that one signal - the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core, so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data. The BE rejects an unknown cipher; it is not a planning-time check - the read gate runs in buildScanRange ahead of any per-file test, not inside buildEncryptionInfo, because such a table's files carry no key metadata and a per-file check would skip them - REST catalogs: upstream fails silently, but StarRocks refuses rather than inheriting that silence Signed-off-by: Annie Yan <anniey@outlook.com>
…ption StarRocks cannot read or write Iceberg tables whose data files are encrypted at rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks. Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME key management: a per-file data encryption key encrypts the Parquet file, and the DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata, protected by the encrypted manifest that carries it. JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to be built: - BE generates the per-file DEK and applies PME through parquet-cpp on write - the DEK travels back to FE, which serializes it into key_metadata at commit - on read FE recovers the DEK from key_metadata and ships it per scan range; BE decrypts the footer and pages itself The metadata half -- manifest encryption and the KEK chain -- is inherited from the Iceberg library exactly as Spark does, since FE planning and commit already go through it. key_metadata serialization delegates to Iceberg's own (package-private) StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is byte-identical to what any other Iceberg engine produces. All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink (DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE signal from one place so they cannot drift. Position-delete files carry their own per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile and records its own keyMetadata. An unencrypted delete file beside encrypted data would disclose which rows were removed, which is what table encryption exists to prevent. Fails closed throughout: - a key that cannot be recovered fails the query; never falls back to plaintext - an encrypted footer with no DEK is refused rather than parsed as plaintext - AES_CTR has no PME equivalent and is rejected at planning, not written as GCM - BE never defaults the DEK length. Only FE can see encryption.data-key-length, so FE resolves it using Iceberg's own constant and default; encryption enabled with no length reaching BE is an error. At commit the key BE returns is checked against that length, so a file encrypted below the table's declared strength is refused - a table that declares encryption while the catalog supplies a plaintext EncryptionManager is refused on both read and write. Those signals come from different places -- the property travels with the metadata for every catalog, the manager comes from the catalog's TableOperations -- so they can disagree, and a plaintext manager would otherwise read as "unencrypted table" - whether a table is encrypted is decided from encryption.key-id on both the write signal and at commit. Deciding it from the manager's type on one side and the property on the other would let BE encrypt a file the commit records no key for, leaving data nothing can decrypt - the encrypted page-header length prefix sits outside the AEAD, so it is bounded (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB cap the plaintext path uses) before it is used to size any buffer Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds the DEK in the clear, which is safe only because Iceberg encrypts the manifest carrying it, and that chain arrives in 1.11.0. Table format v3 is required -- EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore catalogs only for now: in 1.11.0 only HiveTableOperations builds a real EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its properties say (apache/iceberg#13225). Because the decision reads table.encryption() rather than a version or catalog name, no StarRocks change is needed when that lands. arrow/parquet-cpp is built with -DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols were already in libparquet.a. Observability: no signal was removed or renamed. Existing Parquet reader/writer scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and refusals surface as query errors naming the table and the reason, so no new metric was added. Key material is never logged or put in a profile. Signed-off-by: Annie Yan <anniey@outlook.com>
Records the design behind Iceberg table encryption: the trust boundaries and where key material is allowed to exist, the write and read flows across FE and BE, and the key-management chain from the table master key down to the per-file DEK. Written down because the constraints are not visible from the code. Why Iceberg 1.11.0 is a correctness floor rather than a version preference, why table format v3 is required, why the feature is limited to Hive Metastore catalogs today and what changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check exists at all -- a table can declare encryption the catalog will not apply, and the two signals come from different places. Also records the gaps that are deliberately out of scope, so the next person does not have to rediscover which are intentional. The diagrams ship with their PlantUML sources and a render script, because they had already gone stale once: a PNG kept asserting behaviour the code no longer had, and with no source in tree there was nothing to check it against. Re-rendering the two untouched diagrams reproduces their committed PNGs byte for byte, so the sources and the images are known to agree. Corrects four claims that no longer matched the code: - the write path decides from the table's declaration (encryption.key-id), not from the EncryptionManager's type, and all three sinks derive that one signal - the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core, so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data. The BE rejects an unknown cipher; it is not a planning-time check - the read gate runs in buildScanRange ahead of any per-file test, not inside buildEncryptionInfo, because such a table's files carry no key metadata and a per-file check would skip them - REST catalogs: upstream fails silently, but StarRocks refuses rather than inheriting that silence Signed-off-by: Annie Yan <anniey@outlook.com>
…ption StarRocks cannot read or write Iceberg tables whose data files are encrypted at rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks. Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME key management: a per-file data encryption key encrypts the Parquet file, and the DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata, protected by the encrypted manifest that carries it. JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to be built: - BE generates the per-file DEK and applies PME through parquet-cpp on write - the DEK travels back to FE, which serializes it into key_metadata at commit - on read FE recovers the DEK from key_metadata and ships it per scan range; BE decrypts the footer and pages itself The metadata half -- manifest encryption and the KEK chain -- is inherited from the Iceberg library exactly as Spark does, since FE planning and commit already go through it. key_metadata serialization delegates to Iceberg's own (package-private) StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is byte-identical to what any other Iceberg engine produces. All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink (DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE signal from one place so they cannot drift. Position-delete files carry their own per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile and records its own keyMetadata. An unencrypted delete file beside encrypted data would disclose which rows were removed, which is what table encryption exists to prevent. Fails closed throughout: - a key that cannot be recovered fails the query; never falls back to plaintext - an encrypted footer with no DEK is refused rather than parsed as plaintext - AES_CTR has no PME equivalent and is rejected at planning, not written as GCM - BE never defaults the DEK length. Only FE can see encryption.data-key-length, so FE resolves it using Iceberg's own constant and default; encryption enabled with no length reaching BE is an error. At commit the key BE returns is checked against that length, so a file encrypted below the table's declared strength is refused - a table that declares encryption while the catalog supplies a plaintext EncryptionManager is refused on both read and write. Those signals come from different places -- the property travels with the metadata for every catalog, the manager comes from the catalog's TableOperations -- so they can disagree, and a plaintext manager would otherwise read as "unencrypted table" - whether a table is encrypted is decided from encryption.key-id on both the write signal and at commit. Deciding it from the manager's type on one side and the property on the other would let BE encrypt a file the commit records no key for, leaving data nothing can decrypt - the encrypted page-header length prefix sits outside the AEAD, so it is bounded (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB cap the plaintext path uses) before it is used to size any buffer Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds the DEK in the clear, which is safe only because Iceberg encrypts the manifest carrying it, and that chain arrives in 1.11.0. Table format v3 is required -- EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore catalogs only for now: in 1.11.0 only HiveTableOperations builds a real EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its properties say (apache/iceberg#13225). Because the decision reads table.encryption() rather than a version or catalog name, no StarRocks change is needed when that lands. arrow/parquet-cpp is built with -DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols were already in libparquet.a. Observability: no signal was removed or renamed. Existing Parquet reader/writer scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and refusals surface as query errors naming the table and the reason, so no new metric was added. Key material is never logged or put in a profile. Signed-off-by: Annie Yan <anniey@outlook.com>
Records the design behind Iceberg table encryption: the trust boundaries and where key material is allowed to exist, the write and read flows across FE and BE, and the key-management chain from the table master key down to the per-file DEK. Written down because the constraints are not visible from the code. Why Iceberg 1.11.0 is a correctness floor rather than a version preference, why table format v3 is required, why the feature is limited to Hive Metastore catalogs today and what changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check exists at all -- a table can declare encryption the catalog will not apply, and the two signals come from different places. Also records the gaps that are deliberately out of scope, so the next person does not have to rediscover which are intentional. The diagrams ship with their PlantUML sources and a render script, because they had already gone stale once: a PNG kept asserting behaviour the code no longer had, and with no source in tree there was nothing to check it against. Re-rendering the two untouched diagrams reproduces their committed PNGs byte for byte, so the sources and the images are known to agree. Corrects four claims that no longer matched the code: - the write path decides from the table's declaration (encryption.key-id), not from the EncryptionManager's type, and all three sinks derive that one signal - the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core, so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data. The BE rejects an unknown cipher; it is not a planning-time check - the read gate runs in buildScanRange ahead of any per-file test, not inside buildEncryptionInfo, because such a table's files carry no key metadata and a per-file check would skip them - REST catalogs: upstream fails silently, but StarRocks refuses rather than inheriting that silence Signed-off-by: Annie Yan <anniey@outlook.com>
…ption StarRocks cannot read or write Iceberg tables whose data files are encrypted at rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks. Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME key management: a per-file data encryption key encrypts the Parquet file, and the DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata, protected by the encrypted manifest that carries it. JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to be built: - BE generates the per-file DEK and applies PME through parquet-cpp on write - the DEK travels back to FE, which serializes it into key_metadata at commit - on read FE recovers the DEK from key_metadata and ships it per scan range; BE decrypts the footer and pages itself The metadata half -- manifest encryption and the KEK chain -- is inherited from the Iceberg library exactly as Spark does, since FE planning and commit already go through it. key_metadata serialization delegates to Iceberg's own (package-private) StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is byte-identical to what any other Iceberg engine produces. All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink (DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE signal from one place so they cannot drift. Position-delete files carry their own per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile and records its own keyMetadata. An unencrypted delete file beside encrypted data would disclose which rows were removed, which is what table encryption exists to prevent. Fails closed throughout: - a key that cannot be recovered fails the query; never falls back to plaintext - an encrypted footer with no DEK is refused rather than parsed as plaintext - AES_CTR has no PME equivalent and is rejected at planning, not written as GCM - BE never defaults the DEK length. Only FE can see encryption.data-key-length, so FE resolves it using Iceberg's own constant and default; encryption enabled with no length reaching BE is an error. At commit the key BE returns is checked against that length, so a file encrypted below the table's declared strength is refused - a table that declares encryption while the catalog supplies a plaintext EncryptionManager is refused on both read and write. Those signals come from different places -- the property travels with the metadata for every catalog, the manager comes from the catalog's TableOperations -- so they can disagree, and a plaintext manager would otherwise read as "unencrypted table" - whether a table is encrypted is decided from encryption.key-id on both the write signal and at commit. Deciding it from the manager's type on one side and the property on the other would let BE encrypt a file the commit records no key for, leaving data nothing can decrypt - the encrypted page-header length prefix sits outside the AEAD, so it is bounded (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB cap the plaintext path uses) before it is used to size any buffer Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds the DEK in the clear, which is safe only because Iceberg encrypts the manifest carrying it, and that chain arrives in 1.11.0. Table format v3 is required -- EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore catalogs only for now: in 1.11.0 only HiveTableOperations builds a real EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its properties say (apache/iceberg#13225). Because the decision reads table.encryption() rather than a version or catalog name, no StarRocks change is needed when that lands. arrow/parquet-cpp is built with -DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols were already in libparquet.a. Observability: no signal was removed or renamed. Existing Parquet reader/writer scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and refusals surface as query errors naming the table and the reason, so no new metric was added. Key material is never logged or put in a profile. Signed-off-by: Annie Yan <anniey@outlook.com>
Records the design behind Iceberg table encryption: the trust boundaries and where key material is allowed to exist, the write and read flows across FE and BE, and the key-management chain from the table master key down to the per-file DEK. Written down because the constraints are not visible from the code. Why Iceberg 1.11.0 is a correctness floor rather than a version preference, why table format v3 is required, why the feature is limited to Hive Metastore catalogs today and what changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check exists at all -- a table can declare encryption the catalog will not apply, and the two signals come from different places. Also records the gaps that are deliberately out of scope, so the next person does not have to rediscover which are intentional. The diagrams ship with their PlantUML sources and a render script, because they had already gone stale once: a PNG kept asserting behaviour the code no longer had, and with no source in tree there was nothing to check it against. Re-rendering the two untouched diagrams reproduces their committed PNGs byte for byte, so the sources and the images are known to agree. Corrects four claims that no longer matched the code: - the write path decides from the table's declaration (encryption.key-id), not from the EncryptionManager's type, and all three sinks derive that one signal - the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core, so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data. The BE rejects an unknown cipher; it is not a planning-time check - the read gate runs in buildScanRange ahead of any per-file test, not inside buildEncryptionInfo, because such a table's files carry no key metadata and a per-file check would skip them - REST catalogs: upstream fails silently, but StarRocks refuses rather than inheriting that silence Signed-off-by: Annie Yan <anniey@outlook.com>
Generated-by: Codex
|
@singhpk234 @szlta apologies if I’ve been unclear. I definitely agree it would be good to get something into this release. But I don't understand what's being proposed, it seems very non-trivial to me and not quite acceptable to just decide this without any spec change in. Let me be concrete:
I’m struggling to understand the value of this PR with this restrictions and no KMS vending being available to lift them (there's no KMS vending spec in). But I'd argue that hybrid configuration is valuable when KMS authorization is managed externally to the catalog. (I'd also argue that external KMS authorization is not problematic enough to introduce such intrusive enforcements^ but I'm not going to block on this). I think I'm misunderstanding something because I can't imagine that just deciding all this now in this PR for the next release is good - if there's a straightforward implementation, please feel free to implement it on top of this PR, or supersede this PR if that’s easier! Very happy for that. (My point here is that I'm just not understanding unfortunately) |
Here is what i was proposing : c8116f3
To recap, folks advocated we need this in prev release 1.11 #13225 (comment), but we were not able to make it, it been ~4 months where some of us in the working group, talked through scenarios, came up with a rest spec proposal. I also raised this in community syncs, it was said there can be use cases we should support this in the sdk the hybrid mode. I was searching for use cases. A couple of folks also reached out to me to see if i can help in getting this in 1.12, because presently engines such as Apache Spark can't work at all with RESTCatalog, I have been trying to follow up in this pr Jul 27 .... I understand totally it takes time ! With that being said, I don't think there is anything wrong in being ambitious to get this in a upcoming release, though we should not be making implementation decision in haste .... lets do the right thing ! Thank you for all the work so far on this ! |
…ption StarRocks cannot read or write Iceberg tables whose data files are encrypted at rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks. Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME key management: a per-file data encryption key encrypts the Parquet file, and the DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata, protected by the encrypted manifest that carries it. JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to be built: - BE generates the per-file DEK and applies PME through parquet-cpp on write - the DEK travels back to FE, which serializes it into key_metadata at commit - on read FE recovers the DEK from key_metadata and ships it per scan range; BE decrypts the footer and pages itself The metadata half -- manifest encryption and the KEK chain -- is inherited from the Iceberg library exactly as Spark does, since FE planning and commit already go through it. key_metadata serialization delegates to Iceberg's own (package-private) StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is byte-identical to what any other Iceberg engine produces. All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink (DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE signal from one place so they cannot drift. Position-delete files carry their own per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile and records its own keyMetadata. An unencrypted delete file beside encrypted data would disclose which rows were removed, which is what table encryption exists to prevent. Fails closed throughout: - a key that cannot be recovered fails the query; never falls back to plaintext - an encrypted footer with no DEK is refused rather than parsed as plaintext - AES_CTR has no PME equivalent and is rejected at planning, not written as GCM - BE never defaults the DEK length. Only FE can see encryption.data-key-length, so FE resolves it using Iceberg's own constant and default; encryption enabled with no length reaching BE is an error. At commit the key BE returns is checked against that length, so a file encrypted below the table's declared strength is refused - a table that declares encryption while the catalog supplies a plaintext EncryptionManager is refused on both read and write. Those signals come from different places -- the property travels with the metadata for every catalog, the manager comes from the catalog's TableOperations -- so they can disagree, and a plaintext manager would otherwise read as "unencrypted table" - whether a table is encrypted is decided from encryption.key-id on both the write signal and at commit. Deciding it from the manager's type on one side and the property on the other would let BE encrypt a file the commit records no key for, leaving data nothing can decrypt - the encrypted page-header length prefix sits outside the AEAD, so it is bounded (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB cap the plaintext path uses) before it is used to size any buffer Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds the DEK in the clear, which is safe only because Iceberg encrypts the manifest carrying it, and that chain arrives in 1.11.0. Table format v3 is required -- EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore catalogs only for now: in 1.11.0 only HiveTableOperations builds a real EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its properties say (apache/iceberg#13225). Because the decision reads table.encryption() rather than a version or catalog name, no StarRocks change is needed when that lands. arrow/parquet-cpp is built with -DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols were already in libparquet.a. Observability: no signal was removed or renamed. Existing Parquet reader/writer scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and refusals surface as query errors naming the table and the reason, so no new metric was added. Key material is never logged or put in a profile. Signed-off-by: Annie Yan <anniey@outlook.com>
Records the design behind Iceberg table encryption: the trust boundaries and where key material is allowed to exist, the write and read flows across FE and BE, and the key-management chain from the table master key down to the per-file DEK. Written down because the constraints are not visible from the code. Why Iceberg 1.11.0 is a correctness floor rather than a version preference, why table format v3 is required, why the feature is limited to Hive Metastore catalogs today and what changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check exists at all -- a table can declare encryption the catalog will not apply, and the two signals come from different places. Also records the gaps that are deliberately out of scope, so the next person does not have to rediscover which are intentional. The diagrams ship with their PlantUML sources and a render script, because they had already gone stale once: a PNG kept asserting behaviour the code no longer had, and with no source in tree there was nothing to check it against. Re-rendering the two untouched diagrams reproduces their committed PNGs byte for byte, so the sources and the images are known to agree. Corrects four claims that no longer matched the code: - the write path decides from the table's declaration (encryption.key-id), not from the EncryptionManager's type, and all three sinks derive that one signal - the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core, so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data. The BE rejects an unknown cipher; it is not a planning-time check - the read gate runs in buildScanRange ahead of any per-file test, not inside buildEncryptionInfo, because such a table's files carry no key metadata and a per-file check would skip them - REST catalogs: upstream fails silently, but StarRocks refuses rather than inheriting that silence Signed-off-by: Annie Yan <anniey@outlook.com>
…ption StarRocks cannot read or write Iceberg tables whose data files are encrypted at rest, so an Iceberg lakehouse that requires encryption cannot include StarRocks. Iceberg specifies this as Parquet Modular Encryption (AES-GCM) under Standard/PME key management: a per-file data encryption key encrypts the Parquet file, and the DEK is recorded in the Iceberg file's key_metadata as StandardKeyMetadata, protected by the encrypted manifest that carries it. JVM engines inherit the data-file half by upgrading Iceberg -- OutputFileFactory hands an EncryptedOutputFile to parquet-java, and EncryptingFileIO wraps the FileIO on read. StarRocks has its own C++ Parquet reader and writer, so that layer has to be built: - BE generates the per-file DEK and applies PME through parquet-cpp on write - the DEK travels back to FE, which serializes it into key_metadata at commit - on read FE recovers the DEK from key_metadata and ships it per scan range; BE decrypts the footer and pages itself The metadata half -- manifest encryption and the KEK chain -- is inherited from the Iceberg library exactly as Spark does, since FE planning and commit already go through it. key_metadata serialization delegates to Iceberg's own (package-private) StandardKeyMetadata rather than being reimplemented, so a StarRocks-written file is byte-identical to what any other Iceberg engine produces. All three write sinks are covered -- IcebergTableSink (INSERT), IcebergDeleteSink (DELETE) and IcebergRowDeltaSink (UPDATE/MERGE) -- and they derive the FE->BE signal from one place so they cannot drift. Position-delete files carry their own per-file key, matching Iceberg: PositionDeleteWriter takes an EncryptedOutputFile and records its own keyMetadata. An unencrypted delete file beside encrypted data would disclose which rows were removed, which is what table encryption exists to prevent. Fails closed throughout: - a key that cannot be recovered fails the query; never falls back to plaintext - an encrypted footer with no DEK is refused rather than parsed as plaintext - AES_CTR has no PME equivalent and is rejected at planning, not written as GCM - BE never defaults the DEK length. Only FE can see encryption.data-key-length, so FE resolves it using Iceberg's own constant and default; encryption enabled with no length reaching BE is an error. At commit the key BE returns is checked against that length, so a file encrypted below the table's declared strength is refused - a table that declares encryption while the catalog supplies a plaintext EncryptionManager is refused on both read and write. Those signals come from different places -- the property travels with the metadata for every catalog, the manager comes from the catalog's TableOperations -- so they can disagree, and a plaintext manager would otherwise read as "unencrypted table" - whether a table is encrypted is decided from encryption.key-id on both the write signal and at commit. Deciding it from the manager's type on one side and the property on the other would let BE encrypt a file the commit records no key for, leaving data nothing can decrypt - the encrypted page-header length prefix sits outside the AEAD, so it is bounded (in 64-bit arithmetic, against both the remaining chunk bytes and the same 16 MB cap the plaintext path uses) before it is used to size any buffer Requires Iceberg 1.11.0 as a correctness floor, not a preference: key_metadata holds the DEK in the clear, which is safe only because Iceberg encrypts the manifest carrying it, and that chain arrives in 1.11.0. Table format v3 is required -- EncryptionUtil.checkCompatibility rejects encryption.key-id below v3. Hive Metastore catalogs only for now: in 1.11.0 only HiveTableOperations builds a real EncryptionManager, so a REST-catalog table reports itself unencrypted whatever its properties say (apache/iceberg#13225). Because the decision reads table.encryption() rather than a version or catalog name, no StarRocks change is needed when that lands. arrow/parquet-cpp is built with -DPARQUET_REQUIRE_ENCRYPTION=ON and the three PME headers installed; the symbols were already in libparquet.a. Observability: no signal was removed or renamed. Existing Parquet reader/writer scan stats and the Iceberg commit metrics cover the encrypted paths unchanged, and refusals surface as query errors naming the table and the reason, so no new metric was added. Key material is never logged or put in a profile. Signed-off-by: Annie Yan <anniey@outlook.com>
Records the design behind Iceberg table encryption: the trust boundaries and where key material is allowed to exist, the write and read flows across FE and BE, and the key-management chain from the table master key down to the per-file DEK. Written down because the constraints are not visible from the code. Why Iceberg 1.11.0 is a correctness floor rather than a version preference, why table format v3 is required, why the feature is limited to Hive Metastore catalogs today and what changes when apache/iceberg#13225 lands, and why the declared-vs-enforced check exists at all -- a table can declare encryption the catalog will not apply, and the two signals come from different places. Also records the gaps that are deliberately out of scope, so the next person does not have to rediscover which are intentional. The diagrams ship with their PlantUML sources and a render script, because they had already gone stale once: a PNG kept asserting behaviour the code no longer had, and with no source in tree there was nothing to check it against. Re-rendering the two untouched diagrams reproduces their committed PNGs byte for byte, so the sources and the images are known to agree. Corrects four claims that no longer matched the code: - the write path decides from the table's declaration (encryption.key-id), not from the EncryptionManager's type, and all three sinks derive that one signal - the cipher is not configurable and never was. Iceberg 1.11.0 has no cipher table property and its EncryptionAlgorithm enum is referenced nowhere in iceberg-core, so FE always signals AES_GCM_V1 -- the only PME mode that authenticates page data. The BE rejects an unknown cipher; it is not a planning-time check - the read gate runs in buildScanRange ahead of any per-file test, not inside buildEncryptionInfo, because such a table's files carry no key metadata and a per-file check would skip them - REST catalogs: upstream fails silently, but StarRocks refuses rather than inheriting that silence Signed-off-by: Annie Yan <anniey@outlook.com>
|
@singhpk234 thank you for the implementation! I've re-iterated some high-level thoughts in #18081 (review), happy to work on integrating that implementation to this PR (or whatever works) once we're aligned on the approach there - it generally makes sense to me with some discussions I think worth having as I mentioned above - and +1 on not rushing the implementation here. LMKWYT! |
This PR implements client-side support for REST catalog encryption. With it, clients interacting with a REST catalog can read and write encrypted data.
It is similar to #13066, that integrates encryption with the Hive catalog.
cc @rdblue @RussellSpitzer @ggershinsky