Skip to content

feat(obfuscation)!: make json transforms caller-provided - #2548

Merged
gh-worker-dd-mergequeue-cf854d[bot] merged 14 commits into
mainfrom
m/json-obfuscation-api
Sep 24, 2026
Merged

gh-worker-dd-mergequeue-cf854d[bot] merged 14 commits into
mainfrom
m/json-obfuscation-api

Conversation

@webern

@webern webern commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Human Summary

While working on DataDog/saluki#2601, I discovered that I couldn’t use libdatadog’s JSON obfuscator as written. Its transformer was a plain function pointer whose only argument was the JSON value. Unlike a closure, a function pointer cannot capture runtime state, so there was no way to pass Saluki’s dynamically loaded SQL obfuscation configuration without putting it in static or global state. Existing users can configure ordinary SQL obfuscation, but the existing JSON transformer API provides no normal way to pass those SQL settings when transforming SQL embedded in JSON. The provided transformer uses default settings.

In fixing this, I decided to take it a little further. Transform failures are now returned as structured information, scan failures use a typed error, and the hot-path API accepts caller-owned output and scratch buffers so repeated calls can reuse their allocations. Transformations return Cow<str>, which also avoids forcing an allocation when they can return borrowed text.

The coding agent brought in thiserror as a dependency. This is idiomatic and seems like a good call. The scanner's new const fns are clippy::missing_const_for_fn, which the workspace enables; they have no runtime effect. Scanner::new becoming restart is for buffer reuse. The parse-state stack now survives between passes.

What does this PR do?

This separates JSON obfuscation policy from transformation behavior. JsonObfuscatorConfig now contains only serializable, comparable data. Callers provide transforms per call through generic closures that can capture runtime SQL configuration.

The allocating API is joined by obfuscate_into, which reuses caller-owned output and scratch buffers. Scratch capacity is observable and explicitly trimmable so callers can choose their memory-retention policy.

JSON scan errors and transform errors are reported separately. SQL obfuscation now reports an empty result as SqlObfuscationError::EmptyResult, and the crate exports the Agent's exact SQL failure replacement. This also pins non-ASCII identifier behavior and fixes the JSON unquoting fallback to retain the literal when unquoting fails.

Motivation

The stored fn(&str) -> String transformer cannot capture runtime configuration or report errors. This prevents Agent-compatible configured SQL transformation inside JSON values.

DataDog/saluki#2601 currently duplicates libdatadog's JSON scanner to work around that API. A call-scoped closure lets Saluki use the shared scanner while retaining its configured SQL behavior, logging, and failure policy.

Additional Notes

Transform callbacks return Cow<str>, allowing borrowed, static, or owned replacements. The callback is generic and allocation-free; no transformer trait or stored trait object is introduced.

A transform error replaces that value with "?", records the error, and continues. Callers that need the Agent's SQL fallback can observe and log the SQL error inside the callback, then return SQL_OBFUSCATION_FAILURE_REPLACEMENT.

JsonObfuscationScratch retains capacity from the largest prior input until the caller trims or drops it. retained_capacity and trim_to make that policy explicit.

Rebase onto #2490

#2490 moved SqlObfuscateConfig, SqlObfuscationMode and DbmsKind from sql into obfuscation_config and renamed the config struct to SqlConfig, and it began deserializing JsonObfuscator directly from the Agent's /info payload. Resolved as follows:

  • Adopted obfuscation_config::{DbmsKind, SqlConfig, SqlObfuscationMode} everywhere, including the obfuscate_with doc example. No type is reintroduced in sql.
  • JsonObfuscatorConfig keeps #[serde(default)] but not deny_unknown_fields: feat(data-pipeline)!: refactor agent's /info obfuscation config format #2490 dropped it deliberately, and an /info payload from a newer Agent carries fields this struct has no counterpart for. A test pins that forward compatibility.
  • feat(data-pipeline)!: refactor agent's /info obfuscation config format #2490's hand-written PartialEq for JsonObfuscatorConfig, which existed only to skirt the uncomparable transformer field, is replaced by a derive now that the field is gone. The config also derives Eq.
  • obfuscate_resource_for_stats and obfuscate_pb_span now consume the Result from obfuscate_sql: a resource that obfuscates to nothing is left as sent rather than blanked.

Agent /info field names

Checked against the Agent rather than guessed. pkg/trace/api/info.go serves a reduced view:

type reducedJSONObfuscationConfig struct {
	Enabled  bool     `json:"enabled"`
	KeepKeys []string `json:"keep_keys"`
}
...
oconf.Elasticsearch = reducedJSONObfuscationConfig{Enabled: o.ES.Enabled, KeepKeys: o.ES.KeepValues}

So /info sends keep_keys, which is already this crate's field name, and it does not report the transform set at all. The Agent's own obfuscation config (pkg/obfuscate, apm_config.obfuscation.*) calls the same two sets keep_values and obfuscate_sql_values. Rather than rename the Rust fields — transform_keys is no longer SQL-specific here, so obfuscate_sql_values would be a lie — both Agent spellings are accepted as #[serde(alias = ...)], in the same style as #2490's PascalCase aliases. obfuscation_config::tests::test_agent_json_obfuscation_field_names pins all of it.

BREAKING CHANGE: JsonObfuscatorConfig::transformer and JsonStringTransformer are removed. JSON transforms move to JsonObfuscator::obfuscate_with or JsonObfuscator::obfuscate_into. SQL obfuscation functions now return Result.

How to test the change?

The following pass on the rebased branch:

cargo +stable clippy -p libdd-trace-obfuscation --all-targets -- -D warnings
cargo +nightly-2026-07-26 fmt --all -- --check
cargo nextest run -p libdd-trace-obfuscation           # 387 passed
cargo test -p libdd-trace-obfuscation --doc
cargo check -p libdd-trace-stats -p libdd-data-pipeline-core \
  -p libdd-data-pipeline -p libdd-data-pipeline-ffi --all-targets
cargo nextest run -p libdd-data-pipeline -p libdd-trace-stats \
  -E '!test(tracing_integration_tests::)'              # 253 passed

The tracing_integration_tests:: suite needs Docker and was not run locally. Cargo.lock gains only thiserror, which LICENSE-3rdparty.csv already covers, so no regeneration was needed. cargo deny check still reports pre-existing workspace advisory and license-policy failures unrelated to this diff.

References

@webern webern added breaking-change AI Generated PR largely written by AI tools labels Sep 18, 2026
@datadog-datadog-prod-us1

datadog-datadog-prod-us1 Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Tests

✅ All CI checks and tests passed.

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

🎯 Code Coverage (details)
• Patch Coverage: 98.85%
• Overall Coverage: 78.61% (+0.31%)

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: c60e9e5 | Docs | View more details | Give us feedback!

@dd-octo-sts

dd-octo-sts Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Artifact Size Benchmark Report

aarch64-alpine-linux-musl
Artifact Baseline Commit Change
/aarch64-alpine-linux-musl/lib/libdatadog_profiling.so 9.02 MB 9.02 MB 0% (0 B) 👌
/aarch64-alpine-linux-musl/lib/libdatadog_profiling.a 95.97 MB 95.99 MB +.02% (+24.14 KB) 🔍
aarch64-unknown-linux-gnu
Artifact Baseline Commit Change
/aarch64-unknown-linux-gnu/lib/libdatadog_profiling.so 12.19 MB 12.19 MB +0% (+184 B) 👌
/aarch64-unknown-linux-gnu/lib/libdatadog_profiling.a 107.35 MB 107.37 MB +.01% (+21.71 KB) 🔍
libdatadog-x64-windows
Artifact Baseline Commit Change
/libdatadog-x64-windows/debug/dynamic/datadog_profiling_ffi.dll 29.07 MB 29.08 MB +.03% (+9.00 KB) 🔍
/libdatadog-x64-windows/debug/dynamic/datadog_profiling_ffi.lib 96.08 KB 96.08 KB 0% (0 B) 👌
/libdatadog-x64-windows/debug/dynamic/datadog_profiling_ffi.pdb 191.49 MB 191.55 MB +.03% (+64.00 KB) 🔍
/libdatadog-x64-windows/debug/static/datadog_profiling_ffi.lib 815.92 MB 817.32 MB +.17% (+1.40 MB) 🔍
/libdatadog-x64-windows/release/dynamic/datadog_profiling_ffi.dll 9.71 MB 9.72 MB +.11% (+11.00 KB) 🔍
/libdatadog-x64-windows/release/dynamic/datadog_profiling_ffi.lib 96.08 KB 96.08 KB 0% (0 B) 👌
/libdatadog-x64-windows/release/dynamic/datadog_profiling_ffi.pdb 27.50 MB 27.51 MB +.05% (+16.00 KB) 🔍
/libdatadog-x64-windows/release/static/datadog_profiling_ffi.lib 55.60 MB 55.62 MB +.02% (+13.81 KB) 🔍
libdatadog-x86-windows
Artifact Baseline Commit Change
/libdatadog-x86-windows/debug/dynamic/datadog_profiling_ffi.dll 25.42 MB 25.43 MB +.03% (+8.00 KB) 🔍
/libdatadog-x86-windows/debug/dynamic/datadog_profiling_ffi.lib 97.58 KB 97.58 KB 0% (0 B) 👌
/libdatadog-x86-windows/debug/dynamic/datadog_profiling_ffi.pdb 196.70 MB 196.75 MB +.02% (+48.00 KB) 🔍
/libdatadog-x86-windows/debug/static/datadog_profiling_ffi.lib 803.55 MB 802.87 MB --.08% (-699.04 KB) 💪
/libdatadog-x86-windows/release/dynamic/datadog_profiling_ffi.dll 7.52 MB 7.55 MB +.39% (+30.50 KB) 🔍
/libdatadog-x86-windows/release/dynamic/datadog_profiling_ffi.lib 97.58 KB 97.58 KB 0% (0 B) 👌
/libdatadog-x86-windows/release/dynamic/datadog_profiling_ffi.pdb 29.61 MB 29.62 MB +.05% (+16.00 KB) 🔍
/libdatadog-x86-windows/release/static/datadog_profiling_ffi.lib 52.55 MB 52.56 MB +.02% (+14.82 KB) 🔍
x86_64-alpine-linux-musl
Artifact Baseline Commit Change
/x86_64-alpine-linux-musl/lib/libdatadog_profiling.a 85.96 MB 85.98 MB +.02% (+25.81 KB) 🔍
/x86_64-alpine-linux-musl/lib/libdatadog_profiling.so 10.05 MB 10.05 MB +.03% (+4.00 KB) 🔍
x86_64-unknown-linux-gnu
Artifact Baseline Commit Change
/x86_64-unknown-linux-gnu/lib/libdatadog_profiling.a 101.83 MB 101.85 MB +.02% (+23.14 KB) 🔍
/x86_64-unknown-linux-gnu/lib/libdatadog_profiling.so 12.27 MB 12.28 MB +.03% (+4.20 KB) 🔍

@webern
webern force-pushed the m/json-obfuscation-api branch 2 times, most recently from b36d1da to e673c33 Compare September 18, 2026 10:58
@pr-commenter

pr-commenter Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Benchmarks

Comparison

Benchmark execution time: 2026-09-24 12:30:10

Comparing candidate commit c60e9e5 in PR branch m/json-obfuscation-api with baseline commit a5e7164 in branch main.

📊 Benchmarking dashboard

Found 0 performance improvements and 1 performance regressions! Performance is the same for 176 metrics, 0 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:vec_map/as_deduped_map/needs_dedup_1_in_4/128

  • 🟥 execution_time [+631.589ns; +636.434ns] or [+10.333%; +10.413%]

Benchmark execution time: 2026-09-24 12:23:55

Comparing candidate commit c60e9e5 in PR branch m/json-obfuscation-api with baseline commit a5e7164 in branch main.

📊 Benchmarking dashboard

Found 22 performance improvements and 6 performance regressions! Performance is the same for 137 metrics, 11 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:credit_card/is_card_number/ 3782-8224-6310-005

  • 🟩 execution_time [-4.027µs; -3.792µs] or [-5.002%; -4.710%]
  • 🟩 throughput [+615529.253op/s; +655013.709op/s] or [+4.955%; +5.273%]

scenario:credit_card/is_card_number/ 378282246310005

  • 🟩 execution_time [-5.443µs; -5.351µs] or [-7.362%; -7.237%]
  • 🟩 throughput [+1056274.670op/s; +1073344.772op/s] or [+7.809%; +7.936%]

scenario:credit_card/is_card_number/378282246310005

  • 🟩 execution_time [-5.870µs; -5.802µs] or [-8.287%; -8.191%]
  • 🟩 throughput [+1260532.004op/s; +1274728.188op/s] or [+8.928%; +9.029%]

scenario:credit_card/is_card_number/37828224631000521389798

  • 🟩 execution_time [-7.292µs; -7.261µs] or [-13.734%; -13.675%]
  • 🟩 throughput [+2983857.672op/s; +2997897.623op/s] or [+15.844%; +15.918%]

scenario:credit_card/is_card_number_no_luhn/ 378282246310005

  • 🟩 execution_time [-4.207µs; -4.165µs] or [-7.197%; -7.125%]
  • 🟩 throughput [+1313188.844op/s; +1326175.760op/s] or [+7.676%; +7.751%]

scenario:credit_card/is_card_number_no_luhn/378282246310005

  • 🟩 execution_time [-5.043µs; -4.989µs] or [-9.077%; -8.980%]
  • 🟩 throughput [+1776686.439op/s; +1796257.248op/s] or [+9.870%; +9.979%]

scenario:credit_card/is_card_number_no_luhn/37828224631000521389798

  • 🟩 execution_time [-7.307µs; -7.277µs] or [-13.759%; -13.703%]
  • 🟩 throughput [+2990485.817op/s; +3004018.377op/s] or [+15.881%; +15.953%]

scenario:thread_cpu/alloc_free/system/16

  • 🟥 execution_time [+3.507ns; +3.541ns] or [+24.490%; +24.723%]

scenario:thread_cpu/profiler_attached/fast_path_noop/16

  • 🟥 execution_time [+2.851ns; +2.856ns] or [+19.701%; +19.742%]

scenario:thread_cpu/profiler_attached/fast_path_noop/256

  • 🟥 execution_time [+2.850ns; +2.858ns] or [+19.698%; +19.752%]

scenario:thread_cpu/profiler_attached/fast_path_noop/4096

  • 🟥 execution_time [+2.846ns; +2.853ns] or [+19.670%; +19.720%]

scenario:thread_cpu/profiler_attached/fast_path_noop/64

  • 🟥 execution_time [+2.850ns; +2.857ns] or [+19.703%; +19.748%]

scenario:thread_cpu/profiler_attached/fast_path_noop/65536

  • 🟥 execution_time [+2.854ns; +2.863ns] or [+19.727%; +19.787%]

scenario:thread_cpu/profiler_attached/fast_path_system/16

  • 🟩 execution_time [-16.251ns; -16.194ns] or [-36.420%; -36.290%]

scenario:thread_cpu/profiler_attached/fast_path_system/256

  • 🟩 execution_time [-15.978ns; -15.915ns] or [-36.011%; -35.869%]

scenario:thread_cpu/profiler_attached/fast_path_system/4096

  • 🟩 execution_time [-8.441ns; -8.293ns] or [-7.378%; -7.248%]

scenario:thread_cpu/profiler_attached/fast_path_system/64

  • 🟩 execution_time [-15.940ns; -15.873ns] or [-35.943%; -35.792%]

scenario:thread_cpu/profiler_attached/fast_path_system/65536

  • 🟩 execution_time [-9.492ns; -9.369ns] or [-8.499%; -8.389%]

scenario:thread_cpu/profiler_attached/slow_path_system/4096

  • 🟩 execution_time [-9.090ns; -8.989ns] or [-5.894%; -5.829%]

scenario:trace_buffer/8_senders/no_delay

  • 🟩 execution_time [-591.856µs; -399.660µs] or [-9.630%; -6.503%]
  • 🟩 throughput [+90343.467op/s; +138512.274op/s] or [+7.683%; +11.779%]

Unstable benchmarks

These benchmarks have a confidence interval too wide to call a change; treat them as noise rather than signal.

scenario:datadog_sample_span/parent_not_sampled_short_circuit/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+555.735%; -555.735%]

scenario:datadog_sample_span/parent_sampled_short_circuit/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+556.312%; -556.007%]

scenario:flagevaluation_evp/coalescer/typical/100flags_50users_10fields

  • unstable execution_time [-6.890µs; +13.229µs] or [-3.724%; +7.151%]

scenario:glob_matcher/ascii_case_insensitive_match/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+554.044%; -554.940%]

scenario:glob_matcher/ascii_exact_match/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+555.735%; -555.735%]

scenario:glob_matcher/ascii_exact_miss/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+555.251%; -555.507%]

scenario:glob_matcher/ascii_wildcard_backtrack_match/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+557.291%; -556.468%]

scenario:glob_matcher/ascii_wildcard_heavy_backtrack/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+557.442%; -556.539%]

scenario:glob_matcher/ascii_wildcard_question_match/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+558.051%; -556.826%]

scenario:glob_matcher/ascii_wildcard_star_match/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+556.503%; -556.096%]

scenario:glob_matcher/star_short_circuit/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+556.409%; -556.052%]

Candidate

Omitted due to size.

Baseline

Omitted due to size.

@webern
webern force-pushed the m/json-obfuscation-api branch 2 times, most recently from 30e9dc1 to ec9c1c5 Compare September 21, 2026 09:26
JsonObfuscatorConfig stored a `fn(&str) -> String` transformer, so a
caller with runtime configuration to capture could not use it, and the
config was neither serializable nor comparable. Remove the transformer
field and the `JsonStringTransformer` alias: the config is plain data
now, including `transform_keys`, which is no longer `#[serde(skip)]`,
and `PartialEq` is derived rather than written by hand around a field
that could not be compared.

Now that the config is deserialized straight from the Agent's `/info`
payload, accept the Agent's own spellings too: `keep_values` for
`keep_keys` and `obfuscate_sql_values` for `transform_keys`. `/info`
sends `keep_keys` and does not report the transform set at all, so
either spelling parses and neither is required.

The transform is passed per call instead, as a method-scoped generic:

  obfuscate_with(input, |v| ...)          allocating convenience
  obfuscate_into(input, out, scratch, f)  caller-owned buffers

The callback is `for<'a> FnMut(&'a str) -> Result<Cow<'a, str>, E>`, so
it can capture anything, needs no Send/Sync/'static, and can return a
borrowed, static or owned replacement with no forced allocation.
JsonObfuscationScratch holds the reusable working memory, and reports
and trims what it retains so reuse cannot silently pin memory.

Failures are reported in two channels through JsonObfuscationReport: a
JSON scan error, whose partial output is still usable and ends in
`...`, and the errors of individual transforms. Nothing is logged
behind the caller's back.

obfuscate_sql and its wrappers now return Result. An empty result is a
failure, as it is in the Agent, and SQL_OBFUSCATION_FAILURE_REPLACEMENT
exposes the Agent's replacement text so callers do not split stats
aggregation by inventing their own. Fixes #2541, whose non-ASCII
identifier case is now pinned by a test.

Also match the Agent's JSON unquoting fallback, which keeps the quotes
of a value it cannot unquote rather than stripping them, and hand
values that carry no escapes to the callback as slices of the input.
@webern
webern force-pushed the m/json-obfuscation-api branch from ec9c1c5 to c25301a Compare September 21, 2026 18:15
@webern
webern marked this pull request as ready for review September 22, 2026 09:14
@webern
webern requested review from a team as code owners September 22, 2026 09:14

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c25301ab49

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread libdd-trace-obfuscation/src/json/mod.rs Outdated
Comment thread libdd-trace-obfuscation/src/obfuscate.rs Outdated
@webern
webern marked this pull request as draft September 22, 2026 10:08
Use the Agent's non-parsable marker when a non-empty SQL or Cassandra resource obfuscates to nothing. Apply the same policy to protobuf spans, v0.4 spans, and client-side stats so none of those paths can fall back to forwarding comment text as sent.\n\nKeep the JSON transform failure marker distinct and add regression coverage for whitespace-only and comment-only resources.
Escape quotes, backslashes, and control characters before embedding callback replacements in JSON output. Cover both transform entry points and the SQL quoted-identifier configuration that exposed the invalid output.
@webern
webern marked this pull request as ready for review September 22, 2026 12:23

@Eldolfin Eldolfin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good 👍🏽

One thing I saw repeated multiple times in this PR is references of types and files from the agent repo which I think we should avoid because they will get stale and don't add much value IMO

Comment thread libdd-trace-obfuscation/src/sql.rs Outdated
Comment thread libdd-trace-obfuscation/src/sql.rs Outdated
Comment thread libdd-trace-obfuscation/src/sql.rs Outdated
Comment thread libdd-trace-obfuscation/src/obfuscate.rs Outdated
Comment thread libdd-trace-obfuscation/src/json/mod.rs Outdated
`Pass` tracked what it was doing with four booleans, a literal range and a
depth, so states that cannot occur - a key that is also a transformed value, a
wiped value inside a kept subtree, a stale literal range - were representable
and had to be guarded against by hand.

Rename it to `ParserState` and replace those fields with one `ParserPhase`.
Each phase carries only the state it uses: the byte range of the literal in
hand, and the `KeptSubtree` the pass sits inside, if any. Keeping outlives the
phase that reads one key or value, so a transform key inside a kept subtree
still rewrites its own value, as before.

The flag assignments scattered through the loop become transitions:
`begin_next_token`, `begin_literal`, `finish_key` and a `finish_value` that
returns the kept subtree that is still open. The `wiped` flag is now the
`AwaitingValue` to `ObfuscatedValue` transition, and the defensive
`!transforming` guard at a key is satisfied by the phase itself.

No behavior change: the output of every entry point is byte for byte what it
was, including the keep and transform precedence and the top-level literal
quirks, which gain tests here.
@webern
webern requested a review from a team as a code owner September 24, 2026 11:25
@gh-worker-ownership-write-b05516
gh-worker-ownership-write-b05516 Bot removed the request for review from a team September 24, 2026 11:45
@webern
webern requested a review from Eldolfin September 24, 2026 11:46
@gh-worker-dd-mergequeue-cf854d
gh-worker-dd-mergequeue-cf854d Bot merged commit 9eed262 into main Sep 24, 2026
101 checks passed
@gh-worker-dd-mergequeue-cf854d
gh-worker-dd-mergequeue-cf854d Bot deleted the m/json-obfuscation-api branch September 24, 2026 14:24
iunanua added a commit that referenced this pull request Oct 2, 2026
…ibdd-data-pipeline, libdd-li... (#2619)

<!-- release-proposal-inputs: {"crates":"libdd-capabilities-impl,
libdd-common, libdd-data-pipeline, libdd-library-config,
libdd-profiling-heap-allocator, libdd-remote-config, libdd-sampling,
libdd-shared-runtime, libdd-telemetry, libdd-tinybytes,
libdd-trace-utils","main_start_ref":"","level_overrides":"","bypass_standard_checks":false}
-->

# Release proposal for libdd-capabilities-impl, libdd-common,
libdd-data-pipeline, libdd-library-config,
libdd-profiling-heap-allocator, libdd-remote-config, libdd-sampling,
libdd-shared-runtime, libdd-telemetry, libdd-tinybytes,
libdd-trace-utils and their dependencies

This PR contains version bumps based on public API changes and commits
since last release.


### ⚠️ Crates left out of this proposal affected by its major
bumps

These publishable workspace crates are not part of this release but
their dependency requirement was rewritten on this branch while their
published version still requires the old major. If they are a dependency
on your deployment not including them in the release could result in
duplicate packages or symbol incompatibility.

- `libdd-capabilities-impl` `6.0.0` → `7.0.0` affects:
`libdd-crashtracker`, `libdd-live-debugger`, `libdd-tracer-flare`
- `libdd-common` `7.0.0` → `8.0.0` affects: `libdd-crashtracker`,
`libdd-ffe`, `libdd-http-client`, `libdd-ipc`, `libdd-live-debugger`,
`libdd-profiling`, `libdd-tracer-flare`
- `libdd-data-pipeline` `11.0.0` → `12.0.0` affects:
`libdd-live-debugger`
- `libdd-remote-config` `6.0.0` → `7.0.0` affects: `libdd-ffe`,
`libdd-live-debugger`, `libdd-tracer-flare`
- `libdd-telemetry` `9.0.0` → `10.0.0` affects: `libdd-crashtracker`
- `libdd-trace-stats` `10.0.0` → `11.0.0` affects: `libdd-ipc`
- `libdd-trace-utils` `13.0.0` → `14.0.0` affects: `libdd-tracer-flare`

## libdd-capabilities
**Next version:** `4.0.1`
**Semver bump:** `patch`
**Tag:** `libdd-capabilities-v4.0.1`

### Commits

- build: Update workspace to Rust 2024 edition (#2575)

## libdd-common
**Next version:** `8.0.0`
**Semver bump:** `major`
**Tag:** `libdd-common-v8.0.0`

### Commits

- fix(ipc)!: use atomic deadlines for shared limiters (#2604)
- feat(sidecar)!: Authenticate sidecar connections and shared memory
(#2551)
- build: Update workspace to Rust 2024 edition (#2575)
- feat(trace_utils)!: add mutable metadata (#2545)

## libdd-ddsketch
**Next version:** `1.1.3`
**Semver bump:** `patch`
**Tag:** `libdd-ddsketch-v1.1.3`

### Commits

- build: Update workspace to Rust 2024 edition (#2575)

## libdd-profiling-heap-sampler
**Next version:** `1.1.1`
**Semver bump:** `patch`
**Tag:** `libdd-profiling-heap-sampler-v1.1.1`

### Commits

- build: Update workspace to Rust 2024 edition (#2575)

## libdd-tinybytes
**Next version:** `1.1.5`
**Semver bump:** `patch`
**Tag:** `libdd-tinybytes-v1.1.5`

### Commits

- build: Update workspace to Rust 2024 edition (#2575)

## libdd-trace-protobuf
**Next version:** `5.1.0`
**Semver bump:** `minor`
**Tag:** `libdd-trace-protobuf-v5.1.0`

### Commits

- fix(data-pipeline)!: revert changes that made /info un-parsable
(#2586)
- build: Update workspace to Rust 2024 edition (#2575)

## libdd-capabilities-impl
**Next version:** `7.0.0`
**Semver bump:** `major`
**Tag:** `libdd-capabilities-impl-v7.0.0`

### ⚠️ major bump forced due to:

- `libdd-common`: ^7.0.0 → ^8.0.0

### Commits

- feat(sidecar)!: Authenticate sidecar connections and shared memory
(#2551)
- build: Update workspace to Rust 2024 edition (#2575)

## libdd-profiling-heap-allocator
**Next version:** `1.2.0`
**Semver bump:** `minor`
**Tag:** `libdd-profiling-heap-allocator-v1.2.0`

### Commits

- build: Update workspace to Rust 2024 edition (#2575)

## libdd-library-config
**Next version:** `4.2.0`
**Semver bump:** `minor`
**Tag:** `libdd-library-config-v4.2.0`

### Commits

- build: Update workspace to Rust 2024 edition (#2575)

## libdd-trace-normalization
**Next version:** `4.2.0`
**Semver bump:** `minor`
**Tag:** `libdd-trace-normalization-v4.2.0`

### Commits

- build: Update workspace to Rust 2024 edition (#2575)

## libdd-remote-config
**Next version:** `7.0.0`
**Semver bump:** `major`
**Tag:** `libdd-remote-config-v7.0.0`

### ⚠️ major bump forced due to:

- `libdd-common`: ^7.0.0 → ^8.0.0

### Commits

- feat(sidecar)!: Authenticate sidecar connections and shared memory
(#2551)
- build: Update workspace to Rust 2024 edition (#2575)

## libdd-shared-runtime
**Next version:** `6.0.0`
**Semver bump:** `major`
**Tag:** `libdd-shared-runtime-v6.0.0`

### ⚠️ major bump forced due to:

- `libdd-common`: ^7.0.0 → ^8.0.0

### Commits

- build: Update workspace to Rust 2024 edition (#2575)

## libdd-trace-utils
**Next version:** `14.0.0`
**Semver bump:** `major`
**Tag:** `libdd-trace-utils-v14.0.0`

### ⚠️ major bump forced due to:

- `libdd-common`: ^7.0.0 → ^8.0.0

### Commits

- feat(trace-utils): add _dd.sdk.otlp_export and datadog.sdk.semantics
OTLP resource attributes (#2603)
- build: Update workspace to Rust 2024 edition (#2575)
- feat(trace_utils)!: add mutable metadata (#2545)

## libdd-dogstatsd-client
**Next version:** `8.0.0`
**Semver bump:** `major`
**Tag:** `libdd-dogstatsd-client-v8.0.0`

### ⚠️ major bump forced due to:

- `libdd-common`: ^7.0.0 → ^8.0.0

### Commits

- feat(sidecar)!: Authenticate sidecar connections and shared memory
(#2551)
- build: Update workspace to Rust 2024 edition (#2575)

## libdd-telemetry
**Next version:** `10.0.0`
**Semver bump:** `major`
**Tag:** `libdd-telemetry-v10.0.0`

### ⚠️ major bump forced due to:

- `libdd-common`: ^7.0.0 → ^8.0.0

### Commits

- build: Update workspace to Rust 2024 edition (#2575)
- feat(telemetry)!: Use mutable metadata (#2552)

## libdd-sampling
**Next version:** `8.0.0`
**Semver bump:** `major`
**Tag:** `libdd-sampling-v8.0.0`

### ⚠️ major bump forced due to:

- `libdd-common`: ^7.0.0 → ^8.0.0
- `libdd-trace-utils`: ^13.0.0 → ^14.0.0

### Commits

- build: Update workspace to Rust 2024 edition (#2575)

## libdd-trace-obfuscation
**Next version:** `10.0.0`
**Semver bump:** `major`
**Tag:** `libdd-trace-obfuscation-v10.0.0`

### ⚠️ major bump forced due to:

- `libdd-common`: ^7.0.0 → ^8.0.0
- `libdd-trace-utils`: ^13.0.0 → ^14.0.0

### Commits

- fix(data-pipeline)!: revert changes that made /info un-parsable
(#2586)
- build: Update workspace to Rust 2024 edition (#2575)
- feat(obfuscation)!: make json transforms caller-provided (#2548)

## libdd-trace-stats
**Next version:** `11.0.0`
**Semver bump:** `major`
**Tag:** `libdd-trace-stats-v11.0.0`

### ⚠️ major bump forced due to:

- `libdd-common`: ^7.0.0 → ^8.0.0
- `libdd-telemetry`: ^9.0.0 → ^10.0.0
- `libdd-trace-obfuscation`: ^9.0.0 → ^10.0.0
- `libdd-trace-utils`: ^13.0.0 → ^14.0.0

### Commits

- feat(sidecar)!: Authenticate sidecar connections and shared memory
(#2551)
- build: Update workspace to Rust 2024 edition (#2575)
- feat(trace_utils)!: add mutable metadata (#2545)
- fix(stats): fix precedence for http endpoint (#2582)

## libdd-data-pipeline-core
**Next version:** `3.0.0`
**Semver bump:** `major`
**Tag:** `libdd-data-pipeline-core-v3.0.0`

### ⚠️ major bump forced due to:

- `libdd-common`: ^7.0.0 → ^8.0.0
- `libdd-trace-obfuscation`: ^9.0.0 → ^10.0.0
- `libdd-trace-stats`: ^10.0.0 → ^11.0.0
- `libdd-trace-utils`: ^13.0.0 → ^14.0.0

### Commits

- build: Update workspace to Rust 2024 edition (#2575)
- feat(trace_utils)!: add mutable metadata (#2545)

## libdd-data-pipeline
**Next version:** `12.0.0`
**Semver bump:** `major`
**Tag:** `libdd-data-pipeline-v12.0.0`

### ⚠️ major bump forced due to:

- `libdd-common`: ^7.0.0 → ^8.0.0
- `libdd-telemetry`: ^9.0.0 → ^10.0.0
- `libdd-trace-obfuscation`: ^9.0.0 → ^10.0.0
- `libdd-trace-stats`: ^10.0.0 → ^11.0.0
- `libdd-trace-utils`: ^13.0.0 → ^14.0.0

### Commits

- feat(trace-utils): add _dd.sdk.otlp_export and datadog.sdk.semantics
OTLP resource attributes (#2603)
- fix(data-pipeline)!: revert changes that made /info un-parsable
(#2586)
- build: Update workspace to Rust 2024 edition (#2575)
- feat(telemetry)!: Use mutable metadata (#2552)
- feat(trace_utils)!: add mutable metadata (#2545)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: iunanua <18325288+iunanua@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Obfuscation: an empty SQL obfuscation result is reported as success where the Agent fails, and non-ASCII identifier support is untested

2 participants