feat(stats)!: add whole key cardinality limit [APMSP-3568] - #2158
Conversation
This stack of pull requests is managed by Graphite. Learn more about stacking. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a17a2c66d6
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
Artifact Size Benchmark Reportaarch64-alpine-linux-musl
aarch64-unknown-linux-gnu
libdatadog-x64-windows
libdatadog-x86-windows
x86_64-alpine-linux-musl
x86_64-unknown-linux-gnu
|
a17a2c6 to
2655975
Compare
🎉 All green!🧪 All tests passed 🎯 Code Coverage (details) 🔗 Commit SHA: 8d22e4b | Docs | Datadog PR Page | Give us feedback! |
Clippy Allow Annotation ReportComparing clippy allow annotations between branches:
Summary by Rule
Annotation Counts by File
Annotation Stats by Crate
About This ReportThis report tracks Clippy allow annotations for specific rules, showing how they've changed in this PR. Decreasing the number of these annotations generally improves code quality. |
ce821ea to
ebd7e4e
Compare
ebd7e4e to
b04c7f6
Compare
b04c7f6 to
8d22e4b
Compare
|
View all feedbacks in Devflow UI.
The expected merge time in
|
# What does this PR do? Add telemetry for the collapsed stats group based on [this RFC](https://datadoghq.atlassian.net/wiki/spaces/APM/pages/6821151019/PENDING+Cardinality+Limits) # Motivation Follow-up of #2158 # Additional Notes Anything else we should know when reviewing? # How to test the change? Describe here in detail how the change can be validated. Co-authored-by: vianney.ruhlmann <vianney.ruhlmann@datadoghq.com>
…ibdd-data-pipeline, libdd-li... (#2201) # Release proposal for libdd-capabilities-impl, libdd-common, libdd-data-pipeline, libdd-library-config, libdd-remote-config, libdd-sampling, libdd-telemetry, libdd-tinybytes, libdd-trace-utils and their dependencies This PR contains version bumps based on public API changes and commits since last release. ## libdd-capabilities **Next version:** `2.1.0` **Semver bump:** `minor` **Tag:** `libdd-capabilities-v2.1.0` ### Commits - feat(data-pipeline)!: add stdout log trace exporter (#2074) ## libdd-common **Next version:** `5.1.0` **Semver bump:** `minor` **Tag:** `libdd-common-v5.1.0` ### Commits - refactor(clippy): prefer core and alloc imports (#2196) - fix: update rustls-webpki to 0.103.13 (#2187) - fix: update anyhow for unsoundness (#2186) - feat(machine id): Add helpers in ddcommon to fetch the machine UUID l… (#2163) ## libdd-ddsketch **Next version:** `1.1.0` **Semver bump:** `minor` **Tag:** `libdd-ddsketch-v1.1.0` ### Commits - feat(data-pipeline)!: export client-computed span stats as OTLP trace metrics (#2067) - test(ddsketch): add microbenchmarks for add/encode/collapse (#2125) ## libdd-trace-protobuf **Next version:** `4.0.0` **Semver bump:** `major` **Tag:** `libdd-trace-protobuf-v4.0.0` ### Commits - chore!: update protobufs to be in sync with datadog-agent (#2180) - feat(stats)!: add whole key cardinality limit (#2158) - feat(remote-config)!: use the proto file from the agent (#2165) - feat(data-pipeline): OTLP HTTP/protobuf trace export (#2115) ## libdd-capabilities-impl **Next version:** `3.0.0` **Semver bump:** `major` **Tag:** `libdd-capabilities-impl-v3.0.0` ###⚠️ major bump forced due to: - `libdd-common`: ^4.1.0 → ^5.1.0 ### Commits - feat(data-pipeline)!: add stdout log trace exporter (#2074) ## libdd-library-config **Next version:** `3.0.0` **Semver bump:** `major` **Tag:** `libdd-library-config-v3.0.0` ###⚠️ major bump forced due to: - `libdd-trace-protobuf`: ^3.0.2 → ^4.0.0 ### Commits - refactor(clippy): prefer core and alloc imports (#2196) - feat(library-config)!: caller-supplied threadlocal schema and extra process-context attributes (#2162) - fix(otel-thread-ctx): put the threadlocal attributes at the right place in the context (#2167) ## libdd-remote-config **Next version:** `2.0.0` **Semver bump:** `major` **Tag:** `libdd-remote-config-v2.0.0` ###⚠️ major bump forced due to: - `libdd-trace-protobuf`: ^3.0.2 → ^4.0.0 ### Commits - refactor(libdd-remote-config)!: hide Target inner properties so they are not leaked (#2182) - feat(remote-config)!: use the proto file from the agent (#2165) - refactor(rc): reexport Endpoint and Tag common types (#2147) ## libdd-trace-normalization **Next version:** `3.0.0` **Semver bump:** `major` **Tag:** `libdd-trace-normalization-v3.0.0` ###⚠️ major bump forced due to: - `libdd-trace-protobuf`: ^3.0.1 → ^4.0.0 ### Commits - feat(data-pipeline)!: CSS Trace Filters (#1985) ## libdd-shared-runtime **Next version:** `2.0.0` **Semver bump:** `major` **Tag:** `libdd-shared-runtime-v2.0.0` ###⚠️ major bump forced due to: - `libdd-common`: ^4.1.0 → ^5.1.0 ### Commits - feat(shared-runtime)!: SharedRuntime Borrowed & Owned mode (#2061) - feat(shared-runtime)!: use weak waker in trigger [APMSP-3371] (#2050) ## libdd-trace-utils **Next version:** `9.0.0` **Semver bump:** `major` **Tag:** `libdd-trace-utils-v9.0.0` ###⚠️ major bump forced due to: - `libdd-common`: ^4.2.0 → ^5.1.0 - `libdd-trace-protobuf`: ^3.0.2 → ^4.0.0 ### Commits - ci(miri): skip slow miri tests (#2188) - chore!: update protobufs to be in sync with datadog-agent (#2180) - feat(data-pipeline): add agentless export (#2081) - feat(data-pipeline)!: add stdout log trace exporter (#2074) - feat(data-pipeline): OTLP HTTP/protobuf trace export (#2115) - feat(otlp)!: Export OTLP spans with attribute-level OTel compatibility (#2091) - test(trace-utils): add V05 msgpack decode microbenchmark (#2127) - feat(data-pipeline)!: export client-computed span stats as OTLP trace metrics (#2067) - test(trace-utils): add VecMap microbenchmarks (#2126) - chore(stats)!: submit p0 telemetry in stats (#2130) - refactor(change-buffer)!: replace slot index with span_id, fix segment isolation (#2105) - feat(data-pipeline)!: CSS Trace Filters (#1985) - feat(trace-exporter): add v1 span and its encoder (#2039) - fix(trace-utils): mark decoded span maps as deduped (#2110) - feat(trace-utils)!: change buffer implementation (#2055) - feat(native-spans)!: change buffer foundation (#2046) - refactor(span)!: use VecMap for `meta`, `metrics` and `meta_struct` for v04 spans (#2043) - test: fix timeouts on heavily contended scenarios (#2093) ## libdd-telemetry **Next version:** `6.0.0` **Semver bump:** `major` **Tag:** `libdd-telemetry-v6.0.0` ###⚠️ major bump forced due to: - `libdd-common`: ^4.2.0 → ^5.1.0 - `libdd-shared-runtime`: ^1.0.0 → ^2.0.0 ### Commits - ci(miri): skip slow miri tests (#2188) - refactor(libdd-telemetry)!: avoid leaking libdd-common types in the public API (#2152) - feat(shared-runtime)!: SharedRuntime Borrowed & Owned mode (#2061) ## libdd-trace-obfuscation **Next version:** `5.0.0` **Semver bump:** `major` **Tag:** `libdd-trace-obfuscation-v5.0.0` ###⚠️ major bump forced due to: - `libdd-common`: ^4.2.0 → ^5.1.0 - `libdd-trace-protobuf`: ^3.0.2 → ^4.0.0 - `libdd-trace-utils`: ^8.0.0 → ^9.0.0 ### Commits - refactor(clippy): prefer core and alloc imports (#2196) - ci(miri): skip slow miri tests (#2188) - fix: update anyhow for unsoundness (#2186) ## libdd-trace-stats **Next version:** `6.0.0` **Semver bump:** `major` **Tag:** `libdd-trace-stats-v6.0.0` ###⚠️ major bump forced due to: - `libdd-common`: ^4.2.0 → ^5.1.0 - `libdd-shared-runtime`: ^1.0.0 → ^2.0.0 - `libdd-trace-protobuf`: ^3.0.2 → ^4.0.0 - `libdd-trace-utils`: ^8.0.0 → ^9.0.0 ### Commits - chore!: update protobufs to be in sync with datadog-agent (#2180) - feat(stats)!: send telemetry for cardinality limits (#2159) - feat(stats)!: add whole key cardinality limit (#2158) - fix(trace-stats)!: add grpc_method to aggregation key (#2151) - feat(shared-runtime)!: SharedRuntime Borrowed & Owned mode (#2061) - feat(data-pipeline)!: export client-computed span stats as OTLP trace metrics (#2067) - refactor(span)!: use VecMap for `meta`, `metrics` and `meta_struct` for v04 spans (#2043) ## libdd-data-pipeline **Next version:** `7.0.0` **Semver bump:** `major` **Tag:** `libdd-data-pipeline-v7.0.0` ###⚠️ major bump forced due to: - `libdd-common`: ^4.2.0 → ^5.1.0 - `libdd-shared-runtime`: ^1.0.0 → ^2.0.0 - `libdd-telemetry`: ^5.0.1 → ^6.0.0 - `libdd-trace-protobuf`: ^3.0.2 → ^4.0.0 - `libdd-trace-stats`: ^5.0.0 → ^6.0.0 - `libdd-trace-utils`: ^8.0.0 → ^9.0.0 ### Commits - feat(trace_exporter): enable telemetry in stats exporter (#2160) - refactor(libdd-telemetry)!: avoid leaking libdd-common types in the public API (#2152) - feat(stats): emit canonical gRPC status name for OTLP rpc.response.status_code (#2183) - feat(data-pipeline): add agentless export (#2081) - feat(stats)!: send telemetry for cardinality limits (#2159) - feat(stats)!: add whole key cardinality limit (#2158) - fix(trace-stats)!: add grpc_method to aggregation key (#2151) - feat(data-pipeline)!: add stdout log trace exporter (#2074) - feat(shared-runtime)!: SharedRuntime Borrowed & Owned mode (#2061) - feat(data-pipeline): OTLP HTTP/protobuf trace export (#2115) - feat(otlp)!: Export OTLP spans with attribute-level OTel compatibility (#2091) - feat(data-pipeline)!: export client-computed span stats as OTLP trace metrics (#2067) - chore(stats)!: submit p0 telemetry in stats (#2130) - feat(data-pipeline)!: CSS Trace Filters (#1985) - feat(shared-runtime)!: use weak waker in trigger [APMSP-3371] (#2050) - refactor(span)!: use VecMap for `meta`, `metrics` and `meta_struct` for v04 spans (#2043) - feat(stats)!: add endpoint gating to client-side stats [APMSP-3361] (#2040) ## libdd-dogstatsd-client **Next version:** `4.0.0` **Semver bump:** `major` **Tag:** `libdd-dogstatsd-client-v4.0.0` ###⚠️ major bump forced due to: - `libdd-common`: ^4.1.0 → ^5.1.0 ## libdd-sampling **Next version:** `5.0.0` **Semver bump:** `major` **Tag:** `libdd-sampling-v5.0.0` ###⚠️ major bump forced due to: - `libdd-common`: ^4.2.0 → ^5.1.0 - `libdd-trace-utils`: ^8.0.0 → ^9.0.0 [APMSP-3371]: https://datadoghq.atlassian.net/browse/APMSP-3371?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: iunanua <18325288+iunanua@users.noreply.github.com>
# What does this PR do? Add whole-key cardinality limit to the SpanConcentrator. See [this RFC](https://datadoghq.atlassian.net/wiki/spaces/APM/pages/6821151019/PENDING+Cardinality+Limits) for details The telemetry metrics specified in the RFC are implemented in #2159 The per fields limit from the RFC is out of the scope of this PR ## Breaking change The cardinality limit is applied to **all** users of the SpanConcentrator however the limit can be configured manually. # Motivation Avoid unbounded bucket growth when dealing with high cardinality. # Additional Notes ## Limit default value The limit of 7000 entries has been chosen based on the following results **Single Bucket** | Component | Typical | Worst | | --- | --- | --- | | HashMap flat table | 3\.5 MB _(unchanged)_ | 3\.5 MB _(unchanged)_ | | String heap (7 001 × 76 / 122 B) | 0\.5 MB | 0\.9 MB | | DDSketch bins (7 001 × 1 600 / 32 768 B) | 11\.2 MB | 229\.3 MB | | **StatsBucket total** | **~15.2 MB** | **~233.7 MB** | **SpanConcentrator** | Limit | Typical (3 buckets) | Worst (3 buckets) | | --- | --- | --- | | **7 000** | **~45.6 MB** | **~701.2 MB** | ## Trilean change The overflow sentinel key requires us to use a trilean to represent the `is_trace_root` in the aggregation key. Co-authored-by: vianney.ruhlmann <vianney.ruhlmann@datadoghq.com> Signed-off-by: Taegyun Kim <taegyun.kim@datadoghq.com>
# What does this PR do? Add telemetry for the collapsed stats group based on [this RFC](https://datadoghq.atlassian.net/wiki/spaces/APM/pages/6821151019/PENDING+Cardinality+Limits) # Motivation Follow-up of #2158 # Additional Notes Anything else we should know when reviewing? # How to test the change? Describe here in detail how the change can be validated. Co-authored-by: vianney.ruhlmann <vianney.ruhlmann@datadoghq.com> Signed-off-by: Taegyun Kim <taegyun.kim@datadoghq.com>
…1303) ## Overview Follows up on the `dd-trace-rs` v0.5.0 bump ([#1302](#1302)): bumps the shared `libdatadog` git rev (`libdd-capabilities`, `libdd-common`, `libdd-trace-protobuf`, `libdd-trace-utils`, `libdd-trace-normalization`, `libdd-trace-obfuscation`, `libdd-trace-stats`) from `a820699` to `85ce322a`, matching the exact versions `datadog-opentelemetry` v0.5.0 pulls from crates.io. Adds a `[patch.crates-io]` block redirecting those (and the other transitively-pulled `libdd-*` crates) to the same rev, collapsing what would otherwise be two compiled copies of each. This addresses the duplication [Copilot flagged on serverless-components#144](DataDog/serverless-components#144 (comment)) — see `bottlecap/LIBDATADOG_VERSION_ALIGNMENT.md` for the full writeup of options considered and why this one (patch + shared rev bump) was chosen over switching to crates.io versions directly (tried first, reverted — broke type identity for `ReplaceRule`/`Endpoint`/`Span`/etc. across the git-vs-registry boundary throughout the trace pipeline) or leaving the duplication as-is. ### Why not just add this to #1302? #1302 skipped this on purpose. The duplication only cost +256 bytes on the release binary (LTO strips the unreachable copy), while actually fixing it means an untested major-version bump across six libdatadog crates, including `libdd-trace-obfuscation` — the code that redacts sensitive data from spans. Not something to risk on the same PR as an urgent customer fix (SLES-2907, B3 propagation errors). This PR does that bump on its own, with its own tests. **Depends on [DataDog/serverless-components#145](DataDog/serverless-components#145 (which itself builds on #144, still draft) — the `dogstatsd`/`datadog-fips`/`datadog-agent-config` rev pin here must stay in sync with whatever commit that PR lands at. Opened as draft for that reason. JIRA: [SLES-2907](https://datadoghq.atlassian.net/browse/SLES-2907) ### Code changes at the new rev Most of the API changes in `SpanConcentrator` are just things the compiler forces on you, with no behavior change: - `new()` gained an obfuscation-config param. We pass `None` — bottlecap has never done client-side stats obfuscation. - `flush()` now returns `FlushResult { obfuscated_buckets, unobfuscated_buckets }` instead of a flat `Vec`. We concatenate both — obfuscation is off, so everything ends up in `unobfuscated_buckets` anyway. - `pb::TracerPayload` gained `container_debug`, `pb::ClientGroupedStats` gained `additional_metric_tags`. Added the missing fields to 3 struct literals so things compile; both are left at their defaults. `SpanConcentrator::new` also gained a cardinality-override param — the new rev defaults to capping trace-stats aggregation at 7,000 distinct keys per 10-second bucket, added in [DataDog/libdatadog#2158](DataDog/libdatadog#2158) to bound `SpanConcentrator`'s own memory use ([#2158](DataDog/libdatadog#2158) has the math: ~45.6 MB typical, ~700 MB worst case across 3 buckets with the limit on). This is an APM/trace-stats decision, unrelated to Datadog's custom-metrics cardinality limits. The scope of this PR is dependency consolidation, not adopting a new safety feature, so we pass `Some(usize::MAX)` to keep this rev bump behavior-neutral: no Lambda workload will ever produce `usize::MAX` distinct keys in a 10-second window, so stats collapsing never triggers, matching pre-bump behavior exactly. Passing `None` instead would opt into new behavior (collapsing above 7,000 keys into a generic `tracer_blocked_value` bucket for workloads that never hit that before) — worth considering on its own merits given Lambda's memory constraints, but that's a deliberate product/safety decision that deserves its own PR and review, not something to fold into a dependency bump. `test_no_cardinality_limit_applied` covers the unbounded behavior we're keeping. ## Testing - `cargo build --all-targets` — clean - `cargo test --all-targets` — 548/548 pass across all test binaries - `cargo clippy --all-targets -- -D warnings` — clean - `cargo fmt --all -- --check` — clean - `cargo tree --duplicates | grep ^libdd-` — empty (zero `libdd-*` duplicates) - `dd-rust-license-tool check` — passes - Verified against the real `serverless-components#145` commit (not a local `file://` override used during earlier iteration) [SLES-2907]: https://datadoghq.atlassian.net/browse/SLES-2907?atlOrigin=eyJpIjoiNWRkNTljNzYxNjVmNDY3MDlhMDU5Y2ZhYzA5YTRkZjUiLCJwIjoiZ2l0aHViLWNvbS1KU1cifQ --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

What does this PR do?
Add whole-key cardinality limit to the SpanConcentrator. See this RFC for details
The telemetry metrics specified in the RFC are implemented in #2159
The per fields limit from the RFC is out of the scope of this PR
Breaking change
The cardinality limit is applied to all users of the SpanConcentrator however the limit can be configured manually.
Motivation
Avoid unbounded bucket growth when dealing with high cardinality.
Additional Notes
Limit default value
The limit of 7000 entries has been chosen based on the following results
Single Bucket
SpanConcentrator
Trilean change
The overflow sentinel key requires us to use a trilean to represent the
is_trace_rootin the aggregation key.