Skip to content

ci: Introduce ast-grep, rule skill guidance, extern c panic rule - #2595

Open
colin-higgins wants to merge 3 commits into
mainfrom
colin.higgins/thematic-rule-skill
Open

colin-higgins wants to merge 3 commits into
mainfrom
colin.higgins/thematic-rule-skill

Conversation

@colin-higgins

@colin-higgins colin-higgins commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

What does this PR do?

Adds a blocking ast-grep guard that rejects newly introduced pub extern "C" FFI entry points unless their bodies acknowledge panic containment. Matching is done on the Rust AST (not regex). Enforcement is diff-scoped so the ~400 existing accessors are left alone.

Also lands the ast-grep project (pinned runner, Lint job, pre-commit) and an add-lint-rule skill so later review patterns follow the same path.

Changes

  • ast-grep rule ffi-extern-c-panic-containment matches pub / pub unsafe extern "C" function items and accepts:
    • std::panic::catch_unwind
    • wrap_with_ffi_result! / wrap_with_void_ffi_result!
    • the *_no_catch! variants (explicit acknowledgement)
    • catch_panic! (data-pipeline-ffi and similar)
  • // allow(ffi-panic-boundary): <justification> on a line directly above the function is accepted; empty justifications and names that appear only in comments are not.
  • scripts/run-ffi-panic-lint.sh intersects ast-grep hits with git diff -U0 added signature lines. Git decides what is new; ast-grep decides what matched.
  • ./scripts/run-ast-grep.sh is the one local/CI command: rule tests, whole-repo scan (this rule off so legacy hits stay quiet), then the diff-scoped FFI check. The Lint job uses the same script (fetch-depth: 0). Failures emit GitHub Actions annotations.
  • Documents the convention in AGENTS.md / CONTRIBUTING.md and the add-lint-rule skill (including “do not parse Rust with regex”).

Motivation

Panics that cross C ABI boundaries can abort embedding runtimes. Clippy’s panic/unwrap lints only see explicit panic constructs, not panics from callees, so recent reviews kept flagging new uncontained FFI entry points (e.g. PRs #2545, #2551, #2566, #2569). A whole-repo error scan would fail on hundreds of intentional legacy accessors; a diff-scoped ast-grep rule blocks regressions without that rewrite.

This replaces the regex / ad-hoc brace-matcher approach in the earlier ffi-panic-lint helper crate.

Additional Notes

The check is additive and does not modify any existing FFI function or runtime behavior. datadog-sidecar-ffi still uses panic=abort (do not add catch_unwind there); new sidecar entry points should use *_no_catch! or a justified allow.

How to test the change?

# Rule fixtures + snapshots (valid wrappers / allows vs invalid / comment-only / empty allow)
./scripts/run-ast-grep.sh test

# Same command CI runs (tests, scan, diff-scoped FFI check)
./scripts/run-ast-grep.sh

# Expect this to fail: add a new `pub extern "C" fn` with no wrapper or allow, re-run, then revert.

Confirm an unchanged legacy accessor in a touched file does not fail.

@colin-higgins
colin-higgins requested review from a team as code owners September 29, 2026 19:37
@colin-higgins colin-higgins changed the title Introduce ast-grep, rule skill guidance, extern c panic rule ci: Introduce ast-grep, rule skill guidance, extern c panic rule Sep 29, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9a03908415

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread scripts/run-ffi-panic-lint.sh Outdated
Comment thread scripts/run-ffi-panic-lint.sh Outdated
Comment thread .sg/rules/ffi-extern-c-panic-containment.yml
@datadog-prod-us1-6

datadog-prod-us1-6 Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Tests

✅ All CI checks and tests passed.

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

🎯 Code Coverage (details)
• Patch Coverage: 100.00%
• Overall Coverage: 79.00% (+0.02%)

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 95c9eaf | Docs | View more details | Give us feedback!

@dd-octo-sts

dd-octo-sts Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Artifact Size Benchmark Report

aarch64-alpine-linux-musl
Artifact Baseline Commit Change
/aarch64-alpine-linux-musl/lib/libdatadog_profiling.so 9.08 MB 9.08 MB 0% (0 B) 👌
/aarch64-alpine-linux-musl/lib/libdatadog_profiling.a 96.32 MB 96.32 MB 0% (0 B) 👌
aarch64-unknown-linux-gnu
Artifact Baseline Commit Change
/aarch64-unknown-linux-gnu/lib/libdatadog_profiling.so 12.26 MB 12.26 MB 0% (0 B) 👌
/aarch64-unknown-linux-gnu/lib/libdatadog_profiling.a 107.70 MB 107.70 MB 0% (0 B) 👌
libdatadog-x64-windows
Artifact Baseline Commit Change
/libdatadog-x64-windows/debug/dynamic/datadog_profiling_ffi.dll 29.15 MB 29.15 MB 0% (0 B) 👌
/libdatadog-x64-windows/debug/dynamic/datadog_profiling_ffi.lib 97.94 KB 97.94 KB 0% (0 B) 👌
/libdatadog-x64-windows/debug/dynamic/datadog_profiling_ffi.pdb 191.91 MB 191.91 MB +0% (+8.00 KB) 👌
/libdatadog-x64-windows/debug/static/datadog_profiling_ffi.lib 817.64 MB 817.64 MB 0% (0 B) 👌
/libdatadog-x64-windows/release/dynamic/datadog_profiling_ffi.dll 9.75 MB 9.75 MB 0% (0 B) 👌
/libdatadog-x64-windows/release/dynamic/datadog_profiling_ffi.lib 97.94 KB 97.94 KB 0% (0 B) 👌
/libdatadog-x64-windows/release/dynamic/datadog_profiling_ffi.pdb 27.61 MB 27.61 MB 0% (0 B) 👌
/libdatadog-x64-windows/release/static/datadog_profiling_ffi.lib 55.79 MB 55.79 MB 0% (0 B) 👌
libdatadog-x86-windows
Artifact Baseline Commit Change
/libdatadog-x86-windows/debug/dynamic/datadog_profiling_ffi.dll 25.48 MB 25.48 MB 0% (0 B) 👌
/libdatadog-x86-windows/debug/dynamic/datadog_profiling_ffi.lib 99.47 KB 99.47 KB 0% (0 B) 👌
/libdatadog-x86-windows/debug/dynamic/datadog_profiling_ffi.pdb 197.18 MB 197.22 MB +.01% (+40.00 KB) 🔍
/libdatadog-x86-windows/debug/static/datadog_profiling_ffi.lib 804.50 MB 804.50 MB 0% (0 B) 👌
/libdatadog-x86-windows/release/dynamic/datadog_profiling_ffi.dll 7.56 MB 7.56 MB 0% (0 B) 👌
/libdatadog-x86-windows/release/dynamic/datadog_profiling_ffi.lib 99.47 KB 99.47 KB 0% (0 B) 👌
/libdatadog-x86-windows/release/dynamic/datadog_profiling_ffi.pdb 29.73 MB 29.73 MB 0% (0 B) 👌
/libdatadog-x86-windows/release/static/datadog_profiling_ffi.lib 52.72 MB 52.72 MB 0% (0 B) 👌
x86_64-alpine-linux-musl
Artifact Baseline Commit Change
/x86_64-alpine-linux-musl/lib/libdatadog_profiling.a 86.30 MB 86.30 MB 0% (0 B) 👌
/x86_64-alpine-linux-musl/lib/libdatadog_profiling.so 10.08 MB 10.08 MB 0% (0 B) 👌
x86_64-unknown-linux-gnu
Artifact Baseline Commit Change
/x86_64-unknown-linux-gnu/lib/libdatadog_profiling.a 102.14 MB 102.14 MB 0% (0 B) 👌
/x86_64-unknown-linux-gnu/lib/libdatadog_profiling.so 12.32 MB 12.32 MB 0% (0 B) 👌

@pr-commenter

pr-commenter Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Benchmarks

Comparison

Benchmark execution time: 2026-09-29 21:03:35

Comparing candidate commit 95c9eaf in PR branch colin.higgins/thematic-rule-skill with baseline commit 313c069 in branch main.

📊 Benchmarking dashboard

Found 0 performance improvements and 1 performance regressions! Performance is the same for 176 metrics, 0 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:vec_map/as_deduped_map/needs_dedup_1_in_2/8

  • 🟥 execution_time [+22.811ns; +23.004ns] or [+4.925%; +4.966%]

Benchmark execution time: 2026-09-29 20:57:13

Comparing candidate commit 95c9eaf in PR branch colin.higgins/thematic-rule-skill with baseline commit 313c069 in branch main.

📊 Benchmarking dashboard

Found 1 performance improvements and 0 performance regressions! Performance is the same for 165 metrics, 10 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:thread_cpu/alloc_free/system/4096

  • 🟩 execution_time [-8.380ns; -8.195ns] or [-8.965%; -8.768%]

Unstable benchmarks

These benchmarks have a confidence interval too wide to call a change; treat them as noise rather than signal.

scenario:datadog_sample_span/parent_not_sampled_short_circuit/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+554.966%; -555.373%]

scenario:datadog_sample_span/parent_sampled_short_circuit/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+556.120%; -555.916%]

scenario:glob_matcher/ascii_case_insensitive_match/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+555.141%; -555.455%]

scenario:glob_matcher/ascii_exact_match/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+548.016%; -552.117%]

scenario:glob_matcher/ascii_exact_miss/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+554.308%; -555.063%]

scenario:glob_matcher/ascii_wildcard_backtrack_match/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+554.957%; -555.369%]

scenario:glob_matcher/ascii_wildcard_heavy_backtrack/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+558.338%; -556.962%]

scenario:glob_matcher/ascii_wildcard_question_match/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+553.410%; -554.642%]

scenario:glob_matcher/ascii_wildcard_star_match/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+555.735%; -555.735%]

scenario:glob_matcher/star_short_circuit/allocated_bytes

  • unstable execution_time [-0.000ns; +0.000ns] or [+553.535%; -554.701%]

Candidate

Omitted due to size.

Baseline

Omitted due to size.

ast-grep cannot expand macros, so a new c_setters! path produced an
uncontained export with no matching function_item. Also propagate
filtered scan failures and diff github.event.before on push so
origin/main being HEAD does not empty the range.

Co-authored-by: Cursor <cursoragent@cursor.com>
Comment on lines +142 to +144
If a correct rule would fail on a large, intentional legacy set, do
**not** rewrite the repo and do **not** fall back to regex. Copy the
FFI panic rule:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I disagree with this. If we decide a pattern is harmful and lint against it, we should lint the whole codebase.
It is fine to have escape hatches in the codebase, adding local lint skips in comments but we should not have code that causes lint violations go unnoticed until someone decides to touch it for whatever reason.

I think the pattern for adding lints should probably be stacked PR:

  • One adding skips for parts of the code that would fails
  • A stacked PR on top of this adding the lint in the CI

This ways it's easy to know what code is non compliant, and if the lint actually makes sense to apply to the whole codebase

Comment on lines +204 to +216
# Format / Clippy (touched crate)
cargo +nightly-2026-07-26 fmt --all -- --check
cargo +stable clippy -p <crate> --all-targets -- -D warnings

# Workspace dependency declarations
(cd .github/actions && cargo run -p workspace-deps-lint)

# Deps / licenses
cargo deny check
cargo machete --with-metadata --skip-target-dir

# Crypto graph
./scripts/check_crypto_providers.sh

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

these should already be in AGENTS.md so prefer referring to it in the skill. This way we have only one place to maintain

Comment on lines +220 to +226
alt = "|".join(re.escape(name) for name in sorted(macro_names))
rule = f"""
id: ffi-macro-invocation-emits-extern-c
language: rust
rule:
kind: macro_invocation
regex: "(^|::)({alt})!"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A bash script that contains inline python that contains inline ast-grep rules?? Seems a bit too.. sloppy honestly

exit 0
fi

DIFF="$(git diff -U0 --diff-filter=ACMR "$BASE" -- '*.rs' || true)"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This needs --no-ext-diff: with an external diff tool configured (difftastic etc.) no added lines get parsed, so the check silently passes locally.

Suggested change
DIFF="$(git diff -U0 --diff-filter=ACMR "$BASE" -- '*.rs' || true)"
DIFF="$(git --no-pager diff --no-ext-diff --no-color --no-textconv -U0 --diff-filter=ACMR "$BASE" -- '*.rs' || true)"

Comment on lines +29 to +37
kind: visibility_modifier
regex: "^pub$"
- has:
kind: function_modifiers
has:
kind: extern_modifier
has:
kind: string_literal
regex: '^"C"$'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This only matches pub extern "C" fn, so it misses non-pub #[no_mangle] functions, extern fn without an ABI string, and extern "system".

Comment on lines +21 to +24
ignores:
- "**/tests/**"
- "**/examples/**"
- "symbolizer-ffi/**"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we add datadog-sidecar-ffi/** to ignores? AGENTS.md rules out catch_unwind there, so otherwise it'd need ~140 allow comments.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants