Skip to content

chore(ci): benchmark metric tag value allowlists - #2320

Draft
lukesteensen wants to merge 1 commit into
feature/tag-value-allowlistfrom
test/smp-tag-value-allowlist
Draft

chore(ci): benchmark metric tag value allowlists#2320
lukesteensen wants to merge 1 commit into
feature/tag-value-allowlistfrom
test/smp-tag-value-allowlist

Conversation

@lukesteensen

Copy link
Copy Markdown
Contributor

Summary

Add non-gating SMP experiments for the ADP-only metric_tag_value_allowlist feature introduced by #2271.

The experiments run ADP standalone and use one set of Jinja variables to generate both the allow-list configuration and the Lading corpus. This avoids Core Agent noise and prevents the generated rules from drifting away from the metrics, tags, and values sent by Lading.

The matrix measures:

  • idle memory for zero rules, one rule with 10 values, 500 and 10,000 rules with 10 values each, and one rule with 10,000 values;
  • CPU at 100 MiB/s for zero rules, an all-allowed broad-prefix rule, 25% remove and replace mismatch paths, 500 and 10,000 matching prefix rules, and a 10,000-rule full-table miss.

These experiments intentionally have no checks: yet. They belong to the on-demand/nightly full suite so we can establish measurements before choosing quality-gate bounds.

This PR is stacked on #2271. The SMP PR baseline is still the merge base with main, so the comparison includes the parent feature implementation.

Validation

  • make generate-smp-experiments
  • make check-smp-experiments
  • Parsed every generated allow-list and verified rule counts, unique values, and non-overlapping same-tag prefixes.
  • Started ADP locally with the generated 10,000-rule configuration and confirmed the topology became healthy.
  • git diff --check

References

@dd-octo-sts dd-octo-sts Bot added the area/test All things testing: unit/integration, correctness, SMP regression, etc. label Aug 12, 2026
@pr-commenter

pr-commenter Bot commented Aug 12, 2026

Copy link
Copy Markdown

Binary Size Analysis (Agent Data Plane)

Baseline: 0f47357 · Comparison: 2be0a82 · diff
Analysis Configuration: stripped binaries · Pass/Fail Threshold: +5%
Sizes: 41.27 MiB (baseline) vs 41.14 MiB (comparison)
Size Change: -134.50 KiB (-0.32%)

✅ Binary size difference within threshold

Changes by Module
Module File Size Symbols
core -72.61 KiB 4745
serde_json +31.94 KiB 353
[sections] -31.50 KiB 9
tokio -30.57 KiB 1525
http_body_util -25.11 KiB 99
agent_data_plane::dogstatsd_contexts::artifact -24.86 KiB 24
figment -18.86 KiB 134
agent_data_plane::internal::remote_agent +18.84 KiB 119
anon.1ac75c3e2cc4e42e4c4d3caf2e351f47.630.llvm.2617875662619396550 +17.43 KiB 1
anon.6c8cc82583d787f2710e1f8f86e78987.13.llvm.13314861055718887123 -17.17 KiB 1
agent_data_plane_config_system::saluki_only::_ +16.87 KiB 4
agent_data_plane_config::domains::dogstatsd +15.96 KiB 22
comfy_table -14.67 KiB 19
smallvec +13.14 KiB 72
chrono +12.63 KiB 15
agent_data_plane_config_system::saluki_env_overlay::PathRecorder -12.45 KiB 23
agent_data_plane_config::shared::_ -12.38 KiB 16
anon.eaa35ebdc5998855717b6f667ad22404.791.llvm.3148331925466594304 -11.93 KiB 1
anon.1ac75c3e2cc4e42e4c4d3caf2e351f47.669.llvm.2617875662619396550 +11.84 KiB 1
agent_data_plane::internal::env +10.84 KiB 204
Detailed Symbol Changes
    FILE SIZE        VM SIZE    
 --------------  -------------- 
  [NEW] +42.8Ki  [NEW] +42.7Ki    agent_data_plane::cli::run::handle_run_command::_{{closure}}::h6d892181b81751eb
  [NEW] +40.4Ki  [NEW] +40.3Ki    agent_data_plane::cli::run::create_topology::_{{closure}}::h4e6ee5a870242c28
  [NEW] +31.1Ki  [NEW] +30.9Ki    datadog_agent_commons::ipc::client::RemoteAgentClient::from_client_configuration::_{{closure}}::_{{closure}}::_{{closure}}::h15e5c55b2534bd2f
  [NEW] +31.1Ki  [NEW] +30.9Ki    agent_data_plane::cli::dogstatsd::run_dogstatsd_command::_{{closure}}::h6b215381893ce646
  [NEW] +26.7Ki  [NEW] +26.5Ki    _<figment::value::de::ConfiguredValueDe<I> as serde_core::de::Deserializer>::deserialize_struct::hff3ee7185b08af36
  [NEW] +25.9Ki  [NEW] +25.8Ki    _<figment::value::de::ConfiguredValueDe<I> as serde_core::de::Deserializer>::deserialize_struct::hf46c6b57f60ea399
  [NEW] +25.9Ki  [NEW] +25.8Ki    agent_data_plane::state::metrics::rules::get_datadog_agent_remappings::hf494852330a7d7e5
  [NEW] +25.0Ki  [NEW] +24.8Ki    agent_data_plane::internal::remote_agent::run_remote_agent_registration_loop::_{{closure}}::h5e51a25fd3e89d0b
  [NEW] +22.3Ki  [NEW] +22.1Ki    agent_data_plane::main::_{{closure}}::heac4885bfc658838
  [NEW] +20.7Ki  [NEW] +20.5Ki    agent_data_plane::cli::debug::handle_debug_command::_{{closure}}::h9acf3dd4891f6833
  [DEL] -21.0Ki  [DEL] -20.9Ki    agent_data_plane::cli::debug::handle_debug_command::_{{closure}}::ha491094e358ebf93
  [DEL] -21.7Ki  [DEL] -21.6Ki    agent_data_plane::main::_{{closure}}::h6b51b32b5abcd802
  [DEL] -25.1Ki  [DEL] -25.0Ki    agent_data_plane::internal::remote_agent::run_remote_agent_registration_loop::_{{closure}}::ha8741207da6d0025
  -2.2% -25.8Ki  -2.2% -25.8Ki    [section .gcc_except_table]
  [DEL] -26.7Ki  [DEL] -26.5Ki    datadog_agent_commons::ipc::client::RemoteAgentClient::from_client_configuration::_{{closure}}::_{{closure}}::_{{closure}}::h3146495241b183c8
  [DEL] -26.7Ki  [DEL] -26.5Ki    core::ptr::drop_in_place<agent_data_plane::cli::run::handle_run_command::{{closure}}>::hac8ba744a6354ad7
  [DEL] -28.5Ki  [DEL] -28.3Ki    agent_data_plane::dogstatsd_contexts::artifact::for_each_record::h39b2832027986acc
  [DEL] -30.6Ki  [DEL] -30.4Ki    agent_data_plane::cli::dogstatsd::run_dogstatsd_command::_{{closure}}::h8e6a26bba5766166
  [DEL] -40.3Ki  [DEL] -40.2Ki    agent_data_plane::cli::run::create_topology::_{{closure}}::h23953cd93c271a64
  [DEL] -42.2Ki  [DEL] -42.1Ki    agent_data_plane::cli::run::handle_run_command::_{{closure}}::h8c5bc31d01253285
  -1.2%  -137Ki  -1.1% -95.2Ki    [18133 Others]
  -0.3%  -134Ki  -0.3% -92.0Ki    TOTAL

@pr-commenter

pr-commenter Bot commented Aug 12, 2026

Copy link
Copy Markdown

Regression Detector (Agent Data Plane)

Run ID: c080f32d-4887-42d7-85a7-c7a4ad1881fc
Baseline: 0f47357a · Comparison: 2be0a82f · diff

Optimization Goals: ✅ No significant changes detected

Fine details of change detection per experiment (5)

Experiments configured erratic: true are tagged (ignored) and skipped when determining which experiments regressed or improved. Experiments which are detected as erratic at runtime are tagged (erratic) to flag that the run's sample dispersion was high, but their regression / improvement signal still counts.

experiment goal Δ mean % links
quality_gates_rss_dsd_ultraheavy memory ⚪ +1.18 metrics profiles logs
quality_gates_rss_dsd_medium memory ⚪ +0.59 metrics profiles logs
quality_gates_rss_dsd_heavy memory ⚪ +0.57 metrics profiles logs
quality_gates_rss_dsd_low memory ⚪ +0.21 metrics profiles logs
quality_gates_rss_idle memory ⚪ -0.05 metrics profiles logs
Bounds Checks: ❌ Failed (5)
experiment check replicates observed links
quality_gates_rss_dsd_heavy memory_usage 10/10 ✅ 230 MiB ≤ 250 MiB metrics profiles logs
quality_gates_rss_dsd_low memory_usage 10/10 ✅ 51.3 MiB ≤ 60 MiB metrics profiles logs
quality_gates_rss_dsd_medium memory_usage 10/10 ✅ 89.2 MiB ≤ 100 MiB metrics profiles logs
quality_gates_rss_dsd_ultraheavy memory_usage 9/10 ❌ 460 MiB ≤ 420 MiB metrics profiles logs
quality_gates_rss_idle memory_usage 10/10 ✅ 31.7 MiB ≤ 40 MiB metrics profiles logs
Explanation

A change is flagged as a regression when |Δ mean %| > 5.00% in the regressing direction for its optimization goal AND SMP marks the experiment as a regression (is_regression: true). Improvements use the matching criteria for the improving direction. Experiments configured erratic: true (tagged (ignored)) are skipped outright; experiments detected as erratic at runtime (tagged (erratic)) still count, since that flag describes sample dispersion rather than directional certainty. The Δ mean % cell is colored accordingly: 🟢 = improvement, 🔴 = regression, ⚪ = neutral. Reduction in CPU or memory is an improvement; reduction in ingress throughput is a regression. Experiments tagged (no analysis) show ⚠️ n/a: SMP ran them but produced no analysis, usually because a replicate failed and exhausted its retries. Check the SMP report for that experiment's replicate failures.

@lukesteensen lukesteensen changed the title test(smp): benchmark metric tag value allowlists chore(ci): benchmark metric tag value allowlists Aug 12, 2026
@lukesteensen
lukesteensen force-pushed the feature/tag-value-allowlist branch from 2b7632e to 257b400 Compare August 14, 2026 20:59
@lukesteensen
lukesteensen force-pushed the test/smp-tag-value-allowlist branch from 6f35b22 to 2be0a82 Compare August 14, 2026 20:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/test All things testing: unit/integration, correctness, SMP regression, etc.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant