Conversation
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Binary Size Analysis (Agent Data Plane)Baseline: a41cd43 · Comparison: 477292b · diff ✅ Binary size difference within thresholdChanges by Module
Detailed Symbol Changes |
Regression Detector (Agent Data Plane)Run ID: Optimization Goals: ✅ No significant changes detectedFine details of change detection per experiment (5)Experiments configured
Bounds Checks: ✅ Passed (5)
ExplanationA change is flagged as a regression when |Δ mean %| > 5.00% in the regressing direction for its optimization goal AND SMP marks the experiment as a regression ( |
Summary
When ADP handles DogStatsD, the Core Agent's aggregator stops flushing DogStatsD traffic, so its
/telemetryflush series flatline or drop sharply:aggregator.flush_time,aggregator.flush_count, andaggregator.flushfor events and service checks. That endpoint is the migration target for customers moving offgo_expvar(DADP-179), but ADP had no flush timing or per-flush counts at all, so there was nothing to remap. This PR adds that instrumentation and maps it to the Core Agent names. Stacked on #2766 (part 1 of 2).In ADP, events and service checks skip the aggregator and go straight to their encoders, so the encoder's flush stands in for the aggregator's event and service-check flush. Metrics still flush from
dsd_agg.flowchart LR subgraph dsd_agg["dsd_agg (aggregate transform)"] A1["aggregate_last_flush_count{data_type}"] A2["aggregate_last_flush_duration_nanoseconds"] end subgraph enc["dd_events_encode / dd_service_checks_encode"] E1["encoder_flushed_events_total"] E2["encoder_last_flush_events"] E3["encoder_last_flush_duration_nanoseconds"] end A1 --> FC["aggregator__flush_count{flush_type=series|sketches}"] A2 --> FT["aggregator__flush_time{flush_type=main}"] E1 --> F["aggregator__flush{data_type=events|service_checks}"] E2 --> FC2["aggregator__flush_count{flush_type=events|service_checks}"] E3 --> FT2["aggregator__flush_time{flush_type=event|service_check}"]Known differences from the Core Agent:
flush_countvalues are smaller, andflush_timecovers one encode-and-dispatch cycle.flush_timefor metrics is only reported asflush_type="main"(thedsd_aggflush, including handing off to the encoder). The Core Agent'smetric_sketchandchecks_metric_sampletimings include serialization, which ADP does in a separate component, so they stay unmapped.aggregator.flushfor events and service checks only carriesdata_type, the same as the existingdsd_aggseries. A failed encoder flush stops the encoder instead of being counted as an error.Durations are recorded in nanoseconds to match the Core Agent's
flush_time, because a remap rule can rename a metric but not rescale its value. The new encoder telemetry is also emitted by the logs encoder, which uses the same encoder wrapper, but no rule maps it.Test plan
flush_records_last_flush_counts, which runs realdsd_aggflushes and checks that the per-flush series and sketch gauges describe only the most recent flushrecord_flush_tracks_totals_and_last_flushfor the encoder flush telemetryrender_rar_telemetry_remaps_aggregator_flush_telemetry, which checks every new remapped series and its type, and that the logs encoder's series isn't remappedrar_rules_carry_expected_help_textto pinaggregator.flush_countandaggregator.flush_timehelp text to the Core Agent's registered strings🤖 Generated with Claude Code