Skip to content

fix(sidecar): avoid telemetry cache lock inversions - #2562

Merged
gh-worker-dd-mergequeue-cf854d[bot] merged 3 commits into
mainfrom
glopes/deadlock
Sep 29, 2026
Merged

gh-worker-dd-mergequeue-cf854d[bot] merged 3 commits into
mainfrom
glopes/deadlock

Conversation

@cataphract

Copy link
Copy Markdown
Contributor

What does this PR do?

Snapshot telemetry client handles before stats and flush inspect them. This keeps the cache lock out of application and client lock scopes while preserving Stop sequencing.

Guard Stop removal by client identity and cover retirement, snapshot lifetime, and replacement races with regression tests.

How to test the change?

See DataDog/dd-trace-php#4219

Snapshot telemetry client handles before stats and flush inspect them.
This keeps the cache lock out of application and client lock scopes while
preserving Stop sequencing.

Guard Stop removal by client identity and cover retirement, snapshot
lifetime, and replacement races with regression tests.
@cataphract
cataphract requested review from a team as code owners September 22, 2026 18:19
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-22T18:23:12.047376Z 87678dc PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@cataphract

cataphract commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor Author

Working through the strongest version of this alternative: after last_handle.await, process
Stop and take() the client, release the client lock, then acquire the cache lock and remove the
entry only if it is still the same Arc and still contains None.

Abbreviated, I understand that implementation to look like this:

if let Some(last_handle) = last_handle {
    last_handle.await.ok();
}

let (processed, stopped_client) = {
    let mut guard = telemetry_mutex.lock_or_panic();
    let processed = guard
        .as_mut()
        .map(|client| client.process_actions(actions))
        .unwrap_or_default();

    let stopped_client = guard.take();
    (processed, stopped_client)
}; // Release the client lock.

telemetry_clients.remove_if_stopped(service, env, &telemetry_mutex);
drop(stopped_client);

with removal approximately:

fn remove_if_stopped(&self, key: &Key, expected: &ClientHandle) {
    let mut clients = self.inner.lock_or_panic();

    let should_remove = clients.get(key).is_some_and(|entry| {
        Arc::ptr_eq(&entry.client, expected)
            && entry.client.lock_or_panic().is_none()
    });

    if should_remove {
        clients.remove(key);
    }
}

But during the interval between Stop's take() and cache removal, a new batch follows
this path:

cache still contains old Arc
    ↓
first get_or_create() returns that Arc through get_existing_client()
    ↓
caller locks it and observes None
    ↓
caller retries get_or_create()
    ↓
get_existing_client() returns the same cached Arc again
    ↓
caller observes None again and drops the batch

Swapping removal and take() avoids publishing None, but then removal occurs only after the Stop
task has awaited earlier work. During that wait, the client remains discoverable. New actions can
attach themselves after the Stop task in the last_handle chain, then run after Stop has taken the
client and be dropped.

The snapshot approach avoids changing these lifecycle semantics. Stop still removes the entry
synchronously and defers take() until preceding actions have completed, while stats and flush no
longer hold the cache lock when acquiring client locks. It also fixes the separate
cache-to-application inversion in compute_stats(), caused by the temporary cache guard surviving
into the later struct fields. For those reasons, I think snapshotting is the smaller and safer fix.

Snapshotting is not free: each stats or flush operation allocates a Vec and
performs an atomic reference-count increment and decrement for every cached
client. In exchange, the global cache lock is held only long enough to copy the
handles instead of while waiting for and inspecting every client.

@pr-commenter

pr-commenter Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Benchmarks

Comparison

Candidate

Candidate benchmark details

Baseline

Baseline benchmark details

@datadog-official

datadog-official Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Tests

✅ All CI checks and tests passed.

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

🎯 Code Coverage (details)
• Patch Coverage: 92.04%
• Overall Coverage: 78.90% (+0.27%)

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 5d183fe | Docs | View more details | Give us feedback!

@dd-octo-sts

dd-octo-sts Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Artifact Size Benchmark Report

aarch64-alpine-linux-musl
Artifact Baseline Commit Change
/aarch64-alpine-linux-musl/lib/libdatadog_profiling.a 95.92 MB 95.92 MB 0% (0 B) 👌
/aarch64-alpine-linux-musl/lib/libdatadog_profiling.so 9.02 MB 9.02 MB 0% (0 B) 👌
aarch64-unknown-linux-gnu
Artifact Baseline Commit Change
/aarch64-unknown-linux-gnu/lib/libdatadog_profiling.so 12.19 MB 12.19 MB 0% (0 B) 👌
/aarch64-unknown-linux-gnu/lib/libdatadog_profiling.a 107.30 MB 107.30 MB 0% (0 B) 👌
libdatadog-x64-windows
Artifact Baseline Commit Change
/libdatadog-x64-windows/debug/dynamic/datadog_profiling_ffi.dll 29.06 MB 29.06 MB 0% (0 B) 👌
/libdatadog-x64-windows/debug/dynamic/datadog_profiling_ffi.lib 96.08 KB 96.08 KB 0% (0 B) 👌
/libdatadog-x64-windows/debug/dynamic/datadog_profiling_ffi.pdb 191.44 MB 191.43 MB -0% (-8.00 KB) 👌
/libdatadog-x64-windows/debug/static/datadog_profiling_ffi.lib 816.66 MB 816.66 MB 0% (0 B) 👌
/libdatadog-x64-windows/release/dynamic/datadog_profiling_ffi.dll 9.70 MB 9.70 MB 0% (0 B) 👌
/libdatadog-x64-windows/release/dynamic/datadog_profiling_ffi.lib 96.08 KB 96.08 KB 0% (0 B) 👌
/libdatadog-x64-windows/release/dynamic/datadog_profiling_ffi.pdb 27.49 MB 27.49 MB 0% (0 B) 👌
/libdatadog-x64-windows/release/static/datadog_profiling_ffi.lib 55.57 MB 55.57 MB 0% (0 B) 👌
libdatadog-x86-windows
Artifact Baseline Commit Change
/libdatadog-x86-windows/debug/dynamic/datadog_profiling_ffi.dll 25.41 MB 25.41 MB 0% (0 B) 👌
/libdatadog-x86-windows/debug/dynamic/datadog_profiling_ffi.lib 97.58 KB 97.58 KB 0% (0 B) 👌
/libdatadog-x86-windows/debug/dynamic/datadog_profiling_ffi.pdb 196.64 MB 196.63 MB -0% (-8.00 KB) 👌
/libdatadog-x86-windows/debug/static/datadog_profiling_ffi.lib 802.18 MB 802.18 MB 0% (0 B) 👌
/libdatadog-x86-windows/release/dynamic/datadog_profiling_ffi.dll 7.52 MB 7.52 MB 0% (0 B) 👌
/libdatadog-x86-windows/release/dynamic/datadog_profiling_ffi.lib 97.58 KB 97.58 KB 0% (0 B) 👌
/libdatadog-x86-windows/release/dynamic/datadog_profiling_ffi.pdb 29.60 MB 29.60 MB 0% (0 B) 👌
/libdatadog-x86-windows/release/static/datadog_profiling_ffi.lib 52.52 MB 52.52 MB 0% (0 B) 👌
x86_64-alpine-linux-musl
Artifact Baseline Commit Change
/x86_64-alpine-linux-musl/lib/libdatadog_profiling.a 85.92 MB 85.92 MB 0% (0 B) 👌
/x86_64-alpine-linux-musl/lib/libdatadog_profiling.so 10.04 MB 10.04 MB 0% (0 B) 👌
x86_64-unknown-linux-gnu
Artifact Baseline Commit Change
/x86_64-unknown-linux-gnu/lib/libdatadog_profiling.a 101.79 MB 101.79 MB 0% (0 B) 👌
/x86_64-unknown-linux-gnu/lib/libdatadog_profiling.so 12.26 MB 12.26 MB 0% (0 B) 👌

@bwoebi

bwoebi commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

@cataphract I've looked at it more closely and found that the nested case for taking the lock no longer applies anyway:

https://github.com/DataDog/libdatadog/compare/bob/deadlock?expand=1

@cataphract

Copy link
Copy Markdown
Contributor Author

@bwoebi I like the direction of releasing the client lock before going for the the cache lock. I'm not sure, however, that removing snapshotting is the best direction.

  1. This cannot be proven without a realistic benchmark, but I think that memory/cpu are worth it to avoid holding the lock on the whole cache during stats/flush. Perhaps just changing that lock to a read/write lock would give the same advantage though, without the downside.
  2. Because stats keep holding the cache lock throughout and lock clients during that time, bugs like the two fixed here are more likely to show up again.

Other notes:

    // Stop marks clients retired before removing them from the cache.
   client.filter(|client| client.lock_or_panic().is_some())

This doesn't do much. It can still become None right after the check. The caller needs to check anyway.

get_or_create has an existing bug. It's not atomic. Two threads can see no entry in the cache and create two telemetry clients -- but the first to be written to the cache is going to be replaced there right after.

Finally, stepping back a bit... This dance with the two exposed locks (three if we count the application map lock involved in the stats-enqueue deadlock) is not great. I think a better one would be to have only the cache lock and have a separate actor (a tokio task) handling, in a serialized fashion, the actions for that specific client.

In any case, feel free to move forward with your commit, the most important thing is to fix the deadlocks.

@bwoebi

bwoebi commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

This doesn't do much. It can still become None right after the check. The caller needs to check anyway.

Yeah, I'm aware. It's just about making the window a bit smaller. The code isn't optimal and probably should be overhauled anyway eventually.

But yes, merging this now, to avoid the deadlock.

I'm not sure, however, that removing snapshotting is the best direction

Agree, however my intuition would be to not snapshot ... until we prove that snapshotting is the right direction.

@bwoebi

bwoebi commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

/merge

@gh-worker-devflow-routing-ef8351

gh-worker-devflow-routing-ef8351 Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

View all feedbacks in Devflow UI.

2026-09-28 14:39:12 UTC ℹ️ Start processing command /merge


2026-09-28 14:39:18 UTC ℹ️ MergeQueue: Pull request is not mergeable yet

It will be processed automatically as soon as GitHub reports it as mergeable. View in MergeQueue UI.

  • Run /code blockers to see what is blocking it.
  • Run /remove to cancel it.

2026-09-28 14:40:23 UTC ℹ️ MergeQueue: merge request added to the queue

The expected merge time in main is approximately 48m (p90).


2026-09-28 15:19:16 UTC ❌ MergeQueue: This merge request is not mergeable, blocked by github

PR can't be merged according to github policy

@yannham yannham left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed with @cataphract for the longer term, actor-based architecture. A three mutex dance is just asking for deadlocks to happen...

Comment thread datadog-sidecar/src/service/telemetry.rs Outdated
bwoebi and others added 2 commits September 28, 2026 17:16
Queued tasks now own their converted actions and cloned worker handles.
The existing last_handle chain still ensures that preceding batches are enqueued before Stop.

Consequently, take() no longer needs to run inside the spawned task.
@cataphract

Copy link
Copy Markdown
Contributor Author

/merge

@gh-worker-devflow-routing-ef8351

gh-worker-devflow-routing-ef8351 Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

View all feedbacks in Devflow UI.

2026-09-29 10:52:39 UTC ℹ️ Start processing command /merge


2026-09-29 10:52:43 UTC ℹ️ MergeQueue: pull request added to the queue

The expected merge time in main is approximately 48m (p90).


2026-09-29 11:24:36 UTC ℹ️ MergeQueue: This merge request was merged

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants