Skip to content

Release the appliance peer a completed diagnostic allocated - #119

Merged
QuiteYellow merged 3 commits into
mainfrom
probe/close-notify-after-diagnostic
Oct 6, 2026
Merged

QuiteYellow merged 3 commits into
mainfrom
probe/close-notify-after-diagnostic

Conversation

@QuiteYellow

Copy link
Copy Markdown
Owner

What this changes

diagnose_dtls_handshake closed its socket without a shutdown, so every run that completed left a peer entry on the appliance. The firmware frees one when it receives close_notify and has no idle timeout that reclaims it otherwise, so the entry stayed for the life of the appliance.

The connection is a memory BIO, so shutdown() only queues the alert and it has to be read back out and sent. That is what _release_peer does, on the completed path only.

Why, measured

Six completed diagnostics inside six minutes took one of my own appliances off DTLS. The bridge went from 94 ok, 0 err to 0 ok, 14 err, timeouts=14 and stayed there about seven minutes.

The appliance stayed on the network throughout. It answered 8 of 8 pings at 0% loss and served its plaintext 5683 directory first try, so only DTLS was affected. Recovery came from the bridge's own watchdog, which logged unreachable for 175s — forcing session reconnect. So a wedged endpoint clears itself in roughly three minutes, and a power cycle is not the fix.

The docstring already warned that this function "can allocate appliance-side association state". Nothing released it.

Scope

The completed handshake only.

  • Before the handshake finishes there is no session for OpenSSL to shut down.
  • An appliance that answers with a fatal alert drops the peer itself (SSL_CHECK_FAIL calls RemovePeerFromList).
  • A handshake that times out after a HelloVerifyRequest does still leave a peer behind, because HELLO_VERIFY_REQUIRED sits in that function's exclusion list. Fixing that one means sending a plaintext alert mid-handshake and I have not measured what these appliances do with it, so it stays out of this change.

Line numbers above come from Samsung's published TizenRT tree pinned at e590f30ab, the closest match I have to these appliances. The firmware image itself is not readable.

The send is best effort by design. The caller already holds its result, and a peer this fails to release is the behaviour that shipped before, so nothing on this path raises.

Tests

Two added. The first asserts a completed diagnostic's last datagram is an Alert record under the negotiated epoch; it fails with the release disabled and passes with it. The second pins the scope: a handshake that never completed sends no alert.

949 pass on the tracked trees, where main is 947.

Hardware

Eight paced handshakes across my dryer and oven after the change. Each sent 39 bytes of close_notify and cost two ClientHellos, where a handshake running over a peer the appliance has not freed costs three. Every bridge poll window through both runs stayed at 0 err, 0 ping-fail, 0 timeouts.

diagnose_dtls_handshake closed its socket without a shutdown, so every run
that completed left a peer entry on the appliance. The firmware frees one
on close_notify and has no idle timeout that reclaims it otherwise, so the
entry stayed for the life of the appliance.

Measured on 2026-10-04: six completed diagnostics inside six minutes took a
healthy dryer's DTLS endpoint down, the bridge going from 94 ok / 0 err to
0 ok / 14 err with 14 timeouts, and it stayed there about seven minutes.
Throughout it the dryer answered 8/8 pings and served its plaintext 5683
directory first try, so only DTLS was affected. Recovery came from the
bridge's own watchdog, logging "unreachable for 175s -- forcing session
reconnect".

The connection is a memory BIO, so shutdown() only queues the alert and it
has to be read back out and sent. The send is best effort: the caller
already holds its result, and a peer this fails to release is the behaviour
that shipped before, so nothing on this path raises.

Scope is the completed handshake only. Before the handshake finishes there
is no session to shut down, and an appliance that answers with a fatal alert
drops the peer itself, so those paths are unchanged. A handshake that times
out after a HelloVerifyRequest does still leave a peer; that needs its own
change and its own hardware measurement.

Verified on hardware after the fix: eight paced handshakes across both
appliances each sent 39 bytes of close_notify and cost two ClientHellos,
where a run over a peer the appliance had not freed costs three, and every
bridge poll window stayed at 0 err / 0 ping-fail / 0 timeouts.
@QuiteYellow
QuiteYellow merged commit 48798b0 into main Oct 6, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant