Repository navigation
Release the appliance peer a completed diagnostic allocated - #119
Merged
Merged
Conversation
diagnose_dtls_handshake closed its socket without a shutdown, so every run that completed left a peer entry on the appliance. The firmware frees one on close_notify and has no idle timeout that reclaims it otherwise, so the entry stayed for the life of the appliance. Measured on 2026-10-04: six completed diagnostics inside six minutes took a healthy dryer's DTLS endpoint down, the bridge going from 94 ok / 0 err to 0 ok / 14 err with 14 timeouts, and it stayed there about seven minutes. Throughout it the dryer answered 8/8 pings and served its plaintext 5683 directory first try, so only DTLS was affected. Recovery came from the bridge's own watchdog, logging "unreachable for 175s -- forcing session reconnect". The connection is a memory BIO, so shutdown() only queues the alert and it has to be read back out and sent. The send is best effort: the caller already holds its result, and a peer this fails to release is the behaviour that shipped before, so nothing on this path raises. Scope is the completed handshake only. Before the handshake finishes there is no session to shut down, and an appliance that answers with a fatal alert drops the peer itself, so those paths are unchanged. A handshake that times out after a HelloVerifyRequest does still leave a peer; that needs its own change and its own hardware measurement. Verified on hardware after the fix: eight paced handshakes across both appliances each sent 39 bytes of close_notify and cost two ClientHellos, where a run over a peer the appliance had not freed costs three, and every bridge poll window stayed at 0 err / 0 ping-fail / 0 timeouts.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
diagnose_dtls_handshakeclosed its socket without a shutdown, so every run that completed left a peer entry on the appliance. The firmware frees one when it receivesclose_notifyand has no idle timeout that reclaims it otherwise, so the entry stayed for the life of the appliance.The connection is a memory BIO, so
shutdown()only queues the alert and it has to be read back out and sent. That is what_release_peerdoes, on the completed path only.Why, measured
Six completed diagnostics inside six minutes took one of my own appliances off DTLS. The bridge went from
94 ok, 0 errto0 ok, 14 err, timeouts=14and stayed there about seven minutes.The appliance stayed on the network throughout. It answered 8 of 8 pings at 0% loss and served its plaintext 5683 directory first try, so only DTLS was affected. Recovery came from the bridge's own watchdog, which logged
unreachable for 175s — forcing session reconnect. So a wedged endpoint clears itself in roughly three minutes, and a power cycle is not the fix.The docstring already warned that this function "can allocate appliance-side association state". Nothing released it.
Scope
The completed handshake only.
SSL_CHECK_FAILcallsRemovePeerFromList).HELLO_VERIFY_REQUIREDsits in that function's exclusion list. Fixing that one means sending a plaintext alert mid-handshake and I have not measured what these appliances do with it, so it stays out of this change.Line numbers above come from Samsung's published TizenRT tree pinned at
e590f30ab, the closest match I have to these appliances. The firmware image itself is not readable.The send is best effort by design. The caller already holds its result, and a peer this fails to release is the behaviour that shipped before, so nothing on this path raises.
Tests
Two added. The first asserts a completed diagnostic's last datagram is an Alert record under the negotiated epoch; it fails with the release disabled and passes with it. The second pins the scope: a handshake that never completed sends no alert.
949 pass on the tracked trees, where main is 947.
Hardware
Eight paced handshakes across my dryer and oven after the change. Each sent 39 bytes of
close_notifyand cost two ClientHellos, where a handshake running over a peer the appliance has not freed costs three. Every bridge poll window through both runs stayed at0 err, 0 ping-fail, 0 timeouts.