Skip to content

Add the pure-Python DTLS 1.2 ECDHE-PSK engine, wired to nothing - #117

Merged
QuiteYellow merged 6 commits into
mainfrom
feat/dtls-psk-engine
Oct 6, 2026
Merged

QuiteYellow merged 6 commits into
mainfrom
feat/dtls-psk-engine

Conversation

@QuiteYellow

@QuiteYellow QuiteYellow commented Oct 4, 2026 •

Copy link
Copy Markdown
Owner

Draft, and deliberately not mergeable: nothing imports this engine, so merging it would ship a module the library never calls. It is here to be read.

PskAuth cannot work on OpenSSL. The DTLS 1.2 PSK client callback passes the identity as a NUL-terminated char * and returns the key length, so a recovered OwnerPSK identity containing a zero byte reaches the wire truncated. Measured against OpenSSL 4.0.2: 16 bytes in, 8 bytes out, and psk_use_session_cb, which does carry a length, fires zero times for DTLS 1.2 across three configurations. #115 has the background.

This is a DTLS 1.2 client speaking one ciphersuite, TLS_ECDHE_PSK_WITH_AES_128_CBC_SHA256, on cryptography.

What I am asking for

@Jason-Morcos this lands on your PskAuth design: the engine replaces the DTLS transport underneath your auth provider. It now authenticates to KRZ303's oven, so there is working code to look at.

The part worth your time is the handshake state machine and the record layer. The crypto primitives I can check against the appliance's own Mbed TLS and intend to.

Thirteen defects so far. Eight I found by writing attacks against my own code, three more on first contact with the oven, and two by checking the engine against this repository's earlier protocol fixes.

One of the thirteen is the reason this is a draft. A test named test_hvr_ignored_after_cookie_hello asserted that a second HelloVerifyRequest should not reset the handshake transcript. It passed every run, because it encoded the same wrong assumption that produced the defect it was meant to catch: the engine answered the first cookie challenge and silently dropped any later one, which is the shape localthings#504 describes as the likely cause of #20's silent timeouts. A stub for reproducing it had been in the tree since August.

A self-audit cannot find a defect its author has written into the tests, and there is nothing to diff against either. No Python library implements ECDHE-PSK, and the five hand-rolled pure-Python DTLS clients on GitHub are all Hue Entertainment, so they use plain PSK with AEAD records and share none of the hard part.

Already covered

  • 1266 tests pass, on current dependencies and at the floor job's pins
  • handshake, application data both ways and a clean close against a reference server built from the appliance firmware's own Mbed TLS 2.7.8, in both its cookie and no-cookie configurations
  • 331 new cases: the record layer across every padding boundary, every single-bit flip of a record, all 256 padding lengths, wrong epoch, sequence and content type, off-curve points, suite and version downgrade, replay-window edges, and hostile-input probes for each defect above
  • the OpenSSL interop tests @KRZ303 wrote for Fix Python DTLS PSK handshake compatibility and probe validation #116, carried over with the import path changed. Their cookie case pins down something useful: OpenSSL frames its HelloVerifyRequest as DTLS 1.0 while the appliance's Mbed TLS frames it as 1.2, so a client accepting only 1.2 record headers interoperates with one and not the other
  • 60,000 randomised hostile datagrams across four lifecycle phases, and 120 on-path trials corrupting the server direction mid-handshake with no corrupted record accepted

The open design question

Routing PSK through this engine needs a connection seam that AuthenticationProvider does not have. The Protocol is runtime_checkable and typed to OpenSSL.SSL.Context, configure_context is documented public API with three implementations, and dtls_session.py constructs its context inline. Adding a second verb to that Protocol is a change to your design, so I have left it alone. WantRead, ZeroReturn and DtlsError are aliased to pyOpenSSL's exceptions here for the same reason.

That is the next change, and I would rather agree the shape than present one.

Also in this diff

cryptography becomes a declared dependency. protocol/auth.py has imported it directly since August, resolving it only as a transitive dependency of pyOpenSSL. The floor is 38.0 because pyOpenSSL==23.1.0 caps it below 41 and the floor job installs with --no-deps; the full suite passes there on Python 3.11.

A second and much smaller ask, for anyone with PSK hardware

This needs no credential of your own. localthings#432 shows an appliance running the entire handshake against a deliberately invalid identity and failing only at unknown PSK identity, which still exercises the ClientHello, the cookie round trip, the ServerKeyExchange parse, the ECDH and the ClientKeyExchange.

If your appliance negotiates ECDHE-PSK-AES128-CBC-SHA256, a trace of that failed handshake is useful, and the HelloVerifyRequest record header especially. The engine has met one oven and a laptop build of the appliance firmware's own stack. A second device family would show whether that reference server imitates real hardware or only itself.

OpenSSL's DTLS 1.2 PSK client callback passes the identity as a
NUL-terminated char * and returns the key length, so it cannot carry a
recovered OwnerPSK identity containing a zero byte. Measured against
OpenSSL 4.0.2: a 16-byte identity with a zero at offset 8 reaches the
wire as 8 bytes, and psk_use_session_cb (which does carry a length)
fires zero times for DTLS 1.2 across three configurations. See #115.

This adds a client speaking one ciphersuite,
TLS_ECDHE_PSK_WITH_AES_128_CBC_SHA256, on cryptography. Nothing imports
it: PskAuth still raises, and routing it in needs a connection seam that
AuthenticationProvider does not have yet, which is a separate change.

Every wire decision came from the appliance firmware's own Mbed TLS
2.7.8 rather than a specification. MBEDTLS_LIGHT_DEVICE compiles out
both Encrypt-then-MAC and the extended master secret, so the record
layout is classic MAC-then-encrypt with the classic derivation, and
secp256r1 is pinned.

Verified against a reference server built from those same sources in
both its cookie and no-cookie configurations, under the repository's own
unmodified _drive_dtls_handshake, and on one real appliance: KRZ303's
oven authenticated and returned /oic/d with the expected device id.

Tests are 331 cases over the record layer, the key schedule and hostile
input, plus the OpenSSL interop tests @KRZ303 wrote for #116, carried
over with the import path changed. The hostile suite covers defects that
were found and fixed rather than hypotheticals, including a replay
window that advanced on unauthenticated records, application data
accepted at epoch zero, an unclamped window shift that allocated until
the machine died, and a repeated HelloVerifyRequest that was ignored --
the shape localthings#504 describes as the likely cause of #20's silent
timeouts.

cryptography becomes a declared dependency. protocol/auth.py has
imported it directly since August and only ever resolved it as a
transitive dependency of pyOpenSSL. The floor is 38.0 because
pyOpenSSL 23.1.0 caps below 41, and the floor CI job installs with
--no-deps; the full suite passes there on Python 3.11.
tools/check_share_safety.py reads a four-component section number as an
IPv4 address, so "RFC 6347 4.1.2.6" failed the share-safety job. The
house style in dtls_probe.py cites three components, and anti-replay sits
inside 4.1.2, so the reference stays exact.
@QuiteYellow

Copy link
Copy Markdown
Owner Author

Notes for whoever builds the seam, from what @KRZ303's oven run actually forced. Splitting it because only some of it is in this diff.

Already in this branch

Both engine defects their run exposed are fixed here, and both came from #116:

  • Transcript. The initial ClientHello was excluded unconditionally. RFC 6347 excludes it only when a cookie round trip happens, so against a server that skips HelloVerifyRequest the transcript held no ClientHello at all and Finished could never verify. Reproducible locally by building the reference server with mbedtls_ssl_conf_dtls_cookies(&conf, NULL, NULL, NULL): the old engine draws fatal alert 50 and the server reports MBEDTLS_ERR_SSL_BAD_HS_FINISHED.
  • HelloVerifyRequest framing. Only DTLS 1.2 record headers were accepted. The exception for DTLS 1.0 turns out to be an OpenSSL compatibility fix. Their interop test in this branch pins OpenSSL emitting fe ff under cookie exchange, and they have since recorded the oven's own header as 16 fe fd 00 00 00 00 00 00 00 00 00 2f, which frames 1.2 like the reference server. It stays, because OpenSSL is a peer the interop tests exercise on every run.

The repeated-HelloVerifyRequest fix in here is mine. I found it afterwards, while checking this engine against the repository's earlier protocol fixes.

Not in this branch, and a Stage B obligation

Their third fix landed in the standalone probe, so none of it reaches this diff. It still matters, because it says what a real appliance needs from the layer above:

  • An unconnected socket, filtered by target IP. The probe used sock.connect(), which drops replies arriving from a different source port. This repository already fixed that once, in a879ef4, and endpoint.py has used recvfrom since.
  • A stable local UDP port. Their probe grew a PSK_LOCAL_PORT; dtls_session.py already has that option, for the reason recorded there about appliances holding stale peers.
  • CoAP response correlation on message id and token, plus empty ACKs and separate CON replies. The session layer handles this; a probe written from scratch had to learn it.

The design point those three share: the engine here owns no socket. Its whole surface is bio_read and bio_write, and the unmodified _drive_dtls_handshake already drives it. So the seam has to sit above the socket and leave transport to endpoint.py, or it will reintroduce exactly the problems their probe hit. Anything that hands the engine a socket of its own is the wrong shape.

The comment read "DTLS 1.2 servers may frame HelloVerifyRequest as DTLS
1.0", which sounds like a guess and invites deletion. Both halves are now
measured, so record them.

OpenSSL frames it that way under cookie exchange, which the interop tests
pin. KRZ303 recorded the oven's own header as 16 fe fd 00 .. 2f, so the
appliance does not. Removing the exception fails
test_openssl_psk_interoperability[True] outright, and Mbed TLS derives
that record version from the negotiated version, so other appliance
generations are unmeasured rather than known to frame 1.2.

Also notes why the exception is safe to keep: in that state a peer can
already send a 1.2-framed HelloVerifyRequest, so accepting the same
message in a second framing grants nothing further.

Comment only; the executable code is unchanged.
@Jason-Morcos

Copy link
Copy Markdown
Contributor

@QuiteYellow Yes — I'd put the connection factory at the point where DtlsCoapSession.connect() currently constructs SSL.Context / SSL.Connection, and keep AuthenticationProvider.configure_context() intact. Existing third-party providers should still work without gaining a required method on that runtime-checkable Protocol.

My preference is a small internal factory: built-in PskAuth can create the socket-free PSK engine using its private credential material; the existing context-based path remains the fallback for certificate and third-party providers. If we want a public connection-provider extension later, make that a separate opt-in interface. The same factory needs to reach diagnose_dtls_handshake(auth=...), otherwise a credential could work in a session and still fail in the diagnostic. Keep the existing endpoint, fixed local port, cancellation/deadline, CoAP correlation and close machinery above it, as you suggested. No second socket owner.

PskAuth.validate_identity() should describe what the normal session path can actually carry. Once that path is wired and tested, it can accept the full 16 bytes; direct configure_context() calls must still explicitly reject NUL identities rather than truncate them. Our integration already delegates this check to PskAuth, so we don't need to copy the backend choice into HA.

I checked 31eb4b5 and ran its three new test files locally: 331 passed. I also found one reproducible gap in the cookie handling:

Two successive challenges with different cookies First answered Second answered
Both record headers FE FD yes yes
Both record headers FE FF yes no

This is an offline input test with HVR message sequences 0 and 1, both epoch zero, and DTLS 1.2 in the HVR bodies. _handle_record() permits the FE FF exception only in sent_hello; the second challenge arrives in sent_cookie_hello and gets dropped before the repeated-cookie handler. I'd extend that narrow exception to both states and cover changed-cookie and same-cookie/timer cases under both framings. I haven't observed an appliance sending the FE FF variant.

I dug out two useful pieces of our July 12 refrigerator evidence, too:

  • A retained Mbed TLS PSK probe received HVR → ServerHello → ServerKeyExchange → ServerHelloDone, selected 0xc037, then was rejected with fatal 115 (unknown_psk_identity). Its HVR record header was 16 fe fd 00 00 00 00 00 00 00 00 00 2f; the body version was also FE FD. I replayed that recorded server flight into this exact engine offline with synthetic credentials. It reached sent_client_flight and emitted the complete synthetic 16-byte identity containing a zero byte. That's another device family's real flight passing the parser/ECDH path, not a successful authentication with this engine.
  • A separate retained cookie-probe capture has two different FE FD cookies on the same UDP tuple, with the same client random. The second arrived about 31 seconds after the first ClientHello. The client answered both, advancing its handshake sequence 0 → 1 → 2, but no ServerHello followed. So repeated challenges are real on our hardware; answering them isn't by itself a fix for every HVR-only timeout. That capture does not establish the PSK cipher, so I'm keeping it separate from the PSK result above.

I also tried a bounded live WD53 check today: one initial ClientHello, stopping before answering a cookie or sending credentials/CoAP. No response at the configured port, and a subsequent v0.1.22 discovery attempt returned no_ocf_response, so I stopped there. No live authentication claim from that attempt, and no HA unload or ownership changes.

QuiteYellow added a commit that referenced this pull request Oct 5, 2026
_handle_record admitted a DTLS 1.0-framed HelloVerifyRequest only in
sent_hello, while _handle_hello_verify_request answers one in
sent_cookie_hello too. So a server that frames its challenges that way got
its first answered and its second dropped, which is the OpenSSL behaviour
that handler exists to avoid: both sides retransmit to the deadline with no
alert. Found by @Jason-Morcos on #117.

The four existing repeated-cookie tests call _handle_hello_verify_request
directly, so none of them crossed the record layer where the gate sits.
That is the same shape as the audit test which asserted the original
repeated-HVR bug was correct.

The version check also reads only the first message in the record, while
_handle_handshake_fragment walks every message in it. A 1.0-framed record
leading with a HelloVerifyRequest therefore carried a ServerHello straight
into got_server_hello, which the framing test in test_dtls_psk_interop.py
claims is refused -- true only while the ServerHello arrived in its own
record. The exemption now travels with the record and admits one message.

No peer is known to send a second 1.0-framed challenge: an OpenSSL server
asked to re-challenge sends fatal alert 40 and one cookie, and the
appliance frames 1.2. This is a consistency repair, not a measured break.

Eight tests, five through the record layer; three fail without the fix and
two pin behaviour it must not change.
_handle_record admitted a DTLS 1.0-framed HelloVerifyRequest only in
sent_hello, while _handle_hello_verify_request answers one in
sent_cookie_hello too. So a server that frames its challenges that way got
its first answered and its second dropped, which is the OpenSSL behaviour
that handler exists to avoid: both sides retransmit to the deadline with no
alert. Found by @Jason-Morcos on #117.

The three existing repeated-cookie tests call _handle_hello_verify_request
directly, so none of them crossed the record layer where the gate sits.
That is the same shape as the audit test which asserted the original
repeated-HVR bug was correct.

The version check also reads only the first message in the record, while
_handle_handshake_fragment walks every message in it. A 1.0-framed record
leading with a HelloVerifyRequest therefore carried a ServerHello straight
into got_server_hello, which the framing test in test_dtls_psk_interop.py
claims is refused -- true only while the ServerHello arrived in its own
record. The exemption now travels with the record and admits one message.

No peer is known to send a second 1.0-framed challenge: an OpenSSL server
asked to re-challenge sends fatal alert 40 and one cookie, and the
appliance frames 1.2. This is a consistency repair, not a measured break.

Eight tests, five through the record layer; three fail without the fix and
two pin behaviour it must not change.
The engine commit added cryptography>=38.0 to dependencies, but the floor
job pins cbor2, pyOpenSSL and pytest and then installs with --no-deps, so
nothing ever resolved against that floor: cryptography arrived as whatever
pyOpenSSL 23.1.0 happened to pull, around 40.x. A declared floor no job
exercises is not a floor.

Verified at 38.0.0 on Python 3.11 with the pinned cbor2 and pyOpenSSL: the
whole suite passes, which is unsurprising -- AES-CBC, HMAC, EC key
generation and from_encoded_point all long predate it.
@QuiteYellow

Copy link
Copy Markdown
Owner Author

@Jason-Morcos Reproduced the cookie gap, built the factory the way you described, and found a second bug with the same root cause. Pushed as 6bc9065 and d3228a2 here, plus two branches: feat/dtls-connection-seam and feat/psk-engine-wiring.

The cookie gap

Your table reproduces exactly:

record header fefd: after 1st ('sent_cookie_hello','aaaaaaaa') | after 2nd ('sent_cookie_hello','bbbbbbbb', count 1) | reply 87 bytes
record header feff: after 1st ('sent_cookie_hello','aaaaaaaa') | after 2nd ('sent_cookie_hello','aaaaaaaa', count 0) | reply 0 bytes

_handle_record gated the exception on sent_hello while _handle_hello_verify_request accepts both hello states. The suite missed it because all three repeated-cookie tests call _handle_hello_verify_request directly and never cross the record layer, which is the same shape as the audit test that asserted the original repeated-HVR bug was correct.

A second one, same root cause

The version check reads the first message in a record; _handle_handshake_fragment walks every message in it. So a 1.0-framed record leading with a HelloVerifyRequest carried a ServerHello straight into got_server_hello:

FE FF record, HVR alone           -> state 'sent_cookie_hello', server_random seen: False
FE FF record, HVR + ServerHello   -> state 'got_server_hello',  server_random seen: True

That contradicts both the comment calling the exception narrow and test_dtls_psk_interop.py's claim that a 1.0-framed ServerHello is refused, which held only while it arrived in its own record. One change covers both cases: a 1.0-framed record is admitted for one message, in either hello state. Eight tests, five through the record layer, three of which fail without the fix.

I should bound it, because I have not seen a peer do this either. An OpenSSL server asked to re-challenge by rejecting the echoed cookie sends fatal alert 40 and a single cookie. So this is a consistency repair rather than a measured interop break.

The factory

DtlsCoapSession._new_dtls_connection and dtls_probe._diagnostic_connection, reached through a private _create_dtls_connection, with nothing added to the Protocol. Your objection is enforced at dtls_session.py:549, where isinstance(auth, AuthenticationProvider) raises TypeError, so a required method would reject any provider that has not grown one. A test asserts that providers with and without the factory both remain AuthenticationProvider.

Two details beyond what you described. The factory has to own set_connect_state() and set_ciphertext_mtu(), since connect() called both and the engine has neither, so it takes the MTU. And the diagnostic asks the factory before the context path: _validate_diagnostic_auth only duck-checks configure_context, which PskAuth keeps, so that ordering is the only thing routing the credential to the right engine. Both have tests.

The certificate path is a verbatim move. I traced the call sequence on main and on the branch with a real CertificateAuth: 16 identical steps, cancellation checks included.

validate_identity accepts any 16 bytes and is a pure check again, with the measured figures kept. configure_context raises instead of truncating and names the path that does carry the credential. Credential material still reaches no attribute, public or private — test_psk_auth asserts _identity and _key are both absent — so the engine factory is a closure, like the OpenSSL callback beside it.

One correction on the HA side

@mbillow, LocalThings has its own copy of the check. custom_components/localthings/credentials.py:175 raises PskIdentityZeroByte("PSK identity contains a zero byte, which DTLS cannot carry"), merged in localthings#530. So a user holding a NUL-carrying OwnerPSK is refused at config-flow time whatever ships here, and that message stops being accurate once this lands. It is also localthings#435's remaining gap 4, and that list is where someone sizing this work would look, so it needs the update as much as the code does.

On the refrigerator header

16 fe fd 00 00 00 00 00 00 00 00 00 2f is what @KRZ303 recorded from their oven in issuecomment-5979194475. Those 13 bytes do not distinguish a device: COOKIE_LEN is 4 + 28 (ssl_cookie.c:97), so 0x2f = 47 = a 12-byte handshake header, 2 version, 1 length, 32 cookie, on any Mbed TLS server at sequence 0. Your capture showing the same string corroborates it, which I think is the stronger way to put it.

What the firmware says about this path

Reading the fork at e590f30ab and its Mbed TLS 2.7.8, three things compose:

  • CAdecryptSsl calls SetupCipher only when no peer entry exists (ca_adapter_net_ssl.c:2177-2191).
  • The suite list is one file-static array (:323), rebuilt in place (:1542) and passed to mbedtls_ssl_conf_ciphersuites as a pointer (:1586), which ssl_srv.c:2107 reads live while parsing a ClientHello.
  • MBEDTLS_SSL_DTLS_CLIENT_PORT_REUSE is on (config.h:1419), so an epoch-0 ClientHello on an established session from the same address reaches ssl_handle_possible_reconnect (ssl_tls.c:3647): an invalid cookie draws a HelloVerifyRequest and keeps the peer entry, a valid one partially resets the context on that same peer. SetupCipher runs in neither case.

This client offers exactly one suite, so a reused peer entry is binary: ECDHE-PSK is in whatever that shared array last held and gets selected, or the handshake fails with handshake_failure and nothing names the stale entry. A certificate ClientHello offers a list and degrades gently. Inferred from source and untested on hardware, but if it holds then a clean close_notify matters more for PSK than for certificates, so close() now has a test through the engine: one datagram, one alert record at epoch 1, server reporting an orderly close.

One more in case it saves you time. The cookie's first four bytes are mbedtls_time() and the HMAC covers them (ssl_cookie.c:195-212), checked against a 60 s window, so a re-challenge from this stack always carries a changed cookie and the changed-cookie branch is the one that matters. That says a second challenge must differ; it says nothing about why one was issued when the first was answered inside the window, so it leaves your 31-second capture open.

Where it stands

1314 tests pass on 3.11 through 3.14 and at the declared dependency floor. d3228a2 pins cryptography==38.0.0 in the floor job, which had been resolving to whatever pyOpenSSL happened to pull. A session completes a handshake against a real OpenSSL PSK server with an identity leading with a zero byte, with and without cookie exchange, and reads that identity back out of the ClientKeyExchange. No appliance hardware has seen any of it; @KRZ303's oven is the only PSK hardware here.

I closed both loose ends. The mtu one had an answer already in the tree: dtls_probe has validated 576-16384 since it was written, with the message "mtu is outside the safe UDP range", and the session was the only caller carrying no check. That check moves to dtls_handshake, which both already import from, so there is now one rule and three callers.

Why it mattered at all: set_ciphertext_mtu consults the value only when a flight has to fragment. Measured here, OpenSSL takes 100 and 70000 alike and emits the same 212-byte ClientHello, so an unusable number stayed invisible until something large went out. A second backend turned that from invisible into inconsistent, since DtlsPskClient validates 256-65535 of its own. Both providers now refuse the same number where it was supplied. The session's range sits inside the engine's, so the engine's check is unreachable from a session and remains a guard for direct use. 3417c38; nothing in the library or the bridge passes a non-default mtu.

On sequencing, the seam is now #120 against main, ahead of the wiring. It changes no behaviour and the existing suite already covers it, it holds whichever engine ends up shipping, and on its own it makes the engine swap reviewable in isolation — which is also what #28's close says about new auth-path work. The wiring stacks on top, and that is why the mtu fix sits with the wiring: the divergence exists only once a provider carries an engine of its own.

@QuiteYellow
QuiteYellow marked this pull request as ready for review October 5, 2026 19:22
@Jason-Morcos

Copy link
Copy Markdown
Contributor

@QuiteYellow Confirmed on d3228a2: my original two-challenge reproduction now answers the changed cookie with both FE FD and FE FF framing. All 339 engine tests pass locally, including the HVR-plus-ServerHello record case you added. Carrying the restriction through the whole record is the right fix for that second bug.

I also reran the retained refrigerator server-flight replay against this head: it still reaches ClientKeyExchange and emits the full synthetic NUL-containing identity. Same limit as before: offline parsing/ECDH evidence, not a new authenticated hardware run.

Agreed on the header wording. The matching 13 bytes corroborate the framing and cookie layout; they aren't a device fingerprint. The refrigerator evidence comes from the retained capture's provenance and the rest of that exchange, not from that header being unique. And the separate capture where the client answered both cookies but never got a ServerHello remains useful evidence against treating every HVR-only timeout as the same bug.

I've answered the dispatch question on #120: getattr is what I intended, with the hook kept private and the existing Protocol unchanged. For LocalThings, KRZ303’s still-open #575 already proposes delegating that check to PskAuth.validate_identity(). It’s worth carrying that forward and updating its Mbed-specific wording, rather than starting another copy of the fix. My earlier delegation statement referred to our integration, which already uses that API.

One qualification on “a re-challenge always carries a changed cookie”: in the Mbed TLS 2.7.8 writer, the time-based path HMACs the timestamp and client ID using the existing cookie key. Two calls in the same clock second with the same key/client ID can therefore produce the same cookie. The 60-second setting is an expiration window, not forced rotation per challenge. So I’d keep the same-cookie/timer coverage too; our changed cookie at 31 seconds doesn’t identify why the peer challenged again.

The reused-peer/shared-cipher-array explanation is worth testing, but I don’t have a hardware result establishing it as the cause of our old stall. I’d keep that source inference separate from the captured behavior, as you’ve done.

The same-cookie branch said "our answer was lost rather than refused",
which picks one explanation for something consistent with two.
@Jason-Morcos corrected the premise behind it on #117: ssl_cookie_hmac is
pure over a 4-byte timestamp and the client id under a key
mbedtls_ssl_cookie_setup generates once, so a server re-challenging inside
the same clock second emits a byte-identical cookie rather than a rotated
one. An identical cookie therefore says nothing about whether our answer
was lost or the peer challenged again.

Behaviour is unchanged and still right -- the flight timer owns that
retransmission either way, so the handshake must not be renumbered
underneath it. Only the stated reason was wrong.
@QuiteYellow

Copy link
Copy Markdown
Owner Author

@Jason-Morcos You are right about the cookie. ssl_cookie_hmac is pure over the four-byte timestamp and the client id, under a key mbedtls_ssl_cookie_setup generates once, so two challenges inside the same clock second come out byte-identical. The 60 s setting is an expiry window cookie_check applies. Calling it rotation was my error, both branches are reachable on this stack, and the same-cookie and timer coverage stays exactly where it is.

It had also reached code here, which I only found by checking. The same-cookie branch read "our answer was lost rather than refused", picking one explanation for something your correction makes consistent with two: a deterministic cookie means a peer re-challenging inside the same second is byte-identical, so an identical cookie says nothing about which happened. b762aeb drops the claim and keeps the behaviour, since the flight timer owns that retransmission either way.

Your review landed on #120 as b3e2d81. The "stays interruptible" wording is gone. In its place is what the checks actually promise: cancellation set during a provider's work stops the attempt before any socket exists, and a call already blocked on a PEM read runs to completion. Your two cases are tests now rather than something you ran by hand, cancellation inside the factory and a factory consuming the whole deadline, each one failing if a socket is opened. The hook contract is in the docstring in the terms you gave it. "Built-in" is out of the dispatch description, since you have settled that a third-party provider supplying the hook deliberately may use it.

Indeed, @KRZ303's #575 already delegates to PskAuth.validate_identity() and should carry forward, as you say.

@QuiteYellow
QuiteYellow merged commit b693c03 into main Oct 6, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants