Repository navigation
Add the pure-Python DTLS 1.2 ECDHE-PSK engine, wired to nothing - #117
Conversation
OpenSSL's DTLS 1.2 PSK client callback passes the identity as a NUL-terminated char * and returns the key length, so it cannot carry a recovered OwnerPSK identity containing a zero byte. Measured against OpenSSL 4.0.2: a 16-byte identity with a zero at offset 8 reaches the wire as 8 bytes, and psk_use_session_cb (which does carry a length) fires zero times for DTLS 1.2 across three configurations. See #115. This adds a client speaking one ciphersuite, TLS_ECDHE_PSK_WITH_AES_128_CBC_SHA256, on cryptography. Nothing imports it: PskAuth still raises, and routing it in needs a connection seam that AuthenticationProvider does not have yet, which is a separate change. Every wire decision came from the appliance firmware's own Mbed TLS 2.7.8 rather than a specification. MBEDTLS_LIGHT_DEVICE compiles out both Encrypt-then-MAC and the extended master secret, so the record layout is classic MAC-then-encrypt with the classic derivation, and secp256r1 is pinned. Verified against a reference server built from those same sources in both its cookie and no-cookie configurations, under the repository's own unmodified _drive_dtls_handshake, and on one real appliance: KRZ303's oven authenticated and returned /oic/d with the expected device id. Tests are 331 cases over the record layer, the key schedule and hostile input, plus the OpenSSL interop tests @KRZ303 wrote for #116, carried over with the import path changed. The hostile suite covers defects that were found and fixed rather than hypotheticals, including a replay window that advanced on unauthenticated records, application data accepted at epoch zero, an unclamped window shift that allocated until the machine died, and a repeated HelloVerifyRequest that was ignored -- the shape localthings#504 describes as the likely cause of #20's silent timeouts. cryptography becomes a declared dependency. protocol/auth.py has imported it directly since August and only ever resolved it as a transitive dependency of pyOpenSSL. The floor is 38.0 because pyOpenSSL 23.1.0 caps below 41, and the floor CI job installs with --no-deps; the full suite passes there on Python 3.11.
tools/check_share_safety.py reads a four-component section number as an IPv4 address, so "RFC 6347 4.1.2.6" failed the share-safety job. The house style in dtls_probe.py cites three components, and anti-replay sits inside 4.1.2, so the reference stays exact.
|
Notes for whoever builds the seam, from what @KRZ303's oven run actually forced. Splitting it because only some of it is in this diff. Already in this branchBoth engine defects their run exposed are fixed here, and both came from #116:
The repeated-HelloVerifyRequest fix in here is mine. I found it afterwards, while checking this engine against the repository's earlier protocol fixes. Not in this branch, and a Stage B obligationTheir third fix landed in the standalone probe, so none of it reaches this diff. It still matters, because it says what a real appliance needs from the layer above:
The design point those three share: the engine here owns no socket. Its whole surface is |
The comment read "DTLS 1.2 servers may frame HelloVerifyRequest as DTLS 1.0", which sounds like a guess and invites deletion. Both halves are now measured, so record them. OpenSSL frames it that way under cookie exchange, which the interop tests pin. KRZ303 recorded the oven's own header as 16 fe fd 00 .. 2f, so the appliance does not. Removing the exception fails test_openssl_psk_interoperability[True] outright, and Mbed TLS derives that record version from the negotiated version, so other appliance generations are unmeasured rather than known to frame 1.2. Also notes why the exception is safe to keep: in that state a peer can already send a 1.2-framed HelloVerifyRequest, so accepting the same message in a second framing grants nothing further. Comment only; the executable code is unchanged.
|
@QuiteYellow Yes — I'd put the connection factory at the point where My preference is a small internal factory: built-in
I checked
This is an offline input test with HVR message sequences 0 and 1, both epoch zero, and DTLS 1.2 in the HVR bodies. I dug out two useful pieces of our July 12 refrigerator evidence, too:
I also tried a bounded live WD53 check today: one initial ClientHello, stopping before answering a cookie or sending credentials/CoAP. No response at the configured port, and a subsequent v0.1.22 discovery attempt returned |
_handle_record admitted a DTLS 1.0-framed HelloVerifyRequest only in sent_hello, while _handle_hello_verify_request answers one in sent_cookie_hello too. So a server that frames its challenges that way got its first answered and its second dropped, which is the OpenSSL behaviour that handler exists to avoid: both sides retransmit to the deadline with no alert. Found by @Jason-Morcos on #117. The four existing repeated-cookie tests call _handle_hello_verify_request directly, so none of them crossed the record layer where the gate sits. That is the same shape as the audit test which asserted the original repeated-HVR bug was correct. The version check also reads only the first message in the record, while _handle_handshake_fragment walks every message in it. A 1.0-framed record leading with a HelloVerifyRequest therefore carried a ServerHello straight into got_server_hello, which the framing test in test_dtls_psk_interop.py claims is refused -- true only while the ServerHello arrived in its own record. The exemption now travels with the record and admits one message. No peer is known to send a second 1.0-framed challenge: an OpenSSL server asked to re-challenge sends fatal alert 40 and one cookie, and the appliance frames 1.2. This is a consistency repair, not a measured break. Eight tests, five through the record layer; three fail without the fix and two pin behaviour it must not change.
_handle_record admitted a DTLS 1.0-framed HelloVerifyRequest only in sent_hello, while _handle_hello_verify_request answers one in sent_cookie_hello too. So a server that frames its challenges that way got its first answered and its second dropped, which is the OpenSSL behaviour that handler exists to avoid: both sides retransmit to the deadline with no alert. Found by @Jason-Morcos on #117. The three existing repeated-cookie tests call _handle_hello_verify_request directly, so none of them crossed the record layer where the gate sits. That is the same shape as the audit test which asserted the original repeated-HVR bug was correct. The version check also reads only the first message in the record, while _handle_handshake_fragment walks every message in it. A 1.0-framed record leading with a HelloVerifyRequest therefore carried a ServerHello straight into got_server_hello, which the framing test in test_dtls_psk_interop.py claims is refused -- true only while the ServerHello arrived in its own record. The exemption now travels with the record and admits one message. No peer is known to send a second 1.0-framed challenge: an OpenSSL server asked to re-challenge sends fatal alert 40 and one cookie, and the appliance frames 1.2. This is a consistency repair, not a measured break. Eight tests, five through the record layer; three fail without the fix and two pin behaviour it must not change.
The engine commit added cryptography>=38.0 to dependencies, but the floor job pins cbor2, pyOpenSSL and pytest and then installs with --no-deps, so nothing ever resolved against that floor: cryptography arrived as whatever pyOpenSSL 23.1.0 happened to pull, around 40.x. A declared floor no job exercises is not a floor. Verified at 38.0.0 on Python 3.11 with the pinned cbor2 and pyOpenSSL: the whole suite passes, which is unsurprising -- AES-CBC, HMAC, EC key generation and from_encoded_point all long predate it.
30807b1 to
d3228a2
Compare
|
@Jason-Morcos Reproduced the cookie gap, built the factory the way you described, and found a second bug with the same root cause. Pushed as The cookie gapYour table reproduces exactly:
A second one, same root causeThe version check reads the first message in a record; That contradicts both the comment calling the exception narrow and I should bound it, because I have not seen a peer do this either. An OpenSSL server asked to re-challenge by rejecting the echoed cookie sends fatal alert 40 and a single cookie. So this is a consistency repair rather than a measured interop break. The factory
Two details beyond what you described. The factory has to own The certificate path is a verbatim move. I traced the call sequence on
One correction on the HA side@mbillow, LocalThings has its own copy of the check. On the refrigerator header
What the firmware says about this pathReading the fork at
This client offers exactly one suite, so a reused peer entry is binary: ECDHE-PSK is in whatever that shared array last held and gets selected, or the handshake fails with One more in case it saves you time. The cookie's first four bytes are Where it stands1314 tests pass on 3.11 through 3.14 and at the declared dependency floor. I closed both loose ends. The Why it mattered at all: On sequencing, the seam is now #120 against |
|
@QuiteYellow Confirmed on I also reran the retained refrigerator server-flight replay against this head: it still reaches ClientKeyExchange and emits the full synthetic NUL-containing identity. Same limit as before: offline parsing/ECDH evidence, not a new authenticated hardware run. Agreed on the header wording. The matching 13 bytes corroborate the framing and cookie layout; they aren't a device fingerprint. The refrigerator evidence comes from the retained capture's provenance and the rest of that exchange, not from that header being unique. And the separate capture where the client answered both cookies but never got a ServerHello remains useful evidence against treating every HVR-only timeout as the same bug. I've answered the dispatch question on #120: One qualification on “a re-challenge always carries a changed cookie”: in the Mbed TLS 2.7.8 writer, the time-based path HMACs the timestamp and client ID using the existing cookie key. Two calls in the same clock second with the same key/client ID can therefore produce the same cookie. The 60-second setting is an expiration window, not forced rotation per challenge. So I’d keep the same-cookie/timer coverage too; our changed cookie at 31 seconds doesn’t identify why the peer challenged again. The reused-peer/shared-cipher-array explanation is worth testing, but I don’t have a hardware result establishing it as the cause of our old stall. I’d keep that source inference separate from the captured behavior, as you’ve done. |
The same-cookie branch said "our answer was lost rather than refused", which picks one explanation for something consistent with two. @Jason-Morcos corrected the premise behind it on #117: ssl_cookie_hmac is pure over a 4-byte timestamp and the client id under a key mbedtls_ssl_cookie_setup generates once, so a server re-challenging inside the same clock second emits a byte-identical cookie rather than a rotated one. An identical cookie therefore says nothing about whether our answer was lost or the peer challenged again. Behaviour is unchanged and still right -- the flight timer owns that retransmission either way, so the handshake must not be renumbered underneath it. Only the stated reason was wrong.
|
@Jason-Morcos You are right about the cookie. It had also reached code here, which I only found by checking. The same-cookie branch read "our answer was lost rather than refused", picking one explanation for something your correction makes consistent with two: a deterministic cookie means a peer re-challenging inside the same second is byte-identical, so an identical cookie says nothing about which happened. Your review landed on #120 as Indeed, @KRZ303's #575 already delegates to |
Draft, and deliberately not mergeable: nothing imports this engine, so merging it would ship a module the library never calls. It is here to be read.
PskAuthcannot work on OpenSSL. The DTLS 1.2 PSK client callback passes the identity as a NUL-terminatedchar *and returns the key length, so a recovered OwnerPSK identity containing a zero byte reaches the wire truncated. Measured against OpenSSL 4.0.2: 16 bytes in, 8 bytes out, andpsk_use_session_cb, which does carry a length, fires zero times for DTLS 1.2 across three configurations. #115 has the background.This is a DTLS 1.2 client speaking one ciphersuite,
TLS_ECDHE_PSK_WITH_AES_128_CBC_SHA256, oncryptography.What I am asking for
@Jason-Morcos this lands on your
PskAuthdesign: the engine replaces the DTLS transport underneath your auth provider. It now authenticates to KRZ303's oven, so there is working code to look at.The part worth your time is the handshake state machine and the record layer. The crypto primitives I can check against the appliance's own Mbed TLS and intend to.
Thirteen defects so far. Eight I found by writing attacks against my own code, three more on first contact with the oven, and two by checking the engine against this repository's earlier protocol fixes.
One of the thirteen is the reason this is a draft. A test named
test_hvr_ignored_after_cookie_helloasserted that a second HelloVerifyRequest should not reset the handshake transcript. It passed every run, because it encoded the same wrong assumption that produced the defect it was meant to catch: the engine answered the first cookie challenge and silently dropped any later one, which is the shape localthings#504 describes as the likely cause of #20's silent timeouts. A stub for reproducing it had been in the tree since August.A self-audit cannot find a defect its author has written into the tests, and there is nothing to diff against either. No Python library implements ECDHE-PSK, and the five hand-rolled pure-Python DTLS clients on GitHub are all Hue Entertainment, so they use plain PSK with AEAD records and share none of the hard part.
Already covered
The open design question
Routing PSK through this engine needs a connection seam that
AuthenticationProviderdoes not have. The Protocol isruntime_checkableand typed toOpenSSL.SSL.Context,configure_contextis documented public API with three implementations, anddtls_session.pyconstructs its context inline. Adding a second verb to that Protocol is a change to your design, so I have left it alone.WantRead,ZeroReturnandDtlsErrorare aliased to pyOpenSSL's exceptions here for the same reason.That is the next change, and I would rather agree the shape than present one.
Also in this diff
cryptographybecomes a declared dependency.protocol/auth.pyhas imported it directly since August, resolving it only as a transitive dependency of pyOpenSSL. The floor is 38.0 becausepyOpenSSL==23.1.0caps it below 41 and the floor job installs with--no-deps; the full suite passes there on Python 3.11.A second and much smaller ask, for anyone with PSK hardware
This needs no credential of your own. localthings#432 shows an appliance running the entire handshake against a deliberately invalid identity and failing only at
unknown PSK identity, which still exercises the ClientHello, the cookie round trip, the ServerKeyExchange parse, the ECDH and the ClientKeyExchange.If your appliance negotiates
ECDHE-PSK-AES128-CBC-SHA256, a trace of that failed handshake is useful, and the HelloVerifyRequest record header especially. The engine has met one oven and a laptop build of the appliance firmware's own stack. A second device family would show whether that reference server imitates real hardware or only itself.