Skip to content

Add TLS session resumption via SSLSessionCache - #789

Merged
dkropachev merged 6 commits into
scylladb:masterfrom
sylwiaszunejko:tls-ticket
Sep 22, 2026
Merged

dkropachev merged 6 commits into
scylladb:masterfrom
sylwiaszunejko:tls-ticket

Conversation

@sylwiaszunejko

@sylwiaszunejko sylwiaszunejko commented Apr 3, 2026

Copy link
Copy Markdown

What and why

A shard-aware driver opens one TLS connection per shard to every node, and each one currently
pays for a full handshake — certificate exchange plus a signature, which is the expensive part,
especially with certificate authentication. TLS lets a client skip that by replaying a session
established earlier with the same peer (RFC 5077 tickets for TLS 1.2, RFC 8446 PSKs for
TLS 1.3), but OpenSSL never does this on its own: the client has to hold on to the session and
offer it explicitly on the next connection. Neither the stdlib ssl module nor pyOpenSSL
exposes SSL_CTX_sess_set_new_cb, so there is no way around doing it by hand.

This adds that: one SSLSessionCache per Cluster — a new cassandra.ssl_session_cache
module — offered to every connection before its handshake and refreshed after. On by default
whenever ssl_context is set and the reactor can restore a session before the handshake.

cluster = Cluster(ssl_context=ssl_context)                                  # resumption on
cluster = Cluster(ssl_context=ssl_context, ssl_session_cache=None)          # off
cluster = Cluster(ssl_context=ssl_context,
                  ssl_session_cache=SSLSessionCache(max_size=64))           # sized, or shared

The server has to issue tickets for any of this to happen. Scylla does so only when
enable_session_tickets is set in client_encryption_options, which defaults to true from
Scylla 2026.3 and to false in the releases before it. Without it nothing resumes and every
connection performs a full handshake, exactly as it does today.

Design notes

  • A cached session is not consumed by being used. get() leaves the entry in place, and
    each successful handshake stores a fresh session over it. Measured: one session is accepted by
    four concurrent connections on TLS 1.2 and 1.3 against a loopback TLS server, and by every
    pool connection against real Scylla. Treating tickets as single-use (removing on get())
    would mean only the first connection of a per-shard burst resumes — precisely the case this
    ticket is about. RFC 8446's "SHOULD NOT reuse" concerns 0-RTT replay and tracking; the driver
    sends no early data.
  • The session is stored from the ReadyMessage / AuthSuccessMessage handlers, not right after
    the handshake.
    A TLS 1.3 server sends its NewSessionTicket as a post-handshake message;
    confirmed against Scylla that has_ticket is False immediately after connect() and True
    after the first CQL exchange. Storing is idempotent, so nothing needs to track whether it
    already happened, and every failure in this path is logged and dropped — both call sites are
    wrapped in @defunct_on_error, where a raised exception would kill a healthy connection over
    an optimisation.
  • An entry is never offered past the lifetime the peer announced. A ticket carries one, and
    what remains of it after the handshake and CQL startup is what the entry gets, measured from a
    monotonic mark rather than the wall-clock SSLSession.time. A zero lifetime means opposite
    things in the two RFCs, so the negotiated version decides: discard on TLS 1.3, "unspecified"
    on TLS 1.2. RFC 8446 caps a client at seven days regardless. Re-storing a session the peer
    handed back unchanged keeps the deadline it had, so one ticket cannot be kept alive by being
    reused.
  • The SSLContext is part of the cache key, together with the endpoint and the name the
    peer certificate is verified against. A session cannot be replayed onto a different context —
    the stdlib rejects it with ValueError: Session refers to a different SSLContext — which is
    also why the deprecated ssl_options-only path does not participate: each of those
    connections builds its own context. The verified name is in the key because a resumed
    handshake sends no Certificate, so that name is never checked again. The shard-aware port is
    the exception: that endpoint carries the node's key, so the per-shard burst resumes the
    session the control connection established rather than handshaking in full.
  • A session offered on a connection whose TLS handshake then failed is retracted, whether
    the failure surfaces at connect time or, where ssl_options defer the handshake, in the
    reactor. Only a TLS error counts, and only the session that connection offered — a sibling may
    have stored a newer one meanwhile, and that one failed nothing.
  • A cache you supply is yours. shutdown() leaves its entries alone, so clusters can share
    sessions, at the same time or one after another. One the driver created for a cluster is
    reachable only through it and goes when it does.
  • Policy is separated from accessors (_set_tls_session, _get_resumable_tls_session,
    _tls_negotiated_version) so a reactor not using the stdlib ssl module overrides only those
    three; nothing else in the policy touches a socket, and a test holds that by driving it with
    no socket at all.

Not covered

  • asyncio — the handshake happens inside loop.create_connection(..., ssl=...), which offers
    no point at which a session could be restored. AsyncioConnection declares
    supports_tls_session_resumption = False and no cache is created for it. Worth knowing that
    this is the default reactor on Python 3.12 and newer when the libev extension is not
    installed, since asyncore left the standard library there.

Fixes: https://scylladb.atlassian.net/browse/DRIVER-165

Pre-review checklist

  • I have split my patch into logically separate commits.
  • All commit messages clearly explain what they change and why.
  • I added relevant tests for new features and bug fixes.
  • All commits compile, pass static checks and pass test.
  • PR description sums up the changes and reasons why they should be introduced.
  • I have provided docstrings for the public items that I want to introduce.
  • I have adjusted the documentation in ./docs/source/.
  • I added appropriate Fixes: annotations to PR description.

@Lorak-mmk

Copy link
Copy Markdown

This reduces reconnection latency and CPU overhead, especially in
deployments with short-lived connections or frequent reconnects.

Such claims would ideally be supported by benchmarks. Could you try to create some?
I very vaguely remember this feature being postponed because the performance gains were underwhelming (but perhaps memory is failing me).

@sylwiaszunejko

Copy link
Copy Markdown
Author

This reduces reconnection latency and CPU overhead, especially in
deployments with short-lived connections or frequent reconnects.

Such claims would ideally be supported by benchmarks. Could you try to create some? I very vaguely remember this feature being postponed because the performance gains were underwhelming (but perhaps memory is failing me).

That's the goal, but you're right, I don't have any tests to prove that, removed this claim from the PR description. If I manage to create proper benchmarks I will update on that

@mykaul

mykaul commented Apr 3, 2026

Copy link
Copy Markdown

We could, if it helps, only support this for TLS 1.3.

@sylwiaszunejko

Copy link
Copy Markdown
Author

@dkropachev @Lorak-mmk I pushed changes with improvement from older Dmitry's PR, will update PR description soon

@dkropachev dkropachev left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I rechecked the TLS session-resumption path against the current branch. The ssl_options configuration still builds a fresh SSLContext per Connection, and a cached stdlib session from the previous connection is incompatible with that new context. I reproduced the failure locally on Python 3.10.12; the session restore path raises ValueError: Session refers to a different SSLContext. Since the new code only catches AttributeError and ssl.SSLError, reconnects fail instead of falling back to a full handshake, and the regression is enabled by default because Cluster auto-creates SSLSessionCache for ssl_options.

Comment thread cassandra/connection.py Outdated

@dkropachev dkropachev left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two blocking issues from local validation:

  1. Twisted caches a TLS session even after hostname verification has already failed, which lets an untrusted peer populate the resumption cache.
  2. SSLSessionCache accepts max_size <= 0 and then crashes on the first insert (KeyError from popitem() on an empty OrderedDict).

Comment thread cassandra/io/twistedreactor.py Outdated
Comment thread cassandra/connection.py Outdated

@dkropachev dkropachev left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two correctness issues need attention before this lands: the PyOpenSSL TLS 1.3 cache point is too early to capture the resumable session, and the cache can evict a live entry while expired ones remain resident.

Comment thread cassandra/io/twistedreactor.py Outdated
Comment thread cassandra/connection.py Outdated
Comment thread cassandra/io/twistedreactor.py Outdated
Comment thread tests/integration/standard/test_tls_resumption.py Outdated
Comment thread cassandra/cluster.py Outdated
Comment thread tests/integration/standard/test_tls_resumption.py Outdated
@coderabbitai

coderabbitai Bot commented Jul 15, 2026

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Adds SSLSessionCache, a thread-safe bounded LRU cache for TLS sessions. Cluster creates, disables, or accepts a cache and passes it to connections. Connections derive endpoint-specific keys, restore sessions before handshakes, and store sessions after startup or authentication. Reactor implementations declare unsupported resumption. Tests cover cache behavior, TLS 1.2 and TLS 1.3, concurrency, cache isolation, and shard-aware connections.

Possibly related PRs

Suggested reviewers: mykaul, dkropachev, lorak-mmk

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 19.83% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed The description includes a valid Fixes annotation for DRIVER-165.
Out of Scope Changes check ✅ Passed The changes support the stated TLS session resumption objective and add related implementation, reactor declarations, documentation, and tests. No unrelated changes are evident.
Title check ✅ Passed The title clearly and concisely identifies the main change: adding TLS session resumption through SSLSessionCache.
Description check ✅ Passed The description explains the motivation, design, scope, configuration, limitations, tests, issue reference, and completed checklist items. It satisfies the repository template.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai
coderabbitai Bot requested a review from dkropachev July 15, 2026 08:11

@dkropachev dkropachev left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One inline correctness issue. Branch also needs a rebase onto current master; PR is currently conflicting.

Comment thread cassandra/connection.py Outdated
Comment thread cassandra/cluster.py Outdated
Comment thread tests/unit/test_tls_resumption.py
Comment thread tests/integration/standard/test_tls_resumption.py Outdated
@nikagra

nikagra commented Sep 7, 2026

Copy link
Copy Markdown

The design is right where it matters — one cache per Cluster, restore before the handshake, store from Ready/AuthSuccess, context and hostname in the key. None of that is what I want to revisit.

What I would like to settle before the next push is the shape of the bookkeeping around it, because each round so far has been answered with a parameter or a method rather than a change of mechanism:

round what was added
ticket lifetime ignored set(lifetime=)
deadline slides on reuse set(extend=) + _tls_session_was_renewed()
context retention weakref → discard_context()acquire_context/release_context + _context_owners + two Cluster attributes + a shutdown hook
bad entry re-offered discard(key, session) + _tls_session_offered cleared in three places
shard-aware keying _tls_session_cache_key_override set from pool.py

Two changes would collapse most of that.

1. The cache has an owner, and the driver knows which it is. An auto-created cache is a Cluster attribute, so its lifetime is already the cluster's and it needs no refcount. A user-supplied one is the user's object, and the driver deleting rows in it is the part I would question. Recording which of the two it is at construction removes acquire_context, release_context, _context_owners, both Cluster attributes, the shutdown hook and the k[0] is ssl_context scan — along with the cases each of those had to handle: cluster that never connected, shutdown that raised, cache assigned after construction, keys that are not tuples.

2. The accessor boundary should hold. _tls_session_lifetime and _tls_session_was_renewed read self._socket directly, against the note at :1352. @dkropachev's open twisted and eventlet threads are the same boundary from the reactor side.

Order I would suggest: 1, then 2 (which also removes the version-guessing in _get_resumable_tls_session), then move the extend decision inside set(), where the lock is.

Separately: five threads are resolved while the code at each is unchanged — cluster.py, connection.py, pool.py, cluster.py, test_tls_resumption.py. If any of those was a decision rather than an oversight, say so and I will drop it.

Deliberately not raising this round, to keep it to the shape above: docs for an on-by-default change, SSLSessionCache.set/discard edge behaviour, the asyncio hostname duplication, the SniEndPoint key override, unit-suite cost.

@nikagra nikagra left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Detail under the comment above — the two changes there would remove most of what these three sit on.

Comment thread cassandra/connection.py Outdated
Comment thread cassandra/connection.py
Comment thread cassandra/connection.py Outdated
@sylwiaszunejko
sylwiaszunejko force-pushed the tls-ticket branch 2 times, most recently from b540378 to a9124b7 Compare September 8, 2026 10:53

@nikagra nikagra left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both findings below are from the 2026-09-08 rework of the cache core, not the earlier rounds -- the rest of the backlog reads as closed to me. One follow-up on the TLS 1.3 discard thread separately.

Comment thread cassandra/connection.py Outdated
Comment thread cassandra/cluster.py Outdated
Comment thread cassandra/connection.py Outdated
Comment thread cassandra/cluster.py Outdated
Comment thread cassandra/connection.py Outdated

@nikagra nikagra left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Major] 🟠 The description still contradicts the code on the point @dkropachev asked to have updated in #789 (comment):

  • "No TTL. OpenSSL enforces session lifetime itself" — it doesn't, and the driver now does. _tls_session_lifetime() takes the lifetime the server announced, capped at _MAX_TLS_SESSION_LIFETIME, and SSLSessionCache drops an entry past its deadline both on lookup and on eviction. That bullet needs to describe the rule that shipped, including the TLS 1.2 / 1.3 split on a zero lifetime.
  • "Policy is separated from accessors (_get_resumable_tls_session / _set_tls_session)" — three now; _tls_negotiated_version is the one the retention rules read.
  • "Measured: … stdlib and pyOpenSSL" — no pyOpenSSL reactor is left in the tree. Worth saying the measurement predates their removal, or dropping that half.
  • The deferrals are agreed but unlinked: #984 (test capability detection), #985 (multi-ticket caching — #789 (comment) asked for this link explicitly), #986 (tickets after CQL startup). A "Deferred" section naming the three closes two bot threads with it.

The checklist's docstring and docs boxes are also unticked on a PR that adds docs/api/cassandra/ssl_session_cache.rst.


Notes on the process

Separately, and about process rather than code. I have closed six of my seven open threads with this pass; the one left is #789 (comment), which is now on its third restatement. Worth noting why that happens, because the code here is converging and the threads are not: git diff 637f384c 8c74beec -- cassandra/connection.py is five hunks — one import, one docstring cross-reference, the 190-line SSLSessionCache move into its own module, one docstring reword, and one line from the rebase. No functional change at all, yet 15 threads are open and only three of them anchor to current code.

Three things would end this:

  • @coderabbitai pause. 33 of the 100 review threads on this PR are bot-opened (copilot 25, coderabbit 8), several were stale when posted, and they bury the human blockers.
  • Settle the retention policy in DRIVER-165, not here. Lifetime source, cache ownership and failed-handshake behaviour have each been designed twice in review — two complete mechanisms were built and then deleted (SSLContext retention went strong key → weakref key → discard_context → none of it; _decide_tls_session_cache plus three bookkeeping flags → a property). Get an explicit ack from @dkropachev on the policy, then review conformance only.
  • Cut the prose to contract level. cassandra/ssl_session_cache.py is 70 lines of code to 141 of docstring and comment; _is_same_session is 23 lines of docstring over 4 of code. Rationale that is not a caller's contract belongs in DRIVER-165 or a commit message. As it stands every behaviour change forces a paragraph rewrite, which manufactures new review surface — two of my findings this round, and a good share of every round before it, are prose contradicting code rather than defects.

If the next round is not green, the way out is three stacked PRs — (a) ssl_session_cache.py + EndPoint.tls_session_cache_key + tests, no behaviour change; (b) connection wiring; (c) the Cluster surface — one reviewer concern each.

Comment thread cassandra/cluster.py
Comment thread tests/unit/test_cluster.py
Comment thread tests/unit/test_ssl_session_cache.py
Comment thread tests/integration/standard/test_tls_resumption.py
Comment thread cassandra/ssl_session_cache.py Outdated
Comment thread cassandra/connection.py
Comment thread cassandra/connection.py
Comment thread docs/api/index.rst Outdated
Comment thread cassandra/connection.py
Comment thread cassandra/connection.py
Comment thread cassandra/cluster.py Outdated
Comment thread tests/integration/standard/test_tls_resumption.py

@nikagra nikagra left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

First pass over the head commits — the TLS mechanics hold up end to end (session restored pre-handshake, offered-identity retraction, monotonic lifetimes, the TLS 1.2/1.3 split). Nothing below blocks merge.

Comment thread cassandra/pool.py
Comment thread cassandra/cluster.py Outdated
Comment thread tests/unit/test_cluster.py
Comment thread cassandra/cluster.py
Comment thread tests/integration/standard/test_tls_resumption.py Outdated
Comment thread cassandra/cluster.py Outdated

@nikagra nikagra left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. My open thread on the TLS 1.3 discard path is closed -- option 1 landed as written. Two non-blocking notes inline.

Comment thread docs/security.rst
Comment thread cassandra/ssl_session_cache.py
TLS clients can skip the expensive part of a handshake by replaying a
session established earlier with the same peer (RFC 5077 tickets for
TLS 1.2, RFC 8446 PSKs for TLS 1.3), but OpenSSL never does this on its
own: the client has to hold on to the session and offer it explicitly on
the next connection.

Add the storage half of that: a bounded, thread-safe LRU of TLS sessions
keyed by TLS peer identity, plus an EndPoint.tls_session_cache_key
property that produces the key.  The cache is a module of its own,
cassandra.ssl_session_cache, because keeping sessions is all it does: what
may be offered to whom, and for how long, belongs with the connections that
work it out.  A cached session is not consumed by
being used -- one session can be replayed by any number of concurrent
connections -- so get() leaves the entry in place and each successful
handshake stores back over it whatever the peer handed over.

Entries carry the lifetime the caller gives them and are dropped once it
runs out, so a session is never offered past the point the peer said it
would honour it; what that lifetime should be is for the caller to work
out, since it depends on how the session resumes.  Whether a store moves an
existing deadline is decided here rather than asked of the caller: a session
the peer handed back unchanged keeps the deadline it had, since its lifetime
runs from when the peer issued it and not from when it was last replayed, and
re-stamping a full lifetime on every reuse would let one ticket be offered
for as long as connections keep being opened.  Below TLS 1.3 an abbreviated
handshake hands back the offered session itself, so that is the ordinary
case; TLS 1.3 normally issues a fresh one, which is entitled to a lifetime of
its own.  The two are told apart by session id, compared under the same lock
that stores the result, so no concurrent store can land in between.

An entry recognised this way keeps the object holding it as well as its
deadline.  SSLSocket.session builds a new wrapper on every access, so what a
resumed connection stores back is another handle on the credential already
there; swapping one for the other would change nothing but the identity a
caller compares against when it asks for its own session to be dropped.
Keeping it means an entry changes identity only when it changes credential.

A caller can also say which session it offered on the handshake it is
storing the result of, and a store of that same session is skipped where the
entry no longer holds it: another connection has stored a session the peer
reissued to it, or the deadline passed and a lookup dropped the entry.  In
neither case has this caller anything to add, while storing it would put a
deadline running from now on a session the peer issued at some earlier
point -- and would replace a fresher session with an older one.

Neither comparison can say anything about a ticket whose session id is
empty, which RFC 5077 section 3.4 lets a server send and for which
SSLSession offers no ticket to compare instead.  Two of those are reported
as the same session, the conservative reading: the deadline then stays where
it is, where calling them different would re-stamp a full lifetime on what
may well be the ticket already held.  What that costs is resumption rather
than correctness -- a reissued ticket inherits its predecessor's deadline,
and one arriving where the entry has since gone is not stored, which the
next connection puts right by offering nothing and storing afresh.

SNI endpoints add the server name to their key, since they all share a
proxy address and port but are distinct TLS peers.  Client-routes
endpoints key on the node's host_id rather than the proxy address they
happen to resolve to at the moment.

Nothing uses the cache yet.

Refs DRIVER-165
Offer the cached session for the endpoint before the handshake, and
store the negotiated session once the connection is up, so that the next
connection to the same node -- in particular the burst of per-shard
connections a pool opens at once -- can skip the certificate exchange
and signature of a full handshake.

The session is stored from the ReadyMessage / AuthSuccessMessage
handlers rather than right after the handshake.  A TLS 1.3 server sends
its NewSessionTicket as a post-handshake message, so a session read
straight after connect() carries no ticket and would not resume; by the
time the CQL handshake has completed the ticket has been read off the
socket.  Storing is idempotent, so nothing needs to track whether it
already happened, and every failure in this path is logged and dropped:
resumption is an optimisation, and both call sites are wrapped in
@defunct_on_error, where a raised exception would kill a healthy
connection.

How long a session may be offered is worked out here, because it depends on
how the session resumes: a ticket's lifetime is the one the server
announced, while SSLSession.timeout is only the local context's default and
says nothing about what the peer will still accept, so it is used solely for
a session that resumes by id.  RFC 8446 section 4.6.1 also caps the
client at seven days however long the server asked for.  A zero lifetime is
read against the negotiated version, since the two RFCs disagree on it: TLS
1.3 says discard the ticket immediately, while RFC 5077 section 3.3 reserves
zero for "lifetime unspecified" and leaves retention to local policy, so a
TLS 1.2 ticket is kept and timed by the local timeout.  The version also
rules out caching a TLS 1.3 session that carries only an id: resumption
there is the ticket's pre-shared key, while the id a TLS 1.3 handshake
carries is the legacy_session_id_echo a server sends back for the middlebox
compatibility mode of RFC 8446 appendix D.4, which resumes nothing.  OpenSSL
reports no id at all until a NewSessionTicket has been read, so that session
is not one it produces; stating the rule against the negotiated version
rather than against what one library exposes is what makes it hold
regardless.
OpenSSL does not apply either limit on the client's behalf -- it will offer
an expired ticket and let the server refuse it.  What the entry gets is what
is left of the announced lifetime: the peer issued the ticket during the
handshake, while the store happens a startup exchange and perhaps an
authentication later, and that time is the peer's to spend rather than the
entry's.  The age is measured from a monotonic mark taken as the handshake
began, not from SSLSession.time, which is a wall-clock stamp: subtracting
that would let a clock step landing in between decide the answer -- far
enough forward and nothing is cached at all, backward and the age
disappears.  The deadline the cache keeps is monotonic too, so nothing after
the store can skew it either.

A pool reaches a shard-aware node on a second port, which would otherwise
key those connections separately from the one the control connection
established, leaving the whole per-shard burst to handshake in full.  The
endpoint alias that _get_shard_aware_endpoint already builds for that port
therefore carries the node's cache key, so both listeners share one session
and nothing has to be threaded through the connection factory.  The port
stays part of the key by default, so two unrelated TLS servers on one
address still cannot share a session; only an endpoint that names another
node is exempt.

The SSLContext is part of the key because a session cannot be replayed onto
a different one and a cache may be shared by several clusters.  It is held
strongly there: a cached session already keeps its context alive on its own
-- CPython's SSLSession holds a reference to the context it was established
with -- so holding it weakly here would buy nothing.

The key also carries the name wrap_socket() is given, which is the name
the peer certificate is verified against.  A resumed handshake sends no
Certificate, so that name is never checked again; offering a session to a
connection expecting a different name would silently skip hostname
verification for it.  Both the key and wrap_socket() take the name from
one accessor so the two cannot drift apart.

A session offered on a connection whose handshake then failed is dropped from
the cache.  Both RFCs have a server fall back to a full handshake rather than
fail when it will not resume, so this should not happen; but nothing stores a
fresh session for a connection that never came up, so an entry that did
provoke a failure would otherwise be offered again by every later connection
until its lifetime ran out.  The handshake need not fail where it is started:
ssl_options may carry do_handshake_on_connect=False, which leaves it to the
first read or write and so to the reactor, and a failure there arrives at
defunct() instead -- which retracts for the same reason and under the same
rule.  Only a TLS error counts: a refused or reset connection says nothing
about the session.  And only the session this
connection offered goes: connections to one node are opened together, so
another may have stored a session the peer issued in its place, and that one
failed nothing.  A sibling storing the same session back is not that, and
leaves the entry retractable, because the cache keeps the object it already
holds for a session it recognises.  Nor does an endpoint that borrowed its key
retract anything: the shard-aware listener resumes the node's session and
stores back to it, which is the point of sharing the key, but the entry
belongs to the endpoint that owns it.  Dropping it there would cost the node's
other pools and its control connection a handshake apiece, and a pool filling
against that listener would do it again on every retry, while a session that
genuinely cannot resume fails on the owner too and goes from there.

What was offered is kept until the store, which hands it to the cache and
then lets it go.  Telling a session the peer reissued from the one that came
back unchanged needs both sides of the exchange, and where connections to a
node are opened together only the connection that offered one knows the
second.  Letting it go there is what bounds the retraction above: reaching
the store means the CQL handshake completed, so the TLS one stood, and a TLS
error long afterwards has nothing of this connection's to withdraw.

Three accessors are the whole of what a reactor whose TLS does not go
through the stdlib ssl module has to reimplement to take part: the policy
around them asks one for the session to store, one for the negotiated
version and one to restore a session onto a socket, and reads nothing off a
socket itself.  Connections whose SSLContext is
derived from ssl_options do not participate, because a session cannot be
replayed onto a different context and each of those connections builds
its own.  The asyncio reactor opts out entirely: its handshake happens
inside loop.create_connection(), with no point at which a session could
be restored.

Refs DRIVER-165
Create an SSLSessionCache per Cluster whenever TLS is configured through
ssl_context, and hand it to every connection the cluster opens, so that
resumption is on by default with no configuration.  Pass
ssl_session_cache=None to turn it off, or an instance of your own to size
it or share it between clusters.

No cache is created where resumption cannot work: the deprecated
ssl_options-only path, whose per-connection SSLContexts a session cannot
be replayed onto, and reactors that report they cannot restore a session
before the handshake, which today means asyncio -- which is also what the
default connection class resolves to on Python 3.12 and newer with no libev
extension installed, asyncore having left the standard library there.
connection_class is not required to derive from Connection, so one that does
not report the capability at all is treated as lacking it rather than
raising.

Nothing is settled in advance.  ssl_session_cache is a property, answering
against whatever ssl_context and connection_class are in force when it is
read and making a cache on first use where one is wanted -- once, under a
lock of its own, however many threads ask at the same moment, since for a
cluster given TLS after connect() the first to ask are whichever pools are
opening connections.  Both of those are
public attributes, so a decision taken at construction would leave resumption
off on a reactor that does support it, or hand the keyword to a connection
class that does not take it -- and one retaken later needs the first to be
remembered, which is how an explicit None came to be overwritten.  Behind the
property is what the caller set and nothing else: assigning None turns
resumption off and leaves it off, and assigning a cache asks for one as much
as passing it to the constructor does, whenever it is done.

A cache the caller asked for that cannot be used reads back as None, and
connect() says why.  Asking for resumption and silently getting none is
worse than not having it: a cache answered back here would be handed to
every connection -- which a connection class that does not take the keyword
cannot even accept -- and would sit reachable and empty for anyone reading
it back, which is also what a server that issues no tickets looks like.
Only what the caller set is worth interrupting anyone over, since a cache
made for a cluster is only ever made where it can be used; turning TLS off
afterwards is not something to complain about. Where nothing was asked for
there is still something worth finding -- resumption is on by default
wherever it works, so a cluster that configured TLS and will not get it is a
surprise -- and that is a line at debug rather than a warning. A cluster
with no TLS at all is told nothing, having nothing to resume. connect() is
the one place that says any of it, so it is said once without anything
having to record that it was; a cache assigned after that reads back as None
just the same, without a second word about it.
Saying it from the setter instead would speak too early, while the caller is
still configuring -- assigning the cache before the context is an ordering
that owes nobody a warning.

Whose the cache is settles what becomes of it, and nothing has to track that
either: the field behind the property holds what the caller set, and one
made for the cluster is kept apart from it.  One created here is reachable
only through the attribute, so it and the sessions in it go when the cluster
does and shutdown has nothing to do.  One the caller supplied stays the
caller's: shutdown leaves its entries alone, which is what lets clusters
share sessions -- at the same time, or one after another, so that a cluster
replacing an earlier one resumes rather than handshaking in full -- and
keeps the driver from deleting rows in an object it does not own.  Its
entries hold the SSLContext their session was established with, bounded by
the cache's max_size, and clear() is there for a caller who wants them gone
sooner.

Refs DRIVER-165
Stand up a TLS server on loopback and connect to it with the driver's own
socket setup, so the restore-before-handshake and store-after-startup
paths run for real and the result is read back the way OpenSSL reports
it, through SSLSocket.session_reused.  Covers TLS 1.2 and TLS 1.3, the
latter skipped where the local OpenSSL does not offer it -- skipping the
subclass rather than the base, since a skipped base would take its
subclasses with it.

The certificates come from tests/tls_certificates.py rather than from here,
so that the integration suite generates its own the same way -- the
addresses to name are the caller's argument, which is what that suite needs
and this one takes the loopback default for. cryptography is optional, so
that module reports whether it found it and both suites skip on that.

One waits for the monotonic clock to move before opening its second
connection. What a store works out as the deadline is the mark taken when
the handshake began plus the lifetime announced -- the two readings of the
clock cancel -- so two handshakes it cannot tell apart are given the same
deadline, which is right but leaves a comparison of the two nothing to see.
Windows resolves that clock to about ten milliseconds, and two loopback
handshakes fit inside it.

Two of these pin down behaviour that is easy to regress: that four
connections opened at once all resume from the single cached session --
the per-shard burst DRIVER-165 is about -- and that on TLS 1.3 nothing is
cached until the server's NewSessionTicket has actually been read off the
socket.

Refs DRIVER-165
Restart the cluster with client encryption on, warm a session cache with
one cluster, then hand it to a second one and require every connection it
opens to have resumed -- which is the question only a real server can
answer: whether it accepts one session offered concurrently by the whole
batch of per-shard connections.

The cluster is given a shard-aware TLS port, since that is the port those
per-shard connections use and therefore where resumption has to pay off;
Scylla leaves it unset by default.  The certificate names every node
rather than only the contact point, or the driver could not build pools to
the rest of the cluster and the test would quietly examine a single host.
Each Session is held for the duration of a test: Cluster.sessions is a
WeakSet, so a dropped Session takes its pools -- everything worth
inspecting -- with it and leaves only the control connection behind.  The
number of connections collected is asserted before their resumption
flags, so the test cannot pass by examining almost nothing.  A connection
that has lost its socket is left out of the reading rather than counted as
not having resumed -- SSLSocket.session_reused answers None once the socket
is closed, which would read as a resumption failure for a connection that is
simply gone -- and it is that count which keeps leaving one out from
quietly shrinking the sample.  The key and
certificate outlive the run when KEEP_TEST_CLUSTER does: remove_cluster() is
a no-op there, and what it leaves behind still names those files.

Whether there is anything here to test at all depends on the reactor, so the
skip asks the connection class Cluster will instantiate rather than reading
EVENT_LOOP_MANAGER: with no selector set and no libev to import, that class
is the asyncio reactor, which cannot restore a session before the handshake.

Follows the reconfigure-and-remove pattern the other modules here use for
cluster-level options, and generates the server certificate with
cryptography, through the generator the loopback suite also uses, so the
test does not depend on an openssl binary.

Refs DRIVER-165
Resumption is on by default and has prerequisites on both sides, which until
now were described only in the ssl_session_cache docstring -- read by
someone already looking at the attribute, rather than by someone setting up
TLS. Give it a section in the security guide, next to the SSL configuration
it belongs with, covering what the server and the reactor each have to
provide -- and what resumption costs, which is the part a security guide
owes a reader most.

Scylla only issues session tickets when enable_session_tickets is set in
client_encryption_options.  That defaults to true from Scylla 2026.3 and to
false in the releases before it, so on an earlier one a cluster has to be
told to issue them; without them the cache stays empty and every connection
performs a full handshake, with no indication of why.

What it costs is verification. A resumed handshake carries no Certificate
message, so nothing about the server's certificate is checked again while a
cached session is offered -- not its expiry, and not a revocation list the
context carries -- and one that expires or is revoked goes on being accepted
until the entry does, bounded by the lifetime the server announced and the
driver's seven-day cap. The hostname is the exception, and only because the
cache is keyed by the name verified when the session was established, so a
session is never offered where a different name is expected. Someone who
cannot accept that window is told how to turn resumption off.

The reactor and server paragraphs move off the class rather than being
copied, leaving the docstring to the semantics of the attribute -- what it
holds, whose it is, how to turn it off or size it -- and a pointer to the
guide for what has to be true for any of it to happen.

Refs DRIVER-165
@sylwiaszunejko

Copy link
Copy Markdown
Author

CI failures due to flaky tests

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants