fix: avoid quadratic buffering of long SSE lines - #76
Conversation
The async reader is a hand-maintained parallel copy of the sync reader, so it had the same quadratic buffering of a long line. It rebuilt the growing partial line on every chunk. This applies the same fix: collect the fragments in a list, and join them one time when the line ends. CR, LF and CRLF handling is unchanged, and an unterminated final line is still discarded. The new tests mirror the sync tests for fragmented input with empty chunks and for UTF-8 characters split across chunk boundaries. A 16 MiB line in 10,000-byte chunks now reads in about 10ms instead of about 790ms, and the time per doubling of the line size drops from about 4x-6x to about 2x.
Both branches of the pending-fragment block started with the same append, so the append moves above the branch. This also makes the sync and async readers read the same, which matters because they are parallel files that are maintained by hand. The behavior does not change.
|
Thanks for the thorough writeup and the reproducible benchmark. The analysis is right and the sync reader change is correct. I pushed two commits to the branch rather than asking you to round-trip on them. What changed
Worth flagging why this was easy to miss:
Async reader performanceSame shape as your benchmark: one long Python 3.13.0, Intel Core Ultra 7 155H, Linux 7.1.5 x86_64.
Line size excludes the two trailing LF bytes; 1 MiB = 1,048,576 bytes. Doubling the line size costs about 2x after the fix, against 3.5x to 5.5x before it. Do not read too much into the 19.68x on the last row: at 32 MiB the unpatched reader is thrashing the allocator on top of the copying, so that figure is above the underlying quadratic trend rather than a truer measure of it. Correctness verificationI checked both readers as behavior-preserving refactors rather than reading them closely, since the new fast path reorders when partial data is merged. Against the pre-PR implementations, over every possible chunking of every byte string in Two invariants make the fast path safe, for the record. One trade-off worth recordingThe fragment list costs roughly 41 bytes per fragment, so peak memory now depends on chunk size in a way it did not before. Peak above baseline for a 2 MiB line (
Identical at production chunk sizes, and 1-byte chunks are not reachable through |
🤖 I have created a release *beep* *boop* --- ## [1.7.3](1.7.2...1.7.3) (2026-09-17) ### Bug Fixes * avoid quadratic buffering of long SSE lines ([#76](#76)) ([8ede399](8ede399)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
|
@bach-ta thank you again for your contribution. This has been released in v1.7.3 |
Requirements
Related issues
No linked issue.
Describe the solution you've provided
When a long SSE line spans many chunks,
_BufferedLineReader.lines_from()repeatedly concatenates the accumulated partial line with the next fragment. Each concatenation copies the growing prefix, making buffering quadratic in line length for a fixed chunk size.This change collects fragments in a list and joins them once the line ends. It preserves the existing splitting and decoding logic, including CR/LF/CRLF handling and discarding an unterminated final line.
The tests add coverage for fragmented lines with empty chunks, CRLF split across chunks, and UTF-8 characters split across byte boundaries.
Describe alternatives you've considered
Larger HTTP chunks reduce the number of copies but leave the quadratic behavior. A mutable byte buffer is another option; retaining the existing fragments and joining once keeps the change small and preserves the current parsing logic.
Additional context
The benchmark compares the published
launchdarkly-eventsource==1.7.2reader (d10fd70) against this patch (745b247) on Python 3.12.14, macOS 15.7.3, Apple M4 Pro.Both implementations read the same synthetic
data:line in 10,000-byte chunks. Inputs are prepared before timing; the timed region consumes the reader's output, including decoding. Every result is checked against the exact expected list of lines.Each case runs nine times per implementation, in separate fresh processes with an identical warm-up. Case and implementation order are randomized with seed 0. Values below are median wall times, with Q1–Q3 in brackets. Growth is relative to the preceding row, calculated from unrounded medians.
Line size excludes the two trailing LF bytes; 1 MiB = 1,048,576 bytes. Doubling the line size takes approximately 4–7× as long in the original reader and approximately 2× with the patch. A separate short-line control (about 1 MiB total) measured 1.30ms original / 1.32ms patched.
These measurements isolate line reading and decoding, not HTTP or full SDK initialization. Absolute timings will vary by machine.
Reproduction
Save the two scripts below as
bench_patch.shandbench_patch.pyin the same directory. They require macOS or Linux, Bash, Git,tar, and Python withvenvandpip. Package installation requires network access.With a clean local checkout of this PR, run from the directory containing both scripts:
The script compares release 1.7.2 with the checkout's committed
HEADand writes the summary, individual measurements, source snapshots, and environment metadata toreader-benchmark/. It also measures additional chunk sizes and the short-line control. Each run replaces the previous output.bench_patch.shbench_patch.pyValidation
mypy,isort,pycodestyle).Note
Overview
Fixes quadratic-time buffering when SSE lines arrive split across many byte chunks in both sync (
_BufferedLineReader) and async (_AsyncBufferedLineReader) line readers.Instead of prepending each new fragment onto a growing
partial_line(re-copying the prefix every chunk), partial bytes are accumulated inpending_fragmentsandb"".joinruns once when the line terminator appears. CR/LF/CRLF handling and withholding an unterminated final line are unchanged.Tests add coverage for fragments with empty chunks, terminators split across chunks, and long UTF-8 lines split on byte boundaries; the sync mixed-terminator case is adjusted to exercise chunk boundaries more strictly.
Reviewed by Cursor Bugbot for commit 573d3c6. Bugbot is set up for automated code reviews on this repo. Configure here.