Skip to content

Optimize extend_back(); add extend_back_from_slice() - #35

Open
ajasmin wants to merge 2 commits into
andylokandy:masterfrom
ajasmin:bulk-extend-back
Open

ajasmin wants to merge 2 commits into
andylokandy:masterfrom
ajasmin:bulk-extend-back

Conversation

@ajasmin

@ajasmin ajasmin commented Oct 1, 2026

Copy link
Copy Markdown

The previous implementation of extend_back() pushed one element at a time, each push recomputing the wrapped index, checking for a full deque and storing len. This computes the free capacity up front as its (at most two) contiguous regions and fills them with plain slice loops, counting in a local that an AddLenOnDrop guard adds to len on exit, so a panicking iterator keeps the elements already written and the loop body is one unconditional store per element, which LLVM vectorizes for primitive Copy elements from slice iterators.

The storage cast behind as_uninit_slice_mut() moves into a function of the xs field, so extend_back() can borrow the storage and len apart without unsafe code of its own.

Also added: extend_back_from_slice() for T: Copy, one block copy per free region, for element types the optimizer doesn't automatically vectorize (odd sizes, padded structs) and for older compilers. The copy goes through a private copy of std's write_copy_of_slice() (stable since 1.93).

The Wrapping deque gets neither change. Its extend_back() overwrites from the front once full, so filling the free regions is only its first phase, and its Extend impl saturates rather than evicts (#32); the behavior's contract should be settled before touching the one or adding the other.

Should it be wanted, extend_back() could also delegate to extend_back_from_slice() when its iterator is exactly slice.iter().copied(), through a lifetime-erased TypeId check that in some testing was optimised away entirely once monomorphized. It requires reinterpreting the iterator's layout and guarding against that layout changing, which gets complicated and brittle for one call shape, so I left it out.

MSRV

The first commit bumps the MSRV to 1.61 because the existing code was found to be incompatible with an older rustc: the const fns in an impl with a B: Behavior bound need 1.61, so the crate has not built on 1.59 since 3d9a1fd. The README, crate docs and CI matrix move to 1.61, and rust-version is added to Cargo.toml so cargo reports the requirement instead of failing mid-build. Everything here is tested on 1.61.0 and 1.99.0.

CI seems to be disabled on this repo at the time I'm writing this.

Generated code

Measured with arraydeque_codegen.py, which instantiates both methods per element type, compiles with --emit=asm, and classifies each body's copy loops. The script and the tables below were produced with the help of an LLM; the extracted assembly behind every cell is in the gist. before is master 2ccf305, after is this branch, CAP = 4096, generic CPU models.

ARMv9-A (aarch64-unknown-linux-gnu, -C target-feature=+v9a)

element before extend_back() after extend_back(), rustc 1.61 after extend_back(), rustc 1.99 after extend_back_from_slice()
u8 scalar SIMD ×2, SVE SIMD ×2, 32 B/iter memcpy ×2
u16, u32, u64, f32 scalar SIMD ×1 + scalar ×1, SVE SIMD ×2, 32 B/iter memcpy ×2
[u8; 3], padded {u8, u32}, {u64; 3} scalar scalar ×2 scalar ×2 memcpy ×2

x86-64-v3 (x86_64-unknown-linux-gnu, -C target-cpu=x86-64-v3)

element before extend_back() after extend_back(), rustc 1.61 after extend_back(), rustc 1.99 after extend_back_from_slice()
u8 scalar SIMD ×2, 128 B/iter SIMD ×2, 128 B/iter memcpy ×2
u16, u32, u64, f32 scalar SIMD ×1 + scalar ×1, 128 B/iter SIMD ×2, 128 B/iter memcpy ×2
[u8; 3], padded {u8, u32}, {u64; 3} scalar scalar ×2 scalar ×2 memcpy ×2

Before this change, the loop recomputed the wrapped index, checked for a full deque and stored len on every element. The scalar loops after it are one unconditional element store per iteration: rustc lowers those element types to more than one store per element, or to a vector-typed one, and neither LLVM's vectorizer nor its memcpy idiom combines such stores, even in a counted loop. That is what extend_back_from_slice() is for. With rustc 1.61 (LLVM 14), only the first region loop is vectorized for elements wider than a byte; the loop filling the region at the start of storage stays per element. rustc 1.99 vectorizes both.


While an LLM was used to help implement this, I can vouch for every line.

`const fn new` and `const fn capacity` sit in an impl with a
`B: Behavior` bound, which const fns accept only since 1.61
(`const_fn_trait_bound`), so the crate has not built on 1.59 since
3d9a1fd. Update the README, the crate docs and the CI matrix to 1.61,
and add `rust-version` to Cargo.toml so cargo reports the requirement
instead of failing mid-build.

While an LLM was used to help implement this, I can vouch for every line.
The previous implementation of `extend_back()` pushed one element at a
time, each push recomputing the wrapped index, checking for a full
deque and storing `len`. Compute the free capacity up front as its (at
most two) contiguous regions and fill them with plain slice loops,
counting in a local that an `AddLenOnDrop` guard adds to `len` on exit,
so a panicking iterator keeps the elements already written and the loop
body is one unconditional store per element, which LLVM vectorizes for
primitive `Copy` elements from slice iterators.

The storage cast behind `as_uninit_slice_mut()` moves into a function of
the `xs` field, so `extend_back()` can borrow the storage and `len` apart
without unsafe code of its own.

Also add `extend_back_from_slice()` for `T: Copy`: one block copy per
free region, for element types the optimizer doesn't automatically
vectorize (odd sizes, padded structs) and for older compilers. The
copy goes through a private copy of std's `write_copy_of_slice()`
(stable since 1.93).

The `Wrapping` deque gets neither change. Its `extend_back()` overwrites
from the front once full, so filling the free regions is only its first
phase, and its `Extend` impl saturates rather than evicts (andylokandy#32); the
behavior's contract should be settled before touching the one or adding
the other.

While an LLM was used to help implement this, I can vouch for every line.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant