Conversation
`const fn new` and `const fn capacity` sit in an impl with a `B: Behavior` bound, which const fns accept only since 1.61 (`const_fn_trait_bound`), so the crate has not built on 1.59 since 3d9a1fd. Update the README, the crate docs and the CI matrix to 1.61, and add `rust-version` to Cargo.toml so cargo reports the requirement instead of failing mid-build. While an LLM was used to help implement this, I can vouch for every line.
The previous implementation of `extend_back()` pushed one element at a time, each push recomputing the wrapped index, checking for a full deque and storing `len`. Compute the free capacity up front as its (at most two) contiguous regions and fill them with plain slice loops, counting in a local that an `AddLenOnDrop` guard adds to `len` on exit, so a panicking iterator keeps the elements already written and the loop body is one unconditional store per element, which LLVM vectorizes for primitive `Copy` elements from slice iterators. The storage cast behind `as_uninit_slice_mut()` moves into a function of the `xs` field, so `extend_back()` can borrow the storage and `len` apart without unsafe code of its own. Also add `extend_back_from_slice()` for `T: Copy`: one block copy per free region, for element types the optimizer doesn't automatically vectorize (odd sizes, padded structs) and for older compilers. The copy goes through a private copy of std's `write_copy_of_slice()` (stable since 1.93). The `Wrapping` deque gets neither change. Its `extend_back()` overwrites from the front once full, so filling the free regions is only its first phase, and its `Extend` impl saturates rather than evicts (andylokandy#32); the behavior's contract should be settled before touching the one or adding the other. While an LLM was used to help implement this, I can vouch for every line.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The previous implementation of
extend_back()pushed one element at a time, each push recomputing the wrapped index, checking for a full deque and storinglen. This computes the free capacity up front as its (at most two) contiguous regions and fills them with plain slice loops, counting in a local that anAddLenOnDropguard adds tolenon exit, so a panicking iterator keeps the elements already written and the loop body is one unconditional store per element, which LLVM vectorizes for primitiveCopyelements from slice iterators.The storage cast behind
as_uninit_slice_mut()moves into a function of thexsfield, soextend_back()can borrow the storage andlenapart without unsafe code of its own.Also added:
extend_back_from_slice()forT: Copy, one block copy per free region, for element types the optimizer doesn't automatically vectorize (odd sizes, padded structs) and for older compilers. The copy goes through a private copy of std'swrite_copy_of_slice()(stable since 1.93).The
Wrappingdeque gets neither change. Itsextend_back()overwrites from the front once full, so filling the free regions is only its first phase, and itsExtendimpl saturates rather than evicts (#32); the behavior's contract should be settled before touching the one or adding the other.Should it be wanted,
extend_back()could also delegate toextend_back_from_slice()when its iterator is exactlyslice.iter().copied(), through a lifetime-erasedTypeIdcheck that in some testing was optimised away entirely once monomorphized. It requires reinterpreting the iterator's layout and guarding against that layout changing, which gets complicated and brittle for one call shape, so I left it out.MSRV
The first commit bumps the MSRV to 1.61 because the existing code was found to be incompatible with an older rustc: the
const fns in an impl with aB: Behaviorbound need 1.61, so the crate has not built on 1.59 since 3d9a1fd. The README, crate docs and CI matrix move to 1.61, andrust-versionis added toCargo.tomlso cargo reports the requirement instead of failing mid-build. Everything here is tested on 1.61.0 and 1.99.0.CI seems to be disabled on this repo at the time I'm writing this.
Generated code
Measured with
arraydeque_codegen.py, which instantiates both methods per element type, compiles with--emit=asm, and classifies each body's copy loops. The script and the tables below were produced with the help of an LLM; the extracted assembly behind every cell is in the gist.beforeis master2ccf305,afteris this branch,CAP = 4096, generic CPU models.ARMv9-A (
aarch64-unknown-linux-gnu,-C target-feature=+v9a)extend_back()extend_back(), rustc 1.61extend_back(), rustc 1.99extend_back_from_slice()u8u16,u32,u64,f32[u8; 3], padded{u8, u32},{u64; 3}x86-64-v3 (
x86_64-unknown-linux-gnu,-C target-cpu=x86-64-v3)extend_back()extend_back(), rustc 1.61extend_back(), rustc 1.99extend_back_from_slice()u8u16,u32,u64,f32[u8; 3], padded{u8, u32},{u64; 3}Before this change, the loop recomputed the wrapped index, checked for a full deque and stored
lenon every element. The scalar loops after it are one unconditional element store per iteration: rustc lowers those element types to more than one store per element, or to a vector-typed one, and neither LLVM's vectorizer nor its memcpy idiom combines such stores, even in a counted loop. That is whatextend_back_from_slice()is for. With rustc 1.61 (LLVM 14), only the first region loop is vectorized for elements wider than a byte; the loop filling the region at the start of storage stays per element. rustc 1.99 vectorizes both.While an LLM was used to help implement this, I can vouch for every line.