diff --git a/.claude/agents/amx-savant.md b/.claude/agents/amx-savant.md index 4cb1edd3..af344e34 100644 --- a/.claude/agents/amx-savant.md +++ b/.claude/agents/amx-savant.md @@ -33,6 +33,28 @@ run win — then you update the docs in the same change. ("amx-tile")` are NIGHTLY (rust-lang/rust#126622) — you use inline `asm!` with raw `.byte` encodings. `LDTILECFG` is the one mnemonic the assembler accepts. + > ⊘ Superseded twice. (1) Since 1.98.1 (LLVM 22.1.8) the integrated + > assembler accepts every AMX mnemonic inside `asm!` with no target feature, + > and `asm_const` makes the tile index a generic: see `hpc/amx_ops.rs`. The + > `.byte` tables in `amx_matmul.rs` stay as the EMR-validated reference. + > (2) Toolchain is 1.99.0 since 2026-10-10 (floor `rust-version` 1.98.1). +- **1.99.0: `u128` is a legal `xmm_reg` operand on x86_64** (1.98.1: E0658, + "type `u128` cannot be used with this register class in stable"). Measured: + `asm!("vpxor {o}, {a}, {b}", a = in(xmm_reg) a_u128, ..)` assembles and the + value arrives as ONE xmm register. Use it to hand a 128-bit row/register to + a hand-written kernel without a two-GPR split or a stack round-trip. + Boundary caveat, also measured: RETURNING the `u128` through the Rust ABI + moves it to `rax:rdx` (`vmovq` + `vpextrq`), so keep 128-bit values inside + one `#[inline]` kernel. Do not pass a 16-byte value by value through a + function boundary and expect a vector register: under x86-64 SysV a 16-byte + integer aggregate goes in two GPRs (`extern "C"` not separately measured). +- **Gating gaps found by the 2026-10-10 unsafe inventory** + (`.claude/knowledge/unsafe-inventory/`), none fixed yet: + `TDPBF16PS` never checks the AMX-BF16 CPUID bit (gate is TILE + INT8); + `simd_runtime/cpu_ops.rs:74` AMX rung installs kernels needing `avx512f` + + `avx512vnni` without checking them; `int8_gemm_amx_tiled` is a safe `pub fn` + whose `amx_available()` and m/n/k multiple-of-16/64 checks are + `debug_assert!` only. - This host: Emerald Rapids (CPUID model 0xCF), kernel 6.18.5, AMX enabled. - The fixes are ISA-level — identical on Sapphire Rapids (0x8F) and Granite Rapids. Do NOT branch kernel correctness on CPU generation. @@ -106,8 +128,9 @@ Each fix exposes the next signature (SIGSEGV→SIGSEGV→SIGILL→wrong→correc Per `.claude/rules/agent-cargo-hygiene.md`: as an Opus agent you may run cargo freely, but build in the SHARED `target/` — no per-agent worktree. Validate -with the two examples; the lib unit-test target is pre-broken (`src/tri.rs` -type-inference errors, unrelated to AMX), so the examples are the gate. +with the two examples. (The lib unit-test target used to be broken by +`src/tri.rs`; as of 2026-10-10 `cargo test --lib` passes 2558 on v4, so run +it too. Every compile with `CARGO_PROFILE_DEV_DEBUG=0`.) ## When you finish diff --git a/.claude/agents/sentinel-qa.md b/.claude/agents/sentinel-qa.md index d1dc737b..1fe68f6f 100644 --- a/.claude/agents/sentinel-qa.md +++ b/.claude/agents/sentinel-qa.md @@ -12,9 +12,47 @@ model: opus You are SENTINEL_QA for Project NDARRAY Expansion, operating in Extreme Rigor Mode. ## Environment -- Rust 1.94 Stable +- Rust 1.99.0 stable (floor `rust-version` 1.98.1). Every compile with + `CARGO_PROFILE_DEV_DEBUG=0 CARGO_INCREMENTAL=0`, `env -u RUSTFLAGS`, tier + pinned by `--config .cargo/config-v3.toml` / `config-v4.toml`. - Target: `adaworldapi/ndarray` +## Start from the inventory, not from a blank grep +`.claude/knowledge/unsafe-inventory/` has every code-level `unsafe` (1,301 on +2026-10-10) with a verdict (`IRREDUCIBLE`, `NEEDS-TF`, `CONSOLIDATE`, +`SAFE-API`, `UNSAFE-FN-API`, `SAFE-CRATE`, `REMOVABLE`), workaround, SAFETY +status and soundness flags. Audit a change against the rows it touches, and +update the rows in the same change. The verdicts are reading-level: re-read a +site before relying on one. + +## Measured facts that decide verdicts (re-check after a toolchain bump) +- **Value intrinsics are NOT safe to call from a plain fn** on x86_64 or + aarch64, even baseline SSE2/NEON, even with `-Ctarget-cpu`/`-Ctarget-feature` + for the whole crate. Only a `#[target_feature]` caller makes them safe, and a + plain fn calling a safe `#[target_feature]` fn is unsafe again. wasm32 + `simd128` is the one exception. Source: `tools/safe_intrinsic_probe`, + identical on 1.98.1 and 1.99.0. So "remove the `unsafe` around this + intrinsic" is not a valid remediation on those arches; `unused_unsafe` will + tell you the truth, so compile before claiming it. +- **`is_x86_feature_detected!` / `is_aarch64_feature_detected!` return `true` + without asking the CPU when the feature is compiled in** (`cfg!(..) || + runtime`). A runtime check written with them proves nothing on a build that + enables the feature. `src/cpu_guard.rs` reads CPUID/XCR0 for that reason. +- `debug_assert!` is not a bounds check. A safe `pub fn` whose only guard is a + `debug_assert!` is a BLOCK (confirmed instances: `GridBlockMut::row_mut`, + `int8_gemm_amx_tiled`). So is a start-only slice index before a full-width + vector load/store (`dot_i8`, `sgemm_blocked`, `dgemm_blocked`). +- A runtime CPU check must cover EVERY feature the callee enables, not only the + headline one (`simd_int_ops.rs:315/320`, `simd_runtime/cpu_ops.rs:74`). + +## Scope rules +- `src/simd_nightly/*` is unsafe by design (validation backend over + `core::simd`, nightly only). No SAFETY-comment requirement there; an inline + note is fine where it helps a reader. +- MKL / OpenBLAS (`backend/mkl.rs`, `backend/openblas.rs`) are a LAB + COMPARISON only; the production GEMM is the native Rust BlasGraph + reimplementation. Their FFI findings are real but low priority. + ## Trigger Conditions You are invoked when any of these appear: - `unsafe` blocks written or modified diff --git a/.claude/blackboard.md b/.claude/blackboard.md index 4aad00ba..0dbfb623 100644 --- a/.claude/blackboard.md +++ b/.claude/blackboard.md @@ -3435,3 +3435,198 @@ Loose ends: the general strided path still gathers scalar (correct — at row strides ≥ a cache line a hardware gather buys nothing, per the doc); a `stride_bytes == 8` twin for `u64` lanes does not exist yet because no caller compares `u64` lanes. + +## 2026-10-10 — `cpu_guard`: build-vs-CPU check before `main` (SIGILL → message) + +New `src/cpu_guard.rs` (`std`, x86_64): compares `cfg!(target_feature)` against +CPUID + XCR0 read directly, from an `.init_array` / `__mod_init_func` / +`.CRT$XCU` hook, and exits 132 with the missing feature list. API: +`check_build_cpu()`, `assert_build_cpu()`, `missing_build_features()`. +Build-time behaviour unchanged: cross-building any tier on any runner works. + +Finding worth keeping: `is_x86_feature_detected!` returns `true` without asking +the CPU when the feature is compiled in (`cfg!(..) || runtime`). The first +version used it and passed its own check under qemu, then SIGILLed in `main`. + +Measured (examples/cpu_guard_probe.rs, qemu-user-static 8.2): +- v4 build, `-cpu Haswell|Skylake-Server|Icelake-Server|max` → message, exit 132. +- v3 build, `-cpu Haswell` → runs, exit 0 (silence twin). +- v3 build, `-cpu Nehalem` (no AVX) → still SIGILL, inside the guard: the + guard is VEX-encoded like the rest of a v3 build. Documented as not covered. +Gates: clippy `-D warnings` v3 / v4 / aarch64 clean; v4 `--lib` 2558 passed. +Loose ends: aarch64 not covered (`is_aarch64_feature_detected!` has the same +short-circuit; needs `getauxval(AT_HWCAP)`). The ctor is linked even when the +consumer references nothing in `cpu_guard` (checked with `nm`), but that was +measured for one example binary only. + +## 2026-10-10 — Rust 1.98.1 → 1.99.0 (channel only; `rust-version` floor stays 1.98.1) + +Edited: `rust-toolchain.toml` (+ bump log), CI clippy matrix and the three +`dtolnay/rust-toolchain@` steps, both Dockerfiles, CLAUDE.md, README(-DE). +Kept: `rust-version = "1.98.1"` and CI `MSRV`/`BLAS_MSRV` (no 1.99-only API +used). The toolchain file's "move TOGETHER" rule is superseded for this bump, +noted in its bump log. Measurement records that say "on 1.98.1" were left alone. + +1.99 delta fixed (all in test crates): `missing_safety_doc` ×4 on mock +`pub unsafe extern "C" fn cblas_*gemm` (blas-mock-tests), `manual_contains` ×3 +(blas-mock-tests/tests/use-blas.rs), deprecated `std::f64::NAN` ×4 through a +shadowing `use std::f64;` (tests/numeric.rs). + +Gates on 1.99.0, each with its tier: + +| gate | result | +|---|---| +| clippy `--workspace --all-targets -D warnings`, v3 and v4 (minus blas-tests, cesium) | exit 0 | +| CI rows: clippy `--features approx,serde,rayon` / `--features native` | exit 0 | +| `test --lib` v3 (`avx512f=false`) | 2507 passed, 31 ignored | +| `test --lib` v4 | 2558 passed, 32 ignored | +| `test --doc` | 684 + 4 passed | +| masking-parity native (host, avx512f=true) / native v3 (avx512f=false) | PASS / PASS | +| masking-parity wasm / wasm-scalar / neon-qemu / nightly | PASS ×4 | +| codegen-witness avx512 (v4, 6 vpternlog) / avx2 (v3) | PASS / PASS | +| floor: `cargo +1.98.1 check --workspace --exclude blas-tests --all-targets` | exit 0 | + +Pre-existing, NOT 1.99, unchanged: `blas-tests` needs a BLAS backend feature +("Missing backend"); `cesium` lib tests fail clippy identically on 1.98.1 +(constant assertions, no-effect / always-zero ops). The 1.99 hard error +`no_mangle_generic_items` and the `#[repr(simd)]`-on-macro change: nothing hit, +all targets compiled. + +Note for amx-savant, deliberately not acted on: 1.99 stabilizes passing 128-bit +integers through vector registers in x86 `asm!`; candidate for the +byte-encoded inline-asm paths in `hpc/amx_ops.rs`. + +## 2026-10-10 — `unsafe` inventory (1,301 sites) — `.claude/knowledge/unsafe-inventory/` + +903 fork + 398 upstream sites, each with a verdict and workaround (`sites.tsv`), +from a regex pre-pass, a native clippy run, and ten reading-level reviews. The +orchestrator corrected one class with the compiler: `safe_intrinsic_probe` +re-run on 1.99.0 is identical to 1.98.1 (E0133 for value intrinsics in plain fns +on x86_64 and aarch64, even baseline SSE2/NEON; only wasm32 simd128 is safe), +so 212 value-intrinsic blocks rated REMOVABLE were relabelled NEEDS-TF. +Four safe-code out-of-bounds paths were re-read and confirmed: `simd_avx2::dot_i8`, +`GridBlockMut::row_mut` (debug_assert only), `sgemm_blocked`/`dgemm_blocked` +(start-only slice check before 16-lane stores), `int8_gemm_amx_tiled` +(AMX + alignment gates debug_assert only). +Loose ends: none of the 160 flags is fixed yet; 860 sites lack `// SAFETY:`; +sentinel-qa has not audited the verdicts (reading-level, not compiled for +NEON/wasm/feature-gated files). + +## 2026-10-10 — cpu_guard checked against LLVM `getHostCPUFeatures` + +Every CPUID leaf/register/bit in `cpu_guard` matches +`llvm/lib/TargetParser/Host.cpp` (main, fetched 2026-10-10). Two gating +differences fixed to mirror LLVM: (1) on Apple targets AVX-512 OS state is +assumed (Darwin saves it lazily; XCR0 does not show it before first use), so +the guard no longer refuses a correct AVX-512 build on an AVX-512 Mac; +(2) `xsave`/`xsaveopt`/`xsavec`/`xsaves` are gated on OS AVX state as LLVM does. +Decision (operator): pre-AVX CPUs (15+ years) are out of scope, so the +"v3 build on Nehalem still SIGILLs inside the guard" case is not a gap. + +## 2026-10-10 — agent cards expanded; inventory scope rules + +`sentinel-qa` card: environment 1.99/debug-0, starts from the inventory, the +measured facts (value intrinsics need a `#[target_feature]` caller; the +`is_*_feature_detected!` short-circuit; `debug_assert!` is not a bounds check; +runtime checks must cover every callee feature), scope rules (nightly unsafe by +design; MKL/OpenBLAS lab-only). `amx-savant` card: the stale "1.94, only +LDTILECFG mnemonic" claim superseded (1.98.1 accepts all AMX mnemonics, +`amx_ops.rs`); 1.99 `u128`-in-`xmm_reg` asm operand verified (E0658 on 1.98.1), +plus the measured Rust-ABI `u128` return in `rax:rdx`; AMX gating gaps; lib +tests no longer pre-broken. +Inventory: 10 nightly rows → `NIGHTLY-BY-DESIGN`, 32 MKL/OpenBLAS rows tagged +`[lab-only]`. Floor-gated candidates recorded: `Vec::into_parts` in +`OwnedRepr::from` (1.99-only, probed). From the other session's feedback, the +per-population ABI header and Register128-through-extern-"C" items belong to +lance-graph, not ndarray, and were not applied here. + +## 2026-10-10 — Workstreams B + C: `I8x16`/`U8x16` parity, `cmp_gt` scoped + +B: `I8x16::{zero, add, sub, min, max}` added to the AVX-512 polyfill, scalar and +nightly arms (wasm and NEON already had them); `U8x16` with NEON's surface +(`LANES, splat, zero, from_slice, from_array, to_array, copy_to_slice, add, sub, +min, max`) added to AVX-512 polyfill, scalar, wasm and nightly, and exported +with `u8x16` from all six `simd.rs` blocks (it was not exported anywhere before, +NEON included). `add`/`sub` WRAP on every arm, matching `vaddq_s8`. New parity +group `check_i8x16_u8x16_lanes` (codes 0xF00-0xF32) in simd-masking-parity: +PASS on native-v4, native-v3, wasm, wasm-scalar, neon-qemu, nightly. +C (masking-ops-cartographer verdict): register-level `cmp_gt(self, other)` / +`cmp*_mask` exist on EVERY arm at native widths; G7 governs slice-level +predicates, not register methods, so nothing is replicated. Doc paragraph on +NEON `I8x16`/`I16x8::cmp_gt` and wasm `I8x16::cmp_gt`; G7 scope note in +`masking-ops-state.md`. NEON `cmp_gt` transmutes replaced with `vst1q_u8/u16` +into a local array + `// SAFETY:`. +Gates: clippy -D warnings v3, v4, aarch64; nightly check; lib tests v3 2507 / +v4 2558; doctests 684+4. (wasm32 lib clippy fails on `getrandom` in default +deps, pre-existing; the parity arm builds wasm for real.) +SoA 128-bit codegen check (probe over committed HEAD, 1.99, `-Ctarget-cpu`): +no `vpextrq` anywhere. `soa_u64x8_xor_popcnt` via `to_array()`: v3 34 `vmovq` +(8-byte loads + transpose: v3 `U64x8` is a flat `[u64; 8]` polyfill), v4 0. +Through the typed `popcnt()` + `reduce_sum()`: v3 17, v4 1. `U8x64` and the +`I8x16` polyfill: 0. Real AVX2 integer backends would remove the v3 cost. + +## 2026-10-10 — Workstream D: SAFETY-comment truth sweep (comment-only) + +Every comment that cited a pinned tier as if it were a compile-time guarantee +now states the caller obligation, in the form of `U64x8::avx2_halves`: +26 `SAFETY: AVX2 baseline` sites + 4 longer ones in `simd_avx2.rs`; the +`avx2_halves` statement itself (it claimed `.cargo/config.toml` pins v3 — the +default is `target-cpu=native`); `I8x16::saturating_abs` in `simd_avx512.rs` +(SSSE3 is NOT in the x86_64 baseline this file compiles for); the AVX2-arm +selection comment in `simd.rs`; `Fingerprint::as_u8x64` docs (endianness never +depended on a pin). NEON "baseline" comments were left: NEON really is baseline +on aarch64. Diff is comment lines only (checked); fmt + clippy v3/v4 clean. +Pending operator decision, NOT landed: `saturating_abs` cfg(ssse3) + scalar +fallback, which removes the SIGILL path on baseline builds. +Addendum (same day, from sentinel-qa's review of the saturating_abs proposal): +`I8x32::saturating_abs` (simd_avx512.rs) claimed a `#[target_feature(enable = +"avx2")]` annotation on its callers that does not exist; rewritten to the +caller-obligation form. Not a unique hole: every `I8x32` method uses +`_mm256_*`. sentinel-qa verdict on the saturating_abs cfg(ssse3) patch: +CONDITIONAL — sound and both paths agree on all 256 inputs, but land it only +with an exhaustive 256-value test and a standing CI line for the +`-Ctarget-cpu=x86-64` baseline build (otherwise the scalar branch is never +compiled in CI). Patch held for operator approval. +PR #348 codex P1 (cpu_guard table coverage): CONFIRMED and fixed. 11 x86_64 +features some rustc 1.99 `-Ctarget-cpu` model enables were absent from the +guard (sse4a tbm kl widekl sha512 sm3 sm4 avxifma avxvnniint8 avxneconvert +avxvnniint16); bit positions + gating copied from LLVM Host.cpp. A coverage +test pins rustc's full CPU-model feature list (regenerate on toolchain bump). +AMX is not a stable cfg target_feature, so -Ctarget-cpu never enables it. +Codex P1 #2 (guard should require AVX2 on every x86_64 build, because +simd_avx2 is always selected without AVX-512): path is real, but the fix +changes the crate minimum ISA for every consumer -> raised to the operator. + +## Optional to-do: pre-AVX guard (operator, 2026-10-10 — postponed, not scheduled) + +A SIGILL guard for CPUs older than ~15 years, i.e. without AVX (Nehalem and +earlier, also AVX-less Atom/Pentium/Celeron parts), running an AVX-or-later +build. Today `guard_before_main` itself is VEX-encoded on such builds and faults +inside the guard (measured: `x86-64-v3` under `qemu -cpu Nehalem`). + +Open design point, to settle before implementing: the check has to run as +non-VEX code inside a crate compiled with `-Ctarget-cpu=v3/v4/native`, and Rust +can only ADD target features per function, never remove them. The likely shape +is a `global_asm!` routine (plain CPUID + XGETBV + `write`/`exit`, legacy SSE2 +encodings only) that `.init_array` (or `__mod_init_func` / `.CRT$XCU`) points at, +ahead of the Rust guard. Falsifier: the same `x86-64-v3` build under +`qemu -cpu Nehalem` must print the message and exit 132, not SIGILL; and the +existing Haswell runs must stay silent. + +Separate from the still-open AVX2 question on PR #348 (Codex P1: a non-AVX2 +build reaching `simd_avx2` on an AVX-only CPU), which is awaiting the +operator's choice. + +## AVX2 startup floor in cpu_guard (operator decision, 2026-10-10, PR #348 codex P1) + +Operator chose the runtime-check option, as a STARTUP check only: no +`is_x86_feature_detected!` anywhere in the SIMD code, because that would break +compile-time dispatch. Implemented as `cpu_guard::BACKEND_FLOOR` (one row, AVX2 +CPUID bit + YMM OS state, `compiled = true` on every x86_64 build), unioned into +`missing_build_features`. The message says no rebuild can help. + +Measured, baseline build (CI's RUSTFLAGS, no target-cpu), `cpu_guard_probe`: +qemu `-cpu Nehalem` 132, `SandyBridge` 132, `IvyBridge` 132, `Haswell` 0, +`max` 0. Side effect: a baseline build on a PRE-AVX CPU now gets the message +too (the guard is not VEX-encoded there). The optional pre-AVX to-do above +remains only for builds compiled with `target-cpu` v3/v4/native. diff --git a/.claude/knowledge/masking-ops-state.md b/.claude/knowledge/masking-ops-state.md index a6dd9f2e..f5731b7d 100644 --- a/.claude/knowledge/masking-ops-state.md +++ b/.claude/knowledge/masking-ops-state.md @@ -200,6 +200,14 @@ unary-with-constant, and that narrowing is exactly the register in which the constant, record G7 as *deliberately absent* **in the IR's docs**, so a future session does not "fix" it. +> **G7 scope note (2026-10-10, masking-ops-cartographer).** Register methods +> `cmp_gt(self, other)` / `cmp{eq,gt}_mask(self, other)` exist on every arm +> (`simd_avx512.rs` `I8x64::cmp_gt`, `simd_scalar.rs`, `simd_neon.rs` `I8x16`/`I16x8::cmp_gt`, +> `simd_wasm.rs` `I8x16::cmp_gt`, `simd_nightly/i8_types.rs`) and are the mechanism the +> constant predicates use (`simd_masking_ops.rs`: `U8x64::from_array(lanes).cmpgt_mask(threshold_v)`). +> They are NOT G7 and NOT a gap. A missing `I8x16::cmp_gt` name on the AVX-512, scalar or +> nightly arm is not to be "completed" without a caller. + **G8 — `lzcnt_bswap_u64_to_u8` (tree-depth column) — NAMED 2026-09-17, not built.** The V3 facet's tail (bytes 8..16 = tiers 2–5, one aligned `u64` at a compile-time offset — a PEEK, not a gather) is a **tree path**. State the diff --git a/.claude/knowledge/unsafe-inventory/README.md b/.claude/knowledge/unsafe-inventory/README.md new file mode 100644 index 00000000..cdf7f010 --- /dev/null +++ b/.claude/knowledge/unsafe-inventory/README.md @@ -0,0 +1,146 @@ +# `unsafe` inventory — 2026-10-10 (master `6704865` + `cpu_guard`, rustc 1.99.0) + +READ BY: sentinel-qa, savant-architect, simd-savant, anyone removing or adding `unsafe`. + +Every code-level `unsafe` in `src/` (comment lines excluded): **1,301 sites**. +**903 were added by this fork, 398 are inherited from upstream ndarray** (the +file exists in `rust-ndarray/ndarray` master `bd3ade9`). Kinds: 904 blocks, +292 `unsafe fn`, 93 `unsafe impl`, 10 `unsafe trait`, 2 `unsafe extern`. + +Per-site data: [`sites.tsv`](sites.tsv), one row per site: +`file line origin kind scope verdict safety flag workaround ops snippet`. + +## How it was made, and how much to trust it + +1. A regex pre-pass recorded each site's kind and the operations visible in its + body (pointer loads, intrinsics, FFI, `asm!`, transmute, …). +2. A real `cargo clippy --lib` on the native (Sapphire Rapids) build with + `undocumented_unsafe_blocks` + `multiple_unsafe_ops_per_block`. +3. Ten reviewers each read one slice of the code and gave every row a verdict. + **These are reading-level verdicts, not compiled ones.** Spot-checks below. +4. The orchestrator corrected one verdict class with the compiler (next section) + and re-read the four most serious flags at source (all four held). + +`scope=test` is crude (everything after a file's first `#[cfg(test)]`); the NEON +reviewer found rows after `simd_neon.rs:2753` mis-tagged as test. + +## The compiler fact that decides most of the SIMD rows + +`tools/safe_intrinsic_probe`, re-run on **rustc 1.99.0**: + +| probe | result | +|---|---| +| plain fn, `_mm_and_si128` (SSE2, x86_64 baseline) | E0133 | +| plain fn, AVX2 intrinsic, `-Ctarget-cpu=x86-64-v3` | E0133 | +| plain fn, AVX-512 intrinsic, `-Ctarget-cpu=x86-64-v4` | E0133 | +| plain fn, AVX-512 intrinsic, `-Ctarget-feature=+avx512f` | E0133 | +| plain fn calling a SAFE `#[target_feature]` fn (x86 and aarch64) | E0133 | +| plain fn, NEON intrinsic, aarch64 (NEON is baseline) | E0133 | +| plain fn, `v128_and`, wasm32 `+simd128` | **OK** | + +Also confirmed in-tree: removing the `unsafe` around a lone `_mm512_srlv_epi32` +(`simd_avx512.rs:1705`) fails with E0133 under both v3 and v4. + +So on x86_64 and aarch64 a value intrinsic is safe to call ONLY inside a function +that itself carries `#[target_feature(enable = ..)]`. Crate-wide target features +do not count, and calling such a function from a plain one is unsafe again. Two +reviewers had marked 212 value-intrinsic blocks `REMOVABLE` on the 1.87 rule; +they are re-labelled `NEEDS-TF` here. **There is no per-site removal for them on +1.99.** The minimum is the current shape: one `unsafe` inside each polyfill +wrapper, so consumers of `ndarray::simd` write none. Only wasm32 is different. + +## Verdicts + +| verdict | fork | upstream | meaning, and the workaround | +|---|---|---|---| +| `IRREDUCIBLE` | 279 | 259 | FFI (MKL/OpenBLAS), `asm!` / AMX tile config, `unsafe impl Send/Sync`, ndarray's raw-pointer core, `#[target_feature]` calls after a runtime check. Keep; document. | +| `NEEDS-TF` | 212 | 0 | Value intrinsics. See above: no per-site fix on 1.99. | +| `CONSOLIDATE` | 185 | 1 | Genuinely unsafe pointer loads/stores (`_mm*_loadu/storeu`, `vld1q/vst1q`). Route each type's slice methods through its existing `from_array`/`to_array` (fixed-size `&[T; N]`), leaving ONE unsafe load and store per type. Several sites bypass helpers that already exist (e.g. `simd_neon.rs:2431`, `:3182`). | +| `SAFE-API` | 107 | 32 | A named safe std API replaces it: `as_chunks`, `<[T;N]>::try_from`, `to_bits`/`from_ne_bytes`, `copy_from_slice`, `NonNull::from_ref`, `Vec::into_flattened`, checked `from_shape_vec`, `HashMap::get_disjoint_mut` (`hpc/blackboard.rs`, 5 sites), a `dyn Any` downcast for same-`'static`-type transmutes. | +| `UNSAFE-FN-API` | 41 | 105 | `unsafe fn` whose contract is the point (`uget`, `from_shape_ptr`, …). | +| `SAFE-CRATE` | 27 | 1 | Only `bytemuck` / `zerocopy` make it safe. Needs operator approval for the dependency. | +| `NIGHTLY-BY-DESIGN` | 10 | 0 | `src/simd_nightly/*`: validation backend over `core::simd`, nightly only, unsafe by design. No SAFETY-comment requirement; inline notes where they help (operator, 2026-10-10). | +| `REMOVABLE` | 42 | 0 | `unsafe fn` keywords on bodies with no unsafe op: scalar fallbacks and forwarders in `simd_runtime/*`, `bgz17_bridge.rs` (10), `aabb.rs`, `byte_scan.rs`, `bitwise.rs`, … Caveat: where the fn is `#[target_feature]`, the keyword can go but its CALLERS still need `unsafe` (probe row 5); dropping the ATTRIBUTE changes codegen on v3 builds, so measure first. | + +## `// SAFETY:` comments (CLAUDE.md hard rule) + +856 sites lack one (526 fork, 330 upstream; nightly rows are exempt). The clippy run counts 439 blocks and +77 impls on the native x86 build alone (NEON / wasm / nightly / feature-gated +files are not compiled there). Several macros (`simd_avx512.rs:50`, `:61`) cover +many expansions with one comment. + +## Soundness leads (160 flags, 63 of them `feature:`) + +A flag is a lead, not a verdict. The first four were re-read at source and hold. + +**Confirmed: out-of-bounds reachable from safe code** +- `simd_avx2.rs:406` `dot_i8`: loop sized by `a.len()`, `b[base..]` checks only + the start, `_mm256_loadu_si256` reads 32 bytes → read past `b`'s end when + `b.len() < a.len()`. Also no AVX2 check in a safe `pub fn`. +- `hpc/blocked_grid/iter.rs:306` `GridBlockMut::row_mut`: both bounds checks are + `debug_assert!`; its SAFETY comment cites "the bounds check above". +- `backend/kernels_avx512.rs:665/:694` `sgemm_blocked` (and `dgemm_blocked` + `:783/:861`, `:1006`): `&mut c[ir*ldc..]` is start-checked only, then a full + 16-lane store; the stated safety contract covers AVX-512F only, not `c`'s length. +- `hpc/int8_tile_gemm.rs:358` `int8_gemm_amx_tiled`: lengths are asserted, but + `amx_available()` and the m/n/k multiple-of-16/64 checks are `debug_assert!`; + the reviewer reports OOB tile access at `:487`, `:538` when misaligned. + +**Reported, not yet re-read** +- `backend/native.rs:219` passes raw pointers to `matrixmultiply` with no + extent check against m/n/k/ld*. +- Lab-only (32 rows tagged `[lab-only]`): MKL / OpenBLAS are a comparison + harness, not the production path, which is the native Rust BlasGraph GEMM. + Their safe wrappers pass raw pointers unchecked and truncate dims with + `as c_int` (`backend/mkl.rs:199/223/244/263`, `openblas.rs:96/120`); stride-0 + row views give `ld < cols` → `xerbla` (`mkl.rs:407/452/504/553`). Real, low + priority. +- `simd_neon.rs:115/147/214` codebook gathers: start-only checks, output length + `debug_assert!` only. `simd_wasm.rs:1416/1477`: same pattern. +- `simd_avx512.rs:3477/3700`: private fns rely on callers' asserts. +- `hpc/gguf.rs:229`: `Vec` pointer cast to 2-aligned BF16; + `jitson/scan_config.rs:138` `from_raw_parts::` on arbitrary byte pointers. +- `hpc/blackboard.rs:295`, `blocked_grid/iter.rs:171`: raw-pointer `&mut` + re-borrows invalid under Stacked Borrows. +- JIT kernels (`jitson_cranelift/noise_jit.rs:79`, `scan_jit.rs:49`) are not + lifetime-tied to the engine that owns their code memory. +- Upstream: `linalg/impl_linalg.rs:316` `set_len` on uninitialised elements + (`Array::uninit` + `assume_init` at `:368` avoids it); + `impl_methods.rs:3301` reads with only a `size_of` check; + `iterators/mod.rs:1459/1484` trusts `TrustedIterator` exactness (upstream FIXME). +- Correctness, not UB: `simd_amx.rs:367` AVX-512 VNNI dot drops the `n % 64` + tail, so `matvec_dispatch` differs from the vnni2/scalar arms; + `simd_neon.rs:90` `vabdq_s16` wraps above 32767; `simd_neon.rs:3126/3138` + shift counts use only the low byte. + +**ISA-gating (`feature:`)** +- AVX-512BW intrinsics (`I8x64`, `I16x32`, …) live behind an `avx512f`-only gate + (`simd.rs`). Practical exposure: Knights Landing-class only. +- `simd_avx2` is compiled on every x86_64 build and its safe `pub fn`s use AVX2 + with no check. A pre-AVX2 x86_64 build reaching them SIGILLs. `cpu_guard` + covers mismatched BUILDS, not this (the code is compiled into a baseline build). +- `simd_runtime/cpu_ops.rs:74`: the AMX rung checks only `amx_int8` + the OS + grant but installs kernels needing `avx512f` + `avx512vnni`; + `simd_int_ops.rs:315/320` check only the VNNI bit. AMX-BF16 (`TDPBF16PS`) + never checks its own CPUID bit. + +## Toolchain-gated candidates (need the floor raised to 1.99) + +- `Vec::into_parts` / `Vec::from_parts` / `Box::into_non_null` are stable on + 1.99.0 and unstable on 1.98.1 (E0658 `box_vec_non_null`, probed). In + `data_repr.rs`, `OwnedRepr::from` becomes `let (ptr, len, capacity) = + v.into_parts();` with no `ManuallyDrop`/`nonnull_from_vec_data`; + `take_as_vec` stays `unsafe` (`from_parts` is unsafe by contract). +- `u128` as an `xmm_reg` `asm!` operand (1.99.0; E0658 on 1.98.1, probed): + for the hand-written asm kernels, see the `amx-savant` card. + +## Suggested order + +1. Fix the confirmed safe-code OOB paths (assert lengths, or make the fns + `unsafe fn` with a real `# Safety`). Small and local. +2. Complete the runtime-check feature sets (`cpu_ops.rs:74`, `simd_int_ops.rs`, + AMX-BF16). +3. `CONSOLIDATE` through `from_array`/`to_array`: shrinks the pointer-op count + per type to two. +4. SAFETY comments, macro-first. +5. `SAFE-API` swaps; `SAFE-CRATE` only with approval. diff --git a/.claude/knowledge/unsafe-inventory/sites.tsv b/.claude/knowledge/unsafe-inventory/sites.tsv new file mode 100644 index 00000000..8047175f --- /dev/null +++ b/.claude/knowledge/unsafe-inventory/sites.tsv @@ -0,0 +1,1302 @@ +file line origin kind scope verdict safety flag workaround ops snippet +src/aabb.rs 150 FORK block src SAFE-API present - Body is F32x16 polyfill + safe indexing; drop #[target_feature]/unsafe on aabb_intersect_batch_avx512 and call it safely target_feature_call unsafe { +src/aabb.rs 156 FORK block src SAFE-API present - Callee is a plain scalar loop; delete aabb_intersect_batch_sse41 and fall to aabb_intersect_batch_scalar (iterator map) target_feature_call unsafe { +src/aabb.rs 179 FORK fn src REMOVABLE present - No unsafe op in body (F32x16 + indexing); make safe fn, or drop the attribute (polyfill is cfg-selected at compile time) target_feature_call unsafe fn aabb_intersect_batch_avx512(query: &Aabb, candidates: &[Aabb]) -> Vec { +src/aabb.rs 247 FORK fn src REMOVABLE missing - Pure scalar body; delete fn, scalar variant is equivalent (comment says LLVM autovectorizes) target_feature_call unsafe fn aabb_intersect_batch_sse41(query: &Aabb, candidates: &[Aabb]) -> Vec { +src/aabb.rs 291 FORK block src SAFE-API present - Same as 150: ray_aabb_slab_test_avx512 body is polyfill+safe code; remove target_feature/unsafe target_feature_call unsafe { +src/aabb.rs 334 FORK fn src REMOVABLE present - No unsafe op in body; make safe fn or drop attribute target_feature_call unsafe fn ray_aabb_slab_test_avx512(ray: &Ray, aabbs: &[Aabb]) -> (Vec, Vec) { +src/aabb.rs 446 FORK block src SAFE-API present - aabb_expand_batch_sse2 is a scalar loop identical to _scalar; delete it and call aabb_expand_batch_scalar target_feature_call unsafe { +src/aabb.rs 469 FORK fn src REMOVABLE missing - Pure scalar body; delete fn (duplicate of aabb_expand_batch_scalar) target_feature_call unsafe fn aabb_expand_batch_sse2(aabbs: &mut [Aabb], dx: f32, dy: f32, dz: f32) { +src/arraytraits.rs 58 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides ptr_arith unsafe { +src/arraytraits.rs 79 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides ptr_arith unsafe { +src/arraytraits.rs 467 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl Send/Sync on raw-pointer type; sound by inherited-mutability argument - unsafe impl Sync for ArrayBase +src/arraytraits.rs 475 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl Send/Sync on raw-pointer type; sound by inherited-mutability argument - unsafe impl Send for ArrayBase +src/arraytraits.rs 482 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl Send/Sync on raw-pointer type; sound by inherited-mutability argument - unsafe impl Sync for ArrayRef where A: Sync {} +src/arraytraits.rs 484 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl Send/Sync on raw-pointer type; sound by inherited-mutability argument - unsafe impl Send for ArrayRef where A: Send {} +src/arraytraits.rs 557 UPSTREAM block src SAFE-API missing - ArrayViewMut::from_shape(xs.len(), xs) (checked, O(1)) instead of from_shape_ptr ptr_arith unsafe { Self::from_shape_ptr(xs.len(), xs.as_mut_ptr()) } +src/arraytraits.rs 592 UPSTREAM block src SAFE-API missing - <[[A;N]]>::as_flattened_mut() + ArrayViewMut::from_shape(dim, flat); drops from_raw_parts_mut and ptr ptr_arith,from_raw_parts unsafe { +src/arraytraits.rs 618 UPSTREAM block src IRREDUCIBLE present - from_data_ptr/with_strides_dim: moving same allocation into Arc repr; core design - unsafe { ArrayBase::from_data_ptr(data, arr.parts.ptr).with_strides_dim(arr.parts.strides, arr.parts.dim) } +src/backend/kernels_avx512.rs 615 FORK fn src CONSOLIDATE present bounds: only slice START indexed (c[ir*ldc..]) then 16-lane store; sgemm_blocked never checks c.len()>=(m-1)*ldc+n -> OOB write on short c Raw loadu/storeu/masked via slice.as_ptr(); use F32x16::from_slice/copy_to_slice (assert len, used in dot_f32 here) + len asserts - unsafe fn sgemm_ukernel_6x16( +src/backend/kernels_avx512.rs 694 FORK block src SAFE-API present bounds: safe pub fn sgemm_blocked has no assert on c/ldc extent; SAFETY says "buffers sized correctly" but nothing checks If ukernel becomes a safe #[target_feature(avx512f)] fn (caller has same feature, Rust>=1.86) this call needs no unsafe - unsafe { +src/backend/kernels_avx512.rs 783 FORK fn src CONSOLIDATE present bounds: full-width _mm512_storeu_pd after start-only slice check; short c => OOB write Same as sgemm_ukernel_6x16: F64x8::from_slice/copy_to_slice + length asserts; masked tail via scalar loop or one masked-store helper - unsafe fn dgemm_ukernel_6x8( +src/backend/kernels_avx512.rs 861 FORK block src SAFE-API missing bounds: no c-extent assert in safe pub fn dgemm_blocked; no SAFETY comment Same as 694: make dgemm_ukernel_6x8 safe #[target_feature] fn so the call needs no unsafe - unsafe { +src/backend/kernels_avx512.rs 899 FORK block src CONSOLIDATE present - Two _mm512_loadu_si512 are the only unsafe ops; add one load512(&[u8])->__m512i helper (asserts len>=64) or use U64x8::popcnt polyfill x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/backend/kernels_avx512.rs 927 FORK block src CONSOLIDATE present - One loadu is the only unsafe op; same load512 helper as hamming_distance x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/backend/kernels_avx512.rs 1006 FORK fn src UNSAFE-FN-API missing bounds: c[r*ldc..] start-only check then 16-lane store Test-only R x 16 ukernel; raw ptr loads/stores after slice-start check; make safe by asserting slice extents (or F32x16 helpers) - unsafe fn ukernel_rows( +src/backend/kernels_avx512.rs 1063 FORK block src SAFE-API present - Test-local; once ukernels are safe #[target_feature] fns (caller has avx512f) the match needs no unsafe - unsafe { +src/backend/kernels_avx512.rs 1140 FORK block src IRREDUCIBLE present - Runtime-checked (assert is_x86_feature_detected avx512f) call of #[target_feature] fns from a feature-less test fn - unsafe { +src/backend/kernels_avx512.rs 1159 FORK block src IRREDUCIBLE present - Same: target_feature call after runtime avx512f assert (could be avoided by making the test closure body a target_feature fn) - unsafe { sgemm_blocked(m, n, k, 1.0, black_box(&a), k, black_box(&b), n, black_box(&mut c1), n) } +src/backend/kernels_avx512.rs 1164 FORK block src IRREDUCIBLE present - Same: target_feature call after runtime avx512f assert - unsafe { sgemm_blocked_desc(m, n, k, 1.0, black_box(&a), k, black_box(&b), n, black_box(&mut c2), n) } +src/backend/mkl.rs 153 FORK block src IRREDUCIBLE missing trunc: usize->c_int wraps for len>=2^31 (no OOB, wrong result) [lab-only] FFI cblas_sdot; n=min(len) so reads are in-bounds, inc=1 ptr_arith,ffi unsafe { cblas_sdot(n, x.as_ptr(), 1, y.as_ptr(), 1) } +src/backend/mkl.rs 158 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_ddot; n=min(len), inc=1 ptr_arith,ffi unsafe { cblas_ddot(n, x.as_ptr(), 1, y.as_ptr(), 1) } +src/backend/mkl.rs 163 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_saxpy; n=min(x,y len) ptr_arith,ffi unsafe { cblas_saxpy(n, alpha, x.as_ptr(), 1, y.as_mut_ptr(), 1) } +src/backend/mkl.rs 168 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_daxpy; n=min(x,y len) ptr_arith,ffi unsafe { cblas_daxpy(n, alpha, x.as_ptr(), 1, y.as_mut_ptr(), 1) } +src/backend/mkl.rs 172 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_sscal; n=x.len(), inc=1 ptr_arith,ffi unsafe { cblas_sscal(x.len() as c_int, alpha, x.as_mut_ptr(), 1) } +src/backend/mkl.rs 176 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_dscal; n=x.len(), inc=1 ptr_arith,ffi unsafe { cblas_dscal(x.len() as c_int, alpha, x.as_mut_ptr(), 1) } +src/backend/mkl.rs 180 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_snrm2; n=x.len(), inc=1 ptr_arith,ffi unsafe { cblas_snrm2(x.len() as c_int, x.as_ptr(), 1) } +src/backend/mkl.rs 184 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_dnrm2; n=x.len(), inc=1 ptr_arith,ffi unsafe { cblas_dnrm2(x.len() as c_int, x.as_ptr(), 1) } +src/backend/mkl.rs 188 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_sasum; n=x.len(), inc=1 ptr_arith,ffi unsafe { cblas_sasum(x.len() as c_int, x.as_ptr(), 1) } +src/backend/mkl.rs 192 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_dasum; n=x.len(), inc=1 ptr_arith,ffi unsafe { cblas_dasum(x.len() as c_int, x.as_ptr(), 1) } +src/backend/mkl.rs 199 FORK block src IRREDUCIBLE missing bounds: safe pub fn, no check a/b/c lengths vs m,n,k,lda,ldb,ldc -> MKL reads/writes OOB; also dims `as c_int` truncate [lab-only] FFI cblas_sgemm. Add checked_mul extent asserts (a>=(m-1)*lda+k, etc.) and i32::try_from dims so the safe fn is sound ptr_arith,ffi unsafe { +src/backend/mkl.rs 223 FORK block src IRREDUCIBLE missing bounds: safe pub fn, no slice-extent check vs m,n,k,ld*; dims `as c_int` truncate [lab-only] FFI cblas_dgemm; same missing extent asserts and try_from as gemm_f32 ptr_arith,ffi unsafe { +src/backend/mkl.rs 244 FORK block src IRREDUCIBLE missing bounds: safe pub fn, no x/y/a length checks (x.len()>=n, y.len()>=m, lda>=n unchecked) [lab-only] FFI cblas_sgemv; assert a>=(m-1)*lda+n, x.len()>=n, y.len()>=m ptr_arith,ffi unsafe { +src/backend/mkl.rs 263 FORK block src IRREDUCIBLE missing bounds: safe pub fn, no x/y/a length checks [lab-only] FFI cblas_dgemv; same asserts needed ptr_arith,ffi unsafe { +src/backend/mkl.rs 407 FORK block src IRREDUCIBLE missing ld: blas_layout gives ld=rs.max(1) for rows<=1; stride 0 (insert_axis) => ld CBLAS xerbla (may abort process) [lab-only] FFI cblas_sgemm; shapes+strides validated (blas_layout), extents implied by ArrayView invariants; ld edge case below ptr_arith,ffi unsafe { +src/backend/mkl.rs 452 FORK block src IRREDUCIBLE missing ld: ld=rs.max(1) < cols for single-row view with stride 0/1 => xerbla [lab-only] FFI cblas_dgemm via ArrayView2 (same validation as sgemm) ptr_arith,ffi unsafe { +src/backend/mkl.rs 504 FORK block src IRREDUCIBLE present ld: same rows<=1 ld=1 < cols edge case => xerbla [lab-only] FFI cblas_gemm_bf16bf16f32; BF16 is repr(transparent) u16 so cast ok (comment present above block) ptr_arith,ffi unsafe { +src/backend/mkl.rs 553 FORK block src IRREDUCIBLE missing ld: same rows<=1 ld edge case => xerbla [lab-only] FFI cblas_gemm_s8s8s32 with &co offset pointer (single i32, FixOffset) ok ptr_arith,ffi unsafe { +src/backend/mod.rs 222 FORK block src CONSOLIDATE present - One shared fn u16_as_bf16(&[u16])->&[BF16] (BF16 is repr(transparent)) replaces 4 from_raw_parts sites; or bytemuck::cast_slice ptr_arith,from_raw_parts let a_bf16: &[crate::hpc::quantized::BF16] = unsafe { +src/backend/mod.rs 229 FORK block src CONSOLIDATE present - Same u16_as_bf16 helper (SAFETY comment sits before the block, clippy happy) ptr_arith,from_raw_parts unsafe { core::slice::from_raw_parts(b.as_ptr() as *const crate::hpc::quantized::BF16, b.len()) }; +src/backend/mod.rs 252 FORK block src CONSOLIDATE present - Same helper; non-x86 branch of gemm_bf16 (not compiled on host) ptr_arith,from_raw_parts unsafe { core::slice::from_raw_parts(a.as_ptr() as *const crate::hpc::quantized::BF16, a.len()) }; +src/backend/mod.rs 255 FORK block src CONSOLIDATE present - Same helper; non-x86 branch (not compiled on host) ptr_arith,from_raw_parts unsafe { core::slice::from_raw_parts(b.as_ptr() as *const crate::hpc::quantized::BF16, b.len()) }; +src/backend/native.rs 104 FORK block src IRREDUCIBLE present - Runtime tier()==Avx512 check then call #[target_feature(avx512f)] kernel via dispatch! macro - Tier::Avx512 => unsafe { super::kernels_avx512::$name($($arg),*) }, +src/backend/native.rs 123 FORK block src IRREDUCIBLE present - Same dispatch! arm (no-return-type form) - Tier::Avx512 => unsafe { super::kernels_avx512::$name($($arg),*) }, +src/backend/native.rs 142 FORK block src IRREDUCIBLE present - Same, custom-path arm; tier verified avx512f only (callee must need only avx512f) - Tier::Avx512 => unsafe { $a512($($arg),*) }, +src/backend/native.rs 159 FORK block src IRREDUCIBLE present - Same, custom-path no-return arm - Tier::Avx512 => unsafe { $a512($($arg),*) }, +src/backend/native.rs 219 FORK block src SAFE-API present bounds: safe pub fn passes raw ptrs; SAFETY claims "valid slices" but no a/b/c extent vs lda/ldb/ldc check Build ArrayView2::from_shape((m,k).strides((lda,1)),a) (validates extents) and use general_mat_mul instead of raw matrixmultiply::sgemm ptr_arith unsafe { +src/backend/native.rs 505 FORK block src IRREDUCIBLE present - Tier::Avx2 (avx2+fma) runtime-verified call of #[target_feature] fn target_feature_call unsafe { dot_f32_avx2(x, y) } +src/backend/native.rs 517 FORK block src IRREDUCIBLE present - Same (avx2+fma tier) target_feature_call unsafe { dot_f64_avx2(x, y) } +src/backend/native.rs 529 FORK block src IRREDUCIBLE present - Same (avx2+fma tier) target_feature_call unsafe { axpy_f32_avx2(alpha, x, y) } +src/backend/native.rs 541 FORK block src IRREDUCIBLE present - Same (avx2+fma tier) target_feature_call unsafe { axpy_f64_avx2(alpha, x, y) } +src/backend/native.rs 553 FORK block src IRREDUCIBLE present - Same (avx2 tier; fma also verified by detection) target_feature_call unsafe { scal_f32_avx2(alpha, x) } +src/backend/native.rs 564 FORK block src IRREDUCIBLE present - Same target_feature_call unsafe { scal_f64_avx2(alpha, x) } +src/backend/native.rs 575 FORK block src IRREDUCIBLE present - Same target_feature_call unsafe { nrm2_f32_avx2(x) } +src/backend/native.rs 586 FORK block src IRREDUCIBLE present - Same target_feature_call unsafe { nrm2_f64_avx2(x) } +src/backend/native.rs 597 FORK block src IRREDUCIBLE present - Same target_feature_call unsafe { asum_f32_avx2(x) } +src/backend/native.rs 608 FORK block src IRREDUCIBLE present - Same target_feature_call unsafe { asum_f64_avx2(x) } +src/backend/native.rs 620 FORK fn src CONSOLIDATE missing - Hand AVX2 loadu via x.as_ptr().add(i), in-bounds (i+16<=min len); use F32x8::from_slice (asserting) instead x86_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn dot_f32_avx2(x: &[f32], y: &[f32]) -> f32 { +src/backend/native.rs 659 FORK fn src CONSOLIDATE missing - Same with f64: F64x4 from_slice polyfill; loop bounds correct x86_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn dot_f64_avx2(x: &[f64], y: &[f64]) -> f64 { +src/backend/native.rs 696 FORK fn src CONSOLIDATE missing - axpy f32 loadu/storeu in-bounds; F32x8::from_slice + copy_to_slice x86_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn axpy_f32_avx2(alpha: f32, x: &[f32], y: &mut [f32]) { +src/backend/native.rs 716 FORK fn src CONSOLIDATE missing - axpy f64 same; F64x4 helpers x86_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn axpy_f64_avx2(alpha: f64, x: &[f64], y: &mut [f64]) { +src/backend/native.rs 738 FORK fn src CONSOLIDATE missing - scal f32 in-bounds; F32x8 helpers x86_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn scal_f32_avx2(alpha: f32, x: &mut [f32]) { +src/backend/native.rs 756 FORK fn src CONSOLIDATE missing - scal f64 in-bounds; F64x4 helpers x86_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn scal_f64_avx2(alpha: f64, x: &mut [f64]) { +src/backend/native.rs 782 FORK fn src CONSOLIDATE missing - nrm2 f32 in-bounds; F32x8 helpers x86_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn nrm2_f32_avx2(x: &[f32]) -> f32 { +src/backend/native.rs 818 FORK fn src CONSOLIDATE missing - nrm2 f64 in-bounds; F64x4 helpers x86_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn nrm2_f64_avx2(x: &[f64]) -> f64 { +src/backend/native.rs 858 FORK fn src CONSOLIDATE missing - asum f32 in-bounds; F32x8 helpers (abs via & mask or .abs()) x86_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn asum_f32_avx2(x: &[f32]) -> f32 { +src/backend/native.rs 896 FORK fn src CONSOLIDATE missing - asum f64 in-bounds; F64x4 helpers x86_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn asum_f64_avx2(x: &[f64]) -> f64 { +src/backend/openblas.rs 48 FORK block src IRREDUCIBLE present trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_sdot; n=min(len), inc=1 ptr_arith,ffi unsafe { cblas_sdot(n, x.as_ptr(), 1, y.as_ptr(), 1) } +src/backend/openblas.rs 53 FORK block src IRREDUCIBLE present trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_ddot; n=min(len) ptr_arith,ffi unsafe { cblas_ddot(n, x.as_ptr(), 1, y.as_ptr(), 1) } +src/backend/openblas.rs 58 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_saxpy; n=min(len) ptr_arith,ffi unsafe { cblas_saxpy(n, alpha, x.as_ptr(), 1, y.as_mut_ptr(), 1) } +src/backend/openblas.rs 63 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_daxpy; n=min(len) ptr_arith,ffi unsafe { cblas_daxpy(n, alpha, x.as_ptr(), 1, y.as_mut_ptr(), 1) } +src/backend/openblas.rs 67 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_sscal; n=x.len() ptr_arith,ffi unsafe { cblas_sscal(x.len() as c_int, alpha, x.as_mut_ptr(), 1) } +src/backend/openblas.rs 71 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_dscal; n=x.len() ptr_arith,ffi unsafe { cblas_dscal(x.len() as c_int, alpha, x.as_mut_ptr(), 1) } +src/backend/openblas.rs 75 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_snrm2 ptr_arith,ffi unsafe { cblas_snrm2(x.len() as c_int, x.as_ptr(), 1) } +src/backend/openblas.rs 79 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_dnrm2 ptr_arith,ffi unsafe { cblas_dnrm2(x.len() as c_int, x.as_ptr(), 1) } +src/backend/openblas.rs 83 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_sasum ptr_arith,ffi unsafe { cblas_sasum(x.len() as c_int, x.as_ptr(), 1) } +src/backend/openblas.rs 87 FORK block src IRREDUCIBLE missing trunc: len as c_int wraps >=2^31 (no OOB) [lab-only] FFI cblas_dasum ptr_arith,ffi unsafe { cblas_dasum(x.len() as c_int, x.as_ptr(), 1) } +src/backend/openblas.rs 96 FORK block src IRREDUCIBLE present bounds: SAFETY says "caller guarantees" on a SAFE pub fn; nothing validates a/b/c extents vs m,n,k,ld*; dims `as c_int` truncate [lab-only] FFI cblas_sgemm; add extent asserts + i32::try_from so the safe fn is sound ptr_arith,ffi unsafe { +src/backend/openblas.rs 120 FORK block src IRREDUCIBLE missing bounds: safe pub fn, no slice-extent check; no SAFETY comment [lab-only] FFI cblas_dgemm; same missing extent checks ptr_arith,ffi unsafe { +src/backend/openblas.rs 141 FORK block src IRREDUCIBLE missing bounds: safe pub fn, no x/y/a length checks [lab-only] FFI cblas_sgemv; assert a extent, x.len()>=n, y.len()>=m ptr_arith,ffi unsafe { +src/backend/openblas.rs 160 FORK block src IRREDUCIBLE missing bounds: safe pub fn, no x/y/a length checks [lab-only] FFI cblas_dgemv; same asserts needed ptr_arith,ffi unsafe { +src/bitwise.rs 100 FORK fn test REMOVABLE missing - No unsafe op in body (U8x64 polyfill + safe indexing); make safe fn or drop attribute target_feature_call unsafe fn hamming_avx512bw(a: &[u8], b: &[u8]) -> u64 { +src/bitwise.rs 149 FORK fn test REMOVABLE missing - Same: popcount_avx512bw body has no unsafe op target_feature_call unsafe fn popcount_avx512bw(a: &[u8]) -> u64 { +src/bitwise.rs 276 FORK block test IRREDUCIBLE present - Runtime has_avx512_bw_popcnt() then call kernel using VPOPCNTDQ intrinsics (_mm512_popcnt_epi64) - return unsafe { crate::backend::kernels_avx512::hamming_distance(a, b) }; +src/bitwise.rs 280 FORK block test SAFE-API present - Callee hamming_avx512bw has no unsafe op; make it a plain safe fn (U8x64 polyfill) and drop this unsafe target_feature_call return unsafe { hamming_avx512bw(a, b) }; +src/bitwise.rs 292 FORK block test IRREDUCIBLE present - Runtime avx512vpopcntdq check then call kernel using VPOPCNTDQ intrinsic - return unsafe { crate::backend::kernels_avx512::popcount(a) }; +src/bitwise.rs 296 FORK block test SAFE-API present - Callee popcount_avx512bw has no unsafe op; plain safe fn target_feature_call return unsafe { popcount_avx512bw(a) }; +src/bitwise.rs 308 FORK block test IRREDUCIBLE present - Runtime has_avx512_bw_popcnt() then call hamming_batch (target_feature, calls hamming_distance) - return unsafe { crate::backend::kernels_avx512::hamming_batch(query, database, num_rows, row_bytes) }; +src/bitwise.rs 564 FORK block test SAFE-API missing - Test call of hamming_avx512bw; becomes safe once callee is a plain fn target_feature_call let got = unsafe { hamming_avx512bw(&a, &b) }; +src/bitwise.rs 579 FORK block test SAFE-API missing - Test call of popcount_avx512bw; same target_feature_call let got = unsafe { popcount_avx512bw(&a) }; +src/bitwise.rs 595 FORK block test IRREDUCIBLE missing - Test call of VPOPCNTDQ kernel after is_x86_feature_detected - let got = unsafe { crate::backend::kernels_avx512::hamming_distance(&a, &b) }; +src/bitwise.rs 610 FORK block test IRREDUCIBLE missing - Test call of VPOPCNTDQ popcount after runtime detect - let got = unsafe { crate::backend::kernels_avx512::popcount(&a) }; +src/bitwise.rs 629 FORK block test SAFE-API missing - Test call of hamming_avx512bw; safe once callee is plain target_feature_call let bw = unsafe { hamming_avx512bw(&a, &b) }; +src/bitwise.rs 633 FORK block test IRREDUCIBLE missing - Test call of VPOPCNTDQ hamming_distance after runtime detect - let vpc = unsafe { crate::backend::kernels_avx512::hamming_distance(&a, &b) }; +src/bitwise.rs 650 FORK block test SAFE-API missing - Test call of popcount_avx512bw target_feature_call let bw = unsafe { popcount_avx512bw(&a) }; +src/bitwise.rs 654 FORK block test IRREDUCIBLE missing - Test call of VPOPCNTDQ popcount after runtime detect - let vpc = unsafe { crate::backend::kernels_avx512::popcount(&a) }; +src/bitwise.rs 674 FORK block test SAFE-API missing - Test call of hamming_avx512bw target_feature_call assert_eq!(unsafe { hamming_avx512bw(&a, &b) }, expected, "avx512bw large"); +src/bitwise.rs 678 FORK block test IRREDUCIBLE missing - Test call of VPOPCNTDQ hamming_distance after runtime detect - unsafe { crate::backend::kernels_avx512::hamming_distance(&a, &b) }, +src/bitwise.rs 801 FORK block test SAFE-API missing - Test call of hamming_avx512bw target_feature_call assert_eq!(unsafe { hamming_avx512bw(&a, &b) }, 0, "bw identical n={}", n); +src/bitwise.rs 805 FORK block test IRREDUCIBLE missing - Test call of VPOPCNTDQ hamming_distance after runtime detect - unsafe { crate::backend::kernels_avx512::hamming_distance(&a, &b) }, +src/byte_scan.rs 23 FORK fn src REMOVABLE present - Body is scalar loops; delete fn (same as scalar fallback) or make safe target_feature_call pub(crate) unsafe fn byte_find_all_avx2(haystack: &[u8], needle: u8) -> Vec { +src/byte_scan.rs 53 FORK fn src REMOVABLE present - Body is U8x64 polyfill + safe indexing; no unsafe op target_feature_call pub(crate) unsafe fn byte_find_all_avx512(haystack: &[u8], needle: u8) -> Vec { +src/byte_scan.rs 87 FORK fn src REMOVABLE present - Scalar body; delete (scalar fallback equivalent) target_feature_call pub(crate) unsafe fn byte_count_avx2(haystack: &[u8], needle: u8) -> usize { +src/byte_scan.rs 116 FORK fn src REMOVABLE present - U8x64 polyfill body; no unsafe op target_feature_call pub(crate) unsafe fn byte_count_avx512(haystack: &[u8], needle: u8) -> usize { +src/byte_scan.rs 149 FORK block src SAFE-API present - Callee byte_find_all_avx512 has no unsafe op; remove target_feature/unsafe and call directly target_feature_call return unsafe { simd_impl::byte_find_all_avx512(haystack, needle) }; +src/byte_scan.rs 153 FORK block src SAFE-API present - byte_find_all_avx2 is a scalar loop; use the iterator fallback (haystack.iter().position/filter_map) target_feature_call return unsafe { simd_impl::byte_find_all_avx2(haystack, needle) }; +src/byte_scan.rs 188 FORK block src SAFE-API present - Callee byte_count_avx512 has no unsafe op; call it as plain fn target_feature_call return unsafe { simd_impl::byte_count_avx512(haystack, needle) }; +src/byte_scan.rs 192 FORK block src SAFE-API present - byte_count_avx2 is scalar; use haystack.iter().filter(..).count() target_feature_call return unsafe { simd_impl::byte_count_avx2(haystack, needle) }; +src/data_repr.rs 44 UPSTREAM block src IRREDUCIBLE missing - OwnedRepr is a hand-rolled Vec (ptr,len,cap); raw ptr core design ptr_arith,from_raw_parts unsafe { slice::from_raw_parts(self.ptr.as_ptr(), self.len) } +src/data_repr.rs 61 UPSTREAM block src IRREDUCIBLE missing - OwnedRepr is a hand-rolled Vec (ptr,len,cap); raw ptr core design ptr_arith unsafe { self.ptr.add(self.len) } +src/data_repr.rs 83 UPSTREAM fn src UNSAFE-FN-API present - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design uninit pub(crate) unsafe fn set_len(&mut self, new_len: usize) { +src/data_repr.rs 101 UPSTREAM fn src UNSAFE-FN-API present - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design - pub(crate) unsafe fn data_subst(self) -> OwnedRepr { +src/data_repr.rs 122 UPSTREAM block src IRREDUCIBLE missing - OwnedRepr is a hand-rolled Vec (ptr,len,cap); raw ptr core design ptr_arith,from_raw_parts unsafe { Vec::from_raw_parts(self.ptr.as_ptr(), len, capacity) } +src/data_repr.rs 169 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl Send/Sync on raw-pointer type; sound by inherited-mutability argument - unsafe impl Sync for OwnedRepr where A: Sync {} +src/data_repr.rs 170 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl Send/Sync on raw-pointer type; sound by inherited-mutability argument - unsafe impl Send for OwnedRepr where A: Send {} +src/data_traits.rs 38 UPSTREAM trait src IRREDUCIBLE missing - unsafe trait: sealed (private_decl) internal contract of ndarray storage; not implementable downstream (upstream design) ptr_arith pub unsafe trait RawData: Sized { +src/data_traits.rs 54 UPSTREAM trait src IRREDUCIBLE missing - unsafe trait: sealed (private_decl) internal contract of ndarray storage; not implementable downstream (upstream design) - pub unsafe trait RawDataMut: RawData { +src/data_traits.rs 82 UPSTREAM trait src IRREDUCIBLE missing - unsafe trait: sealed (private_decl) internal contract of ndarray storage; not implementable downstream (upstream design) - pub unsafe trait RawDataClone: RawData { +src/data_traits.rs 85 UPSTREAM fn src UNSAFE-FN-API missing - trait method: ptr must lie in self data; could debug_assert via _is_pointer_inbounds, not cheaply checkable generally - unsafe fn clone_with_ptr(&self, ptr: NonNull) -> (Self, NonNull); +src/data_traits.rs 88 UPSTREAM fn src UNSAFE-FN-API missing - trait method: ptr must lie in self data; could debug_assert via _is_pointer_inbounds, not cheaply checkable generally - unsafe fn clone_from_with_ptr(&mut self, other: &Self, ptr: NonNull) -> NonNull { +src/data_traits.rs 101 UPSTREAM trait src IRREDUCIBLE missing - unsafe trait: sealed (private_decl) internal contract of ndarray storage; not implementable downstream (upstream design) - pub unsafe trait Data: RawData { +src/data_traits.rs 144 UPSTREAM trait src IRREDUCIBLE missing - unsafe trait: sealed (private_decl) internal contract of ndarray storage; not implementable downstream (upstream design) - pub unsafe trait DataMut: Data + RawDataMut { +src/data_traits.rs 165 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) ptr_arith unsafe impl RawData for RawViewRepr<*const A> { +src/data_traits.rs 176 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) ptr_arith unsafe impl RawDataClone for RawViewRepr<*const A> { +src/data_traits.rs 177 UPSTREAM fn src UNSAFE-FN-API missing - impl of unsafe trait method clone_with_ptr/clone_from_with_ptr; ptr-offset recompute (pointer must be inside self data) - unsafe fn clone_with_ptr(&self, ptr: NonNull) -> (Self, NonNull) { +src/data_traits.rs 182 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) ptr_arith unsafe impl RawData for RawViewRepr<*mut A> { +src/data_traits.rs 193 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) ptr_arith unsafe impl RawDataMut for RawViewRepr<*mut A> { +src/data_traits.rs 208 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) ptr_arith unsafe impl RawDataClone for RawViewRepr<*mut A> { +src/data_traits.rs 209 UPSTREAM fn src UNSAFE-FN-API missing - impl of unsafe trait method clone_with_ptr/clone_from_with_ptr; ptr-offset recompute (pointer must be inside self data) - unsafe fn clone_with_ptr(&self, ptr: NonNull) -> (Self, NonNull) { +src/data_traits.rs 214 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) ptr_arith unsafe impl RawData for OwnedArcRepr { +src/data_traits.rs 225 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl RawDataMut for OwnedArcRepr +src/data_traits.rs 251 UPSTREAM block src IRREDUCIBLE missing - NonNull::offset into Arc::make_mut-ed Vec after COW; offset computed from old base (upstream core design) ptr_arith unsafe { +src/data_traits.rs 261 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl Data for OwnedArcRepr { +src/data_traits.rs 270 UPSTREAM block src IRREDUCIBLE present - from_data_ptr + with_strides_dim: unsafe ctors reusing same data; invariants unchanged (upstream core) - unsafe { +src/data_traits.rs 280 UPSTREAM block src IRREDUCIBLE present - from_data_ptr + with_strides_dim: unsafe ctors reusing same data; invariants unchanged (upstream core) - Ok(owned_data) => unsafe { +src/data_traits.rs 285 UPSTREAM block src IRREDUCIBLE present - from_data_ptr + with_strides_dim: unsafe ctors reusing same data; invariants unchanged (upstream core) - Err(arc_data) => unsafe { +src/data_traits.rs 305 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl DataMut for OwnedArcRepr where A: Clone {} +src/data_traits.rs 307 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl RawDataClone for OwnedArcRepr { +src/data_traits.rs 308 UPSTREAM fn src UNSAFE-FN-API missing - impl of unsafe trait method clone_with_ptr/clone_from_with_ptr; ptr-offset recompute (pointer must be inside self data) - unsafe fn clone_with_ptr(&self, ptr: NonNull) -> (Self, NonNull) { +src/data_traits.rs 314 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) ptr_arith unsafe impl RawData for OwnedRepr { +src/data_traits.rs 320 UPSTREAM block src SAFE-API missing - ptr.add(len) is one-past-end: use slc.as_ptr_range().end (safe) ptr_arith let end = unsafe { ptr.add(slc.len()) }; +src/data_traits.rs 327 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl RawDataMut for OwnedRepr { +src/data_traits.rs 342 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl Data for OwnedRepr { +src/data_traits.rs 361 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl DataMut for OwnedRepr {} +src/data_traits.rs 363 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl RawDataClone for OwnedRepr +src/data_traits.rs 367 UPSTREAM fn src UNSAFE-FN-API missing - impl of unsafe trait method clone_with_ptr/clone_from_with_ptr; ptr-offset recompute (pointer must be inside self data) ptr_arith unsafe fn clone_with_ptr(&self, ptr: NonNull) -> (Self, NonNull) { +src/data_traits.rs 377 UPSTREAM fn src UNSAFE-FN-API missing - impl of unsafe trait method clone_with_ptr/clone_from_with_ptr; ptr-offset recompute (pointer must be inside self data) ptr_arith unsafe fn clone_from_with_ptr(&mut self, other: &Self, ptr: NonNull) -> NonNull { +src/data_traits.rs 388 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) ptr_arith unsafe impl RawData for ViewRepr<&A> { +src/data_traits.rs 399 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl Data for ViewRepr<&A> { +src/data_traits.rs 416 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl RawDataClone for ViewRepr<&A> { +src/data_traits.rs 417 UPSTREAM fn src UNSAFE-FN-API missing - impl of unsafe trait method clone_with_ptr/clone_from_with_ptr; ptr-offset recompute (pointer must be inside self data) - unsafe fn clone_with_ptr(&self, ptr: NonNull) -> (Self, NonNull) { +src/data_traits.rs 422 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) ptr_arith unsafe impl RawData for ViewRepr<&mut A> { +src/data_traits.rs 433 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl RawDataMut for ViewRepr<&mut A> { +src/data_traits.rs 448 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl Data for ViewRepr<&mut A> { +src/data_traits.rs 465 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl DataMut for ViewRepr<&mut A> {} +src/data_traits.rs 480 UPSTREAM trait src IRREDUCIBLE missing - unsafe trait: sealed (private_decl) internal contract of ndarray storage; not implementable downstream (upstream design) uninit pub unsafe trait DataOwned: Data { +src/data_traits.rs 502 UPSTREAM trait src IRREDUCIBLE missing - unsafe trait: sealed (private_decl) internal contract of ndarray storage; not implementable downstream (upstream design) - pub unsafe trait DataShared: Clone + Data + RawDataClone {} +src/data_traits.rs 504 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl DataShared for OwnedArcRepr {} +src/data_traits.rs 505 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl DataShared for ViewRepr<&A> {} +src/data_traits.rs 507 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) uninit unsafe impl DataOwned for OwnedRepr { +src/data_traits.rs 523 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) uninit unsafe impl DataOwned for OwnedArcRepr { +src/data_traits.rs 539 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) ptr_arith unsafe impl RawData for CowRepr<'_, A> { +src/data_traits.rs 553 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl RawDataMut for CowRepr<'_, A> +src/data_traits.rs 581 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl RawDataClone for CowRepr<'_, A> +src/data_traits.rs 585 UPSTREAM fn src UNSAFE-FN-API missing - impl of unsafe trait method clone_with_ptr/clone_from_with_ptr; ptr-offset recompute (pointer must be inside self data) - unsafe fn clone_with_ptr(&self, ptr: NonNull) -> (Self, NonNull) { +src/data_traits.rs 598 UPSTREAM fn src UNSAFE-FN-API missing - impl of unsafe trait method clone_with_ptr/clone_from_with_ptr; ptr-offset recompute (pointer must be inside self data) - unsafe fn clone_from_with_ptr(&mut self, other: &Self, ptr: NonNull) -> NonNull { +src/data_traits.rs 616 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl<'a, A> Data for CowRepr<'a, A> { +src/data_traits.rs 625 UPSTREAM block src IRREDUCIBLE present - from_data_ptr + with_strides_dim: unsafe ctors reusing same data; invariants unchanged (upstream core) - CowRepr::Owned(data) => unsafe { +src/data_traits.rs 638 UPSTREAM block src IRREDUCIBLE present - from_data_ptr + with_strides_dim: unsafe ctors reusing same data; invariants unchanged (upstream core) - CowRepr::Owned(data) => unsafe { +src/data_traits.rs 647 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) - unsafe impl DataMut for CowRepr<'_, A> where A: Clone {} +src/data_traits.rs 649 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl of sealed storage trait; only the unsafe-trait contract, no unsafe op in the impl line itself (upstream) uninit unsafe impl<'a, A> DataOwned for CowRepr<'a, A> { +src/data_traits.rs 681 UPSTREAM fn src UNSAFE-FN-API present - trait contract: A,B same repr; cannot be checked (generic layout) - unsafe fn data_subst(self) -> Self::Output; +src/data_traits.rs 687 UPSTREAM fn src UNSAFE-FN-API missing - delegates to inherent OwnedRepr::data_subst (unsafe layout reinterpretation of Vec) - unsafe fn data_subst(self) -> Self::Output { +src/data_traits.rs 695 UPSTREAM fn src UNSAFE-FN-API missing layout: Arc>->OwnedRepr raw cast has no compile-time layout check (size/align equality only by caller contract) Arc::from_raw(into_raw as *const OwnedRepr) layout cast; relies on caller guarantee A,B same repr ptr_arith unsafe fn data_subst(self) -> Self::Output { +src/data_traits.rs 703 UPSTREAM fn src UNSAFE-FN-API missing - trait demands unsafe fn; body only builds a ZST marker (ViewRepr/RawViewRepr::new), no unsafe op inside - unsafe fn data_subst(self) -> Self::Output { +src/data_traits.rs 711 UPSTREAM fn src UNSAFE-FN-API missing - trait demands unsafe fn; body only builds a ZST marker (ViewRepr/RawViewRepr::new), no unsafe op inside - unsafe fn data_subst(self) -> Self::Output { +src/data_traits.rs 719 UPSTREAM fn src UNSAFE-FN-API missing - trait demands unsafe fn; body only builds a ZST marker (ViewRepr/RawViewRepr::new), no unsafe op inside - unsafe fn data_subst(self) -> Self::Output { +src/data_traits.rs 727 UPSTREAM fn src UNSAFE-FN-API missing - trait demands unsafe fn; body only builds a ZST marker (ViewRepr/RawViewRepr::new), no unsafe op inside - unsafe fn data_subst(self) -> Self::Output { +src/data_traits.rs 735 UPSTREAM fn src UNSAFE-FN-API missing - trait-required unsafe fn; delegates to variants data_subst - unsafe fn data_subst(self) -> Self::Output { +src/dimension/dynindeximpl.rs 24 UPSTREAM block src SAFE-API missing - &ar[..len as usize]; bounds check is one cmp (debug_assert already there) unchecked unsafe { ar.get_unchecked(..len as usize) } +src/dimension/dynindeximpl.rs 36 UPSTREAM block src SAFE-API missing - &mut ar[..len as usize]; same unchecked unsafe { ar.get_unchecked_mut(..len as usize) } +src/dimension/ndindex.rs 20 UPSTREAM trait src IRREDUCIBLE missing - unsafe trait contract (implementors promise in-bounds offsets / stable AsRef) - pub unsafe trait NdIndex: Debug { +src/dimension/ndindex.rs 27 UPSTREAM impl src IRREDUCIBLE missing - unsafe trait NdIndex impl: index_checked/unchecked must agree & be in-bounds; trait contract - unsafe impl NdIndex for D +src/dimension/ndindex.rs 39 UPSTREAM impl src IRREDUCIBLE missing - unsafe trait NdIndex impl: index_checked/unchecked must agree & be in-bounds; trait contract - unsafe impl NdIndex for () { +src/dimension/ndindex.rs 50 UPSTREAM impl src IRREDUCIBLE missing - unsafe trait NdIndex impl: index_checked/unchecked must agree & be in-bounds; trait contract - unsafe impl NdIndex for (Ix, Ix) { +src/dimension/ndindex.rs 60 UPSTREAM impl src IRREDUCIBLE missing - unsafe trait NdIndex impl: index_checked/unchecked must agree & be in-bounds; trait contract - unsafe impl NdIndex for (Ix, Ix, Ix) { +src/dimension/ndindex.rs 74 UPSTREAM impl src IRREDUCIBLE missing - unsafe trait NdIndex impl: index_checked/unchecked must agree & be in-bounds; trait contract - unsafe impl NdIndex for (Ix, Ix, Ix, Ix) { +src/dimension/ndindex.rs 86 UPSTREAM impl src IRREDUCIBLE missing - unsafe trait NdIndex impl: index_checked/unchecked must agree & be in-bounds; trait contract - unsafe impl NdIndex for (Ix, Ix, Ix, Ix, Ix) { +src/dimension/ndindex.rs 99 UPSTREAM impl src IRREDUCIBLE missing - unsafe trait NdIndex impl: index_checked/unchecked must agree & be in-bounds; trait contract - unsafe impl NdIndex for (Ix, Ix, Ix, Ix, Ix, Ix) { +src/dimension/ndindex.rs 112 UPSTREAM impl src IRREDUCIBLE missing - unsafe trait NdIndex impl: index_checked/unchecked must agree & be in-bounds; trait contract - unsafe impl NdIndex for Ix { +src/dimension/ndindex.rs 123 UPSTREAM impl src IRREDUCIBLE missing - unsafe trait NdIndex impl: index_checked/unchecked must agree & be in-bounds; trait contract - unsafe impl NdIndex for Ix { +src/dimension/ndindex.rs 140 UPSTREAM impl src IRREDUCIBLE missing - unsafe trait NdIndex impl: index_checked/unchecked must agree & be in-bounds; trait contract - unsafe impl NdIndex<$ix_n> for [Ix; $n] { +src/dimension/ndindex.rs 169 UPSTREAM impl src IRREDUCIBLE missing - unsafe trait NdIndex impl: index_checked/unchecked must agree & be in-bounds; trait contract - unsafe impl NdIndex for Dim<[Ix; N]> { +src/dimension/ndindex.rs 186 UPSTREAM impl src IRREDUCIBLE missing - unsafe trait NdIndex impl: index_checked/unchecked must agree & be in-bounds; trait contract - unsafe impl NdIndex for [Ix; N] { +src/dimension/ndindex.rs 209 UPSTREAM impl src IRREDUCIBLE missing - unsafe trait NdIndex impl: index_checked/unchecked must agree & be in-bounds; trait contract - unsafe impl NdIndex for &IxDyn { +src/dimension/ndindex.rs 218 UPSTREAM impl src IRREDUCIBLE missing - unsafe trait NdIndex impl: index_checked/unchecked must agree & be in-bounds; trait contract - unsafe impl NdIndex for &[Ix] { +src/distance.rs 58 FORK fn src REMOVABLE present - Body is F32x8 polyfill + safe code; no unsafe op target_feature_call pub(crate) unsafe fn squared_distances_avx2(query: [f32; 3], points: &[[f32; 3]], out: &mut Vec) { +src/distance.rs 121 FORK block src SAFE-API present - Drop target_feature/unsafe on squared_distances_avx2 (no unsafe op inside) and call it as safe fn target_feature_call unsafe { simd_impl::squared_distances_avx2(query, points, &mut out) }; +src/extension/nonnull.rs 8 UPSTREAM block src SAFE-API present - NonNull::from(v.as_mut_slice()).cast::() (same dangling ptr for empty Vec) - no unsafe ptr_arith,unchecked unsafe { NonNull::new_unchecked(v.as_mut_ptr()) } +src/extension/nonnull.rs 17 UPSTREAM fn src UNSAFE-FN-API present - caller guarantees non-null; could be safe via NonNull::new(ptr).expect() at one cmp, kept as debug-checked ptr_arith,unchecked pub(crate) unsafe fn nonnull_debug_checked_from_ptr(ptr: *mut T) -> NonNull { +src/free_functions.rs 91 UPSTREAM block src SAFE-API missing - Array::from_shape_vec((), vec![x]) with unwrap/expect (trivial check), or Array0 via from_elem-like safe ctor - unsafe { ArrayBase::from_shape_vec_unchecked((), vec![x]) } +src/free_functions.rs 110 UPSTREAM block src SAFE-API present - NonNull::from_ref(x).cast_mut() (const-stable) - references are non-null ptr_arith,unchecked unsafe { NonNull::new_unchecked(x as *const A as *mut A) }, +src/free_functions.rs 146 UPSTREAM block src SAFE-API present - NonNull::from_ref(xs).cast::() (const-stable) ptr_arith,unchecked unsafe { NonNull::new_unchecked(xs.as_ptr() as *mut A) }, +src/free_functions.rs 187 UPSTREAM block src SAFE-API present - NonNull::from_ref(xs).cast::() ptr_arith,unchecked let ptr = unsafe { NonNull::new_unchecked(xs.as_ptr() as *mut A) }; +src/free_functions.rs 269 UPSTREAM block src SAFE-API missing - Vec::into_flattened() (stable 1.80) replaces manual from_raw_parts/forget; then checked from_shape_vec ptr_arith,from_raw_parts unsafe { +src/free_functions.rs 397 UPSTREAM block src SAFE-API missing - ArrayView::from_shape(shape.strides(strides), slice) likely OK incl. zero strides (unverified); else keep as core design ptr_arith unsafe { ArrayView::new(nonnull_debug_checked_from_ptr($arr.as_ptr() as *mut A), shape, strides) } +src/hpc/amx_matmul.rs 113 FORK fn src IRREDUCIBLE present - asm! ldtilecfg; could be safe method on a token proving amx_tile_available() + validating TileConfig palette/rows/colsb raw_load_store,ptr_arith,asm pub unsafe fn tile_loadconfig(config: &TileConfig) { +src/hpc/amx_matmul.rs 126 FORK fn src IRREDUCIBLE present logic: tiles 4-7 silently no-op (_ => {}) - caller thinks accumulator zeroed asm! tilezero (byte-encoded); tmm4-7 arms missing -> use amx_ops::tilezero::() asm pub unsafe fn tile_zero(tile: u8) { +src/hpc/amx_matmul.rs 141 FORK fn src IRREDUCIBLE present - asm! tilerelease; no operands, could be safe once a tile-state token exists asm pub unsafe fn tile_release() { +src/hpc/amx_matmul.rs 169 FORK fn src IRREDUCIBLE present bounds: raw ptr+stride, no row/len check; tile>7 silently no-op asm! tileloadd; wrap as safe fn taking &[u8] + rows + stride and asserting (rows-1)*stride+colsb <= len raw_load_store,ptr_arith,asm pub unsafe fn tile_load(tile: u8, ptr: *const u8, stride: usize) { +src/hpc/amx_matmul.rs 232 FORK fn src IRREDUCIBLE present bounds: raw ptr+stride, no row/len check; tiles 4-7 silently no-op asm! tilestored; safe wrapper taking &mut [u8] with (rows-1)*stride+colsb <= len assert raw_load_store,ptr_arith,asm pub unsafe fn tile_store(tile: u8, ptr: *mut u8, stride: usize) { +src/hpc/amx_matmul.rs 288 FORK fn src IRREDUCIBLE present - asm! tdpbusd (byte-encoded); amx_ops::tdpbusd::<0,2,1>() is the mnemonic form asm pub unsafe fn tile_dpbusd() { +src/hpc/amx_matmul.rs 326 FORK fn src IRREDUCIBLE present - asm! 4x tdpbusd (byte-encoded); needs tmm0-7 configured asm pub unsafe fn tile_dpbusd_2x2() { +src/hpc/amx_matmul.rs 359 FORK fn src IRREDUCIBLE present feature: callers gate on amx_available() = TILE+INT8 only; BF16 bit never checked (SIGILL if INT8 exposed w/o BF16) asm! tdpbf16ps (byte-encoded); needs AMX-BF16 (CPUID 7.0:EDX[22]) asm pub unsafe fn tile_dpbf16ps() { +src/hpc/amx_matmul.rs 576 FORK block src CONSOLIDATE present - one audited fn bf16_as_u16(&[BF16])->&[u16] (also used at :645 twice); or bytemuck::cast_slice if BF16 derives Pod ptr_arith,from_raw_parts let a_u16: &[u16] = unsafe { core::slice::from_raw_parts(a.as_ptr() as *const u16, a.len()) }; +src/hpc/amx_matmul.rs 613 FORK block src IRREDUCIBLE present - calling #[target_feature(avx512bf16)] fn after is_x86_feature_detected! - unsafe { +src/hpc/amx_matmul.rs 645 FORK fn src UNSAFE-FN-API present - target_feature fn; body safe-indexed except 2 from_raw_parts + loadu/transmute; could be safe #[target_feature] fn (1.86+) with helpers x86_intrinsic,ptr_arith,from_raw_parts unsafe fn bf16_gemm_vdpbf16ps(a: &[BF16], b: &[BF16], c: &mut [f32], m: usize, n: usize, k: usize) { +src/hpc/amx_matmul.rs 973 FORK block src IRREDUCIBLE present - call #[target_feature(avx512vnni)] fn after runtime detect (guarded by is_x86_feature_detected) - unsafe { +src/hpc/amx_matmul.rs 983 FORK block src IRREDUCIBLE present - call #[target_feature(avxvnni)] fn after runtime detect - unsafe { +src/hpc/amx_matmul.rs 1057 FORK block test IRREDUCIBLE missing - test: raw asm! tilezero/tilerelease; use amx_ops::tilezero::<0>()/tilerelease() mnemonics instead of .byte raw_load_store,asm unsafe { +src/hpc/amx_matmul.rs 1373 FORK block test IRREDUCIBLE present - test: target_feature call; guarded by have_avx512 early return above - unsafe { crate::backend::kernels_avx512::sgemm_blocked(m, n, k, 1.0, &av, k, &bv, n, &mut c, n) } +src/hpc/amx_ops.rs 118 FORK fn src IRREDUCIBLE present - asm! ldtilecfg; contract in module doc + # Safety; safe form needs permission token + config validation ptr_arith,asm pub unsafe fn ldtilecfg(cfg: *const u8) { +src/hpc/amx_ops.rs 148 FORK fn src IRREDUCIBLE present - asm! sttilecfg; could take &mut TileConfig (align(64), 64 bytes) -> safe given tile permission token ptr_arith,asm pub unsafe fn sttilecfg(cfg: *mut u8) { +src/hpc/amx_ops.rs 175 FORK fn src IRREDUCIBLE present - asm! tilerelease asm pub unsafe fn tilerelease() { +src/hpc/amx_ops.rs 203 FORK fn src IRREDUCIBLE present - asm! tilezero; T<8 enforced by const assert; needs configured tile asm pub unsafe fn tilezero() { +src/hpc/amx_ops.rs 236 FORK fn src IRREDUCIBLE present bounds: base+stride unchecked by design (documented contract) asm! tileloadd; safe wrapper = &[u8] + rows/colsb asserts; raw base+stride unchecked ptr_arith,asm pub unsafe fn tileloadd(base: *const u8, stride: usize) { +src/hpc/amx_ops.rs 266 FORK fn src IRREDUCIBLE present bounds: base+stride unchecked by design (documented contract) asm! tileloaddt1; same wrapper shape as tileloadd ptr_arith,asm pub unsafe fn tileloaddt1(base: *const u8, stride: usize) { +src/hpc/amx_ops.rs 300 FORK fn src IRREDUCIBLE present bounds: base+stride unchecked by design (documented contract) asm! tilestored; safe wrapper = &mut [u8] with extent assert ptr_arith,asm pub unsafe fn tilestored(base: *mut u8, stride: usize) { +src/hpc/amx_ops.rs 333 FORK fn src IRREDUCIBLE present bounds: base+stride unchecked by design (documented contract) asm! tileloaddrs (AMX-MOVRS); needs movrs feature bit ptr_arith,asm pub unsafe fn tileloaddrs(base: *const u8, stride: usize) { +src/hpc/amx_ops.rs 363 FORK fn src IRREDUCIBLE present bounds: base+stride unchecked by design (documented contract) asm! tileloaddrst1 (AMX-MOVRS) ptr_arith,asm pub unsafe fn tileloaddrst1(base: *const u8, stride: usize) { +src/hpc/amx_ops.rs 407 FORK fn src IRREDUCIBLE present - macro-generated asm! tdp* ops; tile distinctness enforced at compile time; tier gate is doc-only asm pub unsafe fn $name() { +src/hpc/amx_ops.rs 514 FORK fn src IRREDUCIBLE present - asm! .byte tmmultf32ps (claimed never executed); tier gate doc-only asm pub unsafe fn tmmultf32ps() { +src/hpc/amx_ops.rs 570 FORK fn src IRREDUCIBLE present - macro-generated asm! tcvtrow* reg-row form; zmm operand asm pub unsafe fn $name(row: u32) -> $ty { +src/hpc/amx_ops.rs 606 FORK fn src IRREDUCIBLE present - macro-generated asm! tcvtrow* imm-row form; ROW<16 const-asserted asm pub unsafe fn $name_imm() -> $ty { +src/hpc/amx_ops.rs 828 FORK fn test REMOVABLE missing - test probe macro: unsafe fn() cast only; drop if wrappers become safe fns with one inner unsafe - ($w as unsafe fn(), concat!("ndarray_amx_probe_", stringify!($w))) +src/hpc/amx_ops.rs 835 FORK fn test REMOVABLE missing - test: unsafe fn() in a param type, no unsafe op; becomes fn() if wrappers are safe fns - fn contains((f, name): (unsafe fn(), &str), needle: &[u8]) -> bool { +src/hpc/amx_ops.rs 848 FORK fn test REMOVABLE missing - test: unsafe fn() in a param type, no unsafe op - fn contains_masked((f, name): (unsafe fn(), &str), needle: &[(u8, u8)]) -> bool { +src/hpc/amx_ops.rs 857 FORK fn test IRREDUCIBLE missing - test probe wrapper (bytes inspected, never executed); move unsafe into body of safe fn - unsafe fn w_tilezero0() { +src/hpc/amx_ops.rs 862 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tilezero7() { +src/hpc/amx_ops.rs 867 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tilerelease() { +src/hpc/amx_ops.rs 872 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tdpbusd_021() { +src/hpc/amx_ops.rs 877 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tdpbf16ps_021() { +src/hpc/amx_ops.rs 882 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tdpbusd_012() { +src/hpc/amx_ops.rs 887 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tdpbssd_012() { +src/hpc/amx_ops.rs 892 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tdpfp16ps_012() { +src/hpc/amx_ops.rs 897 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tcmmimfp16ps_012() { +src/hpc/amx_ops.rs 902 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tdpbf8ps_012() { +src/hpc/amx_ops.rs 907 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tdphf8ps_012() { +src/hpc/amx_ops.rs 912 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tmmultf32ps_012() { +src/hpc/amx_ops.rs 917 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tmmultf32ps_021() { +src/hpc/amx_ops.rs 922 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tdpbsud_012() { +src/hpc/amx_ops.rs 927 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tdpbuud_012() { +src/hpc/amx_ops.rs 932 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tcmmrlfp16ps_012() { +src/hpc/amx_ops.rs 937 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tdpbhf8ps_012() { +src/hpc/amx_ops.rs 942 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) - unsafe fn w_tdphbf8ps_012() { +src/hpc/amx_ops.rs 951 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed); passes &SCRATCH ptr to ldtilecfg ptr_arith unsafe fn w_ldtilecfg() { +src/hpc/amx_ops.rs 956 FORK fn test IRREDUCIBLE missing aliasing: SCRATCH.as_ptr() as *mut u8 from an immutable static handed to sttilecfg - UB if wrapper were ever called test probe wrapper (never executed) ptr_arith unsafe fn w_sttilecfg() { +src/hpc/amx_ops.rs 961 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) raw_load_store,ptr_arith unsafe fn w_tileloadd2() { +src/hpc/amx_ops.rs 966 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) raw_load_store,ptr_arith unsafe fn w_tileloaddt1_2() { +src/hpc/amx_ops.rs 971 FORK fn test IRREDUCIBLE missing aliasing: tilestored to const-cast immutable static - UB if ever executed test probe wrapper (never executed) raw_load_store,ptr_arith unsafe fn w_tilestored0() { +src/hpc/amx_ops.rs 976 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) raw_load_store,ptr_arith unsafe fn w_tileloaddrs2() { +src/hpc/amx_ops.rs 981 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed) raw_load_store,ptr_arith unsafe fn w_tileloaddrst1_2() { +src/hpc/amx_ops.rs 987 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed); only cfg avx512f - unsafe fn w_tcvtrowd2ps_imm_1_3() { +src/hpc/amx_ops.rs 993 FORK fn test IRREDUCIBLE missing - test probe wrapper (never executed); only cfg avx512f - unsafe fn w_tilemovrow_imm_1_5() { +src/hpc/bf16_tile_gemm.rs 176 FORK block src SAFE-CRATE present - bytemuck::cast_slice:: (or zerocopy AsBytes); no std safe u16->u8 view ptr_arith,from_raw_parts unsafe { core::slice::from_raw_parts(self.data.as_ptr() as *const u8, self.data.len() * 2) } +src/hpc/bf16_tile_gemm.rs 214 FORK block src IRREDUCIBLE present feature: gate is amx_available() (TILE+INT8); TDPBF16PS needs AMX-BF16 bit, unchecked (amx_ops::amx_features().bf16 exists) calling AMX tile kernel after amx_available() (CPUID+XCR0+arch_prctl) - unsafe { +src/hpc/bf16_tile_gemm.rs 219 FORK block src IRREDUCIBLE present - calling #[target_feature(avx512bf16)] fn after is_x86_feature_detected! - unsafe { +src/hpc/bf16_tile_gemm.rs 255 FORK fn src UNSAFE-FN-API present bounds: relies on callers' asserts for ptr.add reads (callers do assert; private fn) target_feature fn; lens preconditions (a=16k, b=k*16, c=256) only asserted by callers - add asserts + make safe #[target_feature] fn x86_intrinsic,raw_load_store,ptr_arith,transmute unsafe fn avx512bf16_path(a_bf16: &[u16], b_vnni: &[u16], c: &mut [f32], k: usize) { +src/hpc/bf16_tile_gemm.rs 266 FORK block src CONSOLIDATE present - 13 ops in one block: replace loadu/storeu with helpers over &[f32;16]/&[u16;32] (chunks_exact + try_from) or safe_unaligned_simd x86_intrinsic,raw_load_store,ptr_arith,transmute unsafe { +src/hpc/bf16_tile_gemm.rs 298 FORK fn src IRREDUCIBLE present feature: TDPBF16PS requires AMX-BF16; only INT8-based amx_available() checked upstream AMX tile asm kernel; tile_load/store extents depend on caller-asserted lengths (16k/k16/256) raw_load_store,ptr_arith unsafe fn amx_path(a_bf16: &[u16], b_vnni: &[u16], c: &mut [f32], k: usize) { +src/hpc/bf16_tile_gemm.rs 538 FORK block test IRREDUCIBLE present - test: target_feature call after detection - unsafe { avx512bf16_path(&a, packed.data(), &mut c, k) }; +src/hpc/bf16_tile_gemm.rs 558 FORK block test IRREDUCIBLE present - test: target_feature call after detection - unsafe { avx512bf16_path(&a_bf, packed_f.data(), &mut c_v, k) }; +src/hpc/bgz17_bridge.rs 41 FORK fn src IRREDUCIBLE missing - target_feature fns only coerce to unsafe fn pointers; alt: private safe wrapper fns selected post-detection - type L1Fn = unsafe fn(&[i16; 17], &[i16; 17]) -> u32; +src/hpc/bgz17_bridge.rs 45 FORK fn src REMOVABLE missing - body has no unsafe op (simd wrappers are safe fns); safe #[target_feature] fn (1.86+) coerces to unsafe fn ptr - unverified, not compiled as such target_feature_call unsafe fn l1_avx512(a: &[i16; 17], b: &[i16; 17]) -> u32 { +src/hpc/bgz17_bridge.rs 57 FORK fn src REMOVABLE missing - same as l1_avx512: no unsafe op in body; safe #[target_feature(avx2)] fn target_feature_call unsafe fn l1_avx2(a: &[i16; 17], b: &[i16; 17]) -> u32 { +src/hpc/bgz17_bridge.rs 92 FORK fn src IRREDUCIBLE missing - unsafe fn-pointer type needed for target_feature fns - type L1WeightedFn = unsafe fn(&[i16; 17], &[i16; 17]) -> u32; +src/hpc/bgz17_bridge.rs 98 FORK fn src REMOVABLE missing - no unsafe op in body; safe #[target_feature] fn target_feature_call unsafe fn l1_weighted_avx512(a: &[i16; 17], b: &[i16; 17]) -> u32 { +src/hpc/bgz17_bridge.rs 112 FORK fn src REMOVABLE missing - no unsafe op in body; safe #[target_feature] fn target_feature_call unsafe fn l1_weighted_avx2(a: &[i16; 17], b: &[i16; 17]) -> u32 { +src/hpc/bgz17_bridge.rs 156 FORK fn src IRREDUCIBLE missing - unsafe fn-pointer type needed for target_feature fns - type SignAgreementFn = unsafe fn(&[i16; 17], &[i16; 17]) -> u32; +src/hpc/bgz17_bridge.rs 160 FORK fn src REMOVABLE missing - no unsafe op in body; safe #[target_feature] fn target_feature_call unsafe fn sign_agreement_avx512(a: &[i16; 17], b: &[i16; 17]) -> u32 { +src/hpc/bgz17_bridge.rs 173 FORK fn src REMOVABLE missing - no unsafe op in body; safe #[target_feature] fn target_feature_call unsafe fn sign_agreement_avx2(a: &[i16; 17], b: &[i16; 17]) -> u32 { +src/hpc/bgz17_bridge.rs 211 FORK fn src IRREDUCIBLE missing - unsafe fn-pointer type needed for target_feature fns - type XorBindFn = unsafe fn(&[i16; 17], &[i16; 17]) -> [i16; 17]; +src/hpc/bgz17_bridge.rs 215 FORK fn src REMOVABLE missing - no unsafe op in body; safe #[target_feature] fn target_feature_call unsafe fn xor_bind_avx512(a: &[i16; 17], b: &[i16; 17]) -> [i16; 17] { +src/hpc/bgz17_bridge.rs 228 FORK fn src REMOVABLE missing - no unsafe op in body; safe #[target_feature] fn target_feature_call unsafe fn xor_bind_avx2(a: &[i16; 17], b: &[i16; 17]) -> [i16; 17] { +src/hpc/bgz17_bridge.rs 264 FORK fn src IRREDUCIBLE missing - unsafe fn-pointer type needed for target_feature fns - type InjectNoiseFn = unsafe fn(&[i16; 17], i16, u64) -> [i16; 17]; +src/hpc/bgz17_bridge.rs 282 FORK fn src REMOVABLE missing - no unsafe op in body (PRNG + safe simd ops); safe #[target_feature] fn target_feature_call unsafe fn inject_noise_avx512(dims: &[i16; 17], scale: i16, seed: u64) -> [i16; 17] { +src/hpc/bgz17_bridge.rs 309 FORK fn src REMOVABLE missing - no unsafe op in body; safe #[target_feature] fn target_feature_call unsafe fn inject_noise_avx2(dims: &[i16; 17], scale: i16, seed: u64) -> [i16; 17] { +src/hpc/bgz17_bridge.rs 421 FORK block src IRREDUCIBLE present - call through unsafe fn pointer selected by LazyLock after is_x86_feature_detected; or store safe wrapper fn ptrs - unsafe { L1_KERNEL(&self.dims, &other.dims) } +src/hpc/bgz17_bridge.rs 436 FORK block src IRREDUCIBLE present - call through unsafe fn pointer selected post-detection - unsafe { L1_WEIGHTED_KERNEL(&self.dims, &other.dims) } +src/hpc/bgz17_bridge.rs 445 FORK block src IRREDUCIBLE present - call through unsafe fn pointer selected post-detection - unsafe { SIGN_AGREEMENT_KERNEL(&self.dims, &other.dims) } +src/hpc/bgz17_bridge.rs 455 FORK block src IRREDUCIBLE present - call through unsafe fn pointer selected post-detection - let dims = unsafe { XOR_BIND_KERNEL(&self.dims, &other.dims) }; +src/hpc/bgz17_bridge.rs 501 FORK block src IRREDUCIBLE present - call through unsafe fn pointer selected post-detection - let dims = unsafe { INJECT_NOISE_KERNEL(&self.dims, scale, seed) }; +src/hpc/blackboard.rs 295 FORK block src SAFE-API present aliasing: two &mut HashMap reborrows via raw ptr while first value ref live (SB-invalid) HashMap::get_disjoint_mut([key_a,key_b]) (stable 1.86) panics on dup keys; then downcast_mut each ptr_arith unsafe { +src/hpc/blackboard.rs 321 FORK block src SAFE-API present aliasing: two &mut HashMap reborrows via raw ptr while first value ref live (SB-invalid) HashMap::get_disjoint_mut([key_a,key_b]) (stable 1.86) panics on dup keys; then downcast_mut each ptr_arith unsafe { +src/hpc/blackboard.rs 347 FORK block src SAFE-API present aliasing: two &mut HashMap reborrows via raw ptr while first value ref live (SB-invalid) HashMap::get_disjoint_mut([key_a,key_b]) (stable 1.86) panics on dup keys; then downcast_mut each ptr_arith unsafe { +src/hpc/blackboard.rs 377 FORK block src SAFE-API present aliasing: two &mut HashMap reborrows via raw ptr while first value ref live (SB-invalid) HashMap::get_disjoint_mut([a,b,c]) (stable 1.86), then downcast_mut each; drops raw ptr ptr_arith unsafe { +src/hpc/blackboard.rs 413 FORK block src SAFE-API present aliasing: two &mut HashMap reborrows via raw ptr while first value ref live (SB-invalid) HashMap::get_disjoint_mut([a,b,c]) (stable 1.86), then downcast_mut each; drops raw ptr ptr_arith unsafe { +src/hpc/blocked_grid/base.rs 496 FORK impl src IRREDUCIBLE present - unsafe impl Send for raw-ptr strided view; vanishes only if redesigned around split_at_mut/slices - unsafe impl<'a, T: Send, const BR: usize, const BC: usize> Send for GridBlockMut<'a, T, BR, BC> {} +src/hpc/blocked_grid/base.rs 497 FORK impl src IRREDUCIBLE missing - unsafe impl Sync for raw-ptr view (shares comment above, clippy wants its own) - unsafe impl<'a, T: Sync, const BR: usize, const BC: usize> Sync for GridBlockMut<'a, T, BR, BC> {} +src/hpc/blocked_grid/base.rs 524 FORK block src SAFE-API missing bounds: ptr.add(start) with no check start<=len (block coords unchecked) grid.data.as_mut_ptr().wrapping_add(start) is safe; deref sites stay unsafe. Add assert!(start<=len) ptr_arith let data = unsafe { grid.data.as_mut_ptr().add(start) }; +src/hpc/blocked_grid/base.rs 623 FORK fn src UNSAFE-FN-API present - from_raw: caller owns aliasing/bounds contract; could assert data_len>=BC but ptr validity uncheckable ptr_arith pub(super) unsafe fn from_raw( +src/hpc/blocked_grid/grid_struct_macro.rs 320 FORK impl src SAFE-API present bounds: Send impl requires $fty: Send but holds shared ref -> should be Sync store &'a BlockedGrid instead of *const; Send/Sync then auto-derived, impl deleted - unsafe impl<'a> Send for [<$name L1BlockIter>]<'a> +src/hpc/blocked_grid/grid_struct_macro.rs 322 FORK impl src SAFE-API present - same: &'a refs make Sync automatic - unsafe impl<'a> Sync for [<$name L1BlockIter>]<'a> +src/hpc/blocked_grid/grid_struct_macro.rs 349 FORK block src SAFE-API present - with &'a BlockedGrid fields (copy the refs) the deref is unnecessary - unsafe { &*self.$field }, br, bc, +src/hpc/blocked_grid/grid_struct_macro.rs 488 FORK block src SAFE-API present - disjoint field borrows: let f = &mut self.field per field in macro; borrowck accepts distinct fields - unsafe { &mut *[<$field _ptr>] }, br, bc, +src/hpc/blocked_grid/iter.rs 141 FORK impl src IRREDUCIBLE present - unsafe impl Send on iterator holding *mut BlockedGrid (PhantomData<&mut>); not needed if it held &'a mut - unsafe impl<'a, T: Send, const BR: usize, const BC: usize> Send for BaseBlockIterMut<'a, T, BR, BC> {} +src/hpc/blocked_grid/iter.rs 142 FORK impl src IRREDUCIBLE missing - unsafe impl Sync, same reasoning; lacks own SAFETY comment - unsafe impl<'a, T: Sync, const BR: usize, const BC: usize> Sync for BaseBlockIterMut<'a, T, BR, BC> {} +src/hpc/blocked_grid/iter.rs 171 FORK block src SAFE-API missing aliasing: fresh &mut Grid per next() invalidates earlier GridBlockMut raw ptrs under Stacked Borrows hold Option<&'a mut Grid>/data ptr; current &mut *self.ptr reborrows whole grid each next() - let grid: &'a mut BlockedGrid = unsafe { &mut *self.ptr }; +src/hpc/blocked_grid/iter.rs 330 FORK block src CONSOLIDATE present bounds: safe pub fn row_mut checks only via debug_assert; release OOB if r>=BR (UB) assert! (not debug_assert) start+BC<=data_len and r Send +src/hpc/blocked_grid/super_block.rs 330 FORK impl src IRREDUCIBLE missing - unsafe impl Sync on raw-ptr super block; needs own SAFETY comment - unsafe impl<'a, T: Sync, const BR: usize, const BC: usize, const N: usize> Sync +src/hpc/blocked_grid/super_block.rs 444 FORK impl src IRREDUCIBLE present - unsafe impl Send for TierBlockIterMut (raw *mut T from &mut grid) - unsafe impl<'a, T: Send, const BR: usize, const BC: usize, const N: usize> Send for TierBlockIterMut<'a, T, BR, BC, N> {} +src/hpc/blocked_grid/super_block.rs 486 FORK block src SAFE-API present - self.data.wrapping_add(start) is safe; or slice split_at_mut; add assert start<=data_len ptr_arith data: unsafe { self.data.add(start) }, +src/hpc/cam_pq.rs 203 FORK block src IRREDUCIBLE missing - calling #[target_feature(avx512f)] fn after simd_caps().avx512f runtime check target_feature_call return unsafe { self.distance_batch_avx512(cams) }; +src/hpc/cam_pq.rs 216 FORK fn src CONSOLIDATE missing - gather needs unsafe; wrap in safe gather_table256(&[f32;256],[u8;16])->F32x16 (u8 idx<256 in-bounds), or scalar loop target_feature_call unsafe fn distance_batch_avx512(&self, cams: &[CamFingerprint]) -> Vec { +src/hpc/fingerprint.rs 311 FORK block src SAFE-CRATE present - bytemuck::cast_slice(&self.words) (no std safe u64->u8 view) ptr_arith,from_raw_parts unsafe { std::slice::from_raw_parts(self.words.as_ptr() as *const u8, N * 8) } +src/hpc/fingerprint.rs 318 FORK block src SAFE-CRATE present - bytemuck::cast_slice_mut(&mut self.words) ptr_arith,from_raw_parts unsafe { std::slice::from_raw_parts_mut(self.words.as_mut_ptr() as *mut u8, N * 8) } +src/hpc/fingerprint.rs 715 FORK block test SAFE-CRATE present - bytemuck::cast_ref::<[u64;8],[u8;64]> (N==8 only) ptr_arith unsafe { &*(self.words.as_ptr() as *const [u8; 64]) } +src/hpc/gguf.rs 229 FORK block src SAFE-API present align: Vec ptr cast to *const BF16 (align 2); Vec alignment not guaranteed buf.chunks_exact(2).map(|c|BF16(u16::from_le_bytes(..))) or read into Vec; no cast ptr_arith,from_raw_parts unsafe { std::slice::from_raw_parts(buf.as_ptr() as *const super::quantized::BF16, n_elements) }; +src/hpc/gguf_indexer.rs 463 FORK block src SAFE-CRATE missing bounds: batch_elems<=bf16_buf.len() relies on resize above, not asserted bytemuck::cast_slice_mut(&mut bf16_buf[..batch_elems]) (Vec -> bytes, native endian) ptr_arith,from_raw_parts unsafe { std::slice::from_raw_parts_mut(bf16_buf.as_mut_ptr() as *mut u8, batch_elems * 2) }; +src/hpc/int8_tile_gemm.rs 58 FORK block src IRREDUCIBLE present - calling AMX kernel after amx_available(); lens asserted just above - unsafe { +src/hpc/int8_tile_gemm.rs 79 FORK fn src IRREDUCIBLE present - AMX tile asm kernel; extents rely on caller asserts (16k, k16, 256) raw_load_store,ptr_arith unsafe fn amx_path(a_u8: &[u8], b_vnni: &[i8], c: &mut [i32], k: usize) { +src/hpc/int8_tile_gemm.rs 150 FORK fn src UNSAFE-FN-API present - target_feature(avx512vnni); body fully bounds-checked indexing + loadu/storeu on own scratch; could be safe target_feature fn x86_intrinsic pub unsafe fn int8_gemm_vpdpbusd_zmm(a_u8: &[u8], b_i8: &[i8], c: &mut [i32], m: usize, n: usize, k: usize) { +src/hpc/int8_tile_gemm.rs 250 FORK fn src UNSAFE-FN-API present - target_feature(avxvnni); same shape as zmm version; precondition is CPU feature only x86_intrinsic pub unsafe fn int8_gemm_vpdpbusd_ymm(a_u8: &[u8], b_i8: &[i8], c: &mut [i32], m: usize, n: usize, k: usize) { +src/hpc/int8_tile_gemm.rs 413 FORK block src CONSOLIDATE present soundness: safe pub fn int8_gemm_amx_tiled gates AMX (amx_available) and m/n/k%16/16/64 only via debug_assert - release can SIGILL or OOB tile read/write (m%16 bottom strip in _rb, c stores) 10 ops: tile_load/store on raw ptrs; wrap in checked helper taking slices+extent; one-time shape asserts replace debug_asserts raw_load_store,ptr_arith unsafe { +src/hpc/int8_tile_gemm.rs 487 FORK block src IRREDUCIBLE present bounds: store of 16 rows into par_chunks_mut(16*n) chunk assumes m%16==0 (debug_assert only) rayon worker AMX kernel; per-thread LDTILECFG; c_rows chunk may be short if m%16!=0 (store OOB) raw_load_store,ptr_arith unsafe { +src/hpc/int8_tile_gemm.rs 538 FORK block src CONSOLIDATE present bounds: bottom strip reads/writes 16 rows at m32 and 16 cols at n32 with no m%16/n%16 check in release (OOB if misaligned) 43 ops: one block of raw-ptr tile loads/stores; hoist into checked tile helpers (slice + rows + pitch) so block has no raw ptr math raw_load_store,ptr_arith unsafe { +src/hpc/int8_tile_gemm.rs 757 FORK block test IRREDUCIBLE present - test: target_feature call after detection - unsafe { int8_gemm_vpdpbusd_zmm(&a, &b, &mut got, m, n, k) }; +src/hpc/int8_tile_gemm.rs 816 FORK block test IRREDUCIBLE present - test: target_feature call after detection - unsafe { int8_gemm_vpdpbusd_ymm(&a, &b, &mut got, m, n, k) }; +src/hpc/jitson/scan_config.rs 97 FORK block src IRREDUCIBLE present - extern "C" callback from JIT code builds slice from raw ptr+len; better `unsafe extern "C" fn` from_raw_parts let a_slice = unsafe { core::slice::from_raw_parts(a, len) }; +src/hpc/jitson/scan_config.rs 98 FORK block src IRREDUCIBLE present - extern "C" callback from JIT code builds slice from raw ptr+len; better `unsafe extern "C" fn` from_raw_parts let b_slice = unsafe { core::slice::from_raw_parts(b, len) }; +src/hpc/jitson/scan_config.rs 110 FORK block src IRREDUCIBLE present - extern "C" callback from JIT code builds slice from raw ptr+len; better `unsafe extern "C" fn` from_raw_parts let a_i8 = unsafe { core::slice::from_raw_parts(a, len) }; +src/hpc/jitson/scan_config.rs 111 FORK block src IRREDUCIBLE present - extern "C" callback from JIT code builds slice from raw ptr+len; better `unsafe extern "C" fn` from_raw_parts let b_i8 = unsafe { core::slice::from_raw_parts(b, len) }; +src/hpc/jitson/scan_config.rs 138 FORK block src IRREDUCIBLE present align: from_raw_parts:: needs 4-aligned ptr; JIT passes pointers into arbitrary byte buffers JIT callback raw ptr->slice (f32) from_raw_parts let a_slice = unsafe { core::slice::from_raw_parts(a, len) }; +src/hpc/jitson/scan_config.rs 139 FORK block src IRREDUCIBLE present align: as 138 same; clippy wants separate SAFETY comment from_raw_parts let b_slice = unsafe { core::slice::from_raw_parts(b, len) }; +src/hpc/jitson_cranelift/engine.rs 47 FORK impl src SAFE-API present - store fn addresses as usize (already converted to usize in build()); auto Send, impl deleted - unsafe impl Send for JitEngineBuilder {} +src/hpc/jitson_cranelift/engine.rs 62 FORK fn src UNSAFE-FN-API present - register_fn: ptr later CALLED by JIT code with assumed ABI/signature; uncheckable contract ptr_arith pub unsafe fn register_fn(mut self, name: &str, ptr: *const u8) -> Self { +src/hpc/jitson_cranelift/engine.rs 136 FORK impl src SAFE-API present - store fn_ptr/prefetch_chain ptrs as usize -> KernelCache auto Send+Sync - unsafe impl Send for KernelCache {} +src/hpc/jitson_cranelift/engine.rs 137 FORK impl src SAFE-API present - same as 136 - unsafe impl Sync for KernelCache {} +src/hpc/jitson_cranelift/engine.rs 169 FORK impl src IRREDUCIBLE present lifetime: Send/Sync claim over JITModule not verified; kernels not lifetime-tied to engine unsafe impl Send for JitEngine: asserts JITModule movable; unverified here - unsafe impl Send for JitEngine {} +src/hpc/jitson_cranelift/engine.rs 170 FORK impl src IRREDUCIBLE present - unsafe impl Sync for JitEngine; RUN phase reads only fn ptrs - unsafe impl Sync for JitEngine {} +src/hpc/jitson_cranelift/engine.rs 297 FORK block src CONSOLIDATE present - one shared prefetch(ptr) helper (also packed.rs asm); _mm_prefetch is a hint; cfg(target_feature=sse) redundant on x86_64 x86_intrinsic,ptr_arith unsafe { +src/hpc/jitson_cranelift/engine.rs 525 FORK block test CONSOLIDATE present - test calls unsafe ScanKernel::scan; checked safe wrapper scan_slices(&[u8],&[u8],&mut [u64]) removes it ptr_arith let count = unsafe { +src/hpc/jitson_cranelift/noise_jit.rs 60 FORK impl src SAFE-API present - store fn_ptr as usize or typed fn pointer -> auto Send; impl removable - unsafe impl Send for NoiseKernel {} +src/hpc/jitson_cranelift/noise_jit.rs 63 FORK impl src SAFE-API present - same as 60 - unsafe impl Sync for NoiseKernel {} +src/hpc/jitson_cranelift/noise_jit.rs 79 FORK fn src UNSAFE-FN-API present lifetime: NoiseKernel not tied to engine; dangling call if code memory freed (verify JITModule drop) evaluate takes only f64s; could be safe if kernel keeps JITModule alive (Arc) so fn_ptr validity is an invariant transmute,ffi pub unsafe fn evaluate(&self, x: f64, y: f64, z: f64) -> f64 { +src/hpc/jitson_cranelift/noise_jit.rs 82 FORK extern src IRREDUCIBLE present - transmute *const u8 -> extern "C" fn on JIT code; no safe ptr->fn cast. Transmute once in from_raw into typed fn field transmute,ffi let func: unsafe extern "C" fn(f64, f64, f64) -> f64 = std::mem::transmute(self.fn_ptr); +src/hpc/jitson_cranelift/scan_jit.rs 33 FORK impl src SAFE-API present - fn_ptr as usize/typed fn -> auto Send - unsafe impl Send for ScanKernel {} +src/hpc/jitson_cranelift/scan_jit.rs 34 FORK impl src SAFE-API present - same as 33 - unsafe impl Sync for ScanKernel {} +src/hpc/jitson_cranelift/scan_jit.rs 49 FORK fn src UNSAFE-FN-API present lifetime: ScanKernel not tied to engine; dangling call if code memory freed (verify JITModule drop) scan: raw ptr args; safe via scan_slices(&[u8],&[u8],&mut [u64]) asserting len>=field_len*record_size, out>=top_k ptr_arith pub unsafe fn scan( +src/hpc/jitson_cranelift/scan_jit.rs 54 FORK extern src IRREDUCIBLE present - transmute *const u8 -> unsafe extern "C" fn(5 args); signature must match build_scan_ir; use typed fn field ptr_arith,transmute,ffi let func: unsafe extern "C" fn(*const u8, *const u8, u64, u64, *mut u64) -> u64 = +src/hpc/merkle_tree.rs 93 FORK block src SAFE-CRATE present - bytemuck::cast_slice::; std-only: feed to_ne_bytes per word (slower) ptr_arith,from_raw_parts let bytes = unsafe { core::slice::from_raw_parts(words.as_ptr() as *const u8, words.len() * 8) }; +src/hpc/merkle_tree.rs 121 FORK block src SAFE-CRATE present - bytemuck::cast_slice::; std-only: feed to_ne_bytes per word (slower) ptr_arith,from_raw_parts let bytes = unsafe { core::slice::from_raw_parts(container.as_ptr() as *const u8, 256 * 8) }; +src/hpc/merkle_tree.rs 237 FORK block src SAFE-CRATE present - bytemuck::cast_slice::; std-only: feed to_ne_bytes per word (slower) ptr_arith,from_raw_parts unsafe { core::slice::from_raw_parts(self.bits.as_ptr() as *const u8, BITS_WORDS * 8) } +src/hpc/packed.rs 35 FORK block src CONSOLIDATE present - asm! prefetcht0 -> _mm_prefetch::<_MM_HINT_T0>; share one prefetch helper with engine.rs:297 asm unsafe { +src/hpc/packed.rs 48 FORK block src CONSOLIDATE present - asm! prefetcht1 -> _mm_prefetch::<_MM_HINT_T1> in same helper asm unsafe { +src/hpc/palette_distance.rs 20 FORK fn src UNSAFE-FN-API missing - type alias of unsafe fn pointer; bodies need no feature-specific op, wrap in safe fns holding the runtime check - type NearestFn = unsafe fn(&[Base17], &Base17) -> u8; +src/hpc/palette_distance.rs 24 FORK fn src IRREDUCIBLE missing - #[target_feature(avx512f)] fn selected by is_x86_feature_detected; body is scalar loop over l1() target_feature_call unsafe fn nearest_avx512(entries: &[Base17], query: &Base17) -> u8 { +src/hpc/palette_distance.rs 67 FORK fn src IRREDUCIBLE missing - #[target_feature(avx2)] fn selected by runtime check target_feature_call unsafe fn nearest_avx2(entries: &[Base17], query: &Base17) -> u8 { +src/hpc/palette_distance.rs 182 FORK block src CONSOLIDATE present bounds: i as u8 truncates if entries.len()>256 (Palette.entries is pub Vec) call via safe wrapper that dispatches once; nearest_* could be plain fns - unsafe { NEAREST_KERNEL(&self.entries, query) } +src/hpc/simd_dispatch.rs 219 FORK block src IRREDUCIBLE present - call #[target_feature(avx512bw)] fn; runtime check done once in SimdDispatch::detect (caps.avx512bw) target_feature_call unsafe { super::byte_scan::simd_impl::byte_find_all_avx512(haystack, needle) } +src/hpc/simd_dispatch.rs 225 FORK block src IRREDUCIBLE present - call #[target_feature(avx2)] fn after table-time runtime check; kernels are mostly scalar bodies so unsafe is only the feature call target_feature_call unsafe { super::byte_scan::simd_impl::byte_find_all_avx2(haystack, needle) } +src/hpc/simd_dispatch.rs 231 FORK block src IRREDUCIBLE present - call #[target_feature(avx512bw)] fn; runtime check done once in SimdDispatch::detect (caps.avx512bw) target_feature_call unsafe { super::byte_scan::simd_impl::byte_count_avx512(haystack, needle) } +src/hpc/simd_dispatch.rs 237 FORK block src IRREDUCIBLE present - call #[target_feature(avx2)] fn after table-time runtime check; kernels are mostly scalar bodies so unsafe is only the feature call target_feature_call unsafe { super::byte_scan::simd_impl::byte_count_avx2(haystack, needle) } +src/hpc/simd_dispatch.rs 258 FORK block src IRREDUCIBLE present - call #[target_feature(avx2)] fn after table-time runtime check; kernels are mostly scalar bodies so unsafe is only the feature call target_feature_call unsafe { super::distance::simd_impl::squared_distances_avx2(query, points, &mut out) }; +src/hpc/simd_dispatch.rs 278 FORK block src IRREDUCIBLE present precond: callee docs say count>=32 but wrapper does not check (body handles any count; doc stale) call target_feature(avx2) fn after table-time check; body is scalar code target_feature_call unsafe { super::nibble::nibble_unpack_avx2(packed, count, &mut out) }; +src/hpc/simd_dispatch.rs 285 FORK block src IRREDUCIBLE present precond: callee docs say packed.len()>=16 but wrapper does not check (body handles any len; doc stale) call target_feature(avx2) fn after table-time check; body is scalar code target_feature_call unsafe { super::nibble::nibble_above_threshold_avx2(packed, threshold) } +src/hpc/simd_dispatch.rs 297 FORK block src IRREDUCIBLE present - call #[target_feature(avx2)] fn after table-time runtime check; kernels are mostly scalar bodies so unsafe is only the feature call target_feature_call unsafe { super::spatial_hash::batch_sq_dist_avx2(query, candidates, radius_sq) } +src/hpc/vnni_gemm.rs 55 FORK block src IRREDUCIBLE missing feature: kernel enables avx512bw but caps gate checks only avx512f+avx512vnni (benign on real HW, bw not needed by intrinsics used) call target_feature fn after simd_caps().has_avx512_vnni() (avx512f&&vnni; fn also enables avx512bw) target_feature_call unsafe { int8_gemm_vnni_avx512(a, b, c, m, n, k) } +src/hpc/vnni_gemm.rs 105 FORK fn src UNSAFE-FN-API present bounds: raw ptr.add on c/b_packed has no length check inside; only caller simd_int_ops::gemm_u8_i8 and int8_gemm_vnni assert m*n pub(crate) target_feature; target_feature + no inner length asserts; add asserts then make safe target_feature fn target_feature_call pub(crate) unsafe fn int8_gemm_vnni_avx512(a: &[u8], b: &[i8], c: &mut [i32], m: usize, n: usize, k: usize) { +src/hpc/vnni_gemm.rs 225 FORK fn src UNSAFE-FN-API present bounds: raw ptr.add on c has no length check inside; callers assert m*n pub(crate) target_feature(avxvnni); raw ptr.add on c without own length check; add asserts - pub(crate) unsafe fn int8_gemm_avxvnni_ymm(a: &[u8], b: &[i8], c: &mut [i32], m: usize, n: usize, k: usize) { +src/hpc/vsa.rs 185 FORK block src SAFE-CRATE present - bytemuck::cast_slice(&self.words) ptr_arith,from_raw_parts unsafe { std::slice::from_raw_parts(self.words.as_ptr() as *const u8, VSA_WORDS * 8) } +src/impl_1d.rs 48 UPSTREAM block src IRREDUCIBLE missing - rotate1_front moves elements out of owned data via copy_from_nonoverlapping; AbortIfPanic guard; core ptr design ptr_arith unsafe { +src/impl_clone.rs 16 UPSTREAM block src IRREDUCIBLE present - RawDataClone::clone_with_ptr contract (unsafe trait method); core design - unsafe { +src/impl_clone.rs 29 UPSTREAM block src IRREDUCIBLE missing - RawDataClone::clone_from_with_ptr contract; core design - unsafe { +src/impl_constructors.rs 63 UPSTREAM block src SAFE-API missing - checked from_shape_vec(shape, v).unwrap() - re-validates size already checked (O(ndim)), negligible vs alloc - unsafe { Self::from_shape_vec_unchecked(v.len() as Ix, v) } +src/impl_constructors.rs 334 UPSTREAM block src SAFE-API missing - checked from_shape_vec(shape, v).unwrap() - re-validates size already checked (O(ndim)), negligible vs alloc - unsafe { Self::from_shape_vec_unchecked(shape, v) } +src/impl_constructors.rs 402 UPSTREAM block src SAFE-API missing - checked from_shape_vec(shape, v).unwrap() - re-validates size already checked (O(ndim)), negligible vs alloc - unsafe { Self::from_shape_vec_unchecked(shape, v) } +src/impl_constructors.rs 434 UPSTREAM block src SAFE-API missing - checked from_shape_vec(shape, v).unwrap() - re-validates size already checked (O(ndim)), negligible vs alloc - unsafe { Self::from_shape_vec_unchecked(shape, v) } +src/impl_constructors.rs 438 UPSTREAM block src SAFE-API missing - checked from_shape_vec(shape, v).unwrap() - re-validates size already checked (O(ndim)), negligible vs alloc - unsafe { Self::from_shape_vec_unchecked(shape, v) } +src/impl_constructors.rs 491 UPSTREAM block src CONSOLIDATE missing - single audited from_vec_dim_stride_unchecked after explicit validation (already is the helper) - unsafe { Ok(Self::from_vec_dim_stride_unchecked(dim, strides, v)) } +src/impl_constructors.rs 518 UPSTREAM fn src UNSAFE-FN-API present - pub unsafe: caller guarantees v.len()==shape size; checkable -> already exists as safe from_shape_vec - pub unsafe fn from_shape_vec_unchecked(shape: Sh, v: Vec) -> Self +src/impl_constructors.rs 528 UPSTREAM fn src UNSAFE-FN-API missing - private helper; precondition (dim/strides within v) checkable at O(ndim); core design ptr_arith unsafe fn from_vec_dim_stride_unchecked(dim: D, strides: D, mut v: Vec) -> Self { +src/impl_constructors.rs 542 UPSTREAM fn src UNSAFE-FN-API present - trusted-iterator length contract; pub(crate) - pub(crate) unsafe fn from_shape_trusted_iter_unchecked(shape: Sh, iter: I, map: F) -> Self +src/impl_constructors.rs 604 UPSTREAM block src SAFE-API missing - Vec::resize_with(size, MaybeUninit::uninit) (optimises to no-op) + checked from_shape_vec; no set_len uninit unsafe { +src/impl_constructors.rs 641 UPSTREAM block src IRREDUCIBLE present - raw_view_mut_unchecked on unshared fresh array, then deref_into_view_mut; core design - unsafe { +src/impl_cow.rs 35 UPSTREAM block src IRREDUCIBLE present - from_data_ptr/with_strides_dim: same ptr, repr change View->Cow; core design - unsafe { +src/impl_cow.rs 48 UPSTREAM block src IRREDUCIBLE present - from_data_ptr/with_strides_dim: same ptr, repr change Owned->Cow; core design - unsafe { +src/impl_internal_constructors.rs 27 UPSTREAM fn src UNSAFE-FN-API present - pub(crate); ptr must lie in data; debug_assert pointer_is_inbounds already - pub(crate) unsafe fn from_data_ptr(data: S, ptr: NonNull) -> Self { +src/impl_internal_constructors.rs 53 UPSTREAM fn src UNSAFE-FN-API present - pub(crate); strides/dim must be valid for ptr - pub(crate) unsafe fn with_strides_dim(self, strides: E, dim: E) -> ArrayBase +src/impl_methods.rs 133 UPSTREAM block src SAFE-CRATE missing - bytemuck::cast_slice::(s) (same size/align); core has no safe usize->isize slice cast ptr_arith,from_raw_parts unsafe { slice::from_raw_parts(s.as_ptr() as *const _, s.len()) } +src/impl_methods.rs 153 UPSTREAM block src IRREDUCIBLE missing - ArrayView(Mut)/RawArrayView::new unsafe ctor: ndarray core invariant (ptr/dim/strides), ArrayRef itself upholds it - unsafe { ArrayView::new(*self._ptr(), self._dim().clone(), self._strides().clone()) } +src/impl_methods.rs 158 UPSTREAM block src IRREDUCIBLE missing - ArrayView(Mut)/RawArrayView::new unsafe ctor: ndarray core invariant (ptr/dim/strides), ArrayRef itself upholds it - unsafe { ArrayViewMut::new(*self._ptr(), self._dim().clone(), self._strides().clone()) } +src/impl_methods.rs 207 UPSTREAM block src IRREDUCIBLE missing - from_shape_vec_unchecked w/ self strides on contiguous slice; checked from_shape_vec likely OK but negative/zero-len strides unverified - unsafe { +src/impl_methods.rs 342 UPSTREAM block src IRREDUCIBLE missing - deref of first element after is_empty check (core ptr design); safe alt self.iter().next()/iter_mut().next() with iterator-setup cost ptr_arith Some(unsafe { &*self.as_ptr() }) +src/impl_methods.rs 365 UPSTREAM block src IRREDUCIBLE missing - deref of first element after is_empty check (core ptr design); safe alt self.iter().next()/iter_mut().next() with iterator-setup cost ptr_arith Some(unsafe { &mut *self.as_mut_ptr() }) +src/impl_methods.rs 392 UPSTREAM block src IRREDUCIBLE missing - uget on last index after non-empty check; safe alt self.get(index)/get_mut (re-does bounds check) - Some(unsafe { self.uget(index) }) +src/impl_methods.rs 419 UPSTREAM block src IRREDUCIBLE missing - uget on last index after non-empty check; safe alt self.get(index)/get_mut (re-does bounds check) - Some(unsafe { self.uget_mut(index) }) +src/impl_methods.rs 583 UPSTREAM block src IRREDUCIBLE present - with_strides_dim: unsafe ctor, new dim/strides are subset of old data (core design) - unsafe { self.with_strides_dim(new_strides, new_dim) } +src/impl_methods.rs 667 UPSTREAM block src IRREDUCIBLE missing - ptr.offset(offset) after do_slice/do_collapse_axis/invert_axis computed in-bounds offset (core strided-view design) ptr_arith unsafe { +src/impl_methods.rs 775 UPSTREAM block src IRREDUCIBLE missing - raw ptr from get_ptr/get_mut_ptr (index_checked) turned into reference; core design - unsafe { self.get_ptr(index).map(|ptr| &*ptr) } +src/impl_methods.rs 800 UPSTREAM block src IRREDUCIBLE missing - ptr.offset(offset) where offset from index_checked (bounds validated) ptr_arith .map(move |offset| unsafe { ptr.as_ptr().offset(offset) as *const _ }) +src/impl_methods.rs 811 UPSTREAM block src IRREDUCIBLE missing - raw ptr from get_ptr/get_mut_ptr (index_checked) turned into reference; core design - unsafe { self.get_mut_ptr(index).map(|ptr| &mut *ptr) } +src/impl_methods.rs 842 UPSTREAM block src IRREDUCIBLE missing - ptr.offset(offset) where offset from index_checked (bounds validated) ptr_arith .map(move |offset| unsafe { ptr.offset(offset) }) +src/impl_methods.rs 857 UPSTREAM fn src UNSAFE-FN-API present - pub unsafe fn: index must be in-bounds (debug_bounds_check only in debug); contract is the point - pub unsafe fn uget(&self, index: I) -> &A +src/impl_methods.rs 881 UPSTREAM fn src UNSAFE-FN-API present - pub unsafe fn: index must be in-bounds (debug_bounds_check only in debug); contract is the point - pub unsafe fn uget_mut(&mut self, index: I) -> &mut A +src/impl_methods.rs 906 UPSTREAM block src IRREDUCIBLE missing - ptr::swap of two index_checked offsets; no safe two-&mut swap of overlapping view elements besides via get_mut pair (can alias) ptr_arith,ptr_rw unsafe { +src/impl_methods.rs 929 UPSTREAM fn src UNSAFE-FN-API present - pub unsafe fn: index must be in-bounds (debug_bounds_check only in debug); contract is the point - pub unsafe fn uswap(&mut self, index1: I, index2: I) +src/impl_methods.rs 945 UPSTREAM block src IRREDUCIBLE missing - deref data ptr of 0-d array after ndim==0 assert (always one element) ptr_arith unsafe { &*self.as_ptr() } +src/impl_methods.rs 1029 UPSTREAM block src IRREDUCIBLE present - with_strides_dim: unsafe ctor, new dim/strides are subset of old data (core design) - unsafe { self.with_strides_dim(strides, dim) } +src/impl_methods.rs 1041 UPSTREAM block src IRREDUCIBLE missing - ptr.offset(offset) after do_slice/do_collapse_axis/invert_axis computed in-bounds offset (core strided-view design) ptr_arith self.0.ptr = unsafe { self._ptr().offset(offset) }; +src/impl_methods.rs 1086 UPSTREAM block src SAFE-API present - view[index].clone() / view.get(index): index already bounds-checked above; redundant check negligible - unsafe { view.uget(index).clone() } +src/impl_methods.rs 1098 UPSTREAM block src SAFE-API missing - Array::from_shape_vec(dim, vec![]).unwrap(): dim has zero-length axis, checked ctor trivial cost - unsafe { Array::from_shape_vec_unchecked(dim, vec![]) } +src/impl_methods.rs 1541 UPSTREAM block src IRREDUCIBLE present - with_strides_dim: unsafe ctor, new dim/strides are subset of old data (core design) - unsafe { self.with_strides_dim(Ix1(stride as Ix), Ix1(len)) } +src/impl_methods.rs 1621 UPSTREAM block src SAFE-API present - Array::from_shape_vec(dim, v).expect(..): default strides and len from same array; O(ndim) check - unsafe { +src/impl_methods.rs 1681 UPSTREAM block src IRREDUCIBLE missing - ArrayView(Mut)/RawArrayView::new unsafe ctor: ndarray core invariant (ptr/dim/strides), ArrayRef itself upholds it - unsafe { RawArrayView::new(*self._ptr(), self._dim().clone(), self._strides().clone()) } +src/impl_methods.rs 1687 UPSTREAM block src IRREDUCIBLE missing - ArrayView(Mut)/RawArrayView::new unsafe ctor: ndarray core invariant (ptr/dim/strides), ArrayRef itself upholds it - unsafe { RawArrayViewMut::new(*self._ptr(), self._dim().clone(), self._strides().clone()) } +src/impl_methods.rs 1706 UPSTREAM block src IRREDUCIBLE missing - ArrayView(Mut)/RawArrayView::new unsafe ctor: ndarray core invariant (ptr/dim/strides), ArrayRef itself upholds it - unsafe { RawArrayViewMut::new(self.parts.ptr, self.parts.dim.clone(), self.parts.strides.clone()) } +src/impl_methods.rs 1713 UPSTREAM fn src UNSAFE-FN-API present - pub(crate) unsafe fn: caller must ensure unshared owned data; not checkable - pub(crate) unsafe fn raw_view_mut_unchecked(&mut self) -> RawArrayViewMut +src/impl_methods.rs 1728 UPSTREAM block src IRREDUCIBLE missing - slice::from_raw_parts[_mut] over standard-layout contiguous data (core strided-view -> slice) ptr_arith,from_raw_parts unsafe { Some(slice::from_raw_parts_mut(self._ptr().as_ptr(), self.len())) } +src/impl_methods.rs 1756 UPSTREAM block src IRREDUCIBLE missing - from_raw_parts[_mut](ptr.sub(offset)) over contiguous-any-order memory (core design); offset from low-addr helper ptr_arith,from_raw_parts unsafe { Ok(slice::from_raw_parts_mut(self._ptr().sub(offset).as_ptr(), self.len())) } +src/impl_methods.rs 1771 UPSTREAM block src IRREDUCIBLE missing - slice::from_raw_parts[_mut] over standard-layout contiguous data (core strided-view -> slice) ptr_arith,from_raw_parts unsafe { Some(slice::from_raw_parts(self._ptr().as_ptr(), self.len())) } +src/impl_methods.rs 1781 UPSTREAM block src IRREDUCIBLE missing - slice::from_raw_parts[_mut] over standard-layout contiguous data (core strided-view -> slice) ptr_arith,from_raw_parts unsafe { Some(slice::from_raw_parts_mut(self._ptr().as_ptr(), self.len())) } +src/impl_methods.rs 1795 UPSTREAM block src IRREDUCIBLE missing - from_raw_parts[_mut](ptr.sub(offset)) over contiguous-any-order memory (core design); offset from low-addr helper ptr_arith,from_raw_parts unsafe { Some(slice::from_raw_parts(self._ptr().sub(offset).as_ptr(), self.len())) } +src/impl_methods.rs 1816 UPSTREAM block src IRREDUCIBLE missing - from_raw_parts[_mut](ptr.sub(offset)) over contiguous-any-order memory (core design); offset from low-addr helper ptr_arith,from_raw_parts unsafe { Ok(slice::from_raw_parts_mut(self._ptr().sub(offset).as_ptr(), self.len())) } +src/impl_methods.rs 1898 UPSTREAM block src SAFE-API present - len==0 case: ArrayView::from_shape(shape, &[]) (safe, empty slice) instead of from_shape_ptr(self.as_ptr()) ptr_arith unsafe { +src/impl_methods.rs 1905 UPSTREAM block src IRREDUCIBLE missing - ArrayView::new(ptr, shape, to_strides) after reshape_dim validated strides - Ok(to_strides) => unsafe { +src/impl_methods.rs 1915 UPSTREAM block src IRREDUCIBLE missing - from_shape_trusted_iter_unchecked: TrustedIterator length contract, perf-motivated (avoids push/branch) - unsafe { +src/impl_methods.rs 1993 UPSTREAM block src IRREDUCIBLE present - with_strides_dim after contiguity/len-preserving check (core design) - unsafe { +src/impl_methods.rs 2040 UPSTREAM block src IRREDUCIBLE present - with_strides_dim after contiguity/len-preserving check (core design) - unsafe { +src/impl_methods.rs 2089 UPSTREAM block src IRREDUCIBLE present - with_strides_dim after contiguity/len-preserving check (core design) - unsafe { +src/impl_methods.rs 2096 UPSTREAM block src IRREDUCIBLE present - with_strides_dim after contiguity/len-preserving check (core design) - Ok(to_strides) => unsafe { +src/impl_methods.rs 2106 UPSTREAM block src IRREDUCIBLE missing - from_shape_trusted_iter_unchecked: TrustedIterator length contract, perf-motivated (avoids push/branch) - unsafe { +src/impl_methods.rs 2159 UPSTREAM block src IRREDUCIBLE present - with_strides_dim after contiguity/len-preserving check (core design) - unsafe { cl.with_strides_dim(shape.default_strides(), shape) } +src/impl_methods.rs 2162 UPSTREAM block src SAFE-API missing - ArrayBase::from_shape_vec(shape, v).unwrap(): len==shape.size() by construction, cheap check - unsafe { ArrayBase::from_shape_vec_unchecked(shape, v) } +src/impl_methods.rs 2246 UPSTREAM block src IRREDUCIBLE present - from_data_ptr + with_strides_dim, dims only converted to IxDyn - unsafe { +src/impl_methods.rs 2272 UPSTREAM block src IRREDUCIBLE present - with_strides_dim stays unsafe; the unlimited_transmute D->D2 branch could reuse the safe D2::from_dimension path used below - unsafe { +src/impl_methods.rs 2380 UPSTREAM block src IRREDUCIBLE present - ArrayView::new with broadcast (possibly zero) strides; ok only because view is read-only - unsafe { Some(ArrayView::new(*self._ptr(), dim, broadcast_strides)) } +src/impl_methods.rs 2496 UPSTREAM block src IRREDUCIBLE present - with_strides_dim: unsafe ctor, new dim/strides are subset of old data (core design) - unsafe { self.with_strides_dim(new_strides, new_dim) } +src/impl_methods.rs 2621 UPSTREAM block src IRREDUCIBLE missing - ptr.offset(offset) after do_slice/do_collapse_axis/invert_axis computed in-bounds offset (core strided-view design) ptr_arith unsafe { +src/impl_methods.rs 2703 UPSTREAM block src IRREDUCIBLE present - with_strides_dim: unsafe ctor, new dim/strides are subset of old data (core design) - unsafe { +src/impl_methods.rs 2882 UPSTREAM block src IRREDUCIBLE missing - from_shape_trusted_iter_unchecked: TrustedIterator length contract, perf-motivated (avoids push/branch) - unsafe { +src/impl_methods.rs 2910 UPSTREAM block src IRREDUCIBLE missing - from_shape_trusted_iter_unchecked: TrustedIterator length contract, perf-motivated (avoids push/branch) - unsafe { ArrayBase::from_shape_trusted_iter_unchecked(dim.strides(strides), slc.iter_mut(), f) } +src/impl_methods.rs 2912 UPSTREAM block src IRREDUCIBLE missing - from_shape_trusted_iter_unchecked: TrustedIterator length contract, perf-motivated (avoids push/branch) - unsafe { ArrayBase::from_shape_trusted_iter_unchecked(dim, self.iter_mut(), f) } +src/impl_methods.rs 2988 UPSTREAM block src SAFE-API present - A,B:'static already required: use dyn Any downcast (Box::downcast / Option<&mut dyn Any>) instead of transmute; one alloc or optimi - unsafe { unlimited_transmute::(b) } +src/impl_methods.rs 2995 UPSTREAM block src SAFE-API present - A,B:'static already required: use dyn Any downcast (Box::downcast / Option<&mut dyn Any>) instead of transmute; one alloc or optimi - unsafe { unlimited_transmute::, Array>(output) } +src/impl_methods.rs 3195 UPSTREAM block src IRREDUCIBLE present - &*prev / &mut *curr from raw views inside Zip closure; relies on S:DataMut non-aliasing & ordered Zip - Zip::from(prev).and(curr).for_each(|prev, curr| unsafe { +src/impl_methods.rs 3301 UPSTREAM fn src UNSAFE-FN-API present align: ptr.read() requires align_of::()>=align_of::() but only size_of equality is asserted (use read_unaligned/transmute_copy) private unsafe fn: transmute_copy-like; caller guarantees same type ptr_arith,ptr_rw unsafe fn unlimited_transmute(data: A) -> B { +src/impl_owned_array.rs 77 UPSTREAM block src SAFE-API missing - address subtraction: (as_ptr().addr()-data.as_ptr().addr())/size_of::() (ZST branch already above); or keep offset_from ptr_arith let offset = unsafe { self.as_ptr().offset_from(self.data.as_ptr()) }; +src/impl_owned_array.rs 338 UPSTREAM block src IRREDUCIBLE missing - calls unsafe move_into_uninit (moves out of self; set_len(0) afterwards); core design - unsafe { self.move_into_uninit(new_array.into_maybe_uninit()) } +src/impl_owned_array.rs 388 UPSTREAM block src IRREDUCIBLE present - bitwise move of elements + set_len(0) under AbortIfPanic guard; core design ptr_arith,uninit unsafe { +src/impl_owned_array.rs 421 UPSTREAM block src IRREDUCIBLE missing - OwnedRepr::set_len(0) to release elements; needs unsafe set_len contract uninit unsafe { +src/impl_owned_array.rs 438 UPSTREAM block src IRREDUCIBLE present - raw_view_mut + set_len(0) + drop_unreachable_raw; leak-on-panic by design uninit unsafe { +src/impl_owned_array.rs 485 UPSTREAM block src IRREDUCIBLE missing - OwnedRepr reserve/ptr fixup + set_len; core design uninit unsafe { +src/impl_owned_array.rs 709 UPSTREAM block src IRREDUCIBLE present - append: reserve, ptr_offset fixup, clone-into-tail in memory order with len guard; core design - unsafe { +src/impl_owned_array.rs 761 UPSTREAM block src IRREDUCIBLE missing - SetLenOnDrop::drop set_len(self.len) panic-safety guard; core design uninit unsafe { +src/impl_owned_array.rs 840 UPSTREAM block src IRREDUCIBLE missing - append_to_array: offset_from + reserve + ptr update; core design ptr_arith unsafe { +src/impl_owned_array.rs 866 UPSTREAM fn src UNSAFE-FN-API present - drops unreachable elements; raw view + data ptr/len must match; pub(crate) - pub(crate) unsafe fn drop_unreachable_raw(mut self_: RawArrayViewMut, data_ptr: NonNull, data_len: usize) +src/impl_raw_views.rs 20 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design - pub(crate) unsafe fn new(ptr: NonNull, dim: D, strides: D) -> Self { +src/impl_raw_views.rs 25 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn new_(ptr: *const A, dim: D, strides: D) -> Self { +src/impl_raw_views.rs 70 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith pub unsafe fn from_shape_ptr(shape: Sh, ptr: *const A) -> Self +src/impl_raw_views.rs 98 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith pub unsafe fn deref_into_view<'a>(self) -> ArrayView<'a, A, D> { +src/impl_raw_views.rs 117 UPSTREAM block src IRREDUCIBLE missing - raw view construction/split with ptr offset; core design ptr_arith unsafe { self.parts.ptr.as_ptr().offset(offset) } +src/impl_raw_views.rs 122 UPSTREAM block src IRREDUCIBLE missing - raw view construction/split with ptr offset; core design - let left = unsafe { Self::new_(left_ptr, dim_left, self.parts.strides.clone()) }; +src/impl_raw_views.rs 127 UPSTREAM block src IRREDUCIBLE missing - raw view construction/split with ptr offset; core design - let right = unsafe { Self::new_(right_ptr, dim_right, self.parts.strides) }; +src/impl_raw_views.rs 146 UPSTREAM block src IRREDUCIBLE missing - raw view construction/split with ptr offset; core design - unsafe { RawArrayView::new(ptr, self.parts.dim, self.parts.strides) } +src/impl_raw_views.rs 184 UPSTREAM block src IRREDUCIBLE missing - raw view construction/split with ptr offset; core design ptr_arith unsafe { ptr_re.add(1) } +src/impl_raw_views.rs 205 UPSTREAM block src IRREDUCIBLE missing - raw view construction/split with ptr offset; core design - unsafe { +src/impl_raw_views.rs 223 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design - pub(crate) unsafe fn new(ptr: NonNull, dim: D, strides: D) -> Self { +src/impl_raw_views.rs 228 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn new_(ptr: *mut A, dim: D, strides: D) -> Self { +src/impl_raw_views.rs 273 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith pub unsafe fn from_shape_ptr(shape: Sh, ptr: *mut A) -> Self +src/impl_raw_views.rs 299 UPSTREAM block src IRREDUCIBLE missing - raw view construction/split with ptr offset; core design - unsafe { RawArrayView::new(self.parts.ptr, self.parts.dim, self.parts.strides) } +src/impl_raw_views.rs 311 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith pub unsafe fn deref_into_view<'a>(self) -> ArrayView<'a, A, D> { +src/impl_raw_views.rs 325 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith pub unsafe fn deref_into_view_mut<'a>(self) -> ArrayViewMut<'a, A, D> { +src/impl_raw_views.rs 338 UPSTREAM block src IRREDUCIBLE missing - raw view construction/split with ptr offset; core design - unsafe { +src/impl_raw_views.rs 360 UPSTREAM block src IRREDUCIBLE missing - raw view construction/split with ptr offset; core design - unsafe { RawArrayViewMut::new(ptr, self.parts.dim, self.parts.strides) } +src/impl_raw_views.rs 372 UPSTREAM block src IRREDUCIBLE missing - raw view construction/split with ptr offset; core design - unsafe { +src/impl_ref_types.rs 58 UPSTREAM block src IRREDUCIBLE missing - repr(transparent) pointer cast between RawRef/ArrayRef/ArrayBase deref; core design - unsafe { &*ptr } +src/impl_ref_types.rs 77 UPSTREAM block src IRREDUCIBLE missing - repr(transparent) pointer cast between RawRef/ArrayRef/ArrayBase deref; core design - unsafe { &mut *ptr } +src/impl_ref_types.rs 91 UPSTREAM block src IRREDUCIBLE present - repr(transparent) pointer cast between RawRef/ArrayRef/ArrayBase deref; core design ptr_arith unsafe { &*((self as *const ArrayRef) as *const RawRef) } +src/impl_ref_types.rs 103 UPSTREAM block src IRREDUCIBLE present - repr(transparent) pointer cast between RawRef/ArrayRef/ArrayBase deref; core design ptr_arith unsafe { &mut *((self as *mut ArrayRef) as *mut RawRef) } +src/impl_ref_types.rs 136 UPSTREAM block src IRREDUCIBLE missing - repr(transparent) pointer cast between RawRef/ArrayRef/ArrayBase deref; core design - unsafe { &*ptr } +src/impl_ref_types.rs 153 UPSTREAM block src IRREDUCIBLE missing - repr(transparent) pointer cast between RawRef/ArrayRef/ArrayBase deref; core design - unsafe { &mut *ptr } +src/impl_ref_types.rs 165 UPSTREAM block src IRREDUCIBLE missing - repr(transparent) pointer cast between RawRef/ArrayRef/ArrayBase deref; core design - unsafe { &*ptr } +src/impl_ref_types.rs 177 UPSTREAM block src IRREDUCIBLE missing - repr(transparent) pointer cast between RawRef/ArrayRef/ArrayBase deref; core design - unsafe { &mut *ptr } +src/impl_special_element_types.rs 35 UPSTREAM fn src UNSAFE-FN-API present - assume_init contract (all elements initialised); uninit tracking not checkable transmute,uninit pub unsafe fn assume_init(self) -> ArrayBase<>::Output, D> { +src/impl_views/constructors.rs 60 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides ptr_arith unsafe { +src/impl_views/constructors.rs 115 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith pub unsafe fn from_shape_ptr(shape: Sh, ptr: *const A) -> Self +src/impl_views/constructors.rs 165 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides ptr_arith unsafe { +src/impl_views/constructors.rs 220 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith pub unsafe fn from_shape_ptr(shape: Sh, ptr: *mut A) -> Self +src/impl_views/constructors.rs 233 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { ArrayViewMut::new(self.parts.ptr, self.parts.dim, self.parts.strides) } +src/impl_views/constructors.rs 246 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith pub(crate) unsafe fn new(ptr: NonNull, dim: D, strides: D) -> Self { +src/impl_views/constructors.rs 256 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith pub(crate) unsafe fn new_(ptr: *const A, dim: D, strides: D) -> Self { +src/impl_views/constructors.rs 269 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith pub(crate) unsafe fn new(ptr: NonNull, dim: D, strides: D) -> Self { +src/impl_views/constructors.rs 281 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith pub(crate) unsafe fn new_(ptr: *mut A, dim: D, strides: D) -> Self { +src/impl_views/conversions.rs 34 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { ArrayView::new(self.parts.ptr, self.parts.dim, self.parts.strides) } +src/impl_views/conversions.rs 44 UPSTREAM block src IRREDUCIBLE missing - slice::from_raw_parts on contiguous checked layout; core design ptr_arith,from_raw_parts unsafe { Some(slice::from_raw_parts(self.parts.ptr.as_ptr(), self.len())) } +src/impl_views/conversions.rs 59 UPSTREAM block src IRREDUCIBLE missing - slice::from_raw_parts on contiguous checked layout; core design ptr_arith,from_raw_parts unsafe { Some(slice::from_raw_parts(self.parts.ptr.sub(offset).as_ptr(), self.len())) } +src/impl_views/conversions.rs 68 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { RawArrayView::new(self.parts.ptr, self.parts.dim, self.parts.strides) } +src/impl_views/conversions.rs 156 UPSTREAM block src IRREDUCIBLE present - into_cell_view: cast to MathCell then deref_into_view; layout-equivalence argument - unsafe { +src/impl_views/conversions.rs 176 UPSTREAM fn src UNSAFE-FN-API present - into_maybe_uninit: documented Safety (must not write uninit into borrowed array) uninit pub(crate) unsafe fn into_maybe_uninit(self) -> ArrayViewMut<'a, MaybeUninit, D> { +src/impl_views/conversions.rs 194 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { Baseiter::new(self.parts.ptr, self.parts.dim, self.parts.strides) } +src/impl_views/conversions.rs 204 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { Baseiter::new(self.parts.ptr, self.parts.dim, self.parts.strides) } +src/impl_views/conversions.rs 215 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { Baseiter::new(self.parts.ptr, self.parts.dim, self.parts.strides) } +src/impl_views/conversions.rs 272 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { ArrayView::new(self.parts.ptr, self.parts.dim, self.parts.strides) } +src/impl_views/conversions.rs 277 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { RawArrayViewMut::new(self.parts.ptr, self.parts.dim, self.parts.strides) } +src/impl_views/conversions.rs 282 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { Baseiter::new(self.parts.ptr, self.parts.dim, self.parts.strides) } +src/impl_views/conversions.rs 294 UPSTREAM block src IRREDUCIBLE missing - slice::from_raw_parts on contiguous checked layout; core design ptr_arith,from_raw_parts unsafe { Ok(slice::from_raw_parts_mut(self.parts.ptr.as_ptr(), self.len())) } +src/impl_views/conversions.rs 305 UPSTREAM block src IRREDUCIBLE missing - slice::from_raw_parts on contiguous checked layout; core design ptr_arith,from_raw_parts unsafe { Ok(slice::from_raw_parts_mut(self.parts.ptr.sub(offset).as_ptr(), self.len())) } +src/impl_views/indexing.rs 99 UPSTREAM fn src UNSAFE-FN-API present - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design - unsafe fn uget(self, index: I) -> Self::Output; +src/impl_views/indexing.rs 124 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { &*self.get_ptr(index).unwrap_or_else(|| array_out_of_bounds()) } +src/impl_views/indexing.rs 128 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { self.get_ptr(index).map(|ptr| &*ptr) } +src/impl_views/indexing.rs 142 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn uget(self, index: I) -> &'a A { +src/impl_views/indexing.rs 172 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { +src/impl_views/indexing.rs 190 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { +src/impl_views/indexing.rs 207 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn uget(mut self, index: I) -> &'a mut A { +src/impl_views/splitting.rs 94 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { +src/impl_views/splitting.rs 122 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { +src/impl_views/splitting.rs 144 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { +src/impl_views/splitting.rs 204 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { +src/indexes.rs 148 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design - unsafe fn stride_offset(mut self, stride: Self::Stride, index: usize) -> Self { +src/indexes.rs 197 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design - unsafe fn as_ref(&self, ptr: Self::Ptr) -> Self::Item { +src/indexes.rs 201 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design - unsafe fn uget_ptr(&self, i: &Self::Dim) -> Self::Ptr { +src/iterators/chunks.rs 21 UPSTREAM fn src UNSAFE-FN-API missing - macro DSL for NdProducer::as_ref: ptr must be valid element ptr from this producer (trait contract) - unsafe fn item(&self, ptr) { +src/iterators/chunks.rs 113 UPSTREAM fn src UNSAFE-FN-API missing - macro DSL for NdProducer::as_ref: ptr must be valid element ptr from this producer (trait contract) - unsafe fn item(&self, ptr) { +src/iterators/chunks.rs 195 UPSTREAM block src IRREDUCIBLE missing - ArrayView(Mut)::new from iterator element ptr, chunk dims validated at construction - unsafe { +src/iterators/chunks.rs 217 UPSTREAM block src IRREDUCIBLE missing - ArrayView(Mut)::new from iterator element ptr, chunk dims validated at construction - unsafe { +src/iterators/into_iter.rs 39 UPSTREAM block src IRREDUCIBLE missing - release_all_elements + Baseiter::new: takes ownership of elements for by-value iteration (core) - unsafe { +src/iterators/into_iter.rs 63 UPSTREAM block src IRREDUCIBLE missing - ptr.read() moves element out; each ptr yielded once by Baseiter ptr_arith,ptr_rw self.inner.next().map(|p| unsafe { p.as_ptr().read() }) +src/iterators/into_iter.rs 89 UPSTREAM block src IRREDUCIBLE missing - drop_unreachable_raw: drops elements outside the view in Drop; manual drop-glue (core) - unsafe { +src/iterators/lanes.rs 20 UPSTREAM fn src UNSAFE-FN-API missing - macro DSL for NdProducer::as_ref (trait contract) - unsafe fn item(&self, ptr) { +src/iterators/lanes.rs 72 UPSTREAM fn src UNSAFE-FN-API missing - macro DSL for NdProducer::as_ref (trait contract) - unsafe fn item(&self, ptr) { +src/iterators/macros.rs 7 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl Send/Sync on raw-pointer iterators; bounds mirror slice::Iter/IterMut (A:Sync / A:Send) - unsafe impl<'a, A, D> Send for $name<'a, A, D> +src/iterators/macros.rs 13 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl Send/Sync on raw-pointer iterators; bounds mirror slice::Iter/IterMut (A:Sync / A:Send) - unsafe impl<'a, A, D> Sync for $name<'a, A, D> +src/iterators/macros.rs 25 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl Send/Sync on raw-pointer iterators; bounds mirror slice::Iter/IterMut (A:Sync / A:Send) - unsafe impl<'a, A, D> Send for $name<'a, A, D> +src/iterators/macros.rs 31 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl Send/Sync on raw-pointer iterators; bounds mirror slice::Iter/IterMut (A:Sync / A:Send) - unsafe impl<'a, A, D> Sync for $name<'a, A, D> +src/iterators/macros.rs 55 UPSTREAM fn src UNSAFE-FN-API missing - macro-generated NdProducer unsafe methods (as_ref/uget_ptr): ptr/index in-bounds contract - unsafe fn item(&$self_:ident, $ptr:pat) { +src/iterators/macros.rs 82 UPSTREAM fn src UNSAFE-FN-API missing - macro-generated NdProducer unsafe methods (as_ref/uget_ptr): ptr/index in-bounds contract ptr_arith unsafe fn as_ref(&$self_, $ptr: *mut A) -> Self::Item { +src/iterators/macros.rs 86 UPSTREAM fn src UNSAFE-FN-API missing - macro-generated NdProducer unsafe methods (as_ref/uget_ptr): ptr/index in-bounds contract ptr_arith unsafe fn uget_ptr(&self, i: &Self::Dim) -> *mut A { +src/iterators/mod.rs 55 UPSTREAM fn src UNSAFE-FN-API present - Baseiter::new: dim/strides must be correct for ptr; prose doc, no # Safety heading - pub unsafe fn new(ptr: NonNull, len: D, stride: D) -> Baseiter { +src/iterators/mod.rs 73 UPSTREAM block src IRREDUCIBLE missing - NonNull::offset by stride_offset of in-range index (Baseiter core) ptr_arith unsafe { Some(self.ptr.offset(offset)) } +src/iterators/mod.rs 93 UPSTREAM block src IRREDUCIBLE missing - row-wise ptr.offset loops in fold/rfold; offsets bounded by dim ptr_arith unsafe { +src/iterators/mod.rs 137 UPSTREAM block src IRREDUCIBLE missing - NonNull::offset by stride_offset of in-range index (Baseiter core) ptr_arith unsafe { Some(self.ptr.offset(offset)) } +src/iterators/mod.rs 149 UPSTREAM block src IRREDUCIBLE missing - NonNull::offset by stride_offset of in-range index (Baseiter core) ptr_arith unsafe { Some(self.ptr.offset(offset)) } +src/iterators/mod.rs 163 UPSTREAM block src IRREDUCIBLE missing - row-wise ptr.offset loops in fold/rfold; offsets bounded by dim ptr_arith unsafe { +src/iterators/mod.rs 214 UPSTREAM block src IRREDUCIBLE missing - NonNull::as_ref/as_mut or as_ref(ptr) on ptr yielded by Baseiter (in-bounds, lifetime tied to iterator) - self.inner.next().map(|p| unsafe { p.as_ref() }) +src/iterators/mod.rs 225 UPSTREAM block src IRREDUCIBLE missing - NonNull::as_ref/as_mut or as_ref(ptr) on ptr yielded by Baseiter (in-bounds, lifetime tied to iterator) - unsafe { self.inner.fold(init, move |acc, ptr| g(acc, ptr.as_ref())) } +src/iterators/mod.rs 232 UPSTREAM block src IRREDUCIBLE missing - NonNull::as_ref/as_mut or as_ref(ptr) on ptr yielded by Baseiter (in-bounds, lifetime tied to iterator) - self.inner.next_back().map(|p| unsafe { p.as_ref() }) +src/iterators/mod.rs 239 UPSTREAM block src IRREDUCIBLE missing - NonNull::as_ref/as_mut or as_ref(ptr) on ptr yielded by Baseiter (in-bounds, lifetime tied to iterator) - unsafe { self.inner.rfold(init, move |acc, ptr| g(acc, ptr.as_ref())) } +src/iterators/mod.rs 616 UPSTREAM block src IRREDUCIBLE missing - NonNull::as_ref/as_mut or as_ref(ptr) on ptr yielded by Baseiter (in-bounds, lifetime tied to iterator) - self.inner.next().map(|mut p| unsafe { p.as_mut() }) +src/iterators/mod.rs 627 UPSTREAM block src IRREDUCIBLE missing - NonNull::as_ref/as_mut or as_ref(ptr) on ptr yielded by Baseiter (in-bounds, lifetime tied to iterator) - unsafe { +src/iterators/mod.rs 637 UPSTREAM block src IRREDUCIBLE missing - NonNull::as_ref/as_mut or as_ref(ptr) on ptr yielded by Baseiter (in-bounds, lifetime tied to iterator) - self.inner.next_back().map(|mut p| unsafe { p.as_mut() }) +src/iterators/mod.rs 644 UPSTREAM block src IRREDUCIBLE missing - NonNull::as_ref/as_mut or as_ref(ptr) on ptr yielded by Baseiter (in-bounds, lifetime tied to iterator) - unsafe { +src/iterators/mod.rs 716 UPSTREAM block src IRREDUCIBLE missing - ArrayView(Mut)::new[_] from iterator ptr, inner_len/stride fixed at construction - .map(|ptr| unsafe { ArrayView::new(ptr, Ix1(self.inner_len), Ix1(self.inner_stride as Ix)) }) +src/iterators/mod.rs 737 UPSTREAM block src IRREDUCIBLE missing - ArrayView(Mut)::new[_] from iterator ptr, inner_len/stride fixed at construction - .map(|ptr| unsafe { ArrayView::new(ptr, Ix1(self.inner_len), Ix1(self.inner_stride as Ix)) }) +src/iterators/mod.rs 764 UPSTREAM block src IRREDUCIBLE missing - ArrayView(Mut)::new[_] from iterator ptr, inner_len/stride fixed at construction - .map(|ptr| unsafe { ArrayViewMut::new(ptr, Ix1(self.inner_len), Ix1(self.inner_stride as Ix)) }) +src/iterators/mod.rs 785 UPSTREAM block src IRREDUCIBLE missing - ArrayView(Mut)::new[_] from iterator ptr, inner_len/stride fixed at construction - .map(|ptr| unsafe { ArrayViewMut::new(ptr, Ix1(self.inner_len), Ix1(self.inner_stride as Ix)) }) +src/iterators/mod.rs 839 UPSTREAM fn src UNSAFE-FN-API missing - unsafe offset/as_ref/uget_ptr: index *mut A { +src/iterators/mod.rs 898 UPSTREAM block src IRREDUCIBLE missing - offset(index) guarded by index0 (comment explains), else dangling ptr ptr_arith unsafe { self.iter.offset(self.iter.index) } +src/iterators/mod.rs 1150 UPSTREAM fn src UNSAFE-FN-API missing - unsafe offset/as_ref/uget_ptr: index Self::Item { +src/iterators/mod.rs 1154 UPSTREAM fn src UNSAFE-FN-API missing - unsafe offset/as_ref/uget_ptr: index Self::Ptr { +src/iterators/mod.rs 1187 UPSTREAM block src IRREDUCIBLE present - offset only when len()>0 (comment explains), else dangling ptr ptr_arith unsafe { self.iter.offset(self.iter.index) } +src/iterators/mod.rs 1200 UPSTREAM fn src UNSAFE-FN-API missing - unsafe offset/as_ref/uget_ptr: index Self::Item { +src/iterators/mod.rs 1204 UPSTREAM fn src UNSAFE-FN-API missing - unsafe offset/as_ref/uget_ptr: index Self::Ptr { +src/iterators/mod.rs 1321 UPSTREAM block src IRREDUCIBLE missing - ArrayView(Mut)::new[_] from iterator ptr, inner_len/stride fixed at construction - unsafe { $array::new_(ptr, self.iter.inner_dim.clone(), self.iter.inner_strides.clone()) } +src/iterators/mod.rs 1323 UPSTREAM block src IRREDUCIBLE missing - ArrayView(Mut)::new[_] from iterator ptr, inner_len/stride fixed at construction - unsafe { $array::new_(ptr, self.partial_chunk_dim.clone(), self.iter.inner_strides.clone()) } +src/iterators/mod.rs 1439 UPSTREAM trait src IRREDUCIBLE missing - unsafe trait TrustedIterator: exact-length contract, used for unchecked Vec building (internal) - pub unsafe trait TrustedIterator {} +src/iterators/mod.rs 1446 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl TrustedIterator: marker for exact-size iterators (std slice/Range/Map trivially satisfy) - unsafe impl TrustedIterator for Linspace {} +src/iterators/mod.rs 1448 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl TrustedIterator: marker for exact-size iterators (std slice/Range/Map trivially satisfy) - unsafe impl TrustedIterator for Geomspace {} +src/iterators/mod.rs 1450 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl TrustedIterator: marker for exact-size iterators (std slice/Range/Map trivially satisfy) - unsafe impl TrustedIterator for Logspace {} +src/iterators/mod.rs 1451 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl TrustedIterator: marker for exact-size iterators (std slice/Range/Map trivially satisfy) - unsafe impl TrustedIterator for Iter<'_, A, D> {} +src/iterators/mod.rs 1452 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl TrustedIterator: marker for exact-size iterators (std slice/Range/Map trivially satisfy) - unsafe impl TrustedIterator for IterMut<'_, A, D> {} +src/iterators/mod.rs 1453 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl TrustedIterator: marker for exact-size iterators (std slice/Range/Map trivially satisfy) - unsafe impl TrustedIterator for std::iter::Cloned where I: TrustedIterator {} +src/iterators/mod.rs 1454 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl TrustedIterator: marker for exact-size iterators (std slice/Range/Map trivially satisfy) - unsafe impl TrustedIterator for std::iter::Map where I: TrustedIterator {} +src/iterators/mod.rs 1455 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl TrustedIterator: marker for exact-size iterators (std slice/Range/Map trivially satisfy) - unsafe impl TrustedIterator for slice::Iter<'_, A> {} +src/iterators/mod.rs 1456 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl TrustedIterator: marker for exact-size iterators (std slice/Range/Map trivially satisfy) - unsafe impl TrustedIterator for slice::IterMut<'_, A> {} +src/iterators/mod.rs 1457 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl TrustedIterator: marker for exact-size iterators (std slice/Range/Map trivially satisfy) - unsafe impl TrustedIterator for ::std::ops::Range {} +src/iterators/mod.rs 1459 UPSTREAM impl src IRREDUCIBLE missing soundness: upstream FIXME says size "needs to be checked up front"; to_vec_mapped writes via raw ptr trusting len unsafe impl TrustedIterator for indices iter - unsafe impl TrustedIterator for IndicesIter where D: Dimension {} +src/iterators/mod.rs 1460 UPSTREAM impl src IRREDUCIBLE missing soundness: upstream FIXME says size "needs to be checked up front"; to_vec_mapped writes via raw ptr trusting len unsafe impl TrustedIterator for indices iter - unsafe impl TrustedIterator for IndicesIterF where D: Dimension {} +src/iterators/mod.rs 1461 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl TrustedIterator: marker for exact-size iterators (std slice/Range/Map trivially satisfy) - unsafe impl TrustedIterator for IntoIter where D: Dimension {} +src/iterators/mod.rs 1484 UPSTREAM block src IRREDUCIBLE missing bounds: writes out_ptr without capacity check, relies entirely on TrustedIterator exactness ptr::write + set_len + offset into reserved Vec buffer (hand-rolled collect for TrustedIterator vectorization) ptr_arith,ptr_rw,uninit iter.fold((), |(), elt| unsafe { +src/iterators/windows.rs 67 UPSTREAM fn src UNSAFE-FN-API missing - macro/NdProducer unsafe methods (as_ref/uget_ptr): ptr/index contract - unsafe fn item(&self, ptr) { +src/iterators/windows.rs 115 UPSTREAM block src IRREDUCIBLE missing - ArrayView::new from iterator ptr, window dims fixed at construction - unsafe { +src/iterators/windows.rs 184 UPSTREAM fn src UNSAFE-FN-API missing - macro/NdProducer unsafe methods (as_ref/uget_ptr): ptr/index contract ptr_arith unsafe fn as_ref(&self, ptr: *mut A) -> Self::Item { +src/iterators/windows.rs 188 UPSTREAM fn src UNSAFE-FN-API missing - macro/NdProducer unsafe methods (as_ref/uget_ptr): ptr/index contract ptr_arith unsafe fn uget_ptr(&self, i: &Self::Dim) -> *mut A { +src/lib.rs 2089 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { ArrayView::new(*ptr, dim, strides) } +src/lib.rs 2103 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { self.with_strides_dim(s, d) } +src/linalg/impl_linalg.rs 87 UPSTREAM block src SAFE-API missing - self.iter().zip(rhs.iter()).fold(A::zero(), ...) (len asserted equal above) removes both uget - unsafe { +src/linalg/impl_linalg.rs 114 UPSTREAM block src IRREDUCIBLE missing - FFI: cblas dot via blas_sys; ptr from blas_1d_params (feature blas) ptr_arith unsafe { +src/linalg/impl_linalg.rs 138 UPSTREAM fn src UNSAFE-FN-API missing - ptr.offset for negative stride; could use wrapping_offset/NonNull math; feature blas ptr_arith unsafe fn blas_1d_params(ptr: *const A, len: usize, stride: isize) -> (*const A, blas_index, blas_index) { +src/linalg/impl_linalg.rs 316 UPSTREAM block src SAFE-API missing uninit: Vec::set_len on uninitialised A ("A is Copy so safe" - still UB for float/int validity per Rust rules; beta=0 path) Array::uninit((m,n).set_f(cm)) + mat_mul_impl on raw view + assume_init, as at :368; avoids Vec::set_len uninit unsafe { +src/linalg/impl_linalg.rs 368 UPSTREAM block src IRREDUCIBLE missing - Array1::uninit + raw view + assume_init: deliberate uninit output for gemv with beta=0; core design uninit unsafe { +src/linalg/impl_linalg.rs 455 UPSTREAM block src IRREDUCIBLE missing - FFI: cblas gemm via blas_sys; ptr/stride/layout validated above (feature blas) ptr_arith,ffi unsafe { +src/linalg/impl_linalg.rs 502 UPSTREAM block src IRREDUCIBLE missing - FFI: matrixmultiply::{sgemm,dgemm,cgemm,zgemm} raw-pointer kernel call with validated dims/strides ptr_arith unsafe { +src/linalg/impl_linalg.rs 521 UPSTREAM block src IRREDUCIBLE missing - FFI: matrixmultiply::{sgemm,dgemm,cgemm,zgemm} raw-pointer kernel call with validated dims/strides ptr_arith unsafe { +src/linalg/impl_linalg.rs 540 UPSTREAM block src IRREDUCIBLE missing - FFI: matrixmultiply::{sgemm,dgemm,cgemm,zgemm} raw-pointer kernel call with validated dims/strides ptr_arith unsafe { +src/linalg/impl_linalg.rs 561 UPSTREAM block src IRREDUCIBLE missing - FFI: matrixmultiply::{sgemm,dgemm,cgemm,zgemm} raw-pointer kernel call with validated dims/strides ptr_arith unsafe { +src/linalg/impl_linalg.rs 595 UPSTREAM block src SAFE-API missing - fold over (0..k) with lhs[(i,x)]*rhs[(x,j)] and c[(i,j)] checked indexing, or Zip/outer iterators; small-matrix fallback loop - unsafe { +src/linalg/impl_linalg.rs 653 UPSTREAM block src IRREDUCIBLE missing - general_mat_vec_mul_impl with y raw view derived from &mut ArrayRef1 (valid for write) - unsafe { general_mat_vec_mul_impl(alpha, a, x, beta, y.raw_view_mut()) } +src/linalg/impl_linalg.rs 665 UPSTREAM fn src UNSAFE-FN-API present - raw dest view must be valid for writing; may be uninit iff beta==0 - unsafe fn general_mat_vec_mul_impl(alpha: A, a: &ArrayRef2, x: &ArrayRef1, beta: A, y: RawArrayViewMut) +src/linalg/impl_linalg.rs 765 UPSTREAM block src IRREDUCIBLE missing - assume_init after Zip writes every chunk; init by construction uninit unsafe { out.assume_init() } +src/linalg/impl_linalg.rs 784 UPSTREAM block src SAFE-API missing - (a as &dyn Any).downcast_ref::().copied().expect(..) - A,B are 'static so no unsafe ptr::read ptr_arith,ptr_rw unsafe { ::std::ptr::read(a as *const _ as *const B) } +src/nibble.rs 32 FORK block src SAFE-API present - nibble_unpack_avx2 is scalar array code equal to the scalar path; delete it, call nibble_unpack_scalar target_feature_call unsafe { +src/nibble.rs 59 FORK fn src REMOVABLE present - Scalar body, no unsafe op target_feature_call pub(crate) unsafe fn nibble_unpack_avx2(packed: &[u8], count: usize, out: &mut Vec) { +src/nibble.rs 143 FORK block src SAFE-API present - Callee nibble_sub_clamp_avx512 uses U8x64 polyfill, no unsafe op; plain fn target_feature_call unsafe { +src/nibble.rs 150 FORK block src SAFE-API present - nibble_sub_clamp_avx2 is a scalar loop; delete, use nibble_sub_clamp_scalar target_feature_call unsafe { +src/nibble.rs 170 FORK fn src REMOVABLE missing - Scalar body, no unsafe op target_feature_call unsafe fn nibble_sub_clamp_avx2(packed: &mut [u8], delta: u8) { +src/nibble.rs 197 FORK fn src REMOVABLE present - U8x64 polyfill body, no unsafe op target_feature_call unsafe fn nibble_sub_clamp_avx512(packed: &mut [u8], delta: u8) { +src/nibble.rs 232 FORK block src SAFE-API present - nibble_above_threshold_avx2 is scalar; delete, call the _scalar fn target_feature_call return unsafe { nibble_above_threshold_avx2(packed, threshold) }; +src/nibble.rs 258 FORK fn src REMOVABLE present - Scalar body, no unsafe op target_feature_call pub(crate) unsafe fn nibble_above_threshold_avx2(packed: &[u8], threshold: u8) -> Vec { +src/palette_codec.rs 271 FORK block src SAFE-API present - unpack_generic_avx512 is a scalar shift/mask loop = unpack_indices; delete and call scalar target_feature_call return unsafe { unpack_generic_avx512(packed, bits_per_index, count) }; +src/palette_codec.rs 274 FORK block src SAFE-API missing - unpack_4bit_avx2 is scalar array code; remove, call unpack_indices (no SAFETY comment here) target_feature_call return unsafe { unpack_4bit_avx2(packed, count) }; +src/palette_codec.rs 288 FORK block src SAFE-API present - pack_generic_avx512 is a scalar loop = pack_indices; delete target_feature_call return unsafe { pack_generic_avx512(indices, bits_per_index) }; +src/palette_codec.rs 304 FORK fn src REMOVABLE present - Scalar body, no unsafe op target_feature_call unsafe fn unpack_generic_avx512(packed: &[u64], bits_per_index: usize, count: usize) -> Vec { +src/palette_codec.rs 335 FORK fn src REMOVABLE present - Scalar body, no unsafe op target_feature_call unsafe fn pack_generic_avx512(indices: &[u8], bits_per_index: usize) -> Vec { +src/palette_codec.rs 353 FORK fn src REMOVABLE missing - Body has no unsafe op (uses safe bytemuck_cast_u64_to_u8 helper) target_feature_call unsafe fn unpack_4bit_avx2(packed: &[u64], count: usize) -> Vec { +src/palette_codec.rs 408 FORK block src SAFE-API present - packed.iter().flat_map(u64::to_le_bytes) / per-word to_le_bytes loop; alt bytemuck::cast_slice ptr_arith,from_raw_parts unsafe { core::slice::from_raw_parts(words.as_ptr() as *const u8, words.len() * 8) } +src/palette_codec.rs 427 FORK block src SAFE-API present - bedrock_reorder_xzy_avx512 is scalar; delete and use the scalar loop below it target_feature_call return unsafe { bedrock_reorder_xzy_avx512(states) }; +src/palette_codec.rs 505 FORK fn src SAFE-API present - get_unchecked(max idx 4095) -> plain indexing or states.try_into::<&[u16;4096]>() (bounds elided); target_feature adds nothing unchecked,target_feature_call unsafe fn bedrock_reorder_xzy_avx512(states: &[u16]) -> Vec { +src/parallel/impl_par_methods.rs 90 UPSTREAM block src IRREDUCIBLE missing - SendProducer::new (unconditional Send over raw view for disjoint chunked writes); core rayon design - let splits = unsafe { +src/parallel/impl_par_methods.rs 101 UPSTREAM block src IRREDUCIBLE missing - collect_with_partial unsafe contract: writes to disjoint output chunk - unsafe { +src/parallel/impl_par_methods.rs 117 UPSTREAM block src IRREDUCIBLE missing - assume_init after all chunks written (checked via Partial len) uninit unsafe { +src/parallel/send_producer.rs 13 UPSTREAM fn src UNSAFE-FN-API missing - caller asserts producer is safe to send across threads (disjoint access) - pub(crate) unsafe fn new(producer: T) -> Self { +src/parallel/send_producer.rs 18 UPSTREAM impl src IRREDUCIBLE missing unbounded: Send for any P (pub(crate) only; sound only if constructed via unsafe new) unsafe impl Send over arbitrary P: raw-pointer producer; soundness by disjoint split contract - unsafe impl

Send for SendProducer

{} +src/parallel/send_producer.rs 65 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design - unsafe fn as_ref(&self, ptr: Self::Ptr) -> Self::Item { +src/parallel/send_producer.rs 70 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design - unsafe fn uget_ptr(&self, i: &Self::Dim) -> Self::Ptr { +src/partial.rs 32 UPSTREAM fn src UNSAFE-FN-API present - ptr must be dereferenceable for len elements; has # Safety doc ptr_arith pub(crate) unsafe fn new(ptr: *mut T) -> Self { +src/partial.rs 78 UPSTREAM impl src IRREDUCIBLE missing - unsafe impl Send for raw-pointer owner Partial where T:Send - unsafe impl Send for Partial where T: Send {} +src/partial.rs 83 UPSTREAM block src IRREDUCIBLE missing - drop_in_place on owned slice of initialised prefix; raw-pointer owner ptr_rw,from_raw_parts unsafe { +src/property_mask.rs 101 FORK block src SAFE-API present - test_section_avx512 is U64x8 polyfill + safe indexing; drop unsafe/target_feature (SAFETY text about pointers is stale) target_feature_call unsafe { +src/property_mask.rs 108 FORK block src SAFE-API present - test_section_avx2 is scalar array code; delete, call test_section_scalar target_feature_call unsafe { +src/property_mask.rs 126 FORK block src SAFE-API present - count_section_avx512 has no unsafe op (polyfill + count_ones); plain safe fn target_feature_call return unsafe { self.count_section_avx512(states) }; +src/property_mask.rs 161 FORK fn src REMOVABLE missing - No unsafe op in body; U64x8::from_slice asserts length target_feature_call unsafe fn test_section_avx512(&self, states: &[u64], result: &mut [u64]) { +src/property_mask.rs 205 FORK fn src REMOVABLE missing - No unsafe op in body target_feature_call unsafe fn count_section_avx512(&self, states: &[u64]) -> u32 { +src/property_mask.rs 248 FORK fn src REMOVABLE missing - Scalar body, no unsafe op target_feature_call unsafe fn test_section_avx2(&self, states: &[u64], result: &mut [u64]) { +src/property_mask.rs 316 FORK block src SAFE-API present - count_section_multi_avx512 has no unsafe op; plain safe fn target_feature_call unsafe { +src/property_mask.rs 342 FORK fn src REMOVABLE missing - No unsafe op in body target_feature_call unsafe fn count_section_multi_avx512(masks: &[PropertyMask], states: &[u64]) -> MultiMaskResult { +src/simd.rs 1121 FORK block test CONSOLIDATE present - test-only oracle: loadu/storeu via one safe helper over &[u32;N]; cfg(target_feature=avx2) gate is compile-time x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd.rs 1292 FORK block test CONSOLIDATE present - test-only oracle: loadu/storeu via one safe helper over &[u32;N]; cfg(target_feature=avx2) gate is compile-time x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_amx.rs 217 FORK block src IRREDUCIBLE present - _xgetbv is unsafe (needs xsave); legal once OSXSAVE verified at step 2; no safe wrapper in core::arch - let xcr0: u64 = unsafe { core::arch::x86_64::_xgetbv(0) }; +src/simd_amx.rs 244 FORK block src IRREDUCIBLE present - asm! raw syscall arch_prctl(ARCH_REQ_XCOMP_PERM); no std API (libc::syscall would be FFI). SAFETY comment sits above `let ret`, clippy misses it asm unsafe { +src/simd_amx.rs 324 FORK fn src UNSAFE-FN-API missing - target_feature(avx512vnni) wrapper over value intrinsic; could be a safe #[target_feature] fn (TF1.1), callers still need unsafe off-feature - pub unsafe fn vnni_dpbusd( +src/simd_amx.rs 341 FORK fn src UNSAFE-FN-API missing - only precondition is avx512vnni CPU support; loads via row[off..] slicing are bounds-checked; could be safe target_feature fn; add # Safety x86_intrinsic,raw_load_store,ptr_arith pub unsafe fn vnni_dot_u8_i8(row: &[u8], energy: &[i8]) -> i32 { +src/simd_amx.rs 367 FORK fn src UNSAFE-FN-API missing correctness: calls vnni_dot_u8_i8 which drops n%64 tail -> avx512 path differs from vnni2/scalar for n not multiple of 64 (not UB) precondition = avx512vnni only; indexing is checked (panics, not UB); add # Safety doc - pub unsafe fn vnni_matvec(table: &[u8], energy_i8: &[i8], result: &mut [i32], n: usize) { +src/simd_amx.rs 389 FORK fn src UNSAFE-FN-API present - precondition = avx2+avxvnni at runtime; loads from checked slices; could be safe target_feature fn x86_intrinsic,raw_load_store,ptr_arith pub unsafe fn vnni2_dot_u8_i8(row: &[u8], energy: &[i8]) -> i32 { +src/simd_amx.rs 422 FORK fn src UNSAFE-FN-API missing - precondition = avx2+avxvnni; indexing checked; add # Safety doc - pub unsafe fn vnni2_matvec(table: &[u8], energy_i8: &[i8], result: &mut [i32], n: usize) { +src/simd_amx.rs 474 FORK block src IRREDUCIBLE missing - call target_feature fn after is_x86_feature_detected!(avx512vnni); check complete (avx512vnni implies avx512f in rustc) - unsafe { +src/simd_amx.rs 480 FORK block src IRREDUCIBLE missing - call target_feature fn after avx2&&avxvnni detection; matches callee's enable list exactly - unsafe { +src/simd_avx2.rs 358 FORK block src CONSOLIDATE missing feature: safe pub fn uses AVX2 intrinsics with no cfg(target_feature)/runtime check (SIGILL on non-avx2 x86_64 build) one safe load_u8x32(&[u8;32]) helper (chunks_exact(32)+try_into); transmute __m256i->[i64;4] -> store helper or _mm256_extract_epi64 x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_avx2.rs 410 FORK block src CONSOLIDATE missing bounds: no a.len()==b.len() check; b shorter than a but >base passes b[base..] and loadu reads 32B past end. feature: pub safe fn uses AVX2 unchecked safe load/store helpers over &[u8;32]/&mut [i32;8]; chunks_exact(32) zip for a,b x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_avx2.rs 689 FORK fn src UNSAFE-FN-API present - contract is per-lane signed offset in one allocation; not checkable at runtime without base slice len -> take &[f32] and bounds-check ind... ptr_arith pub unsafe fn gather(indices: I32x16, base_ptr: *const f32) -> Self { +src/simd_avx2.rs 695 FORK block src IRREDUCIBLE present - raw pointer offset+deref under caller contract (gather); safe form needs &[f32] base + checked index ptr_arith o[i] = unsafe { *base_ptr.offset(idx[i] as isize) }; +src/simd_avx2.rs 1507 FORK fn src UNSAFE-FN-API missing no # Safety doc; writes ptr.add(i) with no stated bound take &mut [u8;64] -> safe masked-store loop raw_load_store,ptr_arith pub unsafe fn mask_store(self, ptr: *mut u8, mask: u64) { +src/simd_avx2.rs 1668 FORK block src CONSOLIDATE present - two loadu of the 64B array: take &[u64;8] -> halves via one load helper (or from_array twice, array is exactly 64B) x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_avx2.rs 1681 FORK block src CONSOLIDATE present - two storeu into [u64;8]: one store helper taking &mut [u64;8] x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_avx2.rs 1699 FORK block src SAFE-API present feature: relies on AVX2 being present, not enforced by cfg value intrinsics only (sll/srl/or): make fn #[target_feature(enable="avx2")] (safe fn) or u64 lane rotate_left loop; unsafe is only the f... x86_intrinsic unsafe { +src/simd_avx2.rs 1741 FORK block src SAFE-API present feature: AVX2 assumed, not enforced value intrinsic only (_mm256_mul_epu32): #[target_feature(avx2)] safe fn or (a as u32 as u64)*(b as u32 as u64) lane loop x86_intrinsic,target_feature_call unsafe { Self::from_avx2_halves(_mm256_mul_epu32(a_lo, b_lo), _mm256_mul_epu32(a_hi, b_hi)) } +src/simd_avx2.rs 1773 FORK block src SAFE-API present feature: AVX2 assumed, not enforced value shuffles only (unpack/permute2x128): #[target_feature(avx2)] fn, or array index loops (file already measured LLVM emits vpunpck/vperm) x86_intrinsic let c = unsafe { +src/simd_avx2.rs 1825 FORK block src SAFE-API present feature: AVX2 assumed, not enforced value intrinsic only (sllv/srlv): #[target_feature(avx2)] fn or per-lane << >> loop (counts<64 asserted in debug) x86_intrinsic,target_feature_call unsafe { Self::from_avx2_halves(_mm256_sllv_epi64(lo, clo), _mm256_sllv_epi64(hi, chi)) } +src/simd_avx2.rs 1840 FORK block src SAFE-API present feature: AVX2 assumed, not enforced value intrinsic only (sllv/srlv): #[target_feature(avx2)] fn or per-lane << >> loop (counts<64 asserted in debug) x86_intrinsic,target_feature_call unsafe { Self::from_avx2_halves(_mm256_srlv_epi64(lo, clo), _mm256_srlv_epi64(hi, chi)) } +src/simd_avx2.rs 2067 FORK block src SAFE-API missing feature: AVX2 assumed, not enforced (pub safe fn) value-only set1/setzero: [v;16] array or #[target_feature(avx2)] safe fn; unsafe only feature gate x86_intrinsic Self(unsafe { _mm256_set1_epi16(v as i16) }) +src/simd_avx2.rs 2072 FORK block src SAFE-API missing feature: AVX2 assumed, not enforced (pub safe fn) value-only set1/setzero: [v;16] array or #[target_feature(avx2)] safe fn; unsafe only feature gate x86_intrinsic Self(unsafe { _mm256_setzero_si256() }) +src/simd_avx2.rs 2079 FORK block src CONSOLIDATE present - assert!(len>=16) then s[..16].try_into::<&[u16;16]> -> same load helper as from_array x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm256_loadu_si256(s.as_ptr() as *const __m256i) }) +src/simd_avx2.rs 2084 FORK block src CONSOLIDATE missing - ONE audited loadu over &[u16;16]; this is that constructor, other sites should route through it x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm256_loadu_si256(arr.as_ptr() as *const __m256i) }) +src/simd_avx2.rs 2091 FORK block src CONSOLIDATE present - one storeu helper over &mut [u16;16]; to_array is that helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm256_storeu_si256(arr.as_mut_ptr() as *mut __m256i, self.0) }; +src/simd_avx2.rs 2098 FORK block src CONSOLIDATE missing - copy_to_slice: s[..16].try_into::<&mut [u16;16]> then store helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm256_storeu_si256(s.as_mut_ptr() as *mut __m256i, self.0) }; +src/simd_avx2.rs 2107 FORK block src SAFE-API present - value-only shift; #[target_feature(avx2)] fn or lane loop with imm<16 guard x86_intrinsic Self(unsafe { _mm256_srl_epi16(self.0, _mm_cvtsi32_si128(imm as i32)) }) +src/simd_avx2.rs 2115 FORK block src SAFE-API present - value-only shift; #[target_feature(avx2)] fn or lane loop with imm<16 guard x86_intrinsic Self(unsafe { _mm256_sll_epi16(self.0, _mm_cvtsi32_si128(imm as i32)) }) +src/simd_avx2.rs 2121 FORK block src SAFE-API missing - value-only mullo: lane wrapping_mul loop (LLVM -> vpmullw) or target_feature fn x86_intrinsic Self(unsafe { _mm256_mullo_epi16(self.0, other.0) }) +src/simd_avx2.rs 2139 FORK block src CONSOLIDATE present - value-only const-generic shuffle/blend: wrap in #[target_feature(avx2)] fn; no array-loop equivalent for const IMM x86_intrinsic Self(unsafe { _mm256_permute2x128_si256::(self.0, other.0) }) +src/simd_avx2.rs 2148 FORK block src CONSOLIDATE present - value-only const-generic shuffle/blend: wrap in #[target_feature(avx2)] fn; no array-loop equivalent for const IMM x86_intrinsic Self(unsafe { _mm256_blend_epi32::(self.0, other.0) }) +src/simd_avx2.rs 2158 FORK block src CONSOLIDATE present - value-only cvt chain; wrap in #[target_feature(avx2)] fn or u16->f32 lane loop (vpmovzx+vcvtdq2ps via LLVM) x86_intrinsic crate::simd_avx512::F32x8(unsafe { _mm256_cvtepi32_ps(_mm256_cvtepu16_epi32(_mm256_castsi256_si128(self.0))) }) +src/simd_avx2.rs 2165 FORK block src CONSOLIDATE present - value-only cvt chain; wrap in #[target_feature(avx2)] fn or u16->f32 lane loop (vpmovzx+vcvtdq2ps via LLVM) x86_intrinsic crate::simd_avx512::F32x8(unsafe { +src/simd_avx2.rs 2182 FORK block src SAFE-API missing - lane wrapping_add/sub/mul loop on [u16;16] (LLVM emits vpaddw/vpsubw/vpmullw) or target_feature fn x86_intrinsic Self(unsafe { _mm256_add_epi16(self.0, rhs.0) }) +src/simd_avx2.rs 2189 FORK block src SAFE-API missing - lane wrapping_add/sub/mul loop on [u16;16] (LLVM emits vpaddw/vpsubw/vpmullw) or target_feature fn x86_intrinsic Self(unsafe { _mm256_sub_epi16(self.0, rhs.0) }) +src/simd_avx2.rs 2196 FORK block src SAFE-API missing - lane wrapping_add/sub/mul loop on [u16;16] (LLVM emits vpaddw/vpsubw/vpmullw) or target_feature fn x86_intrinsic Self(unsafe { _mm256_mullo_epi16(self.0, rhs.0) }) +src/simd_avx2.rs 2202 FORK block src SAFE-API missing - lane wrapping_add/sub/mul loop on [u16;16] (LLVM emits vpaddw/vpsubw/vpmullw) or target_feature fn x86_intrinsic self.0 = unsafe { _mm256_add_epi16(self.0, rhs.0) }; +src/simd_avx2.rs 2208 FORK block src SAFE-API missing - lane wrapping_add/sub/mul loop on [u16;16] (LLVM emits vpaddw/vpsubw/vpmullw) or target_feature fn x86_intrinsic self.0 = unsafe { _mm256_sub_epi16(self.0, rhs.0) }; +src/simd_avx2.rs 2215 FORK block src SAFE-API missing - bitwise and/or/xor/not on [u64;4] lanes (u64 ops, LLVM vpand/vpor/vpxor) or target_feature fn x86_intrinsic Self(unsafe { _mm256_and_si256(self.0, rhs.0) }) +src/simd_avx2.rs 2222 FORK block src SAFE-API missing - bitwise and/or/xor/not on [u64;4] lanes (u64 ops, LLVM vpand/vpor/vpxor) or target_feature fn x86_intrinsic Self(unsafe { _mm256_or_si256(self.0, rhs.0) }) +src/simd_avx2.rs 2229 FORK block src SAFE-API missing - bitwise and/or/xor/not on [u64;4] lanes (u64 ops, LLVM vpand/vpor/vpxor) or target_feature fn x86_intrinsic Self(unsafe { _mm256_xor_si256(self.0, rhs.0) }) +src/simd_avx2.rs 2235 FORK block src SAFE-API missing - bitwise and/or/xor/not on [u64;4] lanes (u64 ops, LLVM vpand/vpor/vpxor) or target_feature fn x86_intrinsic self.0 = unsafe { _mm256_and_si256(self.0, rhs.0) }; +src/simd_avx2.rs 2241 FORK block src SAFE-API missing - bitwise and/or/xor/not on [u64;4] lanes (u64 ops, LLVM vpand/vpor/vpxor) or target_feature fn x86_intrinsic self.0 = unsafe { _mm256_or_si256(self.0, rhs.0) }; +src/simd_avx2.rs 2247 FORK block src SAFE-API missing - bitwise and/or/xor/not on [u64;4] lanes (u64 ops, LLVM vpand/vpor/vpxor) or target_feature fn x86_intrinsic self.0 = unsafe { _mm256_xor_si256(self.0, rhs.0) }; +src/simd_avx2.rs 2254 FORK block src SAFE-API missing - bitwise and/or/xor/not on [u64;4] lanes (u64 ops, LLVM vpand/vpor/vpxor) or target_feature fn x86_intrinsic Self(unsafe { _mm256_xor_si256(self.0, _mm256_set1_epi16(-1)) }) +src/simd_avx2.rs 2591 FORK block src CONSOLIDATE present - two loadu of 64B array; same halves-load helper as U64x8::avx2_halves (one audited line) x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_avx2.rs 2604 FORK block src SAFE-API present feature: AVX2 assumed, not enforced value-only min/max tree: #[target_feature(avx2)] fn, or iter().min() (file measured it scalar, so keep intrinsics via target_feature fn) x86_intrinsic unsafe { +src/simd_avx2.rs 2618 FORK block src SAFE-API present feature: AVX2 assumed, not enforced value-only min/max tree: #[target_feature(avx2)] fn, or iter().min() (file measured it scalar, so keep intrinsics via target_feature fn) x86_intrinsic unsafe { +src/simd_avx2.rs 2691 FORK block src SAFE-API present feature: AVX2 assumed, not enforced value-only movemask/cmpgt: #[target_feature(avx2)] fn x86_intrinsic let neg = unsafe { +src/simd_avx2.rs 2724 FORK block src SAFE-API present feature: AVX2 assumed, not enforced value-only movemask/cmpgt: #[target_feature(avx2)] fn x86_intrinsic unsafe { +src/simd_avx2.rs 2828 FORK block src SAFE-API present feature: AVX2 assumed, not enforced (pub safe fn) value-only set1: [v;32] array or target_feature fn x86_intrinsic Self(unsafe { _mm256_set1_epi8(v as i8) }) +src/simd_avx2.rs 2836 FORK block src CONSOLIDATE present - assert!(len>=32) then &s[..32] as &[u8;32] -> from_array x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm256_loadu_si256(s.as_ptr() as *const __m256i) }) +src/simd_avx2.rs 2849 FORK fn src UNSAFE-FN-API present - pub unsafe fn from_ptr: documented contract (32 readable bytes) is the point; unchecked hot-loop load, cannot be checked without a slice x86_intrinsic,raw_load_store,ptr_arith pub unsafe fn from_ptr(ptr: *const u8) -> Self { +src/simd_avx2.rs 2857 FORK block src CONSOLIDATE present - this IS the audited loadu-from-array constructor; other sites should call it x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm256_loadu_si256(arr.as_ptr() as *const __m256i) }) +src/simd_avx2.rs 2865 FORK block src CONSOLIDATE present - this IS the audited storeu-to-array; copy_to_slice/sum_bytes_u64 should call it x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm256_storeu_si256(out.as_mut_ptr() as *mut __m256i, self.0) }; +src/simd_avx2.rs 2874 FORK block src CONSOLIDATE present - s[..32].try_into::<&mut [u8;32]> + store helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm256_storeu_si256(s.as_mut_ptr() as *mut __m256i, self.0) }; +src/simd_avx2.rs 2907 FORK block src SAFE-API present - value-only sad_epu8: target_feature fn, or byte iter sum as u64 x86_intrinsic let sums = unsafe { _mm256_sad_epu8(self.0, _mm256_setzero_si256()) }; +src/simd_avx2.rs 2911 FORK block src CONSOLIDATE missing - use to_array-style store helper over &mut [u64;4] (or _mm256_extract_epi64 x4) x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm256_storeu_si256(tmp.as_mut_ptr() as *mut __m256i, sums) }; +src/simd_avx2.rs 2921 FORK block src SAFE-API present - value-only u8 min/max/adds/subs/avg: lane iter ops (min/max/saturating_*) LLVM vectorizes, or target_feature fn x86_intrinsic Self(unsafe { _mm256_min_epu8(self.0, other.0) }) +src/simd_avx2.rs 2928 FORK block src SAFE-API present - value-only u8 min/max/adds/subs/avg: lane iter ops (min/max/saturating_*) LLVM vectorizes, or target_feature fn x86_intrinsic Self(unsafe { _mm256_max_epu8(self.0, other.0) }) +src/simd_avx2.rs 2939 FORK block src SAFE-API present - value-only cmpeq: target_feature fn or lane compare to bitmask x86_intrinsic let eq = unsafe { _mm256_cmpeq_epi8(self.0, other.0) }; +src/simd_avx2.rs 2942 FORK block src SAFE-API missing - value-only movemask: target_feature fn x86_intrinsic unsafe { _mm256_movemask_epi8(eq) as u32 } +src/simd_avx2.rs 2952 FORK block src SAFE-API present - value-only xor/cmpgt/movemask: target_feature fn x86_intrinsic unsafe { +src/simd_avx2.rs 2966 FORK block src SAFE-API present - value-only movemask/shuffle/unpack/blendv: #[target_feature(avx2)] fn x86_intrinsic unsafe { _mm256_movemask_epi8(self.0) as u32 } +src/simd_avx2.rs 2975 FORK block src SAFE-API present - value-only u8 min/max/adds/subs/avg: lane iter ops (min/max/saturating_*) LLVM vectorizes, or target_feature fn x86_intrinsic Self(unsafe { _mm256_adds_epu8(self.0, other.0) }) +src/simd_avx2.rs 2982 FORK block src SAFE-API present - value-only u8 min/max/adds/subs/avg: lane iter ops (min/max/saturating_*) LLVM vectorizes, or target_feature fn x86_intrinsic Self(unsafe { _mm256_subs_epu8(self.0, other.0) }) +src/simd_avx2.rs 2989 FORK block src SAFE-API present - value-only u8 min/max/adds/subs/avg: lane iter ops (min/max/saturating_*) LLVM vectorizes, or target_feature fn x86_intrinsic Self(unsafe { _mm256_avg_epu8(self.0, other.0) }) +src/simd_avx2.rs 3000 FORK block src SAFE-API present - value-only 16-bit shift: target_feature fn or u16 lane shift loop x86_intrinsic Self(unsafe { _mm256_srl_epi16(self.0, _mm_cvtsi32_si128(imm as i32)) }) +src/simd_avx2.rs 3007 FORK block src SAFE-API present - value-only 16-bit shift: target_feature fn or u16 lane shift loop x86_intrinsic Self(unsafe { _mm256_sll_epi16(self.0, _mm_cvtsi32_si128(imm as i32)) }) +src/simd_avx2.rs 3020 FORK block src SAFE-API present - value-only movemask/shuffle/unpack/blendv: #[target_feature(avx2)] fn x86_intrinsic Self(unsafe { _mm256_shuffle_epi8(self.0, idx.0) }) +src/simd_avx2.rs 3045 FORK block src SAFE-API present - value-only movemask/shuffle/unpack/blendv: #[target_feature(avx2)] fn x86_intrinsic Self(unsafe { _mm256_unpacklo_epi8(self.0, other.0) }) +src/simd_avx2.rs 3052 FORK block src SAFE-API present - value-only movemask/shuffle/unpack/blendv: #[target_feature(avx2)] fn x86_intrinsic Self(unsafe { _mm256_unpackhi_epi8(self.0, other.0) }) +src/simd_avx2.rs 3064 FORK block src SAFE-API present - value-only movemask/shuffle/unpack/blendv: #[target_feature(avx2)] fn x86_intrinsic Self(unsafe { _mm256_blendv_epi8(b.0, a.0, mask.0) }) +src/simd_avx2.rs 3100 FORK block src SAFE-API present - bitwise op on [u64;4] lanes or target_feature fn x86_intrinsic Self(unsafe { _mm256_and_si256(self.0, rhs.0) }) +src/simd_avx2.rs 3110 FORK block src SAFE-API present - bitwise op on [u64;4] lanes or target_feature fn x86_intrinsic Self(unsafe { _mm256_or_si256(self.0, rhs.0) }) +src/simd_avx2.rs 3120 FORK block src SAFE-API present - bitwise op on [u64;4] lanes or target_feature fn x86_intrinsic Self(unsafe { _mm256_xor_si256(self.0, rhs.0) }) +src/simd_avx2.rs 3130 FORK block src SAFE-API present - lane wrapping_add/sub loop (vpaddb/vpsubb) or target_feature fn x86_intrinsic Self(unsafe { _mm256_add_epi8(self.0, rhs.0) }) +src/simd_avx2.rs 3140 FORK block src SAFE-API present - lane wrapping_add/sub loop (vpaddb/vpsubb) or target_feature fn x86_intrinsic Self(unsafe { _mm256_sub_epi8(self.0, rhs.0) }) +src/simd_avx512.rs 50 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe - Self(unsafe { $intr(self.0, rhs.0) }) +src/simd_avx512.rs 61 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe - self.0 = unsafe { $intr(self.0, rhs.0) }; +src/simd_avx512.rs 78 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_setzero_ps() }) +src/simd_avx512.rs 87 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_set1_ps(v) }) +src/simd_avx512.rs 93 FORK block src CONSOLIDATE missing - assert!(len) present; route via from_array w/ <&[T;N]>::try_from(&s[..N]) -> one unsafe load(&[T;N]) helper x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_ps(s.as_ptr()) }) +src/simd_avx512.rs 98 FORK block src CONSOLIDATE missing - this is the leaf load; keep as the single unsafe load(&[T;N]) helper (alt: safe_unaligned_simd) x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_ps(arr.as_ptr()) }) +src/simd_avx512.rs 104 FORK block src CONSOLIDATE missing - this is the leaf store (to_array); keep as the single unsafe store(&mut [T;N]) helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_ps(arr.as_mut_ptr(), self.0) }; +src/simd_avx512.rs 111 FORK block src CONSOLIDATE missing - assert!(len) present; route via to_array/&mut [T;N] -> one unsafe store(&mut [T;N]) helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_ps(s.as_mut_ptr(), self.0) }; +src/simd_avx512.rs 118 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_reduce_add_ps(self.0) } +src/simd_avx512.rs 123 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_reduce_min_ps(self.0) } +src/simd_avx512.rs 128 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_reduce_max_ps(self.0) } +src/simd_avx512.rs 135 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_min_ps(self.0, other.0) }) +src/simd_avx512.rs 140 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_max_ps(self.0, other.0) }) +src/simd_avx512.rs 152 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_fmadd_ps(self.0, b.0, c.0) }) +src/simd_avx512.rs 157 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_sqrt_ps(self.0) }) +src/simd_avx512.rs 164 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_roundscale_ps::<0x08>(self.0) }) +src/simd_avx512.rs 171 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_roundscale_ps::<0x09>(self.0) }) +src/simd_avx512.rs 176 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { +src/simd_avx512.rs 186 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic U32x16(unsafe { _mm512_castps_si512(self.0) }) +src/simd_avx512.rs 191 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_castsi512_ps(bits.0) }) +src/simd_avx512.rs 199 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic I32x16(unsafe { _mm512_cvttps_epi32(self.0) }) +src/simd_avx512.rs 206 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic F32Mask16(unsafe { _mm512_cmp_ps_mask::<_CMP_EQ_OQ>(self.0, other.0) }) +src/simd_avx512.rs 211 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic F32Mask16(unsafe { _mm512_cmp_ps_mask::<_CMP_NEQ_UQ>(self.0, other.0) }) +src/simd_avx512.rs 216 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic F32Mask16(unsafe { _mm512_cmp_ps_mask::<_CMP_LT_OS>(self.0, other.0) }) +src/simd_avx512.rs 221 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic F32Mask16(unsafe { _mm512_cmp_ps_mask::<_CMP_LE_OS>(self.0, other.0) }) +src/simd_avx512.rs 244 FORK fn src UNSAFE-FN-API present - pub unsafe fn gather: idx bounds not checkable cheaply; safe form = gather from &[f32] + min/max idx check (costs a reduce) x86_intrinsic,ptr_arith pub unsafe fn gather(indices: I32x16, base_ptr: *const f32) -> Self { +src/simd_avx512.rs 262 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { +src/simd_avx512.rs 317 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic F32x16(unsafe { _mm512_mask_blend_ps(self.0, false_val.0, true_val.0) }) +src/simd_avx512.rs 332 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_setzero_pd() }) +src/simd_avx512.rs 341 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_set1_pd(v) }) +src/simd_avx512.rs 347 FORK block src CONSOLIDATE missing - assert!(len) present; route via from_array w/ <&[T;N]>::try_from(&s[..N]) -> one unsafe load(&[T;N]) helper x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_pd(s.as_ptr()) }) +src/simd_avx512.rs 352 FORK block src CONSOLIDATE missing - this is the leaf load; keep as the single unsafe load(&[T;N]) helper (alt: safe_unaligned_simd) x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_pd(arr.as_ptr()) }) +src/simd_avx512.rs 358 FORK block src CONSOLIDATE missing - this is the leaf store (to_array); keep as the single unsafe store(&mut [T;N]) helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_pd(arr.as_mut_ptr(), self.0) }; +src/simd_avx512.rs 365 FORK block src CONSOLIDATE missing - assert!(len) present; route via to_array/&mut [T;N] -> one unsafe store(&mut [T;N]) helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_pd(s.as_mut_ptr(), self.0) }; +src/simd_avx512.rs 370 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_reduce_add_pd(self.0) } +src/simd_avx512.rs 375 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_reduce_min_pd(self.0) } +src/simd_avx512.rs 380 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_reduce_max_pd(self.0) } +src/simd_avx512.rs 385 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_min_pd(self.0, other.0) }) +src/simd_avx512.rs 390 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_max_pd(self.0, other.0) }) +src/simd_avx512.rs 400 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_fmadd_pd(self.0, b.0, c.0) }) +src/simd_avx512.rs 405 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_sqrt_pd(self.0) }) +src/simd_avx512.rs 410 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_roundscale_pd::<0x08>(self.0) }) +src/simd_avx512.rs 415 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_roundscale_pd::<0x09>(self.0) }) +src/simd_avx512.rs 420 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { +src/simd_avx512.rs 428 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic U64x8(unsafe { _mm512_castpd_si512(self.0) }) +src/simd_avx512.rs 433 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_castsi512_pd(bits.0) }) +src/simd_avx512.rs 440 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic F64Mask8(unsafe { _mm512_cmp_pd_mask::<_CMP_EQ_OQ>(self.0, other.0) }) +src/simd_avx512.rs 445 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic F64Mask8(unsafe { _mm512_cmp_pd_mask::<_CMP_NEQ_UQ>(self.0, other.0) }) +src/simd_avx512.rs 450 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic F64Mask8(unsafe { _mm512_cmp_pd_mask::<_CMP_LT_OS>(self.0, other.0) }) +src/simd_avx512.rs 455 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic F64Mask8(unsafe { _mm512_cmp_pd_mask::<_CMP_LE_OS>(self.0, other.0) }) +src/simd_avx512.rs 482 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { +src/simd_avx512.rs 512 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic F64x8(unsafe { _mm512_mask_blend_pd(self.0, false_val.0, true_val.0) }) +src/simd_avx512.rs 529 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_set1_epi8(v as i8) }) +src/simd_avx512.rs 535 FORK block src CONSOLIDATE missing - assert!(len) present; route via from_array w/ <&[T;N]>::try_from(&s[..N]) -> one unsafe load(&[T;N]) helper x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_si512(s.as_ptr() as *const _) }) +src/simd_avx512.rs 540 FORK block src CONSOLIDATE missing - this is the leaf load; keep as the single unsafe load(&[T;N]) helper (alt: safe_unaligned_simd) x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_si512(arr.as_ptr() as *const _) }) +src/simd_avx512.rs 546 FORK block src CONSOLIDATE missing - this is the leaf store (to_array); keep as the single unsafe store(&mut [T;N]) helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_si512(arr.as_mut_ptr() as *mut _, self.0) }; +src/simd_avx512.rs 553 FORK block src CONSOLIDATE missing - assert!(len) present; route via to_array/&mut [T;N] -> one unsafe store(&mut [T;N]) helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_si512(s.as_mut_ptr() as *mut _, self.0) }; +src/simd_avx512.rs 559 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { +src/simd_avx512.rs 595 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_min_epu8(self.0, other.0) }) +src/simd_avx512.rs 600 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_max_epu8(self.0, other.0) }) +src/simd_avx512.rs 610 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_cmpeq_epi8_mask(self.0, other.0) } +src/simd_avx512.rs 619 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { +src/simd_avx512.rs 637 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_subs_epu8(self.0, other.0) }) +src/simd_avx512.rs 647 FORK block src NEEDS-TF present feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_avg_epu8(self.0, other.0) }) +src/simd_avx512.rs 656 FORK block src NEEDS-TF present feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_cmpgt_epu8_mask(self.0, other.0) } +src/simd_avx512.rs 664 FORK block src NEEDS-TF present feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_mask_blend_epi8(mask, a.0, b.0) }) +src/simd_avx512.rs 671 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { +src/simd_avx512.rs 694 FORK fn src UNSAFE-FN-API present - ptr contract; safe form: take &mut [u8;64] (type carries length) then one internal unsafe masked store x86_intrinsic,raw_load_store,ptr_arith pub unsafe fn mask_store(self, ptr: *mut u8, mask: u64) { +src/simd_avx512.rs 704 FORK block src NEEDS-TF present feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_adds_epu8(self.0, other.0) }) +src/simd_avx512.rs 722 FORK block src IRREDUCIBLE present - calls #[target_feature(avx512vbmi)] fn after runtime simd_caps() check - unsafe { Self(permute_bytes_vbmi(self.0, idx.0)) } +src/simd_avx512.rs 743 FORK block src NEEDS-TF present feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_movepi8_mask(self.0) } +src/simd_avx512.rs 749 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_unpacklo_epi8(self.0, other.0) }) +src/simd_avx512.rs 755 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_unpackhi_epi8(self.0, other.0) }) +src/simd_avx512.rs 762 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_shuffle_epi8(self.0, idx.0) }) +src/simd_avx512.rs 771 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { +src/simd_avx512.rs 784 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { +src/simd_avx512.rs 801 FORK fn src IRREDUCIBLE present - unsafe fn with target_feature(avx512vbmi); callable only after runtime check; contract is the feature precondition x86_intrinsic unsafe fn permute_bytes_vbmi(v: __m512i, idx: __m512i) -> __m512i { +src/simd_avx512.rs 816 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { +src/simd_avx512.rs 861 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { +src/simd_avx512.rs 893 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_set1_epi32(v) }) +src/simd_avx512.rs 899 FORK block src CONSOLIDATE missing - assert!(len) present; route via from_array w/ <&[T;N]>::try_from(&s[..N]) -> one unsafe load(&[T;N]) helper x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_si512(s.as_ptr() as *const _) }) +src/simd_avx512.rs 904 FORK block src CONSOLIDATE missing - this is the leaf load; keep as the single unsafe load(&[T;N]) helper (alt: safe_unaligned_simd) x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_si512(arr.as_ptr() as *const _) }) +src/simd_avx512.rs 910 FORK block src CONSOLIDATE missing - this is the leaf store (to_array); keep as the single unsafe store(&mut [T;N]) helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_si512(arr.as_mut_ptr() as *mut _, self.0) }; +src/simd_avx512.rs 917 FORK block src CONSOLIDATE missing - assert!(len) present; route via to_array/&mut [T;N] -> one unsafe store(&mut [T;N]) helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_si512(s.as_mut_ptr() as *mut _, self.0) }; +src/simd_avx512.rs 922 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_reduce_add_epi32(self.0) } +src/simd_avx512.rs 927 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_reduce_min_epi32(self.0) } +src/simd_avx512.rs 932 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_reduce_max_epi32(self.0) } +src/simd_avx512.rs 943 FORK block src CONSOLIDATE missing - assert!(len) present; route via from_array w/ <&[T;N]>::try_from(&s[..N]) -> one unsafe load(&[T;N]) helper x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_cvtepi16_epi32(_mm256_loadu_si256(s.as_ptr() as *const __m256i)) }) +src/simd_avx512.rs 949 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_abs_epi32(self.0) }) +src/simd_avx512.rs 955 FORK block src CONSOLIDATE missing - cvt is value-only; only the 256-bit storeu on local [i16;16] is unsafe -> shared store helper x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_avx512.rs 966 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_cmpge_epi32_mask(self.0, _mm512_setzero_si512()) } +src/simd_avx512.rs 992 FORK block src NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_cmpgt_epi32_mask(self.0, other.0) } +src/simd_avx512.rs 997 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_min_epi32(self.0, other.0) }) +src/simd_avx512.rs 1002 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_max_epi32(self.0, other.0) }) +src/simd_avx512.rs 1008 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic F32x16(unsafe { _mm512_cvtepi32_ps(self.0) }) +src/simd_avx512.rs 1055 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { +src/simd_avx512.rs 1066 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { Self(_mm512_sub_epi32(_mm512_setzero_si512(), self.0)) } +src/simd_avx512.rs 1095 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_set1_epi64(v) }) +src/simd_avx512.rs 1101 FORK block src CONSOLIDATE missing - assert!(len) present; route via from_array w/ <&[T;N]>::try_from(&s[..N]) -> one unsafe load(&[T;N]) helper x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_si512(s.as_ptr() as *const _) }) +src/simd_avx512.rs 1106 FORK block src CONSOLIDATE missing - this is the leaf load; keep as the single unsafe load(&[T;N]) helper (alt: safe_unaligned_simd) x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_si512(arr.as_ptr() as *const _) }) +src/simd_avx512.rs 1112 FORK block src CONSOLIDATE missing - this is the leaf store (to_array); keep as the single unsafe store(&mut [T;N]) helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_si512(arr.as_mut_ptr() as *mut _, self.0) }; +src/simd_avx512.rs 1119 FORK block src CONSOLIDATE missing - assert!(len) present; route via to_array/&mut [T;N] -> one unsafe store(&mut [T;N]) helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_si512(s.as_mut_ptr() as *mut _, self.0) }; +src/simd_avx512.rs 1124 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_reduce_add_epi64(self.0) } +src/simd_avx512.rs 1129 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_reduce_min_epi64(self.0) } +src/simd_avx512.rs 1134 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_reduce_max_epi64(self.0) } +src/simd_avx512.rs 1139 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_min_epi64(self.0, other.0) }) +src/simd_avx512.rs 1144 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_max_epi64(self.0, other.0) }) +src/simd_avx512.rs 1149 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_abs_epi64(self.0) }) +src/simd_avx512.rs 1216 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { +src/simd_avx512.rs 1227 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { Self(_mm512_sub_epi64(_mm512_setzero_si512(), self.0)) } +src/simd_avx512.rs 1258 FORK block src NEEDS-TF present feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_set1_epi16(v as i16) }) +src/simd_avx512.rs 1263 FORK block src NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_setzero_si512() }) +src/simd_avx512.rs 1270 FORK block src CONSOLIDATE present - assert!(len) present; route via from_array w/ <&[T;N]>::try_from(&s[..N]) -> one unsafe load(&[T;N]) helper x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_si512(s.as_ptr() as *const _) }) +src/simd_avx512.rs 1276 FORK block src CONSOLIDATE present - this is the leaf load; keep as the single unsafe load(&[T;N]) helper (alt: safe_unaligned_simd) x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_si512(arr.as_ptr() as *const _) }) +src/simd_avx512.rs 1283 FORK block src CONSOLIDATE present - this is the leaf store (to_array); keep as the single unsafe store(&mut [T;N]) helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_si512(arr.as_mut_ptr() as *mut _, self.0) }; +src/simd_avx512.rs 1290 FORK block src CONSOLIDATE missing - assert!(len) present; route via to_array/&mut [T;N] -> one unsafe store(&mut [T;N]) helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_si512(s.as_mut_ptr() as *mut _, self.0) }; +src/simd_avx512.rs 1297 FORK block src NEEDS-TF present feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { +src/simd_avx512.rs 1307 FORK block src NEEDS-TF present feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { +src/simd_avx512.rs 1317 FORK block src NEEDS-TF present feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic U8x64(unsafe { _mm512_packus_epi16(self.0, other.0) }) +src/simd_avx512.rs 1323 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { +src/simd_avx512.rs 1337 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { +src/simd_avx512.rs 1352 FORK block src NEEDS-TF present feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_mullo_epi16(self.0, other.0) }) +src/simd_avx512.rs 1367 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_add_epi16(self.0, rhs.0) }) +src/simd_avx512.rs 1374 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_sub_epi16(self.0, rhs.0) }) +src/simd_avx512.rs 1380 FORK block src NEEDS-TF missing feature: needs avx512bw but simd.rs gates mod on avx512f only (low: F-without-BW is KNL) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic self.0 = unsafe { _mm512_add_epi16(self.0, rhs.0) }; +src/simd_avx512.rs 1415 FORK block src NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm256_fmadd_ps(self.0, a.0, b.0) }) +src/simd_avx512.rs 1435 FORK block src NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm256_movemask_ps(_mm256_cmp_ps::<_CMP_GT_OQ>(self.0, other.0)) as u32 } +src/simd_avx512.rs 1612 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic Self(unsafe { _mm512_set1_epi32(v as i32) }) +src/simd_avx512.rs 1618 FORK block src CONSOLIDATE missing - assert!(len) present; route via from_array w/ <&[T;N]>::try_from(&s[..N]) -> one unsafe load(&[T;N]) helper x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_si512(s.as_ptr() as *const _) }) +src/simd_avx512.rs 1623 FORK block src CONSOLIDATE missing - this is the leaf load; keep as the single unsafe load(&[T;N]) helper (alt: safe_unaligned_simd) x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_si512(arr.as_ptr() as *const _) }) +src/simd_avx512.rs 1629 FORK block src CONSOLIDATE missing - this is the leaf store (to_array); keep as the single unsafe store(&mut [T;N]) helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_si512(arr.as_mut_ptr() as *mut _, self.0) }; +src/simd_avx512.rs 1636 FORK block src CONSOLIDATE missing - assert!(len) present; route via to_array/&mut [T;N] -> one unsafe store(&mut [T;N]) helper x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_si512(s.as_mut_ptr() as *mut _, self.0) }; +src/simd_avx512.rs 1641 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_reduce_add_epi32(self.0) as u32 } +src/simd_avx512.rs 1662 FORK block src NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsic, safe-callable under avx512f (Rust>=1.87); unverified: native clippy gave no unused_unsafe x86_intrinsic unsafe { _mm512_cmpeq_epu32_mask(self.0, other.0) } +src/simd_avx512.rs 1675 FORK block src IRREDUCIBLE present - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_rolv_epi32(self.0, _mm512_set1_epi32(n as i32)) }) +src/simd_avx512.rs 1693 FORK block src IRREDUCIBLE missing - 2 value intrinsics (set1,xor); target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/ x86_intrinsic unsafe { +src/simd_avx512.rs 1705 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_srlv_epi32(self.0, rhs.0) }) +src/simd_avx512.rs 1713 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_sllv_epi32(self.0, rhs.0) }) +src/simd_avx512.rs 1768 FORK block src IRREDUCIBLE present - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_rolv_epi64(self.0, _mm512_set1_epi64(n as i64)) }) +src/simd_avx512.rs 1784 FORK block src IRREDUCIBLE present - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_rorv_epi64(self.0, _mm512_set1_epi64(n as i64)) }) +src/simd_avx512.rs 1805 FORK block src IRREDUCIBLE present - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_mul_epu32(self.0, rhs.0) }) +src/simd_avx512.rs 1822 FORK block src IRREDUCIBLE present - 24 value shuffles in one block; target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod x86_intrinsic unsafe { +src/simd_avx512.rs 1863 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_set1_epi64(v as i64) }) +src/simd_avx512.rs 1869 FORK block src SAFE-CRATE missing - leaf load after assert!(len>=N): this IS the audited helper; safe_unaligned_simd::*loadu(&[T;N]) would remove unsafe x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_si512(s.as_ptr() as *const _) }) +src/simd_avx512.rs 1874 FORK block src CONSOLIDATE missing - from_array: delegate to Self::from_slice(&arr) (single audited load) or transmute via bytemuck x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_si512(arr.as_ptr() as *const _) }) +src/simd_avx512.rs 1880 FORK block src CONSOLIDATE missing - to_array: call self.copy_to_slice(&mut arr) (single audited store) x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_si512(arr.as_mut_ptr() as *mut _, self.0) }; +src/simd_avx512.rs 1887 FORK block src SAFE-CRATE missing - leaf store after assert!(len>=N): this IS the audited helper; safe_unaligned_simd::*storeu(&mut [T;N]) x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_si512(s.as_mut_ptr() as *mut _, self.0) }; +src/simd_avx512.rs 1892 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic unsafe { _mm512_reduce_add_epi64(self.0) as u64 } +src/simd_avx512.rs 1908 FORK block src IRREDUCIBLE present - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic unsafe { _mm512_cmpeq_epu64_mask(self.0, other.0) } +src/simd_avx512.rs 1925 FORK block src IRREDUCIBLE present - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic unsafe { _mm512_cmpgt_epu64_mask(self.0, other.0) } +src/simd_avx512.rs 1945 FORK block src IRREDUCIBLE missing - 2 value intrinsics (set1,xor); target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/ x86_intrinsic unsafe { +src/simd_avx512.rs 1958 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_srlv_epi64(self.0, rhs.0) }) +src/simd_avx512.rs 1967 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_sllv_epi64(self.0, rhs.0) }) +src/simd_avx512.rs 1997 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_set1_epi8(v) }) +src/simd_avx512.rs 2002 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_setzero_si512() }) +src/simd_avx512.rs 2008 FORK block src SAFE-CRATE missing - leaf load after assert!(len>=N): this IS the audited helper; safe_unaligned_simd::*loadu(&[T;N]) would remove unsafe x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_si512(s.as_ptr() as *const _) }) +src/simd_avx512.rs 2013 FORK block src CONSOLIDATE missing - from_array: delegate to Self::from_slice(&arr) (single audited load) or transmute via bytemuck x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_si512(arr.as_ptr() as *const _) }) +src/simd_avx512.rs 2019 FORK block src CONSOLIDATE missing - to_array: call self.copy_to_slice(&mut arr) (single audited store) x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_si512(arr.as_mut_ptr() as *mut _, self.0) }; +src/simd_avx512.rs 2026 FORK block src SAFE-CRATE missing - leaf store after assert!(len>=N): this IS the audited helper; safe_unaligned_simd::*storeu(&mut [T;N]) x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_si512(s.as_mut_ptr() as *mut _, self.0) }; +src/simd_avx512.rs 2031 FORK block src IRREDUCIBLE missing feature: avx512bw intrinsic; cfg gate is avx512f only (simd.rs:~243), SIGILL on F-without-BW CPU (KNL) target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_add_epi8(self.0, other.0) }) +src/simd_avx512.rs 2036 FORK block src IRREDUCIBLE missing feature: avx512bw intrinsic; cfg gate is avx512f only (simd.rs:~243), SIGILL on F-without-BW CPU (KNL) target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_sub_epi8(self.0, other.0) }) +src/simd_avx512.rs 2041 FORK block src IRREDUCIBLE missing feature: avx512bw intrinsic; cfg gate is avx512f only (simd.rs:~243), SIGILL on F-without-BW CPU (KNL) target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_min_epi8(self.0, other.0) }) +src/simd_avx512.rs 2046 FORK block src IRREDUCIBLE missing feature: avx512bw intrinsic; cfg gate is avx512f only (simd.rs:~243), SIGILL on F-without-BW CPU (KNL) target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_max_epi8(self.0, other.0) }) +src/simd_avx512.rs 2052 FORK block src IRREDUCIBLE missing feature: avx512bw intrinsic; cfg gate is avx512f only (simd.rs:~243), SIGILL on F-without-BW CPU (KNL) target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic unsafe { _mm512_cmpgt_epi8_mask(self.0, other.0) } +src/simd_avx512.rs 2086 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_set1_epi8(v) }) +src/simd_avx512.rs 2091 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_setzero_si256() }) +src/simd_avx512.rs 2097 FORK block src SAFE-CRATE missing - leaf load after assert!(len>=N): this IS the audited helper; safe_unaligned_simd::*loadu(&[T;N]) would remove unsafe x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm256_loadu_si256(s.as_ptr() as *const __m256i) }) +src/simd_avx512.rs 2102 FORK block src CONSOLIDATE missing - from_array: delegate to Self::from_slice(&arr) (single audited load) or transmute via bytemuck x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm256_loadu_si256(arr.as_ptr() as *const __m256i) }) +src/simd_avx512.rs 2108 FORK block src CONSOLIDATE missing - to_array: call self.copy_to_slice(&mut arr) (single audited store) x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm256_storeu_si256(arr.as_mut_ptr() as *mut __m256i, self.0) }; +src/simd_avx512.rs 2115 FORK block src SAFE-CRATE missing - leaf store after assert!(len>=N): this IS the audited helper; safe_unaligned_simd::*storeu(&mut [T;N]) x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm256_storeu_si256(s.as_mut_ptr() as *mut __m256i, self.0) }; +src/simd_avx512.rs 2120 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_add_epi8(self.0, other.0) }) +src/simd_avx512.rs 2125 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_sub_epi8(self.0, other.0) }) +src/simd_avx512.rs 2130 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_min_epi8(self.0, other.0) }) +src/simd_avx512.rs 2135 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_max_epi8(self.0, other.0) }) +src/simd_avx512.rs 2142 FORK block src IRREDUCIBLE missing - 2 value intrinsics (avx2 cmpgt+movemask); target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed x86_intrinsic unsafe { _mm256_movemask_epi8(_mm256_cmpgt_epi8(self.0, other.0)) as u32 } +src/simd_avx512.rs 2150 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_add_epi8(self.0, rhs.0) }) +src/simd_avx512.rs 2157 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_sub_epi8(self.0, rhs.0) }) +src/simd_avx512.rs 2163 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic self.0 = unsafe { _mm256_add_epi8(self.0, rhs.0) }; +src/simd_avx512.rs 2169 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic self.0 = unsafe { _mm256_sub_epi8(self.0, rhs.0) }; +src/simd_avx512.rs 2197 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_set1_epi16(v) }) +src/simd_avx512.rs 2202 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_setzero_si512() }) +src/simd_avx512.rs 2208 FORK block src SAFE-CRATE missing - leaf load after assert!(len>=N): this IS the audited helper; safe_unaligned_simd::*loadu(&[T;N]) would remove unsafe x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_si512(s.as_ptr() as *const _) }) +src/simd_avx512.rs 2213 FORK block src CONSOLIDATE missing - from_array: delegate to Self::from_slice(&arr) (single audited load) or transmute via bytemuck x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm512_loadu_si512(arr.as_ptr() as *const _) }) +src/simd_avx512.rs 2219 FORK block src CONSOLIDATE missing - to_array: call self.copy_to_slice(&mut arr) (single audited store) x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_si512(arr.as_mut_ptr() as *mut _, self.0) }; +src/simd_avx512.rs 2226 FORK block src SAFE-CRATE missing - leaf store after assert!(len>=N): this IS the audited helper; safe_unaligned_simd::*storeu(&mut [T;N]) x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm512_storeu_si512(s.as_mut_ptr() as *mut _, self.0) }; +src/simd_avx512.rs 2231 FORK block src IRREDUCIBLE missing feature: avx512bw intrinsic; cfg gate is avx512f only (simd.rs:~243), SIGILL on F-without-BW CPU (KNL) target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_add_epi16(self.0, other.0) }) +src/simd_avx512.rs 2236 FORK block src IRREDUCIBLE missing feature: avx512bw intrinsic; cfg gate is avx512f only (simd.rs:~243), SIGILL on F-without-BW CPU (KNL) target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_sub_epi16(self.0, other.0) }) +src/simd_avx512.rs 2241 FORK block src IRREDUCIBLE missing feature: avx512bw intrinsic; cfg gate is avx512f only (simd.rs:~243), SIGILL on F-without-BW CPU (KNL) target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_min_epi16(self.0, other.0) }) +src/simd_avx512.rs 2246 FORK block src IRREDUCIBLE missing feature: avx512bw intrinsic; cfg gate is avx512f only (simd.rs:~243), SIGILL on F-without-BW CPU (KNL) target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm512_max_epi16(self.0, other.0) }) +src/simd_avx512.rs 2252 FORK block src IRREDUCIBLE missing feature: avx512bw intrinsic; cfg gate is avx512f only (simd.rs:~243), SIGILL on F-without-BW CPU (KNL) target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic unsafe { _mm512_cmpgt_epi16_mask(self.0, other.0) } +src/simd_avx512.rs 2286 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_set1_epi16(v) }) +src/simd_avx512.rs 2291 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_setzero_si256() }) +src/simd_avx512.rs 2297 FORK block src SAFE-CRATE missing - leaf load after assert!(len>=N): this IS the audited helper; safe_unaligned_simd::*loadu(&[T;N]) would remove unsafe x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm256_loadu_si256(s.as_ptr() as *const __m256i) }) +src/simd_avx512.rs 2302 FORK block src CONSOLIDATE missing - from_array: delegate to Self::from_slice(&arr) (single audited load) or transmute via bytemuck x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm256_loadu_si256(arr.as_ptr() as *const __m256i) }) +src/simd_avx512.rs 2308 FORK block src CONSOLIDATE missing - to_array: call self.copy_to_slice(&mut arr) (single audited store) x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm256_storeu_si256(arr.as_mut_ptr() as *mut __m256i, self.0) }; +src/simd_avx512.rs 2315 FORK block src SAFE-CRATE missing - leaf store after assert!(len>=N): this IS the audited helper; safe_unaligned_simd::*storeu(&mut [T;N]) x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm256_storeu_si256(s.as_mut_ptr() as *mut __m256i, self.0) }; +src/simd_avx512.rs 2320 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_add_epi16(self.0, other.0) }) +src/simd_avx512.rs 2325 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_sub_epi16(self.0, other.0) }) +src/simd_avx512.rs 2330 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_min_epi16(self.0, other.0) }) +src/simd_avx512.rs 2335 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_max_epi16(self.0, other.0) }) +src/simd_avx512.rs 2342 FORK block src IRREDUCIBLE missing - 5 value avx2 intrinsics; target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type g x86_intrinsic unsafe { +src/simd_avx512.rs 2359 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_add_epi16(self.0, rhs.0) }) +src/simd_avx512.rs 2366 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_sub_epi16(self.0, rhs.0) }) +src/simd_avx512.rs 2372 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic self.0 = unsafe { _mm256_add_epi16(self.0, rhs.0) }; +src/simd_avx512.rs 2378 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic self.0 = unsafe { _mm256_sub_epi16(self.0, rhs.0) }; +src/simd_avx512.rs 2408 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_set1_ps(v) }) +src/simd_avx512.rs 2414 FORK block src SAFE-CRATE missing - leaf load after assert!(len>=N): this IS the audited helper; safe_unaligned_simd::*loadu(&[T;N]) would remove unsafe x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm256_loadu_ps(s.as_ptr()) }) +src/simd_avx512.rs 2419 FORK block src CONSOLIDATE missing - from_array: delegate to Self::from_slice(&arr) (single audited load) or transmute via bytemuck x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm256_loadu_ps(a.as_ptr()) }) +src/simd_avx512.rs 2425 FORK block src CONSOLIDATE missing - to_array: call self.copy_to_slice(&mut arr) (single audited store) x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm256_storeu_ps(out.as_mut_ptr(), self.0) }; +src/simd_avx512.rs 2432 FORK block src SAFE-CRATE missing - leaf store after assert!(len>=N): this IS the audited helper; safe_unaligned_simd::*storeu(&mut [T;N]) x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm256_storeu_ps(s.as_mut_ptr(), self.0) }; +src/simd_avx512.rs 2437 FORK block src IRREDUCIBLE missing - 8 value AVX/SSE intrinsics (reduce); target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unles x86_intrinsic unsafe { +src/simd_avx512.rs 2454 FORK block src IRREDUCIBLE missing - 3 value intrinsics (abs via mask); target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless x86_intrinsic unsafe { +src/simd_avx512.rs 2465 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_add_ps(self.0, rhs.0) }) +src/simd_avx512.rs 2472 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic self.0 = unsafe { _mm256_add_ps(self.0, rhs.0) }; +src/simd_avx512.rs 2480 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_mul_ps(self.0, rhs.0) }) +src/simd_avx512.rs 2487 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic self.0 = unsafe { _mm256_mul_ps(self.0, rhs.0) }; +src/simd_avx512.rs 2495 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_sub_ps(self.0, rhs.0) }) +src/simd_avx512.rs 2502 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic self.0 = unsafe { _mm256_sub_ps(self.0, rhs.0) }; +src/simd_avx512.rs 2510 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_div_ps(self.0, rhs.0) }) +src/simd_avx512.rs 2517 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic self.0 = unsafe { _mm256_div_ps(self.0, rhs.0) }; +src/simd_avx512.rs 2544 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_set1_pd(v) }) +src/simd_avx512.rs 2550 FORK block src SAFE-CRATE missing - leaf load after assert!(len>=N): this IS the audited helper; safe_unaligned_simd::*loadu(&[T;N]) would remove unsafe x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm256_loadu_pd(s.as_ptr()) }) +src/simd_avx512.rs 2555 FORK block src CONSOLIDATE missing - from_array: delegate to Self::from_slice(&arr) (single audited load) or transmute via bytemuck x86_intrinsic,raw_load_store,ptr_arith Self(unsafe { _mm256_loadu_pd(a.as_ptr()) }) +src/simd_avx512.rs 2561 FORK block src CONSOLIDATE missing - to_array: call self.copy_to_slice(&mut arr) (single audited store) x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm256_storeu_pd(out.as_mut_ptr(), self.0) }; +src/simd_avx512.rs 2568 FORK block src SAFE-CRATE missing - leaf store after assert!(len>=N): this IS the audited helper; safe_unaligned_simd::*storeu(&mut [T;N]) x86_intrinsic,raw_load_store,ptr_arith unsafe { _mm256_storeu_pd(s.as_mut_ptr(), self.0) }; +src/simd_avx512.rs 2573 FORK block src IRREDUCIBLE missing - 6 value intrinsics (reduce); target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/ty x86_intrinsic unsafe { +src/simd_avx512.rs 2585 FORK block src IRREDUCIBLE missing - 3 value intrinsics (abs via mask); target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless x86_intrinsic unsafe { +src/simd_avx512.rs 2596 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_add_pd(self.0, rhs.0) }) +src/simd_avx512.rs 2603 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic self.0 = unsafe { _mm256_add_pd(self.0, rhs.0) }; +src/simd_avx512.rs 2611 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_mul_pd(self.0, rhs.0) }) +src/simd_avx512.rs 2618 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic self.0 = unsafe { _mm256_mul_pd(self.0, rhs.0) }; +src/simd_avx512.rs 2626 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_sub_pd(self.0, rhs.0) }) +src/simd_avx512.rs 2633 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic self.0 = unsafe { _mm256_sub_pd(self.0, rhs.0) }; +src/simd_avx512.rs 2641 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic Self(unsafe { _mm256_div_pd(self.0, rhs.0) }) +src/simd_avx512.rs 2648 FORK block src IRREDUCIBLE missing - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature x86_intrinsic self.0 = unsafe { _mm256_div_pd(self.0, rhs.0) }; +src/simd_avx512.rs 2858 FORK block src SAFE-API missing - self.0.map(i8::saturating_abs) on [i8;16]; LLVM should autovec to vpabsb+vpminub (not verified); else CONSOLIDATE via I8x16 load/store helpe x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_avx512.rs 3078 FORK block src IRREDUCIBLE present - 3 value avx2 intrinsics; target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type g x86_intrinsic unsafe { +src/simd_avx512.rs 3129 FORK block src IRREDUCIBLE present - value intrinsic _mm512_popcnt_epi64 under cfg(target_feature=avx512vpopcntdq); REMOVABLE only if global features make it safe-callable (unve x86_intrinsic unsafe { +src/simd_avx512.rs 3227 FORK block src IRREDUCIBLE present - _mm_prefetch on raw ptr (hint, never derefs); no safe std prefetch on stable x86_intrinsic,ptr_arith unsafe { +src/simd_avx512.rs 3238 FORK block src IRREDUCIBLE present - _mm_prefetch on raw ptr (hint, never derefs); no safe std prefetch on stable x86_intrinsic,ptr_arith unsafe { +src/simd_avx512.rs 3249 FORK block src IRREDUCIBLE present - _mm_prefetch on raw ptr (hint, never derefs); no safe std prefetch on stable x86_intrinsic,ptr_arith unsafe { +src/simd_avx512.rs 3341 FORK fn src UNSAFE-FN-API present - only precondition is avx512bf16; assert!(len>=16) already covers slice. Make safe #[target_feature] fn (target_feature_11) + <&[u16;16]>::tr x86_intrinsic,raw_load_store,ptr_arith,transmute pub unsafe fn from_u16_slice(s: &[u16]) -> Self { +src/simd_avx512.rs 3354 FORK fn src UNSAFE-FN-API present - only precondition is avx512bf16+f; as safe #[target_feature(enable=..)] fn the body needs no unsafe x86_intrinsic pub unsafe fn to_f32x16(self) -> F32x16 { +src/simd_avx512.rs 3372 FORK fn src UNSAFE-FN-API missing - only precondition is avx512bf16; assert covers len>=8. Make safe #[target_feature] fn; doc lacks # Safety x86_intrinsic,raw_load_store,ptr_arith,transmute pub unsafe fn from_u16_slice(s: &[u16]) -> Self { +src/simd_avx512.rs 3381 FORK fn src UNSAFE-FN-API missing - only precondition is avx512bf16+vl; make safe #[target_feature] fn (body then unsafe-free); lacks # Safety doc x86_intrinsic pub unsafe fn to_f32x8(self) -> F32x8 { +src/simd_avx512.rs 3392 FORK fn src UNSAFE-FN-API missing - only precondition is avx512bf16+f; make safe #[target_feature] fn; lacks # Safety doc x86_intrinsic pub unsafe fn to_bf16x16(self) -> BF16x16 { +src/simd_avx512.rs 3403 FORK fn src UNSAFE-FN-API missing - only precondition is avx512bf16+vl; make safe #[target_feature] fn; lacks # Safety doc x86_intrinsic pub unsafe fn to_bf16x8(self) -> BF16x8 { +src/simd_avx512.rs 3439 FORK block src IRREDUCIBLE present - calls #[target_feature(avx512bf16,vl)] fn after is_x86_feature_detected! check target_feature_call unsafe { +src/simd_avx512.rs 3455 FORK block src IRREDUCIBLE present - calls #[target_feature(avx512f)] fn after is_x86_feature_detected!(avx512f) target_feature_call unsafe { +src/simd_avx512.rs 3477 FORK fn src UNSAFE-FN-API missing bounds: output.len()>=input.len() not checked in fn (writes via ptr.add(i)); only callers assert precondition: avx512f + output.len()>=input.len(); latter unchecked inside. Replace add()/get_unchecked with as_chunks + zip + leaf load/sto x86_intrinsic,raw_load_store,ptr_arith,unchecked,target_feature_call unsafe fn convert_bf16_to_f32_avx512f(input: &[u16], output: &mut [f32]) { +src/simd_avx512.rs 3505 FORK block src IRREDUCIBLE missing - calls #[target_feature(avx512bf16,vl)] fn after is_x86_feature_detected! check; add SAFETY comment target_feature_call unsafe { +src/simd_avx512.rs 3522 FORK fn src UNSAFE-FN-API missing - precondition avx512bf16+vl only; slice access is checked (as_chunks, [..len]); transmutes of [T;16]<->vec -> bytemuck/safe target_feature fn x86_intrinsic,transmute,target_feature_call unsafe fn convert_bf16_to_f32_avx512bf16(input: &[u16], output: &mut [f32]) { +src/simd_avx512.rs 3553 FORK fn src UNSAFE-FN-API missing - precondition avx512bf16+vl only; slice access checked (as_chunks); make safe #[target_feature] fn; transmute -> bytemuck x86_intrinsic,transmute,target_feature_call unsafe fn convert_f32_to_bf16_avx512bf16(input: &[f32], output: &mut [u16]) { +src/simd_avx512.rs 3626 FORK fn src UNSAFE-FN-API present - only precondition is avx512f; pure register ops. Make safe #[target_feature(enable="avx512f")] fn; body then unsafe-free x86_intrinsic pub unsafe fn f32_to_bf16_x16_rne(lane: __m512) -> __m256i { +src/simd_avx512.rs 3686 FORK block src IRREDUCIBLE present - calls #[target_feature(avx512f)] fn after is_x86_feature_detected!(avx512f) target_feature_call unsafe { +src/simd_avx512.rs 3700 FORK fn src UNSAFE-FN-API present bounds: output.len()>=input.len() not checked in fn (storeu at output.add(i)); only caller asserts precondition avx512f + output.len()>=input.len(); latter unchecked in fn (ptr.add stores). Use as_chunks/chunks_exact + safe leaf load/store x86_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn convert_f32_to_bf16_avx512f_rne(input: &[f32], output: &mut [u16]) { +src/simd_avx512.rs 3821 FORK block src IRREDUCIBLE present - test: calls #[target_feature(avx512f)] fn after is_x86_feature_detected! check target_feature_call unsafe { convert_bf16_to_f32_avx512f(&input, &mut output) }; +src/simd_avx512.rs 3953 FORK block src CONSOLIDATE present - test loop: loadu_ps/storeu_si256 on ptr.add(i); use chunks_exact(16) + safe load/store helper x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_avx512.rs 3968 FORK block src IRREDUCIBLE present - test: calls #[target_feature(avx512bf16,vl)] fn after is_x86_feature_detected! check target_feature_call unsafe { +src/simd_avx512.rs 4040 FORK block src CONSOLIDATE present - loadu_si256/cvtph_ps/storeu_ps on input[c*16..]; bounds ok (c + helper fn(&[u16;16],&mut [f32;16]) x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_avx512.rs 4083 FORK block src CONSOLIDATE present - loadu_ps/cvtps_ph/storeu_si256 on slices; bounds ok; use as_chunks + helper fn(&[f32;16],&mut [u16;16]) x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_avx512.rs 4345 FORK block src CONSOLIDATE present - loadu_si256/cvtph_ps/storeu_ps on input[c*16..]; bounds ok (c + helper fn(&[u16;16],&mut [f32;16]) x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_avx512.rs 4362 FORK block src CONSOLIDATE present - f16c: loadu_si128/cvtph_ps/storeu_ps on 8-lane chunks; bounds ok; as_chunks::<8> + helper x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_avx512.rs 4395 FORK block src CONSOLIDATE present - loadu_ps/cvtps_ph/storeu_si256 on slices; bounds ok; use as_chunks + helper fn(&[f32;16],&mut [u16;16]) x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_avx512.rs 4412 FORK block src CONSOLIDATE present - f16c: loadu_ps/cvtps_ph/storeu_si128 on 8-lane chunks; bounds ok; as_chunks::<8> + helper x86_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_avx512.rs 5042 FORK block src IRREDUCIBLE present - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature; arg swap x86_intrinsic U64x8(unsafe { _mm512_andnot_si512(other.0, self.0) }) +src/simd_avx512.rs 5069 FORK block src IRREDUCIBLE present - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature; IMM cons x86_intrinsic U64x8(unsafe { _mm512_ternarylogic_epi64::(self.0, b.0, c.0) }) +src/simd_avx512.rs 5089 FORK block src IRREDUCIBLE present - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature; arg swap x86_intrinsic U32x16(unsafe { _mm512_andnot_si512(other.0, self.0) }) +src/simd_avx512.rs 5107 FORK block src IRREDUCIBLE present - target-feature call; module compiled on ALL x86_64 (lib.rs:243 not cfg(avx512f)) so unsafe needed unless mod/type gated on feature; IMM cons x86_intrinsic U32x16(unsafe { _mm512_ternarylogic_epi32::(self.0, b.0, c.0) }) +src/simd_half.rs 369 FORK block src SAFE-CRATE present - bytemuck::cast_slice/_mut (F16 is repr(transparent) u16; needs Pod derive) or have kernel take &[F16] and read .0 ptr_arith,from_raw_parts let src_u16: &[u16] = unsafe { core::slice::from_raw_parts(src.as_ptr() as *const u16, src.len()) }; +src/simd_half.rs 370 FORK block src IRREDUCIBLE missing - call #[target_feature(f16c,avx)] fn after is_x86_feature_detected! f16c+avx - unsafe { +src/simd_half.rs 408 FORK block src SAFE-CRATE present - bytemuck::cast_slice/_mut (F16 is repr(transparent) u16; needs Pod derive) or have kernel take &[F16] and read .0 ptr_arith,from_raw_parts unsafe { core::slice::from_raw_parts_mut(dst.as_mut_ptr() as *mut u16, dst.len()) }; +src/simd_half.rs 409 FORK block src IRREDUCIBLE missing - call #[target_feature(f16c,avx)] fn after is_x86_feature_detected! f16c+avx - unsafe { +src/simd_half.rs 442 FORK fn src IRREDUCIBLE present - asm! stmxcsr/ldmxcsr (MXCSR save/restore) + unaligned ptr loads/stores in target_feature fn; loads could use chunks_exact+helper x86_intrinsic,raw_load_store,ptr_arith,asm unsafe fn cast_f16_to_f32_batch_f16c(src: &[u16], dst: &mut [f32]) { +src/simd_half.rs 498 FORK fn src IRREDUCIBLE present - asm! stmxcsr/ldmxcsr (MXCSR save/restore) + unaligned ptr loads/stores in target_feature fn; loads could use chunks_exact+helper x86_intrinsic,raw_load_store,ptr_arith,asm unsafe fn cast_f32_to_f16_batch_f16c(src: &[f32], dst: &mut [u16]) { +src/simd_half.rs 936 FORK block test IRREDUCIBLE missing - asm! stmxcsr in test to snapshot MXCSR (_mm_getcsr is deprecated and also unsafe) asm unsafe { +src/simd_half.rs 944 FORK block test IRREDUCIBLE missing - asm! stmxcsr in test to snapshot MXCSR (_mm_getcsr is deprecated and also unsafe) asm unsafe { +src/simd_half.rs 959 FORK block test IRREDUCIBLE missing - asm! stmxcsr in test to snapshot MXCSR (_mm_getcsr is deprecated and also unsafe) asm unsafe { +src/simd_half.rs 963 FORK block test IRREDUCIBLE missing - asm! stmxcsr in test to snapshot MXCSR (_mm_getcsr is deprecated and also unsafe) asm unsafe { +src/simd_int_ops.rs 315 FORK block src IRREDUCIBLE present feature: detects only avx512vnni but kernel enables avx512f+avx512bw too; comment assumes implication (true on shipped HW, not checked) call #[target_feature(avx512f,avx512vnni,avx512bw)] fn after is_x86_feature_detected!(avx512vnni) target_feature_call unsafe { crate::hpc::vnni_gemm::int8_gemm_vnni_avx512(a, b, c, m, n, k) }; +src/simd_int_ops.rs 320 FORK block src IRREDUCIBLE present feature: detects only avxvnni; kernel also enables avx2 (implied on shipped HW, not checked) call #[target_feature(avx,avx2,avxvnni)] fn after is_x86_feature_detected!(avxvnni) - unsafe { crate::hpc::vnni_gemm::int8_gemm_avxvnni_ymm(a, b, c, m, n, k) }; +src/simd_neon.rs 36 FORK fn src CONSOLIDATE missing - make safe fn over &[f32;4]; 2 loads via one audited ld(&[f32;4]) helper (or safe_unaligned_simd); rest is value ops neon_intrinsic,raw_load_store,ptr_arith,target_feature_call pub unsafe fn dot_f32x4_neon(a: &[f32; 4], b: &[f32; 4]) -> f32 { +src/simd_neon.rs 49 FORK fn src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] unsafe fn wraps only vfmaq_f32 (value intrinsic): make it a safe fn (unverified, not compiled on host) neon_intrinsic,target_feature_call pub unsafe fn fma_f32x4_neon(acc: float32x4_t, a: float32x4_t, b: float32x4_t) -> float32x4_t { +src/simd_neon.rs 57 FORK fn src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] unsafe fn wraps only vpaddq/vgetq_lane (value intrinsics): make it a safe fn (unverified) neon_intrinsic pub unsafe fn hsum_f32x4(v: float32x4_t) -> f32 { +src/simd_neon.rs 66 FORK fn src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] unsafe fn wraps only vcntq_u8 (value intrinsic): make it a safe fn (unverified) neon_intrinsic pub unsafe fn popcount_u8x16(data: uint8x16_t) -> uint8x16_t { +src/simd_neon.rs 74 FORK fn src CONSOLIDATE missing - safe fn over &[u8;16]; loads via shared U8x16::from_array helper, rest value ops neon_intrinsic,raw_load_store,ptr_arith pub unsafe fn hamming_u8x16(a: &[u8; 16], b: &[u8; 16]) -> u32 { +src/simd_neon.rs 90 FORK fn src CONSOLIDATE missing correctness: vabdq_s16 yields int16 so |a-b|>32767 wraps negative before vpaddlq_s16; use u16 reinterpret + vpaddlq_u16 safe fn; loads of a[..8]/a[8..16] via one &[i16;8] helper (try_into), tail already scalar neon_intrinsic,raw_load_store,ptr_arith,target_feature_call pub unsafe fn base17_l1_neon(a: &[i16; 17], b: &[i16; 17]) -> i32 { +src/simd_neon.rs 115 FORK fn src UNSAFE-FN-API missing bounds: centroids[offset..] only checks offset<=len but vld1q reads 4 f32; output.len()>=dim only debug_assert, vst1q may write OOB in release precondition checkable: slice centroids[off..off+4] / output[c*4..c*4+4] to &[f32;4] then it is a safe fn target_feature_call pub unsafe fn codebook_gather_f32x4_neon( +src/simd_neon.rs 147 FORK fn src UNSAFE-FN-API missing bounds: centroids[idx*dim+c*4..] only checks start<=len, vld1q reads 4 f32; output.len() only debug_assert same as 115: slice to &[f32;4] (off..off+4) and assert output.len()>=dim in release -> safe fn neon_intrinsic,raw_load_store,ptr_arith pub unsafe fn codebook_gather_f32x4_a72(centroids: &[f32], indices: &[u8], dim: usize, output: &mut [f32]) { +src/simd_neon.rs 196 FORK fn src CONSOLIDATE missing - safe fn over &[i8;16]; loads via I8x16::from_array helper, rest value ops (vmull/vpaddl/vaddv) neon_intrinsic,raw_load_store,ptr_arith,target_feature_call pub unsafe fn dot_i8x16_neon(a: &[i8; 16], b: &[i8; 16]) -> i32 { +src/simd_neon.rs 214 FORK fn src UNSAFE-FN-API missing bounds: centroids_i8[base..] only checks start<=len, vld1q reads 16 i8; output_i32.len()>=dim only debug_assert precondition checkable: slice [base..base+16] and out [c*16..+16] to arrays, assert lens in release -> safe fn; name says dotprod but none u - pub unsafe fn codebook_gather_i8_dotprod( +src/simd_neon.rs 261 FORK fn src IRREDUCIBLE missing feature: doc says ARMv8.2 fp16 required but only f16/f32_batch check it; direct pub callers unguarded (FCVTL/FCVTN believed baseline ASIMD, unverified) asm! FCVTL/FCVTN (vcvt_f32_f16 needs unstable f16 type); safe alt f16_to_f32_scalar loop or `half` crate; wrapper itself could be a safe fn wasm_intrinsic,ptr_arith,asm pub unsafe fn f16x4_to_f32x4(input: &[u16; 4]) -> [f32; 4] { +src/simd_neon.rs 279 FORK fn src IRREDUCIBLE missing feature: fp16 guard lives only in batch callers (see 261) asm! FCVTL/FCVTN (vcvt_f32_f16 needs unstable f16 type); safe alt f16_to_f32_scalar loop or `half` crate; wrapper itself could be a safe fn wasm_intrinsic,ptr_arith,asm pub unsafe fn f16x8_to_f32x8(input: &[u16; 8]) -> [f32; 8] { +src/simd_neon.rs 300 FORK fn src IRREDUCIBLE missing feature: fp16 guard lives only in batch callers (see 261) asm! FCVTL/FCVTN (vcvt_f32_f16 needs unstable f16 type); safe alt f16_to_f32_scalar loop or `half` crate; wrapper itself could be a safe fn wasm_intrinsic,ptr_arith,asm pub unsafe fn f32x4_to_f16x4(input: &[f32; 4]) -> [u16; 4] { +src/simd_neon.rs 317 FORK fn src IRREDUCIBLE missing feature: fp16 guard lives only in batch callers (see 261) asm! FCVTL/FCVTN (vcvt_f32_f16 needs unstable f16 type); safe alt f16_to_f32_scalar loop or `half` crate; wrapper itself could be a safe fn wasm_intrinsic,ptr_arith,asm pub unsafe fn f32x8_to_f16x8(input: &[f32; 8]) -> [u16; 8] { +src/simd_neon.rs 415 FORK block src CONSOLIDATE missing - make f16x4_to_f32x4 a safe fn (asm inside, args are &[;4]); this call then needs no unsafe wasm_intrinsic let dst = unsafe { f16x4_to_f32x4(src) }; +src/simd_neon.rs 442 FORK block src CONSOLIDATE missing - make f32x4_to_f16x4 a safe fn (asm inside); this call then needs no unsafe wasm_intrinsic let dst = unsafe { f32x4_to_f16x4(src) }; +src/simd_neon.rs 503 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 512 FORK block src CONSOLIDATE missing - delegate to from_array(s[..N].try_into().unwrap()); one audited vld1q (&[T;N]) site (or safe_unaligned_simd) neon_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_neon.rs 533 FORK block src CONSOLIDATE missing - delegate: s[..N].copy_from_slice(&self.to_array()) so only to_array holds the vst1q neon_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_neon.rs 544 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 569 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { Self([vabsq_f32(self.0[0]), vabsq_f32(self.0[1]), vabsq_f32(self.0[2]), vabsq_f32(self.0[3])]) } +src/simd_neon.rs 574 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) - unsafe { +src/simd_neon.rs 581 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) - unsafe { +src/simd_neon.rs 588 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) - unsafe { +src/simd_neon.rs 595 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 607 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 619 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 724 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 738 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 752 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 766 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) - unsafe { +src/simd_neon.rs 804 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { Self([vnegq_f32(self.0[0]), vnegq_f32(self.0[1]), vnegq_f32(self.0[2]), vnegq_f32(self.0[3])]) } +src/simd_neon.rs 869 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 878 FORK block src CONSOLIDATE missing - delegate to from_array(s[..N].try_into().unwrap()); one audited vld1q (&[T;N]) site (or safe_unaligned_simd) neon_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_neon.rs 899 FORK block src CONSOLIDATE missing - delegate: s[..N].copy_from_slice(&self.to_array()) so only to_array holds the vst1q neon_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_neon.rs 910 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 935 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { Self([vabsq_f64(self.0[0]), vabsq_f64(self.0[1]), vabsq_f64(self.0[2]), vabsq_f64(self.0[3])]) } +src/simd_neon.rs 940 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) - unsafe { +src/simd_neon.rs 947 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) - unsafe { +src/simd_neon.rs 954 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) - unsafe { +src/simd_neon.rs 961 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 973 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 985 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 1101 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 1115 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 1129 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 1143 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) - unsafe { +src/simd_neon.rs 1181 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { Self([vnegq_f64(self.0[0]), vnegq_f64(self.0[1]), vnegq_f64(self.0[2]), vnegq_f64(self.0[3])]) } +src/simd_neon.rs 1307 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vdupq_n_s8(v) }) +src/simd_neon.rs 1312 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vdupq_n_s8(0) }) +src/simd_neon.rs 1318 FORK block src CONSOLIDATE missing - delegate to from_array(s[..N].try_into().unwrap()); one audited vld1q (&[T;N]) site (or safe_unaligned_simd) neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { vld1q_s8(s.as_ptr()) }) +src/simd_neon.rs 1323 FORK block src CONSOLIDATE missing - this IS the load helper (arr by value, valid N elems); make it the single audited vld1q site or use safe_unaligned_simd neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { vld1q_s8(arr.as_ptr()) }) +src/simd_neon.rs 1329 FORK block src CONSOLIDATE missing - this IS the store helper (local array, N elems); make it the single audited vst1q site or use safe_unaligned_simd neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1q_s8(arr.as_mut_ptr(), self.0) }; +src/simd_neon.rs 1336 FORK block src CONSOLIDATE missing - delegate: s[..N].copy_from_slice(&self.to_array()) so only to_array holds the vst1q neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1q_s8(s.as_mut_ptr(), self.0) }; +src/simd_neon.rs 1341 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vaddq_s8(self.0, other.0) }) +src/simd_neon.rs 1345 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vsubq_s8(self.0, other.0) }) +src/simd_neon.rs 1349 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vminq_s8(self.0, other.0) }) +src/simd_neon.rs 1353 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vmaxq_s8(self.0, other.0) }) +src/simd_neon.rs 1359 FORK block src CONSOLIDATE missing - transmute of cmp vector to array: reuse U8x16/U16x8(cmp).to_array() (or bytemuck); vcgtq itself is a value op neon_intrinsic,transmute unsafe { +src/simd_neon.rs 1397 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vdupq_n_s16(v) }) +src/simd_neon.rs 1402 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vdupq_n_s16(0) }) +src/simd_neon.rs 1408 FORK block src CONSOLIDATE missing - delegate to from_array(s[..N].try_into().unwrap()); one audited vld1q (&[T;N]) site (or safe_unaligned_simd) neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { vld1q_s16(s.as_ptr()) }) +src/simd_neon.rs 1413 FORK block src CONSOLIDATE missing - this IS the load helper (arr by value, valid N elems); make it the single audited vld1q site or use safe_unaligned_simd neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { vld1q_s16(arr.as_ptr()) }) +src/simd_neon.rs 1419 FORK block src CONSOLIDATE missing - this IS the store helper (local array, N elems); make it the single audited vst1q site or use safe_unaligned_simd neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1q_s16(arr.as_mut_ptr(), self.0) }; +src/simd_neon.rs 1426 FORK block src CONSOLIDATE missing - delegate: s[..N].copy_from_slice(&self.to_array()) so only to_array holds the vst1q neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1q_s16(s.as_mut_ptr(), self.0) }; +src/simd_neon.rs 1431 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vaddq_s16(self.0, other.0) }) +src/simd_neon.rs 1435 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vsubq_s16(self.0, other.0) }) +src/simd_neon.rs 1439 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vminq_s16(self.0, other.0) }) +src/simd_neon.rs 1443 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vmaxq_s16(self.0, other.0) }) +src/simd_neon.rs 1449 FORK block src CONSOLIDATE missing - transmute of cmp vector to array: reuse U8x16/U16x8(cmp).to_array() (or bytemuck); vcgtq itself is a value op neon_intrinsic,transmute unsafe { +src/simd_neon.rs 1491 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vdupq_n_u8(v) }) +src/simd_neon.rs 1495 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vdupq_n_u8(0) }) +src/simd_neon.rs 1500 FORK block src CONSOLIDATE missing - delegate to from_array(s[..N].try_into().unwrap()); one audited vld1q (&[T;N]) site (or safe_unaligned_simd) neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { vld1q_u8(s.as_ptr()) }) +src/simd_neon.rs 1504 FORK block src CONSOLIDATE missing - this IS the load helper (arr by value, valid N elems); make it the single audited vld1q site or use safe_unaligned_simd neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { vld1q_u8(arr.as_ptr()) }) +src/simd_neon.rs 1509 FORK block src CONSOLIDATE missing - this IS the store helper (local array, N elems); make it the single audited vst1q site or use safe_unaligned_simd neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1q_u8(arr.as_mut_ptr(), self.0) }; +src/simd_neon.rs 1515 FORK block src CONSOLIDATE missing - delegate: s[..N].copy_from_slice(&self.to_array()) so only to_array holds the vst1q neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1q_u8(s.as_mut_ptr(), self.0) }; +src/simd_neon.rs 1519 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vaddq_u8(self.0, other.0) }) +src/simd_neon.rs 1523 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vsubq_u8(self.0, other.0) }) +src/simd_neon.rs 1527 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vminq_u8(self.0, other.0) }) +src/simd_neon.rs 1531 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vmaxq_u8(self.0, other.0) }) +src/simd_neon.rs 1545 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vdupq_n_u16(v) }) +src/simd_neon.rs 1549 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vdupq_n_u16(0) }) +src/simd_neon.rs 1554 FORK block src CONSOLIDATE missing - delegate to from_array(s[..N].try_into().unwrap()); one audited vld1q (&[T;N]) site (or safe_unaligned_simd) neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { vld1q_u16(s.as_ptr()) }) +src/simd_neon.rs 1558 FORK block src CONSOLIDATE missing - this IS the load helper (arr by value, valid N elems); make it the single audited vld1q site or use safe_unaligned_simd neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { vld1q_u16(arr.as_ptr()) }) +src/simd_neon.rs 1563 FORK block src CONSOLIDATE missing - this IS the store helper (local array, N elems); make it the single audited vst1q site or use safe_unaligned_simd neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1q_u16(arr.as_mut_ptr(), self.0) }; +src/simd_neon.rs 1569 FORK block src CONSOLIDATE missing - delegate: s[..N].copy_from_slice(&self.to_array()) so only to_array holds the vst1q neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1q_u16(s.as_mut_ptr(), self.0) }; +src/simd_neon.rs 1573 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vaddq_u16(self.0, other.0) }) +src/simd_neon.rs 1577 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vsubq_u16(self.0, other.0) }) +src/simd_neon.rs 1581 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vminq_u16(self.0, other.0) }) +src/simd_neon.rs 1585 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vmaxq_u16(self.0, other.0) }) +src/simd_neon.rs 1604 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vdupq_n_u32(v) }) +src/simd_neon.rs 1608 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vdupq_n_u32(0) }) +src/simd_neon.rs 1613 FORK block src CONSOLIDATE missing - delegate to from_array(s[..N].try_into().unwrap()); one audited vld1q (&[T;N]) site (or safe_unaligned_simd) neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { vld1q_u32(s.as_ptr()) }) +src/simd_neon.rs 1617 FORK block src CONSOLIDATE missing - this IS the load helper (arr by value, valid N elems); make it the single audited vld1q site or use safe_unaligned_simd neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { vld1q_u32(arr.as_ptr()) }) +src/simd_neon.rs 1622 FORK block src CONSOLIDATE missing - this IS the store helper (local array, N elems); make it the single audited vst1q site or use safe_unaligned_simd neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1q_u32(arr.as_mut_ptr(), self.0) }; +src/simd_neon.rs 1628 FORK block src CONSOLIDATE missing - delegate: s[..N].copy_from_slice(&self.to_array()) so only to_array holds the vst1q neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1q_u32(s.as_mut_ptr(), self.0) }; +src/simd_neon.rs 1632 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vaddq_u32(self.0, other.0) }) +src/simd_neon.rs 1636 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vsubq_u32(self.0, other.0) }) +src/simd_neon.rs 1640 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vminq_u32(self.0, other.0) }) +src/simd_neon.rs 1644 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vmaxq_u32(self.0, other.0) }) +src/simd_neon.rs 1649 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { veorq_u32(self.0, other.0) }) +src/simd_neon.rs 1663 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 1883 FORK block src NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic let q: [u32; 4] = unsafe { +src/simd_neon.rs 1991 FORK block src NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic U32x4(unsafe { +src/simd_neon.rs 2065 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vdupq_n_u64(v) }) +src/simd_neon.rs 2069 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vdupq_n_u64(0) }) +src/simd_neon.rs 2074 FORK block src CONSOLIDATE missing - delegate to from_array(s[..N].try_into().unwrap()); one audited vld1q (&[T;N]) site (or safe_unaligned_simd) neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { vld1q_u64(s.as_ptr()) }) +src/simd_neon.rs 2078 FORK block src CONSOLIDATE missing - this IS the load helper (arr by value, valid N elems); make it the single audited vld1q site or use safe_unaligned_simd neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { vld1q_u64(arr.as_ptr()) }) +src/simd_neon.rs 2083 FORK block src CONSOLIDATE missing - this IS the store helper (local array, N elems); make it the single audited vst1q site or use safe_unaligned_simd neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1q_u64(arr.as_mut_ptr(), self.0) }; +src/simd_neon.rs 2089 FORK block src CONSOLIDATE missing - delegate: s[..N].copy_from_slice(&self.to_array()) so only to_array holds the vst1q neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1q_u64(s.as_mut_ptr(), self.0) }; +src/simd_neon.rs 2093 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vaddq_u64(self.0, other.0) }) +src/simd_neon.rs 2097 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vsubq_u64(self.0, other.0) }) +src/simd_neon.rs 2139 FORK block src NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 2158 FORK block src NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 2175 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vdupq_n_s32(v) }) +src/simd_neon.rs 2179 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vdupq_n_s32(0) }) +src/simd_neon.rs 2184 FORK block src CONSOLIDATE missing - delegate to from_array(s[..N].try_into().unwrap()); one audited vld1q (&[T;N]) site (or safe_unaligned_simd) neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { vld1q_s32(s.as_ptr()) }) +src/simd_neon.rs 2188 FORK block src CONSOLIDATE missing - this IS the load helper (arr by value, valid N elems); make it the single audited vld1q site or use safe_unaligned_simd neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { vld1q_s32(arr.as_ptr()) }) +src/simd_neon.rs 2193 FORK block src CONSOLIDATE missing - this IS the store helper (local array, N elems); make it the single audited vst1q site or use safe_unaligned_simd neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1q_s32(arr.as_mut_ptr(), self.0) }; +src/simd_neon.rs 2199 FORK block src CONSOLIDATE missing - delegate: s[..N].copy_from_slice(&self.to_array()) so only to_array holds the vst1q neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1q_s32(s.as_mut_ptr(), self.0) }; +src/simd_neon.rs 2203 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vaddq_s32(self.0, other.0) }) +src/simd_neon.rs 2207 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vsubq_s32(self.0, other.0) }) +src/simd_neon.rs 2211 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vminq_s32(self.0, other.0) }) +src/simd_neon.rs 2215 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vmaxq_s32(self.0, other.0) }) +src/simd_neon.rs 2229 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vdupq_n_s64(v) }) +src/simd_neon.rs 2233 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vdupq_n_s64(0) }) +src/simd_neon.rs 2238 FORK block src CONSOLIDATE missing - delegate to from_array(s[..N].try_into().unwrap()); one audited vld1q (&[T;N]) site (or safe_unaligned_simd) neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { vld1q_s64(s.as_ptr()) }) +src/simd_neon.rs 2242 FORK block src CONSOLIDATE missing - this IS the load helper (arr by value, valid N elems); make it the single audited vld1q site or use safe_unaligned_simd neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { vld1q_s64(arr.as_ptr()) }) +src/simd_neon.rs 2247 FORK block src CONSOLIDATE missing - this IS the store helper (local array, N elems); make it the single audited vst1q site or use safe_unaligned_simd neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1q_s64(arr.as_mut_ptr(), self.0) }; +src/simd_neon.rs 2253 FORK block src CONSOLIDATE missing - delegate: s[..N].copy_from_slice(&self.to_array()) so only to_array holds the vst1q neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1q_s64(s.as_mut_ptr(), self.0) }; +src/simd_neon.rs 2257 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vaddq_s64(self.0, other.0) }) +src/simd_neon.rs 2261 FORK block src NEEDS-TF missing - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(unsafe { vsubq_s64(self.0, other.0) }) +src/simd_neon.rs 2431 FORK block src CONSOLIDATE present - use Self::from_array(lanes) (existing helper); the 'aligned' remark in the comment is irrelevant to vld1q neon_intrinsic,raw_load_store,ptr_arith Self(unsafe { core::arch::aarch64::vld1q_s8(lanes.as_ptr()) }) +src/simd_neon.rs 2459 FORK block src NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) - Self(unsafe { core::arch::aarch64::vqabsq_s8(self.0) }) +src/simd_neon.rs 2605 FORK block src IRREDUCIBLE present - asm! prfm hint; core::arch::aarch64::_prefetch is unstable here; fn is already safe (ptr never dereferenced) asm unsafe { +src/simd_neon.rs 2620 FORK block src IRREDUCIBLE present - asm! prfm hint; core::arch::aarch64::_prefetch is unstable here; fn is already safe (ptr never dereferenced) asm unsafe { +src/simd_neon.rs 2635 FORK block src IRREDUCIBLE present - asm! prfm hint; core::arch::aarch64::_prefetch is unstable here; fn is already safe (ptr never dereferenced) asm unsafe { +src/simd_neon.rs 2843 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic let q: [u64; 4] = unsafe { +src/simd_neon.rs 2861 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(core::array::from_fn(|p| unsafe { +src/simd_neon.rs 2889 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(core::array::from_fn(|p| unsafe { U64x2(vmull_u32(vmovn_u64(self.0[p].0), vmovn_u64(rhs.0[p].0))) })) +src/simd_neon.rs 2913 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) - U64x2(unsafe { +src/simd_neon.rs 2930 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(core::array::from_fn(|p| unsafe { +src/simd_neon.rs 2949 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) - Self(core::array::from_fn(|p| U64x2(unsafe { vbicq_u64(self.0[p].0, other.0[p].0) }))) +src/simd_neon.rs 2969 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic U64x2(unsafe { +src/simd_neon.rs 3063 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(core::array::from_fn(|p| U64x2(unsafe { vandq_u64(self.0[p].0, r.0[p].0) }))) +src/simd_neon.rs 3072 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(core::array::from_fn(|p| U64x2(unsafe { vorrq_u64(self.0[p].0, r.0[p].0) }))) +src/simd_neon.rs 3081 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(core::array::from_fn(|p| U64x2(unsafe { veorq_u64(self.0[p].0, r.0[p].0) }))) +src/simd_neon.rs 3112 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic U64x2(unsafe { vreinterpretq_u64_u32(vmvnq_u32(vreinterpretq_u32_u64(self.0[p].0))) }) +src/simd_neon.rs 3126 FORK block test NEEDS-TF present correctness: vshlq_u64 uses only low byte of count as signed; counts 128..255 shift right, 256 shifts 0 (doc says >=64 gives 0); debug_assert only [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(core::array::from_fn(|p| U64x2(unsafe { vshlq_u64(self.0[p].0, vreinterpretq_s64_u64(r.0[p].0)) }))) +src/simd_neon.rs 3138 FORK block test NEEDS-TF present correctness: same low-byte signed-count caveat as 3126 (right shift via negated count) [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic U64x2(unsafe { vshlq_u64(self.0[p].0, vnegq_s64(vreinterpretq_s64_u64(r.0[p].0))) }) +src/simd_neon.rs 3147 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic (0..4).all(|p| unsafe { vminvq_u32(vreinterpretq_u32_u64(vceqq_u64(self.0[p].0, other.0[p].0))) == u32::MAX }) +src/simd_neon.rs 3182 FORK block test CONSOLIDATE present - use U32x4::from_array(WEIGHTS).0 (existing helper); rest of quad_mask4 is value ops, so no local unsafe neon_intrinsic,raw_load_store,ptr_arith unsafe { vaddvq_u32(vandq_u32(cmp, vld1q_u32(WEIGHTS.as_ptr()))) as u16 } +src/simd_neon.rs 3241 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic let q: [i32; 4] = unsafe { +src/simd_neon.rs 3253 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 3263 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 3285 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic aarch64_simd::F32x16(core::array::from_fn(|p| unsafe { vcvtq_f32_s32(self.0[p].0) })) +src/simd_neon.rs 3293 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(core::array::from_fn(|p| I32x4(unsafe { vabsq_s32(self.0[p].0) }))) +src/simd_neon.rs 3303 FORK block test CONSOLIDATE present - as_chunks::<4>() on s[..16] + one audited vld1_s16(&[i16;4]) helper (or safe_unaligned_simd) neon_intrinsic,raw_load_store,ptr_arith Self(core::array::from_fn(|p| I32x4(unsafe { vmovl_s16(vld1_s16(s.as_ptr().add(4 * p))) }))) +src/simd_neon.rs 3312 FORK block test CONSOLIDATE present - as_chunks_mut::<4>() on o + one audited vst1_s16(&mut [i16;4]) helper neon_intrinsic,raw_load_store,ptr_arith unsafe { vst1_s16(o.as_mut_ptr().add(4 * p), vmovn_s32(self.0[p].0)) }; +src/simd_neon.rs 3323 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic m |= quad_mask4(unsafe { vcgezq_s32(self.0[p].0) }) << (4 * p); +src/simd_neon.rs 3337 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic m |= quad_mask4(unsafe { vcgtq_s32(self.0[p].0, other.0[p].0) }) << (4 * p); +src/simd_neon.rs 3379 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(core::array::from_fn(|p| I32x4(unsafe { vmulq_s32(self.0[p].0, r.0[p].0) }))) +src/simd_neon.rs 3395 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(core::array::from_fn(|p| I32x4(unsafe { vnegq_s32(self.0[p].0) }))) +src/simd_neon.rs 3404 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(core::array::from_fn(|p| I32x4(unsafe { vandq_s32(self.0[p].0, r.0[p].0) }))) +src/simd_neon.rs 3413 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(core::array::from_fn(|p| I32x4(unsafe { vorrq_s32(self.0[p].0, r.0[p].0) }))) +src/simd_neon.rs 3422 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic Self(core::array::from_fn(|p| I32x4(unsafe { veorq_s32(self.0[p].0, r.0[p].0) }))) +src/simd_neon.rs 3452 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) - Self(core::array::from_fn(|p| I32x4(unsafe { vmvnq_s32(self.0[p].0) }))) +src/simd_neon.rs 3460 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic (0..4).all(|p| unsafe { vminvq_u32(vceqq_s32(self.0[p].0, other.0[p].0)) == u32::MAX }) +src/simd_neon.rs 3480 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_neon.rs 3511 FORK block test NEEDS-TF present - [corrected: E0133 on 1.99 without a #[target_feature] caller, probe-verified] value intrinsics only; NEON is baseline on aarch64 so unsafe can go (unverified, not compiled on host) neon_intrinsic unsafe { +src/simd_nightly/f32_types.rs 92 FORK fn src NIGHTLY-BY-DESIGN n/a (nightly) - contract: 16 offsets valid/aligned; could be safe via gather(indices,&[f32],origin) using slice.get() per lane (costs bounds check) ptr_arith pub unsafe fn gather(indices: I32x16, base_ptr: *const f32) -> Self { +src/simd_nightly/f32_types.rs 97 FORK block src NIGHTLY-BY-DESIGN n/a (nightly) - raw deref covered by fn's # Safety contract; safe alt = checked slice index in a safe gather variant ptr_arith *o = unsafe { *base_ptr.offset(i as isize) }; +src/simd_nightly/u8_types.rs 294 FORK block src NIGHTLY-BY-DESIGN n/a (nightly) - array::from_fn(|i| u16::from_ne_bytes([b[2*i],b[2*i+1]])) / to_ne_bytes (same endianness as transmute); or as_chunks::<2>() transmute let mut words: [u16; 32] = unsafe { core::mem::transmute(bytes) }; +src/simd_nightly/u8_types.rs 298 FORK block src NIGHTLY-BY-DESIGN n/a (nightly) - inverse: flat_map(u16::to_ne_bytes) via from_fn; zero-cost after optimisation (unverified, nightly not compiled) transmute Self::from_array(unsafe { core::mem::transmute(words) }) +src/simd_nightly/u8_types.rs 317 FORK block src NIGHTLY-BY-DESIGN n/a (nightly) - array::from_fn(|i| u16::from_ne_bytes([b[2*i],b[2*i+1]])) / to_ne_bytes (same endianness as transmute); or as_chunks::<2>() transmute let mut words: [u16; 32] = unsafe { core::mem::transmute(bytes) }; +src/simd_nightly/u8_types.rs 321 FORK block src NIGHTLY-BY-DESIGN n/a (nightly) - inverse: flat_map(u16::to_ne_bytes) via from_fn; zero-cost after optimisation (unverified, nightly not compiled) transmute Self::from_array(unsafe { core::mem::transmute(words) }) +src/simd_nightly/u8_types.rs 696 FORK block src NIGHTLY-BY-DESIGN n/a (nightly) - array::from_fn(|i| u16::from_ne_bytes([b[2*i],b[2*i+1]])) / to_ne_bytes (same endianness as transmute); or as_chunks::<2>() transmute let mut words: [u16; 16] = unsafe { core::mem::transmute(bytes) }; +src/simd_nightly/u8_types.rs 700 FORK block src NIGHTLY-BY-DESIGN n/a (nightly) - inverse: flat_map(u16::to_ne_bytes) via from_fn; zero-cost after optimisation (unverified, nightly not compiled) transmute Self::from_array(unsafe { core::mem::transmute(words) }) +src/simd_nightly/u8_types.rs 718 FORK block src NIGHTLY-BY-DESIGN n/a (nightly) - array::from_fn(|i| u16::from_ne_bytes([b[2*i],b[2*i+1]])) / to_ne_bytes (same endianness as transmute); or as_chunks::<2>() transmute let mut words: [u16; 16] = unsafe { core::mem::transmute(bytes) }; +src/simd_nightly/u8_types.rs 722 FORK block src NIGHTLY-BY-DESIGN n/a (nightly) - inverse: flat_map(u16::to_ne_bytes) via from_fn; zero-cost after optimisation (unverified, nightly not compiled) transmute Self::from_array(unsafe { core::mem::transmute(words) }) +src/simd_runtime/add_mul.rs 26 FORK fn src CONSOLIDATE missing - fn-ptr type only needs `unsafe` because #[target_feature] kernels can't coerce to safe fn ptrs; wrap each in safe fn w/ one unsafe, or keep - type AddMulF32Fn = unsafe fn(&mut [f32], &[f32], &[f32]); +src/simd_runtime/add_mul.rs 27 FORK fn src CONSOLIDATE missing - same as line 26 (f64) - type AddMulF64Fn = unsafe fn(&mut [f64], &[f64], &[f64]); +src/simd_runtime/add_mul.rs 78 FORK block src IRREDUCIBLE present - call fn-pointer to #[target_feature] kernel after runtime CPU check; selection in LazyLock checks avx512f | avx2+fma, exactly the kernels' enables (complete) - unsafe { (*ADD_MUL_F32_DISPATCH)(acc, a, b) } +src/simd_runtime/add_mul.rs 84 FORK block src IRREDUCIBLE missing - same as line 78 (f64); selection checks complete - unsafe { (*ADD_MUL_F64_DISPATCH)(acc, a, b) } +src/simd_runtime/add_mul.rs 93 FORK fn src CONSOLIDATE missing - loadu/storeu on chunks n/16 are in-bounds; use one audited load/store helper on &[f32;16] (chunks_exact) or safe_unaligned_simd x86_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn add_mul_f32_avx512(acc: &mut [f32], a: &[f32], b: &[f32]) { +src/simd_runtime/add_mul.rs 112 FORK fn src CONSOLIDATE missing - same shape, 8xf64: helper load(&[f64;8]); bounds ok (n/8 chunks) x86_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn add_mul_f64_avx512(acc: &mut [f64], a: &[f64], b: &[f64]) { +src/simd_runtime/add_mul.rs 135 FORK fn src CONSOLIDATE missing - same shape 8xf32 avx2+fma; helper load(&[f32;8]); bounds ok x86_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn add_mul_f32_avx2_fma(acc: &mut [f32], a: &[f32], b: &[f32]) { +src/simd_runtime/add_mul.rs 154 FORK fn src CONSOLIDATE missing - same shape 4xf64 avx2+fma; helper load(&[f64;4]); bounds ok x86_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn add_mul_f64_avx2_fma(acc: &mut [f64], a: &[f64], b: &[f64]) { +src/simd_runtime/add_mul.rs 177 FORK fn src CONSOLIDATE missing - vld1q/vst1q on &[f32;4] helper; neon baseline on aarch64 so target_feature redundant. Unverified, not compiled on host neon_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn add_mul_f32_neon(acc: &mut [f32], a: &[f32], b: &[f32]) { +src/simd_runtime/add_mul.rs 196 FORK fn src CONSOLIDATE missing - vld1q_f64/vst1q_f64 helper on &[f64;2]. Unverified, not compiled on host neon_intrinsic,raw_load_store,ptr_arith,target_feature_call unsafe fn add_mul_f64_neon(acc: &mut [f64], a: &[f64], b: &[f64]) { +src/simd_runtime/add_mul.rs 217 FORK fn src REMOVABLE missing - body is safe; a safe fn coerces to `unsafe fn` pointer type, drop the keyword - unsafe fn add_mul_f32_scalar(acc: &mut [f32], a: &[f32], b: &[f32]) { +src/simd_runtime/add_mul.rs 224 FORK fn src REMOVABLE missing - same as 217 (f64) - unsafe fn add_mul_f64_scalar(acc: &mut [f64], a: &[f64], b: &[f64]) { +src/simd_runtime/add_mul.rs 239 FORK fn src CONSOLIDATE missing - redundant wrapper (avx512 f32): unsafe target_feature fn coerces to `unsafe fn` ptr; make kernel pub(super) and reference it directly target_feature_call pub(super) unsafe fn add_mul_f32_avx512_safe(acc: &mut [f32], a: &[f32], b: &[f32]) { +src/simd_runtime/add_mul.rs 244 FORK fn src CONSOLIDATE missing - redundant wrapper (avx512 f64): unsafe target_feature fn coerces to `unsafe fn` ptr; make kernel pub(super) and reference it directly target_feature_call pub(super) unsafe fn add_mul_f64_avx512_safe(acc: &mut [f64], a: &[f64], b: &[f64]) { +src/simd_runtime/add_mul.rs 249 FORK fn src CONSOLIDATE missing - redundant wrapper (avx2+fma f32): unsafe target_feature fn coerces to `unsafe fn` ptr; make kernel pub(super) and reference it directly target_feature_call pub(super) unsafe fn add_mul_f32_avx2_fma_safe(acc: &mut [f32], a: &[f32], b: &[f32]) { +src/simd_runtime/add_mul.rs 254 FORK fn src CONSOLIDATE missing - redundant wrapper (avx2+fma f64): unsafe target_feature fn coerces to `unsafe fn` ptr; make kernel pub(super) and reference it directly target_feature_call pub(super) unsafe fn add_mul_f64_avx2_fma_safe(acc: &mut [f64], a: &[f64], b: &[f64]) { +src/simd_runtime/add_mul.rs 259 FORK fn src CONSOLIDATE missing - redundant wrapper (neon f32); reference kernel directly. Unverified, not compiled on host target_feature_call pub(super) unsafe fn add_mul_f32_neon_safe(acc: &mut [f32], a: &[f32], b: &[f32]) { +src/simd_runtime/add_mul.rs 264 FORK fn src CONSOLIDATE missing - redundant wrapper (neon f64); reference kernel directly. Unverified, not compiled on host target_feature_call pub(super) unsafe fn add_mul_f64_neon_safe(acc: &mut [f64], a: &[f64], b: &[f64]) { +src/simd_runtime/add_mul.rs 268 FORK fn src REMOVABLE missing - safe body, forwards to safe-bodied scalar; use scalar fn directly in CpuOps (safe fn coerces) - pub(super) unsafe fn add_mul_f32_scalar_safe(acc: &mut [f32], a: &[f32], b: &[f32]) { +src/simd_runtime/add_mul.rs 272 FORK fn src REMOVABLE missing - same as 268 (f64) - pub(super) unsafe fn add_mul_f64_scalar_safe(acc: &mut [f64], a: &[f64], b: &[f64]) { +src/simd_runtime/casts.rs 41 FORK block src SAFE-CRATE present - bytemuck::cast_slice:: (BF16 is repr(transparent) u16, would need Pod/TransparentWrapper); else one audited as_u16(&[BF16]) helper ptr_arith,from_raw_parts let src_u16: &[u16] = unsafe { core::slice::from_raw_parts(src.as_ptr() as *const u16, src.len()) }; +src/simd_runtime/casts.rs 52 FORK block src SAFE-CRATE present - bytemuck::cast_slice_mut; else shared helper as_u16_mut(&mut [BF16]) with the single unsafe ptr_arith,from_raw_parts let dst_u16: &mut [u16] = unsafe { core::slice::from_raw_parts_mut(dst.as_mut_ptr() as *mut u16, dst.len()) }; +src/simd_runtime/cpu_ops.rs 74 FORK fn src UNSAFE-FN-API missing feature: AMX rung (select_cpu_ops) gates on amx_int8+OS only, installs kernels needing avx512f+avx512vnni unchecked; cpu_ops_for_tier(any) safe-returns unsupported tiers contract = host supports tier; public ptr could be safe fn if cpu_ops_for_tier/for_cpu returned only tiers the host supports - pub vnni_dot_u8_i8: unsafe fn(&[u8], &[i8]) -> i32, +src/simd_runtime/cpu_ops.rs 77 FORK fn src UNSAFE-FN-API missing - same contract as line 74 - pub add_mul_f32: unsafe fn(&mut [f32], &[f32], &[f32]), +src/simd_runtime/cpu_ops.rs 80 FORK fn src UNSAFE-FN-API missing - same contract as line 74 - pub add_mul_f64: unsafe fn(&mut [f64], &[f64], &[f64]), +src/simd_runtime/cpu_ops.rs 427 FORK block test IRREDUCIBLE missing - call through cpu_ops() fn ptr, tier selected by runtime caps; add SAFETY comment - let got = unsafe { (ops.vnni_dot_u8_i8)(&a, &b) }; +src/simd_runtime/cpu_ops.rs 434 FORK block test IRREDUCIBLE missing - same as 427 - unsafe { (ops.add_mul_f32)(&mut acc, &xa, &xb) }; +src/simd_runtime/vnni_dot.rs 21 FORK fn src CONSOLIDATE missing - unsafe fn-ptr type; safe wrappers (one unsafe each) would make dispatch fn-ptr safe - type VnniDotFn = unsafe fn(&[u8], &[i8]) -> i32; +src/simd_runtime/vnni_dot.rs 56 FORK block src IRREDUCIBLE present - fn-ptr call after dispatch closure; checks avx512f+avx512vnni | avx2+avxvnni match kernel enables (complete) - unsafe { (*VNNI_DOT_U8_I8_DISPATCH)(a, b) } +src/simd_runtime/vnni_dot.rs 76 FORK fn src IRREDUCIBLE present - calls simd_amx::vnni_dot_u8_i8 from fn enabling avx512f,avx512vnni (superset of callee); slicing to multiple of 64 in-bounds target_feature_call unsafe fn vnni_dot_u8_i8_avx512_with_tail(a: &[u8], b: &[i8]) -> i32 { +src/simd_runtime/vnni_dot.rs 107 FORK fn src CONSOLIDATE present - redundant wrapper: simd_amx::vnni2_dot_u8_i8 (same enables) coerces to VnniDotFn directly - unsafe fn vnni2_dot_u8_i8_safe_wrapper(a: &[u8], b: &[i8]) -> i32 { +src/simd_runtime/vnni_dot.rs 124 FORK fn src REMOVABLE present - safe body; use simd_amx::vnni_dot_u8_i8_scalar directly (safe fn coerces to unsafe fn ptr) - unsafe fn vnni_dot_u8_i8_scalar_safe_wrapper(a: &[u8], b: &[i8]) -> i32 { +src/simd_runtime/vnni_dot.rs 134 FORK fn src CONSOLIDATE present - redundant forwarding wrapper; make kernel pub(super) and point CpuOps at it target_feature_call pub(super) unsafe fn vnni_dot_u8_i8_avx512_with_tail_safe(a: &[u8], b: &[i8]) -> i32 { +src/simd_runtime/vnni_dot.rs 141 FORK fn src CONSOLIDATE present - redundant forwarding wrapper of vnni2 wrapper - pub(super) unsafe fn vnni2_dot_u8_i8_safe(a: &[u8], b: &[i8]) -> i32 { +src/simd_runtime/vnni_dot.rs 146 FORK fn src REMOVABLE present - safe forwarder of safe fn; drop - pub(super) unsafe fn vnni_dot_u8_i8_scalar_wrapper(a: &[u8], b: &[i8]) -> i32 { +src/simd_runtime/vnni_dot.rs 168 FORK block test IRREDUCIBLE present - test calls target_feature fn after is_x86_feature_detected avx2+avxvnni (checked) - let got = unsafe { vnni2_dot_u8_i8_safe_wrapper(&a, &b) }; +src/simd_scalar.rs 1545 FORK fn src UNSAFE-FN-API missing no # Safety doc; writes ptr.add(i) for each set mask bit with no bounds contract stated take &mut [u8;64] (or masked copy loop) -> safe fn; ptr contract is only "64 writable bytes" so it is checkable by type raw_load_store,ptr_arith pub unsafe fn mask_store(self, ptr: *mut u8, mask: u64) { +src/simd_wasm.rs 102 FORK block src CONSOLIDATE present - unverified, not compiled on host. v128_load pointer op; route via one load(&[T;N]) helper after asserting len (slice -> &[f32;4] chunks) wasm_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_wasm.rs 129 FORK block src CONSOLIDATE present - unverified, not compiled on host. v128_store pointer op; one store(&mut [T;N]) helper wasm_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_wasm.rs 455 FORK block src CONSOLIDATE present - unverified, not compiled on host. v128_load pointer op; route via one load(&[T;N]) helper after asserting len (slice -> &[f32;4] chunks) wasm_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_wasm.rs 482 FORK block src CONSOLIDATE present - unverified, not compiled on host. v128_store pointer op; one store(&mut [T;N]) helper wasm_intrinsic,raw_load_store,ptr_arith unsafe { +src/simd_wasm.rs 782 FORK block src CONSOLIDATE present - unverified, not compiled on host. assert then try_into &[i8;16]; use from_array/to_array helper wasm_intrinsic,raw_load_store,ptr_arith Self(unsafe { v128_load(s.as_ptr() as *const v128) }) +src/simd_wasm.rs 788 FORK block src CONSOLIDATE present - unverified, not compiled on host. this IS the audited load-from-array; other sites call it wasm_intrinsic,raw_load_store,ptr_arith Self(unsafe { v128_load(arr.as_ptr() as *const v128) }) +src/simd_wasm.rs 795 FORK block src CONSOLIDATE present - unverified, not compiled on host. this IS the audited store-to-array wasm_intrinsic,raw_load_store,ptr_arith unsafe { v128_store(arr.as_mut_ptr() as *mut v128, self.0) }; +src/simd_wasm.rs 803 FORK block src CONSOLIDATE present - unverified, not compiled on host. assert then try_into &[i8;16]; use from_array/to_array helper wasm_intrinsic,raw_load_store,ptr_arith unsafe { v128_store(s.as_mut_ptr() as *mut v128, self.0) }; +src/simd_wasm.rs 909 FORK block src CONSOLIDATE present - unverified, not compiled on host. this IS the audited load-from-array; other sites call it wasm_intrinsic,raw_load_store,ptr_arith Self(unsafe { v128_load(a.as_ptr() as *const v128) }) +src/simd_wasm.rs 916 FORK block src CONSOLIDATE present - unverified, not compiled on host. this IS the audited store-to-array wasm_intrinsic,raw_load_store,ptr_arith unsafe { v128_store(a.as_mut_ptr() as *mut v128, self.0) }; +src/simd_wasm.rs 1323 FORK block src CONSOLIDATE present - unverified, not compiled on host. arrays are &[T;16]: call one load helper (from_array) wasm_intrinsic,raw_load_store,ptr_arith let (va, vb) = unsafe { (v128_load(a.as_ptr() as *const v128), v128_load(b.as_ptr() as *const v128)) }; +src/simd_wasm.rs 1342 FORK block src CONSOLIDATE present - unverified, not compiled on host. store into [u8;16] via the store helper (or to_array) wasm_intrinsic,raw_load_store,ptr_arith unsafe { v128_store(tmp.as_mut_ptr() as *mut v128, v) }; +src/simd_wasm.rs 1350 FORK block src CONSOLIDATE present - unverified, not compiled on host. arrays are &[T;16]: call one load helper (from_array) wasm_intrinsic,raw_load_store,ptr_arith let (va, vb) = unsafe { (v128_load(a.as_ptr() as *const v128), v128_load(b.as_ptr() as *const v128)) }; +src/simd_wasm.rs 1363 FORK block src CONSOLIDATE present - unverified, not compiled on host. a.chunks_exact(16) -> &[u8;16] then load helper; removes ptr.add wasm_intrinsic,raw_load_store,ptr_arith let (va, vb) = unsafe { +src/simd_wasm.rs 1386 FORK block src CONSOLIDATE present soundness: safe nested fn derefs raw *const i16 params (only called with in-bounds ptrs) unverified, not compiled on host. abs_diff_block should take &[i16;8] and use load helper wasm_intrinsic,raw_load_store,ptr_arith let (va, vb) = unsafe { (v128_load(a as *const v128), v128_load(b as *const v128)) }; +src/simd_wasm.rs 1399 FORK block src SAFE-API present - unverified, not compiled on host. a[0..8].try_into::<&[i16;8]>() / a[8..16]; no raw ptr needed ptr_arith let hi = unsafe { abs_diff_block(a.as_ptr().add(8), b.as_ptr().add(8)) }; +src/simd_wasm.rs 1416 FORK block src CONSOLIDATE present bounds: safe pub fn, idx*dim+off unchecked (only comment "caller guarantees"); output.len()>=dim only debug_assert -> OOB read/write in release unverified, not compiled on host. centroids[off..off+4]/output[c*4..c*4+4] try_into &[f32;4] + load/store helper; real bounds checks wasm_intrinsic,raw_load_store,ptr_arith let v = unsafe { v128_load(centroids.as_ptr().add(off) as *const v128) }; +src/simd_wasm.rs 1420 FORK block src CONSOLIDATE present bounds: safe pub fn, idx*dim+off unchecked (only comment "caller guarantees"); output.len()>=dim only debug_assert -> OOB read/write in release unverified, not compiled on host. centroids[off..off+4]/output[c*4..c*4+4] try_into &[f32;4] + load/store helper; real bounds checks wasm_intrinsic,raw_load_store,ptr_arith unsafe { v128_store(output.as_mut_ptr().add(c * 4) as *mut v128, acc) }; +src/simd_wasm.rs 1477 FORK block src CONSOLIDATE present bounds: safe pub fn; a.len()>=k only debug_assert -> OOB read in release when k>len unverified, not compiled on host. a[off..off+16] try_into &[i8;16] + load helper wasm_intrinsic,raw_load_store,ptr_arith let (va, vb) = unsafe { +src/simd_wasm.rs 1556 FORK block src CONSOLIDATE present - unverified, not compiled on host. output[base..base+4] try_into &mut [f32;4] + store helper (or scalar chunks_exact_mut) wasm_intrinsic,raw_load_store,ptr_arith unsafe { v128_store(output.as_mut_ptr().add(base) as *mut v128, widened) }; +src/slice.rs 303 UPSTREAM trait src IRREDUCIBLE present - unsafe trait SliceArg: AsRef stability contract; sealed (private_decl) so not implementable downstream - pub unsafe trait SliceArg: AsRef<[SliceInfoElem]> { +src/slice.rs 316 UPSTREAM impl src IRREDUCIBLE missing - sealed unsafe trait impl: SliceArg AsRef-stability contract - unsafe impl SliceArg for &T +src/slice.rs 336 UPSTREAM impl src IRREDUCIBLE missing - sealed unsafe trait impl: SliceArg AsRef-stability contract - unsafe impl SliceArg<$in_dim> for SliceInfo +src/slice.rs 363 UPSTREAM impl src IRREDUCIBLE missing - sealed unsafe trait impl: SliceArg AsRef-stability contract - unsafe impl SliceArg for SliceInfo +src/slice.rs 382 UPSTREAM impl src IRREDUCIBLE missing - sealed unsafe trait impl: SliceArg AsRef-stability contract - unsafe impl SliceArg for [SliceInfoElem] { +src/slice.rs 458 UPSTREAM fn src UNSAFE-FN-API present - pub(doc hidden) new_unchecked; debug build checks dims; release unchecked unchecked pub unsafe fn new_unchecked( +src/slice.rs 482 UPSTREAM fn src UNSAFE-FN-API present - pub unsafe fn new: dims checked, only AsRef-stability contract is unchecked; TryFrom is the safe alt - pub unsafe fn new(indices: T) -> Result, ShapeError> { +src/slice.rs 529 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { +src/slice.rs 545 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { +src/slice.rs 565 UPSTREAM block src IRREDUCIBLE missing - raw-pointer core design: view/slice built from validated NonNull ptr, dim, strides - unsafe { +src/slice.rs 865 UPSTREAM block src IRREDUCIBLE present - s![] macro -> SliceInfo::new_unchecked on macro-constructed indices unchecked unsafe { +src/slice.rs 941 UPSTREAM block src IRREDUCIBLE missing - multi_slice_move: disjointness asserted then raw_view slice_move deref_into_view_mut - unsafe { +src/spatial_hash.rs 275 FORK block src SAFE-API present - batch_sq_dist_avx2 has no unsafe op (F32x8 polyfill); call as plain safe fn target_feature_call return unsafe { batch_sq_dist_avx2(query, candidates, radius_sq) }; +src/spatial_hash.rs 299 FORK fn src REMOVABLE present - No unsafe op in body target_feature_call pub(crate) unsafe fn batch_sq_dist_avx2(query: [f32; 3], candidates: &[[f32; 3]], radius_sq: f32) -> Vec<(usize, f32)> { +src/stacking.rs 63 UPSTREAM block src SAFE-API present - Array::from_shape_vec(res_dim, Vec::with_capacity(n)).unwrap() (empty vec, size 0: check trivially passes) - let mut res = unsafe { +src/stacking.rs 126 UPSTREAM block src SAFE-API present - same as :63 - let mut res = unsafe { +src/zip/mod.rs 97 UPSTREAM block src IRREDUCIBLE missing - ArrayView::new from broadcast result ptr/dim/strides; core design - unsafe { ArrayView::new(res.parts.ptr, res.parts.dim, res.parts.strides) } +src/zip/mod.rs 108 UPSTREAM fn src UNSAFE-FN-API missing - NdProducer-like trait method: ptr must come from uget_ptr in bounds - unsafe fn as_ref(&self, ptr: Self::Ptr) -> Self::Item; +src/zip/mod.rs 109 UPSTREAM fn src UNSAFE-FN-API missing - trait method: index must be in bounds - unsafe fn uget_ptr(&self, i: &Self::Dim) -> Self::Ptr; +src/zip/mod.rs 312 UPSTREAM block src IRREDUCIBLE missing - as_ref(as_ptr()) on single-element zip (non-empty checked) ptr_arith function(acc, unsafe { self.parts.as_ref(self.parts.as_ptr()) }) +src/zip/mod.rs 329 UPSTREAM block src IRREDUCIBLE missing - inner() called with ptrs from uget_ptr(index) in-bounds - unsafe { self.inner(acc, ptrs, inner_strides, size, &mut function) } +src/zip/mod.rs 340 UPSTREAM fn src UNSAFE-FN-API missing - inner loop over raw ptrs with strides; caller guarantees in-bounds - unsafe fn inner( +src/zip/mod.rs 386 UPSTREAM block src IRREDUCIBLE missing - uget_ptr + inner: index from dimension iterator, in-bounds by construction - unsafe { +src/zip/mod.rs 410 UPSTREAM block src IRREDUCIBLE missing - same as :386 (f-order) - unsafe { +src/zip/mod.rs 457 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn offset(self, off: isize) -> Self; +src/zip/mod.rs 458 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn stride_offset(self, index: usize, stride: isize) -> Self { +src/zip/mod.rs 464 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn offset(self, off: isize) -> Self { +src/zip/mod.rs 472 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design - unsafe fn stride_offset(self, stride: Self::Args, index: usize) -> Self; +src/zip/mod.rs 477 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn stride_offset(self, stride: Self::Args, index: usize) -> Self { +src/zip/mod.rs 488 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design - unsafe fn stride_offset(self, stride: Self::Args, index: usize) -> Self { +src/zip/mod.rs 534 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design - unsafe fn as_ref(&self, ptr: Self::Ptr) -> Self::Item { +src/zip/mod.rs 540 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design - unsafe fn uget_ptr(&self, i: &Self::Dim) -> Self::Ptr { +src/zip/mod.rs 716 UPSTREAM fn src UNSAFE-FN-API present - and_unchecked: caller guarantees producer dims match - pub(crate) unsafe fn and_unchecked

(self, p: P) -> Zip<($($p,)* P::Output, ), D> +src/zip/mod.rs 776 UPSTREAM block src IRREDUCIBLE missing - collect_with_partial on output raw view under Partial guard - unsafe { +src/zip/mod.rs 783 UPSTREAM block src IRREDUCIBLE missing - assume_init after full collect uninit unsafe { +src/zip/mod.rs 839 UPSTREAM fn src UNSAFE-FN-API missing - caller guarantees output producer is valid, uninit, same dims - pub(crate) unsafe fn collect_with_partial(self, mut f: F) -> Partial +src/zip/ndproducer.rs 85 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design - unsafe fn as_ref(&self, ptr: Self::Ptr) -> Self::Item; +src/zip/ndproducer.rs 87 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design - unsafe fn uget_ptr(&self, i: &Self::Dim) -> Self::Ptr; +src/zip/ndproducer.rs 102 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design - unsafe fn stride_offset(self, s: Self::Stride, index: usize) -> Self; +src/zip/ndproducer.rs 108 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn stride_offset(self, s: Self::Stride, index: usize) -> Self { +src/zip/ndproducer.rs 116 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn stride_offset(self, s: Self::Stride, index: usize) -> Self { +src/zip/ndproducer.rs 264 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn as_ref(&self, ptr: *mut A) -> Self::Item { +src/zip/ndproducer.rs 268 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn uget_ptr(&self, i: &Self::Dim) -> *mut A { +src/zip/ndproducer.rs 312 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn as_ref(&self, ptr: *mut A) -> Self::Item { +src/zip/ndproducer.rs 316 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn uget_ptr(&self, i: &Self::Dim) -> *mut A { +src/zip/ndproducer.rs 360 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn as_ref(&self, ptr: *const A) -> *const A { +src/zip/ndproducer.rs 364 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn uget_ptr(&self, i: &Self::Dim) -> *const A { +src/zip/ndproducer.rs 409 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn as_ref(&self, ptr: *mut A) -> *mut A { +src/zip/ndproducer.rs 413 UPSTREAM fn src UNSAFE-FN-API missing - unsafe fn: caller upholds pointer/bounds/validity precondition; raw-pointer core design ptr_arith unsafe fn uget_ptr(&self, i: &Self::Dim) -> *mut A { diff --git a/.claude/knowledge/vertical-simd-consumer-contract.md b/.claude/knowledge/vertical-simd-consumer-contract.md index e4e4c54c..be66b6da 100644 --- a/.claude/knowledge/vertical-simd-consumer-contract.md +++ b/.claude/knowledge/vertical-simd-consumer-contract.md @@ -471,8 +471,8 @@ simd_{avx512,avx2,neon,wasm,scalar}.rs peer backends, each owns realization `cargo rustc --release --manifest-path crates/neon-simd-parity/Cargo.toml --target aarch64-unknown-linux-gnu -- --emit=asm`, then count `(and|orr|eor|bic|orn) v*.16b` against `(and|orr|eor|bic) w*,`. -- **`unsafe` at the intrinsic boundary — where and why (measured, 1.98.1, - `tools/safe_intrinsic_probe`).** x86 and aarch64 SIMD intrinsics are safe +- **`unsafe` at the intrinsic boundary — where and why (measured, 1.98.1, re-run on + 1.99.0 2026-10-10 with identical results, `tools/safe_intrinsic_probe`).** x86 and aarch64 SIMD intrinsics are safe fns whose CALL requires the caller to carry the matching `#[target_feature]`; build-config features do not count (rustc says so in the E0133 note), and a safe annotated fn called from a plain fn fails the diff --git a/.github/workflows/ci.yaml b/.github/workflows/ci.yaml index cc98a65d..c8adf320 100644 --- a/.github/workflows/ci.yaml +++ b/.github/workflows/ci.yaml @@ -69,16 +69,18 @@ jobs: # Must equal `rust-toolchain.toml`'s channel. The version is NOT # restated in this comment on purpose: CI here does NOT read the # toolchain file (it uses `dtolnay/rust-toolchain@`), so the - # number lives in three places — the `MSRV`/`BLAS_MSRV` env above, the - # matrix entry below, and each `dtolnay/rust-toolchain@` step. All of - # them move together with the toolchain file, and its bump log is the - # one place that records WHY each move happened. + # number lives in two places — the matrix entry below and each + # `dtolnay/rust-toolchain@` step — which move together with the + # toolchain file. The `MSRV`/`BLAS_MSRV` env above tracks + # `Cargo.toml`'s `rust-version` floor instead, which can lag the + # channel (it did from the 1.99.0 bump). The toolchain file's bump log + # records WHY each move happened. rust: - - "1.98.1" + - "1.99.0" name: clippy/${{ matrix.rust }} steps: - uses: actions/checkout@v4 - - uses: dtolnay/rust-toolchain@1.98.1 + - uses: dtolnay/rust-toolchain@1.99.0 with: components: clippy - uses: Swatinem/rust-cache@v2 @@ -103,7 +105,7 @@ jobs: # Stable rustfmt on the pinned channel (see `rust-toolchain.toml`; # version deliberately not restated here). No nightly dependency since # rustfmt.toml is stable-clean post-PR #133. - - uses: dtolnay/rust-toolchain@1.98.1 + - uses: dtolnay/rust-toolchain@1.99.0 with: components: rustfmt - run: cargo fmt --all --check @@ -236,7 +238,7 @@ jobs: name: hpc-stream-parallel/rayon steps: - uses: actions/checkout@v4 - - uses: dtolnay/rust-toolchain@1.98.1 + - uses: dtolnay/rust-toolchain@1.99.0 - uses: Swatinem/rust-cache@v2 - uses: taiki-e/install-action@nextest - name: cargo check (no rayon — scalar path unchanged) diff --git a/CLAUDE.md b/CLAUDE.md index bb4c8f6f..f3dfc2b9 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ - **What**: High-performance linear algebra with pluggable BLAS backends (Native SIMD, MKL, OpenBLAS) - **Source**: `adaworldapi/rustynum` — reference GEMM, SIMD, and FFI implementations - **Target**: This repo — ndarray fork enhanced with HPC backends -- **Rust**: stable only — the pinned `rust-toolchain.toml` (1.98.1; `rust-version` in `Cargo.toml` is the floor). No nightly features on any default or supported build path. +- **Rust**: stable only — the pinned `rust-toolchain.toml` (1.99.0; `rust-version` in `Cargo.toml` is the floor, 1.98.1). No nightly features on any default or supported build path. - **The one documented exception — `nightly-simd` (opt-in, validation-only, since PR #173):** a Cargo feature that swaps the SIMD realization for `core::simd` (`src/simd_nightly/*`, `#![feature(portable_simd)]`) so the realization matrix can witness that backend too. It is never enabled by default, nothing on stable may depend on it, every stable CI row builds without it, and it is exercised only by the dedicated nightly CI rows (`nightly-simd-polyfill`, the `simd-matrix` nightly row) and `scripts/masking-parity.sh nightly`. Removing the feature would drop that backend from the matrix; enabling it anywhere by default would violate this rule. ## Agent Protocol diff --git a/Cargo.toml b/Cargo.toml index 71dc1d87..93e74c28 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -49,6 +49,10 @@ required-features = ["std"] name = "bf16_rne_exhaustive" required-features = ["std"] +[[example]] +name = "cpu_guard_probe" +required-features = ["std"] + [[example]] name = "splat3d_flex" required-features = ["splat3d"] diff --git a/Dockerfile b/Dockerfile index b8bd2bd4..f6e85c28 100644 --- a/Dockerfile +++ b/Dockerfile @@ -27,8 +27,8 @@ ENV RUSTUP_HOME=/usr/local/rustup \ CARGO_HOME=/usr/local/cargo \ PATH=/usr/local/cargo/bin:$PATH RUN curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | \ - sh -s -- -y --default-toolchain 1.98.1 --profile minimal \ - && rustc --version | grep -q "1.98.1" + sh -s -- -y --default-toolchain 1.99.0 --profile minimal \ + && rustc --version | grep -q "1.99.0" WORKDIR /app diff --git a/Dockerfile.avx512 b/Dockerfile.avx512 index f53f1802..b29e8b10 100644 --- a/Dockerfile.avx512 +++ b/Dockerfile.avx512 @@ -21,8 +21,8 @@ ENV RUSTUP_HOME=/usr/local/rustup \ CARGO_HOME=/usr/local/cargo \ PATH=/usr/local/cargo/bin:$PATH RUN curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | \ - sh -s -- -y --default-toolchain 1.98.1 --profile minimal \ - && rustc --version | grep -q "1.98.1" + sh -s -- -y --default-toolchain 1.99.0 --profile minimal \ + && rustc --version | grep -q "1.99.0" WORKDIR /app diff --git a/README-DE.md b/README-DE.md index f30ca59e..e8fe0104 100644 --- a/README-DE.md +++ b/README-DE.md @@ -1,6 +1,6 @@ # ndarray — HPC-Erweiterung fuer Rust -*Fork von [rust-ndarray/ndarray](https://github.com/rust-ndarray/ndarray) mit 100 HPC-Modulen, 2,534 bestandenen Bibliothekstests und SIMD-Kernels von Intel AMX bis Raspberry Pi NEON. Laeuft auf stabilem Rust 1.98.1 ohne Nightly-Features.* +*Fork von [rust-ndarray/ndarray](https://github.com/rust-ndarray/ndarray) mit 100 HPC-Modulen, 2,534 bestandenen Bibliothekstests und SIMD-Kernels von Intel AMX bis Raspberry Pi NEON. Laeuft auf stabilem Rust 1.99.0 ohne Nightly-Features.* Zaehlungen bei Commit `f2c1aea`: `pub mod`-Eintraege in `src/hpc/mod.rs`; `cargo test --lib` (2,534 bestanden, 32 ignoriert). Wie jede Zahl auf dieser Seite ermittelt wurde: [Belege](#belege-fuer-die-zahlen-auf-dieser-seite). @@ -197,7 +197,7 @@ cargo test --lib ## Anforderungen -- Rust 1.98.1 stable (festgelegt in `rust-toolchain.toml`; kein Nightly, keine instabilen Features) +- Rust 1.99.0 stable (festgelegt in `rust-toolchain.toml`; kein Nightly, keine instabilen Features) - Optional: gcc-aarch64-linux-gnu fuer Pi-Cross-Kompilierung - Optional: Intel MKL oder OpenBLAS (Feature-gesteuert) diff --git a/README.md b/README.md index 7e47cd37..deeec44f 100644 --- a/README.md +++ b/README.md @@ -1,6 +1,6 @@ # ndarray — HPC Expansion for Rust -*Fork of [rust-ndarray/ndarray](https://github.com/rust-ndarray/ndarray) with 100 HPC modules, 2,534 passing library tests, and SIMD kernels from Intel AMX to Raspberry Pi NEON. Runs on stable Rust 1.98.1 without nightly features.* +*Fork of [rust-ndarray/ndarray](https://github.com/rust-ndarray/ndarray) with 100 HPC modules, 2,534 passing library tests, and SIMD kernels from Intel AMX to Raspberry Pi NEON. Runs on stable Rust 1.99.0 without nightly features.* Counts at commit `f2c1aea`: `pub mod` entries in `src/hpc/mod.rs`; `cargo test --lib` (2,534 passed, 32 ignored). How every number on this page was obtained: [Evidence](#evidence-for-the-numbers-on-this-page). @@ -197,7 +197,7 @@ cargo test --lib ## Requirements -- Rust 1.98.1 stable (pinned in `rust-toolchain.toml`; no nightly, no unstable features) +- Rust 1.99.0 stable (pinned in `rust-toolchain.toml`; no nightly, no unstable features) - Optional: gcc-aarch64-linux-gnu for Pi cross-compilation - Optional: Intel MKL or OpenBLAS (feature-gated) diff --git a/crates/blas-mock-tests/src/lib.rs b/crates/blas-mock-tests/src/lib.rs index 4370464f..afd725d1 100644 --- a/crates/blas-mock-tests/src/lib.rs +++ b/crates/blas-mock-tests/src/lib.rs @@ -11,6 +11,12 @@ thread_local! { pub static CALL_COUNT: RefCell = const { RefCell::new(0) }; } +/// Mock of the CBLAS symbol: counts the call in [`CALL_COUNT`] and does no math. +/// +/// # Safety +/// +/// Never dereferences its pointer arguments, so any values are sound. It is +/// `unsafe` only to match the C signature that `cblas-sys` declares. #[rustfmt::skip] #[no_mangle] #[allow(unused)] @@ -33,6 +39,12 @@ pub unsafe extern "C" fn cblas_sgemm( CALL_COUNT.with(|ctx| *ctx.borrow_mut() += 1); } +/// Mock of the CBLAS symbol: counts the call in [`CALL_COUNT`] and does no math. +/// +/// # Safety +/// +/// Never dereferences its pointer arguments, so any values are sound. It is +/// `unsafe` only to match the C signature that `cblas-sys` declares. #[rustfmt::skip] #[no_mangle] #[allow(unused)] @@ -55,6 +67,12 @@ pub unsafe extern "C" fn cblas_dgemm( CALL_COUNT.with(|ctx| *ctx.borrow_mut() += 1); } +/// Mock of the CBLAS symbol: counts the call in [`CALL_COUNT`] and does no math. +/// +/// # Safety +/// +/// Never dereferences its pointer arguments, so any values are sound. It is +/// `unsafe` only to match the C signature that `cblas-sys` declares. #[rustfmt::skip] #[no_mangle] #[allow(unused)] @@ -77,6 +95,12 @@ pub unsafe extern "C" fn cblas_cgemm( CALL_COUNT.with(|ctx| *ctx.borrow_mut() += 1); } +/// Mock of the CBLAS symbol: counts the call in [`CALL_COUNT`] and does no math. +/// +/// # Safety +/// +/// Never dereferences its pointer arguments, so any values are sound. It is +/// `unsafe` only to match the C signature that `cblas-sys` declares. #[rustfmt::skip] #[no_mangle] #[allow(unused)] diff --git a/crates/blas-mock-tests/tests/use-blas.rs b/crates/blas-mock-tests/tests/use-blas.rs index a259c515..60b31ae7 100644 --- a/crates/blas-mock-tests/tests/use-blas.rs +++ b/crates/blas-mock-tests/tests/use-blas.rs @@ -76,9 +76,9 @@ fn test_gen_mat_mul_uses_blas() { let should_use_blas = av.strides().iter().all(|&s| s > 0) && bv.strides().iter().all(|&s| s > 0) && cv.strides().iter().all(|&s| s > 0) - && av.strides().iter().any(|&s| s == 1) - && bv.strides().iter().any(|&s| s == 1) - && cv.strides().iter().any(|&s| s == 1); + && av.strides().contains(&1) + && bv.strides().contains(&1) + && cv.strides().contains(&1); assert_eq!(should_use_blas, ncalls > 0); } } diff --git a/crates/simd-masking-parity/src/lib.rs b/crates/simd-masking-parity/src/lib.rs index bf0ccdcd..ae252aa7 100644 --- a/crates/simd-masking-parity/src/lib.rs +++ b/crates/simd-masking-parity/src/lib.rs @@ -28,7 +28,8 @@ //! index-addressed `masked_group_sum_i32_via` (two-hop zero-fallback); `0xD4x` //! `eq_u32_via_to_mask` (the same index lane, packed as a predicate rather than //! folded into a sum); `0xExx` the `F64x8` lane compares (all six relations -//! over every pair of 16 IEEE edge values: NaN, ±0, ±inf, subnormal). `main.rs` (native / qemu) and +//! over every pair of 16 IEEE edge values: NaN, ±0, ±inf, subnormal); `0xFxx` the 16-lane byte vectors `I8x16` / `U8x16` (wrapping +//! `add`/`sub`, signed/unsigned `min`/`max`, the load/store round trips). `main.rs` (native / qemu) and //! `selfcheck()` (the wasm cdylib export, driven by `run.mjs`) both call //! [`run`]. @@ -37,18 +38,19 @@ use ndarray::simd::{ eq_u32_via_to_mask, eq_u64_to_mask, eq_u8_to_mask, ge_i32_to_mask, ge_i32_to_mask_under, ge_u64_to_mask, ge_u8_to_mask, gt_i32_to_mask, gt_i32_to_mask_under, gt_u64_to_mask, gt_u8_to_mask, le_i32_to_mask, le_i32_to_mask_under, le_u64_to_mask, le_u8_to_mask, lt_i32_to_mask, lt_i32_to_mask_under, lt_u64_to_mask, - lt_u8_to_mask, mask_all, mask_and, mask_and_assign, mask_andnot, mask_andnot_assign, mask_any, mask_gather_u32, mask_gather_u32_under, - mask_not, mask_not_assign, mask_or, mask_or_assign, mask_scatter_or_u32, mask_set_range, mask_shift_morton, - mask_ternlog, mask_ternlog_any, mask_ternlog_assign, mask_ternlog_popcount, mask_xor, mask_xor_assign, - masked_group_sum_i32, masked_group_sum_i32_via, masked_key_run_count_u32, masked_max_i32, masked_min_i32, - masked_strided_group_sum, masked_sum_i32, masked_sum_wrapping_add_i32, ne_i32_to_mask, ne_i32_to_mask_under, - ne_u32_to_mask, ne_u32_to_mask_under, ne_u64_to_mask, ne_u8_to_mask, ternary_match_strided_to_mask, - ternary_match_u32_to_mask, ternary_match_u32_to_mask_under, ternary_match_u64_to_mask, - ternary_match_u64_to_mask_under, ternlog, F64x8, I32x16, KeyRunCarry, MortonDir, U32x16, U64x8, + lt_u8_to_mask, mask_all, mask_and, mask_and_assign, mask_andnot, mask_andnot_assign, mask_any, mask_gather_u32, + mask_gather_u32_under, mask_not, mask_not_assign, mask_or, mask_or_assign, mask_scatter_or_u32, mask_set_range, + mask_shift_morton, mask_ternlog, mask_ternlog_any, mask_ternlog_assign, mask_ternlog_popcount, mask_xor, + mask_xor_assign, masked_group_sum_i32, masked_group_sum_i32_via, masked_key_run_count_u32, masked_max_i32, + masked_min_i32, masked_strided_group_sum, masked_sum_i32, masked_sum_wrapping_add_i32, ne_i32_to_mask, + ne_i32_to_mask_under, ne_u32_to_mask, ne_u32_to_mask_under, ne_u64_to_mask, ne_u8_to_mask, + ternary_match_strided_to_mask, ternary_match_u32_to_mask, ternary_match_u32_to_mask_under, + ternary_match_u64_to_mask, ternary_match_u64_to_mask_under, ternlog, F64x8, I32x16, I8x16, KeyRunCarry, MortonDir, + U32x16, U64x8, U8x16, }; /// Number of check groups [`run`] executes (for the log line only). -pub const CHECKS: usize = 14; +pub const CHECKS: usize = 15; /// The wasm export: identical to [`run`], `extern "C"` so `run.mjs` can call it. #[no_mangle] @@ -62,7 +64,7 @@ pub fn run() -> u32 { check_ternlog_all_tables, check_u64x8_algebra, check_i32x16_compare, check_predicates_to_mask, check_mask_algebra, check_care_match, check_masked_reductions, check_blend, check_morton_shift, check_predicates_under, check_set_range, check_unsigned_compare_to_mask, check_gather_scatter_group, - check_f64x8_compare, + check_f64x8_compare, check_i8x16_u8x16_lanes, ]; for g in groups { if let Err(code) = g() { @@ -483,6 +485,173 @@ fn check_f64x8_compare() -> Result<(), u32> { Ok(()) } +// ── 0xFxx: I8x16 / U8x16 lane arithmetic — wrapping add/sub, min/max ──────── +// +// Every op is checked lane-for-lane against scalar `wrapping_add` / +// `wrapping_sub` / `min` / `max`. Each operand vector has 16 DISTINCT lanes, +// so a lane permutation (a swapped half, a reversed load) cannot pass, and +// the edge values `MIN`/`MAX`/`-1`/`0`/`1` (i8) and `0`/`255`/`1`/`128` (u8) +// sit in pairs that overflow both ways: `MAX+1`, `MIN-1`, `MIN+MIN`, +// `255+1`, `0-1`. A saturating implementation fails here; so does a signed +// `min` on the unsigned type (128 vs 1). Only facade methods every arm ships +// are used: no `cmp_gt`, no `==` on the vector types. + +fn check_i8x16_u8x16_lanes() -> Result<(), u32> { + // Lane pairs that wrap: i8 A[1]+B[1] = MAX+1, A[0]-C[0] = MIN-1, + // A[0]+A[0] = MIN+MIN; u8 C[1]+A[1] = 255+1, A[0]-B[0] = 0-1, and + // A[3]=128 vs D[3]=1 separates unsigned from signed min/max. + let i8_vecs: [[i8; 16]; 4] = [ + [i8::MIN, i8::MAX, -1, 0, 1, 2, -2, 64, -64, 100, -100, 124, -125, 42, 7, -7], + [i8::MAX, 1, i8::MIN, -1, 0, 126, -127, 65, -65, 27, -29, 5, -5, 8, -8, 3], + [1, i8::MIN, 0, i8::MAX, -1, 3, -3, 11, -11, 13, -13, 99, -99, 50, -50, 20], + [-1, i8::MAX, 1, i8::MIN, 0, -2, 2, -127, 126, -64, 63, -100, 99, 9, -9, 17], + ]; + let u8_vecs: [[u8; 16]; 4] = [ + [0, 255, 1, 128, 2, 254, 127, 129, 64, 192, 10, 200, 33, 77, 5, 250], + [1, 0, 255, 127, 253, 3, 128, 126, 191, 65, 246, 56, 34, 78, 6, 249], + [255, 1, 128, 0, 254, 2, 129, 127, 63, 193, 11, 199, 32, 76, 4, 251], + [128, 129, 0, 1, 3, 253, 255, 254, 190, 66, 245, 57, 35, 79, 7, 248], + ]; + // Lane distinctness of the fixtures is a precondition, not an assumption. + for v in &i8_vecs { + for i in 0..16 { + for j in (i + 1)..16 { + if v[i] == v[j] { + return Err(0xF0E); + } + } + } + } + for v in &u8_vecs { + for i in 0..16 { + for j in (i + 1)..16 { + if v[i] == v[j] { + return Err(0xF0E); + } + } + } + } + + let mut i8_wrapped = 0u32; + let mut u8_wrapped = 0u32; + + for a_arr in &i8_vecs { + let a = I8x16::from_array(*a_arr); + if a.to_array() != *a_arr { + return Err(0xF00); + } + for b_arr in &i8_vecs { + let b = I8x16::from_array(*b_arr); + let add = a.add(b).to_array(); + let sub = a.sub(b).to_array(); + let min = a.min(b).to_array(); + let max = a.max(b).to_array(); + for i in 0..16 { + if add[i] != a_arr[i].wrapping_add(b_arr[i]) { + return Err(0xF01); + } + if sub[i] != a_arr[i].wrapping_sub(b_arr[i]) { + return Err(0xF02); + } + if min[i] != a_arr[i].min(b_arr[i]) { + return Err(0xF03); + } + if max[i] != a_arr[i].max(b_arr[i]) { + return Err(0xF04); + } + if a_arr[i].checked_add(b_arr[i]).is_none() || a_arr[i].checked_sub(b_arr[i]).is_none() { + i8_wrapped += 1; + } + } + } + } + // The named wrap cases, explicitly: MAX+1, MIN-1, MIN+MIN. + let w = I8x16::splat(i8::MAX).add(I8x16::splat(1)).to_array(); + let x = I8x16::splat(i8::MIN).sub(I8x16::splat(1)).to_array(); + let y = I8x16::splat(i8::MIN).add(I8x16::splat(i8::MIN)).to_array(); + if w != [i8::MIN; 16] || x != [i8::MAX; 16] || y != [0i8; 16] { + return Err(0xF05); + } + + for a_arr in &u8_vecs { + let a = U8x16::from_array(*a_arr); + if a.to_array() != *a_arr { + return Err(0xF10); + } + for b_arr in &u8_vecs { + let b = U8x16::from_array(*b_arr); + let add = a.add(b).to_array(); + let sub = a.sub(b).to_array(); + let min = a.min(b).to_array(); + let max = a.max(b).to_array(); + for i in 0..16 { + if add[i] != a_arr[i].wrapping_add(b_arr[i]) { + return Err(0xF11); + } + if sub[i] != a_arr[i].wrapping_sub(b_arr[i]) { + return Err(0xF12); + } + if min[i] != a_arr[i].min(b_arr[i]) { + return Err(0xF13); + } + if max[i] != a_arr[i].max(b_arr[i]) { + return Err(0xF14); + } + if a_arr[i].checked_add(b_arr[i]).is_none() || a_arr[i].checked_sub(b_arr[i]).is_none() { + u8_wrapped += 1; + } + } + } + } + // 255+1 and 0-1, explicitly; and unsigned ordering: max(128, 1) is 128. + let w = U8x16::splat(255).add(U8x16::splat(1)).to_array(); + let x = U8x16::splat(0).sub(U8x16::splat(1)).to_array(); + let m = U8x16::splat(128).max(U8x16::splat(1)).to_array(); + if w != [0u8; 16] || x != [255u8; 16] || m != [128u8; 16] { + return Err(0xF15); + } + + // Anti-vacuity: the fixtures really overflowed, in both types. + if i8_wrapped == 0 || u8_wrapped == 0 { + return Err(0xF0F); + } + + // from_slice reads the FIRST 16 of a longer slice; copy_to_slice writes + // exactly the first 16 and leaves the rest untouched. + let long_i8: [i8; 20] = core::array::from_fn(|i| (i as i8).wrapping_mul(13).wrapping_sub(100)); + let want_i8: [i8; 16] = core::array::from_fn(|i| long_i8[i]); + if I8x16::from_slice(&long_i8).to_array() != want_i8 { + return Err(0xF20); + } + let mut out_i8 = [99i8; 20]; + I8x16::from_array(want_i8).copy_to_slice(&mut out_i8); + if out_i8[..16] != want_i8 || out_i8[16..] != [99i8; 4] { + return Err(0xF21); + } + let long_u8: [u8; 20] = core::array::from_fn(|i| (i as u8).wrapping_mul(17).wrapping_add(200)); + let want_u8: [u8; 16] = core::array::from_fn(|i| long_u8[i]); + if U8x16::from_slice(&long_u8).to_array() != want_u8 { + return Err(0xF22); + } + let mut out_u8 = [99u8; 20]; + U8x16::from_array(want_u8).copy_to_slice(&mut out_u8); + if out_u8[..16] != want_u8 || out_u8[16..] != [99u8; 4] { + return Err(0xF23); + } + + // zero() is splat(0), and splat fills every lane. + if I8x16::zero().to_array() != I8x16::splat(0).to_array() || I8x16::zero().to_array() != [0i8; 16] { + return Err(0xF30); + } + if U8x16::zero().to_array() != U8x16::splat(0).to_array() || U8x16::zero().to_array() != [0u8; 16] { + return Err(0xF31); + } + if I8x16::splat(-7).to_array() != [-7i8; 16] || U8x16::splat(201).to_array() != [201u8; 16] { + return Err(0xF32); + } + Ok(()) +} + // ── 0x5xx: predicate → mask, every tail shape, full-overwrite + zero tail ──── fn check_predicates_to_mask() -> Result<(), u32> { diff --git a/examples/cpu_guard_probe.rs b/examples/cpu_guard_probe.rs new file mode 100644 index 00000000..f0aa79a3 --- /dev/null +++ b/examples/cpu_guard_probe.rs @@ -0,0 +1,21 @@ +//! Probe for `ndarray::cpu_guard`: does SIMD work, so a build for a newer CPU +//! than the one running it would otherwise die with SIGILL. +//! +//! Falsifier (needs `qemu-user-static`): +//! ```sh +//! env -u RUSTFLAGS cargo --config .cargo/config-v4.toml build --release --example cpu_guard_probe +//! qemu-x86_64-static -cpu Haswell target/release/examples/cpu_guard_probe +//! ``` +//! Expected: an `error: this binary was compiled for a CPU with [...]` line and +//! exit status 132, not `Illegal instruction`. +use ndarray::simd::F32x16; + +fn main() { + let x: Vec = (0..1024).map(|i| i as f32).collect(); + let mut acc = F32x16::splat(0.0); + let (chunks, _) = x.as_chunks::<16>(); + for chunk in chunks { + acc += F32x16::from_slice(chunk); + } + println!("sum = {}", acc.reduce_sum()); +} diff --git a/rust-toolchain.toml b/rust-toolchain.toml index 06fb5ae5..fcd1dc57 100644 --- a/rust-toolchain.toml +++ b/rust-toolchain.toml @@ -1,5 +1,5 @@ [toolchain] -channel = "1.98.1" +channel = "1.99.0" # The pinned version is the `channel` line ABOVE — deliberately not restated # here. This comment used to open "Pinned to 1.97.1", and this file's own # warning explains why that is a trap: "A stale comment on a version pin is how @@ -22,6 +22,16 @@ channel = "1.98.1" # workspace-wide sweep. The 1.98 delta measured across the # stack is ONE lint, `clippy::chunks_exact_to_as_chunks` # (new, default-on STYLE group) — zero sites in this repo. +# 1.98.1 → 1.99.0 current stable (2026-10-01, LLVM 23). Channel ONLY: +# `rust-version` stays 1.98.1 because no 1.99-only API is +# used, so CI's MSRV rows keep testing that floor. This is +# a deliberate split, superseding the "move TOGETHER" rule +# above for this bump. Measured 1.99 delta, all in test +# crates: `clippy::missing_safety_doc` on 4 mock +# `pub unsafe extern "C"` fns (blas-mock-tests), +# `clippy::manual_contains` x3 (blas-mock-tests), +# deprecated `std::f64::NAN` via `use std::f64;` x4 +# (tests/numeric.rs). # # Never auto-track `stable` — bump explicitly when a future version is # reviewed and the workspace clippy passes clean. diff --git a/src/cpu_guard.rs b/src/cpu_guard.rs new file mode 100644 index 00000000..b1be1990 --- /dev/null +++ b/src/cpu_guard.rs @@ -0,0 +1,470 @@ +//! Build-vs-CPU guard: turn "built for a newer CPU" from a SIGILL into a message. +//! +//! The default build is `-Ctarget-cpu=native` (`.cargo/config.toml`), so the +//! binary is compiled for the instruction set of the machine that BUILT it, and +//! LLVM may emit any of those instructions anywhere, not only inside +//! `crate::simd`. Running such a binary on a CPU that lacks one of them dies +//! on the first such instruction with `SIGILL` (illegal instruction), with no +//! hint of why. That happens when the build host and the run host differ: a +//! release asset built on a CI runner, a Docker image built on one machine and +//! deployed on another. +//! +//! This module compares the target features the crate was COMPILED with +//! (`cfg!(target_feature = ...)`) against what the running CPU REPORTS +//! (CPUID and XCR0 read directly, see below), and names every feature that is +//! missing. x86_64 only for now; aarch64 is documented below as not covered. +//! +//! It also enforces one floor that is NOT a compile-time feature: **AVX2 on +//! every x86_64 build.** `crate::simd` selects its AVX2 realization whenever +//! AVX-512 is absent, and that code calls AVX2 intrinsics without a runtime +//! check, so a build compiled without `avx2` still needs AVX2 the moment it +//! touches the SIMD types. The floor is checked once, here, at startup, and +//! never per call: dispatch stays compile-time (operator, 2026-10-10, after +//! codex review on PR #348). +//! +//! # What it does not do +//! +//! It never changes how anything is built. Cross-building (for example +//! `--config .cargo/config-v4.toml` on a runner without AVX-512) is unaffected: +//! the check runs when the finished program STARTS, never at compile time. It +//! only ever fires for a binary that would otherwise have crashed with SIGILL +//! on this CPU, so there is no override switch to forget. +//! +//! # When it runs +//! +//! * Automatically, before `main`, on Linux / Android / FreeBSD (`.init_array`), +//! macOS / iOS (`__mod_init_func`) and Windows (`.CRT$XCU`), for every binary +//! that links this crate. On failure it prints the missing features to stderr +//! and exits with status 132 (`128 + SIGILL`), the status the crash would +//! have produced. +//! * On demand, through [`check_build_cpu`] (returns the mismatch) or +//! [`assert_build_cpu`] (panics with the same message), for targets without a +//! pre-`main` hook or for callers that want to report it themselves. +//! +//! # Limitation (measured) +//! +//! The guard is compiled with the build's own target features: Rust cannot +//! compile one function for a lower target than the rest of the crate. So the +//! guard itself needs whatever instruction ENCODING the build uses. +//! +//! * Covered: a build for AVX-512 (or for AVX-512 extensions such as VBMI, +//! BF16, FP16, VNNI) run on a CPU with AVX/AVX2. The guard's own code is +//! VEX-encoded, which such CPUs run. Measured: a `x86-64-v4` build under +//! `qemu-x86_64 -cpu Haswell|Skylake-Server|Icelake-Server` prints the +//! missing features and exits 132 (qemu emulates no AVX-512 at all). +//! * Not covered, by decision: a `x86-64-v3`/`v4`/`native` build run on a CPU +//! WITHOUT AVX (pre-2011, e.g. Nehalem). There even scalar code is +//! VEX-encoded, and the guard faults inside itself. Measured: `x86-64-v3` +//! under `-cpu Nehalem` still dies with SIGILL in `guard_before_main`. +//! Covering these CPUs is an OPTIONAL to-do, postponed (operator, +//! 2026-10-10); see `.claude/blackboard.md`, "Optional to-do: pre-AVX +//! guard". +//! * Silent when it should be: the same `x86-64-v3` build under `-cpu Haswell` +//! runs normally (exit 0). +//! * AVX2 floor: a baseline build (no `target-cpu`, CI's flags) under +//! `-cpu Nehalem|SandyBridge|IvyBridge` prints the AVX2 message and exits +//! 132; under `-cpu Haswell|max` it runs normally. Such a build is not +//! VEX-encoded, so on a pre-AVX CPU the guard runs and reports instead of +//! faulting. + +use std::fmt; + +/// One target feature the crate was compiled with but the running CPU lacks. +pub type MissingFeature = &'static str; + +/// The running CPU lacks target features this build was compiled to use. +/// +/// Returned by [`check_build_cpu`]. Its `Display` names every missing +/// feature and how to rebuild. +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct BuildCpuMismatch { + /// Features the build needs that the running CPU does not report: those + /// enabled at compile time, plus the x86_64 AVX2 floor. + pub missing: Vec, +} + +impl fmt::Display for BuildCpuMismatch { + fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { + write!( + f, + "this binary needs CPU features [{}], which this CPU does not support. \ + Running it would crash with SIGILL (illegal instruction).", + self.missing.join(", ") + )?; + if self.missing.contains(&"avx2") { + // Rebuilding cannot help: the AVX2 floor applies to every x86_64 build. + write!( + f, + " ndarray's x86_64 SIMD backend needs AVX2 whatever the build flags, \ + so no build of this program can run on this CPU." + ) + } else { + write!( + f, + " Rebuild for this machine (the default `target-cpu=native` does that when \ + built here), or for a common baseline such as \ + `--config .cargo/config-v3.toml` (x86-64-v3, AVX2)." + ) + } + } +} + +impl std::error::Error for BuildCpuMismatch {} + +/// `(name, compiled-in, present-on-this-CPU)`. +type FeatureRow = (&'static str, bool, fn() -> bool); + +/// Builds the table. `cfg!(target_feature = ...)` needs a string literal, so +/// the rows are generated rather than looped. +#[cfg(target_arch = "x86_64")] +macro_rules! feature_table { + ($($name:tt => $present:expr),* $(,)?) => { + &[$(($name, cfg!(target_feature = $name), (|| $present) as fn() -> bool),)*] + }; +} + +// Why not `is_x86_feature_detected!`: that macro returns `true` WITHOUT asking +// the CPU whenever the feature is enabled at compile time (it expands to +// `cfg!(target_feature = ..) || runtime_check`). For the one question this +// module asks, "was it compiled in AND is it missing?", it therefore always +// answers "present". Measured: a `x86-64-v4` build under +// `qemu-x86_64 -cpu Haswell` passed that check and then died with SIGILL on +// its first EVEX instruction in `main`. So the CPU is asked directly. +#[cfg(target_arch = "x86_64")] +mod cpuid { + use core::arch::x86_64::{__cpuid, __cpuid_count, CpuidResult}; + + fn leaf(eax: u32, ecx: u32) -> CpuidResult { + // CPUID exists on every x86_64 CPU, so `__cpuid_count` is a safe fn. + // `bit` checks the maximum supported leaf before asking for one. + __cpuid_count(eax, ecx) + } + + fn max_basic() -> u32 { + __cpuid(0).eax + } + + fn max_ext() -> u32 { + __cpuid(0x8000_0000).eax + } + + /// Bit `b` of a register of leaf `(eax, ecx)`, or false when the leaf is + /// beyond what the CPU reports. + pub(super) fn bit(eax: u32, ecx: u32, reg: char, b: u32) -> bool { + let supported = if eax >= 0x8000_0000 { + max_ext() >= eax + } else { + max_basic() >= eax + }; + if !supported || (eax == 7 && ecx > 0 && leaf(7, 0).eax < ecx) { + return false; + } + let r = leaf(eax, ecx); + let v = match reg { + 'a' => r.eax, + 'b' => r.ebx, + 'c' => r.ecx, + _ => r.edx, + }; + v >> b & 1 == 1 + } + + /// XCR0: which register states the OS saves on a context switch. AVX and + /// AVX-512 instructions fault (`#UD`, i.e. SIGILL) unless these are set, + /// even when the CPUID feature bit is. + fn xcr0() -> u64 { + if !bit(1, 0, 'c', 27) { + return 0; // OSXSAVE clear: XGETBV itself would fault. + } + let (lo, hi): (u32, u32); + // SAFETY: OSXSAVE (CPUID.1:ECX bit 27) is set, so XGETBV is enabled + // and XCR0 (ECX = 0) is readable at every privilege level. + unsafe { + core::arch::asm!("xgetbv", in("ecx") 0u32, out("eax") lo, out("edx") hi, + options(nomem, nostack, preserves_flags)); + } + (hi as u64) << 32 | lo as u64 + } + + /// The OS saves XMM + YMM state. + pub(super) fn os_avx() -> bool { + xcr0() & 0b110 == 0b110 + } + + /// The OS saves XMM + YMM + opmask + ZMM state. + /// + /// On Apple targets this is assumed: Darwin saves the AVX-512 context + /// lazily, on first use, so XCR0 does not show it until then. LLVM's own + /// host detection (`llvm/lib/TargetParser/Host.cpp`, `HasAVX512Save`) + /// makes the same exception; without it this guard would refuse to start + /// a correct AVX-512 build on an AVX-512 Mac. + pub(super) fn os_avx512() -> bool { + cfg!(target_vendor = "apple") || xcr0() & 0b1110_0110 == 0b1110_0110 + } +} + +// Bit positions and OS-state gating mirror LLVM's `getHostCPUFeatures` +// (`llvm/lib/TargetParser/Host.cpp`, checked 2026-10-10), which is what +// `-Ctarget-cpu=native` itself consults, so "supported" here means what it +// means to the compiler that produced the build. +#[cfg(target_arch = "x86_64")] +const FEATURES: &[FeatureRow] = { + use cpuid::{bit, os_avx as avx_os, os_avx512 as z}; + feature_table!( + "sse3" => bit(1, 0, 'c', 0), + "pclmulqdq" => bit(1, 0, 'c', 1), + "ssse3" => bit(1, 0, 'c', 9), + "fma" => bit(1, 0, 'c', 12) && avx_os(), + "cmpxchg16b" => bit(1, 0, 'c', 13), + "sse4.1" => bit(1, 0, 'c', 19), + "sse4.2" => bit(1, 0, 'c', 20), + "movbe" => bit(1, 0, 'c', 22), + "popcnt" => bit(1, 0, 'c', 23), + "aes" => bit(1, 0, 'c', 25), + "xsave" => bit(1, 0, 'c', 26) && avx_os(), + "avx" => bit(1, 0, 'c', 28) && avx_os(), + "f16c" => bit(1, 0, 'c', 29) && avx_os(), + "rdrand" => bit(1, 0, 'c', 30), + "fxsr" => bit(1, 0, 'd', 24), + "bmi1" => bit(7, 0, 'b', 3), + "avx2" => bit(7, 0, 'b', 5) && avx_os(), + "bmi2" => bit(7, 0, 'b', 8), + "avx512f" => bit(7, 0, 'b', 16) && z(), + "avx512dq" => bit(7, 0, 'b', 17) && z(), + "rdseed" => bit(7, 0, 'b', 18), + "adx" => bit(7, 0, 'b', 19), + "avx512ifma" => bit(7, 0, 'b', 21) && z(), + "avx512cd" => bit(7, 0, 'b', 28) && z(), + "sha" => bit(7, 0, 'b', 29), + "avx512bw" => bit(7, 0, 'b', 30) && z(), + "avx512vl" => bit(7, 0, 'b', 31) && z(), + "avx512vbmi" => bit(7, 0, 'c', 1) && z(), + "avx512vbmi2" => bit(7, 0, 'c', 6) && z(), + "gfni" => bit(7, 0, 'c', 8), + "vaes" => bit(7, 0, 'c', 9) && avx_os(), + "vpclmulqdq" => bit(7, 0, 'c', 10) && avx_os(), + "avx512vnni" => bit(7, 0, 'c', 11) && z(), + "avx512bitalg" => bit(7, 0, 'c', 12) && z(), + "avx512vpopcntdq" => bit(7, 0, 'c', 14) && z(), + "avx512vp2intersect" => bit(7, 0, 'd', 8) && z(), + "avx512fp16" => bit(7, 0, 'd', 23) && z(), + "avxvnni" => bit(7, 1, 'a', 4) && avx_os(), + "avx512bf16" => bit(7, 1, 'a', 5) && z(), + "xsaveopt" => bit(0xD, 1, 'a', 0) && avx_os(), + "xsavec" => bit(0xD, 1, 'a', 1) && avx_os(), + "xsaves" => bit(0xD, 1, 'a', 3) && avx_os(), + "lzcnt" => bit(0x8000_0001, 0, 'c', 5), + "sse4a" => bit(0x8000_0001, 0, 'c', 6), + "tbm" => bit(0x8000_0001, 0, 'c', 21), + "kl" => bit(7, 0, 'c', 23), + "widekl" => bit(7, 0, 'c', 23) && bit(0x19, 0, 'b', 2), + // LLVM gates SHA512/SM3/SM4 on the CPUID bit only (no XCR0 check); + // mirrored as is. + "sha512" => bit(7, 1, 'a', 0), + "sm3" => bit(7, 1, 'a', 1), + "sm4" => bit(7, 1, 'a', 2), + "avxifma" => bit(7, 1, 'a', 23) && avx_os(), + "avxvnniint8" => bit(7, 1, 'd', 4) && avx_os(), + "avxneconvert" => bit(7, 1, 'd', 5) && avx_os(), + "avxvnniint16" => bit(7, 1, 'd', 10) && avx_os(), + ) +}; + +/// Features required on every x86_64 build, whatever it was compiled with: +/// the floor of `crate::simd`'s AVX2 realization (see the module docs). The +/// `compiled` column is `true` by definition. +#[cfg(target_arch = "x86_64")] +const BACKEND_FLOOR: &[FeatureRow] = &[("avx2", true, || cpuid::bit(7, 0, 'b', 5) && cpuid::os_avx())]; + +#[cfg(not(target_arch = "x86_64"))] +const BACKEND_FLOOR: &[FeatureRow] = &[]; + +// aarch64 is NOT covered yet, deliberately: `is_aarch64_feature_detected!` +// has the same compile-time short-circuit, so a table built on it could never +// fire, and a guard that cannot fire is worse than none (it reads as coverage). +// Covering it needs the kernel's HWCAP bits (`getauxval(AT_HWCAP)`), which std +// does not expose. Until then the explicit API returns Ok on aarch64. +#[cfg(not(target_arch = "x86_64"))] +const FEATURES: &[FeatureRow] = &[]; + +/// Target features the crate was compiled with that the running CPU lacks. +/// +/// Empty on a CPU that supports the build, and always empty on architectures +/// this guard does not cover yet (anything but x86_64). +pub fn missing_build_features() -> Vec { + let mut missing = missing_in(FEATURES); + for f in missing_in(BACKEND_FLOOR) { + if !missing.contains(&f) { + missing.push(f); + } + } + missing +} + +fn missing_in(table: &[FeatureRow]) -> Vec { + table + .iter() + .filter(|&&(_, compiled, detected)| compiled && !detected()) + .map(|&(name, _, _)| name) + .collect() +} + +/// Checks that the running CPU supports every target feature this build uses. +/// +/// # Examples +/// +/// ``` +/// // A build made on the machine that runs it always passes. +/// assert!(ndarray::cpu_guard::check_build_cpu().is_ok()); +/// ``` +pub fn check_build_cpu() -> Result<(), BuildCpuMismatch> { + let missing = missing_build_features(); + if missing.is_empty() { + Ok(()) + } else { + Err(BuildCpuMismatch { missing }) + } +} + +/// Panics with a readable message if the running CPU cannot run this build. +/// +/// # Examples +/// +/// ``` +/// ndarray::cpu_guard::assert_build_cpu(); +/// ``` +pub fn assert_build_cpu() { + if let Err(mismatch) = check_build_cpu() { + panic!("{mismatch}"); + } +} + +/// The pre-`main` hook. Prints and exits instead of panicking: unwinding out +/// of a loader-invoked `extern "C"` function would abort with no message. +extern "C" fn guard_before_main() { + if let Err(mismatch) = check_build_cpu() { + use std::io::Write; + let _ = writeln!(std::io::stderr(), "error: {mismatch}"); + std::process::exit(132); + } +} + +// The loader calls every function pointer in these sections before `main`. +// `#[used]` keeps the static even though nothing references it. +#[cfg(any(target_os = "linux", target_os = "android", target_os = "freebsd"))] +#[used] +#[link_section = ".init_array"] +static GUARD_BEFORE_MAIN: extern "C" fn() = guard_before_main; + +#[cfg(any(target_os = "macos", target_os = "ios"))] +#[used] +#[link_section = "__DATA,__mod_init_func"] +static GUARD_BEFORE_MAIN: extern "C" fn() = guard_before_main; + +#[cfg(target_os = "windows")] +#[used] +#[link_section = ".CRT$XCU"] +static GUARD_BEFORE_MAIN: extern "C" fn() = guard_before_main; + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn the_host_runs_its_own_build() { + // Built here, run here: nothing may be missing. + assert_eq!(missing_build_features(), Vec::::new()); + assert!(check_build_cpu().is_ok()); + } + + #[test] + fn the_table_sees_the_features_this_build_was_compiled_with() { + // Anti-vacuity: on x86_64 every build has at least SSE2-era features, + // and a native/v3 build has AVX2, so the "compiled" column must not be + // all false (that would make the guard unable to fire). + #[cfg(all(target_arch = "x86_64", target_feature = "avx2"))] + assert!(FEATURES + .iter() + .any(|&(n, compiled, _)| n == "avx2" && compiled)); + } + + /// Every x86_64 target feature that some `-Ctarget-cpu` model enables on + /// rustc 1.99, minus `sse`/`sse2` (the x86_64 baseline). A feature missing + /// from `FEATURES` is one the guard cannot see, so a build that uses it on + /// a CPU lacking it SIGILLs instead of reporting (codex review, PR #348). + /// Regenerate on a toolchain bump: + /// `for c in $(rustc --print target-cpus | awk 'NR>1{print $1}'); do + /// rustc --print cfg -Ctarget-cpu=$c; done | grep target_feature | sort -u` + #[cfg(target_arch = "x86_64")] + const RUSTC_CPU_MODEL_FEATURES: &[&str] = &[ + "adx", "aes", "avx", "avx2", "avx512bf16", "avx512bitalg", "avx512bw", "avx512cd", "avx512dq", "avx512f", + "avx512fp16", "avx512ifma", "avx512vbmi", "avx512vbmi2", "avx512vl", "avx512vnni", "avx512vp2intersect", + "avx512vpopcntdq", "avxifma", "avxneconvert", "avxvnni", "avxvnniint16", "avxvnniint8", "bmi1", "bmi2", + "cmpxchg16b", "f16c", "fma", "fxsr", "gfni", "kl", "lzcnt", "movbe", "pclmulqdq", "popcnt", "rdrand", "rdseed", + "sha", "sha512", "sm3", "sm4", "sse3", "sse4.1", "sse4.2", "sse4a", "ssse3", "tbm", "vaes", "vpclmulqdq", + "widekl", "xsave", "xsavec", "xsaveopt", "xsaves", + ]; + + #[test] + #[cfg(target_arch = "x86_64")] + fn the_table_covers_every_feature_a_cpu_model_can_enable() { + let missing: Vec<_> = RUSTC_CPU_MODEL_FEATURES + .iter() + .filter(|f| !FEATURES.iter().any(|&(n, _, _)| n == **f)) + .collect(); + assert!(missing.is_empty(), "guard cannot see: {missing:?}"); + assert!(RUSTC_CPU_MODEL_FEATURES.len() > 40, "anti-vacuity"); + } + + #[test] + fn a_compiled_feature_the_cpu_lacks_is_reported() { + // Can-fire: a synthetic table, because the host supports its own build. + let table: &[FeatureRow] = &[ + ("present", true, || true), + ("absent", true, || false), + ("not-compiled", false, || false), + ("also-absent", true, || false), + ]; + assert_eq!(missing_in(table), vec!["absent", "also-absent"]); + } + + #[test] + fn a_feature_not_compiled_in_is_never_reported() { + // Can-stay-silent: a CPU lacking a feature the build does not use is fine. + let table: &[FeatureRow] = &[("a", false, || false), ("b", true, || true)]; + assert!(missing_in(table).is_empty()); + } + + #[test] + #[cfg(target_arch = "x86_64")] + fn avx2_is_required_on_every_x86_64_build() { + // The floor must not depend on the build flags: a baseline build + // (no `avx2` compiled in) still reaches the AVX2 SIMD backend. + assert!(BACKEND_FLOOR + .iter() + .any(|&(n, compiled, _)| n == "avx2" && compiled)); + } + + #[test] + fn the_message_explains_the_avx2_floor_only_when_avx2_is_missing() { + let floor = BuildCpuMismatch { missing: vec!["avx2"] }.to_string(); + assert!(floor.contains("needs AVX2 whatever the build"), "{floor}"); + assert!(!floor.contains("Rebuild"), "rebuilding cannot help: {floor}"); + let other = BuildCpuMismatch { + missing: vec!["avx512f"], + } + .to_string(); + assert!(!other.contains("needs AVX2"), "{other}"); + assert!(other.contains("Rebuild"), "{other}"); + } + + #[test] + fn the_message_names_every_missing_feature() { + let m = BuildCpuMismatch { + missing: vec!["avx512f", "avx512bw"], + }; + let text = m.to_string(); + assert!(text.contains("avx512f, avx512bw"), "{text}"); + assert!(text.contains("SIGILL"), "{text}"); + } +} diff --git a/src/hpc/fingerprint.rs b/src/hpc/fingerprint.rs index 5d9b9cb6..35dfcdd2 100644 --- a/src/hpc/fingerprint.rs +++ b/src/hpc/fingerprint.rs @@ -658,9 +658,9 @@ impl Fingerprint<8> { /// (both use `u64::to_le_bytes` / `u64::from_le_bytes`). /// On a big-endian target the raw memory layout of `[u64; 8]` would put /// the high byte first, contradicting the LE contract and breaking - /// cross-platform SIMD consumers. The `.cargo/config.toml` pins - /// `target-cpu=x86-64-v4`, so all supported targets (x86_64 + aarch64) - /// are little-endian; the cfg gate makes the LE assumption explicit + /// cross-platform SIMD consumers. All supported targets (x86_64, aarch64, + /// wasm32) are little-endian, whatever `target-cpu` is chosen; the cfg gate + /// makes the LE assumption explicit /// rather than implicit. See P2 review on PR #167. /// /// # Design reference diff --git a/src/lib.rs b/src/lib.rs index 0c1bb999..74d2fda3 100644 --- a/src/lib.rs +++ b/src/lib.rs @@ -276,6 +276,12 @@ pub mod simd_amx; #[cfg(feature = "std")] pub mod simd_caps; +/// Build-vs-CPU guard: before `main`, checks that the running CPU supports +/// every target feature this build was compiled with, and prints the missing +/// ones instead of crashing with SIGILL. See [`cpu_guard::check_build_cpu`]. +#[cfg(feature = "std")] +pub mod cpu_guard; + /// Bitwise SIMD primitives — popcount, Hamming distance over byte slices. /// Graduated from `crate::hpc::bitwise::*` (substrate-tier; uses /// `crate::simd::U64x8` polyfill internally). Back-compat re-export in diff --git a/src/simd.rs b/src/simd.rs index 980c2c0a..a3735dad 100644 --- a/src/simd.rs +++ b/src/simd.rs @@ -235,9 +235,9 @@ pub const PREFERRED_I16_LANES: usize = 16; pub use crate::simd_nightly::{ batch_packed_i4_16, f32x16, f32x8, f64x4, f64x8, i16x16, i16x32, i32x16, i32x8, i64x4, i64x8, i8x16, i8x32, i8x64, palette_lookup_u8x8, prefetch_read_t0, prefetch_read_t1, prefetch_read_t2, u16x16, u16x32, u16x8, u32x16, u32x8, - u64x4, u64x8, u8x32, u8x64, u8x8, BF16x16, BF16x8, F16x16, F32Mask16, F32Mask8, F32x16, F32x8, F64Mask4, F64Mask8, - F64x4, F64x8, I16x16, I16x32, I32x16, I32x8, I64x4, I64x8, I8x16, I8x32, I8x64, U16x16, U16x32, U16x8, U32x16, - U32x8, U64x4, U64x8, U8x32, U8x64, U8x8, + u64x4, u64x8, u8x16, u8x32, u8x64, u8x8, BF16x16, BF16x8, F16x16, F32Mask16, F32Mask8, F32x16, F32x8, F64Mask4, + F64Mask8, F64x4, F64x8, I16x16, I16x32, I32x16, I32x8, I64x4, I64x8, I8x16, I8x32, I8x64, U16x16, U16x32, U16x8, + U32x16, U32x8, U64x4, U64x8, U8x16, U8x32, U8x64, U8x8, }; #[cfg(all(target_arch = "x86_64", target_feature = "avx512f", not(feature = "nightly-simd")))] @@ -266,12 +266,13 @@ pub use crate::simd_avx512::{ u32x8, u64x4, u64x8, + u8x16, u8x64, u8x8, F32Mask16, // 512-bit (native AVX-512, __m512/__m512d/__m512i) F32x16, - // 256-bit (AVX2 baseline, __m256/__m256d/__m256i) + // 256-bit (AVX2, __m256/__m256d/__m256i) F32x8, F64Mask8, F64x4, @@ -294,6 +295,7 @@ pub use crate::simd_avx512::{ U32x8, U64x4, U64x8, + U8x16, U8x64, U8x8, }; @@ -316,7 +318,9 @@ pub use crate::simd_avx512::{f32_to_bf16_batch_rne, f32_to_bf16_scalar_rne}; #[cfg(all(target_arch = "x86_64", target_feature = "avx512bf16", not(feature = "nightly-simd")))] pub use crate::simd_avx512::{BF16x16, BF16x8}; -// AVX2 baseline arm — selected by the `x86-64-v3` cargo default. The +// AVX2 arm — selected whenever `avx512f` is not compiled in (a +// `.cargo/config-v3.toml` build, or the default `target-cpu=native` on a host +// without AVX-512). The // predicate is `not(avx512f)` rather than `avx2 + not(avx512f)` so that // an x86-64 baseline build (e.g. a `RUSTFLAGS` env that REPLACES the // `.cargo/config.toml` target-cpu pin) still has a matching arm and @@ -330,8 +334,8 @@ pub use crate::simd_avx512::{BF16x16, BF16x8}; // backend for one compile-time target, so a per-function feature gate is // a second, contradictory selection mechanism. The intrinsic calls in // `simd_avx2.rs` sit inside narrow `unsafe` blocks whose SAFETY -// precondition is the v3 baseline `.cargo/config.toml` pins for every -// x86_64 build; a baseline build compiles this arm but is not a supported +// precondition is AVX2 at run time, a caller obligation (the default +// `target-cpu=native` names no tier); a baseline build compiles this arm but is not a supported // execution target for it (it would SIGILL on the first `vp*` — the // PR #170 failure mode the config pin exists to prevent). #[cfg(all( @@ -341,7 +345,7 @@ pub use crate::simd_avx512::{BF16x16, BF16x8}; ))] pub use crate::simd_avx512::{ batch_packed_i4_16, f32x8, f64x4, i16x16, i8x16, i8x32, palette_lookup_u8x8, prefetch_read_t0, prefetch_read_t1, - prefetch_read_t2, u16x8, u8x8, F32x8, F64x4, I16x16, I8x16, I8x32, U16x8, U8x8, + prefetch_read_t2, u16x8, u8x16, u8x8, F32x8, F64x4, I16x16, I8x16, I8x32, U16x8, U8x16, U8x8, }; #[cfg(all( @@ -385,8 +389,8 @@ pub use crate::simd_neon::aarch64_simd::{f32x16, f64x8, F32Mask16, F32x16, F64Ma // W1a NEON-native types + free functions #[cfg(all(target_arch = "aarch64", not(feature = "nightly-simd")))] pub use crate::simd_neon::{ - batch_packed_i4_16, i8x16, i8x32, palette_lookup_u8x8, prefetch_read_t0, prefetch_read_t1, prefetch_read_t2, u8x8, - I8x16, I8x32, U8x8, + batch_packed_i4_16, i8x16, i8x32, palette_lookup_u8x8, prefetch_read_t0, prefetch_read_t1, prefetch_read_t2, u8x16, + u8x8, I8x16, I8x32, U8x16, U8x8, }; // U16x8 on aarch64 comes from simd_neon (backed by uint16x8_t) #[cfg(all(target_arch = "aarch64", not(feature = "nightly-simd")))] @@ -419,7 +423,8 @@ pub use scalar::{ // so this arm is gated identically. #[cfg(all(target_arch = "wasm32", target_feature = "simd128", not(feature = "nightly-simd")))] pub use crate::simd_wasm::wasm32_simd::{ - f32x16, f64x8, i32x16, i8x16, u32x16, u64x8, F32Mask16, F32x16, F64Mask8, F64x8, I32x16, I8x16, U32x16, U64x8, + f32x16, f64x8, i32x16, i8x16, u32x16, u64x8, u8x16, F32Mask16, F32x16, F64Mask8, F64x8, I32x16, I8x16, U32x16, + U64x8, U8x16, }; // `u32x16`/`U32x16`, `i32x16`/`I32x16` and `u64x8`/`U64x8` come from the // native `wasm32_simd` arm above (the lowercase alias travels with its type — @@ -442,8 +447,8 @@ pub use scalar::{ pub use scalar::{ batch_packed_i4_16, f32x16, f32x8, f64x4, f64x8, i16x16, i16x32, i32x16, i32x8, i64x4, i64x8, i8x16, i8x32, i8x64, palette_lookup_u8x8, prefetch_read_t0, prefetch_read_t1, prefetch_read_t2, u16x16, u16x8, u32x16, u32x8, u64x4, - u64x8, u8x64, u8x8, F32Mask16, F32x16, F32x8, F64Mask8, F64x4, F64x8, I16x16, I16x32, I32x16, I32x8, I64x4, I64x8, - I8x16, I8x32, I8x64, U16x16, U16x32, U16x8, U32x16, U32x8, U64x4, U64x8, U8x64, U8x8, + u64x8, u8x16, u8x64, u8x8, F32Mask16, F32x16, F32x8, F64Mask8, F64x4, F64x8, I16x16, I16x32, I32x16, I32x8, I64x4, + I64x8, I8x16, I8x32, I8x64, U16x16, U16x32, U16x8, U32x16, U32x8, U64x4, U64x8, U8x16, U8x64, U8x8, }; // Scalar BF16 conversion — always available on all platforms diff --git a/src/simd_avx2.rs b/src/simd_avx2.rs index 1364b722..13adc6ae 100644 --- a/src/simd_avx2.rs +++ b/src/simd_avx2.rs @@ -1654,9 +1654,10 @@ impl U64x8 { /// The two 256-bit halves of the 64-byte-aligned array, loaded once. #[inline(always)] fn avx2_halves(self) -> (__m256i, __m256i) { - // SAFETY: this file is the x86-64-v3 backend. `.cargo/config.toml` - // pins `-Ctarget-cpu=x86-64-v3` for the SUPPORTED x86_64 builds that - // select this arm, but that pin is not enforced by the arm's cfg — + // SAFETY: this file is the AVX2 backend. `simd.rs` selects it whenever + // `avx512f` is not compiled in: a `.cargo/config-v3.toml` build, or the + // default `target-cpu=native` on a host without AVX-512. That choice + // is not enforced by the arm's cfg — // a build whose RUSTFLAGS replaced the config compiles this arm too // and is "not a supported execution target for it (it would SIGILL)", // as `simd.rs`'s arm note says. So the obligation is the CALLER's: @@ -1819,8 +1820,8 @@ impl Shl for U64x8 { debug_assert!(rhs.to_array().iter().all(|&n| n < 64), "U64x8 shift counts are a caller contract: < 64"); let (lo, hi) = self.avx2_halves(); let (clo, chi) = rhs.avx2_halves(); - // SAFETY: same obligation as `avx2_halves` — this is the x86-64-v3 - // arm, AVX2 is present on any host that runs it; `_mm256_sllv_epi64` + // SAFETY: same obligation as `avx2_halves` — AVX2 must be present at + // run time, the caller's obligation, not a compile-time guarantee; `_mm256_sllv_epi64` // is an AVX2 instruction operating on the register values only. unsafe { Self::from_avx2_halves(_mm256_sllv_epi64(lo, clo), _mm256_sllv_epi64(hi, chi)) } } @@ -2101,7 +2102,8 @@ impl U16x16 { /// Logical right shift each 16-bit lane by `imm` (matches `U16x32::shr`). #[inline(always)] pub fn shr(self, imm: u32) -> Self { - // SAFETY: AVX2 baseline; `_mm256_srl_epi16` takes a runtime lane count + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`); `_mm256_srl_epi16` takes a runtime lane count // from the low 64 bits of an xmm, so every shift amount works (the // earlier `match {1,2,4,8}` returned zero for all other amounts). Self(unsafe { _mm256_srl_epi16(self.0, _mm_cvtsi32_si128(imm as i32)) }) @@ -2110,7 +2112,8 @@ impl U16x16 { /// Logical left shift each 16-bit lane by `imm` (matches `U16x32::shl`). #[inline(always)] pub fn shl(self, imm: u32) -> Self { - // SAFETY: AVX2 baseline; `_mm256_sll_epi16` takes a runtime lane count + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`); `_mm256_sll_epi16` takes a runtime lane count // (same fix as `shr` — the `match {1,2,4,8}` zeroed all other amounts). Self(unsafe { _mm256_sll_epi16(self.0, _mm_cvtsi32_si128(imm as i32)) }) } @@ -2135,7 +2138,8 @@ impl U16x16 { /// partial sums into add-alignment. #[inline(always)] pub fn permute2x128(self, other: Self) -> Self { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). Self(unsafe { _mm256_permute2x128_si256::(self.0, other.0) }) } @@ -2144,7 +2148,8 @@ impl U16x16 { /// lane combine (with `IMM=0xF0`). #[inline(always)] pub fn blend_epi32(self, other: Self) -> Self { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). Self(unsafe { _mm256_blend_epi32::(self.0, other.0) }) } @@ -2154,14 +2159,16 @@ impl U16x16 { /// per-query `scale·partial` FMA. #[inline(always)] pub fn to_f32x8_lo(self) -> crate::simd_avx512::F32x8 { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). crate::simd_avx512::F32x8(unsafe { _mm256_cvtepi32_ps(_mm256_cvtepu16_epi32(_mm256_castsi256_si128(self.0))) }) } /// Zero-extend the high 8 × u16 lanes to f32 (sibling of `to_f32x8_lo`). #[inline(always)] pub fn to_f32x8_hi(self) -> crate::simd_avx512::F32x8 { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). crate::simd_avx512::F32x8(unsafe { _mm256_cvtepi32_ps(_mm256_cvtepu16_epi32(_mm256_extracti128_si256::<1>(self.0))) }) @@ -2586,8 +2593,8 @@ impl I32x16 { /// The two 256-bit halves of the 64-byte-aligned array, loaded once. #[inline(always)] fn avx2_halves(self) -> (__m256i, __m256i) { - // SAFETY: x86-64-v3 backend, AVX2 is a compile-time property (see - // `U64x8::avx2_halves`); the array is 64 bytes, both loads in bounds. + // SAFETY: AVX2 is a run-time precondition of this arm, the caller's + // obligation (see `U64x8::avx2_halves`); the array is 64 bytes, both loads in bounds. unsafe { let p = self.0.as_ptr() as *const __m256i; (_mm256_loadu_si256(p), _mm256_loadu_si256(p.add(1))) @@ -2794,8 +2801,8 @@ impl I64x8 { // that want REAL AVX2 SIMD speedup over scalar should chunk their data // in 32-byte windows and use U8x32. // -// Requires AVX2 at compile time (project baseline is x86-64-v3, so this -// holds on every supported build). Calling these methods on a baseline +// Requires AVX2 at run time. This arm is compiled into every x86_64 build +// (see `simd.rs`), so the obligation is the caller's. Calling these methods on a baseline // x86_64 build (no AVX2) would SIGILL — same constraint as the rest of // the file's `_mm256_*` users (e.g. the AVX2 popcount at line ~357). // ═══════════════════════════════════════════════════════════════════ @@ -2823,8 +2830,8 @@ impl U8x32 { /// Broadcast a single byte to all 32 lanes. #[inline(always)] pub fn splat(v: u8) -> Self { - // SAFETY: AVX2 is the project baseline (x86-64-v3); calling - // `_mm256_set1_epi8` requires AVX, which AVX2 implies. + // SAFETY: AVX2 must be present at run time (caller obligation, see + // `U64x8::avx2_halves`); `_mm256_set1_epi8` needs AVX, which AVX2 implies. Self(unsafe { _mm256_set1_epi8(v as i8) }) } @@ -2903,7 +2910,8 @@ impl U8x32 { /// counting set bits in popcount-style masks. #[inline(always)] pub fn sum_bytes_u64(self) -> u64 { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). let sums = unsafe { _mm256_sad_epu8(self.0, _mm256_setzero_si256()) }; // sad_epu8 places 4 partial sums (one per 64-bit lane) in u16 slots. // Pull them out and add manually — small N, scalar is fine. @@ -2917,14 +2925,16 @@ impl U8x32 { /// Lane-wise unsigned min. #[inline(always)] pub fn simd_min(self, other: Self) -> Self { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). Self(unsafe { _mm256_min_epu8(self.0, other.0) }) } /// Lane-wise unsigned max. #[inline(always)] pub fn simd_max(self, other: Self) -> Self { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). Self(unsafe { _mm256_max_epu8(self.0, other.0) }) } @@ -2935,7 +2945,8 @@ impl U8x32 { /// at the natural AVX2 width.) #[inline(always)] pub fn cmpeq_mask(self, other: Self) -> u32 { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). let eq = unsafe { _mm256_cmpeq_epi8(self.0, other.0) }; // movemask_epi8 extracts the MSB of each byte. After cmpeq, each // lane is 0xFF (match) or 0x00 (mismatch); MSB matches what we want. @@ -2948,7 +2959,8 @@ impl U8x32 { /// ordering for unsigned compare). #[inline(always)] pub fn cmpgt_mask(self, other: Self) -> u32 { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). unsafe { let bias = _mm256_set1_epi8(i8::MIN); // 0x80 let a_s = _mm256_xor_si256(self.0, bias); @@ -2962,7 +2974,8 @@ impl U8x32 { /// `U8x64::movemask` at AVX2 width). #[inline(always)] pub fn movemask(self) -> u32 { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). unsafe { _mm256_movemask_epi8(self.0) as u32 } } @@ -2971,21 +2984,24 @@ impl U8x32 { /// Per-lane saturating unsigned add: `min(a + b, 255)`. #[inline(always)] pub fn saturating_add(self, other: Self) -> Self { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). Self(unsafe { _mm256_adds_epu8(self.0, other.0) }) } /// Per-lane saturating unsigned sub: `max(a - b, 0)`. #[inline(always)] pub fn saturating_sub(self, other: Self) -> Self { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). Self(unsafe { _mm256_subs_epu8(self.0, other.0) }) } /// Per-lane unsigned rounded average: `(a + b + 1) >> 1`. #[inline(always)] pub fn pairwise_avg(self, other: Self) -> Self { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). Self(unsafe { _mm256_avg_epu8(self.0, other.0) }) } @@ -2995,7 +3011,8 @@ impl U8x32 { /// 8-bit shift; 16-bit shift + mask is the standard idiom.) #[inline(always)] pub fn shr_epi16(self, imm: u32) -> Self { - // SAFETY: AVX2 baseline. `imm` is an arbitrary count; we use the + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). `imm` is an arbitrary count; we use the // vector-count form to avoid the const-generic constraint. Self(unsafe { _mm256_srl_epi16(self.0, _mm_cvtsi32_si128(imm as i32)) }) } @@ -3003,7 +3020,8 @@ impl U8x32 { /// Left shift each 16-bit lane by `imm` bits. #[inline(always)] pub fn shl_epi16(self, imm: u32) -> Self { - // SAFETY: AVX2 baseline. Vector-count form (see shr_epi16). + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). Vector-count form (see shr_epi16). Self(unsafe { _mm256_sll_epi16(self.0, _mm_cvtsi32_si128(imm as i32)) }) } @@ -3016,7 +3034,8 @@ impl U8x32 { /// `permute_bytes` for that, which falls back to scalar.) #[inline(always)] pub fn shuffle_bytes(self, idx: Self) -> Self { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). Self(unsafe { _mm256_shuffle_epi8(self.0, idx.0) }) } @@ -3041,14 +3060,16 @@ impl U8x32 { /// Output: `[a0,b0, a1,b1, ..., a7,b7]` within each 128-bit half. #[inline(always)] pub fn unpack_lo_epi8(self, other: Self) -> Self { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). Self(unsafe { _mm256_unpacklo_epi8(self.0, other.0) }) } /// Interleave high 8 bytes of each 128-bit half (`_mm256_unpackhi_epi8`). #[inline(always)] pub fn unpack_hi_epi8(self, other: Self) -> Self { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). Self(unsafe { _mm256_unpackhi_epi8(self.0, other.0) }) } @@ -3060,7 +3081,8 @@ impl U8x32 { /// 64-bit-bitmask shape of `U8x64::mask_blend`). #[inline(always)] pub fn mask_blend(mask: Self, a: Self, b: Self) -> Self { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). Self(unsafe { _mm256_blendv_epi8(b.0, a.0, mask.0) }) } @@ -3096,7 +3118,8 @@ impl core::ops::BitAnd for U8x32 { type Output = Self; #[inline(always)] fn bitand(self, rhs: Self) -> Self { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). Self(unsafe { _mm256_and_si256(self.0, rhs.0) }) } } @@ -3106,7 +3129,8 @@ impl core::ops::BitOr for U8x32 { type Output = Self; #[inline(always)] fn bitor(self, rhs: Self) -> Self { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). Self(unsafe { _mm256_or_si256(self.0, rhs.0) }) } } @@ -3116,7 +3140,8 @@ impl core::ops::BitXor for U8x32 { type Output = Self; #[inline(always)] fn bitxor(self, rhs: Self) -> Self { - // SAFETY: AVX2 baseline. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). Self(unsafe { _mm256_xor_si256(self.0, rhs.0) }) } } @@ -3126,7 +3151,8 @@ impl core::ops::Add for U8x32 { type Output = Self; #[inline(always)] fn add(self, rhs: Self) -> Self { - // SAFETY: AVX2 baseline. WRAPS — use saturating_add for clamp. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). WRAPS — use saturating_add for clamp. Self(unsafe { _mm256_add_epi8(self.0, rhs.0) }) } } @@ -3136,7 +3162,8 @@ impl core::ops::Sub for U8x32 { type Output = Self; #[inline(always)] fn sub(self, rhs: Self) -> Self { - // SAFETY: AVX2 baseline. WRAPS — use saturating_sub for clamp. + // SAFETY: AVX2 must be present at run time (caller obligation, not a + // compile-time guarantee; see `U64x8::avx2_halves`). WRAPS — use saturating_sub for clamp. Self(unsafe { _mm256_sub_epi8(self.0, rhs.0) }) } } diff --git a/src/simd_avx512.rs b/src/simd_avx512.rs index a8d120b1..4ff2aaf9 100644 --- a/src/simd_avx512.rs +++ b/src/simd_avx512.rs @@ -2778,6 +2778,74 @@ impl I8x16 { s[..16].copy_from_slice(&self.0); } + /// All 16 lanes set to `0`. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::I8x16; + /// assert_eq!(I8x16::zero().to_array(), [0i8; 16]); + /// ``` + #[inline(always)] + pub fn zero() -> Self { + Self([0i8; 16]) + } + + /// Lane-wise **wrapping** addition (matches NEON `vaddq_s8`). + /// + /// Overflow wraps: `i8::MAX + 1 == i8::MIN`. Not saturating. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::I8x16; + /// let r = I8x16::splat(i8::MAX).add(I8x16::splat(1)); + /// assert_eq!(r.to_array(), [i8::MIN; 16]); + /// ``` + #[inline(always)] + pub fn add(self, other: Self) -> Self { + Self(core::array::from_fn(|i| self.0[i].wrapping_add(other.0[i]))) + } + + /// Lane-wise **wrapping** subtraction (matches NEON `vsubq_s8`). + /// + /// Overflow wraps: `i8::MIN - 1 == i8::MAX`. Not saturating. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::I8x16; + /// let r = I8x16::splat(i8::MIN).sub(I8x16::splat(1)); + /// assert_eq!(r.to_array(), [i8::MAX; 16]); + /// ``` + #[inline(always)] + pub fn sub(self, other: Self) -> Self { + Self(core::array::from_fn(|i| self.0[i].wrapping_sub(other.0[i]))) + } + + /// Lane-wise signed minimum (matches NEON `vminq_s8`). + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::I8x16; + /// let r = I8x16::splat(-5).min(I8x16::splat(3)); + /// assert_eq!(r.to_array(), [-5i8; 16]); + /// ``` + #[inline(always)] + pub fn min(self, other: Self) -> Self { + Self(core::array::from_fn(|i| self.0[i].min(other.0[i]))) + } + + /// Lane-wise signed maximum (matches NEON `vmaxq_s8`). + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::I8x16; + /// let r = I8x16::splat(-5).max(I8x16::splat(3)); + /// assert_eq!(r.to_array(), [3i8; 16]); + /// ``` + #[inline(always)] + pub fn max(self, other: Self) -> Self { + Self(core::array::from_fn(|i| self.0[i].max(other.0[i]))) + } + // ── W1a-#1: from_i4_packed_u64 + lane_i8 ──────────────────────────────── /// Unpack 16 signed i4 nibbles from a `u64` into 16 sign-extended `i8` lanes. @@ -2846,9 +2914,12 @@ impl I8x16 { /// ``` #[inline(always)] pub fn saturating_abs(self) -> Self { - // SAFETY: `_mm_abs_epi8` (SSSE3) and `_mm_min_epu8` (SSE2) are available - // on every x86_64 build this file compiles for — the workspace pins - // `x86-64-v3`, which includes SSSE3. The unaligned load/store match the + // SAFETY: `_mm_abs_epi8` needs SSSE3 and `_mm_min_epu8` needs SSE2. SSE2 + // is in the x86_64 baseline; SSSE3 is NOT, and this file compiles for + // every x86_64 build, so SSSE3 at run time is the caller's obligation. + // `target-cpu=native` (default) and `config-v3`/`config-v4` include it; + // a baseline build (e.g. a RUSTFLAGS env replacing the config) compiles + // this and would SIGILL. The unaligned load/store match the // `[i8; 16]` storage. VPABSB returns 0x80 for `i8::MIN` (the bit pattern // of +128, which does not fit in i8); VPMINUB then clamps 0x80 (= 128 // unsigned) down to 0x7f (= 127 = `i8::MAX`), producing the saturating @@ -2872,6 +2943,170 @@ impl core::fmt::Debug for I8x16 { } } +// ─── U8x16 (scalar-storage polyfill for the AVX-512 backend) ───────────────── + +/// 16-lane `u8` vector. On the AVX-512 backend this is a scalar-storage +/// polyfill; on NEON it is backed by `uint8x16_t`. +/// +/// Edge cases and lane layout are identical across backends; only performance +/// differs. `add`/`sub` wrap on overflow, `min`/`max` are unsigned. +#[cfg(target_arch = "x86_64")] +#[derive(Copy, Clone, PartialEq)] +#[repr(align(16))] +pub struct U8x16(pub [u8; 16]); + +#[cfg(target_arch = "x86_64")] +impl U8x16 { + pub const LANES: usize = 16; + + /// Broadcast a single `u8` value to all 16 lanes. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// let v = U8x16::splat(3); + /// assert!(v.to_array().iter().all(|&x| x == 3)); + /// ``` + #[inline(always)] + pub fn splat(v: u8) -> Self { + Self([v; 16]) + } + + /// All 16 lanes set to `0`. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// assert_eq!(U8x16::zero().to_array(), [0u8; 16]); + /// ``` + #[inline(always)] + pub fn zero() -> Self { + Self([0u8; 16]) + } + + /// Load 16 lanes from the first 16 elements of a slice (at least 16 + /// elements required; panics otherwise). + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// let src: Vec = (0..20).collect(); + /// assert_eq!(U8x16::from_slice(&src).to_array()[15], 15); + /// ``` + #[inline(always)] + pub fn from_slice(s: &[u8]) -> Self { + assert!(s.len() >= 16); + let mut a = [0u8; 16]; + a.copy_from_slice(&s[..16]); + Self(a) + } + + /// Load from a fixed-size array. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// let a = [7u8; 16]; + /// assert_eq!(U8x16::from_array(a).to_array(), a); + /// ``` + #[inline(always)] + pub fn from_array(arr: [u8; 16]) -> Self { + Self(arr) + } + + /// Extract all 16 lanes as an array. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// assert_eq!(U8x16::splat(9).to_array(), [9u8; 16]); + /// ``` + #[inline(always)] + pub fn to_array(self) -> [u8; 16] { + self.0 + } + + /// Copy lanes into the first 16 elements of a slice (at least 16 + /// elements required; panics otherwise). + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// let mut out = [0u8; 20]; + /// U8x16::splat(4).copy_to_slice(&mut out); + /// assert_eq!(&out[..16], &[4u8; 16]); + /// assert_eq!(&out[16..], &[0u8; 4]); + /// ``` + #[inline(always)] + pub fn copy_to_slice(self, s: &mut [u8]) { + assert!(s.len() >= 16); + s[..16].copy_from_slice(&self.0); + } + + /// Lane-wise **wrapping** addition (matches NEON `vaddq_u8`). + /// + /// Overflow wraps: `255 + 1 == 0`. Not saturating. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// let r = U8x16::splat(255).add(U8x16::splat(1)); + /// assert_eq!(r.to_array(), [0u8; 16]); + /// ``` + #[inline(always)] + pub fn add(self, other: Self) -> Self { + Self(core::array::from_fn(|i| self.0[i].wrapping_add(other.0[i]))) + } + + /// Lane-wise **wrapping** subtraction (matches NEON `vsubq_u8`). + /// + /// Underflow wraps: `0 - 1 == 255`. Not saturating. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// let r = U8x16::zero().sub(U8x16::splat(1)); + /// assert_eq!(r.to_array(), [255u8; 16]); + /// ``` + #[inline(always)] + pub fn sub(self, other: Self) -> Self { + Self(core::array::from_fn(|i| self.0[i].wrapping_sub(other.0[i]))) + } + + /// Lane-wise unsigned minimum (matches NEON `vminq_u8`). + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// let r = U8x16::splat(200).min(U8x16::splat(3)); + /// assert_eq!(r.to_array(), [3u8; 16]); + /// ``` + #[inline(always)] + pub fn min(self, other: Self) -> Self { + Self(core::array::from_fn(|i| self.0[i].min(other.0[i]))) + } + + /// Lane-wise unsigned maximum (matches NEON `vmaxq_u8`). + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// let r = U8x16::splat(200).max(U8x16::splat(3)); + /// assert_eq!(r.to_array(), [200u8; 16]); + /// ``` + #[inline(always)] + pub fn max(self, other: Self) -> Self { + Self(core::array::from_fn(|i| self.0[i].max(other.0[i]))) + } +} + +#[cfg(target_arch = "x86_64")] +impl core::fmt::Debug for U8x16 { + fn fmt(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result { + write!(f, "U8x16({:?})", &self.0[..]) + } +} + // ─── U8x8 (scalar-storage polyfill for AVX-512 backend) ────────────────────── /// 8-lane `u8` vector. Scalar-storage polyfill used by `palette_lookup_u8x8`. @@ -3067,13 +3302,14 @@ impl I8x32 { /// ``` #[inline(always)] pub fn saturating_abs(self) -> Self { - // SAFETY: _mm256_abs_epi8 (VPABSB) is an AVX2 intrinsic; we are in - // the simd_avx512.rs file which is only compiled for x86_64. The - // `target_feature(enable = "avx2")` annotation on the calling code - // path guarantees AVX2 availability. The raw_abs result for 0x80 - // is 0x80 (bit-pattern +128); VPMINUB then clamps it to 0x7f. - // UNVERIFIED: _mm256_abs_epi8 stability on Rust 1.94 stable — it is - // in std::arch::x86_64 since Rust 1.0 for AVX2 so should compile. + // SAFETY: _mm256_abs_epi8 (VPABSB) and _mm256_min_epu8 are AVX2 + // intrinsics. There is no `#[target_feature]` annotation on any caller + // (an earlier version of this comment claimed one), and this file + // compiles for every x86_64 build, so AVX2 at run time is the caller's + // obligation — the same footing as every other `I8x32` method, which + // all use `_mm256_*` (see `U64x8::avx2_halves` in simd_avx2.rs). The + // raw_abs result for 0x80 is 0x80 (bit-pattern +128); VPMINUB then + // clamps it to 0x7f. #[cfg(target_arch = "x86_64")] unsafe { let raw_abs = core::arch::x86_64::_mm256_abs_epi8(self.0); @@ -3295,6 +3531,9 @@ where pub type i8x16 = I8x16; #[cfg(target_arch = "x86_64")] #[allow(non_camel_case_types)] +pub type u8x16 = U8x16; +#[cfg(target_arch = "x86_64")] +#[allow(non_camel_case_types)] pub type u16x8 = U16x8; #[cfg(target_arch = "x86_64")] #[allow(non_camel_case_types)] diff --git a/src/simd_neon.rs b/src/simd_neon.rs index b95b4e75..c0b7ed30 100644 --- a/src/simd_neon.rs +++ b/src/simd_neon.rs @@ -1354,11 +1354,23 @@ impl I8x16 { } /// Compare-greater-than: returns 16-bit mask. Bit i set where self[i] > other[i]. + /// + /// Register-level compare: it takes a second register, so it is binary in form. + /// The masking-ops predicates call register compares like this one with a + /// broadcast constant; lane-vs-lane predicates (G7) stay deliberately absent at + /// the slice/IR layer, see `.claude/knowledge/masking-ops-state.md` § G7. This + /// method is not that gap. The same operation exists on every backend at its + /// native widths (`I8x64::cmp_gt` etc.); an `I8x16` alias is not to be added to + /// other arms without a caller. #[inline(always)] pub fn cmp_gt(self, other: Self) -> u16 { + let mut arr = [0u8; 16]; + // SAFETY: NEON is baseline on aarch64, so the intrinsics are available; + // `arr` is a local 16-byte array, so the 16-byte store is in bounds; + // `vcgtq_s8` lanes are 0x00/0xFF. unsafe { let cmp = vcgtq_s8(self.0, other.0); // uint8x16_t, 0xFF where true - let arr: [u8; 16] = core::mem::transmute(cmp); + vst1q_u8(arr.as_mut_ptr(), cmp); let mut m: u16 = 0; for i in 0..16 { if arr[i] != 0 { @@ -1444,11 +1456,23 @@ impl I16x8 { } /// Compare-greater-than: returns 8-bit mask. Bit i set where self[i] > other[i]. + /// + /// Register-level compare: it takes a second register, so it is binary in form. + /// The masking-ops predicates call register compares like this one with a + /// broadcast constant; lane-vs-lane predicates (G7) stay deliberately absent at + /// the slice/IR layer, see `.claude/knowledge/masking-ops-state.md` § G7. This + /// method is not that gap. The same operation exists on every backend at its + /// native widths (`I8x64::cmp_gt` etc.); an `I8x16` alias is not to be added to + /// other arms without a caller. #[inline(always)] pub fn cmp_gt(self, other: Self) -> u8 { + let mut arr = [0u16; 8]; + // SAFETY: NEON is baseline on aarch64, so the intrinsics are available; + // `arr` is a local 16-byte array, so the 16-byte store is in bounds; + // `vcgtq_s16` lanes are 0x0000/0xFFFF. unsafe { let cmp = vcgtq_s16(self.0, other.0); // uint16x8_t, 0xFFFF where true - let arr: [u16; 8] = core::mem::transmute(cmp); + vst1q_u16(arr.as_mut_ptr(), cmp); let mut m: u8 = 0; for i in 0..8 { if arr[i] != 0 { @@ -2384,6 +2408,9 @@ neon_int_polyfill!(I16x32, i16, 32, 0i16, u32); pub type i8x16 = I8x16; #[cfg(target_arch = "aarch64")] #[allow(non_camel_case_types)] +pub type u8x16 = U8x16; +#[cfg(target_arch = "aarch64")] +#[allow(non_camel_case_types)] pub type i16x8 = I16x8; #[cfg(target_arch = "aarch64")] #[allow(non_camel_case_types)] diff --git a/src/simd_nightly/mod.rs b/src/simd_nightly/mod.rs index 3e0bd832..321b6a7e 100644 --- a/src/simd_nightly/mod.rs +++ b/src/simd_nightly/mod.rs @@ -45,7 +45,8 @@ pub use masks::{F32Mask16, F32Mask8, F64Mask4, F64Mask8}; pub use u8_types::{U8x32, U8x64}; pub use u_word_types::{U16x16, U16x32, U32x16, U32x8, U64x4, U64x8}; pub use w1a_types::{ - batch_packed_i4_16, palette_lookup_u8x8, prefetch_read_t0, prefetch_read_t1, prefetch_read_t2, I8x16, U16x8, U8x8, + batch_packed_i4_16, palette_lookup_u8x8, prefetch_read_t0, prefetch_read_t1, prefetch_read_t2, I8x16, U16x8, U8x16, + U8x8, }; // Lowercase aliases — match the std::simd convention used by @@ -99,6 +100,8 @@ pub type i64x4 = I64x4; #[allow(non_camel_case_types)] pub type i8x16 = I8x16; #[allow(non_camel_case_types)] +pub type u8x16 = U8x16; +#[allow(non_camel_case_types)] pub type u16x8 = U16x8; #[allow(non_camel_case_types)] pub type u8x8 = U8x8; diff --git a/src/simd_nightly/w1a_types.rs b/src/simd_nightly/w1a_types.rs index 0bd44c95..8e7381ef 100644 --- a/src/simd_nightly/w1a_types.rs +++ b/src/simd_nightly/w1a_types.rs @@ -20,7 +20,7 @@ use core::fmt; use core::simd::cmp::{SimdOrd, SimdPartialEq, SimdPartialOrd}; use core::simd::num::{SimdInt, SimdUint}; -use core::simd::{i8x16 as core_i8x16, u16x8 as core_u16x8, u64x16, u8x8 as core_u8x8, Simd}; +use core::simd::{i8x16 as core_i8x16, u16x8 as core_u16x8, u64x16, u8x16 as core_u8x16, u8x8 as core_u8x8, Simd}; // ── W1a-#1: I8x16 + lane_i8 + from_i4_packed_u64 ──────────────────────────── @@ -78,6 +78,80 @@ impl I8x16 { self.0.copy_to_slice(&mut s[..16]); } + /// All lanes zero. + /// + /// # Examples + /// ```rust + /// # #[cfg(feature = "nightly-simd")] { + /// use ndarray::simd_nightly::I8x16; + /// assert_eq!(I8x16::zero().to_array(), [0i8; 16]); + /// # } + /// ``` + #[inline(always)] + pub fn zero() -> Self { + Self(core_i8x16::splat(0)) + } + + /// Lane-wise **wrapping** addition (`i8::MAX + 1 == i8::MIN`), matching `vaddq_s8`. + /// + /// # Examples + /// ```rust + /// # #[cfg(feature = "nightly-simd")] { + /// use ndarray::simd_nightly::I8x16; + /// assert_eq!(I8x16::splat(i8::MAX).add(I8x16::splat(1)).to_array(), [i8::MIN; 16]); + /// assert_eq!(I8x16::splat(2).add(I8x16::splat(-5)).to_array(), [-3i8; 16]); + /// # } + /// ``` + #[inline(always)] + #[allow(clippy::should_implement_trait)] + pub fn add(self, other: Self) -> Self { + Self(self.0 + other.0) + } + + /// Lane-wise **wrapping** subtraction (`i8::MIN - 1 == i8::MAX`), matching `vsubq_s8`. + /// + /// # Examples + /// ```rust + /// # #[cfg(feature = "nightly-simd")] { + /// use ndarray::simd_nightly::I8x16; + /// assert_eq!(I8x16::splat(i8::MIN).sub(I8x16::splat(1)).to_array(), [i8::MAX; 16]); + /// assert_eq!(I8x16::splat(2).sub(I8x16::splat(5)).to_array(), [-3i8; 16]); + /// # } + /// ``` + #[inline(always)] + #[allow(clippy::should_implement_trait)] + pub fn sub(self, other: Self) -> Self { + Self(self.0 - other.0) + } + + /// Lane-wise signed minimum (same result as [`Self::simd_min`]). + /// + /// # Examples + /// ```rust + /// # #[cfg(feature = "nightly-simd")] { + /// use ndarray::simd_nightly::I8x16; + /// assert_eq!(I8x16::splat(-5).min(I8x16::splat(3)).to_array(), [-5i8; 16]); + /// # } + /// ``` + #[inline(always)] + pub fn min(self, other: Self) -> Self { + Self(self.0.simd_min(other.0)) + } + + /// Lane-wise signed maximum (same result as [`Self::simd_max`]). + /// + /// # Examples + /// ```rust + /// # #[cfg(feature = "nightly-simd")] { + /// use ndarray::simd_nightly::I8x16; + /// assert_eq!(I8x16::splat(-5).max(I8x16::splat(3)).to_array(), [3i8; 16]); + /// # } + /// ``` + #[inline(always)] + pub fn max(self, other: Self) -> Self { + Self(self.0.simd_max(other.0)) + } + /// Unpack 16 signed i4 nibbles from a `u64` into 16 sign-extended `i8` /// lanes: `lane[i] = sign_extend_i4((packed >> (4*i)) & 0xf)`, so /// `0x0..=0x7 → 0..=7` and `0x8..=0xf → -8..=-1`. @@ -195,6 +269,195 @@ impl fmt::Debug for I8x16 { } } +// ── U8x16 (NEON `simd_neon::U8x16` surface parity) ────────────────────────── + +/// 16-lane `u8` vector backed by `core::simd::u8x16`. +/// +/// Mirrors `simd_neon::U8x16`: `add`/`sub` wrap on overflow (`vaddq_u8` / +/// `vsubq_u8` semantics), `min`/`max` are unsigned lane-wise. +/// +/// # Examples +/// ```rust +/// # #[cfg(feature = "nightly-simd")] { +/// use ndarray::simd_nightly::w1a_types::U8x16; +/// let a = U8x16::splat(250); +/// assert_eq!(a.add(U8x16::splat(10)).to_array(), [4u8; 16]); +/// assert_eq!(U8x16::zero().sub(U8x16::splat(1)).to_array(), [255u8; 16]); +/// # } +/// ``` +#[derive(Copy, Clone)] +#[repr(transparent)] +pub struct U8x16(pub core_u8x16); + +impl U8x16 { + /// Number of `u8` lanes. + pub const LANES: usize = 16; + + /// Broadcast a single `u8` value to all 16 lanes. + /// + /// # Examples + /// ```rust + /// # #[cfg(feature = "nightly-simd")] { + /// use ndarray::simd_nightly::w1a_types::U8x16; + /// assert_eq!(U8x16::splat(7).to_array(), [7u8; 16]); + /// # } + /// ``` + #[inline(always)] + pub fn splat(v: u8) -> Self { + Self(core_u8x16::splat(v)) + } + + /// All lanes zero. + /// + /// # Examples + /// ```rust + /// # #[cfg(feature = "nightly-simd")] { + /// use ndarray::simd_nightly::w1a_types::U8x16; + /// assert_eq!(U8x16::zero().to_array(), [0u8; 16]); + /// # } + /// ``` + #[inline(always)] + pub fn zero() -> Self { + Self(core_u8x16::splat(0)) + } + + /// Load the first 16 elements of a slice (panics if `s.len() < 16`). + /// + /// # Examples + /// ```rust + /// # #[cfg(feature = "nightly-simd")] { + /// use ndarray::simd_nightly::w1a_types::U8x16; + /// let s: Vec = (0..20).collect(); + /// assert_eq!(U8x16::from_slice(&s).to_array()[15], 15); + /// # } + /// ``` + #[inline(always)] + pub fn from_slice(s: &[u8]) -> Self { + assert!(s.len() >= 16); + Self(core_u8x16::from_slice(&s[..16])) + } + + /// Load from a fixed-size array. + /// + /// # Examples + /// ```rust + /// # #[cfg(feature = "nightly-simd")] { + /// use ndarray::simd_nightly::w1a_types::U8x16; + /// let a: [u8; 16] = core::array::from_fn(|i| i as u8); + /// assert_eq!(U8x16::from_array(a).to_array(), a); + /// # } + /// ``` + #[inline(always)] + pub fn from_array(arr: [u8; 16]) -> Self { + Self(core_u8x16::from_array(arr)) + } + + /// Extract all 16 lanes as an array. + /// + /// # Examples + /// ```rust + /// # #[cfg(feature = "nightly-simd")] { + /// use ndarray::simd_nightly::w1a_types::U8x16; + /// assert_eq!(U8x16::splat(3).to_array(), [3u8; 16]); + /// # } + /// ``` + #[inline(always)] + pub fn to_array(self) -> [u8; 16] { + self.0.to_array() + } + + /// Copy the 16 lanes into the first 16 elements of a slice + /// (panics if `s.len() < 16`). + /// + /// # Examples + /// ```rust + /// # #[cfg(feature = "nightly-simd")] { + /// use ndarray::simd_nightly::w1a_types::U8x16; + /// let mut out = [0u8; 18]; + /// U8x16::splat(9).copy_to_slice(&mut out); + /// assert_eq!(&out[..16], &[9u8; 16]); + /// assert_eq!(out[16], 0); + /// # } + /// ``` + #[inline(always)] + pub fn copy_to_slice(self, s: &mut [u8]) { + assert!(s.len() >= 16); + self.0.copy_to_slice(&mut s[..16]); + } + + /// Lane-wise **wrapping** addition (`255 + 1 == 0`), matching `vaddq_u8`. + /// + /// # Examples + /// ```rust + /// # #[cfg(feature = "nightly-simd")] { + /// use ndarray::simd_nightly::w1a_types::U8x16; + /// assert_eq!(U8x16::splat(255).add(U8x16::splat(1)).to_array(), [0u8; 16]); + /// assert_eq!(U8x16::splat(2).add(U8x16::splat(3)).to_array(), [5u8; 16]); + /// # } + /// ``` + #[inline(always)] + #[allow(clippy::should_implement_trait)] + pub fn add(self, other: Self) -> Self { + Self(self.0 + other.0) + } + + /// Lane-wise **wrapping** subtraction (`0 - 1 == 255`), matching `vsubq_u8`. + /// + /// # Examples + /// ```rust + /// # #[cfg(feature = "nightly-simd")] { + /// use ndarray::simd_nightly::w1a_types::U8x16; + /// assert_eq!(U8x16::zero().sub(U8x16::splat(1)).to_array(), [255u8; 16]); + /// assert_eq!(U8x16::splat(5).sub(U8x16::splat(3)).to_array(), [2u8; 16]); + /// # } + /// ``` + #[inline(always)] + #[allow(clippy::should_implement_trait)] + pub fn sub(self, other: Self) -> Self { + Self(self.0 - other.0) + } + + /// Lane-wise unsigned minimum. + /// + /// # Examples + /// ```rust + /// # #[cfg(feature = "nightly-simd")] { + /// use ndarray::simd_nightly::w1a_types::U8x16; + /// assert_eq!(U8x16::splat(200).min(U8x16::splat(100)).to_array(), [100u8; 16]); + /// # } + /// ``` + #[inline(always)] + pub fn min(self, other: Self) -> Self { + Self(self.0.simd_min(other.0)) + } + + /// Lane-wise unsigned maximum (`200 > 100`, no signed reinterpretation). + /// + /// # Examples + /// ```rust + /// # #[cfg(feature = "nightly-simd")] { + /// use ndarray::simd_nightly::w1a_types::U8x16; + /// assert_eq!(U8x16::splat(200).max(U8x16::splat(100)).to_array(), [200u8; 16]); + /// # } + /// ``` + #[inline(always)] + pub fn max(self, other: Self) -> Self { + Self(self.0.simd_max(other.0)) + } +} + +impl PartialEq for U8x16 { + fn eq(&self, other: &Self) -> bool { + self.to_array() == other.to_array() + } +} + +impl fmt::Debug for U8x16 { + fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { + write!(f, "U8x16({:?})", &self.to_array()[..]) + } +} + // ── W1a-#3: U16x8 / U8x8 / palette_lookup_u8x8 ───────────────────────────── /// 8-lane `u16` vector backed by `core::simd::u16x8`. diff --git a/src/simd_scalar.rs b/src/simd_scalar.rs index 8a88005a..4bd377c2 100644 --- a/src/simd_scalar.rs +++ b/src/simd_scalar.rs @@ -1746,6 +1746,86 @@ impl I8x16 { } Self(o) } + + /// All lanes zero. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::{I8x16, U8x16}; + /// assert_eq!(I8x16::zero().to_array(), [0; 16]); + /// ``` + #[inline(always)] + pub fn zero() -> Self { + Self([0; 16]) + } + + /// Lane-wise wrapping addition (overflow wraps, matching NEON `vaddq_s8`). + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::{I8x16, U8x16}; + /// let r = I8x16::splat(i8::MAX).add(I8x16::splat(1)); + /// assert!(r.to_array().iter().all(|&x| x == i8::MIN)); + /// ``` + #[inline(always)] + pub fn add(self, other: Self) -> Self { + let mut o = [0 as i8; 16]; + for i in 0..16 { + o[i] = self.0[i].wrapping_add(other.0[i]); + } + Self(o) + } + + /// Lane-wise wrapping subtraction (overflow wraps, matching NEON `vsubq_s8`). + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::{I8x16, U8x16}; + /// let r = I8x16::splat(i8::MIN).sub(I8x16::splat(1)); + /// assert!(r.to_array().iter().all(|&x| x == i8::MAX)); + /// ``` + #[inline(always)] + pub fn sub(self, other: Self) -> Self { + let mut o = [0 as i8; 16]; + for i in 0..16 { + o[i] = self.0[i].wrapping_sub(other.0[i]); + } + Self(o) + } + + /// Lane-wise minimum. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::{I8x16, U8x16}; + /// let r = I8x16::splat(i8::MIN).min(I8x16::splat(i8::MAX)); + /// assert!(r.to_array().iter().all(|&x| x == i8::MIN)); + /// ``` + #[inline(always)] + pub fn min(self, other: Self) -> Self { + let mut o = [0 as i8; 16]; + for i in 0..16 { + o[i] = if self.0[i] < other.0[i] { self.0[i] } else { other.0[i] }; + } + Self(o) + } + + /// Lane-wise maximum. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::{I8x16, U8x16}; + /// let r = I8x16::splat(i8::MIN).max(I8x16::splat(i8::MAX)); + /// assert!(r.to_array().iter().all(|&x| x == i8::MAX)); + /// ``` + #[inline(always)] + pub fn max(self, other: Self) -> Self { + let mut o = [0 as i8; 16]; + for i in 0..16 { + o[i] = if self.0[i] > other.0[i] { self.0[i] } else { other.0[i] }; + } + Self(o) + } } impl core::fmt::Debug for I8x16 { @@ -1754,6 +1834,170 @@ impl core::fmt::Debug for I8x16 { } } +/// 16-lane `u8` vector — scalar fallback for non-NEON, non-x86_64 targets. +/// +/// Mirrors the NEON `U8x16` surface (`simd_neon.rs`); pure safe Rust. +#[derive(Copy, Clone, PartialEq)] +#[repr(align(16))] +pub struct U8x16(pub [u8; 16]); + +impl U8x16 { + pub const LANES: usize = 16; + + /// Broadcast a single `u8` value to all 16 lanes. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::{I8x16, U8x16}; + /// assert_eq!(U8x16::splat(7).to_array(), [7u8; 16]); + /// ``` + #[inline(always)] + pub fn splat(v: u8) -> Self { + Self([v; 16]) + } + + /// Load from a slice (at least 16 elements required; first 16 used). + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::{I8x16, U8x16}; + /// let src: Vec = (0..20).collect(); + /// assert_eq!(U8x16::from_slice(&src).to_array()[15], 15); + /// ``` + #[inline(always)] + pub fn from_slice(s: &[u8]) -> Self { + assert!(s.len() >= 16); + let mut a = [0u8; 16]; + a.copy_from_slice(&s[..16]); + Self(a) + } + + /// Load from a fixed-size array. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::{I8x16, U8x16}; + /// assert_eq!(U8x16::from_array([3u8; 16]).to_array(), [3u8; 16]); + /// ``` + #[inline(always)] + pub fn from_array(arr: [u8; 16]) -> Self { + Self(arr) + } + + /// Extract all 16 lanes as an array. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::{I8x16, U8x16}; + /// assert_eq!(U8x16::splat(1).to_array(), [1u8; 16]); + /// ``` + #[inline(always)] + pub fn to_array(self) -> [u8; 16] { + self.0 + } + + /// Copy lanes into a slice (must have at least 16 elements). + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::{I8x16, U8x16}; + /// let mut out = [0u8; 16]; + /// U8x16::splat(9).copy_to_slice(&mut out); + /// assert_eq!(out, [9u8; 16]); + /// ``` + #[inline(always)] + pub fn copy_to_slice(self, s: &mut [u8]) { + assert!(s.len() >= 16); + s[..16].copy_from_slice(&self.0); + } + + /// All lanes zero. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::{I8x16, U8x16}; + /// assert_eq!(U8x16::zero().to_array(), [0; 16]); + /// ``` + #[inline(always)] + pub fn zero() -> Self { + Self([0; 16]) + } + + /// Lane-wise wrapping addition (overflow wraps, matching NEON `vaddq_u8`). + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::{I8x16, U8x16}; + /// let r = U8x16::splat(u8::MAX).add(U8x16::splat(1)); + /// assert!(r.to_array().iter().all(|&x| x == u8::MIN)); + /// ``` + #[inline(always)] + pub fn add(self, other: Self) -> Self { + let mut o = [0 as u8; 16]; + for i in 0..16 { + o[i] = self.0[i].wrapping_add(other.0[i]); + } + Self(o) + } + + /// Lane-wise wrapping subtraction (overflow wraps, matching NEON `vsubq_u8`). + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::{I8x16, U8x16}; + /// let r = U8x16::splat(u8::MIN).sub(U8x16::splat(1)); + /// assert!(r.to_array().iter().all(|&x| x == u8::MAX)); + /// ``` + #[inline(always)] + pub fn sub(self, other: Self) -> Self { + let mut o = [0 as u8; 16]; + for i in 0..16 { + o[i] = self.0[i].wrapping_sub(other.0[i]); + } + Self(o) + } + + /// Lane-wise minimum. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::{I8x16, U8x16}; + /// let r = U8x16::splat(u8::MIN).min(U8x16::splat(u8::MAX)); + /// assert!(r.to_array().iter().all(|&x| x == u8::MIN)); + /// ``` + #[inline(always)] + pub fn min(self, other: Self) -> Self { + let mut o = [0 as u8; 16]; + for i in 0..16 { + o[i] = if self.0[i] < other.0[i] { self.0[i] } else { other.0[i] }; + } + Self(o) + } + + /// Lane-wise maximum. + /// + /// # Example + /// ```rust,ignore + /// use ndarray::simd::{I8x16, U8x16}; + /// let r = U8x16::splat(u8::MIN).max(U8x16::splat(u8::MAX)); + /// assert!(r.to_array().iter().all(|&x| x == u8::MAX)); + /// ``` + #[inline(always)] + pub fn max(self, other: Self) -> Self { + let mut o = [0 as u8; 16]; + for i in 0..16 { + o[i] = if self.0[i] > other.0[i] { self.0[i] } else { other.0[i] }; + } + Self(o) + } +} + +impl core::fmt::Debug for U8x16 { + fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { + write!(f, "U8x16({:?})", &self.0[..]) + } +} + // ── W1a-#2: I8x32::saturating_abs (scalar) ─────────────────────────────────── impl I8x32 { @@ -2083,6 +2327,8 @@ where #[allow(non_camel_case_types)] pub type i8x16 = I8x16; #[allow(non_camel_case_types)] +pub type u8x16 = U8x16; +#[allow(non_camel_case_types)] pub type u16x8 = U16x8; #[allow(non_camel_case_types)] pub type u8x8 = U8x8; diff --git a/src/simd_wasm.rs b/src/simd_wasm.rs index 623cb36b..6f3622b8 100644 --- a/src/simd_wasm.rs +++ b/src/simd_wasm.rs @@ -803,10 +803,28 @@ pub mod wasm32_simd { unsafe { v128_store(s.as_mut_ptr() as *mut v128, self.0) }; } + /// Lane-wise **wrapping** add (`i8x16_add`): `i8::MAX + 1 == i8::MIN`. + /// Never saturates — matches NEON `vaddq_s8` and the scalar backend. + /// + /// # Examples + /// ```rust,ignore + /// use ndarray::simd::I8x16; + /// let r = I8x16::splat(i8::MAX).add(I8x16::splat(1)); + /// assert_eq!(r.to_array(), [i8::MIN; 16]); + /// ``` #[inline(always)] pub fn add(self, other: Self) -> Self { Self(i8x16_add(self.0, other.0)) } + /// Lane-wise **wrapping** subtract (`i8x16_sub`): `i8::MIN - 1 == i8::MAX`. + /// Never saturates — matches NEON `vsubq_s8` and the scalar backend. + /// + /// # Examples + /// ```rust,ignore + /// use ndarray::simd::I8x16; + /// let r = I8x16::splat(i8::MIN).sub(I8x16::splat(1)); + /// assert_eq!(r.to_array(), [i8::MAX; 16]); + /// ``` #[inline(always)] pub fn sub(self, other: Self) -> Self { Self(i8x16_sub(self.0, other.0)) @@ -826,6 +844,14 @@ pub mod wasm32_simd { } /// Compare-greater-than: returns a 16-bit mask. Bit i set where self[i] > other[i]. + /// + /// Register-level compare: it takes a second register, so it is binary in form. + /// The masking-ops predicates call register compares like this one with a + /// broadcast constant; lane-vs-lane predicates (G7) stay deliberately absent at the + /// slice/IR layer, see `.claude/knowledge/masking-ops-state.md` § G7. This method is + /// not that gap. The same operation exists on every backend at its native widths + /// (`I8x64::cmp_gt` etc.); an `I8x16` alias is not to be added to other arms + /// without a caller. #[inline(always)] pub fn cmp_gt(self, other: Self) -> u16 { i8x16_bitmask(i8x16_gt(self.0, other.0)) @@ -878,6 +904,179 @@ pub mod wasm32_simd { } } + // ════════════════════════════════════════════════════════════════════ + // U8x16 — 16 × u8 backed by one v128 (native byte lane) + // ════════════════════════════════════════════════════════════════════ + + /// 16×u8 backed by one WASM `v128` register. + /// + /// Mirrors the NEON `U8x16` surface exactly: `splat`/`zero`, slice and + /// array load/store, **wrapping** `add`/`sub`, and unsigned `min`/`max`. + /// The value intrinsics are safe under `simd128`; only the pointer + /// load/store need `unsafe`. + #[derive(Copy, Clone)] + #[repr(transparent)] + pub struct U8x16(pub v128); + + impl U8x16 { + pub const LANES: usize = 16; + + /// Broadcast `v` to all 16 lanes. + /// + /// # Examples + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// assert_eq!(U8x16::splat(7).to_array(), [7u8; 16]); + /// ``` + #[inline(always)] + pub fn splat(v: u8) -> Self { + Self(u8x16_splat(v)) + } + + /// All lanes zero. + /// + /// # Examples + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// assert_eq!(U8x16::zero().to_array(), [0u8; 16]); + /// ``` + #[inline(always)] + pub fn zero() -> Self { + Self(u8x16_splat(0)) + } + + /// Load the first 16 elements of `s`. Panics if `s.len() < 16`. + /// + /// # Examples + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// let data: Vec = (0..20).collect(); + /// assert_eq!(U8x16::from_slice(&data).to_array()[15], 15); + /// ``` + #[inline(always)] + pub fn from_slice(s: &[u8]) -> Self { + assert!(s.len() >= 16); + // SAFETY: length checked >= 16, so the 16-byte unaligned load + // (`v128_load` has no alignment requirement) stays in bounds. + Self(unsafe { v128_load(s.as_ptr() as *const v128) }) + } + + /// Build from a `[u8; 16]`. + /// + /// # Examples + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// let a = [3u8; 16]; + /// assert_eq!(U8x16::from_array(a).to_array(), a); + /// ``` + #[inline(always)] + pub fn from_array(arr: [u8; 16]) -> Self { + // SAFETY: a [u8; 16] is exactly 16 bytes; the unaligned load reads + // all of it and nothing beyond. + Self(unsafe { v128_load(arr.as_ptr() as *const v128) }) + } + + /// Copy the 16 lanes out to an array. + /// + /// # Examples + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// assert_eq!(U8x16::splat(9).to_array(), [9u8; 16]); + /// ``` + #[inline(always)] + pub fn to_array(self) -> [u8; 16] { + let mut arr = [0u8; 16]; + // SAFETY: the unaligned store writes exactly 16 bytes into the + // 16-byte array. + unsafe { v128_store(arr.as_mut_ptr() as *mut v128, self.0) }; + arr + } + + /// Store the 16 lanes into the first 16 elements of `s`. Panics if + /// `s.len() < 16`. + /// + /// # Examples + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// let mut out = [0u8; 20]; + /// U8x16::splat(4).copy_to_slice(&mut out); + /// assert_eq!(&out[..16], &[4u8; 16]); + /// assert_eq!(&out[16..], &[0u8; 4]); + /// ``` + #[inline(always)] + pub fn copy_to_slice(self, s: &mut [u8]) { + assert!(s.len() >= 16); + // SAFETY: length checked >= 16, so the 16-byte unaligned store + // stays in bounds. + unsafe { v128_store(s.as_mut_ptr() as *mut v128, self.0) }; + } + + /// Lane-wise **wrapping** add (`u8x16_add`): `255 + 1 == 0`. + /// Never saturates — matches NEON `vaddq_u8`. + /// + /// # Examples + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// let r = U8x16::splat(255).add(U8x16::splat(1)); + /// assert_eq!(r.to_array(), [0u8; 16]); + /// ``` + #[inline(always)] + pub fn add(self, other: Self) -> Self { + Self(u8x16_add(self.0, other.0)) + } + + /// Lane-wise **wrapping** subtract (`u8x16_sub`): `0 - 1 == 255`. + /// Never saturates — matches NEON `vsubq_u8`. + /// + /// # Examples + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// let r = U8x16::zero().sub(U8x16::splat(1)); + /// assert_eq!(r.to_array(), [255u8; 16]); + /// ``` + #[inline(always)] + pub fn sub(self, other: Self) -> Self { + Self(u8x16_sub(self.0, other.0)) + } + + /// Lane-wise unsigned minimum (`u8x16_min`). + /// + /// # Examples + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// let r = U8x16::splat(200).min(U8x16::splat(10)); + /// assert_eq!(r.to_array(), [10u8; 16]); + /// ``` + #[inline(always)] + pub fn min(self, other: Self) -> Self { + Self(u8x16_min(self.0, other.0)) + } + + /// Lane-wise unsigned maximum (`u8x16_max`). Unsigned: `200 > 10`. + /// + /// # Examples + /// ```rust,ignore + /// use ndarray::simd::U8x16; + /// let r = U8x16::splat(200).max(U8x16::splat(10)); + /// assert_eq!(r.to_array(), [200u8; 16]); + /// ``` + #[inline(always)] + pub fn max(self, other: Self) -> Self { + Self(u8x16_max(self.0, other.0)) + } + } + + impl fmt::Debug for U8x16 { + fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { + write!(f, "U8x16({:?})", self.to_array()) + } + } + impl PartialEq for U8x16 { + fn eq(&self, other: &Self) -> bool { + self.to_array() == other.to_array() + } + } + // ════════════════════════════════════════════════════════════════════ // U32x4 / U32x16 — u32 ARX lanes (ChaCha20 / BLAKE), NEON-style. // @@ -1303,6 +1502,8 @@ pub mod wasm32_simd { #[allow(non_camel_case_types)] pub type i8x16 = I8x16; #[allow(non_camel_case_types)] + pub type u8x16 = U8x16; + #[allow(non_camel_case_types)] pub type u32x16 = U32x16; /// Lowercase alias of the native wasm [`I32x16`] (travels with the type). #[allow(non_camel_case_types)] diff --git a/tests/numeric.rs b/tests/numeric.rs index adf8ef99..0dcb5445 100644 --- a/tests/numeric.rs +++ b/tests/numeric.rs @@ -4,7 +4,6 @@ use approx::assert_abs_diff_eq; use ndarray::{arr0, arr1, arr2, array, aview1, Array, Array1, Array2, Array3, Axis}; -use std::f64; #[test] fn test_mean_with_nan_values() { diff --git a/tools/safe_intrinsic_probe/Cargo.toml b/tools/safe_intrinsic_probe/Cargo.toml index e5be9573..a3d1f1df 100644 --- a/tools/safe_intrinsic_probe/Cargo.toml +++ b/tools/safe_intrinsic_probe/Cargo.toml @@ -3,9 +3,9 @@ # after a toolchain bump (the answer is a toolchain property, not a code one): # # cd tools/safe_intrinsic_probe -# RUSTFLAGS="--cfg probe_a" cargo check --target aarch64-unknown-linux-gnu # E0133 on 1.98.1 -# RUSTFLAGS="--cfg probe_c" cargo check --target aarch64-unknown-linux-gnu # E0133 on 1.98.1 -# RUSTFLAGS="--cfg probe_a3 -Ctarget-cpu=x86-64-v4" cargo check # E0133 on 1.98.1 +# RUSTFLAGS="--cfg probe_a" cargo check --target aarch64-unknown-linux-gnu # E0133 on 1.98.1 and 1.99.0 +# RUSTFLAGS="--cfg probe_c" cargo check --target aarch64-unknown-linux-gnu # E0133 on 1.98.1 and 1.99.0 +# RUSTFLAGS="--cfg probe_a3 -Ctarget-cpu=x86-64-v4" cargo check # E0133 on 1.98.1 and 1.99.0 # RUSTFLAGS="--cfg probe_a -Ctarget-feature=+simd128" cargo check --target wasm32-unknown-unknown # OK # # Result matrix and the consequence for the backends: see