weaselbot d040ed63a4
CI / release (arm64, , ubuntu-latest-arm64) (pull_request) Successful in 2m18s
CI / pre-commit (pull_request) Successful in 1m58s
CI / test (-DCMAKE_BUILD_TYPE=Debug -DMSAN_TOOLCHAIN_PATH=/opt/msan, debug) (pull_request) Successful in 3m47s
CI / test (-DCMAKE_CXX_FLAGS=-DUSE_64_BIT=1, 64-bit-versions) (pull_request) Successful in 3m16s
CI / test (-DCMAKE_C_COMPILER=gcc -DCMAKE_CXX_COMPILER=g++, gcc) (pull_request) Successful in 3m12s
CI / test (-DUSE_SIMD_FALLBACK=ON, simd-fallback) (pull_request) Successful in 3m17s
CI / release (amd64, -DMSAN_TOOLCHAIN_PATH=/opt/msan, ubuntu-latest-amd64) (pull_request) Successful in 5m32s
CI / coverage (pull_request) Successful in 3m42s
Move aarch64 Node16 SIMD index loads into assembly
The aarch64 NEON paths in getNodeIndex/getNodeIndexExists,
getChildGeq(Node16*), scan16, and checkMaxBetweenExclusiveImpl<Node16>
loaded the full 16-element Node16::index array (and the Node48
reverseIndex array via scan16) with NEON intrinsics, then masked the
result down to [0, numChildren). Only the in-use slots are initialized;
the unused bytes are indeterminate, so the wide loads were undefined
behavior in C++ (per [basic.indet]) even though the trailing lanes were
discarded. MSan reports this on x86-64; on aarch64 it is the same UB but
MSan's imprecise modeling doesn't flag it at -O0, so there is no red->green
test.

Mirror the existing x86-64 fix (commit 6fed133): implement the index
operations in file-level assembly, where loading and operating on
indeterminate values is well-defined. Add simd_aarch64.S with
find_eq_16, find_ge_16, and mask_in_range_16. AArch64 lacks pmovmskb, so
(like the prior NEON code) these return a 64-bit nibble mask rather than
a 16-bit bitmask; the C++ call sites keep their existing nibble-mask
arithmetic and only swap the inline NEON load/compare for the assembly
call. The childMaxVersion compares stay in C++ NEON intrinsics, matching
x86-64's compare16: those slots are always initialized to zero by the
allocator, so the wide loads are defined.

The assembly functions carry `bti c` landing pads and the same
aeabi_feature_and_bits attributes the compiler emits for
-mbranch-protection=standard, so the object stays BTI/PAC/GCS-compatible
(and warning-free under -z force-bti). CMakeLists.txt builds simd_aarch64.S
into the object library and the SIMD test/bench/fuzz targets on aarch64.

Closes #68
2026-08-02 23:48:28 -04:00
2024-03-06 21:22:30 -08:00
2026-07-14 14:33:51 -04:00
2024-08-21 14:00:00 -07:00
2024-08-05 12:20:38 -07:00
2024-04-04 16:29:26 -07:00
2024-11-15 16:49:21 -08:00
2024-02-20 13:22:22 -08:00
2024-01-30 10:39:43 -08:00
2024-08-30 16:06:43 -07:00
2024-09-09 20:10:55 -07:00
2024-11-10 21:47:26 -08:00
2026-08-02 21:16:27 -04:00
2024-11-19 09:02:16 -08:00

A data structure for optimistic concurrency control on ranges of bitwise-lexicographically-ordered keys.

Intended as an alternative to FoundationDB's skip list.

Hardware for all benchmarks is an AMD Ryzen 9 7900 with (2x32GB) 5600MT/s CL28-34-34-89 1.35V RAM.

$ clang++ --version

Ubuntu clang version 21.1.8 (6ubuntu1)
Target: x86_64-pc-linux-gnu
Thread model: posix
InstalledDir: /usr/lib/llvm-21/bin

Microbenchmark

Skip list

ns/op op/s err% ins/op cyc/op IPC bra/op miss% total benchmark
164.29 6,086,873.38 0.0% 3,107.03 604.19 5.142 558.59 0.0% 1.96 point reads
161.05 6,209,395.38 0.1% 3,036.76 592.21 5.128 539.35 0.0% 1.93 prefix reads
239.55 4,174,539.38 0.1% 3,722.71 880.68 4.227 692.00 0.0% 2.86 range reads
354.75 2,818,919.14 0.7% 4,523.64 1,304.75 3.467 720.22 2.0% 4.23 point writes
345.32 2,895,878.47 0.1% 4,484.57 1,270.31 3.530 705.00 1.8% 4.12 prefix writes
193.48 5,168,547.42 0.1% 2,224.10 711.72 3.125 377.17 3.3% 2.32 range writes
404.89 2,469,777.50 2.4% 6,855.96 1,489.70 4.602 1,227.82 1.3% 0.05 monotonic increasing point writes
134,231.80 7,449.80 1.9% 812,045.25 495,770.40 1.638 151,246.50 0.9% 0.01 worst case for radix tree
37.80 26,454,311.17 0.4% 701.00 139.14 5.038 102.00 0.0% 0.01 create and destroy

Radix tree (this implementation)

ns/op op/s err% ins/op cyc/op IPC bra/op miss% total benchmark
12.89 77,565,115.56 0.1% 244.55 47.43 5.155 34.21 0.6% 0.15 point reads
15.11 66,162,047.76 0.1% 297.79 55.60 5.356 43.23 0.4% 0.18 prefix reads
36.29 27,559,358.29 0.1% 783.16 133.44 5.869 109.52 0.2% 0.43 range reads
20.53 48,719,405.55 0.1% 381.81 75.51 5.057 51.04 0.5% 0.25 point writes
39.37 25,402,042.40 0.1% 685.00 144.83 4.730 106.72 0.3% 0.47 prefix writes
43.78 22,843,841.63 0.1% 800.40 161.06 4.970 127.36 0.1% 0.53 range writes
78.37 12,760,008.75 1.0% 1,452.61 288.24 5.040 278.69 0.1% 0.01 monotonic increasing point writes
322,885.50 3,097.07 1.5% 4,362,382.00 1,183,852.00 3.685 765,301.00 0.1% 0.01 worst case for radix tree
99.99 10,000,718.79 0.4% 1,775.00 367.93 4.824 288.00 0.0% 0.01 create and destroy

"Real data" test

Point queries only. Gc ratio is the ratio of time spent doing garbage collection to time spent adding writes or doing garbage collection. Lower is better.

skip list

Check: 4.62967 seconds, 352.195 MB/s, Add: 3.34177 seconds, 167.771 MB/s, Gc ratio: 37.9399%, Peak idle memory: 5.51852e+06

radix tree

Check: 1.00477 seconds, 1622.8 MB/s, Add: 1.21142 seconds, 462.808 MB/s, Gc ratio: 39.4716%, Peak idle memory: 2.0226e+06

hash table

(The hash table implementation doesn't work on range queries, and its purpose is to provide an idea of how fast point queries can be)

Check: 0.854254 seconds, 1908.74 MB/s, Add: 0.632626 seconds, 886.232 MB/s, Gc ratio: 41.0827%, Peak idle memory: 0
S
Description
A data structure for optimistic concurrency control on ranges of bitwise-lexicographically-ordered keys.
Readme Apache-2.0
27 MiB
v0.0.13
Latest
2024-08-26 21:24:21 +00:00
Languages
C++ 79.9%
TeX 7.7%
CMake 5.5%
Python 3.8%
Assembly 1.8%
Other 1.3%