You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Array-of-strings tuple sketch: key hashing is incompatible with Java (UTF-8 vs UTF-16) #533
@proost, while checking cross-language binary compatibility of the tuple sketches against the snapshots in datasketches-tck, we found that the C++ array-of-strings (AoS) tuple sketch hashes keys differently from Java. For the same keys, C++ and Java produce completely different hashes, so their sketches cannot be meaningfully combined.
Evidence
Comparing the TCK snapshots aos_*_cpp.sk and aos_*_java.sk, generated from the same keys:
Case
Retained (Java / C++)
Hashes in common
aos_1_n10
10 / 10
0
aos_1_n1000
1000 / 1000
0
aos_multikey_n1000
1000 / 1000
0
aos_unicode
3 / 3
0
Each language reads the other's files without error, so nothing fails visibly. But a union of a Java sketch and a C++ sketch built from the same keys double-counts every key, and an intersection comes out empty.
Cause
The seed (0x7A3CCA71) and the , separator match Java, but the bytes that get hashed don't:
finalStrings = stringConcat(strArray); // joined with ','returnhashCharArr(s.toCharArray(), 0, s.length(), PRIME);
XxHash.hashCharArr hashes each char as 2 bytes, little-endian.
C++ (hash_array_of_strings_key in tuple/include/array_of_strings_sketch_impl.hpp) hashes the UTF-8 bytes of each string.
Why C++ and Go should change, not Java
All three implementations need to agree. The cost of changing each one differs a lot:
Java has shipped its AoS sketch for years. Changing its hashing would make every AoS sketch Java users have already saved incompatible with new ones. It would also break Java's compatibility with itself across versions. Java also hashes UTF-16 deliberately: that is how Java stores strings, so it can hash them without first converting them.
Go released its AoS sketch only recently, in v0.2.0. As a pre-1.0 library, it can make breaking changes between minor versions, provided the release notes call them out.
So Java is the reference, and C++ and Go should match it.
Why this issue wasn't caught earlier
Our current cross-language tests only check that the .sk files can be read, not that the hashes match. We discovered this only recently, when we began Cross-Language Binary (CLB) testing.
Proposed fix
In hash_array_of_strings_key, convert each key string from UTF-8 to UTF-16 code units, using surrogate pairs for code points above U+FFFF. Hash the code units as little-endian bytes, with , as 2c 00 between strings. Invalid UTF-8 should throw std::invalid_argument.
Add a test that builds sketches in C++ from the same keys as the Java generators (AosSketchCrossLanguageTest) and requires the hash sets to equal those in the Java .sk files, including aos_unicode.
Update the header comment, which currently says the hashing matches Java.
We have verified this approach: with UTF-8 to UTF-16LE conversion, C++ reproduces the hashes in every Java AoS snapshot exactly. That includes aos_unicode, whose keys contain Korean, Cyrillic and emoji outside the BMP (🔑, 🗝️), and the multikey and 1,000,000-item cases. Apart from the hashing, the format already matches: every multi-entry Java AoS snapshot round-trips through C++ byte for byte.
Optional, while in this code: compact_array_of_strings_tuple_sketch::serialize() has no default serde, unlike deserialize(). Calling it without passing default_array_of_strings_serde<>() compiles but fails at link time with an undefined serde<array<std::string>> symbol.
Timing
The C++ AoS sketch (#476) has not been released yet. We'd like to fix this before 5.3.0, which we plan to release very soon, so that the first release is compatible with Java. Otherwise, fixing it later would invalidate C++ users' saved sketches.
Could you take this on in the next few days? If you're short on time, let us know and we can prepare the PR for your review.
@proost, while checking cross-language binary compatibility of the tuple sketches against the snapshots in datasketches-tck, we found that the C++ array-of-strings (AoS) tuple sketch hashes keys differently from Java. For the same keys, C++ and Java produce completely different hashes, so their sketches cannot be meaningfully combined.
Evidence
Comparing the TCK snapshots
aos_*_cpp.skandaos_*_java.sk, generated from the same keys:aos_1_n10aos_1_n1000aos_multikey_n1000aos_unicodeEach language reads the other's files without error, so nothing fails visibly. But a union of a Java sketch and a C++ sketch built from the same keys double-counts every key, and an intersection comes out empty.
Cause
The seed (
0x7A3CCA71) and the,separator match Java, but the bytes that get hashed don't:tuple/Util.stringArrHash) hashes the joined key as UTF-16 code units:XxHash.hashCharArrhashes eachcharas 2 bytes, little-endian.hash_array_of_strings_keyintuple/include/array_of_strings_sketch_impl.hpp) hashes the UTF-8 bytes of each string.Why C++ and Go should change, not Java
All three implementations need to agree. The cost of changing each one differs a lot:
So Java is the reference, and C++ and Go should match it.
Why this issue wasn't caught earlier
Our current cross-language tests only check that the .sk files can be read, not that the hashes match. We discovered this only recently, when we began Cross-Language Binary (CLB) testing.
Proposed fix
hash_array_of_strings_key, convert each key string from UTF-8 to UTF-16 code units, using surrogate pairs for code points above U+FFFF. Hash the code units as little-endian bytes, with,as2c 00between strings. Invalid UTF-8 should throwstd::invalid_argument.AosSketchCrossLanguageTest) and requires the hash sets to equal those in the Java.skfiles, includingaos_unicode.We have verified this approach: with UTF-8 to UTF-16LE conversion, C++ reproduces the hashes in every Java AoS snapshot exactly. That includes
aos_unicode, whose keys contain Korean, Cyrillic and emoji outside the BMP (🔑, 🗝️), and the multikey and 1,000,000-item cases. Apart from the hashing, the format already matches: every multi-entry Java AoS snapshot round-trips through C++ byte for byte.Optional, while in this code:
compact_array_of_strings_tuple_sketch::serialize()has no default serde, unlikedeserialize(). Calling it without passingdefault_array_of_strings_serde<>()compiles but fails at link time with an undefinedserde<array<std::string>>symbol.Timing
The C++ AoS sketch (#476) has not been released yet. We'd like to fix this before 5.3.0, which we plan to release very soon, so that the first release is compatible with Java. Otherwise, fixing it later would invalidate C++ users' saved sketches.
Could you take this on in the next few days? If you're short on time, let us know and we can prepare the PR for your review.
Go has the same issue; see apache/datasketches-go#191.