Add append_raw_n<N>(const char*) which propagates the key length as a
template parameter, letting the compiler emit a fully inlined memcpy
with a constant size (direct loads/stores) instead of the size-passed-
as-runtime-arg variant which sometimes fell back to a libc memcpy
dispatch.
Use it in the reflection struct atom for both first_key and rest_key
(both are compile-time constants from define_static_string).
Measured (TRUE A/B, 7 alternating rounds in single docker invocation,
two independent runs combined):
CITM: baseline ~4395 -> patched ~4791 MB/s (+9%)
Glaze ~4810 MB/s -> simdjson now at PARITY with Glaze
(within +/- 3% noise across runs)
Twitter: still well ahead of Glaze, no change from baseline
Output is byte-identical, all reflection comprehensive tests pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three small follow-ups on top of #2707, all aimed at closing the
remaining CITM gap:
1. simdjson_really_inline on atom<map>, atom<smart_pointer>,
atom<enum>. The previous PR missed these three. atom<map> in
particular was visible at ~11% of CITM profile time before this
change (CITM has events: std::map<string, CITMEvent>).
2. struct atom passes the key size explicitly to append_raw(c, len)
instead of going through append_raw(const char*) which calls
std::strlen on every key (8 fields × 184 events on CITM).
3. append_raw(const char*) uses std::char_traits<char>::length
(constexpr) instead of std::strlen so the compiler can fold the
length when the pointer is to a compile-time string.
Measured (TRUE A/B, 7 alternating rounds in single docker invocation):
CITM: baseline 4437 -> patched 4686 MB/s (+5.6%)
Twitter: within noise
Output is byte-identical to baseline. All static_reflection_comprehensive_tests
pass.
WIP — pushing for safekeeping while continuing to investigate the
remaining ~5% gap to Glaze on CITM. Not ready for PR yet.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* builder: force-inline atom templates and replace integer writer
Two complementary changes that together speed up reflection-driven JSON
serialization by ~28% on integer-heavy structs (CITM) and ~12% on
string-heavy structs (Twitter).
(1) Add simdjson_really_inline (always_inline) to the hot atom<T>
template overloads (arithmetic, struct, container, optional,
string-like). The constexpr-only declaration was only a hint; the
compiler routinely chose to leave atom<unsigned long> as a real
out-of-line function.
(2) Replace string_builder::append<UInt> body with a forward
cascade-on-magnitude integer writer. The old code computed
digit_count(v) upfront and wrote backward in a loop; the new code is
a straight-line if/else cascade that writes digits forward, no loop,
no helper call.
Either change alone gives only modest gains. Together they unlock the
compiler's cross-call optimization: with all atoms inlined and the
integer writer reduced to straight-line code, the compiler can hoist
b.position into a register across the whole struct serialization, fold
redundant capacity_check calls, and eliminate the strict-aliasing
penalty that otherwise forces b.position/b.capacity reloads after every
char* write.
Output is byte-identical to master on CITM (496682 bytes) and Twitter
(81927 bytes). All static_reflection_comprehensive_tests pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* builder: de-recurse write_uint_jeaiii so always_inline applies on g++/MSVC
g++ ('inlining failed in call to always_inline ...: function not
considered for inlining') and MSVC ('warning C4714: __forceinline not
inlined' under warnings-as-errors) both refuse to inline recursive
functions marked simdjson_really_inline. The original write_uint_jeaiii
called itself in the >=10^4 branches.
Refactor into a non-recursive DAG of helpers: write_lt100, write_lt10000,
write_4_digits, write_lt1e8, write_uint_jeaiii. Each calls strictly
smaller-domain helpers, no cycles. Same straight-line cascade behavior,
same byte output, but every node is now a candidate for always_inline on
all compilers.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* builder: port signed-int append to jeaiii writer and drop digit_count helpers
The signed-integer branch in string_builder::append was structurally
identical to the OLD unsigned branch — same digit_count() upfront +
backward 4-digit batched loop. Port it to use the same forward
write_uint_jeaiii() helper as the unsigned branch (write '-'
unconditionally and advance position only if negative — branchless).
This makes int_log2 / fast_digit_count_32 / fast_digit_count_64 /
digit_count fully unused (verified via grep across include/ and src/);
remove them, ~80 lines of dead code.
The signed write path now benefits from the same compiler-level
optimization (full inlining, capacity-check fusion, no opaque-loop
boundary) as the unsigned path. Signed-int microbench (200 × 100k
values, mixed magnitude and sign): 2272 → 2685 MB/s (+18.2%).
CITM and Twitter benchmarks unchanged in shape and remain byte-identical
to baseline output.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix tests
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Daniel Lemire <daniel@lemire.me>
Each non-first field in a reflection-driven struct serialization was
emitting three separate string_builder::append calls (',', "\"key\"", ':'),
each going through capacity_check + a small write. Combine them into a
single compile-time string per field so each field does one capacity_check
and one memcpy. Same pattern fixed in atom(), append(), and extract_from().
Measured under clang-p2996 -O3 -DNDEBUG -freflection -std=c++26
(median of 5 contemporaneous runs, output byte-identical to baseline):
CITM serialization (496682 bytes):
simdjson_reuse_buffer: 3192 -> 3673 MB/s (+15.1%)
simdjson_static_reflection: 3014 -> 3443 MB/s (+14.2%)
simdjson_to: 2717 -> 3234 MB/s (+19.0%)
simdjson_to_reuse: 2836 -> 3226 MB/s (+13.7%)
Twitter serialization (81927 bytes):
simdjson_reuse_buffer: 6637 -> 7162 MB/s (+7.9%)
simdjson_static_reflection: 5527 -> 5915 MB/s (+7.0%)
simdjson_to: 5070 -> 5350 MB/s (+5.5%)
simdjson_to_reuse: 4976 -> 5303 MB/s (+6.6%)
extract_from (out-of-tree micro-bench, 100 records per call):
User 4-of-9 fields: 1833 -> 1888 MB/s (+3.0%)
Status 3-of-6 fields: 4286 -> 4334 MB/s (+1.1%)
Parsing benchmarks unchanged. tests/builder/static_reflection_comprehensive_tests
still passes (round-trip through extract_from / extract_into is verified).