Most CITM/Twitter realloc cost was due to growing from 1024 in 7-9 doublings
to reach the actual output size. Bumping the initial capacity to 256KB means
Twitter (82KB) fits in one allocation and CITM (496KB) only needs 1 growth.
Measured (TRUE A/B, 7 alternating rounds in single docker, with all prior
follow-ups in this branch + this cap bump):
CITM: 4358 -> 4924 MB/s (+13.0%) - now ~2.3% AHEAD of Glaze
Twitter: 6338 -> 8034 MB/s (+26.8%) - ~50% AHEAD of Glaze
Trade-off: 256KB upfront per string_builder instance. Reasonable for
high-perf JSON serialization but wasteful for tiny one-off messages.
Users serializing small payloads should pass a smaller initial_capacity
to the constructor.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Add append_raw_n<N>(const char*) which propagates the key length as a
template parameter, letting the compiler emit a fully inlined memcpy
with a constant size (direct loads/stores) instead of the size-passed-
as-runtime-arg variant which sometimes fell back to a libc memcpy
dispatch.
Use it in the reflection struct atom for both first_key and rest_key
(both are compile-time constants from define_static_string).
Measured (TRUE A/B, 7 alternating rounds in single docker invocation,
two independent runs combined):
CITM: baseline ~4395 -> patched ~4791 MB/s (+9%)
Glaze ~4810 MB/s -> simdjson now at PARITY with Glaze
(within +/- 3% noise across runs)
Twitter: still well ahead of Glaze, no change from baseline
Output is byte-identical, all reflection comprehensive tests pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>