Compare commits

...

111 Commits

Author SHA1 Message Date
Daniel Lemire 8e85352df4 Merge branch 'master' into parse_string_if_needed 2026-01-21 10:10:42 -05:00
Francisco Geiman Thiesen fc57c09cf0 Add FracturedJson formatting support for DOM serialization (#2580)
* Add FracturedJson formatting support for DOM serialization

Implements FracturedJson formatting as requested in issue #2576.
FracturedJson produces human-readable yet compact JSON output by
intelligently choosing between different layout strategies based on
content complexity, length, and structure similarity.

Key features:
- Four layout modes: inline, compact multiline, table, and expanded
- Structure analysis pass to compute metrics before formatting
- Table formatting for arrays of similar objects with column alignment
- Configurable options for line length, indentation, padding, etc.

New files:
- fractured_json.h: Public API with fractured_json_options struct
- fractured_json-inl.h: Implementation (~1000 lines)
- json_structure_analyzer.h: Structure analysis for layout decisions
- fractured_formatter.h: Formatter class using CRTP pattern

Usage:
  dom::parser parser;
  element doc = parser.parse(json_string);
  std::cout << fractured_json(doc) << std::endl;

  // Or with custom options:
  fractured_json_options opts;
  opts.indent_spaces = 2;
  std::cout << fractured_json(doc, opts) << std::endl;

  // Or format any JSON string (useful with reflection API):
  auto formatted = fractured_json_string(minified_json);

Resolves #2576

* Add comprehensive tests for FracturedJson formatter

Adds 27 test cases covering all aspects of the FracturedJson formatter:

Core functionality tests (13):
- Roundtrip parsing verification
- Inline formatting for simple arrays and objects
- Expanded formatting for complex nested structures
- Compact multiline arrays with configurable items per line
- Table formatting for uniform arrays of objects
- Empty container handling
- All scalar types (string, int, uint, double, bool, null)
- String escaping (quotes, backslashes, control characters)
- Custom indentation options
- Deep nesting (10+ levels)
- Mixed type arrays

Edge case tests (11):
- Unicode strings (Chinese, emoji, Arabic, Russian, accented chars)
- Boundary numbers (INT64_MIN/MAX, UINT64_MAX, DBL_MIN/MAX)
- Nested arrays (arrays of arrays)
- Empty string values
- Keys with special characters (spaces, quotes, colons, etc.)
- Non-uniform arrays (should not trigger table mode)
- Very long strings (500+ chars)
- Large arrays (100 elements)
- Reflection API workflow simulation
- Control characters (tab, newline, CR, null)
- Single element containers

Option tests (3):
- Disable compact multiline mode
- Disable table format mode
- Disable all padding options

* Add FracturedJson integration with builder/reflection API

Extends FracturedJson to work seamlessly with the builder API, enabling
formatted output directly from C++ structs using static reflection.

New functions:
- to_fractured_json_string(obj, opts) - serialize struct to formatted JSON
- to_fractured_json(obj, output, opts) - same with output parameter
- extract_fractured_json<fields...>(obj, opts) - format only specific fields

These functions combine the builder's reflection-based serialization with
FracturedJson formatting in a single convenient call:

  struct User { int id; std::string name; bool active; };
  User user{1, "Alice", true};

  // Minified output (existing):
  auto minified = to_json_string(user);
  // {"id":1,"name":"Alice","active":true}

  // Formatted output (new):
  auto formatted = to_fractured_json_string(user);
  // { "id": 1, "name": "Alice", "active": true }

  // Partial extraction with formatting:
  auto partial = extract_fractured_json<"id", "name">(user);
  // { "id": 1, "name": "Alice" }

New files:
- generic/builder/fractured_json_builder.h - builder integration
- tests/builder/static_reflection_fractured_json_tests.cpp - 7 tests

* Fix INT64_MIN overflow and implement table_similarity_threshold

- Fix undefined behavior when negating INT64_MIN in estimate_number_length()
  and measure_value_length() by returning 20 (the exact length of the
  string representation) directly
- Actually use table_similarity_threshold in check_array_uniformity() by
  calling compute_object_similarity() to compare objects against the first
  object in the array

* Fix -Werror=effc++ member initialization warnings

Initialize all member variables in member initialization lists to
satisfy GCC's -Werror=effc++ flag:
- element_metrics::common_keys - add {} default initializer
- structure_analyzer - add default constructor with member init list
- fractured_formatter - add column_widths_{} to constructor
- fractured_string_builder - add analyzer_{} to constructor

* Add Rule of Five to structure_analyzer class

The class has a pointer member (current_opts_) which triggers
-Werror=effc++ requiring explicit copy/move operations. Delete
copy operations (class shouldn't be copied due to cache) and
default move operations.

* Fix Windows build: wrap std::max to avoid macro conflict

Windows.h defines max/min macros that interfere with std::max/std::min.
Wrapping in parentheses as (std::max)(...) prevents macro expansion.

* Fix GCC 15 false positive -Wfree-nonheap-object warning

GCC 15 on MINGW64 gives a false positive warning in parser_moving_parser()
when the std::vector<std::string> goes out of scope. Suppress this
specific warning with a pragma for GCC builds.

* Fix metrics cache key bug by passing metrics through recursion

The cache was using element addresses as keys, but dom::element objects
are lightweight wrappers that get copied during iteration, causing
different addresses between analysis and formatting phases. This resulted
in cache misses and fallback to empty metrics.

Solution: Store child metrics in the element_metrics struct and pass
them through recursive calls, eliminating the need for address-based
caching entirely.

Changes:
- Add children vector to element_metrics for hierarchical metrics
- Remove metrics_cache_ and related get_metrics/has_metrics methods
- Update all format functions to accept and pass child metrics
- Add public analyze_array/analyze_object overloads for standalone use

* Add ignore patterns for Node.js, Rust, and generated files

Add entries for node_modules, package-lock.json, Rust target
directories, local ablation artifacts, and generated documentation
files.

* Refactor: extract analyze_scalar helper to reduce code duplication

Extract common scalar type handling (STRING, INT64, UINT64, DOUBLE,
BOOL, NULL_VALUE) into a dedicated analyze_scalar method. Each scalar
type shares the same initialization pattern for complexity, child_count,
can_inline, and recommended_layout.

Also simplify boolean formatting in format_scalar to use ternary operator.

* Fix formatting and duplicate error message in amalgamate.py

Reformat cramped is_amalgamator condition to multi-line for readability.
Fix duplicate error message text in _included_filename_root and use
correct variable name (relative_root instead of root).

* Refactor: add count_newlines helper in fractured_json tests

Extract repeated newline counting loop into a reusable static helper
function, used by inline_array_test, inline_object_test, and
expanded_test.

* Revert "Add ignore patterns for Node.js, Rust, and generated files"

This reverts commit 4760ea7cd0.

* various minor changes

---------

Co-authored-by: Daniel Lemire <daniel@lemire.me>
2026-01-20 10:32:24 -05:00
Daniel Lemire 2058b47dfe adding padded string builder (#2592)
* adding padded string builder

* minor rename

* deleting copy constructor

* typo

* [no-ci] tuning documentation.
2026-01-18 20:18:19 -05:00
Daniel Lemire b0486c7fa2 saving. 2026-01-18 11:57:55 -05:00
Daniel Lemire feb1e7feb6 work on the ondemand iterators (#2590)
* work on the ondemand iterators

* guarding two SIMDJSON_ASSUME

* simplify following @jkeiser's comment

* adding safety rails to the iterators

* silencing a warning.

* updating the amalgamation files
2026-01-17 21:15:18 -05:00
Eve Silfanus 2c7fbc1538 Build Performance Optimization (#2588) 2026-01-16 11:42:39 -05:00
Daniel Lemire 8b69401d8a moving the builder files in their own directory (#2578) 2026-01-07 18:08:11 -05:00
Daniel Lemire f504e57e7a Add runtime dispatching for loongarch (#2575)
* loongarch runtime dispatching

* remove CMake config for Loongarch.

* make lasx available

* flipping

* moving...

* fixing minor logic error

* minor fixes

* flipping

* adding hackish header

* better comment and reordering

* adding dispatch
2026-01-02 14:28:23 -05:00
Arthur Chan 135c173053 oss-fuzz: Add unit testing build to oss-fuzz build script (#2574) 2026-01-02 14:28:01 -05:00
Daniel Lemire b4ed3a99a9 fixing typos (#2573) 2025-12-30 17:38:03 -05:00
Daniel Lemire d8e1b36c88 just new small tests. (#2571) 2025-12-23 11:14:09 -05:00
Daniel Lemire c249a1b456 Merge branch 'master' of github.com:simdjson/simdjson 2025-12-22 16:08:30 -05:00
Daniel Lemire 9c4b793c90 adding ref 2025-12-22 16:08:19 -05:00
Daniel Lemire dd92a8414c Add optimization option to pull request template 2025-12-19 23:30:13 -05:00
Dirk Stolle 835bdba123 adjust logo image file names (*_simdjason_* -> *_simdjson_*) (#2569) 2025-12-19 23:21:59 -05:00
Dirk Stolle ad3cd71ca2 fix a few typos (#2568) 2025-12-19 21:56:32 -05:00
Daniel Lemire 860f7e0458 Update logo in README.md 2025-12-17 20:36:08 -05:00
Daniel Lemire 980f2ad3af 4.2.4 2025-12-17 20:33:11 -05:00
Daniel Lemire 7ad9fe63a6 fixing issue 2549 (#2567)
* fixing issue 2549

* saving.
2025-12-17 20:32:36 -05:00
Daniel Lemire 7987418b1f adding the official simdjson logo files 2025-12-13 12:14:40 -05:00
Daniel Lemire 5e871f6724 improving slightly the documentation. 2025-12-12 19:04:34 -05:00
Daniel Lemire 5d16fd5f31 4.2.3 2025-12-12 17:51:39 -05:00
Jake S. Del Mastro 4e9ff03af5 Make it possible to provide custom serializers for range types ( (#2550)
If you provide a custom serializer for range types it is currently never used due to the requires clause for string_builder::append with ranges is overly broad
2025-12-12 17:50:52 -05:00
Daniel Lemire aa7489060a Fix typo in bug report template 2025-12-12 15:24:48 -05:00
Daniel Lemire ae32422891 a few additional tests and removing a bad remark in the documentation... 2025-12-03 19:35:18 -05:00
Liqiang TAO 667d0ed3c7 make code branchless (#2546) 2025-11-18 16:57:06 -05:00
Daniel Lemire b1c31b428d update 2025-11-11 14:21:04 -05:00
Daniel Lemire 56ac56ba32 Merge branch 'master' of github.com:simdjson/simdjson 2025-11-11 14:17:08 -05:00
Muhammad Rizal Nurromdhoni 19549c60ec string_builder range-based append fix (#2544)
* Use std::ranges::range_value_t on range

* Add ranges test
2025-11-11 14:15:00 -05:00
Daniel Lemire 16e99f229b Update iterate_many.md for clarity on JSON processing
Clarified the example JSON format and emphasized the need for efficient processing.
2025-11-10 13:49:13 -05:00
Liqiang TAO 21342a4142 Fix some wrong content in doc (#2542) 2025-11-10 11:24:38 -05:00
Daniel Lemire d0e841d3e9 Release Candidate 4.2.2 (#2539)
* adding documentation.

* release candidate
2025-11-06 12:00:21 -05:00
Daniel Lemire a962652ec3 adding documentation. 2025-11-05 11:42:38 -05:00
hiteshmk05 77d73b068a add: windows wstring support for padded_str (#2537)
* add: windows wstring support for padded_str

* add: padded_string::load for wstring windows

* fix: extra space

* change: file path
2025-11-05 11:29:30 -05:00
Daniel Lemire 19ff7a572d adding concept examples to the compile-time JSON. (#2538) 2025-11-04 15:06:40 -05:00
Daniel Lemire 235dbc5369 4.2.1 2025-11-03 11:04:23 -05:00
Daniel Lemire 2b9c8977af using _json for compile-time JSON strings. (#2536) 2025-11-03 11:03:21 -05:00
Daniel Lemire 3d87bd4abc release 4.2.0 2025-11-02 16:19:39 -05:00
hiteshmk05 a60c0d1e39 fix: cmake error when CMAKE_CXX_FLAGS is empty (#2535) 2025-11-02 11:53:56 -05:00
Francisco Geiman Thiesen 86bfbaada7 Merge pull request #2534 from simdjson/francisco/compile-time-parsing_daniel
Compile-time parsing (C++26)
2025-11-01 22:19:37 -07:00
Daniel Lemire ccac6403d9 fixing support for pre-C++17 2025-11-01 17:28:17 -04:00
Daniel Lemire aade58c3dc tweaks 2025-11-01 17:08:03 -04:00
Daniel Lemire 3319815e25 bringing back compatibility with pre-C++17 2025-11-01 16:46:34 -04:00
hiteshmk05 a24f845bd7 Feature/ondemand wildcard support (#2533)
* Add feature for ondemand-wildcard-JSONQueries

* fix wildcard_test

* fix: extra whitespace
2025-11-01 16:46:08 -04:00
Daniel Lemire 212e2d5857 removing unnecessary changes 2025-10-31 18:56:23 -04:00
Daniel Lemire 547e156e33 update 2025-10-31 18:54:04 -04:00
Daniel Lemire 98a45f7229 fix 2025-10-31 18:48:49 -04:00
Daniel Lemire 8589509d1e Merge branch 'master' into francisco/compile-time-parsing_daniel 2025-10-31 18:47:42 -04:00
Daniel Lemire 2fce4a843d update 2025-10-31 18:46:13 -04:00
Daniel Lemire 62803512e4 saving. 2025-10-31 15:58:51 -04:00
Daniel Lemire 76d9dee854 update 2025-10-31 00:13:00 -04:00
Daniel Lemire bf15f21b0b not great but a start. 2025-10-29 20:10:06 -04:00
Daniel Lemire 32b301893c updating toc 2025-10-29 09:53:39 -04:00
Daniel Lemire 7f1531a1f9 removing garbage. 2025-10-29 09:53:13 -04:00
Daniel Lemire 0a3b555ff7 Merge branch 'master' of github.com:simdjson/simdjson 2025-10-29 09:50:15 -04:00
Daniel Lemire 114d45ad54 some garbage 2025-10-29 00:07:48 -04:00
Daniel Lemire 0112be86b0 4.1.0 (#2532) 2025-10-28 00:10:59 -04:00
Daniel Lemire bf52d8198b 4.1.0 2025-10-27 16:56:06 -04:00
Francisco Geiman Thiesen 58c92d6d82 Adding support for compiled json path + json pointer (reflection based) (#2483)
* Adding compile time json path

* using string_view

* Adding support for compile-time json pointer as well.

* Removing unnecessary comment

* Tests now working, still will re-review.

* Adding documentation on the compile-time json path/pointer parsing feature.

* Adding benchmark showing the significant performance advantage of using compiled paths whenever you have them a priori.

* going for JSONPath (correct wording).

* minor update (mostly doc)

---------

Co-authored-by: Daniel Lemire <daniel@lemire.me>
2025-10-27 16:52:41 -04:00
Daniel Lemire 781a7d6c89 removing an unnecessary branch (#2530)
* removing an unnecessary branch

* fixing typo
2025-10-27 14:19:58 -04:00
Max Marrone 3d0de709a8 Fix outdated references to JsonStream. (#2531) 2025-10-26 16:30:31 -04:00
kevyang 49b86721b4 add missing OUT_OF_CAPACITY error code to error codes array (#2527)
* add missing error code to DLLIMPORTEXPORT

* fix syntax

---------

Co-authored-by: Kevin Yang <kjy@meta.com>
2025-10-22 22:20:08 -04:00
Daniel Lemire def2b6efd2 still not good 2025-10-19 22:23:43 -04:00
Daniel Lemire 87a186fbf1 removing circleci 2025-10-19 20:03:52 -04:00
Daniel Lemire 81f10a01b7 documentation update 2025-10-17 21:00:52 -04:00
Daniel Lemire 36ed7ab48a saving 2025-10-17 20:59:37 -04:00
Daniel Lemire c3d1d62dfe use cpp 2025-10-17 20:45:57 -04:00
Daniel Lemire 67821cb6fd updating the documentation. 2025-10-17 20:33:42 -04:00
Daniel Lemire 3ac287ba3d update dox 2025-10-17 20:28:08 -04:00
Daniel Lemire ec352430a0 JSONPath is now an RFC (#2517)
* JSONPath is now an RFC

* up
2025-10-17 13:35:04 -04:00
Francisco Geiman Thiesen d326f2ce9f Working! 2025-10-10 21:32:10 -07:00
Francisco Geiman Thiesen ca42a49fba Compile-time support for parsing json objects! 2025-10-10 18:40:09 -07:00
Daniel Lemire 8a9daeb0ad Restore Star History Chart in README
Readded the Star History Chart section to the README.
2025-10-10 09:07:28 -04:00
Daniel Lemire 9c5a88f1f3 Update README with star history chart 2025-10-10 09:06:49 -04:00
0xflotus a7f8fb71c5 chore: fix small error in docs (#2497) 2025-10-03 11:03:17 -04:00
Howard Guo a553db4c67 Update workflow name to Ubuntu aarch64 (GCC 13) (#2484) 2025-10-03 09:56:07 -04:00
Jaël Champagne Gareau 1fa1af8c15 Fix yyjson leaks when running ./bench_ondemand (#2485) 2025-10-03 09:55:32 -04:00
Daniel Lemire 5ed1044056 4.0.7 2025-09-30 11:26:03 -04:00
Francisco Geiman Thiesen 62913867ff Merge pull request #2475 from simdjson/francisco/extract_from
Adding extract_from functionality + unit tests
2025-09-29 20:34:35 -07:00
Francisco Geiman Thiesen 6700d48b57 Merge pull request #2480 from simdjson/francisco/extract_from3
minor tweaks... ;-)
2025-09-29 13:54:05 -07:00
Francisco Geiman Thiesen b5577d5e85 Merge branch 'master' into francisco/extract_from 2025-09-29 12:46:48 -07:00
wszqkzqk b84a4ec2b9 Fix: Correct narrowing conversion in lsx string parsing (#2481)
Resolves a build failure on the loong64 architecture caused by a narrowing conversion error.

The compiler, with the -Werror=narrowing flag, was flagging the implicit conversion from 'int' (the return
type of to_bitmask()) to 'uint64_t'.

This is fixed by adding an explicit static_cast to uint64_t in include/simdjson/lsx/stringparsing_defs.h.

Signed-off-by: Zhou Qiankang <wszqkzqk@qq.com>
2025-09-29 11:57:06 -04:00
Daniel Lemire 1638a185f7 minor tweaks... ;-) 2025-09-29 11:12:37 -04:00
Francisco Geiman Thiesen 6fe450f5ce Updateing single_header 2025-09-29 03:44:29 -07:00
Francisco Geiman Thiesen c7b70de070 Merge branch 'master' into francisco/extract_from 2025-09-29 03:23:02 -07:00
Francisco Geiman Thiesen 3279fbd55b Merge pull request #2474 from simdjson/complete_extract_into
this completes the extract_into work.
2025-09-29 03:12:20 -07:00
Francisco Geiman Thiesen 66e64e0e5f Merge branch 'master' into francisco/extract_from 2025-09-29 03:02:12 -07:00
Daniel Lemire 03f81e66af updating single header 2025-09-27 12:26:04 -04:00
Daniel Lemire e3b7eddb37 fixing off-by-one mistake in the documentation (#2477) 2025-09-27 12:19:44 -04:00
Daniel Lemire 99c4ba6e8f tweak 2025-09-26 23:44:42 -04:00
Francisco Geiman Thiesen 6aa7eea334 Adding extract_from functionality + unit tests 2025-09-26 19:48:16 -07:00
Daniel Lemire 6a47cda07f guarding 2025-09-26 22:34:36 -04:00
Daniel Lemire 617c69e104 completing doc 2025-09-26 21:33:10 -04:00
Daniel Lemire 625adceb24 updating single-header 2025-09-26 21:29:28 -04:00
Daniel Lemire bd0e9c1336 tweak 2025-09-26 21:04:53 -04:00
Daniel Lemire 7bf82b02d5 this completes the extra_into work. 2025-09-26 21:00:46 -04:00
Francisco Geiman Thiesen 88a1b3e83b Merge pull request #2471 from simdjson/francisco/extract_into
Adding extract_into functionality + test (targets simdjson >= 4.0 as it relies on reflection)
2025-09-26 01:29:33 -07:00
Francisco Geiman Thiesen c72954eade Addressing reviews. 2025-09-25 21:08:38 -07:00
Francisco Geiman Thiesen 4456a10469 Francisco/using iterators for containers (#2470)
* Using iterators instead of subscript operators and size. This helps us work with a broader range of containers.

* Adding list test

* Using std::ranges::input_range<T> as suggested by moisrex
2025-09-25 17:30:42 -04:00
Francisco Geiman Thiesen d8ed2417ad Adding coverage for extract_into with types that have custom serialization 2025-09-25 03:35:46 -07:00
Francisco Geiman Thiesen 5517df7aee Adding extract_into functionality + test 2025-09-24 22:58:23 -07:00
evbse 6e618b0805 Improve DOM implementation (#2434) 2025-09-21 11:26:33 -06:00
Pavel Novikov bde288a623 Fixed string_builder::operator std::string() (#2465)
* clang format

* fixed `string_builder::operator std::string()`

* fixed variable shadowing error false positive
2025-09-21 11:25:42 -06:00
Daniel Lemire b2932d1b8f release candidate 4.0.6 (#2464) 2025-09-21 08:13:48 -06:00
Daniel Lemire a7811090ef fixing issue 2458 (#2461) 2025-09-20 22:23:09 -06:00
Daniel Lemire 3320885fac Fixing issue 2462 (#2463)
* fun

* progress

* completing the documentation
2025-09-20 22:22:57 -06:00
Daniel Lemire 786c68b158 release 4.0.5 2025-09-18 15:17:32 -06:00
Daniel Lemire 703ef54bd9 allow string reuse (#2454)
* allow string reuse

* portability fix

* using data and not begin
2025-09-18 15:16:38 -06:00
Daniel Lemire 2526068e2f Add mamba link to README 2025-09-18 09:22:03 -06:00
Daniel Lemire 403b8bfb91 let us be careful and not change the API 2024-07-11 19:36:12 -04:00
Daniel Lemire 76f45a0c4b fix: add parse_string_if_needed function 2024-07-11 17:57:05 -04:00
193 changed files with 56852 additions and 23504 deletions
-316
View File
@@ -1,316 +0,0 @@
version: 2.1
# We constantly run out of memory so please do not use parallelism (-j, -j4).
# Reusable image / compiler definitions
executors:
gcc8:
docker:
- image: conanio/gcc8
environment:
CXX: g++-8
CC: gcc-8
CMAKE_BUILD_FLAGS:
CTEST_FLAGS: --output-on-failure
gcc9:
docker:
- image: conanio/gcc9
environment:
CXX: g++-9
CC: gcc-9
CMAKE_BUILD_FLAGS:
CTEST_FLAGS: --output-on-failure
gcc10:
docker:
- image: conanio/gcc10
environment:
CXX: g++-10
CC: gcc-10
CMAKE_BUILD_FLAGS:
CTEST_FLAGS: --output-on-failure
clang10:
docker:
- image: conanio/clang10
environment:
CXX: clang++-10
CC: clang-10
CMAKE_BUILD_FLAGS:
CTEST_FLAGS: --output-on-failure
clang9:
docker:
- image: conanio/clang9
environment:
CXX: clang++-9
CC: clang-9
CMAKE_BUILD_FLAGS:
CTEST_FLAGS: --output-on-failure
clang6:
docker:
- image: conanio/clang60
environment:
CXX: clang++-6.0
CC: clang-6.0
CMAKE_BUILD_FLAGS:
CTEST_FLAGS: --output-on-failure
# Reusable test commands (and initializer for clang 6)
commands:
dependency_restore:
steps:
- restore_cache:
keys:
- cmake-cache-{{ checksum "dependencies/CMakeLists.txt" }}
dependency_cache:
steps:
- save_cache:
key: cmake-cache-{{ checksum "dependencies/CMakeLists.txt" }}
paths:
- dependencies/.cache
install_cmake:
steps:
- run: apt-get update -qq
- run: apt-get install -y cmake
cmake_prep:
steps:
- checkout
- run: mkdir -p build
cmake_build_cache:
steps:
- cmake_prep
- dependency_restore
- run: cmake -DSIMDJSON_DEVELOPER_MODE=ON $CMAKE_FLAGS -DCMAKE_INSTALL_PREFIX:PATH=destination -B build .
- dependency_cache # dependencies are produced in the configure step
cmake_build:
steps:
- cmake_build_cache
- run: cmake --build build
cmake_test:
steps:
- cmake_build
- run: |
cd build &&
tools/json2json -h &&
ctest $CTEST_FLAGS -L acceptance &&
ctest $CTEST_FLAGS -LE acceptance -LE explicitonly
cmake_assert_test:
steps:
- run: |
cd build &&
tools/json2json -h &&
ctest $CTEST_FLAGS -L assert
cmake_test_all:
steps:
- cmake_build
- run: |
cd build &&
tools/json2json -h &&
ctest $CTEST_FLAGS -DSIMDJSON_IMPLEMENTATION="haswell;westmere;fallback" -L acceptance -LE per_implementation &&
SIMDJSON_FORCE_IMPLEMENTATION=haswell ctest $CTEST_FLAGS -L per_implementation -LE explicitonly &&
SIMDJSON_FORCE_IMPLEMENTATION=westmere ctest $CTEST_FLAGS -L per_implementation -LE explicitonly &&
SIMDJSON_FORCE_IMPLEMENTATION=fallback ctest $CTEST_FLAGS -L per_implementation -LE explicitonly &&
ctest $CTEST_FLAGS -LE "acceptance|per_implementation" # Everything we haven't run yet, run now.
cmake_perftest:
steps:
- cmake_build_cache
- run: |
cmake -DSIMDJSON_ENABLE_DOM_CHECKPERF=ON --build build --target checkperf &&
cd build &&
ctest --output-on-failure -R checkperf
# we not only want cmake to build and run tests, but we want also a successful installation from which we can build, link and run programs
cmake_install_test: # this version builds, install, test and then verify from the installation
steps:
- run: cd build && make install
- run: echo -e '#include <simdjson.h>\nint main(int argc,char**argv) {simdjson::dom::parser parser;simdjson::dom::element tweets = parser.load(argv[1]); }' > tmp.cpp && c++ -Ibuild/destination/include -Lbuild/destination/lib -std=c++17 -Wl,-rpath,build/destination/lib -o linkandrun tmp.cpp -lsimdjson && ./linkandrun jsonexamples/twitter.json
cmake_installed_test_cxx20: # assuming that it was installed, this tries to build using C++20
steps:
- run: echo -e '#include <simdjson.h>\nint main(int argc,char**argv) {simdjson::dom::parser parser;simdjson::dom::element tweets = parser.load(argv[1]); }' > tmp.cpp && c++ -Ibuild/destination/include -Lbuild/destination/lib -std=c++20 -Wl,-rpath,build/destination/lib -o linkandrun tmp.cpp -lsimdjson && ./linkandrun jsonexamples/twitter.json
jobs:
# static
justlib-gcc10:
description: Build just the library, install it and do a basic test
executor: gcc10
environment: { CMAKE_FLAGS: -DSIMDJSON_JUST_LIBRARY=ON }
steps: [ cmake_build, cmake_install_test, cmake_installed_test_cxx20 ]
assert-gcc10:
description: Build the library with asserts on, install it and run tests
executor: gcc10
environment: { CMAKE_FLAGS: -DSIMDJSON_GOOGLE_BENCHMARKS=OFF -DCMAKE_CXX_FLAGS_RELEASE=-O3 }
steps: [ cmake_test, cmake_assert_test ]
assert-clang10:
description: Build just the library, install it and do a basic test
executor: clang10
environment: { CMAKE_FLAGS: -DSIMDJSON_GOOGLE_BENCHMARKS=OFF -DCMAKE_CXX_FLAGS_RELEASE=-O3 }
steps: [ cmake_test, cmake_assert_test ]
gcc10-perftest:
description: Build and run performance tests on GCC 10 and AVX 2 with a cmake static build, this test performance regression
executor: gcc10
environment: { CMAKE_FLAGS: -DSIMDJSON_GOOGLE_BENCHMARKS=OFF -DBUILD_SHARED_LIBS=OFF }
steps: [ cmake_perftest ]
gcc10:
description: Build and run tests on GCC 10 and AVX 2 with a cmake static build
executor: gcc10
environment: { CMAKE_FLAGS: -DSIMDJSON_GOOGLE_BENCHMARKS=ON -DBUILD_SHARED_LIBS=OFF }
steps: [ cmake_test, cmake_install_test, cmake_installed_test_cxx20 ]
clang6:
description: Build and run tests on clang 6 and AVX 2 with a cmake static build
executor: clang6
environment: { CMAKE_FLAGS: -DSIMDJSON_GOOGLE_BENCHMARKS=ON -DBUILD_SHARED_LIBS=OFF }
steps: [ cmake_test, cmake_install_test ]
clang10:
description: Build and run tests on clang 10 and AVX 2 with a cmake static build
executor: clang10
environment: { CMAKE_FLAGS: -DSIMDJSON_GOOGLE_BENCHMARKS=ON -DBUILD_SHARED_LIBS=OFF }
steps: [ cmake_test, cmake_install_test, cmake_installed_test_cxx20 ]
# libcpp
libcpp-clang10:
description: Build and run tests on clang 10 and AVX 2 with a cmake static build and libc++
executor: clang10
environment: { CMAKE_FLAGS: -DSIMDJSON_USE_LIBCPP=ON -DBUILD_SHARED_LIBS=OFF }
steps: [ cmake_test, cmake_install_test, cmake_installed_test_cxx20 ]
# sanitize
sanitize-gcc10:
description: Build and run tests on GCC 10 and AVX 2 with a cmake sanitize build
executor: gcc10
environment: { CMAKE_FLAGS: -DCMAKE_BUILD_TYPE=Debug -DBUILD_SHARED_LIBS=ON -DSIMDJSON_SANITIZE=ON, CTEST_FLAGS: --output-on-failure -LE explicitonly }
steps: [ cmake_test ]
sanitize-clang10:
description: Build and run tests on clang 10 and AVX 2 with a cmake sanitize build
executor: clang10
environment: { CMAKE_FLAGS: -DBUILD_SHARED_LIBS=ON -DSIMDJSON_NO_FORCE_INLINING=ON -DSIMDJSON_SANITIZE=ON, CTEST_FLAGS: --output-on-failure -LE explicitonly }
steps: [ cmake_test ]
threadsanitize-gcc10:
description: Build and run tests on GCC 10 and AVX 2 with a cmake sanitize build
executor: gcc10
environment: { CMAKE_FLAGS: -DBUILD_SHARED_LIBS=ON -DSIMDJSON_SANITIZE_THREADS=ON, CTEST_FLAGS: --output-on-failure -LE explicitonly }
steps: [ cmake_test ]
threadsanitize-clang10:
description: Build and run tests on clang 10 and AVX 2 with a cmake sanitize build
executor: clang10
environment: { CMAKE_FLAGS: -DBUILD_SHARED_LIBS=ON -DSIMDJSON_NO_FORCE_INLINING=ON -DSIMDJSON_SANITIZE_THREADS=ON, CTEST_FLAGS: --output-on-failure -LE explicitonly }
steps: [ cmake_test ]
# dynamic
dynamic-gcc10:
description: Build and run tests on GCC 10 and AVX 2 with a cmake dynamic build
executor: gcc10
environment: { CMAKE_FLAGS: -DBUILD_SHARED_LIBS=ON }
steps: [ cmake_test, cmake_install_test ]
dynamic-clang10:
description: Build and run tests on clang 10 and AVX 2 with a cmake dynamic build
executor: clang10
environment: { CMAKE_FLAGS: -DBUILD_SHARED_LIBS=ON }
steps: [ cmake_test, cmake_install_test ]
# unthreaded
unthreaded-gcc10:
description: Build and run tests on GCC 10 and AVX 2 *without* threads
executor: gcc10
environment: { CMAKE_FLAGS: -DSIMDJSON_ENABLE_THREADS=OFF }
steps: [ cmake_test, cmake_install_test ]
unthreaded-clang10:
description: Build and run tests on Clang 10 and AVX 2 *without* threads
executor: clang10
environment: { CMAKE_FLAGS: -DSIMDJSON_ENABLE_THREADS=OFF }
steps: [ cmake_test, cmake_install_test ]
# noexcept
noexcept-gcc10:
description: Build and run tests on GCC 10 and AVX 2 with exceptions off
executor: gcc10
environment: { CMAKE_FLAGS: -DSIMDJSON_EXCEPTIONS=OFF }
steps: [ cmake_test, cmake_install_test ]
noexcept-clang10:
description: Build and run tests on Clang 10 and AVX 2 with exceptions off
executor: clang10
environment: { CMAKE_FLAGS: -DSIMDJSON_EXCEPTIONS=OFF }
steps: [ cmake_test, cmake_install_test ]
#
# Misc.
#
# make (test and checkperf)
arch-haswell-gcc10:
description: Build, run tests and check performance on GCC 10 with -march=haswell
executor: gcc10
environment: { CXXFLAGS: -march=haswell }
steps: [ cmake_test ]
arch-nehalem-gcc10:
description: Build, run tests and check performance on GCC 10 with -march=nehalem
executor: gcc10
environment: { CXXFLAGS: -march=nehalem }
steps: [ cmake_test ]
sanitize-haswell-gcc10:
description: Build and run tests on GCC 10 and AVX 2 with a cmake sanitize build
executor: gcc10
environment: { CXXFLAGS: -march=haswell, CMAKE_FLAGS: -DCMAKE_BUILD_TYPE=Debug -DBUILD_SHARED_LIBS=ON -DSIMDJSON_SANITIZE=ON, CTEST_FLAGS: --output-on-failure -LE explicitonly }
steps: [ cmake_test ]
sanitize-haswell-clang10:
description: Build and run tests on clang 10 and AVX 2 with a cmake sanitize build
executor: clang10
environment: { CXXFLAGS: -march=haswell, CMAKE_FLAGS: -DBUILD_SHARED_LIBS=ON -DSIMDJSON_NO_FORCE_INLINING=ON -DSIMDJSON_SANITIZE=ON, CTEST_FLAGS: --output-on-failure -LE explicitonly }
steps: [ cmake_test ]
workflows:
version: 2.1
build_and_test:
jobs:
# full multi-implementation tests
#- gcc7 tested on GitHub actions
- gcc10 # do not delete this as it tests our performance
- clang6
#- clang10 # this gets tested a lot below
# libc++
- libcpp-clang10
# full single-implementation tests
- sanitize-gcc10
- sanitize-clang10
- threadsanitize-gcc10
- threadsanitize-clang10
- dynamic-gcc10
- dynamic-clang10
- unthreaded-gcc10
- unthreaded-clang10
# no exceptions
- noexcept-gcc10
- noexcept-clang10
# quicker make single-implementation tests
- arch-haswell-gcc10
- arch-nehalem-gcc10
# sanitized single-implementation tests
- sanitize-haswell-gcc10
- sanitize-haswell-clang10
# testing "just the library"
- justlib-gcc10
# testing asserts
- assert-gcc10
- assert-clang10
# TODO add windows: https://circleci.com/docs/2.0/configuration-reference/#windows
+1 -1
View File
@@ -38,7 +38,7 @@ If we cannot reproduce the issue, then we cannot address it. Note that a stack t
It should be possible to trigger the bug by using solely simdjson with our default build setup. If you can only observe the bug within some specific context, with some other software, please reduce the issue first.
**simjson release**
**simdjson release**
Unless you plan to contribute to simdjson, you should only work from releases. Please be mindful that our main branch may have additional features, bugs and documentation items.
+1
View File
@@ -6,6 +6,7 @@ Description
Type of change
- [ ] Bug fix
- [ ] Optimization
- [ ] New feature
- [ ] Refactor / cleanup
- [ ] Documentation / tests
+1 -1
View File
@@ -1,4 +1,4 @@
name: Ubuntu ppc64le (GCC 11)
name: Ubuntu aarch64 (GCC 13)
on:
push:
+8 -2
View File
@@ -2,7 +2,13 @@ name: Doxygen GitHub Pages
on:
release:
types: [created]
# Trigger when a release object is created and when it's published.
# Some GitHub flows create a release object then publish it later; include both.
types: [created, published]
# Also trigger on tag creation pushes so releasing via Git tags still runs the workflow
push:
tags:
- "v*" # common release tag pattern like v1.2.3
# Allows you to run this workflow manually from the Actions tab
workflow_dispatch:
@@ -27,7 +33,7 @@ jobs:
- name: Generate Doxygen Documentation
run: doxygen
- name: Deploy to GitHub Pages
uses: peaceiris/actions-gh-pages@v3
uses: peaceiris/actions-gh-pages@v4
with:
github_token: ${{ secrets.GITHUB_TOKEN }}
publish_dir: doc/api/html
+52 -22
View File
@@ -1,9 +1,18 @@
cmake_minimum_required(VERSION 3.14)
# Build performance optimizations
set(CMAKE_EXPORT_COMPILE_COMMANDS ON CACHE BOOL "Export compile commands for faster IDE integration")
set_property(GLOBAL PROPERTY USE_FOLDERS ON)
# Enable parallel compilation on MSVC
if(MSVC)
add_compile_options(/MP)
endif()
project(
simdjson
# The version number is modified by tools/release.py
VERSION 4.0.4
VERSION 4.2.4
DESCRIPTION "Parsing gigabytes of JSON per second"
HOMEPAGE_URL "https://simdjson.org/"
LANGUAGES CXX C
@@ -20,8 +29,8 @@ string(
# ---- Options, variables ----
# These version numbers are modified by tools/release.py
set(SIMDJSON_LIB_VERSION "27.0.0" CACHE STRING "simdjson library version")
set(SIMDJSON_LIB_SOVERSION "27" CACHE STRING "simdjson library soversion")
set(SIMDJSON_LIB_VERSION "29.0.0" CACHE STRING "simdjson library version")
set(SIMDJSON_LIB_SOVERSION "29" CACHE STRING "simdjson library soversion")
option(SIMDJSON_BUILD_STATIC_LIB "Build simdjson_static library along with simdjson (only makes sense if BUILD_SHARED_LIBS=ON)" OFF)
if(SIMDJSON_BUILD_STATIC_LIB AND NOT BUILD_SHARED_LIBS)
@@ -83,10 +92,36 @@ add_library(simdjson ${SIMDJSON_SOURCES})
add_library(simdjson::simdjson ALIAS simdjson)
set(SIMDJSON_LIBRARIES simdjson)
# Enable precompiled headers for faster builds
if(CMAKE_VERSION VERSION_GREATER_EQUAL "3.16")
target_precompile_headers(simdjson PRIVATE
<algorithm>
<array>
<atomic>
<bit>
<cassert>
<cctype>
<cerrno>
<cstddef>
<cstdint>
<cstdlib>
<cstring>
<memory>
<string>
<utility>
<vector>
)
endif()
if(SIMDJSON_BUILD_STATIC_LIB)
add_library(simdjson_static STATIC ${SIMDJSON_SOURCES})
add_library(simdjson::simdjson_static ALIAS simdjson_static)
list(APPEND SIMDJSON_LIBRARIES simdjson_static)
# Reuse precompiled headers for static library
if(CMAKE_VERSION VERSION_GREATER_EQUAL "3.16")
target_precompile_headers(simdjson_static REUSE_FROM simdjson)
endif()
endif()
set_target_properties(
@@ -112,6 +147,14 @@ simdjson_add_props(
PRIVATE "$<BUILD_INTERFACE:${PROJECT_SOURCE_DIR}/src>"
)
# Optimize linker settings for faster builds
if(MSVC)
target_link_options(simdjson PRIVATE /INCREMENTAL)
if(CMAKE_BUILD_TYPE STREQUAL "Debug")
target_link_options(simdjson PRIVATE /DEBUG:FASTLINK)
endif()
endif()
if(SIMDJSON_STATIC_REFLECTION)
# We would like to require C++26, but no compiler supports that!
# This is a hack:
@@ -140,24 +183,6 @@ if(SIMDJSON_MINUS_ZERO_AS_FLOAT)
simdjson_add_props(target_compile_definitions PRIVATE SIMDJSON_MINUS_ZERO_AS_FLOAT=1)
endif(SIMDJSON_MINUS_ZERO_AS_FLOAT)
if(CMAKE_SYSTEM_PROCESSOR MATCHES "^(loongarch64)$")
option(SIMDJSON_PREFER_LSX "Prefer LoongArch SX" ON)
include(CheckCXXCompilerFlag)
check_cxx_compiler_flag(-mlasx COMPILER_SUPPORTS_LASX)
check_cxx_compiler_flag(-mlsx COMPILER_SUPPORTS_LSX)
if(COMPILER_SUPPORTS_LASX AND NOT SIMDJSON_PREFER_LSX)
simdjson_add_props(
target_compile_options PRIVATE
-mlasx
)
elseif(COMPILER_SUPPORTS_LSX)
simdjson_add_props(
target_compile_options PRIVATE
-mlsx
)
endif()
endif()
# GCC and Clang have horrendous Debug builds when using SIMD.
# A common fix is to use '-Og' instead.
# bug https://gcc.gnu.org/bugzilla/show_bug.cgi?id=54412
@@ -171,7 +196,12 @@ if(
target_compile_options PRIVATE
$<$<CONFIG:DEBUG>:-Og>
)
endif()
# We still want to enable development checks in Debug mode
simdjson_add_props(
target_compile_definitions PUBLIC
SIMDJSON_DEVELOPMENT_CHECKS
)
endif()
if(SIMDJSON_ENABLE_THREADS)
find_package(Threads REQUIRED)
+1 -1
View File
@@ -38,7 +38,7 @@ PROJECT_NAME = simdjson
# could be handy for archiving the generated documentation or if some version
# control system is used.
PROJECT_NUMBER = "4.0.4"
PROJECT_NUMBER = "4.2.4"
# Using the PROJECT_BRIEF tag one can provide an optional one line description
# for a project that appears at the top of each page and should give viewer a
+25
View File
@@ -110,6 +110,24 @@ workflows used by simdjson.
Directory Structure and Source
------------------------------
Before diving into the directory structure, here are key concepts used in the codebase:
- **Amalgamated File**: A file that is conditionally included in the amalgamation process. These are wrapped in `#ifndef SIMDJSON_CONDITIONAL_INCLUDE` blocks and are included based on the target implementation (e.g., ARM64, x86). They include implementation-specific files (e.g., `arm64.h`) and generic files (e.g., under `generic/`). Amalgamated files have associated dependency files (`dependencies.h`) to track includes.
- **Amalgamator File**: A file that orchestrates the inclusion of amalgamated files. Examples: `arm64.h`, `arm64/implementation.h`, `generic/amalgamated.h`. These are not themselves amalgamated but control conditional inclusions.
- **Free Dependency File**: A top-level header that is always included unconditionally. These do not have dependency files and represent the public API (e.g., main headers).
- **Implementation-Specific File**: A file tied to a specific CPU architecture or instruction set (e.g., `arm64/`, `haswell/`). These must be amalgamated.
- **Generic File**: A shared file (under `generic/` or `simdjson/generic/`) that contains common code included once per implementation.
- **Builtin File**: Special files under `simdjson/builtin/` that handle the builtin implementation, a fallback/default implementation used when no optimized implementation is available.
- **Conditional Include Block**: A section wrapped in `#ifndef SIMDJSON_CONDITIONAL_INCLUDE` for editor-only or implementation-specific content.
The script `singleheader/amalgation_helper.py` will generate an HTML report which you can use to visualize the status of each file.
simdjson's source structure, from the top level, looks like this:
* **CMakeLists.txt:** The main build system.
@@ -133,6 +151,12 @@ simdjson's source structure, from the top level, looks like this:
* simdjson/generic/ondemand/*.h: individual On-Demand classes, generically written.
* simdjson/generic/ondemand/dependencies.h: dependencies on common, non-implementation-specific simdjson classes. This will be included before including amalgamated.h.
* simdjson/generic/ondemand/amalgamated.h: all generic ondemand classes for an implementation.
* simdjson/builder.h: the `simdjson::builder` namespace. Includes all public builder classes.
* simdjson/builtin/builder.h: the `simdjson::builtin::builder` namespace.
* simdjson/arm64|fallback|haswell|icelake|ppc64|westmere/builder.h: the `simdjson::<implementation>::builder` namespace. Builder compiled for the specific implementation.
* simdjson/generic/builder/*.h: individual Builder classes, generically written.
* simdjson/generic/builder/dependencies.h: dependencies on common, non-implementation-specific simdjson classes. This will be included before including amalgamated.h.
* simdjson/generic/builder/amalgamated.h: all generic builder classes for an implementation.
* **src:** The source files for non-inlined functionality (e.g. the architecture-specific parser
implementations).
* simdjson.cpp: A "main source" that includes all implementation files from src/. This is
@@ -147,6 +171,7 @@ Other important files and directories:
* **.github/workflows:** Definitions for GitHub Actions (CI).
* **singleheader:** Contains generated `simdjson.h` and `simdjson.cpp` that we release. The files `singleheader/simdjson.h` and `singleheader/simdjson.cpp` should never be edited by hand.
* **singleheader/amalgamate.py:** Generates `singleheader/simdjson.h` and `singleheader/simdjson.cpp` for release (python script). If you add a new implementation (e.g., rvv), you need to edit this file (IMPLEMENTATIONS).
* **singleheader/amalgation_helper.py:** Generates and `amalgamation_report.html` that helps you understand the status of each file.
* **benchmark:** This is where we do benchmarking. Benchmarking is core to every change we make; the
cardinal rule is don't regress performance without knowing exactly why, and what you're trading
for it. Many of our benchmarks are microbenchmarks. We are effectively doing controlled scientific experiments for the purpose of understanding what affects our performance. So we simplify as much as possible. We try to avoid irrelevant factors such as page faults, interrupts, unnecessary system calls. We recommend checking the performance as follows:
+31 -2
View File
@@ -6,7 +6,8 @@
simdjson : Parsing gigabytes of JSON per second
===============================================
<img src="images/logo.png" width="10%" style="float: right">
<img src="images/official_logo/logo_noir/SVG/logo_simdjson_noir.svg" width="40%" style="float: right">
JSON is everywhere on the Internet. Servers spend a *lot* of time parsing it. We need a fresh
approach. The simdjson library uses commonly available SIMD instructions and microparallel algorithms
to parse JSON 4x faster than RapidJSON and 25x faster than JSON for Modern C++.
@@ -62,10 +63,14 @@ Real-world usage
- [WasmEdge](https://wasmedge.org)
- [RonDB](https://github.com/logicalclocks/rondb)
- [GreptimeDB](https://github.com/GreptimeTeam/greptimedb)
- [mamba](https://github.com/mamba-org/mamba)
If you are planning to use simdjson in a product, please work from one of our releases.
Quick Start
-----------
@@ -82,7 +87,7 @@ The simdjson library is easily consumable with a single .h and .cpp file.
```
2. Create `quickstart.cpp`:
```c++
```cpp
#include <iostream>
#include "simdjson.h"
using namespace simdjson;
@@ -112,6 +117,8 @@ Usage documentation is available:
* [Implementation Selection](doc/implementation-selection.md) describes runtime CPU detection and
how you can work with it.
* [API](https://simdjson.github.io/simdjson/) contains the automatically generated API documentation.
* [Compile-Time Parsing](doc/compile_time.md) presents our compile-time parsing function (C++26 only).
Godbolt
-------------
@@ -206,6 +213,21 @@ For the video inclined, <br />
[![simdjson at QCon San Francisco 2019](http://img.youtube.com/vi/wlvKAT7SZIQ/0.jpg)](http://www.youtube.com/watch?v=wlvKAT7SZIQ)<br />
(It was the best voted talk, we're kinda proud of it.)
Citing this work
-----------------
If you use simdjson in published research, please cite the software library. A suitable BibTeX entry is:
```bibtex
@misc{simdjson,
title={{The simdjson library: Parsing Gigabytes of JSON per Second}},
author={Daniel Lemire and Geoff Langdale and John Keiser and Paul Dreik and Francisco Thiesen and others},
year={2019},
howpublished={Software library},
note={https://github.com/simdjson/simdjson}
}
```
Funding
-------
@@ -226,6 +248,13 @@ Contributing to simdjson
Head over to [CONTRIBUTING.md](CONTRIBUTING.md) for information on contributing to simdjson, and
[HACKING.md](HACKING.md) for information on source, building, and architecture/design.
Stars
------
[![Star History Chart](https://api.star-history.com/svg?repos=simdjson/simdjson&type=Date)](https://www.star-history.com/#simdjson/simdjson&Date)
License
-------
+46
View File
@@ -0,0 +1,46 @@
# Accessor Performance Benchmarks (C++26)
These benchmarks compare the performance of runtime vs compile-time JSON accessors.
For the comparison to be meaningful, you must build simdjson with support for
C++26 reflexion. See the `p2996` repository in the main project directory.
## Files
- `accessor_benchmark.h` - Common benchmark framework and test data
- `runtime_accessors.h` - Runtime `at_path()` benchmarks
- `compile_time_accessors.h` - Compile-time `at_path_compiled()` benchmarks (requires C++26 reflection)
## Benchmarks
Each benchmark measures parsing + single field access:
1. **accessor_simple** - Simple field: `.name`
2. **accessor_nested** - Nested field: `.address.city`
3. **accessor_deep** - Deep nested field: `.address.coordinates.lat`
## Building (Linux/macOS)
```bash
cmake -B build -D SIMDJSON_STATIC_REFLECTION=ON -DSIMDJSON_DEVELOPER_MODE=ON
cmake --build build --target=bench_ondemand
```
The `SIMDJSON_STATIC_REFLECTION` will be made unnecessary once mainstream compilers
begin supporting C++26 sufficiently well.
## Running (Linux/macOS)
```bash
# Run all accessor benchmarks
./build/bench_ondemand --benchmark_filter="accessor"
```
## Results
We find that compile-time accessors show performance improvements that scale with path depth:
- Simple fields: ~1.2x faster
- Nested fields: ~1.5x faster
- Deep nested fields: ~1.8x faster
The speedup comes from eliminating runtime path parsing and conversion overhead.
@@ -0,0 +1,132 @@
#pragma once
#include "json_benchmark/file_runner.h"
#include <string>
namespace accessor_performance {
using namespace json_benchmark;
// Test JSON for accessor benchmarks
static const char* TEST_JSON = R"({
"name": "Alice",
"age": 30,
"email": "alice@example.com",
"address": {
"street": "123 Main St",
"city": "Boston",
"state": "MA",
"zip": 12345,
"coordinates": {
"lat": 42.3601,
"lon": -71.0589
}
},
"scores": [95, 87, 92, 88, 91],
"preferences": {
"theme": "dark",
"notifications": {
"email": true,
"push": false,
"sms": true
}
}
})";
// Struct definitions for compile-time validation
#if SIMDJSON_STATIC_REFLECTION
struct Coordinates {
double lat;
double lon;
};
struct Address {
std::string street;
std::string city;
std::string state;
int64_t zip;
Coordinates coordinates;
};
struct Notifications {
bool email;
bool push;
bool sms;
};
struct Preferences {
std::string theme;
Notifications notifications;
};
struct TestData {
std::string name;
int64_t age;
std::string email;
Address address;
std::vector<int64_t> scores;
Preferences preferences;
};
#endif // SIMDJSON_STATIC_REFLECTION
// Single-access benchmark runner: measures ONE field access per iteration
template<typename I>
struct single_access_runner : public file_runner<I> {
std::string result_string;
int64_t result_int{};
double result_double{};
bool result_bool{};
bool setup(benchmark::State &state) {
this->json = simdjson::padded_string(TEST_JSON, strlen(TEST_JSON));
state.SetBytesProcessed(int64_t(state.iterations()) * int64_t(this->json.size()));
return true;
}
bool before_run(benchmark::State &state) {
if (!file_runner<I>::before_run(state)) { return false; }
result_string.clear();
result_int = 0;
result_double = 0.0;
result_bool = false;
return true;
}
bool run(benchmark::State &) {
return this->implementation.run(this->json, result_string, result_int, result_double, result_bool);
}
template<typename R>
bool diff(benchmark::State &state, single_access_runner<R> &reference) {
if (result_string != reference.result_string ||
result_int != reference.result_int ||
result_double != reference.result_double ||
result_bool != reference.result_bool) {
std::cerr << "Accessor benchmark results differ!" << std::endl;
return false;
}
return true;
}
size_t items_per_iteration() {
return 1;
}
};
// Benchmark template definitions
struct runtime_at_path_simple;
template<typename I> simdjson_inline static void accessor_simple(benchmark::State &state) {
run_json_benchmark<single_access_runner<I>, single_access_runner<runtime_at_path_simple>>(state);
}
struct runtime_at_path_nested;
template<typename I> simdjson_inline static void accessor_nested(benchmark::State &state) {
run_json_benchmark<single_access_runner<I>, single_access_runner<runtime_at_path_nested>>(state);
}
struct runtime_at_path_deep;
template<typename I> simdjson_inline static void accessor_deep(benchmark::State &state) {
run_json_benchmark<single_access_runner<I>, single_access_runner<runtime_at_path_deep>>(state);
}
} // namespace accessor_performance
@@ -0,0 +1,56 @@
#pragma once
#if SIMDJSON_EXCEPTIONS && SIMDJSON_STATIC_REFLECTION
#include "accessor_benchmark.h"
namespace accessor_performance {
using namespace simdjson;
struct compile_time_at_path_simple {
ondemand::parser parser{};
bool run(simdjson::padded_string &json, std::string &result_str, int64_t&, double&, bool&) {
auto doc = parser.iterate(json);
std::string_view name;
auto r = ondemand::json_path::at_path_compiled<TestData, ".name">(doc);
if (r.get(name) != SUCCESS) return false;
result_str = name;
return true;
}
};
struct compile_time_at_path_nested {
ondemand::parser parser{};
bool run(simdjson::padded_string &json, std::string &result_str, int64_t&, double&, bool&) {
auto doc = parser.iterate(json);
std::string_view city;
auto r = ondemand::json_path::at_path_compiled<TestData, ".address.city">(doc);
if (r.get(city) != SUCCESS) return false;
result_str = city;
return true;
}
};
struct compile_time_at_path_deep {
ondemand::parser parser{};
bool run(simdjson::padded_string &json, std::string&, int64_t&, double &result_dbl, bool&) {
auto doc = parser.iterate(json);
double lat;
auto r = ondemand::json_path::at_path_compiled<TestData, ".address.coordinates.lat">(doc);
if (r.get(lat) != SUCCESS) return false;
result_dbl = lat;
return true;
}
};
BENCHMARK_TEMPLATE(accessor_simple, compile_time_at_path_simple)->UseManualTime();
BENCHMARK_TEMPLATE(accessor_nested, compile_time_at_path_nested)->UseManualTime();
BENCHMARK_TEMPLATE(accessor_deep, compile_time_at_path_deep)->UseManualTime();
} // namespace accessor_performance
#endif // SIMDJSON_EXCEPTIONS && SIMDJSON_STATIC_REFLECTION
@@ -0,0 +1,53 @@
#pragma once
#if SIMDJSON_EXCEPTIONS
#include "accessor_benchmark.h"
namespace accessor_performance {
using namespace simdjson;
struct runtime_at_path_simple {
ondemand::parser parser{};
bool run(simdjson::padded_string &json, std::string &result_str, int64_t&, double&, bool&) {
auto doc = parser.iterate(json);
std::string_view name;
if (doc.at_path(".name").get(name) != SUCCESS) return false;
result_str = name;
return true;
}
};
struct runtime_at_path_nested {
ondemand::parser parser{};
bool run(simdjson::padded_string &json, std::string &result_str, int64_t&, double&, bool&) {
auto doc = parser.iterate(json);
std::string_view city;
if (doc.at_path(".address.city").get(city) != SUCCESS) return false;
result_str = city;
return true;
}
};
struct runtime_at_path_deep {
ondemand::parser parser{};
bool run(simdjson::padded_string &json, std::string&, int64_t&, double &result_dbl, bool&) {
auto doc = parser.iterate(json);
double lat;
if (doc.at_path(".address.coordinates.lat").get(lat) != SUCCESS) return false;
result_dbl = lat;
return true;
}
};
BENCHMARK_TEMPLATE(accessor_simple, runtime_at_path_simple)->UseManualTime();
BENCHMARK_TEMPLATE(accessor_nested, runtime_at_path_nested)->UseManualTime();
BENCHMARK_TEMPLATE(accessor_deep, runtime_at_path_deep)->UseManualTime();
} // namespace accessor_performance
#endif // SIMDJSON_EXCEPTIONS
+3 -3
View File
@@ -245,7 +245,7 @@ static u32 (*kpc_get_counter_count)(u32 classes);
/// Get counter accumulations.
/// If `all_cpus` is true, the buffer count should not smaller than
/// (cpu_count * counter_count). Otherwize, the buffer count should not smaller
/// (cpu_count * counter_count). Otherwise, the buffer count should not smaller
/// than (counter_count).
/// @see kpc_get_counter_count(), kpc_cpu_count().
/// @param all_cpus true for all CPUs, false for current cpu.
@@ -374,7 +374,7 @@ static int kperf_lightweight_pet_set(u32 enabled) {
// These functions do not require root privileges.
// -----------------------------------------------------------------------------
// KPEP CPU archtecture constants.
// KPEP CPU architecture constants.
#define KPEP_ARCH_I386 0
#define KPEP_ARCH_X86_64 1
#define KPEP_ARCH_ARM 2
@@ -414,7 +414,7 @@ typedef struct kpep_db {
usize fixed_counter_count;
usize config_counter_count;
usize power_counter_count;
u32 archtecture; ///< see `KPEP CPU archtecture constants` above.
u32 architecture; ///< see `KPEP CPU architecture constants` above.
u32 fixed_counter_bits;
u32 config_counter_bits;
u32 power_counter_bits;
+5
View File
@@ -148,4 +148,9 @@ SIMDJSON_POP_DISABLE_WARNINGS
#include "large_amazon_cellphones/simdjson_dom.h"
#include "large_amazon_cellphones/simdjson_ondemand.h"
#include "accessor_performance/runtime_accessors.h"
#if SIMDJSON_STATIC_REFLECTION
#include "accessor_performance/compile_time_accessors.h"
#endif
BENCHMARK_MAIN();
+9 -2
View File
@@ -44,7 +44,10 @@ struct yyjson_base {
struct yyjson : yyjson_base {
bool run(simdjson::padded_string &json, std::vector<uint64_t> &result) {
return yyjson_base::run(yyjson_read(json.data(), json.size(), 0), result);
yyjson_doc *doc = yyjson_read(json.data(), json.size(), 0);
bool b = yyjson_base::run(doc, result);
yyjson_doc_free(doc);
return b;
}
};
BENCHMARK_TEMPLATE(distinct_user_id, yyjson)->UseManualTime();
@@ -52,11 +55,15 @@ BENCHMARK_TEMPLATE(distinct_user_id, yyjson)->UseManualTime();
#if SIMDJSON_COMPETITION_ONDEMAND_INSITU
struct yyjson_insitu : yyjson_base {
bool run(simdjson::padded_string &json, std::vector<uint64_t> &result) {
return yyjson_base::run(yyjson_read_opts(json.data(), json.size(), YYJSON_READ_INSITU, 0, 0), result);
yyjson_doc *doc = yyjson_read_opts(json.data(), json.size(), YYJSON_READ_INSITU, 0, 0);
bool b = yyjson_base::run(doc, result);
yyjson_doc_free(doc);
return b;
}
};
BENCHMARK_TEMPLATE(distinct_user_id, yyjson_insitu)->UseManualTime();
#endif // SIMDJSON_COMPETITION_ONDEMAND_INSITU
} // namespace distinct_user_id
#endif // SIMDJSON_COMPETITION_YYJSON
+10 -2
View File
@@ -36,18 +36,26 @@ struct yyjson_base {
struct yyjson : yyjson_base {
bool run(simdjson::padded_string &json, uint64_t find_id, std::string_view &result) {
return yyjson_base::run(yyjson_read(json.data(), json.size(), 0), find_id, result);
yyjson_doc *doc = yyjson_read(json.data(), json.size(), 0);
bool b = yyjson_base::run(doc, find_id, result);
yyjson_doc_free(doc);
return b;
}
};
BENCHMARK_TEMPLATE(find_tweet, yyjson)->UseManualTime();
#if SIMDJSON_COMPETITION_ONDEMAND_INSITU
struct yyjson_insitu : yyjson_base {
bool run(simdjson::padded_string &json, uint64_t find_id, std::string_view &result) {
return yyjson_base::run(yyjson_read_opts(json.data(), json.size(), YYJSON_READ_INSITU, 0, 0), find_id, result);
yyjson_doc *doc = yyjson_read_opts(json.data(), json.size(), YYJSON_READ_INSITU, 0, 0);
bool b = yyjson_base::run(doc, find_id, result);
yyjson_doc_free(doc);
return b;
}
};
BENCHMARK_TEMPLATE(find_tweet, yyjson_insitu)->UseManualTime();
#endif // SIMDJSON_COMPETITION_ONDEMAND_INSITU
} // namespace find_tweet
#endif // SIMDJSON_COMPETITION_YYJSON
+2
View File
@@ -100,6 +100,7 @@ struct yyjson : yyjson2msgpack {
std::string_view &result) {
yyjson_doc *doc = yyjson_read(json.data(), json.size(), 0);
result = to_msgpack(doc, reinterpret_cast<uint8_t*>(buffer));
yyjson_doc_free(doc);
return true;
}
};
@@ -113,6 +114,7 @@ struct yyjson_insitu : yyjson2msgpack {
yyjson_doc *doc =
yyjson_read_opts(json.data(), json.size(), YYJSON_READ_INSITU, 0, 0);
result = to_msgpack(doc, reinterpret_cast<uint8_t*>(buffer));
yyjson_doc_free(doc);
return true;
}
};
+10 -2
View File
@@ -49,18 +49,26 @@ struct yyjson_base {
struct yyjson : yyjson_base {
bool run(simdjson::padded_string &json, std::vector<point> &result) {
return yyjson_base::run(yyjson_read(json.data(), json.size(), 0), result);
yyjson_doc *doc = yyjson_read(json.data(), json.size(), 0);
bool b = yyjson_base::run(doc, result);
yyjson_doc_free(doc);
return b;
}
};
BENCHMARK_TEMPLATE(kostya, yyjson)->UseManualTime();
#if SIMDJSON_COMPETITION_ONDEMAND_INSITU
struct yyjson_insitu : yyjson_base {
bool run(simdjson::padded_string &json, std::vector<point> &result) {
return yyjson_base::run(yyjson_read_opts(json.data(), json.size(), YYJSON_READ_INSITU, 0, 0), result);
yyjson_doc *doc = yyjson_read_opts(json.data(), json.size(), YYJSON_READ_INSITU, 0, 0);
bool b = yyjson_base::run(doc, result);
yyjson_doc_free(doc);
return b;
}
};
BENCHMARK_TEMPLATE(kostya, yyjson_insitu)->UseManualTime();
#endif // SIMDJSON_COMPETITION_ONDEMAND_INSITU
} // namespace kostya
#endif // SIMDJSON_COMPETITION_YYJSON
+10 -2
View File
@@ -47,18 +47,26 @@ struct yyjson_base {
struct yyjson : yyjson_base {
bool run(simdjson::padded_string &json, std::vector<point> &result) {
return yyjson_base::run(yyjson_read(json.data(), json.size(), 0), result);
yyjson_doc *doc = yyjson_read(json.data(), json.size(), 0);
bool b = yyjson_base::run(doc, result);
yyjson_doc_free(doc);
return b;
}
};
BENCHMARK_TEMPLATE(large_random, yyjson)->UseManualTime();
#if SIMDJSON_COMPETITION_ONDEMAND_INSITU
struct yyjson_insitu : yyjson_base {
bool run(simdjson::padded_string &json, std::vector<point> &result) {
return yyjson_base::run(yyjson_read_opts(json.data(), json.size(), YYJSON_READ_INSITU, 0, 0), result);
yyjson_doc *doc = yyjson_read_opts(json.data(), json.size(), YYJSON_READ_INSITU, 0, 0);
bool b = yyjson_base::run(doc, result);
yyjson_doc_free(doc);
return b;
}
};
BENCHMARK_TEMPLATE(large_random, yyjson_insitu)->UseManualTime();
#endif // SIMDJSON_COMPETITION_ONDEMAND_INSITU
} // namespace large_random
#endif // SIMDJSON_COMPETITION_YYJSON
+10 -3
View File
@@ -62,19 +62,26 @@ struct yyjson_base {
struct yyjson : yyjson_base {
bool run(simdjson::padded_string &json, std::vector<tweet<std::string_view>> &result) {
return yyjson_base::run(yyjson_read(json.data(), json.size(), 0), result);
yyjson_doc *doc = yyjson_read(json.data(), json.size(), 0);
bool b = yyjson_base::run(doc, result);
yyjson_doc_free(doc);
return b;
}
};
BENCHMARK_TEMPLATE(partial_tweets, yyjson)->UseManualTime();
#if SIMDJSON_COMPETITION_ONDEMAND_INSITU
struct yyjson_insitu : yyjson_base {
bool run(simdjson::padded_string &json, std::vector<tweet<std::string_view>> &result) {
return yyjson_base::run(yyjson_read_opts(json.data(), json.size(), YYJSON_READ_INSITU, 0, 0), result);
yyjson_doc *doc = yyjson_read_opts(json.data(), json.size(), YYJSON_READ_INSITU, 0, 0);
bool b = yyjson_base::run(doc, result);
yyjson_doc_free(doc);
return b;
}
};
BENCHMARK_TEMPLATE(partial_tweets, yyjson_insitu)->UseManualTime();
#endif // SIMDJSON_COMPETITION_ONDEMAND_INSITU
} // namespace partial_tweets
#endif // SIMDJSON_COMPETITION_YYJSON
+10 -2
View File
@@ -51,18 +51,26 @@ struct yyjson_base {
struct yyjson : yyjson_base {
bool run(simdjson::padded_string &json, int64_t max_retweet_count, top_tweet_result<StringType> &result) {
return yyjson_base::run(yyjson_read(json.data(), json.size(), 0), max_retweet_count, result);
yyjson_doc *doc = yyjson_read(json.data(), json.size(), 0);
bool b = yyjson_base::run(doc, max_retweet_count, result);
yyjson_doc_free(doc);
return b;
}
};
BENCHMARK_TEMPLATE(top_tweet, yyjson)->UseManualTime();
#if SIMDJSON_COMPETITION_ONDEMAND_INSITU
struct yyjson_insitu : yyjson_base {
bool run(simdjson::padded_string &json, int64_t max_retweet_count, top_tweet_result<StringType> &result) {
return yyjson_base::run(yyjson_read_opts(json.data(), json.size(), YYJSON_READ_INSITU, 0, 0), max_retweet_count, result);
yyjson_doc *doc = yyjson_read_opts(json.data(), json.size(), YYJSON_READ_INSITU, 0, 0);
bool b = yyjson_base::run(doc, max_retweet_count, result);
yyjson_doc_free(doc);
return b;
}
};
BENCHMARK_TEMPLATE(top_tweet, yyjson_insitu)->UseManualTime();
#endif // SIMDJSON_COMPETITION_ONDEMAND_INSITU
} // namespace top_tweet
#endif // SIMDJSON_COMPETITION_YYJSON
+4
View File
@@ -174,6 +174,10 @@ else()
-Werror -Wall -Wextra -Weffc++ -Wsign-compare -Wshadow -Wwrite-strings
-Wpointer-arith -Winit-self -Wconversion -Wno-sign-conversion
)
if(CMAKE_CXX_STANDARD VERSION_GREATER_EQUAL 20)
target_compile_options(simdjson-internal-flags INTERFACE -Wctad-maybe-unsupported)
endif()
endif()
option(SIMDJSON_GLIBCXX_ASSERTIONS "Set _GLIBCXX_ASSERTIONS" OFF)
+3 -2
View File
@@ -17,8 +17,9 @@ editing CMAKE_CXX_FLAGS")
# /EHc used in conjection with /EHs indicates that extern "C" functions
# never throw (terminate-on-throw)
# Here, we disable both with the - argument negation operator
string(REPLACE "/EHsc" "/EHs-c-" CMAKE_CXX_FLAGS ${CMAKE_CXX_FLAGS})
if(CMAKE_CXX_FLAGS)
string(REPLACE "/EHsc" "/EHs-c-" CMAKE_CXX_FLAGS ${CMAKE_CXX_FLAGS})
endif()
# Because we cannot change the flag above on an individual target (yet), the
# definition below must similarly be added globally
add_definitions(-D_HAS_EXCEPTIONS=0)
+250 -112
View File
File diff suppressed because it is too large Load Diff
+6 -3
View File
@@ -1,5 +1,8 @@
We take our documentation seriously. Please start reading the documentation before you attempt to use simdjson. We hope you will enjoy reading us.
* Basics: https://github.com/simdjson/simdjson/blob/master/doc/basics.md is an overview of how to use simdjson and its APIs.
* iterate_many: https://github.com/simdjson/simdjson/blob/master/doc/iterate_many.md describes an interface providing features to work with files or streams containing multiple small JSON documents. As fast and convenient as possible.
* Performance: https://github.com/simdjson/simdjson/blob/master/doc/performance.md shows some more advanced scenarios and how to tune for them.
* [Basics](doc/basics.md) is an overview of how to use simdjson and its APIs.
* [Builder](doc/builder.md) is an overview of how to efficiently write JSON strings using simdjson.
* [Performance](doc/performance.md) shows some more advanced scenarios and how to tune for them.
* [Implementation Selection](doc/implementation-selection.md) describes runtime CPU detection and
how you can work with it.
* [API](https://simdjson.github.io/simdjson/) contains the automatically generated API documentation.
+158 -2
View File
@@ -14,6 +14,7 @@ speed and high convenience.
* [C++26 static reflection](#c--26-static-reflection)
+ [Without `string_buffer` instance](#without--string-buffer--instance)
+ [Without `string_buffer` instance but with explicit error handling](#without--string-buffer--instance-but-with-explicit-error-handling)
+ [Pretty formatted (fractured JSON)](#pretty-formatted-fractured-json)
Overview: string_builder
---------------------------
@@ -60,7 +61,7 @@ The later method (`view()`) is recommended. For performance reasons, we expect
Example: string_builder
---------------------------
```C++
```cpp
struct Car {
std::string make;
std::string model;
@@ -176,9 +177,52 @@ std::vector<std::vector<double>> c = {{1.0, 2.0}, {3.0, 4.0}};
std::string json = simdjson::to_json(c);
```
We also have an overload for when you want to reuse the same `std::string` instance:
```cpp
std::vector<std::vector<double>> c = {{1.0, 2.0}, {3.0, 4.0}};
std::string json;
auto error = simdjson::to_json(c, json);
if(error) { /* there was an error */ }
```
We do recommend that you create and reuse the `string_builder` instance for performance
reasons.
You can also add custom serialization functions using a `tag_invoke` function.
For example, the following
function will allow you to serialize instances of the type `Car`.
```cpp
#include <simdjson>
struct Car {
std::string make;
std::string model;
int64_t year;
std::vector<float> tire_pressure;
};
namespace simdjson {
template <typename builder_type>
void tag_invoke(serialize_tag, builder_type &builder, const Car& car) {
builder.start_object();
builder.append_key_value("make", car.make);
builder.append_comma();
builder.append_key_value("model", car.model);
builder.append_comma();
builder.append_key_value("year", car.year);
builder.append_comma();
builder.append_key_value("tire_pressure", car.tire_pressure);
builder.end_object();
}
} // namespace simdjson
```
C++26 static reflection
------------------------
@@ -240,7 +284,38 @@ with the `simdjson::to_json` template function.
If you know the output size, in bytes, of your JSON string, you may
pass it as a second parameter (e.g., `simdjson::to_json(c, 31123)`).
Sometimes you may want to reuse the same `std::string` instance. We
have an overload for this purpose:
```cpp
Car c = {"Toyota", "Corolla", 2017, {30.0,30.2,30.513,30.79}};
std::string s;
auto error = simdjson::to_json(c, s);
if(error) { /* there was an error */ }
```
You can then also add a third parameter for the expected output size in bytes.
### Extracting just some fields
In some instances, your class might have many fields that you do not want to serialize.
You can achieve this result with the `simdjson::extract_from` template. In the following
example, we serialize only the `year` and `price` fields on the `Car` instance.
```cpp
struct Car {
std::string make;
std::string model;
int year;
double price;
bool electric;
};
Car car{"Ford", "F-150", 2024, 55000.0, false};
// Extract year and price
std::string json_result = simdjson::extract_from<"year", "price">(car);
// Alternatively:
// std::string json_result;
// auto error = extract_from<"year", "price">(car).get(json_result);
// if(error) { /* error handling */ }
```
### Without `string_buffer` instance but with explicit error handling
@@ -249,9 +324,90 @@ pattern:
```cpp
std::string json;
if(simdjson::to(c).get(json)) {
if(simdjson::to_json(c).get(json)) {
// there was an error
} else {
// json contain the serialized JSON
}
```
### Customization
If you want to serialize a value in a custom way, you can do it with a
`tag_invoke` specialization like the following example which will map
the year attribute to a string.
```cpp
#include <simdjson>
struct Car {
std::string make;
std::string model;
int64_t year;
std::vector<float> tire_pressure;
};
namespace simdjson {
template <typename builder_type>
void tag_invoke(serialize_tag, builder_type &builder, const Car& car) {
builder.start_object();
builder.append_key_value("make", car.make);
builder.append_comma();
builder.append_key_value("model", car.model);
builder.append_comma();
builder.append_key_value("year", std::to_string(car.year));
builder.append_comma();
builder.append_key_value("tire_pressure", car.tire_pressure);
builder.end_object();
}
} // namespace simdjson
```
### Pretty formatted (fractured JSON)
In some instances, you may want your JSON to be more readable. For this pupose, we also
support the Fractured JSON standard.
```Cpp
TableTestData data{
{{1, "Alice", true}, {2, "Bob", false}, {3, "Carol", true}, {4, "Dave", false}}
};
fractured_json_options opts;
opts.enable_table_format = true;
opts.min_table_rows = 3;
std::string formatted = simdjson::to_fractured_json_string(data, opts);
```
The result might be as follows.
```json
{
"records": [
{ "active": true , "id": 1, "name": "Alice" },
{ "active": false, "id": 2, "name": "Bob" },
{ "active": true , "id": 3, "name": "Carol" },
{ "active": false, "id": 4, "name": "Dave" }
]
}
```
The `fractured_json_options` struct allows you to customize the formatting behavior. It includes the following options:
- `max_total_line_length` (default: 120): Maximum total characters per line. Content exceeding this will be expanded to multiple lines.
- `max_inline_length` (default: 80): Maximum length for inlined elements. Simple arrays/objects shorter than this may be rendered inline.
- `max_inline_complexity` (default: 2): Maximum nesting depth for inline rendering. Elements with complexity exceeding this will be expanded. Complexity 0 = scalar, 1 = flat array/object, 2 = one level of nesting.
- `max_compact_array_complexity` (default: 1): Maximum complexity for compact array formatting. Arrays with elements of this complexity or less may have multiple items per line.
- `indent_spaces` (default: 4): Number of spaces per indentation level.
- `enable_table_format` (default: true): Enable tabular formatting for arrays of similar objects. When enabled, arrays of objects with identical keys are formatted as aligned tables.
- `min_table_rows` (default: 3): Minimum number of rows to trigger table mode.
- `table_similarity_threshold` (default: 0.8): Similarity threshold for table detection. Objects must share at least this fraction of keys to be formatted as a table.
- `enable_compact_multiline` (default: true): Enable compact multiline arrays. When enabled, arrays of simple elements may have multiple items per line.
- `max_items_per_line` (default: 10): Maximum array items per line in compact mode.
- `simple_bracket_padding` (default: true): Add space inside brackets for simple containers. When true: `{ "key": "value" }`, when false: `{"key": "value"}`.
- `colon_padding` (default: true): Add space after colons. When true: `"key": "value"`, when false: `"key":"value"`.
- `comma_padding` (default: true): Add space after commas in inline content. When true: `[1, 2, 3]`, when false: `[1,2,3]`.
+227
View File
@@ -0,0 +1,227 @@
# Parse json at compile time
* [Introduction](#introduction)
* [Example](#example)
* [Concepts](#concepts)
* [Loading from disk](#loading-from-disk)
* [Limitations (compile-time errors)](#limitations-compile-time-errors)
## Introduction
In some instances, you may want to configure your software at compile-time with a JSON document.
Maybe you have a single code base but many different possible configurations, all resulting in
different software. For example, you might be programming robots, using the same software, but
different robot configurations.
To achieve the desired result, you have a few options. You may start the software and parser a
JSON file at runtime. Or you might convert your JSON data into C++ code that you can compile with
your software.
With C++26, there is another way: parse the JSON file along with your C++ code. In this manner,
the JSON data becomes native C++ data.
The simdjson library supports parsing JSON documents at compile time if you have C++26 support. To
activate C++26 reflection support, you can compile
your code with the `SIMDJSON_STATIC_REFLECTION` macro set:
```cpp
#define SIMDJSON_STATIC_REFLECTION 1
//...
#include "simdjson.h"
```
The `simdjson::compile_time::parse_json` function parses a JSON document at **compile time** and returns a `constexpr` structure reflecting its content. We support the full range of JSON values, which are mapped to C++ types as in
the following table.
| JSON type | C++ type |
|----------------|----------------------------------|
| object | anonymous struct |
| array | `std::array<T, N>` (homogeneous) |
| string | `const char*` (UTF-8) |
| number | `int64_t`, `uint64_t`, `double` |
| `true`/`false` | `bool` |
| `null` | `std::nullptr_t` |
## Example
Suppose you want to parse the following JSON document:
```cpp
{
"port": 8080,
"host": "localhost",
"debug": true
}
```
**Reminder**: In C++, `R"( )"` allows us to write multi-line strings with unescaped quotes.
You can do so, at compile-time, as follows:
```cpp
constexpr auto cfg = R"(
{
"port": 8080,
"host": "localhost",
"debug": true
}
)"_json;
// cfg.port == 8080
// std::string_view(cfg.host) == "localhost"
// cfg.debug == true
```
You can nest objects and arrays:
```cpp
constexpr auto data = R"(
{
"servers": [
{"host": "s1", "port": 3000},
{"host": "s2", "port": 3001}
]
}
)"_json;
// data.servers.size() == 2
// std::string_view(data.servers[0].host) == "s1"
```
Top-level arrays are allowed:
```cpp
constexpr auto arr = R"(
[1, 2, 3]
)"_json;
static_assert(arr.size() == 3);
static_assert(arr[1] == 2);
```
## Concepts
Given that the parsed data is made of structures that depend on the JSON input, you might
want to check that it conforms to your expectation. You can do so with concepts.
Let us consider this example:
```cpp
constexpr auto config = R"(
[
{ "name": "Alice", "age": 30 },
{ "name": "Bob", "age": 25 },
{ "name": "Charlie", "age": 35 }
]
)"_json;
```
You might want to ensure that the result is an array of persons. You can define your
expectation with concepts like so:
```cpp
template <typename T>
concept person = requires(T p) {
std::string_view(p.name); // has name field convertible to string_view
p.age; // has age field
requires std::is_integral_v<decltype(p.age)>; // age is integral
};
/**
* Concept to validate that a type is an array of person objects
*/
template <typename T>
concept array_of_person = requires(T arr) {
arr.size(); // has size method
arr[0]; // can access elements with []
requires person<decltype(arr[0])>; // elements satisfy person concept
};
```
And then a simple static assert with `decltype` is sufficient to check that the expectation is met:
```cpp
constexpr auto config = R"(
[
{ "name": "Alice", "age": 30 },
{ "name": "Bob", "age": 25 },
{ "name": "Charlie", "age": 35 }
]
)"_json;
// Validate that the array satisfies the array_of_person concept
static_assert(array_of_person<decltype(config)>);
```
## Loading from disk
In practice, you may have a JSON file, say `json_data` that you want to parse
at compile time. You may do so as follows.
```c++
constexpr const char json_data[] = {
#embed "test.json"
, 0
};
constexpr auto json = simdjson::compile_time::parse_json<json_data>();
```
## Limitations (compile-time errors)
We have a few limitations which trigger compile-time errors if violated.
- Only JSON objects and arrays are supported at the top level (no primitives).
We will lift this limitation in the future.
- Strings are represented using the `const char*` in UTF-8, but they must not
contain embedded nulls. We would prefer to represent them as std::string or
std::string_view, and hope to do so in the future.
- Heterogeneous arrays are not supported yet. E.g., you need to have arrays of
all integers, or all strings, all floats, all compatible objects, etc.
For example, the following is accepted:
```json
[
{ "name": "Alice", "age": 30 },
{ "name": "Bob", "age": 25 },
{ "name": "Charlie", "age": 35 }
]
```
but the following is not:
```json
[
{ "name": "Alice", "age": 30 },
"Just a string",
42,
{ "name": "Charlie", "age": 35 }
]
```
We may support heterogeneous arrays in the future with std::variant types.
- We parse the first JSON document encountered in the string. Trailing
characters are ignored. Thus if your JSON begins with {"a":1}, everything
after the closing brace is ignored. This limitation will be lifted in the future,
reporting an error.
These limitations are safe in the sense that they result in compile-time errors.
Thus you will not get truncated strings or imprecise floats silently.
Although we are committed to maintaining the functionality in the long run, the
`compile_time::parse_json` function is subject to change.
+445
View File
@@ -0,0 +1,445 @@
# Compile-Time JSONPath and JSON Pointer Accessors
**Note:** This feature requires C++26 Static Reflection support (P2996) and is currently only available with experimental compilers. You must enable it with `-DSIMDJSON_STATIC_REFLECTION=ON` when building.
## Overview
simdjson provides compile-time JSONPath and JSON Pointer accessors that validate paths against struct definitions at compile time and generate optimized accessor code with zero runtime overhead. This combines the safety of compile-time type checking with the performance of pre-parsed, pre-validated access paths.
## Requirements
- C++26 compiler with Static Reflection support (P2996)
- Experimental compiler flags:
- Clang with P2996 support: `-std=c++26 -freflection -fexpansion-statements`
- Build configuration: `-DSIMDJSON_STATIC_REFLECTION=ON`
## How It Works
**Compile Time:**
1. Path string is parsed and converted to access steps
2. Path is validated against struct definition using reflection
3. Field types are checked and verified
4. Optimized accessor code is generated
**Runtime:**
- Direct navigation with no parsing
- No validation overhead
- No string comparisons for path components
- Type-safe extraction
## Two Usage Modes
### Mode 1: With Type Validation (Recommended)
When you provide a struct type, the compiler validates the entire path at compile time:
```cpp
struct User {
std::string name;
int age;
std::vector<std::string> emails;
};
// R"( ... )" is a C++ raw string literal.
const padded_string json = R"({
"name": "Alice",
"age": 30,
"emails": ["alice@example.com", "alice@work.com"]
})"_padded;
ondemand::parser parser;
auto doc = parser.iterate(json);
// Compile-time validation: checks that User has "name" field of type std::string
std::string name;
auto result = ondemand::json_path::at_path_compiled<User, ".name">(doc);
result.get(name); // name = "Alice"
// Compile-time validation: checks that "emails" is array-like with string elements
std::string email;
result = ondemand::json_path::at_path_compiled<User, ".emails[0]">(doc);
result.get(email); // email = "alice@example.com"
```
**Benefits:**
- **Compile-time errors** if path doesn't exist in struct
- **Type safety** - verifies field types match expected types
- **Refactoring protection** - renaming struct fields causes compile errors
**What gets validated:**
- Field existence
- Field types
- Array/container access validity
- Nested struct navigation
### Mode 2: Without Validation
When you omit the struct type, the path is parsed at compile time but not validated:
```cpp
const padded_string json = R"({
"name": "Alice",
"age": 30,
"address": {"city": "Boston"}
})"_padded;
ondemand::parser parser;
auto doc = parser.iterate(json);
// No compile-time validation - path is only parsed
std::string name;
auto result = ondemand::json_path::at_path_compiled<".name">(doc);
result.get(name); // name = "Alice"
std::string_view city;
result = ondemand::json_path::at_path_compiled<".address.city">(doc);
result.get(city); // city = "Boston"
```
**Benefits:**
- Works with dynamic/unknown JSON structures
- Still benefits from compile-time path parsing
- No runtime string parsing overhead
**Use when:**
- JSON structure is not known at compile time
- Working with varied JSON schemas
- Prototyping or exploratory parsing
## JSONPath Syntax
JSONPath uses dot notation and bracket notation for field access:
### Supported Syntax
| Syntax | Description | Example |
|--------|-------------|---------|
| `.field` | Dot notation for field access | `.name`, `.address.city` |
| `["field"]` | Bracket notation with quotes | `["name"]`, `["address"]["city"]` |
| `[index]` | Array index access | `[0]`, `[1]` |
| Mixed | Combination of notations | `.emails[0]`, `["users"][0].name` |
| `$` prefix | Optional root indicator | `$.name`, `$["name"]` |
### Examples
```cpp
struct Address {
std::string city;
int zip;
};
struct Person {
std::string name;
int age;
Address address;
std::vector<std::string> emails;
};
// Dot notation
at_path_compiled<Person, ".name">(doc)
at_path_compiled<Person, ".address.city">(doc)
// Bracket notation
at_path_compiled<Person, "[\"name\"]">(doc)
at_path_compiled<Person, "[\"address\"][\"city\"]">(doc)
// Array access
at_path_compiled<Person, ".emails[0]">(doc)
at_path_compiled<Person, ".emails[1]">(doc)
// Mixed notation
at_path_compiled<Person, ".address[\"zip\"]">(doc)
at_path_compiled<Person, "[\"emails\"][0]">(doc)
// With root indicator
at_path_compiled<Person, "$.name">(doc)
at_path_compiled<Person, "$.address.city">(doc)
```
## JSON Pointer Syntax
JSON Pointer (RFC 6901) uses slash-separated paths:
### Supported Syntax
| Syntax | Description | Example |
|--------|-------------|---------|
| `/field` | Field access | `/name`, `/address/city` |
| `/index` | Array index | `/0`, `/1` |
| `~0` | Escaped `~` | `/field~0name` → field~name |
| `~1` | Escaped `/` | `/field~1name` → field/name |
### Examples
```cpp
struct Car {
std::string make;
std::string model;
int64_t year;
std::vector<double> tire_pressure;
};
// Field access
at_pointer_compiled<Car, "/make">(doc)
at_pointer_compiled<Car, "/model">(doc)
// Array access
at_pointer_compiled<Car, "/tire_pressure/0">(doc)
at_pointer_compiled<Car, "/tire_pressure/1">(doc)
// Root pointer (returns whole document)
at_pointer_compiled<Car, "">(doc)
at_pointer_compiled<Car, "/">(doc)
```
## API Reference
### JSONPath Functions
```cpp
// With type validation
template<typename T, constevalutil::fixed_string Path, typename DocOrValue>
simdjson_result<value> at_path_compiled(DocOrValue& doc_or_val);
// Without validation
template<constevalutil::fixed_string Path, typename DocOrValue>
simdjson_result<value> at_path_compiled(DocOrValue& doc_or_val);
```
### JSON Pointer Functions
```cpp
// With type validation
template<typename T, constevalutil::fixed_string Pointer, typename DocOrValue>
simdjson_result<value> at_pointer_compiled(DocOrValue& doc_or_val);
// Without validation
template<constevalutil::fixed_string Pointer, typename DocOrValue>
simdjson_result<value> at_pointer_compiled(DocOrValue& doc_or_val);
```
### Direct Field Extraction
Extract values directly into variables with compile-time type checking:
```cpp
// JSONPath
template<typename T, constevalutil::fixed_string Path>
struct path_accessor {
template<typename DocOrValue, typename FieldType>
static error_code extract_field(DocOrValue& doc_or_val, FieldType& target);
};
// JSON Pointer
template<typename T, constevalutil::fixed_string Pointer>
struct pointer_accessor {
template<typename DocOrValue, typename FieldType>
static error_code extract_field(DocOrValue& doc_or_val, FieldType& target);
};
```
**Example:**
```cpp
struct User {
std::string name;
int age;
};
ondemand::parser parser;
auto doc = parser.iterate(json);
// Extract directly into variable
std::string name;
ondemand::json_path::path_accessor<User, ".name">::extract_field(doc, name);
int age;
ondemand::json_path::pointer_accessor<User, "/age">::extract_field(doc, age);
```
The compiler verifies that the target variable type matches the field type at the path.
## Complete Examples
### Example 1: Validated Access
```cpp
#include "simdjson.h"
using namespace simdjson;
struct Car {
std::string make;
std::string model;
int64_t year;
std::vector<double> tire_pressure;
};
int main() {
const padded_string json = R"({
"make": "Toyota",
"model": "Camry",
"year": 2018,
"tire_pressure": [40.1, 39.9, 37.7, 40.4]
})"_padded;
ondemand::parser parser;
auto doc = parser.iterate(json);
// Type-validated access
std::string make;
auto result = ondemand::json_path::at_path_compiled<Car, ".make">(doc);
result.get(make); // make = "Toyota"
// Array access with validation
double pressure;
result = ondemand::json_path::at_path_compiled<Car, ".tire_pressure[1]">(doc);
result.get(pressure); // pressure = 39.9
return 0;
}
```
### Example 2: Non-Validated Access
```cpp
#include "simdjson.h"
using namespace simdjson;
int main() {
const padded_string json = R"({
"user": {
"name": "Alice",
"preferences": {
"theme": "dark",
"notifications": true
}
}
})"_padded;
ondemand::parser parser;
auto doc = parser.iterate(json);
// No validation - works with any JSON structure
std::string_view theme;
auto result = ondemand::json_path::at_path_compiled<".user.preferences.theme">(doc);
result.get(theme); // theme = "dark"
bool notifications;
result = ondemand::json_path::at_path_compiled<".user.preferences.notifications">(doc);
result.get(notifications); // notifications = true
return 0;
}
```
### Example 3: Direct Extraction
```cpp
#include "simdjson.h"
using namespace simdjson;
struct Person {
std::string name;
int age;
std::vector<std::string> emails;
};
int main() {
const padded_string json = R"({
"name": "Bob",
"age": 25,
"emails": ["bob@example.com", "bob@work.com"]
})"_padded;
ondemand::parser parser;
auto doc = parser.iterate(json);
// Extract with type validation
std::string name;
ondemand::json_path::path_accessor<Person, ".name">::extract_field(doc, name);
// name = "Bob"
int age;
ondemand::json_path::pointer_accessor<Person, "/age">::extract_field(doc, age);
// age = 25
std::string email;
ondemand::json_path::path_accessor<Person, ".emails[0]">::extract_field(doc, email);
// email = "bob@example.com"
return 0;
}
```
## Error Handling
Compile-time errors occur when:
- Path doesn't exist in struct: `static_assert` failure
- Field type mismatch: `static_assert` failure
- Invalid array access on non-array field: `static_assert` failure
Runtime errors occur when:
- JSON structure doesn't match expected structure
- Array index out of bounds
- Type conversion failures
```cpp
struct User {
std::string name;
int age;
};
// Compile-time error: no "email" field in User
// auto result = ondemand::json_path::at_path_compiled<User, ".email">(doc);
// Compile-time error: age is not an array
// auto result = ondemand::json_path::at_path_compiled<User, ".age[0]">(doc);
// Runtime error if JSON doesn't have "name" field
auto result = ondemand::json_path::at_path_compiled<User, ".name">(doc);
std::string name;
if (result.get(name) != SUCCESS) {
// Handle error
}
```
## Performance
Compile-time accessors provide:
- **Zero path parsing overhead** - paths parsed at compile time
- **Zero validation overhead** - validation done at compile time
- **Direct field access** - no runtime path traversal
- **Type-safe extraction** - no dynamic type checking
Compared to runtime `at_path()` and `at_pointer()`:
- Eliminates runtime path string parsing
- Eliminates runtime path validation
- Generates optimal code path directly
## Limitations
- Requires C++26 compiler with P2996 support (experimental)
- Paths must be compile-time constants (string literals)
- Cannot use runtime-computed paths
- Limited to struct types that support reflection
- Array indices must be compile-time constants in the path
## When to Use
**Use compile-time accessors when:**
- You have well-defined struct types
- JSON structure is known at compile time
- You want maximum type safety
- Performance is critical
**Use runtime `at_path()`/`at_pointer()` when:**
- JSON structure varies or is unknown
- Paths are computed at runtime
- Working with C++20 or earlier
- Flexibility is more important than compile-time checks
## See Also
- [JSON Pointer](basics.md#json-pointer) - Runtime JSON Pointer support
- [JSONPath](basics.md#jsonpath) - Runtime JSONPath support
- [Static Reflection for Deserialization](basics.md#3-using-static-reflection-c26) - Using reflection for full struct deserialization
+52 -32
View File
@@ -41,7 +41,7 @@ The Basics: Loading and Parsing JSON Documents using the DOM front-end
The simdjson library offers a simple DOM tree API, which you can access by creating a
`dom::parser` and calling the `load()` method:
```c++
```cpp
dom::parser parser;
dom::element doc = parser.load(filename); // load and parse a file
```
@@ -49,21 +49,39 @@ dom::element doc = parser.load(filename); // load and parse a file
Or by creating a padded string (for efficiency reasons, simdjson requires a string with
SIMDJSON_PADDING bytes at the end) and calling `parse()`:
```c++
```cpp
dom::parser parser;
dom::element doc = parser.parse("[1,2,3]"_padded); // parse a string, the _padded suffix creates a simdjson::padded_string instance
```
You can also load a `padded_string` from a file.
```cpp
auto json = padded_string::load("twitter.json"); // load JSON file 'twitter.json'.
dom::element doc = parser.parse(json);
```
[You can similarly fetch a file from a URL to a padded string](https://github.com/simdjson/curltostring) using our `simdjson::padded_string_builder`.
(Windows users compiling with C++17 or better may use `wchar_t` strings to support non-ASCII
filenames: `padded_string::load(L"twitter.json")`.)
(Windows users compiling with C++17 or better may use `wchar_t` strings to support non-ASCII
filenames: `padded_string::load(L"twitter.json")`.)
You can copy your data directly on a `simdjson::padded_string` as follows:
```c++
```cpp
const char * data = "my data"; // 7 bytes
simdjson::padded_string my_padded_data(data, 7); // copies to a padded buffer
```
Or as follows...
```c++
```cpp
std::string data = "my data";
simdjson::padded_string my_padded_data(data); // copies to a padded buffer
```
@@ -83,7 +101,7 @@ container-overflow checks, you may encounter sanitizer warnings.
You can safely ignore these warnings. Or you can call `simdjson::pad(std::string&)` to pad the
string with `SIMDJSON_PADDING` spaces: this function returns a `simdjson::padding_string_view` which can be be passed to the parser's iterator function:
```c++
```cpp
std::string json = "[1]";
dom::element doc = parser.parse(simdjson::pad(json));
```
@@ -117,8 +135,9 @@ Once you have an element, you can navigate it with idiomatic C++ iterators, oper
dom::object and dom::array. An exception (`simdjson::simdjson_error`) is thrown if the cast is not possible.
* **Extracting Values (without exceptions):** You can use a variant usage of `get()` with error codes to avoid exceptions. You first declare the variable of the appropriate type (`double`, `uint64_t`, `int64_t`, `bool`, `std::string_view`,
`dom::object` and `dom::array`) and pass it by reference to `get()` which gives you back an error code: e.g.,
```c++
```cpp
simdjson::error_code error;
// _padded returns an simdjson::padded_string instance
simdjson::padded_string numberstring = "1.2"_padded; // our JSON input ("1.2")
simdjson::dom::parser parser;
double value; // variable where we store the value to be parsed
@@ -152,7 +171,8 @@ Once you have an element, you can navigate it with idiomatic C++ iterators, oper
The following code illustrates all of the above:
```c++
```cpp
// R"( ... )" is a C++ raw string literal.
auto cars_json = R"( [
{ "make": "Toyota", "model": "Camry", "year": 2018, "tire_pressure": [ 40.1, 39.9, 37.7, 40.4 ] },
{ "make": "Kia", "model": "Soul", "year": 2012, "tire_pressure": [ 30.1, 31.0, 28.6, 28.7 ] },
@@ -185,7 +205,7 @@ for (dom::object car : parser.parse(cars_json)) {
Here is a different example illustrating the same ideas:
```C++
```cpp
auto abstract_json = R"( [
{ "12345" : {"a":12.34, "b":56.78, "c": 9998877} },
{ "12545" : {"a":11.44, "b":12.78, "c": 11111111} }
@@ -207,7 +227,7 @@ for (dom::object obj : parser.parse(abstract_json)) {
And another one:
```C++
```cpp
auto abstract_json = R"(
{ "str" : { "123" : {"abc" : 3.14 } } } )"_padded;
dom::parser parser;
@@ -221,7 +241,7 @@ C++17 Support
While the simdjson library can be used in any project using C++ 11 and above, field iteration has special support C++ 17's destructuring syntax. For example:
```c++
```cpp
padded_string json = R"( { "foo": 1, "bar": 2 } )"_padded;
dom::parser parser;
dom::object object; // invalid until the get() succeeds
@@ -234,7 +254,7 @@ for (auto [key, value] : object) {
For comparison, here is the C++ 11 version of the same code:
```c++
```cpp
// C++ 11 version for comparison
padded_string json = R"( { "foo": 1, "bar": 2 } )"_padded;
dom::parser parser;
@@ -251,7 +271,7 @@ C++20 Support
simdjson library also supports some C++20 feature including `std::ranges`:
```c++
```cpp
auto cars_json = R"( [
{ "make": "Toyota", "model": "Camry", "year": 2018, "tire_pressure": [ 40.1, 39.9, 37.7, 40.4 ] },
{ "make": "Kia", "model": "Soul", "year": 2012, "tire_pressure": [ 30.1, 31.0, 28.6, 28.7 ] },
@@ -270,7 +290,7 @@ JSON Pointer
The simdjson library also supports [JSON pointer](https://tools.ietf.org/html/rfc6901) through the
`at_pointer()` method, letting you reach further down into the document in a single call:
```c++
```cpp
auto cars_json = R"( [
{ "make": "Toyota", "model": "Camry", "year": 2018, "tire_pressure": [ 40.1, 39.9, 37.7, 40.4 ] },
{ "make": "Kia", "model": "Soul", "year": 2012, "tire_pressure": [ 30.1, 31.0, 28.6, 28.7 ] },
@@ -291,7 +311,7 @@ You can apply a JSON Pointer expression to any node and the path gets interprete
Consider the following example:
```c++
```cpp
auto cars_json = R"( [
{ "make": "Toyota", "model": "Camry", "year": 2018, "tire_pressure": [ 40.1, 39.9, 37.7, 40.4 ] },
{ "make": "Kia", "model": "Soul", "year": 2012, "tire_pressure": [ 30.1, 31.0, 28.6, 28.7 ] },
@@ -313,11 +333,11 @@ JSONPath
------------
The simdjson library supports a subset of [JSONPath](https://datatracker.ietf.org/doc/html/draft-normington-jsonpath-00) through the `at_path()` method, allowing you to reach further into the document in a single call. The subset of JSONPath that is implemented is the subset that is trivially convertible into the JSON Pointer format, using `.` to access a field and `[]` to access a specific index.
The simdjson library supports a subset of [JSONPath](https://www.rfc-editor.org/rfc/rfc9535) (RFC 9535) through the `at_path()` method, allowing you to reach further into the document in a single call. The subset of JSONPath that is implemented is the subset that is trivially convertible into the JSON Pointer format, using `.` to access a field and `[]` to access a specific index.
Consider the following example:
```c++
```cpp
auto cars_json = R"( [
{ "make": "Toyota", "model": "Camry", "year": 2018, "tire_pressure": [ 40.1, 39.9, 37.7, 40.4 ] },
{ "make": "Kia", "model": "Soul", "year": 2012, "tire_pressure": [ 30.1, 31.0, 28.6, 28.7 ] },
@@ -336,7 +356,7 @@ cout << p << endl; // Prints 39.9
We also support the `$` prefix. When you start a JSONPath expression with $, you are indicating that the path starts from the root of the JSON document. E.g.,
```c++
```cpp
auto json = R"( { "c" :{ "foo": { "a": [ 10, 20, 30 ] }}, "d": { "foo2": { "a": [ 10, 20, 30 ] }} , "e": 120 })"_padded;
dom::parser parser;
dom::element doc;
@@ -428,7 +448,7 @@ Error Handling
All simdjson APIs that can fail return `simdjson_result<T>`, which is a &lt;value, error_code&gt;
pair. You can retrieve the value with .get(), like so:
```c++
```cpp
dom::element doc;
auto error = parser.parse(json).get(doc);
if (error) { cerr << error << endl; exit(1); }
@@ -462,7 +482,7 @@ We can write a "quick start" example where we attempt to parse the following JSO
Our program loads the file, selects value corresponding to key "search_metadata" which expected to be an object, and then
it selects the key "count" within that object.
```C++
```cpp
#include <iostream>
#include "simdjson.h"
@@ -490,7 +510,7 @@ triggering exceptions. To do this, we use `["statuses"].at(0)["id"]`. We break t
Observe how we use the `at` method when querying an index into an array, and not the bracket operator.
```C++
```cpp
#include <iostream>
#include "simdjson.h"
@@ -514,7 +534,7 @@ over the content of an array.
This is how the example in "Using the Parsed JSON" could be written using only error code checking:
```c++
```cpp
auto cars_json = R"( [
{ "make": "Toyota", "model": "Camry", "year": 2018, "tire_pressure": [ 40.1, 39.9, 37.7, 40.4 ] },
{ "make": "Kia", "model": "Soul", "year": 2012, "tire_pressure": [ 30.1, 31.0, 28.6, 28.7 ] },
@@ -561,7 +581,7 @@ for (dom::element car_element : cars) {
Here is another example:
```C++
```cpp
auto abstract_json = R"( [
{ "12345" : {"a":12.34, "b":56.78, "c": 9998877} },
{ "12545" : {"a":11.44, "b":12.78, "c": 11111111} }
@@ -594,7 +614,7 @@ for (dom::element elem : array) {
And another one:
```C++
```cpp
auto abstract_json = R"(
{ "str" : { "123" : {"abc" : 3.14 } } } )"_padded;
dom::parser parser;
@@ -608,7 +628,7 @@ Notice how we can string several operations (`parser.parse(abstract_json)["str"]
The next two functions will take as input a JSON document containing an array with a single element, either a string or a number. They return true upon success.
```C++
```cpp
simdjson::dom::parser parser{};
bool parse_double(const char *j, double &d) {
@@ -640,7 +660,7 @@ target_compile_definitions(simdjson PUBLIC SIMDJSON_EXCEPTIONS=OFF)
Users more comfortable with an exception flow may choose to directly cast the `simdjson_result<T>` to the desired type:
```c++
```cpp
dom::element doc = parser.parse(json); // Throws an exception if there was an error!
```
@@ -650,7 +670,7 @@ program from continuing if there was an error.
If one is willing to trigger exceptions, it is possible to write simpler code:
```C++
```cpp
#include <iostream>
#include "simdjson.h"
@@ -671,7 +691,7 @@ inspect or walk over JSON elements. To do that, you can use iterators and the ty
example, here's a quick and dirty recursive function that verbosely prints the JSON document as JSON
(* ignoring nuances like trailing commas and escaping strings, for brevity's sake):
```c++
```cpp
void print_json(dom::element element) {
switch (element.type()) {
case dom::element_type::ARRAY:
@@ -727,7 +747,7 @@ and reuse it. The simdjson library will allocate and retain internal buffers bet
buffers hot in cache and keeping memory allocation and initialization to a minimum. In this manner,
you can parse terabytes of JSON data without doing any new allocation.
```c++
```cpp
dom::parser parser;
// This initializes buffers and a document big enough to handle this JSON.
@@ -770,7 +790,7 @@ without bound:
* You can set a *max capacity* when constructing a parser:
```c++
```cpp
dom::parser parser(1000*1000); // Never grow past documents > 1MB
for (web_request request : listen()) {
dom::element doc;
@@ -786,7 +806,7 @@ without bound:
* You can set a *fixed capacity* that never grows, as well, which can be excellent for
predictability and reliability, since simdjson will never call malloc after startup!
```c++
```cpp
dom::parser parser(0); // This parser will refuse to automatically grow capacity
auto error = parser.allocate(1000*1000); // This allocates enough capacity to handle documents <= 1MB
if (error) { cerr << error << endl; exit(1); }
@@ -817,7 +837,7 @@ When calling `parser.parse` on a pointer (e.g., `parser.parse(my_char_pointer, m
Some users may not be able use our `padded_string` class or to load the data directly from disk (`parser.load`). They may need to pass data pointers to the library. If these users wish to avoid temporary copies and corresponding temporary memory allocations, they may want to call `parser.parse` with the `realloc_if_needed` parameter set to false (e.g., `parser.parse(my_char_pointer, my_length_in_bytes, false)`). In such cases, they need to ensure that there are at least SIMDJSON_PADDING extra bytes at the end that can be safely accessed and read. They do not need to initialize the padded bytes to any value in particular. The following example is safe:
```C++
```cpp
const char *json = R"({"key":"value"})";
const size_t json_len = std::strlen(json);
std::unique_ptr<char[]> padded_json_copy{new char[json_len + SIMDJSON_PADDING]};
@@ -825,7 +845,7 @@ memcpy(padded_json_copy.get(), json, json_len);
memset(padded_json_copy.get() + json_len, 0, SIMDJSON_PADDING);
simdjson::dom::parser parser;
simdjson::dom::element element = parser.parse(padded_json_copy.get(), json_len, false);
````
```
Setting the `realloc_if_needed` parameter `false` in this manner may lead to better performance since copies are avoided, but it requires that the user takes more responsibilities: the simdjson library cannot verify that the input buffer was padded with SIMDJSON_PADDING extra bytes.
+6 -6
View File
@@ -55,7 +55,7 @@ Inspecting the Detected Implementation
You can check what implementation is running with `active_implementation`:
```c++
```cpp
cout << "simdjson v" << SIMDJSON_VERSION << endl;
cout << "Detected the best implementation for your machine: " << simdjson::get_active_implementation()->name();
cout << "(" << simdjson::get_active_implementation()->description() << ")" << endl;
@@ -68,7 +68,7 @@ Querying Available Implementations
You can list all available implementations, regardless of which one was selected:
```c++
```cpp
for (auto implementation : simdjson::get_available_implementations()) {
cout << implementation->name() << ": " << implementation->description() << endl;
}
@@ -76,7 +76,7 @@ for (auto implementation : simdjson::get_available_implementations()) {
And look them up by name:
```c++
```cpp
cout << simdjson::get_available_implementations()["fallback"]->description() << endl;
```
When an implementation is not available, the bracket call `simdjson::get_available_implementations()[name]`
@@ -93,7 +93,7 @@ Manually Selecting the Implementation
If you're trying to do performance tests or see how different implementations of simdjson run, you
can select the CPU architecture yourself:
```c++
```cpp
// Use the fallback implementation, even though my machine is fast enough for anything
simdjson::get_active_implementation() = simdjson::get_available_implementations()["fallback"];
```
@@ -102,7 +102,7 @@ You are responsible for ensuring that the requirements of the selected implement
Furthermore, you should check that the implementation is available before setting it to `simdjson::get_active_implementation()`
by comparing it with the null pointer.
```c++
```cpp
auto my_implementation = simdjson::get_available_implementations()["haswell"];
if (! my_implementation) { exit(1); }
if (! my_implementation->supported_by_runtime_system()) { exit(1); }
@@ -114,7 +114,7 @@ Checking that an Implementation can Run on your System
You should call `supported_by_runtime_system()` to compare the processor's features with the need of the implementation.
```c++
```cpp
for (auto implementation : simdjson::get_available_implementations()) {
if (implementation->supported_by_runtime_system()) {
cout << implementation->name() << ": " << implementation->description() << endl;
+12 -10
View File
@@ -8,7 +8,7 @@ library provides high-speed access to files or streams containing multiple small
{"text":"a"}
{"text":"b"}
{"text":"c"}
...
"..."
```
... you want to read the entries (individual JSON documents) as quickly and as conveniently as possible. Importantly, the input might span several gigabytes, but you want to use a small (fixed) amount of memory. Ideally, you'd also like the parallelize the processing (using more than one core) to speed up the process.
@@ -132,7 +132,7 @@ E.g., `[1,2]{"32":1}` is recognized as two documents.
Some official formats **(non-exhaustive list)**:
- [Newline-Delimited JSON (NDJSON)](https://github.com/ndjson/ndjson-spec/)
- [JSON lines (JSONL)](http://jsonlines.org/)
- [Record separator-delimited JSON (RFC 7464)](https://tools.ietf.org/html/rfc7464) <- Not supported by JsonStream!
- [Record separator-delimited JSON (RFC 7464)](https://tools.ietf.org/html/rfc7464) <- Not supported by simdjson!
- [More on Wikipedia...](https://en.wikipedia.org/wiki/JSON_streaming)
API
@@ -140,8 +140,10 @@ API
Example:
```c++
```cpp
// R"( ... )" is a C++ raw string literal.
auto json = R"({ "foo": 1 } { "foo": 2 } { "foo": 3 } )"_padded;
// _padded returns an simdjson::padded_string instance
ondemand::parser parser;
ondemand::document_stream docs = parser.iterate_many(json);
for (auto doc : docs) {
@@ -197,7 +199,7 @@ and `error()` to check if there were any error.
Let us illustrate the idea with code:
```C++
```cpp
auto json = R"([1,2,3] {"1":1,"2":3,"4":4} [1,2,3] )"_padded;
simdjson::ondemand::parser parser;
simdjson::ondemand::document_stream stream;
@@ -238,7 +240,7 @@ Some users may need to work with truncated streams. The simdjson may truncate do
Consider the following example where a truncated document (`{"key":"intentionally unclosed string `) containing 39 bytes has been left within the stream. In such cases, the first two whole documents are parsed and returned, and the `truncated_bytes()` method returns 39.
```C++
```cpp
auto json = R"([1,2,3] {"1":1,"2":3,"4":4} {"key":"intentionally unclosed string )"_padded;
simdjson::ondemand::parser parser;
simdjson::ondemand::document_stream stream;
@@ -267,7 +269,7 @@ is effectively ignored, as it is set to at least the document size.
Example:
```C++
```cpp
auto json = R"( 1, 2, 3, 4, "a", "b", "c", {"hello": "world"} , [1, 2, 3])"_padded;
ondemand::parser parser;
ondemand::document_stream doc_stream;
@@ -314,7 +316,7 @@ the simdjson library.
Consider a custom class `Car`:
```C++
```cpp
struct Car {
std::string make;
std::string model;
@@ -328,7 +330,7 @@ You may support deserializing directly from a JSON value or document to your own
by defining a single `tag_invoke` function:
```C++
```cpp
namespace simdjson {
// This tag_invoke MUST be inside simdjson namespace
template <typename simdjson_value>
@@ -370,7 +372,7 @@ tag_invoke functions.
Given a stream of JSON documents, you can add them to a data structure
such as a `std::vector<Car>` like so if you support exceptions:
```C++
```cpp
padded_string json =
R"( { "make": "Toyota", "model": "Camry", "year": 2018,
"tire_pressure": [ 40.1, 39.9 ] }
@@ -391,7 +393,7 @@ such as a `std::vector<Car>` like so if you support exceptions:
Otherwise you may use this longer version for explicit handling of errors:
```C++
```cpp
std::vector<Car> cars;
for(auto doc : stream) {
Car c;
+21 -19
View File
@@ -23,7 +23,7 @@ applications with a computation efficiency that is difficult to surpass.
A code example illustrates our API from a programmer's point of view:
```c++
```cpp
ondemand::parser parser;
auto doc = parser.iterate(json);
for (auto tweet : doc["statuses"]) {
@@ -109,7 +109,7 @@ The DOM approach was the only way to parse JSON documents up to version 0.6 of t
Our DOM API looks similar to our On-Demand example, except
it calls `parse` instead of `iterate`:
```c++
```cpp
dom::parser parser;
auto doc = parser.parse(json);
for (auto tweet : doc["statuses"]) {
@@ -157,7 +157,7 @@ examples. To make it short enough to use as an example at all, it has heavily re
a part of the problem (does not get user.screen_name), it has bugs (it does not handle sub-objects
in a tweet at all), and it uses a theoretical, simple event-based API that minimizes ceremony.
```c++
```cpp
struct twitter_callbacks {
bool in_statuses;
bool in_tweet;
@@ -284,14 +284,14 @@ To help visualize the algorithm, we'll walk through the example C++ given at the
This declaration does not allocate any memory; that will happen in the next step.
```c++
```cpp
ondemand::parser parser;
```
2. We then start iterating the JSON document by allocating internal parser buffers, preprocessing
the JSON, and initializing the iterator.
```c++
```cpp
auto doc = parser.iterate(json);
```
@@ -337,14 +337,14 @@ To help visualize the algorithm, we'll walk through the example C++ given at the
3. We iterate over the "statuses" field using a typical C++ iterator, reading past the initial
`{ "statuses": [ {`.
```c++
```cpp
for (ondemand::object tweet : doc["statuses"]) {
```
This shorthand does a lot, and it is helpful to see what it expands to.
Comments in front of each one explain what's going on:
```c++
```cpp
// Validate that the top-level value is an object: check for {. Increase depth to 2 (root > field).
ondemand::object top = doc.get_object();
@@ -396,7 +396,7 @@ To help visualize the algorithm, we'll walk through the example C++ given at the
4. We get the `"text"` field as a string.
```c++
```cpp
std::string_view text = tweet["text"];
```
@@ -435,7 +435,7 @@ To help visualize the algorithm, we'll walk through the example C++ given at the
4. We get the `"screen_name"` from the `"user"` object.
```c++
```cpp
ondemand::object user = tweet["user"];
screen_name = user["screen_name"];
```
@@ -469,7 +469,7 @@ To help visualize the algorithm, we'll walk through the example C++ given at the
5. We get `"retweet_count"` as an unsigned integer.
```c++
```cpp
uint64_t retweets = tweet["retweet_count"];
```
@@ -513,7 +513,7 @@ To help visualize the algorithm, we'll walk through the example C++ given at the
6. We loop to the next tweet.
```c++
```cpp
for (ondemand::object tweet : doc["statuses"]) {
...
}
@@ -521,7 +521,7 @@ To help visualize the algorithm, we'll walk through the example C++ given at the
The relevant parts of the loop are:
```c++
```cpp
while (iter != statuses.end()) {
ondemand::object tweet = *iter;
...
@@ -545,7 +545,7 @@ To help visualize the algorithm, we'll walk through the example C++ given at the
"statuses": [
{ "id": 1, "text": "first!", "user": { "screen_name": "lemire", "name": "Daniel" }, "retweet_count": 40 },
{ "id": 2, "text": "second!", "user": { "screen_name": "jkeiser2", "name": "John" }, "retweet_count": 3 }
^ (depth 3 - root > statuses > tweet)
^ (depth 4 - root > statuses > tweet > field)
],
"search_metadata": { "count": 2 }
}
@@ -566,7 +566,7 @@ To help visualize the algorithm, we'll walk through the example C++ given at the
8. The loop ends. Recall the relevant parts of the statuses loop:
```c++
```cpp
while (iter != statuses.end()) {
ondemand::object tweet = *iter;
...
@@ -610,7 +610,7 @@ When the user requests strings, we unescape them to a single string buffer much
so that users enjoy the same string performance as the core simdjson. We do not write the length to the
string buffer, however; that is stored in the `string_view` instance we return to the user.
```C++
```cpp
ondemand::parser parser;
auto doc = parser.iterate(json);
std::set<std::string_view> default_users;
@@ -645,7 +645,7 @@ from the `unescaped_key()` method has a lifecycle tied to the `parser` instance:
is destroyed or reused with another document, the `std::string_view` instance becomes invalid.
```C++
```cpp
auto doc = parser.iterate(json);
for(auto field : doc.get_object()) {
std::string_view keyv = field.unescaped_key();
@@ -670,9 +670,11 @@ in production systems:
Some care is needed when using the On-Demand API in scenarios where you need to access several sibling arrays or objects because
only one object or array can be active at any one time. Let us consider the following example:
```C++
```cpp
ondemand::parser parser;
// R"( ... )" is a C++ raw string literal.
const padded_string json = R"({ "parent": {"child1": {"name": "John"} , "child2": {"name": "Daniel"}} })"_padded;
// _padded returns an simdjson padded_string instance
auto doc = parser.iterate(json);
ondemand::object parent = doc["parent"];
// parent owns the focus
@@ -688,7 +690,7 @@ in production systems:
A correct usage is given by the following example:
```C++
```cpp
ondemand::parser parser;
const padded_string json = R"({ "parent": {"child1": {"name": "John"} , "child2": {"name": "Daniel"}} })"_padded;
auto doc = parser.iterate(json);
@@ -754,7 +756,7 @@ Some users wish to run at the best possible speed. Under recent Intel and AMD pr
Given that the On-Demand API offer limited runtime dispatching, it matters that your code is compiled against a specific CPU target. You should verify that the code is compiled against the target you expect. Thankfully, the simdjson library will tell you exactly what it detects as an implementation: `icelake` (AVX512 x64 processors), `haswell` (AVX2 x64 processors), `westmere` (SSE4 x64 processors), `arm64` (64-bit ARM), `ppc64` (64-bit POWER), `lasx` (LoongArch), `lsx` (LoongArch), `fallback` (others). Under x64 processors, many programmers will want to target `haswell` whereas under ARM, most programmers will want to target `arm64` (and it should do so automatically). The `fallback` is probably only good for testing purposes, not for deployment.
```C++
```cpp
std::cout << simdjson::builtin_implementation()->name() << std::endl;
```
+3 -3
View File
@@ -132,7 +132,7 @@ Whitespace Characters:
Some official formats **(non-exhaustive list)**:
- [Newline-Delimited JSON (NDJSON)](https://github.com/ndjson/ndjson-spec)
- [JSON lines (JSONL)](http://jsonlines.org/)
- [Record separator-delimited JSON (RFC 7464)](https://tools.ietf.org/html/rfc7464) <- Not supported by JsonStream!
- [Record separator-delimited JSON (RFC 7464)](https://tools.ietf.org/html/rfc7464) <- Not supported by simdjson!
- [More on Wikipedia...](https://en.wikipedia.org/wiki/JSON_streaming)
API
@@ -184,7 +184,7 @@ You may also call the `source()` method to get a `std::string_view` instance on
Let us illustrate the idea with code:
```C++
```cpp
auto json = R"([1,2,3] {"1":1,"2":3,"4":4} [1,2,3] )"_padded;
simdjson::dom::parser parser;
simdjson::dom::document_stream stream;
@@ -225,7 +225,7 @@ Some users may need to work with truncated streams. The simdjson may truncate do
Consider the following example where a truncated document (`{"key":"intentionally unclosed string `) containing 39 bytes has been left within the stream. In such cases, the first two whole documents are parsed and returned, and the `truncated_bytes()` method returns 39.
```C++
```cpp
auto json = R"([1,2,3] {"1":1,"2":3,"4":4} {"key":"intentionally unclosed string )"_padded;
simdjson::dom::parser parser;
simdjson::dom::document_stream stream;
+64 -6
View File
@@ -47,7 +47,7 @@ and reuse it. The simdjson library will allocate and retain internal buffers bet
buffers hot in cache and keeping memory allocation and initialization to a minimum. In this manner,
you can parse terabytes of JSON data without doing any new allocation.
```c++
```cpp
ondemand::parser parser;
// This initializes buffers big enough to handle this JSON.
@@ -71,14 +71,14 @@ Reusing string buffers
We recommend against creating many `std::string` or `simdjson::padded_string` instances to store the JSON content in your application. [Creating many non-trivial objects is convenient but often surprisingly slow](https://lemire.me/blog/2020/08/08/performance-tip-constructing-many-non-trivial-objects-is-slow/). Instead, as much as possible, you should allocate (once or a few times) reusable memory buffers where you write your JSON content. If you have a buffer `json_str` (of type `char*`) allocated for `capacity` bytes and you store a JSON document spanning `length` bytes, you can pass it to simdjson as follows:
```c++
```cpp
auto doc = parser.iterate(padded_string_view(json_str, length, capacity));
```
or simply
```c++
```cpp
auto doc = parser.iterate(json_str, length, capacity);
```
@@ -89,7 +89,7 @@ Server Loops: Long-Running Processes and Memory Capacity
The On-Demand approach also automatically expands its memory capacity when larger documents are parsed. However, for longer processes where very large files are processed (such as server loops), this capacity is not resized down. On-Demand also lets you adjust the maximal capacity that the parser can process:
* You can set an upper bound (*max_capacity*) when construction the parser:
```C++
```cpp
ondemand::parser parser(1000*1000); // Never grows past documents > 1 MB
auto doc = parser.iterate(json);
for (web_request request : listen()) {
@@ -105,7 +105,7 @@ The On-Demand approach also automatically expands its memory capacity when large
The capacity will grow as the parser encounters larger documents up to 1 MB.
* You can also allocate a *fixed capacity* that will never grow:
```C++
```cpp
ondemand::parser parser(1000*1000);
parser.allocate(1000*1000) // Fix the capacity to 1 MB
auto doc = parser.iterate(json);
@@ -253,7 +253,7 @@ long page_size() {
// page boundary.
bool need_allocation(const char *buf, size_t len) {
return ((reinterpret_cast<uintptr_t>(buf + len - 1) % page_size())
+ simdjson::SIMDJSON_PADDING > static_cast<uintptr_t>(page_size()));
+ simdjson::SIMDJSON_PADDING >= static_cast<uintptr_t>(page_size()));
}
simdjson::padded_string_view
@@ -298,3 +298,61 @@ int main() {
return EXIT_SUCCESS;
}
```
Further, whenever you allocate N bytes, memory allocators tend to allocate more memory, without you necessarily knowing about it. Under linux, you can use the `malloc_usable_size` function to see how much memory was actually allocated.
Under an Apple plateform, you can `malloc_size`. The following program illustrates the usage.
```cpp
#include <iostream>
#include <cstddef>
#include <memory>
#include <cstdlib>
#ifdef __APPLE__
#include <malloc/malloc.h> // for malloc_size on macOS
#endif
#ifdef __linux__
#include <malloc.h> // for malloc_usable_size on Linux
#endif
size_t get_usable_size(void* ptr) {
#ifdef __linux__
return malloc_usable_size(ptr);
#elif defined(__APPLE__)
return malloc_size(ptr);
#else
return 0; // Unsupported platform
#endif
}
int main() {
std::cout << "Demonstrating allocation overhead and rounding with operator new\n\n";
#ifdef __linux__
std::cout << "Platform: Linux\n";
#elif defined(__APPLE__)
std::cout << "Platform: macOS (using malloc_size)\n";
#else
std::cout << "Platform: Other/unsupported (usable size will show 0)\n";
#endif
std::cout << "Requested size | Actual usable size\n";
std::cout << "---------------|-------------------\n";
size_t total_requested = 0;
size_t total_usable = 0;
for (size_t requested = 1; requested <= 4096; requested++) {
total_requested += requested;
std::unique_ptr<char[]> ptr(new char[requested]); // Allocate
size_t usable = get_usable_size(ptr.get()); // Get usable size
total_usable += usable;
std::cout << requested << "\t | " << usable << "\n";
}
std::cout << "---------------|-------------------\n";
std::cout << "Total requested: " << total_requested << " bytes\n";
std::cout << "Total usable: " << total_usable << " bytes\n";
std::cout << "Total overhead: " << (total_usable - total_requested) << " bytes\n";
std::cout << "Percentage overhead: "
<< ((total_usable - total_requested) * 100.0 / total_requested) << " %\n";
return EXIT_SUCCESS;
}
```
+1 -6
View File
@@ -8,16 +8,11 @@
* Minifies by first parsing, then minifying.
*/
extern "C" int LLVMFuzzerTestOneInput(const uint8_t *Data, size_t Size) {
auto begin = as_chars(Data);
auto end = begin + Size;
std::string str(begin, end);
simdjson::padded_string str(reinterpret_cast<const char *>(Data), Size);
simdjson::dom::parser parser;
simdjson::dom::element elem;
auto error = parser.parse(str).get(elem);
if (error) { return 0; }
std::string minified = simdjson::minify(elem);
(void)minified;
return 0;
+1 -1
View File
@@ -35,7 +35,7 @@ cmake .. \
-DSIMDJSON_DISABLE_DEPRECATED_API=On \
-DSIMDJSON_FUZZ_LDFLAGS=$LIB_FUZZING_ENGINE
cmake --build . --target all_fuzzers
cmake --build . --target all_fuzzers all_tests
cp fuzz/fuzz_* $OUT
Binary file not shown.

After

Width:  |  Height:  |  Size: 39 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 4.2 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 35 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 93 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 226 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 136 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 46 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 108 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 256 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 136 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 46 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 109 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 258 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 136 KiB

+6
View File
@@ -52,7 +52,13 @@
#include "simdjson/padded_string_view-inl.h"
#include "simdjson/dom.h"
#include "simdjson/builder.h"
#include "simdjson/ondemand.h"
#include "simdjson/convert.h"
#include "simdjson/convert-inl.h"
// Compile-time JSON parsing (C++26 P2996 reflection)
#include "simdjson/compile_time_json.h"
#include "simdjson/compile_time_json-inl.h"
#endif // SIMDJSON_H
+8
View File
@@ -0,0 +1,8 @@
#ifndef SIMDJSON_ARM64_BUILDER_H
#define SIMDJSON_ARM64_BUILDER_H
#include "simdjson/arm64/begin.h"
#include "simdjson/generic/builder/amalgamated.h"
#include "simdjson/arm64/end.h"
#endif // SIMDJSON_ARM64_BUILDER_H
+6 -3
View File
@@ -17,7 +17,8 @@ using namespace simd;
struct backslash_and_quote {
public:
static constexpr uint32_t BYTES_PROCESSED = 32;
simdjson_inline static backslash_and_quote copy_and_find(const uint8_t *src, uint8_t *dst);
// We only copy if dst is non-null.
simdjson_inline backslash_and_quote copy_and_find(const uint8_t *src, uint8_t *dst);
simdjson_inline bool has_quote_first() { return ((bs_bits - 1) & quote_bits) != 0; }
simdjson_inline bool has_backslash() { return bs_bits != 0; }
@@ -34,8 +35,10 @@ simdjson_inline backslash_and_quote backslash_and_quote::copy_and_find(const uin
static_assert(SIMDJSON_PADDING >= (BYTES_PROCESSED - 1), "backslash and quote finder must process fewer than SIMDJSON_PADDING bytes");
simd8<uint8_t> v0(src);
simd8<uint8_t> v1(src + sizeof(v0));
v0.store(dst);
v1.store(dst + sizeof(v0));
if(dst != nullptr) {
v0.store(dst);
v1.store(dst + sizeof(v0));
}
// Getting a 64-bit bitmask is much cheaper than multiple 16-bit bitmasks on ARM; therefore, we
// smash them together into a 64-byte mask and get the bitmask from there.
+14
View File
@@ -0,0 +1,14 @@
#ifndef SIMDJSON_BUILDER_H
#define SIMDJSON_BUILDER_H
#include "simdjson/builtin/builder.h"
namespace simdjson {
/**
* @copydoc simdjson::builtin::builder
*/
namespace builder = builtin::builder;
} // namespace simdjson
#endif // SIMDJSON_BUILDER_H
+2 -2
View File
@@ -20,10 +20,10 @@
#include "simdjson/ppc64.h"
#elif SIMDJSON_BUILTIN_IMPLEMENTATION_IS(westmere)
#include "simdjson/westmere.h"
#elif SIMDJSON_BUILTIN_IMPLEMENTATION_IS(lsx)
#include "simdjson/lsx.h"
#elif SIMDJSON_BUILTIN_IMPLEMENTATION_IS(lasx)
#include "simdjson/lasx.h"
#elif SIMDJSON_BUILTIN_IMPLEMENTATION_IS(lsx)
#include "simdjson/lsx.h"
#else
#error Unknown SIMDJSON_BUILTIN_IMPLEMENTATION
#endif
+40
View File
@@ -0,0 +1,40 @@
#ifndef SIMDJSON_BUILTIN_BUILDER_H
#define SIMDJSON_BUILTIN_BUILDER_H
#include "simdjson/builtin.h"
#include "simdjson/builtin/base.h"
#include "simdjson/generic/builder/dependencies.h"
#define SIMDJSON_CONDITIONAL_INCLUDE
#if SIMDJSON_BUILTIN_IMPLEMENTATION_IS(arm64)
#include "simdjson/arm64/builder.h"
#elif SIMDJSON_BUILTIN_IMPLEMENTATION_IS(fallback)
#include "simdjson/fallback/builder.h"
#elif SIMDJSON_BUILTIN_IMPLEMENTATION_IS(haswell)
#include "simdjson/haswell/builder.h"
#elif SIMDJSON_BUILTIN_IMPLEMENTATION_IS(icelake)
#include "simdjson/icelake/builder.h"
#elif SIMDJSON_BUILTIN_IMPLEMENTATION_IS(ppc64)
#include "simdjson/ppc64/builder.h"
#elif SIMDJSON_BUILTIN_IMPLEMENTATION_IS(westmere)
#include "simdjson/westmere/builder.h"
#elif SIMDJSON_BUILTIN_IMPLEMENTATION_IS(lsx)
#include "simdjson/lsx/builder.h"
#elif SIMDJSON_BUILTIN_IMPLEMENTATION_IS(lasx)
#include "simdjson/lasx/builder.h"
#else
#error Unknown SIMDJSON_BUILTIN_IMPLEMENTATION
#endif
#undef SIMDJSON_CONDITIONAL_INCLUDE
namespace simdjson {
/**
* @copydoc simdjson::SIMDJSON_BUILTIN_IMPLEMENTATION::builder
*/
namespace builder = SIMDJSON_BUILTIN_IMPLEMENTATION::builder;
} // namespace simdjson
#endif // SIMDJSON_BUILTIN_BUILDER_H
+3 -1
View File
@@ -289,7 +289,9 @@ namespace std {
// when the compiler is optimizing.
// We only set SIMDJSON_DEVELOPMENT_CHECKS if both __OPTIMIZE__
// and NDEBUG are not defined.
#if !defined(__OPTIMIZE__) && !defined(NDEBUG)
// We recognize _DEBUG as overriding __OPTIMIZE__ so that if both
// __OPTIMIZE__ and _DEBUG are defined, we still set SIMDJSON_DEVELOPMENT_CHECKS.
#if ((!defined(__OPTIMIZE__) || defined(_DEBUG)) && !defined(NDEBUG))
#define SIMDJSON_DEVELOPMENT_CHECKS 1
#endif // __OPTIMIZE__
#endif // _MSC_VER
File diff suppressed because it is too large Load Diff
+75
View File
@@ -0,0 +1,75 @@
/**
* @file compile_time_json.h
* @brief Compile-time JSON parsing using C++26 reflection with
* std::meta::substitute()
*/
#ifndef SIMDJSON_GENERIC_COMPILE_TIME_JSON_H
#define SIMDJSON_GENERIC_COMPILE_TIME_JSON_H
#if SIMDJSON_STATIC_REFLECTION
#include <algorithm>
#include <array>
#include <charconv>
#include <cstdint>
#include <expected>
#include <meta>
#include <string>
#include <string_view>
#include <vector>
namespace simdjson {
namespace compile_time {
/**
* @brief Compile-time JSON parser. This function parses the provided JSON
* string at compile time and returns a custom struct type representing the JSON
* object.
*
* We have a few limitations which trigger compile-time errors if violated:
* - Only JSON objects and arrays are supported at the top level (no primitives).
* We will lift this limitation in the future.
* - Strings are represented using the const char * in UTF-8, but they must not
* contain embedded nulls. We would prefer to represent them as std::string or
* std::string_view, and hope to do so in the future.
* - Heterogeneous arrays are not supported yet. E.g., you need to have arrays of
* all integers, or all strings, all floats, all compatible objects, etc.
* For example, the following is accepted:
* [
* { "name": "Alice", "age": 30 },
* { "name": "Bob", "age": 25 },
* { "name": "Charlie", "age": 35 }
* ]
* but the following is not:
* [
* { "name": "Alice", "age": 30 },
* "Just a string",
* 42,
* { "name": "Charlie", "age": 35 }
* ]
*
* We may support heterogeneous arrays in the future with std::variant types.
* - We parse the first JSON document encountered in the string. Trailing
* characters are ignored. Thus if your JSON begins with {"a":1}, everything
* after the closing } is ignored. This limitation will be lifted in the future,
* reporting an error.
*
* These limitations are safe in the sense that they result in compile-time errors.
* Thus you will not get truncated strings or imprecise floats silently.
*
* This function is subject to change in the future.
*/
template <constevalutil::fixed_string json_str> consteval auto parse_json();
} // namespace compile_time
} // namespace simdjson
template <simdjson::constevalutil::fixed_string str>
consteval auto operator ""_json() {
return simdjson::compile_time::parse_json<str>();
}
#endif // SIMDJSON_STATIC_REFLECTION
#endif // SIMDJSON_GENERIC_COMPILE_TIME_JSON_H
+5
View File
@@ -13,6 +13,11 @@
#endif
#endif
// C++ 26
#if !defined(SIMDJSON_CPLUSPLUS26) && (SIMDJSON_CPLUSPLUS >= 202402L) // update when the standard is finalized
#define SIMDJSON_CPLUSPLUS26 1
#endif
// C++ 23
#if !defined(SIMDJSON_CPLUSPLUS23) && (SIMDJSON_CPLUSPLUS >= 202302L)
#define SIMDJSON_CPLUSPLUS23 1
+54
View File
@@ -122,11 +122,65 @@ concept optional_type = requires(std::remove_cvref_t<T> obj) {
} -> std::convertible_to<typename std::remove_cvref_t<T>::value_type>;
};
{ static_cast<bool>(obj) } -> std::same_as<bool>; // convertible to bool
{ obj.reset() } noexcept -> std::same_as<void>;
};
// Types we serialize as JSON strings (not as containers)
template <typename T>
concept string_like =
std::is_same_v<std::remove_cvref_t<T>, std::string> ||
std::is_same_v<std::remove_cvref_t<T>, std::string_view> ||
std::is_same_v<std::remove_cvref_t<T>, const char*> ||
std::is_same_v<std::remove_cvref_t<T>, char*>;
// Concept that checks if a type is a container but not a string (because
// strings handling must be handled differently)
// Now uses iterator-based approach for broader container support
template <typename T>
concept container_but_not_string =
std::ranges::input_range<T> && !string_like<T> && !concepts::string_view_keyed_map<T>;
// Concept: Indexable container that is not a string or associative container
// Accepts: std::vector, std::array, std::deque (have operator[], value_type, not string_like)
// Rejects: std::string (string_like), std::list (no operator[]), std::map (has key_type)
template<typename Container>
concept indexable_container = requires {
typename Container::value_type;
requires !concepts::string_like<Container>;
requires !requires { typename Container::key_type; }; // Reject maps/sets
requires requires(Container& c, std::size_t i) {
{ c[i] } -> std::convertible_to<typename Container::value_type>;
};
};
// Variable template to use with std::meta::substitute
template<typename Container>
constexpr bool indexable_container_v = indexable_container<Container>;
} // namespace concepts
/**
* We use tag_invoke as our customization point mechanism.
*/
template <typename Tag, typename... Args>
concept tag_invocable = requires(Tag tag, Args... args) {
tag_invoke(std::forward<Tag>(tag), std::forward<Args>(args)...);
};
template <typename Tag, typename... Args>
concept nothrow_tag_invocable =
tag_invocable<Tag, Args...> && requires(Tag tag, Args... args) {
{
tag_invoke(std::forward<Tag>(tag), std::forward<Args>(args)...)
} noexcept;
};
} // namespace simdjson
#endif // SIMDJSON_SUPPORTS_CONCEPTS
#endif // SIMDJSON_CONCEPTS_H
+38 -2
View File
@@ -5,9 +5,9 @@
#include <string_view>
#include <array>
#if SIMDJSON_CONSTEVAL
namespace simdjson {
namespace constevalutil {
#if SIMDJSON_CONSTEVAL
constexpr static std::array<uint8_t, 256> json_quotable_character = {
1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
@@ -47,7 +47,43 @@ consteval std::string consteval_to_quoted_escaped(std::string_view input) {
out.push_back('"');
return out;
}
#endif // SIMDJSON_CONSTEVAL
#if SIMDJSON_SUPPORTS_CONCEPTS
template <size_t N>
struct fixed_string {
constexpr fixed_string(const char (&str)[N]) {
for (std::size_t i = 0; i < N; ++i) {
data[i] = str[i];
}
}
char data[N];
constexpr std::string_view view() const { return {data, N - 1}; }
constexpr size_t size() const { return N ; }
constexpr operator std::string_view() const { return view(); }
constexpr char operator[](std::size_t index) const { return data[index]; }
constexpr bool operator==(const fixed_string& other) const {
if (N != other.size()) {
return false;
}
for (std::size_t i = 0; i < N; ++i) {
if (data[i] != other.data[i]) {
return false;
}
}
return true;
}
};
template <std::size_t N>
fixed_string(const char (&)[N]) -> fixed_string<N>;
template <fixed_string str>
struct string_constant {
static constexpr std::string_view value = str.view();
};
#endif // SIMDJSON_SUPPORTS_CONCEPTS
} // namespace constevalutil
} // namespace simdjson
#endif // SIMDJSON_CONSTEVAL
#endif // SIMDJSON_CONSTEVALUTIL_H
+4 -2
View File
@@ -79,6 +79,7 @@ inline simdjson_result<ondemand::number> auto_parser<parser_type>::number() noex
return result<ondemand::number>();
}
#if SIMDJSON_EXCEPTIONS
template <typename parser_type>
template <typename T>
inline auto_parser<parser_type>::operator T() noexcept(false) {
@@ -87,6 +88,7 @@ inline auto_parser<parser_type>::operator T() noexcept(false) {
}
return m_doc.get<T>();
}
#endif // SIMDJSON_EXCEPTIONS
template <typename parser_type>
template <typename T>
@@ -109,7 +111,7 @@ inline T to_adaptor<T>::operator()(simdjson_result<ondemand::value> &val) const
template <typename T>
inline auto to_adaptor<T>::operator()(padded_string_view const str) const noexcept {
return auto_parser{str};
return auto_parser<ondemand::parser *>{str};
}
template <typename T>
@@ -119,7 +121,7 @@ inline auto to_adaptor<T>::operator()(ondemand::parser &parser, padded_string_vi
template <typename T>
inline auto to_adaptor<T>::operator()(std::string str) const noexcept {
return auto_parser{pad_with_reserve(str)};
return auto_parser<ondemand::parser *>{pad_with_reserve(str)};
}
template <typename T>
+5 -2
View File
@@ -54,10 +54,11 @@ public:
simdjson_warn_unused simdjson_inline simdjson_result<ondemand::object> object() noexcept;
simdjson_warn_unused simdjson_inline simdjson_result<ondemand::number> number() noexcept;
//template <typename T>
//simdjson_warn_unused simdjson_inline explicit(false) operator simdjson_result<T>() noexcept(is_nothrow_gettable<T>);
#if SIMDJSON_EXCEPTIONS
template <typename T>
simdjson_warn_unused simdjson_inline explicit(false) operator T() noexcept(false);
#endif // SIMDJSON_EXCEPTIONS
template <typename T>
simdjson_warn_unused simdjson_inline std::optional<T> optional() noexcept(is_nothrow_gettable<T>);
@@ -80,6 +81,8 @@ struct to_adaptor {
auto operator()(std::string str) const noexcept;
auto operator()(ondemand::parser &parser, std::string str) const noexcept;
};
// deduction guide
auto_parser(padded_string_view const str) -> auto_parser<ondemand::parser*>;
} // namespace internal
} // namespace convert
+2
View File
@@ -9,6 +9,7 @@
#include "simdjson/dom/object.h"
#include "simdjson/dom/parser.h"
#include "simdjson/dom/serialization.h"
#include "simdjson/dom/fractured_json.h"
// Inline functions
#include "simdjson/dom/array-inl.h"
@@ -19,5 +20,6 @@
#include "simdjson/dom/parser-inl.h"
#include "simdjson/internal/tape_ref-inl.h"
#include "simdjson/dom/serialization-inl.h"
#include "simdjson/dom/fractured_json-inl.h"
#endif // SIMDJSON_DOM_H
+2 -2
View File
@@ -111,7 +111,7 @@ public:
inline simdjson_result<element> at_pointer(std::string_view json_pointer) const noexcept;
/**
* Recursive function which processes the json path of each child element
* Recursive function which processes the JSON path of each child element
*/
inline void process_json_path_of_child_elements(std::vector<element>::iterator& current, std::vector<element>::iterator& end, const std::string_view& path_suffix, std::vector<element>& accumulator) const noexcept;
@@ -126,7 +126,7 @@ public:
* JSONPath queries that trivially convertible to JSON Pointer queries: key
* names and array indices.
*
* https://datatracker.ietf.org/doc/html/draft-normington-jsonpath-00
* https://www.rfc-editor.org/rfc/rfc9535 (RFC 9535)
*
* @return The value associated with the given JSONPath expression, or:
* - INVALID_JSON_POINTER if the JSONPath to JSON Pointer conversion fails
+1 -1
View File
@@ -74,7 +74,7 @@ public:
/**
* Construct an uninitialized document_stream.
*
* ```c++
* ```cpp
* document_stream docs;
* error = parser.parse_many(json).get(docs);
* ```
+1 -1
View File
@@ -408,7 +408,7 @@ public:
* JSONPath queries that trivially convertible to JSON Pointer queries: key
* names and array indices.
*
* https://datatracker.ietf.org/doc/html/draft-normington-jsonpath-00
* https://www.rfc-editor.org/rfc/rfc9535 (RFC 9535)
*
* @return The value associated with the given JSONPath expression, or:
* - INVALID_JSON_POINTER if the JSONPath to JSON Pointer conversion fails
File diff suppressed because it is too large Load Diff
+159
View File
@@ -0,0 +1,159 @@
#ifndef SIMDJSON_DOM_FRACTURED_JSON_H
#define SIMDJSON_DOM_FRACTURED_JSON_H
#include "simdjson/dom/base.h"
#include "simdjson/dom/element.h"
namespace simdjson {
/**
* Configuration options for FracturedJson formatting.
*
* FracturedJson intelligently chooses between different layout strategies
* (inline, compact multiline, table, expanded) based on content complexity,
* length, and structure similarity.
*/
struct fractured_json_options {
/**
* Maximum total characters per line (default: 120).
* Content exceeding this will be expanded to multiple lines.
*/
size_t max_total_line_length = 120;
/**
* Maximum length for inlined elements (default: 80).
* Simple arrays/objects shorter than this may be rendered inline.
*/
size_t max_inline_length = 80;
/**
* Maximum nesting depth for inline rendering (default: 2).
* Elements with complexity exceeding this will be expanded.
* Complexity 0 = scalar, 1 = flat array/object, 2 = one level of nesting.
*/
size_t max_inline_complexity = 2;
/**
* Maximum complexity for compact array formatting (default: 1).
* Arrays with elements of this complexity or less may have multiple
* items per line.
*/
size_t max_compact_array_complexity = 1;
/**
* Number of spaces per indentation level (default: 4).
*/
size_t indent_spaces = 4;
/**
* Enable tabular formatting for arrays of similar objects (default: true).
* When enabled, arrays of objects with identical keys are formatted
* as aligned tables.
*/
bool enable_table_format = true;
/**
* Minimum number of rows to trigger table mode (default: 3).
*/
size_t min_table_rows = 3;
/**
* Similarity threshold for table detection (default: 0.8).
* Objects must share at least this fraction of keys to be formatted
* as a table.
*/
double table_similarity_threshold = 0.8;
/**
* Enable compact multiline arrays (default: true).
* When enabled, arrays of simple elements may have multiple items
* per line.
*/
bool enable_compact_multiline = true;
/**
* Maximum array items per line in compact mode (default: 10).
*/
size_t max_items_per_line = 10;
/**
* Add space inside brackets for simple containers (default: true).
* When true: { "key": "value" }
* When false: {"key": "value"}
*/
bool simple_bracket_padding = true;
/**
* Add space after colons (default: true).
* When true: "key": "value"
* When false: "key":"value"
*/
bool colon_padding = true;
/**
* Add space after commas in inline content (default: true).
* When true: [1, 2, 3]
* When false: [1,2,3]
*/
bool comma_padding = true;
};
/**
* Format JSON using FracturedJson formatting with default options.
*
* FracturedJson produces human-readable yet compact output by intelligently
* choosing between inline, compact multiline, table, and expanded layouts.
*
* dom::parser parser;
* element doc = parser.parse(json_string);
* cout << fractured_json(doc) << endl;
*/
template <class T>
std::string fractured_json(T x);
/**
* Format JSON using FracturedJson formatting with custom options.
*
* dom::parser parser;
* element doc = parser.parse(json_string);
* fractured_json_options opts;
* opts.max_total_line_length = 80;
* cout << fractured_json(doc, opts) << endl;
*/
template <class T>
std::string fractured_json(T x, const fractured_json_options& options);
#if SIMDJSON_EXCEPTIONS
template <class T>
std::string fractured_json(simdjson_result<T> x);
template <class T>
std::string fractured_json(simdjson_result<T> x, const fractured_json_options& options);
#endif
/**
* Format a JSON string using FracturedJson formatting.
*
* This is useful for formatting output from the builder/static reflection API
* or any valid JSON string.
*
* // With static reflection
* MyStruct data = {...};
* auto minified = simdjson::to_json_string(data);
* auto formatted = simdjson::fractured_json_string(minified.value());
*
* // Or with any JSON string
* std::string json = R"({"key":"value"})";
* auto formatted = simdjson::fractured_json_string(json);
*/
inline std::string fractured_json_string(std::string_view json_str);
/**
* Format a JSON string using FracturedJson formatting with custom options.
*/
inline std::string fractured_json_string(std::string_view json_str,
const fractured_json_options& options);
} // namespace simdjson
#endif // SIMDJSON_DOM_FRACTURED_JSON_H
+1 -1
View File
@@ -186,7 +186,7 @@ inline simdjson_result<std::vector<element>> object::at_path_with_wildcard(std::
}
if (i >= json_path.size() || (json_path[i] != '.' && json_path[i] != '[')) {
// expect json path to always start with $ but this isn't currently
// expect JSONPath expressions to always start with $ but this isn't currently
// expected in jsonpathutil.h.
return INVALID_JSON_POINTER;
}
+2 -2
View File
@@ -175,7 +175,7 @@ public:
inline simdjson_result<element> at_pointer(std::string_view json_pointer) const noexcept;
/**
* Recursive function which processes the json path of each child element
* Recursive function which processes the JSON path of each child element
*/
inline void process_json_path_of_child_elements(std::vector<element>::iterator& current, std::vector<element>::iterator& end, const std::string_view& path_suffix, std::vector<element>& accumulator) const noexcept;
@@ -189,7 +189,7 @@ public:
* JSONPath queries that trivially convertible to JSON Pointer queries: key
* names and array indices.
*
* https://datatracker.ietf.org/doc/html/draft-normington-jsonpath-00
* https://www.rfc-editor.org/rfc/rfc9535 (RFC 9535)
*
* @return The value associated with the given JSONPath expression, or:
* - INVALID_JSON_POINTER if the JSONPath to JSON Pointer conversion fails
+2 -5
View File
@@ -204,10 +204,7 @@ public:
*
* ### std::string references
*
* If you pass a mutable std::string reference (std::string&), the parser will seek to extend
* its capacity to SIMDJSON_PADDING bytes beyond the end of the string.
*
* Whenever you pass an std::string reference, the parser will access the bytes beyond the end of
* Whenever you pass an std::string reference, the parser may access the bytes beyond the end of
* the string but before the end of the allocated memory (std::string::capacity()).
* If you are using a sanitizer that checks for reading uninitialized bytes or std::string's
* container-overflow checks, you may encounter sanitizer warnings.
@@ -239,7 +236,7 @@ public:
/** @overload parse(const uint8_t *buf, size_t len, bool realloc_if_needed) */
simdjson_inline simdjson_result<element> parse(const char *buf, size_t len, bool realloc_if_needed = true) & noexcept;
simdjson_inline simdjson_result<element> parse(const char *buf, size_t len, bool realloc_if_needed = true) && =delete;
/** @overload parse(const uint8_t *buf, size_t len, bool realloc_if_needed) */
/** @overload parse(const std::string &) */
simdjson_inline simdjson_result<element> parse(const std::string &s) & noexcept;
simdjson_inline simdjson_result<element> parse(const std::string &s) && =delete;
/** @overload parse(const uint8_t *buf, size_t len, bool realloc_if_needed) */
+1 -1
View File
@@ -8,7 +8,7 @@
namespace simdjson {
inline bool is_fatal(error_code error) noexcept {
return error == TAPE_ERROR || error == INCOMPLETE_ARRAY_OR_OBJECT;
return error == TAPE_ERROR || error == INCOMPLETE_ARRAY_OR_OBJECT || error == OUT_OF_ORDER_ITERATION || error == DEPTH_ERROR;
}
namespace internal {
+4
View File
@@ -265,6 +265,8 @@ struct simdjson_result_base : protected std::pair<T, error_code> {
*/
simdjson_inline T&& value_unsafe() && noexcept;
using value_type = T;
using error_type = error_code;
}; // struct simdjson_result_base
} // namespace internal
@@ -376,6 +378,8 @@ struct simdjson_result : public internal::simdjson_result_base<T> {
*/
simdjson_inline T&& value_unsafe() && noexcept;
using value_type = T;
using error_type = error_code;
}; // struct simdjson_result
#if SIMDJSON_EXCEPTIONS
+8
View File
@@ -0,0 +1,8 @@
#ifndef SIMDJSON_FALLBACK_BUILDER_H
#define SIMDJSON_FALLBACK_BUILDER_H
#include "simdjson/fallback/begin.h"
#include "simdjson/generic/builder/amalgamated.h"
#include "simdjson/fallback/end.h"
#endif // SIMDJSON_FALLBACK_BUILDER_H
@@ -13,7 +13,8 @@ namespace {
struct backslash_and_quote {
public:
static constexpr uint32_t BYTES_PROCESSED = 1;
simdjson_inline static backslash_and_quote copy_and_find(const uint8_t *src, uint8_t *dst);
// We only copy if dst is non-null.
simdjson_inline backslash_and_quote copy_and_find(const uint8_t *src, uint8_t *dst);
simdjson_inline bool has_quote_first() { return c == '"'; }
simdjson_inline bool has_backslash() { return c == '\\'; }
@@ -25,7 +26,9 @@ public:
simdjson_inline backslash_and_quote backslash_and_quote::copy_and_find(const uint8_t *src, uint8_t *dst) {
// store to dest unconditionally - we can overwrite the bits we don't like later
dst[0] = src[0];
if(dst != nullptr) {
dst[0] = src[0];
}
return { src[0] };
}
+2 -2
View File
@@ -17,10 +17,10 @@
#include "simdjson/arm64/begin.h"
#elif SIMDJSON_IMPLEMENTATION_PPC64
#include "simdjson/ppc64/begin.h"
#elif SIMDJSON_IMPLEMENTATION_LSX
#include "simdjson/lsx/begin.h"
#elif SIMDJSON_IMPLEMENTATION_LASX
#include "simdjson/lasx/begin.h"
#elif SIMDJSON_IMPLEMENTATION_LSX
#include "simdjson/lsx/begin.h"
#elif SIMDJSON_IMPLEMENTATION_FALLBACK
#include "simdjson/fallback/begin.h"
#else
@@ -0,0 +1,13 @@
#if defined(SIMDJSON_CONDITIONAL_INCLUDE) && !defined(SIMDJSON_GENERIC_BUILDER_DEPENDENCIES_H)
#error simdjson/generic/builder/dependencies.h must be included before simdjson/generic/builder/amalgamated.h!
#endif
#include "simdjson/generic/builder/json_string_builder.h"
#include "simdjson/generic/builder/json_builder.h"
#include "simdjson/generic/builder/fractured_json_builder.h"
// JSON builder inline definitions
#include "simdjson/generic/builder/json_string_builder-inl.h"
@@ -0,0 +1,14 @@
#ifdef SIMDJSON_CONDITIONAL_INCLUDE
#error simdjson/generic/builder/dependencies.h must be included before defining SIMDJSON_CONDITIONAL_INCLUDE!
#endif
#ifndef SIMDJSON_GENERIC_BUILDER_DEPENDENCIES_H
#define SIMDJSON_GENERIC_BUILDER_DEPENDENCIES_H
// Internal headers needed for builder generics.
// All includes not under simdjson/generic/builder must be here!
// Otherwise, amalgamation will fail.
#include "simdjson/concepts.h"
#include "simdjson/dom/fractured_json.h"
#endif // SIMDJSON_GENERIC_BUILDER_DEPENDENCIES_H
@@ -0,0 +1,117 @@
#ifndef SIMDJSON_GENERIC_FRACTURED_JSON_BUILDER_H
#define SIMDJSON_GENERIC_FRACTURED_JSON_BUILDER_H
#ifndef SIMDJSON_CONDITIONAL_INCLUDE
#include "simdjson/generic/builder/json_builder.h"
#include "simdjson/dom/fractured_json.h"
#endif // SIMDJSON_CONDITIONAL_INCLUDE
#if SIMDJSON_STATIC_REFLECTION
namespace simdjson {
namespace SIMDJSON_IMPLEMENTATION {
namespace builder {
/**
* Serialize an object to a FracturedJson-formatted string.
*
* FracturedJson produces human-readable yet compact JSON output by intelligently
* choosing between different layout strategies (inline, compact multiline, table,
* expanded) based on content complexity, length, and structure similarity.
*
* This function combines the builder's serialization with FracturedJson formatting:
* 1. Serializes the object to minified JSON using reflection
* 2. Parses and reformats using FracturedJson
*
* Example:
* struct User { int id; std::string name; bool active; };
* User user{1, "Alice", true};
* auto result = to_fractured_json_string(user);
* // result.value() == "{ \"id\": 1, \"name\": \"Alice\", \"active\": true }"
*
* @param obj The object to serialize (must be a reflectable type)
* @param opts FracturedJson formatting options
* @param initial_capacity Initial buffer capacity for serialization
* @return The formatted JSON string, or an error
*/
template <class T>
simdjson_warn_unused simdjson_result<std::string> to_fractured_json_string(
const T& obj,
const fractured_json_options& opts = {},
size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) {
// Step 1: Serialize to minified JSON
std::string formatted;
auto error = to_json_string(obj, initial_capacity).get(formatted);
if (error) {
return error;
}
// Step 2: Reformat with FracturedJson
return fractured_json_string(formatted, opts);
}
/**
* Extract specific fields from an object and format with FracturedJson.
*
* Example:
* struct User { int id; std::string name; std::string email; bool active; };
* User user{1, "Alice", "alice@example.com", true};
* auto result = extract_fractured_json<"id", "name">(user);
* // result.value() == "{ \"id\": 1, \"name\": \"Alice\" }"
*
* @param obj The object to serialize
* @param opts FracturedJson formatting options
* @param initial_capacity Initial buffer capacity for serialization
* @return The formatted JSON string containing only the specified fields
*/
template<constevalutil::fixed_string... FieldNames, typename T>
requires(std::is_class_v<T> && (sizeof...(FieldNames) > 0))
simdjson_warn_unused simdjson_result<std::string> extract_fractured_json(
const T& obj,
const fractured_json_options& opts = {},
size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) {
// Step 1: Extract fields to minified JSON
std::string formatted;
auto error = extract_from<FieldNames...>(obj, initial_capacity).get(formatted);
if (error) {
return error;
}
// Step 2: Reformat with FracturedJson
return fractured_json_string(formatted, opts);
}
} // namespace builder
} // namespace SIMDJSON_IMPLEMENTATION
// Global namespace convenience functions
/**
* Serialize an object to a FracturedJson-formatted string.
* Global namespace version for convenience.
*/
template <class T>
simdjson_warn_unused simdjson_result<std::string> to_fractured_json_string(
const T& obj,
const fractured_json_options& opts = {},
size_t initial_capacity = SIMDJSON_IMPLEMENTATION::builder::string_builder::DEFAULT_INITIAL_CAPACITY) {
return SIMDJSON_IMPLEMENTATION::builder::to_fractured_json_string(obj, opts, initial_capacity);
}
/**
* Extract specific fields from an object and format with FracturedJson.
* Global namespace version for convenience.
*/
template<constevalutil::fixed_string... FieldNames, typename T>
requires(std::is_class_v<T> && (sizeof...(FieldNames) > 0))
simdjson_warn_unused simdjson_result<std::string> extract_fractured_json(
const T& obj,
const fractured_json_options& opts = {},
size_t initial_capacity = SIMDJSON_IMPLEMENTATION::builder::string_builder::DEFAULT_INITIAL_CAPACITY) {
return SIMDJSON_IMPLEMENTATION::builder::extract_fractured_json<FieldNames...>(obj, opts, initial_capacity);
}
} // namespace simdjson
#endif // SIMDJSON_STATIC_REFLECTION
#endif // SIMDJSON_GENERIC_FRACTURED_JSON_BUILDER_H
@@ -1,7 +1,3 @@
/**
* This file is part of the builder API. It is temporarily in the ondemand directory
* but we will move it to a builder directory later.
*/
#ifndef SIMDJSON_GENERIC_BUILDER_H
#ifndef SIMDJSON_CONDITIONAL_INCLUDE
@@ -25,30 +21,21 @@ namespace simdjson {
namespace SIMDJSON_IMPLEMENTATION {
namespace builder {
// Concept that checks if a type is a container but not a string (because
// strings handling must be handled differently)
template <typename T>
concept container_but_not_string =
requires(T a) {
{ a.size() } -> std::convertible_to<std::size_t>;
{
a[std::declval<std::size_t>()]
}; // check if elements are accessible for the subscript operator
} && !std::is_same_v<T, std::string> &&
!std::is_same_v<T, std::string_view> && !std::is_same_v<T, const char *>;
template <class T>
requires(container_but_not_string<T>)
requires(concepts::container_but_not_string<T> && !require_custom_serialization<T>)
constexpr void atom(string_builder &b, const T &t) {
if (t.size() == 0) {
auto it = t.begin();
auto end = t.end();
if (it == end) {
b.append_raw("[]");
return;
}
b.append('[');
atom(b, t[0]);
for (size_t i = 1; i < t.size(); ++i) {
atom(b, *it);
++it;
for (; it != end; ++it) {
b.append(',');
atom(b, t[i]);
atom(b, *it);
}
b.append(']');
}
@@ -63,6 +50,7 @@ constexpr void atom(string_builder &b, const T &t) {
}
template <concepts::string_view_keyed_map T>
requires(!require_custom_serialization<T>)
constexpr void atom(string_builder &b, const T &m) {
if (m.empty()) {
b.append_raw("{}");
@@ -91,7 +79,7 @@ constexpr void atom(string_builder &b, const number_type t) {
}
template <class T>
requires(std::is_class_v<T> && !container_but_not_string<T> &&
requires(std::is_class_v<T> && !concepts::container_but_not_string<T> &&
!concepts::string_view_keyed_map<T> &&
!concepts::optional_type<T> &&
!concepts::smart_pointer<T> &&
@@ -99,7 +87,7 @@ template <class T>
!std::is_same_v<T, std::string> &&
!std::is_same_v<T, std::string_view> &&
!std::is_same_v<T, const char*> &&
!std::is_same_v<T, char>)
!std::is_same_v<T, char> && !require_custom_serialization<T>)
constexpr void atom(string_builder &b, const T &t) {
int i = 0;
b.append('{');
@@ -117,6 +105,7 @@ constexpr void atom(string_builder &b, const T &t) {
// Support for optional types (std::optional, etc.)
template <concepts::optional_type T>
requires(!require_custom_serialization<T>)
constexpr void atom(string_builder &b, const T &opt) {
if (opt) {
atom(b, opt.value());
@@ -127,6 +116,7 @@ constexpr void atom(string_builder &b, const T &opt) {
// Support for smart pointers (std::unique_ptr, std::shared_ptr, etc.)
template <concepts::smart_pointer T>
requires(!require_custom_serialization<T>)
constexpr void atom(string_builder &b, const T &ptr) {
if (ptr) {
atom(b, *ptr);
@@ -137,7 +127,7 @@ constexpr void atom(string_builder &b, const T &ptr) {
// Support for enums - serialize as string representation using expand approach from P2996R12
template <typename T>
requires(std::is_enum_v<T>)
requires(std::is_enum_v<T> && !require_custom_serialization<T>)
void atom(string_builder &b, const T &e) {
#if SIMDJSON_STATIC_REFLECTION
constexpr auto enumerators = std::define_static_array(std::meta::enumerators_of(^^T));
@@ -158,10 +148,10 @@ void atom(string_builder &b, const T &e) {
// Support for appendable containers that don't have operator[] (sets, etc.)
template <concepts::appendable_containers T>
requires(!container_but_not_string<T> && !concepts::string_view_keyed_map<T> &&
requires(!concepts::container_but_not_string<T> && !concepts::string_view_keyed_map<T> &&
!concepts::optional_type<T> && !concepts::smart_pointer<T> &&
!std::is_same_v<T, std::string> &&
!std::is_same_v<T, std::string_view> && !std::is_same_v<T, const char*>)
!std::is_same_v<T, std::string_view> && !std::is_same_v<T, const char*> && !require_custom_serialization<T>)
constexpr void atom(string_builder &b, const T &container) {
if (container.empty()) {
b.append_raw("[]");
@@ -196,32 +186,35 @@ void append(string_builder &b, const T &t) {
}
template <concepts::optional_type T>
requires(!require_custom_serialization<T>)
void append(string_builder &b, const T &t) {
atom(b, t);
}
template <concepts::smart_pointer T>
requires(!require_custom_serialization<T>)
void append(string_builder &b, const T &t) {
atom(b, t);
}
template <concepts::appendable_containers T>
requires(!container_but_not_string<T> && !concepts::string_view_keyed_map<T> &&
requires(!concepts::container_but_not_string<T> && !concepts::string_view_keyed_map<T> &&
!concepts::optional_type<T> && !concepts::smart_pointer<T> &&
!std::is_same_v<T, std::string> &&
!std::is_same_v<T, std::string_view> && !std::is_same_v<T, const char*>)
!std::is_same_v<T, std::string_view> && !std::is_same_v<T, const char*> && !require_custom_serialization<T>)
void append(string_builder &b, const T &t) {
atom(b, t);
}
template <concepts::string_view_keyed_map T>
requires(!require_custom_serialization<T>)
void append(string_builder &b, const T &t) {
atom(b, t);
}
// works for struct
template <class Z>
requires(std::is_class_v<Z> && !container_but_not_string<Z> &&
requires(std::is_class_v<Z> && !concepts::container_but_not_string<Z> &&
!concepts::string_view_keyed_map<Z> &&
!concepts::optional_type<Z> &&
!concepts::smart_pointer<Z> &&
@@ -229,7 +222,7 @@ template <class Z>
!std::is_same_v<Z, std::string> &&
!std::is_same_v<Z, std::string_view> &&
!std::is_same_v<Z, const char*> &&
!std::is_same_v<Z, char>)
!std::is_same_v<Z, char> && !require_custom_serialization<Z>)
void append(string_builder &b, const Z &z) {
int i = 0;
b.append('{');
@@ -245,25 +238,35 @@ void append(string_builder &b, const Z &z) {
b.append('}');
}
// works for container
// works for container that have begin() and end() iterators
template <class Z>
requires(container_but_not_string<Z>)
requires(concepts::container_but_not_string<Z> && !require_custom_serialization<Z>)
void append(string_builder &b, const Z &z) {
if (z.size() == 0) {
auto it = z.begin();
auto end = z.end();
if (it == end) {
b.append_raw("[]");
return;
}
b.append('[');
atom(b, z[0]);
for (size_t i = 1; i < z.size(); ++i) {
atom(b, *it);
++it;
for (; it != end; ++it) {
b.append(',');
atom(b, z[i]);
atom(b, *it);
}
b.append(']');
}
template <class Z>
simdjson_result<std::string> to_json_string(const Z &z, size_t initial_capacity = 1024) {
requires (require_custom_serialization<Z>)
void append(string_builder &b, const Z &z) {
b.append(z);
}
template <class Z>
simdjson_warn_unused simdjson_result<std::string> to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) {
string_builder b(initial_capacity);
append(b, z);
std::string_view s;
@@ -272,8 +275,8 @@ simdjson_result<std::string> to_json_string(const Z &z, size_t initial_capacity
}
template <class Z>
simdjson_error to_json(const Z &z, std::string &s) {
string_builder b;
simdjson_warn_unused simdjson_error to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) {
string_builder b(initial_capacity);
append(b, z);
std::string_view view;
if(auto e = b.view().get(view); e) { return e; }
@@ -286,13 +289,88 @@ string_builder& operator<<(string_builder& b, const Z& z) {
append(b, z);
return b;
}
// extract_from: Serialize only specific fields from a struct to JSON
template<constevalutil::fixed_string... FieldNames, typename T>
requires(std::is_class_v<T> && (sizeof...(FieldNames) > 0))
void extract_from(string_builder &b, const T &obj) {
// Helper to check if a field name matches any of the requested fields
auto should_extract = [](std::string_view field_name) constexpr -> bool {
return ((FieldNames.view() == field_name) || ...);
};
b.append('{');
bool first = true;
// Iterate through all members of T using reflection
template for (constexpr auto mem : std::define_static_array(
std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) {
if constexpr (std::meta::is_public(mem)) {
constexpr std::string_view key = std::define_static_string(std::meta::identifier_of(mem));
// Only serialize this field if it's in our list of requested fields
if constexpr (should_extract(key)) {
if (!first) {
b.append(',');
}
first = false;
// Serialize the key
constexpr auto quoted_key = std::define_static_string(constevalutil::consteval_to_quoted_escaped(std::meta::identifier_of(mem)));
b.append_raw(quoted_key);
b.append(':');
// Serialize the value
atom(b, obj.[:mem:]);
}
}
};
b.append('}');
}
template<constevalutil::fixed_string... FieldNames, typename T>
requires(std::is_class_v<T> && (sizeof...(FieldNames) > 0))
simdjson_warn_unused simdjson_result<std::string> extract_from(const T &obj, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) {
string_builder b(initial_capacity);
extract_from<FieldNames...>(b, obj);
std::string_view s;
if(auto e = b.view().get(s); e) { return e; }
return std::string(s);
}
} // namespace builder
} // namespace SIMDJSON_IMPLEMENTATION
// Alias the function template to 'to' in the global namespace
template <class Z>
simdjson_result<std::string> to_json(const Z &z, size_t initial_capacity = 1024) {
return SIMDJSON_IMPLEMENTATION::builder::to_json_string(z, initial_capacity);
simdjson_warn_unused simdjson_result<std::string> to_json(const Z &z, size_t initial_capacity = SIMDJSON_IMPLEMENTATION::builder::string_builder::DEFAULT_INITIAL_CAPACITY) {
SIMDJSON_IMPLEMENTATION::builder::string_builder b(initial_capacity);
SIMDJSON_IMPLEMENTATION::builder::append(b, z);
std::string_view s;
if(auto e = b.view().get(s); e) { return e; }
return std::string(s);
}
template <class Z>
simdjson_warn_unused simdjson_error to_json(const Z &z, std::string &s, size_t initial_capacity = SIMDJSON_IMPLEMENTATION::builder::string_builder::DEFAULT_INITIAL_CAPACITY) {
SIMDJSON_IMPLEMENTATION::builder::string_builder b(initial_capacity);
SIMDJSON_IMPLEMENTATION::builder::append(b, z);
std::string_view view;
if(auto e = b.view().get(view); e) { return e; }
s.assign(view);
return SUCCESS;
}
// Global namespace function for extract_from
template<constevalutil::fixed_string... FieldNames, typename T>
requires(std::is_class_v<T> && (sizeof...(FieldNames) > 0))
simdjson_warn_unused simdjson_result<std::string> extract_from(const T &obj, size_t initial_capacity = SIMDJSON_IMPLEMENTATION::builder::string_builder::DEFAULT_INITIAL_CAPACITY) {
SIMDJSON_IMPLEMENTATION::builder::string_builder b(initial_capacity);
SIMDJSON_IMPLEMENTATION::builder::extract_from<FieldNames...>(b, obj);
std::string_view s;
if(auto e = b.view().get(s); e) { return e; }
return std::string(s);
}
} // namespace simdjson
#endif // SIMDJSON_STATIC_REFLECTION
@@ -1,7 +1,3 @@
/**
* This file is part of the builder API. It is temporarily in the ondemand
* directory but we will move it to a builder directory later.
*/
#include <array>
#include <cstring>
#include <type_traits>
@@ -42,24 +38,25 @@ namespace simdjson {
namespace SIMDJSON_IMPLEMENTATION {
namespace builder {
static SIMDJSON_CONSTEXPR_LAMBDA std::array<uint8_t, 256> json_quotable_character = {
1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0};
static SIMDJSON_CONSTEXPR_LAMBDA std::array<uint8_t, 256>
json_quotable_character = {
1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0};
/**
A possible SWAR implementation of has_json_escapable_byte. It is not used because
it is slower than the current implementation. It is kept here for reference (to show
that we tried it).
A possible SWAR implementation of has_json_escapable_byte. It is not used
because it is slower than the current implementation. It is kept here for
reference (to show that we tried it).
inline bool has_json_escapable_byte(uint64_t x) {
uint64_t is_ascii = 0x8080808080808080ULL & ~x;
@@ -76,7 +73,7 @@ SIMDJSON_CONSTEXPR_LAMBDA simdjson_inline bool
simple_needs_escaping(std::string_view v) {
for (char c : v) {
// a table lookup is faster than a series of comparisons
if(json_quotable_character[static_cast<uint8_t>(c)]) {
if (json_quotable_character[static_cast<uint8_t>(c)]) {
return true;
}
}
@@ -117,7 +114,8 @@ simdjson_inline bool fast_needs_escaping(std::string_view view) {
__m128i running = _mm_setzero_si128();
for (; i + 15 < view.size(); i += 16) {
__m128i word = _mm_loadu_si128(reinterpret_cast<const __m128i *>(view.data() + i));
__m128i word =
_mm_loadu_si128(reinterpret_cast<const __m128i *>(view.data() + i));
running = _mm_or_si128(running, _mm_cmpeq_epi8(word, _mm_set1_epi8(34)));
running = _mm_or_si128(running, _mm_cmpeq_epi8(word, _mm_set1_epi8(92)));
running = _mm_or_si128(
@@ -125,8 +123,8 @@ simdjson_inline bool fast_needs_escaping(std::string_view view) {
_mm_setzero_si128()));
}
if (i < view.size()) {
__m128i word =
_mm_loadu_si128(reinterpret_cast<const __m128i *>(view.data() + view.length() - 16));
__m128i word = _mm_loadu_si128(
reinterpret_cast<const __m128i *>(view.data() + view.length() - 16));
running = _mm_or_si128(running, _mm_cmpeq_epi8(word, _mm_set1_epi8(34)));
running = _mm_or_si128(running, _mm_cmpeq_epi8(word, _mm_set1_epi8(92)));
running = _mm_or_si128(
@@ -141,7 +139,6 @@ simdjson_inline bool fast_needs_escaping(std::string_view view) {
}
#endif
SIMDJSON_CONSTEXPR_LAMBDA inline size_t
find_next_json_quotable_character(const std::string_view view,
size_t location) noexcept {
@@ -156,15 +153,15 @@ find_next_json_quotable_character(const std::string_view view,
SIMDJSON_CONSTEXPR_LAMBDA static std::string_view control_chars[] = {
"\\u0000", "\\u0001", "\\u0002", "\\u0003", "\\u0004", "\\u0005", "\\u0006",
"\\u0007", "\\b", "\\t", "\\n", "\\u000b", "\\f", "\\r",
"\\u0007", "\\b", "\\t", "\\n", "\\u000b", "\\f", "\\r",
"\\u000e", "\\u000f", "\\u0010", "\\u0011", "\\u0012", "\\u0013", "\\u0014",
"\\u0015", "\\u0016", "\\u0017", "\\u0018", "\\u0019", "\\u001a", "\\u001b",
"\\u001c", "\\u001d", "\\u001e", "\\u001f"};
// All Unicode characters may be placed within the quotation marks, except for the
// characters that MUST be escaped: quotation mark, reverse solidus, and the control
// characters (U+0000 through U+001F).
// There are two-character sequence escape representations of some popular characters:
// All Unicode characters may be placed within the quotation marks, except for
// the characters that MUST be escaped: quotation mark, reverse solidus, and the
// control characters (U+0000 through U+001F). There are two-character sequence
// escape representations of some popular characters:
// \", \\, \b, \f, \n, \r, \t.
SIMDJSON_CONSTEXPR_LAMBDA void escape_json_char(char c, char *&out) {
if (c == '"') {
@@ -284,10 +281,11 @@ simdjson_inline void string_builder::clear() noexcept {
namespace internal {
template <typename number_type, typename = typename std::enable_if<
std::is_unsigned<number_type>::value>::type>
simdjson_really_inline int int_log2(number_type x) { return 63 - leading_zeroes(uint64_t(x) | 1); }
simdjson_really_inline int int_log2(number_type x) {
return 63 - leading_zeroes(uint64_t(x) | 1);
}
simdjson_really_inline int fast_digit_count_32(uint32_t x) {
static uint64_t table[] = {
@@ -301,7 +299,6 @@ simdjson_really_inline int fast_digit_count_32(uint32_t x) {
return uint32_t((x + table[int_log2(x)]) >> 32);
}
simdjson_really_inline int fast_digit_count_64(uint64_t x) {
static uint64_t table[] = {9,
99,
@@ -335,28 +332,29 @@ simdjson_really_inline size_t digit_count(number_type v) noexcept {
"We only support 8-bit, 16-bit, 32-bit and 64-bit numbers");
SIMDJSON_IF_CONSTEXPR(sizeof(number_type) <= 4) {
return fast_digit_count_32(static_cast<uint32_t>(v));
} else {
}
else {
return fast_digit_count_64(static_cast<uint64_t>(v));
}
}
static const char decimal_table[200] = {
0x30, 0x30, 0x30, 0x31, 0x30, 0x32, 0x30, 0x33, 0x30, 0x34, 0x30, 0x35,
0x30, 0x36, 0x30, 0x37, 0x30, 0x38, 0x30, 0x39, 0x31, 0x30, 0x31, 0x31,
0x31, 0x32, 0x31, 0x33, 0x31, 0x34, 0x31, 0x35, 0x31, 0x36, 0x31, 0x37,
0x31, 0x38, 0x31, 0x39, 0x32, 0x30, 0x32, 0x31, 0x32, 0x32, 0x32, 0x33,
0x32, 0x34, 0x32, 0x35, 0x32, 0x36, 0x32, 0x37, 0x32, 0x38, 0x32, 0x39,
0x33, 0x30, 0x33, 0x31, 0x33, 0x32, 0x33, 0x33, 0x33, 0x34, 0x33, 0x35,
0x33, 0x36, 0x33, 0x37, 0x33, 0x38, 0x33, 0x39, 0x34, 0x30, 0x34, 0x31,
0x34, 0x32, 0x34, 0x33, 0x34, 0x34, 0x34, 0x35, 0x34, 0x36, 0x34, 0x37,
0x34, 0x38, 0x34, 0x39, 0x35, 0x30, 0x35, 0x31, 0x35, 0x32, 0x35, 0x33,
0x35, 0x34, 0x35, 0x35, 0x35, 0x36, 0x35, 0x37, 0x35, 0x38, 0x35, 0x39,
0x36, 0x30, 0x36, 0x31, 0x36, 0x32, 0x36, 0x33, 0x36, 0x34, 0x36, 0x35,
0x36, 0x36, 0x36, 0x37, 0x36, 0x38, 0x36, 0x39, 0x37, 0x30, 0x37, 0x31,
0x37, 0x32, 0x37, 0x33, 0x37, 0x34, 0x37, 0x35, 0x37, 0x36, 0x37, 0x37,
0x37, 0x38, 0x37, 0x39, 0x38, 0x30, 0x38, 0x31, 0x38, 0x32, 0x38, 0x33,
0x38, 0x34, 0x38, 0x35, 0x38, 0x36, 0x38, 0x37, 0x38, 0x38, 0x38, 0x39,
0x39, 0x30, 0x39, 0x31, 0x39, 0x32, 0x39, 0x33, 0x39, 0x34, 0x39, 0x35,
0x39, 0x36, 0x39, 0x37, 0x39, 0x38, 0x39, 0x39,
0x30, 0x30, 0x30, 0x31, 0x30, 0x32, 0x30, 0x33, 0x30, 0x34, 0x30, 0x35,
0x30, 0x36, 0x30, 0x37, 0x30, 0x38, 0x30, 0x39, 0x31, 0x30, 0x31, 0x31,
0x31, 0x32, 0x31, 0x33, 0x31, 0x34, 0x31, 0x35, 0x31, 0x36, 0x31, 0x37,
0x31, 0x38, 0x31, 0x39, 0x32, 0x30, 0x32, 0x31, 0x32, 0x32, 0x32, 0x33,
0x32, 0x34, 0x32, 0x35, 0x32, 0x36, 0x32, 0x37, 0x32, 0x38, 0x32, 0x39,
0x33, 0x30, 0x33, 0x31, 0x33, 0x32, 0x33, 0x33, 0x33, 0x34, 0x33, 0x35,
0x33, 0x36, 0x33, 0x37, 0x33, 0x38, 0x33, 0x39, 0x34, 0x30, 0x34, 0x31,
0x34, 0x32, 0x34, 0x33, 0x34, 0x34, 0x34, 0x35, 0x34, 0x36, 0x34, 0x37,
0x34, 0x38, 0x34, 0x39, 0x35, 0x30, 0x35, 0x31, 0x35, 0x32, 0x35, 0x33,
0x35, 0x34, 0x35, 0x35, 0x35, 0x36, 0x35, 0x37, 0x35, 0x38, 0x35, 0x39,
0x36, 0x30, 0x36, 0x31, 0x36, 0x32, 0x36, 0x33, 0x36, 0x34, 0x36, 0x35,
0x36, 0x36, 0x36, 0x37, 0x36, 0x38, 0x36, 0x39, 0x37, 0x30, 0x37, 0x31,
0x37, 0x32, 0x37, 0x33, 0x37, 0x34, 0x37, 0x35, 0x37, 0x36, 0x37, 0x37,
0x37, 0x38, 0x37, 0x39, 0x38, 0x30, 0x38, 0x31, 0x38, 0x32, 0x38, 0x33,
0x38, 0x34, 0x38, 0x35, 0x38, 0x36, 0x38, 0x37, 0x38, 0x38, 0x38, 0x39,
0x39, 0x30, 0x39, 0x31, 0x39, 0x32, 0x39, 0x33, 0x39, 0x34, 0x39, 0x35,
0x39, 0x36, 0x39, 0x37, 0x39, 0x38, 0x39, 0x39,
};
} // namespace internal
@@ -392,7 +390,7 @@ simdjson_inline void string_builder::append(number_type v) noexcept {
size_t dc = internal::digit_count(pv);
char *write_pointer = buffer.get() + position + dc - 1;
while (pv >= 100) {
memcpy(write_pointer - 1, &internal::decimal_table[(pv % 100)*2], 2);
memcpy(write_pointer - 1, &internal::decimal_table[(pv % 100) * 2], 2);
write_pointer -= 2;
pv /= 100;
}
@@ -414,12 +412,12 @@ simdjson_inline void string_builder::append(number_type v) noexcept {
pv = 0 - pv; // the 0 is for Microsoft
}
size_t dc = internal::digit_count(pv);
if (negative) {
buffer.get()[position++] = '-';
}
// by always writing the minus sign, we avoid the branch.
buffer.get()[position] = '-';
position += negative ? 1 : 0;
char *write_pointer = buffer.get() + position + dc - 1;
while (pv >= 100) {
memcpy(write_pointer - 1, &internal::decimal_table[(pv % 100)*2], 2);
memcpy(write_pointer - 1, &internal::decimal_table[(pv % 100) * 2], 2);
write_pointer -= 2;
pv /= 100;
}
@@ -471,14 +469,15 @@ string_builder::escape_and_append_with_quotes(char input) noexcept {
}
}
simdjson_inline void string_builder::escape_and_append_with_quotes(const char* input) noexcept {
simdjson_inline void
string_builder::escape_and_append_with_quotes(const char *input) noexcept {
std::string_view cinput(input);
escape_and_append_with_quotes(cinput);
}
#if SIMDJSON_SUPPORTS_CONCEPTS
template<internal::fixed_string key>
simdjson_inline void string_builder::escape_and_append_with_quotes() noexcept {
escape_and_append_with_quotes(internal::string_constant<key>::value);
template <constevalutil::fixed_string key>
simdjson_inline void string_builder::escape_and_append_with_quotes() noexcept {
escape_and_append_with_quotes(constevalutil::string_constant<key>::value);
}
#endif
@@ -505,6 +504,7 @@ simdjson_inline void string_builder::append_raw(const char *str,
#if SIMDJSON_SUPPORTS_CONCEPTS
// Support for optional types (std::optional, etc.)
template <concepts::optional_type T>
requires(!require_custom_serialization<T>)
simdjson_inline void string_builder::append(const T &opt) {
if (opt) {
append(*opt);
@@ -512,22 +512,29 @@ simdjson_inline void string_builder::append(const T &opt) {
append_null();
}
}
template <typename T>
requires(std::is_convertible<T, std::string_view>::value ||
std::is_same<T, const char*>::value )
requires(require_custom_serialization<T>)
simdjson_inline void string_builder::append(T &&val) {
serialize(*this, std::forward<T>(val));
}
template <typename T>
requires(std::is_convertible<T, std::string_view>::value ||
std::is_same<T, const char *>::value)
simdjson_inline void string_builder::append(const T &value) {
escape_and_append_with_quotes(value);
}
#endif
#if SIMDJSON_SUPPORTS_RANGES && SIMDJSON_SUPPORTS_CONCEPTS
// Support for range-based appending (std::ranges::view, etc.)
// Support for range-based appending (std::ranges::view, etc.)
template <std::ranges::range R>
requires (!std::is_convertible<R, std::string_view>::value)
requires(!std::is_convertible<R, std::string_view>::value && !require_custom_serialization<R>)
simdjson_inline void string_builder::append(const R &range) noexcept {
auto it = std::ranges::begin(range);
auto end = std::ranges::end(range);
if constexpr (concepts::is_pair<typename R::value_type>) {
if constexpr (concepts::is_pair<std::ranges::range_value_t<R>>) {
start_object();
if (it == end) {
@@ -540,8 +547,8 @@ simdjson_inline void string_builder::append(const R &range) noexcept {
// Append remaining items with preceding commas
for (; it != end; ++it) {
append_comma();
append_key_value(it->first, it->second);
append_comma();
append_key_value(it->first, it->second);
}
end_object();
} else {
@@ -557,11 +564,10 @@ simdjson_inline void string_builder::append(const R &range) noexcept {
// Append remaining items with preceding commas
for (; it != end; ++it) {
append_comma();
append(*it);
append_comma();
append(*it);
}
end_array();
}
}
@@ -569,7 +575,7 @@ simdjson_inline void string_builder::append(const R &range) noexcept {
#if SIMDJSON_EXCEPTIONS
simdjson_inline string_builder::operator std::string() const noexcept(false) {
return std::string(std::string_view());
return std::string(operator std::string_view());
}
simdjson_inline string_builder::operator std::string_view() const
@@ -598,82 +604,88 @@ simdjson_inline bool string_builder::validate_unicode() const noexcept {
return simdjson::validate_utf8(buffer.get(), position);
}
simdjson_inline void string_builder::start_object() noexcept {
simdjson_inline void string_builder::start_object() noexcept {
if (capacity_check(1)) {
buffer.get()[position++] = '{';
}
}
simdjson_inline void string_builder::end_object() noexcept {
simdjson_inline void string_builder::end_object() noexcept {
if (capacity_check(1)) {
buffer.get()[position++] = '}';
}
}
simdjson_inline void string_builder::start_array() noexcept {
simdjson_inline void string_builder::start_array() noexcept {
if (capacity_check(1)) {
buffer.get()[position++] = '[';
}
}
simdjson_inline void string_builder::end_array() noexcept {
simdjson_inline void string_builder::end_array() noexcept {
if (capacity_check(1)) {
buffer.get()[position++] = ']';
}
}
simdjson_inline void string_builder::append_comma() noexcept {
simdjson_inline void string_builder::append_comma() noexcept {
if (capacity_check(1)) {
buffer.get()[position++] = ',';
}
}
simdjson_inline void string_builder::append_colon() noexcept {
simdjson_inline void string_builder::append_colon() noexcept {
if (capacity_check(1)) {
buffer.get()[position++] = ':';
}
}
template<typename key_type, typename value_type>
simdjson_inline void string_builder::append_key_value(key_type key, value_type value) noexcept {
static_assert(
std::is_same<key_type, const char*>::value ||
std::is_convertible<key_type, std::string_view>::value,
"Unsupported key type");
template <typename key_type, typename value_type>
simdjson_inline void
string_builder::append_key_value(key_type key, value_type value) noexcept {
static_assert(std::is_same<key_type, const char *>::value ||
std::is_convertible<key_type, std::string_view>::value,
"Unsupported key type");
escape_and_append_with_quotes(key);
append_colon();
SIMDJSON_IF_CONSTEXPR(std::is_same<value_type, std::nullptr_t>::value) {
append_null();
} else SIMDJSON_IF_CONSTEXPR(std::is_same<value_type, char>::value) {
}
else SIMDJSON_IF_CONSTEXPR(std::is_same<value_type, char>::value) {
escape_and_append_with_quotes(value);
} else SIMDJSON_IF_CONSTEXPR(std::is_convertible<value_type, std::string_view>::value) {
}
else SIMDJSON_IF_CONSTEXPR(
std::is_convertible<value_type, std::string_view>::value) {
escape_and_append_with_quotes(value);
} else SIMDJSON_IF_CONSTEXPR(std::is_same<value_type, const char*>::value) {
}
else SIMDJSON_IF_CONSTEXPR(std::is_same<value_type, const char *>::value) {
escape_and_append_with_quotes(value);
} else {
}
else {
append(value);
}
}
#if SIMDJSON_SUPPORTS_CONCEPTS
template<internal::fixed_string key, typename value_type>
simdjson_inline void string_builder::append_key_value(value_type value) noexcept {
template <constevalutil::fixed_string key, typename value_type>
simdjson_inline void
string_builder::append_key_value(value_type value) noexcept {
escape_and_append_with_quotes<key>();
append_colon();
SIMDJSON_IF_CONSTEXPR(std::is_same<value_type, std::nullptr_t>::value) {
append_null();
} else SIMDJSON_IF_CONSTEXPR(std::is_same<value_type, char>::value) {
}
else SIMDJSON_IF_CONSTEXPR(std::is_same<value_type, char>::value) {
escape_and_append_with_quotes(value);
} else SIMDJSON_IF_CONSTEXPR(std::is_convertible<value_type, std::string_view>::value) {
}
else SIMDJSON_IF_CONSTEXPR(
std::is_convertible<value_type, std::string_view>::value) {
escape_and_append_with_quotes(value);
} else SIMDJSON_IF_CONSTEXPR(std::is_same<value_type, const char*>::value) {
}
else SIMDJSON_IF_CONSTEXPR(std::is_same<value_type, const char *>::value) {
escape_and_append_with_quotes(value);
} else {
}
else {
append(value);
}
}
@@ -1,7 +1,3 @@
/**
* This file is part of the builder API. It is temporarily in the ondemand directory
* but we will move it to a builder directory later.
*/
#ifndef SIMDJSON_GENERIC_STRING_BUILDER_H
#ifndef SIMDJSON_CONDITIONAL_INCLUDE
@@ -10,28 +6,39 @@
#endif // SIMDJSON_CONDITIONAL_INCLUDE
namespace simdjson {
#if SIMDJSON_SUPPORTS_CONCEPTS
namespace SIMDJSON_IMPLEMENTATION {
namespace builder {
#if SIMDJSON_SUPPORTS_CONCEPTS
// Helper to create string constants
namespace internal {
template <std::size_t N>
struct fixed_string {
constexpr fixed_string(const char (&str)[N]) {
for (std::size_t i = 0; i < N; ++i) {
data[i] = str[i];
}
}
char data[N];
constexpr std::string_view view() const { return {data, N - 1}; }
};
class string_builder;
}}
template <fixed_string str>
struct string_constant {
static constexpr std::string_view value = str.view();
};
} // namespace internal
template <typename T, typename = void>
struct has_custom_serialization : std::false_type {};
inline constexpr struct serialize_tag {
template <typename T>
constexpr void operator()(SIMDJSON_IMPLEMENTATION::builder::string_builder& b, T&& obj) const{
return tag_invoke(*this, b, std::forward<T>(obj));
}
} serialize{};
template <typename T>
struct has_custom_serialization<T, std::void_t<
decltype(tag_invoke(serialize, std::declval<SIMDJSON_IMPLEMENTATION::builder::string_builder&>(), std::declval<T&>()))
>> : std::true_type {};
template <typename T>
constexpr bool require_custom_serialization = has_custom_serialization<T>::value;
#else
struct has_custom_serialization : std::false_type {};
#endif // SIMDJSON_SUPPORTS_CONCEPTS
namespace SIMDJSON_IMPLEMENTATION {
namespace builder {
/**
* A builder for JSON strings representing documents. This is a low-level
* builder that is not meant to be used directly by end-users. Though it
@@ -43,7 +50,9 @@ struct string_constant {
*/
class string_builder {
public:
simdjson_inline string_builder(size_t initial_capacity = 1024);
simdjson_inline string_builder(size_t initial_capacity = DEFAULT_INITIAL_CAPACITY);
static constexpr size_t DEFAULT_INITIAL_CAPACITY = 1024;
/**
* Append number (includes Booleans). Booleans are mapped to the strings
@@ -52,7 +61,7 @@ public:
* represents the number.
*/
template<typename number_type,
typename = typename std::enable_if<std::is_arithmetic<number_type>::value>::type>
typename = typename std::enable_if<std::is_arithmetic<number_type>::value>::type>
simdjson_inline void append(number_type v) noexcept;
/**
@@ -82,7 +91,7 @@ public:
*/
simdjson_inline void escape_and_append_with_quotes(std::string_view input) noexcept;
#if SIMDJSON_SUPPORTS_CONCEPTS
template<internal::fixed_string key>
template<constevalutil::fixed_string key>
simdjson_inline void escape_and_append_with_quotes() noexcept;
#endif
/**
@@ -141,13 +150,18 @@ public:
template<typename key_type, typename value_type>
simdjson_inline void append_key_value(key_type key, value_type value) noexcept;
#if SIMDJSON_SUPPORTS_CONCEPTS
template<internal::fixed_string key, typename value_type>
template<constevalutil::fixed_string key, typename value_type>
simdjson_inline void append_key_value(value_type value) noexcept;
// Support for optional types (std::optional, etc.)
template <concepts::optional_type T>
requires(!require_custom_serialization<T>)
simdjson_inline void append(const T &opt);
template <typename T>
requires(require_custom_serialization<T>)
simdjson_inline void append(T &&val);
// Support for string-like types
template <typename T>
requires(std::is_convertible<T, std::string_view>::value ||
@@ -157,7 +171,7 @@ public:
#if SIMDJSON_SUPPORTS_RANGES && SIMDJSON_SUPPORTS_CONCEPTS
// Support for range-based appending (std::ranges::view, etc.)
template <std::ranges::range R>
requires (!std::is_convertible<R, std::string_view>::value)
requires (!std::is_convertible<R, std::string_view>::value && !require_custom_serialization<R>)
simdjson_inline void append(const R &range) noexcept;
#endif
/**
@@ -257,7 +271,7 @@ private:
#if !SIMDJSON_STATIC_REFLECTION
// fallback implementation until we have static reflection
template <class Z>
simdjson_result<std::string> to_json(const Z &z, size_t initial_capacity = 1024) {
simdjson_warn_unused simdjson_result<std::string> to_json(const Z &z, size_t initial_capacity = simdjson::SIMDJSON_IMPLEMENTATION::builder::string_builder::DEFAULT_INITIAL_CAPACITY) {
simdjson::SIMDJSON_IMPLEMENTATION::builder::string_builder b(initial_capacity);
b.append(z);
std::string_view s;
@@ -265,8 +279,20 @@ simdjson_result<std::string> to_json(const Z &z, size_t initial_capacity = 1024)
if(e) { return e; }
return std::string(s);
}
template <class Z>
simdjson_warn_unused simdjson_error to_json(const Z &z, std::string &s, size_t initial_capacity = simdjson::SIMDJSON_IMPLEMENTATION::builder::string_builder::DEFAULT_INITIAL_CAPACITY) {
simdjson::SIMDJSON_IMPLEMENTATION::builder::string_builder b(initial_capacity);
b.append(z);
std::string_view sv;
auto e = b.view().get(sv);
if(e) { return e; }
s.assign(sv.data(), sv.size());
return simdjson::SUCCESS;
}
#endif
#if SIMDJSON_SUPPORTS_CONCEPTS
#endif // SIMDJSON_SUPPORTS_CONCEPTS
} // namespace simdjson
@@ -40,6 +40,7 @@ public:
simdjson_warn_unused error_code stage1(const uint8_t *buf, size_t len, stage1_mode partial) noexcept final;
simdjson_warn_unused error_code stage2(dom::document &doc) noexcept final;
simdjson_warn_unused error_code stage2_next(dom::document &doc) noexcept final;
simdjson_warn_unused std::pair<const uint8_t *,bool> parse_string_if_needed(const uint8_t *src, uint8_t *dst, bool allow_replacement) const noexcept final;
simdjson_warn_unused uint8_t *parse_string(const uint8_t *src, uint8_t *dst, bool allow_replacement) const noexcept final;
simdjson_warn_unused uint8_t *parse_wobbly_string(const uint8_t *src, uint8_t *dst) const noexcept final;
inline simdjson_warn_unused error_code set_capacity(size_t capacity) noexcept final;
@@ -29,13 +29,13 @@ simdjson_warn_unused simdjson_inline error_code implementation_simdjson_result_b
}
template<typename T>
simdjson_inline error_code implementation_simdjson_result_base<T>::error() const noexcept {
simdjson_warn_unused simdjson_inline error_code implementation_simdjson_result_base<T>::error() const noexcept {
return this->second;
}
template<typename T>
simdjson_inline bool implementation_simdjson_result_base<T>::has_value() const noexcept {
simdjson_warn_unused simdjson_inline bool implementation_simdjson_result_base<T>::has_value() const noexcept {
return this->error() == SUCCESS;
}
@@ -67,17 +67,17 @@ struct implementation_simdjson_result_base {
*
* @param value The variable to assign the value to. May not be set if there is an error.
*/
simdjson_inline error_code get(T &value) && noexcept;
simdjson_warn_unused simdjson_inline error_code get(T &value) && noexcept;
/**
* The error.
*/
simdjson_inline error_code error() const noexcept;
simdjson_warn_unused simdjson_inline error_code error() const noexcept;
/**
* Whether there is a value.
*/
simdjson_inline bool has_value() const noexcept;
simdjson_warn_unused simdjson_inline bool has_value() const noexcept;
#if SIMDJSON_EXCEPTIONS
@@ -138,6 +138,9 @@ struct implementation_simdjson_result_base {
*/
simdjson_inline T&& value_unsafe() && noexcept;
using value_type = T;
using error_type = error_code;
protected:
/** users should never directly access first and second. **/
T first{}; /** Users should never directly access 'first'. **/
+15 -6
View File
@@ -167,7 +167,12 @@ simdjson_inline bool compute_float_64(int64_t power, uint64_t i, bool negative,
// with a returned value of type value128 with a "low component" corresponding to the
// 64-bit least significant bits of the product and with a "high component" corresponding
// to the 64-bit most significant bits of the product.
#if SIMDJSON_STATIC_REFLECTION
simdjson::internal::value128 firstproduct = full_multiplication(i, simdjson::internal::powers_template<>::power_of_five_128[index]);
#else
simdjson::internal::value128 firstproduct = full_multiplication(i, simdjson::internal::power_of_five_128[index]);
#endif
// Both i and power_of_five_128[index] have their most significant bit set to 1 which
// implies that the either the most or the second most significant bit of the product
// is 1. We pack values in this manner for efficiency reasons: it maximizes the use
@@ -200,7 +205,11 @@ simdjson_inline bool compute_float_64(int64_t power, uint64_t i, bool negative,
// with a returned value of type value128 with a "low component" corresponding to the
// 64-bit least significant bits of the product and with a "high component" corresponding
// to the 64-bit most significant bits of the product.
#if SIMDJSON_STATIC_REFLECTION
simdjson::internal::value128 secondproduct = full_multiplication(i, simdjson::internal::powers_template<>::power_of_five_128[index + 1]);
#else
simdjson::internal::value128 secondproduct = full_multiplication(i, simdjson::internal::power_of_five_128[index + 1]);
#endif
firstproduct.low += secondproduct.high;
if(secondproduct.high > firstproduct.low) { firstproduct.high++; }
// As it has been proven by Noble Mushtak and Daniel Lemire in "Fast Number Parsing Without
@@ -361,7 +370,7 @@ simdjson_inline bool is_digit(const uint8_t c) {
return static_cast<uint8_t>(c - '0') <= 9;
}
simdjson_inline error_code parse_decimal_after_separator(simdjson_unused const uint8_t *const src, const uint8_t *&p, uint64_t &i, int64_t &exponent) {
simdjson_warn_unused simdjson_inline error_code parse_decimal_after_separator(simdjson_unused const uint8_t *const src, const uint8_t *&p, uint64_t &i, int64_t &exponent) {
// we continue with the fiction that we have an integer. If the
// floating point number is representable as x * 10^z for some integer
// z that fits in 53 bits, then we will be able to convert back the
@@ -389,7 +398,7 @@ simdjson_inline error_code parse_decimal_after_separator(simdjson_unused const u
return SUCCESS;
}
simdjson_inline error_code parse_exponent(simdjson_unused const uint8_t *const src, const uint8_t *&p, int64_t &exponent) {
simdjson_warn_unused simdjson_inline error_code parse_exponent(simdjson_unused const uint8_t *const src, const uint8_t *&p, int64_t &exponent) {
// Exp Sign: -123.456e[-]78
bool neg_exp = ('-' == *p);
if (neg_exp || '+' == *p) { p++; } // Skip + as well
@@ -478,7 +487,7 @@ static error_code slow_float_parsing(simdjson_unused const uint8_t * src, double
/** @private */
template<typename W>
simdjson_inline error_code write_float(const uint8_t *const src, bool negative, uint64_t i, const uint8_t * start_digits, size_t digit_count, int64_t exponent, W &writer) {
simdjson_warn_unused simdjson_inline error_code write_float(const uint8_t *const src, bool negative, uint64_t i, const uint8_t * start_digits, size_t digit_count, int64_t exponent, W &writer) {
// If we frequently had to deal with long strings of digits,
// we could extend our code by using a 128-bit integer instead
// of a 64-bit integer. However, this is uncommon in practice.
@@ -541,13 +550,13 @@ simdjson_inline error_code write_float(const uint8_t *const src, bool negative,
//
// Our objective is accurate parsing (ULP of 0) at high speed.
template<typename W>
simdjson_inline error_code parse_number(const uint8_t *const src, W &writer);
simdjson_warn_unused simdjson_inline error_code parse_number(const uint8_t *const src, W &writer);
// for performance analysis, it is sometimes useful to skip parsing
#ifdef SIMDJSON_SKIPNUMBERPARSING
template<typename W>
simdjson_inline error_code parse_number(const uint8_t *const, W &writer) {
simdjson_warn_unused simdjson_inline error_code parse_number(const uint8_t *const, W &writer) {
writer.append_s64(0); // always write zero
return SUCCESS; // always succeeds
}
@@ -573,7 +582,7 @@ simdjson_unused simdjson_inline simdjson_result<number_type> get_number_type(con
//
// Our objective is accurate parsing (ULP of 0) at high speed.
template<typename W>
simdjson_inline error_code parse_number(const uint8_t *const src, W &writer) {
simdjson_warn_unused simdjson_inline error_code parse_number(const uint8_t *const src, W &writer) {
//
// Check for minus sign
//
@@ -1,4 +1,4 @@
#if defined(SIMDJSON_CONDITIONAL_INCLUDE) && !defined(SIMDJSON_GENERIC_ONDEMAND_DEPENDENCIES_H)
#if defined(SIMDJSON_CONDITIONAL_INCLUDE) && !defined(SIMDJSON_GENERIC_BUILDER_DEPENDENCIES_H)
#error simdjson/generic/ondemand/dependencies.h must be included before simdjson/generic/ondemand/amalgamated.h!
#endif
@@ -41,13 +41,10 @@
#include "simdjson/generic/ondemand/object_iterator-inl.h"
#include "simdjson/generic/ondemand/parser-inl.h"
#include "simdjson/generic/ondemand/raw_json_string-inl.h"
#include "simdjson/generic/ondemand/serialization-inl.h"
#include "simdjson/generic/ondemand/token_iterator-inl.h"
#include "simdjson/generic/ondemand/value_iterator-inl.h"
#include "simdjson/generic/ondemand/serialization-inl.h"
// JSON builder, ideally they should not be part of the ondemand directory
// but it is convenient for now to have them here.
#include "simdjson/generic/ondemand/json_string_builder.h"
#include "simdjson/generic/ondemand/json_string_builder-inl.h"
#include "simdjson/generic/ondemand/json_builder.h"
// JSON path accessor (compile-time) - must be after inline definitions
#include "simdjson/generic/ondemand/compile_time_accessors.h"
+66 -1
View File
@@ -85,7 +85,7 @@ simdjson_inline simdjson_result<array_iterator> array::begin() noexcept {
simdjson_inline simdjson_result<array_iterator> array::end() noexcept {
return array_iterator(iter);
}
simdjson_inline error_code array::consume() noexcept {
simdjson_warn_unused simdjson_warn_unused simdjson_inline error_code array::consume() noexcept {
auto error = iter.json_iter().skip_child(iter.depth()-1);
if(error) { iter.abandon(); }
return error;
@@ -170,6 +170,67 @@ inline simdjson_result<value> array::at_path(std::string_view json_path) noexcep
return at_pointer(json_pointer);
}
inline simdjson_result<std::vector<value>> array::at_path_with_wildcard(std::string_view json_path) noexcept {
std::vector<value> result;
auto result_pair = get_next_key_and_json_path(json_path);
std::string_view key = result_pair.first;
std::string_view remaining_path = result_pair.second;
// Wildcard case
if(key=="*"){
for(auto element: *this){
if(element.error()){
return element.error();
}
if(remaining_path.empty()){
// Use value_unsafe() because we've already checked for errors above.
// The 'element' is a simdjson_result<value> wrapper, and we need to extract
// the underlying value. value_unsafe() is safe here because error() returned false.
result.push_back(std::move(element).value_unsafe());
}else{
auto nested_result = element.at_path_with_wildcard(remaining_path);
if(nested_result.error()){
return nested_result.error();
}
// Same logic as above.
std::vector<value> nested_matches = std::move(nested_result).value_unsafe();
result.insert(result.end(),
std::make_move_iterator(nested_matches.begin()),
std::make_move_iterator(nested_matches.end()));
}
}
return result;
}else{
// Specific index case in which we access the element at the given index
size_t idx=0;
for(char c:key){
if(c < '0' || c > '9'){
return INVALID_JSON_POINTER;
}
idx = idx*10 + (c - '0');
}
auto element = at(idx);
if(element.error()){
return element.error();
}
if(remaining_path.empty()){
result.push_back(std::move(element).value_unsafe());
return result;
}else{
return element.at_path_with_wildcard(remaining_path);
}
}
}
simdjson_inline simdjson_result<value> array::at(size_t index) noexcept {
size_t i = 0;
for (auto value : *this) {
@@ -228,6 +289,10 @@ simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value> simdj
if (error()) { return error(); }
return first.at_path(json_path);
}
simdjson_inline simdjson_result<std::vector<SIMDJSON_IMPLEMENTATION::ondemand::value>> simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::array>::at_path_with_wildcard(std::string_view json_path) noexcept {
if (error()) { return error(); }
return first.at_path_with_wildcard(json_path);
}
simdjson_inline simdjson_result<std::string_view> simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::array>::raw_json() noexcept {
if (error()) { return error(); }
return first.raw_json();
+15 -4
View File
@@ -5,6 +5,7 @@
#include "simdjson/generic/ondemand/base.h"
#include "simdjson/generic/implementation_simdjson_result_base.h"
#include "simdjson/generic/ondemand/value_iterator.h"
#include <vector>
#endif // SIMDJSON_CONDITIONAL_INCLUDE
namespace simdjson {
@@ -107,7 +108,7 @@ public:
* JSONPath queries that trivially convertible to JSON Pointer queries: key
* names and array indices.
*
* https://datatracker.ietf.org/doc/html/draft-normington-jsonpath-00
* https://www.rfc-editor.org/rfc/rfc9535 (RFC 9535)
*
* @return The value associated with the given JSONPath expression, or:
* - INVALID_JSON_POINTER if the JSONPath to JSON Pointer conversion fails
@@ -117,6 +118,15 @@ public:
*/
inline simdjson_result<value> at_path(std::string_view json_path) noexcept;
/**
* Get all values matching the given JSONPath expression with wildcard support.
* Supports wildcard patterns like "[*]" to match all array elements.
*
* @param json_path JSONPath expression with wildcards
* @return Vector of values matching the wildcard pattern
*/
inline simdjson_result<std::vector<value>> at_path_with_wildcard(std::string_view json_path) noexcept;
/**
* Consumes the array and returns a string_view instance corresponding to the
* array as represented in JSON. It points inside the original document.
@@ -141,7 +151,7 @@ public:
* @returns SUCCESS If the parse succeeded and the out parameter was set to the value.
*/
template <typename T>
simdjson_inline error_code get(T &out)
simdjson_warn_unused simdjson_inline error_code get(T &out)
noexcept(custom_deserializable<T, array> ? nothrow_custom_deserializable<T, array> : true) {
static_assert(custom_deserializable<T, array>);
return deserialize(*this, out);
@@ -166,7 +176,7 @@ protected:
/**
* Go to the end of the array, no matter where you are right now.
*/
simdjson_inline error_code consume() noexcept;
simdjson_warn_unused simdjson_inline error_code consume() noexcept;
/**
* Begin array iteration.
@@ -239,6 +249,7 @@ public:
simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value> at(size_t index) noexcept;
simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value> at_pointer(std::string_view json_pointer) noexcept;
simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value> at_path(std::string_view json_path) noexcept;
simdjson_inline simdjson_result<std::vector<SIMDJSON_IMPLEMENTATION::ondemand::value>> at_path_with_wildcard(std::string_view json_path) noexcept;
simdjson_inline simdjson_result<std::string_view> raw_json() noexcept;
#if SIMDJSON_SUPPORTS_CONCEPTS
// TODO: move this code into object-inl.h
@@ -252,7 +263,7 @@ public:
return first.get<T>();
}
template<typename T>
simdjson_inline error_code get(T& out) noexcept {
simdjson_warn_unused simdjson_inline error_code get(T& out) noexcept {
if (error()) { return error(); }
if constexpr (std::is_same_v<T, SIMDJSON_IMPLEMENTATION::ondemand::array>) {
out = first;
@@ -17,6 +17,10 @@ simdjson_inline array_iterator::array_iterator(const value_iterator &_iter) noex
{}
simdjson_inline simdjson_result<value> array_iterator::operator*() noexcept {
#if SIMDJSON_DEVELOPMENT_CHECKS
SIMDJSON_ASSUME(!has_been_referenced);
has_been_referenced = true;
#endif
if (iter.error()) { iter.abandon(); return iter.error(); }
return value(iter.child());
}
@@ -27,6 +31,9 @@ simdjson_inline bool array_iterator::operator!=(const array_iterator &) const no
return iter.is_open();
}
simdjson_inline array_iterator &array_iterator::operator++() noexcept {
#if SIMDJSON_DEVELOPMENT_CHECKS
has_been_referenced = false;
#endif
error_code error;
// PERF NOTE this is a safety rail ... users should exit loops as soon as they receive an error, so we'll never get here.
// However, it does not seem to make a perf difference, so we add it out of an abundance of caution.
@@ -2,6 +2,7 @@
#ifndef SIMDJSON_CONDITIONAL_INCLUDE
#define SIMDJSON_GENERIC_ONDEMAND_ARRAY_ITERATOR_H
#include <iterator>
#include "simdjson/generic/implementation_simdjson_result_base.h"
#include "simdjson/generic/ondemand/base.h"
#include "simdjson/generic/ondemand/value_iterator.h"
@@ -17,11 +18,17 @@ namespace ondemand {
*
* This is an input_iterator, meaning:
* - It is forward-only
* - * must be called exactly once per element.
* - * must be called at most once per element.
* - ++ must be called exactly once in between each * (*, ++, *, ++, * ...)
*/
class array_iterator {
public:
using iterator_category = std::input_iterator_tag;
using value_type = simdjson_result<value>;
using difference_type = std::ptrdiff_t;
using pointer = void;
using reference = value_type;
/** Create a new, invalid array iterator. */
simdjson_inline array_iterator() noexcept = default;
@@ -65,6 +72,9 @@ public:
simdjson_warn_unused simdjson_inline bool at_end() const noexcept;
private:
#if SIMDJSON_DEVELOPMENT_CHECKS
bool has_been_referenced{false};
#endif
value_iterator iter{};
simdjson_inline array_iterator(const value_iterator &iter) noexcept;
@@ -82,6 +92,12 @@ namespace simdjson {
template<>
struct simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::array_iterator> : public SIMDJSON_IMPLEMENTATION::implementation_simdjson_result_base<SIMDJSON_IMPLEMENTATION::ondemand::array_iterator> {
using iterator_category = std::input_iterator_tag;
using value_type = simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value>;
using difference_type = std::ptrdiff_t;
using pointer = void;
using reference = value_type;
simdjson_inline simdjson_result(SIMDJSON_IMPLEMENTATION::ondemand::array_iterator &&value) noexcept; ///< @private
simdjson_inline simdjson_result(error_code error) noexcept; ///< @private
simdjson_inline simdjson_result() noexcept = default;
@@ -0,0 +1,938 @@
/**
* Compile-time JSON Path and JSON Pointer accessors using C++26 reflection (P2996)
*
* This file validates JSON paths/pointers against struct definitions at compile time
* and generates optimized accessor code with zero runtime overhead.
*
* ## How It Works
*
* **Compile Time**: Path is parsed, validated against struct, types are checked
* **Runtime**: Direct navigation with no parsing or validation overhead
*
* Example:
* ```cpp
* struct User { std::string name; std::vector<std::string> emails; };
*
* std::string email;
* path_accessor<User, ".emails[0]">::extract_field(doc, email);
*
* // Compile time validates:
* // 1. User has "emails" field
* // 2. "emails" is array-like
* // 3. Element type is std::string
* // 4. static_assert(^^std::string == ^^std::string)
*
* // Runtime just navigates:
* // doc.get_object().find_field("emails").get_array().at(0).get(email)
* ```
*
* ## Key Reflection APIs
*
* - `^^Type`: Reflect operator, converts type to std::meta::info
* - `std::meta::nonstatic_data_members_of(type)`: Get all fields of a struct
* - `std::meta::identifier_of(member)`: Get field name as string_view
* - `std::meta::type_of(member)`: Get reflected type of a field
* - `std::meta::is_array_type(type)`: Check if C-style array
* - `std::meta::remove_extent(array)`: Extract element type from array
* - `std::meta::members_of(type)`: Get all members including typedefs
* - `std::meta::is_type(member)`: Check if member is a type (vs field)
*
* All operations execute at compile time in consteval contexts.
*/
#ifndef SIMDJSON_GENERIC_ONDEMAND_COMPILE_TIME_ACCESSORS_H
#ifndef SIMDJSON_CONDITIONAL_INCLUDE
#define SIMDJSON_GENERIC_ONDEMAND_COMPILE_TIME_ACCESSORS_H
#endif // SIMDJSON_CONDITIONAL_INCLUDE
// Arguably, we should just check SIMDJSON_STATIC_REFLECTION since it
// is unlikely that we will have reflection support without concepts support.
#if SIMDJSON_SUPPORTS_CONCEPTS && SIMDJSON_STATIC_REFLECTION
#include <string_view>
#include <cstddef>
#include <array>
namespace simdjson {
namespace SIMDJSON_IMPLEMENTATION {
namespace ondemand {
/***
* JSONPath implementation for compile-time access
* RFC 9535 JSONPath: Query Expressions for JSON, https://www.rfc-editor.org/rfc/rfc9535
*/
namespace json_path {
// Note: value type must be fully defined before this header is included
// This is ensured by including this in amalgamated.h after value-inl.h
using ::simdjson::SIMDJSON_IMPLEMENTATION::ondemand::value;
// Path step types
enum class step_type {
field, // .field_name or ["field_name"]
array_index // [index]
};
// Represents a single step in a JSON path
template<std::size_t N>
struct path_step {
step_type type;
char key[N]; // Field name (empty for array indices)
std::size_t index; // Array index (0 for field access)
constexpr path_step(step_type t, const char (&k)[N], std::size_t idx = 0)
: type(t), index(idx) {
for (std::size_t i = 0; i < N; ++i) {
key[i] = k[i];
}
}
constexpr std::string_view key_view() const {
return {key, N - 1};
}
};
// Helper to create field step
template<std::size_t N>
consteval auto make_field_step(const char (&name)[N]) {
return path_step<N>(step_type::field, name, 0);
}
// Helper to create array index step
consteval auto make_index_step(std::size_t idx) {
return path_step<1>(step_type::array_index, "", idx);
}
// Parse state for compile-time JSON path parsing
struct parse_result {
bool success;
std::size_t pos;
std::string_view error_msg;
};
// Compile-time JSON path parser
// Supports subset: .field, ["field"], [index], nested combinations
template<constevalutil::fixed_string Path>
struct json_path_parser {
static constexpr std::string_view path_str = Path.view();
// Skip leading $ if present
static consteval std::size_t skip_root() {
if (!path_str.empty() && path_str[0] == '$') {
return 1;
}
return 0;
}
// Count the number of steps in the path at compile time
static consteval std::size_t count_steps() {
std::size_t count = 0;
std::size_t i = skip_root();
while (i < path_str.size()) {
if (path_str[i] == '.') {
// Field access: .field
++i;
if (i >= path_str.size()) break;
// Skip field name
while (i < path_str.size() && path_str[i] != '.' && path_str[i] != '[') {
++i;
}
++count;
} else if (path_str[i] == '[') {
// Array or bracket notation
++i;
if (i >= path_str.size()) break;
if (path_str[i] == '"' || path_str[i] == '\'') {
// Field access: ["field"] or ['field']
char quote = path_str[i];
++i;
while (i < path_str.size() && path_str[i] != quote) {
++i;
}
if (i < path_str.size()) ++i; // skip closing quote
if (i < path_str.size() && path_str[i] == ']') ++i;
} else {
// Array index: [0], [123]
while (i < path_str.size() && path_str[i] != ']') {
++i;
}
if (i < path_str.size()) ++i; // skip ]
}
++count;
} else {
++i;
}
}
return count;
}
// Parse a field name at compile time
static consteval std::size_t parse_field_name(std::size_t start, char* out, std::size_t max_len) {
std::size_t len = 0;
std::size_t i = start;
while (i < path_str.size() && path_str[i] != '.' && path_str[i] != '[' && len < max_len - 1) {
out[len++] = path_str[i++];
}
out[len] = '\0';
return i;
}
// Parse an array index at compile time
static consteval std::pair<std::size_t, std::size_t> parse_array_index(std::size_t start) {
std::size_t index = 0;
std::size_t i = start;
while (i < path_str.size() && path_str[i] >= '0' && path_str[i] <= '9') {
index = index * 10 + (path_str[i] - '0');
++i;
}
return {i, index};
}
};
// Compile-time path accessor generator
template<typename T, constevalutil::fixed_string Path>
struct path_accessor {
using value = ::simdjson::SIMDJSON_IMPLEMENTATION::ondemand::value;
static constexpr auto parser = json_path_parser<Path>();
static constexpr std::size_t num_steps = parser.count_steps();
static constexpr std::string_view path_view = Path.view();
// Compile-time accessor generation
// If T is a struct, validates the path at compile time
// If T is void, skips validation
template<typename DocOrValue>
static inline simdjson_result<value> access(DocOrValue& doc_or_val) noexcept {
// Validate path at compile time if T is a struct
if constexpr (std::is_class_v<T>) {
constexpr bool path_valid = validate_path();
static_assert(path_valid, "JSON path does not match struct definition");
}
// Parse the path at compile time to build access steps
return access_impl<parser.skip_root()>(doc_or_val.get_value());
}
// Extract value at path directly into target with compile-time type validation
// Example: std::string name; path_accessor<User, ".name">::extract_field(doc, name);
template<typename DocOrValue, typename FieldType>
static inline error_code extract_field(DocOrValue& doc_or_val, FieldType& target) noexcept {
static_assert(std::is_class_v<T>, "extract_field requires T to be a struct type for validation");
// Validate path exists in struct definition
constexpr bool path_valid = validate_path();
static_assert(path_valid, "JSON path does not match struct definition");
// Get the type at the end of the path
constexpr auto final_type = get_final_type();
// Verify target type matches the field type
static_assert(final_type == ^^FieldType, "Target type does not match the field type at the path");
// All validation done at compile time - just navigate and extract
auto json_value = access_impl<parser.skip_root()>(doc_or_val.get_value());
if (json_value.error()) return json_value.error();
return json_value.get(target);
}
private:
// Get the final type by walking the path through the struct type
template<typename U = T>
static consteval std::enable_if_t<std::is_class_v<U>, std::meta::info> get_final_type() {
auto current_type = ^^T;
std::size_t i = parser.skip_root();
while (i < path_view.size()) {
if (path_view[i] == '.') {
// .field syntax
++i;
std::size_t field_start = i;
while (i < path_view.size() && path_view[i] != '.' && path_view[i] != '[') {
++i;
}
std::string_view field_name = path_view.substr(field_start, i - field_start);
auto members = std::meta::nonstatic_data_members_of(
current_type, std::meta::access_context::unchecked()
);
for (auto mem : members) {
if (std::meta::identifier_of(mem) == field_name) {
current_type = std::meta::type_of(mem);
break;
}
}
} else if (path_view[i] == '[') {
++i;
if (i >= path_view.size()) break;
if (path_view[i] == '"' || path_view[i] == '\'') {
// ["field"] syntax
char quote = path_view[i];
++i;
std::size_t field_start = i;
while (i < path_view.size() && path_view[i] != quote) {
++i;
}
std::string_view field_name = path_view.substr(field_start, i - field_start);
if (i < path_view.size()) ++i; // skip quote
if (i < path_view.size() && path_view[i] == ']') ++i;
auto members = std::meta::nonstatic_data_members_of(
current_type, std::meta::access_context::unchecked()
);
for (auto mem : members) {
if (std::meta::identifier_of(mem) == field_name) {
current_type = std::meta::type_of(mem);
break;
}
}
} else {
// [index] syntax - extract element type
while (i < path_view.size() && path_view[i] >= '0' && path_view[i] <= '9') {
++i;
}
if (i < path_view.size() && path_view[i] == ']') ++i;
current_type = get_element_type_reflected(current_type);
}
} else {
++i;
}
}
return current_type;
}
private:
// Walk path and extract directly into final field using compile-time reflection
template<std::meta::info CurrentType, std::size_t PathPos, typename TargetType>
static inline error_code extract_with_reflection(simdjson_result<value> current, TargetType& target_ref) noexcept {
if (current.error()) return current.error();
// Base case: end of path - extract into target
if constexpr (PathPos >= path_view.size()) {
return current.get(target_ref);
}
// Field access: .field_name
else if constexpr (path_view[PathPos] == '.') {
constexpr auto field_info = parse_next_field(PathPos);
constexpr std::string_view field_name = std::get<0>(field_info);
constexpr std::size_t next_pos = std::get<1>(field_info);
constexpr auto member_info = find_member_by_name(CurrentType, field_name);
static_assert(member_info != ^^void, "Field not found in struct");
constexpr auto member_type = std::meta::type_of(member_info);
auto obj_result = current.get_object();
if (obj_result.error()) return obj_result.error();
auto obj = obj_result.value_unsafe();
auto field_value = obj.find_field_unordered(field_name);
if constexpr (next_pos >= path_view.size()) {
return field_value.get(target_ref);
} else {
return extract_with_reflection<member_type, next_pos>(field_value, target_ref);
}
}
// Bracket notation: [index] or ["field"]
else if constexpr (path_view[PathPos] == '[') {
constexpr auto bracket_info = parse_bracket(PathPos);
constexpr bool is_field = std::get<0>(bracket_info);
constexpr std::size_t next_pos = std::get<2>(bracket_info);
if constexpr (is_field) {
constexpr std::string_view field_name = std::get<1>(bracket_info);
constexpr auto member_info = find_member_by_name(CurrentType, field_name);
static_assert(member_info != ^^void, "Field not found in struct");
constexpr auto member_type = std::meta::type_of(member_info);
auto obj_result = current.get_object();
if (obj_result.error()) return obj_result.error();
auto obj = obj_result.value_unsafe();
auto field_value = obj.find_field_unordered(field_name);
if constexpr (next_pos >= path_view.size()) {
return field_value.get(target_ref);
} else {
return extract_with_reflection<member_type, next_pos>(field_value, target_ref);
}
} else {
constexpr std::size_t index = std::get<3>(bracket_info);
constexpr auto elem_type = get_element_type_reflected(CurrentType);
static_assert(elem_type != ^^void, "Could not determine array element type");
auto arr_result = current.get_array();
if (arr_result.error()) return arr_result.error();
auto arr = arr_result.value_unsafe();
auto elem_value = arr.at(index);
if constexpr (next_pos >= path_view.size()) {
return elem_value.get(target_ref);
} else {
return extract_with_reflection<elem_type, next_pos>(elem_value, target_ref);
}
}
}
// Skip unexpected characters and continue
else {
return extract_with_reflection<CurrentType, PathPos + 1>(current, target_ref);
}
}
// Find member by name in reflected type
static consteval std::meta::info find_member_by_name(std::meta::info type_refl, std::string_view name) {
auto members = std::meta::nonstatic_data_members_of(type_refl, std::meta::access_context::unchecked());
for (auto mem : members) {
if (std::meta::identifier_of(mem) == name) {
return mem;
}
}
}
// Generate compile-time accessor code by walking the path
template<std::size_t PathPos>
static inline simdjson_result<value> access_impl(simdjson_result<value> current) noexcept {
if (current.error()) return current;
if constexpr (PathPos >= path_view.size()) {
return current;
} else if constexpr (path_view[PathPos] == '.') {
constexpr auto field_info = parse_next_field(PathPos);
constexpr std::string_view field_name = std::get<0>(field_info);
constexpr std::size_t next_pos = std::get<1>(field_info);
auto obj_result = current.get_object();
if (obj_result.error()) return obj_result.error();
auto obj = obj_result.value_unsafe();
auto next_value = obj.find_field_unordered(field_name);
return access_impl<next_pos>(next_value);
} else if constexpr (path_view[PathPos] == '[') {
constexpr auto bracket_info = parse_bracket(PathPos);
constexpr bool is_field = std::get<0>(bracket_info);
constexpr std::size_t next_pos = std::get<2>(bracket_info);
if constexpr (is_field) {
constexpr std::string_view field_name = std::get<1>(bracket_info);
auto obj_result = current.get_object();
if (obj_result.error()) return obj_result.error();
auto obj = obj_result.value_unsafe();
auto next_value = obj.find_field_unordered(field_name);
return access_impl<next_pos>(next_value);
} else {
constexpr std::size_t index = std::get<3>(bracket_info);
auto arr_result = current.get_array();
if (arr_result.error()) return arr_result.error();
auto arr = arr_result.value_unsafe();
auto next_value = arr.at(index);
return access_impl<next_pos>(next_value);
}
} else {
return access_impl<PathPos + 1>(current);
}
}
// Parse next field name
static consteval auto parse_next_field(std::size_t start) {
std::size_t i = start + 1;
std::size_t field_start = i;
while (i < path_view.size() && path_view[i] != '.' && path_view[i] != '[') {
++i;
}
std::string_view field_name = path_view.substr(field_start, i - field_start);
return std::make_tuple(field_name, i);
}
// Parse bracket notation: returns (is_field, field_name, next_pos, index)
static consteval auto parse_bracket(std::size_t start) {
std::size_t i = start + 1; // skip '['
if (i < path_view.size() && (path_view[i] == '"' || path_view[i] == '\'')) {
// Field access
char quote = path_view[i];
++i;
std::size_t field_start = i;
while (i < path_view.size() && path_view[i] != quote) {
++i;
}
std::string_view field_name = path_view.substr(field_start, i - field_start);
if (i < path_view.size()) ++i; // skip closing quote
if (i < path_view.size() && path_view[i] == ']') ++i;
return std::make_tuple(true, field_name, i, std::size_t(0));
} else {
// Array index
std::size_t index = 0;
while (i < path_view.size() && path_view[i] >= '0' && path_view[i] <= '9') {
index = index * 10 + (path_view[i] - '0');
++i;
}
if (i < path_view.size() && path_view[i] == ']') ++i;
return std::make_tuple(false, std::string_view{}, i, index);
}
}
public:
// Check if reflected type is array-like (C-style array or indexable container)
// Uses reflection to test: 1) std::meta::is_array_type() for C arrays
// 2) std::meta::substitute() to test concepts::indexable_container concept
static consteval bool is_array_like_reflected(std::meta::info type_reflection) {
if (std::meta::is_array_type(type_reflection)) {
return true;
}
if (std::meta::can_substitute(^^concepts::indexable_container_v, {type_reflection})) {
return std::meta::extract<bool>(std::meta::substitute(^^concepts::indexable_container_v, {type_reflection}));
}
return false;
}
// Extract element type from reflected array or container
// For C arrays: uses std::meta::remove_extent()
// For containers: finds value_type member using std::meta::members_of()
static consteval std::meta::info get_element_type_reflected(std::meta::info type_reflection) {
if (std::meta::is_array_type(type_reflection)) {
return std::meta::remove_extent(type_reflection);
}
auto members = std::meta::members_of(type_reflection, std::meta::access_context::unchecked());
for (auto mem : members) {
if (std::meta::is_type(mem)) {
auto name = std::meta::identifier_of(mem);
if (name == "value_type") {
return mem;
}
}
}
return ^^void;
}
private:
// Check if type has member with given name
template<typename Type>
static consteval bool has_member(std::string_view member_name) {
constexpr auto members = std::meta::nonstatic_data_members_of(^^Type, std::meta::access_context::unchecked());
for (auto mem : members) {
if (std::meta::identifier_of(mem) == member_name) {
return true;
}
}
return false;
}
// Get type of member by name
template<typename Type>
static consteval auto get_member_type(std::string_view member_name) {
constexpr auto members = std::meta::nonstatic_data_members_of(^^Type, std::meta::access_context::unchecked());
for (auto mem : members) {
if (std::meta::identifier_of(mem) == member_name) {
return std::meta::type_of(mem);
}
}
return ^^void;
}
// Check if non-reflected type is array-like
template<typename Type>
static consteval bool is_container_type() {
using BaseType = std::remove_cvref_t<Type>;
if constexpr (requires { typename BaseType::value_type; }) {
return true;
}
if constexpr (std::is_array_v<BaseType>) {
return true;
}
return false;
}
// Extract element type from non-reflected container
template<typename Type>
using extract_element_type = std::conditional_t<
requires { typename std::remove_cvref_t<Type>::value_type; },
typename std::remove_cvref_t<Type>::value_type,
std::conditional_t<
std::is_array_v<std::remove_cvref_t<Type>>,
std::remove_extent_t<std::remove_cvref_t<Type>>,
void
>
>;
// Validate path matches struct definition
static consteval bool validate_path() {
if constexpr (!std::is_class_v<T>) {
return true;
}
auto current_type = ^^T;
std::size_t i = parser.skip_root();
while (i < path_view.size()) {
if (path_view[i] == '.') {
++i;
std::size_t field_start = i;
while (i < path_view.size() && path_view[i] != '.' && path_view[i] != '[') {
++i;
}
std::string_view field_name = path_view.substr(field_start, i - field_start);
bool found = false;
auto members = std::meta::nonstatic_data_members_of(current_type, std::meta::access_context::unchecked());
for (auto mem : members) {
if (std::meta::identifier_of(mem) == field_name) {
current_type = std::meta::type_of(mem);
found = true;
break;
}
}
if (!found) {
return false;
}
} else if (path_view[i] == '[') {
++i;
if (i >= path_view.size()) return false;
if (path_view[i] == '"' || path_view[i] == '\'') {
char quote = path_view[i];
++i;
std::size_t field_start = i;
while (i < path_view.size() && path_view[i] != quote) {
++i;
}
std::string_view field_name = path_view.substr(field_start, i - field_start);
if (i < path_view.size()) ++i;
if (i < path_view.size() && path_view[i] == ']') ++i;
bool found = false;
auto members = std::meta::nonstatic_data_members_of(current_type, std::meta::access_context::unchecked());
for (auto mem : members) {
if (std::meta::identifier_of(mem) == field_name) {
current_type = std::meta::type_of(mem);
found = true;
break;
}
}
if (!found) {
return false;
}
} else {
while (i < path_view.size() && path_view[i] >= '0' && path_view[i] <= '9') {
++i;
}
if (i < path_view.size() && path_view[i] == ']') ++i;
if (!is_array_like_reflected(current_type)) {
return false;
}
auto new_type = get_element_type_reflected(current_type);
if (new_type == ^^void) {
return false;
}
current_type = new_type;
}
} else {
++i;
}
}
return true;
}
};
// Compile-time path accessor with validation
template<typename T, constevalutil::fixed_string Path, typename DocOrValue>
inline simdjson_result<::simdjson::SIMDJSON_IMPLEMENTATION::ondemand::value> at_path_compiled(DocOrValue& doc_or_val) noexcept {
using accessor = path_accessor<T, Path>;
return accessor::access(doc_or_val);
}
// Overload without type parameter (no validation)
template<constevalutil::fixed_string Path, typename DocOrValue>
inline simdjson_result<::simdjson::SIMDJSON_IMPLEMENTATION::ondemand::value> at_path_compiled(DocOrValue& doc_or_val) noexcept {
using accessor = path_accessor<void, Path>;
return accessor::access(doc_or_val);
}
// ============================================================================
// JSON Pointer Compile-Time Support (RFC 6901)
// ============================================================================
// JSON Pointer parser: /field/0/nested (slash-separated)
template<constevalutil::fixed_string Pointer>
struct json_pointer_parser {
static constexpr std::string_view pointer_str = Pointer.view();
// Unescape token: ~0 -> ~, ~1 -> /
static consteval void unescape_token(std::string_view src, char* dest, std::size_t& out_len) {
out_len = 0;
for (std::size_t i = 0; i < src.size(); ++i) {
if (src[i] == '~' && i + 1 < src.size()) {
if (src[i + 1] == '0') {
dest[out_len++] = '~';
++i;
} else if (src[i + 1] == '1') {
dest[out_len++] = '/';
++i;
} else {
dest[out_len++] = src[i];
}
} else {
dest[out_len++] = src[i];
}
}
}
// Check if token is numeric
static consteval bool is_numeric(std::string_view token) {
if (token.empty()) return false;
if (token[0] == '0' && token.size() > 1) return false;
for (char c : token) {
if (c < '0' || c > '9') return false;
}
return true;
}
// Parse numeric token to index
static consteval std::size_t parse_index(std::string_view token) {
std::size_t result = 0;
for (char c : token) {
result = result * 10 + (c - '0');
}
return result;
}
// Count tokens in pointer
static consteval std::size_t count_tokens() {
if (pointer_str.empty() || pointer_str == "/") return 0;
std::size_t count = 0;
std::size_t pos = pointer_str[0] == '/' ? 1 : 0;
while (pos < pointer_str.size()) {
++count;
std::size_t next_slash = pointer_str.find('/', pos);
if (next_slash == std::string_view::npos) break;
pos = next_slash + 1;
}
return count;
}
// Get Nth token
static consteval std::string_view get_token(std::size_t token_index) {
std::size_t pos = pointer_str[0] == '/' ? 1 : 0;
std::size_t current_token = 0;
while (current_token < token_index) {
std::size_t next_slash = pointer_str.find('/', pos);
pos = next_slash + 1;
++current_token;
}
std::size_t token_end = pointer_str.find('/', pos);
if (token_end == std::string_view::npos) token_end = pointer_str.size();
return pointer_str.substr(pos, token_end - pos);
}
};
// JSON Pointer accessor
template<typename T, constevalutil::fixed_string Pointer>
struct pointer_accessor {
using parser = json_pointer_parser<Pointer>;
static constexpr std::string_view pointer_view = Pointer.view();
static constexpr std::size_t token_count = parser::count_tokens();
// Validate pointer against struct definition
static consteval bool validate_pointer() {
if constexpr (!std::is_class_v<T>) {
return true;
}
auto current_type = ^^T;
std::size_t pos = pointer_view[0] == '/' ? 1 : 0;
while (pos < pointer_view.size()) {
// Extract token up to next /
std::size_t token_end = pointer_view.find('/', pos);
if (token_end == std::string_view::npos) token_end = pointer_view.size();
std::string_view token = pointer_view.substr(pos, token_end - pos);
if (parser::is_numeric(token)) {
if (!path_accessor<T, Pointer>::is_array_like_reflected(current_type)) {
return false;
}
current_type = path_accessor<T, Pointer>::get_element_type_reflected(current_type);
} else {
bool found = false;
auto members = std::meta::nonstatic_data_members_of(current_type, std::meta::access_context::unchecked());
for (auto mem : members) {
if (std::meta::identifier_of(mem) == token) {
current_type = std::meta::type_of(mem);
found = true;
break;
}
}
if (!found) return false;
}
pos = token_end + 1;
}
return true;
}
// Recursive accessor
template<std::size_t TokenIndex>
static inline simdjson_result<value> access_impl(simdjson_result<value> current) noexcept {
if constexpr (TokenIndex >= token_count) {
return current;
} else {
constexpr std::string_view token = parser::get_token(TokenIndex);
if constexpr (parser::is_numeric(token)) {
constexpr std::size_t index = parser::parse_index(token);
auto arr = current.get_array().value_unsafe();
auto next_value = arr.at(index);
return access_impl<TokenIndex + 1>(next_value);
} else {
auto obj = current.get_object().value_unsafe();
auto next_value = obj.find_field_unordered(token);
return access_impl<TokenIndex + 1>(next_value);
}
}
}
// Access JSON value at pointer
template<typename DocOrValue>
static inline simdjson_result<value> access(DocOrValue& doc_or_val) noexcept {
if constexpr (std::is_class_v<T>) {
constexpr bool pointer_valid = validate_pointer();
static_assert(pointer_valid, "JSON Pointer does not match struct definition");
}
if (pointer_view.empty() || pointer_view == "/") {
if constexpr (requires { doc_or_val.get_value(); }) {
return doc_or_val.get_value();
} else {
return doc_or_val;
}
}
simdjson_result<value> current = doc_or_val.get_value();
return access_impl<0>(current);
}
// Extract value at pointer directly into target with type validation
template<typename DocOrValue, typename FieldType>
static inline error_code extract_field(DocOrValue& doc_or_val, FieldType& target) noexcept {
static_assert(std::is_class_v<T>, "extract_field requires T to be a struct type for validation");
constexpr bool pointer_valid = validate_pointer();
static_assert(pointer_valid, "JSON Pointer does not match struct definition");
constexpr auto final_type = get_final_type();
static_assert(final_type == ^^FieldType, "Target type does not match the field type at the pointer");
simdjson_result<value> current_value = doc_or_val.get_value();
auto json_value = access_impl<0>(current_value);
if (json_value.error()) return json_value.error();
return json_value.get(target);
}
private:
// Get final type by walking pointer through struct
template<typename U = T>
static consteval std::enable_if_t<std::is_class_v<U>, std::meta::info> get_final_type() {
auto current_type = ^^T;
std::size_t pos = pointer_view[0] == '/' ? 1 : 0;
while (pos < pointer_view.size()) {
std::size_t token_end = pointer_view.find('/', pos);
if (token_end == std::string_view::npos) token_end = pointer_view.size();
std::string_view token = pointer_view.substr(pos, token_end - pos);
if (parser::is_numeric(token)) {
current_type = path_accessor<T, "">::get_element_type_reflected(current_type);
} else {
auto members = std::meta::nonstatic_data_members_of(current_type, std::meta::access_context::unchecked());
for (auto mem : members) {
if (std::meta::identifier_of(mem) == token) {
current_type = std::meta::type_of(mem);
break;
}
}
}
pos = token_end + 1;
}
return current_type;
}
};
// Compile-time JSON Pointer accessor with validation
template<typename T, constevalutil::fixed_string Pointer, typename DocOrValue>
inline simdjson_result<::simdjson::SIMDJSON_IMPLEMENTATION::ondemand::value> at_pointer_compiled(DocOrValue& doc_or_val) noexcept {
using accessor = pointer_accessor<T, Pointer>;
return accessor::access(doc_or_val);
}
// Overload without type parameter (no validation)
template<constevalutil::fixed_string Pointer, typename DocOrValue>
inline simdjson_result<::simdjson::SIMDJSON_IMPLEMENTATION::ondemand::value> at_pointer_compiled(DocOrValue& doc_or_val) noexcept {
using accessor = pointer_accessor<void, Pointer>;
return accessor::access(doc_or_val);
}
} // namespace json_path
} // namespace ondemand
} // namespace SIMDJSON_IMPLEMENTATION
} // namespace simdjson
#endif // SIMDJSON_SUPPORTS_CONCEPTS && SIMDJSON_STATIC_REFLECTION
#endif // SIMDJSON_GENERIC_ONDEMAND_COMPILE_TIME_ACCESSORS_H
@@ -8,7 +8,6 @@
// Internal headers needed for ondemand generics.
// All includes not under simdjson/generic/ondemand must be here!
// Otherwise, amalgamation will fail.
#include "simdjson/concepts.h"
#include "simdjson/dom/base.h" // for MINIMAL_DOCUMENT_CAPACITY
#include "simdjson/implementation.h"
#include "simdjson/padded_string.h"
@@ -7,55 +7,8 @@
#include "simdjson/generic/ondemand/array.h"
#endif // SIMDJSON_CONDITIONAL_INCLUDE
#include <concepts>
namespace simdjson {
namespace tag_invoke_fn_ns {
void tag_invoke();
struct tag_invoke_fn {
template <typename Tag, typename... Args>
requires requires(Tag tag, Args &&...args) {
tag_invoke(std::forward<Tag>(tag), std::forward<Args>(args)...);
}
constexpr auto operator()(Tag tag, Args &&...args) const
noexcept(noexcept(tag_invoke(std::forward<Tag>(tag),
std::forward<Args>(args)...)))
-> decltype(tag_invoke(std::forward<Tag>(tag),
std::forward<Args>(args)...)) {
return tag_invoke(std::forward<Tag>(tag), std::forward<Args>(args)...);
}
};
} // namespace tag_invoke_fn_ns
inline namespace tag_invoke_ns {
inline constexpr tag_invoke_fn_ns::tag_invoke_fn tag_invoke = {};
} // namespace tag_invoke_ns
template <typename Tag, typename... Args>
concept tag_invocable = requires(Tag tag, Args... args) {
tag_invoke(std::forward<Tag>(tag), std::forward<Args>(args)...);
};
template <typename Tag, typename... Args>
concept nothrow_tag_invocable =
tag_invocable<Tag, Args...> && requires(Tag tag, Args... args) {
{
tag_invoke(std::forward<Tag>(tag), std::forward<Args>(args)...)
} noexcept;
};
template <typename Tag, typename... Args>
using tag_invoke_result =
std::invoke_result<decltype(tag_invoke), Tag, Args...>;
template <typename Tag, typename... Args>
using tag_invoke_result_t =
std::invoke_result_t<decltype(tag_invoke), Tag, Args...>;
template <auto &Tag> using tag_t = std::decay_t<decltype(Tag)>;
struct deserialize_tag;
/// These types are deserializable in a built-in way
+121 -24
View File
@@ -141,7 +141,7 @@ simdjson_inline simdjson_result<std::string_view> document::get_string(bool allo
return get_root_value_iterator().get_root_string(true, allow_replacement);
}
template <typename string_type>
simdjson_inline error_code document::get_string(string_type& receiver, bool allow_replacement) noexcept {
simdjson_warn_unused simdjson_inline error_code document::get_string(string_type& receiver, bool allow_replacement) noexcept {
return get_root_value_iterator().get_root_string(receiver, true, allow_replacement);
}
simdjson_inline simdjson_result<std::string_view> document::get_wobbly_string() noexcept {
@@ -167,15 +167,15 @@ template<> simdjson_inline simdjson_result<int64_t> document::get() & noexcept {
template<> simdjson_inline simdjson_result<bool> document::get() & noexcept { return get_bool(); }
template<> simdjson_inline simdjson_result<value> document::get() & noexcept { return get_value(); }
template<> simdjson_inline error_code document::get(array& out) & noexcept { return get_array().get(out); }
template<> simdjson_inline error_code document::get(object& out) & noexcept { return get_object().get(out); }
template<> simdjson_inline error_code document::get(raw_json_string& out) & noexcept { return get_raw_json_string().get(out); }
template<> simdjson_inline error_code document::get(std::string_view& out) & noexcept { return get_string(false).get(out); }
template<> simdjson_inline error_code document::get(double& out) & noexcept { return get_double().get(out); }
template<> simdjson_inline error_code document::get(uint64_t& out) & noexcept { return get_uint64().get(out); }
template<> simdjson_inline error_code document::get(int64_t& out) & noexcept { return get_int64().get(out); }
template<> simdjson_inline error_code document::get(bool& out) & noexcept { return get_bool().get(out); }
template<> simdjson_inline error_code document::get(value& out) & noexcept { return get_value().get(out); }
template<> simdjson_warn_unused simdjson_inline error_code document::get(array& out) & noexcept { return get_array().get(out); }
template<> simdjson_warn_unused simdjson_inline error_code document::get(object& out) & noexcept { return get_object().get(out); }
template<> simdjson_warn_unused simdjson_inline error_code document::get(raw_json_string& out) & noexcept { return get_raw_json_string().get(out); }
template<> simdjson_warn_unused simdjson_inline error_code document::get(std::string_view& out) & noexcept { return get_string(false).get(out); }
template<> simdjson_warn_unused simdjson_inline error_code document::get(double& out) & noexcept { return get_double().get(out); }
template<> simdjson_warn_unused simdjson_inline error_code document::get(uint64_t& out) & noexcept { return get_uint64().get(out); }
template<> simdjson_warn_unused simdjson_inline error_code document::get(int64_t& out) & noexcept { return get_int64().get(out); }
template<> simdjson_warn_unused simdjson_inline error_code document::get(bool& out) & noexcept { return get_bool().get(out); }
template<> simdjson_warn_unused simdjson_inline error_code document::get(value& out) & noexcept { return get_value().get(out); }
template<> simdjson_deprecated simdjson_inline simdjson_result<raw_json_string> document::get() && noexcept { return get_raw_json_string(); }
template<> simdjson_deprecated simdjson_inline simdjson_result<std::string_view> document::get() && noexcept { return get_string(false); }
@@ -245,7 +245,7 @@ simdjson_inline simdjson_result<value> document::operator[](const char *key) & n
return start_or_resume_object()[key];
}
simdjson_inline error_code document::consume() noexcept {
simdjson_warn_unused simdjson_inline error_code document::consume() noexcept {
bool scalar = false;
auto error = is_scalar().get(scalar);
if(error) { return error; }
@@ -347,6 +347,69 @@ simdjson_inline simdjson_result<value> document::at_path(std::string_view json_p
}
}
simdjson_inline simdjson_result<std::vector<value>> document::at_path_with_wildcard(std::string_view json_path) noexcept {
rewind(); // Rewind the document each time at_path_with_wildcard is called
if (json_path.empty()) {
return INVALID_JSON_POINTER;
}
json_type t;
SIMDJSON_TRY(type().get(t));
switch (t) {
case json_type::array:
return (*this).get_array().at_path_with_wildcard(json_path);
case json_type::object:
return (*this).get_object().at_path_with_wildcard(json_path);
default:
return INVALID_JSON_POINTER;
}
}
#if SIMDJSON_SUPPORTS_CONCEPTS && SIMDJSON_STATIC_REFLECTION
template<constevalutil::fixed_string... FieldNames, typename T>
requires(std::is_class_v<T> && (sizeof...(FieldNames) > 0))
simdjson_warn_unused simdjson_inline error_code document::extract_into(T& out) & noexcept {
// Helper to check if a field name matches any of the requested fields
auto should_extract = [](std::string_view field_name) constexpr -> bool {
return ((FieldNames.view() == field_name) || ...);
};
// Iterate through all members of T using reflection
template for (constexpr auto mem : std::define_static_array(
std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) {
if constexpr (!std::meta::is_const(mem) && std::meta::is_public(mem)) {
constexpr std::string_view key = std::define_static_string(std::meta::identifier_of(mem));
// Only extract this field if it's in our list of requested fields
if constexpr (should_extract(key)) {
// Try to find and extract the field
if constexpr (concepts::optional_type<decltype(out.[:mem:])>) {
// For optional fields, it's ok if they're missing
auto field_result = find_field_unordered(key);
if (!field_result.error()) {
auto error = field_result.get(out.[:mem:]);
if (error && error != NO_SUCH_FIELD) {
return error;
}
} else if (field_result.error() != NO_SUCH_FIELD) {
return field_result.error();
} else {
out.[:mem:].reset();
}
} else {
// For required fields (in the requested list), fail if missing
SIMDJSON_TRY((*this)[key].get(out.[:mem:]));
}
}
}
};
return SUCCESS;
}
#endif // SIMDJSON_SUPPORTS_CONCEPTS && SIMDJSON_STATIC_REFLECTION
} // namespace ondemand
} // namespace SIMDJSON_IMPLEMENTATION
} // namespace simdjson
@@ -454,7 +517,7 @@ simdjson_inline simdjson_result<std::string_view> simdjson_result<SIMDJSON_IMPLE
return first.get_string(allow_replacement);
}
template <typename string_type>
simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document>::get_string(string_type& receiver, bool allow_replacement) noexcept {
simdjson_warn_unused simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document>::get_string(string_type& receiver, bool allow_replacement) noexcept {
if (error()) { return error(); }
return first.get_string(receiver, allow_replacement);
}
@@ -490,12 +553,12 @@ simdjson_deprecated simdjson_inline simdjson_result<T> simdjson_result<SIMDJSON_
return std::forward<SIMDJSON_IMPLEMENTATION::ondemand::document>(first).get<T>();
}
template<typename T>
simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document>::get(T &out) & noexcept {
simdjson_warn_unused simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document>::get(T &out) & noexcept {
if (error()) { return error(); }
return first.get<T>(out);
}
template<typename T>
simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document>::get(T &out) && noexcept {
simdjson_warn_unused simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document>::get(T &out) && noexcept {
if (error()) { return error(); }
return std::forward<SIMDJSON_IMPLEMENTATION::ondemand::document>(first).get<T>(out);
}
@@ -505,8 +568,8 @@ template<> simdjson_deprecated simdjson_inline simdjson_result<SIMDJSON_IMPLEMEN
if (error()) { return error(); }
return std::forward<SIMDJSON_IMPLEMENTATION::ondemand::document>(first);
}
template<> simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document>::get<SIMDJSON_IMPLEMENTATION::ondemand::document>(SIMDJSON_IMPLEMENTATION::ondemand::document &out) & noexcept = delete;
template<> simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document>::get<SIMDJSON_IMPLEMENTATION::ondemand::document>(SIMDJSON_IMPLEMENTATION::ondemand::document &out) && noexcept {
template<> simdjson_warn_unused simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document>::get<SIMDJSON_IMPLEMENTATION::ondemand::document>(SIMDJSON_IMPLEMENTATION::ondemand::document &out) & noexcept = delete;
template<> simdjson_warn_unused simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document>::get<SIMDJSON_IMPLEMENTATION::ondemand::document>(SIMDJSON_IMPLEMENTATION::ondemand::document &out) && noexcept {
if (error()) { return error(); }
out = std::forward<SIMDJSON_IMPLEMENTATION::ondemand::document>(first);
return SUCCESS;
@@ -624,6 +687,20 @@ simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value> simdjs
return first.at_path(json_path);
}
simdjson_inline simdjson_result<std::vector<SIMDJSON_IMPLEMENTATION::ondemand::value>> simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document>::at_path_with_wildcard(std::string_view json_path) noexcept {
if (error()) { return error(); }
return first.at_path_with_wildcard(json_path);
}
#if SIMDJSON_STATIC_REFLECTION
template<constevalutil::fixed_string... FieldNames, typename T>
requires(std::is_class_v<T> && (sizeof...(FieldNames) > 0))
simdjson_warn_unused simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document>::extract_into(T& out) & noexcept {
if (error()) { return error(); }
return first.extract_into<FieldNames...>(out);
}
#endif // SIMDJSON_STATIC_REFLECTION
} // namespace simdjson
@@ -658,7 +735,7 @@ simdjson_inline simdjson_result<double> document_reference::get_double() noexcep
simdjson_inline simdjson_result<double> document_reference::get_double_in_string() noexcept { return doc->get_root_value_iterator().get_root_double(false); }
simdjson_inline simdjson_result<std::string_view> document_reference::get_string(bool allow_replacement) noexcept { return doc->get_root_value_iterator().get_root_string(false, allow_replacement); }
template <typename string_type>
simdjson_inline error_code document_reference::get_string(string_type& receiver, bool allow_replacement) noexcept { return doc->get_root_value_iterator().get_root_string(receiver, false, allow_replacement); }
simdjson_warn_unused simdjson_inline error_code document_reference::get_string(string_type& receiver, bool allow_replacement) noexcept { return doc->get_root_value_iterator().get_root_string(receiver, false, allow_replacement); }
simdjson_inline simdjson_result<std::string_view> document_reference::get_wobbly_string() noexcept { return doc->get_root_value_iterator().get_root_wobbly_string(false); }
simdjson_inline simdjson_result<raw_json_string> document_reference::get_raw_json_string() noexcept { return doc->get_root_value_iterator().get_root_raw_json_string(false); }
simdjson_inline simdjson_result<bool> document_reference::get_bool() noexcept { return doc->get_root_value_iterator().get_root_bool(false); }
@@ -709,9 +786,16 @@ simdjson_inline simdjson_result<number> document_reference::get_number() noexcep
simdjson_inline simdjson_result<std::string_view> document_reference::raw_json_token() noexcept { return doc->raw_json_token(); }
simdjson_inline simdjson_result<value> document_reference::at_pointer(std::string_view json_pointer) noexcept { return doc->at_pointer(json_pointer); }
simdjson_inline simdjson_result<value> document_reference::at_path(std::string_view json_path) noexcept { return doc->at_path(json_path); }
simdjson_inline simdjson_result<std::vector<value>> document_reference::at_path_with_wildcard(std::string_view json_path) noexcept { return doc->at_path_with_wildcard(json_path); }
simdjson_inline simdjson_result<std::string_view> document_reference::raw_json() noexcept { return doc->raw_json();}
simdjson_inline document_reference::operator document&() const noexcept { return *doc; }
#if SIMDJSON_SUPPORTS_CONCEPTS && SIMDJSON_STATIC_REFLECTION
template<constevalutil::fixed_string... FieldNames, typename T>
requires(std::is_class_v<T> && (sizeof...(FieldNames) > 0))
simdjson_warn_unused simdjson_inline error_code document_reference::extract_into(T& out) & noexcept {
return doc->extract_into<FieldNames...>(out);
}
#endif // SIMDJSON_SUPPORTS_CONCEPTS && SIMDJSON_STATIC_REFLECTION
} // namespace ondemand
} // namespace SIMDJSON_IMPLEMENTATION
} // namespace simdjson
@@ -808,7 +892,7 @@ simdjson_inline simdjson_result<std::string_view> simdjson_result<SIMDJSON_IMPLE
return first.get_string(allow_replacement);
}
template <typename string_type>
simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document_reference>::get_string(string_type& receiver, bool allow_replacement) noexcept {
simdjson_warn_unused simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document_reference>::get_string(string_type& receiver, bool allow_replacement) noexcept {
if (error()) { return error(); }
return first.get_string(receiver, allow_replacement);
}
@@ -843,12 +927,12 @@ simdjson_inline simdjson_result<T> simdjson_result<SIMDJSON_IMPLEMENTATION::onde
return std::forward<SIMDJSON_IMPLEMENTATION::ondemand::document_reference>(first).get<T>();
}
template <class T>
simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document_reference>::get(T &out) & noexcept {
simdjson_warn_unused simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document_reference>::get(T &out) & noexcept {
if (error()) { return error(); }
return first.get<T>(out);
}
template <class T>
simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document_reference>::get(T &out) && noexcept {
simdjson_warn_unused simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document_reference>::get(T &out) && noexcept {
if (error()) { return error(); }
return std::forward<SIMDJSON_IMPLEMENTATION::ondemand::document_reference>(first).get<T>(out);
}
@@ -865,13 +949,13 @@ simdjson_inline simdjson_result<bool> simdjson_result<SIMDJSON_IMPLEMENTATION::o
return first.is_string();
}
template <>
simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document_reference>::get(SIMDJSON_IMPLEMENTATION::ondemand::document_reference &out) & noexcept {
simdjson_warn_unused simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document_reference>::get(SIMDJSON_IMPLEMENTATION::ondemand::document_reference &out) & noexcept {
if (error()) { return error(); }
out = first;
return SUCCESS;
}
template <>
simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document_reference>::get(SIMDJSON_IMPLEMENTATION::ondemand::document_reference &out) && noexcept {
simdjson_warn_unused simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document_reference>::get(SIMDJSON_IMPLEMENTATION::ondemand::document_reference &out) && noexcept {
if (error()) { return error(); }
out = first;
return SUCCESS;
@@ -959,7 +1043,20 @@ simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value> simdjs
}
return first.at_path(json_path);
}
simdjson_inline simdjson_result<std::vector<SIMDJSON_IMPLEMENTATION::ondemand::value>> simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document_reference>::at_path_with_wildcard(std::string_view json_path) noexcept {
if (error()) {
return error();
}
return first.at_path_with_wildcard(json_path);
}
#if SIMDJSON_STATIC_REFLECTION
template<constevalutil::fixed_string... FieldNames, typename T>
requires(std::is_class_v<T> && (sizeof...(FieldNames) > 0))
simdjson_warn_unused simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::document_reference>::extract_into(T& out) & noexcept {
if (error()) { return error(); }
return first.extract_into<FieldNames...>(out);
}
#endif // SIMDJSON_STATIC_REFLECTION
} // namespace simdjson
#endif // SIMDJSON_GENERIC_ONDEMAND_DOCUMENT_INL_H
+80 -11
View File
@@ -6,6 +6,7 @@
#include "simdjson/generic/ondemand/json_iterator.h"
#include "simdjson/generic/ondemand/deserialize.h"
#include "simdjson/generic/ondemand/value.h"
#include <vector>
#endif // SIMDJSON_CONDITIONAL_INCLUDE
@@ -117,7 +118,7 @@ public:
* @returns INCORRECT_TYPE if the JSON value is not a string. Otherwise, we return SUCCESS.
*/
template <typename string_type>
simdjson_inline error_code get_string(string_type& receiver, bool allow_replacement = false) noexcept;
simdjson_warn_unused simdjson_inline error_code get_string(string_type& receiver, bool allow_replacement = false) noexcept;
/**
* Cast this JSON value to a string.
*
@@ -228,7 +229,7 @@ public:
* @returns SUCCESS If the parse succeeded and the out parameter was set to the value.
*/
template<typename T>
simdjson_inline error_code get(T &out) &
simdjson_warn_unused simdjson_inline error_code get(T &out) &
#if SIMDJSON_SUPPORTS_CONCEPTS
noexcept(custom_deserializable<T, document> ? nothrow_custom_deserializable<T, document> : true)
#else
@@ -402,12 +403,14 @@ public:
simdjson_inline simdjson_result<array_iterator> end() & noexcept;
/**
* Look up a field by name on an object (order-sensitive).
* Look up a field by name on an object (order-sensitive). By order-sensitive, we mean that
* fields must be accessed in the order they appear in the JSON text (although you can
* skip fields). See find_field_unordered() and operator[] for an order-insensitive version.
*
* The following code reads z, then y, then x, and thus will not retrieve x or y if fed the
* JSON `{ "x": 1, "y": 2, "z": 3 }`:
*
* ```c++
* ```cpp
* simdjson::ondemand::parser parser;
* auto obj = parser.parse(R"( { "x": 1, "y": 2, "z": 3 } )"_padded);
* double z = obj.find_field("z");
@@ -446,7 +449,8 @@ public:
* missing case has a non-cache-friendly bump and lots of extra scanning, especially if the object
* in question is large. The fact that the extra code is there also bumps the executable size.
*
* It is the default, however, because it would be highly surprising (and hard to debug) if the
* We default operator[] on find_field_unordered() for convenience.
* It is the default because it would be highly surprising (and hard to debug) if the
* default behavior failed to look up a field just because it was in the wrong order--and many
* APIs assume this. Therefore, you must be explicit if you want to treat objects as out of order.
*
@@ -701,7 +705,7 @@ public:
* JSONPath queries that trivially convertible to JSON Pointer queries: key
* names and array indices.
*
* https://datatracker.ietf.org/doc/html/draft-normington-jsonpath-00
* https://www.rfc-editor.org/rfc/rfc9535 (RFC 9535)
*
* Key values are matched exactly, without unescaping or Unicode normalization.
* We do a byte-by-byte comparison. E.g.
@@ -719,17 +723,64 @@ public:
*/
simdjson_inline simdjson_result<value> at_path(std::string_view json_path) noexcept;
/**
* Get all values matching the given JSONPath expression with wildcard support.
*
* Supports wildcard patterns like "$.array[*]" or "$.object.*" to match multiple elements.
*
* This method materializes all matching values into a vector.
* The document will be consumed after this call.
*
* @param json_path JSONPath expression with wildcards
* @return Vector of values matching the wildcard pattern, or:
* - INVALID_JSON_POINTER if the JSONPath cannot be parsed
* - NO_SUCH_FIELD if a field does not exist
* - INDEX_OUT_OF_BOUNDS if an array index is out of bounds
* - INCORRECT_TYPE if path traversal encounters wrong type
*/
simdjson_inline simdjson_result<std::vector<value>> at_path_with_wildcard(std::string_view json_path) noexcept;
/**
* Consumes the document and returns a string_view instance corresponding to the
* document as represented in JSON. It points inside the original byte array containing
* the JSON document.
*/
simdjson_inline simdjson_result<std::string_view> raw_json() noexcept;
#if SIMDJSON_STATIC_REFLECTION
/**
* Extract only specific fields from the JSON object into a struct.
*
* This allows selective deserialization of only the fields you need,
* potentially improving performance by skipping unwanted fields.
*
* Example:
* ```cpp
* struct Car {
* std::string make;
* std::string model;
* int year;
* double price;
* };
*
* Car car;
* doc.extract_into<"make", "model">(car);
* // Only 'make' and 'model' fields are extracted from JSON
* ```
*
* @tparam FieldNames Compile-time string literals specifying which fields to extract
* @param out The output struct to populate with selected fields
* @returns SUCCESS on success, or an error code if a required field is missing or has wrong type
*/
template<constevalutil::fixed_string... FieldNames, typename T>
requires(std::is_class_v<T> && (sizeof...(FieldNames) > 0))
simdjson_warn_unused simdjson_inline error_code extract_into(T& out) & noexcept;
#endif // SIMDJSON_STATIC_REFLECTION
protected:
/**
* Consumes the document.
*/
simdjson_inline error_code consume() noexcept;
simdjson_warn_unused simdjson_inline error_code consume() noexcept;
simdjson_inline document(ondemand::json_iterator &&iter) noexcept;
simdjson_inline const uint8_t *text(uint32_t idx) const noexcept;
@@ -782,7 +833,7 @@ public:
simdjson_inline simdjson_result<double> get_double_in_string() noexcept;
simdjson_inline simdjson_result<std::string_view> get_string(bool allow_replacement = false) noexcept;
template <typename string_type>
simdjson_inline error_code get_string(string_type& receiver, bool allow_replacement = false) noexcept;
simdjson_warn_unused simdjson_inline error_code get_string(string_type& receiver, bool allow_replacement = false) noexcept;
simdjson_inline simdjson_result<std::string_view> get_wobbly_string() noexcept;
simdjson_inline simdjson_result<raw_json_string> get_raw_json_string() noexcept;
simdjson_inline simdjson_result<bool> get_bool() noexcept;
@@ -826,7 +877,7 @@ public:
* @returns SUCCESS If the parse succeeded and the out parameter was set to the value.
*/
template<typename T>
simdjson_inline error_code get(T &out) &
simdjson_warn_unused simdjson_inline error_code get(T &out) &
#if SIMDJSON_SUPPORTS_CONCEPTS
noexcept(custom_deserializable<T, document> ? nothrow_custom_deserializable<T, document_reference> : true)
#else
@@ -861,6 +912,11 @@ public:
/** @overload template<typename T> error_code get(T &out) & noexcept */
template<typename T> simdjson_inline error_code get(T &out) && noexcept;
simdjson_inline simdjson_result<std::string_view> raw_json() noexcept;
#if SIMDJSON_STATIC_REFLECTION
template<constevalutil::fixed_string... FieldNames, typename T>
requires(std::is_class_v<T> && (sizeof...(FieldNames) > 0))
simdjson_warn_unused simdjson_inline error_code extract_into(T& out) & noexcept;
#endif // SIMDJSON_STATIC_REFLECTION
simdjson_inline operator document&() const noexcept;
#if SIMDJSON_EXCEPTIONS
template <class T>
@@ -901,6 +957,7 @@ public:
simdjson_inline simdjson_result<std::string_view> raw_json_token() noexcept;
simdjson_inline simdjson_result<value> at_pointer(std::string_view json_pointer) noexcept;
simdjson_inline simdjson_result<value> at_path(std::string_view json_path) noexcept;
simdjson_inline simdjson_result<std::vector<value>> at_path_with_wildcard(std::string_view json_path) noexcept;
private:
document *doc{nullptr};
@@ -929,7 +986,7 @@ public:
simdjson_inline simdjson_result<double> get_double_in_string() noexcept;
simdjson_inline simdjson_result<std::string_view> get_string(bool allow_replacement = false) noexcept;
template <typename string_type>
simdjson_inline error_code get_string(string_type& receiver, bool allow_replacement = false) noexcept;
simdjson_warn_unused simdjson_inline error_code get_string(string_type& receiver, bool allow_replacement = false) noexcept;
simdjson_inline simdjson_result<std::string_view> get_wobbly_string() noexcept;
simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::raw_json_string> get_raw_json_string() noexcept;
simdjson_inline simdjson_result<bool> get_bool() noexcept;
@@ -984,6 +1041,12 @@ public:
simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value> at_pointer(std::string_view json_pointer) noexcept;
simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value> at_path(std::string_view json_path) noexcept;
simdjson_inline simdjson_result<std::vector<SIMDJSON_IMPLEMENTATION::ondemand::value>> at_path_with_wildcard(std::string_view json_path) noexcept;
#if SIMDJSON_STATIC_REFLECTION
template<constevalutil::fixed_string... FieldNames, typename T>
requires(std::is_class_v<T> && (sizeof...(FieldNames) > 0))
simdjson_warn_unused simdjson_inline error_code extract_into(T& out) & noexcept;
#endif // SIMDJSON_STATIC_REFLECTION
};
@@ -1010,7 +1073,7 @@ public:
simdjson_inline simdjson_result<double> get_double_in_string() noexcept;
simdjson_inline simdjson_result<std::string_view> get_string(bool allow_replacement = false) noexcept;
template <typename string_type>
simdjson_inline error_code get_string(string_type& receiver, bool allow_replacement = false) noexcept;
simdjson_warn_unused simdjson_inline error_code get_string(string_type& receiver, bool allow_replacement = false) noexcept;
simdjson_inline simdjson_result<std::string_view> get_wobbly_string() noexcept;
simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::raw_json_string> get_raw_json_string() noexcept;
simdjson_inline simdjson_result<bool> get_bool() noexcept;
@@ -1061,6 +1124,12 @@ public:
simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value> at_pointer(std::string_view json_pointer) noexcept;
simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value> at_path(std::string_view json_path) noexcept;
simdjson_inline simdjson_result<std::vector<SIMDJSON_IMPLEMENTATION::ondemand::value>> at_path_with_wildcard(std::string_view json_path) noexcept;
#if SIMDJSON_STATIC_REFLECTION
template<constevalutil::fixed_string... FieldNames, typename T>
requires(std::is_class_v<T> && (sizeof...(FieldNames) > 0))
simdjson_warn_unused simdjson_inline error_code extract_into(T& out) & noexcept;
#endif // SIMDJSON_STATIC_REFLECTION
};
@@ -312,7 +312,10 @@ inline void document_stream::next_document() noexcept {
// Always set depth=1 at the start of document
doc.iter._depth = 1;
// consume comma if comma separated is allowed
if (allow_comma_separated) { doc.iter.consume_character(','); }
if (allow_comma_separated) {
error_code ignored = doc.iter.consume_character(',');
static_cast<void>(ignored); // ignored on purpose
}
// Resets the string buffer at the beginning, thus invalidating the strings.
doc.iter._string_buf_loc = parser->string_buf.get();
doc.iter._root = doc.iter.position();

Some files were not shown because too many files have changed in this diff Show More