Compare commits

..

53 Commits

Author SHA1 Message Date
Daniel Lemire db91f62bd7 updating single 2024-09-27 09:26:51 -04:00
Daniel Lemire 20cd6e30ee fixing incorrect max/min usage 2024-09-27 09:24:23 -04:00
Daniel Lemire e45d8edbd6 minor tweaks to style 2024-09-26 21:03:50 -04:00
Daniel Lemire 6efb8d81bd Merge branch 'builder_development_branch' into general-madness-simpler 2024-09-26 20:49:28 -04:00
Daniel Lemire f562878afc missing file 2024-09-26 20:38:31 -04:00
Daniel Lemire 60aa25a0c2 adding a comment 2024-09-26 20:35:49 -04:00
Daniel Lemire d01662560d simpler madness 2024-09-26 20:35:01 -04:00
Daniel Lemire a63c77a977 typo 2024-09-26 17:20:44 -04:00
M. Bahoosh bbb3cb80a6 Minimal tag_invokes for STL types 2024-09-26 09:27:38 -10:00
Daniel Lemire 43219e50e0 update CI on the builder_development_branch (no code change) (#2262) 2024-09-25 12:33:04 -04:00
M. Bahoosh 49b9860899 Making tag_invoke a "put" as opposed to a "get" (#2256)
* fix: add tests related to issue 2227 (#2229)

* fix: add tests related to issue 2227

* avoiding name clash

* pedantic fix

* deprecate rvalue get on document

* selectively deprecating

* Fix ndjson spec link (#2234)

* fix ndjson spec link

The link in the readme of parse_many links to a casino spam site

* fix link

* [no-ci] Update README.md

* Make simdjson compile again

* Enable SIMDJSON_SINGLEHEADER=OFF in VS Code

With singleheader on, clangd can't find the right
include files.

* Add missing include directives to static build targets of simdjson. (#2240)

* adding a warning

* adding warning regarding SIMDJSON_BUILD_STATIC_LIB

* release candidate

* pedantic viable size

* Making tag_invoke a feeder instead of a producer

* adding missing undef silencer (#2253)

* Ignore pragma once when amalgamating source files (#2248)

With gcc it causes an error in `simdjson.cpp`:
```
simdjson.cpp:548:9: warning: #pragma once in main file
  548 | #pragma once
      |         ^~~~
```

It had previously been commented out in:
https://github.com/simdjson/simdjson/commit/6ef555e6fb79363fae057a9a46b52cd208d9e305

However, this was lost in an upgrade:
https://github.com/simdjson/simdjson/commit/2a4ff7346813b120f2b5b40e95d69352b593cc9c

* Update CI (#2254)

* adding missing undef silencer

* Updating CI

* more fixes

* fix

* big endian fix

* Moving to the new tag_invoke signature

* Fix nlohmann ambiguity on C++23-enabled clang

* Revert "Merge branch 'master' of https://github.com/simdjson/simdjson into builder_development_branch_extra"

This reverts commit 3eeecbab34, reversing
changes made to 6858b208b4.

---------

Co-authored-by: Daniel Lemire <daniel@lemire.me>
Co-authored-by: Sasha Lopoukhine <superlopuh@gmail.com>
Co-authored-by: John Keiser <john@johnkeiser.com>
Co-authored-by: Tan Li Boon <undisputed-seraphim@users.noreply.github.com>
Co-authored-by: tobil4sk <tobil4sk@outlook.com>
2024-09-24 14:47:58 -04:00
Daniel Lemire 0388d79770 Extending the deserialization code with more defaults + docs (#2233)
* Make custom types easier with some predefined cases + docs

* missing include

* adding Ubuntu 24 CXX 20

* using concepts all the way

* minor tweak

* tiny tweak

* tweaks

* more tweaking

* saving

---------

Co-authored-by: Daniel Lemire <dlemire@lemire.me>
2024-08-11 14:14:35 -04:00
Daniel Lemire c0fa1aec66 fix: correct small issues with deserialize (#2232) 2024-08-09 16:07:08 -04:00
M. Bahoosh 6e61b7f6ef Making tag_invoke to support ondemand::document as well + docs (#2228)
* Making `tag_invoke` to support `ondemand::document` as well + docs

* Fix typos and doc update by @lemire

Co-authored-by: Daniel Lemire <daniel@lemire.me>

* Better docs by @lemire

Co-authored-by: Daniel Lemire <daniel@lemire.me>

* Preserving the old, disallowing in the new

I'm disabling `document::get() &&` if the user has provided a `tag_invoke`d version; otherwise, we retain the compatibility.

---------

Co-authored-by: Daniel Lemire <daniel@lemire.me>
2024-08-09 15:25:52 -04:00
M. Bahoosh a76e778804 tag_invoke based custom types (#2219)
* tag_invoke based custom types

Now you can use tag_invoke to add a custom type or a group of custom types.

* Fixing macro usage + Fixing noexcept

* Fixing the usage of #include

We don't need <concepts> at all seems like it

* Fixing tag_invoke impl for MSVC
2024-08-07 09:16:31 -04:00
Daniel Lemire ccf8694510 v3.10.0 2024-08-01 09:32:54 -04:00
Daniel Lemire 9b67497ed0 Allows field::unescape_key to take in a string parameter + additional dev. checks for string overflow (#2224)
* This PR does the following:

1. Upgrade cxxopts.
2. Allows field::unescape_key to take in a string parameter (syntaxic sugar).
3. Adds a dev. check to detect a string buffer overflow (indicating broken code). Note that this is unrecoverable and indicates bad code.

* tweak
2024-08-01 09:31:50 -04:00
Daniel Lemire 0336684df7 [no-ci] Update basics.md 2024-07-31 11:10:01 -04:00
Daniel Lemire c19320dd6e making it more precise 2024-07-31 10:08:04 -04:00
Daniel Lemire 412a5680e8 update 2024-07-31 10:06:47 -04:00
Daniel Lemire a05a56856d fix: use On-Demand throughout. (#2222) 2024-07-29 15:54:21 -04:00
didarpin 58173a6a1f Added the functionality to convert dom::object and dom::array to dom::element. (#2221)
Co-authored-by: didarpin <didarpin@163.com>
2024-07-25 22:26:20 -04:00
Francisco Geiman Thiesen 1721032cfd Merge pull request #2220 from simdjson/adding_macros_for_cpp20_cpp23
fix: add macros to detected C++20 and C++23
2024-07-25 17:44:36 -07:00
Daniel Lemire b73877f95e Update compiler_check.h 2024-07-23 14:25:27 -04:00
Daniel Lemire 09723897e9 fix: add macros to detected C++20 and C++23 2024-07-23 14:23:00 -04:00
Daniel Lemire 49e231b634 [no-ci] Update README.md 2024-07-16 16:05:32 -04:00
Daniel Lemire acdbbab916 adding an example of value capture with std::string_view (#2216)
* adding an example of value capture with std::string_view

* minor fix

* minor fix
2024-07-16 15:51:43 -04:00
Daniel Lemire 5090247c34 adding test (#2214) 2024-07-13 11:40:13 -04:00
Daniel Lemire 4180e05730 [no-ci] fix 'null_ptr' written as 'null_nullptrptr' in the comments 2024-07-11 08:25:27 -04:00
Daniel Lemire feea2bce2c Create config.yml 2024-07-04 16:48:09 -04:00
Daniel Lemire 692f43cd84 Update standard-issue-template.md 2024-07-04 16:46:34 -04:00
Daniel Lemire 3240d55bcc chore: add RelWithDebInfo to our CI tests (#2209)
* chore: add RelWithDebInfo to our tests

* fix

* chore: add various build types to ubuntu ci
2024-07-04 16:26:22 -04:00
Daniel Lemire 66fd28fc00 Marking a few functions as pure (no side-effect) (#2210)
* marking a few trivial functions as pure

* adding other marks

* additional marks

* vs will issue warnings, so don't use [[gnu::pure]] when __clang__ or __GNUC__ is not defined
2024-07-04 16:26:11 -04:00
Daniel Lemire 5f638951c6 Update basics.md 2024-06-27 14:49:57 -04:00
Daniel Lemire 3e94eea939 adding another example (#2206)
Co-authored-by: Daniel Lemire <dlemire@lemire.me>
2024-06-27 13:22:44 -04:00
Tunghohin 0e8311f812 reset moved object's viable_size to 0 in the move ctor of padded_string (#2204) 2024-06-25 18:06:32 -04:00
Daniel Lemire 3620e9d151 version bump 2024-06-11 15:37:22 -04:00
Daniel Lemire d017cd7ca4 Merge branch 'master' of github.com:simdjson/simdjson 2024-06-11 14:08:20 -04:00
Daniel Lemire d9d1ff5856 chore: remove unneeded gcc13 ci tests 2024-06-11 14:07:56 -04:00
Daniel Lemire eb8f2bce14 fix: add test for issue 2199 (#2200)
* fix: add test for issue 2199

* added ubuntu 24 workflow + silencing a warning
2024-06-11 13:55:55 -04:00
Dirk Stolle 2a4ff73468 update string_view lite to version 1.8.0 (#2197)
This is the header as seen for the tag v1.8.0,
commit a47222b9855dd6e6d1eac38acaa495822e2caa69, on
<https://github.com/martinmoene/string-view-lite>.
2024-06-10 10:29:25 -04:00
Daniel Lemire 77fc2b8447 doc: explaining the page trick (#2196)
* doc: explaining the page trick

* simplify

* did as john said

* trying something else

* flipping order

* trying some other order

* hmmm
2024-06-07 22:12:08 -04:00
Daniel Lemire ba8b66a633 Update basics.md 2024-06-07 12:49:32 -04:00
Daniel Lemire 66eec5feaf Update basics.md 2024-06-05 08:55:12 -04:00
Janeczko Jakub 3964f3e5d2 pull size_t from the std namespace (#2191) 2024-06-03 14:08:23 -04:00
Daniel Lemire ee8515122d version bump 2024-05-30 10:53:40 -04:00
halx99 5d35e7ca1f Fix compile error on llvm-19 (#2187) 2024-05-30 10:52:38 -04:00
spershin deefc88b9c Adding path for parsing incomplete json. (#2189)
1. Allows processing inclomplete, damaged, corrupted json to some extent.
2. Pariity with the Presto Java functionality.
3. Protected with SIMDJSON_EXPERIMENTAL_ALLOW_INCOMPLETE_JSON define.
4. Does not interfere with the normal path (can co-exist).
5. Tested in production forkflow.
2024-05-30 10:52:09 -04:00
Yuriy Chernyshov c80dda7c58 Fix building simdjson against libc++ with _LIBCPP_REMOVE_TRANSITIVE_INCLUDES defined (#2184)
```
src/implementation.cpp:193:20: error: no template named 'is_trivially_destructible' in namespace 'std'; did you mean 'is_trivially_move_constructible'?
static_assert(std::is_trivially_destructible<detect_best_supported_implementation_on_first_use>::value, "detect_best_supported_implementation_on_first_use should be trivially destructible");
              ~~~~~^~~~~~~~~~~~~~~~~~~~~~~~~
                   is_trivially_move_constructible
```
2024-05-23 16:54:23 -04:00
Daniel Lemire d2954ef68b Update ubuntu22-gcc13.yml 2024-05-23 16:53:48 -04:00
Daniel Lemire ac719827ff fix: solve issue 2181 (#2182) 2024-05-11 20:44:38 -04:00
Daniel Lemire 6ea77392a7 Update basics.md 2024-05-10 12:21:26 -04:00
pnck e2f879751c fix: issue #2154 (#2178) 2024-05-10 00:33:09 -04:00
82 changed files with 11161 additions and 4948 deletions
+1
View File
@@ -0,0 +1 @@
blank_issues_enabled: false
@@ -18,7 +18,7 @@ We do not make changes to simdjson without clearly identifiable benefits, which
Is your issue:
1. A bug report? If so, please point at a reproducible test. Indicate whether you are willing or able to provide a bug fix as a pull request.
1. A bug report? If so, please point at a reproducible test. Indicate whether you are willing or able to provide a bug fix as a pull request. As a matter of policy, we do not consider a compiler warning to be a bug.
2. A build issue? If so, provide all possible details regarding your system configuration. If we cannot reproduce your issue, we cannot fix it.
+29
View File
@@ -0,0 +1,29 @@
name: Ubuntu ppc64le (GCC 11)
on:
push:
branches:
- master
pull_request:
branches:
- master
jobs:
build:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: uraimo/run-on-arch-action@v2
name: Test
id: runcmd
with:
arch: aarch64
distro: ubuntu_latest
githubToken: ${{ github.token }}
install: |
apt-get update -q -y
apt-get install -y cmake make g++
run: |
cmake -DSIMDJSON_SANITIZE_UNDEFINED=ON -DSIMDJSON_DEVELOPER_MODE=ON -DSIMDJSON_COMPETITION=OFF -B build
cmake --build build -j=2
ctest --output-on-failure --test-dir build
+1 -1
View File
@@ -17,7 +17,7 @@ jobs:
fuzz-seconds: 600
dry-run: false
- name: Upload Crash
uses: actions/upload-artifact@v1
uses: actions/upload-artifact@v4
if: failure() && steps.build.outcome == 'success'
with:
name: artifacts
@@ -24,7 +24,7 @@ jobs:
echo "no trailing whitespace found, good!"
fi
- name: Archive whitespace patch
uses: actions/upload-artifact@v2
uses: actions/upload-artifact@v4
if: always()
with:
name: whitespace-patch
+3
View File
@@ -20,6 +20,9 @@ jobs:
- msystem: "MINGW64"
install: mingw-w64-x86_64-libxml2 mingw-w64-x86_64-cmake mingw-w64-x86_64-ninja mingw-w64-x86_64-clang
type: Debug
- msystem: "MINGW64"
install: mingw-w64-x86_64-libxml2 mingw-w64-x86_64-cmake mingw-w64-x86_64-ninja mingw-w64-x86_64-clang
type: RelWithDebInfo
env:
CMAKE_GENERATOR: Ninja
+3
View File
@@ -22,6 +22,9 @@ jobs:
- msystem: "MINGW64"
install: mingw-w64-x86_64-cmake mingw-w64-x86_64-ninja mingw-w64-x86_64-gcc
type: Debug
- msystem: "MINGW64"
install: mingw-w64-x86_64-cmake mingw-w64-x86_64-ninja mingw-w64-x86_64-gcc
type: RelWithDebInfo
env:
CMAKE_GENERATOR: Ninja
+1 -1
View File
@@ -24,6 +24,6 @@ jobs:
apt-get update -q -y
apt-get install -y cmake make g++
run: |
cmake -DCMAKE_BUILD_TYPE=Release -B build
cmake -DCMAKE_BUILD_TYPE=Release -DSIMDJSON_DEVELOPER_MODE=ON -DSIMDJSON_COMPETITION=OFF -B build
cmake --build build -j=2
ctest --output-on-failure --test-dir build
+1 -1
View File
@@ -24,6 +24,6 @@ jobs:
apt-get update -q -y
apt-get install -y cmake make g++
run: |
cmake -DCMAKE_BUILD_TYPE=Release -B build
cmake -DCMAKE_BUILD_TYPE=Release -DSIMDJSON_DEVELOPER_MODE=ON -DSIMDJSON_COMPETITION=OFF -B build
cmake --build build -j=2
ctest --output-on-failure --test-dir build
+1 -1
View File
@@ -24,6 +24,6 @@ jobs:
apt-get update -q -y
apt-get install -y cmake make g++
run: |
cmake -DCMAKE_BUILD_TYPE=Release -B build
cmake -DCMAKE_BUILD_TYPE=Release -DSIMDJSON_DEVELOPER_MODE=ON -DSIMDJSON_COMPETITION=OFF -B build
cmake --build build -j=2
ctest --output-on-failure --test-dir build
-23
View File
@@ -1,23 +0,0 @@
name: Ubuntu 22.04 CI (GCC 13)
on: [push, pull_request]
jobs:
ubuntu-build:
if: >-
! contains(toJSON(github.event.commits.*.message), '[skip ci]') &&
! contains(toJSON(github.event.commits.*.message), '[skip github]')
runs-on: ubuntu-22.04
steps:
- uses: actions/checkout@v4
- uses: actions/cache@v4
with:
path: dependencies/.cache
key: ${{ hashFiles('dependencies/CMakeLists.txt') }}
- name: Use cmake
run: |
mkdir build &&
cd build &&
CXX=g++-13 cmake -DSIMDJSON_DEVELOPER_MODE=ON .. &&
cmake --build . &&
ctest --output-on-failure -LE explicitonly -j
+22
View File
@@ -0,0 +1,22 @@
name: Ubuntu 24.04 CI (CXX 20)
on: [push, pull_request]
jobs:
ubuntu-build:
if: >-
! contains(toJSON(github.event.commits.*.message), '[skip ci]') &&
! contains(toJSON(github.event.commits.*.message), '[skip github]')
runs-on: ubuntu-24.04
strategy:
matrix:
cxx: [g++-13, clang++-16]
steps:
- uses: actions/checkout@a5ac7e51b41094c92402da3b24376905380afc29 # v4.1.6
- name: Prepare
run: cmake -DSIMDJSON_CXX_STANDARD=20 -DSIMDJSON_DEVELOPER_MODE=ON -B build
env:
CXX: ${{matrix.cxx}}
- name: Build
run: cmake --build build -j=2
- name: Test
run: ctest --output-on-failure --test-dir build
+25
View File
@@ -0,0 +1,25 @@
name: Ubuntu 24.04 CI
on: [push, pull_request]
jobs:
ubuntu-build:
if: >-
! contains(toJSON(github.event.commits.*.message), '[skip ci]') &&
! contains(toJSON(github.event.commits.*.message), '[skip github]')
runs-on: ubuntu-24.04
strategy:
matrix:
shared: [ON, OFF]
cxx: [g++-13, clang++-16]
sanitizer: [ON, OFF]
build_type: [RelWithDebInfo, Debug, Release]
steps:
- uses: actions/checkout@a5ac7e51b41094c92402da3b24376905380afc29 # v4.1.6
- name: Prepare
run: cmake -DCMAKE_BUILD_TYPE=${{matrix.build_type}} -DSIMDJSON_DEVELOPER_MODE=ON -DSIMDJSON_SANITIZE=${{matrix.sanitizer}} -DBUILD_SHARED_LIBS=${{matrix.shared}} -B build
env:
CXX: ${{matrix.cxx}}
- name: Build
run: cmake --build build -j=2
- name: Test
run: ctest --output-on-failure --test-dir build
+11 -15
View File
@@ -13,10 +13,12 @@ jobs:
fail-fast: false
matrix:
include:
- {gen: Visual Studio 17 2022, arch: Win32, shared: ON}
- {gen: Visual Studio 17 2022, arch: Win32, shared: OFF}
- {gen: Visual Studio 17 2022, arch: x64, shared: ON}
- {gen: Visual Studio 17 2022, arch: x64, shared: OFF}
- {gen: Visual Studio 17 2022, arch: Win32, shared: ON, build_type: Release}
- {gen: Visual Studio 17 2022, arch: Win32, shared: OFF, build_type: Release}
- {gen: Visual Studio 17 2022, arch: x64, shared: ON, build_type: Release}
- {gen: Visual Studio 17 2022, arch: x64, shared: OFF, build_type: Debug}
- {gen: Visual Studio 17 2022, arch: x64, shared: OFF, build_type: Release}
- {gen: Visual Studio 17 2022, arch: x64, shared: OFF, build_type: RelWithDebInfo}
steps:
- name: checkout
uses: actions/checkout@v4
@@ -24,21 +26,15 @@ jobs:
run: |
cmake -G "${{matrix.gen}}" -A ${{matrix.arch}} -DSIMDJSON_DEVELOPER_MODE=ON -DSIMDJSON_COMPETITION=OFF -DBUILD_SHARED_LIBS=${{matrix.shared}} -B build
- name: Build Debug
run: cmake --build build --config Debug --verbose
- name: Build Release
run: cmake --build build --config Release --verbose
- name: Run Release tests
run: cmake --build build --config ${{matrix.build_type}} --verbose
- name: Run tests
run: |
cd build
ctest -C Release -LE explicitonly --output-on-failure
- name: Run Debug tests
run: |
cd build
ctest -C Debug -LE explicitonly --output-on-failure
ctest -C ${{matrix.build_type}} -LE explicitonly --output-on-failure
- name: Install
run: |
cmake --install build --config Release
cmake --install build --config ${{matrix.build_type}}
- name: Test Installation
run: |
cmake -G "${{matrix.gen}}" -A ${{matrix.arch}} -B build_install_test tests/installation_tests/find
cmake --build build_install_test --config Release
cmake --build build_install_test --config ${{matrix.build_type}}
+37
View File
@@ -0,0 +1,37 @@
name: VS17-CLANG-CI
on: [push, pull_request]
jobs:
ci:
if: >-
! contains(toJSON(github.event.commits.*.message), '[skip ci]') &&
! contains(toJSON(github.event.commits.*.message), '[skip github]')
name: windows-vs17
runs-on: windows-latest
strategy:
fail-fast: false
matrix:
include:
- {gen: Visual Studio 17 2022, arch: x64, build_type: Debug}
- {gen: Visual Studio 17 2022, arch: x64, build_type: Release}
- {gen: Visual Studio 17 2022, arch: x64, build_type: RelWithDebInfo}
steps:
- name: checkout
uses: actions/checkout@v4
- name: Configure
run: |
cmake -G "${{matrix.gen}}" -A ${{matrix.arch}} -T ClangCL -DSIMDJSON_DEVELOPER_MODE=ON -DSIMDJSON_COMPETITION=OFF -B build
- name: Build
run: cmake --build build --config ${{matrix.build_type}} --verbose
- name: Run tests
run: |
cd build
ctest -C ${{matrix.build_type}} -LE explicitonly --output-on-failure
- name: Install
run: |
cmake --install build --config ${{matrix.build_type}}
- name: Test Installation
run: |
cmake -G "${{matrix.gen}}" -A ${{matrix.arch}} -B build_install_test tests/installation_tests/find
cmake --build build_install_test --config ${{matrix.build_type}}
+9 -13
View File
@@ -13,29 +13,25 @@ jobs:
fail-fast: false
matrix:
include:
- {gen: Visual Studio 17 2022, arch: x64}
- {gen: Visual Studio 17 2022, arch: x64, build_type: Debug}
- {gen: Visual Studio 17 2022, arch: x64, build_type: Release}
- {gen: Visual Studio 17 2022, arch: x64, build_type: RelWithDebInfo}
steps:
- name: checkout
uses: actions/checkout@v4
- name: Configure
run: |
cmake -G "${{matrix.gen}}" -A ${{matrix.arch}} -T ClangCL -DSIMDJSON_DEVELOPER_MODE=ON -DSIMDJSON_COMPETITION=OFF -B build
- name: Build Debug
run: cmake --build build --config Debug --verbose
- name: Build Release
run: cmake --build build --config Release --verbose
- name: Run Release tests
- name: Build
run: cmake --build build --config ${{matrix.build_type}} --verbose
- name: Run tests
run: |
cd build
ctest -C Release -LE explicitonly --output-on-failure
- name: Run Debug tests
run: |
cd build
ctest -C Debug -LE explicitonly --output-on-failure
ctest -C ${{matrix.build_type}} -LE explicitonly --output-on-failure
- name: Install
run: |
cmake --install build --config Release
cmake --install build --config ${{matrix.build_type}}
- name: Test Installation
run: |
cmake -G "${{matrix.gen}}" -A ${{matrix.arch}} -B build_install_test tests/installation_tests/find
cmake --build build_install_test --config Release
cmake --build build_install_test --config ${{matrix.build_type}}
+18 -1
View File
@@ -109,6 +109,23 @@
"numbers": "cpp",
"semaphore": "cpp",
"stop_token": "cpp",
"cfenv": "cpp"
"cfenv": "cpp",
"format": "cpp",
"xlocmes": "cpp",
"xlocmon": "cpp",
"xlocnum": "cpp",
"xloctime": "cpp",
"xutility": "cpp",
"coroutine": "cpp",
"xfacet": "cpp",
"xhash": "cpp",
"xiosbase": "cpp",
"xlocale": "cpp",
"xlocbuf": "cpp",
"xlocinfo": "cpp",
"xmemory": "cpp",
"xstring": "cpp",
"xtr1common": "cpp",
"xtree": "cpp"
}
}
+3 -3
View File
@@ -3,7 +3,7 @@ cmake_minimum_required(VERSION 3.14)
project(
simdjson
# The version number is modified by tools/release.py
VERSION 3.9.2
VERSION 3.10.0
DESCRIPTION "Parsing gigabytes of JSON per second"
HOMEPAGE_URL "https://simdjson.org/"
LANGUAGES CXX C
@@ -20,8 +20,8 @@ string(
# ---- Options, variables ----
# These version numbers are modified by tools/release.py
set(SIMDJSON_LIB_VERSION "22.0.0" CACHE STRING "simdjson library version")
set(SIMDJSON_LIB_SOVERSION "22" CACHE STRING "simdjson library soversion")
set(SIMDJSON_LIB_VERSION "23.0.0" CACHE STRING "simdjson library version")
set(SIMDJSON_LIB_SOVERSION "23" CACHE STRING "simdjson library soversion")
option(SIMDJSON_BUILD_STATIC_LIB "Build simdjson_static library along with simdjson" OFF)
+1 -1
View File
@@ -38,7 +38,7 @@ PROJECT_NAME = simdjson
# could be handy for archiving the generated documentation or if some version
# control system is used.
PROJECT_NUMBER = "3.9.2"
PROJECT_NUMBER = "3.10.0"
# Using the PROJECT_BRIEF tag one can provide an optional one line description
# for a project that appears at the top of each page and should give viewer a
+3 -3
View File
@@ -88,8 +88,8 @@ simdjson's source structure, from the top level, looks like this:
* simdjson/ondemand.h: the `simdjson::ondemand` namespace. Includes all public ondemand classes.
* simdjson/builtin.h: the `simdjson::builtin` namespace. Aliased to the most universal implementation available.
* simdjson/builtin/ondemand.h: the `simdjson::builtin::ondemand` namespace.
* simdjson/arm64|fallback|haswell|icelake|ppc64|westmere/ondemand.h: the `simdjson::<implementation>::ondemand` namespace. on demand compiled for the specific implementation.
* simdjson/generic/ondemand/*.h: individual on demand classes, generically written.
* simdjson/arm64|fallback|haswell|icelake|ppc64|westmere/ondemand.h: the `simdjson::<implementation>::ondemand` namespace. On-Demand compiled for the specific implementation.
* simdjson/generic/ondemand/*.h: individual On-Demand classes, generically written.
* simdjson/generic/ondemand/dependencies.h: dependencies on common, non-implementation-specific simdjson classes. This will be included before including amalgamated.h.
* simdjson/generic/ondemand/amalgamated.h: all generic ondemand classes for an implementation.
* **src:** The source files for non-inlined functionality (e.g. the architecture-specific parser
@@ -99,7 +99,7 @@ simdjson's source structure, from the top level, looks like this:
* *.cpp: other misc. implementations, such as `simdjson::implementation` and the minifier.
* arm64|fallback|haswell|icelake|ppc64|westmere.cpp: Architecture-specific parser implementations.
* generic/*.h: `simdjson::<implementation>` namespace. Generic implementation of the parser, particularly the `dom_parser_implementation`.
* generic/stage1/*.h: `simdjson::<implementation>::stage1` namespace. Generic implementation of the simd-heavy tokenizer/indexer pass of the simdjson parser. Used for the On Demand interface
* generic/stage1/*.h: `simdjson::<implementation>::stage1` namespace. Generic implementation of the simd-heavy tokenizer/indexer pass of the simdjson parser. Used for the On-Demand interface
* generic/stage2/*.h: `simdjson::<implementation>::stage2` namespace. Generic implementation of the tape creator, which consumes the index from stage 1 and actually parses numbers and string and such. Used for the DOM interface.
Other important files and directories:
+2 -1
View File
@@ -46,6 +46,7 @@ Real-world usage
- [Meta Velox](https://velox-lib.io)
- [Google Pax](https://github.com/google/paxml)
- [milvus](https://github.com/milvus-io/milvus)
- [QuestDB](https://questdb.io/blog/questdb-release-8-0-3/)
- [Clang Build Analyzer](https://github.com/aras-p/ClangBuildAnalyzer)
- [Shopify HeapProfiler](https://github.com/Shopify/heap-profiler)
- [StarRocks](https://github.com/StarRocks/starrocks)
@@ -179,7 +180,7 @@ The simdjson library takes advantage of modern microarchitectures, parallelizing
instructions, reducing branch misprediction, and reducing data dependency to take advantage of each
CPU's multiple execution cores.
Our default front-end is called On Demand, and we wrote a paper about it:
Our default front-end is called On-Demand, and we wrote a paper about it:
- John Keiser, Daniel Lemire, [On-Demand JSON: A Better Way to Parse Documents?](http://arxiv.org/abs/2312.17149), Software: Practice and Experience 54 (6), 2024.
+1 -1
View File
@@ -13,7 +13,7 @@ struct nlohmann_json {
auto root = nlohmann::json::parse(json.data(), json.data() + json.size());
for (auto tweet : root["statuses"]) {
if (tweet["id"] == find_id) {
result = tweet["text"];
result = to_string(tweet["text"]);
return true;
}
}
+1 -1
View File
@@ -13,7 +13,7 @@ class OnDemand {
public:
OnDemand() {
if(!displayed_implementation) {
std::cout << "On Demand implementation: " << builtin_implementation()->name() << std::endl;
std::cout << "On-Demand implementation: " << builtin_implementation()->name() << std::endl;
displayed_implementation = true;
}
}
+2 -2
View File
@@ -26,8 +26,8 @@ struct nlohmann_json {
}
}
result.text = top_tweet["text"];
result.screen_name = top_tweet["user"]["screen_name"];
result.text = to_string(top_tweet["text"]);
result.screen_name = to_string(top_tweet["user"]["screen_name"]);
return result.retweet_count != -1;
}
};
+1 -1
View File
@@ -155,6 +155,6 @@ if(SIMDJSON_CXXOPTS)
set_off(CXXOPTS_BUILD_TESTS)
set_off(CXXOPTS_ENABLE_INSTALL)
import_dependency(cxxopts jarro2783/cxxopts 794c975)
import_dependency(cxxopts jarro2783/cxxopts 5965670)
add_dependency(cxxopts)
endif()
+288 -27
View File
@@ -17,6 +17,8 @@ An overview of what you need to know to use simdjson, with examples.
- [Using the parsed JSON](#using-the-parsed-json)
- [Using the parsed JSON: additional examples](#using-the-parsed-json-additional-examples)
- [Adding support for custom types](#adding-support-for-custom-types)
- [1. Specialize `simdjson::ondemand::value::get` to get custom types](#1-specialize-simdjsonondemandvalueget-to-get-custom-types)
- [2. Use `tag_invoke` for custom types (C++20)](#2-use-tag_invoke-for-custom-types-c20)
- [Minifying JSON strings without parsing](#minifying-json-strings-without-parsing)
- [UTF-8 validation (alone)](#utf-8-validation-alone)
- [JSON Pointer](#json-pointer)
@@ -64,20 +66,23 @@ into your project. Then include it in your project with:
using namespace simdjson; // optional
```
You can compile with:
Under most systems, you can compile with:
```
c++ myproject.cpp simdjson.cpp
```
Note:
- We recommend that you use simdjson by copying the single-header `simdjson.h` file along with the source file `simdjson.cpp` directly in your project, as they are part of [every release](https://github.com/simdjson/simdjson/releases) as assets. In this manner, you only have to compile `simdjson.cpp` as any other source file: it works well in every development environment. However, you may also use simdjson as a git submodule ([example](https://github.com/simdjson/cmakedemo)), using FetchContent ([example](https://github.com/simdjson/cmake_demo_single_file)), with ExternalProject_Add ([example](https://github.com/simdjson/cmakedemo_externalproject)) or with CPM ([example](https://github.com/cpm-cmake/CPM.cmake/tree/master/examples/simdjson)).
- Users on macOS and other platforms where default compilers do not provide C++11 compliant by default should request it with the appropriate flag (e.g., `c++ -std=c++11 myproject.cpp simdjson.cpp`).
- The library relies on [runtime CPU detection](implementation-selection.md): avoid specifying an architecture at compile time (e.g., `-march-native`) if you want your binaries to run everywhere.
Using simdjson with package managers
------------------
You can install the simdjson library on your system or in your project using multiple package managers such as MSYS2, the conan package manager, vcpkg, brew, the apt package manager (debian-based Linux systems), the FreeBSD package manager (FreeBSD), and so on. [Visit our wiki for more details](https://github.com/simdjson/simdjson/wiki/Installing-simdjson-with-a-package-manager).
You can install the simdjson library on your system or in your project using multiple package managers such as MSYS2, the conan package manager, vcpkg, brew, the apt package manager (debian-based Linux systems), the FreeBSD package manager (FreeBSD), and so on. E.g., [we provide an complete example with vcpkg](https://github.com/simdjson/simdjson-vcpkg) that works under Windows. [Visit our wiki for more details](https://github.com/simdjson/simdjson/wiki/Installing-simdjson-with-a-package-manager).
The following Linux distributions provide simdjson packages: Alpine, RedHat, Rocky Linux, Debian, Fedora, and Ubuntu.
@@ -156,7 +161,8 @@ For efficiency reasons, simdjson requires a string with a few bytes (`simdjson::
at the end, these bytes may be read but their content does not affect the parsing. In practice,
it means that the JSON inputs should be stored in a memory region with `simdjson::SIMDJSON_PADDING`
extra bytes at the end. You do not have to set these bytes to specific values though you may
want to if you want to avoid runtime warnings with some sanitizers.
want to if you want to avoid runtime warnings with some sanitizers. Advanced users may want to
read the section Free Padding in [our performance notes](performance.md).
The simdjson library offers a tree-like [API](https://en.wikipedia.org/wiki/API), which you can
access by creating a `ondemand::parser` and calling the `iterate()` method. The iterate method
@@ -225,16 +231,16 @@ codepage, and they may call SetFileApisToOEM accordingly.
Documents are iterators
-----------------------
The simdjson library relies on an approach to parsing JSON that we call "On Demand".
The simdjson library relies on an approach to parsing JSON that we call "On-Demand".
A `document` is *not* a fully-parsed JSON value; rather, it is an **iterator** over the JSON text.
This means that while you iterate an array, or search for a field in an object, it is actually
walking through the original JSON text, merrily reading commas and colons and brackets to make sure
you get where you are going. This is the key to On Demand's performance: since it's just an iterator,
you get where you are going. This is the key to On-Demand's performance: since it's just an iterator,
it lets you parse values as you use them. And particularly, it lets you *skip* values you do not want
to use. On Demand is also ideally suited when you want to capture part of the document without parsing it
to use. On-Demand is also ideally suited when you want to capture part of the document without parsing it
immediately (e.g., see [General direct access to the raw JSON string](#general-direct-access-to-the-raw-json-string)).
We refer to "On Demand" as a front-end component since it is an interface between the
We refer to "On-Demand" as a front-end component since it is an interface between the
low-level parsing functions and the user. It hides much of the complexity of parsing JSON
documents.
@@ -340,7 +346,7 @@ floating-point values followed by an integer.
We invite you to keep the following rules in mind:
1. While you are accessing the document, the `document` instance should remain in scope: it is your "iterator" which keeps track of where you are in the JSON document. By design, there is one and only one `document` instance per JSON document.
2. Because On Demand is really just an iterator, you must fully consume the current object or array before accessing a sibling object or array.
2. Because On-Demand is really just an iterator, you must fully consume the current object or array before accessing a sibling object or array.
3. Values can only be consumed once, you should get the values and store them if you plan to need them multiple times. You are expected to access the keys of an object just once. You are expected to go through the values of an array just once.
The simdjson library makes generous use of `std::string_view` instances. If you are unfamiliar
@@ -353,7 +359,7 @@ support for users who avoid exceptions. See [the simdjson error handling documen
* **Validate What You Use:** When calling `iterate`, the document is quickly indexed. If it is
not a valid Unicode (UTF-8) string or if there is an unclosed string, an error may be reported right away.
However, it is not fully validated. On Demand only fully validates the values you use and the
However, it is not fully validated. On-Demand only fully validates the values you use and the
structure leading to it. It means that at every step as you traverse the document, you may encounter an error. You can handle errors either with exceptions or with error codes.
* **Extracting Values:** You can cast a JSON element to a native type:
`double(element)`. This works for `std::string_view`, double, uint64_t, int64_t, bool,
@@ -403,7 +409,7 @@ support for users who avoid exceptions. See [the simdjson error handling documen
* **Field Access:** To get the value of the "foo" field in an object, use `object["foo"]`. This will
scan through the object looking for the field with the matching string, doing a character-by-character
comparison. It may generate the error `simdjson::NO_SUCH_FIELD` if there is no such key in the object, it may throw an exception (see [Error handling](#error-handling)). For efficiency reason, you should avoid looking up the same field repeatedly: e.g., do
not do `object["foo"]` followed by `object["foo"]` with the same `object` instance. For best performance, you should try to query the keys in the same order they appear in the document. If you need several keys and you cannot predict the order they will appear in, it is recommended to iterate through all keys `for(auto field : object) {...}`. Generally, you should not mix and match iterating through an object (`for(auto field : object) {...}`) and key accesses (`object["foo"]`): if you need to iterate through an object after a key access, you need to call `reset()` on the object. Keep in mind that On Demand does not buffer or save the result of the parsing: if you repeatedly access `object["foo"]`, then it must repeatedly seek the key and parse the content. The library does not provide a distinct function to check if a key is present, instead we recommend you attempt to access the key: e.g., by doing `ondemand::value val{}; if (!object["foo"].get(val)) {...}`, you have that `val` contains the requested value inside the if clause. It is your responsibility as a user to temporarily keep a reference to the value (`auto v = object["foo"]`), or to consume the content and store it in your own data structures. If you consume an
not do `object["foo"]` followed by `object["foo"]` with the same `object` instance. For best performance, you should try to query the keys in the same order they appear in the document. If you need several keys and you cannot predict the order they will appear in, it is recommended to iterate through all keys `for(auto field : object) {...}`. Generally, you should not mix and match iterating through an object (`for(auto field : object) {...}`) and key accesses (`object["foo"]`): if you need to iterate through an object after a key access, you need to call `reset()` on the object. Whenever you call `reset()`, you need to keep in mind that though you can iterate over the array repeatedly, values should be consumedonly once (e.g., repeatedly calling `unescaped_key()` on the same key is forbidden). Keep in mind that On-Demand does not buffer or save the result of the parsing: if you repeatedly access `object["foo"]`, then it must repeatedly seek the key and parse the content. The library does not provide a distinct function to check if a key is present, instead we recommend you attempt to access the key: e.g., by doing `ondemand::value val{}; if (!object["foo"].get(val)) {...}`, you have that `val` contains the requested value inside the if clause. It is your responsibility as a user to temporarily keep a reference to the value (`auto v = object["foo"]`), or to consume the content and store it in your own data structures. If you consume an
object twice: `std::string_view(object["foo"]` followed by `std::string_view(object["foo"]` then your code
is in error. Furthermore, you can only consume one field at a time, on the same object. The
value instance you get from `content["bids"]` becomes invalid when you call `content["asks"]`.
@@ -567,7 +573,7 @@ support for users who avoid exceptions. See [the simdjson error handling documen
auto json = R"( { "test":{ "val1":1, "val2":2 } } )"_padded;
auto doc = parser.iterate(json);
size_t count = doc.count_fields(); // requires simdjson 1.0 or better
std::cout << "Number of fields: " << new_count << std::endl; // Prints "Number of fields: 1"
std::cout << "Number of fields: " << count << std::endl; // Prints "Number of fields: 1"
```
Similarly to `count_elements`, you should not let an object instance go out of scope before consuming it after calling
the `count_fields` method. If you access an object inside a document, you can use the `count_fields` method as follow.
@@ -784,6 +790,25 @@ for (ondemand::object points : parser.iterate(points_json)) {
Adding support for custom types
----------------------
There are 2 main ways provided by simdjson to deserialize a value into a custom type:
1. Provide a [**template specialization** for member functions](https://en.cppreference.com/w/cpp/language/template_specialization#Members_of_specializations)
1. Specialize `simdjson::ondemand::document::get` for the whole document
2. Specialize `simdjson::ondemand::value::get` for each value
2. Using `tag_invoke` *(the recommended way if your system supports C++20 or better)*
The differences between the two approaches are as follows:
1. Your compiler needs to support C++20 for `tag_invoke`.
2. The first way limits you to know your type fully, but in the `tag_invoke` way, you can use C++20 concepts and template meta programming as well.
3. Another difference between the two ways is that in `tag_invoke` way you may or may not mark your deserialization function (`tag_invoke` itself) as `noexcept`, but in the
*template specialization* way (the first way) you HAVE TO mark your specialization as `noexcept` which means that you may need to catch exceptions so
that your function does not throw any exception all the while it potentially can (for example any type like `std::string`, `std::vector`, etc. that has an allocator can potentially throw exceptions).
4. Another difference is that the `tag_invoke` way can work on both `document` and `value` but in the first way you have to provide 2 distinct specializations.
5. If you by mistake use both way with each other, the *specialization way* will be used silently due to backward-compatibility.
### 1. Specialize `simdjson::ondemand::value::get` to get custom types
Suppose you have your own types, such as a `Car` struct:
```C++
@@ -817,9 +842,14 @@ type:
```
We may do so by providing additional template definitions to the `ondemand::value` type.
We may start by providing a definition for `std::vector<double>` as follows:
We may start by providing a definition for `std::vector<double>` as follows. Observe
how we guard the code with `#ifndef __cpp_concepts`: that is because the necessary code
is automatically provided by simdjson if C++20 (and concepts) are available.
See [Use `tag_invoke` for custom types](#2-use-tag_invoke-for-custom-types-c20) if you have
C++20 support.
```c++
#ifndef __cpp_concepts // because the code is unnecessary with C++20
template <>
simdjson_inline simdjson_result<std::vector<double>>
simdjson::ondemand::value::get() noexcept {
@@ -831,10 +861,11 @@ simdjson::ondemand::value::get() noexcept {
double val;
error = v.get_double().get(val);
if (error) { return error; }
vec.push_back(val);
try { vec.push_back(val); } catch (...) { return simdjson::UNEXPECTED_ERROR; }
}
return vec;
}
#endif
```
We may then provide support for our `Car` struct:
@@ -852,6 +883,7 @@ simdjson_inline simdjson_result<Car> simdjson::ondemand::value::get() noexcept {
raw_json_string key;
error = field.key().get(key);
if (error) { return error; }
if (key == "make") {
error = field.value().get_string(car.make);
if (error) { return error; }
@@ -863,7 +895,7 @@ simdjson_inline simdjson_result<Car> simdjson::ondemand::value::get() noexcept {
if (error) { return error; }
} else if (key == "tire_pressure") {
error = field.value().get<std::vector<double>>().get(car.tire_pressure);
if (auto error) { return error; }
if (error) { return error; }
}
}
return car;
@@ -889,6 +921,7 @@ struct Car {
std::vector<double> tire_pressure;
};
#ifndef __cpp_concepts // because the code is unnecessary with C++20
template <>
simdjson_inline simdjson_result<std::vector<double>>
simdjson::ondemand::value::get() noexcept {
@@ -904,6 +937,7 @@ simdjson::ondemand::value::get() noexcept {
}
return vec;
}
#endif
template <>
@@ -946,7 +980,7 @@ int main(void) {
ondemand::parser parser;
ondemand::document doc = parser.iterate(json);
for (auto val : doc) {
Car c(val);
Car c(val); // an exception may be thrown
std::cout << c.make << std::endl;
}
direct();
@@ -954,7 +988,7 @@ int main(void) {
}
```
Observe that we require an explicit cast (`Car c(val)` instead of `for (Car c : doc) {`): it is by design.
Observe that we require an explicit cast (`Car c(val)` instead of `for (Car c : doc) {`): it is by design. We require explicit casting.
If you prefer to avoid exceptions, you may modify the `main` function as follows:
@@ -1005,6 +1039,7 @@ struct Car {
std::vector<double> tire_pressure;
};
#ifndef __cpp_concepts // because the code is unnecessary with C++20
template <>
simdjson_inline simdjson_result<std::vector<double>>
simdjson::ondemand::value::get() noexcept {
@@ -1018,7 +1053,7 @@ simdjson::ondemand::value::get() noexcept {
}
return vec;
}
#endif
template <>
simdjson_inline simdjson_result<Car> simdjson::ondemand::document::get() & noexcept {
@@ -1062,6 +1097,178 @@ int main(void) {
}
```
### 2. Use `tag_invoke` for custom types (C++20)
A standard approach to provide support for deserialization is to add a `tag_invoke` function
in your own class:
```C++
/**
* A custom type that we want to parse.
*/
struct Car {
std::string make;
std::string model;
int year = 0;
std::vector<double> tire_pressure;
friend simdjson_result<Car>
tag_invoke(simdjson::deserialize_tag, std::type_identity<Car>, auto &val) {
simdjson::ondemand::object obj;
auto error = val.get_object().get(obj);
if (error) {
return error;
}
Car car{};
// Instead of repeatedly obj["something"], we iterate through the object
// which we expect to be faster.
for (auto field : obj) {
simdjson::ondemand::raw_json_string key;
error = field.key().get(key);
if (error) {
return error;
}
if (key == "make") {
error = field.value().get_string(car.make);
if (error) {
return error;
}
} else if (key == "model") {
error = field.value().get_string(car.model);
if (error) {
return error;
}
} else if (key == "year") {
error = field.value().get(car.year);
if (error) {
return error;
}
} else if (key == "tire_pressure") {
error = field.value().get(car.tire_pressure);
if (error) {
return error;
}
}
}
return car;
}
};
```
Let us explain each argument of `tag_invoke`:
- `simdjson::deserialize_tag`: it is the tag for Customization Point Object (CPO)
- `std::type_identity<T>`: We specify the custom type we want here
- `simdjson::ondemand::value&` or `simdjson::ondemand::document&`: the value or document that we want to deserialize from; you may want to just specify `auto&` to capture both of them.
You can use it like so:
```cpp
simdjson::padded_string json =
R"( { "make": "Toyota", "model": "Camry", "year": 2018,
"tire_pressure": [ 40.1, 39.9 ] })"_padded;
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc = parser.iterate(json);
Car c(doc);
std::cout << c.make << std::endl;
```
Observe how we first get an instance of `document` and then we cast.
You can also handle errors explicitly:
```cpp
simdjson::padded_string json =
R"( { "make": "Toyota", "model": "Camry", "year": 2018,
"tire_pressure": [ 40.1, 39.9 ] })"_padded;
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc = parser.iterate(json);
Car c;
auto error = doc.get(c);
if(error) { std::cerr << simdjson::error_message(error); return false; }
std::cout << c.make << std::endl;
```
You can also read instances of `Car` from an array or an object:
```cpp
simdjson::padded_string json =
R"( [ { "make": "Toyota", "model": "Camry", "year": 2018,
"tire_pressure": [ 40.1, 39.9 ] },
{ "make": "Kia", "model": "Soul", "year": 2012,
"tire_pressure": [ 30.1, 31.0 ] },
{ "make": "Toyota", "model": "Tercel", "year": 1999,
"tire_pressure": [ 29.8, 30.0 ] }
])"_padded;
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc = parser.iterate(json);
for (auto val : doc) {
Car c(val); // an exception may be thrown
std::cout << c.year << std::endl;
}
```
Observe how we first get a generic (`val`) which we cast to `Car`. It is by design: we require
explicit casting. The cast may throw an exception.
Once more, you can handle errors explicitly:
```cpp
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc = parser.iterate(json);
for (auto val : doc) {
Car c;
auto error = val.get(c);
if(error) { std::cerr << simdjson::error_message(error) << std::endl; return false; }
}
```
You can also use the custom `Car` type as part of a template such as `std::vector`:
```cpp
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc = parser.iterate(json);
std::vector<Car> cars(doc);
// visual studio users need an explicit call:
// std::vector<Car> cars = doc.get<std::vector<Car>>();
// because the compiler does not know whether to convert
// doc to an unsigned int or to a vector.
for(Car& c : cars) {
std::cout << c.year << std::endl;
}
```
You can also extend support to arbitrary templates.
Suppose your custom types are `std::unique_ptr<any type>`,
you could add the following `tag_invoke`:
```C++
namespace simdjson {
// This tag_invoke MUST be inside simdjson namespace since we can't make it a friend of unique_ptr
template <typename T>
auto tag_invoke(deserialize_tag, std::type_identity<std::unique_ptr<T>>, auto &val) {
return simdjson_result{std::make_unique<T>(val.template get<T>())};
}
} // namespace simdjson
```
You can also mark your `tag_invoke` function as `noexcept` and we will obey that, unlike the *template specialization* way.
You would use it like this:
```C++
int main() {
auto const json = R"( { "make": "Toyota", "model": "Camry", "year": 2018,
"tire_pressure": [ 40.1, 39.9 ] })"_padded;
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc = parser.iterate(json);
std::unique_ptr<Car> c(doc);
std::cout << c->make << std::endl;
return EXIT_SUCCESS;
}
```
Minifying JSON strings without parsing
----------------------
@@ -1103,9 +1310,9 @@ If you find yourself needing only fast Unicode functions, consider using the sim
JSON Pointer
------------
The simdjson library also supports [JSON pointer](https://tools.ietf.org/html/rfc6901) through the `at_pointer()` method, letting you reach further down into the document in a single call. JSON pointer is supported by both the [DOM approach](https://github.com/simdjson/simdjson/blob/master/doc/dom.md#json-pointer) as well as the On Demand approach.
The simdjson library also supports [JSON pointer](https://tools.ietf.org/html/rfc6901) through the `at_pointer()` method, letting you reach further down into the document in a single call. JSON pointer is supported by both the [DOM approach](https://github.com/simdjson/simdjson/blob/master/doc/dom.md#json-pointer) as well as the On-Demand approach.
**Note:** The On Demand implementation of JSON pointer relies on `find_field` which implies that it does not unescape keys when matching.
**Note:** The On-Demand implementation of JSON pointer relies on `find_field` which implies that it does not unescape keys when matching.
Consider the following example:
@@ -1339,7 +1546,7 @@ Notice how we can retrieve the exact error condition (in this instance `simdjson
from the exception.
We can write a "quick start" example where we attempt to parse the following JSON file and access some data, without triggering exceptions:
```JavaScript
```JSON
{
"statuses": [
{
@@ -1723,7 +1930,7 @@ before printout the data.
}
```
Performance note: the On Demand front-end does not materialize the parsed numbers and other values. If you are accessing everything twice, you may need to parse them twice. Thus the rewind functionality is best suited for cases where the first pass only scans the structure of the document.
Performance note: the On-Demand front-end does not materialize the parsed numbers and other values. If you are accessing everything twice, you may need to parse them twice. Thus the rewind functionality is best suited for cases where the first pass only scans the structure of the document.
Both arrays and objects have a similar method `reset()`. It is similar
to the document `rewind()` method, except that it does not rewind the
@@ -1762,7 +1969,7 @@ for (auto doc : docs) {
```
Unlike `parser.iterate`, `parser.iterate_many` may parse "on demand" (lazily). That is, no parsing may have been done before you enter the loop
Unlike `parser.iterate`, `parser.iterate_many` may parse "On-Demand" (lazily). That is, no parsing may have been done before you enter the loop
`for (auto doc : docs) {` and you should expect the parser to only ever fully parse one JSON document at a time.
As with `parser.iterate`, when calling `parser.iterate_many(string)`, no copy is made of the provided string input. The provided memory buffer may be accessed each time a JSON document is parsed. Calling `parser.iterate_many(string)` on a temporary string buffer (e.g., `docs = parser.parse_many("[1,2,3]"_padded)`) is unsafe (and will not compile) because the `document_stream` instance needs access to the buffer to return the JSON documents.
@@ -2211,12 +2418,39 @@ string representation.
```
You can use `raw_json()` to capture the content of some JSON values as `std::string_view`
instances which can be safely used later. The `std::string_view` instances point inside
the original document and do not depend in any way on simdjson. In the following example,
we store the `std::string_view` instances inside a `std::vector<std::string_view>` instance
and print the out after the parsing is concluded:
```cpp
padded_string json_padded = "{\"a\":[1,2,3], \"b\": 2, \"c\": \"hello\"}"_padded;
std::vector<std::string_view> fields;
ondemand::parser parser;
auto doc = parser.iterate(json_padded);
auto object = doc.get_object();
for (auto field : object) {
fields.push_back(field.value().raw_json());
}
// Output the fields
// Expected output:
// [1,2,3]
// 2
// "hello"
for (std::string_view field_ref : fields) {
std::cout << field_ref << std::endl;
}
```
Storing directly into an existing string instance
-----------------------------------------------------
The simdjson library favours the use of `std::string_view` instances because
it tends to lead to better performance due to causing fewer memory allocations.
However, they are cases where you need to store a string result in a `std::string``
However, they are cases where you need to store a string result in a `std::string`
instance. You can do so with a templated version of the `to_string()` method which takes as
a parameter a reference to a `std::string`.
@@ -2266,7 +2500,7 @@ We built simdjson with thread safety in mind.
The simdjson library is single-threaded except for [`iterate_many`](iterate_many.md) and [`parse_many`](parse_many.md) which may use secondary threads under their control when the library is compiled with thread support.
We recommend using one `parser` object per thread. When using the On Demand front-end (our default), you should access the `document` instances in a single-threaded manner since it
We recommend using one `parser` object per thread. When using the On-Demand front-end (our default), you should access the `document` instances in a single-threaded manner since it
acts as an iterator (and is therefore not thread safe).
The CPU detection, which runs the first time parsing is attempted and switches to the fastest
@@ -2592,11 +2826,37 @@ int main(void) {
}
```
* Example 4: Value capture with `std::string_view` instances
```cpp
void example() {
ondemand::parser parser;
const padded_string json = R"({ "parent": {"child1": {"name": "John"} , "child2": {"name": "Daniel"}} })"_padded;
auto doc = parser.iterate(json);
ondemand::object parent = doc["parent"];
// parent owns the focus
ondemand::object c1 = parent["child1"];
// c1 owns the focus
//
std::string_view as1 = c1["name"];
// We have that as1 == "John", as long as 'parser' and 'json' live
// c2 attempts to grab the focus from parent but fails
ondemand::object c2 = parent["child2"];
// c2 owns the focus, at this point c1 is invalid
std::string_view as2 = c2["name"];
// We have that as2 == "Daniel", as long as 'parser' and 'json' live
std::cout << as1 << " " << as2 << std::endl; // prints John Daniel
}
```
Performance tips
--------
- The On Demand front-end works best when doing a single pass over the input: avoid calling `count_elements`, `rewind` and similar methods.
- Read [our performance notes](performance.md) for advanced topics.
- To better understand the operation of your On-Demand parser, and whether it is performing as well as you think it should be, there is a logger feature built in to simdjson! To use it, define the pre-processor directive `SIMDJSON_VERBOSE_LOGGING` prior to including the `simdjson.h` header, which enables logging in simdjson. Run your code. It may generate a lot of logging output; adding printouts from your application that show each section may be helpful. The log's output will show step-by-step information on state, buffer pointer position, depth, and key retrieval status. Importantly, unless `SIMDJSON_VERBOSE_LOGGING` is defined, logging is entirely disabled and thus carries no overhead.
- The On-Demand front-end works best when doing a single pass over the input: avoid calling `count_elements`, `rewind`, `reset` and similar methods.
- If you are familiar with assembly language, you may use the online tool godbolt to explore the compiled code. The following example may work: [https://godbolt.org/z/xE4GWs573](https://godbolt.org/z/xE4GWs573).
- Given a field `field` in an object, calling `field.key()` is often faster than `field.unescaped_key()` so if you do not need an unescaped `std::string_view` instance, prefer `field.key()`. Similarly, we expect `field.escaped_key()` to be faster than `field.unescaped_key()` even though both return a `std::string_view` instance.
- For release builds, we recommend setting `NDEBUG` pre-processor directive when compiling the `simdjson` library. Importantly, using the optimization flags `-O2` or `-O3` under GCC and LLVM clang does not set the `NDEBUG` directive, you must set it manually (e.g., `-DNDEBUG`).
@@ -2616,10 +2876,11 @@ Performance tips
std::string_view year = data["year"];
std::string_view rating = data["rating"];
```
- To better understand the operation of your On Demand parser, and whether it is performing as well as you think it should be, there is a logger feature built in to simdjson! To use it, define the pre-processor directive `SIMDJSON_VERBOSE_LOGGING` prior to including the `simdjson.h` header, which enables logging in simdjson. Run your code. It may generate a lot of logging output; adding printouts from your application that show each section may be helpful. The log's output will show step-by-step information on state, buffer pointer position, depth, and key retrieval status. Importantly, unless `SIMDJSON_VERBOSE_LOGGING` is defined, logging is entirely disabled and thus carries no overhead.
Further reading
--------
- John Keiser, Daniel Lemire, [On-Demand JSON: A Better Way to Parse Documents?](http://arxiv.org/abs/2312.17149), Software: Practice and Experience (to appear)
- John Keiser, Daniel Lemire, [On-Demand JSON: A Better Way to Parse Documents?](http://arxiv.org/abs/2312.17149), Software: Practice and Experience 54 (6), 2024
+2 -2
View File
@@ -3,7 +3,7 @@ The Document-Object-Model (DOM) front-end
An overview of what you need to know to use simdjson, with examples.
* [DOM vs On Demand](#dom-vs-on-demand)
* [DOM vs On-Demand](#dom-vs-on-demand)
* [The Basics: Loading and Parsing JSON Documents](#the-basics-loading-and-parsing-json-documents-using-the-dom-front-end)
* [Using the Parsed JSON](#using-the-parsed-json)
* [C++17 Support](#c17-support)
@@ -18,7 +18,7 @@ An overview of what you need to know to use simdjson, with examples.
* [Padding and Temporary Copies](#padding-and-temporary-copies)
* [Performance Tips](#performance-tips)
DOM vs On Demand
DOM vs On-Demand
----------------------------------------------
The simdjson library offers two distinct approaches on how to access a JSON document. We support
+30 -30
View File
@@ -9,8 +9,8 @@ Whether we parse JSON or XML, or any other serialized format, there are relative
- Another popular approach is the schema-based deserialization model.
We propose an approach that is as easy to use and often as flexible as the DOM approach, yet as fast and
efficient as the schema-based or event-based approaches. We call this new approach "On Demand". The
simdjson On Demand API offers a familiar, friendly DOM API and
efficient as the schema-based or event-based approaches. We call this new approach "On-Demand". The
simdjson On-Demand API offers a familiar, friendly DOM API and
provides the performance of just-in-time parsing on top of the simdjson superior performance.
To achieve ease of use, we mimicked the *form* of a traditional DOM API: you can iterate over
@@ -18,7 +18,7 @@ arrays, look up fields in objects, and extract native values like `double`, `uin
To achieve performance, we introduced some key limitations that make the DOM API *streaming*:
array/object iteration cannot be restarted, and string/number values can only be parsed once. If
these limitations are acceptable to you, the On Demand API could help you write maintainable
these limitations are acceptable to you, the On-Demand API could help you write maintainable
applications with a computation efficiency that is difficult to surpass.
A code example illustrates our API from a programmer's point of view:
@@ -72,24 +72,24 @@ This streaming approach means that unused fields and values are not parsed or
converted, thus saving space and time. In our example, the `"name"`, `"followers_count"`,
and `"friends_count"` keys and matching values are skipped.
Further, the On Demand API does not parse a value *at all* until you try to convert it (e.g., to `double`,
Further, the On-Demand API does not parse a value *at all* until you try to convert it (e.g., to `double`,
`int`, `string`, or `bool`). In our example, when accessing the key-value pair `"retweet_count": 82`, the parser
may not convert the pair of characters `82` to the binary integer 82. Because the programmer specifies the data
type, we avoid branch mispredictions related to data type determination and improve the performance.
We expect users of an On Demand API to work in terms of a JSON dialect, which is a set of expectations and
We expect users of an On-Demand API to work in terms of a JSON dialect, which is a set of expectations and
specifications that come in addition to the [JSON specification](https://www.rfc-editor.org/rfc/rfc8259.txt).
The On Demand approach is designed around several principles:
The On-Demand approach is designed around several principles:
* **Streaming (\*):** It avoids preparsing values, keeping the memory usage and the latency down.
* **Forward-Only:** To prevent reiteration of the same values and to keep the number of variables down (literally), only a single index is maintained and everything uses it (even if you have nested for loops). This means when you are going through an array of arrays, for example, that the inner array loop will advance the index to the next comma, and the array can just pick it up and look at it.
* **Natural Iteration:** A JSON array or object can be iterated with a normal C++ for loop. Nested arrays and objects are supported by nested for loops.
* **Use-Specific Parsing:** Parsing is always specific to the type required by the programmer. For example, if the programmer asks for an unsigned integer, we just start parsing digits. If there were no digits, we toss an error. There are even different parsers for `double`, `uint64_t` and `int64_t` values. This use-specific parsing avoids the branchiness of a generic "type switch," and makes the code more inlineable and compact.
* **Validate What You Use:** On Demand deliberately validates the values you use and the structure leading to it, but nothing else. The goal is a guarantee that the value you asked for is the correct one and is not malformed: there must be no confusion over whether you got the right value.
* **Validate What You Use:** On-Demand deliberately validates the values you use and the structure leading to it, but nothing else. The goal is a guarantee that the value you asked for is the correct one and is not malformed: there must be no confusion over whether you got the right value.
To understand why On Demand is different, it is helpful to review the major
To understand why On-Demand is different, it is helpful to review the major
approaches to parsing and parser APIs in use today.
### DOM Parsers
@@ -106,7 +106,7 @@ DOM tree is often easy enough that many users use the DOM as-is instead of creat
their own custom data structures.
The DOM approach was the only way to parse JSON documents up to version 0.6 of the simdjson library.
Our DOM API looks similar to our On Demand example, except
Our DOM API looks similar to our On-Demand example, except
it calls `parse` instead of `iterate`:
```c++
@@ -152,7 +152,7 @@ a tweet right now, or is this from some other place in the document
entirely? Though an event-based approach may allow superior performance, it is demanding of the programmer
who must efficiently keep track of its current state within the JSON input.
The following is event-based example of the Twitter problem we have reviewed in the DOM and On Demand
The following is event-based example of the Twitter problem we have reviewed in the DOM and On-Demand
examples. To make it short enough to use as an example at all, it has heavily redacted: it only solves
a part of the problem (does not get user.screen_name), it has bugs (it does not handle sub-objects
in a tweet at all), and it uses a theoretical, simple event-based API that minimizes ceremony.
@@ -257,7 +257,7 @@ stress the branch prediction. Though branch predictors improve with each new gen
the cost of branch mispredictions also tends to increase as pipelines expand, and the processors become
able to schedule longer streams of instructions.
On Demand parsing is tailor-made to solve this problem at the source, parsing values only after the
On-Demand parsing is tailor-made to solve this problem at the source, parsing values only after the
user declares their type by asking for a `double`, an `int`, a `string`, etc. It attempts to do so while
preserving most of the flexibility of DOM parsing.
@@ -297,7 +297,7 @@ To help visualize the algorithm, we'll walk through the example C++ given at the
Since this is the first time this parser has been used, `iterate()` first allocates internal
parser buffers if this is the first time through. When reusing an existing parser, allocation
only happens if the new document is bigger than internal buffers can handle. The On Demand
only happens if the new document is bigger than internal buffers can handle. The On-Demand
API only ever allocates memory in the `iterate()` function call.
The simdjson library then preprocesses the JSON text at high speed, finding all tokens (i.e. the starting
@@ -492,7 +492,7 @@ To help visualize the algorithm, we'll walk through the example C++ given at the
Because of the cast to uint64_t, simdjson knows it's parsing an unsigned integer. This lets
us use a fast parser which *only* knows how to parse digits. It validates that it is an integer
by rejecting negative numbers, strings, and other values based on the fact that they are not the
digits 0-9. This type specificity is part of why parsing with on demand is so fast: you lose all
digits 0-9. This type specificity is part of why parsing with On-Demand is so fast: you lose all
the code that has to understand those other types.
The iterator is advanced to the `}`, and depth decreased back to 3 (root > statuses > tweet).
@@ -597,7 +597,7 @@ To help visualize the algorithm, we'll walk through the example C++ given at the
This means you can very efficiently do things like read a single value from a JSON file, or take
the top N, for example. It also means the things you don't use won't be fully validated. This is
a general principle of On Demand: don't validate what you don't use. We still fully validate
a general principle of On-Demand: don't validate what you don't use. We still fully validate
values you do use, however, as well as the objects and arrays that lead to them, so that you can
be sure you get the information you need.
@@ -654,7 +654,7 @@ for(auto field : doc.get_object()) {
### Iteration Safety
The On Demand API is powerful. To compensate, we add some safeguards to ensure that it can be used without fear
The On-Demand API is powerful. To compensate, we add some safeguards to ensure that it can be used without fear
in production systems:
- If the value fails to be parsed as one type, the program can try to parse it as something else until the program succeeds. Thus
@@ -667,7 +667,7 @@ in production systems:
if it was `nullptr` but did not care what the actual value was--it will iterate. The destructor automates
the iteration.
Some care is needed when using the On Demand API in scenarios where you need to access several sibling arrays or objects because
Some care is needed when using the On-Demand API in scenarios where you need to access several sibling arrays or objects because
only one object or array can be active at any one time. Let us consider the following example:
```C++
@@ -709,36 +709,36 @@ A correct usage is given by the following example:
}
```
### Benefits of the On Demand Approach
### Benefits of the On-Demand Approach
We expect that the On Demand approach has many of the performance benefits of the schema-based approach, while providing a flexibility that is similar to that of the DOM-based approach.
We expect that the On-Demand approach has many of the performance benefits of the schema-based approach, while providing a flexibility that is similar to that of the DOM-based approach.
* Faster than DOM in some cases. Reduced memory usage.
* Straightforward, programmer-friendly interface (arrays and objects).
* Highly expressive, beyond deserialization and pointer queries: many tasks can be accomplished with little code.
### Limitations of the On Demand Approach
### Limitations of the On-Demand Approach
The On Demand approach has some limitations:
The On-Demand approach has some limitations:
* Because it operates in streaming mode, you only have access to the current element in the JSON document. Furthermore, the document is traversed in order so the code is sensitive to the order of the JSON nodes in the same manner as an event-based approach (e.g., SAX). (The one exception to this is field lookup, which is more *performant* when the order of lookups matches the order of fields in the document, but which will still work with out-of-order fields, with a performance hit.)
* The On Demand approach is less safe than DOM: we only validate the components of the JSON document that are used and it is possible to begin ingesting an invalid document only to find out later that the document is invalid. Are you fine ingesting a large JSON document that starts with well formed JSON but ends with invalid JSON content?
* The On-Demand approach is less safe than DOM: we only validate the components of the JSON document that are used and it is possible to begin ingesting an invalid document only to find out later that the document is invalid. Are you fine ingesting a large JSON document that starts with well formed JSON but ends with invalid JSON content?
There are currently additional technical limitations which we expect to resolve in future releases of the simdjson library:
* The simdjson library offers runtime dispatching which allows you to compile one binary and have it run at full speed on different processors, taking advantage of the specific features of the processor. The On Demand API has limited runtime dispatch support. Under x64 systems, to fully benefit from the On Demand API, we recommend that you compile your code for a specific processor. E.g., if your processor supports AVX2 instructions, you should compile your binary executable with AVX2 instruction support (by using your compiler's commands). If you are sufficiently technically proficient, you can implement runtime dispatching within your application, by compiling your On Demand code for different processors.
* The simdjson library offers runtime dispatching which allows you to compile one binary and have it run at full speed on different processors, taking advantage of the specific features of the processor. The On-Demand API has limited runtime dispatch support. Under x64 systems, to fully benefit from the On-Demand API, we recommend that you compile your code for a specific processor. E.g., if your processor supports AVX2 instructions, you should compile your binary executable with AVX2 instruction support (by using your compiler's commands). If you are sufficiently technically proficient, you can implement runtime dispatching within your application, by compiling your On-Demand code for different processors.
* There is an initial phase which scans the entire document quickly, irrespective of the size of the document. We plan to break this phase into distinct steps for large files in a future release as we have done with other components of our API (e.g., `parse_many`).
### Applicability of the On Demand Approach
### Applicability of the On-Demand Approach
At this time we recommend the On Demand API in the following cases:
At this time we recommend the On-Demand API in the following cases:
1. The 64-bit hardware (CPU) used to run the software is known at compile time. If you need runtime dispatching because you cannot be certain of the hardware used to run your software, you will be better served with the core simdjson API. (This only applies to x64 (AMD/Intel). On 64-bit ARM hardware, runtime dispatching is unnecessary.)
2. The used parts of JSON files do not need to be validated and the layout of the nodes follows a strict JSON dialect. If you are receiving JSON from other systems, you might be better served with core simdjson API as it fully validates the JSON inputs and allows you to navigate through the document at will.
3. Speed and efficiency are of the utmost importance. Keep in mind that the core simdjson API is highly efficient so adopting the On Demand API is not necessary for high efficiency.
3. Speed and efficiency are of the utmost importance. Keep in mind that the core simdjson API is highly efficient so adopting the On-Demand API is not necessary for high efficiency.
4. As a developer, you value a clean, flexible and maintainable API.
Good applications for the On Demand API might be:
Good applications for the On-Demand API might be:
* You are working from pre-existing large JSON files that have been vetted. You expect them to be well formed according to a known JSON dialect and to have a consistent layout. For example, you might be doing biomedical research or machine learning on top of static data dumps in JSON.
* Both the generation and the consumption of JSON data is within your system. Your team controls both the software that produces the JSON and the software the parses it, your team knows and control the hardware. Thus you can fully test your system.
@@ -746,13 +746,13 @@ Good applications for the On Demand API might be:
## Checking Your CPU Selection (x64 systems)
The On Demand API uses advanced architecture-specific code for many common processors to make JSON preprocessing and string parsing faster. By default, however, most c++ compilers will compile to the least common denominator (since the program could theoretically be run anywhere). Since On Demand is inlined into your own code, it cannot always use these advanced versions unless the compiler is told to target them.
The On-Demand API uses advanced architecture-specific code for many common processors to make JSON preprocessing and string parsing faster. By default, however, most c++ compilers will compile to the least common denominator (since the program could theoretically be run anywhere). Since On-Demand is inlined into your own code, it cannot always use these advanced versions unless the compiler is told to target them.
On relevant systems, the On Demand API provides some support for runtime dispatching: that is, it will attempt to detect, at runtime, the instructions that your processor supports and optimize the code accordingly. However, it cannot always make full use of the features of your processor.
On relevant systems, the On-Demand API provides some support for runtime dispatching: that is, it will attempt to detect, at runtime, the instructions that your processor supports and optimize the code accordingly. However, it cannot always make full use of the features of your processor.
Some users wish to run at the best possible speed. Under recent Intel and AMD processors, these users should take additional steps to verify that their code is well optimized.
Given that the On Demand API offer limited runtime dispatching, it matters that your code is compiled against a specific CPU target. You should verify that the code is compiled against the target you expect. Thankfully, the simdjson library will tell you exactly what it detects as an implementation: `icelake` (AVX512 x64 processors), `haswell` (AVX2 x64 processors), `westmere` (SSE4 x64 processors), `arm64` (64-bit ARM), `ppc64` (64-bit POWER), `lasx` (LoongArch), `lsx` (LoongArch), `fallback` (others). Under x64 processors, many programmers will want to target `haswell` whereas under ARM, most programmers will want to target `arm64` (and it should do so automatically). The `fallback` is probably only good for testing purposes, not for deployment.
Given that the On-Demand API offer limited runtime dispatching, it matters that your code is compiled against a specific CPU target. You should verify that the code is compiled against the target you expect. Thankfully, the simdjson library will tell you exactly what it detects as an implementation: `icelake` (AVX512 x64 processors), `haswell` (AVX2 x64 processors), `westmere` (SSE4 x64 processors), `arm64` (64-bit ARM), `ppc64` (64-bit POWER), `lasx` (LoongArch), `lsx` (LoongArch), `fallback` (others). Under x64 processors, many programmers will want to target `haswell` whereas under ARM, most programmers will want to target `arm64` (and it should do so automatically). The `fallback` is probably only good for testing purposes, not for deployment.
```C++
std::cout << simdjson::builtin_implementation()->name() << std::endl;
@@ -777,6 +777,6 @@ In these examples, the `-march=haswell` flags targets a haswell processor and th
Instead of specifying a specific microarchitecture, you can let your compiler do the work. The `-march=native` flags says "target the current computer," which is a reasonable default for many applications which both compile and run on the same processor.
Passing `-march=native` to the compiler may make On Demand faster by allowing it to use optimizations specific to your machine. You cannot do this, however, if you are compiling code that might be run on less advanced machines. That is, be mindful that when compiling with the `-march=native` flag, the resulting binary will run on the current system but may not run on other systems (e.g., on an old processor).
Passing `-march=native` to the compiler may make On-Demand faster by allowing it to use optimizations specific to your machine. You cannot do this, however, if you are compiling code that might be run on less advanced machines. That is, be mindful that when compiling with the `-march=native` flag, the resulting binary will run on the current system but may not run on other systems (e.g., on an old processor).
If you are compiling on an ARM or POWER system, you do not need to be concerned with CPU selection during compilation. The `-march=native` flag is useful for best performance on x64 (e.g., Intel) systems but it is generally unsupported on some platforms such as ARM (aarch64) or POWER.
+109 -1
View File
@@ -14,6 +14,7 @@ testing and get the best performance.
* [Number parsing](#number-parsing)
* [Visual Studio](#visual-studio)
* [Power Usage and Downclocking](#power-usage-and-downclocking)
* [Free Padding](#free-padding)
NDEBUG directive
@@ -74,7 +75,7 @@ or simply
Server Loops: Long-Running Processes and Memory Capacity
---------------------------------
The On Demand approach also automatically expands its memory capacity when larger documents are parsed. However, for longer processes where very large files are processed (such as server loops), this capacity is not resized down. On Demand also lets you adjust the maximal capacity that the parser can process:
The On-Demand approach also automatically expands its memory capacity when larger documents are parsed. However, for longer processes where very large files are processed (such as server loops), this capacity is not resized down. On-Demand also lets you adjust the maximal capacity that the parser can process:
* You can set an upper bound (*max_capacity*) when construction the parser:
```C++
@@ -179,3 +180,110 @@ The simdjson library does not generally make use of heavy 256-bit instructions.
the macro `SIMDJSON_AVX512_ALLOWED` to `0` in C++ prior to importing the headers.
You may still be worried about which SIMD instruction set is used by simdjson. Thankfully, [you can always determine and change which architecture-specific implementation is used](implementation-selection.md) by simdjson. Thus even if your CPU supports AVX2, you do not need to use AVX2. You are in control.
Free Padding
-------
For performance reasons, the simdjson library requires that the JSON input contain at least
`simdjson::SIMDJSON_PADDING` bytes at the end of the stream. The value `simdjson::SIMDJSON_PADDING` is
small (e.g., 64 bytes). On modern systems, you can safely read beyond an allocated buffers,
as long as you remain within an allocated page. Pages on modern systems span at least 4 kilobytes,
but can be significantly larger. E.g., Apple systems favour pages spanning 16 kilobytes.
In effect, it means that you can almost always read a few bytes beyond your current buffer---without
allocating extra memory. However, tools such as valgrind or memory sanitizers will flag such behavior as unsafe.
Nevertheless, you can still make sure of this capability in your code if you are an expert
programmer and you are willing to silence sanitizer warnings. The following code provides
a portable example.
The conditional compilation checks for the `_MSC_VER` macro (indicating Microsoft Visual Studio)
and includes platform-specific headers accordingly.
The `page_size()` function determines the default size of a memory page in bytes on the system.
On Windows (when `_WIN32` is defined), it uses `GetSystemInfo()` to retrieve system information and obtain the page size.
On other platforms (non-Windows), it uses `sysconf(_SC_PAGESIZE)` to get the page size.
The function returns the page size.
The `need_allocation()` function checks whether the buffer (given by `buf`) plus the specified length (`len`) is near a page boundary.
If the buffer extends beyond the current page when padded by `simdjson::SIMDJSON_PADDING`, it returns true, indicating that reallocation is needed.
Otherwise, it returns false.
The `get_padded_string_view()` creates a `padded_string_view` from the input buffer.
If reallocation is needed (unlikely case), it allocates a new padded_string and assigns it to `jsonbuffer`.
Otherwise (very likely), it creates a `padded_string_view` directly from the buffer.
The `simdjson::SIMDJSON_PADDING` ensures that there is additional padding for parsing efficiency.
The calling code just needs to provide `jsonbuffer` (an instance of `simdjson::padded_string`)
and pass `get_padded_string_view(buf, len, jsonbuffer)` to `parser.iterate`. Most of the time,
this code will not allocate new memory.
```cpp
#ifdef _WIN32
#include <windows.h>
#include <sysinfoapi.h>
#else
#include <unistd.h>
#endif
#include "simdjson.h"
#include <cstdio>
// Returns the default size of the page in bytes on this system.
long page_size() {
#ifdef _WIN32
SYSTEM_INFO sysInfo;
GetSystemInfo(&sysInfo);
long pagesize = sysInfo.dwPageSize;
#else
long pagesize = sysconf(_SC_PAGESIZE);
#endif
return pagesize;
}
// Returns true if the buffer + len + simdjson::SIMDJSON_PADDING crosses the
// page boundary.
bool need_allocation(const char *buf, size_t len) {
return ((reinterpret_cast<uintptr_t>(buf + len - 1) % page_size()) <
simdjson::SIMDJSON_PADDING);
}
simdjson::padded_string_view
get_padded_string_view(const char *buf, size_t len,
simdjson::padded_string &jsonbuffer) {
if (need_allocation(buf, len)) { // unlikely case
jsonbuffer = simdjson::padded_string(buf, len);
return jsonbuffer;
} else { // no reallcation needed (very likely)
return simdjson::padded_string_view(buf, len,
len + simdjson::SIMDJSON_PADDING);
}
}
int main() {
printf("page_size: %ld\n", page_size());
const char *jsonpoiner = R"(
{
"key": "value"
}
)";
size_t len = strlen(jsonpoiner);
simdjson::padded_string jsonbuffer; // only allocate if needed
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc;
simdjson::error_code error =
parser.iterate(get_padded_string_view(jsonpoiner, len, jsonbuffer))
.get(doc);
if (error) {
printf("error: %s\n", simdjson::error_message(error));
return EXIT_FAILURE;
}
std::string_view value;
error = doc["key"].get_string().get(value);
if (error) {
return EXIT_FAILURE;
}
printf("Value: \"%.*s\"\n", (int)value.size(), value.data());
if (value != "value") {
return EXIT_FAILURE;
}
return EXIT_SUCCESS;
}
```
+1 -1
View File
@@ -25,7 +25,7 @@ IF(${CMAKE_SYSTEM_NAME} MATCHES "Linux")
add_quickstart_test(quickstart2_noexceptions quickstart2_noexceptions.cpp NO_EXCEPTIONS LABELS acceptance)
add_quickstart_test(quickstart2_noexceptions11 quickstart2_noexceptions.cpp NO_EXCEPTIONS CXX_STANDARD c++11)
# On Demand Quick Start
# On-Demand Quick Start
if (SIMDJSON_EXCEPTIONS)
add_quickstart_test(quickstart_ondemand quickstart_ondemand.cpp LABELS quickstart_ondemand acceptance)
add_quickstart_test(quickstart_ondemand11 quickstart_ondemand.cpp CXX_STANDARD c++11 LABELS quickstart_ondemand acceptance)
+22
View File
@@ -13,6 +13,16 @@
#endif
#endif
// C++ 23
#if !defined(SIMDJSON_CPLUSPLUS23) && (SIMDJSON_CPLUSPLUS >= 202302L)
#define SIMDJSON_CPLUSPLUS23 1
#endif
// C++ 20
#if !defined(SIMDJSON_CPLUSPLUS20) && (SIMDJSON_CPLUSPLUS >= 202002L)
#define SIMDJSON_CPLUSPLUS20 1
#endif
// C++ 17
#if !defined(SIMDJSON_CPLUSPLUS17) && (SIMDJSON_CPLUSPLUS >= 201703L)
#define SIMDJSON_CPLUSPLUS17 1
@@ -40,4 +50,16 @@
#endif
#endif
#ifdef __has_include
#if __has_include(<version>)
#include <version>
#endif
#endif
#ifdef __cpp_concepts
#include <utility>
#define SIMDJSON_SUPPORTS_DESERIALIZATION 1
#else // __cpp_concepts
#define SIMDJSON_SUPPORTS_DESERIALIZATION 0
#endif
#endif // SIMDJSON_COMPILER_CHECK_H
+4
View File
@@ -123,6 +123,10 @@ inline simdjson_result<element> array::at(size_t index) const noexcept {
return INDEX_OUT_OF_BOUNDS;
}
inline array::operator element() const noexcept {
return element(tape);
}
//
// array::iterator inline implementation
//
+5
View File
@@ -126,6 +126,11 @@ public:
*/
inline simdjson_result<element> at(size_t index) const noexcept;
/**
* Implicitly convert object to element
*/
inline operator element() const noexcept;
private:
simdjson_inline array(const internal::tape_ref &tape) noexcept;
internal::tape_ref tape;
+2 -2
View File
@@ -224,10 +224,10 @@ simdjson_inline std::string_view document_stream::iterator::source() const noexc
} else {
size_t next_doc_index = stream->batch_start + stream->parser->implementation->structural_indexes[stream->parser->implementation->next_structural_index];
size_t svlen = next_doc_index - current_index();
if(svlen > 1) {
while(svlen > 1 && (std::isspace(start[svlen-1]) || start[svlen-1] == '\0')) {
svlen--;
}
return std::string_view(reinterpret_cast<const char*>(stream->buf) + current_index(), svlen);
return std::string_view(start, svlen);
}
}
+21 -1
View File
@@ -375,6 +375,23 @@ inline simdjson_result<element> element::operator[](const char *key) const noexc
return at_key(key);
}
inline bool is_pointer_well_formed(std::string_view json_pointer) noexcept {
if (simdjson_unlikely(json_pointer[0] != '/')) {
return false;
}
size_t escape = json_pointer.find('~');
if (escape == std::string_view::npos) {
return true;
}
if (escape == json_pointer.size() - 1) {
return false;
}
if (json_pointer[escape + 1] != '0' && json_pointer[escape + 1] != '1') {
return false;
}
return true;
}
inline simdjson_result<element> element::at_pointer(std::string_view json_pointer) const noexcept {
SIMDJSON_DEVELOPMENT_ASSERT(tape.usable()); // https://github.com/simdjson/simdjson/issues/1914
switch (tape.tape_ref_type()) {
@@ -383,7 +400,10 @@ inline simdjson_result<element> element::at_pointer(std::string_view json_pointe
case internal::tape_type::START_ARRAY:
return array(tape).at_pointer(json_pointer);
default: {
if(!json_pointer.empty()) { // a non-empty string is invalid on an atom
if (!json_pointer.empty()) { // a non-empty string can be invalid, or accessing a primitive (issue 2154)
if (is_pointer_well_formed(json_pointer)) {
return NO_SUCH_FIELD;
}
return INVALID_JSON_POINTER;
}
// an empty string means that we return the current node
+4
View File
@@ -153,6 +153,10 @@ inline simdjson_result<element> object::at_key_case_insensitive(std::string_view
return NO_SUCH_FIELD;
}
inline object::operator element() const noexcept {
return element(tape);
}
//
// object::iterator inline implementation
//
+5
View File
@@ -200,6 +200,11 @@ public:
*/
inline simdjson_result<element> at_key_case_insensitive(std::string_view key) const noexcept;
/**
* Implicitly convert object to element
*/
inline operator element() const noexcept;
private:
simdjson_inline object(const internal::tape_ref &tape) noexcept;
+1 -1
View File
@@ -191,7 +191,7 @@ simdjson_inline size_t parser::capacity() const noexcept {
simdjson_inline size_t parser::max_capacity() const noexcept {
return _max_capacity;
}
simdjson_inline size_t parser::max_depth() const noexcept {
simdjson_pure simdjson_inline size_t parser::max_depth() const noexcept {
return implementation ? implementation->max_depth() : DEFAULT_MAX_DEPTH;
}
+1 -1
View File
@@ -527,7 +527,7 @@ public:
*
* @return Maximum depth, in bytes.
*/
simdjson_inline size_t max_depth() const noexcept;
simdjson_pure simdjson_inline size_t max_depth() const noexcept;
/**
* Set max_capacity. This is the largest document this parser can automatically support.
+3 -3
View File
@@ -57,15 +57,15 @@ public:
simdjson_inline void one_char(char c);
simdjson_inline void call_print_newline() {
this->print_newline();
static_cast<formatter*>(this)->print_newline();
}
simdjson_inline void call_print_indents(size_t depth) {
this->print_indents(depth);
static_cast<formatter*>(this)->print_indents(depth);
}
simdjson_inline void call_print_space() {
this->print_space();
static_cast<formatter*>(this)->print_space();
}
protected:
+1 -1
View File
@@ -39,7 +39,7 @@ enum error_code {
INDEX_OUT_OF_BOUNDS, ///< JSON array index too large
NO_SUCH_FIELD, ///< JSON field not found in object
IO_ERROR, ///< Error reading a file
INVALID_JSON_POINTER, ///< Invalid JSON pointer reference
INVALID_JSON_POINTER, ///< Invalid JSON pointer syntax
INVALID_URI_FRAGMENT, ///< Invalid URI fragment
UNEXPECTED_ERROR, ///< indicative of a bug in simdjson
PARSER_IN_USE, ///< parser is already in use.
@@ -4,6 +4,7 @@
// Stuff other things depend on
#include "simdjson/generic/ondemand/base.h"
#include "simdjson/generic/ondemand/deserialize.h"
#include "simdjson/generic/ondemand/value_iterator.h"
#include "simdjson/generic/ondemand/value.h"
#include "simdjson/generic/ondemand/logger.h"
@@ -23,9 +24,13 @@
#include "simdjson/generic/ondemand/object_iterator.h"
#include "simdjson/generic/ondemand/serialization.h"
// Deserialization for standard types
#include "simdjson/generic/ondemand/std_deserialize.h"
// Inline definitions
#include "simdjson/generic/ondemand/array-inl.h"
#include "simdjson/generic/ondemand/array_iterator-inl.h"
#include "simdjson/generic/ondemand/value-inl.h"
#include "simdjson/generic/ondemand/document-inl.h"
#include "simdjson/generic/ondemand/document_stream-inl.h"
#include "simdjson/generic/ondemand/field-inl.h"
@@ -38,5 +43,6 @@
#include "simdjson/generic/ondemand/raw_json_string-inl.h"
#include "simdjson/generic/ondemand/serialization-inl.h"
#include "simdjson/generic/ondemand/token_iterator-inl.h"
#include "simdjson/generic/ondemand/value-inl.h"
#include "simdjson/generic/ondemand/value_iterator-inl.h"
@@ -0,0 +1,114 @@
#if SIMDJSON_SUPPORTS_DESERIALIZATION
#ifndef SIMDJSON_ONDEMAND_DESERIALIZE_H
#ifndef SIMDJSON_CONDITIONAL_INCLUDE
#define SIMDJSON_ONDEMAND_DESERIALIZE_H
#include "simdjson/generic/ondemand/base.h"
#include "simdjson/generic/ondemand/array.h"
#endif // SIMDJSON_CONDITIONAL_INCLUDE
#include <concepts>
namespace simdjson {
namespace tag_invoke_fn_ns {
void tag_invoke();
struct tag_invoke_fn {
template <typename Tag, typename... Args>
requires requires(Tag tag, Args &&...args) {
tag_invoke(std::forward<Tag>(tag), std::forward<Args>(args)...);
}
constexpr auto operator()(Tag tag, Args &&...args) const
noexcept(noexcept(tag_invoke(std::forward<Tag>(tag),
std::forward<Args>(args)...)))
-> decltype(tag_invoke(std::forward<Tag>(tag),
std::forward<Args>(args)...)) {
return tag_invoke(std::forward<Tag>(tag), std::forward<Args>(args)...);
}
};
} // namespace tag_invoke_fn_ns
inline namespace tag_invoke_ns {
inline constexpr tag_invoke_fn_ns::tag_invoke_fn tag_invoke = {};
} // namespace tag_invoke_ns
template <typename Tag, typename... Args>
concept tag_invocable = requires(Tag tag, Args... args) {
tag_invoke(std::forward<Tag>(tag), std::forward<Args>(args)...);
};
template <typename Tag, typename... Args>
concept nothrow_tag_invocable =
tag_invocable<Tag, Args...> && requires(Tag tag, Args... args) {
{
tag_invoke(std::forward<Tag>(tag), std::forward<Args>(args)...)
} noexcept;
};
template <typename Tag, typename... Args>
using tag_invoke_result =
std::invoke_result<decltype(tag_invoke), Tag, Args...>;
template <typename Tag, typename... Args>
using tag_invoke_result_t =
std::invoke_result_t<decltype(tag_invoke), Tag, Args...>;
template <auto &Tag> using tag_t = std::decay_t<decltype(Tag)>;
struct deserialize_tag;
/// These types are deserializable in a built-in way
template <typename> struct is_builtin_deserializable : std::false_type {};
template <> struct is_builtin_deserializable<int64_t> : std::true_type {};
template <> struct is_builtin_deserializable<uint64_t> : std::true_type {};
template <> struct is_builtin_deserializable<double> : std::true_type {};
template <> struct is_builtin_deserializable<bool> : std::true_type {};
template <> struct is_builtin_deserializable<SIMDJSON_IMPLEMENTATION::ondemand::array> : std::true_type {};
template <> struct is_builtin_deserializable<SIMDJSON_IMPLEMENTATION::ondemand::object> : std::true_type {};
template <> struct is_builtin_deserializable<SIMDJSON_IMPLEMENTATION::ondemand::value> : std::true_type {};
template <> struct is_builtin_deserializable<SIMDJSON_IMPLEMENTATION::ondemand::raw_json_string> : std::true_type {};
template <> struct is_builtin_deserializable<std::string_view> : std::true_type {};
template <typename T>
concept is_builtin_deserializable_v = is_builtin_deserializable<T>::value;
template <typename T, typename ValT = SIMDJSON_IMPLEMENTATION::ondemand::value>
concept custom_deserializable = tag_invocable<deserialize_tag, ValT&, T&>;
template <typename T, typename ValT = SIMDJSON_IMPLEMENTATION::ondemand::value>
concept deserializable = custom_deserializable<T, ValT> || is_builtin_deserializable_v<T>;
template <typename T, typename ValT = SIMDJSON_IMPLEMENTATION::ondemand::value>
concept nothrow_custom_deserializable = nothrow_tag_invocable<deserialize_tag, ValT&, T&>;
// built-in types are noexcept and if an error happens, the value simply gets ignored and the error is returned.
template <typename T, typename ValT = SIMDJSON_IMPLEMENTATION::ondemand::value>
concept nothrow_deserializable = nothrow_custom_deserializable<T, ValT> || is_builtin_deserializable_v<T>;
/// Deserialize Tag
inline constexpr struct deserialize_tag {
using value_type = SIMDJSON_IMPLEMENTATION::ondemand::value;
using document_type = SIMDJSON_IMPLEMENTATION::ondemand::document;
// Customization Point for value
template <typename T>
requires custom_deserializable<T, value_type>
[[nodiscard]] constexpr /* error_code */ auto operator()(value_type &object, T& output) const noexcept(nothrow_custom_deserializable<T, value_type>) {
return tag_invoke(*this, object, output);
}
// Customization Point for document
template <typename T>
requires custom_deserializable<T, document_type>
[[nodiscard]] constexpr /* error_code */ auto operator()(document_type &object, T& output) const noexcept(nothrow_custom_deserializable<T, document_type>) {
return tag_invoke(*this, object, output);
}
} deserialize{};
} // namespace simdjson
#endif // SIMDJSON_ONDEMAND_DESERIALIZE_H
#endif // SIMDJSON_SUPPORTS_DESERIALIZATION
@@ -13,7 +13,9 @@
#include "simdjson/generic/ondemand/object-inl.h"
#include "simdjson/generic/ondemand/raw_json_string.h"
#include "simdjson/generic/ondemand/value.h"
#include "simdjson/generic/ondemand/value-inl.h"
#include "simdjson/generic/ondemand/value_iterator-inl.h"
#include "simdjson/generic/ondemand/deserialize.h"
#endif // SIMDJSON_CONDITIONAL_INCLUDE
namespace simdjson {
@@ -167,20 +169,16 @@ template<> simdjson_inline simdjson_result<int64_t> document::get() & noexcept {
template<> simdjson_inline simdjson_result<bool> document::get() & noexcept { return get_bool(); }
template<> simdjson_inline simdjson_result<value> document::get() & noexcept { return get_value(); }
template<> simdjson_inline simdjson_result<raw_json_string> document::get() && noexcept { return get_raw_json_string(); }
template<> simdjson_inline simdjson_result<std::string_view> document::get() && noexcept { return get_string(false); }
template<> simdjson_inline simdjson_result<double> document::get() && noexcept { return std::forward<document>(*this).get_double(); }
template<> simdjson_inline simdjson_result<uint64_t> document::get() && noexcept { return std::forward<document>(*this).get_uint64(); }
template<> simdjson_inline simdjson_result<int64_t> document::get() && noexcept { return std::forward<document>(*this).get_int64(); }
template<> simdjson_inline simdjson_result<bool> document::get() && noexcept { return std::forward<document>(*this).get_bool(); }
template<> simdjson_inline simdjson_result<value> document::get() && noexcept { return get_value(); }
template<> simdjson_inline error_code document::get(array& out) & noexcept { return get_array().get(out); }
template<> simdjson_inline error_code document::get(object& out) & noexcept { return get_object().get(out); }
template<> simdjson_inline error_code document::get(raw_json_string& out) & noexcept { return get_raw_json_string().get(out); }
template<> simdjson_inline error_code document::get(std::string_view& out) & noexcept { return get_string(false).get(out); }
template<> simdjson_inline error_code document::get(double& out) & noexcept { return get_double().get(out); }
template<> simdjson_inline error_code document::get(uint64_t& out) & noexcept { return get_uint64().get(out); }
template<> simdjson_inline error_code document::get(int64_t& out) & noexcept { return get_int64().get(out); }
template<> simdjson_inline error_code document::get(bool& out) & noexcept { return get_bool().get(out); }
template<> simdjson_inline error_code document::get(value& out) & noexcept { return get_value().get(out); }
template<typename T> simdjson_inline error_code document::get(T &out) & noexcept {
return get<T>().get(out);
}
template<typename T> simdjson_inline error_code document::get(T &out) && noexcept {
return std::forward<document>(*this).get<T>().get(out);
}
#if SIMDJSON_EXCEPTIONS
template <class T>
+61 -18
View File
@@ -4,8 +4,11 @@
#define SIMDJSON_GENERIC_ONDEMAND_DOCUMENT_H
#include "simdjson/generic/ondemand/base.h"
#include "simdjson/generic/ondemand/json_iterator.h"
#include "simdjson/generic/ondemand/deserialize.h"
#include "simdjson/generic/ondemand/value.h"
#endif // SIMDJSON_CONDITIONAL_INCLUDE
namespace simdjson {
namespace SIMDJSON_IMPLEMENTATION {
namespace ondemand {
@@ -178,24 +181,39 @@ public:
* @returns A value of the given type, parsed from the JSON.
* @returns INCORRECT_TYPE If the JSON value is not the given type.
*/
template<typename T> simdjson_inline simdjson_result<T> get() & noexcept {
// Unless the simdjson library or the user provides an inline implementation, calling this method should
// immediately fail.
static_assert(!sizeof(T), "The get method with given type is not implemented by the simdjson library. "
"The supported types are ondemand::object, ondemand::array, raw_json_string, std::string_view, uint64_t, "
"int64_t, double, and bool. We recommend you use get_double(), get_bool(), get_uint64(), get_int64(), "
" get_object(), get_array(), get_raw_json_string(), or get_string() instead of the get template."
" You may also add support for custom types, see our documentation.");
template <typename T>
simdjson_inline simdjson_result<T> get() &
#if SIMDJSON_SUPPORTS_DESERIALIZATION
noexcept(custom_deserializable<T, document> ? nothrow_custom_deserializable<T, document> : true)
#else
noexcept
#endif
{
static_assert(std::is_default_constructible<T>::value, "Cannot initialize the specified type.");
T out{};
SIMDJSON_TRY(get<T>(out));
return out;
}
/** @overload template<typename T> simdjson_result<T> get() & noexcept */
template<typename T> simdjson_inline simdjson_result<T> get() && noexcept {
// Unless the simdjson library or the user provides an inline implementation, calling this method should
// immediately fail.
static_assert(!sizeof(T), "The get method with given type is not implemented by the simdjson library. "
"The supported types are ondemand::object, ondemand::array, raw_json_string, std::string_view, uint64_t, "
"int64_t, double, and bool. We recommend you use get_double(), get_bool(), get_uint64(), get_int64(), "
" get_object(), get_array(), get_raw_json_string(), or get_string() instead of the get template."
" You may also add support for custom types, see our documentation.");
/**
* @overload template<typename T> simdjson_result<T> get() & noexcept
*
* We disallow the use tag_invoke CPO on a moved document; it may create UB
* if user uses `ondemand::array` or `ondemand::object` in their custom type.
*
* The member function is still remains specialize-able for compatibility
* reasons, but we completely disallow its use when a tag_invoke customization
* is provided.
*/
template<typename T>
simdjson_inline simdjson_result<T> get() &&
#if SIMDJSON_SUPPORTS_DESERIALIZATION
noexcept(custom_deserializable<T, document> ? nothrow_custom_deserializable<T, document> : true)
#else
noexcept
#endif
{
static_assert(!std::is_same<T, array>::value && !std::is_same<T, object>::value, "You should never hold either an ondemand::array or ondemand::object without a corresponding ondemand::document being alive; that would be Undefined Behaviour.");
return static_cast<document&>(*this).get<T>();
}
/**
@@ -209,7 +227,32 @@ public:
* @returns INCORRECT_TYPE If the JSON value is not an object.
* @returns SUCCESS If the parse succeeded and the out parameter was set to the value.
*/
template<typename T> simdjson_inline error_code get(T &out) & noexcept;
template<typename T>
simdjson_inline error_code get(T &out) &
#if SIMDJSON_SUPPORTS_DESERIALIZATION
noexcept(custom_deserializable<T, document> ? nothrow_custom_deserializable<T, document> : true)
#else
noexcept
#endif
{
#if SIMDJSON_SUPPORTS_DESERIALIZATION
if constexpr (custom_deserializable<T, document>) {
return deserialize(*this, out);
} else {
#endif // SIMDJSON_SUPPORTS_DESERIALIZATION
// Unless the simdjson library or the user provides an inline implementation, calling this method should
// immediately fail.
static_assert(!sizeof(T), "The get method with given type is not implemented by the simdjson library. "
"The supported types are ondemand::object, ondemand::array, raw_json_string, std::string_view, uint64_t, "
"int64_t, double, and bool. We recommend you use get_double(), get_bool(), get_uint64(), get_int64(), "
" get_object(), get_array(), get_raw_json_string(), or get_string() instead of the get template."
" You may also add support for custom types, see our documentation.");
static_cast<void>(out); // to get rid of unused errors
return UNINITIALIZED;
#if SIMDJSON_SUPPORTS_DESERIALIZATION
}
#endif
}
/** @overload template<typename T> error_code get(T &out) & noexcept */
template<typename T> simdjson_inline error_code get(T &out) && noexcept;
@@ -348,10 +348,11 @@ simdjson_inline std::string_view document_stream::iterator::source() const noexc
auto next_index = stream->parser->implementation->structural_indexes[++cur_struct_index];
// normally the length would be next_index - current_index() - 1, except for the last document
size_t svlen = next_index - current_index();
if(svlen > 1) {
const char *start = reinterpret_cast<const char*>(stream->buf) + current_index();
while(svlen > 1 && (std::isspace(start[svlen-1]) || start[svlen-1] == '\0')) {
svlen--;
}
return std::string_view(reinterpret_cast<const char*>(stream->buf) + current_index(), svlen);
return std::string_view(start, svlen);
}
}
cur_struct_index++;
@@ -38,6 +38,14 @@ simdjson_inline simdjson_warn_unused simdjson_result<std::string_view> field::un
return answer;
}
template <typename string_type>
simdjson_inline simdjson_warn_unused error_code field::unescaped_key(string_type& receiver, bool allow_replacement) noexcept {
std::string_view key;
SIMDJSON_TRY( unescaped_key(allow_replacement).get(key) );
receiver = key;
return SUCCESS;
}
simdjson_inline raw_json_string field::key() const noexcept {
SIMDJSON_ASSUME(first.buf != nullptr); // We would like to call .alive() by Visual Studio won't let us.
return first;
@@ -105,6 +113,12 @@ simdjson_inline simdjson_result<std::string_view> simdjson_result<SIMDJSON_IMPLE
return first.unescaped_key(allow_replacement);
}
template<typename string_type>
simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::field>::unescaped_key(string_type &receiver, bool allow_replacement) noexcept {
if (error()) { return error(); }
return first.unescaped_key(receiver, allow_replacement);
}
simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value> simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::field>::value() noexcept {
if (error()) { return error(); }
return std::move(first.value());
+22 -4
View File
@@ -36,21 +36,37 @@ public:
* This consumes the key: once you have called unescaped_key(), you cannot
* call it again nor can you call key().
*/
simdjson_inline simdjson_warn_unused simdjson_result<std::string_view> unescaped_key(bool allow_replacement) noexcept;
simdjson_inline simdjson_warn_unused simdjson_result<std::string_view> unescaped_key(bool allow_replacement = false) noexcept;
/**
* Get the key as a string_view (for higher speed, consider raw_key).
* We deliberately use a more cumbersome name (unescaped_key) to force users
* to think twice about using it. The content is stored in the receiver.
*
* This consumes the key: once you have called unescaped_key(), you cannot
* call it again nor can you call key().
*/
template <typename string_type>
simdjson_inline simdjson_warn_unused error_code unescaped_key(string_type& receiver, bool allow_replacement = false) noexcept;
/**
* Get the key as a raw_json_string. Can be used for direct comparison with
* an unescaped C string: e.g., key() == "test".
* an unescaped C string: e.g., key() == "test". This does not count as
* consumption of the content: you can safely call it repeatedly.
* See escaped_key() for a similar function which returns
* a more convenient std::string_view result.
*/
simdjson_inline raw_json_string key() const noexcept;
/**
* Get the unprocessed key as a string_view. This includes the quotes and may include
* some spaces after the last quote.
* some spaces after the last quote. This does not count as
* consumption of the content: you can safely call it repeatedly.
* See escaped_key().
*/
simdjson_inline std::string_view key_raw_json_token() const noexcept;
/**
* Get the key as a string_view. This does not include the quotes and
* the string is unprocessed key so it may contain escape characters
* (e.g., \uXXXX or \n). Use unescaped_key() to get the unescaped key.
* (e.g., \uXXXX or \n). It does not count as a consumption of the content:
* you can safely call it repeatedly. Use unescaped_key() to get the unescaped key.
*/
simdjson_inline std::string_view escaped_key() const noexcept;
/**
@@ -84,6 +100,8 @@ public:
simdjson_inline simdjson_result() noexcept = default;
simdjson_inline simdjson_result<std::string_view> unescaped_key(bool allow_replacement = false) noexcept;
template<typename string_type>
simdjson_inline error_code unescaped_key(string_type &receiver, bool allow_replacement = false) noexcept;
simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::raw_json_string> key() noexcept;
simdjson_inline simdjson_result<std::string_view> key_raw_json_token() noexcept;
simdjson_inline simdjson_result<std::string_view> escaped_key() noexcept;
@@ -54,6 +54,23 @@ simdjson_inline json_iterator::json_iterator(const uint8_t *buf, ondemand::parse
#endif
}
#ifdef SIMDJSON_EXPERIMENTAL_ALLOW_INCOMPLETE_JSON
simdjson_inline json_iterator::json_iterator(const uint8_t *buf, ondemand::parser *_parser, bool streaming) noexcept
: token(buf, &_parser->implementation->structural_indexes[0]),
parser{_parser},
_string_buf_loc{parser->string_buf.get()},
_depth{1},
_root{parser->implementation->structural_indexes.get()},
_streaming{streaming}
{
logger::log_headers();
#if SIMDJSON_CHECK_EOF
assert_more_tokens();
#endif
}
#endif // SIMDJSON_EXPERIMENTAL_ALLOW_INCOMPLETE_JSON
inline void json_iterator::rewind() noexcept {
token.set_position( root_position() );
logger::log_headers(); // We start again
@@ -337,11 +354,23 @@ simdjson_inline token_position json_iterator::position() const noexcept {
}
simdjson_inline simdjson_result<std::string_view> json_iterator::unescape(raw_json_string in, bool allow_replacement) noexcept {
#if SIMDJSON_DEVELOPMENT_CHECKS
auto result = parser->unescape(in, _string_buf_loc, allow_replacement);
SIMDJSON_ASSUME(!parser->string_buffer_overflow(_string_buf_loc));
return result;
#else
return parser->unescape(in, _string_buf_loc, allow_replacement);
#endif
}
simdjson_inline simdjson_result<std::string_view> json_iterator::unescape_wobbly(raw_json_string in) noexcept {
#if SIMDJSON_DEVELOPMENT_CHECKS
auto result = parser->unescape_wobbly(in, _string_buf_loc);
SIMDJSON_ASSUME(!parser->string_buffer_overflow(_string_buf_loc));
return result;
#else
return parser->unescape_wobbly(in, _string_buf_loc);
#endif
}
simdjson_inline void json_iterator::reenter_child(token_position position, depth_t child_depth) noexcept {
@@ -293,6 +293,9 @@ public:
inline bool balanced() const noexcept;
protected:
simdjson_inline json_iterator(const uint8_t *buf, ondemand::parser *parser) noexcept;
#ifdef SIMDJSON_EXPERIMENTAL_ALLOW_INCOMPLETE_JSON
simdjson_inline json_iterator(const uint8_t *buf, ondemand::parser *parser, bool streaming) noexcept;
#endif // SIMDJSON_EXPERIMENTAL_ALLOW_INCOMPLETE_JSON
/// The last token before the end
simdjson_inline token_position last_position() const noexcept;
/// The token *at* the end. This points at gibberish and should only be used for comparison.
+2 -2
View File
@@ -160,8 +160,8 @@ public:
/**
* Reset the iterator so that we are pointing back at the
* beginning of the object. You should still consume values only once even if you
* can iterate through the object more than once. If you unescape a string within
* the object more than once, you have unsafe code. Note that rewinding an object
* can iterate through the object more than once. If you unescape a string or a key
* within the object more than once, you have unsafe code. Note that rewinding an object
* means that you may need to reparse it anew: it is not a free operation.
*
* @returns true if the object contains some elements (not empty)
+29 -3
View File
@@ -42,6 +42,11 @@ simdjson_warn_unused simdjson_inline error_code parser::allocate(size_t new_capa
_max_depth = new_max_depth;
return SUCCESS;
}
#if SIMDJSON_DEVELOPMENT_CHECKS
simdjson_inline simdjson_warn_unused bool parser::string_buffer_overflow(const uint8_t *string_buf_loc) const noexcept {
return (string_buf_loc < string_buf.get()) || (size_t(string_buf_loc - string_buf.get()) >= capacity());
}
#endif
simdjson_warn_unused simdjson_inline simdjson_result<document> parser::iterate(padded_string_view json) & noexcept {
if (json.padding() < SIMDJSON_PADDING) { return INSUFFICIENT_PADDING; }
@@ -58,6 +63,27 @@ simdjson_warn_unused simdjson_inline simdjson_result<document> parser::iterate(p
return document::start({ reinterpret_cast<const uint8_t *>(json.data()), this });
}
#ifdef SIMDJSON_EXPERIMENTAL_ALLOW_INCOMPLETE_JSON
simdjson_warn_unused simdjson_inline simdjson_result<document> parser::iterate_allow_incomplete_json(padded_string_view json) & noexcept {
if (json.padding() < SIMDJSON_PADDING) { return INSUFFICIENT_PADDING; }
json.remove_utf8_bom();
// Allocate if needed
if (capacity() < json.length() || !string_buf) {
SIMDJSON_TRY( allocate(json.length(), max_depth()) );
}
// Run stage 1.
const simdjson::error_code err = implementation->stage1(reinterpret_cast<const uint8_t *>(json.data()), json.length(), stage1_mode::regular);
if (err) {
if (err != UNCLOSED_STRING)
return err;
}
return document::start({ reinterpret_cast<const uint8_t *>(json.data()), this, true });
}
#endif // SIMDJSON_EXPERIMENTAL_ALLOW_INCOMPLETE_JSON
simdjson_warn_unused simdjson_inline simdjson_result<document> parser::iterate(const char *json, size_t len, size_t allocated) & noexcept {
return iterate(padded_string_view(json, len, allocated));
}
@@ -129,13 +155,13 @@ inline simdjson_result<document_stream> parser::iterate_many(const padded_string
return iterate_many(s.data(), s.length(), batch_size, allow_comma_separated);
}
simdjson_inline size_t parser::capacity() const noexcept {
simdjson_pure simdjson_inline size_t parser::capacity() const noexcept {
return _capacity;
}
simdjson_inline size_t parser::max_capacity() const noexcept {
simdjson_pure simdjson_inline size_t parser::max_capacity() const noexcept {
return _max_capacity;
}
simdjson_inline size_t parser::max_depth() const noexcept {
simdjson_pure simdjson_inline size_t parser::max_depth() const noexcept {
return _max_depth;
}
+19 -5
View File
@@ -13,8 +13,8 @@ namespace SIMDJSON_IMPLEMENTATION {
namespace ondemand {
/**
* The default batch size for document_stream instances for this On Demand kernel.
* Note that different On Demand kernel may use a different DEFAULT_BATCH_SIZE value
* The default batch size for document_stream instances for this On-Demand kernel.
* Note that different On-Demand kernel may use a different DEFAULT_BATCH_SIZE value
* in the future.
*/
static constexpr size_t DEFAULT_BATCH_SIZE = 1000000;
@@ -98,6 +98,9 @@ public:
* - UNCLOSED_STRING if there is an unclosed string in the document.
*/
simdjson_warn_unused simdjson_result<document> iterate(padded_string_view json) & noexcept;
#ifdef SIMDJSON_EXPERIMENTAL_ALLOW_INCOMPLETE_JSON
simdjson_warn_unused simdjson_result<document> iterate_allow_incomplete_json(padded_string_view json) & noexcept;
#endif // SIMDJSON_EXPERIMENTAL_ALLOW_INCOMPLETE_JSON
/** @overload simdjson_result<document> iterate(padded_string_view json) & noexcept */
simdjson_warn_unused simdjson_result<document> iterate(const char *json, size_t len, size_t capacity) & noexcept;
/** @overload simdjson_result<document> iterate(padded_string_view json) & noexcept */
@@ -242,9 +245,9 @@ public:
simdjson_result<document_stream> iterate_many(const char *buf, size_t batch_size = DEFAULT_BATCH_SIZE) noexcept = delete;
/** The capacity of this parser (the largest document it can process). */
simdjson_inline size_t capacity() const noexcept;
simdjson_pure simdjson_inline size_t capacity() const noexcept;
/** The maximum capacity of this parser (the largest document it is allowed to process). */
simdjson_inline size_t max_capacity() const noexcept;
simdjson_pure simdjson_inline size_t max_capacity() const noexcept;
simdjson_inline void set_max_capacity(size_t max_capacity) noexcept;
/**
* The maximum depth of this parser (the most deeply nested objects and arrays it can process).
@@ -252,7 +255,7 @@ public:
* The document's instance current_depth() method should be used to monitor the parsing
* depth and limit it if desired.
*/
simdjson_inline size_t max_depth() const noexcept;
simdjson_pure simdjson_inline size_t max_depth() const noexcept;
/**
* Ensure this parser has enough memory to process JSON documents up to `capacity` bytes in length
@@ -324,6 +327,17 @@ public:
*/
simdjson_inline simdjson_result<std::string_view> unescape_wobbly(raw_json_string in, uint8_t *&dst) const noexcept;
#if SIMDJSON_DEVELOPMENT_CHECKS
/**
* Returns true if string_buf_loc is outside of the allocated range for the
* the string buffer. When true, it indicates that the string buffer has overflowed.
* This is a development-time check that is not needed in production. It can be
* used to detect buffer overflows in the string buffer and usafe usage of the
* string buffer.
*/
bool string_buffer_overflow(const uint8_t *string_buf_loc) const noexcept;
#endif
private:
/** @private [for benchmarking access] The implementation to use */
std::unique_ptr<internal::dom_parser_implementation> implementation{};
@@ -0,0 +1,196 @@
#if SIMDJSON_SUPPORTS_DESERIALIZATION
#ifndef SIMDJSON_ONDEMAND_DESERIALIZE_H
#ifndef SIMDJSON_CONDITIONAL_INCLUDE
#define SIMDJSON_ONDEMAND_DESERIALIZE_H
#include "simdjson/generic/ondemand/base.h"
#include "simdjson/generic/ondemand/array.h"
#endif // SIMDJSON_CONDITIONAL_INCLUDE
#include <concepts>
#include <limits>
#include <list>
#include <memory>
#include <string>
#include <vector>
namespace simdjson {
//////////////////////////////
// Number deserialization
//////////////////////////////
template <std::unsigned_integral T>
error_code tag_invoke(deserialize_tag, auto &val, T &out) noexcept {
using limits = std::numeric_limits<T>;
uint64_t x;
SIMDJSON_TRY(val.get_uint64().get(x));
if (x > (limits::max)()) {
return NUMBER_OUT_OF_RANGE;
}
out = static_cast<T>(x);
return SUCCESS;
}
template <std::floating_point T>
error_code tag_invoke(deserialize_tag, auto &val, T &out) noexcept {
double x;
SIMDJSON_TRY(val.get_double().get(x));
out = static_cast<T>(x);
return SUCCESS;
}
template <std::signed_integral T>
error_code tag_invoke(deserialize_tag, auto &val, T &out) noexcept {
using limits = std::numeric_limits<T>;
int64_t x;
SIMDJSON_TRY(val.get_int64().get(x));
if (x > (limits::max)() || x < (limits::min)()) {
return NUMBER_OUT_OF_RANGE;
}
out = static_cast<T>(x);
return SUCCESS;
}
//////////////////////////////
// STL list deserialization
//////////////////////////////
template <typename T, typename AllocT, typename ValT>
error_code tag_invoke(deserialize_tag, ValT &val,
std::list<T, AllocT> &out) noexcept(false) {
// For better error messages, don't use these as constraints on
// the tag_invoke CPO.
static_assert(
deserializable<T, ValT>,
"The specified type inside the list must itself be deserializable");
static_assert(
std::is_default_constructible_v<T>,
"The specified type inside the list must default constructible.");
SIMDJSON_IMPLEMENTATION::ondemand::array arr;
SIMDJSON_TRY(val.get_array().get(arr));
for (auto v : arr) {
if (auto const err = v.get<T>().get(out.emplace_back()); err) {
// If an error occurs, the empty element that we just inserted gets
// removed. We're not using a temp variable because if T is a heavy type,
// we want the valid path to be the fast path and the slow path be the
// path that has errors in it.
static_cast<void>(out.pop_back());
return err;
}
}
return SUCCESS;
}
//////////////////////////////
// std::string deserialization
//////////////////////////////
template <typename CharT,
typename TraitsT,
typename AllocT,
typename ValT>
error_code tag_invoke(deserialize_tag, ValT &val, std::basic_string<CharT, TraitsT, AllocT> &out) noexcept(false) {
using string_type = std::basic_string<CharT, TraitsT, AllocT>;
if constexpr (std::same_as<string_type, string_type>) {
SIMDJSON_TRY(val.get_string(out));
} else {
// todo: optimize performance
std::string tmp;
SIMDJSON_TRY(val.get_string(tmp));
for (auto const ch : tmp) {
out.push_back(ch);
}
}
return SUCCESS;
}
//////////////////////////////
// STL Vector deserialization
//////////////////////////////
/**
* STL containers have several constructors including one that takes a single
* size argument. Thus, some compilers (Visual Studio) will not be able to
* disambiguate between the size and container constructor. Users should
* explicitly specify the type of the container as needed: e.g.,
* doc.get<std::vector<int>>().
*/
template <typename T, typename AllocT, typename ValT>
error_code tag_invoke(deserialize_tag, ValT &val,
std::vector<T, AllocT> &out) noexcept(false) {
// For better error messages, don't use these as constraints on
// the tag_invoke CPO.
static_assert(
deserializable<T, ValT>,
"The specified type inside the vector must itself be deserializable");
static_assert(
std::is_default_constructible_v<T>,
"The specified type inside the vector must default constructible.");
SIMDJSON_IMPLEMENTATION::ondemand::array arr;
SIMDJSON_TRY(val.get_array().get(arr));
for (auto v : arr) {
if (auto const err = v.get<T>().get(out.emplace_back()); err) {
// If an error occurs, the empty element that we just inserted gets
// removed. We're not using a temp variable because if T is a heavy type,
// we want the valid path to be the fast path and the slow path be the
// path that has errors in it.
static_cast<void>(out.pop_back());
return err;
}
}
return SUCCESS;
}
//////////////////////////////
// std::unique_ptr deserialization
//////////////////////////////
/**
* This CPO (Customization Point Object) will help deserialize into
* `unique_ptr`s.
*
* If constructing T is nothrow, this conversion should be nothrow as well since
* we return MEMALLOC if we're not able to allocate memory instead of throwing
* the the error message.
*
* @tparam T The type inside the unique_ptr
* @tparam Deleter The Deleter of the unique_ptr
* @tparam ValT document/value type
* @param val document/value
* @param out output unique_ptr
* @return status of the conversion
*/
template <typename T, typename Deleter, typename ValT>
error_code tag_invoke(deserialize_tag, ValT &val,
std::unique_ptr<T, Deleter>
&out) noexcept(nothrow_deserializable<T, ValT>) {
// For better error messages, don't use these as constraints on
// the tag_invoke CPO.
static_assert(
deserializable<T, ValT>,
"The specified type inside the unique_ptr must itself be deserializable");
static_assert(
std::is_default_constructible_v<T>,
"The specified type inside the unique_ptr must default constructible.");
auto ptr = new (std::nothrow) T();
if (ptr == nullptr) {
return MEMALLOC;
}
SIMDJSON_TRY(val.template get<T>(*ptr));
out.reset(ptr);
return SUCCESS;
}
} // namespace simdjson
#endif // SIMDJSON_ONDEMAND_DESERIALIZE_H
#endif // SIMDJSON_SUPPORTS_DESERIALIZATION
+42 -9
View File
@@ -80,6 +80,7 @@ simdjson_inline simdjson_result<bool> value::get_bool() noexcept {
simdjson_inline simdjson_result<bool> value::is_null() noexcept {
return iter.is_null();
}
template<> simdjson_inline simdjson_result<array> value::get() noexcept { return get_array(); }
template<> simdjson_inline simdjson_result<object> value::get() noexcept { return get_object(); }
template<> simdjson_inline simdjson_result<raw_json_string> value::get() noexcept { return get_raw_json_string(); }
@@ -90,9 +91,16 @@ template<> simdjson_inline simdjson_result<uint64_t> value::get() noexcept { ret
template<> simdjson_inline simdjson_result<int64_t> value::get() noexcept { return get_int64(); }
template<> simdjson_inline simdjson_result<bool> value::get() noexcept { return get_bool(); }
template<typename T> simdjson_inline error_code value::get(T &out) noexcept {
return get<T>().get(out);
}
template<> simdjson_inline error_code value::get(array& out) noexcept { return get_array().get(out); }
template<> simdjson_inline error_code value::get(object& out) noexcept { return get_object().get(out); }
template<> simdjson_inline error_code value::get(raw_json_string& out) noexcept { return get_raw_json_string().get(out); }
template<> simdjson_inline error_code value::get(std::string_view& out) noexcept { return get_string(false).get(out); }
template<> simdjson_inline error_code value::get(number& out) noexcept { return get_number().get(out); }
template<> simdjson_inline error_code value::get(double& out) noexcept { return get_double().get(out); }
template<> simdjson_inline error_code value::get(uint64_t& out) noexcept { return get_uint64().get(out); }
template<> simdjson_inline error_code value::get(int64_t& out) noexcept { return get_int64().get(out); }
template<> simdjson_inline error_code value::get(bool& out) noexcept { return get_bool().get(out); }
#if SIMDJSON_EXCEPTIONS
template <class T>
@@ -239,6 +247,26 @@ simdjson_inline int32_t value::current_depth() const noexcept{
return iter.json_iter().depth();
}
inline bool is_pointer_well_formed(std::string_view json_pointer) noexcept {
if (simdjson_unlikely(json_pointer.empty())) { // can't be
return false;
}
if (simdjson_unlikely(json_pointer[0] != '/')) {
return false;
}
size_t escape = json_pointer.find('~');
if (escape == std::string_view::npos) {
return true;
}
if (escape == json_pointer.size() - 1) {
return false;
}
if (json_pointer[escape + 1] != '0' && json_pointer[escape + 1] != '1') {
return false;
}
return true;
}
simdjson_inline simdjson_result<value> value::at_pointer(std::string_view json_pointer) noexcept {
json_type t;
SIMDJSON_TRY(type().get(t));
@@ -249,6 +277,10 @@ simdjson_inline simdjson_result<value> value::at_pointer(std::string_view json_p
case json_type::object:
return (*this).get_object().at_pointer(json_pointer);
default:
// a non-empty string can be invalid, or accessing a primitive (issue 2154)
if (is_pointer_well_formed(json_pointer)) {
return NO_SUCH_FIELD;
}
return INVALID_JSON_POINTER;
}
}
@@ -392,6 +424,12 @@ simdjson_inline simdjson_result<bool> simdjson_result<SIMDJSON_IMPLEMENTATION::o
return first.is_null();
}
template<> simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value>::get<SIMDJSON_IMPLEMENTATION::ondemand::value>(SIMDJSON_IMPLEMENTATION::ondemand::value &out) noexcept {
if (error()) { return error(); }
out = first;
return SUCCESS;
}
template<typename T> simdjson_inline simdjson_result<T> simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value>::get() noexcept {
if (error()) { return error(); }
return first.get<T>();
@@ -405,11 +443,6 @@ template<> simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::va
if (error()) { return error(); }
return std::move(first);
}
template<> simdjson_inline error_code simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value>::get<SIMDJSON_IMPLEMENTATION::ondemand::value>(SIMDJSON_IMPLEMENTATION::ondemand::value &out) noexcept {
if (error()) { return error(); }
out = first;
return SUCCESS;
}
simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::json_type> simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value>::type() noexcept {
if (error()) { return error(); }
@@ -443,7 +476,7 @@ simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::number> simdj
template <class T>
simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value>::operator T() noexcept(false) {
if (error()) { throw simdjson_error(error()); }
return static_cast<T>(first);
return first.get<T>();
}
simdjson_inline simdjson_result<SIMDJSON_IMPLEMENTATION::ondemand::value>::operator SIMDJSON_IMPLEMENTATION::ondemand::array() noexcept(false) {
if (error()) { throw simdjson_error(error()); }
+43 -10
View File
@@ -5,12 +5,15 @@
#include "simdjson/generic/ondemand/base.h"
#include "simdjson/generic/implementation_simdjson_result_base.h"
#include "simdjson/generic/ondemand/value_iterator.h"
#include "simdjson/generic/ondemand/deserialize.h"
#endif // SIMDJSON_CONDITIONAL_INCLUDE
#include <type_traits>
namespace simdjson {
namespace SIMDJSON_IMPLEMENTATION {
namespace ondemand {
/**
* An ephemeral JSON value returned during iteration. It is only valid for as long as you do
* not access more data in the JSON document.
@@ -35,16 +38,21 @@ public:
* @returns A value of the given type, parsed from the JSON.
* @returns INCORRECT_TYPE If the JSON value is not the given type.
*/
template<typename T> simdjson_inline simdjson_result<T> get() noexcept {
// Unless the simdjson library or the user provides an inline implementation, calling this method should
// immediately fail.
static_assert(!sizeof(T), "The get method with given type is not implemented by the simdjson library. "
"The supported types are ondemand::object, ondemand::array, raw_json_string, std::string_view, uint64_t, "
"int64_t, double, and bool. We recommend you use get_double(), get_bool(), get_uint64(), get_int64(), "
" get_object(), get_array(), get_raw_json_string(), or get_string() instead of the get template."
" You may also add support for custom types, see our documentation.");
template <typename T>
simdjson_inline simdjson_result<T> get()
#if SIMDJSON_SUPPORTS_DESERIALIZATION
noexcept(custom_deserializable<T, value> ? nothrow_custom_deserializable<T, value> : true)
#else
noexcept
#endif
{
static_assert(std::is_default_constructible<T>::value, "The specified type is not default constructible.");
T out{};
SIMDJSON_TRY(get<T>(out));
return out;
}
/**
* Get this value as the given type.
*
@@ -54,7 +62,32 @@ public:
* @returns INCORRECT_TYPE If the JSON value is not an object.
* @returns SUCCESS If the parse succeeded and the out parameter was set to the value.
*/
template<typename T> simdjson_inline error_code get(T &out) noexcept;
template <typename T>
simdjson_inline error_code get(T &out)
#if SIMDJSON_SUPPORTS_DESERIALIZATION
noexcept(custom_deserializable<T, value> ? nothrow_custom_deserializable<T, value> : true)
#else
noexcept
#endif
{
#if SIMDJSON_SUPPORTS_DESERIALIZATION
if constexpr (custom_deserializable<T, value>) {
return deserialize(*this, out);
} else {
#endif // SIMDJSON_SUPPORTS_DESERIALIZATION
// Unless the simdjson library or the user provides an inline implementation, calling this method should
// immediately fail.
static_assert(!sizeof(T), "The get method with given type is not implemented by the simdjson library. "
"The supported types are ondemand::object, ondemand::array, raw_json_string, std::string_view, uint64_t, "
"int64_t, double, and bool. We recommend you use get_double(), get_bool(), get_uint64(), get_int64(), "
" get_object(), get_array(), get_raw_json_string(), or get_string() instead of the get template."
" You may also add support for custom types, see our documentation.");
static_cast<void>(out); // to get rid of unused errors
return UNINITIALIZED;
#if SIMDJSON_SUPPORTS_DESERIALIZATION
}
#endif
}
/**
* Cast this JSON value to an array.
@@ -106,7 +106,7 @@ public:
* Unescape a valid UTF-8 string from src to dst, stopping at a final unescaped quote. There
* must be an unescaped quote terminating the string. It returns the final output
* position as pointer. In case of error (e.g., the string has bad escaped codes),
* then null_nullptrptr is returned. It is assumed that the output buffer is large
* then null_ptr is returned. It is assumed that the output buffer is large
* enough. E.g., if src points at 'joe"', then dst needs to have four free bytes +
* SIMDJSON_PADDING bytes.
*
@@ -123,7 +123,7 @@ public:
* Unescape a NON-valid UTF-8 string from src to dst, stopping at a final unescaped quote. There
* must be an unescaped quote terminating the string. It returns the final output
* position as pointer. In case of error (e.g., the string has bad escaped codes),
* then null_nullptrptr is returned. It is assumed that the output buffer is large
* then null_ptr is returned. It is assumed that the output buffer is large
* enough. E.g., if src points at 'joe"', then dst needs to have four free bytes +
* SIMDJSON_PADDING bytes.
*
@@ -177,14 +177,14 @@ public:
*
* @return Current capacity, in bytes.
*/
simdjson_inline size_t capacity() const noexcept;
simdjson_pure simdjson_inline size_t capacity() const noexcept;
/**
* The maximum level of nested object and arrays supported by this parser.
*
* @return Maximum depth, in bytes.
*/
simdjson_inline size_t max_depth() const noexcept;
simdjson_pure simdjson_inline size_t max_depth() const noexcept;
/**
* Ensure this parser has enough memory to process JSON documents up to `capacity` bytes in length
@@ -225,11 +225,11 @@ simdjson_inline dom_parser_implementation::dom_parser_implementation() noexcept
simdjson_inline dom_parser_implementation::dom_parser_implementation(dom_parser_implementation &&other) noexcept = default;
simdjson_inline dom_parser_implementation &dom_parser_implementation::operator=(dom_parser_implementation &&other) noexcept = default;
simdjson_inline size_t dom_parser_implementation::capacity() const noexcept {
simdjson_pure simdjson_inline size_t dom_parser_implementation::capacity() const noexcept {
return _capacity;
}
simdjson_inline size_t dom_parser_implementation::max_depth() const noexcept {
simdjson_pure simdjson_inline size_t dom_parser_implementation::max_depth() const noexcept {
return _max_depth;
}
+4 -2
View File
@@ -6,13 +6,13 @@
// Distributed under the Boost Software License, Version 1.0.
// (See accompanying file LICENSE.txt or copy at http://www.boost.org/LICENSE_1_0.txt)
// #pragma once // We remove #pragma once here as it generates a warning in some cases. We rely on the include guard.
#pragma once
#ifndef NONSTD_SV_LITE_H_INCLUDED
#define NONSTD_SV_LITE_H_INCLUDED
#define string_view_lite_MAJOR 1
#define string_view_lite_MINOR 7
#define string_view_lite_MINOR 8
#define string_view_lite_PATCH 0
#define string_view_lite_VERSION nssv_STRINGIFY(string_view_lite_MAJOR) "." nssv_STRINGIFY(string_view_lite_MINOR) "." nssv_STRINGIFY(string_view_lite_PATCH)
@@ -134,6 +134,8 @@
#if nssv_CONFIG_CONVERSION_STD_STRING_FREE_FUNCTIONS
#include <string>
namespace nonstd {
template< class CharT, class Traits, class Allocator = std::allocator<CharT> >
+1
View File
@@ -84,6 +84,7 @@ inline padded_string::padded_string(std::string_view sv_) noexcept
inline padded_string::padded_string(padded_string &&o) noexcept
: viable_size(o.viable_size), data_ptr(o.data_ptr) {
o.data_ptr = nullptr; // we take ownership
o.viable_size = 0;
}
inline padded_string &padded_string::operator=(padded_string &&o) noexcept {
+8
View File
@@ -11,6 +11,9 @@
#include <strings.h>
#endif
// We are using size_t without namespace std:: throughout the project
using std::size_t;
#ifdef _MSC_VER
#define SIMDJSON_VISUAL_STUDIO 1
/**
@@ -148,6 +151,11 @@
#define SIMDJSON_NO_SANITIZE_UNDEFINED
#endif
#if defined(__clang__) || defined(__GNUC__)
#define simdjson_pure [[gnu::pure]]
#else
#define simdjson_pure
#endif
#if defined(__clang__) || defined(__GNUC__)
#if defined(__has_feature)
+3 -3
View File
@@ -4,7 +4,7 @@
#define SIMDJSON_SIMDJSON_VERSION_H
/** The version of simdjson being used (major.minor.revision) */
#define SIMDJSON_VERSION "3.9.2"
#define SIMDJSON_VERSION "3.10.0"
namespace simdjson {
enum {
@@ -15,11 +15,11 @@ enum {
/**
* The minor version (major.MINOR.revision) of simdjson being used.
*/
SIMDJSON_VERSION_MINOR = 9,
SIMDJSON_VERSION_MINOR = 10,
/**
* The revision (major.minor.REVISION) of simdjson being used.
*/
SIMDJSON_VERSION_REVISION = 2
SIMDJSON_VERSION_REVISION = 0
};
} // namespace simdjson
+58 -25
View File
@@ -1,4 +1,4 @@
/* auto-generated on 2024-05-07 18:04:59 -0400. Do not edit! */
/* auto-generated on 2024-07-15 08:51:57 -0400. Do not edit! */
/* including simdjson.cpp: */
/* begin file simdjson.cpp */
#define SIMDJSON_SRC_SIMDJSON_CPP
@@ -40,6 +40,16 @@
#endif
#endif
// C++ 23
#if !defined(SIMDJSON_CPLUSPLUS23) && (SIMDJSON_CPLUSPLUS >= 202302L)
#define SIMDJSON_CPLUSPLUS23 1
#endif
// C++ 20
#if !defined(SIMDJSON_CPLUSPLUS20) && (SIMDJSON_CPLUSPLUS >= 202002L)
#define SIMDJSON_CPLUSPLUS20 1
#endif
// C++ 17
#if !defined(SIMDJSON_CPLUSPLUS17) && (SIMDJSON_CPLUSPLUS >= 201703L)
#define SIMDJSON_CPLUSPLUS17 1
@@ -67,6 +77,18 @@
#endif
#endif
#ifdef __has_include
#if __has_include(<version>)
#include <version>
#endif
#endif
#ifdef __cpp_concepts
#include <utility>
#define SIMDJSON_SUPPORTS_DESERIALIZATION 1
#else // __cpp_concepts
#define SIMDJSON_SUPPORTS_DESERIALIZATION 0
#endif
#endif // SIMDJSON_COMPILER_CHECK_H
/* end file simdjson/compiler_check.h */
/* including simdjson/portability.h: #include "simdjson/portability.h" */
@@ -84,6 +106,9 @@
#include <strings.h>
#endif
// We are using size_t without namespace std:: throughout the project
using std::size_t;
#ifdef _MSC_VER
#define SIMDJSON_VISUAL_STUDIO 1
/**
@@ -221,6 +246,11 @@
#define SIMDJSON_NO_SANITIZE_UNDEFINED
#endif
#if defined(__clang__) || defined(__GNUC__)
#define simdjson_pure [[gnu::pure]]
#else
#define simdjson_pure
#endif
#if defined(__clang__) || defined(__GNUC__)
#if defined(__has_feature)
@@ -527,13 +557,13 @@ SIMDJSON_PUSH_DISABLE_ALL_WARNINGS
// Distributed under the Boost Software License, Version 1.0.
// (See accompanying file LICENSE.txt or copy at http://www.boost.org/LICENSE_1_0.txt)
// #pragma once // We remove #pragma once here as it generates a warning in some cases. We rely on the include guard.
#pragma once
#ifndef NONSTD_SV_LITE_H_INCLUDED
#define NONSTD_SV_LITE_H_INCLUDED
#define string_view_lite_MAJOR 1
#define string_view_lite_MINOR 7
#define string_view_lite_MINOR 8
#define string_view_lite_PATCH 0
#define string_view_lite_VERSION nssv_STRINGIFY(string_view_lite_MAJOR) "." nssv_STRINGIFY(string_view_lite_MINOR) "." nssv_STRINGIFY(string_view_lite_PATCH)
@@ -655,6 +685,8 @@ SIMDJSON_PUSH_DISABLE_ALL_WARNINGS
#if nssv_CONFIG_CONVERSION_STD_STRING_FREE_FUNCTIONS
#include <string>
namespace nonstd {
template< class CharT, class Traits, class Allocator = std::allocator<CharT> >
@@ -2359,7 +2391,7 @@ enum error_code {
INDEX_OUT_OF_BOUNDS, ///< JSON array index too large
NO_SUCH_FIELD, ///< JSON field not found in object
IO_ERROR, ///< Error reading a file
INVALID_JSON_POINTER, ///< Invalid JSON pointer reference
INVALID_JSON_POINTER, ///< Invalid JSON pointer syntax
INVALID_URI_FRAGMENT, ///< Invalid URI fragment
UNEXPECTED_ERROR, ///< indicative of a bug in simdjson
PARSER_IN_USE, ///< parser is already in use.
@@ -5944,7 +5976,7 @@ public:
* Unescape a valid UTF-8 string from src to dst, stopping at a final unescaped quote. There
* must be an unescaped quote terminating the string. It returns the final output
* position as pointer. In case of error (e.g., the string has bad escaped codes),
* then null_nullptrptr is returned. It is assumed that the output buffer is large
* then null_ptr is returned. It is assumed that the output buffer is large
* enough. E.g., if src points at 'joe"', then dst needs to have four free bytes +
* SIMDJSON_PADDING bytes.
*
@@ -5961,7 +5993,7 @@ public:
* Unescape a NON-valid UTF-8 string from src to dst, stopping at a final unescaped quote. There
* must be an unescaped quote terminating the string. It returns the final output
* position as pointer. In case of error (e.g., the string has bad escaped codes),
* then null_nullptrptr is returned. It is assumed that the output buffer is large
* then null_ptr is returned. It is assumed that the output buffer is large
* enough. E.g., if src points at 'joe"', then dst needs to have four free bytes +
* SIMDJSON_PADDING bytes.
*
@@ -6015,14 +6047,14 @@ public:
*
* @return Current capacity, in bytes.
*/
simdjson_inline size_t capacity() const noexcept;
simdjson_pure simdjson_inline size_t capacity() const noexcept;
/**
* The maximum level of nested object and arrays supported by this parser.
*
* @return Maximum depth, in bytes.
*/
simdjson_inline size_t max_depth() const noexcept;
simdjson_pure simdjson_inline size_t max_depth() const noexcept;
/**
* Ensure this parser has enough memory to process JSON documents up to `capacity` bytes in length
@@ -6063,11 +6095,11 @@ simdjson_inline dom_parser_implementation::dom_parser_implementation() noexcept
simdjson_inline dom_parser_implementation::dom_parser_implementation(dom_parser_implementation &&other) noexcept = default;
simdjson_inline dom_parser_implementation &dom_parser_implementation::operator=(dom_parser_implementation &&other) noexcept = default;
simdjson_inline size_t dom_parser_implementation::capacity() const noexcept {
simdjson_pure simdjson_inline size_t dom_parser_implementation::capacity() const noexcept {
return _capacity;
}
simdjson_inline size_t dom_parser_implementation::max_depth() const noexcept {
simdjson_pure simdjson_inline size_t dom_parser_implementation::max_depth() const noexcept {
return _max_depth;
}
@@ -6896,6 +6928,7 @@ static inline uint32_t detect_supported_architectures() {
/* end file internal/isadetection.h */
#include <initializer_list>
#include <type_traits>
namespace simdjson {
@@ -12471,7 +12504,7 @@ simdjson_inline error_code json_structural_indexer::finish(dom_parser_implementa
}
parser.n_structural_indexes = uint32_t(indexer.tail - parser.structural_indexes.get());
/***
* The On Demand API requires special padding.
* The On-Demand API requires special padding.
*
* This is related to https://github.com/simdjson/simdjson/issues/906
* Basically, we want to make sure that if the parsing continues beyond the last (valid)
@@ -13346,7 +13379,7 @@ simdjson_inline bool handle_unicode_codepoint_wobbly(const uint8_t **src_ptr,
* Unescape a valid UTF-8 string from src to dst, stopping at a final unescaped quote. There
* must be an unescaped quote terminating the string. It returns the final output
* position as pointer. In case of error (e.g., the string has bad escaped codes),
* then null_nullptrptr is returned. It is assumed that the output buffer is large
* then null_ptr is returned. It is assumed that the output buffer is large
* enough. E.g., if src points at 'joe"', then dst needs to have four free bytes +
* SIMDJSON_PADDING bytes.
*/
@@ -18690,7 +18723,7 @@ simdjson_inline error_code json_structural_indexer::finish(dom_parser_implementa
}
parser.n_structural_indexes = uint32_t(indexer.tail - parser.structural_indexes.get());
/***
* The On Demand API requires special padding.
* The On-Demand API requires special padding.
*
* This is related to https://github.com/simdjson/simdjson/issues/906
* Basically, we want to make sure that if the parsing continues beyond the last (valid)
@@ -19565,7 +19598,7 @@ simdjson_inline bool handle_unicode_codepoint_wobbly(const uint8_t **src_ptr,
* Unescape a valid UTF-8 string from src to dst, stopping at a final unescaped quote. There
* must be an unescaped quote terminating the string. It returns the final output
* position as pointer. In case of error (e.g., the string has bad escaped codes),
* then null_nullptrptr is returned. It is assumed that the output buffer is large
* then null_ptr is returned. It is assumed that the output buffer is large
* enough. E.g., if src points at 'joe"', then dst needs to have four free bytes +
* SIMDJSON_PADDING bytes.
*/
@@ -24902,7 +24935,7 @@ simdjson_inline error_code json_structural_indexer::finish(dom_parser_implementa
}
parser.n_structural_indexes = uint32_t(indexer.tail - parser.structural_indexes.get());
/***
* The On Demand API requires special padding.
* The On-Demand API requires special padding.
*
* This is related to https://github.com/simdjson/simdjson/issues/906
* Basically, we want to make sure that if the parsing continues beyond the last (valid)
@@ -25777,7 +25810,7 @@ simdjson_inline bool handle_unicode_codepoint_wobbly(const uint8_t **src_ptr,
* Unescape a valid UTF-8 string from src to dst, stopping at a final unescaped quote. There
* must be an unescaped quote terminating the string. It returns the final output
* position as pointer. In case of error (e.g., the string has bad escaped codes),
* then null_nullptrptr is returned. It is assumed that the output buffer is large
* then null_ptr is returned. It is assumed that the output buffer is large
* enough. E.g., if src points at 'joe"', then dst needs to have four free bytes +
* SIMDJSON_PADDING bytes.
*/
@@ -31385,7 +31418,7 @@ simdjson_inline error_code json_structural_indexer::finish(dom_parser_implementa
}
parser.n_structural_indexes = uint32_t(indexer.tail - parser.structural_indexes.get());
/***
* The On Demand API requires special padding.
* The On-Demand API requires special padding.
*
* This is related to https://github.com/simdjson/simdjson/issues/906
* Basically, we want to make sure that if the parsing continues beyond the last (valid)
@@ -32260,7 +32293,7 @@ simdjson_inline bool handle_unicode_codepoint_wobbly(const uint8_t **src_ptr,
* Unescape a valid UTF-8 string from src to dst, stopping at a final unescaped quote. There
* must be an unescaped quote terminating the string. It returns the final output
* position as pointer. In case of error (e.g., the string has bad escaped codes),
* then null_nullptrptr is returned. It is assumed that the output buffer is large
* then null_ptr is returned. It is assumed that the output buffer is large
* enough. E.g., if src points at 'joe"', then dst needs to have four free bytes +
* SIMDJSON_PADDING bytes.
*/
@@ -38442,7 +38475,7 @@ simdjson_inline error_code json_structural_indexer::finish(dom_parser_implementa
}
parser.n_structural_indexes = uint32_t(indexer.tail - parser.structural_indexes.get());
/***
* The On Demand API requires special padding.
* The On-Demand API requires special padding.
*
* This is related to https://github.com/simdjson/simdjson/issues/906
* Basically, we want to make sure that if the parsing continues beyond the last (valid)
@@ -39317,7 +39350,7 @@ simdjson_inline bool handle_unicode_codepoint_wobbly(const uint8_t **src_ptr,
* Unescape a valid UTF-8 string from src to dst, stopping at a final unescaped quote. There
* must be an unescaped quote terminating the string. It returns the final output
* position as pointer. In case of error (e.g., the string has bad escaped codes),
* then null_nullptrptr is returned. It is assumed that the output buffer is large
* then null_ptr is returned. It is assumed that the output buffer is large
* enough. E.g., if src points at 'joe"', then dst needs to have four free bytes +
* SIMDJSON_PADDING bytes.
*/
@@ -44466,7 +44499,7 @@ simdjson_inline error_code json_structural_indexer::finish(dom_parser_implementa
}
parser.n_structural_indexes = uint32_t(indexer.tail - parser.structural_indexes.get());
/***
* The On Demand API requires special padding.
* The On-Demand API requires special padding.
*
* This is related to https://github.com/simdjson/simdjson/issues/906
* Basically, we want to make sure that if the parsing continues beyond the last (valid)
@@ -45341,7 +45374,7 @@ simdjson_inline bool handle_unicode_codepoint_wobbly(const uint8_t **src_ptr,
* Unescape a valid UTF-8 string from src to dst, stopping at a final unescaped quote. There
* must be an unescaped quote terminating the string. It returns the final output
* position as pointer. In case of error (e.g., the string has bad escaped codes),
* then null_nullptrptr is returned. It is assumed that the output buffer is large
* then null_ptr is returned. It is assumed that the output buffer is large
* enough. E.g., if src points at 'joe"', then dst needs to have four free bytes +
* SIMDJSON_PADDING bytes.
*/
@@ -50481,7 +50514,7 @@ simdjson_inline error_code json_structural_indexer::finish(dom_parser_implementa
}
parser.n_structural_indexes = uint32_t(indexer.tail - parser.structural_indexes.get());
/***
* The On Demand API requires special padding.
* The On-Demand API requires special padding.
*
* This is related to https://github.com/simdjson/simdjson/issues/906
* Basically, we want to make sure that if the parsing continues beyond the last (valid)
@@ -51356,7 +51389,7 @@ simdjson_inline bool handle_unicode_codepoint_wobbly(const uint8_t **src_ptr,
* Unescape a valid UTF-8 string from src to dst, stopping at a final unescaped quote. There
* must be an unescaped quote terminating the string. It returns the final output
* position as pointer. In case of error (e.g., the string has bad escaped codes),
* then null_nullptrptr is returned. It is assumed that the output buffer is large
* then null_ptr is returned. It is assumed that the output buffer is large
* enough. E.g., if src points at 'joe"', then dst needs to have four free bytes +
* SIMDJSON_PADDING bytes.
*/
@@ -54557,7 +54590,7 @@ simdjson_inline bool handle_unicode_codepoint_wobbly(const uint8_t **src_ptr,
* Unescape a valid UTF-8 string from src to dst, stopping at a final unescaped quote. There
* must be an unescaped quote terminating the string. It returns the final output
* position as pointer. In case of error (e.g., the string has bad escaped codes),
* then null_nullptrptr is returned. It is assumed that the output buffer is large
* then null_ptr is returned. It is assumed that the output buffer is large
* enough. E.g., if src points at 'joe"', then dst needs to have four free bytes +
* SIMDJSON_PADDING bytes.
*/
+8988 -4661
View File
File diff suppressed because it is too large Load Diff
+1 -1
View File
@@ -263,7 +263,7 @@ simdjson_inline error_code json_structural_indexer::finish(dom_parser_implementa
}
parser.n_structural_indexes = uint32_t(indexer.tail - parser.structural_indexes.get());
/***
* The On Demand API requires special padding.
* The On-Demand API requires special padding.
*
* This is related to https://github.com/simdjson/simdjson/issues/906
* Basically, we want to make sure that if the parsing continues beyond the last (valid)
+1 -1
View File
@@ -143,7 +143,7 @@ simdjson_inline bool handle_unicode_codepoint_wobbly(const uint8_t **src_ptr,
* Unescape a valid UTF-8 string from src to dst, stopping at a final unescaped quote. There
* must be an unescaped quote terminating the string. It returns the final output
* position as pointer. In case of error (e.g., the string has bad escaped codes),
* then null_nullptrptr is returned. It is assumed that the output buffer is large
* then null_ptr is returned. It is assumed that the output buffer is large
* enough. E.g., if src points at 'joe"', then dst needs to have four free bytes +
* SIMDJSON_PADDING bytes.
*/
+1
View File
@@ -7,6 +7,7 @@
#include <internal/isadetection.h>
#include <initializer_list>
#include <type_traits>
namespace simdjson {
+53 -1
View File
@@ -38,6 +38,25 @@ const size_t AMAZON_CELLPHONES_NDJSON_DOC_COUNT = 793;
namespace number_tests {
bool build(const std::string& json) {
simdjson::dom::parser parser;
simdjson::dom::element recvdJson;
auto error = parser.parse(json.c_str(), json.size()).get(recvdJson);
if (error) {
return false;
}
return true;
}
bool issue2213() {
TEST_START();
std::string jsonStr = "[1,2,3,\"4\", {\"a\": 5}]";
for (int i = 0; i < 15; ++i) {
if (!build(jsonStr)) {
TEST_FAIL("The JSON is valid");
}
}
TEST_SUCCEED();
}
bool ground_truth() {
std::cout << __func__ << std::endl;
std::pair<std::string,double> ground_truth[] = {
@@ -397,7 +416,8 @@ namespace number_tests {
}
bool run() {
return bomskip() &&
return issue2213() &&
bomskip() &&
issue2017() &&
truncated_borderline() &&
specific_tests() &&
@@ -950,6 +970,36 @@ namespace dom_api_tests {
return true;
}
bool convert_object_to_element() {
std::cout << "Running " << __func__ << std::endl;
string json(R"({ "a": 1, "b": 2, "c": 3 })");
dom::parser parser;
dom::object object;
dom::element element;
ASSERT_SUCCESS( parser.parse(json).get(object) );
element = object;
ASSERT_EQUAL( element["a"].get_uint64().value_unsafe(), 1 );
ASSERT_EQUAL( element["b"].get_uint64().value_unsafe(), 2 );
ASSERT_EQUAL( element["c"].get_uint64().value_unsafe(), 3 );
return true;
}
bool convert_array_to_element() {
std::cout << "Running " << __func__ << std::endl;
string json(R"([ 1, 10, 100 ])");
dom::parser parser;
dom::array array;
dom::element element;
ASSERT_SUCCESS( parser.parse(json).get(array) );
element = array;
ASSERT_EQUAL( element.at(0).get_uint64().value_unsafe(), 1 );
ASSERT_EQUAL( element.at(1).get_uint64().value_unsafe(), 10 );
ASSERT_EQUAL( element.at(2).get_uint64().value_unsafe(), 100 );
return true;
}
bool string_value() {
std::cout << "Running " << __func__ << std::endl;
string json(R"([ "hi", "has backslash\\" ])");
@@ -1331,6 +1381,8 @@ namespace dom_api_tests {
array_iterator_empty() &&
object_iterator_advance() &&
array_iterator_advance() &&
convert_object_to_element() &&
convert_array_to_element() &&
string_value() &&
numeric_values() &&
boolean_values() &&
+23 -1
View File
@@ -249,6 +249,27 @@ namespace document_stream_tests {
TEST_SUCCEED();
}
bool issue2181() {
TEST_START();
auto json = R"(1 2 34)"_padded;
simdjson::dom::parser parser;
simdjson::dom::document_stream stream;
ASSERT_SUCCESS(parser.parse_many(json).get(stream));
auto i = stream.begin();
size_t count{0};
std::vector<size_t> indexes = { 0, 2, 4 };
std::vector<std::string_view> expected = { "1", "2", "34" };
for(; i != stream.end(); ++i) {
auto doc = *i;
ASSERT_SUCCESS(doc);
ASSERT_TRUE(count < 3);
ASSERT_EQUAL(i.current_index(), indexes[count]);
ASSERT_EQUAL(i.source(), expected[count]);
count++;
}
TEST_SUCCEED();
}
bool issue1310() {
std::cout << "Running " << __func__ << std::endl;
// hex : 20 20 5B 20 33 2C 31 5D 20 22 22 22 22 22 22 22 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20 20
@@ -963,7 +984,8 @@ namespace document_stream_tests {
}
bool run() {
return issue2170() &&
return issue2181() &&
issue2170() &&
skipbom() &&
fuzzaccess() &&
baby_fuzzer() &&
+30
View File
@@ -191,10 +191,40 @@ bool issue1142() {
return true;
}
bool issue2154() { // mistakenly taking value as path should not raise INVALID_JSON_POINTER
#if SIMDJSON_EXCEPTIONS
std::cout << "issue 2154" << std::endl;
auto example_json = R"__(
{
"obj": {
"s": "42",
"n": 42,
"f": 4.2
}
}
)__"_padded;
dom::parser parser;
dom::element example = parser.parse(example_json);
std::string_view sfield = example.at_pointer("/obj/s");
ASSERT_EQUAL(sfield, "42");
int64_t nfield = example.at_pointer("/obj/n");
ASSERT_EQUAL(nfield, 42);
ASSERT_ERROR(example.at_pointer("/obj/X/42").error(), NO_SUCH_FIELD);
ASSERT_ERROR(example.at_pointer("/obj/s/42").error(), NO_SUCH_FIELD);
ASSERT_ERROR(example.at_pointer("/obj/n/42").error(), NO_SUCH_FIELD);
ASSERT_ERROR(example.at_pointer("/obj/f/4.2").error(), NO_SUCH_FIELD);
ASSERT_ERROR(example.at_pointer("/obj/f/4~").error(), INVALID_JSON_POINTER);
ASSERT_ERROR(example.at_pointer("/obj/f/~").error(), INVALID_JSON_POINTER);
ASSERT_ERROR(example.at_pointer("/obj/f/~1").error(), NO_SUCH_FIELD);
#endif
return true;
}
int main() {
if (true
&& demo()
&& issue1142()
&& issue2154()
#ifdef SIMDJSON_ENABLE_DEPRECATED_API
&& legacy_support()
#endif
+32 -26
View File
@@ -2,32 +2,38 @@
link_libraries(simdjson)
include_directories(..)
add_subdirectory(compilation_failure_tests)
add_cpp_test(ondemand_log_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_log_error_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_tostring_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_active_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_array_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_array_error_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_compilation_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_document_stream_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_error_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_error_location_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_json_pointer_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_json_path_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_key_string_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_misc_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_number_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_number_in_string_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_object_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_object_error_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_ordering_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_parse_api_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_readme_examples LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_scalar_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_to_string LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_twitter_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_wrong_type_error_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_iterate_many_csv LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_log_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_log_error_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_tostring_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_active_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_array_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_array_error_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_compilation_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_document_stream_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_error_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_error_location_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_json_pointer_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_json_path_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_key_string_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_misc_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_number_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_number_in_string_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_object_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_object_error_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_ordering_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_parse_api_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_readme_examples LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_scalar_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_to_string LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_twitter_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_wrong_type_error_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_iterate_many_csv LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_custom_types_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_custom_types_document_tests LABELS ondemand acceptance per_implementation)
add_cpp_test(ondemand_stl_types_tests LABELS ondemand acceptance per_implementation)
if(NOT SIMDJSON_SANITIZE)
add_cpp_test(ondemand_cacheline LABELS ondemand acceptance per_implementation)
endif()
if(HAVE_POSIX_FORK AND HAVE_POSIX_WAIT) # assert tests use fork and wait, which aren't on MSVC
add_cpp_test(ondemand_assert_out_of_order_values LABELS assert per_implementation explicitonly ondemand)
+68
View File
@@ -0,0 +1,68 @@
#ifdef _WIN32
#include <windows.h>
#include <sysinfoapi.h>
#else
#include <unistd.h>
#endif
#include "simdjson.h"
#include <cstdio>
// Returns the default size of the page in bytes on this system.
long page_size() {
#ifdef _WIN32
SYSTEM_INFO sysInfo;
GetSystemInfo(&sysInfo);
long pagesize = sysInfo.dwPageSize;
#else
long pagesize = sysconf(_SC_PAGESIZE);
#endif
return pagesize;
}
// Returns true if the buffer + len + simdjson::SIMDJSON_PADDING crosses the
// page boundary.
bool need_allocation(const char *buf, size_t len) {
return ((reinterpret_cast<uintptr_t>(buf + len - 1) % page_size()) <
simdjson::SIMDJSON_PADDING);
}
simdjson::padded_string_view
get_padded_string_view(const char *buf, size_t len,
simdjson::padded_string &jsonbuffer) {
if (need_allocation(buf, len)) { // unlikely case
jsonbuffer = simdjson::padded_string(buf, len);
return jsonbuffer;
} else { // no reallcation needed (very likely)
return simdjson::padded_string_view(buf, len,
len + simdjson::SIMDJSON_PADDING);
}
}
int main() {
printf("page_size: %ld\n", page_size());
const char *jsonpoiner = R"(
{
"key": "value"
}
)";
size_t len = strlen(jsonpoiner);
simdjson::padded_string jsonbuffer; // only allocate if needed
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc;
simdjson::error_code error =
parser.iterate(get_padded_string_view(jsonpoiner, len, jsonbuffer))
.get(doc);
if (error) {
printf("error: %s\n", simdjson::error_message(error));
return EXIT_FAILURE;
}
std::string_view value;
error = doc["key"].get_string().get(value);
if (error) {
return EXIT_FAILURE;
}
printf("Value: \"%.*s\"\n", (int)value.size(), value.data());
if (value != "value") {
return EXIT_FAILURE;
}
return EXIT_SUCCESS;
}
@@ -0,0 +1,192 @@
#include "simdjson.h"
#include "test_ondemand.h"
#include <string>
#include <vector>
#if SIMDJSON_SUPPORTS_DESERIALIZATION
namespace simdjson {
// unique_ptr<T>
template <typename T>
error_code tag_invoke(deserialize_tag, auto &val, std::unique_ptr<T>& out) {
out = std::make_unique<T>(val.template get<T>());
return SUCCESS;
}
} // namespace simdjson
#endif // SIMDJSON_SUPPORTS_DESERIALIZATION
namespace doc_custom_types_tests {
#if SIMDJSON_EXCEPTIONS && defined(__cpp_concepts)
struct Car {
std::string make{};
std::string model{};
int64_t year{};
std::vector<double> tire_pressure{};
friend error_code tag_invoke(simdjson::deserialize_tag,
auto &val, Car& car) {
simdjson::ondemand::object obj;
if (auto const error = val.get_object().get(obj)) {
return error;
}
// Instead of repeatedly obj["something"], we iterate through the object
// which we expect to be faster.
for (auto field : obj) {
simdjson::ondemand::raw_json_string key;
if (auto const error = field.key().get(key)) {
return error;
}
if (key == "make") {
if (auto const error = field.value().get_string(car.make)) {
return error;
}
} else if (key == "model") {
if (auto const error = field.value().get_string(car.model)) {
return error;
}
} else if (key == "year") {
if (auto const error = field.value().get_int64().get(car.year)) {
return error;
}
} else if (key == "tire_pressure") {
if (auto const error = field.value().get<std::vector<double>>().get(
car.tire_pressure)) {
return error;
}
}
}
return error_code::SUCCESS;
}
};
static_assert(simdjson::custom_deserializable<std::unique_ptr<Car>, simdjson::ondemand::value>, "It should be invocable");
static_assert(simdjson::custom_deserializable<std::unique_ptr<Car>, simdjson::ondemand::document>, "Tag_invoke should work with document as well.");
static_assert(simdjson::custom_deserializable<std::vector<Car>, simdjson::ondemand::value>, "It should be invocable");
static_assert(simdjson::custom_deserializable<std::vector<Car>, simdjson::ondemand::document>, "Tag_invoke should work with document as well.");
bool custom_test() {
TEST_START();
auto const json = R"( {
"make": "Toyota",
"model": "Camry",
"year": 2018,
"tire_pressure": [ 40.1, 39.9 ]
} )"_padded;
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc = parser.iterate(json);
// creating car directly from "doc":
const std::unique_ptr<Car> car(doc);
if (car->make != "Toyota") {
return false;
}
if (car->model != "Camry") {
return false;
}
if (car->year != 2018) {
return false;
}
if (car->tire_pressure.size() != 2) {
return false;
}
TEST_SUCCEED();
}
bool readme_test() {
TEST_START();
auto const json = R"( { "make": "Toyota", "model": "Camry", "year": 2018,
"tire_pressure": [ 40.1, 39.9 ] })"_padded;
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc = parser.iterate(json);
std::unique_ptr<Car> c(doc);
std::cout << c->make << std::endl;
TEST_SUCCEED();
}
bool simple_document_test() {
TEST_START();
simdjson::padded_string json =
R"( { "make": "Toyota", "model": "Camry", "year": 2018,
"tire_pressure": [ 40.1, 39.9 ] })"_padded;
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc = parser.iterate(json);
Car c(doc);
std::cout << c.make << std::endl;
ASSERT_EQUAL(c.make, "Toyota");
TEST_SUCCEED();
}
bool custom_test_crazier() {
TEST_START();
simdjson::padded_string json =
R"( [ { "make": "Toyota", "model": "Camry", "year": 2018,
"tire_pressure": [ 40.1, 39.9 ] },
{ "make": "Kia", "model": "Soul", "year": 2012,
"tire_pressure": [ 30.1, 31.0 ] },
{ "make": "Toyota", "model": "Tercel", "year": 1999,
"tire_pressure": [ 29.8, 30.0 ] }
])"_padded;
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc = parser.iterate(json);
#if SIMDJSON_REGULAR_VISUAL_STUDIO
std::vector<Car> cars = doc.get<std::vector<Car>>();
#else
std::vector<Car> cars(doc);
#endif
for(Car& c : cars) {
std::cout << c.year << std::endl;
if (c.make != "Toyota" && c.make != "Kia") {
return false;
}
if (c.model != "Camry" && c.model != "Soul" && c.model != "Tercel") {
return false;
}
if (c.year != 2018 && c.year != 2012 && c.year != 1999) {
return false;
}
if (c.tire_pressure.size() != 2) {
return false;
}
}
TEST_SUCCEED();
}
bool simple_document_test_no_except() {
TEST_START();
simdjson::padded_string json =
R"( { "make": "Toyota", "model": "Camry", "year": 2018,
"tire_pressure": [ 40.1, 39.9 ] })"_padded;
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc = parser.iterate(json);
Car c;
auto error = doc.get(c);
if(error) { std::cerr << simdjson::error_message(error); return false; }
std::cout << c.make << std::endl;
ASSERT_EQUAL(c.make, "Toyota");
TEST_SUCCEED();
}
#endif // SIMDJSON_EXCEPTIONS
bool run() {
return
#if SIMDJSON_EXCEPTIONS && defined(__cpp_concepts)
readme_test() &&
custom_test() &&
simple_document_test() &&
simple_document_test_no_except() &&
custom_test_crazier() &&
#endif // SIMDJSON_EXCEPTIONS
true;
}
} // namespace doc_custom_types_tests
int main(const int argc, char *argv[]) {
return test_main(argc, argv, doc_custom_types_tests::run);
}
@@ -0,0 +1,229 @@
#include "simdjson.h"
#include "test_ondemand.h"
#include <string>
#include <vector>
#if SIMDJSON_SUPPORTS_DESERIALIZATION
template <typename T>
struct is_unique_ptr : std::false_type {
};
template <typename T>
struct is_unique_ptr<std::unique_ptr<T>> : std::true_type {
};
template <typename T>
concept is_unique_ptr_v = is_unique_ptr<T>::value;
namespace simdjson {
// this is to demonstrate that we can use concept; otherwise a simple
// type_identity<unique_ptr<T>> as the second argument would have done the job.
//
// This tag_invoke MUST be inside simdjson namespace
template <typename T>
requires is_unique_ptr_v<T>
auto tag_invoke(deserialize_tag, auto &val, T& out) {
using type = typename T::element_type;
out = std::make_unique<type>(val.template get<type>());
return SUCCESS;
}
} // namespace simdjson
#endif // SIMDJSON_SUPPORTS_DESERIALIZATION
namespace custom_types_tests {
#if SIMDJSON_EXCEPTIONS && SIMDJSON_SUPPORTS_DESERIALIZATION
struct Car {
std::string make{};
std::string model{};
int year{};
std::vector<double> tire_pressure{};
friend simdjson::error_code tag_invoke(simdjson::deserialize_tag, auto &val, Car& car) {
simdjson::ondemand::object obj;
auto error = val.get_object().get(obj);
if (error) {
return error;
}
// Instead of repeatedly obj["something"], we iterate through the object
// which we expect to be faster.
for (auto field : obj) {
simdjson::ondemand::raw_json_string key;
error = field.key().get(key);
if (error) {
return error;
}
if (key == "make") {
error = field.value().get_string(car.make);
if (error) {
return error;
}
} else if (key == "model") {
error = field.value().get_string(car.model);
if (error) {
return error;
}
} else if (key == "year") {
error = field.value().get(car.year);
if (error) {
return error;
}
} else if (key == "tire_pressure") {
error = field.value().get(car.tire_pressure);
if (error) {
return error;
}
}
}
return simdjson::SUCCESS;
}
};
static_assert(simdjson::custom_deserializable<std::unique_ptr<Car>>, "It should be deserializable");
bool custom_uniqueptr_test() {
TEST_START();
simdjson::padded_string json =
R"( [ { "make": "Toyota", "model": "Camry", "year": 2018,
"tire_pressure": [ 40.1, 39.9 ] },
{ "make": "Kia", "model": "Soul", "year": 2012,
"tire_pressure": [ 30.1, 31.0 ] },
{ "make": "Toyota", "model": "Tercel", "year": 1999,
"tire_pressure": [ 29.8, 30.0 ] }
])"_padded;
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc = parser.iterate(json);
for (auto val : doc) {
std::unique_ptr<Car> c(val);
if (c->make != "Toyota" && c->make != "Kia") {
return false;
}
if (c->model != "Camry" && c->model != "Soul" && c->model != "Tercel") {
return false;
}
if (c->year != 2018 && c->year != 2012 && c->year != 1999) {
return false;
}
if (c->tire_pressure.size() != 2) {
return false;
}
}
TEST_SUCCEED();
}
bool custom_test() {
TEST_START();
simdjson::padded_string json =
R"( [ { "make": "Toyota", "model": "Camry", "year": 2018,
"tire_pressure": [ 40.1, 39.9 ] },
{ "make": "Kia", "model": "Soul", "year": 2012,
"tire_pressure": [ 30.1, 31.0 ] },
{ "make": "Toyota", "model": "Tercel", "year": 1999,
"tire_pressure": [ 29.8, 30.0 ] }
])"_padded;
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc = parser.iterate(json);
for (auto val : doc) {
Car c(val);
if (c.make != "Toyota" && c.make != "Kia") {
return false;
}
if (c.model != "Camry" && c.model != "Soul" && c.model != "Tercel") {
return false;
}
if (c.year != 2018 && c.year != 2012 && c.year != 1999) {
return false;
}
if (c.tire_pressure.size() != 2) {
return false;
}
}
TEST_SUCCEED();
}
bool custom_test_crazy() {
TEST_START();
simdjson::padded_string json =
R"( [ { "make": "Toyota", "model": "Camry", "year": 2018,
"tire_pressure": [ 40.1, 39.9 ] },
{ "make": "Kia", "model": "Soul", "year": 2012,
"tire_pressure": [ 30.1, 31.0 ] },
{ "make": "Toyota", "model": "Tercel", "year": 1999,
"tire_pressure": [ 29.8, 30.0 ] }
])"_padded;
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc = parser.iterate(json);
for (auto val : doc) {
Car c(val);
std::cout << c.year << std::endl;
if (c.make != "Toyota" && c.make != "Kia") {
return false;
}
if (c.model != "Camry" && c.model != "Soul" && c.model != "Tercel") {
return false;
}
if (c.year != 2018 && c.year != 2012 && c.year != 1999) {
return false;
}
if (c.tire_pressure.size() != 2) {
return false;
}
}
TEST_SUCCEED();
}
bool custom_no_except() {
TEST_START();
simdjson::padded_string json =
R"( [ { "make": "Toyota", "model": "Camry", "year": 2018,
"tire_pressure": [ 40.1, 39.9 ] },
{ "make": "Kia", "model": "Soul", "year": 2012,
"tire_pressure": [ 30.1, 31.0 ] },
{ "make": "Toyota", "model": "Tercel", "year": 1999,
"tire_pressure": [ 29.8, 30.0 ] }
])"_padded;
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc = parser.iterate(json);
for (auto val : doc) {
Car c;
auto error = val.get(c);
if(error) { std::cerr << simdjson::error_message(error) << std::endl; return false; }
}
TEST_SUCCEED();
}
#endif // SIMDJSON_EXCEPTIONS
bool run() {
return
#if SIMDJSON_EXCEPTIONS && SIMDJSON_SUPPORTS_DESERIALIZATION
custom_test() &&
custom_uniqueptr_test() &&
custom_no_except() &&
#endif // SIMDJSON_EXCEPTIONS
true;
}
} // namespace custom_types_tests
int main(int argc, char *argv[]) {
return test_main(argc, argv, custom_types_tests::run);
}
@@ -217,6 +217,27 @@ namespace document_stream_tests {
TEST_SUCCEED();
}
bool issue2181() {
TEST_START();
auto json = R"(1 2 34)"_padded;
ondemand::parser parser;
ondemand::document_stream stream;
ASSERT_SUCCESS(parser.iterate_many(json).get(stream));
auto i = stream.begin();
size_t count{0};
std::vector<size_t> indexes = { 0, 2, 4 };
std::vector<std::string_view> expected = { "1", "2", "34" };
for(; i != stream.end(); ++i) {
ASSERT_SUCCESS(i.error());
ASSERT_TRUE(count < 3);
ASSERT_EQUAL(i.current_index(), indexes[count]);
ASSERT_EQUAL(i.source(), expected[count]);
count++;
}
TEST_SUCCEED();
}
bool issue1977() {
TEST_START();
std::string json = R"( 1111 })";
@@ -902,6 +923,7 @@ namespace document_stream_tests {
bool run() {
return
issue2181() &&
issue2170() &&
issue2137() &&
skipbom() &&
@@ -385,8 +385,38 @@ namespace json_pointer_tests {
TEST_SUCCEED();
}
#endif
bool issue2154() { // mistakenly taking value as path should not raise INVALID_JSON_POINTER
#if SIMDJSON_EXCEPTIONS
std::cout << "issue 2154" << std::endl;
auto example_json = R"__({
"obj": {
"s": "42",
"n": 42,
"f": 4.2
}
})__"_padded;
ondemand::parser parser;
ondemand::document doc;
ASSERT_SUCCESS(parser.iterate(example_json).get(doc));
std::string_view sfield = doc.at_pointer("/obj/s");
ASSERT_EQUAL(sfield, "42");
int64_t nfield = doc.at_pointer("/obj/n");
ASSERT_EQUAL(nfield, 42);
ASSERT_ERROR(doc.at_pointer("/obj/X/42").error(), NO_SUCH_FIELD);
ASSERT_ERROR(doc.at_pointer("/obj/s/42").error(), NO_SUCH_FIELD);
ASSERT_ERROR(doc.at_pointer("/obj/n/42").error(), NO_SUCH_FIELD);
ASSERT_ERROR(doc.at_pointer("/obj/f/4.2").error(), NO_SUCH_FIELD);
ASSERT_ERROR(doc.at_pointer("/obj/f/4~").error(), INVALID_JSON_POINTER);
ASSERT_ERROR(doc.at_pointer("/obj/f/~").error(), INVALID_JSON_POINTER);
ASSERT_ERROR(doc.at_pointer("/obj/f/~1").error(), NO_SUCH_FIELD);
#endif
return true;
}
bool run() {
return
issue2154() &&
#if SIMDJSON_EXCEPTIONS
json_pointer_invalidation_exceptions() &&
#endif
+24 -6
View File
@@ -5,6 +5,22 @@ using namespace simdjson;
namespace misc_tests {
using namespace std;
#if SIMDJSON_EXCEPTIONS
// user reported an asan error:
bool issue2199() {
TEST_START();
static constexpr std::string_view kJsonString = R"( { "name": "name", "version": 100, } )";
try {
simdjson::padded_string buffer{kJsonString};
simdjson::ondemand::parser parser;
simdjson::ondemand::document document = parser.iterate(buffer);
(void)document;
} catch (simdjson::simdjson_error& /*error*/) {
std::cerr << "Caught simdjson_error" << std::endl;
}
TEST_SUCCEED();
}
#endif
bool issue1981_success() {
auto error_phrase = R"(false)"_padded;
TEST_START();
@@ -532,7 +548,7 @@ namespace misc_tests {
string_view token;
ASSERT_SUCCESS(o["value"].raw_json_token().get(token));
ASSERT_EQUAL(token, "12321323213213213213213213213211223");
return true;
TEST_SUCCEED();
}
simdjson_warn_unused bool big_integer_in_string() {
TEST_START();
@@ -545,7 +561,7 @@ namespace misc_tests {
string_view token;
ASSERT_SUCCESS(o["value"].raw_json_token().get(token));
ASSERT_EQUAL(token, "\"12321323213213213213213213213211223\"");
return true;
TEST_SUCCEED();
}
simdjson_warn_unused bool test_raw_json_token(string_view json, string_view expected_token, int expected_start_index = 0) {
string title("'");
@@ -558,7 +574,7 @@ namespace misc_tests {
ASSERT_EQUAL( token, expected_token );
// Validate the text is inside the original buffer
ASSERT_EQUAL( reinterpret_cast<const void*>(token.data()), reinterpret_cast<const void*>(&json_padded.data()[expected_start_index]));
return true;
TEST_SUCCEED();
}));
// Test values
@@ -576,10 +592,9 @@ namespace misc_tests {
// Validate the text is inside the original buffer
// Adjust for the {"a":
ASSERT_EQUAL( reinterpret_cast<const void*>(token.data()), reinterpret_cast<const void*>(&json_padded.data()[5+expected_start_index]));
return true;
TEST_SUCCEED();
}));
return true;
TEST_SUCCEED();
}
bool raw_json_token() {
@@ -605,6 +620,9 @@ namespace misc_tests {
bool run() {
return
#if SIMDJSON_EXCEPTIONS
issue2199() &&
#endif
skipbom() &&
issue1981_success() &&
issue1981_failure() &&
+16
View File
@@ -700,6 +700,22 @@ namespace object_tests {
}
return got_key;
}));
SUBTEST("ondemand::unescapedkey(std)", test_ondemand_doc(json, [&](auto doc_result) {
ondemand::object object;
bool got_key = false;
ASSERT_SUCCESS( doc_result.get(object) );
for (auto field : object) {
std::string keyv;
ASSERT_SUCCESS( field.unescaped_key(keyv) );
if(keyv == "key") {
int64_t value;
ASSERT_SUCCESS( field.value().get(value) );
ASSERT_EQUAL( value, 1);
got_key = true;
}
}
return got_key;
}));
SUBTEST("ondemand::rawkey", test_ondemand_doc(json, [&](auto doc_result) {
ondemand::object object;
ASSERT_SUCCESS( doc_result.get(object) );
+47 -3
View File
@@ -49,6 +49,7 @@ struct Car {
std::vector<double> tire_pressure;
};
#ifndef __cpp_concepts
template <>
simdjson_inline simdjson_result<std::vector<double>>
simdjson::ondemand::value::get() noexcept {
@@ -66,7 +67,7 @@ simdjson::ondemand::value::get() noexcept {
}
return vec;
}
#endif
template <>
@@ -153,6 +154,25 @@ simdjson_inline simdjson_result<Car> simdjson::ondemand::document::get() & noexc
#if SIMDJSON_EXCEPTIONS
void main_capture() {
padded_string json_padded = "{\"a\":[1,2,3], \"b\": 2, \"c\": \"hello\"}"_padded;
std::vector<std::string_view> fields;
ondemand::parser parser;
auto doc = parser.iterate(json_padded);
auto object = doc.get_object();
for (auto field : object) {
fields.push_back(field.value().raw_json());
}
// Output the fields
// Expected output:
// [1,2,3]
// 2
// "hello"
for (std::string_view field_ref : fields) {
std::cout << field_ref << std::endl;
}
}
int custom_type_on_document() {
padded_string json = R"( { "make": "Toyota", "model": "Camry", "year": 2018,
@@ -174,8 +194,8 @@ int custom_type_with_exceptions() {
])"_padded;
ondemand::parser parser;
ondemand::document doc = parser.iterate(json);
for (auto val : doc) {
Car c(val);
for (auto value : doc) {
Car c(value);
std::cout << c.make << std::endl;
}
return 0;
@@ -1538,6 +1558,29 @@ bool allow_comma_separated_example() {
}
TEST_SUCCEED();
}
bool issue2215() {
TEST_START();
ondemand::parser parser;
const padded_string json = R"({ "parent": {"child1": {"name": "John"} , "child2": {"name": "Daniel"}} })"_padded;
auto doc = parser.iterate(json);
ondemand::object parent = doc["parent"];
// parent owns the focus
ondemand::object c1 = parent["child1"];
// c1 owns the focus
//
std::string_view as1 = c1["name"];
// We have that as1 == "John", as long as 'parser' and 'json' live
// c2 attempts to grab the focus from parent but fails
ondemand::object c2 = parent["child2"];
// c2 owns the focus, at this point c1 is invalid
std::string_view as2 = c2["name"];
// We have that as2 == "Daniel", as long as 'parser' and 'json' live
ASSERT_EQUAL(as1, "John");
ASSERT_EQUAL(as2, "Daniel");
std::cout << as1 << " " << as2 << std::endl;
TEST_SUCCEED();
}
#endif
bool test_load_example() {
TEST_START();
@@ -1899,6 +1942,7 @@ bool run() {
&& current_location_no_error()
&& to_string_example_no_except()
#if SIMDJSON_EXCEPTIONS
&& issue2215()
&& to_string_example()
&& raw_string()
&& number_tests()
@@ -0,0 +1,48 @@
#include "simdjson.h"
#include "test_ondemand.h"
#include <string>
#include <vector>
namespace stl_types {
#if SIMDJSON_EXCEPTIONS && SIMDJSON_SUPPORTS_DESERIALIZATION
bool basic_general_madness() {
TEST_START();
simdjson::padded_string json =
R"( [ { "codes": [1.2, 3.4, 5.6, 7.8, 9.0] },
{ "codes": [2.2, 4.4, 6.6, 8.8, 10.0] } ])"_padded;
simdjson::ondemand::parser parser;
simdjson::ondemand::document doc = parser.iterate(json);
std::vector<std::unique_ptr<float>> codes;
for (auto val : doc) {
simdjson::ondemand::object obj;
SIMDJSON_TRY(val.get_object().get(obj));
obj["codes"].get(codes); // append to it
}
if (codes.size() != 10) {
return false;
}
if (*codes[5] != 2.2f || *codes[7] != 6.6f || *codes[9] != 10.0f) {
return false;
}
TEST_SUCCEED();
}
#endif // SIMDJSON_EXCEPTIONS
bool run() {
return
#if SIMDJSON_EXCEPTIONS && SIMDJSON_SUPPORTS_DESERIALIZATION
basic_general_madness() &&
#endif // SIMDJSON_EXCEPTIONS
true;
}
} // namespace stl_types
int main(int argc, char *argv[]) {
return test_main(argc, argv, stl_types::run);
}
+2 -2
View File
@@ -31,7 +31,7 @@ int main(int argc, const char *argv[]) {
cxxopts::Options options(progName, progUsage);
options.add_options()
("z,ondemand", "Use On Demand front-end.", cxxopts::value<bool>()->default_value("false"))
("z,ondemand", "Use On-Demand front-end.", cxxopts::value<bool>()->default_value("false"))
("d,rawdump", "Dumps the raw content of the tape.", cxxopts::value<bool>()->default_value("false"))
("f,file", "File name.", cxxopts::value<std::string>())
("h,help", "Print usage.")
@@ -92,7 +92,7 @@ int main(int argc, const char *argv[]) {
}
return EXIT_SUCCESS;
#ifdef __cpp_exceptions
} catch (const cxxopts::OptionException& e) {
} catch (const cxxopts::exceptions::option_has_no_value& e) {
std::cout << "error parsing options: " << e.what() << std::endl;
return EXIT_FAILURE;
}
+1 -1
View File
@@ -278,7 +278,7 @@ int main(int argc, const char *argv[]) {
s.repeated_key_byte_count, s.maximum_depth);
return EXIT_SUCCESS;
#ifdef __cpp_exceptions
} catch (const cxxopts::OptionException& e) {
} catch (const cxxopts::exceptions::option_has_no_value& e) {
std::cout << "error parsing options: " << e.what() << std::endl;
return EXIT_FAILURE;
}
+1 -1
View File
@@ -111,7 +111,7 @@ int main(int argc, const char *argv[]) {
}
return EXIT_SUCCESS;
#ifdef __cpp_exceptions
} catch (const cxxopts::OptionException& e) {
} catch (const cxxopts::exceptions::option_has_no_value& e) {
std::cout << "error parsing options: " << e.what() << std::endl;
return EXIT_FAILURE;
}