Files
Niels Lohmann e5a89d671f Fix lint debt: enum-macro NOLINTs, doctest as SYSTEM, no-op analyzer (#5737)
* Drop stale LCOV_EXCL_LINE from the json_pointer out_of_range.410 throw

The comment said the size_type overflow check in array_index() is only
triggered on special platforms like 32-bit, and the throw was excluded
from coverage. On 64-bit platforms the check is true for SIZE_MAX
itself, and unit-json_pointer.cpp has asserted that case four times
since #5395, so the line is executed in the coverage job. Reword the
comment and remove the exclusion marker so the coverage report notices
if the tests stop reaching it.

Part of #5725

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Name all three C-array check aliases in the enum-macro NOLINTs

NLOHMANN_JSON_SERIALIZE_ENUM(_STRICT) suppressed the c-array warning
under modernize-avoid-c-arrays only, but clang-tidy emits the same
diagnostic under the aliases cppcoreguidelines-avoid-c-arrays and
hicpp-avoid-c-arrays too. Any user running those checks got a false
positive at every macro expansion, and our own tests needed a local
NOLINT at each call site to work around it.

Name all three aliases in the four macro comments instead, and drop
the now-redundant c-array names from the five test call-site NOLINTs.
Comment-only change; behavior, the public API, and the ABI do not
change. Ran make amalgamate.

Part of #5725

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Include doctest as a SYSTEM directory instead of disabling warnings for all tests

test_main added -Wno-deprecated and -Wno-float-equal as PUBLIC compile
options for every non-MSVC compiler, so they were applied to every
translation unit, library headers included, and silenced the CI
warnings meant to check the library's own -Wfloat-equal pragmas. The
only code that actually needed the suppression was the vendored
doctest.h, which was included as a normal (non-SYSTEM) directory.

Include thirdparty/doctest as SYSTEM for test_main, matching what
tests/abi/CMakeLists.txt already does, and drop the two suppressions
from both targets. Verified locally that unit-comparison,
unit-conversions and unit-constructor1 compile clean with
-Werror -Weverything and doctest as -isystem, and that CMake still
configures with JSON_BuildTests=ON.

Part of #5725

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove the no-op ci_clang_analyze target

ci_clang_analyze configured the build with the real compiler and only
then wrapped ninja with scan-build. scan-build intercepts compiles by
overriding CC/CXX, but build.ninja already had the compiler path baked
in from the configure step, so every run bypassed the analyzer: CI
logs show "No bugs found" after a normal build, never an analysis.
The job also used Debian's frozen clang-tools-14 rather than the
image's own clang, and CLANG_ANALYZER_CHECKS still named three
valist.* checkers that current clang merged into security.VAList.

ci_clang_tidy already runs every clang-analyzer-* check (via
.clang-tidy's "Checks: '*'") with warnings as errors, so nothing is
lost by removing the dead job. Delete ci_clang_analyze,
CLANG_ANALYZER_CHECKS and the SCAN_BUILD_TOOL lookup from
cmake/ci.cmake, drop it from the ubuntu.yml ci_static_analysis_clang
matrix, and drop the now-unused clang-tools apt package (iwyu stays
for ci_single_binaries). Reword quality_assurance.md and
assurance_case.md, which described the dead job as a working control,
to say the Clang Static Analyzer checks run through clang-tidy.

Verified that `cmake -DJSON_CI=ON` still configures cleanly and that
ci_clang_analyze no longer appears in the generated build or in any
CMake/workflow file.

Part of #5725

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Re-enable portability-template-virtual-member-function; remove redundant forwards

.clang-tidy disabled three checks "to get the CI going" (#4489,
2024-11-13): portability-template-virtual-member-function,
bugprone-use-after-move and its alias hicpp-invalid-access-moved.

portability-template-virtual-member-function only flagged
output_stream_adapter::write_character/write_characters; annotate
both with NOLINT and re-enable the check.

bugprone-use-after-move flagged several double forwards that have no
effect at runtime:
- from_json.hpp calls std::forward<BasicJsonType>(j).at(Idx) inside
  pack expansions; at() has no ref-qualified overloads and always
  returns an lvalue reference, so the forward is a no-op. Replace with
  plain j.at(Idx) in all four places.
- the move constructor forwards the whole object to its base class
  and then reads other's members. That is item 9 of #5724 (together
  with its cppcheck suppressions) and is left to that change.
- input_adapters.hpp forwards the container twice on purpose, so the
  begin/end iterator types match adapter_type; annotate with NOLINT
  and a comment instead of changing behavior.

The check still flags the move constructor (see above) and two sites
in at(KeyType&&) (both overloads, json.hpp, in the throw's
string_t(std::forward<KeyType>(key)) after
find(std::forward<KeyType>(key))). Open PR #5689 rewrites that hunk,
so bugprone-use-after-move (and hicpp-invalid-access-moved)
stay disabled for now, with a comment explaining why; re-enable them
once #5689 and the #5724 move-constructor change have landed.

Also resolve the portability-avoid-pragma-once TODO: single_include
never has #pragma once (amalgamate.py strips it) and every supported
compiler accepts it in include/, so keep it disabled with an
explanatory comment instead of a TODO. Fix the stale "json.hpp,
around line 1265" comment in unit-class_parser.cpp, which now points
at the move constructor's actual line.

Behavior, the public API and the ABI do not change. Verified with
clang-tidy 22.1.8 that portability-template-virtual-member-function
now reports nothing, that bugprone-use-after-move/
hicpp-invalid-access-moved report only the known at(KeyType&&) and
move-constructor sites, and that unit-custom-base-class, unit-constructor1,
unit-conversions, unit-element_access2, unit-class_parser and
unit-diagnostic-positions (JSON_DIAGNOSTIC_POSITIONS=1) compile
under ASan/UBSan and pass with the same assertion counts as before.
Ran make amalgamate.

Part of #5725

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix stale and malformed NOLINT comments

json_sax.hpp named "-warnings-as-errors" in the NOLINT list on the two
JSON_ASSERT(false) lines; that is the suffix clang-tidy appends to a
diagnostic tag under WarningsAsErrors, not a check name, and every
other JSON_ASSERT(false) omits it.

unit-capacity.cpp carried 30 "// NOLINT(misc-const-correctness)"
comments on "json j = ...;" declarations that are all used with
non-const members afterwards, so the check has nothing to report
there.

unit-constructor2.cpp used a blanket "// NOLINT: access after move is
OK here" on a use-after-move that hides every check on the line;
naming bugprone-use-after-move and hicpp-invalid-access-moved keeps
the intent once those checks are re-enabled (#5724).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 10

* Remove stale .clang-tidy entries

-google-runtime-references disabled a check that neither clang-tidy
22.1.8 nor 23.1.2 lists under --list-checks -checks='*'; it was
removed upstream. The commented-out HeaderFilterRegex line has been
unused since the active HeaderFilterRegex was introduced in #2561
(2021).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 11

* Remove the GCC C++20 -Wignored-attributes pragma in json.hpp

The pragma (added in #5164) claimed to work around the C++ modules
redefinition errors of #5103, but #5103 is about hard errors (e.g.
"redefinition of std::__is_constant_evaluated()", conflicting
std::integral_constant) that ignoring a warning cannot suppress; they
are traced to GCC PR 124430 and reproduce with <map> or <string>
instead of json.hpp too. A GCC 16.2 -std=gnu++20 -fmodules build
following #5103's repro steps still fails with the pragma in place,
and a build of all test TUs with GCC_CXXFLAGS (which enable
-Wignored-attributes) and the pragma removed produces no such
warning. The block only hid a warning class from GCC C++20 users
while suggesting #5103 was handled.

Overlaps #5610, whose hunks touch the closing half of this pragma to
insert the json_literals.hpp include.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 9

* Fix stale doxygen comments hidden by the -Wdocumentation pragma

macro_scope.hpp ignores -Wdocumentation and -Wdocumentation-unknown-command
for the whole library, which also hides genuine documentation mistakes:

- detail::unescape() documented "@return unescaped string" but returns
  void and unescapes its argument in place; reworded to
  "@param[in,out] s string to unescape in place" and dropped the
  bogus @return.
- basic_json::get()'s copy-conversion overload wrote "converted to
  @tparam ValueType" inside @return, which Doxygen and Clang parse as
  a second, malformed @tparam; changed to "@a ValueType", matching the
  two other get() overloads a few lines above that already use it.

This narrows the gap the -Wdocumentation pragma needs to cover; fully
replacing the Doxygen-only commands it also hides (item 2c) is left
for after #5267.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 2

* Fix -Wextra-semi-stmt at its actual source, not assert()

clang_flags.cmake blamed the global -Wno-extra-semi-stmt on assert(),
but assert() expands to an expression under glibc and libc++ and does
not trigger this warning. unit-assert_macro.cpp overrides JSON_ASSERT
with "{if (!(x)) ++assert_counter; }", a bare block followed by a
semicolon at every JSON_ASSERT(...) call site in the library; that
was the actual source of 151 of the 208 -Wextra-semi-stmt sites found
in a Clang 22 -Weverything sweep of the test suite with the flag
removed. Switched to the standard do/while(false) macro idiom, which
does not expand to a statement-plus-semicolon, and corrected the
comment to name the remaining source instead: vendored Doctest's
CAPTURE(x) shim, which already ends in a semicolon.

Verified with clang++ -Wextra-semi-stmt (plus the file's other CI
ignores) that unit-assert_macro.cpp now compiles without any
-Wextra-semi-stmt diagnostic.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 8 (step 1 of 2; step 2 covers the CAPTURE() call sites)

* Drop the redundant semicolon from CAPTURE() call sites; remove -Wno-extra-semi-stmt

doctest_compatibility.h defines CAPTURE(x) as DOCTEST_CAPTURE(x); (with
a trailing semicolon baked into the macro), specifically so call sites
do not need to add one themselves; most of the ~267 call sites already
follow that convention. The remaining 64 call sites across 20 files
wrote "CAPTURE(x);" anyway, turning into a statement plus an empty
statement and triggering -Wextra-semi-stmt. Dropped the redundant
semicolon at each of those sites.

With item 6 having already made vendored Doctest a SYSTEM include, and
this the last known source of -Wextra-semi-stmt findings, removed the
flag from clang_flags.cmake entirely.

Verified with clang++ -Wextra-semi-stmt (plus the file's other CI
ignores) that all 20 touched files, plus a file with no CAPTURE() use
(unit-json_pointer.cpp), compile without any -Wextra-semi-stmt
diagnostic.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5725 item 8 (step 2 of 2)

* Switch ci_static_analysis_clang off the frozen LLVM 22 dev image

ubuntu.yml pinned the clang-tidy/clang-tidy-sanitizer/single-binaries
job to silkeh/clang:dev, a tag last pushed 2026-02-18 that reports
"clang version 22.0.0 (...+20251015...)", a pre-release snapshot from
before the LLVM 22 release; the maintainer now updates dev-unstable,
22, and latest instead. Switched to silkeh/clang:22, matching the
other clang jobs on :latest.

Verified with clang-tidy 22.1.8 (the image's actual version) against
this repository's .clang-tidy and library headers what the release
image newly reports compared to :dev:

- readability-redundant-typename fires at ~250 sites across the
  _cpp20-relevant conversion/to_chars headers; the library targets
  C++11 and keeps the typenames, so the check is disabled in
  .clang-tidy, matching how the file already handles checks that
  don't fit a C++11 codebase.
- misc-anonymous-namespace-in-header fires on the two anonymous
  namespaces in from_json.hpp and to_json.hpp; added the alias to
  their existing NOLINT (cert-dcl59-cpp, fuchsia-header-anon-namespaces,
  google-build-namespaces).
- bugprone-std-namespace-modification fires on every addition to
  namespace std: the std::hash, std::formatter and std::swap
  overloads in json.hpp, and the std::tuple_size/std::tuple_element
  specializations in iteration_proxy.hpp (this last file is not named
  in #5725's item 5, found by actually running clang-tidy 22.1.8
  against the current tree). All six are legal, deliberate additions
  to namespace std (explicit/partial specializations of std types, or
  the pre-C++20 std::swap overload); annotated each with the check
  name next to its existing cert-dcl58-cpp NOLINT.
- modernize-avoid-c-style-cast reported nothing new.

Also added clang++-22/21, clang-tidy-22/21, g++-16 and gcov-16 to the
find_program search lists in ci.cmake so a local "maximal warnings"
configure prefers the current toolchain version over an older one on
PATH.

#5725 item 5

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Regenerate cmake/gcc_flags.cmake for GCC 16.2.0

GCC_CXXFLAGS was generated for GCC 15.1.0, but ci_test_gcc and
ci_test_gcc_cxx{11..26} now run in gcc:latest, currently GCC 16.2.0,
so the "maximal warnings" job was missing warnings introduced since
15.1.0 while carrying entries GCC 16 treats as duplicates or no-ops.

Regenerated with https://github.com/nlohmann/gcc_flags (patched
locally to not crash on an option whose "-x c++ <opt> -" probe fails
before it reads stdin, e.g. -Wabi=; the tool otherwise raises
BrokenPipeError instead of recording the option as an error) run
against g++ 16.2.0 in the official gcc:16 Docker image, keeping the
documented -Wno-* exclusions and the same alphabetical placement
scheme as before.

Also added three GCC 16 warnings the generator cannot discover on its
own because it only probes value ranges/lists it finds in the -Q
option name itself, not in the enum choices --help=warnings documents
separately:

- -Wbidi-chars=any, -Wleading-whitespace=spaces: manually verified
  these compile cleanly with g++ 16.2.0.
- -Wstrict-flex-arrays: deliberately NOT added, unlike the other two.
  Without -fstrict-flex-arrays (which the library does not enable, as
  it would change codegen for flexible array members), GCC prints
  "'-Wstrict-flex-arrays' is ignored when '-fstrict-flex-arrays' is
  not present" on every translation unit, and under our -Werror that
  note itself aborts the build. This differs from the harmless
  no-op warnings already kept in the file (-Whsa, -Wsynth,
  -Wunreachable-code, -Wunsafe-loop-optimizations), which emit
  nothing; #5725 item 7 named -Wstrict-flex-arrays as one of the
  flags GCC 16 adds, but did not anticipate this failure mode.

Verified: compiled the library header and a representative set of
test translation units (including ones touched by items 1, 3, 8, 9,
10 of this issue) with the regenerated GCC_CXXFLAGS plus -Werror
under g++ 16.2.0 at -std=c++11 through -std=c++26, with zero warnings;
ran the full local test suite (129/129 passing, unrelated to this
compiler) as a regression check. CI must still confirm the actual
ci_test_gcc / ci_test_standards_gcc targets end to end, since this was
verified with direct g++ invocations rather than through the CMake/
CXXFLAGS environment-variable plumbing in ci.cmake.

#5725 item 7

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Avoid std::basic_string<CharType> for non-character output_adapter CharType

output_adapter<CharType, StringType> defaulted StringType to
std::basic_string<CharType>, and (with JSON_NO_IO undefined) always
declared a std::basic_ostream<CharType>&-taking constructor. For
CharType with no non-deprecated std::char_traits specialization (only
std::uint8_t is ever used this way, by the binary writers), simply
naming either type - as an unused default template argument, or as an
unused, never-called constructor's parameter type - instantiates
std::char_traits<CharType> merely to name it, which some standard
libraries mark deprecated: with the library-wide -Wdocumentation
pragma (item 2's other half, left for a later commit) temporarily
removed, an Apple clang 21 / libc++ TU calling json::to_cbor(j, vec)
with std::vector<std::uint8_t>& got one -Wdeprecated-declarations
warning per binary writer at the old output_adapters.hpp:193.

Replaced the eager std::basic_string<CharType> / std::basic_ostream
<CharType> defaults with a bool-tagged partial specialization (not
std::conditional, which requires naming both branches' types up
front regardless of which is selected, reproducing the same warning)
that only ever names std::basic_string<CharType> / std::basic_ostream
<CharType> when CharType is actually one of char, wchar_t, char16_t,
char32_t, or (with __cpp_lib_char8_t) char8_t. For any other
CharType, output_adapter's StringType and ostream-constructor
parameter fall back to two distinct empty placeholder types, kept
distinct so the two constructor overloads do not collide into a
single redeclaration.

Public API / behavior: passing a std::basic_string<std::uint8_t>& or
std::basic_ostream<std::uint8_t>& directly to a binary writer's
output_adapter now fails to compile instead of compiling with a
deprecation warning; this was neither documented nor tested. All
documented uses (std::vector<CharType>, std::basic_ostream<CharType>
and StringType for character CharType) are unaffected.

Verified with Apple clang 21 / libc++, with the two -Wdocumentation*
"ignored" pragma lines in macro_scope.hpp temporarily removed and
-std=c++11/c++20 plus the project's -Weverything flag set: calling
to_cbor/to_msgpack/to_ubjson/to_bjdata/to_bson/to_bon8 on a
std::vector<std::uint8_t> now produces no char_traits<unsigned char>
(or any other) deprecation warning, while the char-based string- and
ostream-adapter paths, and a to_cbor/from_cbor round trip, still
compile and run correctly; also verified with GCC 16.2.0. Ran the
full local test suite, including the binary-format unit tests
(unit-cbor, unit-msgpack, unit-ubjson, unit-bjdata, unit-bson,
unit-bon8, unit-binary_writer_sinks, unit-binary_formats,
unit-custom-binary-type): 129/129 passing.

#5725 item 2 (step a)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove the library-wide -Wdocumentation pragma; fix what it hid

macro_scope.hpp / macro_unscope.hpp pushed and popped a Clang
diagnostic region over the entire library that ignored -Wdocumentation
and -Wdocumentation-unknown-command. Removed both pragmas and fixed
every finding a full -Wdocumentation (which implies
-Wdocumentation-unknown-command and -Wdocumentation-deprecated-sync)
build reports, so the library now compiles clean under Clang's
documentation checks without a blanket suppression. Overlaps #5267,
which is still open and edits a nearby doc block (json.hpp's
get()/get_impl() @return, already fixed in the item 2 step (b) commit
of this branch); this commit does not touch that block again.

Unknown Doxygen alias commands (Doxyfile removed in #3071, so these
were never rendered by anything) rewritten as plain prose, keeping the
same information:
- @requirement REQ-JSON-01 / REQ-JSON-02 (iter_impl.hpp,
  json_reverse_iterator.hpp): now "This class satisfies the following
  concept requirements (REQ-JSON-0N):".
- @liveexample{prose,example-id} (three sites in json.hpp): kept the
  prose, dropped the command wrapper and the trailing example-id
  (docs/mkdocs/docs/examples/*.cpp still exist and are used directly
  by the rendered docs, not through this in-header alias) and
  unescaped the "\," commas that were only needed for the old alias's
  comma-separated argument syntax.
- @complexity X (json.hpp x4, json_pointer.hpp x2, serializer.hpp x1):
  now "Complexity: X".

Backslash sequences Clang's comment lexer tried to parse as commands,
escaped to render as literal backslashes:
- lexer.hpp get_codepoint(): two `\u` occurrences.
- binary_reader.hpp get_bson_cstr() / get_bson_cstr_bulk(): two
  `\x00` occurrences.
- serializer.hpp: three `\uXXXX` occurrences (constructor @param,
  append_codepoint_to_string_buffer() @brief, and the ensure_ascii
  member comment).

One finding remained after all of the above: Clang reports
"declaration is marked with '@deprecated' command but does not have a
deprecation attribute" on the deprecated sax_parse(span_input_adapter&&, ...)
overload, even though JSON_HEDLEY_DEPRECATED_FOR does expand to
__attribute__((deprecated(...))) for Clang. Several isolated
reproductions of this exact declaration shape - doc comment,
template<>, two stacked __attribute__ macros, an overload set sharing
the name - did not reproduce the warning, so this looks like a
Clang comment/declaration-association quirk specific to this overload
inside the much larger basic_json class template, not an actual
documentation defect. Rather than keep the pragma library-wide for one
Clang false positive, added a tightly scoped
-Wdocumentation-deprecated-sync push/pop around just that overload.

Verified with Apple clang 21 and the project's actual -Weverything
flag set (cmake/clang_flags.cmake) on the full header at -std=c++11
and -std=c++20: zero -Wdocumentation* diagnostics. Also compiled
clean with GCC 16.2.0 (the pragmas are already __clang__-gated, so
this only confirms no unrelated breakage). Ran make check-amalgamation
and the full local test suite: 129/129 passing.

#5725 item 2 (step c)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Take the JSON value by const reference in the array and tuple from_json paths

Review feedback on #5737 (gregmarr): once the no-op std::forward calls
are gone, the forwarding references have no purpose. from_json_fn
passes the value as const BasicJsonType&, so these functions were only
ever instantiated with a const lvalue anyway.

The std::array, std::pair and std::tuple overloads of from_json and
their helpers now take const BasicJsonType& and pass j on unchanged.
Because the deduced BasicJsonType is now the plain type, tuple_type and
the static_assert name const BasicJsonType& explicitly, so the
reference checks are unchanged: get<std::tuple<const std::string&>>()
still works, and get<std::tuple<std::string&>>() still fails the same
static_assert. from_json_tuple_get_impl keeps its forwarding reference,
since tuple_type calls it through std::declval.

Behavior, the public API and the ABI do not change. unit-conversions,
unit-constructor1, unit-udt, unit-udt_macro, unit-regression1/2/3,
unit-deserialization, unit-noexcept, unit-items, unit-allocator,
unit-custom-object-type, unit-ordered_json2 and
unit-brace-init-copy-semantics pass at C++11, C++17 and C++20 with
unchanged assertion counts. Ran make amalgamate.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:37:47 +02:00

2156 lines
76 KiB
C++

// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <array> // array
#include <cstddef> // size_t
#include <cstdint> // uint32_t
#include <cstdio> // snprintf
#include <initializer_list> // initializer_list
#include <string> // char_traits, string
#include <utility> // move
#include <vector> // vector
#include <nlohmann/detail/input/input_adapters.hpp>
#include <nlohmann/detail/input/number_parse.hpp>
#include <nlohmann/detail/input/position_t.hpp>
#include <nlohmann/detail/input/string_scan.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/type_traits.hpp>
#include <nlohmann/detail/string_utils.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
///////////
// lexer //
///////////
template<typename BasicJsonType>
class lexer_base
{
public:
/// token types for the parser
enum class token_type
{
uninitialized, ///< indicating the scanner is uninitialized
literal_true, ///< the `true` literal
literal_false, ///< the `false` literal
literal_null, ///< the `null` literal
value_string, ///< a string -- use get_string() for actual value
value_unsigned, ///< an unsigned integer -- use get_number_unsigned() for actual value
value_integer, ///< a signed integer -- use get_number_integer() for actual value
value_float, ///< an floating point number -- use get_number_float() for actual value
begin_array, ///< the character for array begin `[`
begin_object, ///< the character for object begin `{`
end_array, ///< the character for array end `]`
end_object, ///< the character for object end `}`
name_separator, ///< the name separator `:`
value_separator, ///< the value separator `,`
parse_error, ///< indicating a parse error
end_of_input, ///< indicating the end of the input buffer
literal_or_value ///< a literal or the begin of a value (only for diagnostics)
};
/// return name of values of type token_type (only used for errors)
JSON_HEDLEY_RETURNS_NON_NULL
JSON_HEDLEY_CONST
static const char* token_type_name(const token_type t) noexcept
{
switch (t)
{
case token_type::uninitialized:
return "<uninitialized>";
case token_type::literal_true:
return "true literal";
case token_type::literal_false:
return "false literal";
case token_type::literal_null:
return "null literal";
case token_type::value_string:
return "string literal";
case token_type::value_unsigned:
case token_type::value_integer:
case token_type::value_float:
return "number literal";
case token_type::begin_array:
return "'['";
case token_type::begin_object:
return "'{'";
case token_type::end_array:
return "']'";
case token_type::end_object:
return "'}'";
case token_type::name_separator:
return "':'";
case token_type::value_separator:
return "','";
case token_type::parse_error:
return "<parse error>";
case token_type::end_of_input:
return "end of input";
case token_type::literal_or_value:
return "'[', '{', or a literal";
// LCOV_EXCL_START
default: // catch non-enum values
return "unknown token";
// LCOV_EXCL_STOP
}
}
};
// Detect whether an input adapter can reconstruct already-consumed input on
// demand (see iterator_input_adapter::supports_seek). Adapters that do not
// expose the flag - e.g. file, stream, wide-string, and user-defined adapters -
// are treated as non-seekable streaming input, for which the lexer keeps
// copying every scanned character eagerly. The value is read via tag dispatch
// on is_detected so the flag is only referenced for adapters that provide it.
template<typename InputAdapterType>
using detect_supports_seek = decltype(InputAdapterType::supports_seek);
template<typename InputAdapterType>
constexpr bool input_adapter_supports_seek(std::true_type /*detected*/)
{
return InputAdapterType::supports_seek;
}
template<typename InputAdapterType>
constexpr bool input_adapter_supports_seek(std::false_type /*detected*/)
{
return false;
}
// Detect whether an input adapter reads with one character of lookahead that
// can be left in the input (see input_stream_adapter::supports_lookahead,
// which is only defined with JSON_PRECISE_STREAM_POSITION), detected like
// supports_seek above.
template<typename InputAdapterType>
using detect_supports_lookahead = decltype(InputAdapterType::supports_lookahead);
template<typename InputAdapterType>
constexpr bool input_adapter_supports_lookahead(std::true_type /*detected*/)
{
return InputAdapterType::supports_lookahead;
}
template<typename InputAdapterType>
constexpr bool input_adapter_supports_lookahead(std::false_type /*detected*/)
{
return false;
}
// Detect whether an input adapter exposes a contiguous byte block that the
// lexer can scan directly (see iterator_input_adapter::supports_bulk_scan).
// Adapters without the flag - file, stream, wide-string, user-defined - fall
// back to the character-at-a-time string scanner.
template<typename InputAdapterType>
using detect_supports_bulk_scan = decltype(InputAdapterType::supports_bulk_scan);
template<typename InputAdapterType>
constexpr bool input_adapter_supports_bulk_scan(std::true_type /*detected*/)
{
return InputAdapterType::supports_bulk_scan;
}
template<typename InputAdapterType>
constexpr bool input_adapter_supports_bulk_scan(std::false_type /*detected*/)
{
return false;
}
/*!
@brief lexical analysis
This class organizes the lexical analysis during JSON deserialization.
*/
template<typename BasicJsonType, typename InputAdapterType>
class lexer : public lexer_base<BasicJsonType>
{
using number_integer_t = typename BasicJsonType::number_integer_t;
using number_unsigned_t = typename BasicJsonType::number_unsigned_t;
using number_float_t = typename BasicJsonType::number_float_t;
using string_t = typename BasicJsonType::string_t;
using char_type = typename InputAdapterType::char_type;
using char_int_type = typename char_traits<char_type>::int_type;
/// whether the last read token can be reconstructed from the input adapter
/// on demand (in error paths) instead of being copied on every scanned
/// character; see input_adapter_supports_seek
static constexpr bool lazy_token_string =
input_adapter_supports_seek<InputAdapterType>(is_detected<detect_supports_seek, InputAdapterType> {});
/// whether a simulated unget can be passed on to the input adapter, which
/// then leaves the character in the input; see
/// input_adapter_supports_lookahead
static constexpr bool can_release_lookahead =
input_adapter_supports_lookahead<InputAdapterType>(is_detected<detect_supports_lookahead, InputAdapterType> {});
/// whether string scanning may bulk-consume runs of ordinary characters
/// directly from a contiguous input buffer (SWAR fast path). This requires
/// the token to be reconstructible lazily (lazy_token_string), so bypassing
/// the per-character capture in get() cannot lose error diagnostics.
static constexpr bool bulk_scan =
lazy_token_string
&& input_adapter_supports_bulk_scan<InputAdapterType>(is_detected<detect_supports_bulk_scan, InputAdapterType> {});
public:
using token_type = typename lexer_base<BasicJsonType>::token_type;
explicit lexer(InputAdapterType&& adapter, bool ignore_comments_ = false, bool discard_number_values_ = false) noexcept
: ia(std::move(adapter))
, ignore_comments(ignore_comments_)
, discard_number_values(discard_number_values_)
{}
// deleted because of pointer members
lexer(const lexer&) = delete;
lexer(lexer&&) = default; // NOLINT(hicpp-noexcept-move,performance-noexcept-move-constructor)
lexer& operator=(lexer&) = delete;
lexer& operator=(lexer&&) = default; // NOLINT(hicpp-noexcept-move,performance-noexcept-move-constructor)
~lexer() = default;
private:
/////////////////////
// scan functions
/////////////////////
/*!
@brief get codepoint from 4 hex characters following `\\u`
For input "\\u c1 c2 c3 c4" the codepoint is:
(c1 * 0x1000) + (c2 * 0x0100) + (c3 * 0x0010) + c4
= (c1 << 12) + (c2 << 8) + (c3 << 4) + (c4 << 0)
Furthermore, the possible characters '0'..'9', 'A'..'F', and 'a'..'f'
must be converted to the integers 0x0..0x9, 0xA..0xF, 0xA..0xF, resp. The
conversion is done by subtracting the offset (0x30, 0x37, and 0x57)
between the ASCII value of the character and the desired integer value.
@return codepoint (0x0000..0xFFFF) or -1 in case of an error (e.g. EOF or
non-hex character)
*/
int get_codepoint()
{
// this function only makes sense after reading `\u`
JSON_ASSERT(current == 'u');
int codepoint = 0;
const auto factors = { 12u, 8u, 4u, 0u };
for (const auto factor : factors)
{
get();
if (current >= '0' && current <= '9')
{
codepoint += static_cast<int>((static_cast<unsigned int>(current) - 0x30u) << factor);
}
else if (current >= 'A' && current <= 'F')
{
codepoint += static_cast<int>((static_cast<unsigned int>(current) - 0x37u) << factor);
}
else if (current >= 'a' && current <= 'f')
{
codepoint += static_cast<int>((static_cast<unsigned int>(current) - 0x57u) << factor);
}
else
{
return -1;
}
}
JSON_ASSERT(0x0000 <= codepoint && codepoint <= 0xFFFF);
return codepoint;
}
/*!
@brief check if the next byte(s) are inside a given range
Adds the current byte and, for each passed range, reads a new byte and
checks if it is inside the range. If a violation was detected, set up an
error message and return false. Otherwise, return true.
@param[in] ranges list of integers; interpreted as list of pairs of
inclusive lower and upper bound, respectively
@pre The passed list @a ranges must have 2, 4, or 6 elements; that is,
1, 2, or 3 pairs. This precondition is enforced by an assertion.
@return true if and only if no range violation was detected
*/
bool next_byte_in_range(std::initializer_list<char_int_type> ranges)
{
JSON_ASSERT(ranges.size() == 2 || ranges.size() == 4 || ranges.size() == 6);
add(current);
for (auto range = ranges.begin(); range != ranges.end(); ++range)
{
get();
if (JSON_HEDLEY_LIKELY(*range <= current && current <= *(++range))) // NOLINT(bugprone-inc-dec-in-conditions)
{
add(current);
}
else
{
error_message = "invalid string: ill-formed UTF-8 byte";
return false;
}
}
return true;
}
/// contiguous input: bulk-append the run of ordinary characters and complete
/// well-formed UTF-8 sequences starting at the current read position, leaving
/// the first byte that needs individual handling (the closing quote, an
/// escape, a control character, or an ill-formed UTF-8 byte) for get()
void scan_string_bulk(std::true_type /*bulk*/)
{
// a pending unget must be consumed through the normal path first
if (next_unget)
{
return;
}
const std::size_t remaining = ia.bulk_remaining();
if (remaining == 0)
{
return;
}
const auto* const data = reinterpret_cast<const unsigned char*>(ia.bulk_data());
const std::size_t pos = string_bulk_run(data, remaining);
if (pos == 0)
{
return;
}
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), pos);
ia.bulk_skip(pos);
// the run contains no newline (all bytes < 0x20 are treated as special),
// so only the flat character counters advance
position.chars_read_total += pos;
position.chars_read_current_line += pos;
}
/// streaming input: no bulk fast path
void scan_string_bulk(std::false_type /*bulk*/) const noexcept {}
/*!
@brief scan a string literal
This function scans a string according to Sect. 7 of RFC 8259. While
scanning, bytes are escaped and copied into buffer token_buffer. Then the
function returns successfully, token_buffer is *not* null-terminated (as it
may contain \0 bytes), and token_buffer.size() is the number of bytes in the
string.
@return token_type::value_string if string could be successfully scanned,
token_type::parse_error otherwise
@note In case of errors, variable error_message contains a textual
description.
*/
token_type scan_string()
{
// reset token_buffer (ignore opening quote)
reset();
// we entered the function by reading an open quote
JSON_ASSERT(current == '\"');
while (true)
{
// bulk-consume ordinary characters from contiguous input, then
// handle the next special byte through the switch below
scan_string_bulk(std::integral_constant<bool, bulk_scan> {});
// get the next character
switch (get())
{
// end of file while parsing the string
case char_traits<char_type>::eof():
{
error_message = "invalid string: missing closing quote";
return token_type::parse_error;
}
// closing quote
case '\"':
{
return token_type::value_string;
}
// escapes
case '\\':
{
switch (get())
{
// quotation mark
case '\"':
add('\"');
break;
// reverse solidus
case '\\':
add('\\');
break;
// solidus
case '/':
add('/');
break;
// backspace
case 'b':
add('\b');
break;
// form feed
case 'f':
add('\f');
break;
// line feed
case 'n':
add('\n');
break;
// carriage return
case 'r':
add('\r');
break;
// tab
case 't':
add('\t');
break;
// unicode escapes
case 'u':
{
const int codepoint1 = get_codepoint();
int codepoint = codepoint1; // start with codepoint1
if (JSON_HEDLEY_UNLIKELY(codepoint1 == -1))
{
error_message = "invalid string: '\\u' must be followed by 4 hex digits";
return token_type::parse_error;
}
// check if code point is a high surrogate
if (0xD800 <= codepoint1 && codepoint1 <= 0xDBFF)
{
// expect next \uxxxx entry
if (JSON_HEDLEY_LIKELY(get() == '\\' && get() == 'u'))
{
const int codepoint2 = get_codepoint();
if (JSON_HEDLEY_UNLIKELY(codepoint2 == -1))
{
error_message = "invalid string: '\\u' must be followed by 4 hex digits";
return token_type::parse_error;
}
// check if codepoint2 is a low surrogate
if (JSON_HEDLEY_LIKELY(0xDC00 <= codepoint2 && codepoint2 <= 0xDFFF))
{
// overwrite codepoint
codepoint = static_cast<int>(
// high surrogate occupies the most significant 22 bits
(static_cast<unsigned int>(codepoint1) << 10u)
// low surrogate occupies the least significant 15 bits
+ static_cast<unsigned int>(codepoint2)
// there is still the 0xD800, 0xDC00, and 0x10000 noise
// in the result, so we have to subtract with:
// (0xD800 << 10) + DC00 - 0x10000 = 0x35FDC00
- 0x35FDC00u);
}
else
{
error_message = "invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF";
return token_type::parse_error;
}
}
else
{
error_message = "invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF";
return token_type::parse_error;
}
}
else
{
if (JSON_HEDLEY_UNLIKELY(0xDC00 <= codepoint1 && codepoint1 <= 0xDFFF))
{
error_message = "invalid string: surrogate U+DC00..U+DFFF must follow U+D800..U+DBFF";
return token_type::parse_error;
}
}
// the result of the above calculation yields a proper codepoint
JSON_ASSERT(0x00 <= codepoint && codepoint <= 0x10FFFF);
// translate codepoint into bytes
encode_utf8(static_cast<std::uint32_t>(codepoint), [this](std::uint32_t byte)
{
add(static_cast<char_int_type>(byte));
});
break;
}
// other characters after escape
default:
error_message = "invalid string: forbidden character after backslash";
return token_type::parse_error;
}
break;
}
// invalid control characters
case 0x00:
{
error_message = "invalid string: control character U+0000 (NUL) must be escaped to \\u0000";
return token_type::parse_error;
}
case 0x01:
{
error_message = "invalid string: control character U+0001 (SOH) must be escaped to \\u0001";
return token_type::parse_error;
}
case 0x02:
{
error_message = "invalid string: control character U+0002 (STX) must be escaped to \\u0002";
return token_type::parse_error;
}
case 0x03:
{
error_message = "invalid string: control character U+0003 (ETX) must be escaped to \\u0003";
return token_type::parse_error;
}
case 0x04:
{
error_message = "invalid string: control character U+0004 (EOT) must be escaped to \\u0004";
return token_type::parse_error;
}
case 0x05:
{
error_message = "invalid string: control character U+0005 (ENQ) must be escaped to \\u0005";
return token_type::parse_error;
}
case 0x06:
{
error_message = "invalid string: control character U+0006 (ACK) must be escaped to \\u0006";
return token_type::parse_error;
}
case 0x07:
{
error_message = "invalid string: control character U+0007 (BEL) must be escaped to \\u0007";
return token_type::parse_error;
}
case 0x08:
{
error_message = "invalid string: control character U+0008 (BS) must be escaped to \\u0008 or \\b";
return token_type::parse_error;
}
case 0x09:
{
error_message = "invalid string: control character U+0009 (HT) must be escaped to \\u0009 or \\t";
return token_type::parse_error;
}
case 0x0A:
{
error_message = "invalid string: control character U+000A (LF) must be escaped to \\u000A or \\n";
return token_type::parse_error;
}
case 0x0B:
{
error_message = "invalid string: control character U+000B (VT) must be escaped to \\u000B";
return token_type::parse_error;
}
case 0x0C:
{
error_message = "invalid string: control character U+000C (FF) must be escaped to \\u000C or \\f";
return token_type::parse_error;
}
case 0x0D:
{
error_message = "invalid string: control character U+000D (CR) must be escaped to \\u000D or \\r";
return token_type::parse_error;
}
case 0x0E:
{
error_message = "invalid string: control character U+000E (SO) must be escaped to \\u000E";
return token_type::parse_error;
}
case 0x0F:
{
error_message = "invalid string: control character U+000F (SI) must be escaped to \\u000F";
return token_type::parse_error;
}
case 0x10:
{
error_message = "invalid string: control character U+0010 (DLE) must be escaped to \\u0010";
return token_type::parse_error;
}
case 0x11:
{
error_message = "invalid string: control character U+0011 (DC1) must be escaped to \\u0011";
return token_type::parse_error;
}
case 0x12:
{
error_message = "invalid string: control character U+0012 (DC2) must be escaped to \\u0012";
return token_type::parse_error;
}
case 0x13:
{
error_message = "invalid string: control character U+0013 (DC3) must be escaped to \\u0013";
return token_type::parse_error;
}
case 0x14:
{
error_message = "invalid string: control character U+0014 (DC4) must be escaped to \\u0014";
return token_type::parse_error;
}
case 0x15:
{
error_message = "invalid string: control character U+0015 (NAK) must be escaped to \\u0015";
return token_type::parse_error;
}
case 0x16:
{
error_message = "invalid string: control character U+0016 (SYN) must be escaped to \\u0016";
return token_type::parse_error;
}
case 0x17:
{
error_message = "invalid string: control character U+0017 (ETB) must be escaped to \\u0017";
return token_type::parse_error;
}
case 0x18:
{
error_message = "invalid string: control character U+0018 (CAN) must be escaped to \\u0018";
return token_type::parse_error;
}
case 0x19:
{
error_message = "invalid string: control character U+0019 (EM) must be escaped to \\u0019";
return token_type::parse_error;
}
case 0x1A:
{
error_message = "invalid string: control character U+001A (SUB) must be escaped to \\u001A";
return token_type::parse_error;
}
case 0x1B:
{
error_message = "invalid string: control character U+001B (ESC) must be escaped to \\u001B";
return token_type::parse_error;
}
case 0x1C:
{
error_message = "invalid string: control character U+001C (FS) must be escaped to \\u001C";
return token_type::parse_error;
}
case 0x1D:
{
error_message = "invalid string: control character U+001D (GS) must be escaped to \\u001D";
return token_type::parse_error;
}
case 0x1E:
{
error_message = "invalid string: control character U+001E (RS) must be escaped to \\u001E";
return token_type::parse_error;
}
case 0x1F:
{
error_message = "invalid string: control character U+001F (US) must be escaped to \\u001F";
return token_type::parse_error;
}
// U+0020..U+007F (except U+0022 (quote) and U+005C (backspace))
case 0x20:
case 0x21:
case 0x23:
case 0x24:
case 0x25:
case 0x26:
case 0x27:
case 0x28:
case 0x29:
case 0x2A:
case 0x2B:
case 0x2C:
case 0x2D:
case 0x2E:
case 0x2F:
case 0x30:
case 0x31:
case 0x32:
case 0x33:
case 0x34:
case 0x35:
case 0x36:
case 0x37:
case 0x38:
case 0x39:
case 0x3A:
case 0x3B:
case 0x3C:
case 0x3D:
case 0x3E:
case 0x3F:
case 0x40:
case 0x41:
case 0x42:
case 0x43:
case 0x44:
case 0x45:
case 0x46:
case 0x47:
case 0x48:
case 0x49:
case 0x4A:
case 0x4B:
case 0x4C:
case 0x4D:
case 0x4E:
case 0x4F:
case 0x50:
case 0x51:
case 0x52:
case 0x53:
case 0x54:
case 0x55:
case 0x56:
case 0x57:
case 0x58:
case 0x59:
case 0x5A:
case 0x5B:
case 0x5D:
case 0x5E:
case 0x5F:
case 0x60:
case 0x61:
case 0x62:
case 0x63:
case 0x64:
case 0x65:
case 0x66:
case 0x67:
case 0x68:
case 0x69:
case 0x6A:
case 0x6B:
case 0x6C:
case 0x6D:
case 0x6E:
case 0x6F:
case 0x70:
case 0x71:
case 0x72:
case 0x73:
case 0x74:
case 0x75:
case 0x76:
case 0x77:
case 0x78:
case 0x79:
case 0x7A:
case 0x7B:
case 0x7C:
case 0x7D:
case 0x7E:
case 0x7F:
{
add(current);
break;
}
// U+0080..U+07FF: bytes C2..DF 80..BF
case 0xC2:
case 0xC3:
case 0xC4:
case 0xC5:
case 0xC6:
case 0xC7:
case 0xC8:
case 0xC9:
case 0xCA:
case 0xCB:
case 0xCC:
case 0xCD:
case 0xCE:
case 0xCF:
case 0xD0:
case 0xD1:
case 0xD2:
case 0xD3:
case 0xD4:
case 0xD5:
case 0xD6:
case 0xD7:
case 0xD8:
case 0xD9:
case 0xDA:
case 0xDB:
case 0xDC:
case 0xDD:
case 0xDE:
case 0xDF:
{
if (JSON_HEDLEY_UNLIKELY(!next_byte_in_range({0x80, 0xBF})))
{
return token_type::parse_error;
}
break;
}
// U+0800..U+0FFF: bytes E0 A0..BF 80..BF
case 0xE0:
{
if (JSON_HEDLEY_UNLIKELY(!(next_byte_in_range({0xA0, 0xBF, 0x80, 0xBF}))))
{
return token_type::parse_error;
}
break;
}
// U+1000..U+CFFF: bytes E1..EC 80..BF 80..BF
// U+E000..U+FFFF: bytes EE..EF 80..BF 80..BF
case 0xE1:
case 0xE2:
case 0xE3:
case 0xE4:
case 0xE5:
case 0xE6:
case 0xE7:
case 0xE8:
case 0xE9:
case 0xEA:
case 0xEB:
case 0xEC:
case 0xEE:
case 0xEF:
{
if (JSON_HEDLEY_UNLIKELY(!(next_byte_in_range({0x80, 0xBF, 0x80, 0xBF}))))
{
return token_type::parse_error;
}
break;
}
// U+D000..U+D7FF: bytes ED 80..9F 80..BF
case 0xED:
{
if (JSON_HEDLEY_UNLIKELY(!(next_byte_in_range({0x80, 0x9F, 0x80, 0xBF}))))
{
return token_type::parse_error;
}
break;
}
// U+10000..U+3FFFF F0 90..BF 80..BF 80..BF
case 0xF0:
{
if (JSON_HEDLEY_UNLIKELY(!(next_byte_in_range({0x90, 0xBF, 0x80, 0xBF, 0x80, 0xBF}))))
{
return token_type::parse_error;
}
break;
}
// U+40000..U+FFFFF F1..F3 80..BF 80..BF 80..BF
case 0xF1:
case 0xF2:
case 0xF3:
{
if (JSON_HEDLEY_UNLIKELY(!(next_byte_in_range({0x80, 0xBF, 0x80, 0xBF, 0x80, 0xBF}))))
{
return token_type::parse_error;
}
break;
}
// U+100000..U+10FFFF F4 80..8F 80..BF 80..BF
case 0xF4:
{
if (JSON_HEDLEY_UNLIKELY(!(next_byte_in_range({0x80, 0x8F, 0x80, 0xBF, 0x80, 0xBF}))))
{
return token_type::parse_error;
}
break;
}
// the remaining bytes (80..C1 and F5..FF) are ill-formed
default:
{
error_message = "invalid string: ill-formed UTF-8 byte";
return token_type::parse_error;
}
}
}
}
/*!
* @brief scan a comment
* @return whether comment could be scanned successfully
*/
bool scan_comment()
{
switch (get())
{
// single-line comments skip input until a newline or EOF is read
case '/':
{
while (true)
{
switch (get())
{
case '\n':
case '\r':
case char_traits<char_type>::eof():
#if !JSON_STRICT_NUL_HANDLING
case '\0':
#endif
return true;
default:
break;
}
}
JSON_HEDLEY_UNREACHABLE();
}
// multi-line comments skip input until */ is read
case '*':
{
while (true)
{
switch (get())
{
#if !JSON_STRICT_NUL_HANDLING
case '\0':
#endif
case char_traits<char_type>::eof():
{
error_message = "invalid comment; missing closing '*/'";
return false;
}
case '*':
{
switch (get())
{
case '/':
return true;
default:
{
unget();
continue;
}
}
}
default:
continue;
}
}
JSON_HEDLEY_UNREACHABLE();
}
// unexpected character after reading '/'
default:
{
error_message = "invalid comment; expecting '/' or '*' after '/'";
return false;
}
}
}
/*!
@brief scan a number literal
This function scans a string according to Sect. 6 of RFC 8259.
The function is realized with a deterministic finite state machine derived
from the grammar described in RFC 8259. Starting in state "init", the
input is read and used to determined the next state. Only state "done"
accepts the number. State "error" is a trap state to model errors. In the
table below, "anything" means any character but the ones listed before.
state | 0 | 1-9 | e E | + | - | . | anything
---------|----------|----------|----------|---------|---------|----------|-----------
init | zero | any1 | [error] | [error] | minus | [error] | [error]
minus | zero | any1 | [error] | [error] | [error] | [error] | [error]
zero | done | done | exponent | done | done | decimal1 | done
any1 | any1 | any1 | exponent | done | done | decimal1 | done
decimal1 | decimal2 | decimal2 | [error] | [error] | [error] | [error] | [error]
decimal2 | decimal2 | decimal2 | exponent | done | done | done | done
exponent | any2 | any2 | [error] | sign | sign | [error] | [error]
sign | any2 | any2 | [error] | [error] | [error] | [error] | [error]
any2 | any2 | any2 | done | done | done | done | done
The state machine is realized with one label per state (prefixed with
"scan_number_") and `goto` statements between them. The state machine
contains cycles, but any cycle can be left when EOF is read. Therefore,
the function is guaranteed to terminate.
During scanning, the read bytes are stored in token_buffer. This string is
then converted to a signed integer, an unsigned integer, or a
floating-point number.
@return token_type::value_unsigned, token_type::value_integer, or
token_type::value_float if number could be successfully scanned,
token_type::parse_error otherwise
@note The scanner is independent of the current locale: token_buffer
always holds `.`. Only the std::strtod fallback of convert_number()
depends on the locale, and it looks up the decimal point right
before converting (see detail::convert_float_locale_aware()).
*/
token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated.
{
// reset token_buffer to store the number's bytes
reset();
// the type of the parsed number; initially set to unsigned; will be
// changed if minus sign, decimal point, or exponent is read
token_type number_type = token_type::value_unsigned;
// offset just past the last mantissa byte in token_buffer (i.e. the
// index of 'e'/'E', or the whole token when there is no exponent).
// convert_number() uses it to count significant digits; npos means
// "not seen an exponent yet" and is resolved at scan_number_done
std::size_t mantissa_end = std::string::npos;
// state (init): we just found out we need to scan a number
switch (current)
{
case '-':
{
add(current);
goto scan_number_minus;
}
case '0':
{
add(current);
goto scan_number_zero;
}
case '1':
case '2':
case '3':
case '4':
case '5':
case '6':
case '7':
case '8':
case '9':
{
add(current);
goto scan_number_any1;
}
// all other characters are rejected outside scan_number()
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
scan_number_minus:
// state: we just parsed a leading minus sign
number_type = token_type::value_integer;
switch (get())
{
case '0':
{
add(current);
goto scan_number_zero;
}
case '1':
case '2':
case '3':
case '4':
case '5':
case '6':
case '7':
case '8':
case '9':
{
add(current);
goto scan_number_any1;
}
default:
{
error_message = "invalid number; expected digit after '-'";
return token_type::parse_error;
}
}
scan_number_zero:
// state: we just parse a zero (maybe with a leading minus sign)
switch (get())
{
case '.':
{
add(current);
decimal_point_position = token_buffer.size() - 1;
goto scan_number_decimal1;
}
case 'e':
case 'E':
{
add(current);
goto scan_number_exponent;
}
default:
goto scan_number_done;
}
scan_number_any1:
// state: we just parsed a number 0-9 (maybe with a leading minus sign)
switch (get())
{
case '0':
case '1':
case '2':
case '3':
case '4':
case '5':
case '6':
case '7':
case '8':
case '9':
{
add(current);
goto scan_number_any1;
}
case '.':
{
add(current);
decimal_point_position = token_buffer.size() - 1;
goto scan_number_decimal1;
}
case 'e':
case 'E':
{
add(current);
goto scan_number_exponent;
}
default:
goto scan_number_done;
}
scan_number_decimal1:
// state: we just parsed a decimal point
number_type = token_type::value_float;
switch (get())
{
case '0':
case '1':
case '2':
case '3':
case '4':
case '5':
case '6':
case '7':
case '8':
case '9':
{
add(current);
goto scan_number_decimal2;
}
default:
{
error_message = "invalid number; expected digit after '.'";
return token_type::parse_error;
}
}
scan_number_decimal2:
// we just parsed at least one number after a decimal point
switch (get())
{
case '0':
case '1':
case '2':
case '3':
case '4':
case '5':
case '6':
case '7':
case '8':
case '9':
{
add(current);
goto scan_number_decimal2;
}
case 'e':
case 'E':
{
add(current);
goto scan_number_exponent;
}
default:
goto scan_number_done;
}
scan_number_exponent:
// we just parsed an exponent
number_type = token_type::value_float;
// this label is reached only right after the 'e'/'E' was appended (from
// the zero, any1, and decimal2 states), so the mantissa ends before it
mantissa_end = token_buffer.size() - 1;
switch (get())
{
case '+':
case '-':
{
add(current);
goto scan_number_sign;
}
case '0':
case '1':
case '2':
case '3':
case '4':
case '5':
case '6':
case '7':
case '8':
case '9':
{
add(current);
goto scan_number_any2;
}
default:
{
error_message =
"invalid number; expected '+', '-', or digit after exponent";
return token_type::parse_error;
}
}
scan_number_sign:
// we just parsed an exponent sign
switch (get())
{
case '0':
case '1':
case '2':
case '3':
case '4':
case '5':
case '6':
case '7':
case '8':
case '9':
{
add(current);
goto scan_number_any2;
}
default:
{
error_message = "invalid number; expected digit after exponent sign";
return token_type::parse_error;
}
}
scan_number_any2:
// we just parsed a number after the exponent or exponent sign
switch (get())
{
case '0':
case '1':
case '2':
case '3':
case '4':
case '5':
case '6':
case '7':
case '8':
case '9':
{
add(current);
goto scan_number_any2;
}
default:
goto scan_number_done;
}
scan_number_done:
// unget the character after the number (we only read it to know that
// we are done scanning a number)
unget();
// no exponent was scanned: the mantissa spans the whole token
if (mantissa_end == std::string::npos)
{
mantissa_end = token_buffer.size();
}
return convert_number(number_type, mantissa_end);
}
/*!
@brief convert an already-validated integer token to its value
The digit sequence in [first, last) has been validated by the caller, so a
dedicated parser can avoid the locale/errno overhead of std::strtoull.
@return the token type on success; token_type::uninitialized if @a
number_type is not an integer type or the value does not fit, in
which case the caller falls back to the floating-point conversion
(matching the previous std::strtoull/std::strtoll behavior)
*/
token_type convert_integer(token_type number_type, const char* first, const char* last)
{
if (number_type == token_type::value_unsigned)
{
if (parse_integer_unsigned(first, last, value_unsigned))
{
return token_type::value_unsigned;
}
}
else if (number_type == token_type::value_integer)
{
if (parse_integer_signed(first, last, value_integer))
{
return token_type::value_integer;
}
}
return token_type::uninitialized;
}
/*!
@brief convert the number text in token_buffer to its value and token type
The digit sequence in token_buffer has already been validated (by the
scan_number() state machine or by the contiguous fast path) and holds '.'
as decimal point, independent of the locale. Integers are parsed first and fall
back to floating point on overflow. This is shared so both scanners produce
identical results.
@param[in] mantissa_end offset just past the last mantissa byte in
token_buffer (the index of 'e'/'E', or
token_buffer.size() when there is no exponent);
used to skip Clinger's fast path when it cannot
possibly succeed - see detail::mantissa_fits_clinger()
*/
token_type convert_number(token_type number_type, std::size_t mantissa_end)
{
// accept() only needs to know whether the input is valid, so it sets
// discard_number_values (see json.hpp), and an integer token whose
// digit count shows that it fits is reported without calling
// convert_integer(). A number with up to 18 digits always fits into
// both std::uint64_t and std::int64_t (18 nines is about 1e18, below
// INT64_MAX, which is about 9.2e18). Longer tokens take the exact path
// below, including the fallback to floating point when the value does
// not fit.
//
// With a narrower number_unsigned_t/number_integer_t (e.g.
// std::uint32_t), the exact path would reclassify some of these tokens
// as (finite) floats, while this check reports integers. That does not
// change the result of accept(): it always parses through
// json_sax_acceptor, whose number callbacks discard their argument and
// return true, and the parser rejects neither integers nor finite
// floats. value_unsigned/value_integer are left unset here, so a caller
// that reads the converted value must not set discard_number_values.
//
// On contiguous input, scan_number_bulk_contiguous() converts integer
// tokens itself and does not pass them to this function, unless
// JSON_DIAGNOSTIC_POSITIONS is enabled. This check is therefore only
// reached for input without bulk access (e.g. streams), with
// JSON_DIAGNOSTIC_POSITIONS, or when scan_number_bulk_contiguous()
// falls back to scan_number().
if (discard_number_values)
{
constexpr std::size_t safe_digit_count = 18;
if (number_type == token_type::value_unsigned && token_buffer.size() <= safe_digit_count)
{
return token_type::value_unsigned;
}
if (number_type == token_type::value_integer && token_buffer.size() - 1 <= safe_digit_count)
{
return token_type::value_integer;
}
}
const char* const num_begin = token_buffer.data();
const char* const num_end = num_begin + token_buffer.size();
if (number_type != token_type::value_float)
{
const token_type integer_result = convert_integer(number_type, num_begin, num_end);
if (integer_result != token_type::uninitialized)
{
return integer_result;
}
}
// this code is reached if we parse a floating-point number or if an
// integer conversion above overflowed. Prefer std::from_chars
// (Eisel-Lemire, locale-independent, correctly rounded) when available;
// otherwise the exact Clinger fast path (double only); otherwise the
// locale-aware strtof/strtod/strtold.
if (convert_float_fast(num_begin, num_end, decimal_point_position, mantissa_end, value_float))
{
return token_type::value_float;
}
convert_float_locale_aware(token_buffer, decimal_point_position, value_float);
return token_type::value_float;
}
/*!
@brief contiguous fast path for scanning a number
Parses the whole number token straight from the input buffer, avoiding the
per-character get()/add() of scan_number(). On success it fills token_buffer
(as scan_number() does) and
returns the token type. On anything it does not fully recognize as a
well-formed number it makes no state change and returns
token_type::uninitialized, so the caller falls back to scan_number(), which
then produces the exact diagnostic. @a current is the first digit or the
leading minus (already read); the remaining bytes are taken from the adapter.
*/
token_type scan_number_bulk_contiguous()
{
// a pending unget offsets the buffer position from current; fall back
if (next_unget)
{
return token_type::uninitialized;
}
const std::size_t rem = ia.bulk_remaining();
if (rem == 0)
{
// the first digit is the last input byte; let scan_number() finish
return token_type::uninitialized;
}
// the byte before the next unread one is current (contiguous input)
const char* const data = reinterpret_cast<const char*>(ia.bulk_data()) - 1;
const std::size_t avail = rem + 1;
// validate + classify the number extent (mirrors scan_number()'s grammar)
std::size_t i = 0;
std::size_t dot_index = std::string::npos;
token_type number_type = token_type::value_unsigned;
if (data[0] == '-')
{
number_type = token_type::value_integer;
i = 1;
if (i >= avail)
{
return token_type::uninitialized;
}
}
if (data[i] == '0')
{
++i;
}
else if (data[i] >= '1' && data[i] <= '9')
{
++i;
while (i < avail && data[i] >= '0' && data[i] <= '9')
{
++i;
}
}
else
{
return token_type::uninitialized;
}
if (i < avail && data[i] == '.')
{
number_type = token_type::value_float;
dot_index = i;
++i;
if (i >= avail || !(data[i] >= '0' && data[i] <= '9'))
{
return token_type::uninitialized;
}
while (i < avail && data[i] >= '0' && data[i] <= '9')
{
++i;
}
}
// the mantissa ends here, whether or not an exponent part follows
const std::size_t mantissa_end = i;
if (i < avail && (data[i] == 'e' || data[i] == 'E'))
{
number_type = token_type::value_float;
++i;
if (i < avail && (data[i] == '+' || data[i] == '-'))
{
++i;
}
if (i >= avail || !(data[i] >= '0' && data[i] <= '9'))
{
return token_type::uninitialized;
}
while (i < avail && data[i] >= '0' && data[i] <= '9')
{
++i;
}
}
const std::size_t len = i;
// reset() records where this token starts (for diagnostics), so it has
// to run before the input position advances below
reset();
// An integer token needs no token_buffer: the SAX callbacks for
// number_integer/number_unsigned take only the value, and the overflow
// diagnostic rebuilds the text from the input. Convert straight from the
// input buffer and leave token_buffer empty. (JSON_DIAGNOSTIC_POSITIONS
// derives a number's start position from get_string().size(), so there
// the token still has to be materialized.)
#if !JSON_DIAGNOSTIC_POSITIONS
if (number_type != token_type::value_float)
{
const token_type integer_result = convert_integer(number_type, data, data + len);
if (JSON_HEDLEY_LIKELY(integer_result != token_type::uninitialized))
{
ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1);
position.chars_read_current_line += (len - 1);
return integer_result;
}
// The value does not fit an integer, so this token converts as a
// float. Recording that here keeps convert_number() below from
// repeating the integer attempt that just failed.
number_type = token_type::value_float;
}
#endif
// materialize the token exactly as scan_number() would. reset() already
// cleared token_buffer, so append() fills it (assign() is avoided
// because custom string_t types need not provide it)
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len);
decimal_point_position = dot_index;
ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1);
position.chars_read_current_line += (len - 1);
return convert_number(number_type, mantissa_end);
}
/// contiguous input: try the number fast path, else the byte-path scanner
token_type scan_number_dispatch(std::true_type /*bulk*/)
{
const token_type t = scan_number_bulk_contiguous();
return (t != token_type::uninitialized) ? t : scan_number();
}
/// streaming input: always use the byte-path scanner
token_type scan_number_dispatch(std::false_type /*bulk*/)
{
return scan_number();
}
/*!
@param[in] literal_text the literal text to expect
@param[in] length the length of the passed literal text
@param[in] return_type the token type to return on success
*/
JSON_HEDLEY_NON_NULL(2)
token_type scan_literal(const char_type* literal_text, const std::size_t length,
token_type return_type)
{
JSON_ASSERT(char_traits<char_type>::to_char_type(current) == literal_text[0]);
for (std::size_t i = 1; i < length; ++i)
{
if (JSON_HEDLEY_UNLIKELY(char_traits<char_type>::to_char_type(get()) != literal_text[i]))
{
error_message = "invalid literal";
return token_type::parse_error;
}
}
return return_type;
}
/////////////////////
// input management
/////////////////////
/// reset token_buffer; current character is beginning of token
void reset() noexcept
{
token_buffer.clear();
decimal_point_position = std::string::npos;
#if JSON_DIAGNOSTIC_POSITIONS
// the first character of the token has already been read, hence the -1
token_start_position = position.chars_read_total - 1;
#endif
note_token_start(std::integral_constant<bool, lazy_token_string> {});
}
/// seekable adapter: remember where the current token starts so it can be
/// reconstructed from the input on error; current has already been
/// consumed, hence the -1
void note_token_start(std::true_type /*lazy*/) noexcept
{
token_string_start = ia.get_consumed_count() - 1;
}
/// streaming adapter: start copying the token eagerly, beginning with the
/// already-read first character
void note_token_start(std::false_type /*lazy*/) noexcept
{
token_string.clear();
token_string.push_back(char_traits<char_type>::to_char_type(current));
}
/*
@brief get next character from the input
This function provides the interface to the used input adapter. It does
not throw in case the input reached EOF, but returns a
`char_traits<char>::eof()` in that case. Stores the scanned characters
for use in error messages.
@return character read from the input
*/
char_int_type get()
{
advance_position();
if (next_unget)
{
// only reset the next_unget variable and work with current
next_unget = false;
}
else
{
current = ia.get_character();
}
return track_after_read();
}
/// shared head of get() / get_ignoring_pending_unget(): bump the
/// per-character position counters (line-count-on-'\n' bookkeeping is
/// handled afterwards, in track_after_read(), once `current` is known)
void advance_position() noexcept
{
++position.chars_read_total;
++position.chars_read_current_line;
}
/// shared tail of get() / get_ignoring_pending_unget(): capture the
/// character for error messages (if needed) and update line/column
/// bookkeeping for the character now in `current`
char_int_type track_after_read()
{
// seekable adapters reconstruct the token lazily on error (see
// get_token_string), so the eager per-character copy is skipped
capture_char(std::integral_constant<bool, lazy_token_string> {});
if (current == '\n')
{
++position.lines_read;
// remember the column the newline was read at: chars_read_current_line
// is about to be cleared, and a matching unget() cannot reconstruct it
chars_read_before_newline = position.chars_read_current_line;
position.chars_read_current_line = 0;
}
return current;
}
/*!
@brief like get(), but for call sites that can prove no unget() is pending
get() has to check the `next_unget` flag on every call, because a
previous token may have ended with unget() (e.g. scan_number() always
ungets the character that terminated the number, so the next call to
scan() can see it again). skip_whitespace() reads that first,
possibly-ungotten character via a plain get(), but every further
character it reads is guaranteed to be a fresh read: nothing between
those calls invokes unget(). This variant skips the (otherwise always
false) next_unget branch for those calls; it is not a general
replacement for get().
*/
char_int_type get_ignoring_pending_unget()
{
JSON_ASSERT(!next_unget);
advance_position();
current = ia.get_character();
return track_after_read();
}
/// seekable adapter: nothing to capture, the token is rebuilt on error
void capture_char(std::true_type /*lazy*/) const noexcept {}
/// streaming adapter: copy the scanned character into token_string
void capture_char(std::false_type /*lazy*/)
{
if (JSON_HEDLEY_LIKELY(current != char_traits<char_type>::eof()))
{
token_string.push_back(char_traits<char_type>::to_char_type(current));
}
}
/*!
@brief unget current character (read it again on next get)
We implement unget by setting variable next_unget to true. The input is not
changed - we just simulate ungetting by modifying chars_read_total,
chars_read_current_line, and token_string. The next call to get() will
behave as if the unget character is read again.
*/
void unget()
{
next_unget = true;
--position.chars_read_total;
// in case we "unget" a newline, we have to also decrement the lines_read
// and restore the column that get() cleared when it saw the newline;
// chars_read_current_line == 0 can only mean the last get() read one
if (position.chars_read_current_line == 0)
{
if (position.lines_read > 0)
{
--position.lines_read;
}
// chars_read_before_newline counts the newline itself, which is the
// character being ungotten, hence the -1
position.chars_read_current_line = (chars_read_before_newline > 0)
? chars_read_before_newline - 1
: 0;
}
else
{
--position.chars_read_current_line;
}
uncapture_char(std::integral_constant<bool, lazy_token_string> {});
}
/// adapter without lookahead: nothing to do (see release_lookahead)
void release_lookahead_impl(std::false_type /*can_release*/) const noexcept {}
/// adapter with lookahead: leave the character in the input instead
void release_lookahead_impl(std::true_type /*can_release*/)
{
if (next_unget)
{
// the character is read from the input again rather than replayed
// from current, so the adapter must not step over it
next_unget = false;
ia.release_lookahead();
}
}
/// seekable adapter: nothing was captured, so nothing to undo
void uncapture_char(std::true_type /*lazy*/) const noexcept {}
/// streaming adapter: drop the character copied by the matching get()
void uncapture_char(std::false_type /*lazy*/)
{
if (JSON_HEDLEY_LIKELY(current != char_traits<char_type>::eof()))
{
JSON_ASSERT(!token_string.empty());
token_string.pop_back();
}
}
/// add a character to token_buffer
void add(char_int_type c)
{
token_buffer.push_back(static_cast<typename string_t::value_type>(c));
}
public:
/////////////////////
// value getters
/////////////////////
/// return integer value
constexpr number_integer_t get_number_integer() const noexcept
{
return value_integer;
}
/// return unsigned integer value
constexpr number_unsigned_t get_number_unsigned() const noexcept
{
return value_unsigned;
}
/// return floating-point value
constexpr number_float_t get_number_float() const noexcept
{
return value_float;
}
/// return current string value
string_t& get_string()
{
// a number token holds '.' regardless of the locale (#4084)
return token_buffer;
}
/////////////////////
// diagnostics
/////////////////////
/// return position of last read token
constexpr position_t get_position() const noexcept
{
return position;
}
/*!
@brief pass a pending simulated unget on to the input
unget() only rewinds the lexer's own bookkeeping, so the character that
terminated the last token (e.g. the character after a number) would still
be stepped over when the input adapter is done. Callers that hand the
input back to the user afterwards - operator>> and non-strict sax_parse -
call this once when scanning is done, so that the input is positioned
right after the value.
Adapters without lookahead (see input_adapter_supports_lookahead) are not
handed back to the user, so this is a no-op for them. Without
JSON_PRECISE_STREAM_POSITION, no adapter has lookahead, so this is always a
no-op and the terminating character stays consumed.
Scanning may continue after this call: @a next_unget is cleared, and the
character is read from the input again instead of being replayed from
@a current. A pending unget of EOF needs no special case, because reaching
EOF leaves no lookahead to release.
*/
void release_lookahead()
{
release_lookahead_impl(std::integral_constant<bool, can_release_lookahead> {});
}
#if JSON_DIAGNOSTIC_POSITIONS
/// return the offset of the first character of the last read token; unlike
/// the token's parsed value, this accounts for escape sequences
constexpr std::size_t get_token_start_position() const noexcept
{
return token_start_position;
}
#endif
/// seekable adapter: rebuild the last read token from the input on demand
const std::vector<char_type>& collect_token_chars(std::vector<char_type>& out, std::true_type /*lazy*/) const
{
// a pending unget of a real (non-EOF) character means that character
// was consumed from the input but is not part of the token; EOF is
// never consumed, so it must not be subtracted (mirrors unget())
const bool pending_real_unget = next_unget && current != char_traits<char_type>::eof();
const std::size_t stop = ia.get_consumed_count() - (pending_real_unget ? 1u : 0u);
if (JSON_HEDLEY_LIKELY(stop >= token_string_start))
{
ia.copy_consumed_range(token_string_start, stop, out);
}
return out;
}
/// streaming adapter: the token was copied eagerly while scanning
const std::vector<char_type>& collect_token_chars(std::vector<char_type>& /*out*/, std::false_type /*lazy*/) const
{
return token_string;
}
/// return the last read token (for errors only). Will never contain EOF
/// (an arbitrary value that is not a valid char value, often -1), because
/// 255 may legitimately occur. May contain NUL, which should be escaped.
std::string get_token_string() const
{
std::vector<char_type> reconstructed;
const std::vector<char_type>& chars = collect_token_chars(reconstructed, std::integral_constant<bool, lazy_token_string> {});
// escape control characters
std::string result;
for (const auto c : chars)
{
if (static_cast<unsigned char>(c) <= '\x1F')
{
// escape control characters
std::array<char, 9> cs{{}};
static_cast<void>((std::snprintf)(cs.data(), cs.size(), "<U+%.4X>", static_cast<unsigned char>(c))); // NOLINT(cppcoreguidelines-pro-type-vararg,hicpp-vararg)
result += cs.data();
}
else
{
// add character as is
result.push_back(static_cast<std::string::value_type>(c));
}
}
return result;
}
/// return syntax error message
JSON_HEDLEY_RETURNS_NON_NULL
constexpr const char* get_error_message() const noexcept
{
return error_message;
}
/////////////////////
// actual scanner
/////////////////////
/*!
@brief skip the UTF-8 byte order mark
@return true iff there is no BOM or the correct BOM has been skipped
*/
bool skip_bom()
{
if (get() == 0xEF)
{
// check if we completely parse the BOM
return get() == 0xBB && get() == 0xBF;
}
// the first character is not the beginning of the BOM; unget it to
// process is later
unget();
return true;
}
/// whether `current` is one of the four JSON whitespace characters
bool current_is_whitespace() const noexcept
{
return current == ' ' || current == '\t' || current == '\n' || current == '\r';
}
void skip_whitespace()
{
// the first character may be a pending unget() left over from the
// previous token (see get_ignoring_pending_unget()); every
// subsequent character read by this loop is guaranteed fresh, since
// nothing below calls unget()
get();
if (!current_is_whitespace())
{
return;
}
// this is written as an if-guarded do-while (rather than a plain
// while loop) because that shape is what lets both GCC and Clang
// keep the input adapter's read pointer in a register across
// iterations; the equivalent while-loop measurably defeated that
// optimization in testing, turning long whitespace runs (e.g. the
// indentation of pretty-printed JSON) from a register-only loop
// into one that reloads the pointer from memory every character
do
{
get_ignoring_pending_unget();
}
while (current_is_whitespace());
}
token_type scan()
{
// initially, skip the BOM
if (position.chars_read_total == 0 && !skip_bom())
{
error_message = "invalid BOM; must be 0xEF 0xBB 0xBF if given";
return token_type::parse_error;
}
// read the next character and ignore whitespace
skip_whitespace();
// ignore comments
while (ignore_comments && current == '/')
{
if (!scan_comment())
{
return token_type::parse_error;
}
// skip following whitespace
skip_whitespace();
}
switch (current)
{
// structural characters
case '[':
return token_type::begin_array;
case ']':
return token_type::end_array;
case '{':
return token_type::begin_object;
case '}':
return token_type::end_object;
case ':':
return token_type::name_separator;
case ',':
return token_type::value_separator;
// literals
case 't':
{
std::array<char_type, 4> true_literal = {{static_cast<char_type>('t'), static_cast<char_type>('r'), static_cast<char_type>('u'), static_cast<char_type>('e')}};
return scan_literal(true_literal.data(), true_literal.size(), token_type::literal_true);
}
case 'f':
{
std::array<char_type, 5> false_literal = {{static_cast<char_type>('f'), static_cast<char_type>('a'), static_cast<char_type>('l'), static_cast<char_type>('s'), static_cast<char_type>('e')}};
return scan_literal(false_literal.data(), false_literal.size(), token_type::literal_false);
}
case 'n':
{
std::array<char_type, 4> null_literal = {{static_cast<char_type>('n'), static_cast<char_type>('u'), static_cast<char_type>('l'), static_cast<char_type>('l')}};
return scan_literal(null_literal.data(), null_literal.size(), token_type::literal_null);
}
// string
case '\"':
return scan_string();
// number
case '-':
case '0':
case '1':
case '2':
case '3':
case '4':
case '5':
case '6':
case '7':
case '8':
case '9':
return scan_number_dispatch(std::integral_constant<bool, bulk_scan> {});
#if !JSON_STRICT_NUL_HANDLING
case '\0':
#endif
// end of input; by default, a null byte is also treated as end of
// input for backwards compatibility (see JSON_STRICT_NUL_HANDLING
// to opt into rejecting a null byte in the input instead)
case char_traits<char_type>::eof():
return token_type::end_of_input;
// error
default:
error_message = "invalid literal";
return token_type::parse_error;
}
}
private:
/// input adapter
InputAdapterType ia;
/// whether comments should be ignored (true) or signaled as errors (false)
const bool ignore_comments = false;
/// the current character
char_int_type current = char_traits<char_type>::eof();
/// whether the next get() call should just return current
bool next_unget = false;
/// the start position of the current token
position_t position {};
/// the value chars_read_current_line had when the last newline was read, so
/// that unget() can restore the column instead of leaving it at 0
std::size_t chars_read_before_newline = 0;
/// raw input token string for error messages; only populated for streaming
/// adapters (seekable adapters reconstruct it lazily via token_string_start)
std::vector<char_type> token_string {};
/// start offset of the current token within the input, used to reconstruct
/// the last read token on error for seekable adapters (see collect_token_chars)
std::size_t token_string_start = 0;
#if JSON_DIAGNOSTIC_POSITIONS
/// start offset of the current token within the input, used to report
/// diagnostic positions (see reset())
std::size_t token_start_position = 0;
#endif
/// buffer for variable-length tokens (numbers, strings)
string_t token_buffer {};
/// a description of occurred lexer errors
const char* error_message = "";
// number values
number_integer_t value_integer = 0;
number_unsigned_t value_unsigned = 0;
number_float_t value_float = 0;
/// the position of the decimal point in token_buffer
std::size_t decimal_point_position = std::string::npos;
/// whether the caller only needs the token types and never looks at the
/// converted numeric values; set only by accept(), which parses through
/// json_sax_acceptor. When set, convert_number() skips converting integer
/// tokens whose digit count guarantees that they fit into 64 bits (see
/// there)
const bool discard_number_values = false;
};
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END