mirror of
https://github.com/nlohmann/json.git
synced 2026-09-24 06:10:22 +01:00
Document that a NUL byte in the input is treated as end of input (#5534)
* docs: document that a NUL byte in the input is treated as end of input A NUL byte anywhere in the input - trailing, or embedded ahead of more otherwise well-formed JSON - is currently treated the same as genuine end of input, so parsing silently stops there instead of raising the parse_error.101 any other unexpected byte triggers. This mirrors the NUL-terminated-C-string convention already used when no explicit input length is given (json::parse(const char*) already stops at strlen()), just applied uniformly rather than only when a length is genuinely unavailable. This behavior predates this change and is not being altered here - changing it would be an observable, backwards-incompatible behavior change for any caller that (knowingly or not) depends on it, which is not something to do silently in a patch. Documenting the current, verified behavior as a new FAQ entry instead, so it's an intentional and discoverable part of the contract rather than a surprise. Fixes #5530. Signed-off-by: Niels Lohmann <niels.lohmann@gmail.com> Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4RQ1Ahan5YAGbnAQGjZTY * Add JSON_STRICT_NUL_HANDLING opt-in macro for issue #5530 A NUL byte anywhere in the input is currently treated the same as real end of input, rather than raising parse_error.101 like any other unexpected byte (documented in the previous commit's FAQ entry). A full unconditional fix was tried in PR #5532 but rejected as too risky to ship by default: any caller could depend on the current behavior, even unknowingly (e.g. a zero-padded buffer). On PR #5534, gregmarr proposed a compile-time opt-in flag instead, and the maintainer agreed, wanting it available now and defaulting to the corrected behavior in 4.0.0. This mirrors the existing JSON_BRACE_INIT_COPY_SEMANTICS precedent as closely as sensible: - JSON_STRICT_NUL_HANDLING defaults to 0 (off); the three lexer sites that treat '\0' as EOF/comment-terminator are gated with `#if !JSON_STRICT_NUL_HANDLING` so the default-off behavior is byte-for-byte identical to today's. - input_adapters.hpp's `T (&array)[N]` overload additionally trims a single trailing '\0' from a `char` array (e.g. a string literal like `json::parse("123")`) when the macro is on, so that case keeps working; every other element type (unsigned char, std::uint8_t, ...) always keeps its full extent. This intentionally does *not* reuse the existing strlen()-based pointer overload via SFINAE-excluding `char` from the array overload, as originally sketched for this change: that approach is ambiguous against the newer generic container overload added since PR #5532, and even where it compiles, strlen()-scanning a `char` array that is not NUL-terminated within its bounds reads past the end of the array (confirmed with AddressSanitizer). Trimming only a single trailing byte, without scanning, avoids both problems. - Documented via docs/mkdocs/docs/api/macros/json_strict_nul_handling.md, linked from the macros index/nav/features page, the FAQ entry, and the parse/accept/operator>> reference pages. - Tested in unit-class_parser.cpp and unit-deserialization.cpp, default state unguarded and opt-in state guarded. Since the library itself #undefs the macro at the end of json.hpp (as JSON_BRACE_INIT_COPY_SEMANTICS already does), a plain `#if defined(JSON_STRICT_NUL_HANDLING)` guard after the include never actually triggers; the tests instead capture the command-line value into a test-local macro before including the header. A few pre-existing fixtures elsewhere (std::array<uint8_t, N> sized one larger than their literal, relying on value-initialization to silently add a trailing zero byte) needed the same one-byte adjustment to keep passing under the opt-in behavior. Unlike the precedent, this adds a proper `JSON_StrictNulHandling` CMake option (rather than a raw -DCMAKE_CXX_FLAGS injection) and wires its ci_test_strict_nul_handling target into the ci_cmake_options job matrix in .github/workflows/ubuntu.yml, so the opt-in build is actually exercised in CI -- closing the one gap in the precedent's own CI setup (ci_test_brace_init_copy_semantics is defined but never referenced by any workflow, so it has never actually run). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Clarify where JSON_STRICT_NUL_HANDLING does not reject NUL bytes Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <niels.lohmann@gmail.com> Signed-off-by: Niels Lohmann <mail@nlohmann.me> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
f751547a81
commit
ed513715a8
@@ -8,6 +8,14 @@
|
||||
|
||||
#include "doctest_compatibility.h"
|
||||
|
||||
// capture whether JSON_STRICT_NUL_HANDLING was enabled on the command line
|
||||
// (e.g. -DJSON_STRICT_NUL_HANDLING=1) *before* including json.hpp, since the
|
||||
// library #undefs JSON_STRICT_NUL_HANDLING itself once the header has been
|
||||
// fully processed (see include/nlohmann/detail/macro_unscope.hpp)
|
||||
#if defined(JSON_STRICT_NUL_HANDLING) && (JSON_STRICT_NUL_HANDLING == 1)
|
||||
#define JSON_TEST_STRICT_NUL_HANDLING_ENABLED 1
|
||||
#endif
|
||||
|
||||
#define JSON_TESTS_PRIVATE
|
||||
#include <nlohmann/json.hpp>
|
||||
using nlohmann::json;
|
||||
@@ -545,6 +553,88 @@ TEST_CASE("parser class")
|
||||
}
|
||||
}
|
||||
|
||||
SECTION("NUL byte handling (issue #5530, JSON_STRICT_NUL_HANDLING)")
|
||||
{
|
||||
// by default, a NUL byte anywhere in the input (not inside a quoted
|
||||
// string, which is covered above) is silently treated the same as
|
||||
// real end of input; JSON_STRICT_NUL_HANDLING (off by default, see
|
||||
// docs/mkdocs/docs/api/macros/json_strict_nul_handling.md) makes a
|
||||
// NUL byte an error like any other unexpected byte instead.
|
||||
//
|
||||
// The two sections below are mutually exclusive: this whole test
|
||||
// binary is compiled once, with JSON_STRICT_NUL_HANDLING either
|
||||
// left at its default or forced to 1 (e.g. by the dedicated
|
||||
// ci_test_strict_nul_handling CI target), so only the section
|
||||
// matching the actual, compiled-in behavior can pass.
|
||||
#if !defined(JSON_TEST_STRICT_NUL_HANDLING_ENABLED)
|
||||
SECTION("default behavior (macro not enabled)")
|
||||
{
|
||||
// a NUL byte after a complete value silently truncates the input
|
||||
std::string s = "123";
|
||||
s.push_back('\0');
|
||||
s += "4";
|
||||
CHECK(json::parse(s) == json(123));
|
||||
CHECK(json::accept(s));
|
||||
|
||||
// parsing from a string literal is unaffected either way
|
||||
CHECK(json::parse("123") == json(123));
|
||||
}
|
||||
#endif
|
||||
|
||||
#if defined(JSON_TEST_STRICT_NUL_HANDLING_ENABLED)
|
||||
SECTION("opt-in strict behavior (JSON_STRICT_NUL_HANDLING == 1)")
|
||||
{
|
||||
// a NUL byte after a complete value is now a parse error,
|
||||
// instead of silently truncating the input
|
||||
{
|
||||
std::string s = "123";
|
||||
s.push_back('\0');
|
||||
json _; // NOLINT(readability-identifier-naming)
|
||||
CHECK_THROWS_WITH_AS(_ = json::parse(s),
|
||||
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid literal; last read: '123<U+0000>'; expected end of input",
|
||||
json::parse_error&);
|
||||
CHECK_FALSE(json::accept(s));
|
||||
}
|
||||
|
||||
// a NUL byte where a value is expected is now a parse error,
|
||||
// instead of being treated the same as an empty input
|
||||
{
|
||||
const std::string s(1, '\0');
|
||||
json _; // NOLINT(readability-identifier-naming)
|
||||
CHECK_THROWS_WITH_AS(_ = json::parse(s),
|
||||
"[json.exception.parse_error.101] parse error at line 1, column 1: syntax error while parsing value - invalid literal; last read: '<U+0000>'",
|
||||
json::parse_error&);
|
||||
CHECK_FALSE(json::accept(s));
|
||||
}
|
||||
|
||||
// a NUL byte inside a // comment no longer stops the comment
|
||||
// scan early; scanning continues correctly past it
|
||||
{
|
||||
std::string s = "1 // a";
|
||||
s.push_back('\0');
|
||||
s += "b\n";
|
||||
CHECK(json::parse(s, nullptr, true, true) == json(1));
|
||||
CHECK(json::accept(s, true, true));
|
||||
}
|
||||
|
||||
// a NUL byte inside a /* */ comment no longer stops the
|
||||
// comment scan early either
|
||||
{
|
||||
std::string s = "1 /* a";
|
||||
s.push_back('\0');
|
||||
s += "b */ ";
|
||||
CHECK(json::parse(s, nullptr, true, true) == json(1));
|
||||
CHECK(json::accept(s, true, true));
|
||||
}
|
||||
|
||||
// regression guard: parsing from a string literal (which
|
||||
// carries a compiler-appended trailing '\0') still works,
|
||||
// even though a NUL byte is now rejected everywhere else
|
||||
CHECK(json::parse("123") == json(123));
|
||||
}
|
||||
#endif
|
||||
}
|
||||
|
||||
SECTION("number")
|
||||
{
|
||||
SECTION("integers")
|
||||
@@ -1894,7 +1984,13 @@ TEST_CASE("parser class")
|
||||
|
||||
SECTION("from std::array")
|
||||
{
|
||||
std::array<uint8_t, 5> v { {'t', 'r', 'u', 'e'} };
|
||||
// NOTE: this array is sized to exactly the length of "true" (unlike
|
||||
// the trailing-NUL-tolerant default behavior elsewhere in this file,
|
||||
// see the "NUL byte handling" section above); a size of 5 here would
|
||||
// leave a value-initialized trailing 0x00 element that is only
|
||||
// silently accepted as end-of-input by default and would fail under
|
||||
// JSON_STRICT_NUL_HANDLING
|
||||
std::array<uint8_t, 4> v { {'t', 'r', 'u', 'e'} };
|
||||
json j;
|
||||
json::parser(nlohmann::detail::input_adapter(std::begin(v), std::end(v))).parse(true, j);
|
||||
CHECK(j == json(true));
|
||||
@@ -2035,7 +2131,17 @@ TEST_CASE("parser class")
|
||||
{
|
||||
json _;
|
||||
CHECK_THROWS_WITH_AS(_ = json::parse("/a", nullptr, true, true), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid comment; expecting '/' or '*' after '/'; last read: '/a'", json::parse_error);
|
||||
// "/*" is a string literal, so it carries a compiler-appended trailing
|
||||
// '\0'; by default that NUL is read like any other byte and shows up
|
||||
// in "last read", but JSON_STRICT_NUL_HANDLING trims exactly that one
|
||||
// trailing byte from a char array (see
|
||||
// docs/mkdocs/docs/api/macros/json_strict_nul_handling.md), so it no
|
||||
// longer appears in the message in that state
|
||||
#if defined(JSON_TEST_STRICT_NUL_HANDLING_ENABLED)
|
||||
CHECK_THROWS_WITH_AS(_ = json::parse("/*", nullptr, true, true), "[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid comment; missing closing '*/'; last read: '/*'", json::parse_error);
|
||||
#else
|
||||
CHECK_THROWS_WITH_AS(_ = json::parse("/*", nullptr, true, true), "[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid comment; missing closing '*/'; last read: '/*<U+0000>'", json::parse_error);
|
||||
#endif
|
||||
}
|
||||
|
||||
#if JSON_DIAGNOSTIC_POSITIONS
|
||||
|
||||
@@ -8,6 +8,14 @@
|
||||
|
||||
#include "doctest_compatibility.h"
|
||||
|
||||
// capture whether JSON_STRICT_NUL_HANDLING was enabled on the command line
|
||||
// (e.g. -DJSON_STRICT_NUL_HANDLING=1) *before* including json.hpp, since the
|
||||
// library #undefs JSON_STRICT_NUL_HANDLING itself once the header has been
|
||||
// fully processed (see include/nlohmann/detail/macro_unscope.hpp)
|
||||
#if defined(JSON_STRICT_NUL_HANDLING) && (JSON_STRICT_NUL_HANDLING == 1)
|
||||
#define JSON_TEST_STRICT_NUL_HANDLING_ENABLED 1
|
||||
#endif
|
||||
|
||||
#include <nlohmann/json.hpp>
|
||||
using nlohmann::json;
|
||||
#ifdef JSON_TEST_NO_GLOBAL_UDLS
|
||||
@@ -323,6 +331,23 @@ TEST_CASE("deserialization")
|
||||
CHECK(j == json({"foo", 1, 2, 3, false, {{"one", 1}}}));
|
||||
}
|
||||
|
||||
SECTION("operator>> with a NUL byte after the value (issue #5530)")
|
||||
{
|
||||
// operator>> parses non-strictly (it does not require the whole
|
||||
// stream to be consumed), so a NUL byte following a complete
|
||||
// value is simply left unread on the stream and never reaches
|
||||
// the "expected end of input" check that JSON_STRICT_NUL_HANDLING
|
||||
// affects; this holds regardless of the macro (verified below for
|
||||
// the opt-in state as well)
|
||||
std::string data = "123";
|
||||
data.push_back('\0');
|
||||
std::istringstream ss(data);
|
||||
json j;
|
||||
ss >> j;
|
||||
CHECK(j == json(123));
|
||||
CHECK(ss.good());
|
||||
}
|
||||
|
||||
SECTION("user-defined string literal")
|
||||
{
|
||||
CHECK("[\"foo\",1,2,3,false,{\"one\":1}]"_json == json({"foo", 1, 2, 3, false, {{"one", 1}}}));
|
||||
@@ -405,6 +430,27 @@ TEST_CASE("deserialization")
|
||||
CHECK_THROWS_WITH_AS(ss >> j, "[json.exception.parse_error.101] parse error at line 1, column 29: syntax error while parsing array - unexpected end of input; expected ']'", json::parse_error&);
|
||||
}
|
||||
|
||||
#if defined(JSON_TEST_STRICT_NUL_HANDLING_ENABLED)
|
||||
SECTION("operator>> with a NUL byte where a value is expected (JSON_STRICT_NUL_HANDLING == 1, issue #5530)")
|
||||
{
|
||||
// a trailing NUL byte *after* a complete value is unaffected by the
|
||||
// macro (see the successful-deserialization "operator>> with a NUL
|
||||
// byte after the value" section above): operator>> parses
|
||||
// non-strictly and never reaches the "expected end of input" check
|
||||
// that the macro changes. A NUL byte where a *value* is expected,
|
||||
// however, goes through the same token dispatch as any other input
|
||||
// and is affected: with the macro enabled it now raises
|
||||
// parse_error.101 (like any other unrecognized byte) instead of
|
||||
// being silently treated the same as an empty stream.
|
||||
std::string const data(1, '\0');
|
||||
std::istringstream ss(data);
|
||||
json j;
|
||||
CHECK_THROWS_WITH_AS(ss >> j,
|
||||
"[json.exception.parse_error.101] parse error at line 1, column 1: syntax error while parsing value - invalid literal; last read: '<U+0000>'",
|
||||
json::parse_error&);
|
||||
}
|
||||
#endif
|
||||
|
||||
SECTION("user-defined string literal")
|
||||
{
|
||||
CHECK_THROWS_WITH_AS("[\"foo\",1,2,3,false,{\"one\":1}"_json, "[json.exception.parse_error.101] parse error at line 1, column 29: syntax error while parsing array - unexpected end of input; expected ']'", json::parse_error&);
|
||||
@@ -453,7 +499,11 @@ TEST_CASE("deserialization")
|
||||
|
||||
SECTION("from std::array")
|
||||
{
|
||||
std::array<uint8_t, 5> const v { {'t', 'r', 'u', 'e'} };
|
||||
// sized to exactly the length of "true": a size of 5 would leave
|
||||
// a value-initialized trailing 0x00 element that is only
|
||||
// silently accepted as end-of-input by default and would fail
|
||||
// under JSON_STRICT_NUL_HANDLING
|
||||
std::array<uint8_t, 4> const v { {'t', 'r', 'u', 'e'} };
|
||||
CHECK(json::parse(v) == json(true));
|
||||
CHECK(json::accept(v));
|
||||
|
||||
@@ -549,7 +599,9 @@ TEST_CASE("deserialization")
|
||||
|
||||
SECTION("from std::array")
|
||||
{
|
||||
std::array<uint8_t, 5> v { {'t', 'r', 'u', 'e'} };
|
||||
// sized to exactly the length of "true", see the analogous
|
||||
// "from std::array" section above for why
|
||||
std::array<uint8_t, 4> v { {'t', 'r', 'u', 'e'} };
|
||||
CHECK(json::parse(std::begin(v), std::end(v)) == json(true));
|
||||
CHECK(json::accept(std::begin(v), std::end(v)));
|
||||
|
||||
|
||||
@@ -606,7 +606,11 @@ TEST_CASE("regression tests 2")
|
||||
SECTION("issue #2546 - parsing containers of std::byte")
|
||||
{
|
||||
const char DATA[] = R"("Hello, world!")"; // NOLINT(misc-const-correctness,cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
|
||||
const auto s = std::as_bytes(std::span(DATA));
|
||||
// exclude the trailing '\0' that string-literal initialization adds to
|
||||
// DATA: std::span(DATA) would span the full array extent (including
|
||||
// that NUL), which is only silently accepted as end-of-input by default
|
||||
// and would fail under JSON_STRICT_NUL_HANDLING
|
||||
const auto s = std::as_bytes(std::span(DATA, sizeof(DATA) - 1));
|
||||
const json j = json::parse(s);
|
||||
CHECK(j.dump() == "\"Hello, world!\"");
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user