mirror of
https://github.com/nlohmann/json.git
synced 2026-09-24 06:10:22 +01:00
Document that a NUL byte in the input is treated as end of input (#5534)
* docs: document that a NUL byte in the input is treated as end of input A NUL byte anywhere in the input - trailing, or embedded ahead of more otherwise well-formed JSON - is currently treated the same as genuine end of input, so parsing silently stops there instead of raising the parse_error.101 any other unexpected byte triggers. This mirrors the NUL-terminated-C-string convention already used when no explicit input length is given (json::parse(const char*) already stops at strlen()), just applied uniformly rather than only when a length is genuinely unavailable. This behavior predates this change and is not being altered here - changing it would be an observable, backwards-incompatible behavior change for any caller that (knowingly or not) depends on it, which is not something to do silently in a patch. Documenting the current, verified behavior as a new FAQ entry instead, so it's an intentional and discoverable part of the contract rather than a surprise. Fixes #5530. Signed-off-by: Niels Lohmann <niels.lohmann@gmail.com> Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4RQ1Ahan5YAGbnAQGjZTY * Add JSON_STRICT_NUL_HANDLING opt-in macro for issue #5530 A NUL byte anywhere in the input is currently treated the same as real end of input, rather than raising parse_error.101 like any other unexpected byte (documented in the previous commit's FAQ entry). A full unconditional fix was tried in PR #5532 but rejected as too risky to ship by default: any caller could depend on the current behavior, even unknowingly (e.g. a zero-padded buffer). On PR #5534, gregmarr proposed a compile-time opt-in flag instead, and the maintainer agreed, wanting it available now and defaulting to the corrected behavior in 4.0.0. This mirrors the existing JSON_BRACE_INIT_COPY_SEMANTICS precedent as closely as sensible: - JSON_STRICT_NUL_HANDLING defaults to 0 (off); the three lexer sites that treat '\0' as EOF/comment-terminator are gated with `#if !JSON_STRICT_NUL_HANDLING` so the default-off behavior is byte-for-byte identical to today's. - input_adapters.hpp's `T (&array)[N]` overload additionally trims a single trailing '\0' from a `char` array (e.g. a string literal like `json::parse("123")`) when the macro is on, so that case keeps working; every other element type (unsigned char, std::uint8_t, ...) always keeps its full extent. This intentionally does *not* reuse the existing strlen()-based pointer overload via SFINAE-excluding `char` from the array overload, as originally sketched for this change: that approach is ambiguous against the newer generic container overload added since PR #5532, and even where it compiles, strlen()-scanning a `char` array that is not NUL-terminated within its bounds reads past the end of the array (confirmed with AddressSanitizer). Trimming only a single trailing byte, without scanning, avoids both problems. - Documented via docs/mkdocs/docs/api/macros/json_strict_nul_handling.md, linked from the macros index/nav/features page, the FAQ entry, and the parse/accept/operator>> reference pages. - Tested in unit-class_parser.cpp and unit-deserialization.cpp, default state unguarded and opt-in state guarded. Since the library itself #undefs the macro at the end of json.hpp (as JSON_BRACE_INIT_COPY_SEMANTICS already does), a plain `#if defined(JSON_STRICT_NUL_HANDLING)` guard after the include never actually triggers; the tests instead capture the command-line value into a test-local macro before including the header. A few pre-existing fixtures elsewhere (std::array<uint8_t, N> sized one larger than their literal, relying on value-initialization to silently add a trailing zero byte) needed the same one-byte adjustment to keep passing under the opt-in behavior. Unlike the precedent, this adds a proper `JSON_StrictNulHandling` CMake option (rather than a raw -DCMAKE_CXX_FLAGS injection) and wires its ci_test_strict_nul_handling target into the ci_cmake_options job matrix in .github/workflows/ubuntu.yml, so the opt-in build is actually exercised in CI -- closing the one gap in the precedent's own CI setup (ci_test_brace_init_copy_semantics is defined but never referenced by any workflow, so it has never actually run). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Clarify where JSON_STRICT_NUL_HANDLING does not reject NUL bytes Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <niels.lohmann@gmail.com> Signed-off-by: Niels Lohmann <mail@nlohmann.me> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
f751547a81
commit
ed513715a8
@@ -762,6 +762,21 @@ contiguous_bytes_input_adapter input_adapter(CharT b)
|
||||
template<typename T, std::size_t N>
|
||||
auto input_adapter(T (&array)[N]) -> decltype(input_adapter(array, array + N)) // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
|
||||
{
|
||||
#if JSON_STRICT_NUL_HANDLING
|
||||
// A `char` array from string-literal initialization (e.g. json::parse("123"))
|
||||
// carries a trailing '\0' contributed by the compiler, not by the source
|
||||
// text; drop exactly that one byte so it is not mistaken for real trailing
|
||||
// data. Every other element type (unsigned char, std::uint8_t, ...) keeps
|
||||
// the full extent unconditionally, since a trailing zero byte there is
|
||||
// genuine data (e.g. CBOR/MessagePack). This intentionally does not
|
||||
// strlen()-scan the array (as the pointer overload above does for a
|
||||
// null-delimited string): for a `char` array that is not NUL-terminated
|
||||
// within its bounds, that would read past the end of the array.
|
||||
if (std::is_same<typename std::remove_cv<T>::type, char>::value && N > 0 && array[N - 1] == 0)
|
||||
{
|
||||
return input_adapter(array, array + N - 1);
|
||||
}
|
||||
#endif
|
||||
return input_adapter(array, array + N);
|
||||
}
|
||||
|
||||
|
||||
@@ -952,7 +952,9 @@ class lexer : public lexer_base<BasicJsonType>
|
||||
case '\n':
|
||||
case '\r':
|
||||
case char_traits<char_type>::eof():
|
||||
#if !JSON_STRICT_NUL_HANDLING
|
||||
case '\0':
|
||||
#endif
|
||||
return true;
|
||||
|
||||
default:
|
||||
@@ -970,8 +972,10 @@ class lexer : public lexer_base<BasicJsonType>
|
||||
{
|
||||
switch (get())
|
||||
{
|
||||
case char_traits<char_type>::eof():
|
||||
#if !JSON_STRICT_NUL_HANDLING
|
||||
case '\0':
|
||||
#endif
|
||||
case char_traits<char_type>::eof():
|
||||
{
|
||||
error_message = "invalid comment; missing closing '*/'";
|
||||
return false;
|
||||
@@ -2153,9 +2157,12 @@ scan_number_done:
|
||||
case '9':
|
||||
return scan_number_dispatch(std::integral_constant<bool, bulk_scan> {});
|
||||
|
||||
// end of input (the null byte is needed when parsing from
|
||||
// string literals)
|
||||
#if !JSON_STRICT_NUL_HANDLING
|
||||
case '\0':
|
||||
#endif
|
||||
// end of input; by default, a null byte is also treated as end of
|
||||
// input for backwards compatibility (see JSON_STRICT_NUL_HANDLING
|
||||
// to opt into rejecting a null byte in the input instead)
|
||||
case char_traits<char_type>::eof():
|
||||
return token_type::end_of_input;
|
||||
|
||||
|
||||
@@ -807,3 +807,7 @@ void templated_json_throw(ExceptionType exception)
|
||||
#ifndef JSON_BRACE_INIT_COPY_SEMANTICS
|
||||
#define JSON_BRACE_INIT_COPY_SEMANTICS 0
|
||||
#endif
|
||||
|
||||
#ifndef JSON_STRICT_NUL_HANDLING
|
||||
#define JSON_STRICT_NUL_HANDLING 0
|
||||
#endif
|
||||
|
||||
@@ -27,6 +27,7 @@
|
||||
#undef JSON_DISABLE_ENUM_SERIALIZATION
|
||||
#undef JSON_USE_GLOBAL_UDLS
|
||||
#undef JSON_BRACE_INIT_COPY_SEMANTICS
|
||||
#undef JSON_STRICT_NUL_HANDLING
|
||||
|
||||
#ifndef JSON_TEST_KEEP_MACROS
|
||||
#undef JSON_CATCH
|
||||
|
||||
Reference in New Issue
Block a user