14 Commits
Author SHA1 Message Date
andrew c44b0262bc Merge pull request 'schemagen: reject absent additionalProperties' (#58) from weaselbot/weaseljson:weaselbot/issue-55 into main
Reviewed-on: weaselab/weaseljson#58
2026-07-20 18:06:53 +00:00
andrew 814b24efba Merge pull request 'Fix schemagen integer parser rejecting zero with negative exponents' (#57) from weaselbot/weaseljson:weaselbot/issue-56 into main
Reviewed-on: weaselab/weaseljson#57
2026-07-20 18:00:36 +00:00
weaselbot 86b58e83cb schemagen: reject absent additionalProperties
Per the README contract, an object schema without an explicit
`additionalProperties` was documented as rejected at generation time,
but the code defaulted it to `false` and silently generated a strict
parser that rejects valid documents with extra properties.

Make the code match the documented contract by raising GenError when
`additionalProperties` is absent, with a clear message instructing the
author to set it explicitly to false. Update the README prose section
to list absent additionalProperties alongside true, and fix existing
schemas and tests to declare additionalProperties explicitly.

Closes #55
2026-07-19 22:06:23 -04:00
weaselbot cdff634057 Fix schemagen integer parser rejecting zero with negative exponents
Move the zero-detection check in the generated parseJsonInt64 before the
finalExp < 0 guard. Previously, valid JSON numbers whose mathematical
value is 0 but written with a large negative exponent (e.g. 0e-2,
0.0e-2, -0e-2, 0e-20) were rejected because the negative-finalExp early
return ran before the all-zero-digits branch could set out = 0. Non-zero
values with negative exponents are still correctly rejected.

Closes #56
2026-07-19 22:00:15 -04:00
andrew 16f5d3cd9d Merge pull request 'Handle null parser in WeaselJsonParser_parse' (#54) from null-parser into main
Reviewed-on: weaselab/weaseljson#54
2026-07-15 22:03:59 +00:00
andrew 6beb538b61 Add missing status enum to weaseljson.py 2026-07-15 17:55:22 -04:00
andrew 6520039dc2 Address review feedback 2026-07-15 14:12:41 -04:00
andrew 93203b14f5 Add missing unlikely annotation 2026-07-15 12:31:17 -04:00
andrew 09c0fb72ca Handle null parser in WeaselJsonParser_parse
Closes #51
2026-07-15 12:31:09 -04:00
andrew 70bb33eeaf Merge pull request 'make WeaselJson_OVERFLOW a terminal state like WeaselJson_REJECT' (#53) from weaselbot/weaseljson:weaselbot/issue-52 into main
Reviewed-on: weaselab/weaseljson#53
2026-07-13 20:58:14 +00:00
andrew 42d37d1fd7 Don't list all terminal statuses for weaseljson
Also simplifies codegen slightly presumably. OK and AGAIN should be the only non-terminal statuses ever.
2026-07-13 16:49:16 -04:00
weaselbot 5e462f1477 make WeaselJson_OVERFLOW a terminal state like WeaselJson_REJECT
When a parse step returned WeaselJson_OVERFLOW, the pushdown stack was
left corrupted (frames had been popped before the failing push), but the
overflow status was not made sticky the way WeaselJson_REJECT is. A
subsequent end-of-data call WeaselJsonParser_parse(parser, nullptr, 0)
could then dispatch through the corrupted stack and return WeaselJson_OK,
accepting an incomplete, invalid too-deeply-nested document as valid JSON.

Add an `overflowed` flag, set it whenever a continuation returns
WeaselJson_OVERFLOW, and short-circuit parse() to return
WeaselJson_OVERFLOW on every subsequent call. reset() clears the flag so a
reused parser can accept a different document. This mirrors the existing
stickiness handling for WeaselJson_REJECT.

Closes #52
2026-07-13 16:07:36 -04:00
andrew f5ceb392c2 Merge pull request 'schemagen: add regression test for nullable cyclic $ref targets' (#50) from weaselbot/weaseljson:weaselbot/issue-14 into main
Reviewed-on: weaselab/weaseljson#50
2026-06-30 16:50:07 +00:00
weaselbot be92755cbf schemagen: add regression test for nullable cyclic $ref targets (issue #14)
The generator already preserves nullability for self-referential object
$defs thanks to prior fixes, but issue #14 had no regression coverage.

Add a test using the exact reproduction schema from the issue and verify
that both the outer and recursive `self` fields accept `null`.
2026-06-30 12:42:03 -04:00
10 changed files with 272 additions and 27 deletions
+3 -2
View File
@@ -60,8 +60,9 @@ into the result, so it is non-movable.
## Not supported (rejected at generation time, no fallback) ## Not supported (rejected at generation time, no fallback)
`oneOf` / `anyOf` / `allOf` / `not` / `if`-`then`-`else`, `oneOf` / `anyOf` / `allOf` / `not` / `if`-`then`-`else`,
`patternProperties`, `additionalProperties` with a schema (typed map), `patternProperties`, `additionalProperties` absent or set to `true`,
`additionalProperties: true`, `prefixItems` (tuples), `const`, `additionalProperties` with a schema (typed map), `prefixItems` (tuples),
`const`,
`dependentSchemas`/`dependentRequired`, union `type` lists other than `dependentSchemas`/`dependentRequired`, union `type` lists other than
`["T", "null"]`, non-string enums, and remote (`$ref` to other documents). `["T", "null"]`, non-string enums, and remote (`$ref` to other documents).
+1
View File
@@ -56,6 +56,7 @@
"type": "array", "type": "array",
"items": { "items": {
"type": "object", "type": "object",
"additionalProperties": false,
"required": [ "required": [
"name" "name"
], ],
+180 -12
View File
@@ -32,6 +32,7 @@ class SchemagenEnumTest(unittest.TestCase):
def test_colliding_enum_values_deduplicate(self): def test_colliding_enum_values_deduplicate(self):
schema = { schema = {
"type": "object", "type": "object",
"additionalProperties": False,
"properties": {"role": {"enum": ["foo-bar", "foo_bar"]}}, "properties": {"role": {"enum": ["foo-bar", "foo_bar"]}},
} }
rc, stdout, stderr = self.run_schemagen(schema) rc, stdout, stderr = self.run_schemagen(schema)
@@ -45,6 +46,7 @@ class SchemagenEnumTest(unittest.TestCase):
def test_distinct_enum_values_generate(self): def test_distinct_enum_values_generate(self):
schema = { schema = {
"type": "object", "type": "object",
"additionalProperties": False,
"properties": {"role": {"enum": ["admin", "user", "guest"]}}, "properties": {"role": {"enum": ["admin", "user", "guest"]}},
} }
rc, stdout, stderr = self.run_schemagen(schema) rc, stdout, stderr = self.run_schemagen(schema)
@@ -80,7 +82,7 @@ class SchemagenKeywordTest(unittest.TestCase):
"module", "module",
"import", "import",
] ]
schema = {"type": "object", "properties": {}} schema = {"type": "object", "additionalProperties": False, "properties": {}}
for kw in keywords: for kw in keywords:
schema["properties"][kw] = {"type": "string"} schema["properties"][kw] = {"type": "string"}
rc, stdout, stderr = self.run_schemagen(schema) rc, stdout, stderr = self.run_schemagen(schema)
@@ -164,11 +166,18 @@ class SchemagenCollisionTest(unittest.TestCase):
def test_kind_enum_does_not_duplicate_arr0(self): def test_kind_enum_does_not_duplicate_arr0(self):
schema = { schema = {
"type": "object", "type": "object",
"additionalProperties": False,
"properties": { "properties": {
"arr": {"type": "array", "items": {"type": "string"}}, "arr": {"type": "array", "items": {"type": "string"}},
"obj": {"$ref": "#/$defs/Arr0"}, "obj": {"$ref": "#/$defs/Arr0"},
}, },
"$defs": {"Arr0": {"type": "object", "properties": {}}}, "$defs": {
"Arr0": {
"type": "object",
"additionalProperties": False,
"properties": {},
}
},
} }
out = self.generate_and_compile(schema) out = self.generate_and_compile(schema)
self.assertIn("struct Arr0", out) self.assertIn("struct Arr0", out)
@@ -183,7 +192,13 @@ class SchemagenCollisionTest(unittest.TestCase):
schema = { schema = {
"type": "array", "type": "array",
"items": {"$ref": "#/$defs/Skip"}, "items": {"$ref": "#/$defs/Skip"},
"$defs": {"Skip": {"type": "object", "properties": {}}}, "$defs": {
"Skip": {
"type": "object",
"additionalProperties": False,
"properties": {},
}
},
} }
out = self.generate_and_compile(schema) out = self.generate_and_compile(schema)
self.assertIn("struct Skip", out) self.assertIn("struct Skip", out)
@@ -199,6 +214,7 @@ class SchemagenCollisionTest(unittest.TestCase):
"""Regression test for issue #17: enum and array $defs must be reused.""" """Regression test for issue #17: enum and array $defs must be reused."""
schema = { schema = {
"type": "object", "type": "object",
"additionalProperties": False,
"properties": { "properties": {
"role1": {"$ref": "#/$defs/Role"}, "role1": {"$ref": "#/$defs/Role"},
"role2": {"$ref": "#/$defs/Role"}, "role2": {"$ref": "#/$defs/Role"},
@@ -252,17 +268,14 @@ class SchemagenAdditionalPropertiesTest(unittest.TestCase):
self.assertNotEqual(rc, 0) self.assertNotEqual(rc, 0)
self.assertIn("additionalProperties: true is not supported", stderr) self.assertIn("additionalProperties: true is not supported", stderr)
def test_additional_properties_absent_defaults_to_strict(self): def test_additional_properties_absent_rejected(self):
schema = { schema = {
"type": "object", "type": "object",
"properties": {"name": {"type": "string"}}, "properties": {"name": {"type": "string"}},
} }
rc, stdout, stderr = self.run_schemagen(schema) rc, stdout, stderr = self.run_schemagen(schema)
self.assertEqual(rc, 0, msg=stderr) self.assertNotEqual(rc, 0)
# The generated parser should reject unknown keys. Verify the key-matching self.assertIn("additionalProperties is required for object schemas", stderr)
# helper returns -1 for an unknown key and cbKeyData rejects it.
self.assertIn("int matchKey(Kind k, std::string_view key) const {", stdout)
self.assertNotIn("bool isStrict(Kind k) const", stdout)
def test_additional_properties_false_accepted(self): def test_additional_properties_false_accepted(self):
schema = { schema = {
@@ -310,6 +323,7 @@ class SchemagenCyclicArrayTest(unittest.TestCase):
def test_object_field_to_self_referential_array_rejected(self): def test_object_field_to_self_referential_array_rejected(self):
schema = { schema = {
"type": "object", "type": "object",
"additionalProperties": False,
"properties": {"items": {"$ref": "#/$defs/Items"}}, "properties": {"items": {"$ref": "#/$defs/Items"}},
"$defs": { "$defs": {
"Items": { "Items": {
@@ -326,6 +340,7 @@ class SchemagenCyclicArrayTest(unittest.TestCase):
def test_chain_of_array_refs_rejected(self): def test_chain_of_array_refs_rejected(self):
schema = { schema = {
"type": "object", "type": "object",
"additionalProperties": False,
"properties": {"x": {"$ref": "#/$defs/A"}}, "properties": {"x": {"$ref": "#/$defs/A"}},
"$defs": { "$defs": {
"A": {"type": "array", "items": {"$ref": "#/$defs/B"}}, "A": {"type": "array", "items": {"$ref": "#/$defs/B"}},
@@ -339,6 +354,7 @@ class SchemagenCyclicArrayTest(unittest.TestCase):
def test_non_recursive_array_refs_still_allowed(self): def test_non_recursive_array_refs_still_allowed(self):
schema = { schema = {
"type": "object", "type": "object",
"additionalProperties": False,
"properties": {"roles": {"$ref": "#/$defs/Roles"}}, "properties": {"roles": {"$ref": "#/$defs/Roles"}},
"$defs": { "$defs": {
"Role": {"enum": ["admin", "user"]}, "Role": {"enum": ["admin", "user"]},
@@ -353,6 +369,134 @@ class SchemagenCyclicArrayTest(unittest.TestCase):
self.assertIn("std::optional<std::vector<Role>> roles;", stdout) self.assertIn("std::optional<std::vector<Role>> roles;", stdout)
class SchemagenCyclicNullableObjectTest(unittest.TestCase):
"""Regression tests for issue #14: nullable cyclic $ref targets."""
def setUp(self):
self.repo_root = os.path.dirname(
os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
)
self.include_dir = os.path.join(self.repo_root, "include")
self.lib_src = os.path.join(self.repo_root, "src", "lib.cpp")
self.compiler = shutil.which("c++")
def _compile_harness(self, tmpdir, schema, harness):
schema_path = os.path.join(tmpdir, "schema.json")
with open(schema_path, "w") as fp:
json.dump(schema, fp)
header_path = os.path.join(tmpdir, "gen.h")
result = subprocess.run(
[
sys.executable,
SCRIPT,
schema_path,
"-o",
header_path,
"--namespace",
"test_schema",
],
capture_output=True,
text=True,
check=False,
)
self.assertEqual(result.returncode, 0, msg=result.stderr)
lib_obj = os.path.join(tmpdir, "lib.o")
comp_lib = subprocess.run(
[
self.compiler,
"-std=c++20",
"-I",
self.include_dir,
"-I",
os.path.join(self.repo_root, "third_party", "include"),
"-I",
os.path.join(self.repo_root, "third_party", "valgrind"),
"-c",
self.lib_src,
"-o",
lib_obj,
],
capture_output=True,
text=True,
check=False,
)
self.assertEqual(comp_lib.returncode, 0, msg=comp_lib.stderr)
cpp_path = os.path.join(tmpdir, "test.cpp")
with open(cpp_path, "w") as fp:
fp.write(harness)
exe_path = os.path.join(tmpdir, "test")
comp = subprocess.run(
[
self.compiler,
"-std=c++20",
"-I",
self.include_dir,
"-I",
tmpdir,
lib_obj,
cpp_path,
"-o",
exe_path,
],
capture_output=True,
text=True,
check=False,
)
self.assertEqual(comp.returncode, 0, msg=comp.stderr)
run = subprocess.run([exe_path], capture_output=True, text=True, check=False)
self.assertEqual(run.returncode, 0, msg=run.stdout + run.stderr)
def test_self_referential_nullable_object_accepts_nested_null(self):
"""A nullable object $def must stay nullable on recursive $refs."""
if not self.compiler:
self.skipTest("C++ compiler not available")
schema = {
"type": "object",
"additionalProperties": False,
"properties": {"self": {"$ref": "#/$defs/Self"}},
"$defs": {
"Self": {
"type": ["object", "null"],
"additionalProperties": False,
"properties": {"self": {"$ref": "#/$defs/Self"}},
}
},
}
harness = textwrap.dedent(
"""
#include "gen.h"
#include <cstdio>
#include <cstring>
struct Case { const char *s; WeaselJsonStatus expected; };
int main() {
Case cases[] = {
{ R"({"self": null})", WeaselJson_OK },
{ R"({"self": {"self": null}})", WeaselJson_OK },
{ R"({"self": {"self": {}}})", WeaselJson_OK },
};
for (const auto &c : cases) {
test_schema::RootBuilder b;
char buf[256];
std::strncpy(buf, c.s, sizeof(buf) - 1);
buf[sizeof(buf) - 1] = '\\0';
WeaselJsonStatus st = b.feed(buf, std::strlen(buf));
st = b.finish();
if (st != c.expected) {
std::printf("case %s expected %d got %d\\n", c.s, c.expected, st);
return 1;
}
}
return 0;
}
"""
)
with tempfile.TemporaryDirectory() as tmpdir:
self._compile_harness(tmpdir, schema, harness)
class SchemagenIntegerBoundaryTest(unittest.TestCase): class SchemagenIntegerBoundaryTest(unittest.TestCase):
"""Regression tests for issue #19: integer slot parsing near int64 boundaries.""" """Regression tests for issue #19: integer slot parsing near int64 boundaries."""
@@ -466,6 +610,13 @@ class SchemagenIntegerBoundaryTest(unittest.TestCase):
("123.0", "WeaselJson_OK", 123), ("123.0", "WeaselJson_OK", 123),
("9e18", "WeaselJson_OK", 9000000000000000000), ("9e18", "WeaselJson_OK", 9000000000000000000),
("10e18", "WeaselJson_REJECT", 0), ("10e18", "WeaselJson_REJECT", 0),
# Issue #56: zero written with a negative exponent must be accepted.
("0e-1", "WeaselJson_OK", 0),
("0e-2", "WeaselJson_OK", 0),
("0.0e-2", "WeaselJson_OK", 0),
("-0e-2", "WeaselJson_OK", 0),
("0e-20", "WeaselJson_OK", 0),
("0.000e-5", "WeaselJson_OK", 0),
] ]
cases_src = self._build_cases_array("root", cases) cases_src = self._build_cases_array("root", cases)
harness = textwrap.dedent( harness = textwrap.dedent(
@@ -508,6 +659,9 @@ class SchemagenIntegerBoundaryTest(unittest.TestCase):
('{"age":2.0}', "WeaselJson_OK", 2), ('{"age":2.0}', "WeaselJson_OK", 2),
('{"age":0.0001e4}', "WeaselJson_OK", 1), ('{"age":0.0001e4}', "WeaselJson_OK", 1),
('{"age":0.001}', "WeaselJson_REJECT", 0), ('{"age":0.001}', "WeaselJson_REJECT", 0),
# Issue #56: zero written with a negative exponent must be accepted.
('{"age":0e-2}', "WeaselJson_OK", 0),
('{"age":-0e-20}', "WeaselJson_OK", 0),
( (
'{"age":-9223372036854775808.0}', '{"age":-9223372036854775808.0}',
"WeaselJson_OK", "WeaselJson_OK",
@@ -544,6 +698,7 @@ class SchemagenIntegerBoundaryTest(unittest.TestCase):
) )
schema = { schema = {
"type": "object", "type": "object",
"additionalProperties": False,
"properties": {"age": {"type": "integer"}}, "properties": {"age": {"type": "integer"}},
"required": ["age"], "required": ["age"],
} }
@@ -659,14 +814,22 @@ class SchemagenStringEscapeTest(unittest.TestCase):
def test_newline_in_property_key(self): def test_newline_in_property_key(self):
"""A JSON key containing a newline must become a valid C++ literal.""" """A JSON key containing a newline must become a valid C++ literal."""
schema = {"type": "object", "properties": {"a\nb": {"type": "string"}}} schema = {
"type": "object",
"additionalProperties": False,
"properties": {"a\nb": {"type": "string"}},
}
out = self.generate_and_compile(schema) out = self.generate_and_compile(schema)
self.assertIn('// "a\\nb"', out) self.assertIn('// "a\\nb"', out)
self.assertIn('if (key == "a\\nb") return 0;', out) self.assertIn('if (key == "a\\nb") return 0;', out)
def test_newline_in_enum_value(self): def test_newline_in_enum_value(self):
"""An enum value containing a newline must become a valid C++ literal.""" """An enum value containing a newline must become a valid C++ literal."""
schema = {"type": "object", "properties": {"x": {"enum": ["a\nb"]}}} schema = {
"type": "object",
"additionalProperties": False,
"properties": {"x": {"enum": ["a\nb"]}},
}
out = self.generate_and_compile(schema) out = self.generate_and_compile(schema)
self.assertIn('static constexpr const char *X_names[] = { "a\\nb" };', out) self.assertIn('static constexpr const char *X_names[] = { "a\\nb" };', out)
@@ -674,6 +837,7 @@ class SchemagenStringEscapeTest(unittest.TestCase):
"""Mixed control characters in an enum value must be escaped.""" """Mixed control characters in an enum value must be escaped."""
schema = { schema = {
"type": "object", "type": "object",
"additionalProperties": False,
"properties": { "properties": {
"x": {"enum": ["x\ny\rz\tw\vq\x00\x01"]}, "x": {"enum": ["x\ny\rz\tw\vq\x00\x01"]},
}, },
@@ -686,7 +850,11 @@ class SchemagenStringEscapeTest(unittest.TestCase):
def test_backslash_and_quote_still_escaped(self): def test_backslash_and_quote_still_escaped(self):
"""Existing escaping for backslash and double quote must remain correct.""" """Existing escaping for backslash and double quote must remain correct."""
schema = {"type": "object", "properties": {'a"b\\c': {"type": "string"}}} schema = {
"type": "object",
"additionalProperties": False,
"properties": {'a"b\\c': {"type": "string"}},
}
out = self.generate_and_compile(schema) out = self.generate_and_compile(schema)
self.assertIn('// "a\\"b\\\\c"', out) self.assertIn('// "a\\"b\\\\c"', out)
self.assertIn('if (key == "a\\"b\\\\c") return 0;', out) self.assertIn('if (key == "a\\"b\\\\c") return 0;', out)
+8 -2
View File
@@ -417,7 +417,13 @@ class Builder:
tobj = TObj(name) tobj = TObj(name)
if defname is not None: if defname is not None:
self._building[defname] = (tobj, nullable) self._building[defname] = (tobj, nullable)
ap = node.get("additionalProperties", False) ap = node.get("additionalProperties", None)
if ap is None:
raise GenError(
"additionalProperties is required for object schemas; set it "
"explicitly to false to reject unknown keys (absent "
"additionalProperties is not supported)"
)
if ap is True: if ap is True:
raise GenError("additionalProperties: true is not supported") raise GenError("additionalProperties: true is not supported")
if isinstance(ap, dict): if isinstance(ap, dict):
@@ -1058,10 +1064,10 @@ private:
++trim; ++trim;
}} }}
int64_t finalExp = exp - fracDigits + trim; int64_t finalExp = exp - fracDigits + trim;
if (finalExp < 0) return false;
size_t leadingZeros = 0; size_t leadingZeros = 0;
while (leadingZeros < digits.size() && digits[leadingZeros] == '0') ++leadingZeros; while (leadingZeros < digits.size() && digits[leadingZeros] == '0') ++leadingZeros;
if (leadingZeros == digits.size()) {{ out = 0; return true; }} if (leadingZeros == digits.size()) {{ out = 0; return true; }}
if (finalExp < 0) return false;
if (leadingZeros > 0) digits.erase(0, leadingZeros); if (leadingZeros > 0) digits.erase(0, leadingZeros);
constexpr uint64_t kMaxNeg = 9223372036854775808ULL; constexpr uint64_t kMaxNeg = 9223372036854775808ULL;
+3 -1
View File
@@ -38,6 +38,8 @@ enum WeaselJsonStatus {
WeaselJson_REJECT, WeaselJson_REJECT,
/** json is too deeply nested */ /** json is too deeply nested */
WeaselJson_OVERFLOW, WeaselJson_OVERFLOW,
/** Tried to call parse on a null parser */
WeaselJson_NULL,
}; };
typedef struct WeaselJsonParser WeaselJsonParser; typedef struct WeaselJsonParser WeaselJsonParser;
@@ -65,7 +67,7 @@ void WeaselJsonParser_destroy(WeaselJsonParser *parser);
/** Incrementally parse `len` more bytes starting at `buf`. `buf` may be /** Incrementally parse `len` more bytes starting at `buf`. `buf` may be
* modified. Call with `len` 0 to indicate end of data. `buf` may be null if * modified. Call with `len` 0 to indicate end of data. `buf` may be null if
* `len` is 0. `len` must not be negative; a negative length is treated as a * `len` is 0. `len` must not be negative; a negative length is treated as a
* rejected input. */ * rejected input. Returns WeaselJson_NULL if parser is null */
WeaselJsonStatus WeaselJsonParser_parse(WeaselJsonParser *parser, char *buf, WeaselJsonStatus WeaselJsonParser_parse(WeaselJsonParser *parser, char *buf,
int len); int len);
+3
View File
@@ -46,6 +46,9 @@ WeaselJsonParser_destroy(WeaselJsonParser *parser) {
__attribute__((visibility("default"))) WeaselJsonStatus __attribute__((visibility("default"))) WeaselJsonStatus
WeaselJsonParser_parse(WeaselJsonParser *parser, char *buf, int len) { WeaselJsonParser_parse(WeaselJsonParser *parser, char *buf, int len) {
if (parser == nullptr) [[unlikely]] {
return WeaselJson_NULL;
}
return ((Parser3 *)parser)->parse(buf, len); return ((Parser3 *)parser)->parse(buf, len);
} }
} }
+10 -10
View File
@@ -139,7 +139,7 @@ struct Parser3 {
stackPtr = stack(); stackPtr = stack();
std::ignore = push({N_VALUE, N_WHITESPACE, T_EOF}); std::ignore = push({N_VALUE, N_WHITESPACE, T_EOF});
inKey = false; inKey = false;
rejected = false; terminalStatus = WeaselJson_OK;
utf8Codepoint = 0; utf8Codepoint = 0;
utf16Surrogate = 0; utf16Surrogate = 0;
minCodepoint = 0; minCodepoint = 0;
@@ -162,7 +162,7 @@ struct Parser3 {
NumDfa numDfa; NumDfa numDfa;
Utf8Dfa strDfa; Utf8Dfa strDfa;
bool inKey = false; bool inKey = false;
bool rejected = false; WeaselJsonStatus terminalStatus = WeaselJson_OK;
#ifndef HAS_MUSTTAIL #ifndef HAS_MUSTTAIL
char *stashBufForTrampoline; char *stashBufForTrampoline;
@@ -650,7 +650,7 @@ inline PRESERVE_NONE ContinuationStatus n_string2(Parser3 *self, char *buf,
self->writeBuf[0] = (0b00000111 & codepoint) | 0b11110000; self->writeBuf[0] = (0b00000111 & codepoint) | 0b11110000;
self->writeBuf += 4; self->writeBuf += 4;
} }
} else if (0xdc00 <= codepoint && codepoint <= 0xdfff) { } else if (0xdc00 <= codepoint && codepoint <= 0xdfff) [[unlikely]] {
return WeaselJson_REJECT; return WeaselJson_REJECT;
} else { } else {
if (!(self->flags & WeaselJsonRaw)) { if (!(self->flags & WeaselJsonRaw)) {
@@ -1070,12 +1070,12 @@ constexpr inline struct ContinuationTable {
inline WeaselJsonStatus Parser3::parse(char *buf, int len) { inline WeaselJsonStatus Parser3::parse(char *buf, int len) {
this->dataBegin = this->writeBuf = buf; this->dataBegin = this->writeBuf = buf;
if (this->rejected) [[unlikely]] { if (this->terminalStatus != WeaselJson_OK) [[unlikely]] {
return WeaselJson_REJECT; return this->terminalStatus;
} }
if (len < 0) [[unlikely]] { if (len < 0) [[unlikely]] {
this->rejected = true; this->terminalStatus = WeaselJson_REJECT;
return WeaselJson_REJECT; return WeaselJson_REJECT;
} }
@@ -1085,8 +1085,8 @@ inline WeaselJsonStatus Parser3::parse(char *buf, int len) {
// range. // range.
ContinuationStatus status = ContinuationStatus status =
symbolTables.continuations[top()](this, buf, buf + len); symbolTables.continuations[top()](this, buf, buf + len);
if (status == WeaselJson_REJECT) { if (status > WeaselJson_AGAIN) {
this->rejected = true; this->terminalStatus = WeaselJsonStatus(status);
} }
return WeaselJsonStatus(status); return WeaselJsonStatus(status);
#else #else
@@ -1095,8 +1095,8 @@ inline WeaselJsonStatus Parser3::parse(char *buf, int len) {
while ((result = symbolTables.continuations[top()]( while ((result = symbolTables.continuations[top()](
this, stashBufForTrampoline, buf + len)) == kBounce) this, stashBufForTrampoline, buf + len)) == kBounce)
; ;
if (result == WeaselJson_REJECT) { if (result > WeaselJson_AGAIN) {
this->rejected = true; this->terminalStatus = WeaselJsonStatus(result);
} }
return WeaselJsonStatus(result); return WeaselJsonStatus(result);
#endif #endif
+60
View File
@@ -202,6 +202,62 @@ TEST_CASE("parser3") {
} }
} }
TEST_CASE("overflow state is sticky") {
auto c = noopCallbacks();
// stackSize 3 is exactly big enough to hold reset()'s bootstrap, but too
// small for nested arrays. Overflows must be terminal like rejects: a later
// end-of-data call must never report OK for an incomplete document.
auto *parser = WeaselJsonParser_create(3, &c, nullptr, 0);
REQUIRE(parser != nullptr);
std::string doc = "[[";
REQUIRE(WeaselJsonParser_parse(parser, doc.data(), doc.size()) ==
WeaselJson_OVERFLOW);
// After overflow, the end-of-data call must not return OK.
REQUIRE(WeaselJsonParser_parse(parser, nullptr, 0) != WeaselJson_OK);
// It should keep reporting a terminal failure.
REQUIRE(WeaselJsonParser_parse(parser, nullptr, 0) != WeaselJson_OK);
// Further data chunks must also stay terminal.
std::string more = "]]";
REQUIRE(WeaselJsonParser_parse(parser, more.data(), more.size()) !=
WeaselJson_OK);
WeaselJsonParser_destroy(parser);
}
TEST_CASE("overflow is sticky for nested objects") {
auto c = noopCallbacks();
auto *parser = WeaselJsonParser_create(4, &c, nullptr, 0);
REQUIRE(parser != nullptr);
std::string doc = "{\"a\":{ \"a\":";
REQUIRE(WeaselJsonParser_parse(parser, doc.data(), doc.size()) ==
WeaselJson_OVERFLOW);
REQUIRE(WeaselJsonParser_parse(parser, nullptr, 0) != WeaselJson_OK);
WeaselJsonParser_destroy(parser);
}
TEST_CASE("reset clears overflow state") {
auto c = noopCallbacks();
auto *parser = WeaselJsonParser_create(3, &c, nullptr, 0);
REQUIRE(parser != nullptr);
std::string doc = "[[";
REQUIRE(WeaselJsonParser_parse(parser, doc.data(), doc.size()) ==
WeaselJson_OVERFLOW);
// After reset the parser should accept a minimal document again.
WeaselJsonParser_reset(parser);
std::string copy = "1";
REQUIRE(WeaselJsonParser_parse(parser, copy.data(), copy.size()) ==
WeaselJson_AGAIN);
REQUIRE(WeaselJsonParser_parse(parser, nullptr, 0) == WeaselJson_OK);
WeaselJsonParser_destroy(parser);
}
TEST_CASE("rejected state is sticky") { TEST_CASE("rejected state is sticky") {
auto c = noopCallbacks(); auto c = noopCallbacks();
auto *parser = WeaselJsonParser_create(1024, &c, nullptr, 0); auto *parser = WeaselJsonParser_create(1024, &c, nullptr, 0);
@@ -274,6 +330,10 @@ TEST_CASE("parse rejects negative length") {
WeaselJsonParser_destroy(parser); WeaselJsonParser_destroy(parser);
} }
TEST_CASE("Calling parse with nullptr doesn't crash") {
REQUIRE(WeaselJsonParser_parse(nullptr, nullptr, 0) == WeaselJson_NULL);
}
TEST_CASE("streaming") { testStreaming(json); } TEST_CASE("streaming") { testStreaming(json); }
TEST_CASE("reset clears inKey and transient state") { TEST_CASE("reset clears inKey and transient state") {
+3
View File
@@ -33,6 +33,9 @@ int main(int argc, char **argv) {
case WeaselJson_REJECT: case WeaselJson_REJECT:
case WeaselJson_OVERFLOW: case WeaselJson_OVERFLOW:
return 1; return 1;
case WeaselJson_NULL:
fprintf(stderr, "parse called with a null parser\n");
return 1;
} }
if (l == 0) { if (l == 0) {
return 1; return 1;
+1
View File
@@ -30,6 +30,7 @@ class WeaselJsonStatus(enum.Enum):
AGAIN = 1 AGAIN = 1
REJECT = 2 REJECT = 2
OVERFLOW = 3 OVERFLOW = 3
NULL = 4
class WeaselJsonCallbacksBase: class WeaselJsonCallbacksBase: