schemagen: deduplicate enum constants that collide after sanitization
CI / build (-DCMAKE_C_COMPILER=clang -DCMAKE_CXX_COMPILER=clang++, clang-arm64, ubuntu-latest-arm64, true) (pull_request) Successful in 51s
CI / build (-DCMAKE_C_COMPILER=gcc -DCMAKE_CXX_COMPILER=g++, gcc-arm64, ubuntu-latest-arm64, false) (pull_request) Successful in 51s
CI / pre-commit (pull_request) Successful in 51s
CI / build (-DCMAKE_C_COMPILER=clang -DCMAKE_CXX_COMPILER=clang++, clang-amd64, ubuntu-latest-amd64, true) (pull_request) Successful in 1m29s
CI / build (-DCMAKE_C_COMPILER=gcc -DCMAKE_CXX_COMPILER=g++, gcc-amd64, ubuntu-latest-amd64, false) (pull_request) Successful in 1m22s

Distinct JSON enum values can sanitize to the same C++ identifier
(e.g. "foo-bar" and "foo_bar" both become `foo_bar`), producing an
invalid `enum class` with duplicate constants.

Add `unique_enum_identifiers()` which appends a numeric suffix to
later collisions while preserving enum declaration order, so the
index-to-JSON-value mapping used by the generated parser stays intact.

Also add contrib/schemagen/test_schemagen.py and wire the schemagen
Python tests plus the existing example.schema.json/test_gen.cpp example
into ctest via CMakeLists.txt.

Closes #4
This commit is contained in:
2026-06-18 11:02:47 -04:00
parent 7a8e5f84f2
commit 359f4f4bb6
3 changed files with 114 additions and 1 deletions
+29 -1
View File
@@ -86,6 +86,34 @@ def sanitize(name, fallback="x"):
return s
def unique_enum_identifiers(values):
"""Return a list of sanitized C++ identifiers, one per input value.
Distinct JSON enum values may sanitize to the same C++ token (e.g.
"foo-bar" and "foo_bar" both become "foo_bar"). This helper appends a
numeric suffix to later collisions so the generated `enum class` stays
valid while preserving the original order and therefore the index-to-value
mapping used at parse time.
"""
used = set()
out = []
for v in values:
base = sanitize(v)
if base not in used:
used.add(base)
out.append(base)
continue
n = 1
while True:
candidate = f"{base}_{n}"
if candidate not in used:
used.add(candidate)
out.append(candidate)
break
n += 1
return out
def camel(name):
parts = [p for p in name.replace("-", " ").replace("_", " ").split(" ") if p]
if not parts:
@@ -508,7 +536,7 @@ namespace {ns} {{"""
return ""
out = []
for e in self.b.enums.values():
vals = ", ".join(sanitize(v) for v in e.values)
vals = ", ".join(unique_enum_identifiers(e.values))
out.append(f"enum class {e.name} : int {{ {vals} }};")
return "\n".join(out) + "\n"