schemagen: deduplicate enum constants that collide after sanitization
CI / build (-DCMAKE_C_COMPILER=clang -DCMAKE_CXX_COMPILER=clang++, clang-arm64, ubuntu-latest-arm64, true) (pull_request) Successful in 51s
CI / build (-DCMAKE_C_COMPILER=gcc -DCMAKE_CXX_COMPILER=g++, gcc-arm64, ubuntu-latest-arm64, false) (pull_request) Successful in 51s
CI / pre-commit (pull_request) Successful in 51s
CI / build (-DCMAKE_C_COMPILER=clang -DCMAKE_CXX_COMPILER=clang++, clang-amd64, ubuntu-latest-amd64, true) (pull_request) Successful in 1m29s
CI / build (-DCMAKE_C_COMPILER=gcc -DCMAKE_CXX_COMPILER=g++, gcc-amd64, ubuntu-latest-amd64, false) (pull_request) Successful in 1m22s
CI / build (-DCMAKE_C_COMPILER=clang -DCMAKE_CXX_COMPILER=clang++, clang-arm64, ubuntu-latest-arm64, true) (pull_request) Successful in 51s
CI / build (-DCMAKE_C_COMPILER=gcc -DCMAKE_CXX_COMPILER=g++, gcc-arm64, ubuntu-latest-arm64, false) (pull_request) Successful in 51s
CI / pre-commit (pull_request) Successful in 51s
CI / build (-DCMAKE_C_COMPILER=clang -DCMAKE_CXX_COMPILER=clang++, clang-amd64, ubuntu-latest-amd64, true) (pull_request) Successful in 1m29s
CI / build (-DCMAKE_C_COMPILER=gcc -DCMAKE_CXX_COMPILER=g++, gcc-amd64, ubuntu-latest-amd64, false) (pull_request) Successful in 1m22s
Distinct JSON enum values can sanitize to the same C++ identifier (e.g. "foo-bar" and "foo_bar" both become `foo_bar`), producing an invalid `enum class` with duplicate constants. Add `unique_enum_identifiers()` which appends a numeric suffix to later collisions while preserving enum declaration order, so the index-to-JSON-value mapping used by the generated parser stays intact. Also add contrib/schemagen/test_schemagen.py and wire the schemagen Python tests plus the existing example.schema.json/test_gen.cpp example into ctest via CMakeLists.txt. Closes #4
This commit is contained in:
@@ -86,6 +86,34 @@ def sanitize(name, fallback="x"):
|
||||
return s
|
||||
|
||||
|
||||
def unique_enum_identifiers(values):
|
||||
"""Return a list of sanitized C++ identifiers, one per input value.
|
||||
|
||||
Distinct JSON enum values may sanitize to the same C++ token (e.g.
|
||||
"foo-bar" and "foo_bar" both become "foo_bar"). This helper appends a
|
||||
numeric suffix to later collisions so the generated `enum class` stays
|
||||
valid while preserving the original order and therefore the index-to-value
|
||||
mapping used at parse time.
|
||||
"""
|
||||
used = set()
|
||||
out = []
|
||||
for v in values:
|
||||
base = sanitize(v)
|
||||
if base not in used:
|
||||
used.add(base)
|
||||
out.append(base)
|
||||
continue
|
||||
n = 1
|
||||
while True:
|
||||
candidate = f"{base}_{n}"
|
||||
if candidate not in used:
|
||||
used.add(candidate)
|
||||
out.append(candidate)
|
||||
break
|
||||
n += 1
|
||||
return out
|
||||
|
||||
|
||||
def camel(name):
|
||||
parts = [p for p in name.replace("-", " ").replace("_", " ").split(" ") if p]
|
||||
if not parts:
|
||||
@@ -508,7 +536,7 @@ namespace {ns} {{"""
|
||||
return ""
|
||||
out = []
|
||||
for e in self.b.enums.values():
|
||||
vals = ", ".join(sanitize(v) for v in e.values)
|
||||
vals = ", ".join(unique_enum_identifiers(e.values))
|
||||
out.append(f"enum class {e.name} : int {{ {vals} }};")
|
||||
return "\n".join(out) + "\n"
|
||||
|
||||
|
||||
Reference in New Issue
Block a user