Summary
datamodel-code-generator vulnerable to code injection in via attacker-controlled default_factory schema field
datamodel-code-generator is vulnerable to code injection when generating Python models from an attacker-controlled JSON Schema, OpenAPI, YAML, JSON, Avro, Protobuf, or XSD schema. When a property carries a "default_factory" key, its value is interpolated verbatim, as a raw Python expression, into the generated Field(default_factory=...) / field(default_factory=...) call. Because this assignment is evaluated at class-definition time (i.e. on import of the generated module), an attacker who controls the schema controls a Python expression that runs in the consumer's process. No special CLI flags are required.
Details
The vulnerable chain spans the JSON-Schema-shaped parser and three sink locations (Pydantic v2, dataclass, msgspec):
Source, schema → extras:
src/datamodel_code_generator/parser/jsonschema.py:600-614,DEFAULT_FIELD_KEYSincludes the literal string"default_factory".src/datamodel_code_generator/parser/jsonschema.py:457-459,JsonSchemaObject.__init__stores any non-standard key (includingdefault_factory) inself.extras.src/datamodel_code_generator/parser/jsonschema.py:797-812,get_field_extraspreservesdefault_factorythrough to the field model.
Sinks, extras → generated Python expression:
src/datamodel_code_generator/model/pydantic_base.py:222-249:default_factory = data.pop("default_factory", None) ... if default_factory is not None: field_arguments = [f"default_factory={default_factory}", *field_arguments]The
default_factoryvalue is interpolated raw (norepr(), no validation).src/datamodel_code_generator/model/dataclass.py:211:f"{k}={v if k == 'default_factory' else repr(v)}"Explicit special-case to skip
repr()fordefault_factory.src/datamodel_code_generator/model/msgspec.py:361, same pattern as dataclass.
Because default_factory is in DEFAULT_FIELD_KEYS, no special CLI flag is needed to reach the sink. Any input format that uses the JSON-Schema-shaped parser (jsonschema, openapi, yaml, json, dict, csv), and any input format that converts to it (avro, protobuf, xmlschema), is in scope.
Confirmed PoC matrix
| Input file type | Output model type | Result |
|---|---|---|
jsonschema |
pydantic_v2.BaseModel |
RCE on import |
jsonschema |
dataclasses.dataclass |
RCE on import |
jsonschema |
msgspec.Struct |
RCE on import |
jsonschema |
typing.TypedDict |
safe (TypedDict doesn't render field(); default_factory silently dropped) |
openapi |
pydantic_v2.BaseModel |
RCE on import |
Other JSON-Schema-shaped inputs (yaml, json, dict, csv, avro, protobuf, xmlschema) follow the same code path and are expected to reproduce.
PoC
Self contained Proof of Concept is available at my secret gist: https://gist.github.com/thegr1ffyn/9648b0fe4fcf7d569ac8e61dd11eebaf
Resolution
The fix validates schema-provided default_factory values while extracting JSON Schema field extras. Only the supported factory names dict, list, and set are accepted; any other value now raises a generator error before code generation. Generator-created default factories for supported mutable defaults and optional nested models continue to use the existing code paths.
Submitted by: Hamza Haroon (thegr1ffyn)
Impact
- Who's affected: any developer or CI pipeline that runs
datamodel-codegenagainst a schema they didn't author themselves, third-party API specs, schemas pulled from a registry, vendored upstream.json/.yaml/.avsc/.proto/.xsdfiles, schemas fetched from a remote URL or introspection endpoint, and who imports the generated.py. - What it gains: arbitrary Python code execution in the importer's process at
importtime. The PoC copies/etc/passwdto a tmp file to demonstrate arbitrary read; the same primitive supports any operation the importing process can perform (filesystem write, environment exfiltration, secondary network calls, RCE on CI runners). - What it does NOT need: no special CLI flags, no custom templates, no
--extra-template-data, no--use-schema-description. Default invocation against a malicious schema is sufficient. - What does block it: choosing
--output-model-type typing.TypedDict(which doesn't renderfield()/Field()calls). All other supported output model types are vulnerable.
Untrusted input is evaluated as executable code within the application's runtime environment. Typical impact: arbitrary code execution within the application's privilege context.
CVE-2026-54653 has a CVSS score of 8.8 (High). The vector is network-reachable, no privileges required, and user interaction required. A CVSS score reflects the worst-case severity of the vulnerability, not your specific exposure. Whether this affects your application depends on whether the vulnerable code is present and reachable in your environment. A fixed version is available (0.60.2); upgrading removes the vulnerable code path.
Affected versions
Security releases
Kodem intelligence
Severity tells you how bad this could be in the worst case. It does not tell you whether you are exposed. Exploitability and impact are functions of runtime truth: whether the vulnerable code is present, reachable, and actually executes in your application. A vulnerable package can sit in your dependency tree and never run.
Kodem, an Intelligent Application Security platform, uses runtime intelligence to reveal which vulnerabilities actually execute in production, so teams prioritize the ones that genuinely matter. Kodem's runtime-powered SCA identifies whether this CVE is reachable in your applications.
Already deployed Kodem?
See it in your environmentNew to Kodem? Get a demo →Remediation advice
Upgrade to datamodel-code-generator 0.60.2 or later.
This issue affects datamodel-code-generator versions >= 0.17.0, <= 0.60.1 and is fixed in 0.60.2.
Frequently Asked Questions
- What is CVE-2026-54653? CVE-2026-54653 is a high-severity code injection vulnerability in datamodel-code-generator (pip), affecting versions >= 0.17.0, <= 0.60.1. It is fixed in 0.60.2. Untrusted input is evaluated as executable code within the application's runtime environment.
- How severe is CVE-2026-54653? CVE-2026-54653 has a CVSS score of 8.8 (High). This score reflects the worst-case severity of the vulnerability, not your specific exposure. Whether it represents real risk in your environment depends on whether the vulnerable code is present and reachable.
- Which versions of datamodel-code-generator are affected by CVE-2026-54653? datamodel-code-generator (pip) versions >= 0.17.0, <= 0.60.1 is affected.
- Is there a fix for CVE-2026-54653? Yes. CVE-2026-54653 is fixed in 0.60.2. Upgrade to this version or later.
- Is CVE-2026-54653 exploitable, and should I be worried? Whether CVE-2026-54653 is exploitable in your environment depends on whether the vulnerable code is present and reachable. A CVSS score is a worst-case rating; it does not account for your specific deployment, configuration, or usage patterns. Kodem, an Intelligent Application Security platform, uses runtime intelligence to show which vulnerabilities actually execute in production, so you can focus on the ones that represent real risk. Get a demo
- What actually determines whether CVE-2026-54653 is exploitable, and how bad it is? Exploitability and impact are not fixed properties of a CVE. They depend on runtime truth: whether the vulnerable code is present, reachable, and actually executes in your application. A high CVSS score on a dependency that never runs is not the same as real risk. Kodem, an Intelligent Application Security platform, uses runtime intelligence to reveal which vulnerabilities actually execute in production, so teams prioritize the ones that genuinely matter.
- How do I fix CVE-2026-54653? Upgrade
datamodel-code-generatorto 0.60.2 or later.