Memory corruption via truncated tokenizer result size
Published Jun 24, 2025 · Updated Jun 24, 2025
Heap buffer overflow in ggml-org llama.cpp before b5721 allows local users to corrupt process memory through crafted tokenizer input. The llama_vocab::tokenize adapter casts res.size() from size_t to signed int before comparing it with n_tokens_max; results above INT_MAX become negative, bypass the size guard, and the copy loop writes past the caller's heap buffer. Reaching the overflow requires a tokenized result above INT_MAX and user interaction with crafted text; the published proof uses a Jinja template and a non-BPE model, and code execution has not been demonstrated.
Summary
What happened
Heap buffer overflow in ggml-org llama.cpp before b5721 allows local users to corrupt process memory through crafted tokenizer input. The llama_vocab::tokenize adapter casts res.size() from size_t to signed int before comparing it with n_tokens_max; results above INT_MAX become negative, bypass the size guard, and the copy loop writes past the caller's heap buffer. Reaching the overflow requires a tokenized result above INT_MAX and user interaction with crafted text; the published proof uses a Jinja template and a non-BPE model, and code execution has not been demonstrated.
The record
- CVE
- CVE-2025-52566
- Published
- Jun 24, 2025
- Updated
- Jun 24, 2025
- Vendor
- Georgi Gerganov
- Product
- llama.cpp
- Classifications
- CWE-119, CWE-195
- Attack vector
- local
- Privileges
- unauthenticated
Timeline
How it unfolded
- Jun 24, 2025CVE publishedPublication date reported by the CVE source.
- Jun 24, 2025Record updatedLatest update available in the CVE record.
Exploitability
Present is not the same as exploitable
Compare your product and version with the public record. A matching version still requires validation against your environment.
Is a vulnerable build present?
Compare these published version ranges with your installed build and any vendor patches.
- Affected versionversion=< b5721
What conditions does exploitation require?
What is affected?
Published CVSS scores
CVSS describes severity. EPSS estimates exploitation probability.
Attacks
What attackers are doing with it
Daily unique IPs observed by Shadowserver honeypots for known exploited vulnerabilities (KEVs). Missing observations do not establish an absence of attacks.
Weakness, pattern, technique
Public exploit references
- Upstream tokenizer overflow proof procedureproof of concept · demonstrated
Labels summarize the accepted research assessment. They do not indicate a test against your environment.
Technologies
Your stack
See the directory against your own environment.
Your stack
Check the software in your environment
Book a demo to see how Hinoki identifies affected software and validates exploitability in your environment.
Book a demo