Arbitrary code execution via unsafe benchmark exec
Published Apr 9, 2026 · Updated Apr 9, 2026
Code injection in MetaGPT 0.8.0 and 0.8.1 allows attackers to execute arbitrary Python code through poisoned benchmark inputs. HumanEvalBenchmark and MBPPBenchmark check_solution pass LLM-generated solutions and test cases to Python exec after ineffective sanitization. An attacker must influence an LLM response or evaluation dataset; the aflow optimization loop then runs the payload on the MetaGPT host without review.
Summary
What happened
Code injection in MetaGPT 0.8.0 and 0.8.1 allows attackers to execute arbitrary Python code through poisoned benchmark inputs. HumanEvalBenchmark and MBPPBenchmark check_solution pass LLM-generated solutions and test cases to Python exec after ineffective sanitization. An attacker must influence an LLM response or evaluation dataset; the aflow optimization loop then runs the payload on the MetaGPT host without review.
The record
- CVE
- CVE-2026-5970
- Published
- Apr 9, 2026
- Updated
- Apr 9, 2026
- Vendor
- Foundation Agents
- Product
- MetaGPT
- Classifications
- CWE-94, CWE-74, T1059.006
- Attack vector
- local
- Privileges
- unauthenticated
Timeline
How it unfolded
- Apr 9, 2026CVE publishedPublication date reported by the CVE source.
- Apr 9, 2026Record updatedLatest update available in the CVE record.
Exploitability
Present is not the same as exploitable
Compare your product and version with the public record. A matching version still requires validation against your environment.
Is a vulnerable build present?
Compare these published version ranges with your installed build and any vendor patches.
- Affected versionversion=0.8.0
- Affected versionversion=0.8.1
What conditions does exploitation require?
What is affected?
Published CVSS scores
CVSS describes severity. EPSS estimates exploitation probability.
Attacks
What attackers are doing with it
Daily unique IPs observed by Shadowserver honeypots for known exploited vulnerabilities (KEVs). Missing observations do not establish an absence of attacks.
Weakness, pattern, technique
Public exploit references
- MetaGPT benchmark exec proof of conceptproof of concept · demonstrated
Labels summarize the accepted research assessment. They do not indicate a test against your environment.
Technologies
Your stack
See the directory against your own environment.
Your stack
Check the software in your environment
Book a demo to see how Hinoki identifies affected software and validates exploitability in your environment.
Book a demo