Dataset identity spoofing via truncated deterministic digest
Published Jun 4, 2026 · Updated Jun 4, 2026
Weak hashing in MLflow 3.0 through 3.10.0 allows local users to make different datasets produce the same recorded digest. The client-side digest code hashes only the first 10,000 rows and selected column types, then truncates MD5 output to eight hexadecimal characters. A user able to supply datasets can preserve sampled content or change omitted fields, undermining dataset identity and lineage without altering the recorded digest.
Summary
What happened
Weak hashing in MLflow 3.0 through 3.10.0 allows local users to make different datasets produce the same recorded digest. The client-side digest code hashes only the first 10,000 rows and selected column types, then truncates MD5 output to eight hexadecimal characters. A user able to supply datasets can preserve sampled content or change omitted fields, undermining dataset identity and lineage without altering the recorded digest.
The record
- CVE
- CVE-2026-10803
- Published
- Jun 4, 2026
- Updated
- Jun 4, 2026
- Vendor
- MLflow Project
- Product
- MLflow
- Classifications
- CWE-327, CWE-328, T1565.001
- Attack vector
- local
- Privileges
- authenticated
Timeline
How it unfolded
- Jun 4, 2026CVE publishedPublication date reported by the CVE source.
- Jun 4, 2026Record updatedLatest update available in the CVE record.
Exploitability
Present is not the same as exploitable
Compare your product and version with the public record. A matching version still requires validation against your environment.
Is a vulnerable build present?
Compare these published version ranges with your installed build and any vendor patches.
- Affected versionversion=3.0
- Affected versionversion=3.1
- Affected versionversion=3.10.0
- Affected versionversion=3.2
- Affected versionversion=3.3
- Affected versionversion=3.4
- Affected versionversion=3.5
- Affected versionversion=3.6
- Affected versionversion=3.7
- Affected versionversion=3.8
- Affected versionversion=3.9
What conditions does exploitation require?
What is affected?
Attacks
What attackers are doing with it
Daily unique IPs observed by Shadowserver honeypots for known exploited vulnerabilities (KEVs). Missing observations do not establish an absence of attacks.
Weakness, pattern, technique
Public exploit references
- Deterministic dataset-digest collision reproductionproof of concept · demonstrated
Labels summarize the accepted research assessment. They do not indicate a test against your environment.
Technologies
Your stack
See the directory against your own environment.
Your stack
Check the software in your environment
Book a demo to see how Hinoki identifies affected software and validates exploitability in your environment.
Book a demo