Home/ Blog/ Security news/ Article
Blog · Security news

MLflow's unauthenticated SSRF flaw (CVE-2026-64849) can leak cloud credentials. Patch to 3.15.0 now.

CVE-2026-64849 is an unauthenticated, full-read SSRF in MLflow's webhook delivery. It reaches cloud metadata and leaks instance credentials. Fixed in 3.15.0.

A slightly open metal wall hatch with a mechanical arm holding out a small key

A server-side request forgery bug is usually a nuisance: you can make the server knock on a door, but you rarely get to read what is behind it. CVE-2026-64849 in MLflow is the dangerous version. It is unauthenticated, it hands back the response body, and the door it opens is the cloud instance-metadata service where your machine's credentials live. CISA added it to the Known Exploited Vulnerabilities catalog on August 19, 2026, which means it is being used against real deployments right now. If you run an MLflow tracking server on AWS, GCP, or Azure, treat this as today's work.

The fix is MLflow 3.15.0. Everything before it is affected. The federal deadline to patch is September 2, 2026, and there is no reason for anyone else to move slower than that.

CVE-2026-64849 at a glance
9.3
CVSS base score
CWE-918, full-read SSRF
0
credentials required
unauthenticated test endpoint
Sep 2
federal patch deadline
added to CISA KEV Aug 19, 2026
Source: CISA KEV, NVD, and the MLflow security advisory.

What the flaw actually does

MLflow shipped a webhook feature so the model registry can notify external systems when a model version changes. To help operators check their setup, it exposes a test endpoint, POST /api/2.0/mlflow/webhooks/{id}/test, that fires the webhook on demand. Per the advisory, that endpoint requires no authentication.

When a webhook is registered, MLflow runs its URL check, _validate_webhook_url(), from mlflow/utils/validation.py. That function looks up the hostname, confirms each address it gets back is public, and then discards it. Sending the webhook is a separate code path, mlflow/webhooks/delivery.py. At send time it looks the name up a second time and chases any redirects, and it never pins the address the earlier check approved. The two paths disagree about which IP they are talking to, and that disagreement is the bug.

An attacker registers a webhook pointing at a hostname they control. On the first lookup, their DNS answers with a public IP and the check passes. On the second lookup, milliseconds later when delivery connects, the same hostname answers with 169.254.169.254 or 127.0.0.1. This is classic DNS rebinding: a time-of-check to time-of-use race against a validator that trusts a name instead of a connection. A redirect to an internal URL gets to the same place.

Here is the part that turns a knock into a theft. The test endpoint returns response_status and response_body to the caller. On a cloud host with the metadata service reachable, that body is the instance role's temporary credentials. The attacker does not have to guess or infer anything from timing. They read the keys straight out of the reply.

Why self-hosted ML infrastructure keeps getting hit this way

This is the second self-hosted machine-learning platform in short order to land on the KEV list through a DNS-rebinding SSRF. We wrote up the same failure class in Ray's dashboard, where an internal-only API was reachable through the same rebinding trick. The pattern is not a coincidence.

ML platforms are built to be convenient inside a "trusted" network. They ship features that reach out to other services: webhooks, artifact stores, remote tracking, plugin callbacks. Operators run them on a subnet they think of as internal, often with authentication turned off because "only the data team can reach it." Both assumptions fail the moment one unauthenticated endpoint can be pointed at the metadata service. The network is not a trust boundary when the application will happily make requests on an attacker's behalf.

The specific defect here, validate-the-name-then-connect-to-whatever-it-resolves-to, is architecturally broken against rebinding no matter which vendor writes it. MLflow's own fix is the correct one: a hardened HTTP adapter that inspects the far end of the socket it actually opened, the instant connect() returns and before a single byte of data is exchanged. If that peer is a private or metadata address, the request dies. It also turns off environment proxy handling so a proxy cannot route around the check. You cannot reproduce that logic from outside the app, which is exactly why the containment work has to happen at a different layer.

What to do now

Patch to MLflow 3.15.0. That closes the endpoint. But patching a fleet of tracking servers takes time, and check-time URL validation is defeated by rebinding by design, so do not stop at the upgrade. Move the defense to where the payoff lives:

  • Enforce IMDSv2 and drop IMDSv1. On AWS, require the token-based metadata service and set the hop limit to 1. A rebinding SSRF from an application cannot complete the IMDSv2 token handshake the way it can hit the old unauthenticated endpoint, so even a successful SSRF returns nothing useful. This is the one control that holds while you are still patching, and it survives the next SSRF too.

  • Deny egress from the MLflow host to the metadata IP. Block 169.254.169.254 at the host firewall or with an egress policy. The tracking server has no legitimate reason to talk to its own metadata endpoint.

  • Take the tracking server off any network an attacker can reach. MLflow's tracking server has no authentication by default. If it is exposed beyond the team that needs it, the unauthenticated test endpoint is reachable by anyone who finds it.

  • Scope the instance role down. The credentials this bug exposes are only as valuable as the role behind them. A tracking server that can read one bucket is a smaller loss than one that can assume broad account permissions.

How you would know you were hit

Detection here is not a single log line, it is a correlation. Watch for two things and connect them.

On the MLflow side, look at access logs for calls to /api/2.0/mlflow/webhooks, especially the /test action, and for any webhook whose target hostname resolves to a private, loopback, or link-local address. A webhook configured to reach an internal or metadata range is not a normal operator action. Freshly created webhooks followed immediately by a test call, from a source that is not your CI or your data team, is the shape of an attempt.

On the cloud side, the real evidence is the credential being used somewhere else. If your instance role suddenly makes API calls from an IP or a service that is not the instance, or performs actions the tracking server never does, treat the role as compromised. Turn on and review your provider's metadata-access and role-usage logs. Because a full-read SSRF hands over live credentials, the window between exposure and abuse can be minutes, so the correlation has to be an alert, not a monthly report.

The takeaway

The uncomfortable detail in this one is the test button. A feature added to make webhooks easier to set up is the exact primitive that reads your cloud keys, because it is unauthenticated and it returns the response body. Convenience endpoints that echo a remote response back to the caller deserve the same scrutiny as any credential store. Patch to 3.15.0, enforce IMDSv2 today, and assume the next self-hosted ML tool you deploy has a reach-out feature waiting to be pointed somewhere it should not go.

Topics

Frequently asked questions

What is CVE-2026-64849 in MLflow?

CVE-2026-64849 is an unauthenticated server-side request forgery flaw in MLflow's webhook delivery, rated CVSS 9.3. An attacker can reach internal or cloud metadata services and read the response, including instance credentials. It affects MLflow before 3.15.0 and is fixed in 3.15.0.

How does the MLflow SSRF let an attacker steal cloud credentials?

MLflow validates a webhook URL by resolving the hostname and checking it is public, then discards that address. Delivery re-resolves the name, so DNS rebinding points it at 169.254.169.254. The test endpoint returns the response body, which on a cloud host is the instance role's live credentials.

Which MLflow versions are affected and what is the fix?

Every MLflow version before 3.15.0 is affected. Upgrading to 3.15.0 fixes it by validating the peer IP of the actual connected socket instead of the hostname, which closes the DNS-rebinding window. CISA set a federal patch deadline of September 2, 2026.

Is CVE-2026-64849 being actively exploited?

Yes. CISA added CVE-2026-64849 to its Known Exploited Vulnerabilities catalog on August 19, 2026, which is confirmation that it is being exploited in the wild. Known ransomware use is listed as unknown, but active exploitation means it should be patched immediately.

How can I contain the MLflow SSRF before I finish patching?

Enforce IMDSv2 and set the metadata hop limit to 1 so a rebinding SSRF cannot complete the token handshake. Block egress from the MLflow host to 169.254.169.254, keep the tracking server off attacker-reachable networks, and scope the instance role to the minimum it needs.

How would I detect an attempt against this MLflow flaw?

Watch MLflow access logs for calls to the webhook test endpoint and for webhooks whose target resolves to private, loopback, or metadata addresses. On the cloud side, alert when the instance role is used from an IP or service that is not the instance, which signals the leaked credentials are being abused.

Ready to meet the Guardians?

Deploys fast - agentless for monitoring and cloud, a lightweight agent for deep endpoint security. Just Suriq, standing watch.