A server-side request forgery bug is usually a nuisance: you can make the server knock on a door, but you rarely get to read what is behind it. CVE-2026-64849 in MLflow is the dangerous version. It is unauthenticated, it hands back the response body, and the door it opens is the cloud instance-metadata service where your machine's credentials live. CISA added it to the Known Exploited Vulnerabilities catalog on August 19, 2026, which means it is being used against real deployments right now. If you run an MLflow tracking server on AWS, GCP, or Azure, treat this as today's work.
The fix is MLflow 3.15.0. Everything before it is affected. The federal deadline to patch is September 2, 2026, and there is no reason for anyone else to move slower than that.
What the flaw actually does
MLflow shipped a webhook feature so the model registry can notify external systems when a model version changes. To help operators check their setup, it exposes a test endpoint, POST /api/2.0/mlflow/webhooks/{id}/test, that fires the webhook on demand. Per the advisory, that endpoint requires no authentication.
When a webhook is registered, MLflow runs its URL check, _validate_webhook_url(), from mlflow/utils/validation.py. That function looks up the hostname, confirms each address it gets back is public, and then discards it. Sending the webhook is a separate code path, mlflow/webhooks/delivery.py. At send time it looks the name up a second time and chases any redirects, and it never pins the address the earlier check approved. The two paths disagree about which IP they are talking to, and that disagreement is the bug.
An attacker registers a webhook pointing at a hostname they control. On the first lookup, their DNS answers with a public IP and the check passes. On the second lookup, milliseconds later when delivery connects, the same hostname answers with 169.254.169.254 or 127.0.0.1. This is classic DNS rebinding: a time-of-check to time-of-use race against a validator that trusts a name instead of a connection. A redirect to an internal URL gets to the same place.
Here is the part that turns a knock into a theft. The test endpoint returns response_status and response_body to the caller. On a cloud host with the metadata service reachable, that body is the instance role's temporary credentials. The attacker does not have to guess or infer anything from timing. They read the keys straight out of the reply.
Why self-hosted ML infrastructure keeps getting hit this way
This is the second self-hosted machine-learning platform in short order to land on the KEV list through a DNS-rebinding SSRF. We wrote up the same failure class in Ray's dashboard, where an internal-only API was reachable through the same rebinding trick. The pattern is not a coincidence.
ML platforms are built to be convenient inside a "trusted" network. They ship features that reach out to other services: webhooks, artifact stores, remote tracking, plugin callbacks. Operators run them on a subnet they think of as internal, often with authentication turned off because "only the data team can reach it." Both assumptions fail the moment one unauthenticated endpoint can be pointed at the metadata service. The network is not a trust boundary when the application will happily make requests on an attacker's behalf.
The specific defect here, validate-the-name-then-connect-to-whatever-it-resolves-to, is architecturally broken against rebinding no matter which vendor writes it. MLflow's own fix is the correct one: a hardened HTTP adapter that inspects the far end of the socket it actually opened, the instant connect() returns and before a single byte of data is exchanged. If that peer is a private or metadata address, the request dies. It also turns off environment proxy handling so a proxy cannot route around the check. You cannot reproduce that logic from outside the app, which is exactly why the containment work has to happen at a different layer.
What to do now
Patch to MLflow 3.15.0. That closes the endpoint. But patching a fleet of tracking servers takes time, and check-time URL validation is defeated by rebinding by design, so do not stop at the upgrade. Move the defense to where the payoff lives:
-
Enforce IMDSv2 and drop IMDSv1. On AWS, require the token-based metadata service and set the hop limit to 1. A rebinding SSRF from an application cannot complete the IMDSv2 token handshake the way it can hit the old unauthenticated endpoint, so even a successful SSRF returns nothing useful. This is the one control that holds while you are still patching, and it survives the next SSRF too.
-
Deny egress from the MLflow host to the metadata IP. Block
169.254.169.254at the host firewall or with an egress policy. The tracking server has no legitimate reason to talk to its own metadata endpoint. -
Take the tracking server off any network an attacker can reach. MLflow's tracking server has no authentication by default. If it is exposed beyond the team that needs it, the unauthenticated test endpoint is reachable by anyone who finds it.
-
Scope the instance role down. The credentials this bug exposes are only as valuable as the role behind them. A tracking server that can read one bucket is a smaller loss than one that can assume broad account permissions.
How you would know you were hit
Detection here is not a single log line, it is a correlation. Watch for two things and connect them.
On the MLflow side, look at access logs for calls to /api/2.0/mlflow/webhooks, especially the /test action, and for any webhook whose target hostname resolves to a private, loopback, or link-local address. A webhook configured to reach an internal or metadata range is not a normal operator action. Freshly created webhooks followed immediately by a test call, from a source that is not your CI or your data team, is the shape of an attempt.
On the cloud side, the real evidence is the credential being used somewhere else. If your instance role suddenly makes API calls from an IP or a service that is not the instance, or performs actions the tracking server never does, treat the role as compromised. Turn on and review your provider's metadata-access and role-usage logs. Because a full-read SSRF hands over live credentials, the window between exposure and abuse can be minutes, so the correlation has to be an alert, not a monthly report.
The takeaway
The uncomfortable detail in this one is the test button. A feature added to make webhooks easier to set up is the exact primitive that reads your cloud keys, because it is unauthenticated and it returns the response body. Convenience endpoints that echo a remote response back to the caller deserve the same scrutiny as any credential store. Patch to 3.15.0, enforce IMDSv2 today, and assume the next self-hosted ML tool you deploy has a reach-out feature waiting to be pointed somewhere it should not go.