Exploit Allows Unauthenticated Attackers to Crash NVIDIA GPU Monitoring
A critical flaw in NVIDIA’s DCGM Exporter (CVE-2026-47483) has been identified, allowing unauthorized users to cause the GPU monitoring service to fail, potentially impacting AI and machine learning operations.
The vulnerability was disclosed by Lava
The company assigned a CVSS score of 8.2 and issued a security advisory on July 28, 2026. GPU servers, designed for AI, machine learning, and scientific computing, rely on DCGM Exporter to collect telemetry data from connected GPUs. This tool gathers metrics such as temperature, utilization, memory consumption, power usage, and error logs, along with unique identifiers like GPU model and UUID.
Researchers from Lava noted that this data could reveal critical information
Researchers from Lava noted that this data could reveal critical information about a server’s hardware configuration, workload intensity, and operational health. A team led by Lava researcher Michael Katchinskiy discovered over 2,000 internet-facing DCGM Exporter instances during scans between March and May 2026. These servers represented more than 12,000 unique GPUs, valued at an estimated $100 million.
All exposed endpoints transmitted metrics via unencrypted HTTP
All exposed endpoints transmitted metrics via unencrypted HTTP without requiring authentication. The affected hardware included high-performance NVIDIA GPUs like the Blackwell Ultra B300, H200, and H100, as well as consumer-grade RTX 5090 and 4090 models. The exposed systems spanned 300 organizations, with the United States hosting 5,274 GPUs (44% of the total), followed by Romania (2,054) and China (1,967).
Some instances were hosted on cloud infrastructure managed by providers
Some instances were hosted on cloud infrastructure managed by providers such as Voltage Park, Lambda, Northern Data, and DigitalOcean. Voltage Park accounted for the largest cluster, with 672 Node Exporter hosts and 71 DCGM Exporter hosts. The company confirmed that these deployments were managed by third-party customers and initiated contact with affected parties.
Exploitation vector: Memory exhaustion via profiling endpoints
Approximately 25% of the exposed DCGM Exporter instances also hosted Go’s /debug/pprof/ profiling endpoints. These endpoints, when subjected to a high volume of unauthenticated requests, could consume excessive memory, leading to service crashes. Katchinskiy’s team replicated the issue using NVIDIA’s official DCGM Exporter container without modifications, confirming that the vulnerability stemmed from the default configuration.
Node Exporter reveals broader system details
Node Exporter reveals broader system details In parallel analyses, researchers examined Prometheus Node Exporter, which collects server-level metrics. They identified 12,096 public Node Exporter hosts transmitting data from NVIDIA and Mellanox InfiniBand/RoCE adapters. This data included adapter models, firmware versions, link states, and fabric activity.
Mitigation strategies and recommendations
NVIDIA advises users to upgrade to DCGM Exporter version 4.8.2 or later and disable the enable-pprof flag unless profiling is explicitly required. Modern versions of the tool require manual activation of the profiling endpoint. Katchinskiy emphasized that exposing monitoring services like Node Exporter, DCGM Exporter, and Prometheus to the public internet should be avoided unless strictly necessary.
Best practices include binding exporters to loopback or private interfaces, implementing firewall rules or security groups to restrict access, and securing Prometheus query APIs and target pages. Katchinskiy stressed that publicly accessible monitoring systems can expose far more than basic telemetry, including critical infrastructure details.
Conclusion
As AI infrastructure expands, organizations must ensure visibility into their responsibilities versus those of third-party providers. The findings underscore the importance of securing monitoring tools, which often serve as gateways to sensitive operational data. By addressing misconfigurations and limiting exposure, enterprises can reduce the risk of exploitation in increasingly complex AI environments.
