Skip to content

Read hotspot, VRAM and voltage on NVIDIA through NVAPI - #2128

Open
GribanovIvan wants to merge 3 commits into
flightlessmango:masterfrom
GribanovIvan:feat/nvapi-sensors
Open

GribanovIvan wants to merge 3 commits into
flightlessmango:masterfrom
GribanovIvan:feat/nvapi-sensors

Conversation

@GribanovIvan

Copy link
Copy Markdown

gpu_junction_temp and gpu_mem_temp have always read 0 on GeForce cards and gpu_voltage was never populated for NVIDIA at all, because the backend reads temperatures through NVML and NVML exposes a single sensor there. The sensors exist; NVIDIA's own libnvidia-api.so.1 returns them. This reads them from there.

Closes the NVIDIA half of #760 and #1159.

Tested on

GPU Architecture Driver hotspot VRAM voltage
RTX 3050 Laptop (GA107) Ampere 595.71.05 yes yes yes

Linux 6.18, KDE Wayland, hybrid laptop alongside a Radeon 680M.

The hotspot reading was also confirmed working on a GTX 1060 by another user. I have no details of that system, so I am reporting it only as a second data point on much older hardware, not as something I measured.

Why NVML cannot do it

Probed directly on the 3050:

nvmlDeviceGetTemperature(dev, sensorType=0)    ->  OK, 38 C   (core)
nvmlDeviceGetTemperature(dev, sensorType=1..3) ->  error
nvmlDeviceGetThermalSettings(dev, 0..3)        ->  one sensor: GPU_INTERNAL, target=GPU
nvmlDeviceGetFieldValues(dev, fields 150..240) ->  nothing thermal beyond the core

Nothing else on the system reads them either: nvidia-smi reports only the core, and nvtop exposes a single temp field, both for the same reason.

Implementation

src/nvapi_sensors.{h,cpp} dlopens the library, resolves entry points by id through nvapi_QueryInterface, matches the GPU by PCI bus, and reads thermal sensor 9 (hotspot) and 16 (VRAM) as fixed point with 8 fractional bits, plus voltage rail 0 in microvolts. No elevated privileges, no BAR0 mapping - the driver performs the reads. If the library is missing the module disables itself and nothing else changes. Blackwell reads its hotspot through a different interface, so it is left to report nothing there rather than a wrong number.

Two things in the wiring are load-bearing. Values are seeded in the constructor before the sampling thread starts, or the rows appear one period late and shift the layout. And get_samples_and_copy() reduces metrics by an explicit list of fields, which did not include these three - without adding them the sampled values never reach the HUD at all.

New option

gpu_voltage_pci_dev=0000:01:00.0 restricts the voltage readout to one card. Every GPU reports a voltage, an APU's SoC rail included, and on a hybrid system the second is noise. Unset keeps current behaviour. Uses the same verify_pci_dev() normalisation as pci_dev. The example config line claiming gpu_voltage is AMD-only is removed, since it stops being true here.

Verification

Hotspot was cross-checked against a BAR0 MMIO reader built from the register offsets gddr6-core-junction-vram-temps uses, sampled alongside NVAPI:

NVAPI 49.78   MMIO 49   core (NVML) 43
NVAPI 49.69   MMIO 49   core (NVML) 43
NVAPI 49.66   MMIO 49   core (NVML) 43

Two independent paths agree, both differ from the core, and the MMIO value being the truncated integer part also confirms the fixed-point encoding. Under load the junction sits 12-14 C above the core and tracks it.

VRAM took more work. The thermals call returns 40 unlabelled slots, nine non-zero here, so the memory sensor has to be identified rather than assumed - LACT maps slot 15, which turns out to track the die on this part. Comparing a shader-bound load against a memory-bound one at matched core temperature (70-74 C):

sensor 8    +0.16 C        sensor 14   +0.15 C
sensor 9    -0.07 C        sensor 15   -0.07 C
sensor 13   +0.15 C        sensor 16   +1.24 C

Every die sensor is locked to the core whatever the GPU is doing. Only 16 has an independent heat source, and it reports whole degrees where the die sensors report 1/256 C.

Voltage: 631 mV idle, 1081 mV loaded.

GribanovIvan and others added 2 commits August 28, 2026 22:48
NVML exposes a single temperature sensor on GeForce cards, so gpu_junction_temp
and gpu_mem_temp have always read zero there, and gpu_voltage was never filled
in at all. The sensors do exist - NVIDIA's own libnvidia-api.so.1 returns them.

Add a small module that dlopens that library and reads the junction temperature,
the memory temperature and the core voltage. No elevated privileges and no BAR0
access are involved: the driver performs the reads. If the library is missing the
module quietly disables itself and everything else keeps working as before.

Values are seeded once in the constructor, before the sampling thread starts, so
the first frame already has them; otherwise the rows appear one update later and
shift the layout under them. The metric reduction in get_samples_and_copy() is an
explicit list, so junction_temp, memory_temp and voltage are added to it as well
- without that the sampled values never reach the published struct.

The example config said gpu_voltage only works on AMD GPUs, which was true while
the NVIDIA backend never populated the field; that note goes with this change.

Verified on an RTX 3050 Laptop (GA107, Ampere, GDDR6), driver 595.71.05: the
junction reading tracks load and sits 12-14 C above the core, matching an
independent BAR0 MMIO reader to the degree.

Blackwell reports the hotspot through a different path and is deliberately not
handled - it reports nothing there rather than a wrong number.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every GPU reports a voltage, including an APU's SoC rail, but on a hybrid system
only one of them is usually of interest and the other is noise next to it.

Name a device here and the voltage is shown only for it; leaving it unset keeps
the current behaviour of showing all of them. The address goes through the same
verify_pci_dev() normalisation as the existing pci_dev option.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings August 28, 2026 20:35

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

There are a few confirmed correctness issues (undefined behavior in bit shifting, invalid voltage rendering, and potentially stale NVAPI values when NVML isn’t active) that should be fixed before merging.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

Adds an NVAPI-based sensor path for NVIDIA GPUs to expose hotspot (junction), VRAM temperature, and voltage readings that NVML doesn’t provide on GeForce, and wires the new metrics through sampling/reduction into the HUD and logging. It also introduces a config option to show voltage only for a specific GPU on multi-GPU/hybrid systems.

Changes:

  • Add NvApiSensors (dlopen libnvidia-api.so.1 + nvapi_QueryInterface) to read hotspot/VRAM temps and voltage on NVIDIA.
  • Wire new NVIDIA sensor readings into the metrics sampling path and aggregation so they reach the HUD/logger.
  • Add gpu_voltage_pci_dev option and document it in README + sample config.
File summaries
File Description
src/overlay_params.h Adds gpu_voltage_pci_dev overlay parameter to the config schema.
src/overlay_params.cpp Parses and normalizes gpu_voltage_pci_dev via verify_pci_dev().
src/nvidia.h Adds NvApiSensors member to NVIDIA backend.
src/nvidia.cpp Initializes/seeds NVAPI-backed junction/VRAM/voltage metrics and aggregates them in sampling.
src/nvapi_sensors.h Declares NVAPI sensor reader for hotspot/VRAM temps and voltage.
src/nvapi_sensors.cpp Implements NVAPI dynamic loading, GPU matching, thermal reads, and voltage reads.
src/meson.build Adds nvapi_sensors.cpp to the build.
src/hud_elements.cpp Adds optional voltage display filtering by PCI device.
README.md Documents gpu_voltage_pci_dev option.
data/MangoHud.conf Updates sample config comments and adds gpu_voltage_pci_dev example.
Review details
  • Files reviewed: 10/10 changed files
  • Comments generated: 4
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/hud_elements.cpp
Comment thread src/nvapi_sensors.h Outdated
Comment thread src/nvapi_sensors.cpp Outdated
Comment thread src/nvidia.cpp
Sample NVAPI outside the NVML path. The reads were sitting inside
get_instant_metrics_nvml(), so on a system where only XNVCtrl is available the
thread never refreshed them and the values stayed frozen at whatever the
constructor seeded. NVAPI does not depend on NVML, so it now has its own
function that the sampling loop always calls.

Hide the voltage row when there is no reading. Cards whose voltage interface
does not resolve return -1, which was rendered as "-1 mV". Gate on >= 0 so an
absent sensor is hidden while a genuine 0 from any other driver still shows as
it does today.

Use an unsigned shift when widening the sensor mask. 1 << 31 on a signed int is
undefined, and so is the (1 << bit) - 1 that follows it.

Correct the header comment: read_sensor() divides, so the value is truncated
rather than rounded.
@GribanovIvan

Copy link
Copy Markdown
Author

1060 was tested with 580.105.08 driver

@detiam

detiam commented Sep 28, 2026

Copy link
Copy Markdown

I had a GDDR6X Nvidia GPU (4070 Ti SUPER), and this PR couldn't read gpu_mem_temp of my GPU, so I searched another project that can read my GPU memory temperature, LACT, and found out you need VRAM_INDEX of 15 to read my GDDR6X memory temperature, 16 gives nothing on my GPU.

So it seems we need different VRAM_INDEX for GDDR6, GDDR6X, and GDDR7.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants