Verified Instruction-Level Energy Consumption Measurement for NVIDIA GPUs

Yehia Arafa, Ammar ElWazir, Abdelrahman ElKanishy, Youssef Aly, Ayatelrahman Elsayed, Abdel-Hameed Badawy, Gopinath Chennupati, Stephan Eidenbenz, Nandakishore Santhi

I Introduction

Applications that rely on graphics processor units (GPUs) have increased exponentially over the last decade. GPUs are now used in various fields, from accelerating scientific computing applications to performing fast searches in data-oriented applications. A typical GPU has multiple streaming multiprocessors (SMs). Each can be seen as standalone processors operating concurrently. These SMs are capable of running thousands of threads in parallel. Over the last decade, GPUs’ microarchitecture has evolved to be very complicated. However, the increase in complexity means more processing power. Hence, the recent development of embedded/integrated GPUs and their application in edge/mobile computation have made power and energy consumption a primary metric for evaluating GPUs performance. Especially that researchers have shown that large power consumption has a significant effect on the reliability of the GPUs . Hence, analyzing and predicting the power usage of the GPUs’ hardware components remains an active area of research for many years.

Several monitoring systems (hardware & software) have been proposed in the literature to measure the total power usage of GPUs. However, measuring the energy consumption of the GPUs’ internal hardware components is particularly challenging as the percentage of updates in the microarchitecture can be significant from one GPU generation/architecture to another. Moreover, GPU vendors never publish the data on the actual energy cost of their GPUs’ microarchitecture.

In this paper, we provide an accurate measurement of the energy consumption of almost all the instructions that can execute in modern NVIDIA GPUs. Since the optimizations provided by the CUDA (NVCC) compiler can affect the latency of each instruction . We show the effect of the CUDA compiler’s high-level optimizations on the energy consumption of each instruction. We compute the instructions energy at the PTX granularity, which is independent of the underlying hardware. Thus, the measurement methodology introduced has minimum overhead and is portable across different architectures/generations.

To compute the energy consumption, we use three different software techniques based on the NVIDIA Management Library (NVML) , which query the onboard sensors and read the power usage of the device. We implement two methods using the native NVML API, which we call Sampling Monitoring Approach (SMA), and Multi-Threaded Synchronized Monitoring (MTSM). The third technique uses the newly released CUDA component in the PAPI v.5.7.1 API. Furthermore, we designed a hardware system to measure the power usage of the GPUs in real-time. The hardware measurement is considered as the ground truth to verify the different software measurement techniques.

To the best of our knowledge, we are the first to provide a comprehensive comparison of the energy consumption of each PTX instruction in modern high-end NVIDIA GPGPUs. Furthermore, the compiler optimizations effect on the energy consumption of each instruction has not been explored before in the literature. Also, we are the first to provide an in-depth comparison between different NVML power monitoring software techniques.

In summary, the followings are this paper contributions:

Accurate measurement of the energy consumption of almost all PTX instructions for four high-end NVIDIA GPUs from four different generations (Maxwell, Pascal, Volta, and Turing).

Show the effect of CUDA compiler optimizations levels on the energy consumption of each instruction.

Utilize and Compare three different software techniques (SMA, MTSM, and PAPI) to measure GPU kernels’ energy consumption.

Verify the different software techniques against a custom in-house hardware power measurement on the Volta TITAN V GPU.

The results show that Volta TITAN V GPU has the best energy efficiency among all the other generations for different categories of the instructions. Furthermore, our verification show that MTSM leads to the best results since it integrates the power readings and captures the start and the end of the GPU kernel correctly.

The rest of this paper is organized as follows: Section II provide a brief background on NVIDIA GPUs’ internal architecture. Section III describes the methodology of calculating the instructions energy consumption. Section IV depicts the differences between the software techniques. Section V shows the in-house direct hardware design. In Section VI we present the results. Section VII shows the related work and finally, Section VIII concludes the paper.

II GPGPUs Architecture

GPUs consist of a large number of processors called Streaming Multiprocessor (SMX) in CUDA terminology. These processors are mainly responsible for the computation part. They have several scalar cores, which has some computational resources, including fully pipelined integer Arithmetic Units (ALUs) for performing 32-bit integer instruction, Floating-Point units (FPU32) for performing floating-point operations, and Double-Precision Units (DPU) for 64-bit computations. Furthermore, it includes Special Function Units (SFU) that executes intrinsic instructions, and Load and Store units (LD/ST) for calculations of source and destination memory addresses. In addition to the computational resources, each SMX is coupled with a certain number of warp schedulers, instruction dispatch units, instruction buffer(s), and texture and shared memory units. Each SMX has a private L1 memory, and they all share access to L2 cache memory. The exact number of SMXs on each GPU varies with the GPU’s generation and the computational capabilities.

GPU applications typically consist of one or more kernels that can run on the device. All threads from the same kernel are grouped into a grid. The grid is made up of many blocks; each is composed of groups of 32 threads called warps. Grids and blocks represent a logical view of the thread hierarchy of a CUDA kernel. Warps execute instructions in a SIMD manner, meaning that all threads from the same warp execute the same instruction at any given time.

III Instructions Energy Consumption

We designed special micro-benchmarks to stress the GPU to be able to capture the power usage of each instruction.

We used Parallel-Thread Execution (PTX) to write the micro-benchmarks. PTX is a virtual-assembly language used in NVIDIA’s CUDA programming environment. PTX provides an open-source machine-independent ISA. The PTX ISA itself does not run on the device but rather gets translated to another machine-dependent ISA named Source And Assembly (SASS). SASS is not open. NVIDIA does not allow writing native SASS instructions, unlike PTX, which provides a stable programming model for developers. There have been some research efforts to produce assembly tool-chains by reverse engineering and disassembling the SASS format to achieve better performance. Reading the SASS instructions can be done using CUDA binary utilities (cuobjdump) . The use of PTX helps control the exact sequence of instructions executing without any overhead. Since PTX is a machine-independent, the code is portable across different CUDA runtimes and GPUs.

Figure 1 shows the compilation workflow, which leverages the compilation trajectory of the NVCC compiler. Since the PTX can only contain the code which gets executed on the device (GPU), we pass the instrumented PTX device code to the NVCC compiler for linking at runtime with the host (CPU) CUDA C/C++ code. PTX optimizing assembler (ptxas) is first used to transform the instrumented machine-independent PTX code to a machine-dependent (SASS) instructions then to a CUDA binary file (.cubin). The binary is used to produce a fatbinary file, which gets embedded in the host C/C++ code. An empty kernel gets initialized in the host code, which is then gets replaced by the instrumented PTX kernel, which has the same header and the same name inside the (.fatbin.c). The kernel is executed with one block one thread.

Figure 2 shows an example of the instrumented PTX kernel for the unsigned Div instruction. In our previous work , we presented a similar technique to find the instruction latency. We executed the instruction only once, and red the clk register before and after its execution. The design here is different since we need to capture the change in power usage, which would be unnoticeable if we execute the instruction only once. The key idea here is unrolling a loop and execute the same instruction millions of times and record the power then divide by the number of instructions to get the power consumption of the single instruction. The kernel in Figure 2 shows an example of the micro-benchmark of the unsigned div instruction. We begin by initializing the used registers, lines . Since PTX is a virtual-assembly and gets translated to the SASS, there is no limit on the number of registers to use. Still, in the real SASS assembly, the number of registers is limited and will vary from one generation/architecture to another. When the limit exceeds, register variables will be spilled to memory, causing changes in performance. Line sets the loop count to 1M iterations. The loop body, lines , is composed of 5 back-to-back unsigned div instructions with dependencies, to make sure that the compiler does not optimize any of them. We do a load-add-store operation on the output of the \nth5 div operation and begin the loop with new values each time to force the compiler to execute the instructions. Otherwise, the compiler would run the loop only the first time and squeeze the remaining iterations. We follow the same approach for all the instructions, and the kernel is the same, the only difference is the instruction itself.

GPUs drain power as static power and dynamic power. The static power is a constant power that the GPU consumes to maintain its operation. However, dynamic power is affected by the kernel’s instructions and operations. To eliminate the static power and any overhead dynamic power other than the instruction power consumption, we measure the power and compute the kernel’s energy consumption twice. First, we run the kernel as shown in figure 2, we call that the total energy. Second, while commenting out the back-to-back instructions (lines ), we call that the overhead energy. We then use Eq. 1 to calculate the energy of the instruction. This way, only the real energy of the instruction is calculated.

The kernel is compiled with (–O3) and (–O0) optimization flags. This way, we capture the effect of the CUDA compiler’s higher levels of optimizations on the energy consumption of each PTX instruction. To make sure that in case of (–O3), the compiler does not optimize the instructions and squeeze them, we made sure that the output of the kernel is correct. Line 28 of Figure 2, stores the output of the loop. We read it and validate its correctness. Furthermore, we validate the clk register for each instruction against our previous work .

IV Software Measurement

NVIDIA provides an API named NVIDIA Management Library (NVML) , which offers direct access to the queries exposed via the command line utility, NVIDIA System Management Interface (nvidia-smi). NVML allows developers to query GPU device states such as GPU utilization, clock rates, GPU temperature etc. Additionally, it provides access to the board power draw by querying its instantaneous onboard sensors. The community has widely used NVML since its first release with CUDA v4.1 in 2011. NVIDIA display driver is equipped with NVML, and the SDK offers the API for its use. We use NVML to read the device power usage while running the PTX micro-benchmarks and compute the energy of each instruction. There are several techniques for collecting power usage using NVML. We found that the methods do vary. Therefore, we provide an in-depth comparison of the quality of these techniques on the energy of the individual instructions.

The C-based API provided by NVML can query the power usage of the device and provide an instantaneous power measurement. Therefore, it can be programmed to keep reading the hardware sensor with a certain frequency. This basic approach is popular and was used in other related works . The nvmlDeviceGetPowerUsage() function is used to retrieve the power usage reading for the device, in milliwatts. This function is called and executed by the CPU. We configured the sampling frequency of reading the hardware sensors to its maximum, 66.7 Hz (15 ms window between each call to the function).

We read the power sensor according to the sample interval in the background while the micro-benchmarks are running. Example of the output using this approach are shown in Figures 3(a) and 3(b). The two figures show the power consumption over time for integer Add and unsigned integer Div kernels for the TITAN V (Volta) GPU. The power usage jumps shortly after the launch of the kernel and decreases in steps after the kernel finishes execution until it reaches the steady-state. This is done in 22 sec and 33 sec windows interval for Add and Div respectively. If we calculate the two kernels actual elapsed time, it takes only 0.28 sec and 13 sec for the Add and the Div kernels, respectively. That is, the GPU does something before and after the actual kernel execution. Hence, identifying the window of the kernel is hard and would affect the output as the power consumption varies through time. One solution is to take the maximum reading between the two steady states, but this would be misleading for some kernels, especially the bigger ones. Therefore, we ignore this approach from reporting owing to these issues.

IV-B PAPI API

Performance Application Programming Interface (PAPI) provides an API to access the hardware performance counters found on modern processors. We can read different performance metrics through either a simple programming interface from either C or Fortran programming languages. Researchers have used PAPI as a performance and power monitoring library for different hardware and software components . It is also used as a middleware component in different profiling and tracing tools .

PAPI can work as a high-level wrapper for different components; for example, it uses the Intel RAPL interface to report the power usage and energy consumption for Intel CPUs. Recently, PAPI version 5.7 added the NVML component, which supports both measuring and capping power usage on modern NVIDIA GPU architectures.

The advantage of using PAPI is that the measurements are by default synchronized with the kernel execution. The target kernel is invoked between the papi_start, and the papi_end functions, and a single number, representing the power event we need to measure is returned. The NVML component implemented in PAPI uses the function, getPowerUsage() which query nvmlDeviceGetPowerUsage() function. According to the documentation, this function is called only once when the papi_end is called. Thus, the power returned using this method is an instantaneous power when the kernel finishes execution. Although synchronizing with the kernel solves the SMA issues, taking the instantaneous measurement when the kernel finishes execution can provide non-accurate results especially, for large and irregular kernels as shown in Section VI. Note that PAPI provides an example that works like the SMA approach, which we refrain from this paper.

IV-C Multi-Threaded Synchronized Monitoring (MTSM)

In MTSM, we identify the exact window of the kernel execution. We modified SMA to synchronize the kernel execution. This way, only the power readings of the kernel are recorded. Since the host CPU monitors the NVML API, we use Pthreads for synchronization where one thread calls and monitors the kernel while the other thread records the power.

Algorithm 1 shows the MTSM. We initialize a volatile atomic variable (flag) to zero, which we use later to record the power readings according to the start and end of the target kernel. On line 6 we create a new thread (th1) which executes a function (func1) [line 17] in parallel. This function completes the power monitoring, depending on the atomic flag. This uses the NVML function, nvmlDeviceGetPowerUsage() which returns the device power in milli-watts. The readings of the power during the kernel window are recorded and saved in an array (power_readings), which is used later in computing the kernel energy. In lines , flip the flag value and start computing the elapsed time and the launch kernel, which means starting the power monitoring. At the end of the kernel execution, we record the elapsed time and change the flag. We use the CUDA synchronize function to make sure that the power is recorded correctly. We do not specify any reading sampling frequency for the NVML functions. Although this would give us redundant values, it would be more accurate. With this setup, we found that the power reading frequency is nearly 2kHz2kHz.

Figures 4(a) and 4(b) show the corresponding kernels in Figures 3(a) and 3(b) after identifying the exact kernel execution window. The new graphs are annotated with the start and end of the kernel. We observe that the kernel does not start after the sudden rise in the power from the steady-state, rather after a couple of ms from this sudden increase in power consumption (see add kernel in Figure 4(a) for clarity). After the kernel finishes execution, the power remains high for a small-time, and then it starts descending in steps until it reaches the steady-state again. To compute the kernel’s energy, we calculated the area under the curve for the kernel using Eq. 2. We believe that this approach would provide the most accurate measurement since the power readings of only the kernel are recorded. Computing the energy as the area under the curve is more rigorous than just taking the last power reading multiplied by the time elapsed for the kernel, as is done in PAPI.

We configured MTSM as a shared library that can be linked with the application binaries at runtime. The code is first compiled and then injected or preloaded at runtime using LD_PRELOAD environment variable to any device executable binary file with a kernel that executing on NVIDIA GPUs. The start timing and end timing are automatically triggered by intercepting the CUDA runtime API calls.

V Hardware Measurement

Modern GPUs have two primary sources of power. The first power source is the direct DC power (12 V12~{}V) supply, provided through the side of the card. While the second one is the PCI-E (3.3 V3.3~{}V and 12V12V) power source, provided through the motherboard. We have designed a system to measure each power source in real-time. The hardware measurement is considered as the ground truth to verify the different software measurement techniques.

Figure 5 shows the experimental hardware setup with all the components. To capture the total power, we measure the current and voltage for each power source simultaneously. A clamp meter and a shunt series resistor are used for the current measurement. For voltage measurement, we use a direct probe on the voltage line using an oscilloscope to acquire the signals. Equation 3 is used to calculate the total hardware power drained by the GPU from the two different power sources.

Direct DC Power Supply Source: Power supply provides a 12 V12~{}V voltage through a direct wired link. We use both a 6-pin and 8-pin PCI-E power connectors to deliver a maximum of 300 W300~{}W. Thus, the direct DC power supply source is the main contributor to the card’s power. Figure 5 shows a clamp meter measuring the current of the direct power supply connection. The voltage of the power supply is measured using an oscilloscope probe. The current and voltage are acquired using an oscilloscope, as shown in Figure 5. Therefore, the Direct DC power supply source is calculated using simple multiplication. The third addition term in Eq. 3 shows the calculation of the power which is multiplying Iclamp{\textnormal{{I}}_{clamp}} by VDPS{\textnormal{{V}}_{DPS}}. In which, VDPS{\textnormal{{V}}_{DPS}} is the voltage of the direct power supply.

PCI-E Power Source: Graphics cards are connected to the motherboard through the PCI-E x16 slot connection. 3.3V3.3V and 12V12V voltages are provided through this slot. To accurately measure the power that goes through this slot, an intermediate power sensing technique should be installed between the card and the motherboard. We designed a custom made PCI-E riser board that measures the power supplied through the motherboard. Two in-series shunt resistors are used as a power sensing technique. As shown in Figure 6, each shunt resistor (RS{\textnormal{{R}}_{S}}) is connected in series with 3.3V3.3V and 12V12V separately. Using the series property, the current that flows through the RS{\textnormal{{R}}_{S}} is the same current that goes to the graphics card. Therefore, we measure the voltages VS1{\textnormal{{V}}_{S1}} and VG1{\textnormal{{V}}_{G1}} which are across RS{\textnormal{{R}}_{S}} using oscilloscope. We then divide it with the RS{\textnormal{{R}}_{S}} value. The voltage level is measured using the riser board. We duplicate the same calculation technique for the 3.3V3.3V voltage level, as shown in Eq. 3.

VI Results

We show the energy consumption of each instruction found in the latest PTX ISA, v.6.4 . We report the results of using MTSM and PAPI on four different NVIDIA GPUs from four different generations/architectures; GTX TITAN X: GPU from Maxwell architecture. It has 3584 cores with 151 MHz clock frequency. GTX 1080 Ti: GPU from Pascal architecture. It has 3584 cores with 1481 MHz clock frequency. TITAN V: GPU from Volta architecture. It has 5120 cores with 1200 MHz clock frequency. TITAN RTX: GPU from Turing architecture . It has 4608 cores with 1350 MHz clock frequency.

We used CUDA NVCC compiler v.10.1 to compile and run the codes. CUDA compiler comes equipped with NVML library . Table I shows an enumeration of the energy consumption of the various ALU instructions for the different GPUs. For simplicity, we used each GPU generation to refer it. We denote the (O3) version as Optimized and the (O0) version as Non-Optimized.

The results show that overall Volta GPUs have the lowest energy consumption per instruction among all the tested GPUs. Pascal preceded the Volta while Maxwell and Turing are power hungry devices except for some categories of the instructions.

For Half Precision (FP16) instructions, Volta and Turing have much better results than Pascal. Hence, this confirms that both architectures are suitable for approximate computing applications (e.g. , deep learning, and energy-saving computing). We did not run FP16 instructions on Maxwell as Pascal architecture was the first GPU that offered FP16 support. The same trend can be found in Multi Precision (MP) instructions where Volta and Pascal have better energy consumption compared to the two other generations. MP instructions are essential in a wide variety of algorithms in computational mathematics (i.e. , number theory, random matrix problems, experimental mathematics). Also, it is used in cryptography algorithms and security.

Overall, the energy of Non-Optimized is always more than the Optimized. One reason is that the number of cycles at the (O0) optimization level are more than the (O3) level . This can be because the translation from PTX instruction to native SASS instruction is not one-to-one conversion. Thus, the instruction can take more time to finish execution if it got translated to more than one instruction.

PAPI vs. MTSM: The dominant tendency of the results is that PAPI readings are always more than the MTSM. Although the difference is not significant for small kernels, it can be up to 1 μ\muJ for bigger kernels like Floating Single and Double Precision div instructions.

We verified the different software techniques (MTSM & PAPI) against the hardware setup on Volta TITAN V GPU. Compared to the ground truth hardware measurements, for all the instructions, the average Mean Absolute Percentage Error (MAPE) of MTSM Energy is 6.39 and the mean Root Mean Square Error (RMSE) is 3.97. In contrast, PAPI average MAPE is 10.24 and the average RMSE is 5.04. Figure 7 shows the error of MTSM and PAPI relative to the hardware measurement for some of the instructions. The results prove that MTSM is more accurate than PAPI as it is closer to what has been measured using the hardware.

VII Related Work

Several works in the literature tried to directly measure the instantaneous power usage of the GPUs using various profiling techniques. On the other hand, Researchers have proposed different techniques to indirectly estimate and predict the total GPU’s power/energy consumption. Additional details are discussed by Bridges et al. .

GPU power profiling can be carried out in two different approaches, a software-oriented solution, where the internal power sensors are queried using NVML, and hardware-oriented solutions using external hardware setups.

Software-oriented approaches: Arunkumar et al. used a direct NVML sampling motoring approach running in the background while using a special micro-benchmark to calculate basic compute/memory instructions energy consumption and feed that to their model. They run their evaluation on (Tesla K40) Kepler GPU. They intentionally disabled all compiler optimizations and compiled their micro-benchmarks with (–O3) flag. Burtscher et al. analyzed the power consumption measured by NVML for (Tesla K20) GPU. Kasichayanula et al. used NVML to calculate the energy consumption of some GPU units which drive their model and validate it with a Kill-A-Watt power meter. While these types of hardware power meters are cheap and straightforward to use, they do not give an accurate measurement, especially in HPC settings.

Hardware-oriented approaches: Zhao et al. used an external power meter on an old GPU (GeForce GTX 470) from Fermi architecture, where they designed a micro-benchmark to compute the energy of some PTX instructions and feed that into their model. The authors of validate their roofline model by using PowerMon 2 and a custom PCIe inter-poser to calculate the instantaneous power of (GTX 580) GPU.

Recently, Sen et al. assessed the quality and performance of the power profiling mechanisms using hardware and software techniques. They compared a hardware approach using PowerInsight (a hardware power instrumentation product) to the software NVML approach on a developed matrix multiplication CUDA benchmark.

In a similar spirit, we follow the same line of research. Nevertheless, we focus on the energy consumption of individual instructions while having a detailed comparison of the different software/hardware approaches.

VIII Conclusion & Future Directions

In this paper, we accurately measure the energy consumption of various PTX instructions that execute on NVIDIA GPUs. We also show the effects of different optimization levels of the CUDA (NVCC) compiler on energy consumption of each instruction. We provide an in-depth comparison of various software techniques that query the onboard internal GPU sensors and verify against an in-house custom-designed hardware power measurement. Overall, the paper provides an easy and straightforward way (Multi-Threaded Synchronized Monitoring (MTSM)) that can be used to measure the energy consumption of any NVIDIA GPU kernelThe source code is available on our laboratory page on Github at https://github.com/NMSU-PEARL/GPUs-Energy.. Furthermore, the results give GPU architects and developers a concrete understanding of NVIDIA GPUs’ microarchitecture. This work will help GPU modeling frameworks to have a precise prediction of energy/power consumption of GPUs. Along with GPU/CPU memory and pipeline models, a heterogeneous system can be accurately modeled .

Appendix: Energy Consumption Results

Table I has per-instruction energy breakdown for different generations of NVIDIA GPUs.

References