GPU Dashboard
The GPU dashboard is an AI infrastructure operations dashboard that provides an integrated view of GPU resource composition, utilization efficiency, and issue history on a single screen. When GPU servers grow to many units, checking each server one by one is difficult. The GPU dashboard collects all GPU resources by model and by server and displays them, helping you quickly grasp the utilization status and anomalies.
Supported environment
For an overview of GPU monitoring in server environments, see the WhaTap GPU Monitoring document.
Basic screen guide
The GPU dashboard consists of two tabs.
-
GPU operation status: Collects and displays the GPU resource inventory, utilization efficiency, power consumption, and issues by widget.
-
GPU resource board: Provides an overview of the allocation and utilization status of all GPUs as a summary bar and map.
Option bar
You can adjust the query conditions in the option bar at the top of the screen.
-
Time selection: Specify the time range to query.
-
Filter: Specify filters to query only the GPUs that match conditions such as a specific model or server.
-
View toggle: Switch the display format of the widgets.
-
Preset: Save frequently used combinations of filters and Top 5 metric selections as a preset and load them again.
GPU operation status tab
The GPU operation status tab provides the following widgets.
GPU inventory (by model)
Aggregates the GPUs you have by model and displays them as a donut. When you click a model slice of the donut or a legend item, the GPU Inventory with the filter applied for that model opens in a new tab.
GPU power consumption trend
Displays the GPU power consumption trend over time.
The displayed power consumption is an estimate calculated based on GPU metrics and may differ from actual power meter measurements.
Server severity distribution by GPU model
Displays, for each GPU model, the ongoing event status and representative severity distribution of the servers equipped with that model, as a stacked bar. Occurred x / Total y servers indicates the number of servers, among all servers equipped with the GPU model, that currently have at least one ongoing event. Each server's severity is classified by the highest severity among its ongoing events, and the server counts are shown per Critical, Warning, Info, and Normal. Total n events indicates the total number of events currently ongoing across the servers where events occurred. When you click an item, you can view the list of ongoing events for the servers equipped with that GPU model in a drawer.
Top 5 Server / Top 5 GPU
Displays the top 5 servers (Top 5 Server) and the top 5 GPUs (Top 5 GPU) by a specified metric in separate panels. The default reference metric is GPU utilization for servers and power limit violation rate for GPUs. When you click an item, a drawer showing the recent trend of that target opens. You can change the reference metric in the option bar or a preset.
Xid
Displays the history of Xid events that occurred on GPUs. For how to detect anomalies using Xid events, see the WhaTap GPU Monitoring document.
GPU resource board tab
The GPU resource board tab provides an overview of the allocation and utilization status of all GPUs as a summary bar and map. You can see at a glance how much each GPU is allocated and used, and which GPUs are idle or close to running out of memory.
Summary bar
Summarizes the allocation and utilization status of all GPUs by card. Each card shows the number of physical GPUs (Physical) and MIG instances together.
| Item | Description |
|---|---|
| Total GPU | Total number of GPUs |
| Allocated GPU | Number of allocated GPUs and the allocation rate |
| Idle GPU | Number of idle GPUs that are allocated but not used |
| Active GPU | Number of GPUs in use |
| Effective GPU | Number of GPUs used efficiently |
The summary bar card names are displayed in English (Total GPU, Allocated GPU, Idle GPU, Active GPU, Effective GPU). The idle/active/effective classification is based on GPU utilization: less than 1% is Idle, 1% or more and less than 50% is Active, and 50% or more is Effective.
GPU Resource Map
A scatter plot that arranges all GPUs by utilization. It shows the distribution by dividing each GPU into the Idle, Active, and Effective ranges, and separately marks GPUs that are close to running out of memory due to high memory usage (FB usage 95% or more). When you select a specific area on the map, a GPU Timeline drawer opens where you can check the recent trend of those GPUs.
Utilization status by GPU model
Aggregates GPUs by model and shows the ratio of the Idle, Active, Effective, and Unallocated ranges. You can compare how the allocation and utilization status is distributed by model.
GPU utilization status by server
Displays the GPUs held by each server as a slot grid. Each slot represents one physical GPU, and a GPU split by MIG is displayed by dividing the inside of the slot into instances.
GPU utilization status by server is displayed automatically in environments with 20 or fewer GPU servers.