Skip to main content

GPU Dashboard

The GPU dashboard is an AI infrastructure operations dashboard that provides an integrated view of GPU resource composition, utilization efficiency, and issue history on a single screen. When GPU servers grow to many units, checking each server one by one is difficult. The GPU dashboard collects all GPU resources by model and by server and displays them, helping you quickly grasp the utilization status and anomalies.

Note

Supported environment

For an overview of GPU monitoring in server environments, see the WhaTap GPU Monitoring document.

Basic screen guide

The GPU dashboard consists of two tabs.

  • GPU operation status: Collects and displays the GPU resource inventory, utilization efficiency, power consumption, and issues by widget.

  • GPU resource board: Provides an overview of the allocation and utilization status of all GPUs as a summary bar and map.

Option bar

You can adjust the query conditions in the option bar at the top of the screen.

  • Time selection: Specify the time range to query.

  • Filter: Specify filters to query only the GPUs that match conditions such as a specific model or server.

  • View toggle: Switch the display format of the widgets.

  • Preset: Save frequently used combinations of filters and Top 5 metric selections as a preset and load them again.

GPU operation status tab

The GPU operation status tab provides the following widgets.

GPU inventory (by model)

Aggregates the GPUs you have by model and displays them as a donut. When you click a model slice of the donut or a legend item, the GPU Inventory with the filter applied for that model opens in a new tab.

GPU power consumption trend

Displays the GPU power consumption trend over time.

Caution

The displayed power consumption is an estimate calculated based on GPU metrics and may differ from actual power meter measurements.

Server severity distribution by GPU model

Displays, for each GPU model, the ongoing event status and representative severity distribution of the servers equipped with that model, as a stacked bar. Occurred x / Total y servers indicates the number of servers, among all servers equipped with the GPU model, that currently have at least one ongoing event. Each server's severity is classified by the highest severity among its ongoing events, and the server counts are shown per Critical, Warning, Info, and Normal. Total n events indicates the total number of events currently ongoing across the servers where events occurred. When you click an item, you can view the list of ongoing events for the servers equipped with that GPU model in a drawer.

Top 5 Server / Top 5 GPU

Displays the top 5 servers (Top 5 Server) and the top 5 GPUs (Top 5 GPU) by a specified metric in separate panels. The default reference metric is GPU utilization for servers and power limit violation rate for GPUs. When you click an item, a drawer showing the recent trend of that target opens. You can change the reference metric in the option bar or a preset.

Xid

Displays the history of Xid events that occurred on GPUs. For how to detect anomalies using Xid events, see the WhaTap GPU Monitoring document.

GPU resource board tab

The GPU resource board tab provides an overview of the allocation and utilization status of all GPUs as a summary bar and map. You can see at a glance how much each GPU is allocated and used, and which GPUs are idle or close to running out of memory.

Summary bar

Summarizes the allocation and utilization status of all GPUs by card. Each card shows the number of physical GPUs (Physical) and MIG instances together.

ItemDescription
Total GPUTotal number of GPUs
Allocated GPUNumber of allocated GPUs and the allocation rate
Idle GPUNumber of idle GPUs that are allocated but not used
Active GPUNumber of GPUs in use
Effective GPUNumber of GPUs used efficiently

The summary bar card names are displayed in English (Total GPU, Allocated GPU, Idle GPU, Active GPU, Effective GPU). The idle/active/effective classification is based on GPU utilization: less than 1% is Idle, 1% or more and less than 50% is Active, and 50% or more is Effective.

GPU Resource Map

A scatter plot that arranges all GPUs by utilization. It shows the distribution by dividing each GPU into the Idle, Active, and Effective ranges, and separately marks GPUs that are close to running out of memory due to high memory usage (FB usage 95% or more). When you select a specific area on the map, a GPU Timeline drawer opens where you can check the recent trend of those GPUs.

Utilization status by GPU model

Aggregates GPUs by model and shows the ratio of the Idle, Active, Effective, and Unallocated ranges. You can compare how the allocation and utilization status is distributed by model.

GPU utilization status by server

Displays the GPUs held by each server as a slot grid. Each slot represents one physical GPU, and a GPU split by MIG is displayed by dividing the inside of the slot into instances.

Note

GPU utilization status by server is displayed automatically in environments with 20 or fewer GPU servers.