Job Profiling#
Job profiling measures the resource (CPU, GPU, memory) usage of code run within a job, helping identify performance issues and potential code improvements. Profiling helps increase job efficiency, which is important when working on an HPC cluster. Workflows that run calculations in parallel on multiple CPU cores or utilize specialized hardware such as GPUs are especially important to profile. Further, job profiling tools are useful in conducting scaling studies to increase your workflow’s speed and throughput.
Profiling your jobs helps you:
Right-size job resource requests, reducing queue wait times
Identify code bottlenecks in CPU, memory, or GPU usage
Avoid compute resources sitting idle which blocks other users from running on them
Tip
A job that utilizes ~70% of its requested resources is reasonably efficient. You don’t need to target 100% efficiency — some headroom prevents jobs from failing due to memory spikes or load variation.
Jobstats#
We recommend that users run jobstats on their jobs to profile and fine-tune them. Jobstats is an open-source job monitoring platform developed by Princeton University and used by Northwestern University Research Computing and Data Services. It provides richer efficiency reports than the built-in Slurm utility seff, including per-node and per-GPU breakdowns, and works on both running and completed jobs.
Running Jobstats on a Completed Job#
Run the jobstats command with a job ID to generate an efficiency report:
$ jobstats <job_id>
Example output for a completed job that utilizes multiple GPUs:
================================================================================
Slurm Job Statistics
================================================================================
Job ID: 1234567
User/Account: abc1234/p12345
Job Name: my_multi_gpu_job
State: COMPLETED
Nodes: 1
CPU Cores: 1
CPU Memory: 7GB
GPUs: 3
QOS/Partition: normal/gengpu
Cluster: quest
Start Time: Mon Aug 3, 2026 at 12:00 PM
Run Time: 00:06:58
Time Limit: 00:30:00
Overall Utilization
================================================================================
CPU utilization [|||||||||||||||||||||||||||||||||||||||||||| 88%]
CPU memory usage [|||||||||||||||||||||| 44%]
GPU utilization [||||||||||||||||||||||||||||||||||||||| 78%]
GPU memory usage [|||||||||||||||||||||||||||||||||||||||||| 85%]
Detailed Utilization
================================================================================
CPU utilization per node (CPU time used/run time)
qgpu0516: 00:06:11/00:06:58 (efficiency=88.8%)
CPU memory usage per node - maximum used/allocated
qgpu0516: 3.1GB/7GB
Average GPU utilization per node
qgpu0516 (GPU 0): 85%
qgpu0516 (GPU 2): 62.1%
qgpu0516 (GPU 3): 85.7%
Maximum GPU utilization per node
qgpu0516 (GPU 0): 100%
qgpu0516 (GPU 2): 100%
qgpu0516 (GPU 3): 100%
GPU memory usage per node - maximum used/total
qgpu0516 (GPU 0): 68.0GB/79.6GB (85.3%)
qgpu0516 (GPU 2): 68.0GB/79.6GB (85.3%)
qgpu0516 (GPU 3): 68.0GB/79.6GB (85.3%)
Notes
================================================================================
* Thank you for running jobstats! This tool is in development by
Northwestern Research Computing and Data Services (RCDS) in collaboration
with Princeton University. For more information about jobstats and job
efficiency on Quest, please see our documentation site
https://rcdsdocs.it.northwestern.edu/ or email
quest-help@northwestern.edu.
Reading the Report#
The report has two main parts:
Overall Utilization — Bar-chart summaries of CPU, CPU memory, GPU, and GPU memory usage across the entire job.
Detailed Utilization — Per-node and per-GPU breakdowns. This is especially useful for multi-node or multi-GPU jobs where load imbalance across nodes can indicate a parallelization problem.
Interpretation For Example Job#
For the example job (JobID=1234567) in the previous section, both parts of the jobstats report show that the job utilized the CPU, GPU, and GPU memory resources with reasonably high efficiency (greater than 70% utilization). The CPU memory utilization, however, is low (less than 50%). This implies that the following:
The CPU is being used efficiently and is passing work to the GPUs.
The user over-requested CPU memory for this job. The
#SBATCHparameter that controls CPU memory is--mem. If they submit this job again in the future, they should use--mem=5Gto reduce their memory ask while still giving a small amount of buffer.The program can effectively scale to using three GPU cards, distributing work relatively evenly across them.
However, under
Average GPU utilization per node, GPU cardGPU 2has a lower average utilization percentage thanGPU 0andGPU 3. If it was much lower (0-50%), that would be an indication that the GPU parallelization is not working well.Note that
GPU 0,GPU 2, andGPU 3are the names of the GPU cards on the nodeqgpu0516that this user was allocated for their job.GPU 1could be running the job of another user, or idle and available to run jobs.
GPU memory is not a reservable resource via
#SBATCHparameters. Instead, jobs are allocated GPU memory based on the number of cards they reserve. This job’s GPU memory utilization is near the full capacity of the 3 GPU cards reserved. If the job parameters, for example model or dataset size, were changed substantially, this job could run out of GPU memory.
Interpreting Common Findings#
Finding |
Likely cause |
Suggested action |
|---|---|---|
Low CPU utilization (<50%) |
Code has excessive idle time or load imbalance |
Address sources of idle time in code, execute a scaling study, check if the application supports parallelism and use |
Low CPU memory usage (<50%) |
Memory over-requested |
Lower |
Low GPU utilization (<50%) |
Bottleneck transferring work/data to GPU, load imbalance, or CPU-bound processing |
Check if code is well-configured to use GPUs, execute a scaling study |
Low GPU memory utilization (<50%) |
GPU card type requested or allocated has more vRAM than needed |
Use SLURM directive |
Uneven GPU usage across nodes |
Load imbalance |
Review GPU parallelism, execute a scaling study |
Job used much less than |
Run time much less than time limit |
Reduce |
Jobstats Reporting Options#
To view the guide for running jobstats in your terminal, run jobstats --help. A few key flags are --batch-script and --json. --batch-script shows the batch submission script you used to submit your job to run. --json returns your job’s efficiency report in a .json format that can be easier to save for reference and analysis.
For example, here is the output of jobstats --json for the multi-GPU job (JobID=1234567) above.
{
"gpus": 3,
"nodes": {
"qgpu0516": {
"cpus": 1,
"gpu_total_memory": {
"0": 85520809984,
"2": 85520809984,
"3": 85520809984
},
"gpu_used_memory": {
"0": 72977285120,
"2": 72977285120,
"3": 72977285120
},
"gpu_utilization_avg": {
"0": 85,
"2": 62.1,
"3": 85.7
},
"gpu_utilization_max": {
"0": 100,
"2": 100,
"3": 100
},
"total_memory": 7516192768,
"total_time": 371.1,
"used_memory": 3283456000
}
},
"total_time": 418
}
Memory is reported in bytes and time is reported in seconds. Note that gpu_total_memory and total_memory are the GPU and CPU memory available (or allocated) to your job respectively, whereas gpu_used_memory and used_memory are the GPU and CPU memory utilized by your job. The GPU statistics are reported per-card and labeled by the GPU card index. For this three card job, the user was allocated cards 0, 1, and 3.
Running Jobstats on a Running Job#
Jobstats can also be used to inspect a currently running job. This is useful for catching problems early without waiting for the job to finish.
$ jobstats <job_id>
The output will show utilization metrics accumulated up to the time of the query, with the State field showing RUNNING instead of COMPLETED.
Example output for a running job that utilizes multiple CPU cores:
$ jobstats -d 3456789
================================================================================
Slurm Job Statistics
================================================================================
Job ID: 3456789
User/Account: abc1234/p12345
Job Name: my_4cpu_job
State: RUNNING
Nodes: 1
CPU Cores: 4
CPU Memory: 8GB (2GB per CPU-core)
QOS/Partition: normal/short
Cluster: quest
Start Time: Mon Aug 3, 2026 at 1:00 PM
Run Time: 00:04:54 (in progress)
Time Limit: 01:00:00
Overall Utilization
================================================================================
CPU utilization [|||||||||||||||||||||||||||||||||||||||||||| 88%]
CPU memory usage [ 0%]
Detailed Utilization
================================================================================
CPU utilization per node (CPU time used/run time)
qnode0202: 00:17:28/00:19:39 (efficiency=88.9%)
CPU memory usage per node - maximum used/allocated
qnode0202: 27.4MB/8GB (6.9MB/2GB per core of 4)
Notes
================================================================================
* Thank you for running jobstats! This tool is in development by
Northwestern Research Computing and Data Services (RCDS) in collaboration
with Princeton University. For more information about jobstats and job
efficiency on Quest, please see our documentation site
https://rcdsdocs.it.northwestern.edu/ or email
quest-help@northwestern.edu.
Interpretation For Example Job#
For the example job (JobID=3456789) above, both parts of the jobstats report show that the job has so far utilized the four CPU cores it requested well, but had little-to-no CPU memory utilization. Since this job is currently running and only about five minutes into its (maximum one hour) runtime, we conclude:
The CPU parallelization is working well. Unless this changes significantly throughout the runtime of the job, there is no reason to change the parallelization scheme or number of requested CPUs.
Either the program has not yet reached the stage where it requires more significant CPU memory, or it is not a memory intensive job at all.
If the lack of memory usage is unexpected and persists as the job continues to run, the user should consider canceling the job and debugging.
If the job truly requires only megabytes of memory to run successfully, the user should cancel and resubmit the job with
--mem=1Gor even--mem=100M. This would help with overall efficiency and job scheduling on the system.
Tips for Troubleshooting Jobstats Output#
If the runtime of your job is less than 1-2 minutes, jobstats will not generate a full efficiency report. In this case, jobstats will either provide the output of
sefffor your job, or suggest that you runseffyourself.The database that underlies jobstats as well as the cluster’s scheduler (Slurm) only retains job data for a period of 6 months. If you run jobstats on a job that is 6 months or older, jobstats may throw an error or return output for an unrelated job.
If jobstats can’t find utilization information for your job, try running it with the debugging flags
--forceor--debug. If you continue to encounter issues, please contact the Research Computing and Data Services team at quest-help@northwestern.edu .
Slurm-based Profiling Tools#
Slurm has built-in commands that you can use to check the compute resource utilization and efficiency of a completed job. seff and checkjob generate efficiency reports including CPU and CPU memory information only. Running sacct with the formatting option tresusagein%120 can provide a small amount of information about CPU, CPU memory, GPU, and GPU memory utilization. For more details on those commands, see our page on Slurm Commands and Job Management.
Getting Help#
If you have questions about your job’s efficiency or want help optimizing your resource requests, contact the Research Computing and Data Services team at quest-help@northwestern.edu .