Skip to main content

Command Palette

Search for a command to run...

Familiar Monitoring of TPU Workloads

Updated
•2 min read•View as Markdown
Familiar Monitoring of TPU Workloads

How can I see the usage of my TPU chip? Can I run nvtop on a TPU machine?

These were some of the biggest questions I had when I started working with zt for my research. When running code on a TPU based setup, be that inside of Google Colab or via GCP it is helpful to get a better understanding of the current load on the chips you might be using. When you google "nvtop for TPUs" not much comes up apart from some issues on GitHub. Looking through the pull requests of nvtop you might come across the way to re-compile the code with support for TPU chips.

First, what is a TPU? A Tensor Processing Unit is a specialized chip designed by Google to speed up machine learning workloads, especially neural networks and large matrix operations used in training and inference. I was fortunate enough to work with a large allocation of multiple chips from Google and for the ease of future use, setup a small bash gist: here.

This mini-script will install the necessary requirements and fetch the source code of nvtop. Then it will compile it with TPU support enabled. A very simple way to run it is to do the following:

curl -sSL https://gist.githubusercontent.com/velocitatem/53ff537d3c439c3bd63d8a0f1370801b/raw/c348e534f7c24646c52eabce11cad9ee01611fb1/install_nvtop_tpu.sh | bash

What this does is simply fetch the script and run it on your machine (with the expectation of Ubuntu 22.04). This is not limited to running on provisioned virtual machines, but can also be run in the colab terminal:

![](https://cdn.hashnode.com/uploads/covers/6793605016666de203f20485/4697660f-ef58-4fd8-b827-b421b2f34426.png align="middle")

You can then run some Jax powered script to test it out and monitor the load.