Hands-on GPU sandbox · try it in your browser
Meet the machine
behind the model.
Take the controls of an eight-GPU node. Practice NVIDIA commands, explore GPU memory and recover a model endpoint, with a guide beside your terminal. No hardware setup required.
3 free missions · No account or GPU required
GPU 0: NVIDIA B200
GPU 1: NVIDIA B200
… 6 more devices
BLACKWELL · 8 DEVICES · YOUR TERMINAL
Stop a process. Free its memory. See the machine respond to your command.
Understand what to inspect, why it matters and how to check your result.
Reset any mission. Your commands stay inside your browser sandbox.
The learning path
From first command to incident command.
0 / 24 COMPLETE · SAVED ON THIS DEVICE
First contact
Find your way around Linux, identify GPUs and separate activity from allocation.
First contact
Find your bearings on an eight-GPU node.
~7 min
Idle, but not empty
Find the process holding 56 GiB of GPU memory.
~9 min
The version trap
Separate the driver, its CUDA ceiling and the compiler.
~7 min
From model to endpoint
Read configuration, budget memory and verify the address your client uses.
Bring it online
Take a model service from stopped to healthy.
~10 min
Make it fit
Repair an impossible context and concurrency budget.
~12 min
The silent endpoint
A running process, a client pointed somewhere else.
~10 min
You’re on call
Trace device placement, inspect the fabric and recover a failed endpoint.
Which GPU is zero?
Follow device visibility through a model launch.
~12 min
Read the fabric
Discover what a topology table can and cannot tell you.
~9 min
Night shift
Recover a failed endpoint with fewer instructions.
~15 min
Hardware detective
Correlate NVLink, InfiniBand and ECC evidence. Separate current device state from historical errors.
A missing link
Trace a local link fault without treating topology as a benchmark.
~16 min
The first hop
Locate an InfiniBand problem at the node boundary.
~17 min
After an ECC event
Separate device recovery from application recovery.
~18 min
Yesterday’s error, today’s state
Interpret an error counter without erasing history.
~15 min
What did the check test?
Use DCGM preflight without treating Pass as a certificate.
~16 min
Serving engineer
Repair configuration, trade context for concurrency and verify placement without disrupting neighboring workloads.
The file will not parse
Recover from a configuration syntax failure.
~17 min
Spend the KV budget
Trade concurrency for longer contexts using explicit arithmetic.
~20 min
Free memory, rejected launch
Distinguish available capacity from the configured fraction.
~18 min
Saved is not applied
Observe the boundary between a file and a live process.
~18 min
Find capacity without collateral damage
Place a service around another team’s allocation.
~19 min
Incident command
Diagnose several independent blockers, verify the complete local path and leave a useful handoff.
Two faults, two fabrics
Separate a local GPU link from a network adapter incident.
~24 min
The second blocker
Recover a launch with a bad configuration and a bad target.
~25 min
Follow the client to the device
Verify the full endpoint identity after a migration.
~24 min
The green check’s blind spot
Diagnose a fabric fault that software preflight does not cover.
~22 min
Close the shift
Recover a device, a memory budget and the endpoint contract.
~30 min
Start with Chapter 1 for the fundamentals, then follow the chapters in order. All later chapters are included with active Premium access ↗. Completion is a practice record, not a certification.
Build fluency through repetition
A path you can return to.
Learn a workflow with guidance. Repeat it in Challenge mode, then change the architecture or inject a fault in the open sandbox. Use the linked concepts and courses to explain what you observed.
Guided objectives explain the evidence and the reason for each check.
Challenge mode gives you targets. Choose your own commands and reveal help when needed.
Practice on A100, H100, H200 and B200 profiles. Continue on real hardware for CUDA execution and performance testing.
Future lab topics
The machine has more to teach.
Directions we’re exploring for future labs.
Inside a CUDA kernel
Threads, warps, memory access and the work behind an occupancy number.
Serving under pressure
Batching, KV cache pressure and vLLM scheduling decisions.
Beyond one node
NCCL collectives, distributed training and cluster diagnosis.
Read a GPU profile
Kernel timelines, memory bottlenecks and the signals behind a roofline plot.
Share a GPU with MIG
GPU instances, memory boundaries and workload isolation on a shared device.
Schedule GPUs on Kubernetes
Device plugins, GPU requests and the reasons a training pod stays pending.
GPU emulator / command-line practice
A GPU lab you can open anywhere.
Learn the terminal workflows behind GPU operations and AI infrastructure interviews. Start with the free missions, then connect your observations to the hardware and learning guides.
What is the online GPU emulator?
GPU Lab is a hands-on NVIDIA command sandbox from AI Infra Interviews. Our in-house lab engine connects a browser terminal to GPU inventory, process ownership, memory allocation, virtual files and service state. Guided missions explain what to inspect, which command to use and how to verify your result.
Can I use this GPU simulator without an NVIDIA GPU?
Yes. The command emulator runs in your browser, so you can practice on a laptop without a GPU, cloud account or driver installation. Choose an eight-device A100, H100, H200 or Blackwell B200 node. Profiles use nominal memory capacities; real usable memory depends on hardware and driver reservations.
Which NVIDIA GPU commands can I practice?
Practice nvidia-smi inventory, memory queries, process inspection and NVLink topology. Use nvcc --version to inspect the toolkit, CUDA_VISIBLE_DEVICES to control device selection, and supported ps, kill, systemctl, journalctl, ss and curl commands to diagnose the lab endpoint. The lab also covers NVLink status, ECC evidence, DCGM software preflight and local InfiniBand readouts. The command reference lists the supported forms.
Does the sandbox run CUDA kernels or real vLLM inference?
It emulates command-line workflows and the resulting machine state. The vLLM exercises cover configuration, memory reservation, device placement and endpoint health. Running CUDA kernels, generating model output and measuring throughput or latency require real hardware. The charts display emulator state, with explicitly modeled workload activity.
Can I change GPU load and watch live graphs?
Yes. The free sandbox includes idle, inference, training and stress presets. Select GPUs, adjust workload intensity and follow activity, memory, temperature and power graphs. The process monitor and nvidia-smi read the same machine state. Pause the emulation clock to inspect a sample, or advance it one second at a time. Workload shapes and thermal readings are teaching models, not hardware benchmarks.
Is GPU Lab free to try?
3 beginner missions are free with no account required: node exploration, GPU memory troubleshooting and CUDA version inspection. The open cluster sandbox and command reference are also free. Active Premium access includes 21 serving, hardware-diagnostic and incident-recovery missions, with guided and challenge modes. Mission summaries show the access level before you launch.
Independent educational software. Not affiliated with or endorsed by NVIDIA Corporation. NVIDIA and its product names are trademarks of NVIDIA Corporation.
Third-party software notices
DC Lab Sim components MIT License Copyright (c) 2026 Sean Boerhout Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software. THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. xterm.js Copyright (c) 2017-2019, The xterm.js authors (https://github.com/xtermjs/xterm.js) Copyright (c) 2014-2016, SourceLair Private Company (https://www.sourcelair.com) Copyright (c) 2012-2013, Christopher Jeffrey (https://github.com/chjj/) Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software. THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
