pwdLocate your shell.
The absolute working directory determines how relative paths resolve.
Scope in this emulator
The lab starts in /home/learner.
The operator’s field reference
Supported command forms, diagnostic decisions and the hardware vocabulary behind them. Keep this beside your terminal.
Open the free sandbox →pwdThe absolute working directory determines how relative paths resolve.
The lab starts in /home/learner.
ls -la /workspaceRead names, permissions and whether an entry is a directory.
Permissions and timestamps are teaching fixtures.
cat /etc/os-releaseNAME and VERSION describe the OS, separate from the NVIDIA driver.
This is the virtual node's OS fixture, not your laptop.
cat /etc/systemd/system/llm.serviceUser and ExecStart tell you who launches the process and how.
The lab adds its own JSON configuration bridge; real systemd needs explicit arguments and environment.
cat /models/lab-model/config.jsonLayers, KV heads, head dimension and element size determine bytes per cached token.
An illustrative model, with no weights to download or execute.
echo evidence > /workspace/notes.txtUse cat /workspace/notes.txt to verify the saved text.
Virtual workspace only; notes disappear on reload. There is no upload.
hostnameDevice indices repeat across nodes. Record the hostname alongside GPU identity.
Switch nodes through the sandbox selector; SSH is not emulated.
nvidia-smiCompare utilization, memory allocation and the process table.
Temperature and power follow teaching state; they are not hardware measurements.
nvidia-smi -LA UUID identifies a device more precisely than a context-dependent index.
UUIDs are stable within this lab cluster, not real hardware identifiers.
nvidia-smi --query-gpu=index,uuid,pci.bus_id --format=csvUse the bus ID to connect a device to a kernel event.
PCI addresses repeat on different nodes; preserve the hostname too.
nvidia-smi --versionCompare this with the compiler's version before inferring toolkit availability.
The displayed CUDA version is not an installed-toolkit inventory.
nvcc --versionThe release value belongs to the toolkit that supplied this compiler.
Compilation and kernel execution require a real CUDA environment.
nvidia-smi --query-gpu=index,memory.total,memory.used,memory.free --format=csvUsed plus free equals total in the lab's accounting, in MiB.
Real usable memory includes driver and platform reservations omitted here.
nvidia-smi -q -i 0 -d MEMORYCheck free space on the intended device, not the sum across the node.
Eight devices do not automatically form one addressable allocation pool.
nvidia-smi -q -d UTILIZATIONA process may hold memory while GPU activity is zero.
Memory activity is explicitly unmodeled; allocated bytes do not measure bandwidth use.
nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csvJoin the PID to ps aux before deciding whether a process is yours to stop.
The lab lists GPU processes only; a real system has many additional processes.
ps auxUSER and PID distinguish your notebook from another user's workload.
Host CPU and RSS columns are fixtures; use nvidia-smi for GPU allocation.
kill -TERM 4102Re-query GPU memory after termination. Stop only the intended process.
Requires a learner-owned PID in the current scenario; other users' workloads are protected.
envCUDA_VISIBLE_DEVICES selects and reorders devices seen by CUDA applications.
nvidia-smi continues to show the physical inventory in this lab.
export CUDA_VISIBLE_DEVICES=3,1The next lab model launch uses physical GPU 3. Existing processes keep their placement.
The lab passes this shell environment to its service; real systemd requires explicit configuration.
unset CUDA_VISIBLE_DEVICESNew lab launches return to physical GPU 0 by default.
Changing the environment does not move a running process.
cat /workspace/serve.jsonRead the context cap, sequence cap, memory fraction and listening port.
Edits take effect on the next start. The Files panel edits this virtual file.
systemctl status llm.serviceRead the current state, then inspect the journal for a failed start.
A running process alone does not prove the client can reach its endpoint.
systemctl start llm.serviceA failure reports required memory, free memory and the configured budget.
The lab reserves worst-case KV memory; real vLLM uses a shared KV pool.
systemctl restart llm.serviceConfirm placement and listening port after the restart.
A restart interrupts the lab service and can fail if the new configuration cannot fit.
journalctl -u llm.serviceDistinguish an old failure from the latest successful launch.
The bounded teaching journal retains recent events; it is not a full system journal.
ss -ltnpCompare the LISTEN address and PID with your client's URL.
Only the local lab model endpoint is modeled.
curl -i http://localhost:8000/healthA 200 response requires a running service on that port.
No network request is sent. Health does not test generated answers, throughput or latency.
vllm serve /models/lab-model --max-model-len 4096 --max-num-seqs 16At 128 KiB per token, this lab configuration reserves 8 GiB of KV and 26 GiB overall.
Supported CLI options feed the lab lifecycle; no Python process or language model executes.
nvidia-smi topo -mRead GPU pairs and NVLink paths before reasoning about communication.
Configured topology is not a bandwidth test or a live link-health report.
nvidia-smi nvlink -sFind the device and link reporting Down; correlate with the kernel log.
Lab output reports state only. Real output and recovery requirements vary by architecture.
ibstatRead State, Physical state, Rate and LID for the same port.
No remote fabric discovery; one emulated port per HCA.
ibstatusAn Active port and a negotiated rate describe link readiness.
Neither proves an end-to-end collective is healthy.
iblinkinfoMatch a down adapter to the selected rail in Cluster fabric.
This lab restricts discovery to the selected node's local adapters.
perfqueryLinkDownedCounter preserves a previous outage after the link recovers.
Only HCA 0 is supported; traffic counters are N/A and no benchmark runs.
journalctl -kRecord the PCI identifier, Xid or lab event, and neighboring messages.
Blackwell NVLink faults use an explicit lab event instead of an inapplicable Xid 74.
nvidia-smi -q -i 0 -d ECCVolatile and aggregate counts have different lifetimes.
Lab maintenance clears the modeled current fault and retains history; this is not a vendor recovery procedure.
dcgmi discovery -lCross-check GPU identity with nvidia-smi -L.
Groups, field watches and remote hostengine operation are outside this subset.
dcgmi diag -r 1A pending injected ECC fault fails the check; review kernel and ECC evidence.
A lab Pass does not stress hardware or certify memory, bandwidth or application correctness.
nvidia-smi -L | grep 'GPU 3'Literal matching keeps the original line available for inspection.
grep supports literal -F, -i and -v forms; regular expressions and shell scripts are not executed.
journalctl -u llm.service | tail -n 3Use the full journal if the last lines omit the initiating failure.
At most four filters, 1,024 input characters and 100 history entries.
Choose a symptom, gather evidence, then change the smallest relevant part of the system. Recheck after the change.
High memory.used with low utilization.gpu
nvidia-smi -i 0ps auxIdentify the allocation owner. Stop your disposable workload only when its work can be discarded; otherwise choose an available device or change the workload budget.
Re-query memory after the change, then retry admission.
Zero activity does not release a process's allocations.
A failed llm.service and a memory-reservation error
journalctl -u llm.servicecat /workspace/serve.jsonCompare required memory against both physical free space and the configured fraction. Reduce context or concurrent sequences when the requested reservation is too large.
Restart, inspect the allocation, then call the matching health port.
Increasing the fraction cannot create physical memory.
Active service with curl connection refused
ss -ltnpcat /workspace/serve.jsonUse the listening socket as evidence of the live port. Correct the client, or change the configured port and restart when the service contract requires a different address.
Check the live socket again and obtain a response on the intended URL.
Saving a file does not change a running process.
Allocation appears on a different physical index
envnvidia-smi --query-gpu=index,uuid,memory.used --format=csvCheck the ordering in CUDA_VISIBLE_DEVICES. Configure the intended device for the next launch, then restart; never assume visible index zero means physical GPU zero.
Match the process allocation to the physical inventory after launch.
An environment edit does not migrate an existing allocation.
Down in the per-device link report
nvidia-smi nvlink -sjournalctl -kIdentify the affected device and peer context. In a real environment, follow the architecture-specific vendor workflow before planning reset or maintenance.
After lab maintenance, re-read link state. Historical kernel messages remain.
A connectivity matrix alone cannot confirm current link health.
One local adapter reports Down
ibstatiblinkinfoperfqueryLocate the node and HCA before widening the investigation. The sandbox's rail view maps that local adapter to its logical cluster path.
Restore the lab fault, then check Active state and the retained link-down counter on HCA 0.
An old error count does not prove a fresh outage; a link rate does not measure throughput.
A kernel Xid 48 and uncorrectable ECC count
journalctl -knvidia-smi -q -d ECCPreserve evidence and inspect the vendor workflow for that architecture and error sequence. The lab's Restore control is a teaching maintenance action, not an instruction to reset a production node.
Recheck ECC and preflight results, then explicitly restart the stopped service.
Restoring device readiness does not automatically restart an application.
DCGM preflight reports Pass
dcgmi diag -r 1systemctl status llm.servicess -ltnpContinue checking the application path. Software preflight has a narrower scope than a hardware stress test, model request or latency test.
In the lab, verify allocation and /health. On real infrastructure, add representative application and performance checks.
A green check cannot establish properties it never tested.
Record the full message, device, architecture and related events. Check NVIDIA’s current catalog before taking recovery action. Only the marked examples can be introduced through this lab’s fault controls.
Often an application fault. Investigate application access patterns; software or hardware can also be responsible.
Reference knowledgeInspect the failing access and application. A page fault is not evidence of insufficient free capacity.
Reference knowledgeCorrelate the full event sequence and ECC counters. Recovery depends on architecture and accompanying events.
Fault practice availableRecords remapping activity. Distinguish it from Xid 64, which reports remapping failure.
Reference knowledgeCheck architecture applicability and peer events. This lab uses 74 on Ampere/Hopper, not Blackwell.
Fault practice availableInspect PCIe and system evidence alongside the event. The number alone does not identify the failed component.
Reference knowledgeStart with one GPU’s memory, then widen your view to its neighbors and the network between nodes.
Processes reserve memory on a selected physical device.
Eight GPU endpoints connected by a local switch fabric.
Eight logical rails connect the sandbox’s eight nodes.
| Profile | Generation | Nominal GiB / GPU | Lab NVLink links / GPU | Lab HCA rate |
|---|---|---|---|---|
| B200 | Blackwell | 180 | 18 | 400 Gb/s |
| H200 | Hopper | 141 | 18 | 400 Gb/s |
| H100 | Hopper | 80 | 18 | 400 Gb/s |
| A100 | Ampere | 80 | 12 | 200 Gb/s |
Nominal capacities are teaching inputs. Real usable memory, software versions and fabric configuration depend on the system. The sandbox collapses switch tiers and reports local link state; it does not execute collectives or predict throughput.
A process starts → memory is reserved on its selected GPU → nvidia-smi and the instrument panel show that allocation. A fault changes a link or device → diagnostics and the topology reflect it. Historical counters remain after restoration so you can practice separating old evidence from current state.
Reviewed 29 September 2026. Output and recovery rules vary by driver and architecture.
NVIDIA System Management Interface ↗NVIDIA Xid diagnostic catalog ↗NVIDIA DCGM diagnostics ↗vLLM engine arguments ↗NVIDIA HGX platform components ↗Independent educational software. Not affiliated with or endorsed by NVIDIA Corporation. NVIDIA and its product names are trademarks of NVIDIA Corporation.
DC Lab Sim components MIT License Copyright (c) 2026 Sean Boerhout Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software. THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. xterm.js Copyright (c) 2017-2019, The xterm.js authors (https://github.com/xtermjs/xterm.js) Copyright (c) 2014-2016, SourceLair Private Company (https://www.sourcelair.com) Copyright (c) 2012-2013, Christopher Jeffrey (https://github.com/chjj/) Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software. THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.