AI Infra Interviews logo

The operator’s field reference

Know what the output means.

Supported command forms, diagnostic decisions and the hardware vocabulary behind them. Keep this beside your terminal.

Open the free sandbox →
41 / 41 supported forms
Linux & files
pwd

Locate your shell.

The absolute working directory determines how relative paths resolve.

Scope in this emulator

The lab starts in /home/learner.

Linux & files
ls -la /workspace

Find editable configuration and notes.

Read names, permissions and whether an entry is a directory.

Scope in this emulator

Permissions and timestamps are teaching fixtures.

Linux & files
cat /etc/os-release

Identify the operating system.

NAME and VERSION describe the OS, separate from the NVIDIA driver.

Scope in this emulator

This is the virtual node's OS fixture, not your laptop.

Linux & files
cat /etc/systemd/system/llm.service

Inspect the service launch contract.

User and ExecStart tell you who launches the process and how.

Scope in this emulator

The lab adds its own JSON configuration bridge; real systemd needs explicit arguments and environment.

Linux & files
cat /models/lab-model/config.json

Read the workload's memory assumptions.

Layers, KV heads, head dimension and element size determine bytes per cached token.

Scope in this emulator

An illustrative model, with no weights to download or execute.

Linux & files
echo evidence > /workspace/notes.txt

Keep a short investigation note.

Use cat /workspace/notes.txt to verify the saved text.

Scope in this emulator

Virtual workspace only; notes disappear on reload. There is no upload.

Linux & files
hostname

Confirm which node you are inspecting.

Device indices repeat across nodes. Record the hostname alongside GPU identity.

Scope in this emulator

Switch nodes through the sandbox selector; SSH is not emulated.

GPU inventory
nvidia-smi

Read the node's current GPU summary.

Compare utilization, memory allocation and the process table.

Scope in this emulator

Temperature and power follow teaching state; they are not hardware measurements.

GPU inventory
nvidia-smi -L

List device indices, names and UUIDs.

A UUID identifies a device more precisely than a context-dependent index.

Scope in this emulator

UUIDs are stable within this lab cluster, not real hardware identifiers.

GPU inventory
nvidia-smi --query-gpu=index,uuid,pci.bus_id --format=csv

Join GPU identity to PCI location.

Use the bus ID to connect a device to a kernel event.

Scope in this emulator

PCI addresses repeat on different nodes; preserve the hostname too.

GPU inventory
nvidia-smi --version

Inspect the driver and supported CUDA ceiling.

Compare this with the compiler's version before inferring toolkit availability.

Scope in this emulator

The displayed CUDA version is not an installed-toolkit inventory.

GPU inventory
nvcc --version

Inspect the CUDA compiler in PATH.

The release value belongs to the toolkit that supplied this compiler.

Scope in this emulator

Compilation and kernel execution require a real CUDA environment.

Memory & processes
nvidia-smi --query-gpu=index,memory.total,memory.used,memory.free --format=csv

Compare all devices in a structured view.

Used plus free equals total in the lab's accounting, in MiB.

Scope in this emulator

Real usable memory includes driver and platform reservations omitted here.

Memory & processes
nvidia-smi -q -i 0 -d MEMORY

Inspect one device's framebuffer allocation.

Check free space on the intended device, not the sum across the node.

Scope in this emulator

Eight devices do not automatically form one addressable allocation pool.

Memory & processes
nvidia-smi -q -d UTILIZATION

Separate activity from allocation.

A process may hold memory while GPU activity is zero.

Scope in this emulator

Memory activity is explicitly unmodeled; allocated bytes do not measure bandwidth use.

Memory & processes
nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv

Find the processes reserving GPU memory.

Join the PID to ps aux before deciding whether a process is yours to stop.

Scope in this emulator

The lab lists GPU processes only; a real system has many additional processes.

Memory & processes
ps aux

Check process identity and ownership.

USER and PID distinguish your notebook from another user's workload.

Scope in this emulator

Host CPU and RSS columns are fixtures; use nvidia-smi for GPU allocation.

Memory & processes
kill -TERM 4102

Gracefully stop your disposable notebook when it exists.

Re-query GPU memory after termination. Stop only the intended process.

Scope in this emulator

Requires a learner-owned PID in the current scenario; other users' workloads are protected.

Device placement
env

Inspect launch-time device visibility.

CUDA_VISIBLE_DEVICES selects and reorders devices seen by CUDA applications.

Scope in this emulator

nvidia-smi continues to show the physical inventory in this lab.

Device placement
export CUDA_VISIBLE_DEVICES=3,1

Make physical GPU 3 the first visible CUDA device.

The next lab model launch uses physical GPU 3. Existing processes keep their placement.

Scope in this emulator

The lab passes this shell environment to its service; real systemd requires explicit configuration.

Device placement
unset CUDA_VISIBLE_DEVICES

Remove the shell's visibility override.

New lab launches return to physical GPU 0 by default.

Scope in this emulator

Changing the environment does not move a running process.

Model service
cat /workspace/serve.json

Inspect the proposed serving settings.

Read the context cap, sequence cap, memory fraction and listening port.

Scope in this emulator

Edits take effect on the next start. The Files panel edits this virtual file.

Model service
systemctl status llm.service

Check whether the process supervisor sees an active service.

Read the current state, then inspect the journal for a failed start.

Scope in this emulator

A running process alone does not prove the client can reach its endpoint.

Model service
systemctl start llm.service

Admit the workload against the selected GPU's memory.

A failure reports required memory, free memory and the configured budget.

Scope in this emulator

The lab reserves worst-case KV memory; real vLLM uses a shared KV pool.

Model service
systemctl restart llm.service

Apply edited configuration to a fresh model process.

Confirm placement and listening port after the restart.

Scope in this emulator

A restart interrupts the lab service and can fail if the new configuration cannot fit.

Model service
journalctl -u llm.service

Read the history behind the current service state.

Distinguish an old failure from the latest successful launch.

Scope in this emulator

The bounded teaching journal retains recent events; it is not a full system journal.

Model service
ss -ltnp

Find the port the live service actually uses.

Compare the LISTEN address and PID with your client's URL.

Scope in this emulator

Only the local lab model endpoint is modeled.

Model service
curl -i http://localhost:8000/health

Test the client-facing health route.

A 200 response requires a running service on that port.

Scope in this emulator

No network request is sent. Health does not test generated answers, throughput or latency.

Model service
vllm serve /models/lab-model --max-model-len 4096 --max-num-seqs 16

Launch with explicit serving limits.

At 128 KiB per token, this lab configuration reserves 8 GiB of KV and 26 GiB overall.

Scope in this emulator

Supported CLI options feed the lab lifecycle; no Python process or language model executes.

Links & fabric
nvidia-smi topo -m

Inspect the configured intra-node connectivity matrix.

Read GPU pairs and NVLink paths before reasoning about communication.

Scope in this emulator

Configured topology is not a bandwidth test or a live link-health report.

Links & fabric
nvidia-smi nvlink -s

Inspect current per-GPU link state.

Find the device and link reporting Down; correlate with the kernel log.

Scope in this emulator

Lab output reports state only. Real output and recovery requirements vary by architecture.

Links & fabric
ibstat

Inspect local InfiniBand adapters and port state.

Read State, Physical state, Rate and LID for the same port.

Scope in this emulator

No remote fabric discovery; one emulated port per HCA.

Links & fabric
ibstatus

Read a compact local port summary.

An Active port and a negotiated rate describe link readiness.

Scope in this emulator

Neither proves an end-to-end collective is healthy.

Links & fabric
iblinkinfo

Compare the local adapters' link status.

Match a down adapter to the selected rail in Cluster fabric.

Scope in this emulator

This lab restricts discovery to the selected node's local adapters.

Links & fabric
perfquery

Read the first local HCA's counters.

LinkDownedCounter preserves a previous outage after the link recovers.

Scope in this emulator

Only HCA 0 is supported; traffic counters are N/A and no benchmark runs.

Fault diagnosis
journalctl -k

Read kernel events around a GPU failure.

Record the PCI identifier, Xid or lab event, and neighboring messages.

Scope in this emulator

Blackwell NVLink faults use an explicit lab event instead of an inapplicable Xid 74.

Fault diagnosis
nvidia-smi -q -i 0 -d ECC

Compare current and historical uncorrectable error counters.

Volatile and aggregate counts have different lifetimes.

Scope in this emulator

Lab maintenance clears the modeled current fault and retains history; this is not a vendor recovery procedure.

Fault diagnosis
dcgmi discovery -l

List devices visible to DCGM discovery.

Cross-check GPU identity with nvidia-smi -L.

Scope in this emulator

Groups, field watches and remote hostengine operation are outside this subset.

Fault diagnosis
dcgmi diag -r 1

Run the lab's software preflight checks.

A pending injected ECC fault fails the check; review kernel and ECC evidence.

Scope in this emulator

A lab Pass does not stress hardware or certify memory, bandwidth or application correctness.

Shell filters
nvidia-smi -L | grep 'GPU 3'

Select a specific line from command output.

Literal matching keeps the original line available for inspection.

Scope in this emulator

grep supports literal -F, -i and -v forms; regular expressions and shell scripts are not executed.

Shell filters
journalctl -u llm.service | tail -n 3

Focus on the latest service events.

Use the full journal if the last lines omit the initiating failure.

Scope in this emulator

At most four filters, 1,024 input characters and 100 history entries.

Independent educational software. Not affiliated with or endorsed by NVIDIA Corporation. NVIDIA and its product names are trademarks of NVIDIA Corporation.

Third-party software notices
DC Lab Sim components

MIT License

Copyright (c) 2026 Sean Boerhout

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.


xterm.js

Copyright (c) 2017-2019, The xterm.js authors (https://github.com/xtermjs/xterm.js)
Copyright (c) 2014-2016, SourceLair Private Company (https://www.sourcelair.com)
Copyright (c) 2012-2013, Christopher Jeffrey (https://github.com/chjj/)

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in
all copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
THE SOFTWARE.