◈
LOCAL INFERENCE / TELEMETRY
GPU
Command Center
GPU0 RTX 2080 Ti · GPU1 GTX 1080 · awaiting telemetry
connecting
telemetry stream
GPU0 state
—
0
/min infer ·
0
/min poll
0
fps ·
0
Hz in
✦ Customize
Runtime overview
stable hardware identity · one live telemetry stream
⚡ Live
📊 History
Installed GPUs
Pinned by physical index. Names stay here; the page header stays stable.
connecting…
GPU0 connecting…
GPU1 connecting…
Inference snapshot
same stream · current frame
Inference / min
0
Active slots
0
Live clients
0
Model loaded
—
Gen tok/s
0
GPU Utilization
0
%
VRAM
0
GiB
Mem bus util
0
%
Headroom
0
GiB free
Temperature
0
°C
Fan
0
%
Context Window
0
% full
Used
0
tok
Total
0
tok
Power Draw
0
W
limit
280
W
Clocks
SM
0
MHz
Graphics
0
MHz
Memory
0
MHz
🧠 Loaded Model
idle
0 active
no model loaded
⚙ GPU0 Controls · RTX 2080 Ti
PIN
required for any change
Power limit
280
W
Set
Fan
temp-driven smart curve handles cooling
🌡 Smart curve
Live Clients · connected to :18091
scanning…
GPU0 Tenants · VRAM holders
—
Recent Inference Calls
waiting for inference traffic…
Token Throughput
0
tok/s gen
prompt eval
0
tok/s
Inference
Slots busy
0
/
0
Prompt tokens
0
Model Load Timeline
waiting for a model load/swap…
Requests
0
Tokens generated
0
Avg tok/s
0
Peak tok/s
0
Time
Model
Prompt
Generated
Avg t/s
Peak t/s
Duration
no requests recorded yet — send one and watch it appear