Troubleshooting¶
Diagnostics
Symptom |
Likely cause and action |
|---|---|
Container exits during startup with out-of-memory ( |
Lower |
System becomes unresponsive during very long prefills |
Instability with 128K-token prefills on DGX Spark has been reported (dgx-spark-playbooks#97). Avoid |
Answers degrade, or long outputs repeat themselves |
The FP8 KV cache can affect quality, and vLLM’s DGX Spark guidance says to use it only when memory pressure requires it and quality checks pass. Compare with |
Kernel errors mentioning FP8 block scaling on SM 12.1 |
Seen with CUDA 12.9 builds (vllm#43367). Use a CUDA 13 image; the default is one. |
|
Tool parsing is off or wrong: run |
HTTP 400 mentioning the maximum context length |
The conversation exceeds |
Slow responses for everyone |
Check |
TLS errors in OpenCode |
Set |
|
Use an account that may run Docker. On DGX OS, adding the user to the |
|
Another service uses 8443. Change the |
NVIDIA runtime missing in |
NVIDIA’s troubleshooting step is |
For the HTTP status codes 401, 403, 404, 429, 502 and 504, see the proxy section.
Logs
vLLM:
docker logs codendum-vllm. Prompts are not logged, because--enable-log-requestsis off by default.nginx:
/etc/codendum/proxy/logs/codendum.access.logrecords the time, client address, user id, path, status and durations, never the keys or the request bodies. Errors go tocodendum.error.login the same directory. With the distribution’s nginx, the directory is/var/log/nginx/.