Smoke testΒΆ
On the GB10 host, test vLLM directly:
scripts/smoke-test.sh
Then test through the proxy with a user key. Here the script also checks that
/health, /metrics and the other private endpoints are not reachable, and
that requests without a key are rejected:
read -rsp "API key: " CODENDUM_API_KEY && export CODENDUM_API_KEY
scripts/smoke-test.sh --proxy --base-url https://llm.lab.example:8443
Add --cacert /path/to/ca.pem for a certificate from a private CA.
The checks run in this order:
Health, and the model list containing
coder.A chat completion.
A streamed completion, which must arrive in chunks and end with
[DONE].A tool call with
tool_choice: "auto". This is the path OpenCode uses and the one that exercises theqwen3_coderparser.A tool call with
tool_choice: "required". This goes through structured outputs and does not test the parser.A round trip in which a tool result is sent back and a text answer is expected.
Exit status is 0 when every check passes and 1 when a check fails.
2 means inconclusive: for example, the model legitimately answered without
calling the tool under "auto". An inconclusive result is not a success.
Repeat it, and use the OpenCode loop above to decide. If raw <tool_call>
markup appears in the message text, the parser is misconfigured and the check
fails.