Ornith-1.0-35B-FP8 Vision Path Produces Degenerate Output

#4
by kickxxmatt - opened

Does anyone facing same issue?

Model: deepreinforce-ai/Ornith-1.0-35B-FP8
Inference Backend: vLLM v0.23.1rc1 (custom NVIDIA build)
Hardware: NVIDIA GB10 (DGX Spark, 128GB unified memory)

Issue:
When sending image input to the FP8 variant, the model produces degenerate output consisting entirely of ! characters, regardless of whether thinking mode is enabled or disabled.
Reproduction:
bashBASE64=$(base64 -w 0 image.jpg)
curl -X POST http://localhost:8000/v1/chat/completions
-H "Content-Type: application/json"
-d "{
"model": "ornith-35b",
"messages": [{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,${BASE64}"}},
{"type": "text", "text": "describe this image"}
]}],
"max_tokens": 200,
"chat_template_kwargs": {"enable_thinking": false}
}"
Output:
json{
"choices": [{
"message": {
"content": "!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!",
"reasoning": null
},
"finish_reason": "length"
}],
"usage": {"prompt_tokens": 561, "completion_tokens": 200}
}

docker run -d --restart unless-stopped --name vllm-ornith
--gpus all -p 8000:8000
--ipc=host
-e HF_TOKEN="$HF_TOKEN"
-v ~/.cache/huggingface:/root/.cache/huggingface
vllm-node
vllm serve deepreinforce-ai/Ornith-1.0-35B-FP8
--served-model-name ornith-35b
--host 0.0.0.0 --port 8000
--max-model-len 131072
--gpu-memory-utilization 0.5
--enable-prefix-caching
--enable-auto-tool-choice --tool-call-parser qwen3_xml
--reasoning-parser qwen3
--trust-remote-code

Sign up or log in to comment