{"id":1131,"date":"2026-02-03T09:59:26","date_gmt":"2026-02-03T09:59:26","guid":{"rendered":"https:\/\/tutorial.emka.web.id\/?p=1131"},"modified":"2026-02-03T09:59:26","modified_gmt":"2026-02-03T09:59:26","slug":"how-to-run-qwen-14b-on-amd-mi200-with-vllm","status":"publish","type":"post","link":"https:\/\/emka.web.id\/en\/2026\/02\/how-to-run-qwen-14b-on-amd-mi200-with-vllm.html","title":{"rendered":"How to Run Qwen (14B) on AMD MI200 with vLLM"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">If you are trying to run LLMs on AMD Instinct MI200 (MI210\/MI250) cards, you have probably already experienced the pain of &#8220;HSA errors,&#8221; random segmentation faults, or containers that just hang forever.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We went through the struggle of finding the right Docker image so you don&#8217;t have to. Here is the definitive, battle-tested guide to running an OpenAI-compatible API for Qwen (14B) on ROCm.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" src=\"https:\/\/www.amd.com\/content\/dam\/amd\/en\/images\/backgrounds\/products\/2325906-instinct-mi200-series-hero.jpg\" alt=\"\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The MI200 is a beast, but it uses the <code>gfx90a<\/code> architecture. Most &#8220;bleeding edge&#8221; Docker images today are optimized for the newer MI300 (<code>gfx942<\/code>). If you try to run the latest vLLM (0.11.x) with the default settings, it will crash because the new execution engine (<code>aiter<\/code>) isn&#8217;t fully compatible with MI200 yet.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We are going to use a <strong>stable setup<\/strong> that disables the experimental features and just runs fast.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Prerequisites<\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Host OS:<\/strong> Linux with ROCm kernel drivers installed (<code>rocm-dkms<\/code>).<\/li>\n\n\n\n<li><strong>Docker:<\/strong> Installed and running.<\/li>\n\n\n\n<li><strong>GPU:<\/strong> AMD Instinct MI200 series (MI210, MI250\/X).<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Don&#8217;t use <code>latest<\/code>. Don&#8217;t use <code>0.11.x<\/code>. We are using <strong>vLLM 0.10.1 on ROCm 6.4<\/strong>. It provides the best balance of modern model support (like Qwen 2.5) and stability.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Copy and paste this exact command.<\/p>\n\n\n\n<pre class=\"wp-block-syntaxhighlighter-code\">docker run -it --rm \\\n    --device \/dev\/kfd \\\n    --device \/dev\/dri \\\n    --group-add video \\\n    --ipc=host \\\n    --cap-add=SYS_PTRACE \\\n    --security-opt seccomp=unconfined \\\n    -p 48700:8000 \\\n    -e HUGGING_FACE_HUB_TOKEN=\"your_hf_token_here\" \\\n    -e VLLM_USE_V1=0 \\\n    -v ~\/.cache\/huggingface:\/root\/.cache\/huggingface \\\n    --name qwen-server \\\n    rocm\/vllm:rocm6.4.1_vllm_0.10.1_20250909 \\\n    vllm serve Qwen\/Qwen2.5-14B-Instruct \\\n    --dtype float16 \\\n    --gpu-memory-utilization 0.90 \\\n    --max-model-len 32768 \\\n    --tensor-parallel-size 1 \\\n    --host 0.0.0.0 \\\n    --port 8000\n<\/pre>\n\n\n\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" src=\"https:\/\/tutorial.emka.web.id\/wp-content\/uploads\/sites\/19\/2026\/02\/qwen3-amd-gpu_1-1024x565.jpg\" alt=\"\" class=\"wp-image-1133\" \/><figcaption class=\"wp-element-caption\">Screenshot<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Why these flags matter<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong><code>VLLM_USE_V1=0<\/code><\/strong>: <strong>This is the most important line.<\/strong> The new &#8220;V1&#8221; engine in vLLM crashes on MI200 when loading JIT kernels. We force the legacy engine (V0) for rock-solid stability.<\/li>\n\n\n\n<li><strong><code>--dtype float16<\/code><\/strong>: We don&#8217;t trust <code>auto<\/code> mode on ROCm containers. Explicitly telling it to use float16 prevents initialization stalls.<\/li>\n\n\n\n<li><strong><code>--security-opt seccomp=unconfined<\/code><\/strong>: AMD GPUs need direct memory access that Docker blocks by default. Without this, you get permission errors.<\/li>\n\n\n\n<li><strong>The Model Size (14B)<\/strong>: We chose 14B because the 32B model (at float16) requires ~64GB of VRAM just for weights. On a single GPU, you&#8217;ll hit OOM (Out Of Memory) instantly once you add the KV cache. 14B sits in the &#8220;sweet spot&#8221;\u2014fast, smart, and leaves room for context.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Step 2: Testing the API<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Once the container says <code>Application startup complete<\/code>, your API is live on port <code>48700<\/code>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">You can test it with <code>curl<\/code>. <strong>Note:<\/strong> Be careful with your JSON syntax! Use straight quotes (<code>\"<\/code>), not curly smart quotes (<code>\u201c<\/code>), or the API will throw a 400 Bad Request error.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Here is a test command with a complex System Prompt (as requested in our logs):<\/p>\n\n\n\n<pre class=\"wp-block-syntaxhighlighter-code\">curl http:\/\/localhost:48700\/v1\/chat\/completions \\\n    -H \"Content-Type: application\/json\" \\\n    -H \"Authorization: Bearer any_token_is_fine\" \\\n    -d '{\n        \"model\": \"Qwen\/Qwen2.5-14B-Instruct\",\n        \"messages\": [\n            {\n                \"role\": \"system\",\n                \"content\": \"You are a helpful assistant. Please answer in JSON format.\"\n            },\n            {\n                \"role\": \"user\",\n                \"content\": \"Are you running on an AMD GPU?\"\n            }\n        ],\n        \"temperature\": 0.7\n    }'\n<\/pre>\n\n\n\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" src=\"https:\/\/tutorial.emka.web.id\/wp-content\/uploads\/sites\/19\/2026\/02\/qwen3-amd-gpu_2-1024x319.jpg\" alt=\"\" class=\"wp-image-1134\" \/><figcaption class=\"wp-element-caption\">Screenshot<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Note: If you are using a custom model name (like Qwen3), make sure the <code>\"model\"<\/code> field in your JSON matches exactly what you passed in the Docker command.<\/em><\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Troubleshooting<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>1. It hangs at &#8220;Loading model weights&#8230;&#8221;<\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Cause:<\/strong> ROCm is compiling kernels (JIT) for your specific GPU.<\/li>\n\n\n\n<li><strong>Fix:<\/strong> Wait. On the very first run, this can take 2\u20135 minutes. Subsequent runs will be instant.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>2. <code>RuntimeError: Engine core initialization failed<\/code><\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Cause:<\/strong> You forgot <code>VLLM_USE_V1=0<\/code>.<\/li>\n\n\n\n<li><strong>Fix:<\/strong> Add the env var. The V1 engine tries to use <code>aiter<\/code> libraries optimized for MI300, which segfault on MI200.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>3. <code>HIP out of memory<\/code><\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Cause:<\/strong> Your model is too fat.<\/li>\n\n\n\n<li><strong>Fix:<\/strong> If you absolutely need a 32B or 70B model, you must use <strong>Quantization<\/strong>. Change the docker command to use an AWQ model:Bash<code>vllm serve Qwen\/Qwen2.5-32B-Instruct-AWQ --quantization awq<\/code><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>If you are trying to run LLMs on AMD Instinct MI200 (MI210\/MI250) cards, you have probably already experienced the pain of &#8220;HSA errors,&#8221; random&#8230;<\/p>\n","protected":false},"author":27,"featured_media":1132,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-1131","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/emka.web.id\/en\/wp-json\/wp\/v2\/posts\/1131","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/emka.web.id\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/emka.web.id\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/emka.web.id\/en\/wp-json\/wp\/v2\/users\/27"}],"replies":[{"embeddable":true,"href":"https:\/\/emka.web.id\/en\/wp-json\/wp\/v2\/comments?post=1131"}],"version-history":[{"count":0,"href":"https:\/\/emka.web.id\/en\/wp-json\/wp\/v2\/posts\/1131\/revisions"}],"wp:attachment":[{"href":"https:\/\/emka.web.id\/en\/wp-json\/wp\/v2\/media?parent=1131"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/emka.web.id\/en\/wp-json\/wp\/v2\/categories?post=1131"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/emka.web.id\/en\/wp-json\/wp\/v2\/tags?post=1131"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}