2026年10月3日
cover_7900xtx_125b_1789563895884
Benchmark & deployment guide for Qwen 3.8 125B on dual AMD RX 7900 XTX (48GB VRAM). 38 tok/s daily API & 28 tok/s at 48K context with ROCm 7.14.

Qwen 3.8 125B MoE on Dual RX 7900 XTX: 48K Context at 28 Tok/s Benchmark & Guide

Single Source of Truth (SSOT) & Empirical Baseline:
All configuration flags, VRAM allocations, and benchmark results documented in this guide were gathered on September 16, 2026, from an identical physical test machine: AMD Ryzen 7 3700X (8C/16T), 2 × AMD Radeon RX 7900 XTX 24GB (gfx1100, 48GB VRAM total), 96GB physical DDR4 RAM (~94.2 GiB visible to OS), Ubuntu 24.04 LTS, Linux kernel 6.14, portable TheRock ROCm 7.14 SDK, and nasone32/llama.cpp-RDNA3-7900xtx-opt (Git commit 15995a12d1d530645a4f34c72afdaa30fa680149).
These numbers represent this exact hardware configuration, model quantization, software build, and test workload. They are not theoretical ceilings for all dual-7900 XTX systems, nor performance guarantees for other model architectures.

Deployment Profile Clarification:
FAST-50K (ctx-size=50144) serves as the production performance baseline; STABLE-64K (ctx-size=65536) is a verified on-demand configuration. The benchmark host temporarily ran ctx-size=32768 for daily single-concurrency API sweeps. This does not alter the historical 50K/64K A/B results or imply multi-profile automated traffic routing.

Host RAM Recommendations (64GB vs. 96GB):
All performance metrics were gathered on a 96GB system and cannot be claimed as “measured 64GB figures.” However, when running --parallel 1, --load-mode none, --lazy-mode on-direct, and without competing memory-heavy services, 64GB serves as a viable entry baseline for FAST-50K / 32K profiles. STABLE-64K should be regarded as 96GB recommended, requiring standalone validation on 64GB hosts.


📺 Video Walkthrough & Real-Hardware Benchmark

Watch the companion hardware walkthrough and real-time inference showcase:

YouTube Video Showcase

Watch on YouTube (2K Ultra HD · GeekLab AI)


Can two consumer-grade 24GB AMD Radeon RX 7900 XTX graphics cards turn a 90GB-class Qwen3.8-Flash-Next MoE GGUF model into an enterprise-grade, daily-usable local inference service?

The answer is yes.

However, this is not achieved by brutishly forcing 90GB of weights into 48GB of VRAM. Instead, it relies on an aggressive multi-tier architecture: Dual-GPU Layer Splitting + Host DDR4 RAM Tiering + Kernel Direct I/O Lazy Loading.

Before delving into the setup, let us clarify the headline metrics:
– All generation speeds refer to user-perceived Text Generation / Decode Throughput (TG), not prompt evaluation / prefill (PP).
– In daily single-concurrency API sweeps across prompt sizes of 512 to 8,192 tokens (32K context), median decode throughput reached 38.32 tok/s (P50).
– Under an intense stress benchmark with a fixed 48,009-token prompt and 512 output tokens, the FAST-50K baseline delivered a steady 27.98 tok/s decode speed and 467.03 tok/s prefill speed.
– The on-demand STABLE-64K profile completed the same 48K A/B test with 424.75 tok/s prefill and 24.49 tok/s decode.


Core Technology Stack:

Component Verified Specification Engineering Rationale
Inference Engine nasone32/llama.cpp-RDNA3-7900xtx-opt (15995a12…) Custom RDNA3 kernel optimizations. All parameters and benchmark scores in this guide are bound to this build.
ROCm Runtime TheRock ROCm 7.14 SDK (gfx1100) Isolated portable closure in /opt/qwen38/rocm-7.14. Leaves base system ROCm untouched to avoid dynamic library collision.
Base Model Qwen3.8-Flash-Next-UD-Q3_K_XL (3 GGUF splits) 176.9B total MoE parameters (~3B active per token). Sharded across dual VRAM, host RAM, and direct-IO lazy loading.
Speculative Draft Shared mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf Multi-Token Prediction (MTP) draft module with n-max=2, boosting generation speed by +30% to +40%.
KV Cache Compression Q8_0 / Q8_0 Quantized KV Cache Saves ~50% KV memory overhead with <1% prefill throughput penalty, unlocking stable 50K+ context within 48GB VRAM.
Baseline Profile FAST-50K (ctx-size=50144) Delivers ~9% faster prefill and ~12.5% higher generation throughput over 64K in 48K A/B stress tests.
Extended Profile STABLE-64K (ctx-size=65536) Pre-configured and tested for ultra-long context prompts (>48K tokens) via dedicated startup scripts.
flowchart TD
    Client["OpenAI-Compatible Client<br/>(Cursor / Claude Code / Cline / Cherry)"] -->|Local LAN HTTP /v1| Gateway["Auth Gateway / Reverse Proxy<br/>(Bearer Token & Rate Limiting)"]
    Gateway -->|Localhost Loopback :8082| Server["FAST-50K llama-server<br/>(ROCm 7.14 · gfx1100 · Flash Attention)"]

    subgraph Compute & Memory Topology
        Server --> MTP["MTP Speculative Draft<br/>(Shared Q8_0 · n-max 2)"]
        Server --> GPU0["GPU 0: RX 7900 XTX 24GB<br/>(~23.18 GB VRAM allocated)"]
        Server --> GPU1["GPU 1: RX 7900 XTX 24GB<br/>(~22.93 GB VRAM allocated)"]
        Server --> RAM["Host DDR4 RAM (~96GB)<br/>(27.5GB PLE Embeddings + Page Cache)"]
        Server --> SSD["NVMe PCIe 4.0 SSD<br/>(Kernel Direct I/O Lazy Loading)"]
    end

    GPU0 <-->|Q8_0 Wire All-Reduce| GPU1
    RAM -.->|Direct Memory Access| Server

⚡ 1. Real-World Benchmark Results

Rigorous 48K Long-Context A/B Stress Test

To establish an unshakeable ground truth, we conducted a controlled 3-round benchmark with identical input token IDs, fixed sampling seeds, and zero prefix caching (prompt_n=48009, predicted_n=512, cache_n=0, truncated=false):

Metric / Parameter FAST-50K (Baseline Profile) STABLE-64K (Extended Profile) Variance & Delta
Context Window (ctx-size) 50,144 tokens 65,536 tokens +15,392 tokens
Batch / Micro-Batch Size 4096 / 1024 4096 / 1024 Identical
KV Cache Quantization Q8_0 / Q8_0 Q8_0 / Q8_0 Identical
Speculative Decoding (MTP) Adaptive, n-max=2 Adaptive, n-max=2 Identical
VRAM Fit Target (fit-target) 3800,1024 5000,2048 Larger headroom for 64K scratch
48,009-Token Prefill (PP) 467.03 ± 1.32 tok/s 424.75 ± 1.66 tok/s FAST-50K is ~9.1% faster
512-Token Generation (TG) 27.98 ± 0.20 tok/s 24.49 ± 1.99 tok/s FAST-50K is ~12.5% faster
Total Wall-Clock Latency 121.08 seconds 134.06 seconds -12.98s (-9.7%)

Context Capacity Reality Check:
A 65,536-token window defines the total context boundary. Real-world prompts must reserve adequate budget for system prompts, conversation history, function schemas, and output tokens. It should not be conflated with unbounded “million-word processing.”


Daily 0.5K–8K API Sweeps vs. 48K Stress Tests

In addition to the 48K stress test, we ran single-concurrency OpenAI API sweeps with prompt lengths scaled incrementally from 512 to 8,192 tokens. Unlike the fixed 512-token output benchmark, these requests completed naturally with 119–155 generated tokens:

Benchmark Scenario Prompt Range Output Length Prefill Throughput (PP) Decode Throughput (TG)
API Sweep (FAST-50K Profile) 512 – 8,192 tokens 119 – 155 (Natural) Mean: 544.11 tok/s · P50: 583.64 tok/s Mean: 36.07 tok/s · P50: 34.73 tok/s
API Sweep (32K Profile) 512 – 8,192 tokens 119 – 155 (Natural) Mean: 536.45 tok/s · P50: 593.12 tok/s Mean: 38.05 tok/s · P50: 38.32 tok/s
Stress A/B (FAST-50K) 48,009 tokens 512 (Fixed) 467.03 tok/s 27.98 tok/s

Figure 1: Single-concurrency OpenAI API sweep on FAST-50K; prompts scale from 512 to 8,192 tokens.

Figure 1 | FAST-50K (50,144 context) API sweep. Prefill averaged 544.11 tok/s (P50 583.64 tok/s); generation averaged 36.07 tok/s (P50 34.73 tok/s).

Figure 2: Single-concurrency OpenAI API sweep under 32K context; prompts scale from 512 to 8,192 tokens.

Figure 2 | 32K context API sweep. Prefill P50 reached 593.12 tok/s; generation P50 reached 38.32 tok/s.

A rigorous 24,576-token input + fixed 256-token output test (cache_prompt=false) demonstrated that lowering context from 50K to 32K boosts mean PP from 526.54 to 534.68 tok/s (+1.5%), and TG from 29.53 to 30.39 tok/s (+2.9%).

Why does generation speed drop from 38 tok/s on short prompts to 28 tok/s on 48K prompts?
As the prompt grows, attention KV access costs scale accordingly. Generating each new token requires attending across a vastly larger history buffer. Furthermore, the first round of a short sweep involves CUDA graph capture, kernel warmup, and lazy loading. Quoting an isolated “tok/s” figure without specifying prompt size, output length, concurrency, and cache status is meaningless.


🛠️ 2. Hardware & Host Memory Sizing

Component Tested Platform Specification Technical Importance
GPU 2 × AMD Radeon RX 7900 XTX 24GB (gfx1100) Split across layers; handles attention, router, active experts, and scratch buffers.
Host RAM 96GB Physical DDR4 (94.2 GiB visible) Houses the 27.5GB embedding table (per_layer_token_embd.weight) + mmap file cache.
Inference Core llama.cpp-RDNA3-7900xtx-opt (15995a12…) Build-specific GEMM and memory optimizations for Navi 31 / gfx1100.
Isolated Runtime TheRock ROCm 7.14 SDK Self-contained user-space ROCm distribution; avoids altering system packages.
Main Weights Qwen3.8-Flash-Next-UD-Q3_K_XL (3 GGUF files) Unsloth dynamic quantization preserving perplexity at ~83.8 GiB.
MTP Draft mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf 1.1 GiB draft model predicting 2 speculative tokens per cycle.

Check system hardware and identify GPU nodes:

uname -a
free -h
df -h /
rocminfo | grep -E 'Name:.*gfx1100'
rocm-smi

Both cards must identify as gfx1100. If an integrated GPU (APU) exists, never assume discrete GPUs are 0,1. Execute llama-server --list-devices before declaring HIP_VISIBLE_DEVICES.

64GB vs. 96GB Host RAM Allocation

Host RAM Tier Target Workloads Operational Constraints
64GB 32K Context; FAST-50K (Single concurrency) Disable competing heavy services; lock --parallel 1; validate with 24K–48K prompts while monitoring swap.
96GB Recommended Baseline for FAST-50K; STABLE-64K Ample margin for host-resident experts, OS file page cache, and long-request spikes.
128GB+ STABLE-64K Long-Session Multi-Concurrency Ideal for long-term multi-tenant local AI infrastructure.

Although model files total ~90GB, mmap Direct I/O and dynamic expert routing prevent all 90GB from loading into RAM simultaneously. On 64GB systems, verify cold-start stability and monitor memory headroom:

free -h
swapon --show

During validation, system swap should remain flat, and free memory should cleanly recover after inference completes.


🚀 3. Step-by-Step Production Runbook

Step 1: Install Isolated ROCm 7.14 (Zero System Pollution)

AMD provides standalone portable multi-arch tarballs for TheRock ROCm 7.14:

mkdir -p ~/rocm714-download
cd ~/rocm714-download

wget https://repo.amd.com/rocm/tarball-multi-arch/therock-dist-linux-gfx110X-all-7.14.0.tar.gz

sudo mkdir -p /opt/qwen38/rocm-7.14
sudo chown -R "$USER":"$USER" /opt/qwen38
tar -xf therock-dist-linux-gfx110X-all-7.14.0.tar.gz -C /opt/qwen38/rocm-7.14

Export runtime paths inside deployment scripts (never rely on default shell .bashrc):

export ROCM_PATH=/opt/qwen38/rocm-7.14
export PATH="$ROCM_PATH/bin:$ROCM_PATH/lib/llvm/bin:$PATH"
export LD_LIBRARY_PATH="$ROCM_PATH/lib:$ROCM_PATH/lib/llvm/lib${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}"

Verify dynamic library binding to ensure zero leakage into system ROCm:

# Pre-launch check
env LD_LIBRARY_PATH=/opt/qwen38/rocm-7.14/lib:/opt/qwen38/rocm-7.14/lib/llvm/lib \
  ldd /opt/qwen38/src/llama.cpp-RDNA3-7900xtx-opt/build-rocm-gfx1100-portable/bin/llama-server \
  | grep -Ei 'hip|hsa|rocblas|hipblas'

# Post-launch process mapping audit
grep -E '/opt/qwen38/rocm-7\.14/.*(libhip|libhsa|libroc|libamd)' /proc/<LLAMA_PID>/maps

Step 2: Build RDNA3-Optimized llama.cpp

sudo apt update
sudo apt install -y git cmake ninja-build build-essential pkg-config curl wget python3 python3-pip libssl-dev

mkdir -p /opt/qwen38/{models,src,logs,benchmarks,conf,bin}
cd /opt/qwen38/src

git clone https://github.com/nasone32/llama.cpp-RDNA3-7900xtx-opt.git
cd llama.cpp-RDNA3-7900xtx-opt
git checkout 15995a12d1d530645a4f34c72afdaa30fa680149

cmake -S . -B build-rocm-gfx1100-portable -G Ninja \
  -DCMAKE_BUILD_TYPE=Release \
  -DCMAKE_C_COMPILER="$ROCM_PATH/lib/llvm/bin/clang" \
  -DCMAKE_CXX_COMPILER="$ROCM_PATH/lib/llvm/bin/clang++" \
  -DCMAKE_HIP_COMPILER="$ROCM_PATH/lib/llvm/bin/clang" \
  -DCMAKE_HIP_FLAGS='-mllvm --amdgpu-unroll-threshold-local=600' \
  -DGGML_HIP=ON \
  -DGGML_HIP_GRAPHS=ON \
  -DAMDGPU_TARGETS=gfx1100 \
  -DLLAMA_BUILD_TESTS=ON

cmake --build build-rocm-gfx1100-portable -j "$(nproc)"

Step 3: Download Qwen3.8 Flash Next & MTP Draft Weights

python3 -m venv ~/hf-cli
source ~/hf-cli/bin/activate
pip install -U huggingface_hub

mkdir -p /opt/qwen38/models/Qwen3.8-Flash-Next-GGUF/{UD-Q3_K_XL,MTP}

# Download 3-part base model
hf download unsloth/Qwen3.8-Flash-Next-GGUF \
  UD-Q3_K_XL/Qwen3.8-Flash-Next-UD-Q3_K_XL-00001-of-00003.gguf \
  UD-Q3_K_XL/Qwen3.8-Flash-Next-UD-Q3_K_XL-00002-of-00003.gguf \
  UD-Q3_K_XL/Qwen3.8-Flash-Next-UD-Q3_K_XL-00003-of-00003.gguf \
  --local-dir /opt/qwen38/models/Qwen3.8-Flash-Next-GGUF

# Download MTP draft speculative weights
hf download unsloth/Qwen3.8-Flash-Next-GGUF \
  MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf \
  --local-dir /opt/qwen38/models/Qwen3.8-Flash-Next-GGUF

deactivate

Note: Passing the path of 00001-of-00003.gguf to llama-server automatically loads subsequent parts.


Step 4: FAST-50K Production Launch Script

Save to /opt/qwen38/start-q3-fast50k.sh:

#!/usr/bin/env bash
set -euo pipefail

ROCM_PATH=/opt/qwen38/rocm-7.14
ROOT=/opt/qwen38/src/llama.cpp-RDNA3-7900xtx-opt
SERVER="$ROOT/build-rocm-gfx1100-portable/bin/llama-server"
MODEL=/opt/qwen38/models/Qwen3.8-Flash-Next-GGUF/UD-Q3_K_XL/Qwen3.8-Flash-Next-UD-Q3_K_XL-00001-of-00003.gguf
DRAFT=/opt/qwen38/models/Qwen3.8-Flash-Next-GGUF/MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf

export PATH="$ROCM_PATH/bin:$ROCM_PATH/lib/llvm/bin:$PATH"
export LD_LIBRARY_PATH="$ROCM_PATH/lib:$ROCM_PATH/lib/llvm/lib"
export HIP_VISIBLE_DEVICES=0,1

# ROCm & Inter-GPU All-Reduce Tuning
export GGML_CUDA_ALLREDUCE=internal
export GGML_CUDA_AR_WIRE=q8_0
export GGML_CUDA_AR_Q8_THRESHOLD=1048576
export GGML_CUDA_AR_FUSED_RESIDUAL=1
export GGML_CUDA_GDN_CHUNKED_BF16=1
export GGML_CUDA_AR_P2P=0

exec "$SERVER" \
  --model "$MODEL" \
  --model-draft "$DRAFT" \
  --alias qwen3.8-flash-next-fast50k,qwen3.8-default,qwen3.8 \
  --api-key-file /opt/qwen38/conf/api_keys.txt \
  --metrics \
  --host 127.0.0.1 --port 8082 \
  --ctx-size 50144 \
  --spec-type draft-mtp-adaptive \
  --spec-draft-n-min 1 \
  --spec-draft-n-min-adaptive 1 \
  --spec-draft-n-max 2 \
  --device-draft ROCm0 \
  --split-mode layer \
  --flash-attn on \
  --batch-size 4096 --ubatch-size 1024 \
  --moe-expert-cache 0 \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  --parallel 1 \
  --fit on --fit-target 3800,1024 \
  --load-mode none --lazy-mode on-direct \
  --temp 0.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
  --presence-penalty 0.0 \
  --reasoning on --reasoning-effort medium

Make the script executable:

chmod +x /opt/qwen38/start-q3-fast50k.sh

Step 5: STABLE-64K Extended Profile

Switching to 64K is not merely altering --ctx-size. To prevent out-of-memory errors during ggml_gallocr scratch graph allocations while sustaining ubatch=1024, adjust layer allocation parameters:

--ctx-size 65536
--fit-target 5000,2048

Save these parameters in a separate script (/opt/qwen38/start-q3-stable64k.sh). Maintain clean profile separation: gracefully stop FAST-50K, launch STABLE-64K, perform health probes, execute batch jobs, and revert to FAST-50K.


🔐 4. API Gateway, Security & SSE Streaming

Never expose the backend port 8082 directly to the local network:

  1. Bind to Localhost: llama-server listens strictly on 127.0.0.1:8082.
  2. Gateway Authentication: The frontend gateway (e.g., Nginx, Caddy, or a FastAPI reverse proxy) enforces client Bearer token authentication before relaying requests.
  3. SSE Streaming Integrity: In Server-Sent Events (stream=true), HTTP/1.1 proxies must pass chunked frames without dropping the final data: [DONE] sentinel.

Verify streaming completion via curl:

curl -N http://127.0.0.1:8082/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $(cat /opt/qwen38/conf/api_keys.txt | head -n 1)" \
  -d '{"model":"qwen3.8","messages":[{"role":"user","content":"Reply OK"}],"stream":true,"max_tokens":16}'

The terminal output must conclude cleanly with:

data: [DONE]

🧰 5. Systemd Production Hardening

Manage the service using systemd (/etc/systemd/system/qwen38-fast50k.service):

[Unit]
Description=Qwen3.8 Flash Next FAST-50K Production Service
After=network-online.target
Wants=network-online.target

[Service]
Type=exec
User=xin
Group=xin
SupplementaryGroups=render video
WorkingDirectory=/opt/qwen38
ExecStartPre=/opt/qwen38/bin/preflight-check.sh
ExecStart=/opt/qwen38/start-q3-fast50k.sh
ExecStartPost=/opt/qwen38/bin/wait-health.sh
Restart=on-failure
RestartSec=10
TimeoutStartSec=300
LimitNOFILE=65535
LimitMEMLOCK=infinity

[Install]
WantedBy=multi-user.target

Enable and start the service:

sudo systemctl daemon-reload
sudo systemctl enable --now qwen38-fast50k.service
sudo systemctl status qwen38-fast50k.service

✅ 6. Pre-Flight Acceptance Matrix

Before putting this setup into daily production, verify all 8 criteria:

  • [ ] 1. Isolated Library Verification: Running process dynamically links against /opt/qwen38/rocm-7.14/lib/ with 0 system ROCm leaks.
  • [ ] 2. Systemd Status: qwen38-fast50k.service is active (running) with auto-restart enabled.
  • [ ] 3. Health & Model APIs: curl http://127.0.0.1:8082/health and /v1/models return valid JSON.
  • [ ] 4. Streaming Termination: Streaming chat completions return a definitive data: [DONE] marker.
  • [ ] 5. Gateway LAN Isolation: Direct access to port 8082 from outside localhost is blocked by firewall/bind address.
  • [ ] 6. VRAM Stability Check: rocm-smi confirms ~23.18 GB allocated on GPU 0 and ~22.93 GB on GPU 1 with ~1GB scratch margin.
  • [ ] 7. Swap Zero-Thrashing: swapon --show indicates stable, non-accumulating swap usage during 48K prompt prefill.
  • [ ] 8. Profile Isolation: FAST-50K and STABLE-64K scripts maintain independent configurations and are never run concurrently.

Conclusion

The true value of this dual-Radeon RX 7900 XTX platform is not about chasing theoretical peak numbers. It is about establishing an empirically verified, battle-hardened local AI workstation:
– Defaulting to FAST-50K yields optimal latency (~38 tok/s P50 daily, ~28 tok/s at 48K) for day-to-day coding in Cursor, Claude Code, or autonomous agent loops.
– When an exceptionally large codebase or multi-chapter document must be digested in a single session, the system transitions smoothly to the pre-tested STABLE-64K configuration.

With isolated ROCm 7.14 runtimes, MTP speculative decoding, Q8 KV compression, and systemd service supervision, this setup delivers a viable local alternative to costly multi-card enterprise clusters.


References & Upstream Projects

About The Author

Leave a Reply

Your email address will not be published. Required fields are marked *