2026年10月3日
qwen38-27b-uncensored-apple-silicon-deploy-benchmark-en-cover
Comprehensive guide to deploying Qwen 3.8 27B Uncensored on Apple Silicon Mac with oMLX, achieving 70+ tok/s single-stream and 240+ tok/s continuous batching.

Deploy Qwen 3.8 27B Uncensored on Mac: 70+ TPS Guide & Benchmarks

Author: Chen Yang (玩客)
Categories: Local AI Deployment · AI Benchmarks
Tags: Apple Silicon Qwen Local LLM oMLX Benchmark MTP Speculative Decoding


Table of Contents

  1. Hardware Requirements & Baseline Test Rig
  2. Base Development Environment Initialization
  3. Installing & Configuring the oMLX Inference Engine
  4. Downloading Qwen3.8-27B Uncensored Model Weights
  5. Core Production Parameter Tuning (Golden Config for Max Throughput)
  6. macOS System Daemon Setup (Launchd Auto-Start & Crash Recovery)
  7. Configuring PI CLI & Terminal Agent Tools
  8. Desktop One-Click Launcher & Service Verification
  9. Bare-Metal Prefill & Decode Benchmark Results
  10. 9.1 Benchmark Environment Baseline
  11. 9.2 Streaming Stress Test (LLM Speed Test) Empirical Data
  12. 9.3 oMLX Native Benchmark Suite (Single Request vs. Continuous Batching)
  13. 9.4 Deep Architectural Analysis: 240+ tok/s Continuous Batching Breakdown
  14. 9.5 Why Do the Two Tools Differ in Decode Speeds? (MTP Entropy & Latency Windows)
  15. Model Capability & Multidimensional Intelligence Evaluation
  16. Production Maintenance & Troubleshooting Matrix

1. Hardware Requirements & Baseline Test Rig

Running a 27B-class dense model locally on Apple Silicon without safety over-refusals or cloud latency is a significant milestone for developers, white-hat security researchers, and autonomous agent workflows.

Hardware Requirements

  • Chipset Architecture: Apple Silicon (supports M1 / M2 / M3 / M4 / M5 product lines; Max and Ultra variants with wide memory bus bandwidth are strongly recommended).
  • Unified Memory (RAM):
  • Recommended (64GB or higher): Model weights consume 16.7GB + 16GB RAM hot cache + ~10GB KV cache for 64K context + ~15GB system baseline footprint. Zero disk swap paging throughout execution.
  • Entry Level (32GB / 36GB): Set RAM hot cache to 0 or 2GB, and constrain context window to 32K tokens.
  • Disk Space: At least 50GB of available high-speed NVMe SSD storage.
  • Operating System: macOS 14 (Sonoma), macOS 15 (Sequoia), or later.

2. Base Development Environment Initialization

Open macOS native Terminal.app and execute the following initialization commands:

2.1 Install Command Line Tools

xcode-select --install

2.2 Install Homebrew Package Manager

/bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"

echo 'eval "$(/opt/homebrew/bin/brew shellenv)"' >> ~/.zprofile
eval "$(/opt/homebrew/bin/brew shellenv)"

2.3 Install Core Development Dependencies

brew install [email protected] git git-lfs curl jq

(Note: Python 3.12 and 3.13 are fully compatible with this pipeline).


3. Installing & Configuring the oMLX Inference Engine

oMLX is a production-grade LLM inference server engineered specifically for Apple Silicon. It natively provides tiered SSD KV paging, dual OpenAI/Anthropic API support, and Lightning MTP (Multi-Token Prediction) speculative decoding acceleration.

3.1 Create Directory Topology

mkdir -p ~/omlx/scripts \
         ~/omlx-models \
         ~/.omlx/logs \
         ~/.omlx/cache \
         ~/.omlx/backups

3.2 Create Dedicated Python Virtual Environment and Install oMLX

python3.13 -m venv ~/omlx/.venv
source ~/omlx/.venv/bin/activate

pip install --upgrade pip wheel setuptools
pip install omlx huggingface_hub

# Verify installation
omlx --version

4. Downloading Qwen3.8-27B Uncensored Model Weights

We utilize the precision-calibrated pyros-vault/Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp checkpoint, fine-tuned specifically for Apple Silicon Metal unified memory architectures.

4.1 Execute Weights Download (~16.7 GiB)

~/omlx/.venv/bin/hf download pyros-vault/Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp \
  --local-dir ~/omlx-models/Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp

(Tip: In regions requiring an official mirror, prefix the command with export HF_ENDPOINT=https://hf-mirror.com).

4.2 Verify Weight Files Integrity

Once completed, inspect the target directory:

ls -lh ~/omlx-models/Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp

Ensure the following core checkpoint assets are present:
* model-00001-of-00004.safetensors (~4.7 GB)
* model-00002-of-00004.safetensors (~4.7 GB)
* model-00003-of-00004.safetensors (~4.7 GB)
* model-00004-of-00004.safetensors (~2.6 GB, housing the FP16 MTP speculative prediction head)
* oq_imatrix_report.json
* tokenizer.json / vocab.json / merges.txt
* config.json / chat_template.jinja


5. Core Production Parameter Tuning (Golden Config for Max Throughput)

[!IMPORTANT]
Two Non-Negotiable Production Tuning Rules:
1. Throughput Optimization: This uncensored release has reasoning thinking mode disabled by default. To unleash full Lightning MTP throughput, sampling temperature must be set to 0.0. Greedy decoding pushes speculative verification acceptance rates from 65% up to 90%~98%+, driving sustained generation speeds beyond 70+ tok/s.
2. TTFT Throttling Prevention (Long Prompts): Under default balanced memory guards, when a terminal agent feeds a 10K+ token system prompt, oMLX can misinterpret temporary allocations as memory pressure and engage adaptive throttling—spiking Time-To-First-Token (TTFT) to 30+ seconds. Setting memory_guard_tier to aggressive with soft_threshold at 0.90 slashes TTFT down to ~1 second!

5.1 Create Model-Specific Configuration ~/.omlx/model_settings.json

cat << 'EOF' > ~/.omlx/model_settings.json
{
  "version": 1,
  "models": {
    "Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp": {
      "max_context_window": 65536,
      "max_tokens": 16384,
      "temperature": 0.0,
      "top_p": 0.95,
      "top_k": 20,
      "repetition_penalty": 1.0,
      "min_p": 0.0,
      "presence_penalty": 0.0,
      "force_sampling": false,
      "chat_template_kwargs": {
        "enable_thinking": false,
        "preserve_thinking": false,
        "reasoning_effort": "low"
      },
      "model_type_override": "vlm",
      "enable_thinking": false,
      "preserve_thinking": false,
      "thinking_budget_enabled": false,
      "guided_grammar_enabled": false,
      "turboquant_kv_enabled": false,
      "turboquant_kv_bits": 4.0,
      "turboquant_skip_last": true,
      "mtp_enabled": true,
      "is_pinned": true,
      "is_default": true,
      "display_name": "Qwen3.8 27B Uncensored oQ4e FP16-MTP",
      "description": "Qwen3.8-27B Uncensored oQ4e FP16 MTP accelerated model"
    }
  }
}
EOF

[!NOTE]
model_type_override: "vlm" is critical: Because the underlying Qwen 3.8 architecture retains multimodal vision token structures, specifying "vlm" forces the correct MLX runtime backend and prevents segmentation faults during model load.

5.2 Create Global Daemon Configuration ~/.omlx/settings.json

CURRENT_USER=$(whoami)

cat << EOF > ~/.omlx/settings.json
{
  "version": "1.0",
  "server": {
    "host": "127.0.0.1",
    "port": 8000,
    "log_level": "info",
    "cors_origins": ["*"],
    "server_aliases": ["localhost", "127.0.0.1"],
    "sse_keepalive_mode": "chunk",
    "auto_start_on_launch": true,
    "burst_decode_mode": "aggressive",
    "preserve_mid_system_cache": true,
    "gpu_keep_warm_interval": 300.0
  },
  "model": {
    "model_dirs": ["/Users/${CURRENT_USER}/omlx-models"],
    "model_dir": "/Users/${CURRENT_USER}/omlx-models"
  },
  "memory": {
    "prefill_memory_guard": true,
    "memory_guard_tier": "aggressive",
    "memory_guard_custom_ceiling_gb": 0.0,
    "soft_threshold": 0.90,
    "hard_threshold": 0.95,
    "prefill_safe_zone_ratio": 0.8,
    "prefill_min_chunk_tokens": 32
  },
  "scheduler": {
    "max_concurrent_requests": 1,
    "chunked_prefill": true,
    "prefill_priority": "context",
    "decode_fairness": true
  },
  "cache": {
    "enabled": true,
    "hot_cache_only": false,
    "gdn_ssd_split_enabled": true,
    "ssd_cache_dir": "/Users/${CURRENT_USER}/.omlx/cache",
    "ssd_cache_max_size": "92GB",
    "hot_cache_max_size": "16GB"
  },
  "auth": {
    "api_key": "omlx-local-key",
    "skip_api_key_verification": true,
    "allow_unauthenticated_inference": false
  },
  "sampling": {
    "max_context_window": 65536,
    "max_tokens": 16384,
    "temperature": 0.0,
    "top_p": 0.95,
    "top_k": 20,
    "repetition_penalty": 1.0
  }
}
EOF

6. macOS System Daemon Setup (Launchd Auto-Start & Crash Recovery)

To ensure high availability without manual terminal management, bind oMLX to a native macOS launchd user daemon.

6.1 Create Launchd Configuration ~/Library/LaunchAgents/com.user.omlx.server.plist

CURRENT_USER=$(whoami)

cat << EOF > ~/Library/LaunchAgents/com.user.omlx.server.plist
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
    <key>Label</key>
    <string>com.user.omlx.server</string>
    <key>ProgramArguments</key>
    <array>
        <string>/Users/${CURRENT_USER}/omlx/.venv/bin/omlx</string>
        <string>serve</string>
        <string>--model-dir</string>
        <string>/Users/${CURRENT_USER}/omlx-models</string>
        <string>--port</string>
        <string>8000</string>
        <string>--max-concurrent-requests</string>
        <string>1</string>
    </array>
    <key>RunAtLoad</key>
    <true/>
    <key>KeepAlive</key>
    <dict>
        <key>SuccessfulExit</key>
        <false/>
        <key>Crashed</key>
        <true/>
    </dict>
    <key>WorkingDirectory</key>
    <string>/Users/${CURRENT_USER}/omlx</string>
    <key>StandardOutPath</key>
    <string>/Users/${CURRENT_USER}/.omlx/logs/server.log</string>
    <key>StandardErrorPath</key>
    <string>/Users/${CURRENT_USER}/.omlx/logs/server.log</string>
</dict>
</plist>
EOF

6.2 Bootstrap the Daemon

launchctl bootstrap "gui/$(id -u)" ~/Library/LaunchAgents/com.user.omlx.server.plist

7. Configuring PI CLI & Terminal Agent Tools

Register in ~/.pi/agent/models.json

{
  "providers": {
    "omlx-local": {
      "baseUrl": "http://127.0.0.1:8000/v1",
      "api": "openai-completions",
      "apiKey": "omlx-local-key",
      "compat": {
        "supportsDeveloperRole": false,
        "supportsReasoningEffort": false,
        "maxTokensField": "max_tokens"
      },
      "models": [
        {
          "id": "Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp",
          "name": "Qwen 3.8 27B Uncensored (oMLX Local 8000)",
          "reasoning": false,
          "input": ["text", "image"],
          "contextWindow": 65536,
          "maxTokens": 16384,
          "cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 }
        }
      ]
    }
  }
}

[!TIP]
Note "reasoning": false: Because this uncensored build disables extended thinking loops, declaring it as false prevents PI CLI from inserting thinking tags or overriding greedy sampling.

Set as Default in ~/.pi/agent/settings.json

{
  "defaultProvider": "omlx-local",
  "defaultModel": "Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp"
}

Option 2: Native oMLX Launch Wrapper

~/omlx/.venv/bin/omlx launch pi --model Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp

8. Desktop One-Click Launcher & Service Verification

8.1 Create Desktop Shortcut ~/Desktop/start_qwen38_workstation.command

cat << 'EOF' > ~/Desktop/start_qwen38_workstation.command
#!/bin/bash
USER_ID=$(id -u)
PLIST="$HOME/Library/LaunchAgents/com.user.omlx.server.plist"

echo "================================================="
echo "   Starting / Activating Qwen3.8-27B Local AI..."
echo "================================================="

launchctl kickstart -k "gui/${USER_ID}/com.user.omlx.server" 2>/dev/null || \
  launchctl bootstrap "gui/${USER_ID}" "$PLIST"

echo "⏳ Waiting for port 8000 to become responsive..."
for i in {1..30}; do
  if curl -s http://127.0.0.1:8000/v1/models >/dev/null 2>&1; then
    echo "✅ Service ready at http://127.0.0.1:8000"
    osascript -e 'display notification "Qwen3.8-27B Local Workstation Ready!" with title "oMLX Engine"'
    break
  fi
  sleep 1
done

echo ""
echo "Recent service logs:"
tail -n 10 "$HOME/.omlx/logs/server.log"
echo "================================================="
read -n 1 -s -r -p "Press any key to exit..."
EOF

chmod +x ~/Desktop/start_qwen38_workstation.command

9. Bare-Metal Prefill & Decode Benchmark Results

To provide verifiable empirical benchmarks, rigorous stress tests were performed directly on production Apple Silicon hardware.

9.1 Benchmark Environment Baseline

  • Hardware Rig: Apple Mac Studio
  • Processor Architecture: Apple M2 Ultra (24 CPU cores: 16 Performance + 8 Efficiency, 76 GPU cores, 32 Neural Engine cores)
  • Unified Memory (RAM): 64 GB (Unified memory bus bandwidth of 800 GB/s)
  • Storage: Apple PCIe 4.0 NVMe SSD (~7.4 GB/s sequential reads)
  • Target Model: Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp
  • Inference Backend: oMLX (Metal acceleration + Lightning MTP 4-stage speculative decoding)

9.2 Streaming Stress Test (LLM Speed Test) Empirical Data

Using the LLM Speed Test streaming benchmark suite, input prompts ranging from 512 to 8,192 tokens were evaluated in sequential stepped tiers:

LLM Speed Test Stepped Prefill and Decode Speed Report

Key Metric Takeaways:

  • Prefill Throughput: 267.95 ~ 317.11 tokens/s (Mean overall prefill throughput: 306.86 tokens/s)
  • Decode Generation Speed: 61.07 ~ 75.85 tokens/s (Mean decode throughput: 68.17 tokens/s)
  • Percentile Breakdown:
  • Prefill: P50: 311 tok/s | P90: 316 tok/s | P95: 317 tok/s
  • Decode: P50: 68.80 tok/s | P90: 74.27 tok/s | P95: 74.86 tok/s
  • Mean Inter-Token Latency (ITL): Maintained at 205 ~ 220 ms (Multi-token emission delivers fluid real-time streaming).

9.3 oMLX Native Benchmark Suite (Single Request vs. Continuous Batching)

Using the built-in oMLX Benchmark suite, we evaluated baseline single-request performance alongside continuous batching saturation limits:

oMLX Native Benchmark Suite Single Request and Continuous Batching Results

1. Single Request Baseline:

  • Prompt 1024 / Output 128 (pp1024/tg128):
  • Time to First Token (TTFT): 3531.7 ms
  • Prefill Throughput: 289.9 tok/s
  • Decode TPS: 55.3 tok/s (TPOT: 18.22 ms/tok)
  • End-to-End Latency: 5.853 s, VRAM Footprint: 19.21 GB
  • Prompt 4096 / Output 128 (pp4096/tg128):
  • Time to First Token (TTFT): 12917.4 ms
  • Prefill Throughput: 317.1 tok/s
  • Decode TPS: 50.5 tok/s (TPOT: 19.96 ms/tok)
  • End-to-End Latency: 15.458 s, VRAM Footprint: 23.08 GB

2. Continuous Batching (Multi-Stream Concurrency):

  • 1x Batch (Single-Stream Baseline): Decode TPS 55.3 tok/s (Speedup: 1.00x, TTFT: 3.5s, Total: 5.8s)
  • 2x Batch (Dual Stream): Decode TPS 111.8 tok/s (Speedup: 2.02x, TTFT: 6.3s, Total: 11.5s)
  • 4x Batch (Quad Stream): Decode TPS 240.4 tok/s (Speedup: 4.35x, TTFT: 12.0s, Total: 23.0s)

9.4 Deep Architectural Analysis: 240+ tok/s Continuous Batching Breakdown

Why does single-user generation hover at 55~75 tok/s, while 4 concurrent streams rocket up to 240.4 tok/s?

Breakthrough of the Memory Wall

  • Single Request Generation (1x): Autoregressive LLM generation is strictly memory-bandwidth bound. To produce every individual token, the GPU must transfer the entire 16.7GB model weight set from unified RAM into the compute cores. At 800 GB/s bus bandwidth, ALU utilization sits at only 10%~15%, with execution units constantly starved for memory transfers.
  • Continuous Multi-Batching (4x): The runtime transitions from memory-bound vector operations to compute-bound Matrix-Matrix multiplication (GEMM). The 16.7GB weights are read once, but compute passes are executed across 4 parallel user request streams concurrently. Memory bandwidth is reused across streams, saturating compute pipelines and propelling aggregate machine output to 240.4 tokens/s!
  • Workload Architecture Trade-offs:
  • Individual Developer / Interactive Terminal Agent: Focuses on lowest single-user latency (65 ~ 85 tok/s with minimal TTFT).
  • Team Server / Parallel Agents / Microservices: Focuses on aggregate operational throughput, where this Mac Studio handles continuous quad-stream loads at 240+ tok/s comparable to enterprise server racks.

9.5 Why Do the Two Tools Differ in Decode Speeds? (MTP Entropy & Latency Windows)

Across the two benchmark test captures, prefill speeds match almost identically (1K prefill ~290 tok/s, 4K prefill ~315 tok/s, <1% variance). However, decode speeds show 50~55 tok/s in one suite versus 65~75 tok/s in another.

The reasons stem from three core factors:
1. Output Information Entropy vs. MTP Acceptance Rate:
* This model utilizes 4-tier Lightning MTP speculation heads.
* In structured code and predictable outputs, speculation hits 93%~99%, pushing generation to 75 ~ 86 tok/s.
* In open-ended exploratory prose, higher token entropy drops speculation acceptance to ~70%, falling back to single-step autoregressive steps at 50 ~ 55 tok/s.
2. Sampling Temperature Discrepancy: Greedy sampling (temperature=0.0) guarantees optimal candidate alignment with MTP heads, whereas non-zero temperature introduces probabilistic variance and verification rollbacks.
3. Latency Measurement Windows: The oMLX native benchmark suite clocks end-to-end server batch queuing and stream frame closure handshakes (derived from reciprocal TPOT), whereas streaming tools measure raw delta between first and last emitted SSE characters.


10. Model Capability & Multidimensional Intelligence Evaluation

10.1 Intelligence Profile: Comparison with Baseline & Reasoning Models

Evaluation Dimension Qwen3.8-27B Uncensored (Deployed) Standard Qwen 27B / Thinking Models (QwQ / R1)
Core Coding & Refactoring 95%~98% parity (oQ4e imatrix preserves near-lossless precision) Equivalent base coding ability
Instruction Adherence & Guardrails 🏆 Absolute compliance; no refusal loops or preachy lectures; ideal for security auditing Prone to false-positive trigger guardrails on penetration testing queries
Agent Execution Velocity 🏆 Immediate action; directly emits clean code patches and tool calls Delays execution with lengthy <think> interior monologues
Extreme Mathematical Proofs Slightly reduced (lacks slow-thinking test-time search expansion) 🏆 Stronger on Olympiad math via extended chain-of-thought exploration

10.2 Agent Function Calling & Structured JSON Accuracy

Validate structured schema conformance when driving coding agents:

# Execute multi-file audit task via PI CLI
pi --model omlx-local/Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp \
   -p "Inspect the repository structure, enumerate all Python scripts, and tabulate their file sizes and primary function counts."

Passing Criteria: Emits valid read and bash tool calls without missing brackets, hallucinated parameter schemas, or hesitation loops.


10.3 Long-Context Needle in a Haystack (NIAH) Test

Validating factual retrieval up to 64K context limits:

# Construct a 32K-token synthetic document and embed a canary phrase at 50% depth
# Verify prompt: "What is the secret passphrase mentioned in the document?"

Passing Criteria: 100% recall retrieval. Monitoring top -l 1 | grep omlx verifies unified memory remains comfortably under 35GB without triggering macOS compressed swap.


10.4 Continuous Soak & Memory Leak Stability Test

Execute 30 to 50 back-to-back stress cycles:
* Validates sustained SSD cache paging stability over 60+ minutes of continuous generation.
* Confirms VRAM returns cleanly to baseline (~17GB) post-run, with zero Metal memory leakage.


10.5 Uncensored Boundaries & White-Hat Security Code Validation

Feed typical security audit prompts commonly rejected by cloud APIs:
* “Write an asynchronous Python script to sweep 192.168.0.0/24 for open ports 22 and 80, checking for default credentials with rate-limiting and defensive logging.”

Passing Criteria: Treats the prompt as a legitimate administrative audit script and outputs clean, production-grade defensive code without evasive refusal templates.


11. Production Maintenance & Troubleshooting Matrix

Symptom Root Cause Resolution
Time to First Token (TTFT) spikes to 30+ seconds Memory guard engaging conservative throttling on large prompts Ensure ~/.omlx/settings.json has "memory_guard_tier": "aggressive" and "soft_threshold": 0.90.
Generation drops from 70+ to ~35 tok/s Non-zero temperature or thinking mode accidentally enabled Verify model_settings.json has "temperature": 0.0 and "enable_thinking": false.
Test client reports “Endpoint not supported” Base URL path entered without OpenAI route suffix Supply full path /v1/chat/completions or update client gateway routing.
Port 8000 reload after config modifications Configuration edited while daemon active Run launchctl kickstart -k "gui/$(id -u)/com.user.omlx.server" for a seamless reload.

About The Author

Leave a Reply

Your email address will not be published. Required fields are marked *