Deploy Qwen 3.8 27B Uncensored on Mac: 70+ TPS Guide & Benchmarks
Author: Chen Yang (玩客)
Categories: Local AI Deployment · AI Benchmarks
Tags:Apple SiliconQwenLocal LLMoMLXBenchmarkMTP Speculative Decoding
Table of Contents
- Hardware Requirements & Baseline Test Rig
- Base Development Environment Initialization
- Installing & Configuring the oMLX Inference Engine
- Downloading Qwen3.8-27B Uncensored Model Weights
- Core Production Parameter Tuning (Golden Config for Max Throughput)
- macOS System Daemon Setup (Launchd Auto-Start & Crash Recovery)
- Configuring PI CLI & Terminal Agent Tools
- Desktop One-Click Launcher & Service Verification
- Bare-Metal Prefill & Decode Benchmark Results
- 9.1 Benchmark Environment Baseline
- 9.2 Streaming Stress Test (LLM Speed Test) Empirical Data
- 9.3 oMLX Native Benchmark Suite (Single Request vs. Continuous Batching)
- 9.4 Deep Architectural Analysis: 240+ tok/s Continuous Batching Breakdown
- 9.5 Why Do the Two Tools Differ in Decode Speeds? (MTP Entropy & Latency Windows)
- Model Capability & Multidimensional Intelligence Evaluation
- Production Maintenance & Troubleshooting Matrix
1. Hardware Requirements & Baseline Test Rig
Running a 27B-class dense model locally on Apple Silicon without safety over-refusals or cloud latency is a significant milestone for developers, white-hat security researchers, and autonomous agent workflows.
Hardware Requirements
- Chipset Architecture: Apple Silicon (supports M1 / M2 / M3 / M4 / M5 product lines; Max and Ultra variants with wide memory bus bandwidth are strongly recommended).
- Unified Memory (RAM):
- Recommended (64GB or higher): Model weights consume 16.7GB + 16GB RAM hot cache + ~10GB KV cache for 64K context + ~15GB system baseline footprint. Zero disk swap paging throughout execution.
- Entry Level (32GB / 36GB): Set RAM hot cache to
0or2GB, and constrain context window to 32K tokens. - Disk Space: At least 50GB of available high-speed NVMe SSD storage.
- Operating System: macOS 14 (Sonoma), macOS 15 (Sequoia), or later.
2. Base Development Environment Initialization
Open macOS native Terminal.app and execute the following initialization commands:
2.1 Install Command Line Tools
xcode-select --install
2.2 Install Homebrew Package Manager
/bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"
echo 'eval "$(/opt/homebrew/bin/brew shellenv)"' >> ~/.zprofile
eval "$(/opt/homebrew/bin/brew shellenv)"
2.3 Install Core Development Dependencies
brew install [email protected] git git-lfs curl jq
(Note: Python 3.12 and 3.13 are fully compatible with this pipeline).
3. Installing & Configuring the oMLX Inference Engine
oMLX is a production-grade LLM inference server engineered specifically for Apple Silicon. It natively provides tiered SSD KV paging, dual OpenAI/Anthropic API support, and Lightning MTP (Multi-Token Prediction) speculative decoding acceleration.
3.1 Create Directory Topology
mkdir -p ~/omlx/scripts \
~/omlx-models \
~/.omlx/logs \
~/.omlx/cache \
~/.omlx/backups
3.2 Create Dedicated Python Virtual Environment and Install oMLX
python3.13 -m venv ~/omlx/.venv
source ~/omlx/.venv/bin/activate
pip install --upgrade pip wheel setuptools
pip install omlx huggingface_hub
# Verify installation
omlx --version
4. Downloading Qwen3.8-27B Uncensored Model Weights
We utilize the precision-calibrated pyros-vault/Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp checkpoint, fine-tuned specifically for Apple Silicon Metal unified memory architectures.
4.1 Execute Weights Download (~16.7 GiB)
~/omlx/.venv/bin/hf download pyros-vault/Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp \
--local-dir ~/omlx-models/Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp
(Tip: In regions requiring an official mirror, prefix the command with export HF_ENDPOINT=https://hf-mirror.com).
4.2 Verify Weight Files Integrity
Once completed, inspect the target directory:
ls -lh ~/omlx-models/Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp
Ensure the following core checkpoint assets are present:
* model-00001-of-00004.safetensors (~4.7 GB)
* model-00002-of-00004.safetensors (~4.7 GB)
* model-00003-of-00004.safetensors (~4.7 GB)
* model-00004-of-00004.safetensors (~2.6 GB, housing the FP16 MTP speculative prediction head)
* oq_imatrix_report.json
* tokenizer.json / vocab.json / merges.txt
* config.json / chat_template.jinja
5. Core Production Parameter Tuning (Golden Config for Max Throughput)
[!IMPORTANT]
Two Non-Negotiable Production Tuning Rules:
1. Throughput Optimization: This uncensored release has reasoning thinking mode disabled by default. To unleash full Lightning MTP throughput, sampling temperature must be set to0.0. Greedy decoding pushes speculative verification acceptance rates from 65% up to 90%~98%+, driving sustained generation speeds beyond 70+ tok/s.
2. TTFT Throttling Prevention (Long Prompts): Under defaultbalancedmemory guards, when a terminal agent feeds a 10K+ token system prompt, oMLX can misinterpret temporary allocations as memory pressure and engage adaptive throttling—spiking Time-To-First-Token (TTFT) to 30+ seconds. Settingmemory_guard_tiertoaggressivewithsoft_thresholdat0.90slashes TTFT down to ~1 second!
5.1 Create Model-Specific Configuration ~/.omlx/model_settings.json
cat << 'EOF' > ~/.omlx/model_settings.json
{
"version": 1,
"models": {
"Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp": {
"max_context_window": 65536,
"max_tokens": 16384,
"temperature": 0.0,
"top_p": 0.95,
"top_k": 20,
"repetition_penalty": 1.0,
"min_p": 0.0,
"presence_penalty": 0.0,
"force_sampling": false,
"chat_template_kwargs": {
"enable_thinking": false,
"preserve_thinking": false,
"reasoning_effort": "low"
},
"model_type_override": "vlm",
"enable_thinking": false,
"preserve_thinking": false,
"thinking_budget_enabled": false,
"guided_grammar_enabled": false,
"turboquant_kv_enabled": false,
"turboquant_kv_bits": 4.0,
"turboquant_skip_last": true,
"mtp_enabled": true,
"is_pinned": true,
"is_default": true,
"display_name": "Qwen3.8 27B Uncensored oQ4e FP16-MTP",
"description": "Qwen3.8-27B Uncensored oQ4e FP16 MTP accelerated model"
}
}
}
EOF
[!NOTE]
model_type_override: "vlm"is critical: Because the underlying Qwen 3.8 architecture retains multimodal vision token structures, specifying"vlm"forces the correct MLX runtime backend and prevents segmentation faults during model load.
5.2 Create Global Daemon Configuration ~/.omlx/settings.json
CURRENT_USER=$(whoami)
cat << EOF > ~/.omlx/settings.json
{
"version": "1.0",
"server": {
"host": "127.0.0.1",
"port": 8000,
"log_level": "info",
"cors_origins": ["*"],
"server_aliases": ["localhost", "127.0.0.1"],
"sse_keepalive_mode": "chunk",
"auto_start_on_launch": true,
"burst_decode_mode": "aggressive",
"preserve_mid_system_cache": true,
"gpu_keep_warm_interval": 300.0
},
"model": {
"model_dirs": ["/Users/${CURRENT_USER}/omlx-models"],
"model_dir": "/Users/${CURRENT_USER}/omlx-models"
},
"memory": {
"prefill_memory_guard": true,
"memory_guard_tier": "aggressive",
"memory_guard_custom_ceiling_gb": 0.0,
"soft_threshold": 0.90,
"hard_threshold": 0.95,
"prefill_safe_zone_ratio": 0.8,
"prefill_min_chunk_tokens": 32
},
"scheduler": {
"max_concurrent_requests": 1,
"chunked_prefill": true,
"prefill_priority": "context",
"decode_fairness": true
},
"cache": {
"enabled": true,
"hot_cache_only": false,
"gdn_ssd_split_enabled": true,
"ssd_cache_dir": "/Users/${CURRENT_USER}/.omlx/cache",
"ssd_cache_max_size": "92GB",
"hot_cache_max_size": "16GB"
},
"auth": {
"api_key": "omlx-local-key",
"skip_api_key_verification": true,
"allow_unauthenticated_inference": false
},
"sampling": {
"max_context_window": 65536,
"max_tokens": 16384,
"temperature": 0.0,
"top_p": 0.95,
"top_k": 20,
"repetition_penalty": 1.0
}
}
EOF
6. macOS System Daemon Setup (Launchd Auto-Start & Crash Recovery)
To ensure high availability without manual terminal management, bind oMLX to a native macOS launchd user daemon.
6.1 Create Launchd Configuration ~/Library/LaunchAgents/com.user.omlx.server.plist
CURRENT_USER=$(whoami)
cat << EOF > ~/Library/LaunchAgents/com.user.omlx.server.plist
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>com.user.omlx.server</string>
<key>ProgramArguments</key>
<array>
<string>/Users/${CURRENT_USER}/omlx/.venv/bin/omlx</string>
<string>serve</string>
<string>--model-dir</string>
<string>/Users/${CURRENT_USER}/omlx-models</string>
<string>--port</string>
<string>8000</string>
<string>--max-concurrent-requests</string>
<string>1</string>
</array>
<key>RunAtLoad</key>
<true/>
<key>KeepAlive</key>
<dict>
<key>SuccessfulExit</key>
<false/>
<key>Crashed</key>
<true/>
</dict>
<key>WorkingDirectory</key>
<string>/Users/${CURRENT_USER}/omlx</string>
<key>StandardOutPath</key>
<string>/Users/${CURRENT_USER}/.omlx/logs/server.log</string>
<key>StandardErrorPath</key>
<string>/Users/${CURRENT_USER}/.omlx/logs/server.log</string>
</dict>
</plist>
EOF
6.2 Bootstrap the Daemon
launchctl bootstrap "gui/$(id -u)" ~/Library/LaunchAgents/com.user.omlx.server.plist
7. Configuring PI CLI & Terminal Agent Tools
Option 1: Standard Model Dictionary (Recommended)
Register in ~/.pi/agent/models.json
{
"providers": {
"omlx-local": {
"baseUrl": "http://127.0.0.1:8000/v1",
"api": "openai-completions",
"apiKey": "omlx-local-key",
"compat": {
"supportsDeveloperRole": false,
"supportsReasoningEffort": false,
"maxTokensField": "max_tokens"
},
"models": [
{
"id": "Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp",
"name": "Qwen 3.8 27B Uncensored (oMLX Local 8000)",
"reasoning": false,
"input": ["text", "image"],
"contextWindow": 65536,
"maxTokens": 16384,
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 }
}
]
}
}
}
[!TIP]
Note"reasoning": false: Because this uncensored build disables extended thinking loops, declaring it asfalseprevents PI CLI from inserting thinking tags or overriding greedy sampling.
Set as Default in ~/.pi/agent/settings.json
{
"defaultProvider": "omlx-local",
"defaultModel": "Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp"
}
Option 2: Native oMLX Launch Wrapper
~/omlx/.venv/bin/omlx launch pi --model Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp
8. Desktop One-Click Launcher & Service Verification
8.1 Create Desktop Shortcut ~/Desktop/start_qwen38_workstation.command
cat << 'EOF' > ~/Desktop/start_qwen38_workstation.command
#!/bin/bash
USER_ID=$(id -u)
PLIST="$HOME/Library/LaunchAgents/com.user.omlx.server.plist"
echo "================================================="
echo " Starting / Activating Qwen3.8-27B Local AI..."
echo "================================================="
launchctl kickstart -k "gui/${USER_ID}/com.user.omlx.server" 2>/dev/null || \
launchctl bootstrap "gui/${USER_ID}" "$PLIST"
echo "⏳ Waiting for port 8000 to become responsive..."
for i in {1..30}; do
if curl -s http://127.0.0.1:8000/v1/models >/dev/null 2>&1; then
echo "✅ Service ready at http://127.0.0.1:8000"
osascript -e 'display notification "Qwen3.8-27B Local Workstation Ready!" with title "oMLX Engine"'
break
fi
sleep 1
done
echo ""
echo "Recent service logs:"
tail -n 10 "$HOME/.omlx/logs/server.log"
echo "================================================="
read -n 1 -s -r -p "Press any key to exit..."
EOF
chmod +x ~/Desktop/start_qwen38_workstation.command
9. Bare-Metal Prefill & Decode Benchmark Results
To provide verifiable empirical benchmarks, rigorous stress tests were performed directly on production Apple Silicon hardware.
9.1 Benchmark Environment Baseline
- Hardware Rig: Apple Mac Studio
- Processor Architecture: Apple M2 Ultra (24 CPU cores: 16 Performance + 8 Efficiency, 76 GPU cores, 32 Neural Engine cores)
- Unified Memory (RAM): 64 GB (Unified memory bus bandwidth of 800 GB/s)
- Storage: Apple PCIe 4.0 NVMe SSD (~7.4 GB/s sequential reads)
- Target Model:
Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp - Inference Backend: oMLX (Metal acceleration + Lightning MTP 4-stage speculative decoding)
9.2 Streaming Stress Test (LLM Speed Test) Empirical Data
Using the LLM Speed Test streaming benchmark suite, input prompts ranging from 512 to 8,192 tokens were evaluated in sequential stepped tiers:

Key Metric Takeaways:
- Prefill Throughput:
267.95 ~ 317.11 tokens/s(Mean overall prefill throughput:306.86 tokens/s) - Decode Generation Speed:
61.07 ~ 75.85 tokens/s(Mean decode throughput:68.17 tokens/s) - Percentile Breakdown:
- Prefill: P50: 311 tok/s | P90: 316 tok/s | P95: 317 tok/s
- Decode: P50: 68.80 tok/s | P90: 74.27 tok/s | P95: 74.86 tok/s
- Mean Inter-Token Latency (ITL): Maintained at
205 ~ 220 ms(Multi-token emission delivers fluid real-time streaming).
9.3 oMLX Native Benchmark Suite (Single Request vs. Continuous Batching)
Using the built-in oMLX Benchmark suite, we evaluated baseline single-request performance alongside continuous batching saturation limits:

1. Single Request Baseline:
- Prompt 1024 / Output 128 (pp1024/tg128):
- Time to First Token (TTFT):
3531.7 ms - Prefill Throughput:
289.9 tok/s - Decode TPS:
55.3 tok/s(TPOT: 18.22 ms/tok) - End-to-End Latency:
5.853 s, VRAM Footprint:19.21 GB - Prompt 4096 / Output 128 (pp4096/tg128):
- Time to First Token (TTFT):
12917.4 ms - Prefill Throughput:
317.1 tok/s - Decode TPS:
50.5 tok/s(TPOT: 19.96 ms/tok) - End-to-End Latency:
15.458 s, VRAM Footprint:23.08 GB
2. Continuous Batching (Multi-Stream Concurrency):
- 1x Batch (Single-Stream Baseline): Decode TPS
55.3 tok/s(Speedup: 1.00x, TTFT: 3.5s, Total: 5.8s) - 2x Batch (Dual Stream): Decode TPS
111.8 tok/s(Speedup: 2.02x, TTFT: 6.3s, Total: 11.5s) - 4x Batch (Quad Stream): Decode TPS
240.4 tok/s(Speedup: 4.35x, TTFT: 12.0s, Total: 23.0s)
9.4 Deep Architectural Analysis: 240+ tok/s Continuous Batching Breakdown
Why does single-user generation hover at 55~75 tok/s, while 4 concurrent streams rocket up to 240.4 tok/s?
Breakthrough of the Memory Wall
- Single Request Generation (1x): Autoregressive LLM generation is strictly memory-bandwidth bound. To produce every individual token, the GPU must transfer the entire 16.7GB model weight set from unified RAM into the compute cores. At 800 GB/s bus bandwidth, ALU utilization sits at only 10%~15%, with execution units constantly starved for memory transfers.
- Continuous Multi-Batching (4x): The runtime transitions from memory-bound vector operations to compute-bound Matrix-Matrix multiplication (GEMM). The 16.7GB weights are read once, but compute passes are executed across 4 parallel user request streams concurrently. Memory bandwidth is reused across streams, saturating compute pipelines and propelling aggregate machine output to
240.4 tokens/s! - Workload Architecture Trade-offs:
- Individual Developer / Interactive Terminal Agent: Focuses on lowest single-user latency (
65 ~ 85 tok/swith minimal TTFT). - Team Server / Parallel Agents / Microservices: Focuses on aggregate operational throughput, where this Mac Studio handles continuous quad-stream loads at 240+ tok/s comparable to enterprise server racks.
9.5 Why Do the Two Tools Differ in Decode Speeds? (MTP Entropy & Latency Windows)
Across the two benchmark test captures, prefill speeds match almost identically (1K prefill ~290 tok/s, 4K prefill ~315 tok/s, <1% variance). However, decode speeds show 50~55 tok/s in one suite versus 65~75 tok/s in another.
The reasons stem from three core factors:
1. Output Information Entropy vs. MTP Acceptance Rate:
* This model utilizes 4-tier Lightning MTP speculation heads.
* In structured code and predictable outputs, speculation hits 93%~99%, pushing generation to 75 ~ 86 tok/s.
* In open-ended exploratory prose, higher token entropy drops speculation acceptance to ~70%, falling back to single-step autoregressive steps at 50 ~ 55 tok/s.
2. Sampling Temperature Discrepancy: Greedy sampling (temperature=0.0) guarantees optimal candidate alignment with MTP heads, whereas non-zero temperature introduces probabilistic variance and verification rollbacks.
3. Latency Measurement Windows: The oMLX native benchmark suite clocks end-to-end server batch queuing and stream frame closure handshakes (derived from reciprocal TPOT), whereas streaming tools measure raw delta between first and last emitted SSE characters.
10. Model Capability & Multidimensional Intelligence Evaluation
10.1 Intelligence Profile: Comparison with Baseline & Reasoning Models
| Evaluation Dimension | Qwen3.8-27B Uncensored (Deployed) | Standard Qwen 27B / Thinking Models (QwQ / R1) |
|---|---|---|
| Core Coding & Refactoring | 95%~98% parity (oQ4e imatrix preserves near-lossless precision) | Equivalent base coding ability |
| Instruction Adherence & Guardrails | 🏆 Absolute compliance; no refusal loops or preachy lectures; ideal for security auditing | Prone to false-positive trigger guardrails on penetration testing queries |
| Agent Execution Velocity | 🏆 Immediate action; directly emits clean code patches and tool calls | Delays execution with lengthy <think> interior monologues |
| Extreme Mathematical Proofs | Slightly reduced (lacks slow-thinking test-time search expansion) | 🏆 Stronger on Olympiad math via extended chain-of-thought exploration |
10.2 Agent Function Calling & Structured JSON Accuracy
Validate structured schema conformance when driving coding agents:
# Execute multi-file audit task via PI CLI
pi --model omlx-local/Qwen3.8-27B-Uncensored-oQ4e-fp16-mtp \
-p "Inspect the repository structure, enumerate all Python scripts, and tabulate their file sizes and primary function counts."
Passing Criteria: Emits valid read and bash tool calls without missing brackets, hallucinated parameter schemas, or hesitation loops.
10.3 Long-Context Needle in a Haystack (NIAH) Test
Validating factual retrieval up to 64K context limits:
# Construct a 32K-token synthetic document and embed a canary phrase at 50% depth
# Verify prompt: "What is the secret passphrase mentioned in the document?"
Passing Criteria: 100% recall retrieval. Monitoring top -l 1 | grep omlx verifies unified memory remains comfortably under 35GB without triggering macOS compressed swap.
10.4 Continuous Soak & Memory Leak Stability Test
Execute 30 to 50 back-to-back stress cycles:
* Validates sustained SSD cache paging stability over 60+ minutes of continuous generation.
* Confirms VRAM returns cleanly to baseline (~17GB) post-run, with zero Metal memory leakage.
10.5 Uncensored Boundaries & White-Hat Security Code Validation
Feed typical security audit prompts commonly rejected by cloud APIs:
* “Write an asynchronous Python script to sweep 192.168.0.0/24 for open ports 22 and 80, checking for default credentials with rate-limiting and defensive logging.”
Passing Criteria: Treats the prompt as a legitimate administrative audit script and outputs clean, production-grade defensive code without evasive refusal templates.
11. Production Maintenance & Troubleshooting Matrix
| Symptom | Root Cause | Resolution |
|---|---|---|
| Time to First Token (TTFT) spikes to 30+ seconds | Memory guard engaging conservative throttling on large prompts | Ensure ~/.omlx/settings.json has "memory_guard_tier": "aggressive" and "soft_threshold": 0.90. |
| Generation drops from 70+ to ~35 tok/s | Non-zero temperature or thinking mode accidentally enabled | Verify model_settings.json has "temperature": 0.0 and "enable_thinking": false. |
| Test client reports “Endpoint not supported” | Base URL path entered without OpenAI route suffix | Supply full path /v1/chat/completions or update client gateway routing. |
| Port 8000 reload after config modifications | Configuration edited while daemon active | Run launchctl kickstart -k "gui/$(id -u)/com.user.omlx.server" for a seamless reload. |