2026年10月3日
bonsai2-27b-rtx3060-12gb-en-cover
Deploy Bonsai 2 27B PQ2_0 on an RTX 3060 12GB with PrismML, a local OpenAI-compatible API, systemd service, and safe LAN access.

Deploy Bonsai 2 27B on an RTX 3060 12GB

Tested scope: This guide documents a verified single-user deployment of Bonsai 2 27B PQ2_0 on an NVIDIA GeForce RTX 3060 with 12GB of VRAM. Replace the example paths, users, IP addresses, and network ranges with your own.

Bonsai 2 27B can run on a 12GB consumer GPU, but it needs its matching runtime. This is not a generic stock llama.cpp deployment: Bonsai’s PQ2_0 and PTQ1_0 formats depend on ternary-weight kernels and a Hadamard activation transform supplied by PrismML.

This English guide covers the verified configuration, Ubuntu installation, OpenAI-compatible API, systemd operation, safe LAN access, and the measurement limits that should accompany the performance figures.

Verified configuration and measurement boundary

Bonsai 2 is a 27B-class hybrid-attention model based on Qwen3.8-27B. A small GGUF file does not make it a 7B model; it lowers the storage and memory footprint of the weights.

Item Verified result
GPU NVIDIA GeForce RTX 3060, 12GB VRAM
Model Bonsai 2 27B, PQ2_0 GGUF
Runtime PrismML llama.cpp fork with CUDA
Server profile Full GPU offload, 128K configured context, Q4_0 K/V cache, one parallel slot
API workload 512–8,192-token inputs, fixed 128-token output
Mean prefill throughput 511.70 tokens/s
Mean decode throughput 32.76 tokens/s
Decode P50 / P90 / P95 32.63 / 34.77 / 35.03 tokens/s
Decode range 30.54–35.34 tokens/s

Prefill is input-processing throughput; decode is token-by-token generation throughput. Prefill throughput is not time to first token.

These results are single-concurrency, end-to-end API measurements, not bare-kernel llama-bench figures. The server successfully started with -c 131072, but the largest measured prompt was 8,192 tokens. A 128K configuration can start; that does not prove stable, high-quality inference with a full 128K request. The observed ~10.78GiB VRAM use was a snapshot, not a long-context peak.

Capability sample: useful context, not a leaderboard

The throughput data above came from this RTX 3060 host. Separately, Bonsai 2 PQ2_0 and Qwen3.8-27B Q4_K_M completed the same 102-item EvalScope sample. These results describe basic capability sampling, not this GPU’s speed or a universal model ranking.

Benchmark Samples Metric Bonsai 2 PQ2_0 Qwen3.8 27B Q4_K_M
AIME-2024 10 Accuracy 20.00% 30.00%
C-Eval 20 Accuracy 85.00% 85.00%
IFEval 20 Prompt-level strict 65.00% 70.00%
LiveCodeBench 10 Pass@1 30.00% 40.00%
MMLU-Pro 42 Accuracy 59.52% 80.95%

Bonsai matched the Qwen baseline on the C-Eval sample and trailed it on the sampled math, instruction-following, code, and MMLU-Pro tasks. It is a deployable 27B-class local agent model, not a complete capability replacement for Qwen3.8 27B at Q4_K_M.

Each benchmark had only 10–42 items, for 102 total, and was run on September 17–18, 2026. Preserve the benchmark version, prompt template, sampling parameters, runtime version, raw export, and hardware configuration for a reproducible follow-up.

Use a PrismML-compatible runtime

Stock llama.cpp does not implement the ternary kernels and activation transform needed for Bonsai PQ2_0 and PTQ1_0. Use the PrismML llama.cpp fork or the matching runtime included with the Bonsai Demo.

Component Approximate size Use case
PTQ1_0 language model 5.95GB Tighter VRAM budgets
PQ2_0 language model 7.21GB Format used in this guide; documented as typically faster for prompt processing
mmproj-Q8_0 ~0.63GB Image input only; omit for text-only serving

The 5.95GB figure is for the PTQ1_0 language model alone. It is not a PQ2_0 runtime footprint and does not include an image projector.

Prerequisites

The following commands target Ubuntu 22.04 or 24.04:

nvidia-smi
git --version
curl --version

# Needed only when building from source.
sudo apt update
sudo apt install -y build-essential cmake git curl python3-pip

export APP_DIR="$HOME/bonsai2-serve"
mkdir -p "$APP_DIR"
cd "$APP_DIR"

Option A: use the official demo

The Bonsai Demo is the least fragile path for most users. It installs a matching runtime and model assets. This disables optional Open WebUI and Code Interpreter downloads:

git clone https://github.com/PrismML-Eng/Bonsai-demo.git
cd Bonsai-demo

BONSAI_OPENWEBUI=0 BONSAI_CODE_INTERPRETER=0 ./setup.sh
./scripts/start_llama_server.sh

Verify the default local service:

curl http://127.0.0.1:8080/health
curl http://127.0.0.1:8080/v1/models

Use the demo README as the current source for platform support, binaries, and launch options. It includes macOS, Linux CUDA, Vulkan, ROCm, and CPU paths.

Option B: build and manage the server yourself

Build a CUDA-enabled PrismML runtime:

cd "$APP_DIR"
git clone https://github.com/PrismML-Eng/llama.cpp.git runtime

cmake -S runtime -B runtime/build \
  -DCMAKE_BUILD_TYPE=Release \
  -DGGML_CUDA=ON
cmake --build runtime/build -j "$(nproc)" --target llama-server

The binary will be at $APP_DIR/runtime/build/bin/llama-server. If you do not build it, use a suitable Linux x64 CUDA asset from PrismML releases rather than a stale release tag copied from an old tutorial.

Download the weights

The source repository is prism-ml/Ternary-Bonsai-2-27B-gguf.

python3 -m pip install --user -U "huggingface_hub[cli]"
mkdir -p "$APP_DIR/models"

hf download prism-ml/Ternary-Bonsai-2-27B-gguf \
  Ternary-Bonsai-2-27B-PQ2_0.gguf \
  --local-dir "$APP_DIR/models"

# Optional: required only for image input.
hf download prism-ml/Ternary-Bonsai-2-27B-gguf \
  Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf \
  --local-dir "$APP_DIR/models"

ls -lh "$APP_DIR/models"
sha256sum "$APP_DIR/models"/*.gguf

Record exact GGUF filenames, sizes, and SHA-256 hashes. It makes troubleshooting and benchmark reports much more reliable.

Launch an OpenAI-compatible local API

Bind to loopback first. It is the secure default. Remove --mmproj for a text-only service.

export BIN="$APP_DIR/runtime/build/bin/llama-server"
export MODEL="$APP_DIR/models/Ternary-Bonsai-2-27B-PQ2_0.gguf"
export MMPROJ="$APP_DIR/models/Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf"

"$BIN" \
  -m "$MODEL" \
  --mmproj "$MMPROJ" \
  --host 127.0.0.1 \
  --port 8080 \
  -ngl 999 \
  -fa on \
  -c 131072 \
  -np 1 \
  --cache-type-k q4_0 \
  --cache-type-v q4_0 \
  --jinja \
  --alias bonsai2-27b \
  --temp 1.0 --top-p 0.95 --top-k 20
  • -ngl 999 requests full eligible layer offload; the runtime chooses the actual count.
  • -c 131072 configures a 128K context. It does not expand the supported context or prove full-load stability.
  • -np 1 is an appropriate starting point for a personal 12GB service.
  • --cache-type-k/v q4_0 enables 4-bit K/V cache compression. Validate it with your quality-sensitive prompts.
  • --jinja enables the chat template and tool-call format.

The runtime documentation estimates Q4 K/V cache at about 18KiB per token: roughly 1.8GiB for 100K tokens and 2.25GiB for 128K. Leave extra VRAM headroom for CUDA, activations, the optional projector, and fragmentation.

Make it a systemd service

Create $APP_DIR/scripts/start-bonsai2.sh:

#!/usr/bin/env bash
set -euo pipefail

exec "$APP_DIR/runtime/build/bin/llama-server" \
  -m "$APP_DIR/models/Ternary-Bonsai-2-27B-PQ2_0.gguf" \
  --mmproj "$APP_DIR/models/Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf" \
  --host 127.0.0.1 --port 8080 \
  -ngl 999 -fa on -c 131072 -np 1 \
  --cache-type-k q4_0 --cache-type-v q4_0 \
  --jinja --alias bonsai2-27b \
  --temp 1.0 --top-p 0.95 --top-k 20

Make it executable, then create /etc/systemd/system/bonsai2.service after replacing YOUR_LINUX_USER with a non-root user:

[Unit]
Description=Bonsai 2 27B OpenAI-compatible server
After=network-online.target
Wants=network-online.target

[Service]
Type=simple
User=YOUR_LINUX_USER
WorkingDirectory=%h/bonsai2-serve
ExecStart=%h/bonsai2-serve/scripts/start-bonsai2.sh
Restart=on-failure
RestartSec=3
StandardOutput=journal
StandardError=journal

[Install]
WantedBy=multi-user.target
chmod +x "$APP_DIR/scripts/start-bonsai2.sh"
sudo systemctl daemon-reload
sudo systemctl enable --now bonsai2.service
systemctl status bonsai2.service --no-pager
journalctl -u bonsai2.service -f

Check the API and tool calls

curl -s http://127.0.0.1:8080/v1/models

curl -s http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "bonsai2-27b",
    "messages": [{"role": "user", "content": "Explain ternary quantization in one sentence."}],
    "max_tokens": 128,
    "chat_template_kwargs": {"enable_thinking": false}
  }'

A tool_calls response means the model proposed a call. Your application or agent remains responsible for authorization, argument validation, execution, and returning the tool result.

Share on a LAN only with access control

An OpenAI-compatible llama-server endpoint is not a complete production gateway. Do not expose port 8080 directly to the public internet or bind it to 0.0.0.0 on an unrestricted network.

For a trusted LAN, change the host to 0.0.0.0, restrict access to your real trusted CIDR, restart the service, and test from another LAN device:

# Replace this example with your actual trusted subnet.
sudo ufw allow from 192.168.1.0/24 to any port 8080 proto tcp
sudo systemctl restart bonsai2.service

curl http://SERVER_LAN_IP:8080/health
curl http://SERVER_LAN_IP:8080/v1/models

For access across networks, use a VPN, Tailscale, or an authenticated reverse proxy instead of direct port forwarding.

Reproduce the benchmark honestly

The single-concurrency test used inputs from 512 through 8,192 tokens in 512-token increments, a fixed 128-token output, and a 30-second timeout. Record the GPU, VRAM, driver version, model SHA-256, PrismML release or commit, context, cache type, concurrency, projector use, offload configuration, warmups, and failure rate.

To claim genuine 128K capability, test 32K, 64K, 96K, and 128K prompts separately. Record completion rate, latency, nvidia-smi peak VRAM, and answer quality. A successful 128K configuration alone cannot replace that test.

FAQ

Why is a roughly 7GB model file not automatically much faster than Q4?

Smaller weights reduce storage and memory-bandwidth demands, but Bonsai also performs ternary unpacking and a Hadamard activation transform. The compute-versus-memory balance varies by GPU. On an RTX 3060, the key benefit is making a 27B-class model, a single-user API, and longer contexts plausible within 12GB.

Why disable thinking for the throughput test?

Thinking tokens use output budget and generation time. Disabling them makes connectivity, tool-format, and throughput checks more controlled. Measure complex reasoning workloads separately with end-to-end latency and accuracy.

Conclusion

Bonsai 2 27B on an RTX 3060 12GB matters not because a 7GB file must beat every conventional Q4 model, but because ternary weights plus the PrismML runtime lower the deployment threshold for a 27B-class local model.

Under the tested conditions, it provided an approximately 33 tokens/s single-user, OpenAI-compatible API with optional image input and tool-call formatting. Use the matching runtime, distinguish a 128K configuration from a fully validated 128K workload, and apply access control before sharing the service beyond localhost.

About The Author

Leave a Reply

Your email address will not be published. Required fields are marked *