2026年10月3日
bonsai-27b-pq2-agent-benchmark-en-cover
Can a 7.2GB 27B model handle real code-agent work? A controlled Bonsai 2 PQ2_0 versus Qwen3.8 benchmark of capability and convergence.

7.2GB 27B Agent Benchmark: Bonsai 2 PQ2_0

Test scope: This is a controlled, single-machine comparison of Bonsai 2 27B in PQ2_0 GGUF format and Qwen3.8-27B in Q4_K_M. The results describe these exact builds, settings, tasks, and agent harness—not a universal ranking of either model family.

A 27B-parameter model whose GGUF weights occupy only 7.2GB sounds almost implausible. The practical question is not whether the file fits on disk, though. It is whether a model compressed that aggressively can still finish real work once it is placed in an agent loop.

That is what this benchmark set out to test. Instead of stopping at VRAM figures or token throughput, both models were asked to work through agent tasks: inspect a codebase, reason about a fault, edit files, run tests, respond to failures, and decide when the task was done.

The two model configurations

Model Quantization GGUF weight size Observed runtime VRAM
Bonsai 2 27B PQ2_0 ~7.2GB ~12.2GB
Qwen3.8-27B Q4_K_M ~17.1GB ~23.8GB

Bonsai’s weight file is roughly 42% the size of the Qwen Q4_K_M file, and its observed runtime VRAM usage is close to half. That makes the format interesting for 12GB–16GB consumer GPUs—but only if task capability remains intact.

Methodology: measure work, not just answers

The two models ran on separate GPUs and separate llama-server ports. Both used a 131,072-token context window and the same decoding settings:

temperature = 1.0
top_p       = 0.95
top_k       = 20
min_p       = 0
thinking    = medium

The agent layer was Pi CLI with extra skills, extensions, and prompt templates disabled. Each agent had the normal developer toolset for reading files, editing code, and running commands.

The evaluation had two stages. Stage 1 checked whether extreme compression had visibly damaged baseline reasoning, instruction following, and long-context retrieval. Stage 2 focused on real software tasks and allowed a larger step budget where needed.

Stage 1: no obvious collapse in baseline capability

The first four tasks covered logical reasoning, knowledge comprehension, strict-format instruction following, and locating information in a long input. Both models passed every task.

Task Bonsai 2 Qwen3.8
R1: logical reasoning PASS PASS
K1: knowledge comprehension PASS PASS
I1: strict instruction following PASS PASS
L1: 90K long-context retrieval PASS PASS

The L1 prompt was 90,610 tokens long and placed the required detail among large amounts of similar material. Both models recovered the target information successfully.

This does not prove the two models are equivalent. It does show that PQ2_0 did not make Bonsai immediately unusable for reasoning, recall, or long-context lookup in this test set.

The agent loop exposed a more important difference

Code-agent tasks are unlike one-shot questions. The model must execute a loop:

understand the project
→ inspect files
→ identify the fault
→ modify code
→ run tests
→ react to results
→ decide to stop

The last step matters more than it may seem. An agent can find the right solution, keep looking for hypothetical issues, and exhaust its step budget anyway. That is distinct from failing to solve the task.

In the initial run, this distinction was already visible. Bonsai completed the billing-engine repair (C1), while Qwen encountered a tool-execution problem and timed out. On the cross-file order refactor (C2), the result reversed: Qwen completed quickly, while Bonsai had code that was close to correct but continued checking and exceeded the strict step limit.

That pattern motivated Stage 2: rerun the larger software tasks with task-specific budgets of 20–30 agent steps and inspect how each model reached its outcome.

Stage 2 results

Task Bonsai 2 Qwen3.8
C1: billing-engine bug fix 13 steps / 159.6s / PASS 5 steps / 39.2s / PASS
C2: cross-file order aggregation refactor 14 steps / 105.5s / PASS 7 steps / 76.8s / PASS
P1: CSV streaming export service 26/25 steps / 551.6s / TIMEOUT 7/25 steps / 66.8s / PASS
A1: concurrent inventory-deduction API 19/30 steps / 226.3s / PASS 15/30 steps / 175.7s / PASS

C1: both fixed the code; one took a much longer route

C1 contained three ordinary but consequential faults: converting float values to Decimal for currency handling, applying discount and tax in the wrong order, and rejecting full refunds at a boundary condition.

Qwen followed a compact path: read the code, verify the diagnosis, edit the three issues, run pytest, and stop after a passing result. It used five steps and 39.2 seconds.

Bonsai also found and fixed the defects, but used 13 steps and 159.6 seconds. After its primary patch, it continued to inspect Decimal rounding, construct additional boundary cases, and re-run validation. The end result was correct in both cases; the difference was convergence behavior, not a basic inability to repair the code.

P1: the clearest example of the stop-decision problem

P1 implemented a CSV streaming export service involving HTTP output, CSV escaping, and client-connection lifetime handling.

Qwen passed in seven steps, 66.8 seconds, and 406 thinking tokens. Bonsai completed the core behavior but kept the service running, rechecked output, tested more edge cases, and continued validating CSV behavior. The session reached 26 steps against a 25-step limit and was recorded as a timeout under the benchmark rules:

26 / 25 steps
551.6 seconds
11,383 thinking tokens
TIMEOUT

After the session ended, the hidden validator passed the resulting code. In other words, the implementation was functionally correct, but the agent converted a solved task into a timeout by continuing to work.

A1: complex agent work was still within reach

A1 was the most demanding task: an SQLite inventory-deduction API that could not permit negative inventory and had to deal with lock contention under concurrency.

With a restrictive Stage 1 budget, neither model completed it. With a 30-step budget in Stage 2, both passed. Qwen used transactions and retries; Bonsai used an atomic conditional update approach to prevent inventory from dropping below zero. This is important because it separates a budget-related timeout from a claim that the compressed model cannot handle multi-step engineering tasks.

Thinking-token usage explains the end-to-end gap

Across the four Stage 2 software tasks, the models produced:

Bonsai 2: 22,831 thinking tokens
Qwen3.8:  3,568 thinking tokens

That is about 6.4× as many thinking tokens for Bonsai. The gap appears clearly in individual tasks:

C1: Bonsai 4,372 vs. Qwen 299
P1: Bonsai 11,383 vs. Qwen 406

Raw decode speed alone therefore does not describe agent efficiency. A model can emit tokens quickly and still take longer end to end if it needs many more tokens, tool calls, and self-checking cycles to reach the same solution.

The sessions suggest two different operating styles. Bonsai repeatedly searched for possible remaining mistakes after a working patch and passing tests. Qwen was more completion-oriented: identify the issue, implement the fix, confirm the requirement, and finish. Neither style is automatically “smarter,” but the second one is generally cheaper and less timeout-prone in a bounded agent workflow.

What 7.2GB changes—and what it does not

The strongest case for Bonsai is still deployment size. Moving a 27B-class model from a roughly 17.1GB GGUF file to a 7.2GB file is a reduction of about 58%. In this run, that moved observed runtime VRAM from roughly 23.8GB to 12.2GB.

However, a 7.2GB model file does not mean that 7.2GB of VRAM is enough. Real runtime memory also includes:

KV cache
context length
runtime buffers
batch size
backend overhead

The practical implication is more modest and more useful: a 27B-class model that normally feels comfortable only around 24GB of VRAM may become plausible on 12GB–16GB consumer hardware, subject to the chosen context size and runtime configuration.

Verdict: capability survived; agent efficiency is the trade-off

This benchmark does not show an obvious loss of core reasoning or long-context capability from this extreme quantization. Bonsai 2 completed logical tasks, long-context retrieval, debugging, refactoring, API work, and concurrent data handling.

The cost appeared in agent convergence:

more thinking tokens
more agent steps
longer execution paths
higher timeout risk
weaker stopping decisions

Qwen3.8 Q4_K_M was not winning because it could do work that Bonsai categorically could not. It was winning because it more consistently stopped at the right time after satisfying the task.

For agent deployments, that distinction directly affects token cost, latency, compute usage, and reliability. Bonsai 2 makes a serious case that a 7.2GB 27B model can do real work. Its next challenge is not simply preserving intelligence at extreme compression—it is learning when the work is finished.

Benchmark boundaries

These findings apply to fixed quantization builds, medium thinking, Pi CLI, this task set, and this hardware configuration. They are best read as an observation of how these two models behaved in the same agent framework, rather than a broad claim that one model is universally better.

About The Author

Leave a Reply

Your email address will not be published. Required fields are marked *