A 7.2GB model can finish useful coding tasks, but the harness must distinguish valid completion evidence from repetitive checking.
The difficult part of using a small local coding model is sometimes recognizing when it has finished. In my Bonsai 2 27B PQ2_0 experiments, the agent could produce a correct fix and still continue checking until it exhausted the step budget.
The GGUF weights were about 7.2GB, running through llama-server on an RTX 3060 12GB in the convergence experiments, with Pi as the agent harness, medium thinking, and a configured 131,072-token context window. Weight-file size is not total runtime memory: context, caches, offload choices, and the serving build matter.
This article combines two separate records: controlled convergence experiments and a later three-task exercise. Their task identifiers and measurements are kept separate.
When correct code still times out
A typical loop looked like this:
Understand the request
↓
Edit the code
↓
Run a test → PASS
↓
Inspect more files, repeat the same check
↓
Reach the step limit → TIMEOUT
On the controlled REST API and CSV-export task, P1, the uncontrolled run exceeded 26 steps and took about 551.6 seconds. A hidden validator could accept the work, but the run did not finish within its budget.
This shows a distinction between artifact correctness and agent completion behavior. It does not establish that low-bit quantization caused the loop. The original discussion proposed a confidence-related explanation; that mechanism was not isolated experimentally.
A prompt helped, but left little margin
The first change was a completion policy: trust relevant passing checks, do not repeat the same check without changes, and finish after reviewing the explicit requirements.
Two P1 reruns with prompt-only control reached about 24–25 steps. They completed, but remained near the limit.
Runtime steering then reminded the model to reassess completion after evidence accumulated. Early aggressive steering reduced P1 to roughly 11–12 steps, but that result was not a safe general stopping rule. A different long-chain task exposed a false positive: a superficial successful check was mistaken for evidence that all requested behavior had been delivered.
The fix was to improve the completion evidence, not simply lower the step cap.
A conservative extension
pi-extension-convergence listens to tool events and sends a steering message. It leaves the model able to review unmet requirements rather than terminating the process.
The current documented rules distinguish two cases:
| Evidence | Response |
|---|---|
| Passing automated test plus a verified core business API response | Strong reminder to review requirements and finish |
| Exactly the same passing test command three times without a tracked modification | Soft reminder to break the repeated-check loop |
| One passing test, a generic health probe, directory listing, or exploratory database query | No completion conclusion from that signal alone |
For the business API check, the extension expects strict curl status handling, including --fail and a final HTTP status written with --write-out. This is a recognition rule, not proof that a status code validates all business semantics.
After a tracked source edit, the extension invalidates old evidence and allows steering again. That prevents an earlier successful test from counting as validation of later code.
The final controlled results
The recorded conservative-extension results were:
| Controlled task | Uncontrolled baseline | Final extension run | Hidden validator |
|---|---|---|---|
| A1: long-chain bug and service migration | Timeout, 26+ steps | 19 steps | Pass |
| P1: REST API and CSV export | Timeout, 26+ steps | 16 steps | Pass |
| C1: Python algorithmic repair | 13 steps | 10 steps | Pass |
| C2: TypeScript tenant service | 14 steps | 8 steps | Pass |
P1’s final run took about 216 seconds, versus the 551.6-second baseline. The English cover uses 26+ → 16 steps, with a P1 case-study label, rather than advertising the earlier aggressive 11-step experiment as the final result.
These are recorded task runs, not a broad benchmark showing the same improvement on every repository. The current extension contains additional evidence-validation refinements; these historical timings should not be described as a new measurement of today’s exact commit.
Try it in one Pi workspace
Install into the project rather than changing every workspace first:
mkdir -p .pi/extensions
curl -fsSL \
https://raw.githubusercontent.com/yang2020chen/pi-extension-convergence/main/convergence.ts \
-o .pi/extensions/convergence.ts
Review the extension before enabling it. For repeatable deployment, replace main with a reviewed commit or release reference.
Download the companion policy and append it for a trial session:
curl -fsSL \
https://raw.githubusercontent.com/yang2020chen/pi-extension-convergence/main/prompts/convergence.md \
-o ./convergence-policy.md
pi --append-system-prompt ./convergence-policy.md
The project documents .pi/APPEND_SYSTEM.md for a persistent project policy. It does not use a systemPromptAppend setting in settings.json.
Try a small task with explicit acceptance criteria. Inspect when steering occurs, whether genuinely unfinished work continues, and whether a later edit causes the agent to obtain fresh evidence.
A separate three-task exercise
The later exercise used Bonsai PQ2_0, medium thinking, Pi, and the convergence extension. The agent received goals and boundaries without step-by-step repair instructions. No additional human rescue prompts were recorded during the runs.
| Exercise | Wall time | Assistant turns | Tool calls | Output tokens | Peak context |
|---|---|---|---|---|---|
| Repair two Python application bugs | 7m 57s | 22 | 24 | 10,737 | 25,485 |
| Deploy a GitHub project under port constraints | 9m 31s | 31 | 37 | 21,542 | 38,057 |
| Recover a service in an isolated Linux exercise | 1m 34s | 11 | 14 | 3,940 | 17,269 |
| Total | 19m 02s | 64 | 75 | 36,219 | 38,057 maximum |
These are three additional tasks, not reruns of A1/P1/C1/C2. Without matched no-extension runs, their results do not measure the extension’s causal speedup.
Python repair: clarify the expected behavior
The model found an empty-string indexing error and a completion flag that changed in memory without being persisted. It added regression checks and verified the saved state across a new process.
The evaluation also exposed an ambiguous requirement. A hidden check expected an empty search to return all tasks; the model treated it as invalid input. Both choices can be reasonable, but the intended contract must be stated before claiming that one implementation is incorrect.
Deployment: getting HTTP 200 was only one step
The model repaired a host=localhost error and then handled a Flask debug-reloader process issue during repeatable startup and shutdown testing. It respected the occupied port and used the required replacement port.
The deployment succeeded, but repeated probes still consumed many calls. Completion control reduced one failure mode; it did not eliminate every inefficient action.
Recovery: check the relevant filesystem
The Linux exercise reported a service running out of space even though the root filesystem had about 285GB free. The service directory was on a separate 96MiB tmpfs mount.
The model identified that mount, removed designated cache and rotated-log files, preserved protected data and configuration, and restored the service. The task illustrates why a global disk-space number can mislead a repair agent.
What to measure next
The maximum observed context in the three-task exercise was about 38K, despite configuring 128K. That observation supports testing a smaller window for similar tasks; it does not prove that 64K covers all coding work.
Reported cache-read tokens are repeated traffic across turns, not the size of a single prompt. The original records reported high cache reuse, but a causal cache benefit would need a controlled comparison.
For your next evaluation, record completion quality, wall time, turns, tool calls, retries, human intervention, and the exact harness version. Keep false completion and unnecessary continuation as separate failure categories.
Source and experiment records
- Extension source and current installation guide
- Recorded convergence experiments in Chinese
- Recorded three-task exercise in Chinese
- Existing English Bonsai deployment guide
- Extension license: MIT.
The measurements are retained from the original experiments. This edition updates the documented installation and evidence rules without claiming a fresh model benchmark.