2026年10月3日
English Bonsai Agent cover showing Finish Then Stop and the P1 case study from 26 plus to 16 steps
Make a local coding agent finish: Bonsai 27B convergence experiments, conservative steering rules, and three separately recorded tasks.

A 7.2GB model can finish useful coding tasks, but the harness must distinguish valid completion evidence from repetitive checking.

The difficult part of using a small local coding model is sometimes recognizing when it has finished. In my Bonsai 2 27B PQ2_0 experiments, the agent could produce a correct fix and still continue checking until it exhausted the step budget.

The GGUF weights were about 7.2GB, running through llama-server on an RTX 3060 12GB in the convergence experiments, with Pi as the agent harness, medium thinking, and a configured 131,072-token context window. Weight-file size is not total runtime memory: context, caches, offload choices, and the serving build matter.

This article combines two separate records: controlled convergence experiments and a later three-task exercise. Their task identifiers and measurements are kept separate.

When correct code still times out

A typical loop looked like this:

Understand the request
        ↓
Edit the code
        ↓
Run a test → PASS
        ↓
Inspect more files, repeat the same check
        ↓
Reach the step limit → TIMEOUT

On the controlled REST API and CSV-export task, P1, the uncontrolled run exceeded 26 steps and took about 551.6 seconds. A hidden validator could accept the work, but the run did not finish within its budget.

This shows a distinction between artifact correctness and agent completion behavior. It does not establish that low-bit quantization caused the loop. The original discussion proposed a confidence-related explanation; that mechanism was not isolated experimentally.

A prompt helped, but left little margin

The first change was a completion policy: trust relevant passing checks, do not repeat the same check without changes, and finish after reviewing the explicit requirements.

Two P1 reruns with prompt-only control reached about 24–25 steps. They completed, but remained near the limit.

Runtime steering then reminded the model to reassess completion after evidence accumulated. Early aggressive steering reduced P1 to roughly 11–12 steps, but that result was not a safe general stopping rule. A different long-chain task exposed a false positive: a superficial successful check was mistaken for evidence that all requested behavior had been delivered.

The fix was to improve the completion evidence, not simply lower the step cap.

A conservative extension

pi-extension-convergence listens to tool events and sends a steering message. It leaves the model able to review unmet requirements rather than terminating the process.

The current documented rules distinguish two cases:

Evidence Response
Passing automated test plus a verified core business API response Strong reminder to review requirements and finish
Exactly the same passing test command three times without a tracked modification Soft reminder to break the repeated-check loop
One passing test, a generic health probe, directory listing, or exploratory database query No completion conclusion from that signal alone

For the business API check, the extension expects strict curl status handling, including --fail and a final HTTP status written with --write-out. This is a recognition rule, not proof that a status code validates all business semantics.

After a tracked source edit, the extension invalidates old evidence and allows steering again. That prevents an earlier successful test from counting as validation of later code.

The final controlled results

The recorded conservative-extension results were:

Controlled task Uncontrolled baseline Final extension run Hidden validator
A1: long-chain bug and service migration Timeout, 26+ steps 19 steps Pass
P1: REST API and CSV export Timeout, 26+ steps 16 steps Pass
C1: Python algorithmic repair 13 steps 10 steps Pass
C2: TypeScript tenant service 14 steps 8 steps Pass

P1’s final run took about 216 seconds, versus the 551.6-second baseline. The English cover uses 26+ → 16 steps, with a P1 case-study label, rather than advertising the earlier aggressive 11-step experiment as the final result.

These are recorded task runs, not a broad benchmark showing the same improvement on every repository. The current extension contains additional evidence-validation refinements; these historical timings should not be described as a new measurement of today’s exact commit.

Try it in one Pi workspace

Install into the project rather than changing every workspace first:

mkdir -p .pi/extensions
curl -fsSL \
  https://raw.githubusercontent.com/yang2020chen/pi-extension-convergence/main/convergence.ts \
  -o .pi/extensions/convergence.ts

Review the extension before enabling it. For repeatable deployment, replace main with a reviewed commit or release reference.

Download the companion policy and append it for a trial session:

curl -fsSL \
  https://raw.githubusercontent.com/yang2020chen/pi-extension-convergence/main/prompts/convergence.md \
  -o ./convergence-policy.md

pi --append-system-prompt ./convergence-policy.md

The project documents .pi/APPEND_SYSTEM.md for a persistent project policy. It does not use a systemPromptAppend setting in settings.json.

Try a small task with explicit acceptance criteria. Inspect when steering occurs, whether genuinely unfinished work continues, and whether a later edit causes the agent to obtain fresh evidence.

A separate three-task exercise

The later exercise used Bonsai PQ2_0, medium thinking, Pi, and the convergence extension. The agent received goals and boundaries without step-by-step repair instructions. No additional human rescue prompts were recorded during the runs.

Exercise Wall time Assistant turns Tool calls Output tokens Peak context
Repair two Python application bugs 7m 57s 22 24 10,737 25,485
Deploy a GitHub project under port constraints 9m 31s 31 37 21,542 38,057
Recover a service in an isolated Linux exercise 1m 34s 11 14 3,940 17,269
Total 19m 02s 64 75 36,219 38,057 maximum

These are three additional tasks, not reruns of A1/P1/C1/C2. Without matched no-extension runs, their results do not measure the extension’s causal speedup.

Python repair: clarify the expected behavior

The model found an empty-string indexing error and a completion flag that changed in memory without being persisted. It added regression checks and verified the saved state across a new process.

The evaluation also exposed an ambiguous requirement. A hidden check expected an empty search to return all tasks; the model treated it as invalid input. Both choices can be reasonable, but the intended contract must be stated before claiming that one implementation is incorrect.

Deployment: getting HTTP 200 was only one step

The model repaired a host=localhost error and then handled a Flask debug-reloader process issue during repeatable startup and shutdown testing. It respected the occupied port and used the required replacement port.

The deployment succeeded, but repeated probes still consumed many calls. Completion control reduced one failure mode; it did not eliminate every inefficient action.

Recovery: check the relevant filesystem

The Linux exercise reported a service running out of space even though the root filesystem had about 285GB free. The service directory was on a separate 96MiB tmpfs mount.

The model identified that mount, removed designated cache and rotated-log files, preserved protected data and configuration, and restored the service. The task illustrates why a global disk-space number can mislead a repair agent.

What to measure next

The maximum observed context in the three-task exercise was about 38K, despite configuring 128K. That observation supports testing a smaller window for similar tasks; it does not prove that 64K covers all coding work.

Reported cache-read tokens are repeated traffic across turns, not the size of a single prompt. The original records reported high cache reuse, but a causal cache benefit would need a controlled comparison.

For your next evaluation, record completion quality, wall time, turns, tool calls, retries, human intervention, and the exact harness version. Keep false completion and unnecessary continuation as separate failure categories.

Source and experiment records

The measurements are retained from the original experiments. This edition updates the documented installation and evidence rules without claiming a fresh model benchmark.

About The Author

Leave a Reply

Your email address will not be published. Required fields are marked *