Test a deployed model service with pinned capability samples, deterministic context workloads, and records that make later comparisons inspectable.
A local model generating 30 tokens per second can still be frustrating to use. It might take a long time to produce the first token, struggle to follow instructions, or spend many tool calls on a task that a slower model finishes quickly.
Those are different questions. A useful benchmark needs to distinguish answer quality, startup latency, sustained generation, and the conditions under which the numbers were collected.
Local LLM Standard Test (LLST) brings those measurements into a versioned protocol for an OpenAI-compatible model service. This English edition uses the published v1.0.0 baseline, which supersedes the release-candidate reference figures in the original Chinese article.
Two suites, one recorded protocol
The capability suite uses 102 pinned samples:
| Dataset | Samples | What this slice probes |
|---|---|---|
| MMLU-Pro | 42 | Knowledge and reasoning across subject categories |
| IFEval | 20 | Compliance with explicit instruction constraints |
| AIME-2024 | 10 | Mathematical reasoning |
| C-Eval | 20 | Chinese knowledge and reasoning |
| LiveCodeBench | 10 | Generated code evaluated in a sandbox |
These are small fixed slices, not the complete official datasets. A 10-problem coding result cannot establish a universal model ranking.
The performance suite uses input lengths of 512, 4,096, 16,384, and 28,672 tokens, with a configured output target of 512 tokens. A deterministic seed and workload hashes let you inspect whether two runs received the same recorded input.
Capability scores answer “what did it solve?” Performance measurements answer “how did this serving configuration respond?” Keep the two separate.
What makes a run comparable?
LLST records sample identifiers and hashes for prompts, targets, and records. It also checks dataset snapshots and tokenizer fingerprints. A mismatch can stop the run rather than silently producing a score from different inputs.
Those checks help with drift; they do not certify model quality. In particular, “102/102 verified” means that the execution matched the pinned sample protocol. It does not mean that all 102 answers were correct.
For a useful comparison, also record the model artifact and quantization, serving-engine build, context setting, KV-cache precision, GPU allocation, and machine configuration. A model name alone is insufficient.
Token counts require care. Character counts are not token counts, and different tokenizers can assign different lengths to the same text. Use the tokenizer associated with the server, and retain the workload manifests alongside the results.
Set up your own run
You need a running model API, Python 3.10 or later, and Docker for the code-evaluation sandbox. The project documents Linux and macOS hosts.
git clone https://github.com/yang2020chen/local-llm-standard-test.git
cd local-llm-standard-test
./scripts/install.sh
cp configs/machines/machine.example.yaml configs/machines/machine.local.yaml
Edit the copied machine profile. At minimum, configure the model identity, API endpoints, exact local tokenizer, fingerprint, and runtime context length. The example configuration is a template, not a ready-made fingerprint for your model.
export LLST_API_KEY="your-local-api-key"
The API profile distinguishes the base URL from the performance endpoint. Verify that your server implements the required completion endpoint; an API described as OpenAI-compatible may still differ in supported routes or streaming behavior.
LLST v1.0’s full standard includes the 28,672-token input plus output budget. The current machine template requires at least 29,184 tokens of context. If your service cannot support that tier, describe a reduced test as a custom run rather than labeling it the complete v1.0 standard.
Run the three stages in order:
# Validate configuration, sample identity, tokenizer, and environment
./scripts/run_standard_test.sh --check
# Confirm that the end-to-end path works with a small smoke run
./scripts/run_standard_test.sh --smoke
# Run the full standard protocol
./scripts/run_standard_test.sh
Keep the reports and manifests under the configured output root. The current default layout is outputs/<model_name>/<timestamp>/, rather than the older article’s results/ description.
Read the published baseline as a specific run
The repository’s Baseline #001 is marked VERIFIED STABLE for v1.0.0. Its documented setup uses dual RX 7900 XTX GPUs, Qwen3.8-Flash-Next UD-Q3_K_XL, llama-server, speculative MTP decoding, and a 32K context setting.
The published capability summary reports:
| Dataset | Samples | Recorded score |
|---|---|---|
| MMLU-Pro | 42 | 85.71% |
| IFEval | 20 | 95.00% |
| C-Eval | 20 | 90.00% |
| LiveCodeBench | 10 | 80.00% |
| AIME-2024 | 10 | 30.00% |
The published performance summary reports these figures:
| Input tokens | Target output | Average TTFT | Reported output throughput |
|---|---|---|---|
| 512 | 512 | 1.28 s | 22.64 tok/s |
| 4,096 | 512 | 5.31 s | 16.26 tok/s |
| 16,384 | 512 | 22.13 s | 9.46 tok/s |
| 28,672 | 512 | 40.94 s | 7.29 tok/s |
Use the metric names from the report. Output throughput, first-token delay, and time per output token are different aggregates; do not present this throughput column as an interchangeable decode-only speed figure from another tool.
These are repository-published results, not a new benchmark executed for this English edition. They replace the earlier release-candidate figures, including its higher preliminary IFEval score.
Turn measurements into a deployment decision
The long-context result makes the cost of a large prompt visible. At the highest tier, the service takes about 41 seconds to produce its first token in this recorded configuration.
For interactive work, that suggests testing prompt trimming, retrieval, and summaries against the same workload before buying more hardware. For batch work, a longer initial wait may be acceptable. The decision depends on the actual task and concurrency requirements; this single-request suite does not establish multi-user capacity.
Keep model-generated code inside the project’s sandbox. Review failure reports as part of the result instead of rerunning until only a favorable score remains.
Downloads and reference records
- LLST source and current setup documentation
- Baseline #001 report and manifests
- Original Chinese article
- License: Apache-2.0.
The value of LLST is a comparison you can inspect later: the samples, workload, tokenizer, configuration, and evaluation outputs travel with the headline numbers.