The short version is in the main README. This page lists every recorded run, how it was measured, and its raw data.
Runs A–K and N–S use Qwen2.5 0.5B Instruct, the only model Synapse runs today. Run L measures Microsoft Foundry Local alone on six models from four families.
| Run | Synapse version | Threads | Rounds | Raw data |
|---|---|---|---|---|
| A. 128-token answer, 4 CPU engines | SIMD C# kernels (in progress) | 2 | 1 measured, no warm-up | JSON |
| B. 3-turn dialogue, 4 CPU engines | SIMD C# kernels (in progress) | 2 | 1 measured, no warm-up | JSON |
| C. MLX on the GPU, answer + dialogue | Not included (MLX only) | GPU | 1 warm-up + 3 measured | answer, dialogue |
| D. 8-token smoke test, 3 engines | First scalar C# | 12 | 3 warm-up + 5 measured | JSON |
| E. 8-token memory run, 4 engines | First scalar C# | 8 | 3 warm-up + 5 measured | JSON |
| F. GitHub Actions, 3 operating systems | First scalar C# | 2 | 3 warm-up + 5 measured | workflow run |
| G. llama.cpp alone | Not included | 12 | 3 warm-up + 5 measured | JSON |
| H. 32-token output check | First scalar C# | 8 | 1 measured, no warm-up | JSON |
| I. 8-token, newest (main README) | Rust kernels | 2 | 3 warm-up + 5 measured | JSON |
| J. 8-token, newest | Rust kernels | 8 | 3 warm-up + 5 measured | JSON |
| K. 8-token, newest | SIMD C# kernels | 2 | 3 warm-up + 5 measured | JSON |
| L. Foundry Local, 6 models, answer + dialogue | Not included (Foundry only) | runtime default (about 6 cores) | 1 warm-up + 3 measured | 14 JSON files + 2 default-context probes, results/2026-09-28-m2-pro-foundry-local-* |
| M. Hosted CPU, MLX, and Foundry, 3 operating systems | Rust CPU kernels for Synapse | CPU 2 / external runtime default | CPU smoke 3+5; long/MLX/Foundry 1+3 | GitHub run 36435838291, 22 raw artifacts |
| N. Metal GPU long context, 40k and 131k synthetic tokens | Metal backend (ADR-012), FP32 KV | GPU | 1 measured per row | JSON |
| O. Metal GPU pass-key retrieval, 4k to 120k | Metal backend, FP32 KV (native window) and FP16 KV (YaRN ×4) | GPU | 1 measured per row | native, YaRN, llama.cpp control, corrected |
| P. Tokenizer parity, whole repository | Repo-owned BPE tokenizer (ADR-014) | — | 702 runs, both special modes | JSON |
| Q. Perplexity parity with llama.cpp, 4k to 32k | Metal FP32/FP16 KV, native CPU | GPU / CPU 8 | 1 measured per row | JSON |
| R. Long-context answers, 4 engines, 4k to 32k | Metal FP32/FP16 KV, native CPU | GPU / CPU 8 | 1 measured per case | GPU, CPU |
| S. KV page activation quality (ADR-016) | Native CPU, C# page selection | CPU 8 | 1 measured per row | JSON |
| T. Context sweep, 4k to 32k, tokens + memory + speed (main README) | Metal FP32/FP16 KV with prompt attention reading K/V directly; native CPU | GPU / CPU 8 | 1 warm-up + 2 measured, rotated | JSON, prompt IDs |
| U. Dynamic KV memory, 131k window (ADR-017) | Metal FP16/FP32 KV, growing slots with reservation | GPU | 1 measured per row | JSON |
| V. Second question about a 30k-token document (ADR-018) | Metal FP16 KV, prompt prefix reuse | GPU | 1 measured per row | JSON |
| W. Qwen2.5-7B-1M: layer drop quality and memory (ADR-019) | Metal FP16 KV, head size 128, weight segments | GPU | 1 run per drop set; memory 1–4 runs | qualification 7B, qualification 0.5B, memory |
| Y. Qwen2.5-7B-1M Q4_K_M against Q8_0 and llama.cpp (ADR-021) | Metal FP16 KV, K-quant kernels | GPU | 2 alternating rounds | JSON |
| X. Qwen2.5-7B-1M with a 0.5B draft (ADR-020) | Metal FP16 KV, exact greedy speculation | GPU | 2 alternating rounds | JSON |
Local machine: MacBook Pro, Apple M2 Pro (8 performance + 4 efficiency cores,
19-core GPU), 32 GB, macOS 27.0 arm64. Every CPU sample starts a new process.
Prompts are in scenarios/. Notes on runs A–C:
results/2026-09-28-long-dialogue-and-mlx-evidence.md.
Rules for reading the numbers:
- Runs are separate groups. Do not compare numbers across different runs.
n/ameans the engine does not report that value. It is not zero.- Peak RSS is the whole process: managed, native, and memory-mapped model pages.
- None of these runs is a final ranking. The release gate needs 30 paired runs.
Prompt The capital of France is, median of 5 runs, 4 engines rotated each
round. All engines produced the same 8 tokens in every run. Power state was
not recorded.
| Engine | I. Rust, 2 threads | J. Rust, 8 threads | K. SIMD C#, 2 threads | Whole request (I) | Peak RSS (I) |
|---|---|---|---|---|---|
| Synapse | 136.6 tok/s | 190.6 tok/s | 124.4 tok/s | 216 ms | 555 MiB |
| llama.cpp | 133.7 tok/s | 167.2 tok/s | 141.8 tok/s | 690 ms | 1,256 MiB |
| LLamaSharp | 97.1 tok/s | 147.0 tok/s | 100.1 tok/s | 945 ms | 1,275 MiB |
| dotLLM | 9.7 tok/s | 31.7 tok/s | 9.7 tok/s | 2,227 ms | 1,190 MiB |
First token in run I: Synapse 32 ms, LLamaSharp 32 ms, dotLLM 844 ms. llama.cpp reports only its internal prompt time: 22 ms. llama.cpp speed is its internal generation rate.
Foundry Local SDK 2.0.1 (ONNX Runtime 1.28.0, ONNX Runtime GenAI 0.15.2) with Microsoft's own ONNX packages from the Foundry catalog. These are not the Q8_0 file used in runs A–K, and the SDK does not let us set a thread count: the process used about 6 of the 12 cores. Do not compare these numbers with runs A–K.
Each model and scenario ran in a fresh process. The model stayed loaded for
1 warm-up and 3 measured requests; every request was a new chat session with
the full transcript. Context was capped at 1,024 tokens (see below). Greedy
decoding. AC power; other desktop apps were open (load average about 3.4).
The raw schema still says quality_unreviewed; the manual review below was
performed afterward from the preserved text and reasoning fields and does not
retroactively change measurement eligibility.
128-token answer (capitals-single-long), median of 3:
| Model | Family | Device | File MB | Load ms | Prompt tokens | First token ms | Speed tok/s | Request ms | Peak RSS MiB |
|---|---|---|---|---|---|---|---|---|---|
| Qwen2.5 0.5B | Qwen | CPU | 822 | 832 | 78 | 65 | 231.4 | 614 | 650 |
| Qwen3 0.6B * | Qwen | CPU | 593 | 1,176 | 78 | 100 | 124.3 | 1,122 | 1,493 |
| Phi-3.5-mini | Phi | CPU | 2,590 | 2,878 | 81 | 448 | 42.1 | 3,467 | 3,563 |
| Phi-4-mini | Phi | CPU | 4,915 | 4,353 | 72 | 489 | 28.6 | 4,943 | 4,413 |
| Mistral 7B v0.2 † | Mistral | CPU | 4,167 | 5,570 | 84 | 821 | 25.2 | 5,857 | 5,132 |
| DeepSeek-R1 7B * | DeepSeek | CPU | 6,584 | 5,957 | 68 | 685 | 22.0 | 6,458 | 5,193 |
| Qwen2.5 0.5B | Qwen | GPU (WebGPU) | 700 | 773 | 78 | 45 | 142.3 | 736 | 809 |
Three-turn dialogue (capitals-france-us-uk-3-turns, 64 tokens per turn),
turn 3, median of 3:
| Model | Device | Prompt tokens | First token ms | Speed tok/s | Peak RSS MiB |
|---|---|---|---|---|---|
| Qwen2.5 0.5B | CPU | 293 | 237 | 223.9 | 988 |
| Qwen3 0.6B * | CPU | 293 | 362 | 116.9 | 1,736 |
| Phi-3.5-mini | CPU | 295 | 1,607 | 39.9 | 3,672 |
| Phi-4-mini | CPU | 272 | 1,878 | 28.5 | 4,657 |
| Mistral 7B v0.2 † | CPU | 307 | 3,064 | 24.5 | 5,270 |
| DeepSeek-R1 7B * | CPU | 269 | 2,625 | 21.8 | 5,571 |
| Qwen2.5 0.5B | GPU (WebGPU) | 293 | 113 | 144.1 | 1,008 |
- * Qwen3 and DeepSeek-R1 write reasoning text before the answer. Every generated token is counted; the raw JSON keeps reasoning text apart.
- † Mistral 7B v0.2's chat template has no system role, so the system prompt
opens the first user message (
system_prompt_modein the JSON). - The GPU run used 0.5–0.6 CPU cores instead of 6. Its 128-token answer stopped at 99 tokens.
- Prompt token counts differ because each model has its own tokenizer and chat template.
No model/device variant fully satisfied the 128-token instruction. This makes the throughput table useful for runtime diagnosis, not for a quality-adjusted model ranking.
| Variant | Review of the deterministic measured output |
|---|---|
| Qwen2.5 0.5B CPU | Incorrect and incomplete: it names all three capitals, but calls Fort Knox the federal headquarters and misidentifies Paris landmarks; it hits the 128-token cap before the recap. |
| Qwen2.5 0.5B WebGPU | Mostly factual but incomplete: it stops at 99 tokens without the requested three-line recap. |
| Qwen3 0.6B CPU | Incorrect and incomplete: all 128 tokens are reasoning, it emits no final answer, and the reasoning calls Paris, London, and Madrid France's capitals. |
| Phi-3.5-mini CPU | Incorrect and incomplete: it corrupts Washington, D.C., places Mount Rushmore near it, and reaches the cap during the recap. |
| Phi-4-mini CPU | The three main country-capital sections are substantially correct, but the generated recap is malformed and truncated at the cap. |
| Mistral 7B v0.2 CPU | The France section is correct, but the output reaches the cap during the United States section before covering the United Kingdom or recap. |
| DeepSeek-R1 7B CPU | Incomplete: all 128 tokens are reasoning and no final answer is emitted, although the reasoning recalls the three capitals. |
The three-turn outputs have the same constraint: every CPU turn reaches its 64-token cap. DeepSeek-R1 and Qwen3 remain reasoning-only. Phi-3.5, Phi-4, and Mistral contain mostly correct briefing content but truncate sections; the Qwen2.5 variants additionally contain errors such as placing Paris in the center or south of France and describing Washington transit as free. No instruction-compliance claim is made from those dialogue timings.
Context cap. Every Foundry package sets its context length to the model's
full window, and ONNX Runtime GenAI reserves the KV memory for that whole
window at the first request. With the package default, Qwen3 0.6B reached
4.9 GiB RSS and a 1.2 s first token, and Phi-3.5-mini (131,072 tokens) reached
a 98 GiB compressed footprint and a 12.4 s first token. Raw data:
Qwen3,
Phi-3.5-mini
(one 32-token request each). The runner's fetch therefore lowers only
search.max_length to 1,024 and keeps the original config beside it.
GitHub Actions. The performance workflow runs one isolated job per runner
and model: a model runs only where its file is at most half of the runner's
RAM. That is 3 models on macos-15 (7 GB) and all 6 on ubuntu-24.04 and
windows-2025 (16 GB), 15 jobs in total. All 15 jobs passed in
run 36435838291
and retained separate raw answer/dialogue JSON. Their raw quality state remains
unreviewed; a green job confirms execution and artifact delivery, not answer
correctness.
The performance run passed all 20 jobs and retained 22 raw artifacts: three eight-token CPU matrices, three CPU answer/dialogue pairs, one MLX answer/dialogue pair, and 15 Foundry model/runner answer/dialogue pairs. Each row below is a median of measured rounds on that runner; warm-ups are excluded. CPU wall time includes a fresh process and model load. MLX and Foundry request wall time excludes their resident model load and belongs to separate weight/runtime cohorts.
| CPU runner | 8-token Synapse / llama.cpp wall ms | 8-token Synapse / llama.cpp RSS MiB | 128-token Synapse / llama.cpp wall ms | 128-token Synapse / llama.cpp reported decode tok/s |
|---|---|---|---|---|
| macOS M1 | 526 / 1,671 | 551 / 1,203 | 5,856 / 4,223 | 32.5 / 59.0 |
| Ubuntu x64 | 537 / 635 | 553 / 724 | 4,503 / 3,608 | 36.5 / 47.0 |
| Windows x64 | 557 / 681 | 539 / 575 | 4,936 / 3,330 | 30.3 / 52.0 |
All measured eight-token outputs matched. The 128-token and three-turn CPU
outputs retain quality_status: unreviewed; their timings are not eligible for
a quality-adjusted winner verdict. Hosted MLX produced a 128-token answer at
73.0 tok/s and 1,799 ms request wall on the M1, with three different measured
continuations. Hosted Foundry scheduled 3/6/6 models on Mac/Ubuntu/Windows;
Qwen3 and DeepSeek again spent the 128-token answer budget in reasoning without
a final answer. Foundry and MLX are not ranked against the GGUF CPU subjects.
The workflow's final job downloads every raw artifact and publishes one per-run Markdown table with runner, scenario, turn, subject, measured-round count, median timing/memory, output state, and source artifact. Missing or invalid evidence is listed and makes that reporting job fail.
Method
- Prompt. One prompt per context: the pinned repository haystack, with a
hidden number at 50% depth. The instruction asks for that number, then a
detailed summary (
experiments ... sweep). The exact prompt IDs are inscenarios/context-sweep/. - Runs. Every engine gets the same text, and its prompt token IDs or counts are checked. Each sample is a fresh process: 1 warm-up plus 2 measured rounds, with the engine order rotated. Greedy decoding, up to 128 tokens.
- Columns:
- KV MiB is the cache the context needs, computed from the model geometry;
- peak RSS and peak footprint are sampled from the whole process;
- TTFT and full generation are each engine's own phase clock (llama.cpp perf lines, Synapse CLI timings, and client-side streaming for MLX);
- same tokens as reference counts the leading output tokens shared with llama.cpp Metal FP16 KV, both re-tokenized by the repo tokenizer.
- Merged sources:
- the main sweep;
- an MLX rerun, which replaced its rows because SwiftLM's resident prompt cache had answered repeated prompts: a fresh server per sample now;
- a Synapse Metal rerun after the prompt-attention change.
| context | engine | KV | KV MiB | peak RSS MiB | peak footprint MiB | TTFT s | full generation s | decode tok/s | tokens | correct | same tokens as reference |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 4096 | llamacpp-cpu | f16 | 48 | 1257 | 648 | 8.04 | 9.64 | 79.7 | 128 | 2/2 | 51 |
| 4096 | llamacpp-metal | f16 | 48 | 828 | 125 | 0.85 | 1.75 | 141.4 | 128 | 2/2 | 128 |
| 4096 | llamacpp-metal | f32 | 96 | 876 | 174 | 0.96 | 3.01 | 62.1 | 128 | 2/2 | 128 |
| 4096 | llamacpp-metal | q8_0 | 26 | 808 | 103 | 0.92 | 1.97 | 120.4 | 128 | 2/2 | 59 |
| 4096 | mlx-swiftlm | native | — | 657 | 1138 | 0.81 | 1.70 | 143.1 | 128 | 2/2 | 38 |
| 4096 | synapse-metal | f16 | 48 | 605 | 173 | 1.26 | 2.11 | 148.4 | 128 | 2/2 | 128 |
| 4096 | synapse-metal | f32 | 96 | 602 | 222 | 1.30 | 2.21 | 140.3 | 128 | 2/2 | 128 |
| 4096 | synapse-native | f32 | 96 | 697 | 161 | 15.68 | 17.61 | 65.8 | 128 | 2/2 | 51 |
| 8192 | llamacpp-cpu | f16 | 96 | 1270 | 703 | 32.23 | 34.60 | 53.6 | 128 | 2/2 | 43 |
| 8192 | llamacpp-metal | f16 | 96 | 884 | 181 | 2.47 | 3.47 | 127.3 | 128 | 2/2 | 128 |
| 8192 | llamacpp-metal | f32 | 192 | 981 | 277 | 2.64 | 6.17 | 36.0 | 128 | 2/2 | 128 |
| 8192 | llamacpp-metal | q8_0 | 51 | 841 | 137 | 2.51 | 3.64 | 112.3 | 128 | 2/2 | 52 |
| 8192 | mlx-swiftlm | native | — | 656 | 1898 | 2.05 | 3.16 | 114.4 | 128 | 0/2 | 1 |
| 8192 | synapse-metal | f16 | 96 | 606 | 223 | 3.30 | 4.22 | 137.0 | 128 | 2/2 | 76 |
| 8192 | synapse-metal | f32 | 192 | 607 | 319 | 3.54 | 4.51 | 130.8 | 128 | 2/2 | 128 |
| 8192 | synapse-native | f32 | 192 | 805 | 265 | 59.75 | 62.89 | 40.4 | 128 | 2/2 | 67 |
| 16384 | llamacpp-metal | f16 | 192 | 1004 | 300 | 7.40 | 8.50 | 115.9 | 128 | 0/2 | 128 |
| 16384 | llamacpp-metal | f32 | 384 | 1189 | 487 | 8.10 | 14.23 | 20.7 | 128 | 0/2 | 128 |
| 16384 | llamacpp-metal | q8_0 | 102 | 909 | 204 | 7.55 | 8.88 | 95.2 | 128 | 0/2 | 1 |
| 16384 | mlx-swiftlm | native | — | 657 | 4401 | 5.32 | 6.76 | 88.3 | 128 | 0/2 | 42 |
| 16384 | synapse-metal | f16 | 192 | 606 | 327 | 8.99 | 10.06 | 118.1 | 128 | 0/2 | 128 |
| 16384 | synapse-metal | f32 | 384 | 607 | 521 | 9.89 | 11.11 | 104.7 | 128 | 0/2 | 128 |
| 32768 | llamacpp-metal | f16 | 384 | 1208 | 535 | 26.62 | 26.81 | 92.9 | 19 | 2/2 | 18 |
| 32768 | llamacpp-metal | f32 | 768 | 1610 | 906 | 27.36 | 29.04 | 10.7 | 19 | 2/2 | 18 |
| 32768 | llamacpp-metal | q8_0 | 204 | 1049 | 344 | 25.52 | 25.76 | 75.1 | 19 | 2/2 | 18 |
| 32768 | mlx-swiftlm | native | — | 657 | 13500 | 15.51 | 17.62 | 60.1 | 128 | 0/2 | 0 |
| 32768 | synapse-metal | f16 | 384 | 601 | 534 | 27.63 | 27.82 | 92.2 | 19 | 2/2 | 18 |
| 32768 | synapse-metal | f32 | 768 | 612 | 917 | 37.38 | 37.72 | 53.8 | 19 | 2/2 | 18 |
Reading it
- Engine agreement.
- Synapse Metal wrote the same 128 tokens as llama.cpp at 4k and 16k, and the same whole answer at 32k. At 8k the FP16 output parts after 76 tokens at a near-tie (FP32 KV keeps all 128). llama.cpp's own FP32 and Q8 caches also leave its FP16 output at 32k after 18 tokens.
- At 16k every engine missed the hidden number, and at 8k MLX (with its own 8-bit weights) missed it too. These are model limits, not engine errors.
- Writing speed. Synapse with FP16 KV matches llama.cpp with FP16 KV (92.2 against 92.9 tokens/s at 32k). With FP32 KV, llama.cpp falls to 10.7 tokens/s at 32k, while Synapse keeps 53.8.
- Time to first token. Prompt attention now reads K and V straight from the cache, with the softmax state in registers. Synapse is 1.04× llama.cpp at 32k (27.6 s against 26.6 s), 1.2× at 16k, and 1.5× at 4k. MLX is 1.6–1.8× faster. What is left is the prompt matrix multiply, about 1.55× slower at 512 tokens.
- Memory.
- The footprint counts memory the process owns. Synapse and llama.cpp also map the 644 MiB model file, while MLX copies its weights in. MLX peaks at 13.5 GB at 32k.
- Synapse RSS stays near 610 MiB at every context while its footprint grows with the KV cache: the Metal buffers appear in the footprint, not in RSS.
- Limitations.
- These are diagnostics with 2 measured rounds each, not a paired release verdict.
- Across two Synapse runs of identical decode code, 4k and 8k decode varied by 4–6%.
- Synapse FP32 KV at 32k is sensitive to heat. Repeated runs of the same code gave a 31–37 s first token and 45–73 tokens/s; the table keeps the sweep's own medians.
- CPU rows run up to 8k only.
JSON.
One fresh process per row, /usr/bin/time -l peak footprint.
| Window | Prompt | KV | Synapse | llama.cpp |
|---|---|---|---|---|
| 131,072 (YaRN ×4) | 3,904 tokens | f16 | 234 MiB | 1,645 MiB |
| 131,072 (YaRN ×4) | 3,904 tokens | f32 | 282 MiB | 3,159 MiB |
| 32,768 | 32,640 tokens | f16 | 534 MiB | 535 MiB (run T) |
- The KV cache grows in 1,024-position steps and at least doubles, instead of
being reserved for the whole window.
Generatereserves the prompt plus the maximum answer once. Without that reservation, growing through 32k peaked at 686 MiB. - llama.cpp reserves the whole
-cwindow, so its footprint follows the window, not the prompt.
JSON. Metal, FP16 KV, a 32k window, and the first 30,000 tokens of the pinned haystack. One model asks two questions, 32 new tokens each.
| Prefix reuse | Question | Prompt tokens | Reused | First token |
|---|---|---|---|---|
| off | 1 | 30,031 | 0 | 23.46 s |
| off | 2 | 30,027 | 0 | 23.76 s |
| on | 1 | 30,031 | 0 | 23.49 s |
| on | 2 | 30,027 | 30,015 | 0.13 s |
- With reuse on and off the answers are the same text.
- Reuse is in-process and opt-in (
ModelLoadOptions.ReusePromptPrefix). A persistent prefix cache across processes is the next step. - One run per row: a diagnostic, not a paired verdict.
Quality. Perplexity over 1,024 scored positions of the pinned haystack (experiments layer-drop); dense
12.163. Top-1 is the share of positions whose arg-max equals the dense model's.
| Dropped layers | Perplexity change | Top-1 as dense |
|---|---|---|
| 12 | −3.6% | 86.5% |
| 11 | −2.5% | 85.7% |
| 8 | +0.8% | 84.6% |
| 0 (the first) | +2,467,089% | 0.3% |
| 27 (the last) | +251% | 67.5% |
| 11, 12 (qualified) | −1.1% | 79.8% |
| 8, 11, 12, 14 | +39.2% | 66.4% |
| 8 middle layers | +137.5% | 50.9% |
- A single dropped layer that lowers perplexity still changes about 14% of the predictions.
- The 0.5B has no cheap layer: the best single drop costs +7.6%.
Memory and speed. 3,528-token prompt, 128 tokens, one fresh process per row.
| Variant | Peak RSS | First token | Decode tok/s |
|---|---|---|---|
| Synapse dense | 7,255 MiB | 17.6 s | 13.3–17.0 |
| Synapse, 2 layers dropped (qualified) | 6,780 MiB | 17.5 s | 13.9–15.0 |
| Synapse, 4 layers dropped | 6,326 MiB | 16.0 s | 14.7 |
| Synapse, 8 layers dropped | 5,381 MiB | 13.3 s | 17.9 |
| llama.cpp Q8_0 | 8,096 MiB | 11.5 s | 18.6 |
| llama.cpp Q4_K_M | 4,742 MiB | 12.2 s | 29.2 |
- Metal wraps only the weight segments the kept layers use; before that, a 2-layer drop still peaked at 7,271 MiB.
- Decode varied by ±12% across identical runs on a warm machine, so only the memory column compares run by run.
A 37-token chat prompt and 256 tokens; rounds alternate dense and speculative runs.
| Draft tokens | Decode tok/s, round 1 / 2 | Acceptance | Tokens per 7B pass |
|---|---|---|---|
| — (dense) | 17.4 / 18.3 | — | 1 |
| 3 | 23.5 / 22.5 | 54% | 2.63 |
| 5 | 14.0 / 13.5 | 40% | 3.00 |
| 7 | 14.2 / 13.9 | 32% | 3.27 |
- The output is the 7B's own greedy output. The token identity of the two models is checked before running.
- A 4-token check costs about 1.7 single-token passes; 6–8 tokens cost about 3, so 3 draft tokens is the default.
- On the repository-summary prompt acceptance was 27–46% and speculation was slower than dense.
The same file runs in both engines. 3,528-token prompt, 128 tokens, two alternating rounds.
| Engine and weights | Peak RSS | First token | Decode tok/s |
|---|---|---|---|
| Synapse Q4_K_M | 4,233–4,248 MiB | 19.9–20.1 s | 17.4–18.2 |
| llama.cpp Q4_K_M | 4,846–4,851 MiB | 12.1–12.2 s | 27.0–29.3 |
| Synapse Q8_0 | 7,210–7,252 MiB | 16.9–17.3 s | 15.8–17.9 |
| llama.cpp Q8_0 | 7,866–8,109 MiB | 10.8–11.2 s | 17.3–17.7 |
- For the pinned prompt, Synapse Q4_K_M writes the same text as llama.cpp.
- Perplexity on the same 2,048 tokens is 11.93 for Q4_K_M against 12.16 for Q8_0.
- With the 0.5B draft, Q4_K_M decode drops to 14.4–14.8 tokens/s: acceptance is 40%, and a 4-token check still costs too much.
Details and tables: docs/Features/LongContext.md
and docs/Features/Tokenization.md.
-
P.
synapse tokenizematchedllama-tokenize --no-escapein 702 of 702 runs (351 files, 540,679 tokens). -
Q. Perplexity on identical tokens:
- Synapse against llama.cpp: 12.7724 vs 12.7728 at 4k, 6.8867 vs 6.8868 at 16k, and 3.9686 vs 3.9687 at 32k. Every chunk agreed within 0.01%.
- FP16 KV equals FP32 KV to the fourth decimal.
-
R. 28 exact-answer cases:
- Synapse FP32 KV, Synapse FP16 KV, and llama.cpp answered byte-identically in 28 of 28. MLX, with its own weights, agreed in 23.
- Every engine shares the same misses, which are 0.5B model limits: variable tracking 0 of 4, multi-key 9 of 12, and the single needle at 32k and 50% depth.
- The corrected 120k control (llama.cpp
-fdrops a trailing newline) answers6, like Synapse.
-
S. Query-aware KV page activation, measured by decode-mode perplexity:
- key-bound selection beats the equal-budget random control at every budget;
- with 16-token pages and 2,048 tokens it costs +3.4% at 8k and +6.2% at 16k.
It is CPU-only so far, so no speed or memory claim is made.
| Engine | Load ms | First token ms | Speed tok/s | Process wall ms | Peak RSS MiB |
|---|---|---|---|---|---|
| Synapse | 86 | 424 | 104.5 | 1,819 | 570 |
| llama.cpp | n/a | n/a (prompt 170) | 120.3 | 1,928 | 1,251 |
| LLamaSharp | 745 | 137 | 115.5 | 2,108 | 1,272 |
| dotLLM | 398 | 8,122 | 9.8 | 21,823 | 1,290 |
llama.cpp's CLI reports only its internal prompt and generation times. The longer answers differ between engines after the first sentence and have not been quality-reviewed.
Each turn starts a fresh process with the whole transcript, so there is no KV-cache reuse.
| Turn · prompt tokens | Synapse first token / speed | LLamaSharp first token / speed | llama.cpp prompt / speed | dotLLM first token / speed |
|---|---|---|---|---|
| 1 · 70 | 425 ms / 74 tok/s | 133 ms / 114 tok/s | 161 ms / 140 tok/s | 8,153 ms / 9.7 tok/s |
| 2 · 165 | 841 ms / 100 tok/s | 287 ms / 87 tok/s | 345 ms / 129 tok/s | 17,849 ms / 9.4 tok/s |
| 3 · 274 | 1,391 ms / 95 tok/s | 473 ms / 105 tok/s | 658 ms / 108 tok/s | 29,592 ms / 9.6 tok/s |
Peak RSS in turn 3: Synapse 574 MiB, LLamaSharp 1,274 MiB, llama.cpp 1,263 MiB, dotLLM 1,545 MiB. In turn 3, LLamaSharp stopped at 58 tokens and llama.cpp at 61; the others wrote 64.
| Request | Prompt / output tokens | First token ms | Speed tok/s | Request wall ms | Peak RSS MiB | Prompt tokens served from cache |
|---|---|---|---|---|---|---|
| Answer, 128 tokens | 78 / 128 | 13.8 | 225.5 | 576.5 | 655.6 | 78 |
| Dialogue turn 1 | 77 / 64 | 14.2 | 226.5 | 292.2 | 651.9 | 77 |
| Dialogue turn 2 | 178 / 64 | 40.5 | 225.1 | 320.4 | 651.9 | 77 |
| Dialogue turn 3 | 293 / 64 | 42.3 | 222.7 | 325.6 | 651.9 | 178 |
- The first cold request took 91 ms to the first token. Later requests reused the server's prompt cache.
- Request wall time excludes the one-time model load (629–657 ms).
- MLX uses its own 8-bit weights (Qwen2.5 0.5B MLX 8-bit) and its own chat template, so prompt token counts differ from the CPU runs.
- Archive SHA-256
2ed6b5539b24c5267931d46ea9973775b7d2a9b5ee2f82109afab60f9603675e, weights SHA-2563dd0b6c2983ac5fe35f60ba260b1c7c35e4c38f17f3d5139d0bb477924e7aef4. No Swift build or Python was used.
Prompt The capital of France is, 512-token context, median of 5 runs.
| Engine | Load ms | First token ms | Generation ms | Speed tok/s | Process wall ms | Avg CPU cores | Peak RSS MiB |
|---|---|---|---|---|---|---|---|
| Synapse (scalar) | 56.1 | 328.3 | 633.3 | 24.30 | 691.3 | 10.23 | 559.8 |
| dotLLM | 345.5 | 439.6 | 875.6 | 18.22 | 1,352.5 | 5.66 | 1,190.5 |
| LLamaSharp | 691.2 | 17.2 | 69.0 | 136.16 | 780.9 | 1.76 | 1,137.2 |
All engines produced the same 8 tokens, Paris. It is the largest city in
(Synapse IDs 12095, 13, 1084, 374, 279, 7772, 3283, 304). Synapse's first
token varied from 304.8 to 574.2 ms. Energy was not measured:
powermetrics needs superuser access.
Median of 5 runs, 4 engines rotated each round.
| Engine | Peak RSS MiB | Physical footprint MiB | .NET live heap MiB | Process wall ms |
|---|---|---|---|---|
| Synapse (scalar) | 560.8 | 39.9 | 20.4 | 895.1 |
| dotLLM | 1,187.3 | 663.1 | not measured | 1,350.9 |
| LLamaSharp | 1,274.9 | 606.9 | 1.1 | 962.6 |
| llama.cpp | 1,258.6 | 594.6 | n/a | 681.6 |
RSS and physical footprint are two different macOS views of the same process and must not be added. Synapse maps the model file, so its RSS is far above its footprint. The .NET heap is only part of total memory.
Median of 5 runs. Each operating system is a separate hardware group, not a
leaderboard. Runners: macos-15 Apple M1, 3 vCPU, 7 GB; ubuntu-24.04 and
windows-2025 x64, 4 vCPU, 16 GB.
Writing speed, tokens/s:
| Runner | Synapse | llama.cpp | LLamaSharp | dotLLM |
|---|---|---|---|---|
| macOS 15 ARM64 | 5.2 | 90.1 | 76.8 | 6.5 |
| Ubuntu 24.04 x64 | 4.6 | 54.8 | 39.4 | 16.6 |
| Windows Server 2025 x64 | 4.0 | 45.5 | 33.8 | 14.7 |
Process wall and peak RSS:
| Runner | Synapse | llama.cpp | LLamaSharp | dotLLM |
|---|---|---|---|---|
| macOS 15 ARM64 | 2,543 ms / 555 MiB | 978 ms / 1,195 MiB | 1,046 ms / 1,263 MiB | 2,880 ms / 1,166 MiB |
| Ubuntu 24.04 x64 | 2,815 ms / 565 MiB | 535 ms / 725 MiB | 646 ms / 760 MiB | 1,621 ms / 1,174 MiB |
| Windows Server 2025 x64 | 3,502 ms / 551 MiB | 1,014 ms / 575 MiB | 1,092 ms / 596 MiB | 2,424 ms / 1,153 MiB |
All four engines produced the same 8 tokens on every runner.
- 8 tokens: median process wall 566.2 ms, peak RSS 1,208.3 MiB, internal prompt 14.9 ms, generation 57.9 ms (120.84 tok/s).
- 128 tokens: 47.28 to 99.40 tok/s over 5 runs (median 80.73). The Mac was busy with other work, so the spread is large.
llama-bench: 115.36 ± 5.94 tok/s. It skips tokenization and sampling, so it measures kernels, not a full request.
Status ineligible_quality_mismatch: dotLLM's answer diverged from LLamaSharp
and llama.cpp after a shared start. Synapse has no tokenizer yet, so its
32-token text cannot be compared. These timings are not a speed result.
The runner is a C# tool in
experiments/Synapse.ReferenceBenchmarks.
The performance workflow
shows the exact commands and pinned engine versions.