Synapse is a local-first inference engine whose product code is owned in this monorepo. C# owns the public SDK, typed IRs, portable/reference execution, orchestration, policy, diagnostics, durable metadata adapters, and the first implementation of every path. Rust takes optimized kernels, allocators, hot-KV operations, and direct transfer only when a paired profile proves that boundary is necessary. The portable C# path remains executable and tested.
.NET SDK / CLI / Host
|
+-- in-process local graph runtime (no Orleans/network)
|
+-- Orleans request/control plane (planned D3)
| leases, epochs, region/weight-group placement
v
graph worker ===== direct bounded data ===== graph worker
|
+-- verified Model IR -> Execution IR
+-- C# portable/reference kernels and KV semantics
+-- profiled Rust/native kernels, allocator, hot KV, transfer
ZoneTree (embedded, C#)
+-- package/cache metadata
+-- prefix index and state journal
+-- benchmark/evidence catalog
Orleans (later D3 request/control plane only)
+-- accepts/co-ordinates distributed requests
+-- leases, placement, weight groups, epochs, and recovery
The central FlyBrain execution model is a bounded activation wave over coarse
regions. Model IR answers what may execute; Execution IR decides how and
with which precision; the deployment plan decides where. Region labels
are annotations only. A skippable region declares a graph decision or profile,
its structural/programmed/trained/approximate provenance, and explicit absent
or bypassed output behavior. See
docs/ADR/ADR-002-flybrain-activation-waves.md.
Large tensors, activations, and KV payloads never transit Orleans. There is no grain per token, layer, node, or neuron. Local execution never requires Orleans, a coordination server, or a network. Unsupported models, operators, profiles, or devices fail before generation with structured errors.
dotLLM, LLamaSharp, and direct llama.cpp are launched as separate benchmark subjects. They are not product backends and cannot satisfy Synapse correctness tests. Every comparison records its exact commit/package/native-binary fingerprint and uses the same compatible model inputs. MLX and ONNX Runtime GenAI are candidate external subjects only in qualified hardware/weight-format cohorts. dotLLM source is GPL-3.0 and is not copied into this MIT repository.
The bootstrap doctor proves the installed .NET/Rust toolchains, current CPU
architecture and SIMD capability, memory budget validation, a real Rust child
process, and a durable ZoneTree write/reopen/read. The first inference slice
adds a bounded GGUF v3 reader, a verified 26-region Qwen2 Model IR, and a
managed Qwen2 Q8_0 forward/decode path. Text enters through a repo-owned
byte-level BPE tokenizer read from the GGUF vocabulary (ADR-014), qualified by
whole-corpus parity with llama-tokenize. The first scalar reference
operators and a fixed-shape dense tiny-graph interpreter are covered by FP64
and fail-closed tests. Qwen execution from the IR, paged KV, and the Orleans
topology remain on the critical path. GPU execution (ADR-012) runs a
brand-neutral dense-decoder layout through native/synapse-gpu: Metal is
implemented and tested on Apple silicon; CUDA shares the same step encoder and
ABI but has not run on NVIDIA hardware. Context limits and explicit YaRN
scaling are ADR-013. Quality is measured by teacher-forced scoring on every
backend (perplexity with the llama-perplexity protocol) and by exact-answer
long-context tasks, with every engine given identical tokens (ADR-015). Model
acquisition and import boundaries are in ADR-004 and
docs/Features/ModelPackages.md.