An AOT-compiled language for high-performance AI development.
Status: alpha. Breaking changes are expected. Per-phase status lives in one place, the Quick Roadmap.
AI development runs on two languages: an interpreted one to write in, and C++ or CUDA underneath for anything that has to be fast. Crossing that boundary is where performance and type safety are lost. Neuro is one language on both sides of it.
- Native code, no interpreter. Compiled ahead of time through LLVM 23, with no bytecode VM and
no global interpreter lock. On scalar compute-bound programs it lands in the same range as
clang++ -O3and Rust; see Performance. - Shapes checked by the compiler.
Tensor<T, [d0, d1]>carries its dimensions in the type, so a dimension mismatch is a compile error rather than an exception thrown ninety minutes into a training run. - Ownership without a garbage collector. Move-by-default with
val/mut, borrows, and deterministic destruction, so there is no collector pause in the middle of a training step and no shared mutable state to make parallel tensor code unsafe.
The full reasoning, including what Neuro deliberately is not, is in DESIGN.md.
A single perceptron with ReLU activation, using structs, impl blocks, instance methods,
if-expressions and implicit returns. This file compiles and runs today.
struct Neuron {
weight: f64,
bias: f64
}
impl Neuron {
func new(weight: f64, bias: f64) -> Neuron {
Neuron { weight: weight, bias: bias }
}
// ReLU activation: pass-through if positive, clamp to zero otherwise
func activate(&self, input: f64) -> f64 {
val z = (input * self.weight) + self.bias
if z > 0.0 { z } else { 0.0 }
}
func is_active(&self, input: f64) -> bool {
val z = (input * self.weight) + self.bias
z > 0.0
}
}
func main() -> i32 {
val neuron = Neuron::new(0.5, -0.1)
// Dead region: input too small to overcome the bias
val dead = neuron.activate(0.0) // 0.0 * 0.5 - 0.1 = -0.1, clamped to 0.0
val dead_fires = neuron.is_active(0.0)
println("input 0.0 -> {dead:.2} fires: {dead_fires}")
// Active region: strong enough input fires the neuron
val active = neuron.activate(1.0) // 1.0 * 0.5 - 0.1 = 0.4, passes through
val active_fires = neuron.is_active(1.0)
println("input 1.0 -> {active:.2} fires: {active_fires}")
// The dead region clamps to exactly zero.
if dead > 0.0 { return 1 }
// Scale the active output into the exit code so the result is observed.
return (active * 10.0) as i32 // 0.4 * 10 = 4
}
input 0.0 -> 0.00 fires: false
input 1.0 -> 0.40 fires: true
More programs, each pinned to its exact exit code and printed output, are in examples/; examples/showcase/ holds the ones that combine several features at once.
| Requirement | Version | Notes |
|---|---|---|
| Rust | 1.98.1+ | Install via rustup |
| LLVM 23 | 23.x with dev libraries | Per-platform commands below |
| C linker | any | gcc / clang on Linux and macOS, MSVC on Windows |
This is the only step that differs between systems. Put the export in your shell profile
(~/.bashrc, ~/.zshrc) so it survives a new terminal.
# Arch Linux / CachyOS
sudo pacman -S llvm
export LLVM_SYS_231_PREFIX=/usr
# Ubuntu / Debian
wget -qO- https://apt.llvm.org/llvm.sh | sudo bash -s -- 23
sudo apt-get install -y llvm-23-dev libpolly-23-dev
export LLVM_SYS_231_PREFIX=/usr/lib/llvm-23
# macOS (Homebrew)
brew install llvm
export LLVM_SYS_231_PREFIX="$(brew --prefix llvm)"Windows needs the MSVC toolchain and LLVM's full development archive,
clang+llvm-23.*-x86_64-pc-windows-msvc.tar.xz: the .exe installer ships Clang and LLVM-C.dll
but no llvm-config.exe, no headers and no static libraries, so llvm-sys cannot build against it. The PowerShell walkthrough is in the
installation guide, and
troubleshooting covers the errors that follow from getting it
wrong.
git clone https://github.com/PanzerPeter/Neuro.git
cd Neuro
cargo build --release
cargo test --workspace
cargo install --path compiler/neurc # optional, puts neurc on your PATHThe same four commands run unchanged in PowerShell on Windows.
neurc check examples/basics/hello.nr # type-check only, no binary
neurc run examples/basics/factorial.nr # compile and run, leaving no binary behind
neurc compile examples/basics/factorial.nr # native executable next to the sourceWithout cargo install, prefix each command with cargo run -p neurc --. Flags, --emit obj and
the zero-copy NumPy recipe are in the CLI guide.
Every row is implemented, tested and usable today. Depth lives in the documentation site and docs/; per-release detail is in CHANGELOG.md.
| Feature | Summary |
|---|---|
| Types and inference | i8 through u64, f16 / bf16 / f32 / f64, bool, char, string; literal suffixes, digit separators, as casts, type aliases |
| Control flow | if / else if / else, while, loop, range-for, labelled break / continue, block-as-value, for over any type implementing the prelude's iterator protocol |
| Functions | Recursion, forward references, implicit returns, named arguments with external labels, higher-order functions, |> pipelines, >> composition |
| Generics and traits | Generic functions, structs and impls, const generics, where clauses, turbofish; required and default methods, associated types, operator traits, impl Trait and dyn Trait dispatch. Fully monomorphized |
| Closures | |x: i32| x * x, move closures, (T) -> R function types, compiled to { fn_ptr, env_ptr } with no heap allocation |
| Structs, enums, newtypes | Fields, functional update ..base, impl blocks and trait impls on structs, enums and newtypes with &self / &mut self / consuming receivers; unit, tuple and struct-field variants carrying any sized payload; @derive(Copy, Clone, Debug, PartialEq) |
| Pattern matching | Exhaustive match over variant, literal, or, range and wildcard patterns with if guards, plus val-binding destructuring of structs and arrays |
| Arrays, tuples, collections | [T; N], tuples, zero-copy slices &[T] / &mut [T], and heap-backed Vec<T> / HashMap<K, V> / BTreeMap<K, V> / StringBuilder (reference) |
| Tensors | Tensor<T, [d0, ...]> with shapes checked at compile time: broadcasting, a @ b matmul, slicing, shape generics, named and dynamic axes, reductions, sorting, einsum, .map / .zip / .reduce, .exp() / .log() / .sqrt() / .tanh() / .abs() / .pow(p), and @gpu functions run as NVIDIA or AMD kernels on Linux, with an opt-in CPU fallback chosen at startup, hand-written @kernel functions launched one thread per element, and tensors kept on any GPU with .to(Device::GPU(n)), where float and integer tensor operations run on them too (reference) |
| Automatic differentiation | @grad compiles a reverse-mode derivative beside a function or method, with no gradient tape; wrt: selection, .backward() / .grad() / .zero_grad(), order: 2 Hessians, .detach() / @no_grad, each checked against finite differences (reference) |
| Strings | Immutable fat-pointer string with slices, concatenation, codepoint iteration, interpolation "{x:.2}" and triple-quoted blocks; growable StringBuilder buffer (reference) |
| Errors | Option<T> and Result<T, E> in the implicit prelude as ordinary generic enums; ?? unwraps with a lazy fallback, ? propagates, val-else exits the scope, checked_* arithmetic and float .to_checked::<T>() report what does not fit |
| Ownership | Move-by-default, Copy, borrows with flow-sensitive exclusivity, lifetime elision, deterministic Drop, and pool { } arena blocks (reference) |
| Modules | Every .nr file is a module, mod.nr directories nest, inline module { } blocks group; import with renames and re-export facades, private-by-default visibility, implicit prelude (reference) |
| Toolchain | neurc check / run / compile on inkwell 0.10 and LLVM 23, --emit obj for C and NumPy interop, buffered print / println, and a panic / assert runtime with located diagnostics |
Alpha memory note. Stack values, literals, the owning collections and reassigned bindings are all reclaimed, and so is a heap
stringstored into a struct field, an array or tuple element, a call's argument, a call's return value, or a collection slot. What still leaks is the handful of storing positions whose owner the compiler cannot prove, each of which holds one buffer rather than handing out a dangling one. Full detail, and the reason the analysis answers conservatively, is in the memory model.
neurc compile -O 3 runs LLVM's -O3 pipeline, the same level the C++ and Rust rows are
built at (clang++ -O3, rustc -C opt-level=3), and all three target the generic x86-64 CPU
with no -march=native. The default is -O 0, checked arithmetic with no optimization, so pass
-O 3 before drawing any conclusion about speed.
Best of nine runs on one machine (Core i5-14600K, RTX 5070), lower is better. Reproduce with
cd benchmarks && uv run python run.py --reps 9 --levels 3, which builds every implementation
of each program and refuses to report timings if they disagree on output. Each one is written
the way a programmer of that language would write it for speed. The Python column is the
language itself; the NumPy column is what a Python programmer would write where the work
vectorizes.
| Benchmark | What it stresses | Neuro -O 3 |
clang++ -O3 |
rustc -O3 |
Python 3.14 | NumPy |
|---|---|---|---|---|---|---|
mandelbrot |
scalar f64 in a tight loop |
166 ms | 166 ms | 166 ms | 6668 ms | 1268 ms |
vector_sum |
Vec push, indexed sweep |
24 ms | 26 ms | 24 ms | 6008 ms | 130 ms |
call_overhead |
recursion, call and inline cost | 49 ms | 51 ms | 48 ms | 1409 ms | |
print_lines |
integer holes to standard output | 13 ms | 20 ms | 13 ms | 111 ms | |
format_floats |
f64 holes at a fixed precision |
115 ms | 110 ms | 38 ms | 222 ms | |
int_divide |
guarded / and %, opaque divisor |
95 ms | 88 ms | 89 ms | 1515 ms | |
matmul |
@ on [256, 256] f32 tensors |
3.5 ms | 3.9 ms | 3.8 ms | 770 ms | 61 ms |
Absolute times belong to the machine. Four rows are worth a word. print_lines beats C because
an integer hole renders through a digit loop instead of snprintf, as Rust's does.
format_floats loses to Rust, whose own float formatter is three times faster than the C
library conversion Neuro and C++ both call. int_divide is the one place the compiler spends:
/ and % guard the operand pairs the hardware leaves undefined, and Rust guards the same pairs
for 7% less. matmul runs each @ as a register-blocked, vectorized loop nest, level with the
i-k-j loops the C++ and Rust versions vectorize, and with the same bits as a plain loop.
The GPU benchmarks run their Neuro side as @gpu kernels, against PyTorch on the same GPU and
NumPy on the CPU. Each moves the same data between host and device:
| Benchmark | What it stresses | Neuro @gpu |
PyTorch (CUDA) | NumPy (CPU) |
|---|---|---|---|---|
gpu_relax |
200 fused element-wise steps on [2048, 2048] |
262 ms | 1080 ms | 619 ms |
gpu_reduce |
1000 rounds of a whole and a row .sum() on [2048, 2048] |
747 ms | 1144 ms | 1251 ms |
gpu_matmul |
five [2048, 2048] f32 products |
292 ms | 1137 ms | 217 ms |
gpu_mlp |
twenty batches through a two-layer perceptron | 275 ms | 1133 ms | 158 ms |
These are whole-program times, and both GPU columns start with a fixed cost: about 220 ms for a Neuro program, most of it the CUDA driver starting up, and about 1100 ms for PyTorch's import and CUDA setup. Past that, timed warm inside one process, PyTorch takes 23, 56, 37 and 13 ms against Neuro's roughly 44, 530, 40 and 25. A long reduction folds in 4096 lanes, one GPU thread each, in the order the host folds in too, so the two agree bit for bit. Each GPU thread computes a block of up to 4 × 4 elements of a matrix product, every element summed in order, so the device keeps the host's bits; there is no shared-memory staging and no tensor-core path.
Each numbered phase is a MAJOR-version milestone: completing Phase N ships v(N+1).0.0.
| Phase | Goal | Status |
|---|---|---|
| 1 | Core Language: types, control flow, LLVM backend, ownership and borrow checking, generics, traits, closures, enums and pattern matching, error handling, modules | Complete |
| 2 | Tensors and MLIR: first-class tensor types lowered through MLIR Linalg, the pool allocator, and the value model they need | Complete |
| 3 | Automatic differentiation: a reverse-mode source-to-source transform over Neuro's own typed HIR, @grad(wrt: ...), .backward() / .zero_grad(), higher-order derivatives, elementwise math, .detach() / @no_grad |
Complete |
| 4 | GPU acceleration: MLIR GPU dialects for NVIDIA and AMD, @gpu with an opt-in CPU fallback, @kernel and the KernelOut<T> aliasing model, device tensors on any GPU |
Complete |
| 5 | Language and backend completion: MLIR as the only tensor backend, generic methods, closure capture modes, references in structs, tensors generic over their element type, named dynamic extents, a standard library written in Neuro, @test |
In progress |
| 6 | Neural network library: typed gradients from grad(...), @model, optimizers, layers and attention, .safetensors weights, a DataLoader, training benchmarks against PyTorch |
Planned |
| 7 | Parallelism and scale: scoped parallel { } tasks, channels, multithreaded host kernels, multi-GPU data parallel |
Planned |
| 8 | Interop and deployment: Python extension modules over DLPack, extern func foreign functions, defer |
Planned |
| 9 | Developer experience: debug info, incremental compilation, Language Server Protocol, error recovery, formatter | Planned |
| 10 | Distribution and optimization: the neurpm package manager, prebuilt binaries and installers, a self-updater, optimization passes |
Planned |
Neuro follows Vertical Slice Architecture: code is organized by language feature, not by technical layer.
compiler/
├── infrastructure/ # Shared, zero-business-logic crates
│ ├── ast-types/ # AST node definitions
│ ├── shared-types/ # Primitives shared across slices
│ └── neuro-hir/ # Typed High-Level IR (frontend <-> backend contract)
├── lexical-analysis/ # Tokenizer (logos, Unicode XID)
├── syntax-parsing/ # Pratt + statement parser -> AST
├── module-resolution/ # Multi-file loading, imports, visibility
├── argument-binding/ # Named arguments -> positional calls
├── semantic-analysis/ # Type checker, scope and borrow analysis
├── hir-lowering/ # Type-checked AST -> typed HIR
├── llvm-backend/ # HIR -> object code (inkwell 0.10 / LLVM 23)
├── mlir-backend/ # HIR -> MLIR linalg and GPU kernels
└── neurc/ # CLI compiler driver
Today a .nr file travels: tokens → AST → merged program → type-checked AST → typed HIR →
LLVM object code → system linker. Automatic differentiation runs inside HIR lowering. The
mlir-backend slice lowers the same typed HIR to MLIR linalg; neurc links the tensor
bodies it computes into the program, and on Linux lowers @gpu and @kernel bodies on through the MLIR GPU dialects to NVIDIA or AMD kernels. Stage by stage:
docs/compiler/compilation.md.
| Getting started | Installation, first program, workflow |
| Language reference | One page per feature area, from types to modules |
| CLI guide | Commands, flags, environment variables, interop |
| Compiler internals | Pipeline and one page per slice |
| Editor support | Syntax highlighting for .nr files |
| DESIGN.md | Why the language is shaped this way, and its non-goals |
| CHANGELOG.md | What each release changed |
Everything is published at neuro-lang.netlify.app.
CONTRIBUTING.md has the architecture rules, coding standards, quality gates and pull request process. Open defects are in docs/BUGS.md, and fixing one is the best way to start. Work is most useful in the phase the Quick Roadmap marks in progress.
See also SECURITY.md and the Code of Conduct.
Neuro Shared Source License v2.1. The license covers the compiler, not what you build with it.
You may freely use, study and modify the compiler for any personal or internal purpose, write Neuro programs and distribute or sell the compiled output under any terms you choose, build tools and editor integrations that call into the compiler, and contribute code back. Only redistributing the compiler itself, or a fork of it, as part of a commercial product requires a commercial license.
Neuro is pre-stabilization, and the license guards three risks specific to that: commercial re-packaging before the spec is stable, AI-assisted reproduction of the compiler for a competing product, and misleading forks that fragment an early ecosystem. A permissive license becomes possible once the language stabilizes.
Inspired by Rust (ownership, type system), Python (AI ecosystem simplicity), Swift (ergonomics) and Mojo (AI-first design). Built with inkwell, logos and LLVM.
