Skip to content

Repository files navigation

Neuro Programming Language

An AOT-compiled language for high-performance AI development.

neurc type-checks, compiles and runs a Neuro program that fits a model with @grad, in under a second

License: Neuro Shared Source License v2.1 Documentation LLVM CI

Status: alpha. Breaking changes are expected. Per-phase status lives in one place, the Quick Roadmap.


Why Neuro

AI development runs on two languages: an interpreted one to write in, and C++ or CUDA underneath for anything that has to be fast. Crossing that boundary is where performance and type safety are lost. Neuro is one language on both sides of it.

  • Native code, no interpreter. Compiled ahead of time through LLVM 23, with no bytecode VM and no global interpreter lock. On scalar compute-bound programs it lands in the same range as clang++ -O3 and Rust; see Performance.
  • Shapes checked by the compiler. Tensor<T, [d0, d1]> carries its dimensions in the type, so a dimension mismatch is a compile error rather than an exception thrown ninety minutes into a training run.
  • Ownership without a garbage collector. Move-by-default with val/mut, borrows, and deterministic destruction, so there is no collector pause in the middle of a training step and no shared mutable state to make parallel tensor code unsafe.

The full reasoning, including what Neuro deliberately is not, is in DESIGN.md.


Quick Example

A single perceptron with ReLU activation, using structs, impl blocks, instance methods, if-expressions and implicit returns. This file compiles and runs today.

struct Neuron {
    weight: f64,
    bias: f64
}

impl Neuron {
    func new(weight: f64, bias: f64) -> Neuron {
        Neuron { weight: weight, bias: bias }
    }

    // ReLU activation: pass-through if positive, clamp to zero otherwise
    func activate(&self, input: f64) -> f64 {
        val z = (input * self.weight) + self.bias
        if z > 0.0 { z } else { 0.0 }
    }

    func is_active(&self, input: f64) -> bool {
        val z = (input * self.weight) + self.bias
        z > 0.0
    }
}

func main() -> i32 {
    val neuron = Neuron::new(0.5, -0.1)

    // Dead region: input too small to overcome the bias
    val dead = neuron.activate(0.0)    // 0.0 * 0.5 - 0.1 = -0.1, clamped to 0.0
    val dead_fires = neuron.is_active(0.0)
    println("input 0.0 -> {dead:.2}  fires: {dead_fires}")

    // Active region: strong enough input fires the neuron
    val active = neuron.activate(1.0)  // 1.0 * 0.5 - 0.1 = 0.4, passes through
    val active_fires = neuron.is_active(1.0)
    println("input 1.0 -> {active:.2}  fires: {active_fires}")

    // The dead region clamps to exactly zero.
    if dead > 0.0 { return 1 }

    // Scale the active output into the exit code so the result is observed.
    return (active * 10.0) as i32      // 0.4 * 10 = 4
}
input 0.0 -> 0.00  fires: false
input 1.0 -> 0.40  fires: true

More programs, each pinned to its exact exit code and printed output, are in examples/; examples/showcase/ holds the ones that combine several features at once.


Installation

Requirement Version Notes
Rust 1.98.1+ Install via rustup
LLVM 23 23.x with dev libraries Per-platform commands below
C linker any gcc / clang on Linux and macOS, MSVC on Windows

1. LLVM 23

This is the only step that differs between systems. Put the export in your shell profile (~/.bashrc, ~/.zshrc) so it survives a new terminal.

# Arch Linux / CachyOS
sudo pacman -S llvm
export LLVM_SYS_231_PREFIX=/usr

# Ubuntu / Debian
wget -qO- https://apt.llvm.org/llvm.sh | sudo bash -s -- 23
sudo apt-get install -y llvm-23-dev libpolly-23-dev
export LLVM_SYS_231_PREFIX=/usr/lib/llvm-23

# macOS (Homebrew)
brew install llvm
export LLVM_SYS_231_PREFIX="$(brew --prefix llvm)"

Windows needs the MSVC toolchain and LLVM's full development archive, clang+llvm-23.*-x86_64-pc-windows-msvc.tar.xz: the .exe installer ships Clang and LLVM-C.dll but no llvm-config.exe, no headers and no static libraries, so llvm-sys cannot build against it. The PowerShell walkthrough is in the installation guide, and troubleshooting covers the errors that follow from getting it wrong.

2. Build

git clone https://github.com/PanzerPeter/Neuro.git
cd Neuro
cargo build --release
cargo test --workspace

cargo install --path compiler/neurc   # optional, puts neurc on your PATH

The same four commands run unchanged in PowerShell on Windows.

3. Run something

neurc check examples/basics/hello.nr        # type-check only, no binary
neurc run   examples/basics/factorial.nr    # compile and run, leaving no binary behind
neurc compile examples/basics/factorial.nr  # native executable next to the source

Without cargo install, prefix each command with cargo run -p neurc --. Flags, --emit obj and the zero-copy NumPy recipe are in the CLI guide.


Current Capabilities

Every row is implemented, tested and usable today. Depth lives in the documentation site and docs/; per-release detail is in CHANGELOG.md.

Feature Summary
Types and inference i8 through u64, f16 / bf16 / f32 / f64, bool, char, string; literal suffixes, digit separators, as casts, type aliases
Control flow if / else if / else, while, loop, range-for, labelled break / continue, block-as-value, for over any type implementing the prelude's iterator protocol
Functions Recursion, forward references, implicit returns, named arguments with external labels, higher-order functions, |> pipelines, >> composition
Generics and traits Generic functions, structs and impls, const generics, where clauses, turbofish; required and default methods, associated types, operator traits, impl Trait and dyn Trait dispatch. Fully monomorphized
Closures |x: i32| x * x, move closures, (T) -> R function types, compiled to { fn_ptr, env_ptr } with no heap allocation
Structs, enums, newtypes Fields, functional update ..base, impl blocks and trait impls on structs, enums and newtypes with &self / &mut self / consuming receivers; unit, tuple and struct-field variants carrying any sized payload; @derive(Copy, Clone, Debug, PartialEq)
Pattern matching Exhaustive match over variant, literal, or, range and wildcard patterns with if guards, plus val-binding destructuring of structs and arrays
Arrays, tuples, collections [T; N], tuples, zero-copy slices &[T] / &mut [T], and heap-backed Vec<T> / HashMap<K, V> / BTreeMap<K, V> / StringBuilder (reference)
Tensors Tensor<T, [d0, ...]> with shapes checked at compile time: broadcasting, a @ b matmul, slicing, shape generics, named and dynamic axes, reductions, sorting, einsum, .map / .zip / .reduce, .exp() / .log() / .sqrt() / .tanh() / .abs() / .pow(p), and @gpu functions run as NVIDIA or AMD kernels on Linux, with an opt-in CPU fallback chosen at startup, hand-written @kernel functions launched one thread per element, and tensors kept on any GPU with .to(Device::GPU(n)), where float and integer tensor operations run on them too (reference)
Automatic differentiation @grad compiles a reverse-mode derivative beside a function or method, with no gradient tape; wrt: selection, .backward() / .grad() / .zero_grad(), order: 2 Hessians, .detach() / @no_grad, each checked against finite differences (reference)
Strings Immutable fat-pointer string with slices, concatenation, codepoint iteration, interpolation "{x:.2}" and triple-quoted blocks; growable StringBuilder buffer (reference)
Errors Option<T> and Result<T, E> in the implicit prelude as ordinary generic enums; ?? unwraps with a lazy fallback, ? propagates, val-else exits the scope, checked_* arithmetic and float .to_checked::<T>() report what does not fit
Ownership Move-by-default, Copy, borrows with flow-sensitive exclusivity, lifetime elision, deterministic Drop, and pool { } arena blocks (reference)
Modules Every .nr file is a module, mod.nr directories nest, inline module { } blocks group; import with renames and re-export facades, private-by-default visibility, implicit prelude (reference)
Toolchain neurc check / run / compile on inkwell 0.10 and LLVM 23, --emit obj for C and NumPy interop, buffered print / println, and a panic / assert runtime with located diagnostics

Alpha memory note. Stack values, literals, the owning collections and reassigned bindings are all reclaimed, and so is a heap string stored into a struct field, an array or tuple element, a call's argument, a call's return value, or a collection slot. What still leaks is the handful of storing positions whose owner the compiler cannot prove, each of which holds one buffer rather than handing out a dangling one. Full detail, and the reason the analysis answers conservatively, is in the memory model.


Performance

neurc compile -O 3 runs LLVM's -O3 pipeline, the same level the C++ and Rust rows are built at (clang++ -O3, rustc -C opt-level=3), and all three target the generic x86-64 CPU with no -march=native. The default is -O 0, checked arithmetic with no optimization, so pass -O 3 before drawing any conclusion about speed.

Best of nine runs on one machine (Core i5-14600K, RTX 5070), lower is better. Reproduce with cd benchmarks && uv run python run.py --reps 9 --levels 3, which builds every implementation of each program and refuses to report timings if they disagree on output. Each one is written the way a programmer of that language would write it for speed. The Python column is the language itself; the NumPy column is what a Python programmer would write where the work vectorizes.

Benchmark What it stresses Neuro -O 3 clang++ -O3 rustc -O3 Python 3.14 NumPy
mandelbrot scalar f64 in a tight loop 166 ms 166 ms 166 ms 6668 ms 1268 ms
vector_sum Vec push, indexed sweep 24 ms 26 ms 24 ms 6008 ms 130 ms
call_overhead recursion, call and inline cost 49 ms 51 ms 48 ms 1409 ms
print_lines integer holes to standard output 13 ms 20 ms 13 ms 111 ms
format_floats f64 holes at a fixed precision 115 ms 110 ms 38 ms 222 ms
int_divide guarded / and %, opaque divisor 95 ms 88 ms 89 ms 1515 ms
matmul @ on [256, 256] f32 tensors 3.5 ms 3.9 ms 3.8 ms 770 ms 61 ms

Absolute times belong to the machine. Four rows are worth a word. print_lines beats C because an integer hole renders through a digit loop instead of snprintf, as Rust's does. format_floats loses to Rust, whose own float formatter is three times faster than the C library conversion Neuro and C++ both call. int_divide is the one place the compiler spends: / and % guard the operand pairs the hardware leaves undefined, and Rust guards the same pairs for 7% less. matmul runs each @ as a register-blocked, vectorized loop nest, level with the i-k-j loops the C++ and Rust versions vectorize, and with the same bits as a plain loop.

The GPU benchmarks run their Neuro side as @gpu kernels, against PyTorch on the same GPU and NumPy on the CPU. Each moves the same data between host and device:

Benchmark What it stresses Neuro @gpu PyTorch (CUDA) NumPy (CPU)
gpu_relax 200 fused element-wise steps on [2048, 2048] 262 ms 1080 ms 619 ms
gpu_reduce 1000 rounds of a whole and a row .sum() on [2048, 2048] 747 ms 1144 ms 1251 ms
gpu_matmul five [2048, 2048] f32 products 292 ms 1137 ms 217 ms
gpu_mlp twenty batches through a two-layer perceptron 275 ms 1133 ms 158 ms

These are whole-program times, and both GPU columns start with a fixed cost: about 220 ms for a Neuro program, most of it the CUDA driver starting up, and about 1100 ms for PyTorch's import and CUDA setup. Past that, timed warm inside one process, PyTorch takes 23, 56, 37 and 13 ms against Neuro's roughly 44, 530, 40 and 25. A long reduction folds in 4096 lanes, one GPU thread each, in the order the host folds in too, so the two agree bit for bit. Each GPU thread computes a block of up to 4 × 4 elements of a matrix product, every element summed in order, so the device keeps the host's bits; there is no shared-memory staging and no tensor-core path.


Quick Roadmap

Each numbered phase is a MAJOR-version milestone: completing Phase N ships v(N+1).0.0.

Phase Goal Status
1 Core Language: types, control flow, LLVM backend, ownership and borrow checking, generics, traits, closures, enums and pattern matching, error handling, modules Complete
2 Tensors and MLIR: first-class tensor types lowered through MLIR Linalg, the pool allocator, and the value model they need Complete
3 Automatic differentiation: a reverse-mode source-to-source transform over Neuro's own typed HIR, @grad(wrt: ...), .backward() / .zero_grad(), higher-order derivatives, elementwise math, .detach() / @no_grad Complete
4 GPU acceleration: MLIR GPU dialects for NVIDIA and AMD, @gpu with an opt-in CPU fallback, @kernel and the KernelOut<T> aliasing model, device tensors on any GPU Complete
5 Language and backend completion: MLIR as the only tensor backend, generic methods, closure capture modes, references in structs, tensors generic over their element type, named dynamic extents, a standard library written in Neuro, @test In progress
6 Neural network library: typed gradients from grad(...), @model, optimizers, layers and attention, .safetensors weights, a DataLoader, training benchmarks against PyTorch Planned
7 Parallelism and scale: scoped parallel { } tasks, channels, multithreaded host kernels, multi-GPU data parallel Planned
8 Interop and deployment: Python extension modules over DLPack, extern func foreign functions, defer Planned
9 Developer experience: debug info, incremental compilation, Language Server Protocol, error recovery, formatter Planned
10 Distribution and optimization: the neurpm package manager, prebuilt binaries and installers, a self-updater, optimization passes Planned

Architecture

Neuro follows Vertical Slice Architecture: code is organized by language feature, not by technical layer.

compiler/
├── infrastructure/          # Shared, zero-business-logic crates
│   ├── ast-types/           #   AST node definitions
│   ├── shared-types/        #   Primitives shared across slices
│   └── neuro-hir/           #   Typed High-Level IR (frontend <-> backend contract)
├── lexical-analysis/        # Tokenizer (logos, Unicode XID)
├── syntax-parsing/          # Pratt + statement parser -> AST
├── module-resolution/       # Multi-file loading, imports, visibility
├── argument-binding/        # Named arguments -> positional calls
├── semantic-analysis/       # Type checker, scope and borrow analysis
├── hir-lowering/            # Type-checked AST -> typed HIR
├── llvm-backend/            # HIR -> object code (inkwell 0.10 / LLVM 23)
├── mlir-backend/            # HIR -> MLIR linalg and GPU kernels
└── neurc/                   # CLI compiler driver

Today a .nr file travels: tokens → AST → merged program → type-checked AST → typed HIR → LLVM object code → system linker. Automatic differentiation runs inside HIR lowering. The mlir-backend slice lowers the same typed HIR to MLIR linalg; neurc links the tensor bodies it computes into the program, and on Linux lowers @gpu and @kernel bodies on through the MLIR GPU dialects to NVIDIA or AMD kernels. Stage by stage: docs/compiler/compilation.md.


Documentation

Getting started Installation, first program, workflow
Language reference One page per feature area, from types to modules
CLI guide Commands, flags, environment variables, interop
Compiler internals Pipeline and one page per slice
Editor support Syntax highlighting for .nr files
DESIGN.md Why the language is shaped this way, and its non-goals
CHANGELOG.md What each release changed

Everything is published at neuro-lang.netlify.app.


Contributing

CONTRIBUTING.md has the architecture rules, coding standards, quality gates and pull request process. Open defects are in docs/BUGS.md, and fixing one is the best way to start. Work is most useful in the phase the Quick Roadmap marks in progress.

See also SECURITY.md and the Code of Conduct.


License

Neuro Shared Source License v2.1. The license covers the compiler, not what you build with it.

You may freely use, study and modify the compiler for any personal or internal purpose, write Neuro programs and distribute or sell the compiled output under any terms you choose, build tools and editor integrations that call into the compiler, and contribute code back. Only redistributing the compiler itself, or a fork of it, as part of a commercial product requires a commercial license.

Neuro is pre-stabilization, and the license guards three risks specific to that: commercial re-packaging before the spec is stable, AI-assisted reproduction of the compiler for a competing product, and misleading forks that fragment an early ecosystem. A permissive license becomes possible once the language stabilizes.

Acknowledgments

Inspired by Rust (ownership, type system), Python (AI ecosystem simplicity), Swift (ergonomics) and Mojo (AI-first design). Built with inkwell, logos and LLVM.