中文 | English
FakeLua is an embeddable Lua-subset runtime for high-performance hosts: it compiles scripts to a bytecode VM and optionally to native code via GCC/TCC JIT, ships a large C++ native standard library (net/http/db/crypto/…), and uses an arena allocator with frame reset so there is no GC pause.
FakeLua was designed to address the throughput jitter and memory bloat caused by garbage collection in traditional scripting languages (standard Lua/LuaJIT) when used in high-performance game servers or similar real-time systems.
In a typical real-time high-performance server architecture:
- State & data reside in C++: Core data structures (player state, world maps, monster attributes, physics engine) are all stored in efficient, compact, type-safe C++ on the host side.
- Stateless/shallow-state Lua logic layer: Lua is used only for pure logic processing and business orchestration — reading C++ data and invoking C++ functions. Scripts should not retain large-scale data objects long-term.
To support this positioning, FakeLua does not implement a complex dynamic garbage collector (tri-color marking, generational GC, etc.). Instead, it uses an extremely efficient Arena memory pool (Bump Allocator):
-
Bump allocation ($O(1)$): When creating temporary variables (Table, String, Multi, etc.), FakeLua simply moves an offset pointer within a pre-allocated contiguous memory block. Allocation is nearly as fast as native stack allocation, without
mallocfragmentation or overhead. -
Instant cleanup ($O(1)$): At the end of each frame or request processing,
State::Reset()is called. It invokes destructors in reverse order, but instead of freeing individual blocks, the pool offset pointer is simply reset to zero. There is no complex object graph traversal, no system-levelfreecost or defragmentation — cleanup is instantaneous.
This design allows FakeLua to fully eliminate GC pause impact on frame rates while maintaining JIT native execution speed, keeping memory overhead at a completely predictable, extremely low level.
The same Call API can target any registered backend:
- JIT_GCC: Invokes system GCC (
-O3) to generate high-quality native code. Primary backend for production. - JIT_TCC: Embeds TinyCC for extremely fast compilation. Best for development, debugging, and tests (TCC is fetched automatically by CMake).
- JIT_INTERP: Compiles to FakeLua bytecode and runs on the built-in interpreter (
src/interp/). No external C compiler required; useful for portability, tooling, and mixed JIT↔interp closures.
int ret = 0;
Call(s, JIT_GCC, "add", ret, 10, 20); // Production: GCC (-O3)
Call(s, JIT_TCC, "add", ret, 10, 20); // Dev/test: TCC (fast compile)
Call(s, JIT_INTERP, "add", ret, 10, 20); // Bytecode VM (no host C compiler)The compiler automatically performs type inference and specialization for function math parameters:
- TypeInferencer runs iterative fixed-point inference on each top-level function (leave-one-out) to identify parameters that truly participate in arithmetic (math params).
-
CGen generates
$2^k$ specializations (int64_t/doublecombinations) plus a runtime entry dispatcher that routes to the appropriate specialization based on actual argument types. - Specialized bodies use native C types (
int64_t/double) for arithmetic and generate native Cboolfor comparisons, completely eliminating boxing overhead on hot paths.
-- Example Lua function: recursive Fibonacci
function fib(n)
if n <= 1 then return n end
return fib(n - 1) + fib(n - 2)
endAuto-generated specialized C code:
// 1. Numeric specialization: params/return promoted to native int64_t, no boxing
static int64_t fib_spec_0(int64_t n) {
if (n <= 1) {
return n;
}
return fib_spec_0(n - 1) + fib_spec_0(n - 2);
}
// 2. Generic entry dispatcher: fast type check, zero-overhead routing
static CVar fib_dispatcher(CVar n_var) {
if (LIKELY(n_var.type_ == VAR_INT)) {
return (CVar){.type_ = VAR_INT, .data_.i = fib_spec_0(n_var.data_.i)};
}
// ... dynamic dispatch to double specialization or generic CVar path
}With recursive Fibonacci (n=32) as an example, the GCC backend is 36.6x faster than Lua 5.4, and the TCC backend is 11.2x faster (see benchmark/README.md / 中文).
If a Table constructor can statically infer all its keys at compile time (string literals, explicit/implicit integer indices, booleans, floats), the compiler specializes it as a C struct:
- Struct layout generation: The compiler dynamically generates a C struct layout at compile time, with each specialized key mapped to a fixed-offset member.
- Initialization & deduplication: Constructor initialization fills the JIT specialized struct in a single pass (following Lua's left-to-right order) and checks for duplicate keys at compile time.
- Ultra-fast pointer-offset access: For specialized key reads/writes, pointer offset macros (
FL_SPEC/FL_SET_SPEC) are used directly, completely avoiding hash lookups and key comparisons. - Dynamic fallback: If the key used for read/write is a dynamic variable, it falls back to runtime dynamic dispatch via registered
spec_get/spec_setfunction pointers.
-- Example Lua code: defining and accessing Table fields
local point = { x = 10, y = 20 }
point.x = point.x + 5Auto-generated specialized C struct and pointer-offset access:
// 1. Compile-time key layout inference, auto-generate C struct definition
typedef struct Table_Spec_1 {
CVar x;
CVar y;
} Table_Spec_1;
// 2. On Table initialization, bind specialized struct layout and spec accessors
SET_TABLE_SPEC(point, Table_Spec_1, spec_get_fn, spec_set_fn, 2);
FL_SET_SPEC(Table_Spec_1, point, x, 0, (CVar){.type_ = VAR_INT, .data_.i = 10});
FL_SET_SPEC(Table_Spec_1, point, y, 1, (CVar){.type_ = VAR_INT, .data_.i = 20});
// 3. Field access converted to ultra-fast pointer member offsets (no hash table lookup)
FL_SPEC(Table_Spec_1, point, x) = NativeAdd(FL_SPEC(Table_Spec_1, point, x), (CVar){.type_ = VAR_INT, .data_.i = 5});- Closures & upvalue capture: Static AST analysis automatically derives scope and cross-function capture relationships. Captured variables are heap-boxed (
CVar *), shared across closures in the same scope. - Multi-return & varargs: Functions can
return a, b; C++ side receives viastd::tie(a, b, c). Vararg functions with...fully supported. - Anonymous & higher-order functions:
function(args) body endas values, arbitrary callee calls liketbl[key]()or(fn)(). - Colon method syntax:
obj:method(args)sugar with implicitselfparameter. - Generic
for initerators: Stateless iterators, closure generators, andpairs/ipairswith native C struct-optimized loops. - Per-iteration loop variable capture: Loop variables re-boxed each iteration for independent closure binding.
- Package modules:
package "Name"for namespace isolation, zero-requirecross-module calls. - Complex global initialization: Arbitrary expressions as file-level variable initializers, executed in generated
__fakelua_init(). - NativeObject & C++ interop: Host-side object mapping with group arena batch release, C++ member method binding via
RegisterMethod, colon-syntax calls from Lua. - ECMAScript regex:
string.find/match/gmatch/gsubvia Boost.Regex (supports lookahead, alternation, non-greedy quantifiers — more powerful than Lua patterns). - String algorithms:
string.trim/trim_left/trim_right/split/join/replace/starts_with/ends_with/contains/iequals/icontains/istarts_with/iends_withvia Boost.Algorithm.
- Coroutines: No
coroutine.create/resume/yieldsupport. - Metatables: No
__index,__newindex, metamethods, or operator overloading. require/module: No standard module system (replaced bypackage "Name"mechanism).rawequal/rawget/rawset/rawlen: Meaningless without metatables.- Debug library: No
debug.*standard library. - Implicit type coercion: No string→number conversion in arithmetic (
"10" + 1errors).
FakeLua provides 30+ independent C++ native modules under src/native/ (registered automatically on each State), covering math, string, table, IO, networking, timers, events, random, containers, compression, cryptography, serialization, databases, protobuf, config formats, logging, and subprocesses.
Full API reference: src/native/README.md / 中文
| Category | Modules |
|---|---|
| Core Lua | basic, math, table, string, os, utf8, io, random |
| Runtime / I/O | runtime (runtime.tick()), net (TCP/UDP), http (HTTP/1.1), url, timer, event |
| Data | json, csv, serialize, protobuf, container (Boost.Container deque/vector/list/map/set) |
| Config | yaml, toml, xml, ini |
| Database | mysql (async + pool), redis (async), sqlite (synchronous) |
| Crypto / compress | compress (LZ4/zlib/gzip/Zstd), crypto (OpenSSL digests/ciphers, UUID, CRC-32, xxHash) |
| Process | process (process.run; does not replace os.execute) |
| Logging | log (levels, tagged output, file rotation) |
| Object | object (NativeObject Lua-side API) |
Regex note: string.find/match/gmatch/gsub use ECMAScript regex (boost::regex::ECMAScript), not Lua patterns. See Regex Guide below for migration tips.
| Purpose | Lua Pattern | FakeLua (ECMAScript Regex) |
|---|---|---|
| Digits | %d |
\\d |
| Letters | %a |
[A-Za-z] |
| Alphanumeric | %w |
[A-Za-z0-9] (note \\w additionally includes _) |
| Whitespace | %s |
\\s |
| Escape literal |
%., %%
|
\\.、%
|
| Lazy repeat |
- (e.g. .-) |
? (e.g. .*?) |
| Backreference in replacement |
%1, %0
|
$1, $&
|
Since
\dis not a valid escape in Lua string literals, backslashes in regex patterns must be written as"\\d+". FakeLua does not support[[...]]long strings as a workaround.For scripts that need to be compatible with both standard Lua and FakeLua, use syntax that has the same semantics in both engines, e.g.
[0-9]+instead of%d+.
Key differences:
-
gsubreplacement strings use JS-style notation:$1…$9(capture groups),$&(entire match),$`(text before match),$'(text after match),$$(literal$). Lua's%1/%0are treated as literal characters here. -
Invalid patterns don't throw:
boost::regex_erroris caught and returnsnil, so the script doesn't interrupt. -
string.find'splainparameter has the same semantics as Lua: passingtruedegrades to pure substring search, completely bypassing the regex engine — also the fastest path. -
Performance: The regex path is significantly slower than Lua's native pattern engine; prefer
plainsearch orstring.sub/string.bytebasic operations on hot paths.
- C++23 compiler (GCC 11+ / Clang 16+ / MSVC 2022+)
- CMake 3.5+
- make or ninja
cmake -S . -B build
cmake --build build --parallelOn macOS, first
brew install lua cmakeand add-DCMAKE_PREFIX_PATH="$(brew --prefix)"to the cmake command.
Build only core library and CLI tools (no tests/benchmarks):
cmake --build build --target fakelua flua --parallelcmake -S . -B build -G Ninja
cmake --build build --parallel
ctest --test-dir build -Vcmake -S . -B build -DCMAKE_EXPORT_COMPILE_COMMANDS=ON
cmake --build build --parallel
ctest --test-dir build -V
./build/bin/bench_markUnit tests and benchmarks require the Lua development package (header
lua.hand library files).
- Linux:
sudo apt-get install liblua5.4-devorliblua5.3-dev- macOS:
brew install lua- Windows MSYS2:
pacman -S mingw-w64-x86_64-lua
./build/bin/flua <script.lua> --entry=<func> --jit_type=<0|1|2> --repeat=<N>--entry: Entry function name (defaultmain)--jit_type:0=TCC,1=GCC,2=INTERP (bytecode VM)--repeat: Repeat call count (for performance measurement)--debug: Enable debug mode (defaultfalse; whentrue, outputs generated C source / richer diagnostics)
Comparing Lua 5.4, FakeLua TCC, FakeLua GCC across 11 algorithms (Release -O3 mode):
| Algorithm (typical params) | Lua 5.4 | FakeLua TCC | FakeLua GCC |
|---|---|---|---|
| Fibonacci n=32 | 297.9 ms | 26.7 ms (11.2x↑) | 6.8 ms (36.6x↑) |
| Sum n=5000000 | 33.9 ms | 18.4 ms (1.8x↑) | 1.1 ms (30.4x↑) |
| Popcount n=100000 | 18.2 ms | 3.1 ms (5.9x↑) | 488.0 μs (37.3x↑) |
| BubbleSort n=200 | 1.5 ms | 3.3 ms (0.45x) | 738.8 μs (1.9x↑) |
| Sieve n=5000 | 353.4 μs | 1.0 ms (0.34x) | 219.3 μs (1.8x↑) |
| FloatPoly n=1000000 | — | — | 34.9x↑ (浮点特化,GCC 2x 快于 C++) |
TCC is generally faster than Lua for pure computation; in Table-operation-heavy scenarios, Table struct specialization gives both GCC and TCC a significant boost. Full data available in benchmark/README.md / 中文.
FakeluaStateGuard guard;
State* s = guard.GetState();
CompileFile(s, "script.lua", CompileConfig{.debug_mode = false});
int sum = 0;
Call(s, JIT_GCC, "add", sum, 10, 20); // embed-call a Lua function// Manual management (not recommended — easy to leak)
State* s = FakeluaNewState(StateConfig{});
// ... use s ...
FakeluaDeleteState(s);
// Or RAII style (recommended)
FakeluaStateGuard guard(StateConfig{});
State* s = guard.GetState();
// ... use s ...
// automatically freed| Function | Description |
|---|---|
FakeluaNewState() |
Create FakeLua state |
FakeluaDeleteState() |
Free FakeLua state |
CompileFile() |
Compile a Lua file |
CompileString() |
Compile a Lua code string |
Call() |
Invoke a compiled function |
GetLastRecordedCCode() |
Get the most recently compiled C code |
SetVarInterfaceNewFunc() |
Set custom VarInterface factory |
SetDebugLogLevel(s, level) |
Set this State's debug log level (0=Trace … 6=Off; Lua: log.set_level) |
// Native → FakeLua
CVar v_int = inter::NativeToFakelua(s, 42);
CVar v_str = inter::NativeToFakelua(s, std::string("hello"));
// FakeLua → Native
int native_int = inter::FakeluaToNative<int>(v_int);
std::string native_str = inter::FakeluaToNative<std::string>(v_str);class CustomVar : public VarInterface { /* ... */ };
SetVarInterfaceNewFunc(s, []() { return new CustomVar(); });
// Table-type arguments in Call automatically construct CustomVar instancesLua source
↓
[Lexing] → tokens (flexer)
↓
[Parsing] → AST (bison + syntax_tree)
↓
[File-level stmt check] → reject non-declaration statements (semantic_analysis)
↓
[Preprocessing] → normalized AST (preprocessor)
↓
[Semantic analysis] → analysis result (semantic_analysis)
↓
[Type inference] → type hints (type_inferencer)
↓
┌─────────────────────────────┬──────────────────────────────┐
↓ ↓ ↓
[C code generation] [Bytecode codegen] (shared AST)
(c_gen) (interp/codegen)
↓ ↓
[JIT TCC / GCC] [Interpreter VM]
native code (interp/interpreter)
└───────────── Call(s, JIT_*, …) ─────────────┘
| Module | Responsibility |
|---|---|
lexer/parser |
Lua lexing and parsing |
syntax_tree |
AST representation and traversal |
preprocessor |
Lua syntax normalization (e.g., functiondef hoisting) |
semantic_analysis |
Semantic and control flow analysis |
type_inferencer |
Static type inference and specialization decisions |
c_gen |
C code generation and type-driven optimization |
interp/* |
Bytecode codegen, opcodes, and interpreter VM |
compile_common |
Common type inference and codegen utilities |
jit/* |
TCC/GCC backends plus Vm function registry |
native/* |
Built-in standard libraries (net, http, db, crypto, …) |
state |
FakeLua runtime state management |
var |
Dynamic value CVar and conversion utilities |
A: Certain dynamic features of full Lua (e.g., metatables) are difficult to compile efficiently. The subset focuses on statically analyzable common patterns, achieving near-C performance through type inference and JIT compilation.
A: GCC is the primary production backend (-O3). TCC compiles extremely fast for development and CI. INTERP runs bytecode without a host C compiler — good for constrained environments, tooling, and validating script semantics; hot paths can still call JIT closures when mixed.
A: Yes — use JIT_INTERP (no GCC/TCC required at runtime) or the small TCC backend. Native modules that need OpenSSL/MySQL/etc. are optional at the dependency level for your build.
A: Enable CompileConfig::debug_mode to inspect logs and C code; use GetLastRecordedCCode() to export C code for analysis.
A: Each State is currently thread-local; in multithreaded environments, create an independent State per thread.