A dataframe library for Mojo with the pandas API.
Status: M0 landed and M1 is under way. There is a
DataFramenow, it groups by and joins, and there is still nothing to install. What exists is the layer underneath one, plus the first thing built on it: validity bitmaps, aligned buffers with a size class pool, typed and type erased columns, the StringView layout and the string column built on it, the logical type lattice with its promotion rules, thecomptimedtype dispatch bridge, the compute kernels that run over a column, a stable radix sort with multi-key and null placement, the hash table andfactorizethat group by and join are both built on, thirteen grouped reductions fromsumthroughmedianand distinct count, all seven join kinds from inner through anti and cross, concat and the null handling functions fromis_nullthrough forward fill, an eagerSeriesandDataFrameover all of it that can select, cast, filter, take, slice, sort, group by, join, stack and drop or fill its nulls, and a display layer that renders either of them as a table. It also has the two byte level halves of a CSV reader, a scanner that finds where every field is and parsers that turn a field's bytes into an integer, a float or a boolean, and a variable width string column for a text field to land in, though there is noread_csvon top of them yet and it still cannot read a file. It is tested, fuzzed against a reference model, and checked against numpy and pyarrow in the same process. The specification is twelve documents indocs/specs/, written against Mojo 1.0, pandas 3.0, Polars 1.43 and Arrow 25.0 as of August 2026, and the milestone issues track the rest of the way to a frame you can actually use. If you are looking for a working Mojo dataframe today, you want MojoFrame, which is a research prototype, or Polars, which is not in Mojo but is excellent.
import firepanda as pd
df = pd.read_parquet("trades.parquet")
big = df[df["qty"] > 1000].groupby("symbol")["notional"].sum()That is pandas. It is also, underneath, a lazy columnar query engine that pushed the projection and the predicate down into the Parquet reader, never decoded the columns you did not ask for, and ran the aggregation across every core in your machine.
And this is the same library, from Mojo, with the user's own function compiled into the pipeline rather than called through an interpreter:
from firepanda import read_parquet, col
fn score(price: Float64, qty: Float64) -> Float64:
return price * qty * 0.997
var df = read_parquet("trades.parquet")
.filter(col("qty") > 1000)
.with_column(map2[score](col("price"), col("qty")).alias("net"))
.group_by("symbol").agg(col("net").sum())
.collect()The block above is the target. This is the part that works now, and the output is copied from a real run rather than written by hand.
var df = DataFrame.from_series(columns^)
print(df)
var big = df.filter(greater(qty, threshold)).sort_by("price", descending=True)
print(big)
var widened = big.cast("qty", DType.float64)
print(Series("notional", multiply(
widened.column("qty").as_typed[DType.float64](),
widened.column("price").as_typed[DType.float64](),
))) qty price
0 400 101.25
1 1200 99.5
2 2500 100.0
3 80 98.75
4 1750 <NA>
5 3000 100.125
[6 rows x 2 columns]
qty price
0 3000 100.125
1 2500 100.0
2 1200 99.5
3 1750 <NA>
[4 rows x 2 columns]
0 300375.0
1 250000.0
2 119400.0
3 <NA>
Name: notional, dtype: float64
Group by works too, on one key or several, with a reduction per output column.
var specs = List[AggSpec]()
specs.append(AggSpec("qty", AggKind.SUM, "qty"))
specs.append(AggSpec("price", AggKind.MEAN, "avg_price"))
specs.append(AggSpec("qty", AggKind.COUNT, "trades"))
print(df.group_by(["symbol"], specs)) symbol qty price
0 1 400 101.25
1 2 1200 99.5
2 1 2500 100.0
3 2 80 98.75
4 1 1750 <NA>
5 3 3000 100.125
[6 rows x 3 columns]
symbol qty avg_price trades
0 1 4650 100.625 3
1 2 1280 99.125 2
2 3 3000 100.125 1
[3 rows x 4 columns]
Symbol 1 averages 100.625 over two prices rather than three, because row 4 is null and a mean divides by what is there. trades counts three, because it counts qty and qty has no nulls.
No Parquet, no expression API, and no strings in a frame yet: the string column exists but nothing that takes a frame can hold one. Columns are built by hand, the mask comes from a kernel rather than from df["qty"] > 1000, group by takes a list of specs rather than a chained .groupby("symbol").sum(), and a null and a NaN print differently because in an Arrow layout they are different things.
There is a SQL front end being built in firepanda/sql/, tracked in #304. It reads DuckDB's own PEG grammar, so the syntax it accepts is DuckDB's syntax rather than an approximation of it, and every statement in DuckDB's test corpus goes through both parsers on every change. No query runs yet. What exists is the parser, the AST and the printer.
The part worth reading now is where the line is. The grammar accepts everything DuckDB accepts and firepanda will execute rather less than that, so a query on the far side of the line gets a refusal that names the feature, points at the word in the query, says what firepanda is instead, and links to an issue. It is not a syntax error and it is not a silent wrong answer, because both of those leave you with nothing to go on.
Not Implemented Error: firepanda does not support a slice or a subscript.
LINE 1: SELECT a[1] FROM t
^
firepanda reads a list element with a function rather than with brackets.
See https://github.com/tamnd/firepanda/issues/13
firepanda.sql_support() returns that whole set, so a program can ask what is missing rather than discover it a query at a time. A message with {} in it fills the {} with the word out of your query. Two issues are linked: #13 is a feature with no nearer home, #304 is one that is coming and has not landed.
| Name | firepanda does not support | Instead | Issue |
|---|---|---|---|
operator |
{} as an operator | firepanda holds an operator as the words it was written with and runs the ones it has a kernel for, and this is not one of them. | #13 |
custom-operator |
OPERATOR(...) as a prefix operator | An operator named this way is resolved against the catalog, and firepanda has no catalog of operators to resolve it against. | #13 |
like-escape |
ESCAPE on an operator that has no escaping form | A LIKE and an ILIKE take one. SIMILAR TO does not, which DuckDB says too. | #13 |
field-access |
a field access | It reads and prints back. A field belongs to a struct and a firepanda column holds one scalar, so there is nothing here with a field in it to reach into. | #304 |
subscript |
a slice or a subscript | firepanda reads a subscript or a slice of text whose bounds are whole numbers, as the substring it means. A bound that is an expression, a step, and a slice from the back to the front have no substring that answers them. | #304 |
postfix-operator |
a postfix operator | The two firepanda reads after an operand are a cast and a dotted name, and this is neither. | #13 |
call-modifier |
{} on a call | Both read and print back where they were written. WITHIN GROUP gives the fold an order to see the rows in and EXPORT_STATE asks for its state rather than its answer, and firepanda has no fold that can be told either. | #304 |
call-argument |
{} inside a call | Both read and print back where they were written. An ordered aggregate gives the fold an order to see the rows in and a null treatment gives it a rule for what to do with a null, and firepanda has no fold that can be told either. | #304 |
array-subquery |
ARRAY over a subquery | It reads and prints, and lowering stops on it, since collecting a whole column into one list value is a way of running a subquery that firepanda has no node for. | #13 |
dotted-name |
a dotted name where {} goes | A name with dots in it qualifies one thing by another, and this position takes the name of one thing. Write the last part on its own. | #13 |
quoted-name |
anything but a name here | Only a name fits in this position. Dots in it are fine, since that is how two collations are composed. | #13 |
not-subquery |
NOT in front of a subquery | Write NOT EXISTS or NOT IN, which say which of the two this means. | #13 |
select-clause |
{} in a SELECT | The query node has a slot for each clause firepanda runs, and none for this one yet. | #304 |
select-sample |
a sample on a SELECT | The sample reads and prints, and lowering stops on it, since taking a share of the rows is a row source of its own and firepanda has no node for one yet. | #304 |
table-sample |
a sample on a table | The sample reads and prints, and lowering stops on it, since taking a share of the rows is a row source of its own and firepanda has no node for one yet. | #304 |
table-modifier |
{} on a table | A join, a PIVOT and an UNPIVOT are the three things that go here and all three are read, so this is a fourth one the grammar grew and the transformer has no case for. | #304 |
table-at |
AT on a table | It reads and prints, and lowering turns it down, because it asks for a table as of a version or a moment and firepanda has no storage that keeps either one. | #13 |
join-form |
this kind of join | firepanda runs the joins that name a condition or take none, and POSITIONAL and ASOF joins. NEAREST is not among them, and nor are the mark, single, right_semi and right_anti types JOIN BY can name. | #13 |
with-ordinality |
WITH ORDINALITY | It reads and prints, and lowering stops on it, since it adds a column to what a table function produces and the two are built together. | #13 |
with-using-key |
USING KEY on a WITH | It says to keep one row per key and replace that row as the query runs, which is a way of running a WITH that firepanda has no node for. | #13 |
statement-later |
the {} statement yet | It maps onto something a dataframe already does and it is coming. firepanda runs SELECT today. | #304 |
statement-never |
the {} statement | It asks for a catalog, a transaction or an extension, and firepanda is a dataframe library rather than a database. Read the data with SELECT and do the rest in Mojo. | #13 |
row-value |
a row value | Several expressions in one pair of parentheses make a single value with fields in it, and a firepanda column holds one scalar. Select the parts as separate columns. | #304 |
interval |
an INTERVAL literal | A duration is its own type with its own arithmetic, and firepanda has no column type for one yet. It arrives with the date and time work. | #304 |
special-call |
{} yet | TRY and UNPACK are written the way a call is written and read as calls, and neither one is a function. TRY answers NULL where the expression would have raised and UNPACK spreads a list across the arguments of the call around it, so both are shapes the plan would have to grow. | #304 |
lambda |
a lambda | A function written inside the query has to be compiled along with the query, and firepanda runs the functions it already has. Pass a named one. | #13 |
list-comprehension |
a list comprehension | It runs an expression once for every element, which is a lambda in different brackets, and firepanda runs the functions it already has. | #13 |
named-argument |
an argument passed by name | firepanda matches arguments by position, so f(a := 1) has nowhere to put the name. Pass it in order. | #304 |
columns |
COLUMNS | It stands for however many columns it matches, so the shape of the result is not known until the table is, and firepanda works out the shape first. Name the columns. | #304 |
map-literal |
a MAP literal | A map holds keys and values in one value and a firepanda column holds one scalar. It arrives with the nested types. | #304 |
positional |
a column written as #1 here | firepanda reads #1 as the first column of the FROM in the select list, WHERE, GROUP BY, HAVING and QUALIFY, and as the first output column in a bare ORDER BY. In an ON clause, over a USING or NATURAL join and in a correlated subquery it does not yet, so write the column's name there. | #304 |
default-value |
DEFAULT where a value goes | It stands for whatever a table declares as the default for a column, and that lives in a catalog. firepanda is a dataframe library and has no catalog to ask. | #13 |
unpivot-nulls |
INCLUDE NULLS on an UNPIVOT | The statement spelling of an UNPIVOT has no way to write it, and that is the spelling the node records, so there is nowhere to keep it. EXCLUDE NULLS is the default and is read. | #304 |
unpivot-groups |
more than one FOR group on an UNPIVOT | One UNPIVOT node holds one name column and one set of value columns, so a second group has nowhere to go. | #304 |
quantified-value |
ANY or ALL over a value | A list written out on the right runs. Anything else there is a list DuckDB unnests, and no firepanda column holds a list. Write the list out, or put a SELECT on the right. | #304 |
no-case |
grammar rule {} | The grammar accepts more than firepanda runs, and this is a rule the transformer has no case for. Please file it. | #304 |
aggregate-filter |
FILTER on {} | A filter is read as a CASE around the argument, which answers the same thing for a fold that passes over a null. first, last, arbitrary, list and array_agg keep a null, so it would not. A fold of two arguments takes the CASE around its first when a null there drops the row, and a name that is not a fold has nothing for the CASE to go inside. | #13 |
list-value |
a list written out | A literal in the plan is one scalar, and a list is a value with a length. LIST is a type firepanda has and a column can hold one, so what is missing is the constant rather than the type. | #304 |
modifier-schema |
a name of three parts or more where {} goes | A star modifier names a column of one of the things the FROM brought, so firepanda reads two parts as that thing and that column. Three parts starts at a schema or at a struct, and telling those apart needs the catalog. | #13 |
How much of DuckDB's own test corpus runs is counted per directory of test/sql, with the denominator, by python tools/conformance.py --readme, which ran every file below and wrote this table. A file passes when every statement in it answers what the file expects. "Not run" is a file that needs an extension, a database file or the corpus's data/ directory. A failing file is put in the bucket of its first failure: a refusal by name, a function firepanda lacks, a divergence (an error where DuckDB answers, or the wrong error), a wrong answer, or a crash. Document 11 in the specification has the rules.
| Directory | Passed | Rate | Not run | Unsupported | Function | Divergence | Wrong | Crash |
|---|---|---|---|---|---|---|---|---|
| test/sql/aggregate | 2/157 | 1.2% | 32 | 89 | 30 | 35 | 1 | 0 |
| test/sql/cte | 8/82 | 9.7% | 6 | 62 | 2 | 9 | 0 | 1 |
| test/sql/filter | 0/12 | 0.0% | 5 | 7 | 2 | 2 | 1 | 0 |
| test/sql/join | 7/137 | 5.1% | 12 | 102 | 1 | 27 | 0 | 0 |
| test/sql/limit | 0/10 | 0.0% | 1 | 9 | 1 | 0 | 0 | 0 |
| test/sql/order | 3/31 | 9.6% | 4 | 21 | 0 | 7 | 0 | 0 |
| test/sql/pivot | 0/31 | 0.0% | 4 | 22 | 1 | 7 | 1 | 0 |
| test/sql/projection | 1/17 | 5.8% | 0 | 9 | 0 | 7 | 0 | 0 |
| test/sql/select | 3/9 | 33.3% | 0 | 5 | 0 | 1 | 0 | 0 |
| test/sql/setops | 5/24 | 20.8% | 1 | 14 | 0 | 5 | 0 | 0 |
| test/sql/subquery | 2/92 | 2.1% | 3 | 59 | 3 | 25 | 3 | 0 |
| test/sql/topn | 1/14 | 7.1% | 8 | 12 | 0 | 1 | 0 | 0 |
| test/sql/window | 3/77 | 3.8% | 8 | 44 | 8 | 21 | 1 | 0 |
Every fast dataframe library today is a fast engine in one language with a Python veneer on top. pandas is C and Cython. Polars is Rust. DuckDB is C++. The veneer is where user code lives, and it is why df.apply(lambda ...) falls off a cliff: the moment you write a function the library did not anticipate, you leave the fast language and enter the slow one.
Mojo removes the boundary. The library, the kernels and the user's own hot loop are all the same language, and they all compile. That is the one thing firepanda can offer that a mature library in another language structurally cannot, and it is the reason to build this rather than use Polars.
The honest counterweight is in docs/specs/00-README.md: the argument only pays off for Mojo users. For Python users arriving through pip install firepanda, a Python callable is still a Python callable, and the benchmark tables say so in their own column.
| Memory | Apache Arrow layout throughout. Validity bitmaps, StringView, dictionary encoding. Interchange with pandas, Polars, DuckDB and pyarrow is zero copy through the PyCapsule protocol. |
| Kernels | One generic function per kernel, comptime-parameterized over DType. vectorize writes the remainder loop. No build tags, no runtime feature detection, no hand-written tails. |
| Dispatch | Columns are type-erased with a runtime DType tag; a comptime for over the dtype list generates the dispatch chain. Mojo has no sum types, so this replaces the enum a Rust design would use. |
| Execution | Morsel-driven, work-stealing, radix-partitioned hash aggregation. Workers own their partitions exclusively, because there is no race detector in this language. |
| Planning | Lazy underneath, always. The eager pandas surface builds plans too, so the naive idiom gets projection and predicate pushdown for free without anyone learning an expression API. |
| Front doors | Mojo and Python, neither second class. pip install firepanda must work on a machine with no Mojo toolchain — otherwise the audience is people who already have Mojo, and that audience is too small. |
Read docs/specs/00-README.md first; it is the index and it says what was already decided and why.
| 00 | Index, settled decisions, prerequisites, honesty about scope |
| 01 | State of Mojo 1.0, pandas 3.0, Polars, Arrow and the Mojo ecosystem, with sources |
| 02 | Layers, memory, types, plan, optimizer, execution |
| 03 | Compile-time monomorphization and the generated dispatch table |
| 04 | The two front doors, and every deliberate divergence from pandas |
| 05 | Kernel shape, the hash table, strings, the GPU path |
| 06 | The full pandas 3.0 conformance checklist, milestone-tagged |
| 07 | PythonModuleBuilder, Arrow PyCapsule, wheels, and the ABI problem |
| 08 | M0 through M11, exit criteria, and four points to stop and reassess |
| 09 | Testing, and living without a race detector |
| 10 | What gets measured and against whom |
| 11 | The tree, and why Mojo 1.0's import rules decide it |
There are no time estimates in these documents, deliberately. Milestones are ordered by dependency and by risk; a week count invites a reader to add them up and treat the total as a delivery date.
Distribution. Mojo 1.0 guarantees source compatibility within 1.x and explicitly does not guarantee ABI stability, and a wheel is a binary artifact. The plan is to vendor the runtime and pin hard — which first requires confirming the runtime may be redistributed at all. That question is the first task of M3, before any binding code gets written, because a negative answer changes the entire distribution strategy.
Code size. Monomorphizing every kernel over every dtype produces on the order of a few thousand instantiated function bodies. Compile time and stripped binary size are graphed in CI from the first commit, with named thresholds and three graduated responses, because finding this out late is how the strategy fails.
- tamnd/firepanda-bench — the performance comparison against pandas, Polars, DuckDB, cuDF and MojoFrame. Losses get published next to the wins; for a project whose pitch is performance, the credibility of the numbers is the asset.
- tamnd/kuma — the sibling specification for the same problem in Go. Several documents here are written against it as a contrast.
See CONTRIBUTING.md. At specification stage the most useful contribution is disagreement: if something in docs/specs/ is wrong about Mojo 1.0, about pandas 3.0, or about what an engine of this shape costs to build, an issue saying so is worth more than any amount of code written against a bad premise.
Claims in the specification that could not be confirmed against Modular's own documentation are marked [verify]. Confirming or refuting one of those is a genuinely valuable pull request.
Apache-2.0. See LICENSE.