Skip to content

Repository files navigation

cyclelife

Predicting how many cycles a lithium-ion cell will last, and its state of health along the way, from its first 100 cycles. The data are the public MIT-Stanford fast-charging dataset: 124 commercial A123 APR18650M1A cells (1.1 Ah) cycled to 80 % of nominal capacity under 72 fast-charging policies, with cycle lives from 148 to 2,237 cycles (K. A. Severson, P. M. Attia et al., "Data-driven prediction of battery cycle life before capacity degradation", Nature Energy 4, 383-391, 2019, doi:10.1038/s41560-019-0356-8). The cells are LFP cathode with a graphite anode, so results do not transfer directly to silicon or silicon-graphite anodes, whose degradation physics differ. This is a careful reimplementation with an honest validation design, not a new method.

Splits are by cell, never by cycle

This is the point of the repository. Every split here, for testing and for cross-validation, is a partition of cells. A cell's cycles are never on both sides of a split.

Splitting a table of cycles at random leaks, and the leak inflates scores. Consecutive cycles of one cell are near-duplicates: a model that trained on cycle 500 of a cell and is then tested on its cycle 501 is not being tested. Kapoor and Narayanan's taxonomy of leakage lists this as type L3.2, non-independence between train and test samples, and calls the case where "train and test samples come from the same people or units" "unfortunately common" [2]. Schaeffer et al., in a review of cycle-life prediction illustrated with this dataset, recommend grouped cross-validation so that "all of the data belonging to a group is assigned to the same subset" [3]; their groups are batches and operating conditions, and a cell is the smallest such group.

Measured here, on the state-of-health task: same cells, same models, and as many held-out rows either way. Split by cycle, every one of the 84 cells ends up on both sides, and the error looks about half what it is:

State of health, RMSE [percentage points] (MAPE)

Split by cell: primary test cells Split by cell: secondary test cells Split by cycle: held-out rows (leaks)
Elastic net 4.52 (3.26%) 2.60 (1.91%) 2.45 (1.99%)
Gradient boosting 2.48 (1.85%) 3.27 (2.60%) 1.15 (0.81%)

Splitting by cycle also tunes by cycle, as a notebook that splits this way would; the leaky elastic net then chose the smallest penalty on its grid, so if anything its leaky score is understated.

State-of-health error, split by cell and split by cycle

How the code makes the mistake hard to make:

  • CellSplit cannot be built if its two sides share a cell. split_rows sends each row to the side its cell is on and checks its own output. cell_folds holds out whole cells in cross-validation, so tuning cannot leak either. (src/cyclelife/split.py)
  • tests/test_split.py fails if a cell ever has rows on both sides: on 40 random cycle tables, on cross-validation folds, and on a deliberately cycle-level split, which must raise.
  • The comparison above is the only code that splits by cycle. It bypasses cyclelife.split on purpose, in the open (soh_experiment in src/cyclelife/experiments.py), and a test checks that it really does leak, so the comparison is not empty.

Results

Every number and figure in this section comes from the real data, produced by nox -s data train figures (2026-09-29, Windows, Python 3.13.9; package versions in results/metrics.json). Everything CI runs uses synthetic fixtures, and a test checks that the tables below are exactly the ones rendered from results/metrics.json.

Cycle life

Target log10(cycle life). Inputs: the 13 candidate features of the paper's "discharge" model, all from cycles 1 to 100: six statistics of ΔQ100-10(V), the change in the discharge capacity-voltage curve between cycles 10 and 100, and seven features of the capacity fade curve. Trained on the paper's 41 training cells; hyperparameters chosen by 4-fold cross-validation over the training cells, repeated 10 times. The primary test set is 43 cells from the same two batches, the secondary test set the 40 cells of a third batch, run later.

Cycle life, RMSE [cycles]

Train Primary test Primary test, without b2c1 Secondary test
Paper, Table 1, "discharge" model 76 91 86 173
Elastic net, 13 discharge features (this repo) 71 94 89 170
Gradient boosting, same features (this repo) 7 152 152 257

Cycle life, MAPE [%]

Train Primary test Primary test, without b2c1 Secondary test
Paper, Table 1, "discharge" model 9.8 13.0 10.1 8.6
Elastic net, 13 discharge features (this repo) 9.4 12.7 9.5 9.6
Gradient boosting, same features (this repo) 0.6 13.3 11.0 18.9

Predicted against observed cycle life

Residuals against observed cycle life

Against the paper

The elastic net here replicates the "discharge" row of Table 1 of Severson et al. [1]: the same 13 candidate features, the same target, the same cells in the same split. Its RMSE is within 5 cycles of the paper's on every set, and its percentage error within a point. It does not beat the paper, and is not meant to.

Why it is not identical, from what is known to what is not:

  1. Two recording glitches. b1c0 at cycle 12 and b1c18 at cycle 40 report 1.54 Ah and 2.88 Ah on 1.1 Ah cells. b1c0's cycle 12 has a 407-minute gap in its time record, where its neighbours have none longer than 0.1 minutes: consistent with one of the two automatic restarts of the cycler's computer that the dataset's notes for this batch report ("As such, there are some time gaps in the data"). b1c18's cycle 40 holds 279 separate discharge segments where every other cycle holds one. Used raw, the two values put the max-capacity feature of those two primary-test cells 226 and 886 training standard deviations from the mean, the model predicts a life of zero cycles for both, and the primary-test RMSE is 311 cycles. Here a value is a glitch when it is more than 0.05 Ah from the median of the five cycles centred on it.

    Across all 124 cells and 100,501 cycles, that rule flags 56 values. In cycles 2 to 100, the only ones the features read, it repairs these two and no other. It drops 13 of the 86,922 state-of-health target rows. The other 41 are the first batch's cycle-1 placeholders, recorded as 0 Ah, which nothing reads. The replication does not depend on the threshold: from 0.03 to 0.10 Ah the features change on the same two cycles, to the same values, and the errors are identical. The paper does not mention these values. Its discharge model keeps this feature and still scores 91 cycles, so it presumably worked from data without them, but that is an inference.

Elastic net, cycle-life RMSE [cycles] against the glitch threshold

Glitch threshold Repaired in the features (cycles 2 to 100) Dropped from the targets (cycle 101 to end of life) Flagged in any cycle Primary test Primary test, without b2c1 Secondary test
none (raw values) 0 0 0 311 313 170
0.03 Ah 2 16 59 94 89 170
0.05 Ah (used) 2 13 56 94 89 170
0.10 Ah 2 11 54 94 89 170
  1. Model selection. The paper's elastic net ran in MATLAB, with four-fold cross-validation and Monte Carlo sampling, and kept 6 of the 13 features. This one is scikit-learn's, with a grid over the penalty and four-fold cross-validation repeated 10 times over reshuffled training cells. It keeps 8 features, including all 6 of the paper's.
  2. The data may differ slightly from the paper's. The official MATLAB loader's cleaning and cycle-life rule, applied to the files on data.matr.io today, give a mean cycle life of 801.6 and a standard deviation of 379.7 over the 124 cells; the paper reports 806 and 377. The cycle lives computed here equal that loader's on every cell. Where the difference comes from is not known.

The gradient-boosted model fits the training cells almost exactly (7 cycles) and does worse than the linear model on both test sets: 41 cells are too few for it. On state of health, with tens of thousands of rows, it does better than the linear model on the primary test cells and worse on the third batch.

State of health

One row per cell and cycle, from cycle 101 to the end of life; the target is discharge capacity over the 1.1 Ah nominal. Both models see the 13 early-cycle features and the cycle number. The linear model also gets the cycle number times each feature, without which it could shift a cell's curve but not change its slope. That is 86,909 rows; 13 more, whose capacity is a recording glitch, are dropped rather than repaired, since a model scored on a glitch is scored on noise. Results are in the first table of this README.

Reproduce

pip install nox
nox -s data train figures

data downloads 8,269,341,808 bytes (8.27 GB) from data.matr.io, the dataset's canonical host, checks each file against the size the host reports and a recorded SHA-256, and parses them into a cache of about 100 MB (schema in docs/cache_schema.md). Set CYCLELIFE_DATA_DIR to keep the data outside the repository. Nothing downloaded is committed. train takes about 13 minutes on an 8-core laptop; figures a few seconds. The data are released under CC BY 4.0 by the dataset's authors.

nox              # lint, strict types, tests: what CI runs
nox -s tests     # the tests on every installed Python from 3.11 to 3.13

CI runs ruff, mypy --strict and the tests on Linux, Windows and macOS with Python 3.11, 3.12 and 3.13, on synthetic fixtures only. It never downloads the dataset.

What the tests caught

pytest     # 100 tests, including the docstring examples
nox        # lint, strict types and tests: what CI runs

Problems that really happened while building this, each with what found it:

  1. Two impossible capacities broke the replication. The first full run gave a primary-test RMSE of 311 cycles against the paper's 91, while the training set (71 against 76) and the secondary test set (170 against 173) matched. Found by the replication check against Table 1: comparing each set separately located the problem in one set, and ranking its cells by squared error found b1c0 and b1c18 predicted at zero cycles, from the two glitches described above. Fixed in its own commit, and guarded by test_a_capacity_spike_is_repaired_before_the_fade_features.
  2. Two cells without a cycle life. In the third batch, b3c23 and b3c32 store NaN as their cycle life. Both are dropped by the official cleaning, but the loader converted every cycle life to an integer before cleaning, and would have stopped on them. Found by listing every cell of the three files before the first full parse; the loader now keeps NaN until cleaning and refuses a kept cell without a life.
  3. A docstring example that was never true. The example for cell_folds showed which cells land in which fold, an output that depends on the random generator. Found by the doctest run (--doctest-modules), which executes every example; the example now shows properties that hold for any seed.

What this does not do

  • It does not transfer to other chemistries. All 124 cells are one LFP/graphite model, at one chamber temperature (30 °C), discharged the same way (4C). What varies is the charging policy, and a few test details between the three batches.
  • It does not re-derive the Q(V) curves. They are the dataset authors' Qdlin: each discharge curve fitted with a spline and sampled at 1,000 voltages from 3.5 V to 2.0 V.
  • It gives no uncertainty on a prediction.
  • It does not use temperature or internal resistance, the extra data streams of the paper's "full" model, whose analysis also drops four cells whose temperature probes lost contact.

References

  1. K. A. Severson, P. M. Attia, N. Jin, N. Perkins, B. Jiang, Z. Yang, M. H. Chen, M. Aykol, P. K. Herring, D. Fraggedakis, M. Z. Bazant, S. J. Harris, W. C. Chueh, R. D. Braatz. Data-driven prediction of battery cycle life before capacity degradation. Nature Energy 4, 383-391 (2019). doi:10.1038/s41560-019-0356-8. Data: data.matr.io, CC BY 4.0. Loading notebooks: rdbraatz/data-driven-prediction-of-battery-cycle-life-before-capacity-degradation.
  2. S. Kapoor, A. Narayanan. Leakage and the reproducibility crisis in machine-learning-based science. Patterns (2023). doi:10.1016/j.patter.2023.100804. Quoted from the preprint, arXiv:2207.07048.
  3. J. Schaeffer, G. Galuppini, J. Rhyu, P. A. Asinger, R. Droop, R. Findeisen, R. D. Braatz. Cycle life prediction for lithium-ion batteries: machine learning and more. American Control Conference (2024). arXiv:2404.04049.

Provenance

Built with an AI coding assistant. I chose what the repository should do and how it is validated, and decided between the options proposed during the build: the paper's own split with its elastic net as a replication anchor, checked before anything was built on it; state of health as a per-cycle trajectory, with the by-cycle split shown as the labelled wrong way; and a leakage claim that rests on this repository's own measurement and on published sources rather than on named notebooks.

About

Lithium-ion cycle life and state of health from the first 100 cycles of the Severson et al. 2019 dataset: split by cell, the paper's elastic net replicated, and a measurement of how much a by-cycle split flatters the score.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages