Live · Battery Data Intelligence Platform

Battery data,
research-ready.

BatteryLake unifies the world's fragmented battery cycling datasets into one standardized, quality-assessed foundation — so models are compared on science, not on preprocessing luck.

batterylake.org / datasets ?quality=ready
LCO 2007_NASA_PCoE_LCO_18650_1C_1C_25TNASA Prognostics CoE 34 cells 0.96
LFP 2019_Stanford_MIT_TRI_LFP_18650_MultiC_30TStanford · MIT · TRI 124 cells 0.94
NMC 2024_Imperial_Kirkaldy_NMC_21700_MultiC_MultiTImperial College London 21 cells 0.92
LFP 2026_NTU_Ampace-Samsung_LFP-NMC_21700_2C_2C_25TNTU EEE · Singapore 16 cells 0.95
SOH 94.2%cycle 187 · B0005
Quality gate passed26 validation rules
0
Curated datasets
0
Worldwide institutions
0
Years span
2007-2026
0
Total cycles
0
Total cells
0 GB
Data volume
How it works
From raw lab exports to comparable results
Every dataset travels the same auditable path, so downstream numbers mean the same thing everywhere.
Curate & Ingest 40+ public and lab datasets, tracked with DOI-level provenance.
Standardize One ETL pipeline, one schema: metadata, time-series, cycle summaries.
Quality Gate 26 physical-plausibility and schema rules score every dataset.
Benchmark & Share Reproducible SOH / RUL runs with pinned splits and unified metrics.
Platform
One workspace for the full research loop
Browse, benchmark, and audit — each surface is built on the same standardized data core.
Benchmarks

Reproducible experiments, not preprocessing folklore

  • Random, temporal, and cross-cell split protocols
  • 7 reference models — Ridge to Transformer and PINN
  • RMSE · MAE · MAPE reported on identical folds
Open the benchmark workbench
Split protocol — cross-cellseed 42
Train
62%
Val
18%
Test
20%
RMSE0.0121
MAE0.0094
MAPE1.08%
Quality

An auditable gate before any model sees the data

  • Physical plausibility: voltage windows, energy balance
  • Signal-level QC maps for V(t), I(t), T(t), Qd(n)
  • JSON reports wired into the benchmark pipeline
See the quality methodology
Quality gate — NASA PCoEready
0.96overall
Voltage range 2.0–4.5 V
Coulombic efficiency 95–105%
Temperature drift — 1 review
APIs

The same platform, programmable

  • RESTful endpoints with token auth
  • Parquet, CSV, and JSON exports
  • Embedded DOI citations in every response
Explore the developer console
GET api.batterylake.org/v1/datasets
import requests

resp = requests.get(
  "https://api.batterylake.org/v1/datasets",
  params={"chemistry": "LFP", "quality": "ready"}
)

for ds in resp.json()["items"]:
  print(ds["ref_name"], ds["quality_score"])

Start with research-ready battery data today.

Browse the catalog, download standardized packages, or contribute your lab's datasets to the community.

Platform Usage
Total Pageviews —
Dataset Downloads —
Skill Uses —
Design by
Cloud Application and Platform Group
Nanyang Technological University Singapore
In partner with
SODA Group
Agency for Science, Technology and Research Singapore
Category
Calendar Aging
Characterization
Cycle Aging
EIS
Field Data
Field Fault Diagnosis
Relaxation
SOC Estimation
SOH Estimation
Thermal Runaway
Chemistry
LCO
LFP
NCA
NMC
Domain
EV
Grid
Lab
Form
18650
21700
Cyl
Pouch
Prismatic
Profile
CC/CV
Dynamic
Multi-rate
0 datasets

Benchmark results

Source
Experiment scope
Task definition

Models evaluated · Unranked

  • Linear Regression
  • Random Forest
  • XGBoost
  • LSTM
  • Transformer
  • CNN
  • PINN

Loading prediction curves…

Run your own benchmark

Choose Evaluation Task
SOH Estimation
01Select Dataset
Category
Cycle Aging
Chemistry
LCO
LFP
NCA
NMC
Domain
EV
Grid
Lab
Form
18650
21700
Cyl
Pouch
Prismatic
Profile
CC/CV
Dynamic
Multi-rate
02Select Input Signals
Configure Data Partitioning
Train
—
Validation
—
Test
—
Excluded not used in train / validation / test 0 cells
Drag cells between Train / Validation / Test or into Excluded to leave them out
Select Benchmark Models
Execute Local Training Package
01Download Training Package

Get the configured dataset, model architecture, and runtime environment.

bt_benchmark_soh_lstm.zip
02Run in Your Terminal

Unzip the package, open your terminal, and execute the following command to start local training.

unzip bt_benchmark_soh_lstm.zip
cd bt_benchmark_soh_lstm
mkdir -p data
# copy the processed dataset folder OR a .zip of it into data/
# example: cp -R /path/to/2019_Stanford_MIT_TRI_LFP_18650_MultiC_30T data/
# example: cp /path/to/2019_Stanford_MIT_TRI_LFP_18650_MultiC_30T.zip data/
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
bash run_benchmark.sh
03Upload Output Folder

When training finishes, upload the generated output folder containing your results.

Drag & drop the output folder here, or click to browse Expected: metrics.json, predictions.csv, and predictions/<cell>.csv (one file per cell, every cycle)
Waiting for local output folder
Benchmark Evaluation Results
Performance Highlights
Best Model—
Test RMSE—
Per-Cell Prediction Trajectory
True SOH Pred SOH
Step 1 of 6 - Benchmark Task

How the Studio Works

Inside the Calibration Engine
PIML Objective
ℒ(θ) = ℒdata + λ ℒphysics
Inputs
Electrical Data Thermal Data Operating Conditions Physics Constraints
Battery measurements feeding a physics-informed neural network

Studio Workspace

Start by selecting a dataset

01

Select Dataset

02

Feature Identification

Voltage (V)
Cycle / time
Current (A)
Cycle / time
Temperature (°C)
Cycle / time
Capacity (Ah)
Cycle / time
Internal Resistance (mΩ)
Cycle / time
SOC Range (%)
Cycle / time
03

Calibrate Battery Digital Twin

Training LossTrain Val
10⁰10⁻¹10⁻²10⁻³05001kIterations
0%Twin build
Calibration Progress
Select and confirm a dataset to build its digital twin
04

Synthetic Data Generation Setup

05

Synthetic Data Generation Results

Build a digital twin, adjust the setup, then generate data.
Framework
PyTorch
scikit-learn
XGBoost
Status
Available
Experimental
Recommended
Requires GPU
Task
RUL
SOH
Type
Attention Model
Convolutional Neural Network
Ensemble
Gradient Boosting
Physics-Informed Neural Network
Recurrent Neural Network
Statistical
— models
0
Completed
0
In Progress
0
Pending
0
Skipped
IDDatasetStatusMetaSeriesSummaryQC
Format pattern · Connected
BatteryLake reference name sequencer
Inspecting Field 1 of 7
Assembled reference name
2007_NASA_PCoE_LCO_18650_1C_1C_25T
Interaction: automatic field walkthrough; hover or focus to inspect one field.7 fields · underscore-delimited

Examples

Datasetref_name
NASA PCoE2007_NASA_PCoE_LCO_18650_1C_1C_25T
Stanford-MIT-TRI2019_Stanford_MIT_TRI_LFP_18650_MultiC_30T
NTU EEE Internal2026_NTU_Ampace-Samsung_LFP-NMC_21700_2C_2C_25T
Imperial 217002024_Imperial_Kirkaldy_NMC_21700_MultiC_MultiT

Field Vocabulary

Sources (25+)
NASA_PCoE · CALCE_UMD · Stanford_MIT_TRI · Oxford_Howey · RWTH_Aachen · NTUEEE · SNL · HNEI · UL_Purdue · XJTU · KIT · Tsinghua · CMU_Bills · Stanford_Onori · TUM · Beihang · Imperial · ISU_ILCC · EVERLASTING_4TU · Mendeley · Figshare · Zenodo ...
Chemistries
LFP · NMC · NMC811 · LCO · NCA · LiIon (generic) · MultiChem
Form Factors
18650 · 21700 · Pouch · Prismatic · Cyl · Auto · EV-BMS

Getting Started

Clone the repository and install the conda environment.

git clone https://github.com/tianwen1209/BatteryLake-Benchmark-DataPrep.git
cd BatteryLake-Benchmark-DataPrep
conda create -n batterylake python=3.10
conda activate batterylake
pip install -r requirements.txt

Dataset Registry

All datasets are tracked in dataset_registry.csv with columns for dataset identity, DOI, source URL, assigned owner, ref_name, processing status, QC status, last update, and notes.

Evaluation

The benchmark evaluation framework uses evaluate.py and dataset_interface.py to run baseline models across selected datasets and split protocols.

Team

Research and engineering contributors

YW
Prof. Yonggang Wen
Principal Investigator
WH
Dr Wang Hao
Research Fellow
ZT
Zhu Tianwen
PhD Candidate
CY
Cai Yezi
Frontend Engineer
Affiliation

Nanyang Technological University

Singapore

Repository

BatteryLake-Benchmark-DataPrep

github.com/tianwen1209/BatteryLake-Benchmark-DataPrep

Licence

BatteryLake provides access to battery datasets, metadata, benchmarking resources, and related research materials collected from multiple sources.

Individual datasets, algorithms, and other resources may come with their own licences and citation requests. Please honour these requirements. BatteryLake will display the applicable licence and citation information whenever it is known.

Before downloading, redistributing, modifying, or using a resource, please review the licence and usage terms provided on its dataset or resource page. When source-specific terms are available, those terms take precedence over the general information provided on this page.

BatteryLake does not grant additional rights to third-party content beyond those provided by the original authors, institutions, or data owners.

Citation

If you use BatteryLake, its curated datasets, standardized outputs, benchmarking resources, or platform tools in academic work, please cite the BatteryLake paper below.

BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and Benchmarking
Tianwen Zhu, Hao Wang, and Yonggang Wen, 2026.
BibTeX
@misc{zhu2026batterylakeagenticphysicsgroundedcuration,
      title         = {BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and Benchmarking},
      author        = {Tianwen Zhu and Hao Wang and Yonggang Wen},
      year          = {2026},
      eprint        = {2607.09762},
      archivePrefix = {arXiv},
      primaryClass  = {cs.AI},
      url           = {https://arxiv.org/abs/2607.09762}
}

When using an individual dataset, algorithm, model, or other contributed resource, please also cite its original authors and follow any citation instructions shown on the corresponding resource page.

Citing both BatteryLake and the original resource helps ensure that the platform contributors, dataset creators, and research institutions receive appropriate credit.

No file selected.
Sample CSVs to try uploading
Example Result dataset_03_quality_assesment.json · Assessed 14:32 · example values Ready 1 warning
0.97
Completeness
Missing channels2.1%
0.95
Consistency
Sequence issues3 flags
0.92
Accuracy
Physical checks5 / 6 pass
1.00
Validity
Schema errors0
0.96
Overall Score
Ready 1 warning
Physical Plausibility Checks
High-signal checks that determine whether a dataset can safely enter preprocessing and benchmark training.
5 passed - 1 warning
  • Voltage Range Validation
    All cell voltages within 2.0V-4.5V nominal operating range for the stated chemistry.
  • Energy Balance Check
    Charge/discharge energy integral consistency; coulombic efficiency remains within 95-105% per cycle.
  • Capacity Monotonicity
    Degradation trajectory follows expected non-increasing trend with allowable recovery windows.
  • Temperature Consistency
    Cell surface temperature must remain within 5°C of stated test condition; one dataset needs review.
  • Timestamp Integrity
    Monotonically increasing timestamps with no negative intervals or unreasonable gaps above 24h.
  • Current Direction Consistency
    Charge and discharge current signs follow one convention throughout the dataset.
Diagnostic Output
quality_report.json
{
  "dataset_id": "dataset_03",
  "quality_score": {
    "completeness": 0.97,
    "consistency": 0.95,
    "accuracy": 0.92,
    "validity": 1.00
  },
  "overall": 0.96,
  "gate": "ready_with_warning",
  "checks": [
    { "name": "voltage_range", "passed": true },
    { "name": "temperature_consistency", "status": "review" },
    { "name": "capacity_mono", "passed": true }
  ],
  "generated_at": "2026-04-28T12:00:00Z"
}

01Set Up the Processing Skill

01

Download the skill package

Save batterylake-processing.zip to your BatteryLake data repository.

View SKILL.md Processing standard
See package contents

02Choose a Dataset and Send the Prompt to Your Agent

  1. 01Choose a dataset
    inRaw_Dataset/…
    outProcessed_Dataset/…
  2. 02Paste the prompt to your agent
    
              

    The agent reads the standard and the dataset TODO, then continues from the first unfinished gate.

  3. 03Watch it work
    • Progress lives in status.jsonone line per gate, plus TODO.md
    • Interrupted runs resumefinished steps are skipped when fingerprints match
    • Raw files stay untouchedundecodable data is reported, never invented
    Check the result below

While processing, your agent runs these five checks

  1. InventoryList every file and archive member, hash each source.inventoried
  2. SemanticsConfirm cells, units, clocks, cycle boundaries and labels.adapter design confirmed
  3. ConversionWrite every source field to the canonical layer.converted
  4. FidelityCompare raw and standard values record by record.canonical_validated
  5. EquivalenceRaw and standard loaders train to the same result.benchmark_verified

Your agent produces these files and follows these rules

What the skill produces
click an item
Rules the skill never skips
click to expand
  • Identify files by content Magic bytes, complete payloads and EOF are checked. A .pkl may hold several streams; extensions are only a hint.
  • No cleaning in the fidelity layer No resampling, interpolation, clipping, float32 down-casting or dropped diagnostics. Those belong to a named task view.
  • Capacities are not interchangeable Health capacity, partial-DoD throughput, RPT capacity, rated capacity and author SOH stay separate. Unclear labels are withheld, not guessed.
  • EOL is an observed event The last row is not EOL. Sparse RPTs may only bound an interval, and right-censoring is recorded instead of a fake RUL.
  • No leakage across splits The same physical cell, duplicate copies and overlapping windows stay on one side; scalers are fit on training data only.
  • Gates advance only with evidence inventoried, converted, canonical_validated and benchmark_verified are separate; none is inferred from the one before.

03Upload status.json to Check Benchmark Readiness

Try a real one:
Example dataset_21 needs_semantic_review
Developer console
Explore endpoints as a working platform surface, not a brochure.
Choose a route from the library, inspect its parameters, preview the request, then run a simulated call to see the exact response shape the platform returns.
Live routes0
Formats0
Auth modeToken
Dataset catalog endpoint
Designed for search, filtering, and benchmark preparation. The response returns only datasets that pass the required ETL and quality gates.
Available
GET/v1/datasets?chemistry=LCO&quality=ready
chemistryLCO | LFP | NMC
formatCSV | Parquet
split_readytrue
dataset_id
cell_count
cycle_count
quality_score
response · application/json200 OK

          
Application workbench
Applications are grouped by what the user is trying to do, not by abstract product names.
3 surfaces
BenchmarkExperiment Builder

Pick datasets, split cells, select models, and run comparable SOH/RUL benchmarks.

FeaturesFormula Studio

Upload custom signal formulas and compute cycle-level model inputs from raw curves.

ExplainInterpretation Lab

Inspect feature importance, attention maps, and degradation signals after training.

Open API
Data Download API
Access harmonized, quality-assessed datasets via RESTful endpoints with embedded citations, metadata, and provenance tracking.
  • RESTful endpoints for dataset listing and download
  • Parquet, CSV, and JSON export formats
  • Embedded DOI citations and provenance
  • Selective cell/cycle range queries
Open API
Benchmark API
Run standardized ML benchmarks across curated datasets with reproducible train/val/test splits and unified evaluation metrics.
  • SOH estimation and RUL prediction tasks
  • 5+ baseline models (LR, RF, XGBoost, LSTM, Transformer)
  • Random, temporal, and cross-cell split protocols
  • RMSE, MAE, MAPE unified metric suite
Coming Soon
Feature Engineering
Build custom feature extraction pipelines with interactive notebooks. Extract dQ/dV features, health indicators, and domain-specific signals for advanced analytics.
  • dQ/dV peak detection and tracking
  • Health indicator extraction (IC, DV)
  • Statistical feature pipeline
  • EIS parameter fitting
Coming Soon
Model Interpretation
Understand model predictions with explainability tools. Visualize attention maps, feature importance, and identify key battery aging patterns driving predictions.
  • SHAP feature importance analysis
  • Attention map visualization
  • Gradient-based attribution
  • Integrated gradients for deep models

Platform Roadmap

Data Download API

Open API — harmonized dataset access with provenance and citations.

Available
Benchmark API

Open API — reproducible SOH / RUL experiment execution.

In Progress
Feature Engineering

End-to-end application — dQ/dV, health indicators, formula features.

Planned
Model Interpretation

End-to-end application — SHAP, attention maps, attribution.

Planned
  1. 01Describe0 / 12 required fields
  2. 02Raw datano link yet
  3. 03Submit0 / 5 ready
01

Describe the dataset

Only what the catalog needs and what the processing skill needs to convert your files. Your draft stays in this browser.

Dataset metadata
For the processing skill
02

Provide the raw data

Host the original cycler exports where the team can download them and paste the link. Keep the original file names; do not resample or clean.

03

Submit

Opens a GitHub issue prefilled with your metadata, notes, checklist and data link; attach nothing else. The team takes it from there.

    Readiness
    0 / 5
    Open GitHub issue

    After submission the team runs the batterylake-processing skill on your files. You get a status.json and a quality report, and the dataset appears in the catalog under your reference name with your DOI and license.

    BatteryLake AI Assistant
    BatteryLake assistant
    Online
    Hi! I answer questions about BatteryLake: catalog numbers, individual datasets, the processing skill, status.json, benchmarks and citation. Ask in English or 中文. Now