⭐ All the results are available through this interactive website: Benchmark Explorer
CleanImp is an end-to-end benchmark for evaluating the impact of time series imputation on downstream tasks. The technical details are described in the paper CleanImp: Benchmarking the Impact of Time Series Imputation on Downstream Quality [Experiment, Analysis & Benchmark] (under review for PVLDB 27).
The benchmark follows the complete experimental pipeline introduced in the CleanImp paper:
Time Series → Contamination → Imputation → Downstream Model → Evaluation
Each benchmark execution creates a unique experiment directory. The directory tree below describes each generated folder and file:
framework/
│
├── _caching/ # Cache used to avoid recomputing expensive pipelines
│ └── ... # Imputed matrices and downstream predictions
│
├── imputegap_assets/
│ └── benchmark/
│ └── <experiment>/ # Unique directory created for the benchmark run
│ │
│ ├── <dataset>/ # Results grouped by evaluated dataset
│ │ └── <pattern>/ # Results grouped by missingness pattern
│ │ └── error/ # Upstream/downstream metric results
│ │ │ ├── _metrics_subplot.jpg
│ │ │ │ # Metric evolution across missingness rates
│ │ │ └── report_<pattern>_<dataset>.txt
│ │ │ # Detailed results for this dataset/
│ │ └── recovery/ # Individal imputation plot for each rate
│ │
│ ├── _heatmaps/ # Aggregated results used for heatmap analyses
│ │ └── _benchmarking_*.txt # Benchmark results represented as heatmaps
│ │
│ |── _summary_cleanimp_*.xlsx # ★ Complete summary of metrics and evaluated pipelines ★
│ ├── experimentation_setup.txt # Exact configuration of the benchmark run
│ ├── report_cleanimp_benchmark_*.log # Complete benchmark results in log format
│ └── runtime.log # Runtime information for the experiment
│
└── *_log.txt # General execution logs
Techniques and Datasets | Benchmark Parameterization | Prerequisites | Benchmark Execution
The tables below provide a compact overview of the options available when configuring the benchmark.
| Imputation Algorithms | ||||
|---|---|---|---|---|
BRITS |
BayOTIDE |
BitGraph |
CDRec |
CSDI |
DeepMVI |
DynaMMo |
GAIN |
GPT4TS |
GRIN |
GROUSE |
HKMFT |
IIM |
IterativeSVD |
MICE |
MPIN |
MRNN |
MeanImpute |
MissForest |
MissNet |
Moment |
NuwaTS |
PRISTI |
ROSL |
SAITS |
SPIRIT |
STMVL |
SVT |
SoftImpute |
TKCM |
TRMF |
TimesNet |
XGBOOST |
| Forecasting Models | ||||
|---|---|---|---|---|
arima |
chronos |
croston |
deepar |
dlinear |
exp-smoothing |
hw-add |
lightgbm |
lstm |
ltsf |
moment |
nbeats |
nlinear |
patchtst |
prophet |
transformer |
xgboost |
| Classification Models | ||||
|---|---|---|---|---|
arsenal |
catch22 |
cboss |
cif |
cnn |
itde |
knn |
lstm |
proxstump |
shapedtw |
signature |
stc |
svc |
tsf |
tsfresh |
weasel |
| Forecasting Datasets | ||||
|---|---|---|---|---|
airq |
atm |
beijing_traffic |
climate |
czelan |
economics |
electricity |
etth1 |
etth2 |
human_access |
ili |
nn5 |
nyse |
paris |
wike2000 |
wind_speed |
| Classification Datasets | ||||
|---|---|---|---|---|
Adiac |
ArrowHead |
Beef |
BirdChicken |
CBF |
Car |
CinCECGTorso |
Colposcopy |
Computers |
CricketX |
DistalPhalanxOutlineAgeGroup |
DistalPhalanxOutlineCorrect |
DistalPhalanxTW |
ECGFiveDays |
EOGHorizontalSignal |
EOGVerticalSignal |
Earthquakes |
EthanolLevel |
FaceFour |
FacesUCR |
Fish |
FreezerRegularTrain |
FreezerSmallTrain |
GunPoint |
GunPointAgeSpan |
GunPointMaleVersusFemale |
GunPointOldVersusYoung |
Ham |
Haptics |
Herring |
InlineSkate |
InsectEPGRegularTrain |
LargeKitchenAppliances |
Lightning2 |
Lightning7 |
Meat |
MedicalImages |
MoteStrain |
OSULeaf |
OliveOil |
PhalangesOutlinesCorrect |
PowerCons |
ProximalPhalanxOutlineAgeGroup |
ProximalPhalanxOutlineCorrect |
RefrigerationDevices |
Rock |
ScreenType |
SemgHandGenderCh2 |
SemgHandMovementCh2 |
SemgHandSubjectCh2 |
ShapeletSim |
SharePriceIncrease |
SmallKitchenAppliances |
SonyAIBORobotSurface1 |
Strawberry |
SwedishLeaf |
SyntheticControl |
ToeSegmentation1 |
ToeSegmentation2 |
Trace |
TwoLeadECG |
TwoPatterns |
UMD |
Wafer |
Wine |
Worms |
WormsTwoClass |
Yoga |
The datasets can be found in the following directory: https://github.com/eXascaleInfolab/CleanImp/tree/main/framework/datasets/
The list of the runner’s parameters is the following:
| Parameters/Arguments | Values | Description |
|---|---|---|
--task |
{imputation, forecasting, classification} |
Selects the benchmark task to execute. |
--imp_algs |
List of Algorithms | Selects one or more imputation algorithms. Use all to include every available algorithm. |
--datasets_type |
{forecasting, classification} |
Selects the dataset category used for the benchmark. |
--datasets_list |
List of Datasets | Selects one or more datasets. Use all to include every available dataset. |
--patterns |
{mcar, seqn, blks} |
Selects one or more missingness patterns. Use all to include every available pattern. |
--miss_rate |
{0.1, 0.2, 0.4, 0.6, 0.8} |
Sets one or more missing-value rates. Use all to evaluate the predefined rates. |
--metrics |
{RMSE, MAE, MI, CORRELATION} / {SMAPE, MSE, MAE} / {F1, ACCURACY, RECALL} |
Selects the evaluation metrics for the chosen task. Use all to include every available metric. |
--downstream_mod |
List of Models | Selects the downstream model used to measure the impact of imputation. |
--horizon |
{12, 24, 48} |
Sets the forecasting horizon as the number of future timestamps. (only forecasting) |
--con_by_class |
--no-con_by_class |
Sets if the contamination is performed separately for each class or not. (only classification) |
--imp_by_class |
--no-imp_by_class |
Sets if the imputation is performed separately for each class or not. (only classification) |
--caching |
--no-caching |
Enables or disables caching of intermediate results to avoid repeated computations. |
--plots |
--no-plots |
Enables or disables benchmark plot generation. |
--verbose |
--no-verbose |
Enables or disables detailed execution output. |
With the “caching” tag, you can store the imputed matrix and, for downstream tasks, the classification or prediction results. This allows you to rerun the benchmark without having to recompute pipelines that have already been executed.
CleanImp is implemented in Python and relies on ImputeGAP for time series contamination and imputation. Please start by cloning the GitHub repository:
git clone https://github.com/eXascaleInfolab/CleanImp
cd CleanImp/To set up your environment and prepare the development installation for C++, please run this script of installation:
source cleanimp_install.sh
cd framework/The computed results and the plots of the benchmark will be saved in: ./imputegap_assets/benchmark/*
- We show how to produce the imputation results on the forecasting datasets
To produce the imputation results with one forecasting dataset (Paris), one imputation algorithm (SAITS), one missingness pattern (MCAR), a missing rate (20%), one metric (RMSE), and a forecasting horizon (12 timestamps), run the following command:
python cleanimp_bench.py \
--task imputation \
--imp_algs SAITS \
--datasets_type forecasting \
--datasets_list paris \
--patterns mcar \
--miss_rate 0.2 \
--metrics RMSE \
--horizon 12To produce the imputation results with two imputation algorithms (MICE and MeanImpute), two datasets (Paris and ILI), two missingness patterns (MCAR and SeqN), two missingness rates (0.1 and 0.8), and a forecasting horizon (24 timestamps), run the following command:
python cleanimp_bench.py \
--task imputation \
--imp_algs MICE MeanImpute \
--datasets_type forecasting \
--datasets_list paris ili \
--patterns mcar seqn \
--miss_rate 0.1 0.8 \
--metrics all \
--horizon 24To produce the full imputation results (left-hand part of Fig. 6 in the paper) using all available configurations, replace the parameter values with all:
python cleanimp_bench.py \
--task imputation \
--imp_algs all \
--datasets_type forecasting \
--datasets_list all \
--patterns all \
--miss_rate all \
--metrics all \
--horizon 12To produce the imputation results for the classification datasets (right-hand part of Fig 6 in the paper), change the datasets_type tag to classification:
python cleanimp_bench.py \
--task imputation \
--imp_algs all \
--datasets_type classification \
--datasets_list all \
--patterns all \
--miss_rate all \
--metrics allTo produce the impact of imputation with a forecasting model (chronos) with one forecasting dataset (Paris), one imputation algorithm (SAITS), one missingness pattern (MCAR), one missing rate (20%), all downstream metrics, and a forecasting horizon (12 timestamps), run the following command:
python cleanimp_bench.py \
--task forecasting \
--downstream_mod chronos \
--imp_algs SAITS \
--datasets_list paris \
--patterns mcar \
--miss_rate 0.2 \
--metrics all \
--horizon 12To produce the impact of imputation with a forecasting model (chronos) with two imputation algorithms (MICE and MeanImpute), two datasets (Paris and ILI), two missingness patterns (MCAR and SeqN), two missingness rates (0.1 and 0.8), all downstream metrics, and a forecasting horizon (24 timestamps), run the following command:
python cleanimp_bench.py \
--task forecasting \
--downstream_mod chronos \
--imp_algs MICE MeanImpute \
--datasets_list paris ili \
--patterns mcar seqn \
--miss_rate 0.1 0.8 \
--metrics all \
--horizon 24To produce the impact of imputation with a forecasting model with all possible configurations (Fig. 7, 9 and 10 in the paper),
replace the parameter values with all and adjust the forecasting model and horizon accordingly:
python cleanimp_bench.py \
--task forecasting \
--downstream_mod chronos \
--imp_algs all \
--datasets_list all \
--patterns all \
--miss_rate all \
--metrics all \
--horizon 12To produce the impact of imputation with a classification model (arsenal) with one classification dataset (Computers), one imputation algorithm (GRIN), one missingness pattern (MCAR), one missing rate (20%), and all downstream metrics, run the following command:
python cleanimp_bench.py \
--task classification \
--downstream_mod arsenal \
--datasets_list Computers \
--imp_algs GRIN \
--patterns mcar \
--miss_rate 0.2 \
--metrics allTo produce the impact of imputation with a classification model (arsenal) with two classification datasets (Computers and Car), two imputation algorithms (MeanImpute and MICE), two missingness patterns (MCAR and SeqN), two missing rates (20% and 80%), and all downstream metrics, run the following command:
python cleanimp_bench.py \
--task classification \
--downstream_mod arsenal \
--datasets_list Computers Car \
--imp_algs MeanImpute MICE \
--patterns mcar seqn \
--miss_rate 0.1 0.8 \
--metrics allTo produce the impact of imputation with a classification model with all possible configurations (Fig. 11, 13 and 14 in the paper),
replace the parameter values with all and adjust the classification model and setup accordingly:
python cleanimp_bench.py \
--task classification \
--downstream_mod arsenal \
--datasets_list all \
--imp_algs all \
--patterns all \
--miss_rate all \
--metrics all \
--cont_by_class \
--imp_by_classTo produce a breakdown of the results by feature, run the following commands (scripts are underway).