This repository allows one to perform a reweighting analysis between two sets of Nuisance input files.
The code is tested against Python 3.14. Grid submission through the HTCondor snakemake profile (described below) requires Snakemake v8 or later (which brought the executor-plugin interface that profile relies on), so make sure whichever environment you build satisfies that.
If you DO have conda installed, you can add --use-conda at the end of every snakemake command,
and snakemake will build the per-rule environment described in environment.yaml the first time
it is needed. Note this only manages the environment each job runs in - the environment you invoke
snakemake from yourself still needs to contain Snakemake and, for condor grid submissions, must
satisfy the Snakemake v8+ requirement above.
If you DO NOT have conda installed, skip --use-conda entirely and instead build the environment
yourself from environment.yaml and activate it before running snakemake, e.g.:
micromamba create -f environment.yaml
micromamba activate snakemake-envOr, you can create a virtual Python environment using against requirements.txt:
python3.14 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtUse the snakemake commands WITHOUT --use-conda at the end.
Each route installs the same set of packages: numpy, matplotlib, hep_ml, torch/pytorch,
mplhep, pandas, uproot, awkward, snakemake (>=8), snakemake-executor-plugin-cluster-generic,
htcondor/python-htcondor, xgboost, pot, tensorboard, pyarrow and scipy.
After cloning the repository, one should first obtain input samples. These should be Nuisance FlatTrees, and the original and target samples must live in separate directories. These directories can contain multiple samples, but only with only one ROOT file at each leaf in the directory tree. The original and target directory structures must match, and reweightings will be trained for every matching pair of files. NEUT and GENIE inputs generated by Hank Hua can be found at the following locations:
- SCARF:
/work4/ppd/scarf1488/t2knova/fds/generated/t2knova/ - Imperial LX Machines:
/vols/dune/jmm224/data/T2KNOvA/
In config.yaml, you can change several parameters for your analysis :
In the inputs section, you can specify
original(target)_dir: the directory containing your original (target) input samples (the samples you are reweighting from (to))original(target)_tree: the name of the FlatTree inside your original (target) input ROOT files
In the analysis section, you can specify
modes: the interaction mode numbers to include (these correspond to Nuisance modes)topologies: the topology breakdown- A separate training is run for each listed topology
- Choose from
[_C0pi, _C1pipm, _C1pi0, _CNpi, _COther], _=[N,C] - Topologies with insufficient events to be split into test, training and validation samples will be skipped
train_percentage: the fraction of each sample to be used for trainingval_percentagethe fraction of each samples to be used for validation- the remaining events are used for testing
downsampling: the fraction of events to keep from the input files, to speed up the analysis for testing
In the parameters section, you can specify
reweighting: the parameters used to train the reweighting modelsmetric_sets: parameter sets to be used for the reweighting metrics- eg.
8D: ["Enu_true", "PLep", "CosLep", "Q2", "q0", "q3", "PTlep", "Eav"]
- eg.
extra: additional parameters for which histograms will be plotted showing the reweighiting performance.- Note that all the parameters in
metric_setsandreweightingare automatically plotted.
- Note that all the parameters in
binning_file: the path to a json file containing the binnings for each parameter inallto be used for performance plots- A default quantile-based binning is used for any variables without specified binnings
In the models section, you can specify
model_list: list of models to train- Choose from
[binning, XGB, unnormXGB]for nowbinningis a binned reweightingXGBis a BDT-based reweighting using XGBoostunnormXGBis the same, but trained on unnormalised classes
- Choose from
grid_file: path to a json file containing the full set of hyperparameters to test for each modelselection_metric: the metric by which to select the best hyperparameter set for each model
In the swd_bootstrapping section, you can specify
runs: the number of times to call "Bootstrap_swd.py" per sample/topology/parameter set combinationn_samples: the number of samples to draw per run (total number of bootstraps isrunsxn_samples)n_directions: the number of directions to sample in the parameter space to calculate the SWD
In the output section, you can specify
tag: a string used by Snakemake to identify all saved outputs from this run
The hyperparameters of the models are fine-tuned by the workflow itself: one run of Train_model.py is carried out
per point of the grid given in the models/grid_file file, and Gather_metrics.py then finds, for each model,
the set of hyperparameters giving the best value of the metric given in the models/selection_metric section. The
chosen sets are written in set_hyperparameters/{tag}/{sample}/{topology}/hyperparameters.json and are the ones the
final training uses. Nothing has to be written by hand: to change the hyperparameters that are scanned, change the
grid file.
Once all the above steps are completed, you are ready to run the analysis. To run a complete analysis, including samples creation, model training, metrics evaluation and plotting, enter the following command in your terminal:
snakemake all --cores 8 (--use-conda (only if you use conda))This runs every sample found in the input directories, for every topology listed in the config file. To run a single sample/topology combination, you can also ask for one output file directly, for example:
snakemake saved_metrics/{tag}/{sample}/{topology}/custom_{Dim}D/metrics.json --cores 8 (--use-conda (only if you use conda))where {tag} is the tag given in the output/tag section, {sample} is the relative path of the sample directory
(eg. /FHC/numu/H2O), {topology} is one of the topologies listed in the analysis/topologies section (ex : CC0pi),
and {Dim} is the number of parameters listed in the parameters/reweighting section of config.yaml.
The metrics are written by Compute_metrics.py in saved_metrics, and the plots by Make_plots.py in saved_figures:
asking for one of them does not run the other. Once finished, you can explore the different 'saved' folders containing
the samples, models, metrics and plots.
Note on dry runs (snakemake all -n): the topology counting is a snakemake checkpoint, which means the jobs
that depend on it (all the trainings and metrics) can only be listed once the counts exist. A dry run creates no
file, so it stops at the counting jobs. To see the whole list of jobs, first run the (cheap) counting on every
sample, then ask for the dry run:
snakemake count_all_topologies --cores 8 (--use-conda (only if you use conda))
snakemake all -nThe counting is not repeated afterwards, as its output files are then up to date.
Congrats, you ran your first analysis!
The outputs of the analysis are saved in the following directories:
saved_samples/{tag}/{sample}/: the original and target samples split intotest,train,val, saved as parquet files partitioned by topology
saved_models/{tag}/{sample}/{topology}/: the best trained models for each topology and model type, saved as.jsonfiles for XGB and unnormXGB, and.pklfiles for binned reweightingsaved_models/{tag}/{sample}/{topology}/{model}/{run_id}/: all the trained models for each topology and model type, saved as.jsonfiles for XGB and unnormXGB, and.pklfiles for binned reweighting
saved_swd_distribution/{tag}/{sample}/{topology}/{param_set}/: the null SWD distribution for the target test sample in the parameter space defined byparam_set, saved as.npyfilessaved_metrics/{tag}/{sample}/{topology}/: the metrics and hyperparameters for the best trained models for each topology and model type, saved as a.jsonfile.saved_metrics/{tag}/{sample}/{topology}/{model}/{run_id}/: the metrics and hyperparameters for all the trained models for each topology and model type, saved as a.csvfile.
saved_figures/{tag}/{sample}/{topology}/: the plots showing the reweighting performance of each model for every parameter (1Dhist.pdf) and the training history for the BDT-based reweightings saved as.pdffiles.
Snakemake v8+ can submit jobs to a HTCondor grid using the cluster-generic executor plugin and the htcondor
python package. If you have created your own environment via environment.yaml or requirements.txt, these
will be installed already. If you are instead using conda, they must be available in the environment you invoke
snakemake from. You can install them with the following command:
pip install --user htcondor snakemake-executor-plugin-cluster-genericYou will need to create a htcondor snakemake profile for your HTCondor grid. An example profile for the
Imperial College batch system can be installed from https://github.com/Charlotte-Knight/htcondor-ic.
Then, you can run Snakemake with the --profile option, e.g.:
snakemake all --profile htcondorFor more information on the cluster-generic executor plugin, see
https://snakemake.github.io/snakemake-plugin-catalog/plugins/executor/cluster-generic.html
Reads input files, counts the number of events in each topology, and writes the counts to a json file if the topology has enough events to be split into training, validation and test samples. Samples passing this check are split accordingly, stripped to the parameters listed in the config, and saved as parquet files partitioned by topology. This is a checkpoint because the DAG depends on the number of topologies passing this check, which is only known after the checkpoint has run. This rule is run once for each sample.
Bootstraps swd_bootstrapping/n_samples samples from the target test sample, calculates the SWD between each sample and the
target test sample, and saves the SWD distribution as a .npy file. This rule is run swd_bootstrapping/runs times for each
sample/topology/parameter set combination. These are marked as temp, since they are later aggregated into a single file.
Local rule that aggregates the SWD distributions from run_bootstrap into a single .npy file. This rule is run once per
sample/topology/parameter set combination
Local rule that writes a .json file for each hyperparameter set listed in the grid file. These are marked as temp, since
they contain no information not contained in the grid file.
Trains and saves a single model with a single hyperparameter set. This rule is run once for each hyperparameter set for each
model, for each sample/topology combination. The trained model is saved as a .json file for XGB and unnormXGB, and as a
.pkl file for binned reweighting.
rule compute_metrics (relies on outputs of initialize_analysis, train_model and aggregate_bootstrap)
Computes the relevant metrics for a single trained model, and saves them to a .csv file. This is run once for each trained
model for each sample/topology combination. The metrics computed are:
- SWD between the reweighted original and target test sample for the reweighting parameter set and each metric set
- p-values for each of those SWDs
- Total
$\chi^2$ /dof between the reweighted original and target test sample in each of the same parameter spaces - p-values for each of those
$\chi^2$ /dof values - 1D
$\chi^2$ /dof between the reweighted original and target test sample for every parameter included in the config - p-values for each of those 1D
$\chi^2$ /dof values
These metrics are also computed between the original and target test samples themselves for comparison.
Local rule that finds the best-performing model for every model in a given sample/topology, according to the metric specified
in the config. The hyperparameters and metrics for each model are saved in a single .json file. A symlink is created to each
of the best-performing models in saved_models/{tag}/{sample}/{topology}/{model}. This rule is run once for each
sample/topology combination.
Generates plots showing the reweighting performance of each model for every parameter included in the config. For BDT-based
reweighting models, plots of the training history are also generated. This rule is run once for each sample/topology
combination, and the plots are saved as .pdf files.
Carries out the entire workflow, by requesting the outputs of the initialize_analysis checkpoint for every sample, and the outputs of choose_model and make_plots for every sample/topology combination. This rule is run once during the workflow.
Carries out the intialize_analysis checkpoint for every sample by requesting its output. This rule can be run before a dry run to see the full DAG, since the number of topologies passing the event count check is only known after the checkpoint has run.