Skip to content

Latest commit

 

History

134 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

This repository allows one to perform a reweighting analysis between two sets of Nuisance input files.

Environment setup

The code is tested against Python 3.14. Grid submission through the HTCondor snakemake profile (described below) requires Snakemake v8 or later (which brought the executor-plugin interface that profile relies on), so make sure whichever environment you build satisfies that.

If you DO have conda installed, you can add --use-conda at the end of every snakemake command, and snakemake will build the per-rule environment described in environment.yaml the first time it is needed. Note this only manages the environment each job runs in - the environment you invoke snakemake from yourself still needs to contain Snakemake and, for condor grid submissions, must satisfy the Snakemake v8+ requirement above.

If you DO NOT have conda installed, skip --use-conda entirely and instead build the environment yourself from environment.yaml and activate it before running snakemake, e.g.:

micromamba create -f environment.yaml
micromamba activate snakemake-env

Or, you can create a virtual Python environment using against requirements.txt:

python3.14 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

Use the snakemake commands WITHOUT --use-conda at the end.

Each route installs the same set of packages: numpy, matplotlib, hep_ml, torch/pytorch, mplhep, pandas, uproot, awkward, snakemake (>=8), snakemake-executor-plugin-cluster-generic, htcondor/python-htcondor, xgboost, pot, tensorboard, pyarrow and scipy.

Inputs

After cloning the repository, one should first obtain input samples. These should be Nuisance FlatTrees, and the original and target samples must live in separate directories. These directories can contain multiple samples, but only with only one ROOT file at each leaf in the directory tree. The original and target directory structures must match, and reweightings will be trained for every matching pair of files. NEUT and GENIE inputs generated by Hank Hua can be found at the following locations:

  • SCARF: /work4/ppd/scarf1488/t2knova/fds/generated/t2knova/
  • Imperial LX Machines: /vols/dune/jmm224/data/T2KNOvA/

Configuration

In config.yaml, you can change several parameters for your analysis :

In the inputs section, you can specify

  • original(target)_dir: the directory containing your original (target) input samples (the samples you are reweighting from (to))
  • original(target)_tree: the name of the FlatTree inside your original (target) input ROOT files

In the analysis section, you can specify

  • modes: the interaction mode numbers to include (these correspond to Nuisance modes)
  • topologies: the topology breakdown
    • A separate training is run for each listed topology
    • Choose from [_C0pi, _C1pipm, _C1pi0, _CNpi, _COther], _=[N,C]
    • Topologies with insufficient events to be split into test, training and validation samples will be skipped
  • train_percentage: the fraction of each sample to be used for training
  • val_percentage the fraction of each samples to be used for validation
    • the remaining events are used for testing
  • downsampling: the fraction of events to keep from the input files, to speed up the analysis for testing

In the parameters section, you can specify

  • reweighting: the parameters used to train the reweighting models
  • metric_sets: parameter sets to be used for the reweighting metrics
    • eg. 8D: ["Enu_true", "PLep", "CosLep", "Q2", "q0", "q3", "PTlep", "Eav"]
  • extra: additional parameters for which histograms will be plotted showing the reweighiting performance.
    • Note that all the parameters in metric_sets and reweighting are automatically plotted.
  • binning_file: the path to a json file containing the binnings for each parameter in all to be used for performance plots
    • A default quantile-based binning is used for any variables without specified binnings

In the models section, you can specify

  • model_list: list of models to train
    • Choose from [binning, XGB, unnormXGB] for now
      • binning is a binned reweighting
      • XGB is a BDT-based reweighting using XGBoost
      • unnormXGB is the same, but trained on unnormalised classes
  • grid_file: path to a json file containing the full set of hyperparameters to test for each model
  • selection_metric: the metric by which to select the best hyperparameter set for each model

In the swd_bootstrapping section, you can specify

  • runs: the number of times to call "Bootstrap_swd.py" per sample/topology/parameter set combination
  • n_samples: the number of samples to draw per run (total number of bootstraps is runs x n_samples)
  • n_directions: the number of directions to sample in the parameter space to calculate the SWD

In the output section, you can specify

  • tag: a string used by Snakemake to identify all saved outputs from this run

Hyperparameters

The hyperparameters of the models are fine-tuned by the workflow itself: one run of Train_model.py is carried out per point of the grid given in the models/grid_file file, and Gather_metrics.py then finds, for each model, the set of hyperparameters giving the best value of the metric given in the models/selection_metric section. The chosen sets are written in set_hyperparameters/{tag}/{sample}/{topology}/hyperparameters.json and are the ones the final training uses. Nothing has to be written by hand: to change the hyperparameters that are scanned, change the grid file.

Running the analysis

Once all the above steps are completed, you are ready to run the analysis. To run a complete analysis, including samples creation, model training, metrics evaluation and plotting, enter the following command in your terminal:

snakemake all --cores 8  (--use-conda  (only if you use conda))

This runs every sample found in the input directories, for every topology listed in the config file. To run a single sample/topology combination, you can also ask for one output file directly, for example:

snakemake saved_metrics/{tag}/{sample}/{topology}/custom_{Dim}D/metrics.json --cores 8  (--use-conda  (only if you use conda))

where {tag} is the tag given in the output/tag section, {sample} is the relative path of the sample directory (eg. /FHC/numu/H2O), {topology} is one of the topologies listed in the analysis/topologies section (ex : CC0pi), and {Dim} is the number of parameters listed in the parameters/reweighting section of config.yaml.

The metrics are written by Compute_metrics.py in saved_metrics, and the plots by Make_plots.py in saved_figures: asking for one of them does not run the other. Once finished, you can explore the different 'saved' folders containing the samples, models, metrics and plots.

Note on dry runs (snakemake all -n): the topology counting is a snakemake checkpoint, which means the jobs that depend on it (all the trainings and metrics) can only be listed once the counts exist. A dry run creates no file, so it stops at the counting jobs. To see the whole list of jobs, first run the (cheap) counting on every sample, then ask for the dry run:

snakemake count_all_topologies --cores 8  (--use-conda  (only if you use conda))
snakemake all -n

The counting is not repeated afterwards, as its output files are then up to date.

Congrats, you ran your first analysis!

Outputs

The outputs of the analysis are saved in the following directories:

Samples

  • saved_samples/{tag}/{sample}/: the original and target samples split into test, train, val, saved as parquet files partitioned by topology

Models

  • saved_models/{tag}/{sample}/{topology}/: the best trained models for each topology and model type, saved as .json files for XGB and unnormXGB, and .pkl files for binned reweighting
  • saved_models/{tag}/{sample}/{topology}/{model}/{run_id}/: all the trained models for each topology and model type, saved as .json files for XGB and unnormXGB, and .pkl files for binned reweighting

Metrics

  • saved_swd_distribution/{tag}/{sample}/{topology}/{param_set}/: the null SWD distribution for the target test sample in the parameter space defined by param_set, saved as .npy files
  • saved_metrics/{tag}/{sample}/{topology}/: the metrics and hyperparameters for the best trained models for each topology and model type, saved as a .json file.
  • saved_metrics/{tag}/{sample}/{topology}/{model}/{run_id}/: the metrics and hyperparameters for all the trained models for each topology and model type, saved as a .csv file.

Figures

  • saved_figures/{tag}/{sample}/{topology}/: the plots showing the reweighting performance of each model for every parameter (1Dhist.pdf) and the training history for the BDT-based reweightings saved as .pdf files.

Grid submissions

Snakemake v8+ can submit jobs to a HTCondor grid using the cluster-generic executor plugin and the htcondor python package. If you have created your own environment via environment.yaml or requirements.txt, these will be installed already. If you are instead using conda, they must be available in the environment you invoke snakemake from. You can install them with the following command:

pip install --user htcondor snakemake-executor-plugin-cluster-generic

You will need to create a htcondor snakemake profile for your HTCondor grid. An example profile for the Imperial College batch system can be installed from https://github.com/Charlotte-Knight/htcondor-ic.

Then, you can run Snakemake with the --profile option, e.g.:

snakemake all --profile htcondor

For more information on the cluster-generic executor plugin, see https://snakemake.github.io/snakemake-plugin-catalog/plugins/executor/cluster-generic.html

Snakemake rule descriptions

checkpoint initialize_analysis

Reads input files, counts the number of events in each topology, and writes the counts to a json file if the topology has enough events to be split into training, validation and test samples. Samples passing this check are split accordingly, stripped to the parameters listed in the config, and saved as parquet files partitioned by topology. This is a checkpoint because the DAG depends on the number of topologies passing this check, which is only known after the checkpoint has run. This rule is run once for each sample.

rule run_bootstrap (relies on outputs of initialize_analysis)

Bootstraps swd_bootstrapping/n_samples samples from the target test sample, calculates the SWD between each sample and the target test sample, and saves the SWD distribution as a .npy file. This rule is run swd_bootstrapping/runs times for each sample/topology/parameter set combination. These are marked as temp, since they are later aggregated into a single file.

rule aggregate_bootstrap (relies on outputs of run_bootstrap)

Local rule that aggregates the SWD distributions from run_bootstrap into a single .npy file. This rule is run once per sample/topology/parameter set combination

rule prepare_hps (relies on outputs of initialize_analysis)

Local rule that writes a .json file for each hyperparameter set listed in the grid file. These are marked as temp, since they contain no information not contained in the grid file.

rule train_model (relies on outputs of initialize_analysis and prepare_hps)

Trains and saves a single model with a single hyperparameter set. This rule is run once for each hyperparameter set for each model, for each sample/topology combination. The trained model is saved as a .json file for XGB and unnormXGB, and as a .pkl file for binned reweighting.

rule compute_metrics (relies on outputs of initialize_analysis, train_model and aggregate_bootstrap)

Computes the relevant metrics for a single trained model, and saves them to a .csv file. This is run once for each trained model for each sample/topology combination. The metrics computed are:

  • SWD between the reweighted original and target test sample for the reweighting parameter set and each metric set
  • p-values for each of those SWDs
  • Total $\chi^2$/dof between the reweighted original and target test sample in each of the same parameter spaces
  • p-values for each of those $\chi^2$/dof values
  • 1D $\chi^2$/dof between the reweighted original and target test sample for every parameter included in the config
  • p-values for each of those 1D $\chi^2$/dof values

These metrics are also computed between the original and target test samples themselves for comparison.

rule choose_models (relies on outputs of compute_metrics)

Local rule that finds the best-performing model for every model in a given sample/topology, according to the metric specified in the config. The hyperparameters and metrics for each model are saved in a single .json file. A symlink is created to each of the best-performing models in saved_models/{tag}/{sample}/{topology}/{model}. This rule is run once for each sample/topology combination.

rule make_plots (relies on the outputs of initialize_analysis and choose_models)

Generates plots showing the reweighting performance of each model for every parameter included in the config. For BDT-based reweighting models, plots of the training history are also generated. This rule is run once for each sample/topology combination, and the plots are saved as .pdf files.

rule all

Carries out the entire workflow, by requesting the outputs of the initialize_analysis checkpoint for every sample, and the outputs of choose_model and make_plots for every sample/topology combination. This rule is run once during the workflow.

rule count_all_topologies

Carries out the intialize_analysis checkpoint for every sample by requesting its output. This rule can be run before a dry run to see the full DAG, since the number of topologies passing the event count check is only known after the checkpoint has run.

About

An analysis framework to carry out BDT-based reweighting between neutrino event samples in the form of Nuisance FlatTrees.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages