This directory contains the main Python code for the Coffea4bees project, built on top of the barista framework for high-energy physics analyses.
This folder provides all analysis, skimming, machine learning, plotting, and workflow automation tools for 4b physics analyses. It is the main entry point for running and developing new analysis features.
To run analysis or skimming from this directory:
- Clone the
baristarepository and then this repository.
git clone ssh://git@gitlab.cern.ch:7999/cms-cmu/barista.git
cd barista
git clone ssh://git@gitlab.cern.ch:7999/cms-cmu/coffea4bees.git- Run the main analysis script:
python runner.py --help- Explore subfolders for specialized tasks (skimming, ML, plotting, etc.).
- analysis/: Main analysis processors, helpers, metadata, tools, and tests for physics analysis.
- analysis_dask/: Dask-based analysis modules and configuration for distributed processing.
- archive/: Archived datasets, plots, and skims from previous runs.
- classifier/: Machine learning models, utilities, and scripts for classification tasks.
- examples/: Example scripts for analysis and meta-data rescue.
- jet_clustering/: Jet clustering algorithms, studies, and synthetic data generation.
- metadata/: Central metadata configurations including unified datasets, cross-sections, triggers, and friend trees.
- plots/: Plotting scripts, styles, and metadata for visualizations.
- scripts/: Shell scripts for running, testing, and automating analysis jobs.
- skimmer/: Processors for filtering NanoAOD files and saving skimmed (picoAOD) files.
- stats_analysis/: Statistical analysis scripts and Combine framework integration.
- workflows/: Snakemake workflows and rules for automating analysis pipelines.
For more details about each component, refer to the README.md file in the respective folder.
The metadata/ directory has been reorganized into a centralized structure to support seamless, unified execution across different runs:
Contains all data and MC dataset definition YAML files (e.g. TT.yml, GluGluToHHTo4B.yml, data.yml).
- Unified Cross-Sections: To prevent key collisions and support running Run 2 and Run 3 analysis scripts under the same dataset keys, all cross-sections (
xs) are defined as run-dependent dictionaries:If a dataset is specific to only one Run, the other Run's cross-section is set to a placeholderxs: Run2: <run2_cross_section_value> Run3: <run3_cross_section_value>
1. - Dataset Archive (
metadata/datasets/archive/): Holds older dataset definitions and versions (e.g.HIG-24-010,Run3_archive).
Contains friend tree configurations (friends_HH4b.yml, friends_empty.yml) and their corresponding active JSON lookups:
- Active JSON files for trigger weights and classifier friend trees (e.g.
trigweights_2024_v1p2.json,data_SvBfriend.json, etc.) live directly inmetadata/friends/. - Unused/legacy friend trees are placed in
metadata/friends/archive/.
This package supports running workflows on REANA. The REANA workflow is triggered manually via the GitLab CI pipeline or automatically every Saturday.
Workflow outputs (plots, files) are available at https://plotsalgomez.webtest.cern.ch/HH4b/reana/.
Each output folder is named with the REANA job execution date and the corresponding Git commit hash. Folders are only copied here if the REANA job completes successfully.