Thank you for contributing. This document covers development setup, code standards, testing, and how to add new feature blocks or model types.
git clone https://github.com/your-org/fungal-classifier.git
cd fungal-classifier
conda env create -f environment.yml
conda activate fungal-classifier
pip install -e ".[dev]"Install pre-commit hooks:
pip install pre-commit
pre-commit install- Python 3.11+
- Formatter:
ruff format(enforced by pre-commit) - Linter:
ruff check— no warnings allowed in CI - Type hints: required for all public functions
- Docstrings: NumPy-style for all public classes and functions
- Line length: 100 characters
Run checks locally:
make lint
make format# Fast unit tests (no slow marker)
make test-fast
# All tests including integration
make test
# Single test file
pytest tests/test_features.py -vAll PRs must pass the full test suite. Integration tests (@pytest.mark.slow) run on CI but not on every local commit.
Test coverage target: >80% for fungal_classifier/ core modules.
- Create
fungal_classifier/features/{block_name}.py - Implement a
build_{block_name}_matrix(paths, ...) -> pd.DataFramefunction - Add it to
fungal_classifier/features/__init__.py - Add the block handler to
scripts/01_build_features.py - Add default config section to
configs/default.yaml - Write unit tests in
tests/test_features.py - Add a section to
docs/feature_engineering.md
Convention: All feature matrices must:
- Have
genome_idas the index name - Return
np.float32dtype - Accept a
min_genome_freqparameter and filter rare features - Log their shape at INFO level
- Create
fungal_classifier/models/{model_name}.py - Implement sklearn-compatible
fit(X, y)andpredict(X)/predict_proba(X) - Add to
fungal_classifier/models/__init__.py - Add a
--model-typeoption toscripts/02_train.pyif appropriate - Write tests in
tests/test_models.py
- Branch from
main:git checkout -b feature/my-feature - Write code and tests
- Run
make test && make lint - Update docs if you changed behaviour or added features
- Open a PR with a clear description of what changed and why
- A maintainer will review within a week
Please include:
- Python version and OS
- Conda environment (
conda list) - Minimal reproducible example
- Full error traceback
Planned additions:
-
features/secretome.py— signal peptide and GPI anchor features -
features/synteny.py— gene cluster synteny features -
models/graph_model.py— GNN over genome-scale metabolic networks - Multi-label classification (genomes with multiple ecological roles)
- Active learning module for prioritising new genome sequencing
- Integration with NCBI datasets API for auto-downloading annotation