Summarize command - #64
Merged
Merged
Conversation
The -e flag is not needed.
Get ready for better text Ensure we mark duplicates and only take the best sample Ensure we exclude samples not in master metadata
Make sample_type mandatory. We need it for summary analysis and it is best if people just include it when creating the sample sheet. It should not be much extra work. We maybe want to make it optional if we can derive it from the sample name, to be discussed.
For now just don't group them by alt alleles. We might have a different set of alt alleles when we call the same mutation in different experiments, as we might have a triallelic side. As we in the end group by amino acid change, this leads to problems. We need to better discuss how to hanle csq calling correctly for multiple changes.
This helps to only look at the sequenced samples
Green and red should be more intuitve what they mean
It makes sense that we don't report contamination if we have low coverage, as this could actually just be caused by low coverage. But if we are over the abs threshold, we can be sure, it's contamination
Exclude columns with to many entries and which are numeric. Maybe we want to filter more in the future, or have a way to provide a list
They are big and we don't need them for the analysis
berndbohmeier
requested review from
danieljbridges
and
a balanced review from Copilot
August 28, 2026 13:51
Contributor
There was a problem hiding this comment.
🔵 Needs a closer look
Settings loading, environment dependencies, and destructive output replacement introduce concrete runtime and reliability failures.
Review details
Suppressed comments (5)
Previously missed (3) — in code that hasn't changed since the last review.
environments/dev.yml:20
- The development environment omits the newly required
pydanticdependency, while CI installs the package with--no-deps. A clean development/CI environment can therefore fail when the summarize implementation importssummary_settings; keep this environment aligned with the runtime and conda package dependency lists.
- statsmodels
- seaborn
src/nomadic/summarize/main.py:279
- This output name contains a comma instead of the separator used by every other VCF artifact, producing
summary.variants.annotated,vcf.gz. Use the standard.annotated.vcf.gzname so downstream tooling and users can identify it consistently.
src/nomadic/util/summary_settings.py:14 - In Pydantic v2, an
Optional[...]annotation still defines a required field unless it has a default. As written,map: {shape_name_key: ...}(or a map block that supplies only one display option) raises a validation error even though both values are modeled and consumed as optional. Give both fieldsNonedefaults.
This issue also appears on line 25 of the same file.
src/nomadic/util/summary_settings.py:27
- An empty or comment-only settings file makes
yaml.safe_loadreturnNone, so unpacking it withSettings(**data)raisesTypeError. Since every top-level setting has a default, such a file should load as the default settings object.
src/nomadic/summarize/main.py:122 - The command deletes the existing valid summary before validating the experiments or completing the replacement. Any later input, reference, bcftools, or analysis failure therefore destroys the last usable summary and leaves only a partial output. Build in a sibling temporary directory and atomically replace the old summary only after successful completion.
- Files reviewed: 202/213 changed files
- Comments generated: 0 new
- Review effort level: Balanced
We don't want to load pandas in cli modules, so that startup is not slow. The test that checks this was failing
I just missed, that there was a file for exceptions already
This is because nt changes inside genes are more trustworthy. We add it so that users can filter them
Transposed to make it easier to read in throughput file
Contributor
There was a problem hiding this comment.
🟡 Changes recommended
Several correctness issues currently break mapping-only realtime runs, metadata remapping, empty-variant summaries, and documented configuration paths.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review details
- Files reviewed: 211/226 changed files
- Comments generated: 8
- Review effort level: Balanced
Comment on lines
+183
to
+197
| sample_names = df["barcode"].str.replace( | ||
| r"\ ", | ||
| " ", | ||
| regex=False, | ||
| ) # reverse escaping of spaces | ||
|
|
||
| sample_parts = sample_names.str.split( | ||
| SAMPLE_SEPERATOR, | ||
| n=1, | ||
| expand=True, | ||
| regex=False, | ||
| ) | ||
|
|
||
| df.insert(0, "expt_name", sample_parts[0]) | ||
| df["barcode"] = sample_parts[1] |
Comment on lines
+297
to
+300
| with open(settings_path, "r") as f: | ||
| settings = json.load(f) | ||
| caller = settings["caller"] | ||
| reference_name = str(settings["reference_name"]) |
Comment on lines
+13
to
+14
| center: Optional[tuple[float, float]] | ||
| zoom_level: Optional[int] |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Edit: This is now ready.
UI:
answers which samples are done, which are failing. Gives an idea on how much the sequencing is already done
gives an idea about how good the sequencing is working. Which experiments have good data, which have high coverage, which are contaminated, etc.
Two ways to show prevalence, grouped by columns in the metadata file