Skip to content
Merged

Fixes #222

Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
112 commits
Select commit Hold shift + click to select a range
f80b261
Register cellposev4 in benchmark run scripts
dariarom94 Jul 19, 2026
1a2fa09
fix anndata version mismatch with txsim
dariarom94 Jul 19, 2026
82add80
add segger to workflow (test)
dariarom94 Jul 19, 2026
53e1728
duplicates when FOV stiching cleaned up
dariarom94 Jul 19, 2026
1186b7a
chunks issue atera
dariarom94 Jul 20, 2026
18644d7
segger update image
dariarom94 Jul 20, 2026
ecb302d
claude fix for segger
dariarom94 Jul 20, 2026
7d66898
Merge branch 'main' into fixes
dariarom94 Jul 20, 2026
d400ebe
atera version fix
dariarom94 Jul 20, 2026
64d7b4e
wf for the custom rnaseq scripts
dariarom94 Jul 20, 2026
3edfbf1
adjust the loader image name
dariarom94 Jul 20, 2026
cbd2f12
adjust the memory
dariarom94 Jul 20, 2026
184260e
troubleshootig edges
dariarom94 Jul 20, 2026
9fa9a33
Merge branch 'main' into fixes
dariarom94 Jul 20, 2026
36631c4
segger update
dariarom94 Jul 21, 2026
0626127
cell type label correction
dariarom94 Jul 21, 2026
3186435
fix boundaries
dariarom94 Jul 21, 2026
d8a7d93
Merge branch 'main' into fixes
dariarom94 Jul 21, 2026
3505718
OOM fixes
dariarom94 Jul 21, 2026
d6e110a
fix code
dariarom94 Jul 21, 2026
4660f26
RCTD
dariarom94 Jul 21, 2026
5abd651
segger to RAPIDS
dariarom94 Jul 21, 2026
fe2e90a
Merge branch 'main' into fixes
dariarom94 Jul 21, 2026
0b23474
fix rctd
dariarom94 Jul 22, 2026
196ff1f
segger debug (torchvision)
dariarom94 Jul 22, 2026
4be7bd4
Merge branch 'main' into fixes
dariarom94 Jul 22, 2026
b8d3d7b
save the xenium version
dariarom94 Jul 22, 2026
202ac49
add atera to datasets
dariarom94 Jul 22, 2026
14be8d0
Add gene efficiency correction as a separate pipeline stage (#183)
dariarom94 Jul 22, 2026
0cf0243
moscot to pca and segger troubleshooting
dariarom94 Jul 22, 2026
d7afb84
added fastreseg
dariarom94 Jul 23, 2026
f87a1d9
segger bug new fix
dariarom94 Jul 23, 2026
123e112
fastreseg to workflow
dariarom94 Jul 23, 2026
7549589
add fastreseg test
dariarom94 Jul 23, 2026
9fa9604
Merge branch 'main' into fixes
dariarom94 Jul 23, 2026
ff04467
optimized fastreseg build
dariarom94 Jul 23, 2026
7ffc514
Merge branch 'main' into fixes
dariarom94 Jul 23, 2026
a7404d8
add s3 paths
dariarom94 Jul 23, 2026
19e5f83
troubleshoot comseg/segger
dariarom94 Jul 24, 2026
aaca151
segger update
dariarom94 Jul 25, 2026
c9bdb91
data loader bug
dariarom94 Jul 25, 2026
1df9834
Merge branch 'main' into fixes
dariarom94 Jul 25, 2026
4467d32
rctd adjustment (raw counts)
dariarom94 Jul 26, 2026
58912e4
fix segger and comseg
dariarom94 Jul 26, 2026
0a5aa99
optimize cosmx
dariarom94 Jul 26, 2026
298e666
Merge branch 'main' into fixes
dariarom94 Jul 26, 2026
c921937
parameter test for cellpose4
dariarom94 Jul 26, 2026
57c2d79
add atera
dariarom94 Jul 26, 2026
cf67e09
add a test in pciseq and dynamic memory for bruker
dariarom94 Jul 27, 2026
9f71692
add test to vizgen data
dariarom94 Jul 28, 2026
d795f33
Merge branch 'main' into fixes
dariarom94 Jul 28, 2026
fefaadc
param sweep
dariarom94 Jul 28, 2026
4dc08d3
add params to segmentation
dariarom94 Jul 28, 2026
e9505f1
adjust segger mem
dariarom94 Jul 29, 2026
3a155be
update fastreseg to tacco
dariarom94 Jul 29, 2026
acdd6c7
Add annotation + expression-correction parameter sweeps (rctd, ssam, …
dariarom94 Jul 30, 2026
fa462b8
Add moscot + split parameter sweeps (annotation, expression correction)
dariarom94 Jul 30, 2026
fd37aee
singler: read par['celltype_key'] instead of hardcoding "cell_type"
dariarom94 Jul 30, 2026
3449bd0
fastreseg
dariarom94 Jul 30, 2026
331d219
Merge branch 'main' into fixes
dariarom94 Jul 30, 2026
5961d0a
adjust labels
dariarom94 Jul 30, 2026
857e16a
bruker nsclc
dariarom94 Jul 31, 2026
54f023e
adjust bruker nsclc loader
dariarom94 Jul 31, 2026
6e8b9ea
setup
dariarom94 Jul 31, 2026
4eec493
Merge branch 'main' into fixes
dariarom94 Jul 31, 2026
44ad7af
method correction
dariarom94 Jul 31, 2026
b351148
Merge branch 'main' into fixes
dariarom94 Jul 31, 2026
2b8177c
pin anndata
dariarom94 Aug 1, 2026
6c16707
mirror nsclc
dariarom94 Aug 1, 2026
5618fb8
sync the vizgen files
dariarom94 Aug 2, 2026
b68c100
adjust mem for allen brain
dariarom94 Aug 2, 2026
98dbc69
claude notes
dariarom94 Aug 2, 2026
fa23294
claude notes
dariarom94 Aug 2, 2026
ac9492c
fix mirror script
dariarom94 Aug 2, 2026
26d292e
Merge branch 'main' into fixes
dariarom94 Aug 2, 2026
1ee371a
adapt fastreseg requirements
dariarom94 Aug 3, 2026
c53016a
merscope kuppe script update
dariarom94 Aug 3, 2026
d479db6
fix nsclc loader
dariarom94 Aug 3, 2026
6c54220
adjust mem for new test resources
dariarom94 Aug 4, 2026
9494fc5
adjust the nsclc loader for test resources
dariarom94 Aug 4, 2026
a459afc
change processor to avoid spatialdata 0.8.0 bug
dariarom94 Aug 4, 2026
8f7b0c1
pin spatialdata version
dariarom94 Aug 4, 2026
f18296b
memory fix
dariarom94 Aug 5, 2026
fc96dcb
Merge branch 'main' into fixes
dariarom94 Aug 5, 2026
e75042f
subsampling code
dariarom94 Aug 5, 2026
594bb25
remove unexisting dataset
dariarom94 Aug 5, 2026
90207aa
fix data extraction bag
dariarom94 Aug 5, 2026
a94ec38
transcript assignment edits
dariarom94 Aug 5, 2026
905b166
segger: simplify transcript-assignment OOB handling to an edge clamp
dariarom94 Aug 5, 2026
b368d5e
modify proseg (param sweep)
dariarom94 Aug 5, 2026
0f57e2f
add scale0
dariarom94 Aug 6, 2026
5f09575
Merge branch 'main' into fixes
dariarom94 Aug 6, 2026
c64f7d2
fix code bug
dariarom94 Aug 6, 2026
7317d57
expand test dataset space
dariarom94 Aug 6, 2026
fe3ad75
Merge branch 'main' into fixes
dariarom94 Aug 6, 2026
a0a3d43
fix stardist params
dariarom94 Aug 6, 2026
6669752
process_dataset: opt-in tissue-centered crop (fixes ABCA whole-brain …
dariarom94 Aug 6, 2026
cc07559
memory adjust
dariarom94 Aug 6, 2026
fcda921
stardist tiling error
dariarom94 Aug 6, 2026
f26d173
Merge branch 'main' into fixes
dariarom94 Aug 6, 2026
4af3df3
fix labels
dariarom94 Aug 6, 2026
27ec16d
params for test datasets
dariarom94 Aug 6, 2026
76f4e80
add singler param sweep
dariarom94 Aug 6, 2026
6171f0c
yamls with params - wude
dariarom94 Aug 6, 2026
d6fec4d
tacco params introduced
dariarom94 Aug 7, 2026
325856f
configure sweeps
dariarom94 Aug 7, 2026
a0c96b2
more memory for sim. metrics
dariarom94 Aug 7, 2026
17cd092
param sweeps
dariarom94 Aug 7, 2026
37afb43
fix bug for allen brain
dariarom94 Aug 10, 2026
011301c
mapmycells params
dariarom94 Aug 12, 2026
25fd4c8
pareto params
dariarom94 Aug 12, 2026
6298b6e
Merge branch 'main' into fixes
dariarom94 Aug 12, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
34 changes: 34 additions & 0 deletions scripts/run_benchmark/param_sweep/mapmycells_params.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
# Parameter sweep for the mapmycells (Allen cell_type_mapper / MapMyCells) cell-type
# annotation method. Committed source of truth for run_test_mapmycells_nebius.sh (read
# from GitHub via a raw URL, since the Nebius compute env pulls the repo but cannot see
# the launch host's local files).
#
# Consumed by the run_benchmark workflow via the `method_parameters_yaml` setting
# (src/workflows/run_benchmark/main.nf). For every method the workflow builds:
# * one "default" variant using the `default:` args below, and
# * one extra variant per value in each `sweep:` list, with that ONE arg overridden.
# The benchmark varies a SINGLE parameter at a time (a "star" around the default, not
# a full grid), so total mapmycells variants = 1 default + sum(sweep list lengths) = 5.
#
# See src/methods_cell_type_annotation/mapmycells/NOTES.md ("Optimization / tuning") for
# the rationale.
#
# ⚠️ bootstrap_iteration / bootstrap_factor were JUST ADDED to config.vsh.yaml, so the
# build/main container must be rebuilt + pushed before this sweep runs (check-component).
parameters:
mapmycells:
# Baseline (fast anchor). NB the SHIPPED component default is (1, 1.0) = no
# bootstrapping. Here bootstrap_factor is pinned to the tool default 0.9 (not 1.0)
# ON PURPOSE: with factor 1.0 every bootstrap iteration is identical, which would
# make the bootstrap_iteration sweep below inert. 0.9 makes the axis meaningful.
default:
bootstrap_iteration: 1
bootstrap_factor: 0.9
sweep:
# bootstrap_iteration: how many times each query cell is re-mapped on a random
# 90% (bootstrap_factor) subset of markers before the majority vote. 1 (the
# default variant above) = single pass; the Allen tool default for bootstrapped
# correlation mapping is 100. Walks the shipped speed-tuned value up toward the
# tool default. 1 is omitted here (covered by the default variant); 100 is the
# slowest / most-robust endpoint.
bootstrap_iteration: [10, 25, 50, 100]
32 changes: 32 additions & 0 deletions scripts/run_benchmark/param_sweeps_full/baysor_params.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
# Parameter sweep for the baysor transcript assignment method - FOLLOW-UP (narrowed) sweep.
#
# AUTO-GENERATED - do not hand-edit; regenerate by rendering
# results/param_sweeps/Yaml_to_heatmap_transcript_assignment.qmd
# from the sweep results in ~/projects/txsim_results/param_sweep_tests/transcript_assignment/baysor
#
# Selection: the 3 parameter families that reach the Pareto front (max cells,
# max negative marker purity) in the most datasets, tie-broken by how many datasets
# they beat `default` in and then by Euclidean distance from `default`. Within each
# family: every value on the front in >=1 dataset, plus the highest-distance value,
# topped up by distance to 3 values where the family has that many.
#
# Consumed by the run_benchmark workflow via the `method_parameters_yaml` setting
# (src/workflows/run_benchmark/main.nf): one "default" variant from `default:`,
# plus one variant per value in each `sweep:` list, varying a SINGLE arg at a time.
# Total baysor variants = 1 default + 6 swept values.
parameters:
baysor:
default:
force_2d: true
min_molecules_per_cell: 50
scale: -1.0
scale_std: "25%"
n_clusters: 4
prior_segmentation_confidence: 0.8
sweep:
# on front in 2/3 datasets, beats default in 3; picked: 0.2 (front + max distance), 0.5 (distance top-up)
prior_segmentation_confidence: [0.2, 0.5]
# on front in 1/3 datasets, beats default in 3; picked: 10 (front + max distance), 25 (distance top-up)
min_molecules_per_cell: [10, 25]
# on front in 0/3 datasets, beats default in 1; picked: 8 (max distance), 6 (distance top-up)
n_clusters: [6, 8]
38 changes: 38 additions & 0 deletions scripts/run_benchmark/param_sweeps_full/clustermap_params.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
# Parameter sweep for the clustermap transcript assignment method - FOLLOW-UP (narrowed) sweep.
#
# AUTO-GENERATED - do not hand-edit; regenerate by rendering
# results/param_sweeps/Yaml_to_heatmap_transcript_assignment.qmd
# from the sweep results in ~/projects/txsim_results/param_sweep_tests/transcript_assignment/clustermap
#
# Selection: the 3 parameter families that reach the Pareto front (max cells,
# max negative marker purity) in the most datasets, tie-broken by how many datasets
# they beat `default` in and then by Euclidean distance from `default`. Within each
# family: every value on the front in >=1 dataset, plus the highest-distance value,
# topped up by distance to 3 values where the family has that many.
#
# Consumed by the run_benchmark workflow via the `method_parameters_yaml` setting
# (src/workflows/run_benchmark/main.nf): one "default" variant from `default:`,
# plus one variant per value in each `sweep:` list, varying a SINGLE arg at a time.
# Total clustermap variants = 1 default + 8 swept values.
parameters:
clustermap:
default:
window_size: 700
xy_radius: 40
z_radius: 0
fast_preprocess: false
gauss_blur: true
sigma: 1.0
pct_filter: 0.0
LOF: false
contamination: 0
min_spot_per_cell: 5
dapi_grid_interval: 5
cell_num_threshold: 0.1
sweep:
# on front in 0/3 datasets, beats default in 2; picked: 60 (max distance), 30 (distance top-up), 20 (distance top-up)
xy_radius: [20, 30, 60]
# on front in 0/3 datasets, beats default in 2; picked: 0.01 (max distance), 0.05 (distance top-up), 0.2 (distance top-up)
cell_num_threshold: [0.01, 0.05, 0.2]
# on front in 0/3 datasets, beats default in 2; picked: 0.05 (max distance), 0.1 (distance top-up)
pct_filter: [0.05, 0.1]
32 changes: 32 additions & 0 deletions scripts/run_benchmark/param_sweeps_full/comseg_params.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
# Parameter sweep for the comseg transcript assignment method - FOLLOW-UP (narrowed) sweep.
#
# AUTO-GENERATED - do not hand-edit; regenerate by rendering
# results/param_sweeps/Yaml_to_heatmap_transcript_assignment.qmd
# from the sweep results in /Users/daria.romanovskaia/projects/txsim_results/param_sweep_tests/transcript_assignment/comseg
#
# Selection: the 3 parameter families that reach the Pareto front (max cells,
# max negative marker purity) in the most datasets, tie-broken by how many datasets
# they beat `default` in and then by Euclidean distance from `default`. Within each
# family: every value on the front in >=1 dataset, plus the highest-distance value,
# topped up by distance to 3 values where the family has that many.
#
# Consumed by the run_benchmark workflow via the `method_parameters_yaml` setting
# (src/workflows/run_benchmark/main.nf): one "default" variant from `default:`,
# plus one variant per value in each `sweep:` list, varying a SINGLE arg at a time.
# Total comseg variants = 1 default + 7 swept values.
parameters:
comseg:
default:
mean_cell_diameter: 15.0
max_cell_radius: 25.0
alpha: 0.5
min_rna_per_cell: 5
norm_vector: false
allow_disconnected_polygon: true
sweep:
# on front in 1/3 datasets, beats default in 1; picked: 10.0 (front + max distance), 20.0 (on front)
mean_cell_diameter: [10.0, 20.0]
# on front in 1/3 datasets, beats default in 0; picked: 1.0 (front + max distance), 0.75 (on front), 0.25 (on front)
alpha: [0.25, 0.75, 1.0]
# on front in 0/3 datasets, beats default in 0; picked: 20 (max distance), 10 (distance top-up)
min_rna_per_cell: [10, 20]
30 changes: 30 additions & 0 deletions scripts/run_benchmark/param_sweeps_full/fastreseg_params.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
# Parameter sweep for the fastreseg transcript assignment method - FOLLOW-UP (narrowed) sweep.
#
# AUTO-GENERATED - do not hand-edit; regenerate by rendering
# results/param_sweeps/Yaml_to_heatmap_transcript_assignment.qmd
# from the sweep results in /Users/daria.romanovskaia/projects/txsim_results/param_sweep_tests/transcript_assignment/fastreseg
#
# Selection: the 3 parameter families that reach the Pareto front (max cells,
# max negative marker purity) in the most datasets, tie-broken by how many datasets
# they beat `default` in and then by Euclidean distance from `default`. Within each
# family: every value on the front in >=1 dataset, plus the highest-distance value,
# topped up by distance to 3 values where the family has that many.
#
# Consumed by the run_benchmark workflow via the `method_parameters_yaml` setting
# (src/workflows/run_benchmark/main.nf): one "default" variant from `default:`,
# plus one variant per value in each `sweep:` list, varying a SINGLE arg at a time.
# Total fastreseg variants = 1 default + 0 swept values.
parameters:
fastreseg:
default:
molecular_distance_cutoff: 2.7
flagCell_lrtest_cutoff: 5
svmClass_score_cutoff: -2
cutoff_spatialMerge: 0.5
sweep:
# on front in 1/3 datasets, beats default in 0; picked:
cutoff_spatialMerge: []
# on front in 1/3 datasets, beats default in 0; picked:
flagCell_lrtest_cutoff: []
# on front in 1/3 datasets, beats default in 0; picked:
molecular_distance_cutoff: []
32 changes: 32 additions & 0 deletions scripts/run_benchmark/param_sweeps_full/moscot_params.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
# Parameter sweep for the moscot cell type annotation method - FOLLOW-UP (narrowed) sweep.
#
# AUTO-GENERATED - do not hand-edit; regenerate by rendering
# results/param_sweeps/Yaml_to_heatmap_cell_type_annotation.qmd
# from the sweep results in /Users/daria.romanovskaia/projects/txsim_results/param_sweep_tests/cell_type_annotation/moscot
#
# Selection: the 3 parameter families that reach the Pareto front (max per-cell-type
# co-expression similarity, max negative marker purity) in the most datasets, tie-broken by
# how many datasets they beat `default` in and then by Euclidean distance from `default`.
# Within each family: every value on the front in >=1 dataset, plus the highest-distance
# value, topped up by distance to 3 values where the family has that many.
#
# Consumed by the run_benchmark workflow via the `method_parameters_yaml` setting
# (src/workflows/run_benchmark/main.nf): one "default" variant from `default:`,
# plus one variant per value in each `sweep:` list, varying a SINGLE arg at a time.
# Total moscot variants = 1 default + 7 swept values.
parameters:
moscot:
default:
alpha: 0.8
epsilon: 0.01
tau: 0.3
rank: 500
batch_size: 1024
mapping_mode: "max"
sweep:
# on front in 0/3 datasets, beats default in 2; picked: sum (max distance)
mapping_mode: ["sum"]
# on front in 0/3 datasets, beats default in 2; picked: 0.1 (max distance), 0.001 (distance top-up), 0.05 (distance top-up)
epsilon: [0.001, 0.05, 0.1]
# on front in 0/3 datasets, beats default in 1; picked: 0.5 (max distance), 0.7 (distance top-up), 0.9 (distance top-up)
alpha: [0.5, 0.7, 0.9]
31 changes: 31 additions & 0 deletions scripts/run_benchmark/param_sweeps_full/pciseq_params.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
# Parameter sweep for the pciseq transcript assignment method - FOLLOW-UP (narrowed) sweep.
#
# AUTO-GENERATED - do not hand-edit; regenerate by rendering
# results/param_sweeps/Yaml_to_heatmap_transcript_assignment.qmd
# from the sweep results in /Users/daria.romanovskaia/projects/txsim_results/param_sweep_tests/transcript_assignment/pciseq
#
# Selection: the 3 parameter families that reach the Pareto front (max cells,
# max negative marker purity) in the most datasets, tie-broken by how many datasets
# they beat `default` in and then by Euclidean distance from `default`. Within each
# family: every value on the front in >=1 dataset, plus the highest-distance value,
# topped up by distance to 3 values where the family has that many.
#
# Consumed by the run_benchmark workflow via the `method_parameters_yaml` setting
# (src/workflows/run_benchmark/main.nf): one "default" variant from `default:`,
# plus one variant per value in each `sweep:` list, varying a SINGLE arg at a time.
# Total pciseq variants = 1 default + 8 swept values.
parameters:
pciseq:
default:
InsideCellBonus: 2
MisreadDensity: 1e-05
nNeighbors: 3
Inefficiency: 0.2
rGene: 20
sweep:
# on front in 3/3 datasets, beats default in 2; picked: 0.001 (front + max distance), 1.0E-4 (on front), 1.0E-6 (distance top-up)
MisreadDensity: [1.0E-6, 1.0E-4, 0.001]
# on front in 2/3 datasets, beats default in 2; picked: 0 (front + max distance), 1 (on front), 6 (distance top-up)
InsideCellBonus: [0, 1, 6]
# on front in 1/3 datasets, beats default in 2; picked: 40 (front + max distance), 10 (distance top-up)
rGene: [10, 40]
30 changes: 30 additions & 0 deletions scripts/run_benchmark/param_sweeps_full/proseg_params.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
# Parameter sweep for the proseg transcript assignment method - FOLLOW-UP (narrowed) sweep.
#
# AUTO-GENERATED - do not hand-edit; regenerate by rendering
# results/param_sweeps/Yaml_to_heatmap_transcript_assignment.qmd
# from the sweep results in /Users/daria.romanovskaia/projects/txsim_results/param_sweep_tests/transcript_assignment/proseg
#
# Selection: the 3 parameter families that reach the Pareto front (max cells,
# max negative marker purity) in the most datasets, tie-broken by how many datasets
# they beat `default` in and then by Euclidean distance from `default`. Within each
# family: every value on the front in >=1 dataset, plus the highest-distance value,
# topped up by distance to 3 values where the family has that many.
#
# Consumed by the run_benchmark workflow via the `method_parameters_yaml` setting
# (src/workflows/run_benchmark/main.nf): one "default" variant from `default:`,
# plus one variant per value in each `sweep:` list, varying a SINGLE arg at a time.
# Total proseg variants = 1 default + 7 swept values.
parameters:
proseg:
default:
cell_compactness: 0.04
nuclear_reassignment_prob: 0.2
diffusion_probability: 0.2
ncomponents: 10
sweep:
# on front in 1/3 datasets, beats default in 3; picked: 0.03 (front + max distance), 0.06 (distance top-up), 0.02 (distance top-up)
cell_compactness: [0.02, 0.03, 0.06]
# on front in 1/3 datasets, beats default in 2; picked: 0.5 (front + max distance), 0.05 (distance top-up)
nuclear_reassignment_prob: [0.05, 0.5]
# on front in 1/3 datasets, beats default in 1; picked: 15 (on front), 5 (max distance)
ncomponents: [5, 15]
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
# Parameter sweep for the resolvi_correction expression correction method - FOLLOW-UP (narrowed) sweep.
#
# AUTO-GENERATED - do not hand-edit; regenerate by rendering
# results/param_sweeps/Yaml_to_heatmap_expression_correction.qmd
# from the sweep results in /Users/daria.romanovskaia/projects/txsim_results/param_sweep_tests/expression_correction/resolvi_correction
#
# Selection: the 3 parameter families that reach the Pareto front (max per-cell-type
# co-expression similarity, max negative marker purity) in the most datasets, tie-broken by
# how many datasets they beat `default` in and then by Euclidean distance from `default`.
# Within each family: every value on the front in >=1 dataset, plus the highest-distance
# value, topped up by distance to 3 values where the family has that many.
#
# Consumed by the run_benchmark workflow via the `method_parameters_yaml` setting
# (src/workflows/run_benchmark/main.nf): one "default" variant from `default:`,
# plus one variant per value in each `sweep:` list, varying a SINGLE arg at a time.
# Total resolvi_correction variants = 1 default + 4 swept values.
parameters:
resolvi_correction:
default:
celltype_key: "cell_type"
n_hidden: 32
encode_covariates: false
downsample_counts: true
sweep:
# on front in 0/3 datasets, beats default in 3; picked: 128 (max distance), 64 (distance top-up)
n_hidden: [64, 128]
# on front in 0/3 datasets, beats default in 2; picked: false (max distance)
downsample_counts: [false]
# on front in 0/3 datasets, beats default in 1; picked: true (max distance)
encode_covariates: [true]
Original file line number Diff line number Diff line change
Expand Up @@ -28,10 +28,10 @@ set -e
resources_s3=/scratch/task_ist_preprocessing/datasets
publish_dir="/scratch/results/runs/$(date +%Y-%m-%d_%H-%M-%S)_cellposev4_full_sweep"

# The sweep lives in a committed file, read from GitHub at runtime. $params_branch
# defaults to the branch you are on; the file must be pushed there on GitHub.
# The sweep lives in a committed file, read from GitHub at runtime from the `main`
# branch; the file must be committed AND pushed to main on GitHub before launching.
params_repo="openproblems-bio/task_ist_preprocessing"
params_branch="$(git rev-parse --abbrev-ref HEAD)"
params_branch="main"
params_url="https://raw.githubusercontent.com/${params_repo}/${params_branch}/scripts/run_benchmark/param_sweeps_full/cellposev4_params.yaml"

cat > /tmp/params_settings_cellposev4_full.yaml << HERE
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -28,10 +28,10 @@ set -e
resources_s3=/scratch/task_ist_preprocessing/datasets
publish_dir="/scratch/results/runs/$(date +%Y-%m-%d_%H-%M-%S)_stardist_full_sweep"

# The sweep lives in a committed file, read from GitHub at runtime. $params_branch
# defaults to the branch you are on; the file must be pushed there on GitHub.
# The sweep lives in a committed file, read from GitHub at runtime from the `main`
# branch; the file must be committed AND pushed to main on GitHub before launching.
params_repo="openproblems-bio/task_ist_preprocessing"
params_branch="$(git rev-parse --abbrev-ref HEAD)"
params_branch="main"
params_url="https://raw.githubusercontent.com/${params_repo}/${params_branch}/scripts/run_benchmark/param_sweeps_full/stardist_params.yaml"

cat > /tmp/params_settings_stardist_full.yaml << HERE
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -28,10 +28,10 @@ set -e
resources_s3=/scratch/task_ist_preprocessing/datasets
publish_dir="/scratch/results/runs/$(date +%Y-%m-%d_%H-%M-%S)_watershed_full_sweep"

# The sweep lives in a committed file, read from GitHub at runtime. $params_branch
# defaults to the branch you are on; the file must be pushed there on GitHub.
# The sweep lives in a committed file, read from GitHub at runtime from the `main`
# branch; the file must be committed AND pushed to main on GitHub before launching.
params_repo="openproblems-bio/task_ist_preprocessing"
params_branch="$(git rev-parse --abbrev-ref HEAD)"
params_branch="main"
params_url="https://raw.githubusercontent.com/${params_repo}/${params_branch}/scripts/run_benchmark/param_sweeps_full/watershed_params.yaml"

cat > /tmp/params_settings_watershed_full.yaml << HERE
Expand Down
30 changes: 30 additions & 0 deletions scripts/run_benchmark/param_sweeps_full/segger_params.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
# Parameter sweep for the segger transcript assignment method - FOLLOW-UP (narrowed) sweep.
#
# AUTO-GENERATED - do not hand-edit; regenerate by rendering
# results/param_sweeps/Yaml_to_heatmap_transcript_assignment.qmd
# from the sweep results in /Users/daria.romanovskaia/projects/txsim_results/param_sweep_tests/transcript_assignment/segger
#
# Selection: the 3 parameter families that reach the Pareto front (max cells,
# max negative marker purity) in the most datasets, tie-broken by how many datasets
# they beat `default` in and then by Euclidean distance from `default`. Within each
# family: every value on the front in >=1 dataset, plus the highest-distance value,
# topped up by distance to 3 values where the family has that many.
#
# Consumed by the run_benchmark workflow via the `method_parameters_yaml` setting
# (src/workflows/run_benchmark/main.nf): one "default" variant from `default:`,
# plus one variant per value in each `sweep:` list, varying a SINGLE arg at a time.
# Total segger variants = 1 default + 6 swept values.
parameters:
segger:
default:
n_epochs: 20
prediction_graph_buffer_ratio: 0.05
prediction_mode: "cell"
node_representation_dim: 128
sweep:
# on front in 1/3 datasets, beats default in 2; picked: 0.1 (on front), 0.25 (max distance), 0.5 (distance top-up)
prediction_graph_buffer_ratio: [0.1, 0.25, 0.5]
# on front in 0/3 datasets, beats default in 2; picked: 60 (max distance), 40 (distance top-up)
n_epochs: [40, 60]
# on front in 0/3 datasets, beats default in 1; picked: nucleus (max distance)
prediction_mode: ["nucleus"]
Loading
Loading