Resolve atlases by name to their own files; add roi_signature - #3
Merged
Merged
Conversation
…s categories
An allen26 study was being analysed with allen32's partition. Several atlases
share one directory (allen32, allen26 and allen64 all live in allen/), and every
lookup went "the file with the fixed name in the atlas directory":
roi_categories.yaml, roi_mapping.json and allen_labels.nii.gz are allen32's.
Measured on the FORGE treatment run (allen26, Monte Carlo ROI operators):
- the R region tier replaced the study's categories with that file, so
Frontal-Anterior and Olfactory matched no parcel and vanished from every
region-level table, and Deep Subcortical was built from 4 of its 8 parcels
- the ROI mosaics drew allen32's label volume, so the six merged parcels had
no voxels and rendered blank (6/26)
Nothing warned. ROI-level results were never affected.
An atlas is now resolved by NAME through source-localization's registry.yaml,
the place an atlas is actually defined, into an AtlasSpec carrying its own
labels, mapping, categories, brain volume and brain mask. base.py stores the
spec where _atlas_dir lived, and every atlas function accepts a spec wherever
it accepted a directory, so the existing call sites pass it through unchanged.
An atlas that is not registered names its files under atlas_files:. A name that
cannot be resolved raises instead of falling back to whatever shares a
directory, which is the fallback that caused this.
R now receives the EFFECTIVE categories (the study's map, a profile's narrowing,
or the atlas default) through BaseAnalysis._r_config_data(), and a single
resolve_roi_categories() in stats_utils.R, replacing seven copies of the
overwrite block, lets them win. The --roi-categories file is only a fallback.
The 10x voxel convention is now read from the NIfTI header, as
source-localization does. The filename rule treated Antwerp's true-unit label
file as inflated and shrank its affine 10x on the default-affine path; this
moves Antwerp studies' vertex ROI labels and mosaics to their correct places.
Tests: atlas resolution for every registered atlas (and the allen26 regression
itself), the header rule, explicit atlas_files, the effective categories handed
to R, and an R-side test that the study's categories beat a conflicting file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RibM8Zep2YEjjUbc2LLkgj
vertex_signature was the only source-side decoding module, so a study with ROI output alone had no source half of the source-vs-sensor comparison, and Monte Carlo ROI operators never build the vertex estimate it needs. roi_signature runs the same classifiers, cross-validation and permutation test as electrode_signature on per-parcel relative band power, with the identical feature estimator, so where both run on the same epochs their accuracies are like for like. It subclasses electrode_signature, whose table, pickle and figure names now come from the module name rather than a hard-coded prefix. electrode_signature used to find its source partner by globbing the whole results tree for vertex_signature_results.csv and taking the first hit. In the FORGE treatment study that was a stale vertex table from the previous analysis version, which would have been merged against the new run's sensor results as if it were its source side. It now looks only in its own paradigm, preferring roi_signature, and the comparison table records which source module it used. When the sensor signature ran in another paradigm (it reads the raw recordings, so one run can serve every source arm), roi_signature takes sensor_paradigm: and renders the comparison itself. Tests: a full roi_signature lifecycle including figures regenerated from disk, and the in-paradigm lookup (a stale table elsewhere is never used). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RibM8Zep2YEjjUbc2LLkgj
A signature run fits a tiny model -- tens of subjects by tens of features -- LOOCV x (1 + n_permutations) times. With the default multithreaded BLAS each of those fits paid for starting and synchronising a thread pool, and that cost dominated. Measured on the FORGE treatment electrode features (35 subjects x 30 electrodes, 21 LOOCV passes): band default threads one thread accuracy Alpha 127.8 s 1.4 s 0.486 in both Low Gamma 105.4 s 1.7 s 0.600 in both That is why one electrode_signature run was heading for several days, and why v1's logistic fits averaged ~16 minutes with a 15-hour outlier. run_signature is now wrapped in threadpoolctl.threadpool_limits(1). threadpoolctl comes with scikit-learn, which the signature modules already need, and the wrapper falls back to a plain call if it is missing. Results do not change: the tests check that the limit is really applied inside a fit and that the limited and unlimited calls agree. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RibM8Zep2YEjjUbc2LLkgj
The entry (and aa6ce30's message) said reading the 10x convention from the header moves Antwerp studies' vertex ROI labels and mosaics to their correct positions. Measured, it does neither: - load_vertex_roi_labels, the vertex path that used the corrected affine, has never produced a label. It reads the mapping file's top-level keys ("atlas_name", "n_rois", ...) as label ids and raises on every atlas; vertex_network catches that and falls back to spatial node labels. - the mosaics use the affine only for mm axis ranges and slice labels, and the analysis modules never draw them on Antwerp. No statistic changes. ROI extraction and cluster/NBS region labels use the raw affine, which was always right. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RibM8Zep2YEjjUbc2LLkgj
2 of 3 tasks
alexedmon1
added a commit
that referenced
this pull request
Sep 10, 2026
Bumps the package metadata to 0.7.0 and dates the CHANGELOG section that collected everything since v0.6.0: the September audit remediation, per-atlas resolution (#3), roi_signature, and single-threaded signature fits. source_analytics.__version__ comes from git describe, so the tag on this commit is what outputs will be stamped with; the metadata bump keeps an installed (non-git) copy reporting the same version. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RibM8Zep2YEjjUbc2LLkgj
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
pipeline.atlasis looked up in source-localization'sregistry.yaml(where an atlas is actually defined) into anAtlasSpec: labels, mapping, categories, brain volume, brain mask. Studies can name an unregistered atlas's files underatlas_files:. An unresolvable name raises instead of guessing.roi_categories(the study's map, a profile's narrowing, or the atlas default) viaBaseAnalysis._r_config_data(), and oneresolve_roi_categories()instats_utils.Rlets them win. It replaces seven copies of a block that overwrote them with the atlas-directory file.roi_signature: ROI-level decoding with the same classifiers, CV and feature estimator aselectrode_signature, so it works on any ROI output, including Monte Carlo operators that never build a vertex estimate.electrode_signaturenow pairs only with a source signature in its own paradigm.Why
Found on the FORGE treatment re-run (allen26, Monte Carlo ROI operators). allen32, allen26 and allen64 share
allen/, and every lookup took the fixed-name file in that directory, which is allen32's. Nothing warned about any of the following:electrode_signaturecomparisonROI-level results were never affected.
Behaviour changes
Atlas_3DRoisLeftRight.Labels.niihas stored true units since source-localization 2026-03-12 but was still shrunk 10× on the default-affine path. No statistic changes, since ROI extraction and cluster/NBS region labels use the raw affine. The only effect is the mm axis ranges and slice labels ofplot_brain_roi_mosaic/plot_brain_roiwhen drawn on Antwerp, which no analysis module does.aa6ce30's message says this moves Antwerp studies' vertex ROI labels and mosaics. It doesn't.load_vertex_roi_labelshas never run: it reads the mapping file's top-level keys as label ids and raises on every atlas, andvertex_networkswallows the error and uses spatial labels. That is a separate bug, left for the vertex split.signature_source_vs_sensor.csvgains asource_modulecolumn.find_atlas_diris kept for backward compatibility.Compatibility
v0.4.0(own worktree) and is unaffected.autifonyinstalls this repo as an editable path, so it picks up these changes as soon asmainupdates. For its Antwerp data that is cosmetic only; no statistic changes.Test plan
PYTHONPATH=src pytest tests: 219 passed (199 existing + 20 new)--roi-categoriesfile no longer beats the study config; the file still serves as a fallbackroi_psd,roi_aperiodic,roi_connectivity,roi_cross_freq,roi_directedon FORGE treatment (cartesian_mc,shell_mc) and confirm 10/10 categoriesPerformance (added in this PR)
run_signaturenow runs with BLAS/OpenMP limited to one thread (threadpoolctl). Each fit is tiny, and the thread-pool overhead was dominating the runtime:Results are unchanged (tested). A signature module that took days now takes about an hour.
🤖 Generated with Claude Code
https://claude.ai/code/session_01RibM8Zep2YEjjUbc2LLkgj