Skip to content

Resolve atlases by name to their own files; add roi_signature - #3

Merged
alexedmon1 merged 4 commits into
mainfrom
feat/atlas-registry-roi-signature
Sep 10, 2026
Merged

alexedmon1 merged 4 commits into
mainfrom
feat/atlas-registry-roi-signature

Conversation

@alexedmon1

@alexedmon1 alexedmon1 commented Sep 10, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • Atlases resolve by name to their own files. pipeline.atlas is looked up in source-localization's registry.yaml (where an atlas is actually defined) into an AtlasSpec: labels, mapping, categories, brain volume, brain mask. Studies can name an unregistered atlas's files under atlas_files:. An unresolvable name raises instead of guessing.
  • R keeps the study's categories. Python hands R the effective roi_categories (the study's map, a profile's narrowing, or the atlas default) via BaseAnalysis._r_config_data(), and one resolve_roi_categories() in stats_utils.R lets them win. It replaces seven copies of a block that overwrote them with the atlas-directory file.
  • The 10× voxel convention is read from the NIfTI header, as source-localization does. The filename rule was wrong for Antwerp.
  • New roi_signature: ROI-level decoding with the same classifiers, CV and feature estimator as electrode_signature, so it works on any ROI output, including Monte Carlo operators that never build a vertex estimate. electrode_signature now pairs only with a source signature in its own paradigm.

Why

Found on the FORGE treatment re-run (allen26, Monte Carlo ROI operators). allen32, allen26 and allen64 share allen/, and every lookup took the fixed-name file in that directory, which is allen32's. Nothing warned about any of the following:

Symptom Measured
Region-level tables (psd, aperiodic, PAC, …) 8 of 10 categories: Frontal-Anterior and Olfactory matched no parcel and vanished
Deep Subcortical built from 4 of its 8 parcels
ROI mosaics 6 of 26 parcels (the merged ones) had no voxels and drew blank
electrode_signature comparison its results-tree-wide glob would have merged a stale v1 vertex table as the new run's source side

ROI-level results were never affected.

Behaviour changes

  • allen26/allen64 studies: region-level tables and mosaics change (to correct).
  • Antwerp, cosmetic only: Atlas_3DRoisLeftRight.Labels.nii has stored true units since source-localization 2026-03-12 but was still shrunk 10× on the default-affine path. No statistic changes, since ROI extraction and cluster/NBS region labels use the raw affine. The only effect is the mm axis ranges and slice labels of plot_brain_roi_mosaic / plot_brain_roi when drawn on Antwerp, which no analysis module does.
    • Correction: commit aa6ce30's message says this moves Antwerp studies' vertex ROI labels and mosaics. It doesn't. load_vertex_roi_labels has never run: it reads the mapping file's top-level keys as label ids and raises on every atlas, and vertex_network swallows the error and uses spatial labels. That is a separate bug, left for the vertex split.
  • signature_source_vs_sensor.csv gains a source_module column.
  • find_atlas_dir is kept for backward compatibility.

Compatibility

  • MS1 is pinned to v0.4.0 (own worktree) and is unaffected.
  • autifony installs this repo as an editable path, so it picks up these changes as soon as main updates. For its Antwerp data that is cosmetic only; no statistic changes.

Test plan

  • PYTHONPATH=src pytest tests: 219 passed (199 existing + 20 new)
  • Every registered atlas resolves to its own files; the allen26 mosaic atlas carries all 26 parcels
  • R: a conflicting --roi-categories file no longer beats the study config; the file still serves as a fallback
  • After merge: re-run statistics/summary/figures for roi_psd, roi_aperiodic, roi_connectivity, roi_cross_freq, roi_directed on FORGE treatment (cartesian_mc, shell_mc) and confirm 10/10 categories

Performance (added in this PR)

run_signature now runs with BLAS/OpenMP limited to one thread (threadpoolctl). Each fit is tiny, and the thread-pool overhead was dominating the runtime:

band (FORGE electrode features, 21 LOOCV passes) default threads one thread accuracy
Alpha 127.8 s 1.4 s 0.486 in both
Low Gamma 105.4 s 1.7 s 0.600 in both

Results are unchanged (tested). A signature module that took days now takes about an hour.

🤖 Generated with Claude Code

https://claude.ai/code/session_01RibM8Zep2YEjjUbc2LLkgj

alexedmon1 and others added 4 commits September 10, 2026 10:24
…s categories

An allen26 study was being analysed with allen32's partition. Several atlases
share one directory (allen32, allen26 and allen64 all live in allen/), and every
lookup went "the file with the fixed name in the atlas directory":
roi_categories.yaml, roi_mapping.json and allen_labels.nii.gz are allen32's.
Measured on the FORGE treatment run (allen26, Monte Carlo ROI operators):

  - the R region tier replaced the study's categories with that file, so
    Frontal-Anterior and Olfactory matched no parcel and vanished from every
    region-level table, and Deep Subcortical was built from 4 of its 8 parcels
  - the ROI mosaics drew allen32's label volume, so the six merged parcels had
    no voxels and rendered blank (6/26)

Nothing warned. ROI-level results were never affected.

An atlas is now resolved by NAME through source-localization's registry.yaml,
the place an atlas is actually defined, into an AtlasSpec carrying its own
labels, mapping, categories, brain volume and brain mask. base.py stores the
spec where _atlas_dir lived, and every atlas function accepts a spec wherever
it accepted a directory, so the existing call sites pass it through unchanged.
An atlas that is not registered names its files under atlas_files:. A name that
cannot be resolved raises instead of falling back to whatever shares a
directory, which is the fallback that caused this.

R now receives the EFFECTIVE categories (the study's map, a profile's narrowing,
or the atlas default) through BaseAnalysis._r_config_data(), and a single
resolve_roi_categories() in stats_utils.R, replacing seven copies of the
overwrite block, lets them win. The --roi-categories file is only a fallback.

The 10x voxel convention is now read from the NIfTI header, as
source-localization does. The filename rule treated Antwerp's true-unit label
file as inflated and shrank its affine 10x on the default-affine path; this
moves Antwerp studies' vertex ROI labels and mosaics to their correct places.

Tests: atlas resolution for every registered atlas (and the allen26 regression
itself), the header rule, explicit atlas_files, the effective categories handed
to R, and an R-side test that the study's categories beat a conflicting file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RibM8Zep2YEjjUbc2LLkgj
vertex_signature was the only source-side decoding module, so a study with ROI
output alone had no source half of the source-vs-sensor comparison, and Monte
Carlo ROI operators never build the vertex estimate it needs.

roi_signature runs the same classifiers, cross-validation and permutation test
as electrode_signature on per-parcel relative band power, with the identical
feature estimator, so where both run on the same epochs their accuracies are
like for like. It subclasses electrode_signature, whose table, pickle and
figure names now come from the module name rather than a hard-coded prefix.

electrode_signature used to find its source partner by globbing the whole
results tree for vertex_signature_results.csv and taking the first hit. In the
FORGE treatment study that was a stale vertex table from the previous analysis
version, which would have been merged against the new run's sensor results as
if it were its source side. It now looks only in its own paradigm, preferring
roi_signature, and the comparison table records which source module it used.
When the sensor signature ran in another paradigm (it reads the raw recordings,
so one run can serve every source arm), roi_signature takes sensor_paradigm:
and renders the comparison itself.

Tests: a full roi_signature lifecycle including figures regenerated from disk,
and the in-paradigm lookup (a stale table elsewhere is never used).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RibM8Zep2YEjjUbc2LLkgj
A signature run fits a tiny model -- tens of subjects by tens of features --
LOOCV x (1 + n_permutations) times. With the default multithreaded BLAS each
of those fits paid for starting and synchronising a thread pool, and that cost
dominated. Measured on the FORGE treatment electrode features (35 subjects x
30 electrodes, 21 LOOCV passes):

  band        default threads   one thread   accuracy
  Alpha           127.8 s         1.4 s      0.486 in both
  Low Gamma       105.4 s         1.7 s      0.600 in both

That is why one electrode_signature run was heading for several days, and why
v1's logistic fits averaged ~16 minutes with a 15-hour outlier.

run_signature is now wrapped in threadpoolctl.threadpool_limits(1).
threadpoolctl comes with scikit-learn, which the signature modules already
need, and the wrapper falls back to a plain call if it is missing. Results do
not change: the tests check that the limit is really applied inside a fit and
that the limited and unlimited calls agree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RibM8Zep2YEjjUbc2LLkgj
The entry (and aa6ce30's message) said reading the 10x convention from the
header moves Antwerp studies' vertex ROI labels and mosaics to their correct
positions. Measured, it does neither:

  - load_vertex_roi_labels, the vertex path that used the corrected affine,
    has never produced a label. It reads the mapping file's top-level keys
    ("atlas_name", "n_rois", ...) as label ids and raises on every atlas;
    vertex_network catches that and falls back to spatial node labels.
  - the mosaics use the affine only for mm axis ranges and slice labels, and
    the analysis modules never draw them on Antwerp.

No statistic changes. ROI extraction and cluster/NBS region labels use the raw
affine, which was always right.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RibM8Zep2YEjjUbc2LLkgj
@alexedmon1
alexedmon1 merged commit cd20e41 into main Sep 10, 2026
2 checks passed
alexedmon1 added a commit that referenced this pull request Sep 10, 2026
Bumps the package metadata to 0.7.0 and dates the CHANGELOG section that
collected everything since v0.6.0: the September audit remediation, per-atlas
resolution (#3), roi_signature, and single-threaded signature fits.

source_analytics.__version__ comes from git describe, so the tag on this commit
is what outputs will be stamped with; the metadata bump keeps an installed
(non-git) copy reporting the same version.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RibM8Zep2YEjjUbc2LLkgj
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant