Put all column info in one place
a la InputCsvFile, but for data columns
To see what I mean, search for a column name eg TubePlacementInstant and see how many places it appears:
- pyarrow schemas
- pseudonymisation status
- SQL query
- unit test
maybe only the first two could be replaced by this new thing though?
Put caboodle, emap, etc. config envs into their own files and then include multiple files in env_files as needed for each container
would avoid repeating data in all the config files
consider combining waveform-monitoring and waveform-janitoring
one was going to be ro, one was going to be rw, but they turned out more similar in practice than I expected. Likely could merge them.
Ensure snakemake is the only arbiter of file paths
See rule ehr_lookup for a positive example of this. ehr_for_csn does not determine the output file path. It is determined by the snakefile in conjunction with InputCsvFile.
Put all column info in one place
a la InputCsvFile, but for data columns
To see what I mean, search for a column name eg
TubePlacementInstantand see how many places it appears:maybe only the first two could be replaced by this new thing though?
Put caboodle, emap, etc. config envs into their own files and then include multiple files in env_files as needed for each container
would avoid repeating data in all the config files
consider combining
waveform-monitoringandwaveform-janitoringone was going to be ro, one was going to be rw, but they turned out more similar in practice than I expected. Likely could merge them.
Ensure snakemake is the only arbiter of file paths
See rule ehr_lookup for a positive example of this. ehr_for_csn does not determine the output file path. It is determined by the snakefile in conjunction with InputCsvFile.