Skip to content

Refactoring ideas #107

Description

@jeremyestein

Put all column info in one place

a la InputCsvFile, but for data columns

To see what I mean, search for a column name eg TubePlacementInstant and see how many places it appears:

  • pyarrow schemas
  • pseudonymisation status
  • SQL query
  • unit test

maybe only the first two could be replaced by this new thing though?

Put caboodle, emap, etc. config envs into their own files and then include multiple files in env_files as needed for each container

would avoid repeating data in all the config files

consider combining waveform-monitoring and waveform-janitoring

one was going to be ro, one was going to be rw, but they turned out more similar in practice than I expected. Likely could merge them.

Ensure snakemake is the only arbiter of file paths

See rule ehr_lookup for a positive example of this. ehr_for_csn does not determine the output file path. It is determined by the snakefile in conjunction with InputCsvFile.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions