Skip to content

Repository files navigation

AGU25 Attendees: We're not naming names (cough Doug cough), but someone copied the wrong example order into their poster. If you would like to see the correct version, you can find it here:

$H_2$ Generation from Olvine, Ortho and Clino

Eleanor

Aqueous speciation modeling has historically focused on specific, well defined systems, and is ideal for laboratory settings or the study of a small number of real-world systems. Standard tools such as Geochemist's Workbench and EQ3/6 exist to fill this niche, alongside more recent additions such as the WORM Portal. However, existing tools are not ideal for the study of systems that are underspecified (e.g. have incomplete composition), have high degrees of uncertainty (e.g. imprecise characterization), or for understanding the broader equilibrium landscape of systems of interest (e.g. serpentinizing systems in general). Eleanor is a powerful open-source modeling framework based on EQ3/6 which fills this gap, providing the process and data orchestration features necessary to facilitate large-scale aqueous speciation modeling. Eleanor includes a standalone executable which accepts a problem specification in YAML, TOML or JSON format, samples fully-defined systems for speciation via the EQ3/6-based "kernel", validates the results, and stores the data in a Postgres. Eleanor’s modular design allows the user to swap the EQ3/6-based kernel with one of their own.

Dependencies

NOTE: We support both Linux and MacOS systems. You might have some luck using the Linux Subsystem for Windows, but we don't pretend to support it.

Eleanor requires python>=3.14 and one external runtime dependency:

  1. A slightly modified version of EQ3/6 found at 39alpha/eq3_6. Future versions will likely add other kernels based on other speciation tools, but EQ3/6 is what we have now.

A PostgreSQL server is required only if you use the postgres output sink (recommended for large-scale runs).

There are two dev dependencies required for installation:

  1. gfortran - you should be able to install gfortran with your system's package manager (e.g. homebrew)
  2. meson - I recommend installing this via uv. See Install below.

Install

I highly recommend using uv to install eleanor:

uv tool install meson # If you haven't installed meson already
uv tool install git+https://github.com/39alpha/eleanor

CLI overview

Top-level commands:

  • eleanor run — run a simulation workload from an order file.
  • eleanor doctor — print install and plugin diagnostics.
  • eleanor gen config|order — emit starter config/order templates.
  • eleanor postgres schema|bulkload|migrate|dump — postgres-specific helper commands.

eleanor run

eleanor run [OPTIONS] ORDER SIMULATION_SIZE

Common options:

  • -c, --config / -d, --database: select config and optionally override the postgres database name.
  • -n, --num-workers: worker count for the selected executor backend.
  • --executor KIND: override the executor kind from config (built-ins: serial, multiprocessing; plugins may add more, e.g. mpi).
  • --chunks-per-worker: override executor.chunks_per_worker from config.
  • --batch-size: navigator batch size passed into navigate(...).
  • --max-nav-attempts: maximum attempts per navigation point before giving up.
  • --tag: add a tag to the order. Repeatable; tags merge with those the order file declares, deduplicated.
  • --seed: override the order's random seed for this run. See Reproducibility.
  • --null-sink: bypass every configured output sink and discard writes via NullSink.
  • --bulk-load / --no-bulk-load: enable/disable postgres bulk-load optimization for this run, on every configured postgres sink. Bulk-load mode is on by default. See Bulk-load controls.
  • -p, --progress: show progress bars (disabled automatically by --verbose).
  • -v, --verbose: verbose output. Also reports, per sink, how many points it was handed and how many it committed — so one sink dropping points alongside one that did not is visible.
  • -s, --scratch: persist scratch artifacts for all simulations regardless of error status.
  • --timing: report a wall-clock attribution of the dispatch loop to stderr when the run finishes.

Reproducibility and the run seed

Every order carries a top-level seed:

name: h2-generation
creator: alice
seed: 4242
# ... the rest of the order

Omit it and Eleanor generates one. Either way the navigators sample from a generator seeded with it, so a run re-sampled under the same seed visits the same variable-space points.

Whether the seed is recorded depends on the sink. The postgres sink persists the whole order in orders.raw, seed included, so a stored run can be re-sampled from what it wrote — see Recovering an order from the database. The csv sink persists only the rows your query projects, so a seed you did not declare is gone when the run ends — project order.seed as a column, or pass --seed and record it yourself.

--seed overrides what the order file declares, without editing it:

eleanor run --seed 4242 -c config.yaml -d eleanor_db order.yaml 50000

Two caveats:

  • Points are reproducible; their order in the output is not. Sampling happens once, in the parent process, but with a parallel executor the workers commit chunks as they finish.
  • Every run starts the generator over. Each run of an order draws from the head of the stream, so running one twice under the same seed visits the same points twice. Pass a different --seed to sample elsewhere in the variable space.

Recovering an order from the database

eleanor postgres dump order writes a stored order back out as an order file, so a run recorded months ago can be re-run without hunting for the original YAML:

eleanor postgres dump order e49c9daa-... -c config.yaml -d eleanor_db -o order.yaml
eleanor run -c config.yaml -d eleanor_db order.yaml 50000

The dump reproduces the order as Eleanor parsed it — seed, kernel settings, reactants and suppressions included — so re-running it at the original simulation size visits the same variable-space points. Its id is not part of the file: every run allocates a fresh one.

Without -o the order goes to stdout. The format is YAML unless --to json is given or the -o path ends in .json:

eleanor postgres dump order e49c9daa-... -c config.yaml -d eleanor_db --to json | jq .kernel

Optional properties that are empty — tags, notes, species, reactants, suppressions and constraints — are left out, since an order file may simply omit them. Pass -e/--keep-empty to emit them anyway, which is useful as a starting point for editing.

Extracting a point's scratch archive

eleanor postgres dump scratch unpacks the scratch archive stored for one variable-space point into a directory, so the kernel's own input and output files can be inspected:

eleanor postgres dump scratch cf165798-... -c config.yaml -d eleanor_db -o scratch/

The archive holds the files the kernel read and wrote while speciating that point. When the point failed it also carries a traceback.txt describing the error, which is usually the reason to reach for this command. -o/--outdir defaults to the current directory and is created if it does not exist; the archive's own layout is preserved beneath it.

Not every point has scratch to extract. By default a run only stores it for points the kernel failed on, so asking for a point that succeeded reports no scratch found for variable space point. Pass -s/--scratch to eleanor run to store it for every point instead — in that case no traceback.txt is included, since there was no error to record.

Built-in output sinks

Built-in output sink types are:

  • postgres
  • csv
  • memory
  • null

Select a sink in your config under output.kind, with sink-specific settings as flat keys alongside kind. For one-off dry runs, --null-sink on eleanor run overrides config output without editing files.

Writing to several sinks at once

output also accepts a list, in which case the run drives every sink in it. The kernel still runs once per point — the compute graph is fanned out to the sinks inside the worker — so N sinks cost far less than N runs:

output:
  - kind: postgres
    database: {host: localhost, database: eleanor_db, username: alice}
  - kind: csv
    name: export
    filename: summary.csv
    id_columns: [order_id, point_id]
    query:
      row_scope: vs_points[*]
      columns:
        - {path: vs_point.exit_code, name: exit_code}

Each sink is addressed by name, which defaults to its kind. Two sinks of the same kind — two CSVs writing different projections to different files — are fine as long as you name them, and duplicate names are rejected. The name is what labels the sink's progress bar and what Eleanor.run keys its returned ids by.

They must also write to different places. Two sinks aimed at one store have nothing correlating their counters or their buffers, so they corrupt each other: two CSVs on one file interleave rows and overwrite each other's _schema.yaml, and two postgres sinks on one database write every point twice and can recreate its indexes mid-bulk-load. Distinct names do not make that safe, so Eleanor rejects such a run at startup.

Every sink keeps its own id space and its own progress bar. Three consequences worth knowing:

  • There is no cross-sink atomicity. An interrupt, or any sink failing, can leave one sink holding a chunk the others do not. A failure in any sink aborts the whole run rather than continuing with the survivors.
  • Rows from different sinks cannot be joined. The csv sink's point_id counter and the postgres sink's vs_point sequence are unrelated.
  • Cost scales with the sink count. Each serial sink adds a writer thread and a bounded queue, and the Postgres subtransaction pressure described below applies per postgres sink.

Identity columns on the csv sink

The csv sink projects each result through an EQL query, but identity is the sink's own, not part of the object graph the query walks — so ids are requested in settings rather than as query paths. id_columns accepts order_id (the run's UUID) and point_id (a per-run VS-point counter), and prepends them to the header in the order given:

output:
  kind: csv
  filename: rows.csv
  id_columns: [order_id, point_id]
  query:
    row_scope: vs_points[*]
    columns:
      - order.name
      - vs_point.temperature
      - vs_point.exit_code
# header: order_id,point_id,name,temperature,exit_code

Omit id_columns for a file with no identity columns. Every row of one VS point shares its point_id, so a query emitting several rows per point repeats the value — point_id identifies the point, not the row.

Appending to an existing csv file

Pointing the csv sink at a file that already exists appends to it: the run gets its own order_id, its own point_id counter, and its rows go on the end. The header must match the configured query's columns exactly, and the file's _schema.yaml sidecar must be present — a CSV without one is rejected rather than guessed at.

The sidecar also records the version of Eleanor that wrote the file, and appending under a different version is rejected — as is a sidecar recording no version at all, which is what any file written by v0.21.1 or earlier carries. Two Eleanor versions are not guaranteed to produce numerically identical results, and a CSV has nowhere to note which build wrote which row, so a mixed-version file cannot be untangled after the fact. Point the run at a new file instead.

Only the release component is compared: a .devN+g<hash> suffix is stripped first, so unreleased builds of the same version can append to one another's files. 0.21.2.dev1 and 0.21.2.dev9 match; 0.21.1 and 0.21.2 do not.

This is stricter than the postgres sink, which happily accepts runs from several versions because orders.eleanor_version keeps each one attributable. The constraint is the CSV's missing provenance, not the version difference.

Parallelism, batch size, and Postgres subtransaction pressure

The postgres sink uses one outer transaction per batch plus one savepoint per variable-space point. With multiprocessing, each worker has its own connection and can hold up to batch_size in-flight savepoints. Operationally, subtransaction pressure scales roughly with:

batch_size × num_workers

At sufficiently high values this can trigger Postgres SubtransSLRULock contention. If you see this, reduce --batch-size, reduce worker count (-n / backend configuration), or both.

Bulk-load controls

For postgres outputs, bulk-load mode drops secondary indexes/constraints during ingestion and recreates them at finalize. It is on by default, so a plain run already bulk-loads; --no-bulk-load (or bulk_load_optimization: false in the sink block) keeps the indexes and constraints in place:

eleanor run --no-bulk-load -c config.yaml -d eleanor_db order.yaml 200000

You can also control this window explicitly:

eleanor postgres bulkload drop -y -c config.yaml -d eleanor_db
eleanor postgres bulkload recreate -c config.yaml -d eleanor_db

If a bulk-load run is interrupted before finalize, use eleanor postgres bulkload recreate to restore constraints/indexes.

Shell Completion

Bash

eval "$(_ELEANOR_COMPLETE=bash_source eleanor)"

Zsh

eval "$(_ELEANOR_COMPLETE=zsh_source eleanor)"

Fish

_ELEANOR_COMPLETE=fish_source eleanor | source

About

Large-scale aqueous speciation modeling

Topics

Resources

Code of conduct

Contributing

Stars

1 star

Watchers

2 watching

Forks

Releases

Contributors

Languages