Skip to content

Cli queries - #11

Open
kriben wants to merge 21 commits into
mainfrom
cli-queries
Open

kriben wants to merge 21 commits into
mainfrom
cli-queries

Conversation

@kriben

@kriben kriben commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator

No description provided.

Task instantiation and input assembly ran outside the runner's guarded
block, so an invalid configuration value (or a raising task constructor)
propagated a raw exception out of Runner.run(), leaving the job RUNNING
with no TASK_FAIL/JOB_FAIL events.

Build the task and its input inside the guarded block and move the
assembly into _build_task_input(). A constructor failure fails the job
and emits JOB_FAIL; TASK_FAIL is skipped because there is no task
instance to report.
SIGALRM/ITIMER_REAL is a single process-wide timer. An inner Runner
started by a workflow_task re-armed it for its own timeouts and then
cleared it with setitimer(0), so the wrapping task's timeout and the
outer job deadline were silently lost for the rest of the task.

Route every armed timeout through a process-wide _AlarmScheduler that
keeps all active timers, always arms the nearest pending expiry, raises
the error of the timer that fired, and restores the previous SIGALRM
handler once the last timer is removed. Registration from a non-main
thread is rejected up front, since setitimer would otherwise deliver
the signal to the main thread.
Job only verified that declared config fields had values for root
tasks. A dependent task with config_fields and no configuration passed
validation; with a single whole-output dependency the runner then fed
it the upstream output unchanged, failing later inside the task.

Check every task, including mapped ones, and validate map sources
first so their more specific errors still take precedence.

The text_analysis example declared config_fields for ScoreReadability
without supplying values and only worked because of this gap; drop
them to match the Python definition.
A task depending on a single output field (depends_on=(task, "field"))
receives that field's value as its whole input, so the runner silently
dropped any configured values and validation never flagged them. In
YAML, input values for such a task were ignored.

Reject the combination at build time and point to the named-dependency
form (depends_on={"<input_field>": (task, "field")}), which does merge
configuration values.
The per-task and per-item alarm stayed armed through output-type
checks and the TASK_COMPLETE / MAP_ITEM_COMPLETE (and failure) hooks.
If it fired inside a hook, _emit swallowed the TaskTimeoutError as a
HookError warning and the task was still recorded as completed.

Disarm immediately after run() returns or raises, and validate mapped
item input before arming so validation failures are not timed either.
The run-level finally remains as a safety net.
build() returned the builder's internal Workflow, so calling add_task()
(or build() again) afterwards mutated an already validated workflow
without re-validating it.

Snapshot the builder state into a new Workflow on every build(). Fan-in
dependency dicts are copied too, because validation may rewrite their
entries when unwrapping mapped outputs. The builder remains usable and
each build() yields a separately validated workflow.
import_class only caught ModuleNotFoundError. A SyntaxError, an
ImportError, or any exception raised by module-level code escaped the
loader (and the CLI) as a raw traceback, and a missing dependency
imported *by* the task module was misreported as the task module
itself not existing.

Report "Cannot import module" only when the missing module is the
requested module or one of its parent packages; wrap every other
import-time failure as "Error while importing module ..." with the
original exception chained.
The duplicate-key check in _UniqueKeyLoader tested membership of each
constructed key in a set, so a sequence or mapping key (e.g. "? [a, b]")
raised a bare TypeError. That bypassed the loader's yaml.YAMLError
handling and surfaced as a traceback instead of a ConfigLoadError.

Raise a ConstructorError ("found unhashable key"), matching PyYAML's
own error for this case.
Only a TypeError from the hook constructor was converted, and parameter
coercion ran outside the guarded block. Any other exception from a
hook's __init__ (e.g. a ValueError for an invalid parameter), or a
TypeError from coercing a non-string to Path, escaped the loader raw.

Guard coercion and construction together, catch any exception, and
include its type in the message with the original chained.
The enclosing workflow files were tracked in a frozenset and sorted
alphabetically for the error message, so for chains of three or more
files the reported "a -> b -> c" did not match the actual reference
path.

Track ancestors as an ordered tuple, outermost first, and report the
chain as it was followed.
There was no taskmaestro/__main__.py and cli.py had no __main__ guard,
so "python -m taskmaestro" failed and "python -m taskmaestro.cli"
silently exited 0 without doing anything. Only the installed console
script worked.

Add both entry points; each exits with main()'s status.
When rendering a workflow_task as a Mermaid subgraph, edges into it were
redirected to the first inner root without config_fields, which could be
a mapped task. Mapped roots are fed by the JobConfiguration, and
workflow_task derives its input type from the other root, so the diagram
pointed outer edges at the wrong node.

Add Workflow.input_root_names() (roots that are neither configured nor
mapped) and use it in the visualizer, workflow_task and Job root-input
validation, so all three agree on which tasks consume the job input.
input.yaml gave the image as "taskmaestro.png", resolved against the
current working directory, so the YAML example only worked when run
from the repository root; "taskmaestro run" from the example folder
failed with "Image not found".

LoadImage now resolves relative paths against the example's directory,
and both input.yaml and the Python mode give the path relative to it.
Absolute paths are unchanged.
print_report read result.result unconditionally, so a failed run ended
in an AttributeError on None instead of showing why the workflow failed.

Print the status, failed task and error when the job did not complete,
as the release_pipeline and resinsight examples already do, and replace
the type: ignore with a cast now that the result is known to be set.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant