Skip to content

Fix demo helper scripts that print null, fail on first run or watch the wrong entity - #78

Open
bburda wants to merge 64 commits into
mainfrom
fix/demo-helper-scripts
Open

bburda wants to merge 64 commits into
mainfrom
fix/demo-helper-scripts

Conversation

@bburda

@bburda bburda commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Description

The helper scripts a user runs by hand printed null, failed on the first run after start, or watched the wrong entity. CI did not see this, because the smoke tests only exercised the gateway API. This PR fixes the scripts and adds checks to the existing smoke tests so each problem fails CI.

sensor_diagnostics

  • check-demo.sh reads the data resources the sensors really publish (/sensors/scan, /sensors/imu, /sensors/fix) and the message under .data. Section 8 prints each parameter with its value and type from the configuration detail.
  • With an active fault, the fault detail, snapshot and bulk-data sections now run against the app that reported the fault.
  • It waits up to 30 s for linked sensor data after start. A failed /faults read exits non-zero and does not say "no faults".

multi_ecu_aggregation

  • The container scripts change parameters through the ECU's own gateway (PUT /apps/<app>/configurations/<param>). The first ros2 param set in a fresh container failed with Node not found.
  • A refused write fails the script and names the parameter on stderr, so the Scripts API returns it in error.message.
  • path_planner runs its planning timer in its own callback group on a two-thread executor. Its parameter service now answers during an injected delay, so the planning restore-normal takes about 3 s (75 s before).
  • The Dockerfile copies container_scripts/ after the workspace build, so a script change does not rebuild the workspace.

moveit_pick_place

  • move-arm.sh uses the local ROS 2 CLI only when the arm action is visible locally. Before, ros2 node list returning 0 on an empty graph sent every host with ROS 2 sourced down the local path.
  • It reports the real goal status, exits non-zero when a goal did not succeed, and resends a goal that got no response. docker exec no longer needs a TTY.
  • check-entities.sh and check-faults.sh show real values, and check-faults.sh reports a failed fault read.
  • The README says that the pick-and-place loop can preempt manual moves.

turtlebot3_integration

  • check-entities.sh and check-faults.sh show real values and report a failed fault read. The README API examples print real values.
  • setup-triggers.sh and watch-triggers.sh watch anomaly-detector, which reports the navigation faults. The hint is now inject-localization-failure.sh. The gateway does not send trigger events for faults reported from a sub-path of a node (/bridge/anomaly_detector/goal_status), so inject-nav-failure.sh produces no event. That is a gateway issue and is not changed here.
  • restore-normal puts AMCL back on the robot's pose read from Gazebo, so localization works again after inject-localization-failure. It fails when a write is refused.
  • The Dockerfile copies container_scripts/ after the workspace build.

ota_nav2_sensor_fix

  • run-demo.sh names the robot the image builds (RB-Theron).

Tests
The new checks compare script output with direct API reads, run the multi-ECU scripts as the first ones on a fresh stack, and drive real failures (a configuration lock that answers 409, a stopped fault manager that answers 503). Every new check was first run against the old scripts or a broken copy of the new ones, and failed. The EXIT traps that resume paused processes keep the real exit status, so a gateway that never starts still fails the job.

Related Issue

closes #77

Checklist

  • Tested locally
  • README updated (if needed)

Data resources are keyed by full topic id (/sensors/scan, not scan) and
the reading is nested under .data, so sections 5-7 printed null for
every field. Section 8 read .value from the configurations list, which
never carries a value; read each parameter's own detail endpoint
instead. The fault collection has no entity_id/code fields, so the
snapshot/bulk-data walkthrough always fell back to "skip" - resolve the
owning App from the fault's reporting_sources instead. Also drop the
Components/Apps columns that read fields the API never returns
(area, namespace) for ones that do (description, component id), and
fix the diagnostic_bridge id typo in the README (live id is hyphenated).

Covered by a new smoke_test.sh section that runs check-demo.sh live
against a fault and asserts no null fields.
check-entities.sh and check-faults.sh read fields the SOVD entity and
fault responses never carry (area, category, is_located_on, hosted_by,
code, reporter_id, message, timestamp), so every labeled field printed
null. Read the real fields instead (fault_code, severity_label,
reporting_sources, x-medkit.component_id, ...), matching the shape
moveit_pick_place/check-faults.sh already uses.

setup-triggers.sh and watch-triggers.sh both watched
apps/diagnostic-bridge, which reports nothing for this demo - faults
arrive from apps/anomaly-detector. Fixed both to the live entity id.
The nav-failure inject script never brings a fault to CONFIRMED state
under the gateway's own OnChange notification path (verified live:
inject-localization-failure fires reliably, inject-nav-failure never
does across repeated attempts), so the hinted inject script and README
example now point at inject-localization-failure.sh.

Covered by a new smoke_test_turtlebot3.sh section that runs both
scripts against a live fault and asserts no null fields, and a new
smoke_test_navigation.sh section that drives setup-triggers.sh /
watch-triggers.sh / inject-localization-failure.sh and asserts the SSE
stream delivers an event.
run-demo.sh's text called the simulated robot a TurtleBot3. The image
builds the Robotnik RB-Theron description (Dockerfile.gateway), not
TurtleBot3.
move-arm.sh always printed a success line and exited 0, even when the
controller aborted the goal (the pick-and-place loop competes for the same
action). It now parses the action's own final status, exits non-zero and
prints a failure line when the goal did not succeed, and ./move-arm.sh demo
runs every step regardless of earlier failures while still exiting non-zero
overall. The local-vs-container branch also checked only that `ros2 node
list` succeeded, which is true even on an empty, disconnected graph; it now
confirms the target action is actually listed. The container exec no longer
requests a TTY, so the script also works from a pipe or CI.

check-entities.sh displayed several fields (component area, app category,
app location, function category/host, fault code/reporter) that the API
never populates for this demo, always printing null. Those columns are
dropped or replaced with the real equivalents (app-to-component links via
x-medkit, real fault field names) so the explorer only prints real data.

README.md documents the new goal-preemption behavior and corrects the
manipulation-monitor entity id in the triggers section (the demo scripts
already use the hyphenated id; the docs still had the old underscored one).
The container scripts changed parameters with the ros2 CLI, whose first
invocation in a container races an unstarted ros2 daemon and fails with
"Node not found". Switch every inject and restore-normal script to the
gateway's configuration API (PUT .../configurations/<param>), which has no
daemon to warm up, matching how sensor_diagnostics already does this.

path-planner blocks its single-threaded executor for the full injected
planning delay on every cycle, so its parameter service can stay busy long
after the delay was set. restore-normal on the planning ECU retries its
writes there until they land. The smoke test verifies the injection took
effect through the resulting fault instead of a live parameter read-back,
since a read racing that busy window does not reliably recover within any
bounded wait, while a later write does.

Move the container_scripts COPY in the Dockerfile below the colcon build
step: the scripts are not a build input, and putting them last means a
script-only change no longer invalidates the compiled packages.

Add an end-to-end section to the multi-ECU smoke test that runs every
inject and restore-normal script on a fresh stack, verifies the parameters
they touch actually change and reset, checks that all injected faults are
gone after restore, and drives one real write failure to prove it is
reported by name.
… direct reads

The check-demo.sh test only looked for null fields, so it passed when
sections 5-8 printed nothing, dropped parameters or printed fixed values.
It now compares sections 5-7 with direct reads of the same topics, within
8 sigma of the live noise configuration, and requires a second run a
second later to print a new sample. Section 8 must list exactly the
parameters the configurations endpoint lists, each with the value and
ROS type its detail endpoint returns.

check-demo.sh printed the constant type "parameter" for every LiDAR
parameter. It now prints the ROS type from the detail endpoint.

The test strips ANSI colours by piping the script output into sed, which
needs no shellcheck disable.
…ion check

A LiDAR read carries 360 ranges and 360 intensities, which buried the printed values in the failure line.
The script output is piped straight into sed instead of being captured
first and fed back through a here-string, which shellcheck flags as
SC2001.
The trigger test ran inject-localization-failure.sh directly and passed
on any event, so a wrong inject hint in setup-triggers.sh or an event for
another fault went unnoticed. It now runs the command setup-triggers.sh
prints and requires an event carrying LOCALIZATION_UNCERTAINTY.

The inject leaves AMCL with a uniform particle cloud, and the cleanup only
deleted fault records, so a second run on the same stack failed its
localization checks. The test now sets the AMCL pose from the spawn point
plus odometry, runs restore-normal.sh, forces AMCL updates for two
detector intervals without LOCALIZATION_UNCERTAINTY coming back, and
drives both goals again, which returns the robot to the spawn point.

The README notes that restore-normal.sh does not re-localize AMCL.
…t severity

The localization check read LOCALIZATION_UNCERTAINTY as CONFIRMED at
ERROR severity from the fault list. A fault keeps the highest severity it
ever had, also when a FAILED event reactivates it after a clear, so once
the injected localization failure raised it at ERROR, every later WARN
crossing during a healthy drive read as ERROR and failed the check.

The check now compares AMCL's published position spread with the
detector's configured ERROR threshold, which is what the detector
compares when it reports at ERROR severity.
The EXIT trap ran resume_pick_place_loop before print_summary, and
print_summary takes the script's exit status from $?. The resume call
always succeeds, so a run that ended early with a non-zero status and no
recorded failure (the gateway never became healthy, for example) was
reported as "All smoke tests passed!" and exited 0.

The trap now saves $? on entry and restores it right before
print_summary. errexit is turned off inside the trap, because under
set -e restoring a non-zero status would end the trap before the summary
runs.
…lues

The move-arm.sh checks decided the real outcome of a goal from the
script's own exit code, so a script that swapped success and failure
still passed. They now read the status the action client printed itself
("Goal finished with status: ...") and require: SUCCEEDED gives exit 0
and exactly the success line; any other status gives a non-zero exit and
exactly one failure line that names that status.

For ./move-arm.sh demo, each of pick, place and ready (home) must have
exactly one result line, in its own step, matching that step's real
status, and the command must exit non-zero when a step really failed.

The goals are now preempted on purpose instead of by racing two
processes: a probe inside the container waits until the arm controller
and MoveGroup are idle, then answers the next goal the controller starts
with a competing goal. Pausing pick_place_loop now also waits for the arm
to go idle, since a MoveGroup goal sent before the pause still runs.

The check-entities.sh check passed on output that held only the headings.
It now compares what the script prints with direct API reads: every
known component with its name and description, every app and its
component, and each active fault's code, severity, status and sources.
The check-entities.sh and check-faults.sh section injected
inject-nav-failure.sh as soon as the gateway answered. On a fresh stack
Nav2 is not active yet, bt_navigator rejects the goal and
NAVIGATION_GOAL_ABORTED never appears, so the section failed on every
fresh start, which is how CI runs it. It now waits up to 180 s for
bt-navigator and planner-server to report the active lifecycle state.

The file header no longer claims the test injects no fault.
…mulated pose

inject-localization-failure re-initialises AMCL to a uniform particle
cloud, and restore-normal only cancelled goals, reset velocity limits and
cleared faults. AMCL stayed on the scattered cloud, so the fault came back
and navigation ran on a wrong pose estimate.

restore-normal now reads the robot's pose from the running Gazebo
simulation and sets it through AMCL's set_initial_pose operation before
it clears the faults. The map frame of the demo is the Gazebo world
frame, so no spawn pose is assumed. The script fails when it cannot
re-localize.

The inject script's comment said the goal is typically rejected; Nav2
accepts it and drives from the scattered estimate. The README describes
the new restore step.
…uild

The container scripts are not build inputs, but they were copied before
colcon build, so any script change rebuilt the workspace layer.
…CL against the simulation

The navigation test set the AMCL pose itself before running
restore-normal.sh, so it proved the test could recover the demo, not that
the demo recovers itself. It now relies on restore-normal.sh alone and
compares AMCL's estimate with the robot's pose in Gazebo, within 0.2 m and
0.2 rad, since a cloud that converged on the wrong place is no longer
reported as uncertain.
…s it

The arm controller can fail to deliver the goal response to a freshly
started `ros2 action send_goal` whose endpoints are not yet discovered.
It logs "Failed to send goal response ... client will not receive
response" and drops the goal without running it, and the CLI then waits
for that response forever. move-arm.sh hung there with no result.

Each CLI run is now limited to 30 s by `timeout` (inside the container
on the docker branch, since killing `docker exec` leaves the CLI
running). When a run got no goal response at all, the goal never ran, so
the script sends it again, up to three times. An accepted or rejected
goal is never sent twice; if its result does not arrive in time the
script reports it as failed. PYTHONUNBUFFERED keeps the CLI's lines when
the timeout ends it.

The smoke test simulates the lost response with a ros2 on PATH that
drops the first goal and hands the next to the real CLI in the
container, and requires exactly one resend and the real SUCCEEDED
result. The unreachable-ros2 check now reads its output from a file
instead of piping a large string into `grep -q` under pipefail.
…delay

path-planner waited out planning_delay_ms inside its timer callback on a
single-threaded executor, so every parameter request queued behind the
delayed cycle. The gateway timed those requests out and then reported the
node unavailable for its negative-cache window, which made restore-normal
after inject-planning-delay take more than a minute.

Run the planning timer in its own callback group on a two-thread executor,
so the parameter services answer while a cycle waits. The delay still holds
back each path and still raises the SLOW_PLANNING diagnostic. A new
planning_delay_ms also ends the wait of the cycle in progress, so a restore
takes effect at once instead of after one more delayed cycle whose stale
path keeps behavior-planner reporting faults after they were cleared.
All six container scripts write parameters through one put_config helper.
A refused write is named on stderr together with its HTTP status. The
Scripts API reports stderr as the error message of a failed execution and
drops stdout, so the failure lines the scripts printed to stdout never
reached the caller, who only saw "Script exited with code 1".

path-planner now answers parameter requests during an injected delay, so
the planning restore-normal no longer retries its writes for minutes.

README: document how the scripts change parameters and report a failed
write, and drop the troubleshooting note about a slow planning restore.
…essage

Read the parameter writes of each container script from its put_config
calls and check all of them through the gateway with plain GETs: the
written value after each inject, the launch value (ECU params file, else
the node declaration) after each restore-normal. Static checks fail when a
script writes a parameter in a form the test does not read, when an inject
does not change its parameter, when restore-normal does not write it back,
and when a container script is not executed by the test.

For the planning delay, check the PATH_PLANNER fault, that paths come one
delay apart while it is injected and at the planning rate after restore,
and that restore-normal completes within 15 s. Run the host wrappers:
inject-cascade-failure.sh, then restore-normal.sh must exit 0 within 30 s
and leave every parameter at its launch value and no fault.

Replace the write-failure check, which ran its own curl snippet in the
container, with a committed script run through the Scripts API: another
client locks gripper-controller's configurations, actuation restore-normal
must fail and its error message must name exactly the two refused writes.
The lock is released and the ECU restored afterwards.
One read of the Gazebo pose topic returned the model twice. restore-normal
would then post two JSON documents to set_initial_pose and fail, and the
navigation test could not parse the simulated pose. Both now use the first
match only.
…, time the bound

Before every restore-normal whose result is checked, move each parameter it
writes away from its launch value through the gateway (a bool flips, a
double moves by 0.001, an integer by 1; injected values stay), and check
that it is away. The launch-value checks after the restore then fail for
any write the script skips, not only for the injected parameters.

The write-failure test runs one script. Check that every container script
carries the same put_config and the same ERRORS exit check as that script,
initialises ERRORS once and makes all its writes before the check. The
write-failure test also checks that the refused writes left their
parameters unchanged and that the other writes landed, and the test ends by
checking that no fault is left.

Time the planning restore-normal from before the POST until its completed
status is read, and require that time to be within 15 s. The wait runs to
four times the bound so a slow restore reports its time. The host
restore-normal.sh bound is measured in milliseconds the same way.
…y goal

The demo check counted a step with no final action status as a failed
step and accepted any "Failed: <label> (" line for it, and it discarded
the preempt probe's exit code. A move-arm.sh that printed three failure
lines without sending a goal passed. Every checked goal must now reach a
final status the action client printed itself. The probe must exit 0
after it saw the goal it preempted end with a non-SUCCEEDED status, and
that status must be the one move-arm.sh's own goal reports (the ready
goal, and the demo's pick step).

The fault setup waited for the arm controller's action server and for
each goal response and result with no limit, inside `|| true`, so a
healthy gateway with a missing controller hung the test forever and left
the probe running in the container. The fault setup is now a mode of the
same probe, and every wait in it is bounded: 30 s for the server, 60 s
for an idle arm, 10 s for a goal response or result, and a 240 s
`timeout` around the whole process in the container. A probe that runs
out records a FAIL, and the pick-and-place loop is resumed after it.
move-arm.sh runs are limited to 400 s, check-entities.sh to 120 s, and
the real CLI behind the fake ros2 to 60 s. The fault setup also waits for
a fault that manipulation_monitor reports, not any fault.
…emo.sh

/health answers about a second before the gateway links the sensor nodes.
check-demo.sh started in that window printed null for every LiDAR, IMU
and GPS field and still reported the demonstration complete. It now waits
up to 30 s for the sensor data and the LiDAR configurations it prints,
and stops with a message and exit code 1 when they do not appear.
…e ranges at default noise

A first check-demo.sh run starts as soon as /health answers, polled every
0.2 s, and must print no null and either real values or the stop message
with a non-zero exit.

The section 5-7 comparisons ran while the injected LiDAR noise was 0.5 m,
where 8 sigma covers the whole 0.12..3.5 m range, so invented ranges
passed. They now run before the fault injection, at the configured noise,
and the LiDAR check fails when 8 sigma is not below a tenth of the range
span. The run with the fault active keeps the section 8 and fault detail
checks.
…t with direct reads

The test only required no null and the injected fault code, so a script
printing wrong components or a made-up severity passed. It now compares
the printed component ids, each app's component and, for every active
fault, the code, severity label, status, sources and occurrence count
with direct API reads. The fault list is read before and after each
script run and the printed faults must equal one of the two reads.
…live AMCL updates

The inject now starts from a parked pose. The test asserts in Gazebo that
the parked pose is farther from spawn than the agreement tolerance, so a
restore-normal.sh that assumed the spawn pose fails the agreement check.

The forced AMCL updates discarded every response. Each update must now
return 200 and be followed by an AMCL pose with a newer stamp, or the
check that LOCALIZATION_UNCERTAINTY stays absent fails.
The header pointed at a comment in smoke_test_turtlebot3.sh that no longer
exists, and that test now injects a navigation failure itself. The reason
that holds is that the thresholds count FAILED and PASSED events, and only
a direct report sends an exact number of them.
turtlebot3 check-faults.sh and sensor_diagnostics check-demo.sh read
/faults without checking the response. While the fault manager is not
available the gateway answers 503 with an error body, and both scripts
then reported no active faults and exited 0. They now name the failed
read with its HTTP status and exit 1.
The smoke tests pause the demo's fault manager with SIGSTOP, so GET
/faults answers 503 as it does before the fault manager is up, and run
check-faults.sh (turtlebot3) or check-demo.sh (sensor_diagnostics). The
script must exit non-zero with the failed-read message and must not
claim there are no faults. The fault manager is then resumed and /faults
must answer again within 30 s. The EXIT trap also resumes it and keeps
the script's exit status.
@bburda
bburda marked this pull request as ready for review September 28, 2026 18:24
@bburda bburda self-assigned this Sep 28, 2026
printf '%s\n' "${output}"
# An accepted or rejected goal has its answer. Only a goal that got
# no response at all never ran, so only that one is sent again.
if grep -qE '^(Goal accepted with ID|Goal was rejected)' <<< "${output}"; then

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Everything without a "Goal accepted/rejected" line is treated as a lost goal, including docker: command not found, a wrong CONTAINER and The passed action type is invalid; those fail instantly, yet the loop prints "No goal response within 30 s, sending the goal again" twice and ends with status: UNKNOWN. Resend only when the output shows the CLI got as far as Sending goal:, and report any other output as the failure it is.

Comment thread demos/moveit_pick_place/move-arm.sh Outdated
# demo looks identical to being inside the container. Checking that the
# action itself is listed avoids that false positive.
can_reach_action_locally() {
command -v ros2 &> /dev/null && ros2 action list 2> /dev/null | grep -qFx "${ACTION}"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Inside the container this hits the cold-start effect from #77 item 2: with no ros2 daemon up, the first ros2 action list answers from a direct node after a short discovery spin and can be empty, so the script falls through to docker exec, which does not exist in the container, while README line 76 still promises it works from inside. Retry the listing once, treat a missing docker as local, or drop the in-container claim.

Comment thread tests/smoke_test.sh Outdated
}

if [ "$EARLY_RC" -ne 0 ]; then
if grep -q "Sensor data not available" <<< "$EARLY_PLAIN"; then

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This branch also passes when check-demo.sh gives up after its 30 s wait, which is exactly the first-run failure #77 reports, so a gateway that links in 31 s keeps the check green. wait_for_runtime_linking right below proves linking completes, so require EARLY_RC -eq 0 here, or record the linking time and fail when it exceeds DATA_WAIT_SEC.

if [ "$waited" -ge "$DATA_WAIT_SEC" ]; then
echo_error "Sensor data not available at ${GATEWAY_URL} after ${DATA_WAIT_SEC}s."
echo " Check that the sensor nodes are running, then retry."
exit 1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

After inject-failure.sh the IMU stops publishing on purpose, so this gate times out and the script exits before the fault sections, which is the one moment a user runs it to see the fault. Report the missing sensor and continue to the fault sections instead of exiting.

Comment thread demos/sensor_diagnostics/check-demo.sh Outdated
FIRST_FAULT=$(echo "$FAULTS_JSON" | jq -r '.items[0].fault_code')
REPORTING_SOURCE=$(echo "$FAULTS_JSON" | jq -r '.items[0].reporting_sources[0] // empty')
FIRST_ENTITY=$(curl -s "${API_BASE}/apps" | jq -r --arg node "$REPORTING_SOURCE" \
'.items[] | select(.["x-medkit"].ros2.node == $node) | .id' | head -n 1)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The anomaly detector reports with source_id = "/processing/anomaly_detector/" + source (anomaly_detector_node.cpp:213), e.g. /processing/anomaly_detector/imu_sim, while the app's node path is /processing/anomaly_detector, so this exact match fails for every directly reported IMU/GPS fault and sections 10-12 are skipped when one is first in the list. Match on a node-path prefix with a / boundary ($node == $src or ($src | startswith($node + "/"))).

angle_max: .data.angle_max,
range_min: .data.range_min,
range_max: .data.range_max,
sample_ranges: .data.ranges[:5]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The || echo fallback right below is dead: jq exits 0 on a 404 body, so before the scan is linked this section prints "angle_min": null and the "Gazebo may still be starting" hint never appears. Use jq -e '.data | select(. != null) | {...}' or check %{http_code} so the fallback actually runs.


echo_step "6. Faults"
curl -s "${API_BASE}/faults" | jq '.items[] | {code: .code, severity: .severity, reporter: .reporter_id}'
curl -s "${API_BASE}/faults" | jq '.items[] | {code: .fault_code, severity: .severity_label, status: .status, sources: .reporting_sources}'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same dead fallback in the joint-states block above: a 404 for data/joint_states prints "joint_names": null and the hint never runs, because jq exits 0 on the error body. Use jq -e '.data | select(. != null) | {...}'.

ros2 param set /actuation/joint_driver failure_probability 0.0 || ERRORS=$((ERRORS + 1))
put_config joint-driver inject_overheat false
put_config joint-driver drift_rate 0.0
put_config joint-driver failure_probability 0.0

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The writes now fail the script, but the fault clear below still swallows failure (-f ... || true on both DELETEs), so with the fault manager unavailable (the 503 state this PR's own tests drive) the execution completes with "status": "restored" while the ECU's faults stay latched; the turtlebot3 restore in this PR reports that case. Capture %{http_code} of the final DELETE and fail on non-2xx; same in perception-ecu and planning-ecu restore-normal.


# Terminal 2: Inject a fault - the trigger fires in Terminal 3!
./inject-nav-failure.sh
./inject-localization-failure.sh

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"any new or updated faults reported by the anomaly detector" overstates it: faults reported under a source sub-path (NAVIGATION_GOAL_ABORTED / NAVIGATION_GOAL_CANCELED via /goal_status) produce no trigger event, so ./inject-nav-failure.sh stays silent on the watcher. Say so here and name inject-localization-failure.sh as the inject that fires.

restore-normal cleared the ECU's faults with `curl -sf ... || true` twice
and then always printed "restored". While the ECU's fault manager does not
answer, the gateway rejects DELETE /faults with 503, so the execution
completed as restored while the faults stayed.

The second clear now decides the result: a status other than 2xx exits 1
and names the clear on stderr ("FAIL: clear faults (HTTP 503)"), which the
Scripts API returns as the execution error. The first clear stays best
effort. Both clears get a 30 s curl timeout.

The smoke test stops each ECU's fault manager with SIGSTOP and checks,
through the Scripts API and with a direct run of the script, that
restore-normal fails and names the clear. It covers the fault manager
stopped for both clears and stopped only between the first and the final
clear, and that restore-normal completes again once it resumes.
…l ros2 without docker

move-arm.sh sent a goal again on any output without an accepted or
rejected line, so a failure that never sent the goal (no such container,
no action server, an invalid action type) ran three times and printed
"No goal response within 30 s" even when the command failed at once.
Now only output with "Sending goal:" and no response is sent again, and
the resend line names the 30 s wait only when the timeout ended the
run. A goal that was never sent is reported once with the command's
output and "Failed: <pose> (goal not sent: ...)", and the script exits
non-zero.

Without a docker CLI the script now uses the local ros2 directly, and it
no longer runs `docker ps` before it knows it needs it. Inside the
container a cold `ros2 action list` missed the arm action about half the
time, the script then fell through to `docker exec`, and every run
printed "docker: command not found". Where docker is available, the
local-reach check lists the actions a second time before it falls back
to `docker exec`.

The smoke test runs the not-sent cases for real: a missing container
from the host, and in the container no docker and no ros2, a ROS domain
without the action server, and a shadowed control_msgs. Each must be
reported once, non-zero, with no resend line. A copy of move-arm.sh run
in the container with the ros2 daemon stopped before every run must get
the goal accepted in all eight runs, with no "command not found" line.
Fake ros2 CLIs cover the listing that misses once and a CLI that drops
the goal at once after "Sending goal:".
check-entities.sh piped the joint_states read straight into jq. jq
exits 0 on an error body, on empty input and on a reply without data,
so the "not available" fallback never ran and the script printed
"joint_names": null. With joint_state_broadcaster unloaded the gateway
answers 200 with an empty data object, which hit exactly that. The
script now prints the values only when the reply has joint names, and
the hint otherwise.

The smoke test unloads joint_state_broadcaster, waits until the gateway
serves no joint names, and requires check-entities.sh to finish with
the hint in section 5 and no null anywhere; it then loads the
controller again and waits for the data to return. With data, section
5 must show the joint names the API returns. The exit trap reloads the
controller if a run stops halfway.
… and sub-path fault sources

The run started right after /health must exit 0 and print values in
sections 5-8; a stop with "Sensor data not available" no longer passes.

A new section restarts the demo, fails the IMU before anything reads it
and requires check-demo.sh to name the IMU as having no data, print no
null, still print LiDAR and GPS values and run sections 9-12. It times the
readiness wait at the default and with DATA_WAIT_SEC 0 and 2 against the
limit plus one request.

Another section reports faults through the fault manager's report_fault
service: one from /processing/anomaly_detector/imu_sim, which sections
10-12 must show on apps/anomaly-detector without null, and one from
/processing/anomaly_detector_extra, which must not map to any App.
…sub-path fault sources

check-demo.sh waited for a message from every sensor and gave up with
exit 1, so after inject-failure.sh the fault sections never ran. The wait
counted passes, and a read of a sensor without a message blocks for the
gateway's sample timeout, so a 30 s wait took 66 s.

The wait now uses wall-clock time: up to DATA_WAIT_SEC (default 30,
settable) for the gateway to link each sensor node and up to 5 s for its
first message, each request bounded to 3 s. It names the sensors it waits
for and the ones left without data, whose sections print a message
and no null fields.

Sections 10-12 now pick the App whose node is the fault's reporting
source or a path segment above it, so anomaly detector faults reported as
/processing/anomaly_detector/<sensor> get the detail and bulk-data
sections. Section 12 says so when no rosbag is listed, and section 8 when
no configuration is.
…trigger across a navigation failure

With the Gazebo server stopped, /scan keeps its publisher but gets no
message. Before anything reads /scan, the test stops the simulator,
confirms the scan read has no data, and requires check-entities.sh to
print the LiDAR hint and no null. The later run with data must print the
scan's angle and range limits as a direct read returns them.

The run also creates the trigger with setup-triggers.sh and watches it
with watch-triggers.sh across inject-nav-failure. A fault reported under
the anomaly detector's own node path must reach the watcher, and no
NAVIGATION_GOAL_* event may, as the README states.
…tatus

The watcher that stops a fault manager between the two clears exited 1
when restore-normal ended without a pause it saw, and under errexit the
bare `wait` on it aborted the suite before the setup failure was recorded,
skipping the remaining ECUs and sections. The miss is now recorded as that
setup failure and the run continues; the fault manager is resumed either
way. The watcher prints "armed" before it polls and the script starts only
after that, and it exits as soon as the script ends without a pause.

The failure message check accepted any three digits, so a message naming a
status the clear never got passed. The test now measures what a fault
clear gets from the ECU gateway in the same state, requires it to be non-2xx,
and requires the error and the stderr of a direct run to name exactly that
status.
When only the second clear fails, the first one may already have removed
the ECU's faults, so "the faults stay" was not always true. The README now
says that the second clear decides the result, that the error names the
status it got, and that the ECU may still hold faults.
jq exits 0 on an error body and on empty input, so the fallback after
`|| echo` never ran and a scan read without data printed "angle_min":
null. check-entities.sh now prints the scan values only when the read
carries data, and the hint otherwise.
The README said the trigger fires on any new or updated fault. It fires
for the localization fault inject-localization-failure.sh causes, and
not for the navigation goal faults inject-nav-failure.sh causes, which
only show in the fault list.
The in-container move-arm.sh runs stopped the ros2 daemon and ignored
the result, so a run with the daemon still up counted as a cold run. A
warm daemon lets an implementation with a single `ros2 action list`
pass. Each run now records `ros2 daemon status` in a file in the
container; move-arm.sh starts only if it reports the daemon not
running, and a run without that proof is recorded as a failure.
run_move_arm_inside runs move-arm.sh only when its setup succeeds.

The check-entities.sh check with data compared only the joint names. It
now also requires positions and velocities as arrays of numbers, one
per joint name the API returns. The values move with the arm, so they
are checked by type and count only.
…adiness wait is needed on every run

The first check-demo.sh run after /health races the linking, and on a
fast start a copy without any wait finds all data at its first read and
passes. A new section restarts the demo, stops lidar_sim with SIGSTOP
before anything reads its topic and resumes it 2 s into the run. The run
must exit 0, print values in sections 5-8 and say it waits for lidar-sim.
The EXIT trap resumes the node as well.

The DATA_WAIT_SEC sweep adds 08, a whole number of seconds with a leading
zero, which must wait at least 3 s and at most 8 s plus one request.
The digit-only check accepts 08, and bash then reads it as an invalid
octal number: the deadline stayed empty and check-demo.sh skipped the
readiness wait. The value is now converted in base 10 after the check.
…n inject and require a successful scan read

The watcher was guarded only by a 2 s sleep before inject-nav-failure,
and its liveness was proven after the navigation fault. It is now proven
before the inject as well: a fault reported from the anomaly detector's
own node path is repeated every 2 s until its event arrives. Only a
watcher live before and after the inject counts for "no NAVIGATION_GOAL_*
event"; otherwise that check fails as not checked.

The no-scan setup discarded the HTTP status, and an error body has no
.data either, so a failed read counted as the no-data state. The setup now
requires a successful read with an empty data object.
… no-data setup

The README says an accepted goal is never sent again and a missing
result ends as "Failed: <pose> (status: UNKNOWN)", but no test reached
that path: every fake lost the goal before "Goal accepted with ID", and
every real accepted goal reached a final status. The fake ros2 has a
mode that prints "Goal accepted with ID" and then hangs; move-arm.sh
must send the goal once, print no resend line, and exit non-zero with
status UNKNOWN.

The joint_states no-data setup accepted any read without joint names,
including a 404, a 500 or a connection failure, so the hint check could
run against an error response. It now passes only when
the read answers HTTP 200 with an empty data object, and it reports the
status and body otherwise.
…w and resume one while the wait goes on

The link wait was never engaged: natural linking takes about 2 s and the
held LiDAR stays linked, so a script that waits only 5 s passed. A new
section restarts the demo, holds the linked IMU, takes lidar_sim off the
ROS graph and checks the gateway has unlinked it. The IMU resumes 7 s and
lidar_sim returns with the launch's parameters 9 s into the run. The run
must wait for the lidar-sim link, exit 0 and print values in sections
5-8, and must not report the IMU without a message once it sent data
while the wait went on. The EXIT trap resumes and restores both nodes.

The failed-IMU run also checks that the time printed for the IMU matches
the measured wait. The SIGSTOP helpers now take the node name.
…message window

A sensor was dropped once its 5 s window ended, even while the wait went
on for another sensor. The summary then said it had no message for the
whole wait and that its data section showed no values, and the section,
which reads it again, printed values when it had started in the meantime.

Such a sensor stops holding the wait and is still read on each pass, so
a late first message is noticed. The summary gives the time each sensor
was read for and says the data sections read every sensor again.
…ts use

Under smoke_lib.sh's `set -euo pipefail`, a reader that exits before the
end of its input (awk `exit`, `head`) sends SIGPIPE to the writer, the
pipeline returns 141, and an assignment from it ends the suite. The
sensor smoke test died this way in readiness_wait_output, an awk that
stops at section 1 behind the full run output.

readiness_wait_output, the reported-IMU-seconds read and the TurtleBot3
trigger id read now use one awk on a here-string, so no writer sits
behind the early exit. The early readers left are in fail and echo
message arguments, whose status is not used.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Helper scripts in several demos print null, fail on a fresh start or watch the wrong entity

2 participants