Skip to content

fault_manager: key faults by (fault_code, owner) so per-entity faults stay distinct and clears do not mask #698

Description

@bburda

Summary

The fault store keys a fault by fault_code alone (fault_code TEXT PRIMARY KEY) and folds each report's
source_id into a reporting_sources set on that one row, which carries a single status, debounce_counter
and occurrence_count. When more than one entity raises the same fault_code, the reports collapse onto one
record.

Two consequences follow:

  • Per-entity lifecycle is lost. A code raised for two different subjects becomes one record with one status,
    so independent histories, occurrence counts and freeze frames cannot be represented.
  • A clear masks the others. ClearFault and the EVENT_PASSED heal path both match by fault_code and act on
    the whole row, so a clear or heal from one reporter ends the shared record while another reporter's condition
    is still active. ClearFault.srv carries no source, so the clear cannot be scoped.

Example: two nodes report COMMS_LOST. The first recovers and clears its fault with ClearFault. The record
is now CLEARED although the second node is still disconnected, and nothing on the API shows that it is.

The gateway inherits the problem. A per-entity fault route (/apps/{id}/faults/{code}) cannot address the
entity's own record when another entity shares the code.


Proposed solution

Make the fault identity (fault_code, owner), where owner is the source_id the report arrived with, the
entity the fault is about.

  • Key the store on (fault_code, owner) and track status, debounce, occurrence count, freeze frames, snapshots,
    rosbag links and audit rows per record.
  • Scope ClearFault, GetFault, GetSnapshots, GetRosbag and the EVENT_PASSED heal per record by adding a
    trailing source_id to those requests. An empty value means unscoped and is refused as ambiguous when several
    records share the code. Fault.msg and ReportFault.srv keep their shape, and reporting_sources carries the
    record's owner.
  • list_faults returns one item per record, and muting hides a record rather than a code.
  • The gateway carries the owner on the wire and resolves a per-entity fault route to the entity's own record,
    answering 409 when the entity holds several records of the code.
  • Existing stores migrate on open, taking each row's owner from its first reporting source, with an
    unrecoverable row owned by legacy.

Acceptance:

  • Two entities raising the same fault_code are two independent records with independent lifecycle, debounce,
    severity, freeze frames and audit rows.
  • A clear or heal from one owner leaves another owner's active fault CONFIRMED.
  • An unscoped clear of a code with several records is refused and clears nothing.
  • A store written before the change opens and migrates without losing rows.

Additional context

This is a breaking change: service requests gain a trailing field, two records replace one aggregate, and an
unscoped clear of a shared code is refused. Plugins that implement FaultProvider need a way to receive the
record owner, which means a plugin API version bump. Faults must carry a stable, unique owner id per subject, so
a provider that reports for external devices should supply a stable id per device rather than a shared constant.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions