Summary
The fault store keys a fault by fault_code alone (fault_code TEXT PRIMARY KEY) and folds each report's
source_id into a reporting_sources set on that one row, which carries a single status, debounce_counter
and occurrence_count. When more than one entity raises the same fault_code, the reports collapse onto one
record.
Two consequences follow:
- Per-entity lifecycle is lost. A code raised for two different subjects becomes one record with one status,
so independent histories, occurrence counts and freeze frames cannot be represented.
- A clear masks the others.
ClearFault and the EVENT_PASSED heal path both match by fault_code and act on
the whole row, so a clear or heal from one reporter ends the shared record while another reporter's condition
is still active. ClearFault.srv carries no source, so the clear cannot be scoped.
Example: two nodes report COMMS_LOST. The first recovers and clears its fault with ClearFault. The record
is now CLEARED although the second node is still disconnected, and nothing on the API shows that it is.
The gateway inherits the problem. A per-entity fault route (/apps/{id}/faults/{code}) cannot address the
entity's own record when another entity shares the code.
Proposed solution
Make the fault identity (fault_code, owner), where owner is the source_id the report arrived with, the
entity the fault is about.
- Key the store on
(fault_code, owner) and track status, debounce, occurrence count, freeze frames, snapshots,
rosbag links and audit rows per record.
- Scope
ClearFault, GetFault, GetSnapshots, GetRosbag and the EVENT_PASSED heal per record by adding a
trailing source_id to those requests. An empty value means unscoped and is refused as ambiguous when several
records share the code. Fault.msg and ReportFault.srv keep their shape, and reporting_sources carries the
record's owner.
list_faults returns one item per record, and muting hides a record rather than a code.
- The gateway carries the owner on the wire and resolves a per-entity fault route to the entity's own record,
answering 409 when the entity holds several records of the code.
- Existing stores migrate on open, taking each row's owner from its first reporting source, with an
unrecoverable row owned by legacy.
Acceptance:
- Two entities raising the same
fault_code are two independent records with independent lifecycle, debounce,
severity, freeze frames and audit rows.
- A clear or heal from one owner leaves another owner's active fault CONFIRMED.
- An unscoped clear of a code with several records is refused and clears nothing.
- A store written before the change opens and migrates without losing rows.
Additional context
This is a breaking change: service requests gain a trailing field, two records replace one aggregate, and an
unscoped clear of a shared code is refused. Plugins that implement FaultProvider need a way to receive the
record owner, which means a plugin API version bump. Faults must carry a stable, unique owner id per subject, so
a provider that reports for external devices should supply a stable id per device rather than a shared constant.
Summary
The fault store keys a fault by
fault_codealone (fault_code TEXT PRIMARY KEY) and folds each report'ssource_idinto areporting_sourcesset on that one row, which carries a singlestatus,debounce_counterand
occurrence_count. When more than one entity raises the samefault_code, the reports collapse onto onerecord.
Two consequences follow:
so independent histories, occurrence counts and freeze frames cannot be represented.
ClearFaultand theEVENT_PASSEDheal path both match byfault_codeand act onthe whole row, so a clear or heal from one reporter ends the shared record while another reporter's condition
is still active.
ClearFault.srvcarries no source, so the clear cannot be scoped.Example: two nodes report
COMMS_LOST. The first recovers and clears its fault withClearFault. The recordis now CLEARED although the second node is still disconnected, and nothing on the API shows that it is.
The gateway inherits the problem. A per-entity fault route (
/apps/{id}/faults/{code}) cannot address theentity's own record when another entity shares the code.
Proposed solution
Make the fault identity
(fault_code, owner), whereowneris thesource_idthe report arrived with, theentity the fault is about.
(fault_code, owner)and track status, debounce, occurrence count, freeze frames, snapshots,rosbag links and audit rows per record.
ClearFault,GetFault,GetSnapshots,GetRosbagand theEVENT_PASSEDheal per record by adding atrailing
source_idto those requests. An empty value means unscoped and is refused as ambiguous when severalrecords share the code.
Fault.msgandReportFault.srvkeep their shape, andreporting_sourcescarries therecord's owner.
list_faultsreturns one item per record, and muting hides a record rather than a code.answering 409 when the entity holds several records of the code.
unrecoverable row owned by
legacy.Acceptance:
fault_codeare two independent records with independent lifecycle, debounce,severity, freeze frames and audit rows.
Additional context
This is a breaking change: service requests gain a trailing field, two records replace one aggregate, and an
unscoped clear of a shared code is refused. Plugins that implement
FaultProviderneed a way to receive therecord owner, which means a plugin API version bump. Faults must carry a stable, unique owner id per subject, so
a provider that reports for external devices should supply a stable id per device rather than a shared constant.