You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
When rosbag capture is on (snapshots.rosbag.enabled, off by default), the fault manager records a rosbag when a fault confirms. The default format is MCAP. Today these recordings are kept or deleted by three global settings:
max_total_storage_mb: after each new recording, the oldest recordings are deleted until the total fits. Fault code, severity and status play no role, so a recording of an active critical fault can be deleted before a newer recording of a healed warning.
max_bags_per_fault: one cap for every fault code. It keeps the newest N, so the first occurrence (often the one with the root cause) is lost.
auto_cleanup: deletes on clear, and only when max_bags_per_fault is 1. Nothing happens when a fault heals.
On a robot that works offline for a long time, disk space is limited and some faults matter much more than others. The configuration must say, per fault code, per pattern and per severity: how many recordings to keep, how important they are, what to do when the fault heals or is cleared, and how long to keep them. A new recording of an important fault must always get disk space, by deleting less important recordings.
Offload is not part of this. The robot keeps and deletes recordings by its own rules, and a client that wants copies downloads them through the Bulk Data API.
Some related gaps block this today:
The quota adds up the sizes written into the rows when each bag closed. It never checks the disk again. Rows of bags that are gone (the default storage_path is the temp directory) keep counting until a cleanup deletes the row, or until a download or a bulk-data listing finds the directory missing. Other data on the same filesystem is not seen.
With the default policy, the capture queue drops the newest capture when it is full, even when that capture is the most important one.
The fault manager reads its parameters once at startup and has no parameter callback, so a later change has no effect.
Proposed solution
Configuration
All settings are parameters of the fault manager node, under snapshots.rosbag.retention:
fault_manager:
ros__parameters:
healing_enabled: true # needed for the HEALED rulessnapshots:
rosbag:
enabled: truemax_total_storage_mb: 4000# unchanged, the budget for all recordingsmax_bags_per_fault: 1# unchanged, default of max_bags when no rule sets itauto_cleanup: true # unchanged, default of on_cleared when no rule sets itretention:
min_free_disk_mb: 1024sweep_interval_sec: 60defaults:
priority: 10on_healed: keepon_cleared: keepmax_age_sec:
healed: 604800cleared: 86400severity:
CRITICAL:
priority: 90max_bags: 5keep_first: truemax_age_sec:
healed: 0cleared: 0ERROR:
priority: 60max_bags: 3WARN:
priority: 30max_bags: 2INFO:
priority: 5on_healed: deletefault_codes:
MOTOR_OVERHEAT:
priority: 95max_bags: 10keep_first: truepattern_order: ["gnss", "diag"]patterns:
gnss:
match: "GNSS_.*"max_bags: 10keep_first: truediag:
match: "DIAG_.*"priority: 1on_cleared: delete
Without any retention parameter, the behaviour is the same as today.
Parameter
Default
Meaning
retention.min_free_disk_mb
0
Free space that must stay on the filesystem of storage_path. Below it, recordings are deleted in eviction order. 0 turns the check off.
retention.sweep_interval_sec
60
How often age limits and free space are checked. The timer exists only when a rule has a max age or min_free_disk_mb is above 0.
retention.defaults.*
none
Rule fields used when no other level sets them.
retention.severity.<SEVERITY>.*
none
Rule per severity: INFO, WARN, ERROR, CRITICAL.
retention.fault_codes.<CODE>.*
none
Rule per fault code. <CODE> is the full code, dots included.
retention.patterns.<name>.*
none
Rule for fault codes that match the regex in match. <name> is only a label.
retention.pattern_order
[]
Pattern names in the order they are tried. Every pattern must be listed.
Fields of a rule:
Field
Values
Default when no level sets it
Meaning
priority
integer 0-1000
0
Higher is more important. Decides what is deleted first when space is needed.
max_bags
integer >= 0
max_bags_per_fault
Recordings kept per fault code. 0 means no cap, only the space limits apply.
keep_first
true / false
false
With max_bags >= 2, keep the first recording plus the max_bags - 1 newest. Ignored with a warning when max_bags is below 2.
on_healed
keep / delete
keep
What happens to the fault's recordings when it becomes HEALED. HEALED needs healing_enabled: true (default false).
on_cleared
keep / delete
delete when auto_cleanup is true and max_bags is 1, else keep (today's behaviour)
What happens to the fault's recordings when it becomes CLEARED.
max_age_sec
keys active, healed, cleared; seconds, 0 = no limit
all 0
Delete the fault's recordings when the fault is in that status and its last FAILED report is older than the limit. active means PREFAILED, PREPASSED or CONFIRMED. A fault reported once and never again ages out while it is still CONFIRMED.
match
regex
none
Patterns only. The fault codes this pattern applies to.
Rules
Matching
The rule for a fault is built from up to four levels: its fault_codes entry, the first pattern in pattern_order that matches the code, the severity entry, and defaults. Each field comes from the first level that sets it. The rule uses the fault's current severity and status. The fault manager keeps the highest severity ever reported for a code, so a fault that escalates moves all its recordings to the higher rule.
Example: MOTOR_OVERHEAT with severity ERROR gets priority: 95, max_bags: 10, keep_first: true from its code entry, and on_healed, on_cleared and max_age_sec from defaults.
A recording shared by a burst of faults takes the highest priority among its faults. It counts as active if any of its faults is active, else as healed if any of them is healed.
flowchart LR
F["Fault code,<br/>severity, status"] --> C["fault_codes.CODE"]
C -->|field not set| P["first match<br/>in pattern_order"]
P -->|field not set| S["severity.SEVERITY"]
S -->|field not set| D["defaults"]
D -->|field not set| B["built-in default"]
Loading
Disk space
Space is needed when the total is above max_total_storage_mb or free space is below min_free_disk_mb. Whole recordings are then deleted, one at a time: lowest priority first; inside one priority, recordings of cleared faults, then healed, then active; inside that, oldest first. Deletion stops when the limits are met or nothing more can be deleted, and the second case is logged.
A new recording never deletes a recording with a higher priority. This also holds for the space check before the recording opens. If no space can be made, the new recording is dropped and the drop is logged. When other data fills the disk and no recording is being made, there is no priority limit.
max_bags limits the recordings per fault code. Over the cap, the oldest recording is removed, or the second oldest with keep_first.
When the capture queue is full, the capture with the lowest priority is dropped. Pending captures run in priority order.
All recordings are written under storage_path, so they sit on one filesystem, and min_free_disk_mb checks that filesystem. At startup, rows of bags that no longer exist are removed. If none of the recorded bags is found, storage_path is treated as unavailable (for example a disk that is not mounted): nothing is removed and an error is logged.
flowchart TD
A([Fault confirmed]) --> Q{Capture queue full?}
Q -->|no| R[Capture waits in the queue,<br/>highest priority runs first]
Q -->|yes| L{New capture has the<br/>lowest priority?}
L -->|yes| X1([Capture dropped, logged])
L -->|no| O[Drop the lowest-priority<br/>pending capture]
O --> R
R --> W[Write the recording, store it,<br/>apply max_bags and keep_first]
W --> S{Over max_total_storage_mb or<br/>below min_free_disk_mb?}
S -->|no| K([Recording kept])
S -->|yes| C{Another recording with the<br/>same or lower priority?}
C -->|yes| E[Delete that whole recording:<br/>lowest priority, then cleared, healed,<br/>active, then oldest]
E --> S
C -->|no| X2([New recording deleted, logged])
Loading
Status and age
on_healed and on_cleared run at the status change. This includes faults that the fault manager turns from HEALED into CLEARED at startup when healing is disabled.
max_age_sec counts from the fault's last FAILED report (last_occurred). A new FAILED report starts the age again.
The recordings of a fault follow its status. Rules 3-5 can delete a recording in any status.
stateDiagram-v2
direction LR
[*] --> Active: fault confirmed
Active --> Healed: heals<br/>(healing_enabled)
Active --> Cleared: cleared
Healed --> Cleared: cleared
Healed --> Active: fails again
Cleared --> Active: fails again
Active --> Deleted: max_age_sec.active
Healed --> Deleted: on_healed delete<br/>or max_age_sec.healed
Cleared --> Deleted: on_cleared delete<br/>or max_age_sec.cleared
Deleted --> [*]
Loading
Changes
Each retention parameter can be changed while the node runs, for example with PUT /api/v1/apps/fault_manager/configurations/snapshots.rosbag.retention.fault_codes.MOTOR_OVERHEAT.priority on the gateway. A bad value is rejected with HTTP 400 and the reason. A valid value applies to existing recordings at once. Adding or removing a rule needs a restart. An invalid retention setting at startup stops the fault manager with an error, because a partial policy could delete recordings that the configuration meant to keep.
Additional context
Related: #620 (more than one recording per fault code), #649, #650, #651.
Summary
When rosbag capture is on (
snapshots.rosbag.enabled, off by default), the fault manager records a rosbag when a fault confirms. The default format is MCAP. Today these recordings are kept or deleted by three global settings:max_total_storage_mb: after each new recording, the oldest recordings are deleted until the total fits. Fault code, severity and status play no role, so a recording of an active critical fault can be deleted before a newer recording of a healed warning.max_bags_per_fault: one cap for every fault code. It keeps the newest N, so the first occurrence (often the one with the root cause) is lost.auto_cleanup: deletes on clear, and only whenmax_bags_per_faultis1. Nothing happens when a fault heals.On a robot that works offline for a long time, disk space is limited and some faults matter much more than others. The configuration must say, per fault code, per pattern and per severity: how many recordings to keep, how important they are, what to do when the fault heals or is cleared, and how long to keep them. A new recording of an important fault must always get disk space, by deleting less important recordings.
Offload is not part of this. The robot keeps and deletes recordings by its own rules, and a client that wants copies downloads them through the Bulk Data API.
Some related gaps block this today:
storage_pathis the temp directory) keep counting until a cleanup deletes the row, or until a download or a bulk-data listing finds the directory missing. Other data on the same filesystem is not seen.Proposed solution
Configuration
All settings are parameters of the fault manager node, under
snapshots.rosbag.retention:Without any
retentionparameter, the behaviour is the same as today.retention.min_free_disk_mb0storage_path. Below it, recordings are deleted in eviction order.0turns the check off.retention.sweep_interval_sec60min_free_disk_mbis above0.retention.defaults.*retention.severity.<SEVERITY>.*INFO,WARN,ERROR,CRITICAL.retention.fault_codes.<CODE>.*<CODE>is the full code, dots included.retention.patterns.<name>.*match.<name>is only a label.retention.pattern_order[]Fields of a rule:
priority0-10000max_bags>= 0max_bags_per_fault0means no cap, only the space limits apply.keep_firsttrue/falsefalsemax_bags >= 2, keep the first recording plus themax_bags - 1newest. Ignored with a warning whenmax_bagsis below 2.on_healedkeep/deletekeepHEALED.HEALEDneedshealing_enabled: true(defaultfalse).on_clearedkeep/deletedeletewhenauto_cleanupis true andmax_bagsis1, elsekeep(today's behaviour)CLEARED.max_age_secactive,healed,cleared; seconds,0= no limit0activemeansPREFAILED,PREPASSEDorCONFIRMED. A fault reported once and never again ages out while it is stillCONFIRMED.matchRules
Matching
fault_codesentry, the first pattern inpattern_orderthat matches the code, theseverityentry, anddefaults. Each field comes from the first level that sets it. The rule uses the fault's current severity and status. The fault manager keeps the highest severity ever reported for a code, so a fault that escalates moves all its recordings to the higher rule.MOTOR_OVERHEATwith severityERRORgetspriority: 95,max_bags: 10,keep_first: truefrom its code entry, andon_healed,on_clearedandmax_age_secfromdefaults.flowchart LR F["Fault code,<br/>severity, status"] --> C["fault_codes.CODE"] C -->|field not set| P["first match<br/>in pattern_order"] P -->|field not set| S["severity.SEVERITY"] S -->|field not set| D["defaults"] D -->|field not set| B["built-in default"]Disk space
max_total_storage_mbor free space is belowmin_free_disk_mb. Whole recordings are then deleted, one at a time: lowest priority first; inside one priority, recordings of cleared faults, then healed, then active; inside that, oldest first. Deletion stops when the limits are met or nothing more can be deleted, and the second case is logged.max_bagslimits the recordings per fault code. Over the cap, the oldest recording is removed, or the second oldest withkeep_first.storage_path, so they sit on one filesystem, andmin_free_disk_mbchecks that filesystem. At startup, rows of bags that no longer exist are removed. If none of the recorded bags is found,storage_pathis treated as unavailable (for example a disk that is not mounted): nothing is removed and an error is logged.flowchart TD A([Fault confirmed]) --> Q{Capture queue full?} Q -->|no| R[Capture waits in the queue,<br/>highest priority runs first] Q -->|yes| L{New capture has the<br/>lowest priority?} L -->|yes| X1([Capture dropped, logged]) L -->|no| O[Drop the lowest-priority<br/>pending capture] O --> R R --> W[Write the recording, store it,<br/>apply max_bags and keep_first] W --> S{Over max_total_storage_mb or<br/>below min_free_disk_mb?} S -->|no| K([Recording kept]) S -->|yes| C{Another recording with the<br/>same or lower priority?} C -->|yes| E[Delete that whole recording:<br/>lowest priority, then cleared, healed,<br/>active, then oldest] E --> S C -->|no| X2([New recording deleted, logged])Status and age
on_healedandon_clearedrun at the status change. This includes faults that the fault manager turns fromHEALEDintoCLEAREDat startup when healing is disabled.max_age_seccounts from the fault's last FAILED report (last_occurred). A new FAILED report starts the age again.The recordings of a fault follow its status. Rules 3-5 can delete a recording in any status.
stateDiagram-v2 direction LR [*] --> Active: fault confirmed Active --> Healed: heals<br/>(healing_enabled) Active --> Cleared: cleared Healed --> Cleared: cleared Healed --> Active: fails again Cleared --> Active: fails again Active --> Deleted: max_age_sec.active Healed --> Deleted: on_healed delete<br/>or max_age_sec.healed Cleared --> Deleted: on_cleared delete<br/>or max_age_sec.cleared Deleted --> [*]Changes
PUT /api/v1/apps/fault_manager/configurations/snapshots.rosbag.retention.fault_codes.MOTOR_OVERHEAT.priorityon the gateway. A bad value is rejected with HTTP 400 and the reason. A valid value applies to existing recordings at once. Adding or removing a rule needs a restart. An invalid retention setting at startup stops the fault manager with an error, because a partial policy could delete recordings that the configuration meant to keep.Additional context
Related: #620 (more than one recording per fault code), #649, #650, #651.