Skip to content

fault_manager: configurable rosbag retention per fault code, pattern and severity #705

Description

@bburda

Summary

When rosbag capture is on (snapshots.rosbag.enabled, off by default), the fault manager records a rosbag when a fault confirms. The default format is MCAP. Today these recordings are kept or deleted by three global settings:

  • max_total_storage_mb: after each new recording, the oldest recordings are deleted until the total fits. Fault code, severity and status play no role, so a recording of an active critical fault can be deleted before a newer recording of a healed warning.
  • max_bags_per_fault: one cap for every fault code. It keeps the newest N, so the first occurrence (often the one with the root cause) is lost.
  • auto_cleanup: deletes on clear, and only when max_bags_per_fault is 1. Nothing happens when a fault heals.

On a robot that works offline for a long time, disk space is limited and some faults matter much more than others. The configuration must say, per fault code, per pattern and per severity: how many recordings to keep, how important they are, what to do when the fault heals or is cleared, and how long to keep them. A new recording of an important fault must always get disk space, by deleting less important recordings.

Offload is not part of this. The robot keeps and deletes recordings by its own rules, and a client that wants copies downloads them through the Bulk Data API.

Some related gaps block this today:

  • The quota adds up the sizes written into the rows when each bag closed. It never checks the disk again. Rows of bags that are gone (the default storage_path is the temp directory) keep counting until a cleanup deletes the row, or until a download or a bulk-data listing finds the directory missing. Other data on the same filesystem is not seen.
  • With the default policy, the capture queue drops the newest capture when it is full, even when that capture is the most important one.
  • The fault manager reads its parameters once at startup and has no parameter callback, so a later change has no effect.

Proposed solution

Configuration

All settings are parameters of the fault manager node, under snapshots.rosbag.retention:

fault_manager:
  ros__parameters:
    healing_enabled: true              # needed for the HEALED rules
    snapshots:
      rosbag:
        enabled: true
        max_total_storage_mb: 4000     # unchanged, the budget for all recordings
        max_bags_per_fault: 1          # unchanged, default of max_bags when no rule sets it
        auto_cleanup: true             # unchanged, default of on_cleared when no rule sets it
        retention:
          min_free_disk_mb: 1024
          sweep_interval_sec: 60
          defaults:
            priority: 10
            on_healed: keep
            on_cleared: keep
            max_age_sec:
              healed: 604800
              cleared: 86400
          severity:
            CRITICAL:
              priority: 90
              max_bags: 5
              keep_first: true
              max_age_sec:
                healed: 0
                cleared: 0
            ERROR:
              priority: 60
              max_bags: 3
            WARN:
              priority: 30
              max_bags: 2
            INFO:
              priority: 5
              on_healed: delete
          fault_codes:
            MOTOR_OVERHEAT:
              priority: 95
              max_bags: 10
              keep_first: true
          pattern_order: ["gnss", "diag"]
          patterns:
            gnss:
              match: "GNSS_.*"
              max_bags: 10
              keep_first: true
            diag:
              match: "DIAG_.*"
              priority: 1
              on_cleared: delete

Without any retention parameter, the behaviour is the same as today.

Parameter Default Meaning
retention.min_free_disk_mb 0 Free space that must stay on the filesystem of storage_path. Below it, recordings are deleted in eviction order. 0 turns the check off.
retention.sweep_interval_sec 60 How often age limits and free space are checked. The timer exists only when a rule has a max age or min_free_disk_mb is above 0.
retention.defaults.* none Rule fields used when no other level sets them.
retention.severity.<SEVERITY>.* none Rule per severity: INFO, WARN, ERROR, CRITICAL.
retention.fault_codes.<CODE>.* none Rule per fault code. <CODE> is the full code, dots included.
retention.patterns.<name>.* none Rule for fault codes that match the regex in match. <name> is only a label.
retention.pattern_order [] Pattern names in the order they are tried. Every pattern must be listed.

Fields of a rule:

Field Values Default when no level sets it Meaning
priority integer 0-1000 0 Higher is more important. Decides what is deleted first when space is needed.
max_bags integer >= 0 max_bags_per_fault Recordings kept per fault code. 0 means no cap, only the space limits apply.
keep_first true / false false With max_bags >= 2, keep the first recording plus the max_bags - 1 newest. Ignored with a warning when max_bags is below 2.
on_healed keep / delete keep What happens to the fault's recordings when it becomes HEALED. HEALED needs healing_enabled: true (default false).
on_cleared keep / delete delete when auto_cleanup is true and max_bags is 1, else keep (today's behaviour) What happens to the fault's recordings when it becomes CLEARED.
max_age_sec keys active, healed, cleared; seconds, 0 = no limit all 0 Delete the fault's recordings when the fault is in that status and its last FAILED report is older than the limit. active means PREFAILED, PREPASSED or CONFIRMED. A fault reported once and never again ages out while it is still CONFIRMED.
match regex none Patterns only. The fault codes this pattern applies to.

Rules

Matching

  1. The rule for a fault is built from up to four levels: its fault_codes entry, the first pattern in pattern_order that matches the code, the severity entry, and defaults. Each field comes from the first level that sets it. The rule uses the fault's current severity and status. The fault manager keeps the highest severity ever reported for a code, so a fault that escalates moves all its recordings to the higher rule.
    • Example: MOTOR_OVERHEAT with severity ERROR gets priority: 95, max_bags: 10, keep_first: true from its code entry, and on_healed, on_cleared and max_age_sec from defaults.
  2. A recording shared by a burst of faults takes the highest priority among its faults. It counts as active if any of its faults is active, else as healed if any of them is healed.
flowchart LR
    F["Fault code,<br/>severity, status"] --> C["fault_codes.CODE"]
    C -->|field not set| P["first match<br/>in pattern_order"]
    P -->|field not set| S["severity.SEVERITY"]
    S -->|field not set| D["defaults"]
    D -->|field not set| B["built-in default"]
Loading

Disk space

  1. Space is needed when the total is above max_total_storage_mb or free space is below min_free_disk_mb. Whole recordings are then deleted, one at a time: lowest priority first; inside one priority, recordings of cleared faults, then healed, then active; inside that, oldest first. Deletion stops when the limits are met or nothing more can be deleted, and the second case is logged.
  2. A new recording never deletes a recording with a higher priority. This also holds for the space check before the recording opens. If no space can be made, the new recording is dropped and the drop is logged. When other data fills the disk and no recording is being made, there is no priority limit.
  3. max_bags limits the recordings per fault code. Over the cap, the oldest recording is removed, or the second oldest with keep_first.
  4. When the capture queue is full, the capture with the lowest priority is dropped. Pending captures run in priority order.
  5. All recordings are written under storage_path, so they sit on one filesystem, and min_free_disk_mb checks that filesystem. At startup, rows of bags that no longer exist are removed. If none of the recorded bags is found, storage_path is treated as unavailable (for example a disk that is not mounted): nothing is removed and an error is logged.
flowchart TD
    A([Fault confirmed]) --> Q{Capture queue full?}
    Q -->|no| R[Capture waits in the queue,<br/>highest priority runs first]
    Q -->|yes| L{New capture has the<br/>lowest priority?}
    L -->|yes| X1([Capture dropped, logged])
    L -->|no| O[Drop the lowest-priority<br/>pending capture]
    O --> R
    R --> W[Write the recording, store it,<br/>apply max_bags and keep_first]
    W --> S{Over max_total_storage_mb or<br/>below min_free_disk_mb?}
    S -->|no| K([Recording kept])
    S -->|yes| C{Another recording with the<br/>same or lower priority?}
    C -->|yes| E[Delete that whole recording:<br/>lowest priority, then cleared, healed,<br/>active, then oldest]
    E --> S
    C -->|no| X2([New recording deleted, logged])
Loading

Status and age

  1. on_healed and on_cleared run at the status change. This includes faults that the fault manager turns from HEALED into CLEARED at startup when healing is disabled.
  2. max_age_sec counts from the fault's last FAILED report (last_occurred). A new FAILED report starts the age again.

The recordings of a fault follow its status. Rules 3-5 can delete a recording in any status.

stateDiagram-v2
    direction LR
    [*] --> Active: fault confirmed
    Active --> Healed: heals<br/>(healing_enabled)
    Active --> Cleared: cleared
    Healed --> Cleared: cleared
    Healed --> Active: fails again
    Cleared --> Active: fails again
    Active --> Deleted: max_age_sec.active
    Healed --> Deleted: on_healed delete<br/>or max_age_sec.healed
    Cleared --> Deleted: on_cleared delete<br/>or max_age_sec.cleared
    Deleted --> [*]
Loading

Changes

  1. Each retention parameter can be changed while the node runs, for example with PUT /api/v1/apps/fault_manager/configurations/snapshots.rosbag.retention.fault_codes.MOTOR_OVERHEAT.priority on the gateway. A bad value is rejected with HTTP 400 and the reason. A valid value applies to existing recordings at once. Adding or removing a rule needs a restart. An invalid retention setting at startup stops the fault manager with an error, because a partial policy could delete recordings that the configuration meant to keep.

Additional context

Related: #620 (more than one recording per fault code), #649, #650, #651.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions