Release of IncidentRelay 2.0, a system for organizing on-call duties and alert routing

The release of the IncidentRelay 2.0 project has been announced, which develops an open system for organizing on-call duties, alert routing, and incident management, deployed on a self-hosted server. The project is aimed at SRE, DevOps, and infrastructure teams that require a local alternative to cloud-based duty management platforms. The code is written in Python and is distributed under the MIT license.

IncidentRelay accepts events from Prometheus Alertmanager, Grafana Alerting, Zabbix, Sentry, LibreNMS, RMON, AWS SNS/CloudWatch, Datadog, Uptime Kuma, and arbitrary webhook handlers. After normalization, the event is linked to the service, team, and rotation, routing rules and escalation policies are applied, after which the notification is sent to the current on-call staff. Supported delivery methods include Mattermost, Slack, Telegram, Discord, Microsoft Teams, email, webhook, browser/PWA push, and voice call providers.

The main innovation of version 2.0 is the Event Orchestration mechanism, which provides a separate processing layer for events between incoming integrations and the alert lifecycle. Key changes include:

  • A visual rules editor for orchestration has been added. Rules can be applied globally or to individual services, grouped into nested conditions, and sequentially modify priority, severity level, tags, team, route, grouping method, notification and escalation policies;
  • Actions for suppressing, dropping, and pausing event processing, extracting values using regular expressions and JSON Path, splitting strings, creating variables, and transforming values have been implemented;
  • Orchestration configuration is divided into editable drafts and immutable published versions. Configuration validation, publishing, rollback to previous versions, and adding comments to changes are supported;
  • Modes for disabled, shadow, and active are provided. In shadow mode, the results of rule execution are saved for analysis but do not impact actual routing. For the transition period, compatibility modes legacy, hybrid, and orchestration are available;
  • Event simulation and replay tools have been added. They allow you to verify rules on a normalized event or the original integration payload without creating a real alert. Execution tracing, Explain data, and shadow mode metrics are provided for result analysis;
  • Reusable webhook actions have been implemented, executed asynchronously after the event is processed. Headers are stored in encrypted form, secrets are hidden in the API and logs, and access to private networks is disabled by default. Timeouts, retries, and a list of allowed internal addresses can be configured;
  • Integration with Uptime Kuma has been added, receiving standard webhook messages about monitor statuses. Normalization of UP and DOWN states, importance assignment, tag handling, and a new incoming route type uptime_kuma have been implemented;
  • For silences and maintenance window configurations, apply_to_existing and reactivate_on_end parameters have been introduced. The first allows applying suppression to already open alerts, while the second determines whether to resume processing them after the suppression period ends. Notifications, reminders, and escalation chains are paused and restored;
  • Separate read and modify permissions for personal API tokens have been introduced for groups, teams, users, routes, channels, services, rotations, policies, maintenance windows, heartbeat checks, SSO, and orchestrations. Legacy aggregate permissions resources:read, resources:write, and * are retained for backward compatibility;
  • An audit log interface with filtering and pagination has been added. The log records operations related to orchestrations, webhook actions, silences, and maintenance windows, with confidential values removed from the displayed data;
  • A dark theme and user-configurable language and appearance settings have been added to the web interface. French localization has been introduced, translations of orchestration and maintenance sections have been expanded, and the editing of SSO group mapping rules has been fixed;
  • The operation of heartbeat checks has been improved: the re-creation of overdue event notifications has been eliminated, the recovery signal now properly closes the alert and sends a resolution notification, and the representation of timestamps has been unified;
  • Documentation for deployment in Kubernetes using Helm has been prepared. A separate Slack Socket Mode handler has been added to the Helm chart, necessary for the operation of interactive confirmation buttons and incident closure.
  • The protection of outgoing HTTP requests, regular expressions, and sensitive data has been enhanced, centralized processing of UTC and time zones has been implemented, rotation schedule calculations have been revised, and coverage with automated tests has been expanded.

It is recommended to create a backup of the database before the update. Necessary schema changes are performed using the standard IncidentRelay migration mechanism. For gradual implementation of Event Orchestration, developers advise initially using shadow and hybrid modes, checking the execution traces of rules, and only then switching orchestration to active mode.

Source: opennet.ru

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster