The first stable release of IncidentRelay, a system for organizing on-call duties and alert routing

After five months of development, the release of IncidentRelay 1.1 is now live. This project enhances an open system for organizing on-call shifts, routing notifications, and incident management, which can be launched on a self-hosted server. IncidentRelay 1.1 is marked as the first stable release (the 1.0 branch was in beta). The project is aimed at SRE, DevOps, and infrastructure teams that need a locally deployable alternative to SaaS services for on-call management, escalation policies, and incident response. The project code is written in Python and is distributed under the MIT license.

IncidentRelay receives events from monitoring systems, maps them to services, teams, and rotations, and then delivers notifications to responsible on-call staff or teams. The system features on-call schedules, rotations, shift overrides, incident acknowledgment states (ACK/Resolve), reminders, escalations, temporary replacements for on-call staff, scheduled maintenance timing, and suppression of noisy alerts.

The main advantage of the project is complete control over the infrastructure and routing logic. IncidentRelay is deployed in its own environment, uses its own database, and allows for the clear separation of incoming routes, teams, rotations, and delivery channels. Incoming tokens belong to routes rather than channels, making it easier to understand which external source is authorized to send events to a specific team.

IncidentRelay supports receiving events from Prometheus Alertmanager, Grafana Alerting, Zabbix, Sentry, LibreNMS, RMON, AWS SNS/CloudWatch, and arbitrary webhooks. Notification channels include Mattermost, Slack, Telegram, Discord, Microsoft Teams, email, webhooks, browser/PWA push, and voice call providers. In Mattermost and Telegram, notifications can contain actions for acknowledgment and resolution, allowing incidents to be handled without switching to a separate interface.

The project can be launched via Docker Compose, an RPM package for Red Hat-based distributions, manually through systemd, or in Kubernetes using a Helm chart. For smaller installations, SQLite can be used, while for production systems and higher loads, PostgreSQL is recommended.

The new version includes the following changes:

  • Multi-level on-call rotations have been added with time limitations, layer priorities, and the consideration of temporary replacements;
  • A duty calendar, CalDAV, and ICS subscriptions for external calendars have been introduced;
  • Incident escalation policies with multi-step escalation chains have been implemented;
  • Alert groups, event aggregation, delayed notifications, and manual merging of related alerts have been added;
  • Windows for scheduled maintenance and "silent" alerts have been introduced;
  • The ability to add comments to alerts has been implemented;
  • Incident priorities (P1-P5) and automatic priority escalation based on importance have been added;
  • A service catalog, service dependencies, SLI/SLO, Service Impact History, and Business Services have been implemented;
  • Explain Trace has been introduced for routing breakdown: why an alert was included or excluded from the team, grouped, suppressed, or sent to a specific channel;
  • Checks (Heartbeats/dead-man-switch) for watchdog, backup, ETL, and other tasks where the lack of an expected signal is an issue have been added.



Source: opennet.ru
Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster