Skip to content

SLA Management

SLA Management lets you codify your uptime and availability expectations as named policies, evaluate actual device and network metrics against those expectations continuously, and act on any deviation before it becomes a user-visible problem.

Policies attach to a scope - an organization, site, site group, device group, SSID, camera, or NVR - and carry a set of metric thresholds. Every five minutes Celery evaluates every active policy for your organization and raises a breach when a threshold is violated. You acknowledge, track, and resolve breaches from the same interface.


The Celery beat task sla-evaluate-all runs every five minutes on the metrics queue (rate-limited to 2 per minute). It calls SLAMonitoringService.evaluate_all_policies for your organization, which:

  1. Loads every active policy and its thresholds dict (metric_name → threshold_value).
  2. For each non-null threshold, computes the deviation percentage against the current actual.
  3. Calls _is_threshold_violated and, on a violation, persists an SLABreach row and publishes sla.breach.created on the event bus.
  4. On recovery, publishes sla.breach.resolved and updates the breach status automatically.

You can also trigger evaluation manually at any time (see Endpoints).

The sla.breach.created event is always published at HIGH priority. The sla.breach.acknowledged event priority reflects breach severity: critical → CRITICAL, warning → NORMAL (default fallback). The sla.breach.resolved event is published at NORMAL priority. Breach severity is either warning (deviation ≤ 20%) or critical (deviation > 20%).


Every policy targets exactly one scope. When scope is organization, scope_id must be omitted. For all other scopes except ssid, scope_id is the UUID of the target entity and is required. The ssid scope is an exception - scope_id may be omitted (the SSID name is stored in the separate scope_name field instead).

Scope scope_id required Typical use
organization No Org-wide baseline thresholds
site Yes - site UUID Per-location SLA
site_group Yes - site group UUID Regional / campus grouping
device_group Yes - device group UUID Per-fleet (e.g. all APs)
ssid Optional Wi-Fi availability per SSID name
camera Yes - camera UUID Per-camera uptime
nvr Yes - NVR UUID Per-NVR availability

Navigate to Enterprise → SLA in the UI and click New Policy, or use the API directly.

Required body fields:

Field Type Notes
name string Human-readable label
scope string One of the scope values above
scope_id UUID or omit Required for all scopes except organization
thresholds object { "uptime_percent_min": 99.9, "latency_ms_max": 50 } - strictly typed; valid keys: uptime_percent_min, latency_ms_max, packet_loss_percent_max, health_score_min, client_satisfaction_min, error_rate_max

Optional fields:

Field Type Notes
description string Free-text explanation
status string active (default), disabled, or draft; set disabled to suspend evaluation without deleting - PATCH only, not accepted on POST

Example request:

POST /api/v1/sla/policies
Content-Type: application/json
Authorization: Bearer <token>
{
"name": "Core Network Uptime",
"scope": "site",
"scope_id": "a1b2c3d4-...",
"thresholds": {
"uptime_percent_min": 99.5,
"packet_loss_percent_max": 1.0,
"latency_ms_max": 100
}
}

All paths are under the prefix /api/v1/sla. Both config:read and config:write permissions are required at the appropriate tier - see the RBAC reference for the role-to-permission mapping.

Method Path Purpose Permission
GET /api/v1/sla/summary Org-wide compliance summary (pass site_id to scope) config:read
GET /api/v1/sla/policies List policies - filter by site_id, scope, status; limit≤200 config:read
POST /api/v1/sla/policies Create a policy config:write
GET /api/v1/sla/policies/{policy_id} Retrieve one policy config:read
PATCH /api/v1/sla/policies/{policy_id} Update (scope re-verified on change) config:write
DELETE /api/v1/sla/policies/{policy_id} Delete a policy config:write
Method Path Purpose Permission
GET /api/v1/sla/breaches List breaches - filter by site_id, policy_id, status; limit≤200 config:read
POST /api/v1/sla/breaches/{breach_id}/acknowledge Acknowledge with an optional note config:write
POST /api/v1/sla/evaluate Manually trigger evaluation for your org config:write

When a policy is violated, the platform:

  1. Persists an SLABreach row with the policy, scope, metric, actual value, threshold, and severity.
  2. Publishes sla.breach.created on the event bus - any connected alert rule or notification provider will fire if configured to listen for this event type.
  3. Updates the breach status to resolved and publishes sla.breach.resolved automatically once the metric recovers.

Acknowledge a breach to record that a human has reviewed it:

POST /api/v1/sla/breaches/{breach_id}/acknowledge
Content-Type: application/json
{
"notes": "Investigating upstream ISP packet loss."
}

The acknowledgement publishes sla.breach.acknowledged on the event bus.

Filtering the breach list:

GET /api/v1/sla/breaches?status=active&site_id=<uuid>&limit=50

Valid status values: active, acknowledged, resolved.


GET /api/v1/sla/summary (optionally scoped with ?site_id=<uuid>) returns an org-wide view showing:

  • Total active policies
  • Policies currently in breach
  • Breach counts by severity
  • Recent breach history

Use this endpoint to power a management dashboard or feed into a periodic report.


The report engine at /api/v1/sla/reports generates on-demand SLA compliance reports in PDF or CSV format.

POST /api/v1/sla/reports/generate
Content-Type: application/json
{
"period_start": "2026-05-01T00:00:00Z",
"period_end": "2026-06-01T00:00:00Z",
"policy_ids": ["<uuid>", "<uuid>"],
"format": "pdf",
"title": "May 2026 SLA Report"
}

Constraints enforced by the server:

  • period_start must be before period_end
  • Period length must be 366 days or less
  • format must be pdf or csv

Once generated, download the file with:

GET /api/v1/sla/reports/{report_id}/download

Reports render before their database row is published. Missing files and legacy text files labeled as PDF require regeneration and return 409 to an authorized reader. Downloads never substitute JSON for the requested file. The resolved path must stay inside the report organization’s directory. Historical reports must also be within the reader’s current site grants; legacy reports without recorded site scope are unavailable to site-limited readers until regenerated.

Method Path Purpose Permission
POST /api/v1/sla/reports/generate Generate on demand config:read
GET /api/v1/sla/reports List generated reports (limit≤200) config:read
GET /api/v1/sla/reports/{report_id}/download Download the rendered file config:read
Method Path Purpose Permission
GET /api/v1/sla/report-schedules List schedules config:read
POST /api/v1/sla/report-schedules Create a schedule config:write
PUT /api/v1/sla/report-schedules/{schedule_id} Update a schedule config:write
DELETE /api/v1/sla/report-schedules/{schedule_id} Delete a schedule config:write

Schedule timing is configured through the schedule API. New schedules default to timezone: "UTC" and time_of_day: "09:00". Supply an installed IANA timezone, such as America/New_York, and a 24-hour HH:MM clock time. Invalid zones or clock values return 422; corrupt stored timing pauses execution rather than falling back to the server’s timezone.

Frequency Day selection Completed reporting period
weekly day_of_week: Monday=0 through Sunday=6; default Monday Seven local calendar days ending at midnight on the scheduled day
monthly day_of_month: 1-31; default 1; clamp to the month’s last day Previous complete calendar month
quarterly Selected month-day in January, April, July and October Previous complete calendar quarter

A day-of-month of 31 runs on February’s last day and returns to the 31st in March; it does not drift to a fixed 30-day cadence. When a clock time does not exist at a daylight-saving transition, the run shifts forward by that transition’s gap. When a time repeats, it runs at the first occurrence only. The local clock time remains fixed across offset changes; UTC next_run_at can change accordingly.

Worker delays do not extend the reporting window. If several runs were missed, the worker generates one report for the latest due interval and advances to the next future occurrence. It does not backfill every missed report. Use on-demand report generation for older intervals. An explicit next_run_at can override the first run; subsequent runs follow the calendar settings.

For example, this creates a Monday report at 08:30 New York time for the previous seven local days. Add explicit sla_policy_ids when using a site-limited account:

{
"name": "Weekly operations",
"frequency": "weekly",
"day_of_week": 0,
"timezone": "America/New_York",
"time_of_day": "08:30",
"enabled": true
}

The schedule authorization migration pauses existing schedules with no recorded approval (enabled=false, execution_error=approval_required). It preserves their names, policy filters, recipients and timestamps. The upgrade does not assume that a historical creator still has authority to generate those reports.

The calendar migration additionally pauses previously enabled schedules with execution_error=timing_review_required. Existing approval records, filters and next-run timestamps are preserved. Since legacy schedules did not store an IANA timezone, their clock time is copied from next_run_at in UTC (or 09:00 when no first run existed). Review the intended local timezone, weekday/month-day and reporting interval before re-enabling. Downgrading does not resume schedules.

An authorized operator should list schedules with GET /api/v1/sla/report-schedules, review each effective policy filter and its intended scope, and update it with PUT /api/v1/sla/report-schedules/{id}. Send {"enabled": true} only after that review; include corrected sla_policy_ids if necessary. An empty policy list means a dynamic organization-wide report and is unavailable to site-limited operators. A successful edit records the editing operator as the approver, clears the execution error, and re-enabling computes a future run time. The original creator remains in the audit fields. Disabled schedules remain disabled when edited without enabled: true.

Each execution verifies that the approver exists, is active, remains in the organization (or is a platform administrator), and still has config:write. It enforces the permissions and site scope approved for the schedule, checks current site grants, and verifies the explicit policy scope has not changed. Deleting the last site grant cannot expand a previously limited approval into organization-wide access. Promotion of the approver does not widen the recorded site scope. Missing or inaccessible explicit policies pause the schedule instead of producing a partial or broader report. The API exposes execution_error:

Value Operator action
approval_required Review and approve the legacy or malformed schedule.
timing_review_required Review timezone, clock time, calendar day and completed reporting period, then explicitly re-enable.
invalid_schedule_timing Correct the invalid stored timezone, clock or cadence and re-enable.
actor_unavailable Review the disabled/deleted account; have an active authorized operator approve the schedule.
permission_revoked Review the role change and have an authorized operator approve the intended work.
organization_changed Review the approver’s organization and assign a responsible operator in the correct organization.
policy_scope_changed Review deleted/moved policies, resource sites and current grants; correct and approve the filter.
generation_failed Inspect worker logs and database/report-volume health. The schedule remains enabled and due for retry.

Approval is a durable delegation by the operator, bounded by the approving credential’s permissions. Execution does not separately reauthenticate the original browser session or check the issuing API key’s expiry/revocation. Disabling the operator or removing the required role permission stops execution. A permission change during an already-running render is not an atomic revocation of that in-flight work.

Workers lock due schedules and skip rows claimed by another worker. A savepoint rolls back a failed schedule before the batch continues. Successful report rows and next-run timestamps commit together at the end of the worker batch. This is not a complete durable job/delivery ledger: interruption of the outer transaction can roll back that batch, and files written before interruption need retention and orphan cleanup. The batch is limited to 100 due schedules per invocation.

Scheduled reports generate downloadable files. File generation does not itself establish email acceptance, historical backfill or a durable ledger for every reporting operation. Email notices use the separate outbox described below.

Email notices are opt-in: schedules default to delivery_enabled: false, including upgraded schedules. To approve notices through the schedule API, configure a current SLA_REPORT_DELIVERY_EPOCH UUID, an HTTPS FRONTEND_URL, and a same-organization, enabled, verified SMTP or Gmail SMTP provider with certificate-verified STARTTLS. Set delivery_provider_id, supply 1–50 recipient addresses, and explicitly enable delivery. Every recipient must resolve to an active user in the schedule’s organization with permission and site coverage for the report. Arbitrary external mailing lists are not supported by this path.

Approval binds the provider configuration, HTTPS origin, recipients, scope and recovery epoch. Before sending, the worker checks current authorization and approval again. A provider, origin, recipient or permission change may require reapproval or cancel the notice. Recipient email preferences and quiet hours apply; provider rate-limit or availability failures defer sending rather than bypassing those checks.

The sla-deliver-scheduled-reports task runs every minute on metrics. A live scheduler and a worker consuming that queue are both required. Notices contain a sign-in link only: no report data, attachments, credentials or bearer download link is included. The recipient must still pass current authorization when opening the report.

Outbox state Meaning
pending / retry Waiting for a permitted attempt; delivery is not established.
sending The claim was committed before SMTP; do not assume the message can be safely retried.
accepted The SMTP server acknowledged submission; this is not an inbox or read receipt.
unknown Submission or restored state is ambiguous; no automatic replay occurs.
failed A terminal refusal, invalid configuration or retry/age limit stopped this notice.
cancelled / skipped Authorization, operator cancellation or recipient preferences prevented sending.

Operators with config:write and the report’s organization/site access can inspect GET /api/v1/sla/reports/{report_id}/deliveries. They can cancel an unclaimed pending or retry notice through POST /api/v1/sla/reports/{report_id}/deliveries/{delivery_id}/cancel. Accepted or uncertain messages cannot be recalled through that endpoint. Transport retries are bounded to five attempts and seven days; only outcomes known to be retryable are retried. Lost acknowledgements, expired sending leases and mismatched recovery epochs remain unknown for operator reconciliation.

After database recovery, follow the coordinated recovery guide for a new delivery epoch, sender fencing and explicit reapproval. Do not rewrite unknown outbox rows into retries or infer exactly-once delivery. Synthetic SMTP tests do not establish acceptance by your real relay or recipients; validate those separately before enabling notices.


Permission Who can hold it What it unlocks
config:read viewer and above (site-scoped by grant) Read policies, breaches, summary, reports
config:write site_admin and above Create/edit/delete policies; acknowledge breaches; trigger evaluation; create reports and schedules

Role-to-permission mapping is defined in the Roles and permissions reference.


Connecting SLA to other enterprise features

Section titled “Connecting SLA to other enterprise features”
  • Alert Rules - create a rule that fires on sla.breach.created to route breach notifications to any configured notification channel.
  • Health Dashboard - the SLAComplianceCard component on the health dashboard (/health) pulls from GET /api/v1/sla/summary and shows a high-level compliance rate alongside device health scores.
  • Event Correlation - SLA breach events appear in the correlation engine’s event stream; you can write a correlation rule that groups multiple simultaneous breaches into a single incident.
  • Notifications - configure an SMTP, Slack webhook, or Teams webhook provider under Enterprise → Notification Providers, then attach it to an alert rule targeting SLA breach events.

  • Every scope_id supplied at create or update time is verified to be owned by your organization before the operation is committed. Cross-tenant probes receive a 404.
  • All policy and breach reads are org-scoped - you cannot read another organization’s SLA data.
  • The manual /evaluate endpoint derives the target organization from the authenticated user’s JWT. It cannot be directed at another organization’s policies.
  • Multi-tenant isolation is enforced at the application layer. There is no PostgreSQL row-level security - isolation is the responsibility of the service and endpoint code. Passing internal tests does not replace an independent security assessment.

  • Alert Rules - wire breach events to notifications and auto-resolve workflows
  • Health Dashboard (/health) - view SLA compliance alongside device health scores
  • Event Correlation - group breach events into incidents for coordinated response
  • Notification Providers - configure SMTP, Slack, Teams, and webhook delivery

All product names, logos, and brands are property of their respective owners. FreeSDN is an independent project and is not affiliated with or endorsed by the vendors it integrates with. See Trademarks.