SLA Management
SLA Management lets you codify your uptime and availability expectations as named policies, evaluate actual device and network metrics against those expectations continuously, and act on any deviation before it becomes a user-visible problem.
Policies attach to a scope - an organization, site, site group, device group, SSID, camera, or NVR - and carry a set of metric thresholds. Every five minutes Celery evaluates every active policy for your organization and raises a breach when a threshold is violated. You acknowledge, track, and resolve breaches from the same interface.
How evaluation works
Section titled “How evaluation works”The Celery beat task sla-evaluate-all runs every five minutes on the metrics queue (rate-limited to 2 per minute). It calls SLAMonitoringService.evaluate_all_policies for your organization, which:
- Loads every active policy and its
thresholdsdict (metric_name → threshold_value). - For each non-null threshold, computes the deviation percentage against the current actual.
- Calls
_is_threshold_violatedand, on a violation, persists anSLABreachrow and publishessla.breach.createdon the event bus. - On recovery, publishes
sla.breach.resolvedand updates the breach status automatically.
You can also trigger evaluation manually at any time (see Endpoints).
The sla.breach.created event is always published at HIGH priority. The sla.breach.acknowledged event priority reflects breach severity: critical → CRITICAL, warning → NORMAL (default fallback). The sla.breach.resolved event is published at NORMAL priority. Breach severity is either warning (deviation ≤ 20%) or critical (deviation > 20%).
Policy scopes
Section titled “Policy scopes”Every policy targets exactly one scope. When scope is organization, scope_id must be omitted. For all other scopes except ssid, scope_id is the UUID of the target entity and is required. The ssid scope is an exception - scope_id may be omitted (the SSID name is stored in the separate scope_name field instead).
| Scope | scope_id required |
Typical use |
|---|---|---|
organization |
No | Org-wide baseline thresholds |
site |
Yes - site UUID | Per-location SLA |
site_group |
Yes - site group UUID | Regional / campus grouping |
device_group |
Yes - device group UUID | Per-fleet (e.g. all APs) |
ssid |
Optional | Wi-Fi availability per SSID name |
camera |
Yes - camera UUID | Per-camera uptime |
nvr |
Yes - NVR UUID | Per-NVR availability |
Creating a policy
Section titled “Creating a policy”Navigate to Enterprise → SLA in the UI and click New Policy, or use the API directly.
Required body fields:
| Field | Type | Notes |
|---|---|---|
name |
string | Human-readable label |
scope |
string | One of the scope values above |
scope_id |
UUID or omit | Required for all scopes except organization |
thresholds |
object | { "uptime_percent_min": 99.9, "latency_ms_max": 50 } - strictly typed; valid keys: uptime_percent_min, latency_ms_max, packet_loss_percent_max, health_score_min, client_satisfaction_min, error_rate_max |
Optional fields:
| Field | Type | Notes |
|---|---|---|
description |
string | Free-text explanation |
status |
string | active (default), disabled, or draft; set disabled to suspend evaluation without deleting - PATCH only, not accepted on POST |
Example request:
POST /api/v1/sla/policiesContent-Type: application/jsonAuthorization: Bearer <token>
{ "name": "Core Network Uptime", "scope": "site", "scope_id": "a1b2c3d4-...", "thresholds": { "uptime_percent_min": 99.5, "packet_loss_percent_max": 1.0, "latency_ms_max": 100 }}Endpoints
Section titled “Endpoints”All paths are under the prefix /api/v1/sla. Both config:read and config:write permissions are required at the appropriate tier - see the RBAC reference for the role-to-permission mapping.
Policy management
Section titled “Policy management”| Method | Path | Purpose | Permission |
|---|---|---|---|
| GET | /api/v1/sla/summary |
Org-wide compliance summary (pass site_id to scope) |
config:read |
| GET | /api/v1/sla/policies |
List policies - filter by site_id, scope, status; limit≤200 |
config:read |
| POST | /api/v1/sla/policies |
Create a policy | config:write |
| GET | /api/v1/sla/policies/{policy_id} |
Retrieve one policy | config:read |
| PATCH | /api/v1/sla/policies/{policy_id} |
Update (scope re-verified on change) | config:write |
| DELETE | /api/v1/sla/policies/{policy_id} |
Delete a policy | config:write |
Breach management
Section titled “Breach management”| Method | Path | Purpose | Permission |
|---|---|---|---|
| GET | /api/v1/sla/breaches |
List breaches - filter by site_id, policy_id, status; limit≤200 |
config:read |
| POST | /api/v1/sla/breaches/{breach_id}/acknowledge |
Acknowledge with an optional note | config:write |
| POST | /api/v1/sla/evaluate |
Manually trigger evaluation for your org | config:write |
Tracking and acknowledging breaches
Section titled “Tracking and acknowledging breaches”When a policy is violated, the platform:
- Persists an
SLABreachrow with the policy, scope, metric, actual value, threshold, and severity. - Publishes
sla.breach.createdon the event bus - any connected alert rule or notification provider will fire if configured to listen for this event type. - Updates the breach status to
resolvedand publishessla.breach.resolvedautomatically once the metric recovers.
Acknowledge a breach to record that a human has reviewed it:
POST /api/v1/sla/breaches/{breach_id}/acknowledgeContent-Type: application/json
{ "notes": "Investigating upstream ISP packet loss."}The acknowledgement publishes sla.breach.acknowledged on the event bus.
Filtering the breach list:
GET /api/v1/sla/breaches?status=active&site_id=<uuid>&limit=50Valid status values: active, acknowledged, resolved.
Compliance summary
Section titled “Compliance summary”GET /api/v1/sla/summary (optionally scoped with ?site_id=<uuid>) returns an org-wide view showing:
- Total active policies
- Policies currently in breach
- Breach counts by severity
- Recent breach history
Use this endpoint to power a management dashboard or feed into a periodic report.
Reports and schedules
Section titled “Reports and schedules”The report engine at /api/v1/sla/reports generates on-demand SLA compliance reports in PDF or CSV format.
Generating a report
Section titled “Generating a report”POST /api/v1/sla/reports/generateContent-Type: application/json
{ "period_start": "2026-05-01T00:00:00Z", "period_end": "2026-06-01T00:00:00Z", "policy_ids": ["<uuid>", "<uuid>"], "format": "pdf", "title": "May 2026 SLA Report"}Constraints enforced by the server:
period_startmust be beforeperiod_end- Period length must be 366 days or less
formatmust bepdforcsv
Once generated, download the file with:
GET /api/v1/sla/reports/{report_id}/downloadReports render before their database row is published. Missing files and legacy text files labeled as PDF require regeneration and return 409 to an authorized reader. Downloads never substitute JSON for the requested file. The resolved path must stay inside the report organization’s directory. Historical reports must also be within the reader’s current site grants; legacy reports without recorded site scope are unavailable to site-limited readers until regenerated.
Report endpoints
Section titled “Report endpoints”| Method | Path | Purpose | Permission |
|---|---|---|---|
| POST | /api/v1/sla/reports/generate |
Generate on demand | config:read |
| GET | /api/v1/sla/reports |
List generated reports (limit≤200) |
config:read |
| GET | /api/v1/sla/reports/{report_id}/download |
Download the rendered file | config:read |
Schedule endpoints
Section titled “Schedule endpoints”| Method | Path | Purpose | Permission |
|---|---|---|---|
| GET | /api/v1/sla/report-schedules |
List schedules | config:read |
| POST | /api/v1/sla/report-schedules |
Create a schedule | config:write |
| PUT | /api/v1/sla/report-schedules/{schedule_id} |
Update a schedule | config:write |
| DELETE | /api/v1/sla/report-schedules/{schedule_id} |
Delete a schedule | config:write |
Calendar settings
Section titled “Calendar settings”Schedule timing is configured through the schedule API. New schedules default to
timezone: "UTC" and time_of_day: "09:00". Supply an installed IANA timezone,
such as America/New_York, and a 24-hour HH:MM clock time. Invalid zones or
clock values return 422; corrupt stored timing pauses execution rather than
falling back to the server’s timezone.
| Frequency | Day selection | Completed reporting period |
|---|---|---|
weekly |
day_of_week: Monday=0 through Sunday=6; default Monday |
Seven local calendar days ending at midnight on the scheduled day |
monthly |
day_of_month: 1-31; default 1; clamp to the month’s last day |
Previous complete calendar month |
quarterly |
Selected month-day in January, April, July and October | Previous complete calendar quarter |
A day-of-month of 31 runs on February’s last day and returns to the 31st in March;
it does not drift to a fixed 30-day cadence. When a clock time does not exist at
a daylight-saving transition, the run shifts forward by that transition’s gap.
When a time repeats, it runs at the first occurrence only. The local clock time
remains fixed across offset changes; UTC next_run_at can change accordingly.
Worker delays do not extend the reporting window. If several runs were missed,
the worker generates one report for the latest due interval and advances to the
next future occurrence. It does not backfill every missed report. Use on-demand
report generation for older intervals. An explicit next_run_at can override the
first run; subsequent runs follow the calendar settings.
For example, this creates a Monday report at 08:30 New York time for the previous
seven local days. Add explicit sla_policy_ids when using a site-limited account:
{ "name": "Weekly operations", "frequency": "weekly", "day_of_week": 0, "timezone": "America/New_York", "time_of_day": "08:30", "enabled": true}Approval, execution and upgrade behavior
Section titled “Approval, execution and upgrade behavior”The schedule authorization migration pauses existing schedules with no recorded
approval (enabled=false, execution_error=approval_required). It preserves
their names, policy filters, recipients and timestamps. The upgrade does not
assume that a historical creator still has authority to generate those reports.
The calendar migration additionally pauses previously enabled schedules with
execution_error=timing_review_required. Existing approval records, filters and
next-run timestamps are preserved. Since legacy schedules did not store an IANA
timezone, their clock time is copied from next_run_at in UTC (or 09:00 when no
first run existed). Review the intended local timezone, weekday/month-day and
reporting interval before re-enabling. Downgrading does not resume schedules.
An authorized operator should list schedules with GET /api/v1/sla/report-schedules,
review each effective policy filter and its intended scope, and update it with
PUT /api/v1/sla/report-schedules/{id}. Send {"enabled": true} only after that
review; include corrected sla_policy_ids if necessary. An empty policy list
means a dynamic organization-wide report and is unavailable to site-limited
operators. A successful edit records the editing operator as the approver,
clears the execution error, and re-enabling computes a future run time. The
original creator remains in the audit fields. Disabled schedules remain disabled
when edited without enabled: true.
Each execution verifies that the approver exists, is active, remains in the
organization (or is a platform administrator), and still has config:write.
It enforces the permissions and site scope approved for the schedule, checks
current site grants, and verifies the explicit policy scope has not changed.
Deleting the last site grant cannot expand a previously limited approval into
organization-wide access. Promotion of the approver does not widen the recorded
site scope. Missing or inaccessible explicit policies pause the schedule instead
of producing a partial or broader report. The API exposes execution_error:
| Value | Operator action |
|---|---|
approval_required |
Review and approve the legacy or malformed schedule. |
timing_review_required |
Review timezone, clock time, calendar day and completed reporting period, then explicitly re-enable. |
invalid_schedule_timing |
Correct the invalid stored timezone, clock or cadence and re-enable. |
actor_unavailable |
Review the disabled/deleted account; have an active authorized operator approve the schedule. |
permission_revoked |
Review the role change and have an authorized operator approve the intended work. |
organization_changed |
Review the approver’s organization and assign a responsible operator in the correct organization. |
policy_scope_changed |
Review deleted/moved policies, resource sites and current grants; correct and approve the filter. |
generation_failed |
Inspect worker logs and database/report-volume health. The schedule remains enabled and due for retry. |
Approval is a durable delegation by the operator, bounded by the approving credential’s permissions. Execution does not separately reauthenticate the original browser session or check the issuing API key’s expiry/revocation. Disabling the operator or removing the required role permission stops execution. A permission change during an already-running render is not an atomic revocation of that in-flight work.
Workers lock due schedules and skip rows claimed by another worker. A savepoint rolls back a failed schedule before the batch continues. Successful report rows and next-run timestamps commit together at the end of the worker batch. This is not a complete durable job/delivery ledger: interruption of the outer transaction can roll back that batch, and files written before interruption need retention and orphan cleanup. The batch is limited to 100 due schedules per invocation.
Scheduled reports generate downloadable files. File generation does not itself establish email acceptance, historical backfill or a durable ledger for every reporting operation. Email notices use the separate outbox described below.
Scheduled email notices
Section titled “Scheduled email notices”Email notices are opt-in: schedules default to delivery_enabled: false, including
upgraded schedules. To approve notices through the schedule API, configure a current
SLA_REPORT_DELIVERY_EPOCH UUID, an HTTPS FRONTEND_URL, and a same-organization,
enabled, verified SMTP or Gmail SMTP provider with certificate-verified STARTTLS.
Set delivery_provider_id, supply 1–50 recipient addresses, and explicitly enable
delivery. Every recipient must resolve to an active user in the schedule’s organization
with permission and site coverage for the report. Arbitrary external mailing lists are
not supported by this path.
Approval binds the provider configuration, HTTPS origin, recipients, scope and recovery epoch. Before sending, the worker checks current authorization and approval again. A provider, origin, recipient or permission change may require reapproval or cancel the notice. Recipient email preferences and quiet hours apply; provider rate-limit or availability failures defer sending rather than bypassing those checks.
The sla-deliver-scheduled-reports task runs every minute on metrics. A live scheduler
and a worker consuming that queue are both required. Notices contain a sign-in link only:
no report data, attachments, credentials or bearer download link is included. The
recipient must still pass current authorization when opening the report.
| Outbox state | Meaning |
|---|---|
pending / retry |
Waiting for a permitted attempt; delivery is not established. |
sending |
The claim was committed before SMTP; do not assume the message can be safely retried. |
accepted |
The SMTP server acknowledged submission; this is not an inbox or read receipt. |
unknown |
Submission or restored state is ambiguous; no automatic replay occurs. |
failed |
A terminal refusal, invalid configuration or retry/age limit stopped this notice. |
cancelled / skipped |
Authorization, operator cancellation or recipient preferences prevented sending. |
Operators with config:write and the report’s organization/site access can inspect
GET /api/v1/sla/reports/{report_id}/deliveries. They can cancel an unclaimed pending
or retry notice through POST /api/v1/sla/reports/{report_id}/deliveries/{delivery_id}/cancel.
Accepted or uncertain messages cannot be recalled through that endpoint. Transport retries
are bounded to five attempts and seven days; only outcomes known to be retryable are
retried. Lost acknowledgements, expired sending leases and mismatched recovery epochs
remain unknown for operator reconciliation.
After database recovery, follow the coordinated recovery guide for a new delivery epoch, sender fencing and explicit reapproval. Do not rewrite unknown outbox rows into retries or infer exactly-once delivery. Synthetic SMTP tests do not establish acceptance by your real relay or recipients; validate those separately before enabling notices.
Permissions reference
Section titled “Permissions reference”| Permission | Who can hold it | What it unlocks |
|---|---|---|
config:read |
viewer and above (site-scoped by grant) | Read policies, breaches, summary, reports |
config:write |
site_admin and above | Create/edit/delete policies; acknowledge breaches; trigger evaluation; create reports and schedules |
Role-to-permission mapping is defined in the Roles and permissions reference.
Connecting SLA to other enterprise features
Section titled “Connecting SLA to other enterprise features”- Alert Rules - create a rule that fires on
sla.breach.createdto route breach notifications to any configured notification channel. - Health Dashboard - the
SLAComplianceCardcomponent on the health dashboard (/health) pulls fromGET /api/v1/sla/summaryand shows a high-level compliance rate alongside device health scores. - Event Correlation - SLA breach events appear in the correlation engine’s event stream; you can write a correlation rule that groups multiple simultaneous breaches into a single incident.
- Notifications - configure an SMTP, Slack webhook, or Teams webhook provider under Enterprise → Notification Providers, then attach it to an alert rule targeting SLA breach events.
Security notes
Section titled “Security notes”- Every
scope_idsupplied at create or update time is verified to be owned by your organization before the operation is committed. Cross-tenant probes receive a 404. - All policy and breach reads are org-scoped - you cannot read another organization’s SLA data.
- The manual
/evaluateendpoint derives the target organization from the authenticated user’s JWT. It cannot be directed at another organization’s policies. - Multi-tenant isolation is enforced at the application layer. There is no PostgreSQL row-level security - isolation is the responsibility of the service and endpoint code. Passing internal tests does not replace an independent security assessment.
Next steps
Section titled “Next steps”- Alert Rules - wire breach events to notifications and auto-resolve workflows
- Health Dashboard (
/health) - view SLA compliance alongside device health scores - Event Correlation - group breach events into incidents for coordinated response
- Notification Providers - configure SMTP, Slack, Teams, and webhook delivery
All product names, logos, and brands are property of their respective owners. FreeSDN is an independent project and is not affiliated with or endorsed by the vendors it integrates with. See Trademarks.