What exactly does a site reliability engineer do?

A site reliability engineer keeps production services reliable using software engineering: instrumenting systems so problems are visible, writing alerts that fire on real failures, finding the cause of recurring incidents and fixing it in code or configuration, planning capacity and rehearsing recovery. At Clouditive each of those fixes is verified in the running system before we call it done.

An alert that has never fired is an assumption. A failure that comes back every day gets absorbed as normal until it costs a release or a report.

SRE / Observability Foundations costs USD 5,000 for 2 weeks: metrics, logs and traces wired, 3 SLOs defined, alerts tested until they fire, and a runbook.

What you get, step by step

Everything in the package, in the order it lands. The price and the duration stay the same.

  1. Step 1: The signals

    Metrics, logs and traces wired.

  2. Step 2: The SLOs

    3 SLOs defined.

  3. Step 3: Alerts and runbook

    Alerts tested until they fire, and a runbook.

  4. 2 weeks

    USD 5,000

Want to keep going after the package? Add engineers by the hour at the published rates, with a 6-month minimum term. How staff augmentation works.

Observability and SRE for a fintech SaaS platform

Investment analytics / Fintech

Status: Production platform, work ongoing

  1. dev
  2. 1-hour soak
  3. gates
  4. production
  • ~3xmore production releases per week
  • 23% → 4%production release failures
  • 0 → 275production alerting rules
  • 72 srollback, rehearsed in dev

Drawn from the case. The product is a client system and is not shown.

Read the case

For MPI's portfolio-analytics SaaS platform, on AWS EKS, our work includes:

  • An observability stack of VictoriaMetrics, Loki, Tempo, Pyroscope and Grafana, with alerts tested until they fire.
  • The log store is no longer recreated about three times a day by node consolidation.
  • Scaling down no longer interrupts most running reports at once, and long reports finish cleanly when capacity shrinks.
  • A large-report out-of-memory failure diagnosed and fixed.
  • Cluster and cloud cost audits.

Prefer to pay by the hour?

Add engineers to your own team at the published rates instead of buying a package.

  • You interviewYou meet the engineer who will do the work.
  • 6-month minimumBilled per hour worked.
  • Free replacementIf they leave or don't fit, we replace them and cover the handover at no cost.
See how staff augmentation works
  • Lead / ArchitectUSD55–⁠60per hour
  • SeniorUSD45–⁠50per hour
  • MidUSD35–⁠40per hour
  • JuniorUSD30per hour

USD per hour, drawn to one scale

Frequently asked questions

What does observability do?

It answers why a system behaves the way it does from the signals it emits (metrics, logs, traces and profiles), without shipping new code to find out.

Is observability a part of DevOps?

Yes. You cannot own how software runs without seeing it run. In our work it sits with SRE: instrumentation, alerts and the fixes that follow.

What are the top 3 observability tools?

We do not rank vendors. What we have run in production: VictoriaMetrics, Loki, Tempo, Pyroscope and Grafana at MPI, and OpenTelemetry to Google Cloud at Autonomah.

What does an alert tested until it fires mean?

We trigger the failure condition on purpose and check the alert arrives. An alert that has never fired is an assumption, not a control.

What does SRE / Observability Foundations cost?

USD 5,000 for 2 weeks, fixed price: metrics, logs and traces wired, 3 SLOs defined, alerts tested until they fire, and a runbook.

Is SRE still in demand?

We sell SRE because production systems keep needing it: at MPI the work covered alerts, autoscaling, log storage, out-of-memory failures and cost.

How much does an SRE cost per hour?

Our published rates: USD 55–⁠60 for a Lead or Architect, 45–⁠50 Senior, 35–⁠40 Mid and 30 Junior, with a 6-month minimum for staff augmentation.

What SRE / Observability Foundations changes for the business

Each deliverable, in the terms a CEO or a board asks about:

  • Metrics, logs and traces wired: when the product slows down or fails, your team sees where and why instead of guessing.
  • 3 SLOs defined: three written targets for how reliable the product must be, so "the app goes down too often" becomes a number you can track.
  • Alerts tested until they fire: your team hears about a failure from an alert, not from a customer, and every alert has been triggered on purpose to show it arrives.
  • A runbook: written first steps for each alert, so whoever is on call can act without the person who built the system.
  • At MPI, large reports that ran out of memory were diagnosed and fixed, and scaling down no longer interrupts most running reports at once.

What are the four pillars of observability?

Observability is usually described through metrics, logs and traces, with continuous profiling often added as a fourth signal. Metrics show that something changed, logs show what happened, traces show where a request spent its time, and profiles show which code used the CPU or memory. At MPI each signal has its own store (VictoriaMetrics, Loki, Tempo, Pyroscope), read together in Grafana.

On Autonomah, telemetry goes through OpenTelemetry to Google Cloud, with a per-call cost ledger and every production deploy checked to run the exact build that was tested.

On Nyravorn, an online game, we rehearse backup and restore, and smoke scripts check that player data survives a server restart.

Is SRE just DevOps?

No. DevOps is a way of working that joins building and running software; SRE is a specific engineering discipline inside that idea, focused on how reliable production is. A DevOps engineer usually owns the path to production: pipelines and infrastructure as code. An SRE owns what happens after: signals, alerts, incidents, capacity and recovery. On MPI we do both for the same platform.

Is AI replacing SRE? In our work, the engineer directs the work and AI tools assist: our platform lead at MPI works that way. Alerts are still tested until they fire and fixes still verified in the running system.

An open-source alternative to Datadog

OpenTelemetry collects metrics, logs and traces once, and open-source stores such as Mimir, Loki and Tempo, read in Grafana, keep them without a per-host commercial license. For a US transportation services company, moving metrics and logs off a commercial APM cut observability spend by about USD 5,000 a month, and dashboards went from 18 to 135. The work is setup and upkeep instead of a subscription.

Tell us your case.

We reply to every request within 1 business day. We sign an NDA before the call if you ask.