What does observability do?
It answers why a system behaves the way it does from the signals it emits (metrics, logs, traces and profiles), without shipping new code to find out.
Is observability a part of DevOps?
Yes. You cannot own how software runs without seeing it run. In our work it sits with SRE: instrumentation, alerts and the fixes that follow.
What are the top 3 observability tools?
We do not rank vendors. What we have run in production: VictoriaMetrics, Loki, Tempo, Pyroscope and Grafana at MPI, and OpenTelemetry to Google Cloud at Autonomah.
What does an alert tested until it fires mean?
We trigger the failure condition on purpose and check the alert arrives. An alert that has never fired is an assumption, not a control.
What does SRE / Observability Foundations cost?
USD 5,000 for 2 weeks, fixed price: metrics, logs and traces wired, 3 SLOs defined, alerts tested until they fire, and a runbook.
Is SRE still in demand?
We sell SRE because production systems keep needing it: at MPI the work covered alerts, autoscaling, log storage, out-of-memory failures and cost.
How much does an SRE cost per hour?
Our published rates: USD 55–60 for a Lead or Architect, 45–50 Senior, 35–40 Mid and 30 Junior, with a 6-month minimum for staff augmentation.
What SRE / Observability Foundations changes for the business
Each deliverable, in the terms a CEO or a board asks about:
- Metrics, logs and traces wired: when the product slows down or fails, your team sees where and why instead of guessing.
- 3 SLOs defined: three written targets for how reliable the product must be, so "the app goes down too often" becomes a number you can track.
- Alerts tested until they fire: your team hears about a failure from an alert, not from a customer, and every alert has been triggered on purpose to show it arrives.
- A runbook: written first steps for each alert, so whoever is on call can act without the person who built the system.
- At MPI, large reports that ran out of memory were diagnosed and fixed, and scaling down no longer interrupts most running reports at once.
What are the four pillars of observability?
Observability is usually described through metrics, logs and traces, with continuous profiling often added as a fourth signal. Metrics show that something changed, logs show what happened, traces show where a request spent its time, and profiles show which code used the CPU or memory. At MPI each signal has its own store (VictoriaMetrics, Loki, Tempo, Pyroscope), read together in Grafana.
On Autonomah, telemetry goes through OpenTelemetry to Google Cloud, with a per-call cost ledger and every production deploy checked to run the exact build that was tested.
On Nyravorn, an online game, we rehearse backup and restore, and smoke scripts check that player data survives a server restart.
Is SRE just DevOps?
No. DevOps is a way of working that joins building and running software; SRE is a specific engineering discipline inside that idea, focused on how reliable production is. A DevOps engineer usually owns the path to production: pipelines and infrastructure as code. An SRE owns what happens after: signals, alerts, incidents, capacity and recovery. On MPI we do both for the same platform.
Is AI replacing SRE? In our work, the engineer directs the work and AI tools assist: our platform lead at MPI works that way. Alerts are still tested until they fire and fixes still verified in the running system.
An open-source alternative to Datadog
OpenTelemetry collects metrics, logs and traces once, and open-source stores such as Mimir, Loki and Tempo, read in Grafana, keep them without a per-host commercial license. For a US transportation services company, moving metrics and logs off a commercial APM cut observability spend by about USD 5,000 a month, and dashboards went from 18 to 135. The work is setup and upkeep instead of a subscription.