The engagement

Client: MPI

Status
Production platform, work ongoing
Timeline
Ongoing since January 2025

What changed since the new release pipeline?

For fund buyers and sellers. One Clouditive platform engineer, working with an AI agent, sits alongside the client's DevOps lead and about 8 engineers.

Each figure below was measured on the client's own systems: GitLab, Argo CD and Kargo, and its metrics store.

Production releases went from about 1 a week to nearly 4. Median time from merge to production fell from over a week to a few days, and production release failures fell from 23% to 4%. Production went from zero alert rules to 275, and one action replaced seven manual steps and a certification checklist. In October 2026 the first release through the new promotion pipeline ran all 12 service images verified to match.

  • ~3xreleases per week

    About 3x more production releases per week: from 1.15 a week (April to July 2026) to 3.7 to 4.0 a week (last 30 days)

  • 23% → 4%release failures

    Production release failures down from 23% to 4%

  • 275alerting rules

    From zero production alerts to 275 alerting rules

  • 72 srollback in dev

    Rehearsed in dev: rollback in 72 s, going forward again in 124 s, both verified

Also on record

  • Median time from merge to production down from 7 to 9.5 days to 3.4 to 4.3 days; the first release through the new pipeline reached production in about 2 hours, including a 1-hour soak
  • Seven manual steps and a certification checklist replaced by one action
  • First production release through the new promotion pipeline; all 12 running service images verified to match the release (October 2026)
  • Automatic checks can now block a bad release before production
  • Large-report out-of-memory failure diagnosed and fixed

What we did, by area

Areas of work

  • Platform engineering
  • GitOps delivery
  • Infrastructure as code
  • Observability and SRE
  • Cost audits
  • Security checks
GitOps delivery

New code reaches dev on its own. Production waits an hour and must pass health checks.

Promotion pipelines run on Kargo over Argo CD. Dev deploys happen automatically. Production waits for a 1-hour soak, an hour in which the release runs under watch, and for verification gates.

The gates are edge smoke checks, pod-restart and readiness checks, and an HTTP 5xx-ratio gate that stops the release if the share of server errors rises. Database migrations run in a pre-sync hook, before the new version goes live, and the hook takes a snapshot of the database first.

Rollback is a re-promotion of the previous release, rehearsed in dev.

Internal developer portal

One Backstage portal for services, deploys and cost

Services, deploys, cost, vulnerabilities and CI health sit in one portal.

  • Service and API catalogs generated from OpenAPI
  • Dependency map and deploy status
  • FinOps cost per service
  • Image vulnerabilities
  • CI health with 30-day flake and retry statistics
  • Profiling diffs between deploys
Infrastructure as code

Terraform and OpenTofu with Atlantis, moving to Crossplane

Infrastructure changes go through Terraform/OpenTofu with Atlantis. The move to Crossplane runs as zero-downtime cutovers, each with a rollback.

Observability and SRE

Alerts tested until they fire, and recurring failures fixed

The stack is VictoriaMetrics, Loki, Tempo, Pyroscope and Grafana. Each alert was tested until it fired.

  • Node consolidation no longer recreates the log store about three times a day.
  • Autoscaling no longer drops most report pods at once.
  • From zero production alerts to 275 alerting rules.
  • Long report jobs drain gracefully.
  • A large-report out-of-memory failure was diagnosed and fixed.
Cost and security

Cost audits and a vulnerability check on every code change

Cluster and cloud cost audits.

A SOC 2-oriented vulnerability gate runs on every merge request, with suppressions that expire. Secrets live in a managed secrets store.

Team

One platform engineer, working with an AI agent, inside the client's team

One Clouditive platform engineer works with an AI agent alongside the client's DevOps lead and about 8 of its engineers. The engagement has run since January 2025 and continues today.

Stack

  • AWS EKS
  • Argo CD
  • Kargo
  • Helm
  • Karpenter
  • Istio ambient
  • Kong
  • GitLab CI
  • Terraform / OpenTofu
  • Atlantis
  • Crossplane
  • Backstage
  • VictoriaMetrics
  • Loki
  • Tempo
  • Pyroscope
  • Grafana
  • Node.js / TypeScript
  • Python

Services used

The services behind this work.

On the client side

Development lead on the client side

Published without a name or a quote, with the client's approval.

Frequently asked questions

What does Clouditive do for this fintech platform?

Platform engineering, DevOps and SRE for a portfolio-analytics SaaS platform: GitOps delivery, an internal developer portal, infrastructure as code, observability, cost audits and security gates.

How do releases reach production?

Through Kargo promotion pipelines on Argo CD. Dev deploys are automatic. Production waits an hour under watch and must pass automatic health and error-rate checks first.

How does rollback work?

By re-promoting the previous release. It was rehearsed in the dev environment: 72 seconds back and 124 seconds forward, both verified. Only the dev rehearsal has been measured.

How are database migrations protected?

A snapshot of the database is taken automatically before each migration runs.

What does the developer portal show?

Service and API catalogs from OpenAPI, a dependency map, deploy status, FinOps cost, image vulnerabilities, 30-day CI flake and retry statistics, and profiling diffs between deploys.

Which observability stack does the platform use?

VictoriaMetrics, Loki, Tempo, Pyroscope and Grafana, with each alert tested until it fired.

Where do the portal's cost figures come from?

Straight from the cloud billing API.

How is security checked?

Every merge request is scanned by a SOC 2-oriented vulnerability check, with exceptions that expire, and secrets are kept in a managed secrets store.

What is the stack?

AWS EKS with separate dev and prod accounts, Argo CD, Kargo, Helm, Karpenter, Istio ambient, Kong, GitLab CI, Terraform/OpenTofu, Atlantis, Crossplane, Backstage, Node.js/TypeScript and Python.

Tell us your case.

We reply to every request within 1 business day. We sign an NDA before the call if you ask.