Argo CD keeps each Kubernetes environment matching what Git declares. Kargo decides what Git should declare next: it packages a release as “freight” and moves it to the next stage only after checks pass. Together they give you automatic deploys to development, a soak period, verification gates, migrations inside the release and rollback by promoting the previous freight again.
We built this pipeline for MPI, a portfolio-analytics SaaS platform on AWS EKS. The gates that can actually say no are the part I would copy first.
What does Argo CD do, and what doesn’t it?
Argo CD calls itself “a declarative, GitOps continuous delivery tool for Kubernetes”. It watches a Git repository that holds the desired state of an application (Helm charts, manifests, values) and reconciles the cluster toward it. If someone changes the cluster by hand, it shows the drift and can put things back.
Argo CD leaves one question open: which version belongs in which environment. If production should run the image that spent an hour healthy in development, something has to write that version into production’s desired state at the right moment. Without a promotion tool that something is a person editing a values file, or a CI job with a long script and an approval by chat message. I’ve seen both. The chat message is the weaker control, and everybody involved knows it.
What does Kargo add on top?
Kargo’s documentation calls it “an unopinionated continuous promotion platform” that moves new code and configuration through the stages of an application’s lifecycle “using GitOps principles”. The model has three nouns.
A warehouse watches repositories (container images, Git, Helm charts) for new revisions. Freight is a “meta-artifact” that pins the specific revisions of images and manifests that have to travel together, so one release is one piece of freight. A stage is a promotion target, usually an environment. Stages link into a pipeline, and freight reaches production only after it was verified in the stages before.
Kargo writes the promotion to Git and Argo CD syncs it. Git stays the record of what runs where, and every promotion is a commit you can read.
How does a release get from development to production?
On MPI there are two environments in separate AWS accounts. GitLab CI builds and pushes the images, and Kargo sees the new revisions and creates freight. Development promotes new freight without waiting for anyone, and Argo CD syncs it. Then the freight sits in development for one hour while the platform watches it.
Before production is allowed, the freight has to pass smoke checks at the edge, pod restart and readiness checks, and an HTTP 5xx ratio gate. If one fails, the freight isn’t eligible. If they pass, one promote action runs a 79-step automated template. That single action replaced seven manual steps and a four-part certification checklist. We haven’t measured the time it saves.
The last step is checking that the image digests running in production are the ones in the release. That is how the first production release through this path was verified: all 12 running service images matched. It reached production in about 2 hours, including the 1-hour soak. It’s one release, so it shows what the path allows.
Which gates actually stop a bad release?
A gate earns its place only if it can say no. We kept three, because they are cheap to run and each catches a different failure.
Edge smoke checks call the public endpoints through the real ingress path. They catch the release that starts fine but can’t serve traffic through the gateway. Pod restart and readiness checks catch crash loops and services that never become ready, which a green sync alone doesn’t show. The 5xx ratio compares server errors against traffic during the soak, and catches the release that serves, but serves errors.
With these in place, production can now block a bad release. Test that property before you trust any gate: run it once against a case that should fail.
How do you run database migrations safely in GitOps?
Migrations are where GitOps pipelines usually fall back to a manual window. We keep them inside the release. The migration runs as a Kubernetes Job annotated argocd.argoproj.io/hook: PreSync, so it executes before the new manifests are applied, and the platform takes an automatic snapshot of the database before it starts. Argo CD’s docs are clear about the failure case: if any PreSync hook fails, the whole sync stops and is marked as failed. The old version keeps running against the database it already understands.
A minimal hook looks like this:
apiVersion: batch/v1
kind: Job
metadata:
name: db-migrate
annotations:
argocd.argoproj.io/hook: PreSync
argocd.argoproj.io/hook-delete-policy: BeforeHookCreation
spec:
backoffLimit: 0
template:
spec:
restartPolicy: Never
containers:
- name: migrate
image: registry.example.com/app@sha256:<digest>
command: ["npm", "run", "migrate"]
backoffLimit: 0 matters. A migration that fails once should stop the release, so nothing retries against a half-migrated schema. And migrations still have to be backward compatible with the running version, because a rollback does not undo them.
How does rollback by re-promotion work?
You don’t undo a deploy in this model. You promote the previous freight to the stage again. Kargo writes the old revisions back to Git, Argo CD syncs them and the same gates run.
We rehearsed it end to end in development: about 75 seconds to roll back, about two minutes to roll forward again, and 15 to 17 minutes until verification turned green. Those are development numbers, and nobody has measured a production rollback there.
The rehearsal is the point. A rollback path nobody has run is a hypothesis.
What changed in the delivery numbers?
Measured on MPI’s platform in October 2026:
| Metric | Before | After |
|---|---|---|
| Production releases per week | About 1 (April to July 2026) | Nearly 4 (last 30 days) |
| Median lead time, merge to production | 7 to 10 days (April to July) | 3 to 4 days (September) |
| Failed production release pipelines | 23% (77 of 334, all time) | 4% (1 of 25 since July 2026) |
The last row counts failed pipelines. DORA’s change fail rate counts releases that hurt users, and the DORA metrics post explains the difference.
When is Kargo more than you need?
If you have one environment, or one service that deploys straight from CI and you’re happy with that, Argo CD alone is enough. Kargo pays off when several services must move together through two or more environments, when you want soak time and gates between them, and when “what runs in production and how did it get there” has to be answered from Git instead of from someone’s memory.
Start with the gates. Write down what would make you refuse a release today and turn each item into a check that can fail. Then wire the promotion. If your CI/CD and infrastructure as code need rebuilding first, that’s our DevOps service, and the two-week CI/CD & IaC package is USD 4,000. Every price is on the pricing page.
Today, list what a person checks by hand before a production release. Each item is a candidate gate.
Sources, accessed 2026-10-07: Argo CD documentation, Argo CD, Sync Phases and Waves, Kargo core concepts.
