not accepting clients
← back to blog

promotions with kargo

kargo models the sequence and requirements for promoting a version between environments, layering on top of argocd without replacing its sync.

an argocd application maps one source - a path in a repository at a revision - to one destination. it reconciles continuously, reports drift, and heals. what it does not have is a notion of sequence. development, staging, and production are three unrelated applications pointed at three directories, and nothing in argocd knows that the image running in staging is supposed to be the one that already passed development.

so the sequence gets encoded somewhere else. usually it is a person editing an image tag in a values file and opening a pull request, or a ci job doing the same edit with yq and a service account. either way the promotion history lives in commit messages, the answer to “what is in staging that is not in production” requires a diff, and the rule that staging must be healthy for an hour first is a convention rather than a control. kargo adds that missing layer as a set of custom resources, and leaves argocd doing exactly what it already does.

the short version

a Warehouse watches artifact sources - image repositories, git repositories, helm chart repositories - and each time it finds a new combination it freezes that combination into an immutable Freight resource. Stage resources form a graph by declaring which freight they accept and where it may come from: directly from a warehouse, or only from an upstream stage. moving freight into a stage creates a Promotion, which runs an ordered list of steps - clone the declarative repo, set the image, commit, push, tell argocd to sync - and then verifies the result. argocd never learns about any of this. it sees commits landing in branches it already tracks.

the resource model

kargo installs as a helm chart into its own namespace and runs alongside argocd rather than inside it.

helm install kargo \
  oci://ghcr.io/akuity/kargo-charts/kargo \
  --namespace kargo \
  --create-namespace \
  --set api.adminAccount.passwordHash=$hashed_pass \
  --set api.adminAccount.tokenSigningKey=$signing_key \
  --wait

a Project is the top-level grouping. creating one provisions a namespace of the same name, and every warehouse, stage, promotion, and freight for that pipeline lives in it.

apiVersion: kargo.akuity.io/v1alpha1
kind: Project
metadata:
  name: checkout

this is the boundary for rbac. permission to promote to production is a rolebinding in the checkout namespace, not a global argocd rbac line.

discovering artifacts

a Warehouse is a set of subscriptions. it polls each one on spec.interval and records what it finds.

apiVersion: kargo.akuity.io/v1alpha1
kind: Warehouse
metadata:
  name: checkout
  namespace: checkout
spec:
  interval: 2m
  freightCreationPolicy: Automatic
  subscriptions:
    - image:
        repoURL: ghcr.io/organisation/checkout
        imageSelectionStrategy: SemVer
        constraint: ^1.0.0
        strictSemvers: true
        discoveryLimit: 20
    - chart:
        repoURL: https://charts.organisation.dev
        name: checkout
        semverConstraint: ^2.0.0

imageSelectionStrategy decides what “newest” means. SemVer orders tags by version and ignores anything unparseable, which is the right answer for released artifacts. NewestBuild orders by the image’s creation timestamp and is the only strategy that works with opaque tags like commit shas, at the cost of one registry call per tag to read the manifest. Digest follows a mutable tag such as main and produces new freight whenever the digest behind it changes.

each time the resolved set of artifacts changes, the warehouse computes a canonical representation of it, hashes it, and creates a Freight resource whose metadata.name is that hash. freight is immutable and it is a set, not a single artifact: one freight can pin an image, a chart version, and a manifest commit together, and they travel as a unit.

kubectl get freight -n checkout

the hash is unreadable, so kargo also attaches a generated alias in the kargo.akuity.io/alias label - frozen-tauntaun, mortal-dragonfly - and every cli command accepts it in place of the hash.

stages and where freight may come from

a Stage declares what it accepts through spec.requestedFreight. the origin names the warehouse the freight must have come from, and sources decides whether it may arrive directly or only after an upstream stage has had it.

apiVersion: kargo.akuity.io/v1alpha1
kind: Stage
metadata:
  name: development
  namespace: checkout
spec:
  requestedFreight:
    - origin:
        kind: Warehouse
        name: checkout
      sources:
        direct: true

direct: true means new freight is eligible for development the moment the warehouse produces it. downstream stages set sources.stages instead.

apiVersion: kargo.akuity.io/v1alpha1
kind: Stage
metadata:
  name: staging
  namespace: checkout
spec:
  requestedFreight:
    - origin:
        kind: Warehouse
        name: checkout
      sources:
        stages:
          - development

this is the constraint that was previously a convention. freight becomes available to staging only once development holds it, is healthy, and has passed its verification. there is no path by which a fresh image reaches staging without going through development, because the only thing kargo will let a promotion carry is freight that satisfies requestedFreight.

the graph is defined entirely by these references, so a fan-out to three regional prod stages is three stages each listing staging as their source. there is no separate pipeline object.

hotfixes need an escape hatch, and kargo’s is an explicit command:

kargo approve --project checkout --freight-alias frozen-tauntaun --stage prod

that records the freight in status.approvedFor and makes it promotable to prod without an upstream stage. the approval is an object that stays on record, so the shortcut can be audited later.

the promotion template

spec.promotionTemplate is where the stage says what promoting into it does. it is an ordered list of steps, each a named built-in with a config block, and expressions in ${{ }} give steps access to freight, context, and the outputs of earlier steps.

apiVersion: kargo.akuity.io/v1alpha1
kind: Stage
metadata:
  name: staging
  namespace: checkout
spec:
  vars:
    - name: declarativeRepo
      value: https://github.com/organisation/declarative.git
    - name: imageRepo
      value: ghcr.io/organisation/checkout
  requestedFreight:
    - origin:
        kind: Warehouse
        name: checkout
      sources:
        stages:
          - development
  promotionTemplate:
    spec:
      steps:
        - uses: git-clone
          config:
            repoURL: ${{ vars.declarativeRepo }}
            checkout:
              - branch: main
                path: ./src
              - branch: stage/${{ ctx.stage }}
                create: true
                path: ./out
        - uses: kustomize-set-image
          as: set-image
          config:
            path: ./src/overlays/${{ ctx.stage }}
            images:
              - image: ${{ vars.imageRepo }}
        - uses: kustomize-build
          config:
            path: ./src/overlays/${{ ctx.stage }}
            outPath: ./out/manifests.yaml
        - uses: git-commit
          as: commit
          config:
            path: ./out
            message: ${{ outputs['set-image'].commitMessage }}
        - uses: git-push
          config:
            path: ./out
            branch: stage/${{ ctx.stage }}

the two checkouts are the important part of the shape. ./src is the source of truth that humans edit; ./out is a branch containing nothing but rendered manifests for this one stage. the promotion renders kustomize itself and commits the output, so what argocd syncs is a flat yaml file with no overlays to resolve.

this is the rendered manifests pattern, and it is worth the extra branch for two reasons. an argocd application pointed at rendered output cannot behave differently from what was reviewed, because there is no build step left to behave differently. and the diff between two stages becomes a diff of two files instead of a comparison of two overlay trees.

kustomize-set-image reads the image tag from the freight being promoted rather than from a parameter, which is why the step needs no version anywhere in its config. the same template is correct for every promotion into this stage.

handing off to argocd

pushing the commit is enough to make argocd act, since it is already tracking stage/staging. but a promotion that finishes at git-push finishes before anything has rolled out, and kargo would mark it successful while the cluster is still on the old revision. the argocd-update step closes that gap.

        - uses: argocd-update
          config:
            apps:
              - name: checkout-${{ ctx.stage }}
                namespace: argocd
                sources:
                  - repoURL: ${{ vars.declarativeRepo }}
                    desiredRevision: ${{ outputs.commit.commit }}

the step triggers a refresh and then registers a health check against the stage: the application must be synced and healthy at that specific commit, the one the git-commit step produced. an application that is healthy on the previous revision does not satisfy it. this is what makes the stage’s health mean “the thing i promoted is running” rather than “argocd is not complaining”.

kargo will not touch an application that has not opted in. the application needs an annotation naming the project and stage allowed to drive it:

apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: checkout-staging
  namespace: argocd
  annotations:
    kargo.akuity.io/authorized-stage: "checkout:staging"
spec:
  project: checkout
  source:
    repoURL: https://github.com/organisation/declarative.git
    targetRevision: stage/staging
    path: .
  destination:
    server: https://kubernetes.default.svc
    namespace: checkout-staging
  syncPolicy:
    automated:
      prune: true
      selfHeal: true

without the annotation the step fails rather than silently skipping. it is the consent boundary between the two systems: kargo holding cluster-admin does not implicitly mean kargo may promote to prod.

argocd-update can also set updateTargetRevision: true and write the revision into the application spec directly. that removes the rendered branch, and with it the property that git contains a complete record of what each stage was running. the git-first shape is the one to reach for unless the application is not backed by a repository at all.

automatic and manual promotion

by default every promotion is deliberate. auto-promotion is opt-in per stage, and it lives in a ProjectConfig resource that shares its name with the project.

apiVersion: kargo.akuity.io/v1alpha1
kind: ProjectConfig
metadata:
  name: checkout
  namespace: checkout
spec:
  promotionPolicies:
    - stageSelector:
        name: development
      autoPromotionEnabled: true
    - stageSelector:
        name: staging
      autoPromotionEnabled: true

development and staging advance on their own as freight becomes eligible. production is absent from the list, so it only moves when someone asks:

kargo promote --project checkout --freight-alias frozen-tauntaun --stage production

the split between Project and ProjectConfig exists so that permission to change what auto-promotes is not the same permission as changing the project itself.

verification

a stage being healthy is a weaker claim than a stage being correct. spec.verification runs argo rollouts AnalysisTemplate resources after the promotion reports healthy, and freight only becomes eligible downstream if they succeed.

  verification:
    analysisTemplates:
      - name: checkout-smoke

the template is an ordinary argo rollouts analysis, which means it can be a prometheus query rather than a test script.

apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: checkout-smoke
  namespace: checkout
spec:
  metrics:
    - name: error-rate
      interval: 1m
      count: 10
      successCondition: result[0] < 0.01
      failureLimit: 1
      provider:
        prometheus:
          address: http://prometheus.monitoring:9090
          query: |
            sum(rate(http_requests_total{job="checkout",namespace="checkout-staging",code=~"5.."}[5m]))
            /
            sum(rate(http_requests_total{job="checkout",namespace="checkout-staging"}[5m]))

ten samples a minute apart with a failure limit of one means the freight has to hold under a real error budget for ten minutes before production will accept it. the soak time that used to be a rule in a runbook is now the reason a promotion is not available yet.

credentials

kargo needs read access to the repositories a warehouse subscribes to and write access to the declarative repository a promotion pushes to. credentials are plain secrets, identified by a label and matched to repositories by url.

apiVersion: v1
kind: Secret
metadata:
  name: declarative-repo
  namespace: checkout
  labels:
    kargo.akuity.io/cred-type: git
stringData:
  repoURL: https://github.com/organisation/declarative.git
  username: kargo
  password: ghp_...

repoURL can be a regular expression covering many repositories, in which case the secret also needs repoURLIsRegex: "true". a secret in the project namespace applies only to that project; a secret in the global credentials namespace configured at install time applies everywhere, which is convenient and worth being deliberate about.

these are separate from argocd’s own repository credentials. both systems talk to the same repositories and neither reads the other’s secrets, so a rotated token has to be rotated twice.

what to watch out for

warehouse polling is not the pipeline’s clock. spec.interval decides how quickly new artifacts are noticed, not how quickly they move. lowering it to seconds increases registry api calls without shortening a promotion, and registries rate limit on manifest reads - which NewestBuild and Digest do far more of than SemVer.

image tags that are not semver silently produce nothing. with imageSelectionStrategy: SemVer and strictSemvers: true, any tag kargo cannot parse is ignored. a pipeline tagging images with commit shas will run a warehouse that finds no freight at all and reports no error, because from its point of view there is simply nothing matching. check status.discoveredArtifacts on the warehouse before assuming the subscription is wrong.

freight is a set, and promoting it promotes all of it. a warehouse subscribing to an image and a chart produces freight that pins both. that is usually what is wanted, and it means a chart release with no new image still creates new freight and still moves through every stage. subscriptions belong in the same warehouse only when the artifacts should always be deployed together; otherwise use separate warehouses and list both in requestedFreight.

the authorized-stage annotation is per stage, not per project. an application annotated for checkout:staging cannot be updated by the production stage. this is the desired behaviour, but a copied application manifest with a stale annotation fails at the argocd-update step, after the commit has already been pushed. the promotion is left partially applied: git is updated, argocd may sync it anyway through its own automation, and kargo reports the promotion as failed.

selfheal fights promotion steps that write to the cluster. if argocd-update is used with updateTargetRevision: true against an application with selfHeal: true, argocd may revert the spec change before it syncs. the git-first shape avoids this entirely, since the change lands in the repository argocd is reconciling towards rather than in the application object.

verification runs after health, not instead of it. a stage that never reaches a healthy state never starts its analysis, so a failing readiness probe shows up as a promotion stuck in Running rather than a verification failure. kubectl describe stage and the promotion’s status.message distinguish the two.

deleting a stage does not delete what it deployed. stages are not owners of the workloads they promote to; the argocd application is. removing a stage stops promotions and leaves the running deployment in place, which is the safe default and an easy way to accumulate orphaned environments.

references

[1] kargo documentation. “key kargo concepts.”
docs.kargo.io/concepts

[2] kargo documentation. “working with warehouses.”
docs.kargo.io/user-guide/how-to-guides/working-with-warehouses

[3] kargo documentation. “working with stages.”
docs.kargo.io/user-guide/how-to-guides/working-with-stages

[4] kargo documentation. “argocd-update promotion step.”
docs.kargo.io/user-guide/reference-docs/promotion-steps/argocd-update

[5] kargo documentation. “working with projects.”
docs.kargo.io/user-guide/how-to-guides/working-with-projects

[6] kargo documentation. “managing secrets.”
docs.kargo.io/operator-guide/security/managing-secrets

[7] argo rollouts documentation. “analysis, progressive delivery.”
argo-rollouts.readthedocs.io/en/stable/features/analysis

[8] argo cd documentation. “automated sync policy.”
argo-cd.readthedocs.io/en/stable/user-guide/auto_sync

# ask the author

a question
about this
post?

direct line