CI/CD Workflows: A Practical Playbook with Pipeline Examples, Metrics and a Migration Plan
Published

Modern software teams can no longer afford slow release cycles or manual deployments. Users expect fast updates, businesses require continuous innovation, and developers need confidence that their code works reliably. This is where CI/CD workflows, also called CI/CD pipelines, come into play.
Most explainers stop at definitions and abstract benefits. This guide goes further: after a short introduction, it gives you copy-paste pipelines for GitLab CI and GitOps, the Terraform and Kubernetes code to run them, measurable tuning advice, where to place security checks, and a six-week migration plan you can actually follow.
What are CI/CD workflows or pipelines?
CI or Continuous Integration
Continuous Integration is the practice of frequently merging code changes into a shared repository, where every change is automatically built and tested. The goal is to catch bugs early, before they reach production.
Continuous Delivery or Continuous Deployment
In Continuous Delivery, once code passes the tests it is automatically packaged and prepared for release, but the deployment to production still requires a human approval. In Continuous Deployment, every validated build is deployed to production automatically, without manual intervention.
In practice, most teams use both concepts under the umbrella of CI/CD pipelines. These pipelines help teams catch issues early, move faster and release better software by removing manual work.
What a practical CI/CD workflow looks like
A usable CI/CD workflow has three layers: automation for fast feedback (CI), safe release mechanics (CD), and observability that closes the loop. In practice this becomes a short chain of automated steps that runs for each change: pull the code, run fast unit tests, build a predictable artifact (usually a container image), run integration and security checks, and then promote the artifact through delivery gates to production.
Concretely, a good pipeline starts with a merge request that triggers linting and unit tests in parallel. A green build produces an immutable image, tagged with the commit SHA, that is scanned for vulnerabilities and pushed to a registry. The deployment manifests are then updated in Git (GitOps), and an orchestrator rolls the new image out progressively while telemetry evaluates health signals and triggers a rollback if needed. Keeping each stage small and observable is what combines speed with safety.
A Production-Ready GitLab CI Pipeline
The first example is a classic pipeline for a Dockerized application: lint, test, build, scan, then deploy to staging automatically and to production with a manual approval. It builds the Next.js application behind workflows.guru, but only the commands in the .node template are language specific, so you can swap them for Python, Go or Java. If you are new to GitLab pipelines, start with our GitLab CI tutorial first.
1stages:2 - test3 - build4 - scan5 - deploy67variables:8 DOCKER_IMAGE: "$CI_REGISTRY_IMAGE:$CI_COMMIT_SHA"910# Run pipelines for merge requests, the default branch and tags only11workflow:12 rules:13 - if: $CI_PIPELINE_SOURCE == "merge_request_event"14 - if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH15 - if: $CI_COMMIT_TAG1617# Shared setup for Node.js jobs: dependencies are cached per lockfile18.node:19 image: node:2220 cache:21 key:22 files:23 - yarn.lock24 paths:25 - .yarn-cache/26 before_script:27 - yarn install --frozen-lockfile --cache-folder .yarn-cache --prefer-offline2829lint:30 extends: .node31 stage: test32 script:33 - yarn lint3435unit_tests:36 extends: .node37 stage: test38 script:39 # Configure your test runner to write a JUnit report to junit.xml40 - yarn test --ci41 artifacts:42 when: always43 reports:44 junit: junit.xml4546build_image:47 stage: build48 image: docker:2949 services:50 - docker:29-dind51 variables:52 DOCKER_TLS_CERTDIR: "/certs"53 script:54 - echo "$CI_REGISTRY_PASSWORD" | docker login -u "$CI_REGISTRY_USER" --password-stdin "$CI_REGISTRY"55 - docker build -t "$DOCKER_IMAGE" .56 - docker push "$DOCKER_IMAGE"5758image_scan:59 stage: scan60 image:61 name: aquasec/trivy:0.75.062 entrypoint: [""]63 variables:64 TRIVY_USERNAME: "$CI_REGISTRY_USER"65 TRIVY_PASSWORD: "$CI_REGISTRY_PASSWORD"66 script:67 - trivy image --exit-code 1 --severity HIGH,CRITICAL --ignore-unfixed "$DOCKER_IMAGE"6869# kubectl gets its cluster access from the GitLab agent for Kubernetes70# or from a KUBECONFIG file variable defined in the project settings71.deploy:72 stage: deploy73 image:74 name: alpine/k8s:1.36.575 entrypoint: [""]76 script:77 - kubectl set image deployment/workflows-guru workflows-guru="$DOCKER_IMAGE"78 - kubectl rollout status deployment/workflows-guru --timeout=180s7980deploy_to_staging:81 extends: .deploy82 environment:83 name: staging84 url: https://staging.workflows.guru85 rules:86 - if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH8788deploy_to_production:89 extends: .deploy90 environment:91 name: production92 url: https://www.workflows.guru93 resource_group: production94 rules:95 - if: $CI_COMMIT_TAG96 when: manual
Lint and unit tests share the test stage, so they run in parallel and give feedback in the time of the slower of the two. The image is tagged with the commit SHA, which makes every artifact immutable and traceable to its source. Trivy fails the pipeline only for high and critical vulnerabilities that have a fix available, which keeps the gate strict without blocking on noise you cannot act on.
Staging is deployed automatically from the default branch, and kubectl rollout status makes the job fail if the new pods never become ready. Production is only deployed from a tag and requires a click, which is a simple form of human in the loop approval. The resource_group guarantees that two production deployments never run at the same time.
A GitOps Pipeline with GitLab CI and Argo CD
In modern setups, CI builds the artifact and CD is handled by a GitOps controller such as Argo CD or Flux, not by the CI server. GitLab CI never talks to the cluster: it updates a separate GitOps repository that contains the Kubernetes manifests, and Argo CD applies whatever that repository describes. Every deployment becomes a Git commit, so you get an audit trail for free and a rollback is a simple git revert.
You need a GitOps repository (for example gitlab.com:workflows-guru/website-gitops) with a Kustomize overlay per environment, a deploy key with write access to that repository stored in the GITOPS_DEPLOY_KEY CI variable, and Argo CD watching the repository.
1stages:2 - test3 - build4 - scan5 - update-gitops67variables:8 DOCKER_IMAGE: "$CI_REGISTRY_IMAGE:$CI_COMMIT_SHA"9 GITOPS_REPO: "git@gitlab.com:workflows-guru/website-gitops.git"10 OVERLAY_DIR: "kustomize/overlays/staging"11 # Image name used in the base Deployment manifest12 BASE_IMAGE: "registry.gitlab.com/workflows-guru/website"1314test:15 stage: test16 image: node:2217 script:18 - yarn install --frozen-lockfile19 - yarn test2021build_image:22 stage: build23 image: docker:2924 services:25 - docker:29-dind26 variables:27 DOCKER_TLS_CERTDIR: "/certs"28 script:29 - echo "$CI_REGISTRY_PASSWORD" | docker login -u "$CI_REGISTRY_USER" --password-stdin "$CI_REGISTRY"30 - docker build -t "$DOCKER_IMAGE" .31 - docker push "$DOCKER_IMAGE"3233image_scan:34 stage: scan35 image:36 name: aquasec/trivy:0.75.037 entrypoint: [""]38 variables:39 TRIVY_USERNAME: "$CI_REGISTRY_USER"40 TRIVY_PASSWORD: "$CI_REGISTRY_PASSWORD"41 script:42 - trivy image --exit-code 1 --severity HIGH,CRITICAL --ignore-unfixed "$DOCKER_IMAGE"4344update_gitops_repo:45 stage: update-gitops46 image: alpine:3.2247 rules:48 - if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH49 before_script:50 - apk add --no-cache git openssh-client kustomize51 - mkdir -p ~/.ssh && chmod 700 ~/.ssh52 - echo "$GITOPS_DEPLOY_KEY" | tr -d '\r' > ~/.ssh/id_ed2551953 - chmod 600 ~/.ssh/id_ed2551954 - ssh-keyscan gitlab.com >> ~/.ssh/known_hosts55 script:56 - git clone "$GITOPS_REPO" gitops57 - cd "gitops/$OVERLAY_DIR"58 - kustomize edit set image "$BASE_IMAGE:$CI_COMMIT_SHA"59 - git config user.email "ci@workflows.guru"60 - git config user.name "GitLab CI"61 - git commit -am "Deploy $CI_PROJECT_NAME $CI_COMMIT_SHORT_SHA"62 - git push origin main
The last job uses kustomize edit set image instead of editing YAML with sed, so the update is always valid and only touches the image reference. After the job runs, the staging overlay in the GitOps repository looks like this:
1# kustomize/overlays/staging/kustomization.yaml2apiVersion: kustomize.config.k8s.io/v1beta13kind: Kustomization4namespace: workflows-guru-staging5resources:6 - ../../base7images:8 - name: registry.gitlab.com/workflows-guru/website9 newTag: 3f2a1c9e8b7d6a5f4e3d2c1b0a9f8e7d6c5b4a39
Finally, an Argo CD Application tells Argo CD which repository path to watch and where to deploy it. With automated sync enabled, every commit pushed by the pipeline is applied to the cluster within minutes, and selfHeal reverts any manual change made directly in the cluster.
1apiVersion: argoproj.io/v1alpha12kind: Application3metadata:4 name: workflows-guru-staging5 namespace: argocd6spec:7 project: default8 source:9 repoURL: git@gitlab.com:workflows-guru/website-gitops.git10 targetRevision: main11 path: kustomize/overlays/staging12 destination:13 server: https://kubernetes.default.svc14 namespace: workflows-guru-staging15 syncPolicy:16 automated:17 prune: true18 selfHeal: true19 syncOptions:20 - CreateNamespace=true
Argo CD reports the health of every resource it syncs, but it does not roll back a bad release on its own. For automatic rollbacks based on metrics, add a progressive delivery controller such as Argo Rollouts or Flagger, as described later in this guide.
Running GitLab Runners on Kubernetes with Terraform
Shared runners are convenient, but running your own runners on Kubernetes gives you control over runner sizing, caching, network access and cost. The Terraform code below installs the official gitlab-runner Helm chart into any cluster you can reach with a kubeconfig, so it works the same on EKS, GKE, AKS or an on-premises cluster.
Since GitLab 16, runners are created in the GitLab UI (Settings, CI/CD, Runners, New project runner), which gives you a runner authentication token starting with glrt-. Runner tags are configured there as well. The old registration tokens are deprecated, so the code below uses the new token.
terraform/main.tf
1terraform {2 required_version = ">= 1.5"3 required_providers {4 helm = {5 source = "hashicorp/helm"6 version = "~> 3.0"7 }8 }9}1011provider "helm" {12 kubernetes = {13 config_path = var.kubeconfig_path14 }15}1617resource "helm_release" "gitlab_runner" {18 name = "gitlab-runner"19 namespace = var.namespace20 create_namespace = true21 repository = "https://charts.gitlab.io"22 chart = "gitlab-runner"23 version = var.chart_version24 cleanup_on_fail = true25 timeout = 6002627 values = [templatefile("${path.module}/gitlab-runner-values.yaml.tpl", {28 gitlab_url = var.gitlab_url29 concurrent = var.concurrent30 privileged = var.privileged31 })]3233 set_sensitive = [34 {35 name = "runnerToken"36 value = var.runner_token37 }38 ]39}4041output "runner_release" {42 value = "${helm_release.gitlab_runner.namespace}/${helm_release.gitlab_runner.name}"43}
terraform/variables.tf
1variable "kubeconfig_path" {2 type = string3 default = "~/.kube/config"4 description = "Path to the kubeconfig of the target cluster."5}67variable "namespace" {8 type = string9 default = "workflows-guru-runner"10 description = "Namespace where GitLab Runner and its job pods run."11}1213variable "gitlab_url" {14 type = string15 default = "https://gitlab.com/"16}1718variable "runner_token" {19 type = string20 sensitive = true21 description = "Runner authentication token (glrt-...) created in the GitLab UI."22}2324variable "chart_version" {25 type = string26 default = "0.93.0"27 description = "Pin the chart and upgrade deliberately."28}2930variable "concurrent" {31 type = number32 default = 433 description = "Maximum number of jobs this runner executes at the same time."34}3536variable "privileged" {37 type = bool38 default = true39 description = "Required for Docker-in-Docker builds. Disable if you build images rootless."40}
terraform/gitlab-runner-values.yaml.tpl
1gitlabUrl: "${gitlab_url}"2concurrent: ${concurrent}3rbac:4 create: true5runners:6 config: |7 [[runners]]8 [runners.kubernetes]9 namespace = "{{ .Release.Namespace }}"10 image = "alpine:3.22"11 privileged = ${privileged}12 cpu_request = "500m"13 memory_request = "1Gi"
The template is rendered by Terraform first (the ${...} placeholders) and then by Helm (the {{ ... }} expression). Apply it from the terraform/ folder, passing the token through an environment variable so that it never ends up in your shell history or in a committed file:
1cd terraform2export TF_VAR_runner_token="glrt-XXXXXXXXXXXXXXXX"3terraform init4terraform apply
The Kubernetes executor starts a fresh pod for every job, so builds are isolated and the runner scales with your cluster autoscaler. Docker-in-Docker requires privileged = true, which gives job pods broad access to the node. If that is not acceptable, keep privileged builds on a dedicated node pool or build images rootless with BuildKit, then set privileged to false.
Kubernetes Manifests: Deployment, Service and Ingress
The pipelines above deploy to a workflows-guru Deployment. These are the base manifests for a typical web application, ready to be placed in the base folder of the GitOps repository. Update the image, the host and the resource sizing to match your application.
k8s/deployment.yaml
1apiVersion: apps/v12kind: Deployment3metadata:4 name: workflows-guru5 labels:6 app: workflows-guru7spec:8 replicas: 39 selector:10 matchLabels:11 app: workflows-guru12 template:13 metadata:14 labels:15 app: workflows-guru16 spec:17 containers:18 - name: workflows-guru19 # The tag is replaced by CI with the commit SHA, never use :latest20 image: registry.gitlab.com/workflows-guru/website:0.1.021 ports:22 - containerPort: 300023 resources:24 requests:25 cpu: "100m"26 memory: "128Mi"27 limits:28 cpu: "500m"29 memory: "512Mi"30 livenessProbe:31 httpGet:32 path: /api/health33 port: 300034 initialDelaySeconds: 3035 periodSeconds: 1036 timeoutSeconds: 337 failureThreshold: 338 readinessProbe:39 httpGet:40 path: /api/ready41 port: 300042 initialDelaySeconds: 543 periodSeconds: 544 timeoutSeconds: 2
k8s/service.yaml
1apiVersion: v12kind: Service3metadata:4 name: workflows-guru5 labels:6 app: workflows-guru7spec:8 type: ClusterIP9 selector:10 app: workflows-guru11 ports:12 - name: http13 port: 8014 targetPort: 300015 protocol: TCP
k8s/ingress.yaml
This example assumes the NGINX Ingress Controller and, optionally, cert-manager to issue TLS certificates automatically from Let's Encrypt.
1apiVersion: networking.k8s.io/v12kind: Ingress3metadata:4 name: workflows-guru5 annotations:6 nginx.ingress.kubernetes.io/proxy-body-size: "10m"7 nginx.ingress.kubernetes.io/proxy-read-timeout: "60"8 cert-manager.io/cluster-issuer: "letsencrypt-prod"9spec:10 ingressClassName: nginx11 tls:12 - hosts:13 - www.workflows.guru14 secretName: workflows-guru-tls15 rules:16 - host: www.workflows.guru17 http:18 paths:19 - path: /20 pathType: Prefix21 backend:22 service:23 name: workflows-guru24 port:25 number: 80
The readiness probe is what makes rolling updates safe: Kubernetes only sends traffic to a new pod once it reports ready, and kubectl rollout status in the pipeline waits for exactly that. Keep liveness and readiness endpoints separate, so that a slow dependency removes a pod from the load balancer instead of restarting it.
Measurable Tuning: Faster and Cheaper Pipelines
It is tempting to chase raw speed, but the right approach is cost-aware tuning driven by measurement. Start by tracking three numbers: pipeline lead time (from commit to a green, deployable artifact), pipeline flakiness (the percentage of runs that fail for reasons unrelated to the code change), and time to recover (from failure detection to a safe rollback). These numbers tell you where to invest; without them, every optimization is a guess.
Caching is usually the cheapest win. Keying the dependency cache on the lockfile, as in the pipeline above, means dependencies are only downloaded again when they actually change. Depending on the language and on the cache hit rate, this commonly removes 30 to 70 percent of the build time. Watch the hit rate: a cache that is invalidated on every commit costs upload time and saves nothing.
Parallelism is the second lever. Independent jobs belong in the same stage, and GitLab's needs keyword turns stages into a DAG so a job starts as soon as its own dependencies finish. Large test suites can be split across several jobs with parallel: N. Move slow end-to-end tests out of the merge request path and run them on the default branch or on a nightly schedule.
Runner sizing is the third lever. Ephemeral runners that scale with demand avoid paying for idle machines while still absorbing peaks. Instrument the queue time (how long jobs wait for a runner) separately from the execution time: a long queue means you need more runners, a long execution means you need bigger runners or faster jobs. Once you measure both, set targets for lead time and flakiness and review them like any other SLO.
Security and Compliance: Where to Place Checks
Security must be part of the pipeline, not an afterthought. The most effective approach is layered: fast and cheap checks run early on every merge request, while deeper and slower scans run later, but always before a production promotion.
In the merge request stage, run linters, secret detection and software composition analysis (SCA) on your dependencies, including license checks. These run in seconds and catch the most common problems, such as a committed API key or a vulnerable library version. After the build, scan the container image itself, as the Trivy job does above, because base images bring their own vulnerabilities. Full static analysis (SAST) and dynamic testing (DAST) against a staging environment take longer and fit best on the default branch, before the production deployment.
Secrets should never appear in the repository or in job logs. Store them in a secrets manager or in masked and protected CI variables, give each job only the credentials it needs, and prefer short-lived credentials such as OIDC tokens over long-lived keys. For regulated releases, combine protected branches and tags with a manual approval on the production environment, so that every production change has a reviewer and an audit trail.
Progressive Delivery and Observability
Progressive delivery turns deployments from all-or-nothing events into measured and reversible operations. With a canary release, the new version first receives a small share of the traffic, for example 5 percent, then 25 and 50 percent, before it replaces the old version completely. With a blue/green deployment, the new version runs next to the old one and the traffic switches all at once, which makes the rollback instantaneous. Feature flags go one step further and separate deploying code from releasing a feature to users.
Each promotion step should be gated by health signals: request error rate, latency percentiles and, when possible, a business metric such as checkout success. Tools like Argo Rollouts and Flagger query your metrics provider at every step, promote the release automatically when the thresholds hold, and roll back automatically when they do not. A typical rule is to roll back if the error rate of the canary exceeds the stable version by more than one percentage point, or if its p95 latency regresses by more than 20 percent.
Observability is not optional here. Pipeline telemetry, deployment success rate and real user metrics must feed the promotion logic, otherwise a canary is just a slower deployment.
Popular CI/CD Tools
There is no shortage of tools in the CI/CD ecosystem. The patterns in this guide apply to all of them; only the syntax changes.
GitHub Actions is the natural choice for teams already on GitHub, with YAML workflows and a huge marketplace of pre-built actions. Our GitHub Actions tutorial walks through a complete example. GitLab CI/CD offers easy pipeline definitions, a built-in container registry, tight integration with GitLab repositories and scalable self-hosted runners. Jenkins is highly customizable through thousands of plugins and supports very complex workflows, but requires more maintenance than hosted alternatives.
CircleCI is known for fast builds with caching, parallelism and good Docker support, and integrates easily with GitHub and Bitbucket. Travis CI offers a simple setup that has long been popular with open-source projects. Azure DevOps Pipelines is a powerful, enterprise-ready option for Microsoft-centric teams working with Azure, .NET and Windows, although its interface can feel complex for beginners.
Bamboo, from Atlassian, integrates closely with Jira and Bitbucket. TeamCity, from JetBrains, provides powerful build management for large teams with complex needs. AWS CodePipeline is a fully managed, pay-as-you-go service with native integrations for workloads running on AWS, but is less flexible than GitHub Actions or GitLab. On the delivery side, Argo CD and Flux are the reference GitOps controllers for Kubernetes, and you can orchestrate more advanced workflows inside the cluster with Argo Workflows.
How to Migrate: a Six-Week Plan
Teams do not get CI/CD right by flipping a switch. A phased migration reduces risk and creates early advocates who help with the rest of the rollout.
Weeks 1 and 2: start small. Select one low-risk service, document the current manual steps, and create a pipeline that runs the basic tests and produces an immutable artifact. Record baseline metrics now: deployment frequency, lead time for changes, change failure rate and time to restore service. You will need them to prove progress.
Weeks 3 and 4: make it fast and safe. Add caching, parallel jobs and the security checks described above. Measure the lead time and the flakiness, then fix the flaky tests or move them out of the fast path.
Week 5: automate delivery. Introduce GitOps for deployments and a progressive delivery mechanism with automated metric checks. Define the rollback rules and the alerts before the first production release.
Week 6: expand and govern. Onboard more services using the first pipeline as a template, add governance with branch protections and approval policies, and keep tracking the four delivery metrics alongside pipeline flakiness. In larger organizations, share a KPI dashboard and a short status update with stakeholders from the first sprint, so that the progress is visible.
Pitfalls and How to Avoid Them
The most common mistake is trying to automate everything at once. If your test suite takes hours, moving it into CI without refactoring only makes the feedback loop slower. Split the tests into tiers, keep the fast tier on every commit, and push long-running or expensive tests to scheduled pipelines.
The second mistake is not measuring. If you cannot see the lead time or the flakiness, you are blind to regressions caused by changes to the pipeline itself. Instrument the pipeline runs and the runners, and put those metrics on the dashboards developers already look at.
The third mistake is treating the pipeline as a one-off project. Base images, runner versions and scanners age quickly; pin their versions, upgrade them deliberately, and review the pipeline like any other production code.
Wrapping Up
High-level explainers cover the what and the why of CI/CD. What drives adoption is the how: reproducible pipeline templates, measurable tuning, security checks placed at the right stage, and a staged migration plan. Without CI/CD, every feature rollout is risky; with it, automated tests stop regressions before production and standardized pipelines make every team follow the same build, test and release process.
Whether you choose GitHub Actions, GitLab, Jenkins or Argo CD, the important thing is adopting a pipeline mindset: automate early, automate often, and measure everything. The result is faster releases, fewer production incidents, and more developer time spent on the product instead of the plumbing.