At Edelta Corporation, I inherited a deployment process that was entirely manual. Developers SSH'd into EC2 instances, pulled code, ran npm install, and prayed. For 16+ client portals across different environments, this was unsustainable.
Here's how I designed the CI/CD architecture that now deploys all of them automatically.
The Starting Point
Before automation:
- Deployments took 45+ minutes per portal
- No rollback strategy — failures meant another manual deployment
- No visibility into what was deployed where
mainbranch was not always production-ready- Environment variables were managed ad hoc across servers
The Target Architecture
The Workflow File
Here's the core deploy.yml pattern I used for each portal:
name: Build & Deploy
on:
push:
branches: [main]
jobs:
build-and-push:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Configure AWS Credentials
uses: aws-actions/configure-aws-credentials@v4
with:
aws-access-key-id: ${{ secrets.AWS_ACCESS_KEY_ID }}
aws-secret-access-key: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
aws-region: ap-south-1
- name: Login to Amazon ECR
id: login-ecr
uses: aws-actions/amazon-ecr-login@v2
- name: Build, tag, and push Docker image
env:
ECR_REGISTRY: ${{ steps.login-ecr.outputs.registry }}
IMAGE_TAG: ${{ github.sha }}
run: |
docker build -t $ECR_REGISTRY/${{ vars.ECR_REPO }}:$IMAGE_TAG .
docker push $ECR_REGISTRY/${{ vars.ECR_REPO }}:$IMAGE_TAG
echo "image=$ECR_REGISTRY/${{ vars.ECR_REPO }}:$IMAGE_TAG" >> $GITHUB_OUTPUT
deploy:
needs: build-and-push
runs-on: ubuntu-latest
steps:
- name: Download task definition
run: |
aws ecs describe-task-definition \
--task-definition ${{ vars.ECS_TASK }} \
--query taskDefinition > task-definition.json
- name: Update ECS task definition with new image
id: task-def
uses: aws-actions/amazon-ecs-render-task-definition@v1
with:
task-definition: task-definition.json
container-name: ${{ vars.CONTAINER_NAME }}
image: ${{ needs.build-and-push.outputs.image }}
- name: Deploy to ECS
uses: aws-actions/amazon-ecs-deploy-task-definition@v1
with:
task-definition: ${{ steps.task-def.outputs.task-definition }}
service: ${{ vars.ECS_SERVICE }}
cluster: ${{ vars.ECS_CLUSTER }}
wait-for-service-stability: true
- name: Notify Slack on success
if: success()
uses: slackapi/slack-github-action@v1
with:
payload: |
{"text": "✅ *${{ github.repository }}* deployed successfully by ${{ github.actor }} — commit: ${{ github.sha }}"}
env:
SLACK_WEBHOOK_URL: ${{ secrets.SLACK_WEBHOOK }}
- name: Notify Slack on failure
if: failure()
uses: slackapi/slack-github-action@v1
with:
payload: |
{"text": "🚨 *${{ github.repository }}* deployment FAILED — ${{ github.sha }} — <${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}|View Run>"}
env:
SLACK_WEBHOOK_URL: ${{ secrets.SLACK_WEBHOOK }}
Managing Secrets Across 16+ Portals
The challenge with multiple portals is secret sprawl. My approach:
- Shared secrets (AWS credentials, Slack webhook) → GitHub Organization Secrets, available to all repos
- Per-portal secrets (DB URLs, API keys) → Repository Secrets scoped to each portal repo
- Environment-specific config (staging vs production URLs) → GitHub Variables (non-sensitive), not secrets
This separation keeps the pipeline DRY while maintaining security boundaries.
Zero-Downtime Rollouts with ECS
AWS ECS rolling deployments with wait-for-service-stability: true ensures:
- Old containers stay alive until new ones pass health checks
- If health checks fail, ECS rolls back automatically
- The GitHub Action step fails, triggering the Slack failure notification
We configured health checks with:
- Interval: 30 seconds
- Unhealthy threshold: 2 consecutive failures
- Path:
/api/healthendpoint returning{ status: 'ok', version: process.env.IMAGE_TAG }
The version in the health response made it trivial to verify which commit was actually running in production.
Results
| Metric | Before | After |
|---|---|---|
| Deployment time | 45+ mins | ~3.5 mins |
| Failed deployments reaching production | ~2 per month | 0 since rollout |
| Rollback time | 30–60 mins | Automatic (ECS) |
| Deployment visibility | None | Slack + GitHub UI |
| Secrets management | Scattered | Centralized + scoped |
The CI/CD investment paid back its build time within the first week of operation. If you're still manually deploying any service in 2025, this is the week to fix that.