Open to cloud & platform roles

Hi, I'm Shubham

Cloud Engineer

I build and operate production cloud infrastructure that stays up, scales, and costs less. Three years operating production AWS from Remote, India — Terraform, ECS Fargate, CI/CD, and the observability that catches things first.

Shubham DwivediSD
Years in production cloud
3+Years in production cloud
Microservices supported
30+Microservices supported
Uptime maintained
99.9%Uptime maintained
Monthly AWS cost saved
15%Monthly AWS cost saved

About

// who you're hiring

I work on the infrastructure layer that business-critical applications depend on. Day to day that means AWS, Terraform, containers on ECS Fargate, CI/CD pipelines, and the observability tooling that tells us when something is wrong before a customer notices.

Most of my work sits at the intersection of reliability and automation. I have supported 30+ microservices in production, maintained 99.9% uptime, run disaster recovery and failover validation, and been on the hook for root-cause analysis when deployments or connectivity break.

I care a lot about removing manual toil. Reusable Terraform modules cut provisioning from hours to minutes, deployment validation and rollback safeguards cut deployment-related interruptions by 30%, and standardized runbooks brought MTTR down by 20%.

More recently I have been building AI-assisted engineering workflows — a Terraform Engineering Copilot that takes infrastructure requests from Jira through to a reviewed Pull Request, and a FinOps Dashboard Agent that analyzes AWS spend and surfaces rightsizing opportunities.

I also work closely with security and audit: enforcing IAM least-privilege, supporting external audits, and keeping infrastructure change auditable by default.

  • AWS infrastructure & high availability
  • Infrastructure as Code with Terraform
  • CI/CD and release reliability
  • Observability & incident response
  • Cloud cost optimization (FinOps)
  • AI-driven engineering automation

Career Path

// three years, production only

  1. Tata Consultancy Services (TCS)

    Cloud Engineer · Remote, India

    Dec 2024 – Present
    • Managed AWS infrastructure supporting 30+ microservices across ECS Fargate, EFS, NLB, Lambda, IAM, and CloudWatch — maintaining 99.9% uptime for business-critical workloads.
    • Investigated deployment failures, connectivity issues, and performance bottlenecks using logs, metrics, and application telemetry; partnered with dev teams on root-cause analysis and service restoration.
    • Introduced deployment validation and rollback safeguards in production, reducing deployment-related service interruptions by 30%.
    • Enforced IAM least-privilege and security controls, supported 5+ external audits, and executed disaster recovery and failover validation across 30+ AWS-hosted services.
    • Built a Terraform Engineering Copilot (GitHub Copilot, Terraform, Jira, Node.js, PowerShell) to automate infrastructure validation and PR workflows, cutting manual engineering effort by 50%.
    • Developed a FinOps Dashboard Agent to analyze AWS spend and surface optimization opportunities, driving 15% monthly AWS cost savings.
    AWSECS FargateTerraformDockerCloudWatchIAMNode.js
  2. Tata Consultancy Services (TCS)

    Cloud Engineer · Pune, India

    Oct 2023 – Dec 2024
    • Developed reusable Terraform modules for EC2, S3, and RDS provisioning, reducing configuration errors by 40% and shortening provisioning from hours to minutes.
    • Designed and maintained Bitbucket CI/CD pipelines with automated testing and release validation, reducing manual release effort by 50%.
    • Troubleshot deployment failures and standardized runbooks, reducing MTTR by 20%.
    TerraformAWS EC2S3RDSBitbucket PipelinesCI/CD

Tech Stack

// the arsenal for running production infrastructure

  • AWS
  • ECS / Fargate
  • Lambda
  • EFS
  • NLB
  • IAM
  • CloudFormation
  • GCP
  • Terraform
  • Infrastructure as Code
  • Docker
  • Kubernetes
  • Helm
  • Ansible
  • CI/CD
  • AWS
  • ECS / Fargate
  • Lambda
  • EFS
  • NLB
  • IAM
  • CloudFormation
  • GCP
  • Terraform
  • Infrastructure as Code
  • Docker
  • Kubernetes
  • Helm
  • Ansible
  • CI/CD
  • Python
  • Bash
  • Node.js
  • PowerShell
  • Linux / Unix
  • Git
  • Bitbucket
  • GitLab
  • Jenkins
  • Elasticsearch
  • GitHub Copilot
  • Cursor
  • AI Automation
  • Infrastructure Automation
  • Self-Service Tooling
  • Technical Documentation
  • Python
  • Bash
  • Node.js
  • PowerShell
  • Linux / Unix
  • Git
  • Bitbucket
  • GitLab
  • Jenkins
  • Elasticsearch
  • GitHub Copilot
  • Cursor
  • AI Automation
  • Infrastructure Automation
  • Self-Service Tooling
  • Technical Documentation
  • CloudWatch
  • Grafana
  • Prometheus
  • OpenTelemetry
  • Tracing
  • Logging
  • Metrics
  • Incident Response
  • Disaster Recovery
  • Troubleshooting
  • Issue Triage
  • TCP/IP
  • DNS
  • HTTP
  • Load Balancing
  • IAM Least-Privilege
  • Audit Support
  • Scalability
  • Performance
  • CloudWatch
  • Grafana
  • Prometheus
  • OpenTelemetry
  • Tracing
  • Logging
  • Metrics
  • Incident Response
  • Disaster Recovery
  • Troubleshooting
  • Issue Triage
  • TCP/IP
  • DNS
  • HTTP
  • Load Balancing
  • IAM Least-Privilege
  • Audit Support
  • Scalability
  • Performance

Cloud & DevOps

AWSECS / FargateLambdaEFSNLBIAMCloudFormationGCPTerraformInfrastructure as CodeDockerKubernetesHelmAnsibleCI/CD

Systems & Programming

PythonBashNode.jsPowerShellLinux / UnixGitBitbucketGitLabJenkinsElasticsearch

Observability & Reliability

CloudWatchGrafanaPrometheusOpenTelemetryTracingLoggingMetricsIncident ResponseDisaster RecoveryTroubleshootingIssue Triage

Networking & Security

TCP/IPDNSHTTPLoad BalancingIAM Least-PrivilegeAudit SupportScalabilityPerformance

AI & Automation

GitHub CopilotCursorAI AutomationInfrastructure AutomationSelf-Service ToolingTechnical Documentation

Selected Projects

// internal tools that removed real toil

Terraform Engineering Copilot

A repository-aware AI workflow that reads Terraform repositories and generates infrastructure changes aligned with existing repo standards.

Takes an infrastructure request from Jira analysis all the way through Terraform validation, Git automation, and human approval, ending in a Bitbucket Pull Request.

  • Reduced manual engineering effort by 50%
  • Human-in-the-loop approval before every merge
GitHub CopilotTerraformJiraNode.jsPowerShellBitbucket

FinOps Dashboard Agent

An AI-powered agent that analyzes cloud cost data, identifies underutilized resources, and surfaces infrastructure optimization opportunities.

Provides actionable cloud cost analytics and visualization so teams can make proactive rightsizing and cloud-spend decisions instead of reacting to the monthly bill.

  • Helped drive 15% monthly AWS cost savings
  • Continuous rightsizing recommendations
AWS Cost ExplorerPythonData VisualizationAI Automation

Education

// credentials

  • Jaypee Institute of Information Technology

    Bachelor of Technology — Electronics and Communication · Noida, India

    Jul 2019 – Jun 2023
  • Claude Certified Developer – Foundations

    Anthropic

    Aug 2026
  • Oracle Certified Cloud Architect – Associate

    Oracle

    Oct 2025

Playground

// a 20-second taste of life on-call

3AM On-Call

Imagine your phone buzzing at 3am: servers are misbehaving and it's your job to keep the website online. Each tile below is a server. When one lights up, something is wrong — tap it to fix it before it goes down. You get 20 seconds, and the alerts come faster as the clock runs out.

  • Outage — server is down · 1 point
  • Traffic spike — sudden rush of users · 3 points

Scoreboard

Score

0

Best

0

Outages fixed

0

Spikes handled

0

Missed

0

Uptime

100.0%

Ready20.0s

20 seconds. Keep the fleet alive.

You are the01234567890123456789th visitor

Contact

// socials

Open to cloud, platform, and SRE roles — and always happy to talk Terraform, reliability, or cloud cost.

shubhamdwivedi034@gmail.com