All systems operational

Aswin Baby

Site Reliability Engineer

Site Reliability / Senior DevOps Engineer and Certified Kubernetes Administrator with 10+ years operating highly available production infrastructure on AWS, Azure and OpenStack. I take production on-call, resolve incidents, and turn the lessons into automation, runbooks and permanent fixes.

Thrissur, Kerala, India IST (UTC+5:30) Remote since 2020
Experience
10+ yrs
Operating production infrastructure end to end
CI/CD
50%
Faster builds across Linux, Windows and macOS pipelines
Self-service
60%
Fewer ops tickets after shipping a developer portal
Cost
40%
Infrastructure cost cut by migrating VMware estates
01

Selected work

A few systems I've designed, built and kept running.

0 downtime

Exam results at state scale

Highly available, auto-scaling AWS infrastructure that published state-level exam results, absorbing millions of concurrent requests during extreme traffic spikes without going down.

AWS ALBAuto ScalingAurora Multi-AZHAProxyNginx
−60% tickets

Self-service developer portal

A Python portal with REST APIs that lets development and QA teams provision environments, read logs and trigger deployments on their own, without filing an ops ticket.

PythonFlaskFastAPITerraform
LLM + RAG

AI-assisted observability

Alert analysis and automated runbook generation built on Prometheus, Alertmanager and Grafana, backed by LLM-powered RAG pipelines in Python for faster operational insight.

PrometheusAlertmanagerGrafanaPythonRAG
days → ~1 hr

Marketplace image pipeline

Moved marketplace image builds to Packer, taking image creation from days to about an hour, and Dockerised the unit test suite on EKS as a self-service testing platform.

PackerDockerEKSJenkins
02

Experience

FileCloud

Senior DevOps Engineer · Remote

Dec 2020 — Present

Reliability, incident response & observability

  • Work a regular production on-call rotation for customer-facing services — owning alert triage, incident response and resolution, then driving the follow-up automation, runbooks and permanent fixes that cut repeat alerts and MTTR.
  • Built centralised observability with Prometheus, Thanos, Grafana, Loki and ELK: long-term metric storage, multi-cluster dashboards and centralised log search across all environments.
  • Tuned Alertmanager routing and alert rules to remove noisy pages and route actionable alerts to owning teams; integrated Sentry for application error tracking to shorten root-cause analysis.
Show more from FileCloud
  • Developed AI-assisted observability workflows using Prometheus, Alertmanager, Grafana and LLM-powered RAG pipelines (Python) for alert analysis, automated runbook generation and operational insight.
  • Designed disaster recovery and backup strategies — Velero cluster backups, database point-in-time recovery, cross-region replication — against defined RTO/RPO targets, validated by restore drills.

Cloud infrastructure, Kubernetes & data layer

  • Built and operated production Kubernetes clusters on AWS EKS and Azure AKS — highly available control planes, autoscaling, resource governance, Helm releases and GitOps delivery with ArgoCD.
  • Led the Infrastructure as Code initiative, building reusable Terraform modules that provision consistent Dev, Staging and Production estates plus on-demand environments for development and QA.
  • Administered the data and messaging layer behind production services — PostgreSQL, MySQL, MongoDB, Redis and RabbitMQ — covering provisioning, clustering, replication health, upgrades, backups and performance tuning.
  • Executed infrastructure and service migrations across hybrid environments (AWS, Azure, OpenStack, MacStadium, Linode) with minimal downtime.
  • Diagnosed and resolved complex Linux issues across RHEL, Debian and Ubuntu fleets — networking, storage, systemd services, resource exhaustion and application performance bottlenecks.
  • Implemented secrets management with HashiCorp Vault and Azure Key Vault, OPA/Gatekeeper and RBAC cluster policies, and OIDC single sign-on; supported SOC 2 and ISO 27001 audits.

Automation, CI/CD & developer enablement

  • Built and enhanced CI/CD for Linux, Windows and macOS builds with Jenkins, GitHub Actions, Bitbucket Pipelines, Ansible and Terraform — cutting build times by 50% and making deployments and rollbacks routine.
  • Built a self-service developer portal in Python (Flask/FastAPI) so development and QA teams provision environments, access logs and trigger deployments independently — reducing ops tickets by 60%.
  • Developed reusable CLI tools in Python and Go to standardise infrastructure operations; migrated marketplace image builds to Packer, cutting image creation from days to about an hour.

Fingent Technology Solutions

DevOps & IT Infrastructure Engineer · Kochi

Dec 2018 — Dec 2020
  • Architected and deployed cloud infrastructure on AWS, OpenStack and Azure for enterprise clients — landing zones, multi-account governance, IAM, VPC networking and firewall design.
  • Migrated production from VMware to an OpenStack private cloud, and Dev/QA estates from VMware to Proxmox clusters, reducing infrastructure cost by 40%.
  • Containerised legacy applications with Docker and ran them on Amazon ECS and Kubernetes using reusable Helm charts.
  • Automated configuration and provisioning across 100+ Linux servers with Ansible; built Jenkins pipelines with Bash, Python and PowerShell automation.

Supportsages Consultancy Services

Systems Engineer · Kochi

Jun 2016 — Nov 2018
  • Architected a highly available, auto-scaling AWS infrastructure to publish state-level exam results, absorbing millions of concurrent requests with zero downtime.
  • Designed high-availability web stacks with HAProxy, Nginx and PHP-FPM on AWS load balancers, Auto Scaling Groups and Multi-AZ Aurora.
  • Managed PostgreSQL, MySQL and MongoDB on RDS, Aurora and VMs — provisioning, replication, backups and performance tuning.
  • Automated 100+ bare-metal and virtual Linux servers with Ansible; containerised monolithic applications with Docker.
03

Toolbox

cloudCloud

AWSEC2EKSECSRDS/AuroraS3VPCIAMRoute 53AzureOpenStackLinode

k8sContainers & orchestration

KubernetesDockerAKSHelmArgoCDKVMProxmoxVMware

iacInfrastructure as code

TerraformAnsiblePacker

ciCI/CD

JenkinsGitHub ActionsBitbucket PipelinesArgoCDGitSonarQubeSnyk

o11yObservability

PrometheusGrafanaAlertmanagerLokiThanosSentryELKCloudWatchZabbix

dataDatabases & messaging

PostgreSQLMongoDBMySQLRedisRabbitMQ

langLanguages

PythonBashGoPowerShell

sysLinux & networking

RHELDebian/UbuntusystemdNginxHAProxyVPC networkingOIDCRBAC

secSecurity & compliance

HashiCorp VaultAzure Key VaultOPA/GatekeeperSOC 2ISO 27001
04

Certifications & education

CKA
Certified Kubernetes AdministratorCNCF · Linux Foundation
SAA
AWS Certified Solutions Architect – AssociateAmazon Web Services · SAA-C03
104
Microsoft Azure AdministratorMicrosoft · AZ-104
900
Microsoft Azure FundamentalsMicrosoft · AZ-900
2012 — 2016 B.Tech, Electronics & Communication Engineering

Jyothi Engineering College, Thrissur

Open to conversations

Let's keep your systems up.

Looking for someone to own reliability, Kubernetes, or the platform your developers build on? I'd like to hear about it.