Exam results at state scale
Highly available, auto-scaling AWS infrastructure that published state-level exam results, absorbing millions of concurrent requests during extreme traffic spikes without going down.
Site Reliability Engineer
Site Reliability / Senior DevOps Engineer and Certified Kubernetes Administrator with 10+ years operating highly available production infrastructure on AWS, Azure and OpenStack. I take production on-call, resolve incidents, and turn the lessons into automation, runbooks and permanent fixes.
A few systems I've designed, built and kept running.
Highly available, auto-scaling AWS infrastructure that published state-level exam results, absorbing millions of concurrent requests during extreme traffic spikes without going down.
A Python portal with REST APIs that lets development and QA teams provision environments, read logs and trigger deployments on their own, without filing an ops ticket.
Alert analysis and automated runbook generation built on Prometheus, Alertmanager and Grafana, backed by LLM-powered RAG pipelines in Python for faster operational insight.
Moved marketplace image builds to Packer, taking image creation from days to about an hour, and Dockerised the unit test suite on EKS as a self-service testing platform.
Senior DevOps Engineer · Remote
DevOps & IT Infrastructure Engineer · Kochi
Systems Engineer · Kochi
Jyothi Engineering College, Thrissur
Looking for someone to own reliability, Kubernetes, or the platform your developers build on? I'd like to hear about it.