I’m a Site Reliability Engineer (SRE) with 2+ years of experience ensuring high availability, scalability, and performance of production systems in enterprise, SLA-driven environments. I specialize in incident management, observability, production support, and automation, with a strong focus on proactive monitoring and preventing production outages before they impact users.
Impact I deliver:
- Reduced MTTR by 20-30% through efficient incident triaging, RCA, and log analysis. - Automated health checks, alerting, and operational workflows, saving 9.5+ hours/week. - Identified recurring production issues and failure patterns using Dynatrace, Splunk and Snaplogic reducing repeat incidents and improving stability. - Supported CI/CD pipelines (Azure DevOps / GitHub Actions) by monitoring deployments, validating builds, and assisting in faster rollback/recovery in case of failures.
I’m actively seeking opportunities to contribute to building scalable, resilient, and highly available cloud systems, while continuing to grow deeper in SRE / DevOps engineering, cloud-native technologies, and automation practices.
Open to Work
Target role
DevOps / Platform Engineer
Availability
Open to opportunities
Experience
Beginner
Working preference
Hybrid · Gurgaon
DevopsSRECI/CDPythonBash ScriptingLinux/Windows ServerObservabilityDockerKubernetesAWS/Azure servicesTerraformAnsibleServiceNowETL/SSISSQL jobITIL/ITSMIncident managementChange managementApplication support