Deploys and supports distributed, multi-tiered systems at scale while ensuring high availability and fault tolerance across multiple environments. Builds and operates resilient platforms in Amazon Web Services (AWS) using Elastic Compute Cloud (EC2), Simple Storage Service (S3), and Auto Scaling Groups for dynamic resource management. Designs, develops, and executes performance tests using Java-based frameworks, Apache JMeter, k6, and Rush-hour to validate system behavior under day-to-day traffic patterns. Defines and implements observability practices to monitor system health, latency, and error rates through metrics, logs, and distributed tracing using Datadog, Grafana, Splunk, and the Elasticsearch, Logstash, and Kibana (ELK) stack. Automates operational workflows with Python and Shell scripting to enhance efficiency and reduce manual tasks. Supports consistent build, deployment, and orchestration processes using cloud computing and DevOps technologies -- Continuous Integration and Continuous Delivery (CI/CD) pipelines and Kubernetes. Supports Site Reliability Engineering (SRE) functions by establishing Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets, and implementing proactive monitoring and incident response strategies. Builds and refines methodologies for performance, load, stress, and chaos testing and develops analytics and reports aligned with business needs to improve system resilience and optimization.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Principal