Lead DevOps Engineer (Application & IT Operations)

Exadel Inc (Website)
$90 - $100Remote

About The Position

Exadel is seeking a Lead DevOps Engineer (Application & IT Operations) to join their team. This role involves mentoring AppOps engineers, owning production reliability, leading incident response, and architecting operational observability. The ideal candidate will have extensive experience in DevOps and AppOps, with a focus on operating critical applications in enterprise cloud environments (Azure and/or AWS).

Requirements

  • Extensive, lead-level experience in DevOps engineering and AppOps, with a focus on operating critical applications in enterprise cloud environments (Azure and/or AWS).
  • Deep technical knowledge of cloud infrastructure services and components, including networking, load balancers, DNS, SSL/TLS certificates, storage, and messaging services.
  • Hands-on expertise with enterprise observability stacks (such as Datadog, Grafana, Prometheus, ELK/OpenSearch, and OpenTelemetry), alert engineering, and log/metric/trace analysis.
  • Solid practical understanding of continuous integration and continuous deployment (CI/CD) pipelines, multi-environment application lifecycles (Dev, QA, UAT, Prod), and validation strategies in lower/production environments.
  • Proficiency in deployment strategies (including blue/green, rolling, canary) and traffic management.
  • Advanced scripting and automation skills using PowerShell, Bash, or Python to develop runbooks, health checks, self-healing, and remediation workflows.
  • Strong command of Infrastructure as Code (IaC), specifically with Terraform (modules, workspaces), Azure ARM/Bicep, or AWS CloudFormation, including environment drift detection and policy-as-code.
  • Practical understanding of security and compliance protocols in operations, including secrets and key management, vulnerability remediation, least-privilege access, and audit readiness.
  • Proven track record of engineering leadership, stakeholder management, technical mentorship, and leading incident response / RCA processes in global, fast-paced environments.
  • Excellent communication and advisory skills, with the ability to translate technical risks into clear business metrics for stakeholder decision-making.
  • Strong ownership mindset and the ability to ensure 24x7 application reliability and operational excellence.
  • Bachelor’s degree in Computer Science, Information Systems, or a related field, or equivalent practical experience.

Nice To Haves

  • Preferred Certifications: Azure Administrator/Architect, AWS SysOps/DevOps Professional, ITIL Foundation (or higher), SRE Foundation, Terraform Associate/Professional, or other industry-recognized DevOps/SRE credentials.

Responsibilities

  • Mentor AppOps engineers, providing technical guidance, conducting code/review for automation scripting, and developing on-call operational excellence.
  • Own production reliability for critical applications by defining, tracking, and enforcing SLOs, SLAs, error budgets, and capacity/performance baselines.
  • Lead major incident response and production triage, driving clear business and technical communications, and ensuring data-driven root cause analysis (RCA) with long-term preventative actions.
  • Direct release, deployment, and change operations by coordinating application deployments, assessing operational risks, enforcing readiness gates, ensuring compliance with client change processes, and validating post-deployment health to improve change success rates.
  • Architect and maintain operational observability by designing and implementing enterprise dashboards, alert strategies, log/trace pipelines, and runbook automation for rapid system diagnosis and recovery.
  • Establish and continuously improve operational standards, guardrails, and runbooks, while automating repeatable workflows and repetitive tasks to systematically reduce manual toil and improve operational efficiency.
  • Partner cross-functionally with Engineering, CloudOps, Security, and Compliance teams to resolve issues, improve service quality, and consult on resiliency patterns (including circuit breakers, bulkheads, graceful degradation, and retries) and performance tuning.
  • Plan and execute capacity management, scaling strategies, and Disaster Recovery (DR)/Business Continuity Planning (BCP) readiness, including failover testing and simulated scenario exercises.
  • Champion security-by-default and compliance alignment in operations by enforcing secrets hygiene, patch/vulnerability remediation, certificate/DNS management, least-privilege access, and general security standard adherence.
  • Drive service reviews with stakeholders, publishing key operational KPIs (such as MTTR, change success rate, and incident rate) to lead continuous improvement roadmaps.
  • Monitor application health, availability, and performance across all environments, proactively identifying anomalies, resolving application environment issues, and optimizing runtime behavior.
  • Participate in on-call rotation responsibilities with the Service Delivery and Operations Team.

Benefits

  • Comprehensive health, dental, and vision coverage
  • Life and disability insurance
  • Retirement savings programs
  • Paid time off
  • Paid holidays
  • Other wellness or voluntary benefit programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service