Expert Automation & Observability Engineer

Ensono
$140,000 - $180,000Hybrid

About The Position

At Ensono, our Purpose is to be a relentless ally, disrupting the status quo and unleashing our clients to Do Great Things! We enable our clients to achieve key business outcomes that reshape how our world runs. As an expert technology adviser and managed service provider with cross-platform certifications, Ensono empowers our clients to keep up with continuous change and embrace innovation. We can Do Great Things because we have great Associates. The Ensono Core Values unify our diverse talents and are woven into how we do business. These five traits are the key to achieving our purpose: Honesty, Reliability, Curiosity, Collaboration, and Passion. About the role and what you'll be doing: We are seeking an Expert Observability Engineer to serve as the strategic technical lead and architect for our enterprise Observability, APM, and Telemetry ecosystems. You will lead the transformation from decentralized, reactive monitoring to a unified, automated, and proactive observability framework. Operating across hybrid cloud, Kubernetes, and legacy environments, you will design scalable architectures, drive Site Reliability Engineering (SRE) practices, and lead First-of-a-Kind (FOAK) technology implementations to ensure maximum service reliability. We want all new Associates to succeed in their roles at Ensono. That's why we've outlined the job requirements below. To be considered for this role, it's important that you meet all Required Qualifications. If you do not meet all of the Preferred Qualifications, we still encourage you to apply.

Requirements

  • Observability & APM: IBM Instana, Grafana (Enterprise & Alloy), Prometheus, OpenTelemetry, Telegraf, InfluxDB.
  • Legacy/Traditional Monitoring: SolarWinds, Netcool, Elastic/Splunk.
  • Cloud & Containerization: Kubernetes, Docker, OpenShift, AWS/Azure/GCP.
  • Infrastructure: Linux (RHEL), Windows Server, VMware, Citrix VDI, load balancers, and edge proxies.
  • Automation & DevOps: Ansible, Terraform, Python, Bash, Webhooks, CI/CD (GitHub Actions/GitLab/Jenkins).
  • ITSM/Operations: ServiceNow, ITIL 4, advanced Major Incident Management.
  • 12+ years of total IT experience, with a minimum of 5 to 7 years functioning as a Lead Architect, SRE, or Principal Observability Engineer in a massive enterprise environment.
  • Proven track record of migrating organizations from legacy monitoring to proactive, automated observability platforms.
  • Hands-on expertise in building scalable, secure telemetry pipelines and time-series databases.
  • Extensive experience leading FOAK rollouts and complex vendor/operations transition (KT) programs.

Nice To Haves

  • CKA (Certified Kubernetes Administrator), Cloud Architect (AWS/Azure), or specific APM/Observability vendor certifications.

Responsibilities

  • Architect and govern a unified observability framework covering metrics, logs, traces, and events using IBM Instana, Grafana, OpenTelemetry, Telegraf, and InfluxDB.
  • Lead First-of-a-Kind (FOAK) implementations—evaluating new observability tech and converting them into secure, repeatable, production-ready patterns.
  • Define enterprise standards for telemetry pipelines, data retention, high-cardinality controls, and observability cost management.
  • Define and govern Service Level Indicators (SLIs), Objectives (SLOs), and error budgets.
  • Serve as the senior technical escalation point, leading major P1/P2 incident war rooms and conducting evidence-based Root Cause Analysis (RCA).
  • Drastically reduce MTTD/MTTR and alert noise through event correlation, dynamic thresholds, and dependency mapping.
  • Drive Observability-as-Code and infrastructure automation using Ansible, Terraform, Python, and GitOps.
  • Automate the deployment, configuration, and self-healing workflows for monitoring agents and telemetry collectors.
  • Integrate observability platforms seamlessly with ITSM (ServiceNow), Netcool, and CI/CD pipelines.
  • Design deep observability for Docker, Kubernetes, microservices, and multi-cloud environments (Azure/AWS/GCP).
  • Correlate application APM telemetry with Kubernetes control planes, pods, nodes, and infrastructure dependencies.
  • Ensure secure-by-design telemetry pipelines (RBAC, TLS, secrets management, and image scanning).
  • Lead complex Knowledge Transfer (KT) programs, vendor transitions, and operational readiness handovers for global 24x7 teams.
  • Mentor cross-functional engineering teams and influence enterprise technology roadmaps.

Benefits

  • Unlimited Paid Days Off
  • Three health plan options
  • 401k with company match
  • Eligibility for dental, vision, short and long-term disability, life and AD&D coverage, and flexible spending accounts
  • Family Forming Benefit including fertility coverage and adoption/surrogacy reimbursement
  • Paid childbearing and paternal leave
  • Education Reimbursement, Student Loan Assistance or 529 College Funding
  • Sabbatical leave
  • Wellness program
  • Flexible work schedule
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service