Director, Enterprise Observability and Automation

6090-Johnson & Johnson Services Legal Entity•Raritan, NJ
•Hybrid

About The Position

At Johnson & Johnson, we believe health is everything. Our strength in healthcare innovation empowers us to build a world where complex diseases are prevented, treated, and cured, where treatments are smarter and less invasive, and solutions are personal. Through our expertise in Innovative Medicine and MedTech, we are uniquely positioned to innovate across the full spectrum of healthcare solutions today to deliver the breakthroughs of tomorrow, and profoundly impact health for humanity. Learn more at jnj.com As guided by Our Credo, Johnson & Johnson is responsible to our employees who work with us throughout the world. We provide an inclusive work environment where each person is considered as an individual. At Johnson & Johnson, we respect the diversity and dignity of our employees and recognize their merit.

Requirements

  • Bachelor's degree in computer science, engineering, information systems, or a related field; an advanced degree is preferred.
  • 10+ years of progressive experience in software engineering, platform engineering, site reliability engineering, enterprise operations, or a related technology discipline.
  • 4+ years of experience leading teams and/or people leaders in a global, matrixed environment.
  • Demonstrated experience defining and implementing enterprise-scale observability strategies across business-critical applications and services.
  • Proven experience implementing operational automation, orchestration, AIOps, or AI-driven solutions.
  • Experience improving reliability, reducing mean time to detect and restore, simplifying tool landscapes, and optimizing technology spend.
  • Experience working in a highly regulated environment and translating security, privacy, quality, and compliance expectations into practical engineering controls.
  • Expertise in OpenTelemetry, distributed tracing, metrics, logs, events, digital experience monitoring, and large-scale telemetry pipelines.
  • Strong knowledge of cloud-native architecture, public cloud platforms, Kubernetes, APIs, microservices, and modern software delivery practices.
  • Experience operating AI/ML or large-language-model solutions in production, including evaluation, monitoring, MLOps/LLMOps, guardrails, and model-risk controls.
  • Experience with observability and monitoring platforms such as Splunk, AppDynamics, Grafana, Telegraph, Clickhouse, cloud-native monitoring, or comparable technologies.
  • Strong understanding of Incident, Problem, Change, and Knowledge Management processes and their integration with enterprise observability and automation.
  • Exceptional communication, documentation, and stakeholder-management skills, with the ability to translate technical complexity into risk, value, investment, and business decisions for senior leaders.
  • Ability to lead through critical incidents, service disruption, ambiguity, and competing enterprise priorities.
  • Strong analytical, problem-solving, and continuous-improvement mindset.
  • Demonstrated commitment to inclusive leadership, talent development, collaboration, and Our Credo values.

Nice To Haves

  • Experience with LLM observability, evaluation frameworks, agentic-AI runtime controls, retrieval-augmented generation, and AI guardrails.
  • Familiarity with responsible-AI frameworks and evolving regulations and standards, including NIST AI RMF and the EU AI Act.
  • Experience managing large technology portfolios, enterprise observability spend, and strategic suppliers.
  • Experience in healthcare, life sciences, medical technology, or another quality- and compliance-intensive industry.
  • Relevant certifications in ITIL, Splunk, ServiceNow, cloud platforms, AI, machine learning, data analytics, or automation.

Responsibilities

  • Define and execute a multi-year enterprise observability and automation strategy aligned with J&J Technology priorities, business outcomes, cybersecurity expectations, and the needs of Innovative Medicine, MedTech, and enterprise functions.
  • Define a vendor-neutral target architecture spanning telemetry collectors, agents, gateways, routing, processing, and backend platforms.
  • Establish standards for data models, tagging and metadata, retention tiers, data residency, personally identifiable information handling, and cross-signal correlation.
  • Establish lifecycle monitoring for AI, machine-learning, and agentic solutions, including performance, drift, latency, cost, quality, explainability, bias, safety signals, and human oversight.
  • Advance anomaly detection, intelligent alerting, event correlation, automated root-cause analysis, predictive operations, and remediation to reduce operational noise and improve resilience.
  • Partner with product, platform, and business technology leaders to establish service-level objectives, service-level indicators, error budgets, and experience measures for critical products and services.
  • Own the enterprise observability architecture, standards, roadmap, investment portfolio, and strategic vendor relationships across metrics, logs, traces, events, and digital experience telemetry.
  • Collaborate with Information Security & Risk Management, privacy, quality, regulatory, legal, data, and responsible-AI partners to embed practical controls that support security, compliance, fairness, responsibility, and transparency.
  • Build, lead, and develop high-performing teams spanning observability engineering, site reliability engineering, AI/ML platform engineering, and AIOps while fostering inclusion, accountability, and talent growth.
  • Strengthen incident, problem, change, and knowledge-management practices through effective on-call operations, post-incident reviews, actionable problem management, and continuous learning.
  • Provide leadership with clear insight into service health, AI performance, reliability risk, adoption, value realization, and investment priorities.
  • Drive simplification, interoperability, reuse, cost optimization, automation adoption, and measurable improvements in operational maturity.

Benefits

  • medical
  • dental
  • vision
  • life insurance
  • short- and long-term disability
  • business accident insurance
  • group legal insurance
  • consolidated retirement plan (pension)
  • savings plan (401(k))
  • long-term incentive program
  • Vacation –120 hours per calendar year
  • Sick time - 40 hours per calendar year
  • Holiday pay, including Floating Holidays –13 days per calendar year
  • Work, Personal and Family Time - up to 40 hours per calendar year
  • Parental Leave – 480 hours within one year of the birth/adoption/foster care of a child
  • Condolence Leave – 30 days for an immediate family member: 5 days for an extended family member
  • Caregiver Leave – 10 days
  • Volunteer Leave – 4 days
  • Military Spouse Time-Off – 80 hours
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service