About The Position

We have an exciting opportunity to support our Technology team as a Senior Engineer, Observability, based in Shelton, CT. This is a senior individual-contributor role. The Senior Engineer, Observability is responsible for implementing and managing observability across Subway's infrastructure and applications — servers, networks, cloud services, databases, restaurant point-of-sale and kiosk estates, mobile apps, and web platforms. The role owns configuration, maintenance, and documentation of our observability platforms, and partners with technology stakeholders through instrumentation planning, service reviews, and operational enablement so that monitoring keeps pace with the business. Subway's observability estate is large and actively changing: business-event telemetry now runs in production across middleware, order, payment, and restaurant-device flows; a network and security logging domain has recently been onboarded; and incident routing is moving onto PagerDuty. Alongside building coverage, this role carries real responsibility for signal quality and for the cost of the telemetry we collect and query. If you feel that this role is for you, and you are successful with your application, be ready to be Bold, Empowered, Accountable, and ready to have Fun in a fast-paced and agile working environment.

Requirements

  • Bachelor's degree in Computer Science or a related field.
  • 6+ years of related technical experience.
  • Expert in Dynatrace, with demonstrable skills across distributed tracing, log management, business events, metrics, DQL, OpenPipeline, dashboards and notebooks, alerting, anomaly detection, and automation workflows.
  • Working familiarity with additional observability and alerting tooling such as SolarWinds, PagerDuty, Fluent Bit / Calyptia, Azure Monitor, AWS CloudWatch, and Splunk.
  • Practical understanding of observability for distributed retail and digital estates — point-of-sale and in-restaurant devices, third-party delivery and ordering integrations, payment flows, and mobile and web front ends.
  • Track record of designing alerts that catch real failures with a low false-positive rate, and of retiring alerts that do not earn their keep.
  • Comfortable reasoning about missing data, low-volume and seasonal signals, and the difference between an outage and a degradation.
  • Experience managing observability consumption against a commercial commitment — ingest, retention, and query cost — and making evidence-based decisions about what telemetry is worth collecting.
  • Familiarity with tagging and attribution models that let cost be allocated to owning teams.
  • Hands-on experience with AWS and/or Azure services to support instrumentation and management of observability for cloud-hosted applications and services, including serverless and container workloads.
  • Strong interpersonal and communication skills, written and verbal, with the ability to work effectively across technology teams, development groups, architecture stakeholders, and non-technical business partners.
  • Able to explain what telemetry does and does not show, and to hold a clear position in a post-incident review.
  • Proficient in Windows and Linux environments; scripting in PowerShell and Bash, and comfort with at least one general-purpose language for automation.
  • Skilled in administering observability platforms and their integrations, including ServiceNow for incident creation, PagerDuty for routing and on-call, and cloud-native monitoring services.
  • Working knowledge of SNMP, WMI, ICMP, NetFlow, syslog, and REST APIs for data collection and integration.
  • Version control and configuration-as-code practice (Git-based workflows); familiarity with Azure DevOps for backlog and delivery tracking.
  • Naturally curious and committed to continuous learning and improvement.
  • Proactive in identifying opportunities for optimization and innovation in monitoring strategy.
  • Rigorous about evidence — willing to verify a number before acting on it, and to correct the record when it turns out to be wrong.
  • Comfortable working with ambiguity and competing priorities across concurrent platform changes.

Responsibilities

  • Design, configure, and maintain observability and alerting platforms — Dynatrace (primary), SolarWinds, PagerDuty, Fluent Bit / Calyptia log pipelines, AWS CloudWatch, Azure Monitor, and Splunk. This covers instrumentation, dashboards and notebooks, alerting and anomaly detection, business-event processing and creation, log pipeline configuration and analysis, and synthetic and real-user monitoring.
  • Practice detection engineering, not just alert configuration — select the right detection pattern for each failure mode, including silence and dead-man's-switch coverage for feeds that can stop, expected-versus-actual and seasonal-baseline comparison, and success-rate and ratio controls for partial degradation.
  • Backtest new alerts against historical incidents before enabling them and prefer one dimensioned alert over many per-geography copies.
  • Improve signal quality and reduce noise — review and tune alert configurations, thresholds, dwell times, and routing in collaboration with stakeholders; identify and retire duplicate, unowned, and non-actionable alerts; and make sure every production alert has an owner and a documented response.
  • Own telemetry cost and consumption governance — understand how each observability capability is billed, attribute consumption to owning teams through cost and product tagging, monitor ingest and query volumes against commitments, and identify optimization opportunities in retention, sampling, synthetic frequency, and query efficiency.
  • Build and maintain automation — author Dynatruck workflows and automation apps and manage observability configuration as code in source control with a promotion path from lower environments to production.
  • Collaborate with developers, architects, and platform teams to define observability strategies aligned to business and operational goals. This includes integrating observability into CI/CD and infrastructure-as-code workflows and supporting business observability through end-to-end transaction tracing and event instrumentation.
  • Provide technical leadership and enablement — guide the work of internal engineers and external partners engaged on observability, set technical direction and quality expectations, and coach teams through knowledge transfer, documentation, and training. This role leads through technical influence rather than direct people management.
  • Support incident and problem management — perform proactive monitoring and incident response for production systems; participate in major-incident reviews and problem records; correlate anomalies against planned change windows before treating them as faults; and translate post-incident findings into concrete detection improvements.
  • Manage platform vendor relationships day to day — raise and drive vendor support cases to resolution, evaluate release notes and advisories for impact, and coordinate agent and platform upgrades with change management.
  • Develop, document, and maintain observability processes and response procedures, ensuring operational readiness and participating in service reviews so that coverage continues to meet stakeholder needs.

Benefits

  • Insurance Plans (Medical/Life)
  • 401K
  • Competitive Bonus
  • Mobility Allowance
  • Tuition Reimbursement
  • Company Holidays
  • Volunteering time
  • And Many More…
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service