Senior Site Reliability Engineer

FordPalo Alto, CA
$150,200 - $283,500Hybrid

About The Position

In this position, we are seeking a highly skilled Senior Software Engineer to join our Messaging & Observability team, responsible for the mission-critical middleware and tooling that every engineering team at Ford relies on. You will design, build, and operate the messaging middleware platform, data management and replication tooling, and observability infrastructure that gives visibility into services running across our Kubernetes (EKS) environments — while also contributing to and supporting application development efforts as needed. You will balance strong data-driven decision-making with the ability to move forward effectively when information is incomplete.

Requirements

  • Bachelor's degree in Computer Science, Electrical/Computer Engineering, or related field (or equivalent experience)
  • 5+ years of progressive experience in cloud-based software development.
  • 5+ years of experience designing, deploying, and supporting cloud-based solutions in production environments.
  • 5+ years of experience supporting mission-critical, always-on applications with high reliability, availability, and performance requirements.

Nice To Haves

  • 5+ years of hands-on expertise with GCP or AWS and cloud-native services including Pub/Sub, MSK, GCS, BigQuery, and container orchestration platforms such as GKE, EKS, or Kubernetes.
  • Strong expertise in messaging technologies: Kafka, Kafka Connect, Pub/Sub, and related data replication/management tooling.
  • Strong expertise in observability technologies: OpenTelemetry, Vector, Prometheus, VictoriaMetrics, Grafana, Dynatrace.
  • Experience with infrastructure as code and deployment automation using Terraform, as well as CI/CD tools such as Cloud Build, ArgoCD, and Tekton.
  • Experience operating and supporting critical, shared middleware infrastructure in 24x7, "always-on" production environments used by multiple downstream teams.
  • Experience developing backend applications/services (beyond middleware/infra) is a plus, given occasional need to support broader app development efforts.

Responsibilities

  • Design, develop, maintain, and support the messaging middleware platform used by all engineering teams, along with associated data management and replication tooling.
  • Build and evolve observability tooling and pipelines (metrics, logs, traces) that provide visibility into services running in Kubernetes/EKS across the organization.
  • Own end-to-end delivery of middleware and observability services, including the platform infrastructure that supports them.
  • Design, develop, and operate high-performance, cloud-based and microservices-driven platforms at scale.
  • Build resilient backend services and tooling using technologies such as Java, Spring Boot, Kafka, PostgreSQL, gRPC, REST, and Kubernetes.
  • Develop and deploy services and tooling on AWS and/or GCP, leveraging managed services (e.g., MSK, Pub/Sub, EKS, GKE).
  • Contribute to application-level development efforts when needed, applying the same engineering rigor used for middleware and platform work.
  • Implement and maintain robust CI/CD pipelines with a strong focus on security, reliability, and efficiency.
  • Apply TDD and DevOps best practices using tools such as Jenkins, ArgoCD, SonarQube, Fossa, and GitHub.
  • Champion automation and operational consistency across environments for messaging, observability, and application infrastructure.
  • Write clean, maintainable, and well-tested code that meets high quality standards.
  • Perform load, stress, and performance testing on messaging middleware, data replication tooling, and application services to ensure scalability and reliability.
  • Own service health for the messaging and observability platforms by proactively identifying performance bottlenecks, throughput limits, and system risks.
  • Implement comprehensive monitoring, alerting, and performance management strategies for the messaging middleware and dependent services.
  • Ensure messaging and observability platforms consistently meet SLA and reliability targets, minimizing impact to all downstream teams.
  • Collaborate with team members and downstream consumers to establish and evolve best practices that reduce operational risk across shared infrastructure.
  • Proactively identify opportunities to adopt emerging messaging, observability, and cloud-native technologies to improve system efficiency and reliability.
  • Lead or contribute to refactoring initiatives to improve middleware and application performance, scalability, and maintainability.
  • Anticipate future challenges through data-informed decision-making and pragmatic engineering judgment.
  • Work within Agile development environments, partnering closely with product managers and cross-functional teams that depend on the messaging and observability platforms, as well as teams building applications on top of them.
  • Translate business and platform requirements into incremental, production-ready solutions.
  • Participate in and lead design discussions, contributing to architectural decisions and technical standards for shared middleware, observability infrastructure, and consuming applications.
  • Influence technical direction while fostering collaboration across teams that consume the platform.
  • Communicate complex technical concepts clearly to both technical and non-technical stakeholders.
  • Drive to sound conclusions even when data is incomplete or ambiguous.
  • Stay ahead of emerging industry trends in messaging middleware, observability tooling, and cloud-native application development.
  • Align technical solutions with the needs of internal engineering teams and business outcomes, ensuring long-term value creation.

Benefits

  • Immediate medical, dental, vision and prescription drug coverage
  • Flexible family care days, paid parental leave, new parent ramp-up programs, subsidized back-up child care and more
  • Family building benefits including adoption and surrogacy expense reimbursement, fertility treatments, and more
  • Vehicle discount program for employees and family members and management leases
  • Tuition assistance
  • Established and active employee resource groups
  • Paid time off for individual and team community service
  • A generous schedule of paid holidays, including the week between Christmas and New Year’s Day
  • Paid time off and the option to purchase additional vacation time.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service