About The Position

First Due is seeking a highly motivated Platform Site Reliability Engineer (SRE) / DevOps Engineer to help build, scale, and maintain the infrastructure, deployment pipelines, and operational systems that power our platform. This individual contributor role is responsible for improving system reliability, performance, security, deployment processes, and scalability while partnering closely with Engineering, Product, and Security teams. The ideal candidate is passionate about automation, operational excellence, cloud infrastructure, software development and deployment processes, observability, and continuous improvement. They thrive in fast-paced environments and enjoy solving complex technical challenges that directly impact customers and business outcomes. This candidate will be an excellent fit if they are not afraid to drive consensus on standards, develop the golden path, and partner with senior engineering leadership to influence or mandate others to adopt the standards.

Requirements

  • 5+ years of experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, Platform Engineering, or a related role.
  • Strong experience managing and supporting production cloud environments.
  • Experience building and maintaining CI/CD pipelines and deployment automation.
  • Strong understanding of infrastructure-as-code principles and tooling.
  • Experience troubleshooting production systems and resolving complex operational issues.
  • Knowledge of networking, security, scalability, and high-availability architectures.
  • Strong scripting or automation experience using languages such as Python, Bash, PowerShell, or similar.
  • Experience working closely with software engineering teams throughout the software development lifecycle.
  • Strong analytical, problem-solving, and communication skills.
  • Ability to balance operational stability with delivery speed and business priorities.

Nice To Haves

  • Experience supporting high-growth SaaS platforms.
  • Experience with AWS, Azure, or Google Cloud Platform.
  • Experience with Kubernetes, Docker, and containerized environments.
  • Experience with Terraform, Pulumi, CloudFormation, or other infrastructure-as-code tools.
  • Experience with GitHub Actions, GitLab CI/CD, Jenkins, CircleCI, or similar automation platforms.
  • Experience with observability platforms such as Datadog, New Relic, Grafana, Prometheus, Splunk, or similar tools.
  • Experience with Database Operations, Database Scacling, Query Optimization, and BCDR
  • Experience with implementing Zero-downtime deployments
  • Familiarity with security frameworks, compliance requirements, and operational governance practices.
  • Experience building highly available, mission-critical applications and services.
  • Experience supporting distributed and remote engineering organizations.

Responsibilities

  • Design, implement, and maintain scalable, secure, and highly available cloud infrastructure.
  • Build and manage infrastructure-as-code solutions to support repeatable and reliable deployments.
  • Continuously improve platform reliability, resiliency, scalability, and performance.
  • Partner with engineering teams to ensure services are designed and operated with reliability and observability in mind.
  • Support disaster recovery, backup, and business continuity initiatives.
  • Design, implement, and maintain CI/CD pipelines to enable efficient and reliable software delivery.
  • Automate manual operational tasks and improve deployment processes across environments.
  • Partner with engineering teams to streamline development workflows and reduce operational overhead.
  • Improve release management practices and deployment strategies.
  • Champion DevOps best practices across the organization.
  • Monitor production systems and proactively identify opportunities to improve availability and performance.
  • Participate in incident response, troubleshooting, root cause analysis, and post-incident reviews.
  • Develop and maintain monitoring, alerting, logging, and observability solutions.
  • Help define and monitor service-level objectives (SLOs), service-level indicators (SLIs), and reliability metrics.
  • Support on-call rotations and assist with critical production incidents when necessary.
  • Partner with Security and Engineering teams to implement and maintain secure infrastructure practices.
  • Support vulnerability remediation, patch management, access controls, and security monitoring.
  • Assist with compliance initiatives and infrastructure controls.
  • Ensure systems adhere to security, privacy, and operational best practices.
  • Work closely with Engineering, Product, QA, and Security teams to support platform initiatives.
  • Document infrastructure, operational processes, and technical standards.
  • Identify opportunities to improve reliability, efficiency, developer experience, and operational excellence.
  • Contribute to technical discussions, architectural reviews, and long-term platform strategy.

Benefits

  • competitive pay
  • medical
  • dental
  • vision coverage
  • FSA/HSA
  • 401(k)
  • flexible PTO
  • a fully remote workplace
  • a technology stipend
  • opportunities for advancement
  • other benefits and perks
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service