Member of Technical Staff, Infrastructure Engineer

Edison ScientificSan Francisco, CA
$175,000 - $240,000Onsite

About The Position

Edison Scientific is seeking a Member of Technical Staff, Infrastructure Engineer to design, scale, and operate the core platform infrastructure for AI scientist agents. This role is crucial for building the foundation of autonomous scientific discovery, requiring expertise in resilient scheduling, lifecycle management, and resource orchestration beyond typical cloud-native workloads. The engineer will establish infrastructure best practices, collaborate with backend, ML, and research teams, and deliver a production-grade environment to accelerate scientific progress. This senior-level engineering position emphasizes technical ownership, understanding complex systems, making architectural trade-offs, and building scalable foundations.

Requirements

  • Typically, 4+ years of professional infrastructure or platform engineering experience, with hands-on Kubernetes experience in production environments.
  • Experience with cloud infrastructure (AWS EKS, GCP GKE, or Azure AKS) and associated networking, storage, and IAM primitives.
  • Proficiency in at least one systems or backend language for operator development and infrastructure tooling.
  • Hands-on experience with infrastructure-as-code tools (Terraform, Pulumi, or Crossplane) and GitOps workflows.
  • Strong working knowledge of container networking (CNI plugins, service mesh, network policies), storage (CSI, persistent volumes, StatefulSets), and security (RBAC, Pod Security Standards, secrets management).
  • Ability to operate autonomously, make sound technical judgments, and drive projects from concept through production.

Nice To Haves

  • Security experience with gvisor/kata containers, snapshots

Responsibilities

  • Architect, implement, and operate Kubernetes clusters that support thousands of concurrent, persistent resources (agents, jobs, services) with high availability and efficient resource utilization.
  • Drive the strategy for cluster scaling, node pool management, autoscaling policies, and resource quota frameworks to handle rapid workload growth.
  • Design and implement robust scheduling, placement, and affinity strategies to optimize cost, performance, and fault tolerance for heterogeneous workloads (CPU, GPU, memory-intensive).
  • Establish and uphold best practices around observability, monitoring, alerting, and incident response for infrastructure systems (Prometheus, Grafana, Datadog, or similar).
  • Own storage and networking strategy within Kubernetes — including persistent volume management, CSI drivers, service mesh, network policies, and ingress architecture.
  • Troubleshoot complex, cross-system infrastructure issues and guide others through effective debugging and remediation in distributed environments.
  • Collaborate closely with backend, ML, and research teams to understand workload requirements and translate them into reliable infrastructure patterns.

Benefits

  • Competitive salary and equity
  • Full healthcare coverage; we pay 100% of premiums for you and your dependents
  • Support for growing families, including a yearly new parent stipend and fertility coverage through Carrot
  • Mental health support through Rula, our in-network therapist and psychiatrist network with fast availability
  • 12 weeks of paid parental leave for maternity, paternity, and adoption
  • Pet care support with a yearly employer-funded stipend for your animal companions
  • Commuter benefits so you can pay for transit and parking with pre-tax dollars
  • 401(k) company matching
  • $300 health and wellness benefit quarterly
  • Lunch is on us every day you're in the office, and dinner is on us when you're working late
  • Regular team off-sites and company events
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service