About The Position

Salesforce is seeking a Senior/Lead Software Engineer for Site Reliability within the Agentforce Operations team. This role involves shaping infrastructure, process automation, and engineering practices for mission-critical public-sector workflows and the Agentforce Operations Super Agent team. The focus is on optimizing scalability, resilience, and performance across the entire stack to support significant growth. Responsibilities include diagnosing and prioritizing architectural problems, designing and implementing solutions for complex system challenges, and raising the engineering quality bar. The role also involves mentoring engineers and collaborating across teams to understand customer needs and make pragmatic tradeoffs. The company is leveraging AI to automate operational tasks and accelerate infrastructure delivery.

Requirements

  • Proven Production Experience: Experience supporting mission-critical production systems and helping high-growth products navigate the technical debt and architectural changes required to scale reliably.
  • Technical Breadth: Strong proficiency in Kubernetes, Terraform/OpenTofu, and AWS/GCP/Azure.
  • Coding Mastery: Ability to write and review production-level code in Golang, TypeScript, or Python—you view automation as a software engineering problem.
  • Systems Expert: Deep understanding of distributed systems, including how to debug complex interactions between microservices, databases, and AI agents.
  • Low-Ego Collaboration: Experience working within a senior team of Principal engineers, capable of both leading specific initiatives and supporting the broader group’s technical vision.
  • Mentorship Mindset: Enthusiasm for learning and growing as an engineer, and for helping your peers do the same.
  • AI Fluency: A demonstrated ability to use modern AI development tools to move faster and build more reliable systems.
  • Drive: Ability to work independently and collaboratively in a fast-paced startup environment.
  • Excellent Communicator: Strong written and verbal communication skills, with the ability to communicate effectively with people from varied technical backgrounds.
  • 5+ years of experience in SRE, Production Engineering, or Backend Engineering with a heavy focus on operations and infrastructure.
  • Mastery of Golang, GraphQL, and PSQL, with demonstrated experience delivering high-performance optimizations.
  • Experience with RDS, Redis/Elasticache, and EKS.
  • Deep expertise in multiple areas of backend engineering such as data modeling, db performance tuning, state management, API design, transactionality, concurrency, memory management, fault tolerance, and scaling, with the ability to pick up new technologies quickly.
  • Repeated ownership of foundational system work in complex distributed systems, from diagnosing architectural problems to designing solutions to driving adoption across a team.
  • Excellent writing and speaking skills and ability to lead real-time technical-design discussions.
  • Demonstrated ownership and accountability, with a consistent track record of high standards, continuous improvement, and low-ego collaboration.

Nice To Haves

  • Advanced Degree in Computer Science or equivalent practical experience.
  • Compliance and Regulated Industries: Experience building products for regulated industries, particularly the public sector or environments with strong security, compliance, and data-sovereignty requirements.
  • Familiarity with developing for classified or limited-connectivity environments, including Department of Defense Impact Levels such as IL6.
  • Advanced knowledge of microservice orchestration and durability patterns, including hands-on experience with Temporal for workflow reliability and service mesh for secure, observable service-to-service communication in high-growth SaaS environments.
  • Deep knowledge of networking, security, and identity management within major cloud providers.
  • Experience building AI products for supply chain, logistics, manufacturing, or operational workflows.
  • B.S. in Computer Science (M.S. preferred)
  • Exposure to Temporal, Istio, or Typesense
  • Experience with CI/CD systems, specifically Jenkins and Spinnaker.
  • Exposure to the supply chain, logistics, or manufacturing industry.
  • Familiarity with the Salesforce platform
  • Experience with workflow engines

Responsibilities

  • Own the reliability roadmap for major product areas, evolving startup-speed architectures into highly available, globally scalable systems for mission-critical customer-operated cloud environments.
  • Partner with engineers to refine our infrastructure strategy, contributing senior-level perspectives on system design, capacity planning, bottleneck identification, and security.
  • Define, maintain and evolve our tools for deployments, focusing on making deployments as “push-button” and robust as possible.
  • Support the scaling and deployment of our AI/ML infrastructure, ensuring the compute, deployment and platform capabilities required to run AI features reliably and cost-effectively.
  • Design and build the internal systems and processes for testing and releasing the product, along with the tooling and automation that enable customers and partners to deploy, upgrade, diagnose, and run Missionforce Operations successfully with minimal direct engineering involvement.
  • Design, review, and strengthen identity and access controls, networking, workload isolation, secret management, deployment security, and policy enforcement to ensure the product is secure by default.
  • Lean into the future of engineering by using AI tools to automate routine operational tasks and accelerate infrastructure delivery.
  • Diagnose and prioritize architectural problems across our distributed backend systems, contributing technical perspective to help the team invest in the right foundational work at the right time.
  • Design and implement solutions to hard system problems including application performance optimization, infrastructure management, and data model improvements.
  • Raise and maintain our engineering quality bar through pragmatic standards, guardrails, and team practices that prevent regressions and keep our product healthy and performant.
  • Collaborate with peers across customer, product, and engineering to deeply understand customer needs, make pragmatic tradeoffs that maximize impact, and align on engineering reality.
  • Mentor engineers as a senior peer and force multiplier, sharing expertise across backend architecture and system design while helping shape stronger engineering judgment, more effective collaboration, and a higher-performing team.

Benefits

  • time off programs
  • medical
  • dental
  • vision
  • mental health support
  • paid parental leave
  • life and disability insurance
  • 401(k)
  • employee stock purchasing program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service