Senior Cloud Platform Engineer (SMTS)

Salesforce•San Francisco, CA
•$148,500 - $246,000•Hybrid

About The Position

As a Senior Member of Technical Staff (SMTS) within our Monitoring Cloud team, you will be a key owner and operator of the systems that keep Salesforce reliable. You won't just be "using" tools; you will be productizing infrastructure to ensure our monitoring capabilities evolve at the scale of our multi-cloud footprint. Your mission is to bridge the gap between high-level feature design and deep-system stability. From automating the "paved path" across AWS and GCP to securing air-gap environments for our most sensitive customers, you will ensure our monitoring stack is invisible, resilient, and intelligent. This is an AI-first engineering role. You will use AI-assisted development tools (e.g., Claude Code) as the default for every inner-loop activity, code authoring, Terraform and Kubernetes scaffolding, test generation, refactoring, log/trace analysis, runbook drafting, and documentation. We expect AI to compound your throughput on routine implementation so you can focus your human judgment on architecture, security, on-call response, and customer outcomes.

Requirements

  • 5+ years Proven track record in Distributed systems, API platforms, Infrastructure Engineering, Observability or DevOps at scale.
  • Proficiency with Kubernetes (K8s) and Terraform.
  • Hands-on experience managing infrastructure in AWS and/or GCP.
  • Proficiency in programming languages(eg: java, python etc)
  • Experience managing or extending monitoring tools (e.g., Grafana), messaging systems (kafka etc), elastic search, caching frameworks
  • Security First: Understanding of authN/authZ security protocols, particularly in managing isolated or restricted network environments.
  • AI-assisted development fluency: demonstrated use of AI coding assistants (e.g., Claude Code) as part of a daily engineering workflow, able to prompt effectively, critically evaluate generated code, and integrate AI into IaC, testing, and automation pipelines.

Responsibilities

  • Design and implement automation frameworks using Terraform and Kubernetes to manage monitoring infrastructure.
  • Standardize "paved path" deployments across AWS and GCP, eliminating manual configuration errors and ensuring global consistency.
  • Use AI-assisted tooling as the default for authoring, refactoring, and reviewing IaC modules, Helm charts, and automation scripts while directing intent, validating output, and owning the final result.
  • Own the lifecycle of the Monitoring Cloud stack, including version upgrades and performance tuning.
  • Productize core components (e.g., Grafana, custom Terraform providers) to make them consumable as reliable services by internal engineering teams.
  • Leverage AI for upgrade planning, release-note analysis, migration scaffolding, and boilerplate-heavy productization work (API wiring, schema plumbing, SDK generation), while retaining accountability for design and rollout.
  • Deploy and manage the full monitoring stack within highly isolated, air-gapped environments.
  • Ensure that our most secure customer segments receive the same level of observability and reliability as our public cloud offerings.
  • Apply AI assistance during development of the artifacts that ship into these environments; operate them in-network with the disciplined, human-driven workflows these environments require.
  • Participate in the team’s on-call rotation, providing the deep technical expertise required to maintain strict SLAs and availability targets.
  • Conduct root-cause analysis (RCA) for complex system failures and implement long-term preventative fixes.
  • Address support requests with a “customer first” mindset
  • Use AI as a co-pilot during incident response and RCA: summarizing logs, correlating traces, proposing hypotheses, and drafting status updates and postmortem while the engineer remains the accountable responder and decision-maker.
  • Design and deliver platform features that adhere to enterprise standards while pioneering AI-driven development practices to accelerate delivery and enhance system intelligence.
  • Contribute to and evolve the team's AI-assisted development playbook: prompts, agents, skills, evaluation harnesses, and guardrails that let the team ship faster without sacrificing quality or security.

Benefits

  • time off programs
  • medical
  • dental
  • vision
  • mental health support
  • paid parental leave
  • life and disability insurance
  • 401(k)
  • employee stock purchasing program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service