Senior Site Reliability Engineer

AdobeNew York, NY
$139,000 - $257,550Remote

About The Position

The Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses access to hundreds of millions of curated, royalty-free assets that integrate directly with Adobe's desktop, mobile, and web apps, as well as Behance's creative community and Content's advanced services. Our infrastructure spans multiple cloud platforms and is orchestrated by cloud-native, containerized systems as we build the framework for Adobe's next wave of AI-powered products. The team blends startup energy with the resources of a large software company.

Requirements

  • BSc in Computer Science or equivalent experience in a practical setting
  • Python, plus familiarity with PHP, Node.js, or Ruby
  • Familiarity with deploying/operating ML inference pipelines (SageMaker, OpenAI, Bedrock, or equivalent) in production
  • Experience operating cloud-native compute over broad, scalable infrastructures — EC2 Auto Scaling Groups and Kubernetes in AWS; Azure/GCP a plus.
  • Experience with vulnerability/patch management and AMI/golden-image lifecycle automation at fleet scale
  • Strong debugging skills on distributed systems
  • Willingness to participate in on-call rotation

Nice To Haves

  • Production ML operations experience at any scale
  • GPU-backed compute and cost/performance tuning experience
  • Exposure to agentic AI workflows — LangGraph or similar, LLM gateway/routing across providers
  • Relational database operations experience (upgrades, IAM auth, performance tuning)

Responsibilities

  • Design, build, and operate large-scale, distributed, fault-tolerant systems, and the IaC, CI/CD, and automation tooling behind them
  • Own patch, vulnerability, and golden-image lifecycle management across the fleet — triage, remediate, automate
  • Contribute to multi-quarter initiatives: compute rationalization, cloud re-platforming, ML platform migration
  • Build and operate ML inference infrastructure — model serving, GPU workloads, language model gateway and routing — contributing to a broader systems portfolio
  • Help set infrastructure standards and security guardrails for agentic AI, and build the agent tooling we used to run our own operations
  • Partner across software, ML, and platform teams to bake reliability in from the start
  • Share on-call and fix challenges wherever problems show up — web services, databases, pipelines, ML systems

Benefits

  • comprehensive benefits programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service