Staff Platform Engineer

Robots and Pencils
Remote

About The Position

We’re looking for a Staff Platform Engineer to define and lead platform engineering strategy across complex, multi-environment cloud systems. This role is ideal for an experienced engineer who can own infrastructure architecture end-to-end, drive DevSecOps and compliance practices, and serve as a technical leader on the engagements they support. In this role, you will work as a key technical contributor on a cross-functional team, defining standards and owning platform reliability, performance, and security at scale. You’ll mentor engineers, partner with leadership on infrastructure direction, and lead complex migrations and modernization initiatives. At Robots & Pencils, we design AI systems for a human world. Our name says it all. Robots and pencils means engineering paired with creativity, because every agent we ship has to work for real people in real workflows. That balance is baked into how we operate. Every role here contributes directly to that mission. Here, you shape how AI systems integrate into enterprise operations, how teams move at real velocity, and how products create measurable impact for clients and the people they serve. We ship production-ready AI in 30 to 45 days. That pace demands people who take ownership, lead with craft, and care deeply about what they put their name on.

Requirements

  • 7+ years of professional DevOps or platform engineering experience, with experience leading complex platform initiatives
  • Expert scripting and programming skills (e.g., Python, Go, Java, Bash)
  • Deep cloud expertise across at least one major platform
  • Expert Kubernetes and container orchestration skills
  • Expert IaC skills across multiple tools
  • Strong CI/CD architecture experience at scale
  • Strong DevSecOps experience including secrets management, compliance, and auditing
  • Experience with networking, IAM, security architecture, and zero-trust principles in cloud environments
  • Experience with service mesh, distributed systems, and microservices architecture
  • Strong experience with AI/ML platform infrastructure, including model serving and deployment, GPU workload orchestration, LLM gateway and observability, vector store infrastructure, and CI/CD for AI/ML systems
  • Demonstrated leadership and technical mentoring experience across a team or organization
  • Strong stakeholder communication skills, with the ability to translate technical depth across audiences
  • Demonstrable, day-to-day usage and expert knowledge of AI-forward tools such as Claude and Cursor
  • Excellent problem-solving skills and the ability to navigate highly ambiguous technical and business challenges with sound judgment

Nice To Haves

  • Cloud certifications (e.g., AWS DevOps Engineer Professional, CKA, Azure DevOps Engineer) or FinOps experience is a plus
  • Designing and provision HPC cluster infrastructure using CI/CD pipeline across AWS, CoreWeave, GCP, and OCI
  • Experience with HPC job schedulers and workload managers such as Slurm or equivalent for job submission and queue management

Responsibilities

  • Define DevOps strategy and lead infrastructure architecture across multi-environment, multi-region cloud systems
  • Architect and own scalable Kubernetes platforms and containerized infrastructure at scale
  • Own infrastructure as code strategy and standards across environments
  • Lead DevSecOps implementation including secrets management, compliance, auditing, IAM, and zero-trust networking
  • Drive platform reliability, performance SLAs, and cost optimization across production systems
  • Lead complex cloud migrations and platform modernization initiatives
  • Own observability strategy and production reliability practices
  • Lead the design and operation of AI/ML platform infrastructure, including model serving and deployment, GPU workload orchestration, LLM gateway and observability, vector store infrastructure, and CI/CD for AI/ML systems
  • Bring an AI-forward mindset to your daily work, using tools like Claude, Cursor, and other modern AI assistants to ship higher-quality work at pace
  • Partner with engineering, product, and leadership to align platform strategy with business and delivery goals
  • Communicate complex infrastructure decisions and tradeoffs clearly to technical and non-technical stakeholders
  • Lead design reviews, architecture discussions, and release readiness assessments
  • Establish platform engineering standards and best practices on the engagements you support
  • Mentor junior and mid-level engineers, helping them grow their craft, confidence, and impact
  • Act as a technical escalation point on complex infrastructure and platform challenges
  • Evaluate emerging tools and technologies, recommending patterns that improve platform reliability and developer experience
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service