AI Engineer - US

ScaleOps
•Remote

About The Position

ScaleOps is redefining autonomous cloud and AI infrastructure. We're on a mission to free DevOps and platform engineers from manual resource management so they can focus on innovation, not tuning resources. The results: maximized performance and a reduction of cloud costs by up to 80%. As the category leader in Autonomous Cloud and AI Infrastructure Resource Management, we're trusted by leading enterprises including Adobe, Wiz, Epic Games, Northwestern Mutual, Coinbase, DocuSign, and Fortune 100 companies to autonomously manage their most critical production environments. Backed by Insight Partners, Lightspeed Venture Partners, and other leading VCs with over $210M in funding, ScaleOps is the leading player in a massive and growing market. We are building the autonomous infrastructure management platform that will power the next decade of enterprise compute.

Requirements

  • Significant software engineering experience (typically 4+ years) with strong Python skills and solid backend engineering fundamentals.
  • Experience building and operating production systems in cloud environments.
  • Practical experience bringing LLM-based systems into production, including handling latency, cost control, and failure modes.
  • Familiarity with additional agentic frameworks (e.g., LangChain, MetaGPT) and evaluation frameworks.
  • Strong ownership and the ability to operate independently while collaborating closely across teams, with the motivation to grow into technical leadership as the group expands.

Nice To Haves

  • Experience enabling LLMs to consume structured or operational data (configurations, logs, metrics)
  • Experience with retrieval systems (RAG) or vector databases.

Responsibilities

  • Design and build autonomous AI agents that analyze infrastructure in real time and make intelligent decisions.
  • Work with modern agentic frameworks (LangGraph, PydanticAI) and conversational AI to create multi-agent systems - including troubleshooting, optimization, FinOps, and how-to agents.
  • Leverage core LLM capabilities (tool-use, memory, retrieval) to operate safely in production.
  • Develop MCPs to expose ScaleOps capabilities to AI agents that reason over infrastructure environments, metrics, configurations, and cost signals.
  • Build integrations with tools like Slack, Jira, and AI-powered IDEs (Cursor, Windsurf) to deliver context-aware insights, from "why is this pod not scheduling?" to "how can we reduce costs by 30% safely?"
  • Build and deploy machine learning models that learn from infrastructure patterns - detecting the right resource policies for workloads, predicting optimal scaling triggers, and recommending GPU configurations.
  • Own the complete ML pipeline from training to production, ensuring models are reliable, monitored, and continuously improving.
  • Build and embed internal AI tools to accelerate engineering, development, research, and support.
  • Develop AI-powered tools that help Sales and Support teams demonstrate value instantly - agents that analyze customer infrastructure, generate cost optimization reports automatically, and turn technical data into clear business recommendations.
  • Own AI systems from concept to production, ensuring they're fast (sub-2-second responses), reliable, safe, and cost-effective.
  • Build evaluation frameworks to measure quality, implement security controls, and balance performance tradeoffs in production.
  • Define AI architecture and best practices as a founding member of the AI team.
  • Make key technical decisions - choosing frameworks, designing multi-agent systems, establishing data governance - and shape how ScaleOps evolves from AI-enhanced internal tools to customer-facing AI products.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service