DreamWorks Technology - Senior AI Software Engineer

NBCUniversal•Glendale, CA
•$155,000 - $190,000•Remote

About The Position

DreamWorks Animation's Studio Platform Engineering team builds the systems and infrastructure that enable our artists and technology teams to operate at scale. We own the Feature Animation platform, streamline software delivery through modern CI/CD practices, and lead the Studio’s observability ecosystem. You’ll play a key role in designing and evolving the next generation of resilient, cloud-native, AI-enabled platforms that drive innovation across the Studio. We are seeking a Senior AI Software Engineer with a strong background in DevOps, Kubernetes infrastructure, and Python development to build, scale, and secure our AI infrastructure used for data workflows and management. In this role, you will bridge the gap between traditional infrastructure operations and advanced AI engineering. Instead of just consuming AI models, you will engineer the enterprise-grade ecosystem that hosts, monitors, scales, and protects our next generation of resilient, cloud-native, AI-enabled platforms that drive innovation across the Studio in an 'Artists First' methodology.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field
  • 8+ years of experience in Software, DevOps or Platform Engineering, with 1+ years building AI systems for real workflows.
  • Hands-on experience in cloud-native technologies and architectures (Docker, Kubernetes).
  • Proven experience developing and deploying AI agents, tool-calling systems, and multi-agent systems
  • Proficiency in revision control and DevOps best practices (Git)
  • Expert Linux experience (Red Hat, CentOS) and proficient with one or more Unix shell scripting languages (Bash, C-Shell)
  • Proficient in either Python, Java, or JavaScript development
  • Operational support experience, including platform infrastructure monitoring and troubleshooting using tools like Datadog, Prometheus, or similar

Nice To Haves

  • Strong, hands-on experience managing and troubleshooting production workloads running on Kubernetes
  • Hands-on experience developing or deploying custom application servers, Retrieval-Augmented Generation (RAG) models and Model Context Protocols (MCP) servers
  • Deep understanding of LLM model integration and tuning
  • Experience monitoring and mitigating the specific operational costs associated with token usage and GPU compute utilization
  • Deep expertise in at least one major cloud ecosystem (AWS, Azure, OCI or GCP) and infrastructure provisioning
  • Experience with designing automation that reduces manual work in complex, multi-stakeholder environments
  • Experience with AMD Inference Microservice Platform (AIMS) or NVIDIA NIM a plus
  • Excellent communication skills and the ability to translate fluently between artists, production management, and engineers
  • Genuine passion for animation, storytelling, or creative productions

Responsibilities

  • Collaborate across technology and production teams including principal engineers to shape the best practices of Platform as a Service implementation hosting stateless applications and AI infrastructure
  • Influence and contribute to the Agentic Engineering and MCP Server Development efforts
  • Major contributor to the development and adoption of governance, security, and observability best practices for AI and agentic systems, helping establish standards for secure, reliable, responsible, and observable AI agent development and operations
  • Leverage Kubernetes orchestration and operators to automate stateless application deployments and administration across multiple environments (on-premises and in the cloud)
  • Leverage Enterprise grade Kubernetes software stack solution for developing, deploying, and running AI workloads on a Kubernetes platform across development and production environments
  • Identify areas for process and efficiency improvement within Platform Services; recommend solutions and assist in overseeing implementation. Actively facilitate continuous improvement.
  • Monitor and maintain platform infrastructure, utilizing tools like Datadog for performance tracking, alerts, and capacity management
  • Participate in continuous organizational cyber security vulnerability mitigation and compliance efforts
  • Define and document standard run books and operating procedures. Create and maintain system information and architecture diagrams
  • Develop and maintain microservices, tools and automation
  • Provide software troubleshooting and break-fix support for production

Benefits

  • medical, dental and vision insurance
  • 401(k)
  • paid leave
  • tuition reimbursement
  • a variety of other discounts and perks
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service