About The Position

Join phData, a remote-first data and AI consultancy company with employees across the United States, Latin America, and India. We partner with industry leaders, including Snowflake, AWS, Anthropic, Glean, and dbt, to solve the complex data and AI challenges that slow large enterprises. We're growing fast, and we give our people real ownership over their work. We hire top performers and trust them to deliver results. We are looking for a Senior DevOps Engineer (Cloud) to join our Elastic Operations (Managed Services) team. In this role, you will lead the technical delivery and operational reliability of modern cloud data and AI platforms, including Snowflake, Anthropic, and solutions on AWS and Azure, for our managed services clients. You won't just support these platforms, you'll use AI tools every day to troubleshoot faster, automate more, and keep documentation current. You will collaborate closely with clients, Elastic Operations leadership, data engineering and analytics teams, and other cross-functional stakeholders to deliver high-quality solutions and advance phData’s delivery excellence.

Requirements

  • 6+ years of experience in DevOps, cloud data platform operations, and/or production support for data and analytics workloads.
  • Working knowledge of SQL with the ability to write, debug, and optimize queries.
  • Experience providing operational support for a cloud-native data warehouse (Snowflake and/or Amazon Redshift) across a large user base.
  • Hands-on experience with Relational Database Management Systems such as Oracle or Microsoft SQL Server.
  • Experience working in a production support environment, including monitoring and supporting scheduled data jobs and pipelines (ETL/ELT).
  • Working knowledge of Unix/Linux environments and common system administration concepts.
  • Basic proficiency in writing and optimizing Python programs for scripting and automation.
  • Experience with cloud-native data technologies and services on AWS and/or Azure (e.g., S3, ADLS, Azure Data Factory).
  • Familiarity with ITIL processes and working in SLA-driven support environments.
  • Strong troubleshooting, performance tuning, and incident/problem management skills.
  • Strong client-facing written and verbal communication skills, with the ability to clearly explain issues and solutions.
  • Openness to learning new technology stacks and helping train and up-skill other team members.
  • Active daily use of AI coding tools (Cursor, GitHub Copilot, Claude, ChatGPT, or equivalent) with demonstrated judgment about when to trust, verify, and correct AI-generated output.
  • Ability to describe specific engineering tasks completed with AI assistance.
  • Familiarity with CI/CD tooling (e.g., GitHub, Bitbucket) and exposure to infrastructure-as-code tools such as Terraform.
  • Experience delivering projects for external or internal clients in a professional services or consulting environment.
  • Ability to break down complex problems into structured, actionable steps and drive them through to completion.
  • Strong written and verbal communication skills in English.
  • Demonstrated ability to work effectively with distributed and cross-functional teams.
  • Proven track record of taking ownership, managing multiple priorities, and delivering high-quality work with minimal supervision.

Nice To Haves

  • Production support experience and/or certifications with core data platforms such as Snowflake, AWS, Azure, or Databricks.
  • Experience with cloud and distributed data storage technologies such as Amazon S3, Azure Data Lake Storage (ADLS), or similar cloud storage services.
  • Exposure to supporting AI platforms and workloads (e.g., Anthropic Claude, LLM-powered applications) in production environments.
  • Experience with data integration and streaming technologies such as Spark, Kafka, Matillion, Fivetran, dbt, AWS Database Migration Service, or Azure Data Factory.
  • Experience with workflow management and orchestration tools such as Apache Airflow, AWS Managed Airflow.
  • Strong expertise in scripting (preferably Python) to automate repetitive operational tasks.
  • Hands-on experience with automated database deployment frameworks such as Flyway or Liquibase.
  • Prior experience working in a managed services or Elastic Operations environment supporting multiple clients.

Responsibilities

  • Own and drive end-to-end operational support and incident management for modern cloud data platforms (e.g., Snowflake, AI workloads on Anthropic, data lakes, analytics platforms) across multiple client environments.
  • Translate business and data requirements into resilient, cost-effective operational solutions that align with phData methodologies, standards, and best practices.
  • Monitor and support production data jobs and pipelines (ETL/ELT), ensuring timely resolution of failures and minimizing business impact.
  • Respond to pager incidents, perform deep root-cause analysis, and implement preventative measures across customer processes and workflows.
  • Ensure engagements are delivered on time, within scope, and with measurable business value for clients.
  • Collaborate with cross-functional partners such as data engineering, analytics, solutions architecture, and client stakeholders to deliver successful client engagements.
  • Provide technical leadership during incident bridges, troubleshooting sessions, and design or runbook reviews, especially across Snowflake and cloud platform services.
  • Ensure high quality in deliverables through thorough documentation, knowledge articles, runbooks, and adherence to governance and change management processes.
  • Partner with practice and account leaders to identify opportunities to improve operational maturity, standardize patterns, and enhance client satisfaction across a large user base.
  • Contribute to internal initiatives such as building and enhancing operational playbooks, automation scripts, monitoring standards, and training materials for Elastic Operations.
  • Mentor peers by sharing best practices, leading knowledge-sharing sessions, and helping up-skill team members on new technologies and tools.
  • Represent phData with professionalism in all interactions, communicating clearly with both technical and non-technical stakeholders.
  • Use AI tools daily to accelerate troubleshooting, automation, documentation, and knowledge capture, validating output carefully before it reaches production systems.
  • Act as a trusted advisor to senior client stakeholders on operational strategy, reliability, and performance optimization of their cloud data platforms.
  • Lead complex incident and problem-management efforts, coordinating across multiple teams and driving long-term remediation and platform improvements.
  • Mentor and coach junior DevOps engineers, fostering a culture of learning, feedback, and continuous improvement.
  • Help define and refine Elastic Operations standards, reusable assets, and delivery frameworks for managed services.

Benefits

  • 401k plan with company match
  • Dental and Vision insurance
  • Home Office Equipment Stipend
  • Annual stipend for Learning and Development
  • Competitive comp, excellent benefits, 4 weeks PTO plan plus 10 Holidays (and other cool perks)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service