Data Engineer II (Data Operations & Infrastructure)

ICW GroupSan Diego, CA
$95,379 - $160,850Hybrid

About The Position

The Data Engineer II – Data Operations & Infrastructure owns the operational health of ICW Group's data platform stack. This role receives production workloads from development teams and is responsible for their ongoing run and maintain lifecycle—monitoring, incident response, platform administration, governance, and small operational enhancements. The role ensures uptime, reliability, data quality, and cost efficiency across the data ecosystem.

Requirements

  • Bachelor's degree in Computer Science, Applied Mathematics, Engineering, or any other technology-related field required, or equivalent combination of education and experience.
  • Minimum 3 years of experience in Data Engineering, data platform support or production support of data systems required.
  • Experience administering a data lakehouse or data warehouse platform with hands-on RBAC, user management, and platform governance responsibilities required.
  • Must have a solid foundation in data engineering fundamentals including understanding of ETL/ELT patterns, data pipeline architecture, data modeling concepts, and SQL proficiency sufficient to troubleshoot, validate, and maintain production data systems.
  • Strong experience administering a data lakehouse or data warehouse platform (e.g., Snowflake, Databricks, Amazon Redshift, Google BigQuery, Azure Synapse, or similar). Must have hands-on knowledge of role-based access control (RBAC), user/role management, resource/compute configuration, performance monitoring, cost governance, and security policies.
  • Solid understanding of data lakehouse/warehouse architecture—storage layers, compute separation, medallion/multi-layer patterns, partitioning, clustering, and optimization strategies.
  • Hands-on experience with dbt (dbt Cloud or dbt Core) including job monitoring, environment configuration, and execution troubleshooting.
  • Experience with Apache Airflow (or similar orchestration tools) including DAG monitoring, failure triage, infrastructure configuration, and troubleshooting.
  • Experience with Fivetran (or similar managed ingestion tools) for connector monitoring and administration.
  • Proficiency with AWS CloudWatch for monitoring, logging, alarming, and dashboarding of data infrastructure components.
  • Hands-on experience with AWS Lambda in the context of data infrastructure (e.g., event-driven triggers, lightweight operational functions, alerting integrations).
  • Experience with Data Observability tools for infrastructure monitoring, log management, and alerting.
  • Ability to build and maintain observability frameworks for data platforms—including SLA monitoring, pipeline health dashboards, and incident alerting.
  • Understanding of incident management practices—root cause analysis, runbook development, post-mortems, and escalation processes.
  • Working knowledge of database & data warehouse concepts, including MS-SQL, PostgreSQL, and cloud-native warehouses.
  • Understanding of ETL/ELT patterns, data modeling concepts, and data pipeline architecture sufficient to troubleshoot and maintain existing pipelines.
  • Working knowledge of SQL for troubleshooting, data validation, and small enhancement work.
  • Familiarity with Python or similar scripting languages for operational automation and small pipeline modifications.
  • Experience with AWS technologies like S3, EC2, Glue, Lambda, CloudWatch, and related services.
  • Strong operational mindset—proactive in identifying potential issues before they become incidents; focused on reliability and uptime.
  • Strong documentation skills—ability to create and maintain runbooks, SOPs, and knowledge base articles for production support.
  • A self-starter mentality that thrives in a rapidly changing, fast-paced environment and tolerates ambiguity while demonstrating problem-solving.
  • Strong analytical and time management skills. Self-motivated and able to handle tasks with minimal supervision.
  • Must be organized, detail-oriented, and able to multi-task. Ability to work well under pressure and deliver results with tight deadlines and under changing priorities.
  • Ability to cross-collaborate with multiple teams and offer value-added solutions to meet objectives.
  • Strong verbal and written communication skills.

Nice To Haves

  • Insurance experience a plus.
  • Data architecture or data engineering related certifications strongly desired.
  • AWS Cloud Practitioner or more advanced AWS certification preferred.
  • Cloud practitioner or more advanced cloud certification preferred.
  • Certifications in any major data warehouse/lakehouse platform preferred.
  • Observability or monitoring tool certification a plus.
  • Snowflake-specific experience is a plus.
  • Familiarity with DOMO administration and production support (or willingness to learn quickly).
  • Exposure to or willingness to learn data technologies: Spark, Kafka, Talend, etc.

Responsibilities

  • Takes ownership of data pipelines and platform workloads after handoff from development teams, becoming the primary point of accountability for their production health.
  • Provides day-to-day production support for on-premises and cloud-based data pipelines, ensuring continuity of operations, timely incident resolution, and adherence to defined SLAs.
  • Monitors, triages, and resolves data pipeline failures, performance degradations, and platform alerts across the data tech stack.
  • Performs root cause analysis on pipeline failures and data issues, implementing fixes to prevent recurrence.
  • Validates runbooks and handoff documentation from development teams; identifies gaps and works with developers to close them prior to full operational ownership.
  • Maintains and enhances production support documentation including runbooks, troubleshooting guides, SOPs, and knowledge transfer materials.
  • Administers the data warehouse/lakehouse environment including user/role management, warehouse configuration, resource monitoring, cost optimization, and security policies.
  • Manages and maintains the transformation layer infrastructure including environment management, job scheduling, model execution monitoring, and performance troubleshooting.
  • Administers and monitors orchestration tools to ensure reliable scheduling, execution health, and pipeline SLAs are met.
  • Supports data ingestion tool administration including monitoring sync health, managing schema changes, and troubleshooting ingestion failures.
  • Supports business intelligence platform administration, with expanded ownership as the platform matures.
  • Oversees data quality monitoring tools, ensuring rules are maintained, alerts are actionable, and data quality SLAs are met.
  • Performs platform hygiene and decluttering—removing deprecated objects, unused connectors, stale pipelines, and orphaned resources across the data tech stack.
  • Leverages cloud monitoring services to monitor data infrastructure components, configure alarms, maintain dashboards, and set up log-based alerting for pipeline health.
  • Maintains and troubleshoots serverless functions that support data workflows from a data infrastructure perspective (overall environment/cloud infrastructure setup is owned by the Cloud Infrastructure team).
  • Implements and manages application performance monitoring and observability tooling to provide end-to-end visibility into data platform performance, resource utilization, and anomaly detection.
  • Builds and maintains operational dashboards and alerting frameworks to proactively identify and resolve data infrastructure issues before they impact downstream consumers.
  • Implements small operational enhancements to existing data pipelines—configuration updates, performance tuning, parameter changes, and minor fixes as defined by the team.
  • Supports data ingestion workflows including monitoring connector health and adjusting configurations for operational needs.
  • Supports regulatory reporting data initiatives from a data operations and pipeline reliability perspective.
  • Identifies and implements operational improvements—automating manual processes, improving alerting coverage, and reducing mean time to resolution (MTTR).
  • Works closely with Data Engineering development teams to receive production handoffs, validate runbooks, and ensure smooth transitions from development to operations.
  • Partners with Enterprise Architecture, Technology, Cloud Infrastructure, and Project teams to ensure operational consistency while maintaining data governance requirements.
  • Contributes to data governance and data quality best practices including operational reviews, incident post-mortems, and continuous improvement initiatives.

Benefits

  • generous medical, dental, and vision plans
  • 401K retirement plans and company match
  • Bonus potential for all positions
  • Paid Time Off
  • Paid holidays throughout the calendar year
  • Support for continued learning
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service