Data Engineer - Enterprise Data Engineering & Analytics

UT MD Anderson Cancer CenterHouston, TX
Onsite

About The Position

The Data Engineer role is a pivotal position within the Enterprise Data Engineering & Analytics Department, supporting the design, build, and operationalization of integrated data pipelines and analytics solutions that enable MD Anderson’s digital business initiatives. The Data Engineer works across the Context Engine framework to deliver end-to-end data engineering solutions while partnering closely with Enterprise Data Engineering & Analytics teams and other institutional stakeholders. The Data Engineer contributes to the mission of MD Anderson Cancer Center, a leading institution focused on cancer care, research, education, and prevention. In this role, the Data Engineer helps advance enterprise analytics capabilities by ensuring secure, governed, and reusable data assets that accelerate insights and improve time-to-solution across MD Anderson.

Requirements

  • Bachelor's Degree
  • 2 years Clinical, relevant healthcare information technology, or relevant business experience.
  • Must obtain at least one Epic Data Model certification (Clinical, Access, or Revenue) issued by Epic. within 180 Days
  • Must pass pre-employment skills test as required and administered by Human Resources.

Nice To Haves

  • Advanced education in analytics or computer science
  • Hands-on experience building data pipelines in healthcare or research environments
  • At least 2 years of Cloud Data Management Framework (Foundry/Fabric) experience
  • Hands-on use of Large Language Models (LLMs) in real-world projects
  • Python or Spark development
  • Analytics delivery
  • Epic certification
  • Ability to collaborate across technical and clinical teams
  • Bachelor's in computer science
  • Master's degree Business Analytics, Computer Science, Information Technology, Data Science, or related.
  • With preferred degree, no experience required.
  • May substitute required education with years of related experience on a one to one basis.
  • 3-5 years creating data pipelines in a healthcare research environment
  • Experience building and maintaining analytical reports and dashboards
  • Problem solving skills and ability to translate business/clinical requirements into reliable data models
  • Analytics & reporting- cloud data management solutions like Foundry, Fabric etc
  • Data pipeline & ETL development -hands on experience designing, building and maintaining pipelines using python/spark
  • Hands-on use of Large Language Models (LLMs) in real-world projects, such as integrating generative AI solutions into applications, workflows, or analytics platforms.
  • Candidates should be familiar with prompt engineering, model evaluation, and responsible AI practices.
  • Experience collaborating with cross-functional teams to deploy and scale LLM-powered features is highly desirable.
  • Preferred certifications: EPIC Cogito, Clarity, Caboodle, Clinical Data Model, etc.
  • Prefer a candidate in Houston Texas or surrounding area.

Responsibilities

  • Participate in end-to-end solution delivery that increases information capabilities and realizes data value across the institution
  • Build and test end-to-end data pipelines across ingestion, curation, transformation, modeling, and consumption within the Context Engine framework
  • Integrate data governance processes across data provenance, security, data quality, ontology, and metadata management
  • Participate in planning, architecture, analysis, design, and build of data pipelines in partnership with IS, Data Offices, and Data Governance teams
  • Contribute to existing data pipelines spanning acquisition, integration, and consumption for defined use cases
  • Build data curation pipelines including profiling, specification creation, cleansing, transforming, standardizing, mastering, harmonizing, validating, and aggregating data
  • Monitor and support data quality across the Context Engine
  • Incorporate repeatable solution designs and data models to support reuse and scalability
  • Promote effective data management practices and understanding of analytics across the enterprise
  • Adhere to IS division standard operating procedures and all MD Anderson policies
  • Maintain build standards and governance oversight sign-off aligned with institutional data strategy
  • Participate in documentation preparation for enhancements or new technology
  • Perform quality control, testing, and peer review of analytics builds
  • Support system updates, releases, change control processes, and after-hours support as required
  • Train data scientists, analysts, end users, and data consumers on data pipelining and preparation techniques
  • Assist in establishing training plans and curricula for Context Engine tools
  • Provide institutional, department, and one-on-one training on EDEA deliverables
  • Support liaison relationships with customers and OneIS partners to deliver effective technical solutions
  • Explore and promote modern tools, techniques, and architectures to automate data preparation and integration tasks
  • Improve productivity by reducing manual and error-prone processes
  • Model OneIS values through integrity, partnership, quality, and continuous improvement

Benefits

  • Employer-paid medical coverage starting day one for employees working 30+ hours/week
  • Optional group dental, vision, life, AD&D, and disability insurance.
  • Accruals for PTO and Extended Illness Bank
  • Paid holidays
  • Wellness
  • Childcare
  • Other leave options.
  • Tuition Assistance Program after six months of service
  • Access to extensive wellness, fitness, and employee resource groups.
  • Defined-benefit pension through the Teachers Retirement System
  • Voluntary retirement plans
  • Employer-paid life and reduced salary protection programs.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service