Data Engineer

The Carlyle Group Employee Co.New York, NY
Onsite

About The Position

We are seeking a highly skilled Data Engineer to join our dynamic team. The ideal candidate will be responsible for creating robust data pipelines from various data vendors to gold tables, primarily for our Machine Learning (ML) team, utilizing Snowflake and Databricks platforms. The role demands expertise in Python, deep familiarity with financial data sources, and the ability to deploy complex data pipelines efficiently. This position requires a proactive approach to analyzing, aggregating, and enriching financial data from both private and public companies.

Requirements

  • A bachelor's degree, required
  • At least 6 years of experience in data engineering or a related discipline, with a proven track record of success, required
  • Expertise in Python and SQL, with a strong foundation in data manipulation and analysis.
  • Proficient with Databricks/PySpark and dbt for data warehousing and data transformation tasks.
  • Experience with workflow orchestration tools e.g. Airflow, Temporal
  • Experience working with large language models (LLMs) especially prompt engineering, retrieval-augmented generation (RAG)s, and/or vector databases.
  • Demonstrated experience in designing and implementing complex data systems from the ground up.
  • Proficient in handling large-scale data projects, including data cleaning, ETL, and information retrieval.
  • Excellent communication skills required, both verbal and written.

Nice To Haves

  • Concentration in Computer Science, Math, Physics, STEM or other engineering related field, preferred
  • Experience in the financial services or private equity industry, preferred
  • Knowledge of fundamental principles of machine learning, feature engineering, and knowledge graphs are pluses.
  • Previous experience in a product development or financial services environment is highly desirable.

Responsibilities

  • Build, scale, and maintain robust data solutions to support the firm's objectives.
  • Implement and optimize high-performance data pipelines -- extraction, loading, transformation, and orchestration – that are designed for scalability, reliability, maintainability, and speed.
  • Lead software development projects end to end involving large language models (LLMs), retrieval-augmented generation (RAG) frameworks, and other AI technologies.
  • Champion modern software engineering practices as CI/CD, infrastructure-as-code, containerization, and cloud-native deployments
  • Collaborate closely with business stakeholders to transform use cases into production-ready services and solutions, owning the system from concept to production.
  • Implement rigorous testing and monitoring practices to maintain superior data quality and integrity.
  • Mentor and develop junior team members, fostering a culture of excellence and continuous learning within the team.
  • Be willing to travel up to 20% of the time to collaborate with distributed team members across locations.

Benefits

  • retirement benefits
  • health insurance
  • life insurance and disability
  • paid time off
  • paid holidays
  • family planning benefits
  • various wellness programs
  • annual discretionary incentive program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service