Data Engineer (Hybrid)

RTXSanta Isabel, PR
Hybrid

About The Position

Collins Aerospace is seeking an experienced Python and PySpark Developer to design, build, and optimize our next-generation big data pipelines. In this role, you will handle large-scale datasets, optimize distributed computing clusters, and bridge the gap between raw data ingestion and production-ready analytics. The ideal candidate thrives on optimizing cluster performance, resolving data skewness, and writing clean, maintainable Python code. This is a hybrid role in Puerto Rico. Applicants must reside in Puerto Rico. Do you want to be part of a new, exciting initiative to combine foundational IT with new digital technologies? Our Digital Technology team is driving business efficiencies and a better customer experience by connecting technologies, people, information and processes. From making aircraft more electric, intelligent and integrated to building new software platforms such as Internet of Things, big data, artificial intelligence, and blockchain, there’s no better place to be right now than in digital. If you’re an agile thinker who enjoys utilizing modern technology to make big improvements, then you’re a perfect fit for this team. Join Collins Aerospace to help us revolutionize the aerospace industry today!

Requirements

  • Typically requires a University Degree and minimum 5 years prior relevant experience or an Advanced Degree in a related field and minimum 3 years of experience
  • Must be a U.S. Citizen.
  • Strong proficiency in Python (OOP, concurrency, data structures) and advanced SQL.
  • Deep production experience with Apache Spark / PySpark (Data Frames, Spark SQL, RDDs).
  • Hands-on experience with cloud data platforms like AWS (EMR, Glue), Azure (Databricks), or GCP.
  • Experience working with Snowflake, Big Query, Redshift, or Synapse.
  • Proficient with Git, Docker, and automated deployment pipelines.

Nice To Haves

  • PySpark MLlib or deploying Machine Learning models to production.
  • Familiarity with streaming technologies like Apache Kafka or Spark Structured Streaming.
  • Databricks Certified Data Engineer or Apache Spark Developer certifications.

Responsibilities

  • Design and deploy robust batch and streaming ETL/ELT pipelines using PySpark and Python.
  • Optimize Spark jobs by tuning configurations, managing partitioning, and resolving data skew or OOM (Out of Memory) errors.
  • Implement modern data lakehouse architectures using Delta Lake, Iceberg, or Hudi.
  • Build and maintain complex workflow DAGs using orchestration tools like Apache Airflow.
  • Develop backend Python services or REST APIs (e.g., Fast API, Flask) to expose processed data to downstream applications.
  • Write clean, modular, and unit-tested code while participating in rigorous code reviews

Benefits

  • Medical, dental, and vision insurance
  • Three weeks of vacation for newly hired employees
  • Generous 401(k) plan that includes employer matching funds
  • Participation in the Employee Scholar Program (ESP)
  • Life insurance and disability coverage
  • Employee Assistance Plan, including up to 8 free counseling sessions
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service