Staff Data Engineer

ArcherSan Jose, CA
$175,000 - $215,000Remote

About The Position

Archer is an aerospace company based in San Jose, California building an all-electric vertical takeoff and landing aircraft with a mission to advance the benefits of sustainable air mobility. We are designing, manufacturing, and operating an all-electric aircraft that can carry four passengers while producing minimal noise. Our sights are set high and our problems are hard, and we believe that diversity in the workplace is what makes us smarter, drives better insights, and will ultimately lift us all to success. We are dedicated to cultivating an equitable and inclusive environment that embraces our differences, and supports and celebrates all of our team members. Archer is best known for Midnight, our electric aircraft. This role isn't on that program. The AI Products Org is building software for the general aviation industry. As a Staff Data Engineer on the AI Platform team, you will design, build, and operate the data infrastructure that powers large-scale model training and inference. You will own the pipelines, storage systems, and data quality mechanisms that sit upstream of our ML platform, ensuring models train on clean, high-throughput, well-governed data.

Requirements

  • 5+ years of professional data engineering experience excluding internships.
  • BS/MS/PhD in Computer Science, Data Engineering, Software Engineering, or a related field.
  • Hands-on experience building production pipelines with tools like Apache Spark, Flink, Airflow, dbt, or similar batch/streaming frameworks.
  • Deep proficiency with columnar formats (i.e. Parquet), open table formats (Apache Iceberg or Apache Paimon), and object storage systems (S3 or equivalent).
  • Experience with high-throughput messaging systems such as Apache Pulsar or Kafka for real-time data ingestion.
  • Strong SQL skills; experience with StarRocks for large-scale analytical queries and real-time analytics over the lakehouse.
  • Familiarity with AWS data services (S3, Glue, EMR) and containerized workloads (Docker/Kubernetes) in production environments. Airflow/Prefect/Dagster

Nice To Haves

  • Experience with CDC (Change Data Capture) replication from transactional systems (i.e. PostgreSQL → Lakehouse via Debezium or Airbyte).
  • Exposure to audio or time-series data pipelines, including preprocessing for ASR or speech model training.
  • Prior experience in aerospace, aviation, or other safety-critical domains where data lineage and auditability are non-negotiable.

Responsibilities

  • Design and maintain high-throughput, fault-tolerant ingestion and transformation pipelines that feed training workloads at scale, with a focus on latency, throughput, and correctness.
  • Build and operate the data lakehouse — defining table formats (Iceberg, Paimon, Parquet), partitioning strategies, and compaction policies optimized for ML consumption patterns.
  • Instrument pipelines with data quality checks, lineage tracking, and anomaly detection so that model failures trace back to data problems quickly.
  • Partner with ML engineers to define feature stores, dataset versioning, and experiment-to-production data contracts; integrate with tools like MLflow for dataset and artifact tracking.
  • Work closely with AI researchers, platform engineers, and software engineers to understand data access patterns, optimize query performance, and unblock training runs.

Benefits

  • pay-for-performance culture
  • reward performance that supports the Company’s business strategy
  • base pay between $175,00.00 - $215,000.00
  • working with and providing reasonable accommodations to job applicants with physical or mental disabilities, and those with sincerely held religious beliefs
  • Archer's Candidate Privacy Policy
  • Archer is unable to provide work visa sponsorship for this position at the present time.
  • Equal Opportunity employer committed to diversity and inclusivity in the workplace.
  • All aspects of employment are decided on the basis of merit, qualifications, and business needs.
  • We do not discriminate based upon race, color, religion, sex, sexual orientation, age, national origin, disability status, protected veteran status, gender identity or any other characteristic protected by federal, state or local laws.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service