Information Technology_USA - USA_Engineer

Real SoftJacksonville, FL
Onsite

About The Position

We are seeking a highly skilled Data Engineer with expertise in advanced architecture, system design, distributed computing, streaming, and cloud infrastructure. The role involves leading the overall platform vision, ensuring systems can handle scale, and setting coding standards. Responsibilities include core programming, database management, pipeline orchestration, DevOps practices, and leadership. The ideal candidate will mentor junior engineers, estimate project timelines, and translate business needs into technical specifications.

Requirements

  • Python
  • Pyspark
  • AWS Lambda
  • S3
  • AWS glue
  • Kinesis
  • SQL
  • Apache Flink
  • Confluent Kafka
  • Java (optional)
  • Apache Spark or Ray
  • Kafka, Kinesis, or Flink
  • AWS cloud expertise
  • Python or Scala
  • Snowflake, Redshift, or BigQuery
  • DynamoDb or Cassandra
  • Apache Airflow or Prefect
  • Docker
  • Kubernetes
  • Terraform
  • 10+ years of experience

Nice To Haves

  • AI-LLM
  • Gen-AI
  • Advanced Java Concepts
  • Ray

Responsibilities

  • Advanced Architecture & System Design: Responsible for the overall platform vision and ensuring systems do not break under scale.
  • Distributed Computing: Mastery of frameworks like Apache Spark or Ray for massive-scale parallel data processing.
  • Streaming & Event-Driven Architecture: Deep understanding of real-time pipeline design using Kafka, Kinesis, or Flink.
  • Cloud Infrastructure: Expertise in at least one major public cloud (AWS), specifically understanding storage/compute decoupling and cost optimization.
  • Core Programming & Database Management: Set coding standards and review code, requiring complete fluency in the fundamentals.
  • SQL: Advanced mastery for metrics computation, window functions, and query performance tuning across relational and columnar databases (e.g., Snowflake, Redshift, BigQuery).
  • Scripting Languages: High proficiency in Python or Scala for writing reusable pipeline code and interacting with APIs.
  • Data Storage: Deep familiarity with both columnar/analytical stores and NoSQL databases (e.g., DynamoDb, Cassandra).
  • Pipeline Orchestration & DevOps: Ensuring pipelines run smoothly, idempotently, and securely in production.
  • Workflow Orchestration: Ability to architect Directed Acyclic Graphs (DAGs) in tools like Apache Airflow or Prefect.
  • CI/CD & Infrastructure as Code (IaC): Applying software engineering principles to data by using Docker, Kubernetes, and Terraform.
  • Data Governance & Security: Implementing Role-Based Access Control (RBAC), data masking, and compliance frameworks.
  • Leadership & Soft Skills: Mentor junior engineers, estimate project timelines, and translate ambiguous business needs into concrete technical specifications.
  • Mentorship & Code Review: Fostering a collaborative development environment and enforcing style guidelines.
  • System Observability: Building logging, monitoring, and alerting mechanisms so the team knows exactly when and why pipelines fail.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service