Principal Software Engineer

PeerIslands•Southlake, TX
•Hybrid

About The Position

We are seeking a Senior Software Engineer to develop automation tooling using LLMs and agent frameworks for legacy code analysis, documentation generation, and migration. This role involves integrating outputs with engineering review and validation processes. You will analyze and reverse-engineer legacy database and mainframe workloads to identify dependencies and functional logic for modernization. Responsibilities include developing parsers and utilities to extract metadata and lineage from legacy codebases, designing and building cloud-native migration services, and implementing scalable data platforms and pipelines using Apache Spark (PySpark) and Databricks/Google Cloud Dataproc. You will also architect data models, redesign relational schemas for MongoDB, build configuration-driven ETL frameworks, develop data validation and auditing frameworks, orchestrate data workflows using Apache Airflow, and implement Change Data Capture and reverse ETL mechanisms.

Requirements

  • Bachelor’s or foreign equivalent degree in Computer Science, Computer or Electronic Engineering, or a related field.
  • 5 years of progressive, post-baccalaureate experience in the job offered or as a Software Engineer/Developer, Data Engineer, Programmer Analyst, or in a related/similar position.
  • 5 years in data engineering using Apache Spark (PySpark) and Databricks or Google Cloud Dataproc.
  • Experience with databases and data modeling using SQL, NoSQL or MongoDB.
  • Experience with schema design and optimization.
  • Experience with cloud services such as Microsoft Azure, GCP or AWS for data/compute.
  • Experience with Python software development with Agentic AI, data processing, backend services, and TDD.
  • 2 years with Large Language Models (LLMs), LangChain and LangGraph agent frameworks.
  • Hybrid role, ability to work from home.

Responsibilities

  • Develop automation tooling using LLMs (Large Language Models) and agent frameworks to assist with legacy code analysis, documentation generation, and migration/pipeline scaffolding.
  • Integrate outputs with engineering review, validation, and testing.
  • Analyze and reverse-engineer legacy database and mainframe workloads to identify dependencies, data flows, and functional logic required for modernization and migration.
  • Develop parsers and analysis utilities to extract metadata, lineage, and relationships from legacy codebases to support migration planning, documentation, and implementation.
  • Design and build cloud-native migration services to convert legacy procedures and batch logic into modern microservices and standardized data processing jobs.
  • Design and implement scalable data platforms and pipelines to support enterprise eligibility and operational data processing using Apache Spark (PySpark) and Databricks/Google Cloud Dataproc.
  • Architect data models and storage patterns.
  • Redesign relational schemas into denormalized nested document models suitable for MongoDB to improve downstream query efficiency and application performance.
  • Build configuration-driven Extract, Transform, Load (ETL) frameworks that generate Spark jobs from declarative specifications.
  • Develop data validation, auditing, and threshold-based control frameworks to detect data discrepancies and enforce quality gates across pipeline stages.
  • Orchestrate and monitor data workflows using Apache Airflow (Google Cloud Composer), including dependency management, retries, alerts, and operational controls.
  • Implement Change Data Capture and reverse ETL mechanisms to synchronize changes from MongoDB to data lake storage in near real time for downstream analytics and reporting.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service