About The Position

The Infosys Data and Analytics (DNA) unit is at the forefront of transforming data into actionable insights, driving business growth and operational efficiency. We specialize in leveraging advanced AI and analytics to create innovative solutions that address complex business challenges. Our team is dedicated to pioneering the future of data-driven decision-making, enabling organizations to unlock new opportunities and achieve sustainable success. Join us to be part of a dynamic team that is revolutionizing the way businesses harness the power of data and AI. At Infosys DNA, you'll have the opportunity to work with cutting-edge technologies, collaborate with industry experts, and contribute to transformative projects that shape the future of business. We are committed to fostering a culture of continuous learning and growth, ensuring that our team members thrive in a dynamic and supportive environment. If you're passionate about AI and eager to make a significant impact, the Infosys DNA unit is the perfect place for you to grow and excel.

Requirements

  • experience in data engineering, big data processing, or a closely related role.
  • Strong hands-on proficiency in Python, SQL, PySpark, and Apache Spark.
  • Experience with Apache Airflow for workflow orchestration and scheduling.
  • Hands-on experience with AWS services such as S3, Glue, Redshift, or Lambda.
  • Experience working with Snowflake and/or Databricks.
  • Proficiency in data transformation, data validation, performance tuning, and troubleshooting.
  • Working experience with relational databases such as PostgreSQL or MySQL.
  • Ability to process and integrate large datasets across multiple file formats and sources.
  • Strong analytical, problem-solving, collaboration, and communication skills
  • Use MLflow for experiment tracking and model lifecycle management where applicable.

Nice To Haves

  • Experience with supervised machine learning, feature engineering, model training, and model evaluation.
  • Hands-on experience with Pandas, NumPy, Scikit-learn, and Matplotlib.
  • Experience with MLflow for model tracking, deployment, and lifecycle management.
  • Experience implementing data-quality frameworks such as Great Expectations.
  • Knowledge of fraud detection, predictive analytics, financial data, billing, or invoice-processing use cases.
  • Experience integrating Spark workflows with cloud object storage and optimizing distributed processing.
  • AWS Certified Data Engineer - Associate or a comparable cloud/data engineering certification

Responsibilities

  • Contribute to the requirements elicitation process by documenting assigned parts of business requirements, in line with guidance provided
  • Facilitate software application design discussions, and document design decisions to guide the technical team towards building software solutions
  • Participate in coding and integrate new features or updates into existing applications, with a focus on maintaining system stability
  • Conduct code reviews, do changes to the codebase and maintain code repositories
  • Implement test strategies, analyse results, and coordinate bug fixes to uphold the software quality standards
  • Develop user training programs, documentation, and support frameworks to ensure a smooth transition to new software applications
  • Actively participate in resolving production issues and recommend preventive strategies to enhance system reliability
  • Maintain detailed records of code, testing techniques, and support activities to enrich the knowledge base and assist other similar projects
  • Design, develop, and maintain scalable data pipelines using PySpark, Python, SQL, and Apache Spark.
  • Experience designing, developing, and maintaining scalable ETL or ELT data pipelines.
  • Build and optimize batch-processing workflows for high-volume structured and semi-structured datasets.
  • Develop ETL and data-integration solutions using Databricks, AWS Glue, Apache Airflow, and cloud storage services.
  • Ingest, standardize, transform, and validate data from sources such as CSV, JSON, Parquet, transactional systems, and cloud platforms.
  • Create automated data-quality checks to improve completeness, consistency, and accuracy of downstream data.
  • Optimize Spark workloads and complex SQL queries, including CTEs and window functions, to improve processing and reporting performance.
  • Develop and manage Airflow DAGs for workflow scheduling, orchestration, monitoring, and operational reliability.
  • Prepare data for machine learning through cleansing, feature selection, feature engineering, and scalable transformation pipelines.
  • Support supervised machine learning model development, hyperparameter tuning, evaluation, deployment, monitoring, and retraining.
  • Collaborate with data scientists, analysts, business teams, and engineering stakeholders to deliver reliable data products.
  • Troubleshoot pipeline failures, resolve processing bottlenecks, and continuously improve performance and automation.

Benefits

  • Medical/Dental/Vision/Life Insurance
  • Long-term/Short-term Disability
  • Health and Dependent Care Reimbursement Accounts
  • Insurance (Accident, Critical Illness , Hospital Indemnity, Legal)
  • 401(k) plan and contributions dependent on salary level
  • Paid holidays plus Paid Time Off
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service