Data Engineer (PySpark)

CapgeminiMississauga, ON
CA$80,000 - CA$90,000

About The Position

Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d like, where you’ll be supported and inspired by a collaborative community of colleagues around the world, and where you’ll be able to reimagine what’s possible. Join us and help the world’s leading organizations unlock the value of technology and build a more sustainable, more inclusive world.

Requirements

  • 5-8 years of relevant experience
  • Experience in systems analysis and programming of software applications
  • Experience in managing and implementing successful projects
  • Working knowledge of consulting/project management techniques/methods
  • Ability to work under pressure and manage deadlines or unexpected changes in expectations or requirements
  • In-depth understanding of HDFS architecture, data storage, and fault tolerance mechanisms. Experience with HDFS commands and administration.
  • Solid understanding of YARN resource management and job scheduling.
  • Fundamental understanding of MapReduce programming paradigm, even if primary development is in Spark/Flink. Knowledge of Zookeeper for distributed coordination services.
  • Strong proficiency in Spark Core, Spark SQL, Spark Streaming, and Spark GraphX (beneficial).
  • Expert-level programming skills in Python and Pyspark, specifically for developing Spark applications.
  • Experience with Spark performance optimization techniques (e.g., caching, partitioning, shuffle optimizations, memory management).
  • Experience with PySpark for data processing.
  • Familiarity with data manipulation libraries (Pandas, NumPy).
  • Scripting for automation and data orchestration.
  • Complex query writing, subqueries, window functions, and performance tuning.
  • HBase (for real-time access to large datasets within Hadoop).
  • Cassandra, MongoDB, or similar.
  • Familiarity with RDBMS concepts and SQL for data integration.
  • Understanding of dimensional modeling, fact and dimension tables, star/snowflake schemas.

Nice To Haves

  • Spark GraphX (beneficial)

Responsibilities

  • Monitor and control all phases of development process and analysis, design, construction, testing, and implementation as well as provide user and operational support on applications to business users
  • Utilize in-depth specialty knowledge of applications development to analyze complex problems/issues, provide evaluation of business process, system process, and industry standards, and make evaluative judgement
  • Recommend and develop security measures in post implementation analysis of business usage to ensure successful system design and functionality
  • Consult with users/clients and other technology groups on issues, recommend advanced programming solutions, and install and assist customer exposure systems
  • Ensure essential procedures are followed and help define operating standards and processes
  • Serve as advisor or coach to new or lower level analysts
  • Has the ability to operate with a limited level of direct supervision.
  • Can exercise independence of judgement and autonomy.
  • Acts as SME to senior stakeholders and /or other team members.

Benefits

  • Paid time off based on employee grade (A-F), defined by policy: Vacation: 12-25 days, depending on grade, Company paid holidays, Personal Days, Sick Leave
  • Medical, dental, and vision coverage (or provincial healthcare coordination in Canada)
  • Retirement savings plans (e.g., 401(k) in the U.S., RRSP in Canada)
  • Life and disability insurance
  • Employee assistance programs
  • Other benefits as provided by local policy and eligibility
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service