Your Work Shapes the World at Caterpillar Inc. When you join Caterpillar, you're joining a global team who cares not just about the work we do – but also about each other. We are the makers, problem solvers, and future world builders who are creating stronger, more sustainable communities. We don't just talk about progress and innovation here – we make it happen, with our customers, where we work and live. Together, we are building a better world, so we can all enjoy living in it. The Data Pipeline Engineer is responsible for developing, maintaining, and improving the data infrastructure that enables autonomy, robotics, and machine learning development across multiple programs. This role designs and supports scalable solutions for data ingestion, storage, discovery, processing, and annotation, ensuring that high-quality robotics data is accessible and usable throughout its lifecycle. The engineer collaborates closely with software, autonomy, AI/ML, and platform teams to build reliable data workflows, manage cloud-based storage and compute environments, and develop tools that accelerate data-driven development. Key responsibilities include supporting log processing systems, improving data integrity and discoverability, enabling large-scale compute workloads, automating data management processes, and contributing to platforms that support model training, validation, simulation, and analytics. Success in this role requires strong software engineering skills, experience with cloud infrastructure and data systems, a commitment to operational excellence, and the ability to work effectively within cross-functional teams to deliver scalable and maintainable solutions. The Data Pipeline team's mission is to provide robust data infrastructure, compute resources, data management capabilities, annotation workflows, and development tools that accelerate autonomy and robotics development.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior