TetraScience is seeking a Lead Software Platform Engineer to work at the intersection of distributed systems and MLOps. This role involves owning and scaling the company's AI and data infrastructure, which serves as a product for customers and a foundation for other engineering teams. The engineer will architect cloud-based services and MLOps infrastructure to enable production-grade AI/ML workflows, collaborating with Applied AI engineers, data engineers, and platform teams. The position requires acting as the technical design authority for how models, LLMs, and agents operate in production, covering the entire model and prompt lifecycle, evaluation and observability, security, and cost/latency controls for economically viable AI at scale. This is a high-impact role for an engineer experienced in shipping AI/ML infrastructure as a multi-tenant product, with responsibilities including architecting the AI/ML platform's service and API surface, managing the end-to-end model and prompt lifecycle, designing the inference substrate for real-time and batch workloads, integrating AI models and LLMs using architectures like RAG, and architecting the agentic layer. The role also involves designing security into the AI platform, building evaluation and quality infrastructure, establishing observability, designing for reproducibility in regulated environments, contributing to infrastructure-as-code and deployment automation, owning production readiness, acting as an SME and design authority, and setting technical direction on emerging AI infrastructure.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed