Software Engineer, Systems Generalist

Thinking Machines LabSan Francisco, CA
$350,000 - $475,000Onsite

About The Position

Thinking Machines is seeking generalist infrastructure and systems engineers to build the systems powering their foundation models and support internal teams in research and product development. The role involves architecting and scaling core infrastructure, solving complex distributed systems problems, and building robust, scalable platforms. The engineer will work across the full technical stack, collaborating with researchers to accelerate experiments, improve infrastructure efficiency, and enable key insights. This is an evergreen role, meaning applications are continuously reviewed for current and future opportunities. Applicants are encouraged to reapply after gaining more experience, but not more than once every six months. Separate postings for specific project or team needs may also be available.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • Proficiency in at least one backend language (Python or Rust).
  • Experience operating large‑scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).
  • Comfort operating across the stack and owning projects end-to-end.
  • Ability to thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
  • A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.

Nice To Haves

  • Strong debugging across application, OS, and network layers.
  • Proficiency in Python or Rust (or similar), containers, and modern CI.
  • Experience with Kubernetes, controllers/operators, or performance profiling.
  • Familiarity with GPU/ML workflows or large‑scale data/eval pipelines.

Responsibilities

  • Architecting and scaling the core infrastructure.
  • Solving complex distributed systems problems.
  • Building robust, scalable platforms.
  • Working across the full technical stack.
  • Collaborating with researchers to accelerate experiments and improve infrastructure efficiency.
  • Building systems and running large Kubernetes clusters with GPU workloads.
  • Building infrastructure to support Tinker.
  • Designing and optimizing data pipelines using tools like Spark and other modern data infrastructure technologies.
  • Building scalable, reliable data infrastructure while embedding governance best practices.
  • Building tooling, systems, and frameworks to ensure well-configured, optimized developer environments.

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support as needed
  • Visa sponsorship
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service