Developer Infrastructure Engineer (Silicon)

MatX•Mountain View, CA
•Hybrid

About The Position

MatX's mission is to make the world’s best AI models run as efficiently as allowed by physics, bringing the world years ahead in AI quality and availability. The role covers three areas of equal weight: Branch and release methodology (which branches exist, what each one is for, how a change reaches each one, and how that is verified), CI performance and debugging (keeping a large, EDA-heavy test suite fast, affordable, and reliable, and finding the cause when it is not), and Measurement and diagnostics (instrumenting the above so that the state of the system is visible without anyone having to ask). You would own the specialized work and improve the tooling and documentation that everyone else relies on. Success means people need your direct help less over time, not more. Some common problems include operating the long-lived branches, making "what is the current version of X" mechanical, and keeping CI fast and reliable as it grows. This section describes how the role operates day to day. It suits some people well and others poorly, so it is worth reading closely. Ownership and completion: You would receive objectives rather than instructions. In return, a task is complete only when the change is pushed, CI is green, and the resulting state has been checked. Handing back work with the final verification still pending is the main thing that causes friction. Verification: You will be asked how you know something, and the question is not a criticism. Asserting something unchecked, such as that a resource does not exist, a pool is too small, or a limit cannot be avoided, is what draws pushback. The expected answer to "how do you know?" is the command you ran and its output. "I have not verified that yet" is always an acceptable answer. Durable fixes: The best outcome from a request is usually that the request does not need to be made again. Updating the runbook or adding the check that would have caught the problem is part of the work, not extra scope.

Requirements

  • Deep experience in one of the three areas (branch and release methodology, CI performance and debugging, measurement and diagnostics) plus working knowledge of the other two.
  • Version control internals: ability to explain what a rebase does to commit identity, recover a branch that was deleted, verify by content that a cherry-pick landed, and describe what a squash merge queue does to history.
  • Experience building tooling on top of Git rather than only using it.
  • CI systems experience: merge queues, required checks, event-trigger semantics, app-based authentication, concurrency controls, and self-hosted runners.
  • Ability to identify the cause and quantify it for a slow or flaky pipeline.
  • Ability to read a Bazel query and a build profile.
  • Ability to work in Python and shell against REST and GraphQL APIs.
  • Judgment under a schedule: ability to tell leadership what is and is not in a milestone, including saying "not yet" when it is true, and preferring a change that can be reverted to a migration that cannot.

Nice To Haves

  • Developer-infrastructure work under hard deadlines in another industry (silicon, aerospace, games)
  • Experience with repository structure at scale, whether monorepo or polyrepo
  • Familiarity with EDA flows.

Responsibilities

  • Own the branch structure and make "did this change land where it was supposed to?" answerable from a report rather than by inspection.
  • Make "what is the current version of X" mechanical, ensuring several views of the same deliverable agree: a build from source, an archived release artifact, and whatever a downstream consumer is actually running. Drift detection should run continuously and its results should be visible without anyone asking.
  • Keep CI fast and reliable as it grows, including critical-path analysis, remote-execution behavior under real resource limits, runner-pool sizing and cost, and flake triage that ends in a fix rather than a re-run.
  • Handle everyday collaboration: open a clean PR, review one, take feedback, and decline a change that is wrong.
  • Own the specialized work related to release branch structure, freeze enforcement, version pinning, merge-queue and runner behavior.
  • Improve the tooling and documentation that everyone else relies on.
  • Ensure a task is complete only when the change is pushed, CI is green, and the resulting state has been checked.
  • Update the runbook or add the check that would have caught the problem as part of the work.

Benefits

  • 4 weeks PTO (accrued)
  • 12 company Holidays
  • Up to 3 weeks remote work
  • Company-subsidized Medical (Kaiser or Anthem) for employees & dependents
  • Guardian Dental and Vision insurances for employee & dependents
  • Life insurance (employee only)
  • HSA and FSA offerings via Lively
  • Roth IRA or 401K (or both) retirement plans
  • Up to 5% company contribution to 401K
  • 100% company-paid life insurance (up to $300K)
  • Long-term disability insurances
  • $1500 Professional Development Budget (per year)
  • Onsite team lunch & dinner Monday - Friday
  • Company Uber account for commute
  • Reimbursement for train rides for commute
  • $50/mo to use on the perk you value most
  • $35/mo for cellular reimbursement
  • $40/mo for wifi reimbursement
  • 100% paid mental health benefit via SpringHealth and Guardian EAP
  • Up to 12 weeks paid parental leave regardless of path to parenthood
  • 10 weeks pregnancy disability leave
  • Flexible return-to-work hours
  • Benepass reproductive health & parental benefit
  • Up to $20K/month in AI Resources
  • Dedicated internal AI Tooling Team
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service