Senior AI Engineer

Hinge HealthSan Francisco, CA
CA$136,800 - CA$205,200Hybrid

About The Position

We are seeking a Senior AI Developer to join our computer vision (CV) Engineering team in Montreal, QC. In this role, you will design, build, and operate the cloud infrastructure, deployment systems, and supporting services that power our next generation of multi-modal models and reasoning agents. This role is focused on productionizing LLMs, VLMs, and multi-modal reasoning systems for millions of active users. You will own the infrastructure and operational foundations for deploying, scaling, validating, monitoring, and continuously improving these systems in production. This includes model serving architectures, agent orchestration services, evaluation and validation pipelines, telemetry, observability, and developer tooling. You will work closely with ML scientists, CV engineers, platform engineers, and SRE partners to translate advanced research systems into secure, reliable, and scalable product capabilities. This role offers an exciting opportunity to help expand the scope of the CV Engineering team into cloud-native deployment of frontier models and reasoning agents. Hinge Health Hybrid Model We believe that remote work and in-person work have their own advantages and disadvantages, and we want to be able to leverage the best of both worlds. Employees in hybrid roles are required to be in the office 3 days/week.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related field.
  • 3+ years of experience developing and operating cloud-based services, infrastructure, and APIs in production, ideally on AWS.
  • Experience deploying and operating production ML systems, especially LLMs, VLMs, or other large-scale model systems, including release management, evaluation, and monitoring.
  • Experience with observability and operational tooling such as monitoring, logging, tracing, and alerting platforms.
  • Demonstrated experience with CI/CD and production deployment workflows across development, staging, and production environments.

Nice To Haves

  • Experience using agentic development workflows, including AI-assisted coding and review (e.g. via Claude), plus reusable skills or agents to improve velocity and quality.
  • Experience with LLM and agent development tooling such as LangSmith, LangChain, LangGraph, or MLflow.
  • Experience with GPU-backed inference systems, model serving optimization, and scaling for latency-sensitive applications.
  • Experience deploying or integrating hosted model APIs such as Anthropic, Gemini, or Bedrock.
  • Experience building validation and telemetry systems for generative AI, including regression testing, quality scoring, and production monitoring.
  • Experience with containerized services and orchestration technologies such as Docker, Kubernetes, ECS, or EKS.
  • Experience with workflow orchestration tools such as Temporal or Step Functions.
  • Experience with Databricks or similar platforms for data, experimentation, evaluation, or ML platform operations.
  • Experience with IAM, secrets management, encryption, and compliance-minded cloud controls.
  • Experience with infrastructure as code, especially Terraform, and agent infrastructure such as orchestration, tool use, execution control, memory/state handling, and guardrails.

Responsibilities

  • Deploy and Operate Multi-Modal Models at Scale: Build and maintain production systems for serving LLMs, VLMs, and other multi-modal models with high reliability, low latency, and cost efficiency.
  • Build Validation and Evaluation Systems: Work with ML Scientists and QA to create robust offline and online evaluation pipelines for reasoning quality, model behavior, hallucination risk, policy compliance, latency, and regression detection.
  • Validate and Monitor Quality: Ensure reasoning systems are working as expected in long tail cases in-the-wild, ensuring that exceptions are caught early.
  • Develop Reasoning Agent Infrastructure: Work with other engineers and SRE to develop the orchestration, state management, tool execution, guardrails, and supporting backend services and infrastructure required to run reasoning agents safely and effectively in production.
  • Establish Telemetry and Observability: Design dashboards, traces, logs, alerts, and performance analytics for model inference and agent workflows using modern monitoring platforms.
  • Improve Reliability, Performance, and Cost: Optimize throughput, capacity, fallback behavior, model routing, and infrastructure utilization to support millions of active users.
  • Adopt and Evangelize Best Practices: Evaluate emerging tools, frameworks, and deployment patterns for LLM and agent ops, and help the team standardize on effective practices.
  • Ensure Security and Compliance: Implement security controls, access patterns, and operational safeguards that protect user data and support responsible production use of generative AI systems.

Benefits

  • Inclusive healthcare and benefits: Comprehensive medical, dental, and vision coverage, with additional support for gender-affirming care, family and fertility planning, and travel reimbursements where care isn’t locally accessible.
  • Planning for the future: Traditional and Roth 401(k) options with a 2% company match.
  • Modern life stipends: Flexible stipends to support your learning, development, and modern life needs.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service