Backend Engineer_AI Gateway

SBT Global, Inc.Englewood Cliffs, NJ
Onsite

About The Position

We are building an AI Gateway — a high-performance, secure intermediary layer that routes, inspects, and governs all traffic between enterprise clients and LLM/AI service providers. This role requires deep expertise in network programming, async Python, and cloud infrastructure to design and operate a latency-sensitive, policy-driven gateway that handles massive concurrent streaming connections at scale.

Requirements

  • 3+ years of strong hands-on experience with asyncio, aiohttp/anyio, low-level TCP/UDP sockets, WebSocket & SSE servers, TLS handshake, connection-pooling, keep-alive, and zero-copy buffering.
  • 3+ years of experience with FastAPI for designing, developing, and testing production-grade REST/Streaming APIs, including pragmatic use of Pydantic, dependency injection, and OpenAPI docs.
  • Experience with Docker, AKS/EKS/GKE deployment.
  • Experience with HPA/VPA scaling, rolling updates, and service-mesh (e.g., Istio/Envoy) basics.
  • Production use of Redis Streams, RabbitMQ or Kafka.
  • Advanced PostgreSQL query tuning, partitioning, PgBouncer pooling.
  • Experience with HashiCorp Vault (dynamic secrets, transit encryption).
  • Experience with L4/L7 load balancing, TLS termination, reverse-proxy (NGINX/Envoy), mTLS between services, and automated cert lifecycle (ACME/Let’s Encrypt).
  • Experience with Terraform modules for multi-env infra.
  • Experience with GitHub Actions/GitLab CI pipelines with container image scanning and canary/blue-green deployments.
  • Experience with integration with CASB/DLP/WSS or similar edge-security platforms.
  • Experience with centralized logging & audit trails.
  • Experience configuring Azure AD/Entra ID, Okta, SAML 2.0 or OIDC SSO flows.
  • Experience managing Conditional Access & SCIM provisioning.
  • Experience building prompt-engineering workflows, PII anonymisation.
  • Experience monitoring Azure OpenAI (or comparable) usage, rate limits & cost.
  • Experience with custom webhook/REST integrations, circuit-breaker patterns, rate-limiting, auto-retry and health-checking for API Gateways.
  • Experience with end-to-end monitoring with Datadog, Prometheus/Grafana or Azure Monitor.
  • Experience with proactive cost-optimisation and 24/7 incident response (SLA/SLO).

Nice To Haves

  • Enterprise Security Gateways
  • IAM & SSO
  • AI/LLM Pipelines
  • API-Gateway Operations
  • Observability & FinOps

Responsibilities

  • Design, develop and test production-grade REST/Streaming APIs using FastAPI.
  • Implement and manage chunked transfer, bidirectional streaming, back-pressure, and high-throughput data paths for real-time LLM inference.
  • Deploy and manage containerized applications using Docker and Kubernetes (AKS/EKS/GKE), including HPA/VPA scaling, rolling updates, and service-mesh basics.
  • Utilize Redis Streams, RabbitMQ, or Kafka for decoupled task queues, retries, and back-off logic.
  • Perform advanced PostgreSQL query tuning, partitioning, and PgBouncer pooling.
  • Manage secrets and keys using HashiCorp Vault for dynamic secrets and transit encryption.
  • Implement L4/L7 load balancing, TLS termination, reverse-proxy (NGINX/Envoy), mTLS between services, and automated certificate lifecycle management (ACME/Let’s Encrypt).
  • Develop Terraform modules for multi-environment infrastructure.
  • Implement CI/CD pipelines using GitHub Actions/GitLab CI with container image scanning and canary/blue-green deployments.
  • Integrate with CASB/DLP/WSS or similar edge-security platforms.
  • Configure centralized logging and audit trails.
  • Configure Azure AD/Entra ID, Okta, SAML 2.0 or OIDC SSO flows, including Conditional Access and SCIM provisioning.
  • Build prompt-engineering workflows, PII anonymization, and monitor Azure OpenAI usage, rate limits, and costs.
  • Implement custom webhook/REST integrations, circuit-breaker patterns, rate-limiting, auto-retry, and health-checking for API Gateways.
  • Implement end-to-end monitoring with Datadog, Prometheus/Grafana, or Azure Monitor.
  • Perform proactive cost optimization and participate in 24/7 incident response (SLA/SLO).
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service