About The Position

We are sharing a specialised consulting opportunity for experienced Senior Technical Architects with strong expertise in cloud infrastructure, distributed systems, platform engineering, DevOps, SRE, production architecture, and resilience engineering to contribute to an advanced AI training and cloud-infrastructure evaluation project. Selected professionals will design realistic cloud-infrastructure tasks and reinforcement-learning environments that test an AI system's ability to design, deploy, secure, scale, troubleshoot, and recover production-grade systems. The work requires substantial hands-on production ownership, strong systems judgement, and the ability to create reproducible environments, deterministic tests, and rigorous reference solutions. No prior experience in AI is required.

Requirements

  • Senior-level experience in technical architecture, cloud infrastructure, platform engineering, DevOps, systems engineering, or SRE
  • Demonstrated ownership of production infrastructure or a production platform
  • Strong knowledge of distributed systems, scalable APIs, queues, autoscaling, and durable storage
  • Practical IAM, private-networking, and service-security experience
  • Strong observability, SLO, deployment, rollback, and disaster-recovery experience
  • Ability to write infrastructure automation or testing tools
  • Strong debugging skills in containerised environments

Nice To Haves

  • Experience with Terraform or OpenTofu is advantageous
  • Experience with AWS, Azure, GCP, Kubernetes, or multi-cloud environments is valuable
  • Background in internal developer platforms, edge infrastructure, chaos engineering, fault injection, or resilience testing is beneficial
  • No prior AI-training or model-evaluation experience is required

Responsibilities

  • Design realistic infrastructure scenarios spanning distributed systems, networking, security, scalability, and reliability
  • Evaluate architectures involving scalable APIs, queues, durable storage, autoscaling, and service coordination
  • Model partial failures, degraded services, and realistic production constraints
  • Assess trade-offs across performance, availability, security, and operational complexity
  • Define clear system requirements and measurable success criteria
  • Build reproducible, containerised technical environments
  • Create reference implementations and intentionally defective variants
  • Develop deterministic integration, load, security, failure-injection, deployment, and recovery tests
  • Validate infrastructure configuration, topology, and runtime behaviour
  • Support repeatable provisioning, execution, and teardown workflows
  • Design scenarios involving IAM, least privilege, private networking, and service-to-service security
  • Incorporate logging, metrics, tracing, SLOs, and operational telemetry
  • Evaluate rolling deployments, rollback strategies, disaster recovery, and resilience
  • Develop realistic troubleshooting and incident scenarios
  • Write automation or testing tools supporting environment setup and validation
  • Create reinforcement-learning environments for multi-step infrastructure reasoning
  • Build golden solutions and defective or adversarial variants
  • Review peer-created tasks for ambiguity, unrealistic assumptions, or validation gaps
  • Improve benchmark difficulty, reproducibility, and grading reliability
  • Maintain strong engineering standards across project deliverables

Benefits

  • Independent contractor engagement
  • Fully remote
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service