Software Engineer - AI Systems (Go)

Stanbridge UniversityIrvine, CA
Remote

About The Position

Stanbridge University is seeking Software Engineers – AI Systems (Go) to design, build, and operate production AI systems that perform complex, multi-stage work reliably and at scale. This is a hands-on engineering role focused on a challenging class of problems: long-running AI workflows that call models, tools, and external APIs; maintain state across extended executions; produce structured content and generated media; interact with human reviewers; recover from partial failures; and consistently deliver accurate results to users. The core engineering challenges extend well beyond prompt development. These systems must account for provider failures and rate limits, interrupted workflows, changing state, non-deterministic model behavior, incorrect or unsupported outputs, variable latency and cost, and deployments occurring while work is in progress. The successful candidate will combine strong software and distributed-systems engineering judgment with practical experience building and operating LLM-backed applications, AI agents, or agent-based systems in production. You will join at a stage where significant architecture remains to be designed and built. Engineers in this role will have substantial ownership over the patterns, services, infrastructure, and engineering standards that shape the University's AI systems.

Requirements

  • Substantial professional experience developing and operating production software, with strong experience in Go (Golang) or demonstrated depth in another backend language with the ability to become productive in Go quickly.
  • Hands-on experience shipping an LLM-backed application, AI agent, or agent-based system into production and supporting it after deployment.
  • Strong understanding of AI application architecture, including how models, agents, tools, APIs, data sources, services, and application logic interact.
  • Demonstrated distributed-systems engineering knowledge, including concurrency, asynchronous processing, queues, idempotency, retries, partial failure, state management, and recovery.
  • Experience designing systems that remain reliable when individual services, providers, or workflow stages fail.
  • Demonstrated testing discipline for systems involving non-deterministic behavior.
  • Strong experience designing and consuming HTTP APIs.
  • Experience with relational databases, SQL, and persistent application state.
  • Experience building, deploying, monitoring, and troubleshooting backend services in production environments.
  • Understanding of software architecture, testing, debugging, observability, and production engineering practices.
  • Ability to independently own ambiguous technical problems from investigation through production implementation.
  • Strong analytical judgment and the ability to balance reliability, quality, performance, complexity, and cost.

Nice To Haves

  • Experience with AI agent frameworks, orchestration patterns, or custom agent architectures, including an understanding of when a framework may not be appropriate.
  • Experience with retrieval-augmented generation (RAG), embeddings, vector databases, semantic search, or knowledge-retrieval architectures.
  • Experience developing evaluation and observability systems for LLM applications, including tracing, regression suites, quality dashboards, or automated evaluation.
  • Experience with judge models, structured-output validation, evidence checking, or other AI quality-control mechanisms.
  • Experience designing human-in-the-loop workflows involving review, approval, intervention, or modification of active workflow state.
  • Experience with document processing, headless-browser rendering, text-to-speech, image generation, video generation, or other media pipelines.
  • Experience with event-driven architectures, durable job queues, and asynchronous processing at scale.
  • Experience with containers, cloud infrastructure, CI/CD, and production deployment environments.
  • Understanding of prompt injection, authorization boundaries, data isolation, and security considerations when untrusted content is processed by AI systems.
  • Experience developing systems in environments where the accuracy of generated output carries significant operational, regulatory, compliance, or safety implications.

Responsibilities

  • Design, develop, test, deploy, and operate production-quality software and backend services, primarily using Go (Golang).
  • Architect and build AI agents and agent-based systems capable of using tools, interacting with APIs, maintaining state and context, and executing complex multi-step workflows.
  • Design durable workflows that can checkpoint, resume, retry, recover, and safely continue execution following partial failures or system interruptions.
  • Determine how models, agents, tools, APIs, data sources, services, and deterministic application logic should divide responsibilities within an AI system, including recognizing when an AI model is not the appropriate solution.
  • Design integrations across multiple AI model providers, including routing, failover, rate-limit management, health monitoring, and degradation strategies.
  • Build safeguards including structured-output validation, error handling, retries, quality gates, evidence validation, and automated recovery mechanisms.
  • Develop testing strategies for non-deterministic AI behavior using techniques such as deterministic fixtures, recorded and replayed interactions, regression suites, evaluation harnesses, and quality baselines.
  • Develop observability capabilities including tracing, metrics, logging, evaluation data, and replayable execution histories to support production debugging and performance analysis.
  • Design human-in-the-loop workflows incorporating review, approval, intervention, and modification of workflow state.
  • Build and maintain APIs, backend services, relational data models, job-processing infrastructure, and supporting application components.
  • Develop secure methods for ingesting and processing user-supplied documents and other external content.
  • Optimize systems for reliability, latency, throughput, scalability, output quality, and cost per execution.
  • Build reusable engineering patterns and shared components that simplify the addition of new agents, model providers, tools, workflows, and output types.
  • Own technical problems from initial investigation and architecture through implementation, deployment, monitoring, troubleshooting, and ongoing production operation.
  • Collaborate directly with product stakeholders and domain experts to translate qualitative requirements into measurable system behavior and technical solutions.
  • Contribute to architectural decisions and engineering standards for AI-powered applications across the University.

Benefits

  • Health Care Plan (Medical, Dental & Vision)
  • Retirement Plan (401k)
  • Exciting university events
  • Seasonal motivational health and wellness challenges
  • Work/Life Balance initiatives
  • Onsite wellness program / Staff Chiropractor
  • Life Insurance (Basic, Voluntary & AD&D)
  • Paid Time Off (Vacation, Sick & Public Holidays)
  • Family Leave (Maternity, Paternity)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service