Senior Software Engineer, Artificial Intelligence/LLM

Beacon AI•San Carlos, CA
•$173,000 - $225,000•Hybrid

About The Position

Beacon AI is seeking a Senior Software Engineer specializing in Artificial Intelligence and Large Language Models (LLMs). The company is a fast-moving team focused on building an AI platform to enhance aviation safety, efficiency, and capability. Backed by significant investment, Beacon AI has secured numerous Department of Defense contracts and partnered with major airlines. The company culture emphasizes small, focused teams that own their work, ship quickly, and learn rapidly, fostering innovation in human-AI collaboration within aviation. The role involves end-to-end ownership of LLM-powered product features, including designing retrieval and tool-calling flows, developing services, building evaluations and guardrails, and monitoring production performance (cost, latency, quality). Collaboration with ML/infra teammates on embeddings, indexing, and model hosting, as well as with product teammates on user experience, is expected. The role requires a strong focus on reliability within a safety-critical domain.

Requirements

  • Senior engineers who can own a feature or service end-to-end.
  • 5-8 years experience, including some in production ML/LLM systems.
  • Make independent architecture and eval decisions within your service's scope.
  • Partner directly with product and infra.
  • Shipped LLM apps: You’ve put LLM features in front of users and improved them with data.
  • Strong builder: Comfortable writing production code, tests, and docs. You keep things simple and observable.
  • RAG and tools depth: You understand embeddings, chunking, vector search tradeoffs, and function calling.
  • Quality mindset: You design evals, define success metrics, and iterate based on evidence.
  • Cost and latency aware: You track p95, hit SLAs, and reduce cost without hurting quality.
  • Clear communicator: You explain tradeoffs and align partners across product, infra, and security.
  • Ownership: You can take a feature from design through production with minimal oversight.
  • Due to U.S. export control regulations, we can only hire U.S. Persons (U.S. citizens, Green Card holders, lawful permanent residents, or individuals granted asylum or refugee status).
  • All work must be performed in the United States.

Nice To Haves

  • Experience with Bedrock, OpenSearch Serverless, pgvector, Pinecone, or Weaviate.
  • Prompt versioning, guardrails, and provider routing in production.
  • Multimodal work with time series or video.
  • Familiarity with GPU inference, Triton, or TensorRT-LLM.
  • Aviation or other safety-critical domain exposure.
  • DevOps basics for CI/CD, IaC, and secure secrets handling.

Responsibilities

  • Ship LLM-powered product features end-to-end.
  • Design retrieval and tool-calling flows.
  • Write services that run LLM features.
  • Build evals and guardrails for LLM features.
  • Monitor cost, latency, and quality in production.
  • Partner with ML/infra teammates on embeddings, indexing, and model hosting.
  • Partner with product teammates on user experience and outcomes.
  • Build user-facing LLM features.
  • Design and implement retrieval-augmented generation and tool-calling flows using frameworks like LangChain or equivalent primitives.
  • Deliver robust JSON and schema-bound outputs with validation, retries, and fallbacks.
  • Add function calling to integrate with internal tools, search, routing, and data services.
  • Own the service layer, shipping APIs and workers in Python or TypeScript with clear contracts, streaming, and backoff.
  • Add caching, request shaping, prompt templates, and context packing to control latency and cost.
  • Integrate with AWS Bedrock, OpenAI, Anthropic, or self-hosted endpoints.
  • Collaborate with infrastructure teammates to develop chunking, embeddings, and indexing capabilities for documents, time series, and multimedia.
  • Choose and tune vector backends such as OpenSearch, pgvector, or Pinecone.
  • Keep knowledge bases fresh with data syncs from S3, Aurora, DynamoDB, and external sources.
  • Create offline evals and golden sets for prompts, retrievers, and tools.
  • Stand up online metrics for task success, hallucination rate, retrieval precision/recall, p95 latency, and cost per request.
  • Run A/B tests and prompt/version rollouts with guardrails and canaries.
  • Implement content and policy checks, PII detection and redaction, access controls, and auditing.
  • Design human-in-the-loop paths for sensitive actions.
  • Handle aviation data with care and follow internal security standards.
  • Add tracing, logs, and dashboards for model calls, token usage, errors, and saturation.
  • Debug tricky failures across retrieval, prompts, tools, and providers.
  • Transform an internal knowledge base into a low-latency RAG service, complete with explicit schemas and evaluations.
  • Add tool-calling to automate a repetitive cockpit or ops workflow with guardrails and audit trails.
  • Reduce the cost per request through improved chunking, caching, and prompt refactoring, while maintaining task success rates.

Benefits

  • Healthcare: 100% of employee medical premiums covered; 25% for dependents
  • Time Off: 3 weeks PTO plus 13+ paid company holidays
  • 401(k): Offered (no current employer match, but we are committed to enhancing this benefit in the future)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service