Principal AI Data Architect

American IT SystemsAtlanta, GA
Remote

About The Position

We are hiring a Principal AI Data Architect — a hands-on, senior individual contributor responsible for designing, building, governing, and evolving the single source of truth that powers all AI initiatives across the organization. This platform will serve as the foundational backbone for: Conversational AI assistants, Dashboard intelligence, Autonomous AI agents, RAG-powered applications, and Predictive ML models.

Requirements

  • 15+ years in data engineering/architecture
  • 3–5+ years in AI/ML/LLM data platforms
  • Experience designing enterprise-scale AI platforms
  • Strong background in regulated industries
  • Expert: Python, SQL, PySpark
  • Expert: Kafka, Databricks, Delta Lake, Snowflake
  • Expert: AWS (S3, Glue, EKS, Bedrock, Kinesis, Redshift)
  • Expert: Docker, Kubernetes, Terraform, CI/CD
  • Strong: LangChain, LlamaIndex
  • Strong: LLM APIs (OpenAI, Bedrock, Claude, HuggingFace)
  • Strong: Vector DBs (Pinecone, FAISS, ChromaDB, OpenSearch)
  • Strong: Knowledge graphs (Neo4j)
  • Working Knowledge: MLflow, FastAPI
  • Working Knowledge: Observability tools (Grafana, CloudWatch)
  • Working Knowledge: Data lineage and metadata tools

Nice To Haves

  • Degree in Computer Science or related field
  • Experience in presales / solution architecture
  • Background in financial services, SaaS, or regulated industries
  • Familiarity with: MCP (Model Context Protocol)
  • Familiarity with: Agent frameworks (LangGraph, AutoGen, CrewAI)
  • Experience with AI observability systems

Responsibilities

  • Architect and own the enterprise AI data platform.
  • Design multi-domain data models (lakehouse, data mesh, event-driven).
  • Own full data stack: Streaming (Kafka, Spark Structured Streaming), Batch (Databricks, PySpark, Delta Lake), Cloud (AWS, Azure).
  • Eliminate data silos and ensure a unified data layer.
  • Modernize legacy ETL and DWH systems to cloud-native architectures.
  • Design semantic layer with: Ontologies, taxonomies, entity relationships.
  • Build and maintain knowledge graphs (e.g., Neo4j).
  • Define feature store and semantic data contracts.
  • Ensure metadata management, lineage, and auditability.
  • Design embedding pipelines and vector stores: Pinecone, FAISS, ChromaDB, OpenSearch.
  • Define retrieval data contracts for AI systems.
  • Optimize for: Precision, recall, latency, and cost.
  • Build ML and LLMOps pipelines: Training data pipelines, Feature engineering, Model registry (MLflow).
  • Implement CI/CD for AI systems: Validation, deployment, rollback, monitoring.
  • Support LLM fine-tuning workflows: RLHF pipelines, Data curation and filtering.
  • Establish best practices: Versioning, A/B testing, canary releases.
  • Design data services for: Conversational AI (Low-latency APIs for chatbots and copilots), BI & Dashboard Assistants (Semantic query layer and text-to-SQL), Autonomous AI Agents (Tool APIs, memory/state management), Predictive ML Models (Feature pipelines and real-time serving).
  • Create AI Experimentation secure sandbox environments.
  • Implement RBAC and attribute-based access controls.
  • Enforce agent-specific permissions.
  • Ensure: PII masking, Encryption, Audit logging, Compliance (SOX, GDPR, SOC2, AML/KYC).
  • Define schema governance and versioning.
  • Maintain audit trails and data provenance.
  • Build observability for AI agents: Inputs, outputs, reasoning traces.
  • Define evaluation metrics: Accuracy, hallucination rate, relevance.
  • Create feedback loops to improve data quality.
  • Define SLAs for data freshness and AI accuracy.
  • Implement human-in-the-loop workflows.
  • Define reference architecture and standards.
  • Establish: Testing frameworks, CI/CD pipelines, Infrastructure-as-Code (Terraform).
  • Lead design reviews and architecture governance.
  • Conduct internal workshops and enablement sessions.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service