Senior AI & Data Consultant

TruistAtlanta, GA
Onsite

About The Position

The Senior AI and Data Consultant is a senior, hands-on technical leader within the AI and Data Support Operations organization. This teammate is accountable for elevating the reliability, resiliency, and operational excellence of critical enterprise platforms across hybrid cloud and on prem environments. Acting as both a hands on cloud expert and a cross domain influencer, the Data Consultant drives systemic improvements in observability, automation, AIOps adoption, fault tolerance, and incident management within AWS. The role partners closely with Application Development, Infrastructure, Production Support, Platform Delivery, Architecture, Cybersecurity, Risk, and Business technology teams to uplift operational practices and deliver stable, predictable, and scalable services. The AIDC delivers measurable impact through deep expertise in cloud technologies, modern operational tooling and enterprise-scale incident/problem management.

Requirements

  • Experience with Infrastructure as Code (IaC) products (Terraform, Cloudformation) for rapid and consistent deployment of virtual infrastructure.
  • Proficiency with AWS Identity and Access management suite and how it will integrate with Truist Active Directory.
  • Expert level knowledge of AWS virtual OS images and management of core components such as EC2 instances and AMIs.
  • Familiarity with cloud networking concepts, namely VPCs, load balances and security groups for access to virtual assets.
  • Experience with meta tags used for grouping of assets into subcategories.
  • Fundamental understanding of monitoring, alerting and observability tools in AWS.
  • Lead major and high-severity incident response efforts, focusing on diagnosing technical root causes therein, and driving multi-team technical resolution.
  • Drive problem management to closure, ensuring systemic fixes replace recurring operational risks.
  • Establish and maintain standardized incident playbooks, escalation paths, and communication frameworks.
  • Architect and deliver automation solutions that eliminate toil, reduce MTTR, and increase service resilience.
  • Implement intelligent alerting, anomaly detection, and event correlation leveraging AI and AIOps tools.
  • Guide and enforce SLO/SLI adoption across product teams, ensuring metrics inform decision-making and prioritization.
  • Infrastructure as Code (IaC) creation and updates for deploying assets in Amazon Web Services tenant.
  • Deploy patches and device hardening configurations to improve security posture on AI and Data servers.
  • Enhance telemetry coverage across logs, metrics, traces, and events using platforms such as Dynatrace and Splunk.
  • Define and standardize enterprise observability practices, dashboards, and KPIs.
  • Ensure operational readiness of applications and platforms through resiliency testing, chaos engineering, and failure-mode validation.
  • Partner with Delivery, Architecture, Security, and Risk teams to embed reliability and resilience into design and execution.
  • Act as a change agent to elevate operational maturity and drive transformative improvements across Wholesale.
  • Develop, maintain, and enforce runbooks, response playbooks, and automated recovery patterns.
  • Contribute to enterprise SRE frameworks, templates, and maturity models.
  • Promote consistent adoption of best practices across domains and lines of business.
  • Coach and mentor Associate, Professional, and Senior SREs to build technical depth and operational discipline.
  • Provide thought leadership in SRE methodologies, cloud-native operational patterns, and automated reliability engineering.
  • Bachelor’s degree in Computer Science, Data Science, AI, Software Engineering, or related field.
  • Minimum of 7 years of professional experience in AI and data.
  • Strong knowledge of AI models, data architectures, and analytics methodologies.

Nice To Haves

  • Experience managing AWS services and virtual infrastructure.
  • Experience enabling large-scale SRE transformations or modernization initiatives.
  • Familiarity with chaos engineering, resilience assessments, and service failure modeling.
  • Exposure to hybrid-cloud and multi-cloud operational frameworks.

Responsibilities

  • Cloud Operational Support and Modernization
  • Incident & Problem Management Leadership
  • Reliability Engineering & Automation
  • Observability & Operational Excellence
  • Cross-Functional Leadership & Influence
  • Standardization & Documentation
  • Mentorship & Technical Development

Benefits

  • medical
  • dental
  • vision
  • life insurance
  • disability
  • accidental death and dismemberment
  • tax-preferred savings accounts
  • 401k plan
  • vacation
  • sick days
  • paid holidays
  • defined benefit pension plan
  • restricted stock units
  • deferred compensation plan
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service