Java Senior Software Engineer —Vice President

CitiIrving, TX
$125,760 - $188,640Hybrid

About The Position

We are seeking a highly skilled and results-driven Senior Software Engineer with deep expertise in observability, AIOps, and Java/Python application engineering to join our team and play a pivotal role in advancing the firm's Operating Model AI (OMAI) strategy. This role is central to our transition from traditional operational models to a predictive, efficient, and scalable AI-driven framework. The successful candidate will design, build, and ship production-grade applications, standardize telemetry across the estate using Open Telemetry (OTel), harness AIOps platforms to intelligently correlate and remediate operational events, and integrate AI/ML capabilities into observability pipelines delivering a unified, intelligent view of our systems. We are looking for an active, hands-on software engineer who thrives not only in designing, but in implementing the solution as well, takes pride in shipping production-grade code, and brings an energizing, collaborative presence to everything they do.

Requirements

  • 6+ years of hands-on software engineering experience
  • Proven, hands-on Java development skills designing and delivering robust production applications using Spring Boot or equivalent frameworks, independently and at pace
  • Practical Python experience for application development, automation, and tooling
  • Demonstrated expertise in OpenTelemetry (OTel) — instrumenting Java and Python applications and standardizing metrics, logs, and distributed traces across services
  • Exposure to tools like BigPanda for alert correlation, event management, and automated remediation workflows
  • Proficiency with Google Cloud Observability — Cloud Monitoring, Cloud Logging, Cloud Trace, and related tooling — for cloud-native monitoring and performance analysis
  • Strong understanding of SRE principles — including SLIs, SLOs, error budgets, and incident management
  • Solid grasp of RESTful API design and hands-on experience integrating monitoring and automation tools with backend services
  • Strong version control and CI/CD experience — Git, pipeline tooling, and modern release practices
  • Familiarity with containerization and orchestration technologies — Docker and Kubernetes — in cloud-native environments
  • Excellent problem-solving skills — able to analyze complex issues, identify root causes, and implement effective, lasting solutions

Nice To Haves

  • Experience integrating AI/ML or agentic capabilities into operational workflows to enable predictive and self-healing systems (highly desirable)

Responsibilities

  • Design, develop, and deploy Java and Python applications across microservices and distributed architectures, with observability engineered in from the start
  • Standardize telemetry (metrics, logs, and traces) across applications using OpenTelemetry (OTel) to build a consistent, AI-ready data foundation
  • Advance AIOps capabilities by leveraging BigPanda to correlate operational alerts, suppress noise, and build automated remediation workflows
  • Harness Google Cloud Observability tooling — Cloud Monitoring, Cloud Logging, Cloud Trace — for advanced cloud-native monitoring and performance analysis
  • Integrate AI/ML into observability and automation pipelines to enable predictive failure detection and self-healing systems
  • Analyze operational data to identify high-value automation opportunities that maximize reliability, efficiency, and engineering velocity
  • Define and own SLIs, SLOs, and error budgets in partnership with reliability and product engineering teams
  • Partner with AI/ML, platform, and architecture teams to co-innovate and deploy scalable automation solutions aligned to the OMAI strategy
  • Act as Subject Matter Expert (SME) for observability and AIOps — providing hands-on technical guidance and mentorship across the team
  • Participate in all SDLC phases — analysis, design, construction, testing, deployment, and maintenance
  • Proactively troubleshoot and resolve complex application and infrastructure issues, driving long-term stability through collaborative root cause analysis
  • Stay ahead of the curve on emerging practices in observability, AIOps, cloud engineering, and agentic AI

Benefits

  • medical, dental & vision coverage
  • 401(k)
  • life, accident, and disability insurance
  • wellness programs
  • paid time off packages, including planned time off (vacation), unplanned time off (sick leave), and paid holidays
  • Access to Citi's global learning and development resources, including professional certifications in cloud engineering, SRE, and AI/ML disciplines.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service