Cloud Platform Senior Consultant - Remote

AllstateUSA - TX (Remote), TX
$90,700 - $153,925Remote

About The Position

Arity is part of Allstate Corporation, which means we share the same innovative drive that keeps us a step ahead of our customers' evolving needs. We collect and analyze enormous amounts of data to provide cutting-edge solutions for companies invested in transportation. Our engineers are fueled by a passion to impact the future of mobility. As part of an Agile team, they have the freedom to innovate and the opportunity to see projects through from start to finish. Using a variety of languages and a modern technology stack, our engineers advance sensor technology, enterprise engineering, and platform development. Our team collaborates across a global engineering organization while maintaining trust, transparency, and empathy for the end user. The Operational Data Management (ODM) team within Engineering owns the reliability, performance, and scalability of Arity's data infrastructure. We operate mission-critical database and streaming platforms—including PostgreSQL, Redis/Valkey, Amazon Redshift, Google BigQuery, Amazon MSK (Kafka), and Google Pub/Sub—as well as analytics layers such as Starburst Galaxy and AWS Athena. We partner closely with application development teams to tune, troubleshoot, and optimize applications that depend on these technologies, ensuring the data platforms powering Arity's mobility insights remain highly available and performant at scale. Arity is seeking a Cloud Platform Senior Consultant to join ODM. This is a fully remote position. You will design, build, deploy, and operate cloud-native data infrastructure across AWS and Google Cloud Platform, with hands-on work across databases, data streaming, and distributed systems. You will help ensure the platforms that ingest, store, and serve billions of miles of driving data remain resilient, observable, and cost-efficient—directly enabling Arity's products and the customers who rely on them. The ideal candidate combines solid cloud platform skills with strong database and streaming fundamentals, production-grade Python experience, and a collaborative, ownership-minded approach to production support and performance improvement with application teams. A willingness to learn new technologies quickly and deliver practical solutions in production is highly valued.

Requirements

  • 3–5 years of software engineering or infrastructure experience, with at least 2 years in SRE, DevOps, or platform engineering operating production systems at scale.
  • Hands-on experience designing, deploying, and managing cloud infrastructure on AWS and/or Google Cloud Platform, including networking, identity, and security fundamentals.
  • Production experience operating data streaming platforms; hands-on work with Apache Kafka (including Amazon MSK or Confluent Kafka) and a solid understanding of partitions, consumer groups, delivery semantics, and backpressure.
  • Production experience with relational and NoSQL databases; PostgreSQL required, plus familiarity with distributed data stores.
  • Strong Python and Shell scripting for automation, custom tooling, and operational solutions that go beyond basic scripts.
  • Strong experience with infrastructure-as-code (e.g., Terraform), CI/CD (e.g., Jenkins, Git), Ansible, and container orchestration (e.g., Kubernetes) in production environments.
  • Experience implementing and automating monitoring, logging, and alerting for distributed systems (e.g., Prometheus, Grafana, CloudWatch, Datadog, or equivalent), including automated runbooks where applicable.
  • Proven ability to contribute to root cause analysis for production incidents spanning infrastructure, databases, streaming pipelines, and application code layers.
  • Strong problem-solving, communication, and documentation skills with a track record of ownership in on-call and incident management environments.
  • Working understanding of distributed systems principles including high availability, fault tolerance, consistency models, and disaster recovery.

Nice To Haves

  • Production experience with Apache Cassandra, including cluster operations, performance tuning, and troubleshooting in distributed environments.
  • Knowledge of DynamoDB, ElastiCache, Apache NiFi, or self-managed Apache Flink, including checkpoint management, flow design, and streaming job troubleshooting.
  • Advanced experience with Apache Kafka, Google Pub/Sub, and operating streaming workloads across both AWS and GCP.
  • Experience administering or optimizing Starburst Galaxy, Trino, or AWS Athena for large-scale analytics workloads.
  • Experience building AI agents, Model Context Protocol (MCP) servers, or LLM-based tooling to automate DevOps, observability, or operational workflows.
  • Familiarity with Java/Spring Boot application troubleshooting, JVM diagnostics (heap dumps, GC tuning, thread dumps, connection pool analysis), or reading Golang production code for performance analysis.
  • Experience with data pipeline orchestration (e.g., Apache Airflow, dbt), event-driven architectures, or enterprise PaaS platforms (e.g., Cloud Foundry).
  • AWS or Google Cloud professional-level certifications, performance benchmarking, query plan analysis, database capacity planning, APM/distributed tracing, or open-source contributions in database, streaming, or infrastructure projects.

Responsibilities

  • Build, operate, and optimize data streaming infrastructure using Amazon MSK (Kafka) or Google Pub/Sub to support real-time and batch data pipelines.
  • Design, deploy, and manage highly available database and caching platforms—including PostgreSQL, Redis, Valkey, Amazon Redshift, and Google BigQuery—across multi-cloud environments.
  • Develop and maintain infrastructure-as-code, CI/CD pipelines, and cloud automation using Terraform, Python and industry-standard tooling to enable repeatable, secure deployments.
  • Implement monitoring, alerting, and observability for data platform services to proactively detect and resolve issues before they impact customers.
  • Partner with application development teams to troubleshoot, tune, and optimize application performance, query patterns, and data access layers backed by team-managed platforms.
  • Administer and optimize analytics and query engines including Starburst Galaxy and AWS Athena to deliver performant, cost-effective access to large-scale datasets.
  • Participate in incident response, root cause analysis, and post-incident reviews for production database and streaming systems; drive remediation and preventive improvements.
  • Share on-call rotation to provide support for mission-critical data infrastructure.
  • Review application source code, identify performance or reliability issues, and collaborate on targeted fixes or optimization guidance with development teams.
  • Evaluate emerging tools and automation approaches—including AI-assisted workflows—to improve operational efficiency and developer experience.
  • Contribute to capacity planning, disaster recovery, security hardening, and cost optimization initiatives across the data platform estate.

Benefits

  • Allstate provides a comprehensive technology setup, including a laptop, monitors, headset, keyboard, and mouse.
  • Employees eligible to work from home also receive a monthly connectivity reimbursement to help offset internet costs.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service