Sr. Data Engineer

AllocatePalo Alto, CA
Hybrid

About The Position

Allocate is seeking a Senior Data Engineer to contribute to the development of its data infrastructure, which supports analytics, reporting, and data-driven product features. As a fintech startup focused on making private market investing more accessible, the company possesses significant financial and investment data. The existing foundational architecture and strategy were established by the data lead, and the company is now looking for a senior engineer to expand, scale, and strengthen it. This role involves close collaboration with the data lead to model core financial entities, integrate internal and external data sources, and construct the pipelines and infrastructure necessary for informed decision-making by engineering and product teams, as well as for shipping product features. This is a hybrid position based in the Palo Alto, CA office, requiring collaboration with the backend (C#/.NET) and frontend (Node/Vue.js) teams to integrate data pipelines into the platform. The ideal candidate is a hands-on engineer eager to perform high-impact data work in a collaborative startup setting.

Requirements

  • 5+ years of hands-on experience in data engineering or related fields, including designing and building large-scale data pipelines and storage solutions.
  • Experience taking projects through the full lifecycle from design to production deployment.
  • Strong experience working with AWS cloud services for data, including S3, EC2, ECS, EKS, Athena, Redshift, Glue, and Step Functions.
  • Proficiency in SQL and relational database design, including designing efficient schemas and optimizing queries/indexes for performance.
  • Fluency in at least one major programming language used in data engineering (Python preferred).
  • Experience preparing datasets for training, working with feature stores, or integrating ML model outputs into applications.
  • Solid understanding of containerization and deployment, including experience using Docker and Kubernetes (or AWS EKS).
  • Comfort setting up CI/CD pipelines for automated testing and deployment of data pipelines or ML models.
  • Strong analytical and problem-solving skills, with attention to detail regarding data correctness and a knack for troubleshooting data discrepancies or bottlenecks.
  • Experience working in a regulated SEC environment, handling sensitive investor/financial data, and building with auditability, least-privilege, and data governance in mind.
  • Excellent communication skills and a collaborative mindset.
  • Comfort mentoring peers and driving technical projects to completion.
  • A positive attitude toward continuous learning and improvement, with a growth mindset and adaptability.
  • Bachelor’s degree in Computer Science, a similar technical field of study, or equivalent practical experience.

Nice To Haves

  • Experience setting up infrastructure-as-code (Terraform/CloudFormation) for AWS services.
  • Experience building or working with data warehouses or lakehouses (e.g., Snowflake, Databricks Delta Lake).
  • Familiarity with graph databases (Neo4j, AWS Neptune, etc.) and knowledge graph schemas.
  • Pandas/PySpark experience.
  • Experience with TypeScript/Node.js in data contexts.
  • Ability to work across languages, such as writing a data API in C# or Node.js while crafting Python scripts for data processing.
  • Knowledge of vector embeddings and experience with vector databases (Postgres pgvector, Chroma, Pinecone, etc.).
  • Familiarity with frameworks for building AI agents or retrieval-augmented generation (e.g., LangChain, LlamaIndex).
  • Experience with workflow managers (Airflow, Prefect, dbt, or similar).
  • Experience with modern web technologies.

Responsibilities

  • Build and Extend Data Architecture: Construct and expand Allocate's data lakehouse on AWS, integrating data lake storage and warehouse technologies for diverse financial datasets. Contribute to the knowledge graph modeling key relationships (investors, funds, companies, etc.) and to the vector database integration for semantic search and retrieval across AI agents, models, and providers.
  • Develop Data Pipelines: Create robust ETL/ELT pipelines for ingesting, cleaning, and transforming data from various sources (internal application data and third-party APIs). Ensure support for both batch processing and real-time data streaming to facilitate up-to-date analytics and recommendations. Build pipelines with a focus on scalability (to handle increasing data volume and complexity) and reliability (with proper error handling and monitoring).
  • Enable AI/ML Capabilities: Collaborate with the data science and engineering teams to provision data and infrastructure for machine learning models and AI features. This includes preparing training datasets, setting up feature stores, and orchestrating workflows that provide LLM-based agents with necessary context (e.g., retrieving relevant data via vector similarity search). Implement systems to serve AI model outputs (like recommendations) back into the product in real time.
  • Engineering Excellence and Collaboration: Partner with the data lead and the broader engineering team to deliver data and AI infrastructure. Enhance quality through code reviews, testing, and adherence to best practices. Assist engineers using data in their services. Work in cross-functional squads to integrate data-driven features into the product roadmap and share expertise with peers as the team grows.
  • Infrastructure and DevOps: Collaborate with DevOps engineers to deploy and maintain data services. Containerize and orchestrate data tools (using Docker/Kubernetes on AWS EKS) for production use. Implement CI/CD pipelines for data workflows to automate testing and deployment of changes to data processing or models. Monitor the health and performance of data platforms (setting up alerts, dashboards) and be prepared to troubleshoot and resolve production issues to ensure uptime of critical data and AI services.
  • Continuous Improvement: Stay current with advancements in data engineering and AI, including new AWS offerings and open-source ML tools. Evaluate and recommend new technologies, such as assessing the potential benefits of stream processing platforms (Kafka/Kinesis) or orchestration tools (Airflow) for pipeline reliability. Encourage rethinking processes and innovating to build a world-class, intelligent platform.

Benefits

  • Medical, dental, and vision.
  • 401(k)
  • Responsible vacation time (PTO)
  • Total compensation may also include a discretionary performance-based bonus.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service