Software Developer/Engineer (Mid Level experience)

TechArmyPhiladelphia, PA
Hybrid

About The Position

We are seeking a skilled Software Developer/Engineer with mid-level experience to focus on on-premise Large Language Model (LLM) and Vector Database implementation. This contract role requires hands-on experience deploying open-source LLMs, proficiency in Python for LLM-related tasks, and expertise in vector databases and Retrieval-Augmented Generation (RAG) pipelines. A strong understanding of data privacy, security, and enterprise requirements is essential. The role involves creating a reference architecture, a working prototype, and providing documentation and knowledge transfer to internal teams.

Requirements

  • Hands-on experience deploying open-source LLMs such as Meta Llama 3 and Mistral / Mixtral in on-prem or private environments.
  • Strong proficiency in Python for LLM inference, prompt engineering, and integration.
  • Experience with CPU-based inference, model quantization, and performance tuning.
  • Practical experience with open-source vector databases such as Qdrant, Chroma, Milvus, or pgvector.
  • Proven implementation of Retrieval-Augmented Generation (RAG) pipelines.
  • Experience generating and managing embeddings and metadata filtering.
  • Understanding of data privacy, air-gapped deployments, and enterprise security requirements.
  • Experience implementing access controls and audit logging.

Nice To Haves

  • Experience with LangChain or LlamaIndex.
  • Exposure to Rust, Go, or C++ for high-performance services.
  • Familiarity with Docker and Kubernetes for on-prem deployments.
  • Knowledge of inference frameworks (e.g., vLLM, llama.cpp, Hugging Face Transformers).
  • Prior work in regulated or enterprise environments.

Responsibilities

  • Deploying open-source LLMs such as Meta Llama 3 and Mistral / Mixtral in on-prem or private environments.
  • Utilizing Python for LLM inference, prompt engineering, and integration.
  • Optimizing LLM performance through CPU-based inference, model quantization, and performance tuning.
  • Working with open-source vector databases such as Qdrant, Chroma, Milvus, or pgvector.
  • Implementing Retrieval-Augmented Generation (RAG) pipelines.
  • Generating and managing embeddings and metadata filtering.
  • Understanding and implementing data privacy, air-gapped deployments, and enterprise security requirements.
  • Implementing access controls and audit logging.
  • Creating a reference architecture and deployment guidance.
  • Developing a working prototype (LLM + vector DB + RAG).
  • Providing documentation and knowledge transfer to internal teams.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service