Helios is building a new kind of company to solve America’s hardest problems, starting with the government interaction layer. Government shapes every consequential market, but the infrastructure connecting public institutions and private organizations remains fragmented, manual, and difficult to navigate. Helios is rebuilding that layer. Our core platform, Proxi, gives organizations the intelligence they need to understand what government is doing, why it matters, and what to do next. From that foundation, we design and deploy secure, mission-specific systems for government agencies, enterprises, and institutions operating in complex and highly regulated environments. We bring together frontier AI, deep public-sector expertise, and forward-deployed execution. Our team includes leaders and builders from the White House, U.S. Department of State, Datadog, and Microsoft. We are backed by leading institutional investors and trusted by organizations working on high-stakes problems across government and industry. MISSION Build and operate the acquisition and processing infrastructure that turns global public-sector and open-source data into reliable, searchable intelligence. This role owns the path from source discovery and collection through normalization, extraction, enrichment, and publication. You will expand our platform to include new countries, languages, institutions, and source formats covering many additional areas of open source intelligence to advance Proxi’s natural capabilities. We are looking for a senior engineer who has built and operated large-scale acquisition or document-processing systems in production. Expected areas of expertise: Web crawlers and connectors for continuously changing international sources. Fault-tolerant pipelines with reliable scheduling, recovery, replay, and backfill capabilities. Processing complex documents, structured files, images, audio, and video across languages. Stable schemas, source lineage, correction propagation. Operating secure, observable, scalable, and cost-efficient processing systems across cloud and restricted environments. Document conversion, OCR, layout analysis, and structural extraction across complex file formats. Multilingual audio, video, and document processing with reliable alignment and attribution. Scalable inference infrastructure for batch and real-time processing across CPU and GPU workloads. End-to-end provenance, versioning, correction propagation, and reproducible reprocessing. Secure handling of untrusted content with rigorous quality evaluation, observability, and cost controls.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed