Sunset turns sensitive internal enterprise data into de-identified datasets without destroying the structure and meaning that make the data valuable. The data does not arrive in one clean modality. It spans messages, documents, tables, files, images, metadata, and provider-specific structures, with important context distributed across all of them. You will improve how well our system understands and protects that data. Your initial scope will be a prioritized subset of named-entity recognition, entity and identity resolution, structured extraction, classification, semantic review, or other model-backed parts of the de-identification pipeline. We do not expect one person to be an expert in every modality. The goal is measurable improvement in the areas you own: better precision, recall, F1, high-risk coverage, and preserved data utility across the failure modes that matter. This is an applied, production-facing ML role. You will study errors, form hypotheses, build datasets and experiments, improve or replace models, and ship the result into a live pipeline. Evaluation, reproducibility, observability, and safe releases matter because they let us identify, ship, and verify meaningful model improvements in production.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed