This role is for one of our clients. The project builds training data for languages that parsing and vision-language models rarely see, specifically Korean, Japanese, and five Indic scripts. Each task involves taking a real, publicly available PDF page and producing a complete structural map of that page, paired with a faithful transcription of every text region in the original script. The dataset focuses on material models handle worst: handwriting, dense multi-column layouts, tables, diagrams, and mixed-script pages. Documents are drawn from various sources like newspapers, textbooks, examinations, flyers, forms, manuals, menus, brochures, notices, and worksheets to reflect the real diversity of Korean documents. All delivered work is human-authored, with component identification, typing, reading order, and transcription performed by people.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Part-time
Career Level
Mid Level
Education Level
No Education Listed