About The Position

This role is for one of our clients. The project builds training data for languages that parsing and vision-language models rarely see, specifically Korean, Japanese, and five Indic scripts. Each task involves taking a real, publicly available PDF page and producing a complete structural map of that page, paired with a faithful transcription of every text region in the original script. The dataset focuses on material models handle worst: handwriting, dense multi-column layouts, tables, diagrams, and mixed-script pages. Documents are drawn from various sources like newspapers, textbooks, examinations, flyers, forms, manuals, menus, brochures, notices, and worksheets to reflect the real diversity of Korean documents. All delivered work is human-authored, with component identification, typing, reading order, and transcription performed by people.

Requirements

  • Native Korean speaker with full command of Hangul, including hanja where it appears in older or formal documents
  • Worked in bilingual transcription, translation, editorial work, or AI training data, ideally with reviewer experience
  • Exact: character-level accuracy matters more here than speed, and a single wrong jamo is a defect
  • Systematic: apply a taxonomy consistently across hundreds of pages rather than improvising per document
  • Comfortable with unfamiliar layouts: multi-column newspapers, exam papers, handwritten forms

Nice To Haves

  • AI training data: annotation, labeling, grading, or bilingual evaluation for training datasets
  • Transcription and localization: MTPE, subtitling, bilingual QA, OCR correction or post-editing
  • Document production: typesetting, copy-editing, proofreading, or digitization of Korean-language material
  • Script and encoding: Unicode normalization, Korean input methods, Hangul jamo composition, and hanja handling

Responsibilities

  • Open and check a task: pages are provided, so you do not source documents yourself. We find the PDFs and upload them for you. Before annotating, confirm the page is in Korean, is legible, has real content, and shows no personal details
  • Annotate structure: identify and bound every meaningful region of the page - document title, section heading, paragraph, list, table, figure, diagram, caption, formula, question, answer field - and assign each a component type and a reading-order index
  • Record relationships: link each region to the figure or table it belongs to through a parent component identifier
  • Transcribe faithfully: reproduce all text exactly as it appears in Hangul, including any hanja and handwritten content, flagging any region where the source is not legible
  • Capture page metadata: language, document type, source, page dimensions, and flags for tables, formulas and handwriting
  • Review a colleague's work: every task is reviewed end to end by a second Korean expert, and experienced annotators take on that review
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service