ResearchPod Summary
CzechDocs is a new multiway parallel dataset designed to address the persistent challenge of preserving document structure and markup tags during machine translation. As translation workflows increasingly involve complex formats like HTML, DOCX, and PDF, maintaining the integrity of these elements is critical for downstream publication and accessibility. The dataset focuses on civic integration materials for foreign residents in the Czech Republic, covering languages such as Ukrainian, Vietnamese, English, and Russian.
The dataset consists of 316 parallel document mutations derived from 77 unique sources, including government portals, educational websites, and public service information. The authors employed a rigorous preprocessing pipeline, including manual segment-level alignment for HTML files, to ensure that language versions share consistent structure. With over 60,000 segments and a high density of markup tags (averaging 2.1 tags per segment), the corpus is specifically tailored for evaluating how well translation systems handle non-textual document elements.
The authors compare two primary strategies for format-preserving translation:
Results indicate that while LLMs are capable of handling markup directly, especially when prompted explicitly, the detag-and-project method remains a robust alternative that often yields higher scores on tagged text. The authors provide a validation split for immediate research use and have reserved a larger test split for a future shared task.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.