ResearchPod Summary
Historical newspaper digitization is hindered by the complex, nested, and heterogeneous nature of page layouts. The authors investigate two distinct strategies for automated document understanding: a modular bottom-up pipeline that combines specialized open-source models, and a novel top-down transformer architecture designed to explicitly model the hierarchical structure of newspaper pages.
The study proposes two complementary methods:
The authors also introduce the Finlam La Liberté dataset, a new resource specifically curated for evaluating hierarchical information retrieval in historical newspapers.
The research demonstrates that both the modular pipeline and the Tiramisu architecture are effective at reconstructing complex newspaper hierarchies. The modular approach offers flexibility and interpretability by leveraging established, robust components. In contrast, the Tiramisu architecture provides a unified, end-to-end solution that models the document hierarchy directly, allowing for parallelized processing of sections and articles. This work provides valuable tools for digital humanities and large-scale document digitization, offering scalable methods to convert raw images into structured, machine-readable formats.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.