ResearchPod Summary
This study investigates the feasibility of using multimodal Large Language Models (LLMs) to automate the labor-intensive process of creating library catalogue records. The researchers aimed to determine if AI could process scanned images of historical dissertations to generate complete, structured metadata in standard formats like MARC, JSON, and BIBFRAME. Using a collection of 87 16th and 17th-century European dissertations from the Bodleian Libraries, the authors compared the performance of several OpenAI GPT and Google Gemini models. The study employed a rigorous experimental design, testing 12 prompt variations per model to evaluate how different instructions and input strategies (single vs. multiple images) influenced output quality.
The researchers performed both quantitative and qualitative assessments. Quantitative evaluation relied on three metrics: Jaccard similarity (word overlap), semantic similarity (using vector embeddings), and the BLEU score (N-gram precision). These metrics allowed for the comparison of over 5,000 AI-generated outputs against human-curated catalogue cards. Qualitatively, the team manually inspected the outputs to identify subtle inconsistencies and formatting behaviors that standard metrics might overlook, ensuring that the AI's performance was evaluated beyond mere statistical success.
The experiments revealed that GPT-4.1-mini generally provided the most consistent and accurate results across the tested formats. Interestingly, the study found that using only the title page image (the "single" image approach) performed nearly as well as processing the entire document, significantly reducing computational costs. While the models successfully generated structured data, the qualitative analysis highlighted that AI behavior differs from human cataloguers; the models produced subtle variations in field content and structure that require careful monitoring, even in the absence of overt "hallucinations."
Automating the creation of catalogue records could drastically reduce the time and cost associated with digitizing large library collections. This research provides a practical framework for libraries to evaluate and deploy AI tools for metadata generation, demonstrating that while current models are highly capable, they require specific prompt engineering and quality assurance processes to meet professional archival standards.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.