نوع مقاله : مقاله پژوهشی
نویسندگان
1 گروه علم اطلاعات و دانششناسی، دانشکده مدیریت، دانشگاه تهران، تهران، ایران.
2 گروه هوش مصنوعی و رباتیک، دانشکده مهندسی کامپیوتر، دانشگاه علم و صنعت ایران، تهران، ایران
کلیدواژهها
عنوان مقاله English
نویسندگان English
Introduction
Information organization is a cornerstone of every information-retrieval system, since query effectiveness depends on how information is organized and represented. Digital archives, especially historical ones such as the Islamic Revolution Document Center (IRDC) of Iran, face a persistent challenge: traditional retrieval systems cannot provide comprehensive, meaning-based access to document content. Established in 1981 to compile the history of the Islamic Revolution and collect the documents and memoirs of Imam Khomeini's movement, the IRDC holds more than 4.5 million written-document pages, about 31,000 hours of oral-history interviews, nearly 372,000 photographs, more than 13 million news titles, and around 90,000 domestic and foreign articles. Although much of this material has been digitized through optical character recognition (OCR) and automatic speech recognition (ASR), the center's retrieval system relies mainly on descriptive metadata and full-text search, so that key entities (persons, places, events, and concepts) remain hidden, limiting comprehensive access. This calls for a redesign based on knowledge-graph and ontology approaches that can extract and organize entities semantically. This research models a domain ontology for the IRDC digital archive through a hybrid approach combining text analysis with the reuse of existing ontologies, establishing a foundation for a knowledge graph that supports semantic retrieval and AI systems for managing Iran's historical documents.
Research Questions
1. What are the entities and classes of the subject domain of the IRDC digital-archive documents, and how can they be represented within a domain ontology?
2. What should the hierarchical structure and the relationships among the ontology classes be?
3. Which ontologies and classification systems are most suitable for reuse in building the IRDC domain ontology?
Literature Review
Research on ontology- and knowledge-graph-based knowledge organization falls into several strands: automatic or hybrid extraction of structured information from text (Emami, 2022; Moradi et al., 2015; Gutierrez et al., 2016; Wimalasuriya & Dou, 2010); knowledge-graph development, including a Persian knowledge graph (Sajjadi & Minaei Bidgoli, 2019) and FarsBase (Asgari-Bidhendi et al., 2019); knowledge-graph-based question answering (Yang et al., 2023; Tong et al., 2019); human-driven ontology design in specialized domains (Torabi et al., 2022; Sadeghi Niaraki et al., 2018; Servati et al., 2017); knowledge-graph-based retrieval and search engines (Homayouni et al., 2018; Li, 2019; Wang et al., 2019); ontology-based retrieval optimization (Jafari Pavarsi et al., 2020; Asim et al., 2019); and thesaurus enrichment for ontology construction (Hosseini Beheshti & Ejei, 2015; Wu, 2018). A review of this body of work shows that, although these studies advance meaning-based organization, most rely on relatively structured or general sources, including Wikipedia, web tables, scientific databases, and technical texts, rather than real archival environments marked by heterogeneous sources, diverse formats, and semantic ambiguity, and they emphasize technical extraction methods over practical application in operational systems. The present study therefore occupies a distinct, applied, problem-driven position, designing and implementing, for the first time, a domain ontology for contemporary Iranian history within a real digital-archive context and addressing the gap in the domestic literature.
Methodology
A mixed-method design was adopted. Informational entities were identified through semi-automatic and automatic text analysis, and the ontology was designed following the best practices of Arp et al. (2015) for building domain ontologies on Basic Formal Ontology (BFO): defining the scope, identifying universals from existing ontologies and domain texts, arranging them in a general-to-specific hierarchy, ensuring logical coherence and human-readable definitions, and formalizing the result in a computer-usable language. The scope was limited to contemporary Iranian history, that is, the second Pahlavi period, the Islamic Revolution and its aftermath, and the Islamic Republic to the present, with fine-grained classes for persons, parties, governance periods, organizations, places, and related events.
The research population comprised all digitized resources of the archive, from which a purposive sample was selected in consultation with the center's managers and experts. Two tools supported entity recognition: the open-source annotation software INCEpTION for manual and semi-automatic tagging, and a large language model (Grok v3.0) for automatic named-entity recognition, whose errors were corrected under human supervision. Existing ontologies and classifications were analyzed with Protégé and SPARQL queries over RDF/OWL files, reusing DBpedia, the Persian knowledge-graph ontology (FarsBase), BIBFRAME, and the Library of Congress classification for the history of Iran. The hierarchy was built around an is_a backbone with single inheritance, the open-world assumption, and the objectivity principle, using DBpedia as the base. Finally, the classes were validated through the Nominal Group Technique with eleven managers and experts across two three-hour sessions, applying four criteria: hierarchical logic, historical and cultural accuracy, alignment and conceptual equivalence, and completeness and simplicity.
Results
The combination of text analysis with the reuse of DBpedia, BIBFRAME, and the Library of Congress Iranian-history classification produced 535 terms, organized under main classes including agent, person, organization, family, event, work, government period, subject concept, place, architectural structure, and ethnic group. After validation by the expert group and majority consensus, 457 classes were confirmed, 70 were added, 53 were relocated within the hierarchy, 19 were removed, 5 terms were renamed, and 2 classes were both relocated and renamed.
Most of the additions concerned the agent–person and place branches, and a dedicated subject-concept branch was created for intellectual, political, social, and economic currents, schools of thought, and ideologies. The architectural-structure branch was merged into the place branch as spatial subdivisions, and several unconventional or overlapping class labels were corrected or removed. Regarding the third question, no domain ontology for contemporary Iranian history was found domestically; the Library of Congress expansion was preferred over the Dewey-based one because it subsumes the Dewey classes, DBpedia was confirmed as the core general ontology (the Persian knowledge graph differing in only ten classes), and BIBFRAME served as the reference for the work-related classes.
Discussion
The study yielded a validated, extensible base ontology for contemporary Iranian history grounded in the IRDC archive. Its results agreed with earlier Persian-language extraction studies and with hybrid ontology-based extraction approaches, but its data derive from authentic, semi-structured historical documents rather than structured or general texts, with emphasis on a native conceptual structure suited to the Iranian historical domain rather than on extraction algorithms. The scale of the expert-driven revision showed that general structures such as DBpedia or the Persian knowledge graph cannot accurately cover native Iranian-history concepts without extensive localization, consistent with Sajjadi and Minaei Bidgoli (2019) and Asgari-Bidhendi et al. (2019). Unlike purely conceptual ontology-design studies, this work was validated through practical feedback from archival experts in a real institutional setting, making its output directly applicable within the digital-archive system.
Conclusion
By combining text analysis with the reuse of established ontologies and rigorous expert validation, this research produced a base domain ontology for contemporary Iranian history within the real context of the IRDC digital archive. The ontology provided a foundation for a historical knowledge graph, enhanced semantic retrieval, and served as a logical basis for artificial-intelligence systems that manage the country's historical documents, filling a gap in practical ontology design for this domain. Future research should develop interoperable domain ontologies in other subject areas, implement and evaluate an operational knowledge graph, advance more automated extraction of concepts and relations with large language models, and study relation-weighting algorithms.
کلیدواژهها English