DOI QR코드

DOI QR Code

시맨틱 추론 규칙을 이용한 대규모 언어 자원의 품질 고도화 방안 - 국립중앙도서관 주제명표목표를 중심으로 -

Method for Enhancing the Quality of Large-Scale Lexical Resources Using Semantic Inference Rules: Focusing on the National Library of Korea Subject Headings

  • 정도헌 (덕성여자대학교 글로벌융합대학 문헌정보학전공)
  • 투고 : 2024.11.23
  • 심사 : 2024.12.16
  • 발행 : 2024.12.30

초록

본 연구는 대규모 언어 자원에 관한 연구 동향을 바탕으로, 자동화된 시맨틱 추론 기법을 이용한 거대 언어 자원의 효율적 품질 제고 방안을 제시하고 응용 가능성을 제안하고자 한다. 이를 위해, 우선 언어 자원 내 용어 간의 다양한 관계를 분석하여 도출한 정오 사례를 바탕으로 공통의 시맨틱 추론 규칙을 정의하였다. 정의된 추론 규칙을 바탕으로 거대한 언어 자원의 네트워크를 고속 탐색하고 오류를 검출하는 스프레딩 알고리즘 기반의 시맨틱 추론 엔진을 개발하였다. 국립중앙도서관의 주제명표목표에 대한 실험을 통해, 본 연구에서 제안한 자동화 기법과 웹 기반 관리 시스템을 활용하여 대용량 데이터의 품질 고도화 작업을 효율적으로 수행할 수 있음을 확인하였다. 본 연구는 대규모 언어 자원의 품질 고도화를 위한 시맨틱 추론 기법을 새롭게 제안한 점, 복잡한 용어 네트워크에서 발생하는 논리적 오류 사례를 분석하고 체계화하는 최초의 시도였다는 점에서 의의가 있으며, 일련의 과정을 통해 인공 지능 시대의 인간과 기계의 협업 방식을 논의하였다는 데 의의가 있다.

The purpose of this study is to propose an efficient quality enhancement method for large-scale lexical resources using automated semantic inference techniques, based on current research trends in large lexical resources, and to suggest practical applications. To achieve this, common semantic inference rules were first defined by analyzing various relational cases among terms within lexical resources and identifying correct and erroneous patterns. Using these defined inference rules, a semantic inference engine based on a spreading algorithm was developed, enabling rapid network traversal and error detection across very large lexical resources. Through experiments on the Subject Headings of the National Library of Korea, it was confirmed that the automated methods and web-based management system proposed in this study enable effective quality enhancement of large-scale data. The study is significant in that it proposes a novel semantic inference approach for enhancing the quality of large-scale lexical resources, as well as the first attempt to analyze and organize logical error cases arising within complex term networks. Furthermore, it is meaningful in discussing methods of human-machine collaboration in the era of artificial intelligence.

키워드

과제정보

본 연구는 2023년도 덕성여자대학교 교내연구비 지원에 의해 이루어졌음(3000008144).

참고문헌

  1. Baek, Ji-Won & Chung, Yeon-Kyoung (2014). A study on improving access & retrieval system of the National Library of Korea subject headings. Journal of the Korean Society for Information Management, 31(1), 31-51. https://doi.org/10.3743/KOSIM.2014.31.1.031
  2. Choi, Woon Kyung & Chung, Yeon-Kyoung (2014). A study on improvements for high quality in National Library of Korea subject headings List. Journal of the Korean Society for Library and Information Science, 48(1), 75-95. https://doi.org/10.4275/KSLIS.2014.48.1.075
  3. Han, Hui-Jeong, Kim, Tae-Young, Doo, Hyo-Chul, & Oh, Hyo-Jung (2017). Automatic extraction and utilization of technical term dictionaries using definition patterns in technical documents. Journal of the Korean Society for Information Management, 34(4), 81-99. https://doi.org/10.3743/KOSIM.2017.34.4.081
  4. Heo, Go Eun (2019). Network analysis between uncertainty words based on Word2Vec and WordNet. Journal of the Korean Society for Library and Information Science, 53(3), 247-271. https://doi.org/10.4275/KSLIS.2019.53.3.247
  5. Jeong, Do-Heon (2018). Generating and controlling an interlinking network of technical terms to enhance data utilization. Journal of the Korean Society for Information Management, 35(1), 157-182. https://doi.org/10.3743/KOSIM.2018.35.1.157
  6. Jeong, Do-Heon & Choi, Hee-Yoon (2006). Building and analysis of semantic network on S&T multilingual terminology. Journal of Information Management. 37(4), 25-47. https://doi.org/10.1633/JIM.2006.37.4.025
  7. Kim, Sundong, Kang, Minseo, & Lee, Jae-Gil (2014). A method of automatic schema evolution on DBpedia Korea. Proceedings of the Korea Information Processing Society Conference, 21(1), 741-744.
  8. Lee, HyeKyung & Lee, Yong-Gu (2023). A study on the current status of National Library of Korea subject headings list through utilization analysis of subject headings. Journal of the Korean Society for Information Management, 40(2), 157-182. https://doi.org/10.3743/KOSIM.2023.40.2.157
  9. National Library of Korea (2021). National Library of Korea subject headings guidelines. National Bibliography Department, National Library of Korea. Available: https://librarian.nl.go.kr/LI/contents/L20201000000.do
  10. Yeo, Ji-Suk, Yang, Kiduk, Ito, Hiroko, & Lee, HyeKyung (2022). A study on the enhancement of korean diaspora-related subject headings: focusing on korean-related terminology in the National Library of Korea subject headings. Journal of Korean Library and Information Science Society, 53(1), 103-124. https://doi.org/10.16981/kliss.53.1.202203.103
  11. An, Y., Wang, Q., Liu, J., Liu, K., Lyu, Y., Wu, H., She, Q., & Li, S. (2019). Enhancing pre-trained language representations with rich knowledge for machine reading comprehension. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2346-2357. https://doi.org/10.18653/v1/P19-1226
  12. Hu, X.-B., Wang, M., Hu, D., Leeson, M. S., Hines, E. L., & Di Paolo, E. (2012). A ripple-spreading algorithm for the k shortest paths problem. In Proceedings of the 2012 Third Global Congress on Intelligent Systems, 202-208. https://doi.org/10.1109/GCIS.2012.96
  13. Jeong, D. H., Hwang, M., & Sung, W. K. (2011). Generating knowledge map for acronym-expansion recognition. In the Proceedings on U- and E-Service Science and Technology (UNESST 2011), 287-293. https://doi.org/10.1007/978-3-642-27210-3_38
  14. Marciniak, J. (2020). Wordnet as a backbone of domain and application conceptualizations in systems with multimodal data. Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020), 1-6. Available: https://aclanthology.org/2020.mmw-1.5
  15. Mendes, P. N., Jakob, M., Garcia-Silva, A., & Bizer, C. (2011). DBpedia spotlight: shedding light on the web of documents. Proceedings of the 7th International Conference on Semantic Systems, 1-8. https://doi.org/10.1145/2063518.2063519
  16. Miller, G. A. (1995). WordNet: a lexical database for english. Communications of the ACM, 38(11), 39-41. https://doi.org/10.1145/219717.219748
  17. Navigli, R., & Ponzetto, S. P. (2012). BabelNet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network. Artificial Intelligence, 193, 217-250. https://doi.org/10.1016/j.artint.2012.07.001
  18. Nothman, J., Ringland, N., Radford, W., Murphy, T., & Curran, J. R. (2013). Learning multilingual named entity recognition from Wikipedia. Artificial Intelligence, 194, 151-175. https://doi.org/10.1016/j.artint.2012.03.006
  19. Paulheim, H. (2017). Knowledge graph refinement: A survey of approaches and evaluation methods. Semantic Web, 8(3), 489-508. https://doi.org/10.3233/SW-160218
  20. Paulheim, H., & Gangemi, A. (2015). Serving DBpedia with DOLCE - more than Just adding a cherry on top. The Semantic Web - ISWC 2015 (Lecture Notes in Computer Science, vol. 9366), 180-196. https://doi.org/10.1007/978-3-319-25007-6_11
  21. Saedi, C., Branco, A., Rodrigues, J. A., & Silva, J. (2018). WordNet embeddings. In Proceedings of the Third Workshop on Representation Learning for NLP, 122-131. https://doi.org/10.18653/v1/W18-3016