DOI QR코드

DOI QR Code

대규모 언어 모델(LLM)의 포괄적 성능 비교 평가를 위한 평가 지표 및 데이터셋 개발: 폐쇄형 LLM과 공개형 LLM의 비교를 중심으로

Proposing Benchmark Datasets for Comprehensive Evaluation and Comparison of LLMs (Large Language Models): Comparing Open-source LLMs with Closed-source LLMs

  • Seungho Jeong (MIS, College of Business Administration, Seoul National University) ;
  • Dohun Kim (MIS, College of Business Administration, Seoul National University) ;
  • Jinsoo Park (MIS, College of Business Administration, Seoul National University)
  • 투고 : 2024.06.05
  • 심사 : 2024.06.20
  • 발행 : 2024.08.31

초록

2020년 OpenAI가 1,750억 파라미터 규모의 GPT-3를 공개한 이후 간단한 작업부터 복잡한 작업에 이르기까지 다양한 다운스트림 작업에 대응하는 대규모 언어 모델(LLM)의 개발이 가속화되고 있다. LLM이 개발되고 고도화됨에 따라 LLM의 성능을 객관적으로 평가할 수 있는 평가 지표와 데이터셋이 개발되어 활용되고 있다. 이러한 데이터셋은 다양한 분야에 대해 LLM을 객관적으로 평가함에 있어 좋은 성과를 거두었으나, 규모 측면에서 개인이나 소규모 기관에서 활용하기 어렵고 실용적 측면에서 실제 사용자가 체감하는 바와 다소의 괴리를 가지고 있다. 이에 본 연구에서는 사용자의 활용 패턴을 반영하여 비교적 작은 양의 데이터를 활용해 LLM을 평가할 수 있는 평가 지표 및 데이터셋을 제시한다. 더 나아가, 가중치를 일반에 공개하는 공개형 대규모 언어 모델의 개발이 가속화되고 고성능의 공개형 LLM이 출시되고 있음에 따라 연구 수행 시점인 2024년 4월 기준 최신의 폐쇄형 LLM 4종과 공개형 LLM 6종에 대한 평가를 시행하고 폐쇄형 LLM과 공개형 LLM의 비교 평가 결과에 대해 논의한다. 연구 결과 새롭게 개발한 데이터셋이 작은 규모에도 불구하고 기존 데이터셋과 유사한 경향성을 보이는 것으로 나타났다. 상식 추론 및 글 스타일 변환과 같은 간단한 작업에서는 공개형 LLM이 폐쇄형 LLM과 대등하거나 우세한 성능을 보였으나 수학, 코딩, 이미지 질의응답 등의 복잡한 작업에서는 큰 성능 격차를 보임을 확인하였으며, 더 나아가 비교적 작은 규모의 LLM이 규모 대비 좋은 성능을 보임을 확인하였다.

The development of large language models (LLMs) has accelerated since OpenAI released GPT-3, which demonstrated generalizability and capability for various downstream tasks, thanks to its 175 billion parameters. Various metrics and datasets for LLM evaluation have been developed to objectively assess LLMs' performance. Although existing evaluation metrics and datasets have widely been used across various fields, their large scale hinders their use in small organizations or by individuals. Furthermore, there is degree of discrepancy between evaluation results and actual user experiences. The study proposes evaluation metrics and datasets with relatively small amounts of data while reflecting real-world user experiences. In the process of testing the proposed metrics and datasets, the research evaluates and compares four closed-LLMs and six open-LLMs, which are latest as of April 2024. The results show that proposing datasets exhibited trends similar to existing datasets despite its smaller size, and furthermore, well reflected actual user experiences. Moreover, open-LLMs performed similar, or indeed, better than closed-LLMs in simple tasks while closed-LLMs performed significantly better in complex tasks such as mathematics, coding, and vision question-answering.

키워드

참고문헌

  1. Achiam, J., S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. Aleman, D. Almeida, J. Altenschmidt, S. Altman, and S. Anadkat, “GPT-4 technical report”, arXiv preprint arXiv:2303.08774, 2023.
  2. Ahuja, K., H. Diddee, R. Hada, M. Ochieng, K. Ramesh, P. Jain, A. Nambi, and T. Ganu, “MEGA: Multilingual evaluation of generative AI”, arXiv preprint arXiv:2303.12528, 2023.
  3. Anthropic, “Introducing the next generation of Claude”, 2024, Available at https://www.anthropic.com/news/claude-3-family.
  4. Austin, J., A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, and C. Cai, “Program synthesis with large language models”, arXiv preprint arXiv:2108.07732, 2021.
  5. Azerbayev, Z., H. Schoelkopf, K. Paster, M. Santos, S. McAleer, A. Jiang, J. Deng, S. Biderman, and S. Welleck, “Llemma: An open language model for mathematics”, arXiv preprint arXiv: 2310.10631, 2023.
  6. Bhardwaj, R. and S. Poria, “Red-teaming large language models using chain of utterances for safety-alignment”, arXiv preprint arXiv:2308.09662, 2023.
  7. Bommasani, R., D. Hudson, E. Adeli, and R. Altman, “On the opportunities and risks of foundation models”, arXiv preprint arXiv:2108.07258, 2021.
  8. Bommasani, R., K. Klyman, S. Longpre, B. Xiong, S. Kapoor, N. Maslej, A. Narayanan, and P. Liang, “Foundation model transparency reports”, arXiv preprint arXiv:2402.16268, 2024.
  9. Brown, T., B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, and A. Askell, “Language models are few-shot learners”, Advances in Neural Information Processing Systems, Vol.33, 2020, pp. 1877-1901.
  10. Carolan, K., L. Fennelly, and A. Smeaton, “A review of multi-modal large language and vision models”, arXiv preprint arXiv:2404.01322, 2024.
  11. Chang, Y., X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, and Y. Wang, “A survey on evaluation of large language models”, ACM Transactions on Intelligent Systems and Technology, 2023. https://doi.org/10.1145/3641289
  12. Chen, M., J. Tworek, H. Jun, Q. Yuan, H. Pinto, J. Kaplan, H. Edwards, and Y. Burda, “Evaluating large language models trained on code”, arXiv preprint arXiv:2107.03374, 2021.
  13. Chowdhery, A., S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. Chung, C. Sutton, and S. Gehrmann, “PaLM: Scaling language modeling with pathways”, Journal of Machine Learning Research, Vol.24, No.240, 2023, pp. 1-113.
  14. Clark, J., E. Choi, M. Collins, D. Garrette, T. Kwiatkowski, V. Nikolaev, and J. Palomaki, “TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages”, Transactions of the Association for Computational Linguistics, Vol.8, 2020, pp. 454-470. https://doi.org/10.1162/tacl_a_00317
  15. Cobbe, K., V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, and M. Plappert, “Training verifiers to solve math word problems”, arXiv preprint arXiv:2110.14168, 2021.
  16. Costa-jussà, M., J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, and J. Maillard, “No language left behind: Scaling human-centered machine translation”, arXiv preprint arXiv:2207.04672, 2022.
  17. Databricks, “DBRX”, 2024, Available at https://github.com/databricks/dbrx.
  18. Devlin, J., M. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding”, arXiv preprint arXiv:1810.04805, 2018
  19. Dongqi, P. and V. Demberg, “ChatGPT vs Human-authored text: Insights into controllable text summarization and sentence style transfer”, arXiv preprint arXiv:2306.07799, 2023.
  20. Fu, C., P. Chen, Y. Shen, Y. Qin, M. Zhang, X. Lin, J. Yang, X. Zheng, K. Li, X. Sun, Y. Wu, and R. Ji, “MME: A comprehensive evaluation benchmark for multimodal large language models”, arXiv preprint arXiv:2306.13394. 2024.
  21. Gehman, S., S. Gururangan, M. Sap, Y. Choi, N. Smith, “RealToxicityPrompts: Evaluating neural toxic degeneration in language models”, In Findings of the Association for Computational Linguistics: EMNLP 2020, 2020, pp. 3356-3369.
  22. Gowda, T., Z. Zhang, C. Mattmann, and J. May, “Many-to-English machine translation tools, data, and pretrained models”, In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, 2021, pp 306-316.
  23. Hendrycks, D., C. Burns, S. Basart, A. Critch, J. Li, D. Song, and J. Steinhardt, “Aligning AI with shared human values”, arXiv preprint arXiv:2008.02275, 2020a.
  24. Hendrycks, D., C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding”, arXiv preprint arXiv:2009.03300, 2020b.
  25. Hendrycks, D., C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt, “Measuring mathematical problem solving with the MATH dataset”, arXiv preprint arXiv:2103.03874, 2021.
  26. Huang, J. and K. Chang, “Towards reasoning in large language models: A survey”, In 61st Annual Meeting of the Association for Computational Linguistics, ACL 2023, 2023, pp. 1049-1065.
  27. Iyer, S., I. Konstas, A. Cheung, and L. Zettlemoyer, “Summarizing source code using a neural attention model”, In 54th Annual Meeting of the Association for Computational Linguistics, 2016, pp. 2073-2083, Association for Computational Linguistics, August 2016.
  28. Jiang, A., A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. Chaplot, D. Casas, E. Hanna, and F. Bressand, “Mixtral of experts”, arXiv preprint arXiv:2401.04088, 2024.
  29. Kaplan, J., S. McCandlish, T. Henighan, T. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models”, arXiv preprint arXiv:2001.08361, 2021.
  30. Kasneci, E., K. Seßler, S. Küchemann, M. Bannert, D. Dementieva, F. Fischer, U. Gasser, and G. Groh, “ChatGPT for good? On opportunities and challenges of large language models for education”, Learning and Individual Differences, Vol.103, 2023, p. 102274. https://doi.org/10.1016/j.lindif.2023.102274
  31. Kim, D., C. Park, S. Kim, W. Lee, W. Song, Y. Kim, H. Kim, Y. Kim, H. Lee, and J. Kim, “SOLAR 10.7B: Scaling large language models with simple yet effective depth up-scaling”, arXiv preprint arXiv:2312.15166, 2023.
  32. Kim, S. and U. Wong, “ChatGPT Impacts on Academia”, In 2023 International Conference on System Science and Engineering (ICSSE), July 2023, pp. 422-426.
  33. Koubâa, A., W. Boulila, L. Ghouti, A. Alzahem, and S. Latif, “Exploring chatgpt capabilities and limitations: A survey”, IEEE Access, 2023.
  34. Lee, K., T. Hong, H. Ahn, T. Kim, and C. Koo, “Special topic: The impact of chatgpt in society, business, and academia”, Asia Pacific Journal of Information Systems, Vol.33, No.4, 2023, pp. 957-976. https://doi.org/10.14329/apjis.2023.33.4.957
  35. Li, B., R. Wang, G. Wang, Y. Ge, Y. Ge, and Y. Shan, “SEED-Bench: Benchmarking multimodal llms with generative comprehension”, arXiv preprint arXiv:2307.16125, 2023.
  36. Li, H., Y. Zhang, F. Koto, Y. Yang, H. Zhao, Y. Gong, N. Duan, and T. Baldwin, “CMMLU: Measuring massive multitask language understanding in Chinese”, arXiv preprint arXiv:2306.09212, 2023.
  37. Liu, J., C. Xia, Y. Wang, and L. Zhang, “Is your code generated by ChatGPT really correct? Rigorous evaluation of large language models for code generation”, Advances in Neural Information Processing Systems, Vol.36, 2024.
  38. Liu, N., K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang, “Lost in the middle: How language models use long contexts”, Transactions of the Association for Computational Linguistics, Vol.12, 2024, pp. 157-173. https://doi.org/10.1162/tacl_a_00638
  39. Lu, P., H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K. Chang, M. Galley, and J. Gao, “MathVista: Evaluating mathematical reasoning of foundation models in visual contexts”, arXiv preprint arXiv:2310.02255, 2023.
  40. McIntosh, T., T. Susnjak, T. Liu, P. Watters, and M. Halgamuge, “Inadequacies of large language model benchmarks in the era of generative artificial intelligence”, arXiv preprint arXiv:2402.09880, 2024.
  41. Mesnard, T., C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, and L. Sifre, “Gemma: Open models based on gemini research and technology” arXiv preprint arXiv:2403.08295, 2024.
  42. Meta-llama, "Llama 3", 2024, Available at https://github.cor/meta-llama/llama3.
  43. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, and A. Ray, “Training language models to follow instructions with human Feedback”, Advances in Neural Information Processing Systems, Vol.35, 2022, pp. 27730-27744.
  44. Peng, B., C. Li, P. He, M. Galley, and J. Gao, “Instruction tuning with GPT-4”, arXiv preprint arXiv:2304.03277, 2024.
  45. Reid, M., N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J. Alayrac, and R. Soricut, “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context”, arXiv preprint arXiv:2403.05530, 2024.
  46. Roziere, B., J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. Tan, Y. Adi, and J. Liu, “Code llama: Open foundation models for code”, arXiv preprintarXiv:2308.12950, 2023.
  47. Sakaguchi, K., R. Bras, C. Bhagavatula, and Y. Choi, “WinoGrande: An adversarial winograd schema challenge at scale”, Communications of the ACM, Vol.64, No.9, 2021, pp. 99-106. https://doi.org/10.1145/3474381
  48. Schellaert, W., R. Hamon, F. Martinez-Plumed, and J. Hernandez-Orallo, “A proposal for scaling the scaling laws”, In Proceedings of the First edition of the Workshop on the Scaling Behavior of Large Language Models (SCALE-LLM 2024), March 2024, pp. 1-8.
  49. Son, G., H. Lee, S. Kim, S. Kim, N. Muennighoff, T. Choi, C. Park, K. Yoo, and S. Biderman, “KMMLU: Measuring massive multitask language understanding in Korean”, arXiv preprint arXiv:2402.11548, 2024.
  50. Srivastava, A., A. Rastogi, A. Rao, A. Shoeb, A. Abid, A. Fisch, A. Brown, A. Santoro, and A. Gupta, “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models”, arXiv preprint arXiv:2206.04615, 2022.
  51. Taori, R., I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. Hashimoto, “Alpaca: A strong, replicable instruction-following model”, Stanford Center for Research on Foundation Models, Vol.3, No.6, p. 7, 2023.
  52. Touvron, H., T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, and B. Roziere, “LLaMA: Open and efficient foundation language models”, arXiv preprint arXiv:2302.13971, 2023.
  53. Wang, Y., Y. Pan, M. Yan, Z. Su, and T. H. Luan, “A survey on ChatGPT: AI-generated contents, challenges, and solutions”, IEEE Open Journal of the Computer Society, 2023. https://doi.org/10.1109/OJCS.2023.3300321
  54. Xu, B., T. Li, J. Zheng, M. Naseriparsa, Z. Zhao, H. Lin, and F. Xia, “MET-Meme: A multimodal meme dataset rich in metaphors”, In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, July 2022, pp. 2887-2899.
  55. Young, A., B. Chen, C. Li, C. Huang, G. Zhang, G. Zhang, H. Li, and J. Zhu, “Yi: Open foundation models by 01.AI”, arXiv preprint arXiv:2403.04652, 2024.
  56. Yuan, Z., H. Yuan, C. Tan, W. Wang, and S. Huang, “How well do large language models perform in arithmetic tasks?”, arXiv preprint arXiv:2304.02015, 2023.
  57. Yue, X., Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, and Y. Sun, “MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI”, arXiv preprint arXiv: 2311.16502, 2023.
  58. Zellers, R., A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi, “HellaSwag: Can a machine really finish your sentence?”, arXiv preprint arXiv:1905.07830, 2019.
  59. Zhang, W., S. Aljunied, C. Gao, Y. Chia, and L. Bing, 'M3Exam: A multilingual, multimodal, multilevel benchmark for examining large lan-guage models') Advances in Neural Information Processing Systems, Vol.36, 2024.
  60. Zhao, W. X., K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, Y. Du, and C. Yang, “A survey of large language models”, arXiv preprint arXiv:2303.18223, 2023.
  61. Zhuang, Y., Y. Yu, K. Wang, H. Sun, and C. Zhang, 'ToolQA: A dataset for LLM question answering with external tools', Advances in Neural Information Processing Systems, Vol.36, 2024.