참고문헌
- Achiam, J., S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. Aleman, D. Almeida, J. Altenschmidt, S. Altman, and S. Anadkat, “GPT-4 technical report”, arXiv preprint arXiv:2303.08774, 2023.
- Ahuja, K., H. Diddee, R. Hada, M. Ochieng, K. Ramesh, P. Jain, A. Nambi, and T. Ganu, “MEGA: Multilingual evaluation of generative AI”, arXiv preprint arXiv:2303.12528, 2023.
- Anthropic, “Introducing the next generation of Claude”, 2024, Available at https://www.anthropic.com/news/claude-3-family.
- Austin, J., A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, and C. Cai, “Program synthesis with large language models”, arXiv preprint arXiv:2108.07732, 2021.
- Azerbayev, Z., H. Schoelkopf, K. Paster, M. Santos, S. McAleer, A. Jiang, J. Deng, S. Biderman, and S. Welleck, “Llemma: An open language model for mathematics”, arXiv preprint arXiv: 2310.10631, 2023.
- Bhardwaj, R. and S. Poria, “Red-teaming large language models using chain of utterances for safety-alignment”, arXiv preprint arXiv:2308.09662, 2023.
- Bommasani, R., D. Hudson, E. Adeli, and R. Altman, “On the opportunities and risks of foundation models”, arXiv preprint arXiv:2108.07258, 2021.
- Bommasani, R., K. Klyman, S. Longpre, B. Xiong, S. Kapoor, N. Maslej, A. Narayanan, and P. Liang, “Foundation model transparency reports”, arXiv preprint arXiv:2402.16268, 2024.
- Brown, T., B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, and A. Askell, “Language models are few-shot learners”, Advances in Neural Information Processing Systems, Vol.33, 2020, pp. 1877-1901.
- Carolan, K., L. Fennelly, and A. Smeaton, “A review of multi-modal large language and vision models”, arXiv preprint arXiv:2404.01322, 2024.
- Chang, Y., X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, and Y. Wang, “A survey on evaluation of large language models”, ACM Transactions on Intelligent Systems and Technology, 2023. https://doi.org/10.1145/3641289
- Chen, M., J. Tworek, H. Jun, Q. Yuan, H. Pinto, J. Kaplan, H. Edwards, and Y. Burda, “Evaluating large language models trained on code”, arXiv preprint arXiv:2107.03374, 2021.
- Chowdhery, A., S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. Chung, C. Sutton, and S. Gehrmann, “PaLM: Scaling language modeling with pathways”, Journal of Machine Learning Research, Vol.24, No.240, 2023, pp. 1-113.
- Clark, J., E. Choi, M. Collins, D. Garrette, T. Kwiatkowski, V. Nikolaev, and J. Palomaki, “TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages”, Transactions of the Association for Computational Linguistics, Vol.8, 2020, pp. 454-470. https://doi.org/10.1162/tacl_a_00317
- Cobbe, K., V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, and M. Plappert, “Training verifiers to solve math word problems”, arXiv preprint arXiv:2110.14168, 2021.
- Costa-jussà, M., J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, and J. Maillard, “No language left behind: Scaling human-centered machine translation”, arXiv preprint arXiv:2207.04672, 2022.
- Databricks, “DBRX”, 2024, Available at https://github.com/databricks/dbrx.
- Devlin, J., M. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding”, arXiv preprint arXiv:1810.04805, 2018
- Dongqi, P. and V. Demberg, “ChatGPT vs Human-authored text: Insights into controllable text summarization and sentence style transfer”, arXiv preprint arXiv:2306.07799, 2023.
- Fu, C., P. Chen, Y. Shen, Y. Qin, M. Zhang, X. Lin, J. Yang, X. Zheng, K. Li, X. Sun, Y. Wu, and R. Ji, “MME: A comprehensive evaluation benchmark for multimodal large language models”, arXiv preprint arXiv:2306.13394. 2024.
- Gehman, S., S. Gururangan, M. Sap, Y. Choi, N. Smith, “RealToxicityPrompts: Evaluating neural toxic degeneration in language models”, In Findings of the Association for Computational Linguistics: EMNLP 2020, 2020, pp. 3356-3369.
- Gowda, T., Z. Zhang, C. Mattmann, and J. May, “Many-to-English machine translation tools, data, and pretrained models”, In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, 2021, pp 306-316.
- Hendrycks, D., C. Burns, S. Basart, A. Critch, J. Li, D. Song, and J. Steinhardt, “Aligning AI with shared human values”, arXiv preprint arXiv:2008.02275, 2020a.
- Hendrycks, D., C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding”, arXiv preprint arXiv:2009.03300, 2020b.
- Hendrycks, D., C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt, “Measuring mathematical problem solving with the MATH dataset”, arXiv preprint arXiv:2103.03874, 2021.
- Huang, J. and K. Chang, “Towards reasoning in large language models: A survey”, In 61st Annual Meeting of the Association for Computational Linguistics, ACL 2023, 2023, pp. 1049-1065.
- Iyer, S., I. Konstas, A. Cheung, and L. Zettlemoyer, “Summarizing source code using a neural attention model”, In 54th Annual Meeting of the Association for Computational Linguistics, 2016, pp. 2073-2083, Association for Computational Linguistics, August 2016.
- Jiang, A., A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. Chaplot, D. Casas, E. Hanna, and F. Bressand, “Mixtral of experts”, arXiv preprint arXiv:2401.04088, 2024.
- Kaplan, J., S. McCandlish, T. Henighan, T. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models”, arXiv preprint arXiv:2001.08361, 2021.
- Kasneci, E., K. Seßler, S. Küchemann, M. Bannert, D. Dementieva, F. Fischer, U. Gasser, and G. Groh, “ChatGPT for good? On opportunities and challenges of large language models for education”, Learning and Individual Differences, Vol.103, 2023, p. 102274. https://doi.org/10.1016/j.lindif.2023.102274
- Kim, D., C. Park, S. Kim, W. Lee, W. Song, Y. Kim, H. Kim, Y. Kim, H. Lee, and J. Kim, “SOLAR 10.7B: Scaling large language models with simple yet effective depth up-scaling”, arXiv preprint arXiv:2312.15166, 2023.
- Kim, S. and U. Wong, “ChatGPT Impacts on Academia”, In 2023 International Conference on System Science and Engineering (ICSSE), July 2023, pp. 422-426.
- Koubâa, A., W. Boulila, L. Ghouti, A. Alzahem, and S. Latif, “Exploring chatgpt capabilities and limitations: A survey”, IEEE Access, 2023.
- Lee, K., T. Hong, H. Ahn, T. Kim, and C. Koo, “Special topic: The impact of chatgpt in society, business, and academia”, Asia Pacific Journal of Information Systems, Vol.33, No.4, 2023, pp. 957-976. https://doi.org/10.14329/apjis.2023.33.4.957
- Li, B., R. Wang, G. Wang, Y. Ge, Y. Ge, and Y. Shan, “SEED-Bench: Benchmarking multimodal llms with generative comprehension”, arXiv preprint arXiv:2307.16125, 2023.
- Li, H., Y. Zhang, F. Koto, Y. Yang, H. Zhao, Y. Gong, N. Duan, and T. Baldwin, “CMMLU: Measuring massive multitask language understanding in Chinese”, arXiv preprint arXiv:2306.09212, 2023.
- Liu, J., C. Xia, Y. Wang, and L. Zhang, “Is your code generated by ChatGPT really correct? Rigorous evaluation of large language models for code generation”, Advances in Neural Information Processing Systems, Vol.36, 2024.
- Liu, N., K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang, “Lost in the middle: How language models use long contexts”, Transactions of the Association for Computational Linguistics, Vol.12, 2024, pp. 157-173. https://doi.org/10.1162/tacl_a_00638
- Lu, P., H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K. Chang, M. Galley, and J. Gao, “MathVista: Evaluating mathematical reasoning of foundation models in visual contexts”, arXiv preprint arXiv:2310.02255, 2023.
- McIntosh, T., T. Susnjak, T. Liu, P. Watters, and M. Halgamuge, “Inadequacies of large language model benchmarks in the era of generative artificial intelligence”, arXiv preprint arXiv:2402.09880, 2024.
- Mesnard, T., C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, and L. Sifre, “Gemma: Open models based on gemini research and technology” arXiv preprint arXiv:2403.08295, 2024.
- Meta-llama, "Llama 3", 2024, Available at https://github.cor/meta-llama/llama3.
- Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, and A. Ray, “Training language models to follow instructions with human Feedback”, Advances in Neural Information Processing Systems, Vol.35, 2022, pp. 27730-27744.
- Peng, B., C. Li, P. He, M. Galley, and J. Gao, “Instruction tuning with GPT-4”, arXiv preprint arXiv:2304.03277, 2024.
- Reid, M., N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J. Alayrac, and R. Soricut, “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context”, arXiv preprint arXiv:2403.05530, 2024.
- Roziere, B., J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. Tan, Y. Adi, and J. Liu, “Code llama: Open foundation models for code”, arXiv preprintarXiv:2308.12950, 2023.
- Sakaguchi, K., R. Bras, C. Bhagavatula, and Y. Choi, “WinoGrande: An adversarial winograd schema challenge at scale”, Communications of the ACM, Vol.64, No.9, 2021, pp. 99-106. https://doi.org/10.1145/3474381
- Schellaert, W., R. Hamon, F. Martinez-Plumed, and J. Hernandez-Orallo, “A proposal for scaling the scaling laws”, In Proceedings of the First edition of the Workshop on the Scaling Behavior of Large Language Models (SCALE-LLM 2024), March 2024, pp. 1-8.
- Son, G., H. Lee, S. Kim, S. Kim, N. Muennighoff, T. Choi, C. Park, K. Yoo, and S. Biderman, “KMMLU: Measuring massive multitask language understanding in Korean”, arXiv preprint arXiv:2402.11548, 2024.
- Srivastava, A., A. Rastogi, A. Rao, A. Shoeb, A. Abid, A. Fisch, A. Brown, A. Santoro, and A. Gupta, “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models”, arXiv preprint arXiv:2206.04615, 2022.
- Taori, R., I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. Hashimoto, “Alpaca: A strong, replicable instruction-following model”, Stanford Center for Research on Foundation Models, Vol.3, No.6, p. 7, 2023.
- Touvron, H., T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, and B. Roziere, “LLaMA: Open and efficient foundation language models”, arXiv preprint arXiv:2302.13971, 2023.
- Wang, Y., Y. Pan, M. Yan, Z. Su, and T. H. Luan, “A survey on ChatGPT: AI-generated contents, challenges, and solutions”, IEEE Open Journal of the Computer Society, 2023. https://doi.org/10.1109/OJCS.2023.3300321
- Xu, B., T. Li, J. Zheng, M. Naseriparsa, Z. Zhao, H. Lin, and F. Xia, “MET-Meme: A multimodal meme dataset rich in metaphors”, In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, July 2022, pp. 2887-2899.
- Young, A., B. Chen, C. Li, C. Huang, G. Zhang, G. Zhang, H. Li, and J. Zhu, “Yi: Open foundation models by 01.AI”, arXiv preprint arXiv:2403.04652, 2024.
- Yuan, Z., H. Yuan, C. Tan, W. Wang, and S. Huang, “How well do large language models perform in arithmetic tasks?”, arXiv preprint arXiv:2304.02015, 2023.
- Yue, X., Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, and Y. Sun, “MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI”, arXiv preprint arXiv: 2311.16502, 2023.
- Zellers, R., A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi, “HellaSwag: Can a machine really finish your sentence?”, arXiv preprint arXiv:1905.07830, 2019.
- Zhang, W., S. Aljunied, C. Gao, Y. Chia, and L. Bing, 'M3Exam: A multilingual, multimodal, multilevel benchmark for examining large lan-guage models') Advances in Neural Information Processing Systems, Vol.36, 2024.
- Zhao, W. X., K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, Y. Du, and C. Yang, “A survey of large language models”, arXiv preprint arXiv:2303.18223, 2023.
- Zhuang, Y., Y. Yu, K. Wang, H. Sun, and C. Zhang, 'ToolQA: A dataset for LLM question answering with external tools', Advances in Neural Information Processing Systems, Vol.36, 2024.