K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
Published in EMNLP 2026 Findings, 2026
K-BrowseComp is a web-browsing agent benchmark grounded in Korean contexts, consisting of 400 problems. The 300-problem K-BrowseComp-Verified subset is manually constructed and validated by native Korean speakers. On the verified subset, frontier LLMs reach only 30.00–45.67%, a substantial drop from BrowseComp, while Korean LLMs released through Korea’s Proprietary AI Foundation Model program obtain only 0.00–10.33%.
| Paper (arXiv) | Code |
Recommended citation: Nahyun Lee, Dongkeun Yoon, Guijin Son, Geewook Kim, Dayoon Ko, Jeonghun Park, Haneul Yoo, Jaewon Cho, Junghun Park, Changyoon Lee, Kyochul Jang, Jaeyeon Kim, Eunsu Kim, Woojin Cho, Seungone Kim. (2026). "K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts." Findings of EMNLP 2026.
Download Paper
