Sitemap

A list of all the posts and pages found on the site. For you robots out there is an XML version available for digesting as well.

Pages

Posts

portfolio

publications

Can Language Models Laugh at YouTube Short-form Videos?

Published in EMNLP 2023, 2023

A video humor explanation benchmark to evaluate LLMs understanding of complex multimodal tasks through multimodal-filtering pipeline.

Recommended citation: Dayoon Ko, Sangho Lee, Gunhee Kim. (2023). "Can Language Models Laugh at YouTube Short-form Videos?" EMNLP 2023.
Download Paper

GrowOVER: How Can LLMs Adapt to Growing Real-World Knowledge?

Published in ACL 2024, 2024

Continuously updated QA & dialogue benchmarks to assess whether LLMs can handle evolving knowledge, enabling RAG systems to adapt without retraining.

Recommended citation: Dayoon Ko, Jinyoung Kim, Hahyeon Choi, Gunhee Kim. (2024). "GrowOVER: How Can LLMs Adapt to Growing Real-World Knowledge?" ACL 2024.
Download Paper

DynamicER: Resolving Emerging Mentions to Dynamic Entities for RAG

Published in EMNLP 2024, 2024

This work addresses challenges in resolving temporally evolving mentions to entities, improving retrieval and enhancing RAG accuracy in dynamic environments.

Recommended citation: Jinyoung Kim, Dayoon Ko, Gunhee Kim. (2024). "DynamicER: Resolving Emerging Mentions to Dynamic Entities for RAG." EMNLP 2024.
Download Paper

When Should Dense Retrievers Be Updated in Evolving Corpora? Detecting Out-of-Distribution Corpora Using GradNormIR

Published in ACL 2025 Findings, 2025

A method to detect when dense retrievers need updating in evolving corpora using gradient norms, enabling efficient adaptation to distribution shifts.

Recommended citation: Dayoon Ko, Jinyoung Kim, Sohyeon Kim, Jinhyuk Kim, Jaehoon Lee, Seonghak Song, Minyoung Lee, Gunhee Kim. (2025). "When Should Dense Retrievers Be Updated in Evolving Corpora? Detecting Out-of-Distribution Corpora Using GradNormIR." Findings of ACL 2025.
Download Paper

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates

Published in ACL 2025, 2025

MAC benchmark for evaluating adversarial compositionality of multimodal models, revealing vulnerabilities in vision-language models like CLIP to text-based adversarial attacks.

Recommended citation: Jaewoo Ahn, Heeseung Yun, Dayoon Ko, Gunhee Kim. (2025). "Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates." ACL 2025.
Download Paper

Hybrid Deep Searcher: Scalable Parallel and Sequential Search Reasoning

Published in ICLR 2026, 2026

A scalable search agent that dynamically integrates parallel and sequential strategies for multi-hop QA with RAG, trained on HDS-QA dataset.

Recommended citation: Dayoon Ko, Jihyuk Kim, Haeju Park, Sohyeon Kim, Dahyun Lee, Yongrae Jo, Gunhee Kim, Moontae Lee, Kyungjae Lee. (2026). "Hybrid Deep Searcher: Scalable Parallel and Sequential Search Reasoning." ICLR 2026.
Download Paper

When Is Enough Not Enough? Illusory Completion in Search Agents

Published in COLM 2026 Workshop, 2026

A correct final answer does not show whether a search agent verified every constraint; we evaluate verification along agent trajectories with the Epistemic Ledger.

Recommended citation: Dayoon Ko, Jihyuk Kim, Sohyeon Kim, Haeju Park, Dahyun Lee, Gunhee Kim, Moontae Lee, Kyungjae Lee. (2026). "When Is Enough Not Enough? Illusory Completion in Search Agents." COLM 2026 Workshop.
Download Paper

K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts

Published in EMNLP 2026 Findings, 2026

A web-browsing agent benchmark grounded in Korean contexts, revealing a substantial performance drop for frontier and Korean LLMs.

Recommended citation: Nahyun Lee, Dongkeun Yoon, Guijin Son, Geewook Kim, Dayoon Ko, Jeonghun Park, Haneul Yoo, Jaewon Cho, Junghun Park, Changyoon Lee, Kyochul Jang, Jaeyeon Kim, Eunsu Kim, Woojin Cho, Seungone Kim. (2026). "K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts." Findings of EMNLP 2026.
Download Paper

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

Published in arXiv 2026, 2026

A literature inspiration retrieval benchmark grounded in researchers' firsthand knowledge of their own projects.

Recommended citation: Sohyeon Kim, Yoonho Lee, Bo Liu, Dayoon Ko, Rulin Shao, Seungone Kim, Graham Neubig, Pang Wei Koh, Aakanksha Chowdhery, Akari Asai, Omar Khattab, Yejin Choi, Gunhee Kim, Chelsea Finn. (2026). "ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research." arXiv preprint arXiv:2610.02202.
Download Paper

CompLit: Scientific Literature Search Benchmarks Must Cover Implicit, Cumulative, and Unmet Needs

Published in arXiv 2026, 2026

A scientific literature search benchmark across all eight arXiv domains, with one diagnostic setting each for implicit, cumulative, and unmet needs.

Recommended citation: Dayoon Ko, Jihyuk Kim, Soyeong Jeong, Young-Jun Lee, Dahyun Lee, Juyeon Kim, Gunhee Kim, Moontae Lee, Kyungjae Lee. (2026). "CompLit: Scientific Literature Search Benchmarks Must Cover Implicit, Cumulative, and Unmet Needs." arXiv preprint.

talks

teaching