About Me

Hi! I’m Dayoon Ko 😊, a Ph.D. candidate in Computer Science and Engineering at Seoul National University, advised by Prof. Gunhee Kim. Currently, I’m a visiting researcher at UC Berkeley, working with Prof. Sewon Min.

I’m broadly interested in how large language models can find, verify, and use information reliably in the noisy, fast-changing multimodal environments where people actually use them.

Right now, I’m working on on-device multimodal retrieval, where models must search over the photos, videos, and documents people keep on their own devices, under tight memory and compute budgets. I have worked on search agents, scaling their search reasoning, checking whether they verify what they claim, and evaluating them in realistic settings, as well as on keeping LLMs and RAG systems up to date as real-world knowledge evolves.

Outside of research, I enjoy dancing šŸ’ƒ or doing CrossFit šŸ‹šŸ»ā€ā™€ļø. Staying active keeps my brain happy!

šŸ”„ Recent News

[Oct 2026] Started a visiting research position at UC Berkeley with Prof. Sewon Min's group! 🐻
[Oct 2026] "ScholarCatalyst", a benchmark for retrieving papers that inspire new research, is out on arXiv!
[Sep 2026] "K-BrowseComp" has been accepted at EMNLP 2026 Findings! šŸŽ‰
[Sep 2026] "When Is Enough Not Enough? Illusory Completion in Search Agents" has been accepted at a COLM 2026 Workshop!

Selected Publications

ScholarCatalyst
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
Sohyeon Kim, Yoonho Lee, Bo Liu, Dayoon Ko, Rulin Shao, Seungone Kim, Graham Neubig, Pang Wei Koh, Aakanksha Chowdhery, Akari Asai, Omar Khattab, Yejin Choi, Gunhee Kim, Chelsea Finn
arXiv 2026
A literature inspiration retrieval benchmark grounded in researchers' firsthand knowledge of their own projects: 184 researchers who led 207 recent CS projects verified 894 research questions, labeling which prior papers did or could have advanced their work.
K-BrowseComp
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
Nahyun Lee, Dongkeun Yoon, Guijin Son, Geewook Kim, Dayoon Ko, Jeonghun Park, Haneul Yoo, Jaewon Cho, Junghun Park, Changyoon Lee, Kyochul Jang, Jaeyeon Kim, Eunsu Kim, Woojin Cho, Seungone Kim
EMNLP 2026 Findings
A 400-problem web-browsing agent benchmark grounded in Korean contexts, with a 300-problem subset verified by native Korean speakers. Frontier LLMs reach only 30–46% on the verified subset, a substantial drop from BrowseComp.
Illusory Completion
When Is Enough Not Enough? Illusory Completion in Search Agents
Dayoon Ko, Jihyuk Kim, Sohyeon Kim, Haeju Park, Dahyun Lee, Gunhee Kim, Moontae Lee, Kyungjae Lee
COLM 2026 Workshop
A correct final answer does not show whether a search agent verified every constraint. We introduce the Epistemic Ledger to evaluate verification along agent trajectories, revealing illusory completion: agents conclude while constraints remain assumed, refuted, or unchecked, even in their correct answers.
HybridDeepSearcher
Hybrid Deep Searcher: Scalable Parallel and Sequential Search Reasoning
Dayoon Ko, Jihyuk Kim, Haeju Park, Sohyeon Kim, Dahyun Lee, Yongrae Jo, Gunhee Kim, Moontae Lee, Kyungjae Lee
ICLR 2026
A scalable search agent that dynamically integrates parallel and sequential search strategies for multi-hop QA with RAG. We introduce the HDS-QA training dataset and achieve significant improvements.
GradNormIR
When Should Dense Retrievers Be Updated in Evolving Corpora? Detecting Out-of-Distribution Corpora Using GradNormIR
Dayoon Ko, Jinyoung Kim, Sohyeon Kim, Jinhyuk Kim, Jaehoon Lee, Seonghak Song, Minyoung Lee, Gunhee Kim
ACL 2025 Findings
We propose GradNormIR, an unsupervised method that detects out-of-distribution shifts in document collections using gradient norms, enabling timely updates of dense retrievers without manual intervention.
MAC
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
Jaewoo Ahn, Heeseung Yun, Dayoon Ko, Gunhee Kim
ACL 2025
We introduce MAC benchmark for evaluating the robustness of pre-trained multimodal models against adversarial text updates, revealing vulnerabilities in vision-language models like CLIP.
DynamicER
DynamicER: Resolving Emerging Mentions to Dynamic Entities for RAG
Jinyoung Kim, Dayoon Ko, Gunhee Kim
EMNLP 2024
This work addresses challenges in resolving temporally evolving mentions to entities. Resolving mentions is key to improving retrieval, enhancing RAG accuracy in dynamic environments.
GrowOVER
GrowOVER: How Can LLMs Adapt to Growing Real-World Knowledge?
Dayoon Ko, Jinyoung Kim, Hahyeon Choi, Gunhee Kim
ACL 2024
We propose QA & dialogue benchmarks that are continuously and automatically updated to assess whether LLMs can handle evolving knowledge. By making LLMs evaluate their confidence, we enable RAG systems to adapt to new knowledge without retraining.
ExFunTube
Can Language Models Laugh at YouTube Short-form Videos?
Dayoon Ko, Sangho Lee, Gunhee Kim
EMNLP 2023
A video humor explanation benchmark via a multimodal-filtering pipeline to evaluate LLMs' understanding of complex multimodal tasks like humor. We generate several frame captions and filter them based on video segments to enhance LLMs with vision capabilities.

Experiences

Visiting Researcher
UC Berkeley
Hosted by Professor Sewon Min
October 2026 - Present
Research Intern
LG AI Research, Superintelligence Lab
Worked on agentic reasoning and retrieval systems for large language models
March 2025 - September 2026

Education

M.S./Ph.D. in Computer Science and Engineering
Seoul National University
Advisor: Professor Gunhee Kim
September 2022 - Present
B.S. in Computer Science and Engineering
Yonsei University
GPA: 4.12/4.30 (Rank: 1/28)
March 2018 - August 2022