I am an AI researcher interested in building machines that perceive, understand, and act in the real world — in the belief that vision and intelligence are inseparable. I received my Ph.D. from KAIST under Prof. Jaegul Choo and an M.S. under Prof. In So Kweon, collaborating along the way with Carnegie Mellon University, NAVER AI Lab, Lunit AI, and Qualcomm AI Research. I aim to build reliable and accessible AI systems that can benefit people across diverse social and economic backgrounds.
Research Experiences
-
NAVER AI Lab Aug 2026 - PresentAI Researcher (Freelance)
Working on latent world models that can represent environment dynamics and predict future states.
Using these world models to simulate possible futures and reason about action consequences for agents and robots.
Working with Sangdoo Yun, Dongyoon Han, and Byeongho Heo -
NAVER AI Lab Apr 2025 - Oct 2025AI Research Intern -
Carnegie Mellon University Aug 2024 - Feb 2025Visiting Scholar in Computer Science (Korean Government Fellowship)
Collaborating with Prof. Yonatan Bisk's group -
-
-
Hyundai MOBIS Feb 2021 - Jun 2022Research Engineer with Industry-University Scholarship
Publications
-
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models- Letting Multimodal LLMs shift their visual focus, word by word, during generation.
- Matching dense-attention accuracy on image and video benchmarks with up to 90% fewer visual KV-cache entries.
-
RL makes MLLMs see better than SFT- Finding that RL post-training reshapes an MLLM's vision encoder toward stronger visual representations than SFT.
- Turning this analysis into PIVOT, a training recipe whose vision encoder outperforms more heavily trained encoders at under 1% of standard vision-pretraining compute.
-
MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning- Showing that a 93× smaller model (125M language model) matches MLLM captioning quality, making streaming captioning practical for video chatbots and navigation robots.
- Improving caption reliability by letting the model revisit the image and refine details from its initial description.
-
Is user feedback always informative? Retrieval Latent Defending for Semi-Supervised Domain Adaptation without Source Data- Identifying negatively biased feedback: user feedback concentrates on a model's errors, so fine-tuning on it directly can make deployed models worse.
- Building Retrieval Latent Defending, which rebalances this signal during adaptation; validated on classification, segmentation, and medical imaging, and used at Lunit to improve its medical AI systems.
-
Test-time Adaptation in the Dynamic World with Compound Domain Knowledge Management- Demonstrating the challenges self-driving models face in dynamically changing environments.
- Developing a framework that distributes and manages compound domain knowledge for fast and robust adaptation.
-
-
EcoTTA: Memory-Efficient Continual Test-time Adaptation via Self-distilled Regularization- Reducing the memory cost of continual test-time adaptation by up to 86% by training lightweight meta networks, fitting adaptation onto edge devices.
- Preventing catastrophic forgetting during long-term adaptation by using the original model as a guide without duplicating it.
Patents
-
Test-time adaptation via self-distilled regularization
Junha Song and Sungha Choi. 15 May, 2024. [U.S. Patent No. US20240160926A1]
Research Interests
I live by a simple motto: 'live life with full passion, and work with the greatest people.' The directions below reflect my current focus, but my interests are by no means limited to them.
-
Visual representation learning & World model
- Current vision encoders still fall short of providing precise visual states for robotic and interactive agents. I am interested in building visual representations that enable agents to perceive, predict, and interact reliably in real environments. A foundation for this exploration is classical self-supervised learning (SSL) in vision—jigsaw puzzles, rotation prediction, contrastive learning, and masked image modeling—which has produced strong general-purpose visual features. In IJCAI'23, I surveyed masked autoencoders as a generative paradigm for visual pretraining. My ICLR'26 work analyzes how visual representations evolve during MLLM training, showing that RL can boost the vision encoder effectively. Going forward, I aim to pursue SSL within unified multimodal models (UMMs) that jointly perform visual understanding and generation. I view UMMs as natural candidates for world models, which demand visual representations rich enough to anticipate how the world unfolds. This aligns with recent work on unified multimodal pretraining and the broader agenda of AMI Labs.
-
Multimodal LLMs and Agentic AI & RL
- Multimodal LLMs are emerging as capable visual assistants, yet they often miss fine-grained details and fail to ground their responses in the visual evidence. In CVPR'26, I studied multimodal self-refinement for revisiting visual evidence after an initial description. The arXiv'26 Gaze Attention enables MLLMs to attend to task-relevant visual regions during generation, avoiding redundant visual tokens. The ICLR'26 work examines how to post-train MLLMs for better visual understanding, showing that RL is more effective than SFT. Looking ahead, I want to extend these efforts to agentic AI systems, where multimodal models must perceive accurately while supporting tool use and planning. I am particularly interested in RL methods that make visual perception reliable enough for agentic decision-making.
-
Efficient AI systems
- Strong AI systems demand attention not only to capability but also to efficiency—training cost, inference latency, and memory usage. These constraints determine where AI can actually be deployed, especially in on-device and edge settings. My CVPR'23 EcoTTA work addressed memory-efficient continual test-time adaptation. In CVPR'26, I showed a lightweight captioning model can rival larger MLLMs, enabling on-device deployment. More recently, my arXiv'26 Gaze Attention work shrinks the visual KV cache by up to 90% while matching dense-attention baselines. Looking ahead, I want to extend these efforts to agentic systems, where RL, tool routing, and long-context processing introduce new computational and memory bottlenecks.
-
Robust and self-improving AI systems
- My earlier robustness research focused on domain adaptation in classical vision settings, improving classification and segmentation under distribution shift. My CVPR'23 studied test-time adaptation that stays stable over long-term deployment by mitigating catastrophic forgetting. RA-L/ICRA'24 extended this to dynamic environments, where domains can shift abruptly. ECCV'24 showed that user feedback is naturally biased toward a model's wrong predictions, which can degrade adaptation rather than improve it. These projects shaped my interest in self-improving systems that learn from changing environments. I want to revisit this direction with modern multimodal LLMs and agentic AI, in line with the broader push toward open-ended, self-improving AI pursued by Recursive Superintelligence.
Education
-
Korea Advanced Institute of Science and Technology (KAIST) Aug 2023 - Aug 2026Ph.D. degree in Graduate School of AI
Advisor: Prof. Jaegul Choo -
Korea Advanced Institute of Science and Technology (KAIST) Feb 2021 - Feb 2023M.S. degree in the Division of Future Vehicle
Advisor: Prof. In So Kweon
Grade: 3.9 / 4.3 (Percent: 95.56/100) -
Kookmin University (Seoul, South Korea) Feb 2015 - Feb 2021B.S. degree in Automotive Engineering
Grade: 4.39 / 4.5 (Rank: 1/121 | Percent: 98.7/100 | Major: 4.43)
National Science and Engineering Scholarship (Full tuition) from Korea Student Aid Foundation
Mandatory Military Service for 21 Months
Awards and Honors
- Intensive 2-Week Guest Lecturer, Introduction to Deep Learning, LG Innotek (2024)
- Best Master's Thesis Award, Korea Advanced Institute of Science and Technology (KAIST) (2023)
- Lecture planning consultant, Fast Campus (2022)
- Industry-University Scholarship (Full tuition support for M.S. program), Hyundai Mobis (2021–2022)
- National Science and Engineering Scholarship (Full tuition support for B.S. program), Korea Scholarship Foundation (2019–2021)
- Future Transport Design Award and Honorable Judge Award, 'Vehicle monitoring over internet toward digital twins', Cloud Programming World Cup, Japan (2019)
- Capstone Awards, Korean Society of Automotive Engineers (2019)
Projects
- Development of real-time masking/unmasking system for personal video information for public services such as CCTV (article), Korea Ministry of Science and ICT (2021 - 2023)
- Development of segmentation networks robust to environment variance, Hyundai Mobis (2021)
- Satellite image precision object detection, Korea Agency for Defense Development (ADD) (2020)
- Detection of Surrounding Vehicles using Deep Neural Network and Fusion of Panoramic Camera and Lidar Sensor, Korea Foundation for the Advancement of Science and Creativity (KORAC), Korea (2019)
B.S. Research Experiences
- "Style Transfer Maps from Satellite Images by using Generative Model", Korean Institute of Communications and Information Science (KICS) (2020)
- "Improvement of LiDAR and IMU-based autonomous driving performance in right-angle corner situations", Korean Sociey of Automotive Engineers (2019)
- Research Intern at Machine Intelligence Lab, Kookmin University (Dec 2019 - Oct 2020)
- Research Intern at Intelligence and Interaction Lab, Kookmin University (Feb 2019 - Nov 2019)
Skills
- Programming language: Python, C++
- Machine Learning Librarie: Pytorch, Tensorflow
- Application development: Robot Operating System (ROS)
- Sensor utilization: Camera, RGB-D Camera, LiDAR, GPS/IMU