
Leader, NAVER AI Lab
Guest Associate Professor, SNU AIIS
Jin-Hwa Kim has been the Leader of Generation Research at NAVER AI Lab, working since August 2021, and a Guest Associate Professor at the Artificial Intelligence Institute of Seoul National University (SNU AIIS) since August 2022. He has studied neural 3d, world models, multimodal generation, ethical and safe AI, and other related topics. In 2018, he received a Ph.D. from Seoul National University under the supervision of Professor Byoung-Tak Zhang for the work on “Multimodal Deep Learning for Visually-grounded Reasoning.” In September 2017, he received 2017 Google Ph.D. Fellowship in Machine Learning, Ph.D. Completion Scholarship by Seoul National University, and the VQA Challenge 2018 runners-up at the CVPR 2018 VQA Challenge and Visual Dialog Workshop. He was Research Intern at Facebook AI Research (Menlo Park, CA) mentored by Yuandong Tian, Devi Parikh, and Dhruv Batra, from January to May in 2017. He worked for SK Telecom (August 2018 to July 2021) and SK Communications (January 2011 to October 2012).
3 Aug 2026
📣 Welcoming Dr. Jungho Lee to Generation Research, NAVER AI Lab!
I'm excited to share that Jungho Lee has joined our team as a Research Scientist, starting on August 3. 🤗
Jungho earned his Ph.D. in Electrical and Electronics Engineering from Yonsei University, where his work focused on 3D computer vision — building systems that can perceive, reconstruct, and understand the 3D world from real-world visual observations. His research has been recognized at top-tier venues including CVPR, ICCV, ECCV, and AAAI, including a CVPR 2025 Oral presentation.
At Generation Research, Jungho will help push forward one of our most ambitious efforts: building a physical world model designed to understand and simulate real-world environments at city scale. This work sits at the heart of AI Lab's mission, spanning spatial intelligence to embodied AI.
3D vision and world modeling are foundational to how AI will understand and interact with physical space, and Jungho's expertise is a great fit for where we're headed. Welcome aboard!
J.H.
Seoul World Model: Grounding World Simulation Models in a Real-World Metropolis
ECCV 2026 Spotlight
TetraSDF: Precise Mesh Extraction with Multi-resolution Tetrahedral Grid
ECCV 2026 Spotlight
Relaxed Rigidity with Ray-based Grouping for Dynamic Gaussian Splatting
ECCV 2026
Kinematics-Driven Gaussian Shape Deformation for Blurry Monocular Dynamic Scenes
ICML 2026
CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
CVPR 2026
MVCustom: Multi-View Customized Diffusion via Geometric Latent Rendering and Completion
ICLR 2026
MoAI: Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
ICLR 2026
HyperCLOVA X 8B Omni
arXiv preprint, 2026
HyperCLOVA X 32B Think
arXiv preprint, 2026
Pygmalion Effect in Vision: Image-to-Clay Translation for Reflective Geometry Reconstruction
arXiv preprint, 2025
Surface-Based Visibility-Guided Uncertainty for Continuous 3D Active Neural Reconstruction
AAAI 2026 Workshop on AI with Biased or Scarce Data, Oral
Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion Models
NeurIPS 2025
OCK: Unsupervised Dynamic Video Prediction with Object-Centric Kinematics
ICCV 2025
Background-aware Moment Detection for Video Moment Retrieval
WACV 2025
TropicalNeRF: Polyhedral Complex Derivation from Piecewise Trilinear Networks
NeurIPS 2024
A Simple Remedy for Dataset Bias via Self-Influence: A Mislabeled Sample Perspective
NeurIPS 2024
Direct Unlearning Optimization for Robust and Safe Text-to-Image Models
NeurIPS 2024
Effective Rank Analysis and Regularization for Enhanced 3D Gaussian Splatting
NeurIPS 2024
Text2Chart31: Instruction Tuning for Chart Generation with Automatic Feedback
EMNLP 2024 Oral
Factorized Multi-Resolution HashGrid for Efficient Neural Radiance Fields: Execution on Edge-Devices
IEEE Robotics and Automation Letters (RA-L) 2024; ICRA 2025
TimeChara: Evaluating Point-in-Time Character Hallucination of Role-Playing Large Language Models
ACL 2024
Synergistic Integration of Coordinate Network and Tensorial Feature for Improving Neural Radiance Fields from Sparse Inputs
ICML 2024
PruNeRF: Segment-Centric Dataset Pruning via 3D Spatial Consistency
ICML 2024
Vision-Language Generative Model for View-Specific Chest X-ray Generation
CHIL 2024
Geometry-Aware Score Distillation via 3D Consistent Noising and Gradient Consistency Modeling
arXiv preprint, 2024
Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation
ICLR 2024
HyperCLOVA X Technical Report
arXiv preprint, 2024
Dense Text-to-Image Generation with Attention Modulation
ICCV 2023
Text-Conditioned Sampling Framework for Text-to-Image Generation with Masked Generative Models
ICCV 2023
3D-aware Blending with Generative NeRFs
ICCV 2023
Robust Camera Pose Refinement for Multi-Resolution Hash Encoding
ICML 2023
Query-Efficient Black-Box Red Teaming via Bayesian Optimization
ACL 2023
The Dialog Must Go On: Improving Visual Dialog via Generative Self-Training
CVPR 2023
Panoramic Image-to-Image Translation
arXiv preprint, 2023
Semi-Parametric Video-Grounded Text Generation
arXiv preprint, 2023
AlphaTuning: Quantization-Aware Parameter-Efficient Adaptation of Large-Scale Pre-Trained Language Models
Findings of EMNLP 2022
Modal-Specific Pseudo Query Generation for Video Corpus Moment Retrieval
EMNLP 2022
SelecMix: Debiased Learning by Contradicting-pair Sampling
NeurIPS 2022
Mutual Information Divergence: A Unified Metric for Multimodal Generative Models
NeurIPS 2022
Understanding Cross-Domain Few-Shot Learning Based on Domain Similarity and Few-Shot Difficulty
NeurIPS 2022
ReFine: Re-randomization before Fine-tuning for Cross-domain Few-shot Learning
CIKM 2022
Why Knowledge Distillation Amplifies Gender Bias and How to Mitigate from the Perspective of DistilBERT
NAACL 2022 Workshop (GeBNLP)
Logit Mixing Training for More Reliable and Accurate Prediction
IJCAI 2022
Reasoning Visual Dialog with Sparse Graph Learning and Knowledge Transfer
Findings of EMNLP 2021
Semi-orthogonal Embedding for Efficient Unsupervised Anomaly Segmentation
arXiv preprint, 2021
Multi-step Estimation for Gradient-based Meta-learning
arXiv preprint, 2020
CoDraw: Collaborative Drawing as a Testbed for Grounded Goal-driven Communication
ACL 2019
Relational Bilinear Attention Networks with the Assumption of Linear Compositionality
CVPR 2019 Workshop (VQA & Dialog)
Korean Localization of Visual Question Answering for Blind People
NeurIPS 2019 Workshop (AI for Social Good)
Multimodal Dual Attention Memory for Video Story Question Answering
ECCV 2018
Bilinear Attention Networks
NeurIPS 2018
Visual Explanations from Hadamard Product in Multimodal Deep Networks
NIPS 2017 Workshop (ViGIL)
Large-Scale Text Classification with Deep Neural Networks
KIISE Transactions on Computing Practices, Vol. 23, No. 5, 2017
Hadamard Product for Low-rank Bilinear Pooling
ICLR 2017
Overcoming Catastrophic Forgetting by Incremental Moment Matching
NeurIPS 2017 Spotlight
Multimodal Residual Learning for Visual QA
NeurIPS 2016
Large-scale Text Classification with Recurrent Neural Networks
Korea Computer Congress 2016
Active Vision from Image-Text Multimodal System Learning
Journal of KIISE, Vol. 43, No. 7, 2016
TrimZero: A Torch Recurrent Module for Efficient Natural Language Processing
KIIS Spring Conference 2016
rnn: Recurrent Library for Torch
arXiv preprint, 2015
Multimodal foundation models — multimodal LLMs and multimodal generative AI — together with physical world models that represent, simulate, and predict the 3D/4D physical world under actions, followed by a paper-driven seminar. Co-taught with Sangdoo Yun.
Multimodal generative AI from vision and language to audio and speech, with neural graphics (NeRFs, 3DGS) as an emerging topic in multi-view representation learning. Co-taught with Sangdoo Yun.
Core theories and applications of multimodal deep learning across vision, language, audio, and speech, extending into neural graphics (NeRFs). Co-taught with Sangdoo Yun and Jiyoung Lee.
Core theories and applications of multimodal deep learning across vision, language, audio, and speech, with neural graphics (NeRFs) as time permits. Co-taught with Jiyoung Lee.
Area Chair for ICLR 2027, NeurIPS 2024-2026, ICML 2025-2026, and ACL ARR 2024 June
Organizer for ICML 2026 Expo Talks and Panels — Seoul World Model
Reviewer for NeurIPS 2018-2023, ICLR 2019, 2021-2023, ICML 2019-2021, 2024, and COLING 2022
Reviewer for Neural Networks (2020)
Reviewer for IEEE Transactions on Neural Networks and Learning Systems (2019)
Topic Editor for “Identifying, Analyzing, and Overcoming Challenges in Vision and Language Research” in the Frontiers Research Topics
Program Committee for the “1st workshop on Video-Language Models” at NeurIPS 2024
Co-organizer for ICML 2023 Social — ML in Korea
Program Committee for the 2nd workshop on “Video Turing Test” at ECCV 2020
Invited Speaker for the 3rd Workshop on “Closing the Loop Between Vision and Language” at ICCV 2019
Program Committee for the 1st Workshop on “Video Turing Test” at ICCV 2019
Program Committee for Workshop on “SiVL” at ECCV 2018
None of the work above happened alone — it grew out of papers written together, ideas borrowed and returned, and the people who shaped how I think about research. Without these collaborations, there would be little to show for the years behind this page.
ColleaguesDongyoon Han, Gayoung Lee, Jungho Lee, Junho Kim, Sangdoo Yun
Interns MentoredHwan Heo, Hyunin Cho, Inho Kong, Injae Kim, Jaeseok Jeong, Jaeseong Lee, Jangho Park, Jiwook Kim, Jun-Seong Kim, Junha Hyung, Junoh Lee, Junyoung Seo, Min-Seop Kwak, Mingyu Kim, Minjung Shin, Minkyung Kwon, Seonghun Oh, Soohyun Kim, Suhyeon Lee, Susung Hong, Yong-Hyun Park
Collaborating ProfessorsByoung-Tak Zhang, Devi Parikh, Dhruv Batra, Hyunwoo J. Kim, Jun-Yan Zhu, Mohamed Elhoseiny, Seungryong Kim, Tae-Hyun Oh, Youngjung Uh, Yuandong Tian
Personal CollaboratorsDamien Teney, Jiasen Lu, Kushal Kafle, Sang-Woo Lee
A personal site for research, teaching, and writing under one roof. Bauhaus geometric discipline and a trace of Greek architectural proportion, rendered in a bold, direct primary-color palette.
Enjoys computer games — Diablo 2, Elden Ring, Lineage, Overwatch, Stellar Blade.
Amateur badminton player for about 10 years.
Collects deep learning memorabilia: a name card from Andrew Ng, a copy of Deep Learning signed by Ian Goodfellow, an official ICML 2016 poster, a bulldozer Lego set (Set 856, 1979–1981; the one behind the "Blender synthetic" renders), and a first-edition printing of Diederik P. Kingma's PhD thesis.
Once dreamed of becoming a graphic designer: won the 1st Korea Flash Award — Motion Graphics Finalist (Yoon Design Lab) at COEX in 2001. A crewmate from that era, Brand Visual Director and CEO of VBstudio, went on to become the most successful of the group after serving as head of the Brand Design Team at YG Entertainment.