My research explores how generative modeling can be extended to intelligent agents that perceive, reason, and act in the physical world.
My previous work spans multimodal generation and efficient inference across vision, language, and speech, including real-time interactive agents, speech generation, and image generation. Building on this foundation, I am currently working on robot foundation models and agentic systems that connect multimodal understanding with physical action.
My long-term goal is to develop efficient and generalizable embodied agents that can adapt to diverse tasks and interact reliably with the real world.
Previously, I was a research engineer at KRAFTON. I received my M.S. in Artificial Intelligence from KAIST, where I was advised by Prof. Juho Lee, and my B.S. in Industrial Engineering from Seoul National University.
(* denotes equal contribution)
For a complete list of publications, please see my Google Scholar profile.
Notes on world models, generative modeling, and Physical AI will show up here.