Member image
KiHyun Nam
남기현
Ph.D. Candidate
School of Electrical Engineering, KAIST

About me

I am a third-year Ph.D. candidate at KAIST advised by Professor Joon Son Chung, and I earned my M.S. from KAIST. Since June 2026, I have also been a Ph.D. Research Intern with the Qualcomm Speech Team in South Korea. My research focuses on personalized voice agents that understand the person they are speaking with and tailor their interactions to that person.

I envision a future in which voice-first AI agents and physical AI become part of everyday life through robots, smart glasses, and other platforms that enable interaction without screens or manual controls. In these screenless settings, voice can be the most natural, and sometimes the only, way to communicate. As people entrust agents with increasingly personal aspects of their lives and work, I believe that understanding who is speaking is as important as understanding what is said. The same words may call for different responses depending on the speaker, their relationship to the agent, and the situation. I aim to build agents that interpret each participant's speech appropriately while keeping track of the requests and conversational context of the user they are assisting.

My research at KAIST has evolved from robust speech and speaker representation learning to diffusion-based modeling and audio-text alignment. Building on this foundation, I now explore speaker-aware Audio-LLMs that connect speaker understanding with language-based reasoning.

At Qualcomm, I am researching streaming speaker-aware Speech LLMs. I also work on full-duplex Speech LLMs that understand and adapt to their conversation partners. I am particularly interested in agents that listen while speaking and respond to incoming speech according to who is speaking and the conversational context. My goal is for speaker understanding to shape both response content and decisions about when to listen, speak, and respond.

I welcome conversations and collaborations on speech and speaker understanding, personalized voice agents, and voice interaction with robots and wearables. Please feel free to get in touch.

Experience

Ph.D. Research Intern, Qualcomm Speech Team, S. Korea

Jun. 2026 - Present

Deep Learning Research Intern, NAVER Clova Speech (now NAVER CLOUD), S. Korea

Sep. 2019 - Feb. 2020

Deep Learning Research Intern, NAVER Clova Speech (now NAVER CLOUD), S. Korea

Mar. 2021 - Sep. 2021

Education

Ph.D. in School of Electrical Engineering, KAIST

Sept. 2024 - Present

Advisor: Joon Son Chung (Multimodal AI Lab)

M.S. in School of Electrical Engineering, KAIST

Aug. 2022 - Aug. 2024

Advisor: Joon Son Chung (Multimodal AI Lab)

B.S. in Computer Science, Hankuk University of Foreign Studies (HUFS)

Mar. 2015 - Aug. 2022

Selected Awards

2024
  • NIST 2024 Speaker Recognition Evaluation – 1st Place (Audio Track) / 4th Place (Audio‑Visual Track) – Collaboration with Microsoft, KAIST MMAI Lab, PolyU, NUS and UEF

Publications

2026

Thumbnail
SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning
KiHyun Nam, J. W. Heo, S. Bae, H. J. Yu, and J. S. Chung
Preprint, 2026. Paper
Thumbnail
Diffusion‑Link: Diffusion Probabilistic Model for Bridging the Audio‑Text Modality Gap
KiHyun Nam*, J. M. Choi*, H. K. Lee, J. W. Heo, and J. S. Chung
ICASSP, 2026. Paper

2025

Thumbnail
SEED: Speaker Embedding Enhancement Diffusion Model
KiHyun Nam, J. W. Heo, J. W. Jung, G. Park, C. Jung, H. J. Yu, and J. S. Chung
INTERSPEECH, 2025. Paper Code

2024

Thumbnail
Disentangled Representation Learning for Environment‑agnostic Speaker Recognition
KiHyun Nam, H. S. Heo, J. W. Jung, and J. S. Chung
INTERSPEECH, 2024. Paper Project Page Code
Thumbnail
Rethinking Session Variability: Leveraging Session Embeddings for Session Robustness in Speaker Verification
H. S. Heo, KiHyun Nam, B. J. Lee, Y. Kwon, M. Lee, Y. J. Kim, and J. S. Chung
ICASSP, 2024. Paper
Thumbnail
TalkNCE: Improving Active Speaker Detection with Talk‑Aware Contrastive Learning
C. Jung*, S. Lee*, KiHyun Nam, K. Rho, Y. J. Kim, Y. Jang, and J. S. Chung
ICASSP, 2024. Paper
Thumbnail
VoxMM: Rich Transcription of Conversations in the Wild
D. Kwak*, J. Jung*, KiHyun Nam, Y. Jang, J. W. Jung, S. Watanabe, and J. S. Chung
ICASSP, 2024. Paper Dataset

2023

Thumbnail
Disentangled Representation Learning for Multilingual Speaker Recognition
KiHyun Nam*, Y. Kim*, J. Huh, H. S. Heo, J. W. Jung, and J. S. Chung
INTERSPEECH, 2023. Paper Project Page

2020

Thumbnail
ClovaCall: Korean Goal‑Oriented Dialog Speech Corpus for Automatic Speech Recognition of Contact Centers
J. Ha*, KiHyun Nam*, J. Kang, S. Lee, S. Yang, H. Jung, H. Kim, E. Kim, S. Kim, H. A. Kim, K. Doh, C. K. Lee, N. Sung, S. Kim
INTERSPEECH, 2020. Paper Code

Contact

  nkh.mmai (at) kaist.ac.kr
  Room 3103, N24 (LG Innovation Hall)

KAIST logo