Reinforcement Learning & Multi-Agent RL
Decentralized decision-making under partial observability, hierarchical coordination, and scalable cooperative learning.
Ph.D. Student
KAIST AI · OSI Lab
I am a Ph.D. student at the KAIST Kim Jaechul Graduate School of AI and a member of the OSI Lab, advised by Prof. Se-Young Yun. I received my M.S. in Computer Engineering from Arizona State University under the supervision of Prof. Theodore P. Pavlic.
My research focuses on reinforcement learning and multi-agent decision making, with particular interests in learning from human preferences and improving efficient inference and reasoning in large language models.
Our paper MELD: Multilingual Ensemble via Logical Debate was accepted to COLM 2026.
Our paper CUDA: Capturing Uncertainty and Diversity in Preference Feedback Augmentation was published in TMLR.
Our paper Bridging the Gap between Theory of Mind and Action in LLMs was accepted to the ICLR 2026 Agents in the Wild Workshop.
Our paper Dual Preference Learning for Multi-Agent Reinforcement Learning was published in IEEE Access.
I was recognized as a Top Reviewer for NeurIPS 2025.
I received the Encouragement Award at the 2025 Graduate Student Defense Academic Conference.
Our paper MA²E: Addressing Partial Observability in Multi-Agent Reinforcement Learning with Masked Auto-Encoder appeared at ICLR 2025.
Our X-band RADAR study received the 2024 KIMST Outstanding Paper Award.
References link to the related publications below.
Decentralized decision-making under partial observability, hierarchical coordination, and scalable cooperative learning.
Learning intent from limited or uncertain feedback, preference augmentation, and reward-free alignment.
Test-time computation, multilingual debate, efficient search, and Theory-of-Mind-grounded action.
Explainable decision support, course-of-action recommendation, and local-sensing multi-robot awareness.
C = conference · W = workshop · J = journal · K = domestic publication
[C1][W1]
Sehyeok Kang*, Jaejun Ryu*, Se-Young Yun
COLM 2026 · Accepted Earlier version: ACL 2026 MeLLM Workshop
[C2][W6]
Sehyeok Kang*, Yongsik Lee*, Gahee Kim, Song Chong, Se-Young Yun
ICLR 2025 Earlier version: ICLR 2024 GenAI4DM Workshop
[C3]
Minu Kim*, Yongsik Lee*, Sehyeok Kang, Jihwan Oh, Song Chong, Se-Young Yun
NeurIPS 2024
[C4]
Sehyeok Kang, Taeyeong Choi, Theodore P. Pavlic
IEEE ACSOS 2020 · Short paper
[C5]
Taeyeong Choi, Sehyeok Kang, Theodore P. Pavlic
IEEE ICRA 2020
[J1][W4]
Sehyeok Kang*, Jaewook Jeong*, Se-Young Yun
TMLR 2026 Earlier oral version: ICML 2025 MOFA Workshop
[J2]
Sehyeok Kang, Minu Kim, Jihwan Oh, Se-Young Yun
IEEE Access, vol. 14, 2026 Online 2025
[W2]
Sehyeok Kang, Jihwan Oh, Se-Young Yun
ICLR 2026 Agents in the Wild Workshop
[W3]
Sehyeok Kang, Youngjin Ko, Hyeonjun Kim, Myungjoo Kang, Se-Young Yun
IROS 2025 AIR4S Workshop
[W5]
Sehyeok Kang, Yongsik Lee, Se-Young Yun
ICML 2024 Models of Human Feedback for AI Alignment Workshop
Workshop labels [W1], [W4], and [W6] are paired with the corresponding archival versions above.
Value-Based Adaptive MCTS for Efficient LLM Reasoning · Journal of the Korea Society of Digital Industry and Information Management, 2025.
An Analysis of How Authority Bias Impacts Decision-Making in LLM Agent Collaboration · Journal of the Korea Society of Digital Industry and Information Management, 2025.
Incorporating Human Intent into Preference-Based Multi-Agent Reinforcement Learning · KIMST General and Fall Conferences, 2025.
Meta-Reinforcement Learning for Rapid Adaptation in Military Wargames · Graduate Student Defense Academic Conference, 2025.
A Study on Solving the Sparse-Reward Problem in Multi-Agent Reinforcement Learning Using Preference-Based Learning · KIMST Annual Conference, 2024.
A Study on the Analysis of Defense Data from Multiple Sources Using Data Science: Focusing on Case Studies · Journal of the Military Operations Research Society of Korea, 2023.
X-Band RADAR Reflected Signal Measurement of Gallium-Based Liquid Metal · Journal of the Korea Institute of Military Science and Technology, 2023.
CAMICAP: A Study on the Performance Improvement Algorithm of RICAP-Based Data Augmentation Techniques Using Grad-CAM · Journal of Korean Institute of Intelligent Systems, 2022.
Research of a Method of Generating an Adversarial Sample Using Grad-CAM · Journal of Korea Multimedia Society, 2022.
Shooting Sound Analysis Using Convolutional Neural Networks and Long Short-Term Memory · Journal of the Acoustical Society of Korea, 2022.
Research on Unidentified Tank Classification Using Few-Shot Learning · KICS Summer Conference, 2022.
Development of a Deep-Learning-Based Battlefield Gun-Noise Analysis Model · KICS Summer Conference, 2022.
Ho-Gil Kim, Sehyeok Kang, et al. · Cyber Electronic Warfare · Golden Pine Books, 2022.