I received my PhD in Artificial Intelligence from KAIST in 2026, advised by Professors Sung Ju Hwang and Juho Lee in the Machine Learning and Artificial Intelligence (MLAI) Lab. My research focuses on AI safety, responsible AI, and evaluation. I also collaborate closely with Kenji Kawaguchi. I was a recipient of the 2023 Apple Scholars in AI/ML PhD Fellowship. Here is my CV.
During my PhD, I interned with Apple Machine Learning Research in Seattle, working on synthetic data generation for tool-calling language models. I also interned at Krafton and Mila, where I worked with Yoshua Bengio, Minsu Kim, Moksh Jain, and Esmeralda S. Whitammer. I interned at Apple Cambridge with Anders Johannsen and Jianpeng Cheng, and the National University of Singapore with Kenji Kawaguchi.
📖 Education
- 2022–2026, PhD in Artificial Intelligence, KAIST. Advisors: Sung Ju Hwang and Juho Lee.
- 2020–2022, MS in Artificial Intelligence, KAIST.
- 2011–2018, BA in Library and Information Science, Yonsei University.
💻 Work Experience
- 2025.10–2026.09, Research internship, Apple Machine Learning Research, Seattle. Host: Raviteja Vemulapalli.
- 2025.07–2025.10, Research internship, Krafton, Seoul.
- 2024.01–2024.06, Research internship, Mila, Montreal. Advisor: Yoshua Bengio.
- 2023.05–2023.09, Research internship, Apple Cambridge. Host: Anders Johannsen.
- 2022.07–2022.09, Remote research internship, National University of Singapore. Advisor: Kenji Kawaguchi.
📝 Publications
-
Simulate to Generalize: Scaling Stateful Supervision for API-calling Agents using LLM World Models
[paper]
Seanie Lee, Sanjoy Chowdhury, Chao Jiang, Cheng-Yu Hsieh, Ting-Yao Hu, Alexander T. Toshev, Oncel Tuzel, and Raviteja Vemulapalli
arXiv 2026 -
THINKSAFE: Self-Generated Safety Alignment for Reasoning Models
[paper] [code]
Seanie Lee*, Sangwoo Park*, Yumin Choi, Gyeongman Kim, Minki Kang, Jihun Yun, Dongmin Park, Jongho Park, and Sung Ju Hwang
NeurIPS 2026 -
It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs
[paper] [code]
Sangwoo Park*, Woongyeong Yeo*, Seanie Lee, Yumin Choi, Hyomin Lee, Kangsan Kim, Jinheon Baek, Seong Joon Oh, and Sung Ju Hwang
NeurIPS 2026 -
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
[paper] [code]
Hyomin Lee, Sangwoo Park, Yumin Choi, Sohyun An, Seanie Lee, and Sung Ju Hwang
EMNLP 2026 -
Rethinking Reward Models for Multi-Domain Test-Time Scaling
[paper] [code]
Dong Bok Lee*, Seanie Lee*, Sangwoo Park, Minki Kang, Jinheon Baek, Dongki Kim, Dominik Wagner, Jiongdao Jin, Heejun Lee, Tobias Bocklet, Jinyu Wang, Jingjing Fu, Sung Ju Hwang, Jiang Bian, and Lei Song
TMLR 2026 -
HoliSafe: Holistic Safety Benchmarking and Modeling with Safety Meta Token for Vision-Language Model
[paper] [code]
Youngwan Lee, Kangsan Kim, Kwanyong Park, Ilcahe Jung, Soojin Jang, Seanie Lee, Yong-Ju Lee, and Sung Ju Hwang
CVPR Findings 2025 -
FedSVD: Adaptive Orthogonalization for Private Federated Learning with LoRA
[paper] [code]
Seanie Lee*, Sangwoo Park*, Dong Bok Lee*, Dominik Wagner, Haebin Seong, Tobias Bocklet, Juho Lee, and Sung Ju Hwang
NeurIPS 2025 -
Distilling LLM Agent into Small Models with Retrieval and Code Tools
[paper] [code]
Minki Kang, Jongwon Jeong, Seanie Lee, Jaewoong Cho, and Sung Ju Hwang
NeurIPS Spotlight 2025 -
Reliable Decision-Making via Calibration-Oriented Retrieval-Augmented Generation
[paper] [code]
Chaeyun Jang, Deukhwan Cho, Seanie Lee, Hyungi Lee, and Juho Lee
NeurIPS 2025 -
Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training
[paper] [code]
Brian R. Bartoldson, Siddarth Venkatraman, James Diffenderfer, Moksh Jain, Tal Ben-Nun, Seanie Lee, Minsu Kim, Johan Obando-Ceron, Yoshua Bengio, and Bhavya Kailkhura
NeurIPS 2025 -
SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models
[paper] [code]
Seanie Lee*, Dong Bok Lee*, Dominik Wagner, Minki Kang, Haebin Seong, Tobias Bocklet, Juho Lee, and Sung Ju Hwang
ACL Findings 2025 -
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
[paper]
Dominik Wagner, Ilja Baumann, Natalie Engert, Seanie Lee, Elmar Nöth, Korbinian Riedhammer, and Tobias Bocklet
Interspeech 2025 -
HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
[paper] [code]
Seanie Lee*, Haebin Seong*, Dong Bok Lee, Minki Kang, Xiaoyin Chen, Dominik Wagner, Yoshua Bengio, Juho Lee, and Sung Ju Hwang
ICLR 2025 -
Learning Diverse Attacks on Large Language Models for Robust Red-teaming and Safety Tuning
[paper] [code]
Seanie Lee, Minsu Kim, Lynn Cherif, David Dobre, Juho Lee, Sung Ju Hwang, Kenji Kawaguchi, Gauthier Gidel, Yoshua Bengio, Nikolay Malkin, and Moksh Jain
ICLR 2025 -
Optimized Speculative Sampling for GPU Hardware Accelerators
[paper] [code]
Dominik Wagner, Seanie Lee, Ilja Baumann, Philipp Seeberger, Korbinian Riedhammer, and Tobias Bocklet
EMNLP 2024 -
Drug Discovery with Dynamic Goal-aware Fragment
[paper] [code]
Seul Lee, Seanie Lee, Kenji Kawaguchi, and Sung Ju Hwang
ICML 2024 -
Effective and Efficient Conversation Retrieval for Dialogue State Tracking with Implicit Text Summaries
[paper]
Seanie Lee, Jianpeng Cheng, Joris Driesen, Alexandru Coca, and Anders Johannsen
NAACL 2024 -
Self-Supervised Dataset Distillation for Transfer Learning
[paper] [code]
Dong Bok Lee*, Seanie Lee*, Joonho Ko, Kenji Kawaguchi, Juho Lee, and Sung Ju Hwang
ICLR 2024 -
DiffusionNAG: Task-guided Neural Architecture Generation with Diffusion Models
[paper] [code]
Sohyun Ahn*, Hayeon Lee*, Jaehyeong Jo, Seanie Lee, and Sung Ju Hwang
ICLR 2024 -
Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive Tasks
[paper] [code]
Minki Kang, Seanie Lee, Jinheon Baek, Kenji Kawaguchi, and Sung Ju Hwang
NeurIPS 2023 -
Scalable Set Encoding with Universal Mini-Batch Consistency and Unbiased Full Set Gradient Approximation
[paper] [code]
Jeffrey Willette*, Seanie Lee*, Bruno Andreis, Kenji Kawaguchi, Juho Lee, and Sung Ju Hwang
ICML 2023 -
Margin-based Neural Network Watermarking
[paper]
Byungjoo Kim, Suyoung Lee, Seanie Lee, Sooel Son, and Sung Ju Hwang
ICML 2023 -
Self-Supervised Set Representation Learning for Unsupervised Meta-Learning
[paper]
Dong Bok Lee*, Seanie Lee*, Kenji Kawaguchi, Yunji Kim, Jihwan Bang, Jung-Woo Ha, and Sung Ju Hwang
ICLR 2023 -
Self-Distillation for Further Pre-training of Transformers
[paper] [code]
Seanie Lee, Minki Kang, Juho Lee, Sung Ju Hwang, and Kenji Kawaguchi
ICLR 2023 -
Set-based Meta-Interpolation for Few-Task Meta-Learning
[paper]
Seanie Lee*, Bruno Andreis*, Kenji Kawaguchi, and Sung Ju Hwang
NeurIPS 2022 -
On Divergence Measures for Bayesian Pseudocoresets
[paper]
Balhae Kim, Jungwon Choi, Seanie Lee, Yoonho Lee, Jung-Woo Ha, and Juho Lee
NeurIPS 2022 -
Set Based Stochastic Subsampling
[paper]
Bruno Andreis, Seanie Lee, A. Tuan Nguyen, Juho Lee, Eunho Yang, and Sung Ju Hwang
ICML 2022 -
Sequential Reptile: Inter-Task Gradient Alignment for Multilingual Learning
[paper]
Seanie Lee*, Hae Beom Lee*, Juho Lee, and Sung Ju Hwang
ICLR 2022 -
Learning to Perturb Word Embeddings for Out-of-distribution QA
[paper] [code]
Seanie Lee*, Minki Kang*, Juho Lee, and Sung Ju Hwang
ACL 2021 -
Contrastive Learning with Adversarial Perturbations for Conditional Text Generation
[paper] [code]
Seanie Lee*, Dong Bok Lee*, and Sung Ju Hwang
ICLR 2021 -
Meta-GMVAE: Mixture of Gaussian VAE for Unsupervised Meta-Learning
[paper] [code]
Dong Bok Lee, Dongchan Min, Seanie Lee, and Sung Ju Hwang
ICLR 2021 -
Generating Diverse and Consistent QA pairs from Contexts with Information-Maximizing Hierarchical Conditional VAEs
[paper] [code] [video]
Dong Bok Lee*, Seanie Lee*, WooTae Jeong, Donghwan Kim, and Sung Ju Hwang
ACL 2020 -
g2pM: A Neural Grapheme-to-Phoneme Conversion Package for Mandarin Chinese Based on a New Open Benchmark Dataset
[paper] [code]
Kyubyong Park* and Seanie Lee*
Interspeech 2020
🎖 Honors and Awards
- 2023, Apple Scholars in AI/ML PhD Fellowship.
- 2022, Google Travel Grant for NeurIPS 2022.
- 2019, Silver Medal, Named Entity Recognition in the NAVER NLP Challenge.
💬 Invited Talks
- 2025.05, Seminar, Korea University, Seoul. Synthetic Data Generation for LLM Safeguards.
- 2025.04, Seminar, Hanyang University, Seoul. Synthetic Data Generation for LLM Safeguards.
- 2023.10, Tech talk, Nuremberg Institute of Technology Georg Simon Ohm. Scalable Set Encoding with Universal Mini-Batch Consistency and Unbiased Full Set Gradient Approximation.
- 2023.05, Tech talk, Samsung SDS. Scalable Set Encoding with Universal Mini-Batch Consistency and Unbiased Full Set Gradient Approximation.
- 2020.12, Tech talk, NAVER. Generating Diverse and Consistent QA pairs from Contexts with Information-Maximizing Hierarchical Conditional VAEs.