TL;DR  ·  who I am

Byung-Kwan Lee

Research Scientist at NVIDIA Building frontier VLMs that are efficient.

01DistillationVLsI · GenRecal · Masters · Hide to See
02PruningMAD · Image Token Pruning
03RLUnified RL and Imitation Learning · ZPPO
04AgentAXPO · RACE
05HarnessMid-Harness · SpatialClaw
0Papers
0First & Lead
0U.S. Patents
0On-going
More about me
Byung-Kwan Lee

Byung-Kwan Lee

Ph.D., School of Electrical Engineering at KAIST

His research explores knowledge distillation, model pruning, reinforcement learning, multimodal agent, and agent harness.

“Standing on the shoulders of giants.” — Isaac Newton

Research Interest

Building efficient, high-performing Vision-Language Models (VLMs), with focus on:

University Collaboration or Intern & Full-time

Feel free to reach out for collaboration by email if:

  • Your research interests align with mine.
  • You already have a draft idea to discuss and develop together.

For university collaboration, alignment is the primary criterion.

We are actively looking for talented internship and full-time candidates. Preferred qualifications:

  • Strong first author publications at top-tier main-track conferences (not workshops); co-first author papers are discouraged.
  • Deep expertise in one focused area: in my view, five-to-ten first author papers at top-tier venues.
  • Top-tier venues: CVPR, ICCV, ECCV, NeurIPS, ICLR, ICML, ACL, EMNLP. I am not familiar with other conferences and journals.

If you meet these criteria, please send me an email with your resume/CV attached.

Work Experience

NVIDIA Research Scientist Oct. 2025 — Current
  • LEAD Mid-Harness: Scaling Actions Between Model and Harness [Completed] — Action-level test-time scaling for terminal agents that samples and verifies candidate commands at the model–harness boundary before execution, lifting trajectory success without changing the generator or harness.
    Under Review U.S. Provisional Patent Filed Internal Tech Transfer
  • LEAD ZPPO: Zone of Proximal Policy Optimization [Completed] — RL post-training that keeps the teacher inside the prompt rather than the policy gradient, using BCQ/NCQ question reformulations and a prompt replay buffer to lift small students on hard questions without policy drift.
    Under Review U.S. Provisional Patent Filed
  • LEAD AXPO: Agent eXplorative Policy Optimization [Completed] — Agentic RL training that recovers tool usage through tool-call resampling, improving multimodal reasoning performance against larger baselines.
    NeurIPS 2026 U.S. Provisional Patent Filed Internal Tech Transfer
  • LEAD Masking Teacher and Reinforcing Student [Completed] — Mask-progressive RL distillation that gradually unmasks teacher weights and uses offline RL with accuracy & distillation rewards.
    CVPR 2026 U.S. Non-Provisional Patent Filed
  • LEAD Unified RL & Imitation Learning for VLMs [Completed] — Combines RL with adversarial imitation, using an LLM-based discriminator and multi-teacher guidance to build lightweight yet powerful VLMs.
    NeurIPS 2025 U.S. Patent Upgraded to Non-Provisional
  • LEAD GenRecal [Completed] — Cross-architecture VLM distillation via a Recalibrator that aligns heterogeneous token representations regardless of vocabulary, token splits, or index ordering.
    ECCV 2026 U.S. Patent Upgraded to Non-Provisional Internal Tech Transfer
  • LEAD VLsI: Verbalized Layers-to-Interactions [Completed] — Hardened the verbalizer-based layer-wise distillation into a production-ready recipe and led its internal tech transfer into NVIDIA's VLM development pipeline.
    CVPR 2025 U.S. Patent Upgraded to Non-Provisional Internal Tech Transfer
NVIDIA Research Intern Oct. 2024 — Oct. 2025
  • LEAD Unified RL & Imitation Learning for VLMs [Initiated] — First formulation of the RL + adversarial imitation training pipeline; initial experiments and team setup.
    U.S. Provisional Patent Filed
  • LEAD GenRecal [Initiated] — Initial design of the Recalibrator framework for cross-tokenizer VLM distillation; first proof-of-concept and internal demo.
    U.S. Provisional Patent Filed
  • LEAD VLsI: Verbalized Layers-to-Interactions [Initiated] — Layer-wise distillation using intermediate verbalizers, enabling small VLMs (2B/7B) to align with large VLMs' reasoning progression and outperform GPT-4V.
    CVPR 2025 U.S. Provisional Patent Filed

Education

KAIST Mar. 2020 — Aug. 2025
Ph.D., School of Electrical Engineering GPA 3.77 / 4.3
Dissertation: Building High-performing, Efficient-size Vision Language Models: Merge, Modify, and Distill  [Link]  [Degree Certificate]
KAIST Mar. 2018 — Feb. 2020
M.S., The Cho Chun Sik Graduate School of Green Transportation GPA 3.72 / 4.3
Thesis: Training Encoder-Attention through Fully-Connected CRFs for Efficient End-to-End Lane Detection Model  [Link]  [Degree Certificate]
Hanyang University Mar. 2014 — Feb. 2018
B.S., Mathematics and Electronic Engineering GPA 3.86 / 4.5
Thesis: Learning Rank of Evaluated Songs for Classifying Unknown Music by Convolutional Neural Network  [Link]  [Degree Certificate]

Publications

Overall21 Accepts · 3 Pending · 5 Tech Reports
CV5 CVPR · 2 ICCV · 3 ECCV · 1 ICIP
ML6 NeurIPS · 1 ICLR
NLP1 ACL · 1 EMNLP
Journal1 Pattern Recognition
In Progress16 On-going Projects
Publication Profile
  1. Mid-Harness
    Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
    Minki Kang, Ryo Hachiuma, Shaokun Zhang, Subhashree Radhakrishnan, Yonggan Fu, Jindong Jiang, Mingjie Liu, Ehsan Hosseini-Asl, Yi Dong, Yu-Chiang Frank Wang, Byung-Kwan Lee
    Under Review   [Paper][Project]
    U.S. Patent Application Filed (Provisional), NVIDIA Research
  2. ZPPO
    Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
    Byung-Kwan Lee, Ximing Lu, Shizhe Diao, Minki Kang, Saurav Muralidharan, Karan Sapra, Andrew Tao, Yu-Chiang Frank Wang, Ryo Hachiuma
    Under Review   [Paper][Project]
    U.S. Patent Application Filed (Provisional), NVIDIA Research
  3. RACE
    When to Switch: Reliable Action-Chunk Extension for Vision-Language-Action Models
    Seonghoon Yu, Dongwon Kim, HyungRok Jung, Yoonjae Baek, Byung-Kwan Lee, Suha Kwak, Jeany Son
    Under Review   [Paper][Code]
  4. AXPO
    Agent Explorative Policy Optimization for Multimodal Agentic Reasoning
    Minki Kang, Shizhe Diao, Ryo Hachiuma, Sung Ju Hwang, Pavlo Molchanov, Yu-Chiang Frank Wang, Byung-Kwan Lee
    Neural Information Processing Systems (NeurIPS), 2026   [Paper][Project]
    U.S. Patent Application Filed (Provisional), NVIDIA Research
  5. SpatialClaw
    SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
    Seokju Cho, Ryo Hachiuma, Abhishek Badki, Hang Su, Byung-Kwan Lee, Chan Hee Song, Sifei Liu, Subhashree Radhakrishnan, Seungryong Kim, Yu-Chiang Frank Wang, Min-Hung Chen
    Neural Information Processing Systems (NeurIPS), 2026   [Paper][Project][Code]
    U.S. Patent Application Filed (Provisional), NVIDIA Research
  6. Hide to See
    Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation
    Seonghoon Yu, Dongjun Nam, Byung-Kwan Lee†, Jeany Son†
    Neural Information Processing Systems (NeurIPS), 2026   [Paper][Project][Code]
  7. DSTP
    Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding
    Jiwan Kim, Kibum Kim, Wonjoong Kim, Byung-Kwan Lee, Chanyoung Park
    European Conference on Computer Vision (ECCV), 2026   [Paper][Project]
  8. GenRecal
    GenRecal: Generation after Recalibration from Large to Small Vision Language Models
    Byung-Kwan Lee, Ryo Hachiuma, Yong Man Ro, Yu-Chiang Frank Wang, Yueh-Hua Wu
    European Conference on Computer Vision (ECCV), 2026   [Paper][Project]
    U.S. Patent Application Filed (Non-Provisional), NVIDIA Research
  9. MTRS
    Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
    Byung-Kwan Lee, Yu-Chiang Frank Wang, Ryo Hachiuma
    IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026   [Paper]
    U.S. Patent Application Filed (Non-Provisional), NVIDIA Research
  10. R-TAP
    Recursive Think-Answer Process for LLMs and VLMs
    Byung-Kwan Lee*, Youngchae Chee*, Yong Man Ro
    IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 2026   [Paper][Project]
  11. RefineBench
    RefineBench: Evaluating Refinement Capability in Language Models
    Young-Jun Lee*, Seungone Kim*, Byung-Kwan Lee, Minkyeong Moon, Yechan Hwang, Jong Myoung Kim, Graham Neubig, Sean Welleck, Ho-Jin Choi
    International Conference on Learning Representations (ICLR), 2026   [Paper][Project]
    Best Runner-Up Award (Oral, Top 1%), Multi-Turn Interactions in LLMs Workshop @ NeurIPS 2025   [Link]
  12. Thanos
    Enhancing Conversational Agents with Skill-of-Mind-Infused Large Language Model
    Young-Jun Lee, Byung-Kwan Lee, Dokyong Lee, Junyoung Youn, Kyeong-Jin Oh, Yechan Hwang, Ho-Jin Choi
    Technical Report   [Paper][Code][HF Model]
  13. RIL
    Unified Reinforcement and Imitation Learning for Vision-Language Models
    Byung-Kwan Lee, Ryo Hachiuma, Yong Man Ro, Yu-Chiang Frank Wang, Yueh-Hua Wu
    Neural Information Processing Systems (NeurIPS), 2025   [Paper][Project]
    U.S. Patent Application Filed (Non-Provisional), NVIDIA Research
  14. MultiVerse
    MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
    Young-Jun Lee, Byung-Kwan Lee, Jianshu Zhang, Yechan Hwang, Byungsoo Ko, Han-Gyu Kim, Dongyu Yao, Xuankun Rong, Eojin Joo, Seung-Ho Han, Bowon Ko, Ho-Jin Choi
    IEEE/CVF International Conference on Computer Vision (ICCV), 2025   [Paper][Project]
    Workshop for Knowledge-Intensive Multimodal Reasoning, ICCV 2025   [Link]
  15. Multi-vision Sensor Understanding
    Are Vision-Language Models Truly Understanding Multi-vision Sensor?
    Sangyun Chung*, Youngjoon Yu*, Youngchae Chee, Se Yeon Kim, Byung-Kwan Lee, Yong Man Ro
    Technical Report   [Paper][Code]
  16. VLsI
    VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models
    Byung-Kwan Lee, Ryo Hachiuma, Yu-Chiang Frank Wang, Yong Man Ro†, Yueh-Hua Wu†
    IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025   [Paper][Project]
    U.S. Patent Application Filed (Non-Provisional), NVIDIA Research
  17. Phantom
    Phantom of Latent for Large Language and Vision Models
    Byung-Kwan Lee, Sangyun Chung, Chae Won Kim, Beomchan Park, Yong Man Ro
    Technical Report   [Paper][Code][HF Model]
  18. SPARK
    SPARK: Multi-Vision Sensor Perception and Reasoning Benchmark for Large-scale Vision-Language Models
    Youngjoon Yu*, Sangyun Chung*, Byung-Kwan Lee, Yong Man Ro
    Technical Report   [Paper][Code][HF Dataset]
  19. TroL
    TroL: Traversal of Layers for Large Language and Vision Models
    Byung-Kwan Lee, Sangyun Chung, Chae Won Kim, Beomchan Park, Yong Man Ro
    Empirical Methods in Natural Language Processing (EMNLP), 2024   [Paper][Code][HF Model]
  20. Meteor
    Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models
    Byung-Kwan Lee, Chae Won Kim, Beomchan Park, Yong Man Ro
    Neural Information Processing Systems (NeurIPS), 2024   [Paper][Code][HF Model]
    31st Samsung HumanTech Paper Awards in Computer Science & Engineering
  21. MoAI
    MoAI: Mixture of All Intelligence for Large Language and Vision Models
    Byung-Kwan Lee, Beomchan Park, Chae Won Kim, Yong Man Ro
    European Conference on Computer Vision (ECCV), 2024   [Paper][Code][HF Model]
  22. CoLLaVO
    CoLLaVO: Crayon Large Language and Vision mOdel
    Byung-Kwan Lee, Beomchan Park, Chae Won Kim, Yong Man Ro
    Findings of the Association for Computational Linguistics (ACL), 2024   [Paper][Code][HF Model]
    2024 KCC XAI Workshop, Best Paper Awards
  23. Causal Unsupervised Segmentation
    Causal Unsupervised Semantic Segmentation
    Junho Kim*, Byung-Kwan Lee*, Yong Man Ro
    Journal of Pattern Recognition   [Paper][Code]
  24. Adversarial Double ML
    Mitigating Adversarial Vulnerability through Causal Parameter Estimation by Adversarial Double Machine Learning
    Byung-Kwan Lee*, Junho Kim*, Yong Man Ro
    IEEE/CVF International Conference on Computer Vision (ICCV), 2023   [Paper][Code]
  25. C2Cap
    Mitigating Dataset Bias in Image Captioning through CLIP Confounder-free Captioning Network
    YeonJu Kim, Junho Kim, Byung-Kwan Lee, Sebin Shin, Yong Man Ro
    IEEE International Conference on Image Processing (ICIP), 2023   [Paper][Code]
  26. Causal Adversarial Instruments
    Demystifying Causal Features on Adversarial Examples and Causal Inoculation for Robust Network by Adversarial Instrumental Variable Regression
    Junho Kim*, Byung-Kwan Lee*, Yong Man Ro
    IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023   [Paper][Code]
  27. Masking Adversarial Damage
    Masking Adversarial Damage: Finding Adversarial Saliency for Robust and Sparse Network
    Byung-Kwan Lee*, Junho Kim*, Yong Man Ro
    IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022   [Paper][Code]
  28. Information Bottleneck
    Distilling Robust and Non-Robust Features in Adversarial Examples by Information Bottleneck
    Junho Kim*, Byung-Kwan Lee*, Yong Man Ro
    Neural Information Processing Systems (NeurIPS), 2021   [Paper][Code]
  29. Hierarchical Bayesian Defense
    Towards Adversarial Robustness of Bayesian Neural Network through Hierarchical Variational Inference
    Byung-Kwan Lee, Youngjoon Yu, Yong Man Ro
    Technical Report   [Paper][Code]

Reviewer Experience

Journal

Conference

Invited Talks & Awards

NVIDIA Stock

Current Price & Today Change TradingView
Loading NVDA price...
Loading NVDA daily chart...