Yang Sui | 隋阳

I am a Member of Technical Staff on the Microsoft AI (MAI) Superintelligence Team, where I work on AI infrastructure and machine learning systems (MLSys) for large-scale multimodal models, including LLMs and diffusion models, across training, reinforcement learning (RL), and inference.

I was a Postdoctoral Associate in the Department of Computer Science at Rice University, where I worked with Dr. Xia "Ben" Hu and the DATA Lab on systems and algorithms for large language models, diffusion models, and multimodal LLMs.

Before completing my Ph.D. at Rutgers University, where I was advised by Prof. Bo Yuan, I gained research and engineering experience across industry. In spring 2024, I was a Research Intern on the Creative Vision team at Snap Research, where I developed BitsFusion, a 1.99-bit weight-quantization method for text-to-image diffusion models. In 2022, I was a Research Intern at Media Lab, Tencent America, studying the efficiency and robustness of learned image compression and transformer models. In 2019, I was an Algorithm Engineer at JD, working on face verification and recognition. In 2018, as an R&D Intern at Baidu, I helped initiate Paddle-Lite, the mobile and edge inference framework for PaddlePaddle. The project was presented at the NeurIPS Expo, Baidu Create, and Wave Summit+.

Outside work, I enjoy basketball, soccer, DOTA/DOTA2, and World of Warcraft. I am a fan of Tracy McGrady, Stephen Curry, Lionel Messi, and PIS (YaphetS).

Email  /  Google Scholar  /  GitHub   

profile photo
Research Interests

  1. AI Infrastructure & Machine Learning Systems:

    • Large-scale training: Multidimensional parallelism across data, tensor, pipeline, sequence/context, and MoE experts; FSDP/ZeRO and distributed optimizers; mixed-precision training; activation checkpointing and offloading; fused kernels; and communication-computation overlap.

    • Reinforcement learning: PPO/GRPO-style post-training; distributed actor, rollout, reference, critic, and reward-model orchestration; colocated and disaggregated execution; asynchronous rollouts; dynamic batching and sequence packing; weight synchronization; CPU/GPU offloading; and resource-aware scheduling and load balancing.

    • Inference and serving: Prefill-decode (PD) disaggregation; continuous and chunked batching; paged KV caches and prefix caching; KV-cache and expert offloading; CUDA graphs; kernel fusion; speculative and parallel decoding; request scheduling; and distributed model serving.

    • System co-design: Algorithm-system and algorithm-hardware co-design for scalable, high-performance AI.

  2. Efficient AI:

    • Model compression: Quantization, pruning, low-rank approximation, knowledge distillation, and structured sparsity.

    • Token and reasoning efficiency: Token compression, efficient reasoning, dynamic computation, sparse computation, and memory-efficient attention.

News
  • 08/2026: One paper was accepted to EMNLP 2026.
  • 04/2026: One paper was accepted to ACL 2026 Findings.
  • 12/2025: I joined the Microsoft AI (MAI) Superintelligence Team as a Member of Technical Staff.
  • 09/2025: Two papers were accepted to the Thirty-ninth Annual Conference on Neural Information Processing Systems (NeurIPS 2025).
  • 08/2025: One paper was accepted to TMLR 2025.
  • 05/2025: One paper was accepted to TMLR 2025.
  • 04/2025: Our work DFloat11: Lossless Compression for LLM was featured by 新智元 and 机器之心.
  • 04/2025: Our survey Stop Overthinking was featured by 新智元.
  • 03/2025: We released Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models, which surveys methods for improving the computational efficiency of LLM reasoning.
  • 02/2025: Three papers were accepted to CVPR 2025. TopV was my first project as an advisor—congratulations to Cheng and the team!
  • 10/2024: I delivered a guest lecture, "Model Compression: Pruning, Quantization, and Recent Advances," for CSCE 689: Special Topics in Generative AI at Texas A&M University.
  • 10/2024: I joined the Department of Computer Science at Rice University as a Postdoctoral Associate.
  • 09/2024: One paper was accepted to NeurIPS 2024.
  • 09/2024: One paper was accepted to EMNLP 2024 Findings.
  • 09/2024: One paper was accepted to ASP-DAC 2025.
  • 07/2024: One paper was accepted to BMVC 2024.
  • 07/2024: One paper was accepted to ECCV 2024.
  • 05/2024: One paper was accepted to IEEE Transactions on Neural Networks and Learning Systems (TNNLS).
  • 04/2024: I received the Paul Panayotatos Scholarship at Rutgers University.
  • 03/2024: One paper was accepted to DAC 2024.
  • 02/2024: I joined the Creative Vision team at Snap Research as a Research Intern in Santa Monica.
  • 12/2023: Two papers were accepted as posters at DCC 2024.
  • 10/2023: One paper was accepted to HPCA 2024.
  • 09/2023: One paper was accepted to ICCAD 2023.
  • 07/2023: I was invited to present "Efficient Diffusion Models and Large Language Models: Quantization, Pruning, and LoRA." (Video)
  • 07/2023: One paper was selected for a spotlight presentation at the ICML 2023 Neural Compression Workshop.
  • 06/2023: One paper was accepted to IROS 2023.
  • 03/2023: One paper was accepted to ISCA 2023.
  • 02/2023: One paper was accepted to DAC 2023.
  • 02/2023: One paper received the Best Paper Runner-Up Award and was selected for an oral presentation at the AAAI 2023 DCAA Workshop.
  • 11/2022: Two papers were accepted to AAAI 2023 for oral presentation.
  • 10/2022: I presented a poster at the IBM IEEE CAS/EDS 5th AI Compute Symposium at the IBM Thomas J. Watson Research Center in Yorktown Heights, New York.
  • 05/2022: I joined Media Lab, Tencent America as a remote Research Intern.
  • 03/2022: One paper was accepted to CVPR 2022.
  • 09/2021: One paper was accepted to NeurIPS 2021.
  • 09/2021: One paper was accepted to ICCAD 2021.
  • 03/2021: One paper was accepted to ISCA 2021.
  • 02/2021: One paper was accepted to CVPR 2021.
Publications & Preprints
(*: Equal Contribution; ‡: Corresponding Author/Project Lead)

2026

SampKV saliency-aware mixed-precision KV cache overview SampKV: Saliency-Aware Mixed-Precision KV Cache for Efficient Diffusion Language Model Inference
Zhenrui Wang, Haodong Chang, Yang Sui‡, Rongjian Liang, Bo Yuan, Jiang Hu
[EMNLP 2026] The 2026 Conference on Empirical Methods in Natural Language Processing
TIDE expert-routing and I/O-aware expert-offload overview TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload
Zhiben Chen, Youpeng Zhao, Yang Sui, Jun Wang, Yuzhang Shang
[arXiv 2026] arXiv preprint
arXiv
GitHub
Token-compression motivation across image, audio, and video inputs A Survey of Token Compression for Efficient Multimodal Large Language Models
Kele Shao*, Keda Tao*, Kejia Zhang, Sicheng Feng, Mu Cai, Yuzhang Shang, Haoxuan You, Can Qin, Yang Sui, Huan Wang
[TMLR 2026] Transactions on Machine Learning Research
arXiv
OpenReview
GitHub
AutoL2S two-stage long-short reasoning training pipeline AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models
Feng Luo, Yu-Neng Chuang, Guanchu Wang, Hoang Anh Duy Le, Shaochen Zhong, Hongyi Liu, Jiayi Yuan, Yang Sui, Vladimir Braverman, Vipin Chaudhary, Xia Hu
[ACL 2026 Findings] Findings of the Association for Computational Linguistics: ACL 2026
arXiv
GitHub
TensorGS tensor-core acceleration framework for 3D Gaussian splatting Accelerating 3D Gaussian Splatting using Tensor Cores
Sheng Li, Yang Sui, Yue Wu, Zhuoran Song, Bo Yuan, Xulong Tang, Yue Dai
[arXiv 2026] arXiv preprint
arXiv
TAPE temporal-aware pruning framework for diffusion-based video generation Temporal Aware Pruning for Efficient Diffusion-based Video Generation
Sheng Li, Yang Sui, Junhao Ran, Bo Yuan, Yue Dai, Xulong Tang
[arXiv 2026] arXiv preprint
arXiv
ATA attention-guided and action-guided inference framework ATA: Bridging Implicit Reasoning with Attention-Guided and Action-Guided Inference for Vision-Language Action Models
Cheng Yang, Jianhao Jiao, Lingyi Huang, Jinqi Xiao, Zhexiang Tang, Yu Gong, Yibiao Ying, Yang Sui, Jintian Lin, Wen Huang, Bo Yuan
[ICRA 2026] IEEE International Conference on Robotics and Automation
arXiv

2025

3DSP 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float
Tianyi Zhang, Mohsen Hariri, Shaochen Zhong, Vipin Chaudhary, Yang Sui, Xia Hu, Anshumali Shrivastava
[NeurIPS 2025] The Thirty-ninth Annual Conference on Neural Information Processing Systems
GitHub
DFloat11 Models in Hugging Face
Reddit
新智元
机器之心
arXiv
3DSP HoliTom: Holistic Token Merging for Fast Video Large Language Models
Kele Shao, Keda Tao, Can Qin, Haoxuan You, Yang Sui, Huan Wang
[NeurIPS 2025] The Thirty-ninth Annual Conference on Neural Information Processing Systems
arXiv
GitHub
3DSP Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Hanjie Chen, Xia Hu
[TMLR 2025] Transactions on Machine Learning Research
LinkedIn Recommendation
LinkedIn Recommendation
X
新智元
Daily Paper in Hugging Face
arXiv
GitHub
3DSP TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model
Cheng Yang, Yang Sui‡, Jinqi Xiao, Lingyi Huang, Yu Gong, Chendi Li, Jinghua Yan, Yu Bai, Ponnuswamy Sadayappan, Xia Hu, Bo Yuan
[CVPR 2025] The IEEE/CVF Computer Vision and Pattern Recognition Conference
arXiv
3DSP SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device
Yushu Wu, Zhixing Zhang, Yanyu Li, Yanwu Xu, Anil Kag, Yang Sui, Huseyin Coskun, Ke Ma, Aleksei Lebedev, Ju Hu, Dimitris Metaxas, Yanzhi Wang, Sergey Tulyakov, Jian Ren
[CVPR 2025] The IEEE / CVF Computer Vision and Pattern Recognition Conference
arXiv
Project Page
3DSP DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
Keda Tao, Can Qin, Haoxuan You, Yang Sui, Huan Wang
[CVPR 2025] The IEEE / CVF Computer Vision and Pattern Recognition Conference
arXiv
GitHub
LowDiff progressive low-resolution conditioning architecture LowDiff: Efficient Diffusion Sampling with Low-Resolution Condition
Jiuyi Xu, Qing Jin, Meida Chen, Andrew Feng, Yang Sui, Yangming Shi
[arXiv 2025] arXiv preprint
arXiv
VidKV mixed-precision KV-cache quantization framework Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models
Keda Tao, Haoxuan You, Yang Sui, Can Qin, Huan Wang
[arXiv 2025] arXiv preprint
arXiv
GitHub
iTAP incremental task-graph partitioning example iTAP: An Incremental Task Graph Partitioner for Task-parallel Static Timing Analysis
Boyang Zhang, Che Chang, Cheng-Hsiang Chiu, Dian-Lun Lin, Yang Sui, Chih-Chun Chang, Yi-Hua Chung, Wan-Luan Lee, Zizheng Guo, Yibo Lin, Tsung-Wei Huang
[ASP-DAC 2025] 30th Asia and South Pacific Design Automation Conference
DOI
PDF
EcoSpa coupled-sparsity training overview EcoSpa: Efficient Transformer Training with Coupled Sparsity
Jinqi Xiao, Cheng Luo, Lingyi Huang, Cheng Yang, Yang Sui, Huy Phan, Xiao Zang, Yibiao Ying, Zhexiang Tang, Anima Anandkumar, Bo Yuan
[NeurIPS 2025 Workshop] The First Workshop on Efficient Reasoning
arXiv
OpenReview
Who Routes the Router project preview Who Routes the Router: Rethinking the Evaluation of LLM Routing Systems
Jiayi Yuan*, Yifan Lu*, Rixin Liu, Yu-Neng Chuang, Hongyi Liu, Shaochen Zhong, Yang Sui, Guanchu Wang, Jiarong Xing, Xia Hu
[NeurIPS 2025 Workshop] Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling
OpenReview
GitHub
3DSP DisDet: Exploring Detectability of Backdoor Attack on Diffusion Models
Yang Sui, Huy Phan, Jinqi Xiao, Tianfang Zhang, Zijie Tang, Cong Shi, Yan Wang, Yingying Chen, Bo Yuan
[TMLR 2025] Transactions on Machine Learning Research
arXiv
3DSP Confident or Seek Stronger: Exploring Uncertainty-Based On-device LLM Routing From Benchmarking to Generalization
Yu-Neng Chuang, Leisheng Yu, Guanchu Wang, Lizhe Zhang, Zirui Liu, Xuanting Cai, Yang Sui, Vladimir Braverman, Xia Hu
[arXiv]
arXiv

2024

3DSP BitsFusion: 1.99 bits Weight Quantization of Diffusion Model
Yang Sui, Yanyu Li, Anil Kag, Yerlan Idelbayev, Junli Cao, Ju Hu, Dhritiman Sagar, Bo Yuan, Sergey Tulyakov, Jian Ren
[NeurIPS 2024] The Thirty-eighth Annual Conference on Neural Information Processing Systems
Daily Paper in Hugging Face
LinkedIn Recommendation
Reddit Discussion
arXiv
Project Page
Video
3DSP MoE-I²: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition
Cheng Yang*, Yang Sui*, Jinqi Xiao, Lingyi Huang, Yu Gong, Yuanlin Duan, Wenqi Jia, Miao Yin, Yu Cheng, Bo Yuan
[EMNLP 2024 Findings] The 2024 Conference on Empirical Methods in Natural Language Processing
3DSP Transferable Learned Image Compression-Resistant Adversarial Perturbations
Yang Sui, Zhuohang Li, Ding Ding, Xiang Pan, Xiaozhong Xu, Shan Liu, Zhenzhong Chen
[BMVC 2024] The 35th British Machine Vision Conference, 2024
arXiv
3DSP Clean & Compact: Efficient Data-Free Backdoor Defense with Model Compactness
Huy Phan, Jinqi Xiao, Yang Sui, Tianfang Zhang, Zijie Tang, Cong Shi, Yan Wang, Yingying Chen, Bo Yuan
[ECCV 2024] The 18th European Conference on Computer Vision ECCV 2024
PDF
3DSP Co-Exploring Sparsification and Low-Rank Decomposition for Compact DNNs
Yang Sui, Miao Yin, Yu Gong, Bo Yuan
[TNNLS] IEEE Transactions on Neural Networks and Learning Systems
PDF
Neural SLAM algorithm and hardware co-design overview Algorithm and Hardware Co-Design for Energy-Efficient Neural SLAM
Lingyi Huang, Cheng Yang, Yu Gong, Yang Sui, Xiao Zang, Anthony Goeckner, Qi Zhu, Bo Yuan
[DAC 2024] Proceedings of the 61st ACM/IEEE Design Automation Conference
Paper
3DSP MOPED: Efficient Motion Planning Engine with Flexible Dimension Support
Lingyi Huang, Yu Gong, Yang Sui, Xiao Zang, Bo Yuan
[HPCA 2024] In Proceedings of the IEEE International Symposium on High-Performance Computer Architecture, 2024
PDF

2023

3DSP In-Sensor Radio Frequency Computing for Energy-Efficient Intelligent Radar
Yang Sui, Minning Zhu, Lingyi Huang, Chung-Tse Michael Wu, Bo Yuan
[ICCAD 2023] In Proceedings of the IEEE/ACM International Conference on Computer-Aided Design, 2023
PDF
3DSP Corner-to-Center Long-range Context Model for Efficient Learned Image Compression
Yang Sui, Ding Ding, Xiang Pan, Xiaozhong Xu, Shan Liu, Bo Yuan, Zhenzhong Chen
[JVCI] Journal of Visual Communication and Image Representation
PDF
3DSP Reconstruction Distortion of Learned Image Compression with Imperceptible Perturbations
Yang Sui, Zhuohang Li, Ding Ding, Xiang Pan, Xiaozhong Xu, Shan Liu, Zhenzhong Chen
[DCC 2024] In Proceedings of the Data Compression Conference, 2024
[NCW@ICML 2023] Neural Compression: From Information Theory to Applications
Spotlight presentation
PDF
Website
3DSP DynGMP: Graph Neural Network-based Motion Planning in Unpredictable Dynamic Environments
Wenjin Zhang, Xiao Zang, Lingyi Huang, Yang Sui, Jingjin Yu, Yingying Chen, Bo Yuan
[IROS 2023] In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems, 2023
PDF
3DSP ETTE: Efficient Tensor-Train-based Computing Engine for Deep Neural Networks
Yu Gong, Miao Yin, Lingyi Huang, Jinqi Xiao, Yang Sui, Chunhua Deng, Bo Yuan
[ISCA 2023] In Proceedings of the 50th International Symposium on Computer Architecture, 2023
PDF
3DSP DSPIMM: Digital Sparse In-Memory Matrix Vector Multiplier for Communication Applications
Amitesh Sridharan, Fan Zhang, Yang Sui, Bo Yuan, Deliang Fan
[DAC 2023] In Proceedings of the 60th ACM/IEEE Design Automation Conference, 2023
PDF
3DSP Towards Sparse and Low-rank Neural Networks with Hybrid Compression
Yang Sui, Wanzhao Yang, Miao Yin, Yu Gong, Bo Yuan
[DCAA@AAAI 2023] DCAA, The First Workshop on DL-Hardware Co-Design for AI Acceleration
Website
Award

Best Paper Runner-Up Award

3DSP CSTAR: Towards Compact and STructured Deep Neural Networks with Adversarial Robustness
Huy Phan, Miao Yin, Yang Sui, Bo Yuan, Saman Zonouz
[AAAI 2023] In Proceedings of the AAAI Conference on Artificial Intelligence, 2023
PDF

Oral

3DSP HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural Networks
Jinqi Xiao, Chengming Zhang, Yu Gong, Miao Yin, Yang Sui, Lizhi Xiang, Dingwen Tao, Bo Yuan
[AAAI 2023] In Proceedings of the AAAI Conference on Artificial Intelligence, 2023
PDF

Oral

3DSP ELRT: Efficient Low-Rank Training for Compact Convolutional Neural Networks
Yang Sui, Miao Yin, Yu Gong, Jinqi Xiao, Huy Phan, Bo Yuan
[arXiv]
PDF

2022

3DSP Algorithm and Hardware Co-Design of Energy-Efficient LSTM Networks for Video Recognition with Hierarchical Tucker Tensor Decomposition
Yu Gong, Miao Yin, Lingyi Huang, Chunhua Deng, Yang Sui, Bo Yuan
[arXiv]
PDF
3DSP HODEC: Towards Efficient High-Order DEcomposed Convolutional Neural Networks
Miao Yin, Yang Sui, Wanzhao Yang, Xiao Zang, Yu Gong, Bo Yuan
[CVPR 2022] In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022
PDF

2021

3DSP CHIP: CHannel Independence-based Pruning for Compact Neural Networks
Yang Sui, Miao Yin, Yi Xie, Huy Phan, Saman Zonouz, Bo Yuan
[NeurIPS 2021] Advances in Neural Information Processing Systems 34, 2021
PDF
3DSP Algorithm and Hardware Co-design for Deep Learning-powered Channel Decoder: A Case Study
Boyang Zhang*, Yang Sui*, Lingyi Huang, Siyu Liao, Chunhua Deng, Bo Yuan
[ICCAD 2021] In Proceedings of the IEEE/ACM International Conference On Computer Aided Design, 2021
PDF
3DSP GoSPA: An Energy-efficient High-performance Globally Optimized SParse Convolutional Neural Network Accelerator
Chunhua Deng, Yang Sui, Siyu Liao, Xuehai Qian, Bo Yuan
[ISCA 2021] In Proceedings of the ACM/IEEE 48th Annual International Symposium on Computer Architecture, 2021
PDF
3DSP Towards Efficient Tensor Decomposition-Based DNN Model Compression With Optimization Framework
Miao Yin, Yang Sui, Siyu Liao, Bo Yuan
[CVPR 2021] In Proceedings of The IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021
PDF
Projects

2018

3DSP Paddle-Lite (Paddle-Mobile)
Initial contributors: Yang Sui, Ruilong Liu, Jiaying Zhao, Wang Liu, Yonghui Li.
[Baidu] The authors contributed almost equally to this work.
Paddle-Lite GitHub

Paddle-Lite, the successor to Paddle-Mobile, is an open-source deep learning inference framework for mobile, embedded, and IoT devices. It supports PaddlePaddle models and models converted from other frameworks. The project was presented at the NeurIPS Expo, Baidu Create, and Wave Summit+.

Previous "Efficient Deep Learning Reading Group" Sessions:

Professional Services
  • Program Committee and Reviewing:
    • NeurIPS'22, 23, 24
    • ICLR'24
    • ICML'22, 23, 24
    • CVPR'22, 23, 24
    • ICCV'23
    • ECCV'22, 24
    • KDD'23
    • AAAI'22, 23, 24, 25
    • IROS'23
    • TNNLS
Teaching Experience
  • Teaching Assistant at Rutgers University
    • 14:332:351 - Programming Methodology II, Fall 2020
      Instructor: Prof. Saman Zonouz   

    • 14:332:351 - Programming Methodology II, Fall 2023
      Instructor: Prof. Yao Liu   

Supervised and Collaborated Students
  • Daniel Gu, undergraduate student at Rice University
    Topic: LLM quantization.

  • Eric Chien, master's student at Rice University
    Topic: Multimodal LLMs.

  • Cheng Yang, Ph.D. student at Rutgers University
    Topic: Token pruning for multimodal LLMs.
    [CVPR 2025] TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model.

  • Wenjin Zhang, Ph.D. at Rutgers University
    Topic: Quantization for model compression.

  • Justin Ding, master's student at Rutgers University
    Topic: Model pruning for a Graduate Special Problems project.

  • Linqi Xiao, master's student at Rutgers University
    Topic: Error-correcting codes for a Graduate Special Problems project.

  • Srinihar Bondalapati, master's student at Rutgers University
    Topic: Quantization for model compression.

  • Yue Wang, master's student at Rutgers University
    Topic: Dataset distillation and model pruning for a Graduate Special Problems project.

  • Veena Vrushi, undergraduate student at Rutgers University
    Topic: Deep learning through the Project SUPER research program.

  • Vijay Maddila, graduate student at Rutgers University
    Topic: Large language models.

  • Ayan Patel, high school student at High Technology High School
    Topic: Deep learning.

  • Zhiyu Chen, master's student at Rutgers University
    Topic: Large language models for a Graduate Special Problems project.

Talks
  • Efficient Diffusion Models and Large Language Models: Quantization, Pruning, and LoRA. (Video)
    July, 2023.

  • Model Compression: Pruning, Quantization, and Recent Advances.
    Texas A&M University, CSCE 689 Special Topics: Generative AI, Instructor: Dr. Zhengzhong Tu, October, 2024.

  • Efficient Multimodal LLMs and Efficient Large Reasoning Models.
    Rice University, COMP 652 001: Natural Language Processing, Instructor: Dr. Hanjie Chen, April, 2025.

  • Efficient MLLMs and LRMs: Token Pruning in Multimodal Large Language Models and Efficient Reasoning in Large Reasoning Models.
    Texas A&M University, CSCE 753 Computer Vision and Robot Perception, Instructor: Dr. Zhengzhong Tu, April, 2025.

Honors & Awards
Hobbies & Interests

Outside work, I follow basketball, soccer, Formula 1, and snooker. I am a fan of Tracy McGrady, Stephen Curry, Lionel Messi, and PIS (YaphetS).

I have played DOTA/DOTA2, World of Warcraft, and Warcraft III for many years. My favorite classes and races include Balance Druid and Havoc Demon Hunter in World of Warcraft, and Night Elf in Warcraft III. A few memorable gaming milestones:

  • World of Warcraft
    • Ranked among the world's top 10 Balance Druid DPS players for Heroic Flamebender Ka'graz in Blackrock Foundry, as reported by Warcraft Logs in 2015.
    • Led a 15-player team that completed the Heroic Highmaul raid in Warlords of Draenor in 2014.
  • DOTA
    • Member of the five-player DOTA team at Jilin City No. 1 High School, 2011.
  • Hearthstone Legend
    • Reached Legend rank and a ladder ranking of 147 in 2016.
  • Diablo III
    • Reached ladder rank 698 as a Witch Doctor in Season 9, 2017.

I enjoy R&B and classical music, especially works by Chopin, Bach, and Paganini.

Kim Tae-yeon was an important role model for me in high school and a source of encouragement during a difficult period of my life.

Xiaolan—whose name was inspired by "Detective Conan"—is my playful and affectionate grey-and-white cat. She even appears in my NeurIPS 2021 paper, "CHIP." Say hi, Xiaolan!



*Last updated on 08/2026*
Inspired by Jon Barron