Biography

I am a Ph.D. candidate in Computer Science at the University of Maryland, College Park, advised by Prof. Ang Li. I am currently a Student Researcher at Google and previously worked at ByteDance Seed, Tencent AI Lab, and JD Explore Academy, where I focused on efficient model training and large-scale natural language systems.

My research spans theory, algorithm, system, and application in the pursuit of efficient foundation models. I study how computation, parameters, and modalities interact inside large models to uncover the principles that govern their capacity, redundancy, and generalization. Based on these insights, I develop methods for model compression, adaptive inference, parameter-efficient fine-tuning, and modality-aware optimization, with the goal of making powerful models substantially more efficient without sacrificing capability.

I further explore system- and hardware-aware techniques that make these advances practical in real-world deployment. More recently, I have been extending this line of work to unified multimodal architectures across language and vision, with applications in dense retrieval and generation. Broadly, my goal is to bridge model understanding, algorithm design, and efficient deployment, enabling foundation models that are not only stronger, but also more economical, reliable, and widely usable.

News

Show earlier news (2022 – 2024)...

Research Experience

Google DeepMind & Ads
Student Researcher · Mountain View, CA
Efficient Post-Training
ByteDance Seed
Research Intern · San Jose, CA
Multimodal Foundation Models
Tencent AI Lab
Research Intern · Bellevue, WA
Sparse and Efficient Large Language Models
JD Explore Academy
Research Intern · Beijing, China
Efficient and Adaptive Methods for NLP

Selected Publications

1. Shwai He, Chaorui Deng, Ang Li, Shen Yan, "Understanding and Harnessing Sparsity for Unified Multimodal Models", Transactions on Machine Learning Research (TMLR). Project Paper Code
BibTeX
@article{he2026understanding,
  title={Understanding and Harnessing Sparsity for Unified Multimodal Models},
  author={He, Shwai and Deng, Chaorui and Li, Ang and Yan, Shen},
  journal={Transactions on Machine Learning Research (TMLR)},
  year={2026}
}
2. Shwai He, Guoheng Sun, Haichao Zhang, Yun Fu, Ang Li, "Demystifying When Pruning Works via Representation Hierarchies", Proceedings of the Forty-third International Conference on Machine Learning (ICML 2026). Project Paper Code 💭 Reflection
BibTeX
@inproceedings{he2026demystifying,
  title={Demystifying When Pruning Works via Representation Hierarchies},
  author={He, Shwai and Sun, Guoheng and Zhang, Haichao and Fu, Yun and Li, Ang},
  booktitle={Proceedings of the 43rd International Conference on Machine Learning (ICML)},
  year={2026}
}
3. Shwai He, Ang Li, "Disentangling Representation Evolution in Transformers through Directional Decomposition", Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026).
4. Shwai He, Weilin Cai, Jiayi Huang, Ang Li, "Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts", Proceedings of the Fourteenth International Conference on Learning Representations (ICLR 2026). Project Paper Code 💭 Reflection
BibTeX
@inproceedings{he2026capacity,
  title={Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts},
  author={He, Shwai and Cai, Weilin and Huang, Jiayi and Li, Ang},
  booktitle={Proceedings of the 14th International Conference on Learning Representations (ICLR)},
  year={2026}
}
5. Shwai He*, Guoheng Sun*, Zheyu Shen, Ang Li, "Uncovering the Redundancy in Transformers via a Unified Study of Layer Dropping", Transactions on Machine Learning Research (TMLR). Project Paper Code HF 💭 Reflection
BibTeX
@article{he2026uncovering,
  title={Uncovering the Redundancy in Transformers via a Unified Study of Layer Dropping},
  author={He, Shwai and Sun, Guoheng and Shen, Zheyu and Li, Ang},
  journal={Transactions on Machine Learning Research (TMLR)},
  year={2026}
}
6. Yibin Lei, Shwai He*, Ang Li, Andrew Yates, "Making Large Language Models Efficient Dense Retrievers", Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026). Paper Code Models
BibTeX
@inproceedings{lei2026making,
  title={Making Large Language Models Efficient Dense Retrievers},
  author={Lei, Yibin and He, Shwai and Li, Ang and Yates, Andrew},
  booktitle={Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL)},
  year={2026}
}
7. Shwai He, Tao Ge, Guoheng Sun, Bowei Tian, Xiaoyang Wang, Dong Yu, "Router-Tuning: A Simple and Effective Approach for Dynamic Depth", Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025). Project Paper Code 💭 Reflection
BibTeX
@inproceedings{he2025router,
  title={Router-Tuning: A Simple and Effective Approach for Dynamic Depth},
  author={He, Shwai and Ge, Tao and Sun, Guoheng and Tian, Bowei and Wang, Xiaoyang and Yu, Dong},
  booktitle={Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP)},
  year={2025}
}
8. Shwai He*, Daize Dong*, Liang Ding, Ang Li, "Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques", Transactions on Machine Learning Research (TMLR). Paper Code
BibTeX
@article{he2025towards,
  title={Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques},
  author={He, Shwai and Dong, Daize and Ding, Liang and Li, Ang},
  journal={Transactions on Machine Learning Research (TMLR)},
  year={2025}
}
9. Shwai He, Run-Ze Fan, Liang Ding, Li Shen, Tianyi Zhou, Dacheng Tao, "Merging Experts into One: Improving Computational Efficiency of Mixture of Experts", Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP 2023 Oral). Paper Code 💭 Reflection
BibTeX
@inproceedings{he2023merging,
  title={Merging Experts into One: Improving Computational Efficiency of Mixture of Experts},
  author={He, Shwai and Fan, Run-Ze and Ding, Liang and Shen, Li and Zhou, Tianyi and Tao, Dacheng},
  booktitle={Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP)},
  year={2023}
}
10. Shwai He, Liang Ding, Daize Dong, Boan Liu, Fuqiang Yu, Dacheng Tao, "PAD-Net: An Efficient Framework for Dynamic Networks", Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL 2023). Paper Code 💭 Reflection
BibTeX
@inproceedings{he2023pad,
  title={PAD-Net: An Efficient Framework for Dynamic Networks},
  author={He, Shwai and Ding, Liang and Dong, Daize and Liu, Boan and Yu, Fuqiang and Tao, Dacheng},
  booktitle={Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL)},
  year={2023}
}
11. Shwai He, Liang Ding, Daize Dong, Miao Zhang, Dacheng Tao, "SparseAdapter: An Easy Approach for Improving the Parameter-Efficiency of Adapters", Findings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP 2022). Paper Code 💭 Reflection
BibTeX
@inproceedings{he2022sparseadapter,
  title={SparseAdapter: An Easy Approach for Improving the Parameter-Efficiency of Adapters},
  author={He, Shwai and Ding, Liang and Dong, Daize and Zhang, Miao and Tao, Dacheng},
  booktitle={Findings of the Association for Computational Linguistics: EMNLP 2022},
  year={2022}
}
12. Shwai He, Chenbo Jiang, Daize Dong, Liang Ding, "SD-Conv: Towards the Parameter-Efficiency of Dynamic Convolution", IEEE/CVF Winter Conference on Applications of Computer Vision, 2023 (WACV 2023). Paper
BibTeX
@inproceedings{he2023sdconv,
  title={SD-Conv: Towards the Parameter-Efficiency of Dynamic Convolution},
  author={He, Shwai and Jiang, Chenbo and Dong, Daize and Ding, Liang},
  booktitle={IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)},
  year={2023}
}
13. Shwai He, Shi Gu, "Multi-modal Attention Network for Stock Movements Prediction", the AAAI-22 Workshop on Knowledge Discovery from Unstructured Data in Financial Service (KDF 2022). Paper
BibTeX
@article{he2022multimodal,
  title={Multi-modal Attention Network for Stock Movements Prediction},
  author={He, Shwai and Gu, Shi},
  journal={AAAI-22 Workshop on Knowledge Discovery from Unstructured Data in Financial Service},
  year={2022}
}
14. Weilin Cai, Le Qin, Shwai He, Junwei Cui, Ang Li, Jiayi Huang, "DualSparse-MoE: Coordinating Tensor/Neuron-Level Sparsity with Expert Partition and Reconstruction", Proceedings of the Forty-third International Conference on Machine Learning (ICML 2026). Paper
BibTeX
@inproceedings{cai2026dualsparse,
  title={DualSparse-MoE: Coordinating Tensor/Neuron-Level Sparsity with Expert Partition and Reconstruction},
  author={Cai, Weilin and Qin, Le and He, Shwai and Cui, Junwei and Li, Ang and Huang, Jiayi},
  booktitle={Proceedings of the 43rd International Conference on Machine Learning (ICML)},
  year={2026}
}
15. Haichao Zhang, Wenhao Chai, Shwai He, Ang Li, Yun Fu, "Dense Video Understanding with Inter-tokenization Acceleration", European Conference on Computer Vision (ECCV 2026).
BibTeX
@inproceedings{zhang2026dense,
  title={Dense Video Understanding with Inter-tokenization Acceleration},
  author={Zhang, Haichao and Chai, Wenhao and He, Shwai and Li, Ang and Fu, Yun},
  booktitle={European Conference on Computer Vision (ECCV)},
  year={2026}
}
16. Guoheng Sun, Ziyao Wang, Bowei Tian, Meng Liu, Zheyu Shen, Shwai He, Yexiao He, Wanghao Ye, Yiting Wang, Ang Li, "CoIn: Counting the Invisible Reasoning Tokens in Commercial Opaque LLM APIs", Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026). Paper
BibTeX
@inproceedings{sun2026coin,
  title={CoIn: Counting the Invisible Reasoning Tokens in Commercial Opaque LLM APIs},
  author={Sun, Guoheng and Wang, Ziyao and Tian, Bowei and Liu, Meng and Shen, Zheyu and He, Shwai and He, Yexiao and Ye, Wanghao and Wang, Yiting and Li, Ang},
  booktitle={Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP)},
  year={2026}
}
17. Chenbo Jiang, Jie Yang, Shwai He, Yu-Kun Lai, Lin Gao, "NeuralSlice: Neural 3D Triangle Mesh Reconstruction via Slicing 4D Tetrahedral Meshes", Proceedings of the 40th International Conference on Machine Learning, 2023 (ICML 2023). Paper Code
BibTeX
@inproceedings{jiang2023neuralslice,
  title={NeuralSlice: Neural 3D Triangle Mesh Reconstruction via Slicing 4D Tetrahedral Meshes},
  author={Jiang, Chenbo and Yang, Jie and He, Shwai and Lai, Yu-Kun and Gao, Lin},
  booktitle={Proceedings of the 40th International Conference on Machine Learning (ICML)},
  year={2023}
}
18. Changtong Zan, Keqin Peng, Liang Ding, Baopu Qiu, Boan Liu, Shwai He, Qingyu Lu, Zheng Zhang, Chuang Liu, Weifeng Liu, Yibing Zhan, Dacheng Tao, "Vega-MT: The JD Explore Academy Translation System for WMT", The Conference on Machine Translation, 2022 (WMT 2022). Paper
BibTeX
@inproceedings{zan2022vega,
  title={Vega-MT: The JD Explore Academy Translation System for WMT},
  author={Zan, Changtong and Peng, Keqin and Ding, Liang and Qiu, Baopu and Liu, Boan and He, Shwai and Lu, Qingyu and Zhang, Zheng and Liu, Chuang and Liu, Weifeng and Zhan, Yibing and Tao, Dacheng},
  booktitle={Proceedings of the Seventh Conference on Machine Translation (WMT)},
  year={2022}
}

Teaching

  • 2025: Teaching Assistant for CMSC 250 (Discrete Structures) and CMSC 320 (Introduction to Data Science)
  • 2024: Teaching Assistant for CMSC 351 (Algorithms)

Last updated: August 21, 2026