Skip to content

Runxin Xu许润昕

Researcher at DeepSeek
Large language models · Reasoning · Post-training

runxinxu AT gmail DOT com

Portrait of Runxin Xu

About

I am a researcher at DeepSeek, where I have been deeply involved in the DeepSeek model series — V1 / V2 / V3 / V3.1 / V3.2 / V4 / V4.1, the R1 reasoning models, and the Math, Coder, and MoE lines.

My long-term research interest lies in AGI: pushing the boundaries of machine intelligence with methods that are simple, scalable, and effective. I keep reminding myself to re-read The Bitter Lesson.

Before DeepSeek, I was a master’s student at the Institute of Computational Linguistics, School of EECS, Peking University, advised by Baobao Chang and Zhifang Sui. Prior to that I received my bachelor’s degree from Shanghai Jiao Tong University.

Selected Publications

  1. arXiv 2026

    DeepSeek-AI, …, Runxin Xu, …

    Single-token decode FLOPs versus context length across DeepSeek generations
  2. arXiv 2026

    DeepSeek-AI, …, Runxin Xu, …

    Overall architecture of the DeepSeek-V4 series
  3. Nature 2025

    DeepSeek-AI, …, Runxin Xu (core contributor), …

    Cover Article
    AIME accuracy of DeepSeek-R1-Zero rising over RL training
  4. arXiv 2025

    DeepSeek-AI, …, Runxin Xu, …

    DeepSeek Sparse Attention instantiated under MLA
  5. ICLR 2025

    Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu, Baobao Chang

    Competition hierarchy and mathematical domains covered by Omni-MATH
  6. arXiv 2024

    DeepSeek-AI, …, Runxin Xu, …

    Benchmark performance of DeepSeek-V3 and its counterparts
  7. arXiv 2024

    DeepSeek-AI, …, Runxin Xu, …

    DeepSeek-V2 architecture: MLA and DeepSeekMoE
  8. arXiv 2024

    DeepSeek-AI, …, Runxin Xu, …

    DeepSeek-Coder-V2 performance on math and code benchmarks
  9. arXiv 2024

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, Daya Guo

    Top-1 accuracy on the competition-level MATH benchmark over time
  10. arXiv 2024

    DeepSeek-AI, …, Runxin Xu, …

    Training loss curves under different learning-rate schedulers
  11. ACL 2024

    Damai Dai, Chengqi Deng, Chenggang Zhao, Runxin Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, Wenfeng Liang

    Conventional top-2 routing vs. fine-grained expert segmentation and shared expert isolation
  12. ACL 2024

    Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, Zhifang Sui

    Automatic outcome annotation versus automatic process annotation
  13. ACL 2024

    Lei Li, Yuqi Wang, Runxin Xu, Peiyi Wang, Xiachong Feng, Lingpeng Kong, Qi Liu

    Single-figure and multiple-figure caption pairs in ArXivCap

A full list is available on Google Scholar.

Experience

  1. Aug 2023 — Present

    Researcher. Large language models on the path to AGI.

  2. Nov 2022 — Mar 2023

    Quantitative researcher. Advised by Bowei Ma and An Ju.

  3. ByteDance Search
    Jan 2022 — Sep 2022

    Search engine for Douyin Mall. Advised by Shian Chen, Zhe Chen, and Pengcheng Yang.

  4. Mar 2021 — Dec 2021

    Effective and efficient language models. Advised by Songfang Huang and Fuli Luo.

  5. Nov 2019 — Jan 2021

    Information extraction and machine translation. Advised by Lei Li, Mingxuan Wang, and Jun Cao.

  6. Microsoft C+AI
    Jul 2019 — Oct 2019

    Built the vscode-maven extension. Advised by Rome Li, Jinbo Wang, and Yan Zhang.

Education

  1. Sep 2020 — Jun 2023

    M.S., School of EECS. Advised by Baobao Chang and Zhifang Sui at the Institute of Computational Linguistics.

  2. Sep 2016 — Jun 2020

    B.Eng., School of Cyber Science and Engineering.

Honours & Awards