Dingdong Wang

πŸ‘‹ About Me

Hi, I am Dingdong Wang (ηŽ‹δΈε†¬).

I am a Ph.D. candidate in Systems Engineering and Engineering Management (SEEM) at πŸ‡­πŸ‡° The Chinese University of Hong Kong (CUHK), advised by Prof. Helen Meng (Fellow of IEEE and ISCA). My research focuses on speech-native and multimodal foundation models that can hear, understand, reason, and interact. Concretely, I work on speech representation, instruction-following, spoken language reasoning, and end-to-end dialogue.

If you are interested in collaboration, feel free to drop me an email.

I am currently looking for full-time job opportunities and am broadly interested in multimodal LLMs, audio processing, agentic AI, and world models.

πŸŽ“ Education

The Chinese University of Hong Kong (CUHK)
Ph.D., Systems Engineering and Engineering Management (SEEM)
08/2023 – 07/2027 Hong Kong, China
  • Supervisor: Prof. Helen Meng (Fellow of IEEE and ISCA).
  • Research area: Speech and multimodal large language models, with a focus on spoken language understanding, reasoning, and end-to-end dialogue.
The Chinese University of Hong Kong (CUHK)
Bachelor of Science Β· First Class Honors
09/2018 – 07/2022 Hong Kong, China
  • Major: Statistics.
  • Awards: Department of Statistics Scholarship; Dean's List.

πŸ”₯ News

  • [2026/08] 3 papers are accepted by EMNLP 2026.
  • [2026/05] The survey of audio reasoning in multimodal foundation models is released on arXiv.
  • [2026/04] One paper accepted to ICML 2026.
  • [2026/02] EmotionThinker selected for πŸ† Oral Presentation.
  • [2026/01] Two first-author papers, MMSU and EmotionThinker, accepted to ICLR 2026.
  • [2025/12] Wan-Move tops Hugging Face Weekly Papers (Dec 7–13).
  • [2025/09] Wan-Move accepted to NeurIPS 2025.
  • [2025/08] One first-author paper, Speech Discrete Tokens or Continuous Features? accepted to EMNLP 2025 Main Conference.
  • [2025/05] One first-author paper, InSerter accepted to ACL 2025 Main Conference.
  • [2025/05] SocialCC and Language-Codec accepted to ACL 2025 Main Conference.
  • [2024/09] SimpleSpeech wins the πŸ† Best Student Paper Award at INTERSPEECH 2024.

πŸ“ Selected Publications (Full List)

MMSU benchmark example tasks
ICLR 2026

MMSU: A Massive Multi-Task Spoken Language Understanding and Reasoning Benchmark

Dingdong Wang, Junan Li, Jincenzi Wu, Dongchao Yang, Xueyuan Chen, Tianhua Zhang, Helen Meng

EmotionThinker RL reward architecture
ICLR 2026

EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning

Dingdong Wang, Shujie Liu, Tianhua Zhang, Youjun Chen, Jinyu Li, Helen Meng

  • First RL-enhanced SpeechLLM for interpretable SER: reformulates speech emotion recognition as deep reasoning.
  • GRPO-PTR: progressive trust-aware reasoning reward for RL, improving both accuracy and reasoning quality.
InSerter interleaved pre-training pipeline
ACL 2025 Β· Main

InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training

Dingdong Wang*, Jin Xu*, Ruihang Chu, Zhifang Guo, Xiong Wang, Jincenzi Wu, Dongchao Yang, Shengpeng Ji, Junyang Lin

  • Unsupervised pre-training: interleaved speech-text representation alignment.
  • SpeechInstructBench: first speech-oriented instruction following benchmark.
  • SOTA performance: + ~10% top-1 accuracy.
Discrete tokens vs. continuous features architecture comparison
EMNLP 2025 Β· Main

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs

Dingdong Wang, Junan Li, Mingyu Cui, Dongchao Yang, Xueyuan Chen, Helen Meng

  • Controlled comparison: SSL-based discrete vs. continuous features (HuBERT/WavLM) across 6 spoken language understanding tasks.
SimpleSpeech SQ-Codec and training pipeline
INTERSPEECH 2024

SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Dongchao Yang, Dingdong Wang, Haohan Guo, Xueyuan Chen, Xixin Wu, Helen Meng

  • SQ-Codec: scalar-quantization speech codec mapping audio into a finite, compact latent space.
  • Simple and scalable: non-autoregressive, trained on 4k hours of speech-only data with no alignment information needed.
  • Significant gains in speech quality and generation speed over prior large-scale TTS models.

πŸ“‹ Academic Service

Reviewer

  • IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP)
  • Conference on Neural Information Processing Systems (NeurIPS)
  • International Conference on Machine Learning (ICML)
  • International Conference on Learning Representations (ICLR)
  • Association for the Advancement of Artificial Intelligence (AAAI)
  • Annual Meeting of the Association for Computational Linguistics (ACL)
  • IEEE International Conference on Multimedia and Expo (ICME)
  • European Conference on Computer Vision (ECCV)
  • International Joint Conference on Artificial Intelligence (IJCAI)