Our lab is committed to cutting-edge research in speech generation, spoken dialogue systems, and spatial audio generation. We strive to develop intelligent, natural, and immersive audio technologies that advance human–machine interaction and multimedia experiences.
MM-Speech
Popular repositories Loading
-
SwanSphere
SwanSphere Public[ICML 2026] Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
Python 120
-
DuplexSurvey
DuplexSurvey Public[EMNLP 2026 Main] Speaking While Listening: a survey and empirical audit of full-duplex spoken dialogue systems — L0–L3 architectural hierarchy, T×I×R interaction ontology, and a five-state decisio…
-
DualAxisRM
DualAxisRM Public[ACL 2026] Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models
-
DiTReducio
DiTReducio Public[ACL 2026] DiTReducio: A Training-Free Acceleration for DiT-Based TTS viaProgressive Calibration
Repositories
- AudioGen-Paradigm-Bench Public
A Reproducible Study of Autoregressive, Diffusion and Flow-Matching Generative Models for Speech, Singing Voice and Sound Effects
- VoxZip Public
[ACM MM 2026] VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference
- CSAVocoder Public
[EMNLP 2026] CSAVocoder: A Causal Spatial Audio Vocoder Towards Real-Time Spatial Audio Generation
- DuplexSurvey Public
[EMNLP 2026 Main] Speaking While Listening: a survey and empirical audit of full-duplex spoken dialogue systems — L0–L3 architectural hierarchy, T×I×R interaction ontology, and a five-state decision machine, with a curated list of models, datasets, and benchmarks.
- Controllable-TTS-Survey Public
This repository organizes an ongoing survey and benchmark study on multi-constraint controllable text-to-speech.
- SwanSphere Public
[ICML 2026] Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
- RoleJudge Public
[ACM MM 2026] Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…