Skip to content
View ShawnPi233's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report ShawnPi233

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
ShawnPi233/README.md

Bingsong Bai

I am a Foundation Model Algorithm Engineer at ModelBest (面壁智能,VoxCPM). I received my Master's degree from the School of Artificial Intelligence, Beijing University of Posts and Telecommunications (BUPT). My research interests lie in the intersection of Large Speech Models (LSM), Automatic Speech Recognition (ASR), Singing Voice Conversion (SVC), and Expressive Text-to-Speech (TTS).

I have gained extensive industry experience through research and engineering internships at Zhipu AI (智谱AI语音输入法), Tencent Music Entertainment (腾讯音乐天琴实验室), Momo (陌陌) and Kunlun Inc (昆仑万维).

🔥 News

  • 2026.03: 🚀 Joined ModelBest as a Large Speech Foundation Model Researcher.
  • 2026.01: 🎉 One paper (SynParaSpeech) accepted by ICASSP 2026 as the first author!
  • 2025.12: 🎉 One paper (HQ-SVC) accepted by AAAI 2026 as the first author!
  • 2025.10: 🚀 Joined Zhipu AI as a Speech Large Model Research Intern.
  • 2025.07: 🎸 Joined Tencent Music (QQ Music) focusing on multi-speaker conversational podcast TTS.
  • 2025.03: 👫 Joined Momo focusing on paralinguistic TTS and understanding.
  • 2024.06: 🎉 One paper (SPA-SVC) accepted by Interspeech 2024 as the first author.

📑 Selected Research Papers

📎 For a full list of publications, please visit my Google Scholar.

HQ-SVC: High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios, Bingsong Bai, et al., AAAI 2026. [CCF-A]

SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding, Bingsong Bai, et al., ICASSP 2026. [CCF-B]

SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion, Bingsong Bai, et al., Interspeech 2024. [CCF-B]

🗣 Large Speech Models & TTS

Pinned Loading

  1. HQ-SVC HQ-SVC Public

    Official Repository of Paper: "Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios"(AAAI 2026)

    Python 114 9

  2. SynParaSpeech SynParaSpeech Public

    Official Repository of Paper: "SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding" (ICASSP 2026)

    JavaScript 72 4