Stars
yuekaizhang / speculators
Forked from vllm-project/speculatorsA unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
MiMo: Unlocking the Reasoning Potential of Language Model – From Pretraining to Posttraining
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice…
中国专利.skill:专利点挖掘与交底书(发明/实用/外观)编写,通俗解读专利,嗅探政策动向,辅助审查答复。
Automatically Update Text-to-speech (TTS) Papers Daily using Github Actions (Update Every 12th hours)
Zonos2 is a leading open-weight text-to-speech MoE.
QuantaAlpha transforms how you discover quantitative alpha factors by combining LLM intelligence with evolutionary strategies. Just describe your research direction, and watch as factors are automa…
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
The AI that really does things. Any OS. Any Platform. The lobster way. 🦞
A powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics, and features robust zero-shot text-to-speech
[INTERSPEECH 2026 Oral]Official code for "Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis"
CosyVoice_DPO_NOTES: Supercharge Your Cosyvoice model with Cutting-Edge DPO Fine-Tuning!
Official Repository of Paper: "SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding" (ICASSP 2026)
High-quality speech synthesis with LoRA fine-tuning on index-tts, enhancing prosody and naturalness for single and multi-speaker voices.
This repository presents an evaluation framework for speech-to-speech (S2S) models, following the methodology described in the EmphAsses paper (de Seyssel et al., 2023).
Official code for "EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting"
Wan: Open and Advanced Large-Scale Video Generative Models
Text-audio foundation model from Boson AI
Official code for "F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization"
Easily train a good VC model with voice data <= 10 mins!
Evaluation Metrics Used For The Performance Evaluation of Voice Conversion (VC) Models
Multilingual Automatic Speech Recognition with word-level timestamps and confidence
Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads