-
Tianjin University
- Tianjin, China
- https://linan2.github.io/
Stars
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
Google Scholar skills for Claude Code — search, citation tracking, full-text access, and Zotero export via Chrome DevTools MCP
「摘星」· 微信助手 — AI Agent 智能助手 / AI 群聊摘要 / 关键词即时提醒 / 公众号摘要与文章实时推送 / RAG 语义检索 / 聊天、朋友圈归档 / 收藏导出 。本地运行,数据不出本机。
AuK: An Open-Source Foundational Model for Speech Generation and Editing
Wespeaker implementations for speaker recognition and verification: U3-xi,Uncertainty aware AAM Softmax, Score normalization, calibration, and unofficial RecXi (NeurIPS 2023) source code.
A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表
FireRedTTS3: Multilingual and Multi-Dialect Voice Cloning with Instruction-Guided Voice Design and Speech Editing
DeepSeek Harness: Everything is a Plugin.
Java 离线语音识别(ASR)、语音合成(TTS)、声纹识别、VAD、说话人分离、降噪、KWS 关键词唤醒 SDK。基于 sherpa-onnx,Spring Boot / Solon 自动装配,JDK 8+ 兼容。
Robust Speech Recognition via Large-Scale Weak Supervision
Open-Source Turn-Taking Detection Model and Dataset for Full-Duplex Spoken Dialogue Systems
Reference implementation of an end-to-end voice agent built using the NVIDIA Nemotron models
Hy3 (295B A21B), a leading reasoning and agent model in its size, with great cost efficiency.
A Fully Self-Hosted Solution for Full-Duplex Voice Interaction
Adaptive Flow-Matching for Target Speaker Extraction
Plug-and-play streaming semantic VAD for real-time full-duplex spoken dialogue systems.
A curated list of full-duplex spoken dialogue models & benchmarks
A curated list of models, benchmarks, tools and guides for audio editing
Speed-optimized streaming neural speech enhancement network
Lightweight coding agent that runs in your terminal
利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.
Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR
Robust Speech Recognition Across Languages, Dialects, and Complex Acoustic Scenarios