Highlights
- Pro
Lists (4)
Sort Name ascending (A-Z)
Stars
CuBlaze is a CUDA implementation of the InfoNCE Loss (Information Noise-Contrastive Estimation) for self-supervised contrastive learning. This implementation provides full batch processing with nat…
PyTorch implementation of "Supervised Contrastive Learning" (and SimCLR incidentally)
PyTorch implementation of Contrastive Learning methods
Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inference and support Web deplo…
A Deep Learning Python Toolkit for Healthcare Applications.
YASA (Yet Another Spindle Algorithm): a Python package to analyze polysomnographic sleep recordings.
PyTorch implementation of VALL-E(Zero-Shot Text-To-Speech), Reproduced Demo https://lifeiteng.github.io/valle/index.html
State-of-the-art audio codec with 90x compression factor. Supports 44.1kHz, 24kHz, and 16kHz mono/stereo audio.
unofficial implementation of the High Fidelity Neural Audio Compression
A list of publicly available room impulse response datasets and scripts to download them.
collection of awesome research in brain decoding, including interaction with multi-modalities, theories, and foundation models.
An evolving, large-scale and multi-domain ASR corpus for low-resource languages with automated crawling, transcription and refinement
✨✨Latest Advances on Multimodal Large Language Models
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Andr…
Some fast-ish algorithms for batch text search in moderate-sized collections, intended for data cleanup
My Python scripts for crawling paper related on speech processing.
Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.
Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation Pytorch's Implement
Python scripts to create noisy and reverberant 2-speaker mixture audio with Libri-Light and WHAM
Kaldi-compatible online & offline feature extraction with PyTorch, supporting CUDA, batch processing, chunk processing, and autograd - Provide C++ & Python API
Unofficial PyTorch implementation of Google AI's VoiceFilter system
Implementing the paper -
This is the official repository for M2UGen
Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction (LLM-TSE)
A PyTorch implementation of "TasNet: Surpassing Ideal Time-Frequency Masking for Speech Separation" (see recipes in aps framework https://github.com/funcwj/aps)