cublas
Here are 128 public repositories matching this topic...
LUPINE is a GPU over IP bridge allowing GPUs on remote machines to be attached to CPU-only machines.
-
Updated
Sep 9, 2026 - C++
Safe rust wrapper around CUDA toolkit
-
Updated
Aug 12, 2026 - Rust
Several optimization methods of half-precision general matrix multiplication (HGEMM) using tensor core with WMMA API and MMA PTX instruction.
-
Updated
Sep 8, 2024 - Cuda
🚀🚀🚀 This repository lists some awesome public CUDA, cuda-python, cuBLAS, cuDNN, CUTLASS, TensorRT, TensorRT-LLM, Triton, TVM, MLIR, PTX and High Performance Computing (HPC) projects.
-
Updated
Aug 2, 2025
CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
-
Updated
Mar 30, 2026 - Cuda
Hooked CUDA-related dynamic libraries by using automated code generation tools.
-
Updated
Dec 12, 2023 - C
GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT.
-
Updated
Jun 18, 2026 - Cuda
Deep Learning library using GPU(CUDA/cuBLAS)
-
Updated
Sep 18, 2021 - Elixir
A Deep Learning framework with very few dependencies, Written in Rust
-
Updated
Feb 14, 2025 - Rust
Algorithms implemented in CUDA + resources about GPGPU
-
Updated
Jan 18, 2022 - Cuda
Harness the power of GPU acceleration for fusing visual odometry and IMU data with an advanced Unscented Kalman Filter (UKF) implementation. Developed in C++ and utilizing CUDA, cuBLAS, and cuSOLVER, this system offers unparalleled real-time performance in state and covariance estimation for robotics and autonomous system applications.
-
Updated
Mar 21, 2024 - Cuda
bilibili视频【CUDA 12.x 并行编程入门(C++版)】配套代码
-
Updated
Aug 12, 2024 - Cuda
code for benchmarking GPU performance based on cublasSgemm and cublasHgemm
-
Updated
May 20, 2022 - Cuda
Add this topic to your repo
To associate your repository with the cublas topic, visit your repo's landing page and select "manage topics."