CV Related
[CVPR 2024 Highlight] Putting the Object Back Into Video Object Segmentation
This is an automatic full segmentation tool based on Segment-Anything-2 and Segment-Anything-1. Our tool performs automatic full segmentation of the video, enabling the tracking of each object and …
Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
Based on GroundingDino and SAM, use semantic strings to segment any element in an image. The comfyui version of sd-webui-segment-anything.
ComfyUI nodes to use segment-anything-2
Fast and flexible image augmentation library. Paper about the library: https://www.mdpi.com/2078-2489/11/2/125
[CAAI AIR'24] Bilateral Reference for High-Resolution Dichotomous Image Segmentation
[CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
An open source implementation of CLIP.
Semantic segmentation models with 500+ pretrained convolutional and transformer-based backbones.
[ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"
[ECCV 2024] Official implementation of the paper "Semantic-SAM: Segment and Recognize Anything at Any Granularity"
real time face swap and one-click video deepfake with only a single image
[ECCV 2024] The official code of paper "Open-Vocabulary SAM".
[CVPR 2025] MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation
[CVPR 2024 🔥] GeoChat, the first grounded Large Vision Language Model for Remote Sensing
Detectron2 is a platform for object detection, segmentation and other visual recognition tasks.
We write your reusable computer vision tools. 💜