YouTube clip generator — download, transcribe, find viral hooks, render clips with styled subtitles.
YouTube URL → Download → Transcribe → AI Hook Research → Styled Clip Render
- Download video via go-ytdlp (auto-installs binary)
- Transcribe audio via Groq Whisper API (cloud, fast) or local whisper (CPU fallback)
- Correct transcript via LLM (fixes speech recognition errors)
- Find hooks — LLM analyzes transcript for viral-worthy segments (15-90s)
- Render clips via moviego with:
- Aspect ratio crop (9:16, 1:1, 16:9)
- LLM-directed visual style (caption layout, colors, zoom, fade)
- Word-by-word animated subtitles
# Clone
git clone https://github.com/igun997/ngeclip.git
cd ngeclip
# Setup
cp .env.example .env
# Edit .env with your API keys
# Install deps
make deps
# Run
make build
./ngeclip auto "https://youtube.com/watch?v=VIDEO_ID"| Command | Description |
|---|---|
ngeclip auto [url] |
Full pipeline with TUI progress + hook picker |
ngeclip tui |
Pick existing video/hooks from output dir |
ngeclip process [url] |
Full pipeline, render all hooks (no TUI) |
ngeclip download [url] |
Download only |
ngeclip transcribe [file] |
Transcribe only |
ngeclip hooks [transcript] |
Find hooks only |
ngeclip render [video] [hooks] |
Render all hooks |
ngeclip pick [video] [hooks] |
TUI pick then render |
Copy .env.example to .env:
# Output
NGECLIP_OUTPUT_DIR=./output
NGECLIP_ASPECT_RATIO=9:16 # 9:16, 1:1, 16:9
# LLM (OpenAI-compatible: OpenAI, Groq, Ollama, vLLM, 9Router, etc)
NGECLIP_LLM_BASE_URL=http://localhost:20128/v1
NGECLIP_LLM_API_KEY=your-key-here
NGECLIP_LLM_MODEL=groq/llama-3.3-70b-versatile
# Transcription (Groq whisper - cloud, free, fast)
NGECLIP_GROQ_API_KEY=your-key-here
NGECLIP_GROQ_BASE_URL=https://api.groq.com/openai/v1/audio/transcriptions
NGECLIP_GROQ_MODEL=whisper-large-v3-turbo
NGECLIP_TRANSCRIBE_BACKEND=groq # groq or local
# Local whisper (fallback)
NGECLIP_WHISPER_MODEL=base # tiny, base, small, medium, large
# Clip settings
NGECLIP_MAX_CLIP_DURATION=90
NGECLIP_MIN_CLIP_DURATION=15
# Subtitle styling
NGECLIP_SUBTITLE_FONT=Arial
NGECLIP_SUBTITLE_SIZE=48- Go 1.25+
- FFmpeg on PATH
- LLM API (any OpenAI-compatible endpoint)
- Groq API key (free at https://console.groq.com) — or local whisper
yt-dlp and whisper are auto-installed on first run.
cmd/ngeclip/ CLI entry (cobra commands + TUI auto flow)
internal/
clip/ Pipeline orchestrator
downloader/ go-ytdlp wrapper (auto-installs binary)
transcriber/
groq.go Groq Whisper cloud transcription
transcriber.go Local whisper fallback
corrector.go LLM transcript correction (streaming)
install.go Auto-install whisper in venv
renderer/
renderer.go moviego: crop, zoom, fade, subtitle composite
style_director.go LLM decides visual style per clip
hooks/
finder.go LLM hook research (OpenAI-compatible)
tui/
pipeline.go Progress bar for auto pipeline
picker.go File picker (videos, hooks)
hook_picker.go Hook selection with toggle
pkg/
config/ .env + env var config
models/ Shared types
make help # show all commands
make build # compile binary
make run URL=... # auto pipeline
make deps # install dependencies
make check # verify tools
make clean # remove artifacts
make test # run tests
make lint # golangci-lint
make env # create .env from exampleBefore rendering each clip, the LLM decides:
- Caption layout:
word_center(one word at a time),bottom_bar, orkaraoke - Colors: text, stroke, highlight for active word
- Motion: Ken Burns zoom (direction + amount)
- Fade: in/out durations
- Color grade: brightness, contrast, saturation boost
Falls back to sensible defaults if LLM unavailable.
clip_00_Michael_Jackson_Masih_Hidup.mp4
For screen-share / presentation videos, ngeclip uses a grid-labeled vision measurement approach to detect and composite speaker + screen panels into portrait layout.
Hook frames → 12×12 labeled grid overlay → VLM structured measurement → normalized rects → interpolated composite
- Extract start/middle/end hook frames
- Overlay deterministic 12×12 grid (columns A-L, rows 1-12) via ffmpeg drawgrid
- Prompt VLM to identify screen, speaker, and caption-safe regions using grid cell labels (e.g.
A1:F6) - Convert cell labels → normalized
[0,1]rectangles - Validate confidence ≥ 0.75, reject overlapping/malformed measurements
- Interpolate panel geometry between 3 temporal samples
- Render top-screen / bottom-speaker composite with measured crop regions
- Fall back to adaptive portrait crop if measurement invalid
| Contribution | vs. Prior Art |
|---|---|
| Grid-cell spatial vocabulary for VLMs — discrete A1:L12 labels instead of raw pixel coordinates | Prior work uses bounding box coords (hallucination-prone) or trained object detectors (need labeled data). Grid cells constrain VLM output to a small, validatable alphabet. |
| 3-sample temporal interpolation — measure once at start/mid/end, interpolate at render | Per-frame detection is expensive; static crop ignores speaker movement. This balances cost vs. smoothness. |
| Confidence-gated fallback — hard rejection threshold with graceful degradation | Most VLM reframing systems lack explicit confidence gates; they trust model output or fail silently. |
| Planning-time vision, render-time geometry — VLM runs once during planning, renderer is pure math | Separates expensive inference from deterministic rendering. Enables offline QA of measurements before committing to render. |
| Paper | Year | Relation |
|---|---|---|
| Reframe Anything: LLM Agent for Open World Video Reframing | 2024 | LLM/VLM-driven reframing agent; closest overall concept |
| ChatDirector: Space-Aware Scene Rendering and Speech-Driven Layout Transition | CHI 2024 | Dynamic speaker+content layout composition in video conferencing |
| TalkDirector: Real-time Multimodal Slide Augmentation via Adaptive Presenter Integration | 2025 | Multimodal inference for presenter placement relative to slides |
| AutoFlip: Open Source Framework for Intelligent Video Reframing | Google 2020 | Content-aware reframing with fallback; our adaptive crop parallels this |
| VoiceVision: Speaker-Aware Cropping for Multi-Speaker Videos | 2025 | Active speaker detection + dynamic crop |
| Motion-based Video Retargeting with Optimized Crop-and-Warp | SIGGRAPH 2010 | Temporal interpolation of crop keyframes — foundational for our 3-sample approach |
| Accurate Screen Detection in Presentation Videos using YOLOv7 | 2024 | Screen region detection via object detection (vs. our grid+VLM approach) |
| TRISHUL: Region Identification and Screen Hierarchy for VLM GUI Agents | CVPR-W 2025 | VLM spatial grounding with structured output; validates grid-prompting approach |
| Dynamic Beauty: Composition-Aware Video Reframing | ACM MM 2025 | End-to-end composition-quality-aware reframing |
| Structuring Lecture Videos by Projection Screen Localization | IEEE TPAMI 2014 | Foundational screen+presenter tracking in lecture videos |
MIT