Skip to content
igun997Public

About

Headless Engine of Clippers

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

5 Commits

Folders and files

Repository files navigation

ngeclip

YouTube clip generator — download, transcribe, find viral hooks, render clips with styled subtitles.

What it does

YouTube URL → Download → Transcribe → AI Hook Research → Styled Clip Render
  1. Download video via go-ytdlp (auto-installs binary)
  2. Transcribe audio via Groq Whisper API (cloud, fast) or local whisper (CPU fallback)
  3. Correct transcript via LLM (fixes speech recognition errors)
  4. Find hooks — LLM analyzes transcript for viral-worthy segments (15-90s)
  5. Render clips via moviego with:
    • Aspect ratio crop (9:16, 1:1, 16:9)
    • LLM-directed visual style (caption layout, colors, zoom, fade)
    • Word-by-word animated subtitles

Quick start

# Clone
git clone https://github.com/igun997/ngeclip.git
cd ngeclip

# Setup
cp .env.example .env
# Edit .env with your API keys

# Install deps
make deps

# Run
make build
./ngeclip auto "https://youtube.com/watch?v=VIDEO_ID"

Commands

Command Description
ngeclip auto [url] Full pipeline with TUI progress + hook picker
ngeclip tui Pick existing video/hooks from output dir
ngeclip process [url] Full pipeline, render all hooks (no TUI)
ngeclip download [url] Download only
ngeclip transcribe [file] Transcribe only
ngeclip hooks [transcript] Find hooks only
ngeclip render [video] [hooks] Render all hooks
ngeclip pick [video] [hooks] TUI pick then render

Configuration

Copy .env.example to .env:

# Output
NGECLIP_OUTPUT_DIR=./output
NGECLIP_ASPECT_RATIO=9:16          # 9:16, 1:1, 16:9

# LLM (OpenAI-compatible: OpenAI, Groq, Ollama, vLLM, 9Router, etc)
NGECLIP_LLM_BASE_URL=http://localhost:20128/v1
NGECLIP_LLM_API_KEY=your-key-here
NGECLIP_LLM_MODEL=groq/llama-3.3-70b-versatile

# Transcription (Groq whisper - cloud, free, fast)
NGECLIP_GROQ_API_KEY=your-key-here
NGECLIP_GROQ_BASE_URL=https://api.groq.com/openai/v1/audio/transcriptions
NGECLIP_GROQ_MODEL=whisper-large-v3-turbo
NGECLIP_TRANSCRIBE_BACKEND=groq    # groq or local

# Local whisper (fallback)
NGECLIP_WHISPER_MODEL=base         # tiny, base, small, medium, large

# Clip settings
NGECLIP_MAX_CLIP_DURATION=90
NGECLIP_MIN_CLIP_DURATION=15

# Subtitle styling
NGECLIP_SUBTITLE_FONT=Arial
NGECLIP_SUBTITLE_SIZE=48

Requirements

  • Go 1.25+
  • FFmpeg on PATH
  • LLM API (any OpenAI-compatible endpoint)
  • Groq API key (free at https://console.groq.com) — or local whisper

yt-dlp and whisper are auto-installed on first run.

Architecture

cmd/ngeclip/           CLI entry (cobra commands + TUI auto flow)
internal/
  clip/                Pipeline orchestrator
  downloader/          go-ytdlp wrapper (auto-installs binary)
  transcriber/
    groq.go            Groq Whisper cloud transcription
    transcriber.go     Local whisper fallback
    corrector.go       LLM transcript correction (streaming)
    install.go         Auto-install whisper in venv
  renderer/
    renderer.go        moviego: crop, zoom, fade, subtitle composite
    style_director.go  LLM decides visual style per clip
  hooks/
    finder.go          LLM hook research (OpenAI-compatible)
  tui/
    pipeline.go        Progress bar for auto pipeline
    picker.go          File picker (videos, hooks)
    hook_picker.go     Hook selection with toggle
pkg/
  config/              .env + env var config
  models/              Shared types

Makefile

make help         # show all commands
make build        # compile binary
make run URL=...  # auto pipeline
make deps         # install dependencies
make check        # verify tools
make clean        # remove artifacts
make test         # run tests
make lint         # golangci-lint
make env          # create .env from example

How the LLM style director works

Before rendering each clip, the LLM decides:

  • Caption layout: word_center (one word at a time), bottom_bar, or karaoke
  • Colors: text, stroke, highlight for active word
  • Motion: Ken Burns zoom (direction + amount)
  • Fade: in/out durations
  • Color grade: brightness, contrast, saturation boost

Falls back to sensible defaults if LLM unavailable.

Example output

clip_00_Michael_Jackson_Masih_Hidup.mp4

Dynamic Presentation Stack — Method & Related Work

For screen-share / presentation videos, ngeclip uses a grid-labeled vision measurement approach to detect and composite speaker + screen panels into portrait layout.

How it works

Hook frames → 12×12 labeled grid overlay → VLM structured measurement → normalized rects → interpolated composite
  1. Extract start/middle/end hook frames
  2. Overlay deterministic 12×12 grid (columns A-L, rows 1-12) via ffmpeg drawgrid
  3. Prompt VLM to identify screen, speaker, and caption-safe regions using grid cell labels (e.g. A1:F6)
  4. Convert cell labels → normalized [0,1] rectangles
  5. Validate confidence ≥ 0.75, reject overlapping/malformed measurements
  6. Interpolate panel geometry between 3 temporal samples
  7. Render top-screen / bottom-speaker composite with measured crop regions
  8. Fall back to adaptive portrait crop if measurement invalid

What's original

Contribution vs. Prior Art
Grid-cell spatial vocabulary for VLMs — discrete A1:L12 labels instead of raw pixel coordinates Prior work uses bounding box coords (hallucination-prone) or trained object detectors (need labeled data). Grid cells constrain VLM output to a small, validatable alphabet.
3-sample temporal interpolation — measure once at start/mid/end, interpolate at render Per-frame detection is expensive; static crop ignores speaker movement. This balances cost vs. smoothness.
Confidence-gated fallback — hard rejection threshold with graceful degradation Most VLM reframing systems lack explicit confidence gates; they trust model output or fail silently.
Planning-time vision, render-time geometry — VLM runs once during planning, renderer is pure math Separates expensive inference from deterministic rendering. Enables offline QA of measurements before committing to render.

Related papers

Paper Year Relation
Reframe Anything: LLM Agent for Open World Video Reframing 2024 LLM/VLM-driven reframing agent; closest overall concept
ChatDirector: Space-Aware Scene Rendering and Speech-Driven Layout Transition CHI 2024 Dynamic speaker+content layout composition in video conferencing
TalkDirector: Real-time Multimodal Slide Augmentation via Adaptive Presenter Integration 2025 Multimodal inference for presenter placement relative to slides
AutoFlip: Open Source Framework for Intelligent Video Reframing Google 2020 Content-aware reframing with fallback; our adaptive crop parallels this
VoiceVision: Speaker-Aware Cropping for Multi-Speaker Videos 2025 Active speaker detection + dynamic crop
Motion-based Video Retargeting with Optimized Crop-and-Warp SIGGRAPH 2010 Temporal interpolation of crop keyframes — foundational for our 3-sample approach
Accurate Screen Detection in Presentation Videos using YOLOv7 2024 Screen region detection via object detection (vs. our grid+VLM approach)
TRISHUL: Region Identification and Screen Hierarchy for VLM GUI Agents CVPR-W 2025 VLM spatial grounding with structured output; validates grid-prompting approach
Dynamic Beauty: Composition-Aware Video Reframing ACM MM 2025 End-to-end composition-quality-aware reframing
Structuring Lecture Videos by Projection Screen Localization IEEE TPAMI 2014 Foundational screen+presenter tracking in lecture videos

License

MIT

About

Headless Engine of Clippers

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages