15 Apr 2026
reel-forge — fully autonomous short-form video agent
So I built a thing — a fully autonomous short-form video agent.
Give it a topic and a brand identity. It researches it, writes a script, generates voiceover, creates the visuals, adds captions, and spits out a finished 9:16 video. No touching a single frame manually.
The pipeline has 6 phases:
Research — DuckDuckGo + LLM summarisation
Planning — hook, segments, CTA, tone guidance
Script — narration + visual prompts per segment
Assets — Kokoro TTS for voice, AnimateDiff for video
Render — Whisper captions + FFmpeg composite → 1080×1920
The fun engineering challenge: fitting TTS and video generation on an 8GB GPU that's also running a display. Had to release VRAM between phases just to keep it alive. "Local-first" and "GPU-intensive" are not a fun combination.
Each phase is resumable too — so if your GPU dies mid-render you pick up where you left off rather than starting over.
Also ships with an MCP server, so an AI agent can drive the whole pipeline without any human in the loop.
Is it perfect? No. Does it go from prompt to finished video autonomously? Yes.