15 Apr 2026

reel-forge — fully autonomous short-form video agent

So I built a thing — a fully autonomous short-form video agent.

Give it a topic and a brand identity. It researches it, writes a script, generates voiceover, creates the visuals, adds captions, and spits out a finished 9:16 video. No touching a single frame manually.

The pipeline has 6 phases:

Research — DuckDuckGo + LLM summarisation
Planning — hook, segments, CTA, tone guidance
Script — narration + visual prompts per segment
Assets — Kokoro TTS for voice, AnimateDiff for video
Render — Whisper captions + FFmpeg composite → 1080×1920

The fun engineering challenge: fitting TTS and video generation on an 8GB GPU that's also running a display. Had to release VRAM between phases just to keep it alive. "Local-first" and "GPU-intensive" are not a fun combination.

Each phase is resumable too — so if your GPU dies mid-render you pick up where you left off rather than starting over.

Also ships with an MCP server, so an AI agent can drive the whole pipeline without any human in the loop.

Is it perfect? No. Does it go from prompt to finished video autonomously? Yes.

buildinpublicgenerativeaipythonllmagenticmcp

Project · Originally on LinkedIn

newer "Never scaffold. Always fully implement." — dev philosophy