Running today · built and operated by Johnny Bodegas

An idea goes in.
A finished, checked video
comes out.

I built a working AI video engine where agents write and direct, code handles timing and quality, and a person approves instead of editing. That's the system Elva is hiring for, and I've already shipped one.

✍️
IntakeBrief, source page, brand colours, picturesHuman gate
🧠
Script agentBeats in a chosen framework, unbacked claims flaggedHuman gate
🎙️
VoiceMeasured word by word; it sets the clockHuman gate
🎬
Director agentPicks templates, places each picture on its lineHuman gate
✅
Checks + render21 rules, then the final film and a quality reportHuman gate
Proof first

Not a prototype. A working engine.

Every item in this section is in the working code today. I checked each one against the source before writing this page.

661automated tests passing
21quality rules run before any person reviews
21scene templates the director chooses from
8pre-release checks, incl. a 6,000-frame endurance render
👁️

A director that sees the pictures Works today

Every uploaded picture is read once by a vision model for its text and a one-line description. The director places the invoice screen on the line about invoices, not in upload order.

⏱️

Word-timed cards Works today

Each on-screen card is anchored to the moment its word is actually spoken, measured from the voice. Re-record the voice and the cards re-time themselves.

🎨

Versioned house style Works today

Templates, thresholds, voices, palettes and sound effects are versioned records. A new style is a new version, not a new editor.

🔗

Only what changed goes stale Works today

Approvals are tied to the exact inputs. Edit one line upstream and only the pieces that depend on it lose their approval.

🌐

Briefs grounded in real sources Works today

Paste a URL: the page becomes a dated snapshot of facts, and the brand palette is read from the site's own stylesheets for a person to confirm.

🛟

One missing asset never kills a film Works today

A missing or failed picture renders as a clean fallback card. A model that ignores a rule is asked again, then the rule is applied for it.

Script frameworks

Problem–Agitate–SolveAIDAContrarian value-firstSPARKThe Sequence

Video types

Reel · 5–90 sShort · 5–180 sVSL · 181–1,200 s16:99:16

Script source

Agent writes itBring your own scriptTwo-voice scripts

See it run, live, on a real brief.

One film through every gate in about six minutes, including a rule failing on purpose and an approval going stale.

Book a live demo of one of my pipelines →
Fit for Elva

Your brief, already in production.

Elva wants a system that takes an idea and references, analyses them, builds a narrative, picks visuals and audio, and outputs a structured video. Here's where I've already built each piece.

What Elva needs
What I've built
Idea + references as input
An intake with topic, brief, source URL snapshot, facts, proof screenshots, pictures and logos, all in one approved brief that everything downstream reads from.
Content analysis
Pages stripped to facts, brand colours read from stylesheets, and every picture read by a vision model for what it says and shows.
Narrative generation
A script agent that writes beats in a chosen framework, with strict output formats, a re-ask on bad output, and flags for claims it can't back up.
Editing orchestration
Scene timing cut by code from the measured voice; a director agent fills each scene from 21 templates; 21 rules check the result.
Visual + audio selection
Template and picture choice per scene, opt-in image generation, voice casting, sound effects, and loudness mastering to a set target.
Tool + API integration
ElevenLabs voice with word timing, image models through an AI gateway, and a second project (my studio) holding adapters for video, image, voice, music and stock-footage providers.
System prompts + behavioural rules
Agent instructions, frameworks and thresholds stored as versioned data, enforced by code rather than requested in prompts.
Quality + reliability
661 tests, 8 pre-release checks, repeatable renders (same inputs, same frames), and a quality report on every render.
How I build: I build the engine as one app, so the whole pipeline lives in one place instead of hopping between separately connected tools. Today the production engine calls voice, image and video providers directly through their APIs. Planned in v3 MCP function calls will let the agents it spawns oversee and review each pipeline step, and if you prefer an n8n-style or other orchestration setup, I can build that too.
Quality at volume

Built for 100 videos a day: machines review first, a human approves last.

At 100 videos a day, nobody can check every card and scene by hand. So every version has to pass the machine checks before it reaches me, and I give the final verdict as the approver, not the editor.

🧪

21 automatic rules Works today

They run on every version of every storyboard. Any error blocks that scene's approval and the final render, and the final render checks again and refuses to start if errors remain. Six of the rules are warnings that flag work for a closer look.

🚦

8 pre-release checks Works today

Free disk space, type checks on the app and the engine, the full test suite (661 tests), a scan for hard-coded off-brand colours, a repeatability check, a clean build, and a 6,000-frame endurance render.

🔗

Only what changed goes stale Works today

Approvals are tied to the exact inputs. Change one line and only the pieces that depend on it need approving again; everything else keeps its sign-off.

What the rules catch Works today
What counts as a fail
Scenes held too long
A scene runs past the length cap for its section of the video (separate caps for short and long form)
Silent or snapping motion
Something enters with no sound effect within 0.1 s, or text snaps on instead of revealing over at least 0.2 s
Dead motion and missing pattern breaks
The screen sits idle for more than about a quarter of a second, or there's no pattern break every 8–12 seconds
Picture out of sync with the voice
On-screen words don't match what's being said, a card isn't tied to a spoken word, or the timeline drifts more than half a second from the measured voice
Off-brand or broken build
Colours hard-coded outside the house style, a missing sound file, gaps or overlaps between scenes, missing required content, picture-in-picture or picture-card rule breaks
Weak spots (warnings)
Degraded voice timing, the same visual repeated three or more times, placeholder copy, too little variety in pattern breaks
Sound and pacing targets
Loudness mastered to −14 LUFS (±0.5) with peaks at or below −1 dBTP; first motion within the first 0.1 s and first spoken word within 0.6 s; too much dead air (over 12% short form, 20% long form) fails
Next, in version 3 Planned Agent reviewers, including a vision review of every storyboard's contact sheet, apply their own pass criteria on top of the rules. Storyboards that pass everything are approved automatically, so I only see the flagged scenes and the final verdict.
Where it's going

Version 3: from one engine to a factory.

Two working projects, combined into one system built for volume. And it adapts: every client has their own workflow, formats and approval steps, and the engine is built to be customised to the pipeline you already run.

v1

Video Production Engine Running

  • The deterministic core: timeline renderer, voice timing, 21 rules, hash-tied approvals
  • Kinetic-type explainers, reels and VSLs
  • One operator, one machine
v2

Creative Video Studio Built

  • 13 video types defined, from documentary montage to talking head and clip factory
  • Six visual styles, well over a hundred provider adapters, cost caps per run
  • Already calls the engine to check and render
v3

The combined factory Planned

  • Engine as the core, studio as the tool library
  • Agents for routing, script, direction, generation and quality review
  • Three human gates, storyboard reviewed only by exception
  • 4–6 render workers; designed for up to 100 videos a day

Adapts to your pipeline

The version 3 plan is a set of building blocks, not a fixed product. Each one can be shaped around your needs:

0Harden & verify
1Shared core for many brands
2Agent layer
3Generation router with spend caps
4Quality review & auto-approval
5Scale-out rendering
6Batch, scheduling & new formats
Version 3 is a plan, not a product yet. How long a build takes depends on scope, so I don't quote a number before I've seen yours.

Send me your video pipeline brief.

I'll review it and give you the exact build time, live in a demo or right after. That includes building from scratch if it doesn't fit my current engine.

Want to see v1 handle your kind of brief?

Book a live demo of one of my pipelines →
About

Johnny Bodegas

I design and build AI production systems end to end: the agents, the rules that keep them honest, the rendering, and the review screens. I'm based in the Philippines and work European hours.