Lab Notebook · OpenClaw Media Fleet · July 2026

We gave our OpenClaw agents eyes, voices, and a six-DGX media fleet.

Cloud subscriptions handle intelligence and orchestration. Six NVIDIA DGX Sparks handle image generation, face lock, TTS, and approved voice cloning — so agents can produce media inside the same workflows that already run the business.

DGX Sparks
Local media production capacity
2
Media skills
Fleet Image + Fleet Voice
Hybrid
Cloud + local
Subscriptions for brains, Sparks for assets
Clone
Face + voice
Reusable identities, not one-off luck

Since then

For the current control-plane story — six Sparks, MCP, subscription gas tanks, and per-agent RAG memory — start with the newer operator brief: AI Fleet Router.

Field note · 2026-07-27

Pretty generations are easy. Production media is the hard part.

A chatbot can suggest a visual or draft a voiceover script. That is not the same as an AI Employee that can take a brief, choose the right model, lock a face or voice, generate the asset, and deliver it back into Discord, a page build, or a campaign folder.

We did not want another tab full of AI toys. We wanted OpenClaw agents that can actually help run multimodal production without a scavenger hunt across five apps.

Jeff's OpenClaw home lab rack with local infrastructure
Home lab context. The media fleet sits inside the same infrastructure posture that already powers local models, routers, and agent work.

The breakthrough was the system, not one model.

Subscriptions, skills, and Sparks each do a different job. Connected correctly, they turn media from a side quest into an agent capability.

🧠
SUBSCRIPTIONS

Cloud intelligence

Reasoning, creative direction, prompt craft, script optimization, and tool orchestration stay on the models we already pay for.

🛠️
OPENCLAW SKILLS

Reusable media workflows

Fleet Image and Fleet Voice give every approved agent the same interview → optimize → generate → deliver loop instead of ad-hoc commands.

6 DGX SPARKS

Local media capacity

Image gen, face/character lock, TTS, and approved voice cloning run on private fleet hardware with room for parallel work.

One conversation. Two production lanes.

The same OpenClaw agent can route a request into image generation, voice production, or both — without leaving the workflow.

🖼️
IMAGE

Generate campaign-ready visuals

Agents turn a rough request into a production prompt, pick FLUX schnell for drafts, FLUX dev for hero quality, or SDXL when the look needs it, then ship the job through ComfyUI on the fleet.

🎙️
AUDIO

Clone approved voices and narrate

With a clean source clip and permission, agents create a reusable voice ID, annotate pacing and emotion tags, generate WAV audio, and deliver it where the work is happening.

🔁
IDENTITY

Lock faces and characters

Instant face clone via PuLID/InstantID, saved characters by name, and optional full-character LoRA training when a persona needs higher fidelity across dozens of generations.

The media stack, in pictures.

Real assets from the VA Staffer library: the infrastructure and reference systems behind repeatable media production.

What happens after “make this.”

The interesting part is not the button click. It is the chain of decisions the agent handles around it.

🎯

Understand the brief

Audience, goal, tone, subject, format, delivery channel, and whether this needs a new clone or a saved identity.

✍️

Optimize the prompt or script

Image prompts get scene, lighting, and quality cues. Voice scripts get emotion tags, pace, and temperature/speed settings before anything generates.

🧭

Route to the fleet

The skill hits the fleet router. Available Spark capacity handles generation instead of tying every media job to one laptop or one SaaS tab.

🔍

Generate and inspect

The agent creates the visual or audio, checks it against the brief, and adjusts identity strength, model, or script when the first pass is off.

📤

Deliver back into the work

Named files, attachments in the right conversation, and handoff into the page, campaign, video edit, or approval thread — not a dead-end download folder.

“Subscriptions give the agents access to great intelligence. The DGX fleet gives that intelligence somewhere to do real media work.”

That split is intentional. Cloud models are excellent at thinking. Local Sparks are excellent at repeated generation, cloning, and private media capacity under our control.

Cloud intelligence and local production are better together.

This is a hybrid system by design. Each side does the work it is best suited to do.

Subscriptions

What cloud still owns

  • Strong reasoning and creative direction
  • Better scripts, prompts, and production plans
  • Tool use and orchestration inside OpenClaw
  • Access to leading models without hosting every one ourselves
Six DGX Sparks

What the fleet owns

  • Dedicated capacity for image and audio workloads
  • Local control over voice assets and reusable character identities
  • Room to test, route, and improve models across the rack
  • Less dependence on per-generation SaaS for repeatable production
Fleet monitor showing local backends and model routing health
Router posture. Agents call skills. Skills hit the fleet. The fleet decides where the media job runs.

Why six machines instead of one demo box?

A single workstation can prove the idea. A fleet turns it into an operating capability.

📦
CAPACITY

Parallel media jobs

Image and audio work can be heavy. Multiple Sparks give the router more places to send work when several agents need media at once.

🎯
SPECIALIZATION

Separate the workloads

Voice cloning does not have to fight hero image generation for the same resources. Different nodes can serve different models cleanly.

🛡️
RESILIENCE

Upgrade without freeze

Fleet routing gives room to maintain, test, and improve without treating one machine as the entire production system.

What this unlocks in the business

Media stops being a separate department of tabs. It becomes another skill an AI Employee can run with human taste still in charge.

📣

Brand voice on demand

Approved cloned voices for promos, announcements, lesson intros, and internal updates — with pacing and emotion control, not flat robot reads.

👤

Persona consistency

Founder likeness, team personas, and mascots stay recognizable across images instead of drifting into lookalike territory every generation.

🚀

Content velocity

Agents can prepare assets while writing the page, building the funnel, or packaging the campaign — same conversation, fewer handoffs.

🔐

Private production capacity

Sensitive voice samples and recurring brand characters can live on infrastructure we control, with cloud still used where it wins.

What this is — and what it is not

This is a production system for AI Employees: subscriptions for intelligence, OpenClaw skills for process, and a six-Spark fleet for media generation and cloning.

It is not autopilot brand taste. Humans still approve likeness, voice use, final creative direction, and what ships publicly. That boundary is what keeps the system useful instead of chaotic.

Related reading

More from the lab if you want the surrounding infrastructure story.

Want an AI Employee that can think and produce media?

The goal was never more AI tools. The goal was an AI Employee that can move from idea to useful asset without handing the operator a scavenger hunt.

Get an AI Employee Hire Beau
Beau, VA Staffer's AI Employee
Built by Beau

This page was created by Beau, VA Staffer's AI Employee.

Beau is Jeff's AI Employee for pages, assets, drafts, deployment, and support materials. He helps the team move faster by turning ideas into real deliverables that can be edited, deployed, and improved over time.