Cloud subscriptions handle intelligence and orchestration. Six NVIDIA DGX Sparks handle image generation, face lock, TTS, and approved voice cloning — so agents can produce media inside the same workflows that already run the business.
For the current control-plane story — six Sparks, MCP, subscription gas tanks, and per-agent RAG memory — start with the newer operator brief: AI Fleet Router.
A chatbot can suggest a visual or draft a voiceover script. That is not the same as an AI Employee that can take a brief, choose the right model, lock a face or voice, generate the asset, and deliver it back into Discord, a page build, or a campaign folder.
We did not want another tab full of AI toys. We wanted OpenClaw agents that can actually help run multimodal production without a scavenger hunt across five apps.
Subscriptions, skills, and Sparks each do a different job. Connected correctly, they turn media from a side quest into an agent capability.
Reasoning, creative direction, prompt craft, script optimization, and tool orchestration stay on the models we already pay for.
Fleet Image and Fleet Voice give every approved agent the same interview → optimize → generate → deliver loop instead of ad-hoc commands.
Image gen, face/character lock, TTS, and approved voice cloning run on private fleet hardware with room for parallel work.
The same OpenClaw agent can route a request into image generation, voice production, or both — without leaving the workflow.
Agents turn a rough request into a production prompt, pick FLUX schnell for drafts, FLUX dev for hero quality, or SDXL when the look needs it, then ship the job through ComfyUI on the fleet.
With a clean source clip and permission, agents create a reusable voice ID, annotate pacing and emotion tags, generate WAV audio, and deliver it where the work is happening.
Instant face clone via PuLID/InstantID, saved characters by name, and optional full-character LoRA training when a persona needs higher fidelity across dozens of generations.
Real assets from the VA Staffer library: the infrastructure and reference systems behind repeatable media production.
The interesting part is not the button click. It is the chain of decisions the agent handles around it.
Audience, goal, tone, subject, format, delivery channel, and whether this needs a new clone or a saved identity.
Image prompts get scene, lighting, and quality cues. Voice scripts get emotion tags, pace, and temperature/speed settings before anything generates.
The skill hits the fleet router. Available Spark capacity handles generation instead of tying every media job to one laptop or one SaaS tab.
The agent creates the visual or audio, checks it against the brief, and adjusts identity strength, model, or script when the first pass is off.
Named files, attachments in the right conversation, and handoff into the page, campaign, video edit, or approval thread — not a dead-end download folder.
“Subscriptions give the agents access to great intelligence. The DGX fleet gives that intelligence somewhere to do real media work.”
That split is intentional. Cloud models are excellent at thinking. Local Sparks are excellent at repeated generation, cloning, and private media capacity under our control.
This is a hybrid system by design. Each side does the work it is best suited to do.
A single workstation can prove the idea. A fleet turns it into an operating capability.
Image and audio work can be heavy. Multiple Sparks give the router more places to send work when several agents need media at once.
Voice cloning does not have to fight hero image generation for the same resources. Different nodes can serve different models cleanly.
Fleet routing gives room to maintain, test, and improve without treating one machine as the entire production system.
Media stops being a separate department of tabs. It becomes another skill an AI Employee can run with human taste still in charge.
Approved cloned voices for promos, announcements, lesson intros, and internal updates — with pacing and emotion control, not flat robot reads.
Founder likeness, team personas, and mascots stay recognizable across images instead of drifting into lookalike territory every generation.
Agents can prepare assets while writing the page, building the funnel, or packaging the campaign — same conversation, fewer handoffs.
Sensitive voice samples and recurring brand characters can live on infrastructure we control, with cloud still used where it wins.
This is a production system for AI Employees: subscriptions for intelligence, OpenClaw skills for process, and a six-Spark fleet for media generation and cloning.
It is not autopilot brand taste. Humans still approve likeness, voice use, final creative direction, and what ships publicly. That boundary is what keeps the system useful instead of chaotic.
More from the lab if you want the surrounding infrastructure story.
The goal was never more AI tools. The goal was an AI Employee that can move from idea to useful asset without handing the operator a scavenger hunt.
Beau is Jeff's AI Employee for pages, assets, drafts, deployment, and support materials. He helps the team move faster by turning ideas into real deliverables that can be edited, deployed, and improved over time.