Describe an idea, the assistant sharpens the prompt, and 0G renders the images and video.
The assistant asks what you left out and explains why each answer changes the result. Then Z-Image-Turbo and MiniMax H3 render it through 0G Compute, over a route you can verify on-chain.
DeepSeek V4 Flash runs a short interview instead of taking a prompt. One question at a time, always with the reason it matters, always with real options you can ignore. Five answers later you have a brief worth rendering, whether that is one still or a whole video.
A matte ceramic mug for a product page. Four angles.
What is it sitting on, and against?
WhyBackgrounds are where generated images fall apart. Naming the surface and the backdrop removes most of the mush.
Worn woodWarm, tactile, lived-in
MarbleCool, clean, premium
Seamless paperOne flat colour, no context
In its real placeA kitchen, a desk, a street
Worn oak. Warm, a bit beaten up.
How should it feel?
WhyMood ties the grade, the palette, and the styling together so a set of images reads as one set.
WarmInviting, lived-in, unhurried
ClinicalPrecise, cool, technical
PremiumQuiet, dark, expensive
PlayfulSaturated, bold, direct
Brief complete. Nothing renders until you approve.
A matte ceramic mug for a product page. Four angles.
What is it sitting on, and against? Backgrounds are where generated images fall apart. Naming the surface and the backdrop removes most of the mush. Answer: Worn wood.
How should it feel? Mood ties the grade, the palette, and the styling together so a set of images reads as one set. Answer: Warm.
Free-type anything instead. The options are shortcuts, not a menu.
Brief so far
subject
A product
framing
Product, centred
light
Window, from the left
surface
Worn oak
mood
Warm
Prompt for Z-Image-TurboReady
Matte ceramic mug, centred with room to breathe, soft morning window light from the left, long falloff across the body, visible glaze texture, 4:5. On worn oak against a neutral backdrop.Warm and unhurried.
Subject, framing, light, surface, mood. The answers compose into a single prompt, and the same brief renders a consistent set instead of five unrelated images.
01What is the subject, exactly?
02How close are we?
03Where is the light coming from?
04What is it sitting on, and against?
05How should it feel?
Four images from one brief · 840 credits · ~8s each
Generating a video
Five questions about a sequence
Subject, hook, setting, tone, payoff. The answers become a shot list. Every beat carries a frame prompt for z-image and a motion direction for MiniMax H3.
01Who or what is on camera?
02What happens in the first second?
03Where are we?
04How should it feel?
05What do you want them to do at the end?
3 shots · 12s · 5,480 credits · MiniMax-H3
Nothing renders untilYou approve the briefYou see the credit costYou can still edit every line
Features
Built for a set.Not one lucky render.
A brief you can read, images and video that match each other, and a line item for every request that ran.
One reference, one consistent set
z-image carries a locked reference across every image in a brief, so the fourth angle is the same product, and the same person, as the first.
01
02
03
04
Reference locked. Change the angle without changing the object.
Guided prompting
Every question the assistant asks comes with the reason it matters, so the next brief you write is better than this one.
How close are we?
Framing decides what the image is about. A macro of one detail and a wide of the whole thing are two different arguments.
MacroProductPortrait
Approve the frame, then move it
A still is a finished deliverable on its own. When you want motion, that exact image becomes the locked first frame. You know what is moving before you pay for it.
z-image · 40 crbecomes the first frame of
MiniMax H3 · 6s
Brand kit
Palette, type, and tone saved once. Every brief the assistant writes inherits them, whether the output is a still or a sequence.
Queue variants of one brief and 0G Compute routes them upstream in parallel.
Ceramic mug · 4 anglesdone
Skincare range · flat lay64%
Shoe drop · 20s reveal22%
Re-render one image without touching the rest.
A receipt for every request
0G's broker serves each request over a hardware-attested route. Utsuro writes the model, the work, the cost and that route to 0G Chain as a receipt you can open in the explorer.
Recent on-chain inference receipts
Tx
Model
Work
Cost
Route
0x9f3a…c41b
z-image-turbo
4 images · 4:5
160 cr
Attested
0x1d70…8ae2
minimax-h3
8s motion · 24fps
2,080 cr
Attested
0x4c92…07f5
deepseek-v4-flash
1.4k tokens · interview
3 cr
Attested
Settled per request. What a receipt proves is the cost and the route that served it. The model itself runs at the upstream, under that provider’s policy.
The stack
Three models.One decentralized network.
Utsuro does not host GPUs. Every model is reached through 0G Compute on Verified Routing: the route from 0G’s broker to the upstream model is hardware-attested and provable on-chain, and every request settles per call.
Image generation
Z-Image-Turbo
https://router-api.0g.ai/v1/z-image-turbo
Photoreal stills at speed. Generate finished images on their own, or use one as the locked first frame of a video so you know exactly what will move.
Cost
60 cr / image
Latency
~8s
Photoreal
Text-to-image
1024 × 1024
Two per request
Verified route
Verified Routing. The route to this model is attested on-chain
Video generation
MiniMax-H3
https://router-api.0g.ai/v1/minimax-h3
Motion with intent. Takes a prompt, or an image plus a direction, and renders footage with coherent camera work instead of drifting slideshow.
The first video model on the 0G Private Computer. Weights went public the same morning it landed here; Artificial Analysis ranked it #2 text-to-video and #3 image-to-video on day one.
Cost
300 cr / second
Latency
minutes
Text-to-video
Image-to-video
4–15s
2K
Verified route
Verified Routing. The route to this model is attested on-chain
Creative assistant
DeepSeek V4 Flash
https://router-api.0g.ai/v1/deepseek-v4-flash
The part that makes the other two worth using. It asks what you left out, explains why it matters, and turns a half-formed idea into a prompt worth rendering.
The 0731 release is now the official one, swapped in under the same deepseek-v4-flash id. Terminal Bench 2.1 went from 61.8 to 82.7, on a 13B-active MoE.
Cost
2 cr / 1k tokens
Latency
~600ms
Guided prompting
Shot direction
Composition notes
Rewrites and variants
Streaming
Verified Routing. The route to this model is attested on-chain
All three models run on Verified Routing: the path from 0G's broker to the upstream is hardware-attested and provable on-chain. The inference itself runs at the upstream, under that provider's policy. What you can check is which route served your request.
0G Storage
Outputs you actually hold
Images and video are pinned and content-addressed, not parked in a bucket that disappears when a subscription lapses.
0G Chain
Costs you can audit
Each request settles on-chain at its listed rate. Every credit you spend has a receipt with a hash behind it, and no month-end number you have to take on trust.
0G Mainnet · chain 16661. All three of these landed on 0G the same day they shipped, and you do not have to take that on faith: the created timestamps are public at router-api.0g.ai/v1/models. Credits are bought with a card and settled on-chain behind the scenes. You never need a wallet to generate anything.
How it works
Four steps.One of them is talking.
One line is enough
~10 seconds
No format, no keyword soup, no negative prompts. Say what you want and whether you want it still or moving. If you do not say, the assistant asks.
Your idea
A matte ceramic mug for a product page. Four angles.
StillVideo
What you get
Every frame herewas made with Utsuro.
MiniMax H3
z-image
z-image
z-image
MiniMax H3
MiniMax H3
MiniMax H3
Stills and video from the same three models. Three of the tiles below are the video files themselves, playing as they came out of MiniMax H3, not frames pulled from them. Each tile carries the model that made it and what it cost in credits.
Ceramic mug · three-quarterz-image40 cr
Mug to camera · counter light · 5sMiniMax H31,300 cr
Images and video. Product stills, flat lays, brand portraits, hero frames: z-image renders those on their own, and plenty of people never generate anything else. When you do want motion, MiniMax H3 takes a prompt, or a still you already approved, and moves it with real camera direction. Both come out of the same interview.
Do images and video cost the same?
No, and the gap is large. An image is 40 credits. A second of video is 260, so ten seconds of motion is 2,600. The assistant session costs a few hundred either way. A four-image product set lands around 760 credits total, under a dollar. A fifteen-second, three-shot video lands near 4,600.
Who owns what I generate?
You do. Utsuro claims no rights over your briefs, your images, or your finished video. We do not train on your projects, and nothing you make appears anywhere public unless you put it there yourself.
Can I use the output commercially?
Yes, on every plan including Starter. Commercial use is included, there is no separate licence tier, and there is no watermark on anything you generate. The usual limits still apply: no real person's likeness without their permission, no trademarks you do not own.
Why is a blockchain involved at all?
Three concrete reasons. 0G Compute reaches every model on Verified Routing: the route from 0G's broker to the upstream is hardware-attested and provable on-chain, so you can check which route served a given request rather than taking a vendor's word for it. What that does not mean: the inference itself runs at the upstream provider, under that provider's policy. The tier attests the path, not the enclave the model ran in, and we are not going to claim otherwise. 0G Storage pins your outputs so they are content-addressed and still yours if a subscription lapses. 0G Chain settles each request at its listed rate, so your bill is a list of receipts instead of a number you have to trust. You buy credits with a card and never need a wallet.
What is actually new here?
MiniMax H3 is the first video model on the 0G Private Computer. Text-to-video, or image-to-video from a first frame you have already approved, at 2K, 4 to 15 seconds, billed per generated second. Its weights went public the same morning it landed, and Artificial Analysis ranked it #2 text-to-video and #3 image-to-video on day one. Alongside it, DeepSeek V4 Flash 0731 is now the official release, swapped in under the same deepseek-v4-flash id, with Terminal Bench 2.1 up from 61.8 to 82.7 on a 13B-active MoE. 0G took all three of that morning's launches the same day; the created timestamps are public at router-api.0g.ai/v1/models.
What if the result is wrong?
Most of the time you catch it before it costs anything, because you approve the brief first. Rewrite the prompt, change the light, cut a beat, or send the assistant back with a note. After a render, re-roll a single image or a single shot without touching the rest of the set. You only pay for what actually ran.
Start free
Bring the idea.It will ask for the rest.
One sentence is enough to start. Ninety seconds later you have a brief you would have been proud to write, and images or video to match.