Guide a new image with the edges or depth of a reference so the composition stays controlled while the visual treatment changes.

ControlNet separates two things that a plain prompt keeps tangled: where things are, and what they look like. ControlNet reads a structural map out of your reference — an edge trace or a depth field — and holds the generator to that geometry while the prompt decides material, lighting, palette, and subject. That is the whole idea, and it explains both the strengths and the limits below.
Because ControlNet constrains structure rather than content, it is the right tool when you already know the layout. A floor plan sketch that should become a render, a product silhouette that should change material, a pose that should be worn by a different character, a wordmark that should keep its letterforms in a new style — all of these are structure-first problems.
It is the wrong tool when the thing you want preserved is identity or texture. ControlNet has no concept of “the same person” or “the same fabric”. ControlNet knows edges and distances. Asking it to keep a face is the single most common reason people conclude that ControlNet does not work, when in fact they wanted a different workflow.
Picking the wrong map is a more common failure than picking the wrong prompt.
| Canny (edges) | Depth | |
|---|---|---|
| What it keeps | Outlines and internal lines | Spatial arrangement, near/far relationships |
| Best for | Products, line art, typography, architecture elevations | Interiors, landscapes, scenes with clear foreground and background |
| Fails when | The reference is soft or low-contrast — no edges to extract | The scene is flat, so depth carries almost no information |
| Typical mistake | Strength so high the output traces the reference | Expecting it to hold fine detail, which it never does |
| Verify by | Following every outline in the output | Checking that objects sit at the right distances |
Inspect the extracted map before you judge any output. Most disappointing results are visible in the map first.
01
Match the aspect ratio you want to output, and give it enough resolution to extract clean structure. A blurry reference produces a noisy map, and ControlNet reproduces that noise as wobbly geometry.
02
Ask what must survive. If the answer is an outline, use canny. If it is a spatial layout, use depth. Do not default to one because it worked last time.
03
The reference already states the composition. Spend the prompt on material, lighting, palette, and subject — the parts ControlNet is deliberately leaving open.
04
Output ignoring the reference means raise it. Output tracing the reference means lower it. This one dial resolves most ControlNet problems on its own.
A rough line drawing becomes a finished visual with the proportions intact. The clearest demonstration of what ControlNet is for.
One product silhouette, many finishes — brushed aluminium, matte ceramic, worn leather — with the shape identical across the set.
ControlNet depth keeps rooms and facades spatially correct while daylight, season, and finish change between variants.
Keep a pose and re-cast the character. ControlNet holds the skeleton; the prompt supplies who is standing in it.
Canny holds letterforms well enough to restyle a wordmark and still read it. Check every character afterwards.
Reuse one ControlNet reference across a batch so every frame shares a layout — useful for card sets, icon families, and storyboards.
Each one leaves composition to ControlNet and spends the words on everything else.
Canny from a product sketch — render as brushed aluminium under soft studio light, charcoal seamless background.
Depth from an empty room photo — furnish as a warm Scandinavian living room, late afternoon daylight.
Canny from a logo outline — restyle as chrome with sharp specular highlights, keep every letter readable.
Depth from a street photo — same layout at night, wet asphalt, neon signage, no people.
Canny from a character line drawing — finish as cel-shaded anime with flat colour and clean line art.
Depth from a landscape — same terrain in winter, overcast, snow cover, muted palette.
Canny from an architectural elevation — photoreal facade in weathered concrete and glass.
Depth from a tabletop scene — replace every object with ceramics, keep the arrangement exact.
Do
Don't
01Strength is the dial that fixes most problems — try it before rewriting anything.
02Canny holds outlines; depth holds distances. Neither holds identity.
03A hand-drawn reference often controls ControlNet better than a photo, because its edges are unambiguous.
04Reuse one reference across a batch to get a consistent series for free.
05After ControlNet locks the layout, do local fixes in a general image editor rather than another controlled run.
It extracts a structural map from your reference — an edge drawing or a depth field — and feeds that map to the generator alongside the prompt. ControlNet does not copy the reference's pixels; it copies its geometry, then lets the prompt decide everything else.
ControlNet canny keeps outlines, so use it when the silhouette and internal lines matter: product shapes, architecture, line art, typography. Depth keeps the spatial arrangement, so use it when you want the same layout in three dimensions but different surfaces. ControlNet behaves very differently between the two.
Usually the control strength is too low, or the extracted map is nearly empty. A soft, low-contrast photo produces almost no canny edges, and ControlNet then has nothing to hold. Check the extracted map before blaming the prompt.
The opposite ControlNet problem — control strength too high, or a prompt that mostly restates the reference. ControlNet is meant to constrain composition, not content. Lower the strength and make the prompt describe something the reference is not.
Not reliably. ControlNet preserves pose and proportion, not likeness. If you need the same face, work from an image-to-image or reference-driven workflow instead; if you need the same pose on a different person, ControlNet is exactly right.
Canny holds letterforms surprisingly well, which makes ControlNet useful for restyling a wordmark while keeping it readable. Verify every character afterwards — a broken edge in the map becomes a broken letter in the output.
Match the output aspect ratio and give it enough pixels to extract clean structure. A small or blurry reference yields a noisy map, and ControlNet faithfully reproduces that noise as wobbly geometry.
Yes, and the order helps. Use ControlNet to lock the composition first, then take the approved frame into a general image editor for local fixes. Trying to do both in one prompt is where most runs go wrong.
ControlNet constrains geometry. When the thing that must survive is a subject rather than a layout, an image-to-image workflow is the better starting point.
Bring a reference with clear structure, pick canny or depth deliberately, and spend the prompt on what should change.
Open the generatorUse a reference image as a structural guide while changing the subject, material, lighting, or visual style. ControlNet is useful when a normal prompt keeps drifting away from the composition you already approved.
Canny guidance follows visible boundaries, so it works well for product silhouettes, architecture, line art, and compositions with clean edges. Depth guidance follows near and far relationships, which is more useful for rooms, people, vehicles, and scenes where perspective matters more than every contour.
Start with one control signal. If the result is too rigid, simplify the source or reduce the amount of detail in the prompt. If the result drifts, use a cleaner reference with a strong subject outline and fewer competing background elements.
The control image already describes much of the layout. Use the prompt to name the new subject, material, environment, lighting, color palette, and finish. Add preserve instructions for the parts that must remain recognizable, such as camera angle, pose, product outline, or room perspective.
Avoid asking the model to keep the structure and completely replace every object at the same time. Change one visual layer first, compare the result to the reference, then make a second pass if the composition still holds.
Use ControlNet when geometry is the priority: pose, edges, depth, perspective, or placement. Use image to image when identity, texture, color, or a localized edit matters more than reproducing the exact structure.
Review hard edges, repeated lines, fingers, product labels, and contact shadows at full size. A result can match the broad map while still breaking small geometry, so keep the source beside the output during approval.
Optional cookies help us understand usage and measure ads. Essential cookies stay on so Voor works.