Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

Image Generation, In Practice

Reference images, inpainting, character LoRAs, upscaling, and the licence nobody read.

Image Generation, In Practice

Reference images, inpainting, character LoRAs, upscaling, and the licence nobody read.

Level
Nothing assumed
Lessons
69
Reading time
547 min
Price
Free, no sign-up to read

The craft of getting a specific, finished image out of a model on a deadline. Choosing between hosted tools and open weights on cost, control and licensing. Preset platforms like Higgsfield and what they take away. Reference images and denoise strength, inpainting and outpainting, pose and depth conditioning, character consistency, upscalers that invent detail, the finishing pass in a raster editor, and what you must tell a client.

Opens after the Vectors, Logos and Things That Get Printed exam

Sign in, finish that course, and pass its exam. You can read this syllabus meanwhile.

Go to Vectors, Logos and Things That Get Printed

Download the textbook (PDF) · free to print and teach from, with the exam paper and every answer at the back.

Module 1

11 lessons · 88 min

The kit, and what each piece of it costs you

Before any picture, a set of decisions that are easy to get wrong and expensive to reverse: hosted service or your own machine, which files a "model" actually consists of, which interface to spend your learning time on, how to make a large model fit a small graphics card, and why the same prompt produces different pictures on two platforms. This block ends with a setup you can reproduce in six months.

By the end you can

Assemble a working image-generation setup from the hardware you actually own — hosted service, local install or a free cloud GPU — and state for your chosen route the VRAM it needs, the seconds and rupees per image it costs, the licence it carries, and the exact files on disk that make up the model

  1. 1The generator you pick is a licensing decisionLocked — this takes you to what opens it. 9 minThe licence on a model's weights governs whether you may use it commercially, and it is a completely separate question from who owns the picture it made.
  2. 2A preset is somebody else's workflow, frozenLocked — this takes you to what opens it. 8 minRecipe-driven tools trade control for speed, and the bill arrives at the first revision, when the thing the client wants changed is the thing the preset already decided.
  3. 3Three routes to a GPU, and the arithmetic for eachLocked — this takes you to what opens it. 9 minThe choice between a hosted service, your own card and a free cloud notebook is decided by how many images you make per month and how much control you need, and the crossover is usually somewhere around a few thousand images.
  4. 4A model is four kinds of file, and mixing them up wastes an eveningLocked — this takes you to what opens it. 8 minA checkpoint, a VAE, a LoRA and an embedding are different file types with different sizes and different folders, and almost every "it loaded but the image is grey mush" problem is one of them in the wrong place.
  5. 5Four front ends, and which one deserves your learning timeLocked — this takes you to what opens it. 8 minThe interface you learn decides which problems are easy, and the right choice is the simplest one that can express the workflow you actually need — which for finished client work usually means a node graph eventually.
  6. 6The first hour, and the four errors everybody getsLocked — this takes you to what opens it. 8 minAlmost every failed local install is one of four specific errors, each with a one-line cause, and knowing them turns a lost weekend into twenty minutes.
  7. 7Steps, samplers and the seconds you should actually expectLocked — this takes you to what opens it. 8 minDoubling the step count past roughly 30 buys almost nothing on a standard model, while a distilled model does a usable image in 4 to 8 steps, so measure iterations per second and spend your time where the curve is still steep.
  8. 8Making a big model fit a small cardLocked — this takes you to what opens it. 8 minQuantisation and offloading trade quality or speed for memory in known amounts — fp8 costs almost nothing, Q4 costs a little detail, and CPU offload costs seconds rather than quality.
  9. 9Every platform's filter is different, and that is a production riskLocked — this takes you to what opens it. 7 minA refusal is produced by a keyword layer, an output classifier and an account-level policy stacked together, all tuned for the platform's liability rather than your brief, which is why the same prompt passes on one service and fails on another.
  10. 10The setup that still works in six monthsLocked — this takes you to what opens it. 7 minDiffusion output depends on the exact version of the interface, the model hash and the sampler implementation, so recording those four or five facts with each approved image is what makes a client revision six months later a ten-minute job.
  11. 11You are paying to make decisions you could make cheaplyLocked — this takes you to what opens it. 8 minChoose the frame on cheap low-step renders and spend full quality only on the winner; after about six full-quality attempts the remaining distance is editing work, not generation.

Module 2

10 lessons · 77 min

From a request to a specification you can generate against

The work that decides whether a job goes well happens before the first prompt. This block turns "we need something for the launch" into a document: the exact pixel dimensions the image has to land in, the aspect ratio the model can natively produce, a reference pack where every picture has a stated purpose, a written list of what must never appear, and an honest note about which parts of the brief a generator cannot deliver at all.

By the end you can

Convert a one-line request into a written image specification — final pixel dimensions, the native generation ratio and the crop-safe area inside it, colour and brand constraints, a purpose-labelled reference pack and an explicit forbidden list — and state which elements will have to be photographed or composited rather than generated

  1. 12Start at the placement, not the pictureLocked — this takes you to what opens it. 8 minThe final placement fixes the dimensions, the crop, the amount of detail that will ever be visible and the file format, and every one of those decisions is cheaper to make before generating than after.
  2. 13The ratios the model was trained on, and what happens outside themLocked — this takes you to what opens it. 8 minModels are trained at a specific set of resolutions, and generating far outside those buckets produces duplicated subjects and stretched anatomy, so you pick the nearest trained ratio and crop to the delivery size afterwards.
  3. 14Designing for the crop you will not controlLocked — this takes you to what opens it. 8 minPlan the frame as concentric zones — a core that survives every crop, a middle that survives most, and an outer band that exists to be thrown away — because the image will be re-cropped by people and platforms who never see your original.
  4. 15Six references, each with a jobLocked — this takes you to what opens it. 7 minA mood board of vibes cannot be acted on, but six references each labelled with the single thing it is there to demonstrate can be used directly as conditioning and as the basis for sign-off.
  5. 16Write down what must not appearLocked — this takes you to what opens it. 7 minA negative prompt is a weak statistical nudge rather than a guarantee, so anything that would be genuinely damaging in the final image belongs on a written checklist that a person verifies before delivery.
  6. 17A hex value is not something a diffusion model can hitLocked — this takes you to what opens it. 8 minDiffusion models cannot be instructed to produce an exact colour, so brand colour is achieved by generating in a neutral or near palette and then grading or compositing the exact value in afterwards.
  7. 18One hero or twelve pieces is a different jobLocked — this takes you to what opens it. 7 minA single image is a search problem and a set is a consistency problem, and they need different budgets, different methods and a decision made before any generating starts.
  8. 19The frames where a phone camera winsLocked — this takes you to what opens it. 7 minAnything that must be exactly itself — a real product, real premises, real staff, food you are selling — is faster, cheaper and safer to photograph than to generate, and recognising those frames early is a professional skill rather than a failure.
  9. 20The one page that ends the revision argumentLocked — this takes you to what opens it. 7 minA signed one-page spec that names dimensions, ratio, forbidden content, the reference pack and what will be photographed rather than generated converts later disputes from matters of taste into matters of record.
  10. 21What you can sell, and what you must discloseLocked — this takes you to what opens it. 10 minBeing licensed to use a model, owning its output, and having to disclose that output is generated are three separate questions, and only the ownership one is genuinely unsettled.

Module 3

10 lessons · 74 min

Getting to a frame worth keeping

Exploration is where most of the time goes and where most of it is wasted. This block replaces prompt roulette with a method: a prompt built from named slots, the photographic vocabulary that measurably changes an image and the vocabulary that does nothing, a cheap contact sheet you judge on a grid, fixed seeds so a comparison means something, and a written rule for when to stop generating and start editing.

By the end you can

Run a structured exploration that reaches a usable frame inside a stated budget — a prompt assembled in named slots, a low-cost contact sheet judged on a grid, seed-fixed single-variable comparisons — and apply an explicit test for whether a frame is worth taking into repair or should be abandoned

  1. 22Build the prompt in slots, not as a sentence you keep rewritingLocked — this takes you to what opens it. 8 minWriting a prompt as seven named slots makes it debuggable, because when the image is wrong you can change one slot and know which part of the description caused the change.
  2. 23The camera vocabulary that works, and the words that do nothingLocked — this takes you to what opens it. 8 minFocal length, aperture, camera height and shot size measurably change composition because they appeared in the captions of millions of training photographs, while quality words like "8k" and "masterpiece" mostly do nothing on modern models.
  3. 24Name the light or accept the model's averageLocked — this takes you to what opens it. 8 minLight is the single highest-leverage slot in a prompt, and leaving it unspecified is the main reason generated images look generically generated.
  4. 25Describing a look without naming a living artistLocked — this takes you to what opens it. 8 minA look is more reliably reproduced by naming its process, era, materials and printing than by naming an artist, and the process description also avoids the ethical and commercial problem of trading on a living person's name.
  5. 26Sixteen cheap frames beat four expensive onesLocked — this takes you to what opens it. 7 minJudging a grid of many low-cost renders at once is faster and more accurate than judging full-quality images one at a time, because composition is visible at thumbnail size and commitment bias is not.
  6. 27Fix the seed or learn nothingLocked — this takes you to what opens it. 7 minA seed fixes the starting noise, so changing it and the prompt together means you cannot attribute the difference to either, and seed-locked comparison is what turns guessing into testing.
  7. 28Change one thing, or you have learned nothingLocked — this takes you to what opens it. 7 minA comparison is only informative when one variable moves, and the habit of changing prompt, sampler and steps together is why so much accumulated knowledge about image generation turns out to be wrong.
  8. 29Recognising the twenty minutes that will never convergeLocked — this takes you to what opens it. 7 minSome requests fail for structural reasons rather than for want of a better prompt, and the professional skill is recognising them inside three attempts and switching method instead of continuing.
  9. 30Interrogating a picture for vocabularyLocked — this takes you to what opens it. 7 minAn interrogator recovers usable vocabulary from a reference image but never recovers the image, because it is describing appearance rather than inverting the generation process.
  10. 31What "good enough to take forward" actually meansLocked — this takes you to what opens it. 7 minA frame is worth repairing when its composition, light and structure are right, because those are the things repair cannot create, while hands, small objects and texture always can be fixed.

Module 4

10 lessons · 82 min

Steering with pictures instead of adjectives

Words run out. Everything a brief asks for that a prompt cannot deliver — this composition, this perspective, this palette, this pose, this part of the frame and not that part — is delivered by giving the model an image instead. This block covers the ladder of denoise strengths, the five-minute sketch, colour blocks painted in a free editor, a depth map rendered from Blender primitives, the difference between the control types, prompting regions separately, and how to stack two controls without producing something that looks traced.

By the end you can

Take control of composition by conditioning on an image — a photograph, a biro sketch, a flat colour block, a depth render or a pose skeleton — choose the conditioning type that preserves what the brief requires and discards what it does not, and state the strength and end-step you used and the failure each setting was avoiding

  1. 32Stop describing it and show it the pictureLocked — this takes you to what opens it. 8 minDenoise strength decides how much of your reference survives; the useful range is roughly 0.4 to 0.6, and everything above 0.75 throws the reference away.
  2. 33The model is guessing your compositionLocked — this takes you to what opens it. 9 minConditioning makes composition an input rather than an outcome, and letting the control end around 60% of the steps is what stops the result looking traced.
  3. 34Three passes down the denoise ladderLocked — this takes you to what opens it. 8 minRunning several img2img passes at decreasing denoise strength converges on a finished image more reliably than one pass at a middling value, because each pass only has to make a small correction.
  4. 35A biro sketch is the most reliable composition tool you ownLocked — this takes you to what opens it. 8 minFive minutes of rough drawing photographed on a phone gives the model a composition it will follow, and it requires no drawing skill because the conditioning reads structure rather than quality of line.
  5. 36Paint flat shapes, then let the model make them realLocked — this takes you to what opens it. 8 minA crude painting of flat colour shapes run through img2img at low denoise places objects and palette exactly, because the colour distribution of the input survives the parts of the schedule that decide layout.
  6. 37Get the perspective right by building itLocked — this takes you to what opens it. 9 minRendering a depth map from a few primitives in Blender gives the model correct perspective and object placement that no sketch or prompt can supply, and it is the standard professional route for interiors, architecture and product staging.
  7. 38Canny, lineart, depth, softedge, pose, normalLocked — this takes you to what opens it. 8 minEach control type preserves one property of the input and discards the rest, so the choice is made by naming what must survive from the source and what must be free to change.
  8. 39Different words for different parts of the frameLocked — this takes you to what opens it. 8 minRegional prompting assigns separate text to separate areas of the image, which solves attribute binding failures that no amount of rephrasing a single prompt can fix.
  9. 40Transferring a look from an image without training anythingLocked — this takes you to what opens it. 8 minAn image-prompt adapter injects a reference image's appearance through the model's attention layers, giving style transfer in seconds rather than the hours a LoRA costs, at the price of much less control over what exactly is transferred.
  10. 41Two controls, without the plastic lookLocked — this takes you to what opens it. 8 minStacking conditioning signals removes the model's freedom in proportion to the total constraint, so the sum of the strengths matters more than any individual value and should stay near 1.2 to 1.5.

Module 5

10 lessons · 82 min

Repair, extend, assemble

The base render is raw material. This block is the work between that and a deliverable: a triage order that fixes structure before detail, the specific settings that repair hands and faces, the only reliable way to put words in a picture, extending a frame when the crop changes, removing an object without leaving a ghost, putting a real photographed product into a generated scene so the shadow and grain agree, and enlarging a canvas past what the model can render in one pass.

By the end you can

Take a flawed base render to a deliverable by masked repair — fix hands, faces and small objects at full resolution with stated denoise values, extend the frame without a visible seam, remove an object cleanly, and composite a photographed item into a generated scene with matching perspective, contact shadow, optics and grain

  1. 42The finished image is never one generationLocked — this takes you to what opens it. 9 minGenerate a base plate, then repair it region by region; the "inpaint at full resolution" setting is what gives a small detail enough pixels to be drawn correctly.
  2. 43Fix it in the order that stops you doing it twiceLocked — this takes you to what opens it. 8 minRepair has a correct order — structure, then faces and hands, then edges, then small objects, then grade and grain last — because each stage changes the pixels the next stage would otherwise have worked on.
  3. 44The three repairs everybody needs, with settingsLocked — this takes you to what opens it. 8 minHands, eyes and teeth fail because they are small, high-variance structures rendered at too few pixels, so the fix is always to give the region more pixels rather than to describe it better.
  4. 45Type is set, not generatedLocked — this takes you to what opens it. 8 minGenerated text fails at unusual words, small sizes and long strings because a diffusion model renders letterforms as visual texture rather than assembling glyphs, so any text that must be correct is set in a type tool over a clean plate.
  5. 46Extending a frame when the crop changesLocked — this takes you to what opens it. 8 minOutpainting extends an image in overlapping steps of a few hundred pixels, and the overlap is what lets the model match light, perspective and texture rather than inventing a discontinuous new scene.
  6. 47Four ways to remove an object, and when each is rightLocked — this takes you to what opens it. 8 minClone, heal, content-aware fill and inpainting fail in different ways, so the choice depends on whether the area behind the object is a texture, a structure or a scene that has to be invented.
  7. 48The real product in a generated worldLocked — this takes you to what opens it. 9 minA photographed product composited into a generated scene is convincing only when perspective, light direction, contact shadow, reflection and grain all agree, and the contact shadow is the one that gives it away.
  8. 49Making two sources agreeLocked — this takes you to what opens it. 8 minEvery image carries a signature of noise, sharpness, depth of field and chromatic behaviour, and a composite reads as fake when two sources disagree on that signature rather than on content.
  9. 50The fault the client sees on a big screenLocked — this takes you to what opens it. 8 minA composite or repair fails at its boundary, and the three specific faults — colour contamination, a bright or dark halo, and over-feathering — each have a distinct cause and a distinct fix.
  10. 51Getting to poster size without a second horizonLocked — this takes you to what opens it. 8 minA large final image is produced by generating at native size and enlarging in a separate pass, and where extra generated detail is needed the canvas is processed as overlapping tiles with a low denoise so no tile can invent a new scene.

Module 6

9 lessons · 72 min

The same thing, twelve times

One good image is a search problem; twelve that belong together is an engineering one. This block builds the character sheet you generate from, the practicalities of training a LoRA on free hardware including the dataset rules that decide whether it works, the decision between an adapter and a trained model, why a product is harder than a face, the constants that hold a visual system together across a set, a naming scheme that survives the revision round, and a worked run of twelve images from start to delivery.

By the end you can

Deliver a set of images that hold one character, one product and one visual system — build a character sheet, choose between an image adapter and a trained LoRA on a stated cost-and-quality basis, prepare a training set that avoids the four common dataset faults, and hold light, palette, lens and finish constant across a campaign

  1. 52The same face twice is the hard partLocked — this takes you to what opens it. 9 minOne-photo identity adapters get you a family resemblance; a specific real person needs a LoRA trained on varied photographs, or their actual photographed face composited in an editor.
  2. 53Build the reference before you build the imagesLocked — this takes you to what opens it. 8 minA character sheet turns an identity from something you hope recurs into an asset you condition on, and building it first is what makes every later frame a matter of reuse rather than of luck.
  3. 54Training a character LoRA on hardware you can borrowLocked — this takes you to what opens it. 9 minA usable character LoRA trains from 15 to 30 images in under an hour on free cloud hardware, and the settings that matter are rank, learning rate and total steps, with overfitting visible as the training images reappearing in the output.
  4. 55Four dataset faults, and the LoRA each one ruinsLocked — this takes you to what opens it. 8 minA LoRA learns whatever is constant across its training images, so variety in everything except the subject is the property that decides whether it works, and no training setting compensates for a set that lacks it.
  5. 56The decision, with the numbersLocked — this takes you to what opens it. 7 minAn adapter gives roughly seventy per cent of an identity in under a minute and a trained LoRA gives most of the rest in about an hour, so the decision turns on how many images the identity has to survive and how closely it will be examined.
  6. 57A product is harder than a faceLocked — this takes you to what opens it. 8 minA face is forgiven small variation and a product is not, because a viewer compares the rendered product against the real one they can buy, so compositing a photograph is almost always the correct answer.
  7. 58Five constants that do more than the faceLocked — this takes you to what opens it. 8 minCoherence in a set is produced by holding light, lens, palette, crop and finish constant, and these are cheap to enforce and far more visible to a viewer than small drift in a character's features.
  8. 59A file scheme that survives the revision roundLocked — this takes you to what opens it. 7 minA naming convention that encodes project, frame, version and status lets you find the approved file in seconds six months later, and its absence is what turns a small revision into an afternoon.
  9. 60A worked run, with the budget and the failureLocked — this takes you to what opens it. 8 minA twelve-image campaign is about thirty hours of work distributed unevenly — a third on establishing the system, most of the rest on repair and consistency, and very little on generation itself.

Module 7

9 lessons · 72 min

Finishing and handing over

The last few hours decide whether the work looks professional. This block covers the resolution the job actually needs rather than the number people quote, what happens to colour between a screen and a press, making type legible over a busy photograph, the twelve-point check that catches what a client would otherwise find, exporting without visible compression damage, writing alt text for a generated image, and handling the revision that turns out to be a rebuild.

By the end you can

Take an approved frame to a delivered file — reach the output resolution the medium genuinely requires, prepare colour for screen or press, set type that stays legible over a photographic background, run a named twelve-point QA pass, and hand over a package with disclosure, alt text and archived source files

  1. 61Upscaling either restores detail or invents itLocked — this takes you to what opens it. 8 minDiffusion-based upscalers re-generate rather than sharpen, so never point one at text, a logo, a product detail or a face you need to stay recognisable.
  2. 62Nothing goes to a client straight from the modelLocked — this takes you to what opens it. 8 minCurves for contrast, a colour pass to match the set, fine grain to fake a sensor, and real type over any generated words — that is the ten minutes that separates output from a deliverable.
  3. 63300 dpi is one number applied to everythingLocked — this takes you to what opens it. 8 minRequired resolution is set by viewing distance rather than by a universal figure, so a billboard needs far fewer pixels per inch than a brochure and most "high res please" requests are satisfied by a file you already have.
  4. 64The blue on your screen is not going on the pressLocked — this takes you to what opens it. 8 minScreens emit light and presses absorb it, so a set of vivid screen colours cannot be printed at all, and the conversion has to be seen and approved before delivery rather than discovered on the printed sheet.
  5. 65Making words readable on a photographLocked — this takes you to what opens it. 8 minType over a photograph fails on contrast rather than on typeface, and the fixes are all about creating a controlled surface for the words rather than about choosing better fonts.
  6. 66Twelve things to check before you sendLocked — this takes you to what opens it. 8 minA written check performed at 100% on the final file catches the specific faults that generated images produce and that the person who made the image has stopped being able to see.
  7. 67Exporting without wrecking it at the last stepLocked — this takes you to what opens it. 8 minExport is where visible damage is introduced by compression, colour profile loss and resampling, and each has a specific setting that prevents it.
  8. 68Describing the image, including what it isLocked — this takes you to what opens it. 8 minAlt text describes what an image conveys in its context rather than listing what is in it, and whether to mention that an image is generated depends on whether that fact matters to understanding it.
  9. 69When "just change one thing" is a rebuildLocked — this takes you to what opens it. 8 minRevisions divide into adjustments, regional changes and structural changes, and the professional skill is saying which category a request falls into before agreeing to it.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly