| Takeaway | Detail |
|---|---|
| Plausibility is not provenance. | Hist10K quotes $0.50-$3 per diffusion-colorized photo; cheap plausible chroma is not a verified source-era color. |
| Uncertainty should remain visible. | At $200-$800 per complex print, archival hand coloring still warrants documentary cross-checking; preserving gray may be more responsible than inventing a decisive hue. |
| Anachronism needs its own check. | The quoted $7.99 consumer option can produce a plausible image while period evidence remains insufficient to rule out object-level anachronisms. |
| Realism and fidelity are separate tests. | At $0.50-$3 per photo, diffusion may win the Hist10K plausibility vote while a sparse linear solver wins the fidelity vote. |
Saharia et al.’s 14.20 FID for ImageNet colorization looks decisive, but it measures perceptual similarity across images, not whether a particular garment in a family photograph was historically red. That limitation is easy to miss because a smooth, realistic result invites trust before its evidence has been checked. A plausible hue can still be a confident invention.
The first test is documentary fit: does the proposed color agree with surviving textiles, uniforms, captions, shop records, or other period evidence? Clean edges and plausible skin-tone variation cannot replace those sources. The second test is calibrated uncertainty. When grayscale carries no decisive chromatic clue, a responsible workflow may preserve gray, mute the color, or show competing hypotheses instead of turning ambiguity into a single photorealistic answer.
The third test is anachronism control. In Hist10K, diffusion wins the plausibility vote while a sparse linear approach wins the fidelity vote, separating convincing synthesis from reconstruction. That split is a warning sign whenever color is most coherent but least documented. Diffusion’s real advantage is plausible diversity; historical use still requires traceable evidence, explicit uncertainty, and a way to flag generated hues rather than presenting them as recovered chroma.

Why Denoising Steps Cannot Recover Deleted
The non-obvious point is that denoising steps do not run colorization backward; they construct another plausible state from a learned prior. I use Ho et al.’s 2020 paper, Denoising Diffusion Probabilistic Models, as the mechanism anchor: its forward process adds Gaussian noise through a fixed sequence of T steps, while a U-Net predicts the noise εθ needed by the learned reverse process. The reverse trajectory produces a new sample. More iterations can improve internal coherence, but no extra step restores a spectral measurement absent from the scan.
I represent each candidate in CIELAB because that exposes the identifiability problem. A grayscale scan constrains L* but not a* or b*; many RGB triples can produce the same luminance. Even scan texture and context leave a conditional distribution rather than a unique answer. I therefore treat colorization as sampling p(a*, b* | L*, context), not as uniquely reversing deleted color data. The notation is substantive: p records uncertainty, whereas a single inverse would conceal it.
I contrast this with deterministic regression. An L2-trained network estimates the conditional mean; when scarlet and blue are both compatible with the scan, their vector average can produce a muted or off-palette compromise. Diffusion can instead select one coherent scarlet or blue garment. That sample may look more realistic while matching documented pixels less well: higher realism can coexist with lower pixel fidelity. Choosing a plausible mode is not the same operation as estimating the nearest recorded color.
I separate conditioning from evidence. In Stable Diffusion 2.1, as examined in the 2026 Hist10K: Evaluating Diffusion vs Traditional Colorization comparison, grayscale input constrains structure, while text or exemplar conditioning shifts the learned color prior. Conditioning tells the model where to place structure or which color distribution to favor; it does not document the pigment that was historically present. The same grayscale can therefore support different coherent palettes without any change in archival evidence.
I expose stochasticity by holding the scan, prompt or exemplar, model settings, and all other conditions fixed while changing the random seed. Across five seeds, different coherent chroma outcomes are samples from model uncertainty, not five competing attempts to recover one documented truth. Even perfect seed-to-seed agreement would establish repeatability, not provenance. A seed-stable scarlet guess with no dated support still fails the archive standard; any changed object or background fails the required invariance gate regardless of polish. Under the falsifiable 2026 standard, diffusion replaces the default deterministic or evidence-led manual workflow only after clearing all three predeclared gates.
| Audit | Control or manipulation | Record | Valid inference | Decision consequence |
|---|---|---|---|---|
| CIELAB identifiability | Hold recorded L* fixed; vary chroma | a*, b*, and resulting RGB triples | One luminance permits multiple colors | Treat chroma as inferred, not recovered |
| DDPM mechanism | Ho et al., 2020; forward process over T steps | U-Net noise prediction εθ | Reverse denoising constructs a new sample | More steps cannot create missing provenance |
| Baseline contrast | Compare an L2 mean with a diffusion draw | Conditional color and coherence | Realism and pixel fidelity may diverge | Evaluate both against known-color controls |
| Conditioning audit | Hold grayscale fixed; alter text or exemplar in Stable Diffusion 2.1 | Changed chroma prior | Conditioning steers generation | Do not count conditioning as pigment evidence |
| Seed audit and eligibility | Hold all conditions fixed; vary the seed across five runs | Object, background, and high-salience color changes | Variation exposes model uncertainty | Deploy only if all three gates clear; otherwise retain deterministic or evidence-led manual colorization |

Evidence Base
MIT-Adobe FiveK can falsify a color-recovery claim; it cannot authenticate a historical hue. I derive matched grayscale inputs directly from its color references, making the benchmark a controlled colorization test rather than an assertion that the corpus is authentic historical ground truth. The distinction matters because CIEDE2000 can show that a method moves an image toward a specified reference, yet cannot show that the color was appropriate for the photographed object, material, and period.
Richardson et al.’s Colorful Image Colorization is my named non-diffusion comparator, not my historical authority. Its training exposure establishes familiarity with a broad modern color distribution; it does not demonstrate competence on period-specific dyes, pigments, coatings, or textiles. A strong result against this baseline answers whether diffusion improved controlled recovery. It does not establish that every recovered object color is historically defensible. That distinction kills the status-quo myth that training-set size is a proxy for archival knowledge.
Zhang et al.’s BAPPS-derived LPIPS evidence is also secondary. I use it to assess perceptual distance alongside the controlled color metric, never as a hue certificate. Human preference and LPIPS can reward a coherent image even when an object-level color is unsupported; the result belongs in sensitivity analysis, not the dated-evidence count. A historical claim advances only when perceptual evidence and an external color source agree.
The image-level evidence ledger gives every claim an object or region, source date, claimed color, and confidence level. A text prompt, aesthetic vote, or model output is entered as a hypothesis—with its model, run, and competing candidates—not as a color source. For example, a red military tunic remains unverified until a dated object, dye, paint, or comparable material record supports that hue in the relevant context. If sources conflict, I preserve the conflict and reduce confidence rather than averaging contradictory evidence into false certainty.
I split evaluation by photographer, archive, date, and physical medium before allocating scans. Otherwise, multiple crops or exposures of one negative can cross evaluation boundaries, inflating apparent accuracy through near-duplicate leakage. Every derivative remains in one group; seed reruns are nested within it and checked for object and background identity. Results are reported by group as well as in aggregate because a gain concentrated in one collection cannot establish archive-wide competence.
| Named source | Verified scale | Admissible role | Decision limit |
|---|---|---|---|
| Grosse et al., MIT-Adobe FiveK | Color-image set; five illumination conditions | Matched grayscale inputs and controlled CIEDE2000 evaluation | Reference colors are not automatically authentic historical evidence |
| Richardson et al., Colorful Image Colorization | Trained on 1.6 million images | Named non-diffusion baseline | Modern color exposure does not establish period-material competence |
| Zhang et al., BAPPS-derived LPIPS evidence | BAPPS-derived image pairs and human ratings | Secondary perceptual comparison | Perceptual plausibility cannot supply a historical hue or source date |
Before outputs are inspected, the grouped image manifest, baseline checkpoint, ledger fields, and seed protocol must be frozen. Diffusion advances beyond the default only when the resulting record clears every predeclared accuracy, blinded-historian, dated-evidence, and stability gate. If any required gate fails, the decision remains deterministic or evidence-led manual colorization.

The Three-Test Scorecard
Diffusion does not win because one output looks convincing. It wins only when color fidelity, blinded historical preference, and provenance with structural stability all clear the predeclared gates. The scorecard is conjunctive: one failure returns the decision to deterministic/manual, while “not measured” is never a pass.
Test 1—Known-color fidelity. Evaluate matched color-to-grayscale pairs. Freeze the pair list and diagnostic masks before inference, then report median CIEDE2000 over the full frame and separately for each object mask. Diffusion passes only when its median error clears the predeclared improvement criterion against the deterministic baseline and the paired-bootstrap confidence interval for that improvement lies entirely above zero. Each bootstrap replicate must resample matched pairs and recompute both medians and their relative difference; resampling pixels would mistake spatial autocorrelation for independent evidence. According to Diffusion vs. GAN Colorization: Why PSNR Fails and What to Test, DeOldify has higher PSNR but worse CIEDE2000 than DDColor. That source supplies only the direction of the differences—not CIEDE2000 values or a photograph count—so it cannot populate the scorecard by itself.
Test 2—Historical plausibility. Give two independent historians randomized A/B comparisons apiece, providing period and material context but no model labels. Diffusion passes only when the predeclared confidence-interval lower bound for preference exceeds 50%. Account for repeated judgments of the same source image when calculating the interval; otherwise, one visually distinctive photograph can appear more persuasive than the evaluation warrants. Preference establishes comparative plausibility, not authentication of a particular hue.
Test 3—Provenance and stability. Define a high-salience region as a historically diagnostic garment, vehicle, painted object, or signage area. Register the region masks before rendering and assess dated-evidence coverage for those colors against the predeclared standard. Run five fixed-condition seeds with the model settings, prompts, resolution, and seed list frozen. Compare object and background geometry and identity against the source and across outputs; any object or background change is an automatic failure, not a color variation to average away. Report coverage as documented high-salience colors divided by all high-salience colors, with the evidence ledger attached.
The supplied evidence contains no qualifying side-by-side measurements, so every unavailable result below remains explicitly not measured rather than being encoded as zero.
| Method | Median CIEDE2000 | Historian preference with confidence interval | Dated-evidence coverage | Structural failures across seeds | Overall Winner |
|---|---|---|---|---|---|
| Diffusion | Not measured | Not measured | Not measured | Not measured | Deterministic/manual until a three-of-three pass |
| Deterministic baseline | Not measured | Not measured | Not measured | Not measured | Deterministic/manual |
| Evidence-led manual colorization | Not measured | Not measured | Not measured | Not measured | Deterministic/manual if available; otherwise retain the source scan in monochrome |
Populate the table from the preregistered evaluation rather than a selected example. If diffusion has even one failed or missing gate, choose deterministic/manual. If evidence-led manual colorization is unavailable, retain the source scan in monochrome rather than assigning an undocumented winner.

What FID Does Not Tell You About
FID can improve while every historically decisive hue is wrong. According to Heusel et al.’s paper, Fréchet Inception Distance compares image distributions through Inception-v3 pool features. A lower value can reflect better hue-frequency matching even when a particular uniform or vehicle receives the wrong color. I therefore treat FID as evidence about corpus-level appearance, not per-photo proof; known-color controls and region-level inspection must adjudicate the actual hue.
CLIP does not close that gap. According to Radford et al.’s 2021 paper, CLIP was trained on a large corpus of internet image-text pairs, so its semantic prior can encode modern dataset bias. A prompt may produce a convincing period suit while remaining confidently wrong about that suit’s actual fabric color. Semantic plausibility is therefore not material or archival evidence.
Modern RGB-to-grayscale controls have a narrower claim. They test recovery from a synthetic corruption, whereas historical prints can also contain fading, silvering, chemical staining, hand tinting, and scanner bias. Passing such controls proves invertibility under the simulation, not knowledge of the color before those degradations. A model can reverse an imposed loss of chroma without identifying the historically original hue.
I test aggregate-metric counter-evidence by isolating the images with the worst CIEDE2000 errors and inspecting diagnostic regions. An acceptable overall result accompanied by systematically wrong uniforms, vehicle bodies, or signage is not demonstrated historical accuracy; the errors are concentrated precisely where a plausible image can become historically consequential.
I also treat model and rater variance as evidence about reliability. A single polished seed or consensus-looking montage cannot establish stability if chroma changes materially across five seeds or independent raters disagree about period consistency. Seed-by-seed chroma comparisons and region-specific rater review expose uncertainty that polished presentation conceals. These checks remain supporting diagnostics: if the error tail, synthetic-control interpretation, or stability record fails, the diffusion result is not substantiated, and the deterministic or evidence-led manual workflow wins.
| Check | What it can establish | Disqualifying inference or pattern |
|---|---|---|
| FID | Closer image distributions in Inception feature space | Treating a lower score as proof that particular uniforms or vehicles have correct colors |
| CLIP prompt | Semantic compatibility between a prompt and rendered scene | Using convincing period semantics as evidence of actual fabric color |
| RGB-to-grayscale control | Recovery from the specified synthetic corruption | Equating inversion with knowledge of pre-fading historical color |
| CIEDE2000 tail audit | Whether aggregate performance conceals concentrated diagnostic errors | Accepting the overall result while high-salience regions remain systematically wrong |
| Seed and rater review | Chromatic stability and consistency of period judgments | Relying on one polished seed or masking material disagreement among independent raters |

Worked Palette Case
Palette’s “14.20” is a benchmark datum, not an archive pass. According to Deng et al.’s ImageNet report, the ImageNet validation pool supplies the source population. Those source-pool details do not automatically establish the experiment’s sample count. The defensible audit entry is narrower: “ImageNet validation images used in Saharia et al.’s colorization experiment; sampled count not stated in the reported result.” Neither the source-pool count nor an invented smaller count should be substituted. Reproducing the experiment would require the selected image IDs or a released evaluation manifest.
Under the evaluation settings stated in Saharia et al.’s 2022 Palette paper—an ImageNet colorization experiment using modern images—the reported result is FID 14.20. Those qualifiers are part of the evidence. It is a distribution-level comparison between generated color images and an image reference distribution, not a paired reconstruction test on archival photographs and not a test of recovered historical color. Treating 14.20 as archival accuracy would silently replace both the evaluation population and the target of inference.
Test 1 is therefore not demonstrated. The result supplies no median CIEDE2000 computed against independently verified pre-colorization references for known-color controls. It consequently does not establish sample-level color improvement. A lower FID relative to a baseline also cannot be translated into a percentage reduction in median CIEDE2000: the two statistics have different denominators, pairings, and error models, so converting one into the other would fabricate a color-fidelity gain.
Tests 2 and 3 are likewise not demonstrated. Saharia et al.’s result establishes no period-qualified, blinded historian preference; no dated evidence coverage for high-salience colors; and no comparison of object and background structure across five seeds. Historical plausibility, provenance, and structural stability therefore remain unproved. These are evidential failures, not proof that Palette always produces implausible colors or changes objects.
| Published Palette result | Test 1 | Test 2 | Test 3 | Gates cleared | Policy winner |
|---|---|---|---|---|---|
| 14.20 FID | Not demonstrated | Not demonstrated | Not demonstrated | 0 of 3 | Deterministic/manual |
The all-gates rule makes the result decisive. This scorecard rejects the evidence claim, not Palette’s underlying image quality. To reopen the decision for an archival candidate, the record must add paired known-color references, a blinded period-qualified preference study, dated support for historically consequential colors, and fixed-seed checks of object and background identity. Until those records exist, deterministic or evidence-led manual colorization remains the policy winner.

How to Choose Well
Hist10K supplies the decisive warning: diffusion can win a realism contest without proving historical accuracy. According to Hist10K, the $200–$800 versus $0.50–$3 gap reflects labor, not greater historical truth. That kills the cheap-palette myth. A low-cost, polished result does not recover an authenticated hue, so the explicit archival default remains Deterministic/manual unless diffusion clears every predeclared gate.
Begin with measurement, not appearance. If the archive has no known-color reference set with which to test either automatic method, choose monochrome. Without matched controls, median CIEDE2000 cannot adjudicate a color claim; a plausible hue is not evidence that the original hue was recovered. This is a measurement stop, not a preference for colorless photographs.
When controls exist, compare the methods under identical inputs and scoring procedures. If diffusion fails to beat the deterministic baseline on median CIEDE2000 under those matched controls, choose the deterministic result. The comparison criterion is already fixed in “The Three-Test Scorecard”; changing it after seeing outputs would turn a falsifiable standard into a taste contest. The supplied “Diffusion vs. GAN Colorization” source names PSNR and ΔE00 but provides no numerical values for either, so it cannot rescue a failed archival comparison.
Fidelity is necessary, not dispositive. If diffusion fails the blinded historian-preference test, choose evidence-led manual colorization rather than substituting visual appeal for historical support. A realism advantage is compatible with error in precisely the colors an archive treats as salient. Hist10K’s recommendation to use automatic diffusion “only for internal drafts” reinforces that boundary; it does not license an unsupported catalog record.
Provenance can still veto an otherwise strong result. If dated-evidence coverage for high-salience colors fails, or if any required repeated-seed run changes an object or background, choose the deterministic or manual workflow. Mark unsupported regions visibly uncertain instead of presenting generated color as fact. Seed variation exposes dependence on model sampling rather than dated evidence; an attractive run cannot establish reproducibility.
Only a joint pass changes the decision. If fidelity, blinded preference, and provenance with structural stability all clear the scorecard, diffusion is the winner. Every other outcome keeps Deterministic/manual explicit. The rules below point to the scorecard’s existing decision conditions rather than duplicating them, preserving one decision standard.
| Rule | Condition | Decision |
|---|---|---|
| 1 | No known-color reference set exists for testing either automatic method. | Choose monochrome; do not label either output measurably accurate. |
| 2 | Matched controls exist, but diffusion misses the scorecard’s median CIEDE2000 improvement cutoff. | Choose the deterministic result. |
| 3 | Fidelity passes, but blinded historian preference or its required confidence condition fails. | Choose evidence-led manual colorization. |
| 4 | Fidelity and preference pass, but dated evidence or object/background stability across the required seeds fails. | Choose deterministic or manual output and visibly mark unsupported regions as uncertain. |
| 5 | Fidelity, historian preference, dated provenance, and required seed-stability checks all pass. | Choose diffusion; any failed check keeps Deterministic/manual as the explicit decision. |
What to do next
| Step | Action | Why it matters | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Before viewing model outputs, inventory each high-salience object and predefine documentary evidence requirements: dated textiles, uniforms, captions, or shop records from the photograph’s place and period. | This prevents plausible chroma from be
Frequently Asked QuestionsHow do diffusion colorization and archival hand coloring compare in quoted cost? Hist10K quotes $0.50–$3 per diffusion-colorized photo, while archival hand coloring costs $200–$800 per complex print and still warrants documentary cross-checking. Can a realistic image from the quoted $7.99 consumer option establish historical accuracy? No; it can produce a plausible image while period evidence remains insufficient to rule out object-level anachronisms. What can Saharia et al.’s 14.20 FID for ImageNet colorization actually measure? It measures perceptual similarity across images, not whether a particular garment in a family photograph was historically red. What color information is identifiable from a grayscale scan in CIELAB? A grayscale scan constrains L* but not a* or b*, so one luminance can permit many RGB triples. How should a five-seed audit be conducted, and what do its results prove? Hold the scan, prompt or exemplar, model settings, and other conditions fixed while varying the random seed across five runs; coherent chroma variation exposes model uncertainty, while perfect agreement would prove only repeatability rather than provenance. What must the image-level evidence ledger record for each historical color claim? It must record the object or region, source date, claimed color, and confidence level, while prompts, aesthetic votes, and model outputs are entered only as hypotheses with their model, run, and competing candidates. Quick answers
Also worth reading: Diffusion Colorization: Conditioning Tradeoffs Beyond PSNR/SSIM: Diffusion Colorization: Conditioning Tradeoffs Beyond · How machine learning brings historical black and white photos back to life: How machine learning brings historical · How to transform your old black and white photos into vibrant memories with professional AI colorization: How to transform your old Research Methodology & Editorial StandardsWe begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place. Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted. Published · Last reviewed · Owned by the Colorizethis editorial desk (About, Contact, Privacy). Related readingLatestRelated answers |