| Takeaway | Detail |
|---|---|
| CIE Lab wins because it is more authentic, not more colorful. | The benchmark's ΔE2000 result favored CIE Lab by 10%, while RGB's extra chroma clustered in silver-mirroring damage. |
| The scoring formula rewards fidelity over saturation. | Composite score uses 15% CLIPScore, 15% aesthetic prediction, 20% ImageReward, 20% HPS, and 10% from X-IQE terms. |
| Text hallucination is the clearest failure sign for RGB. | The CIE Lab output had a 4% text hallucination rate, and human vision's lower color acuity lets RGB's false chroma hide damage. |
| The 'more colorful' RGB appearance is a trap. | RGB outputs add chroma to silver-mirroring damage; the 2026 benchmark's 20% reward-model weight deprioritizes that cosmetic gloss. |
At 10% lower ΔE2000 error, the 2026 benchmark hands the albumen-print restoration win to CIE Lab diffusion—even though the RGB output looks more colorful. The counterintuitive result comes from where that color lives: almost all of RGB's extra chroma lands in silver-mirroring damage, turning a defect into a false hue. The scoring rubric gives 15% of its weight to CLIPScore and 15% to aesthetic prediction, with 20% each to ImageReward and HPS, so pigment-faithful reconstruction outranks saturated pixels.
The 4% text hallucination rate makes the difference concrete. RGB models often reconstruct faded print edges as letters or texture; CIE Lab suppresses that hallucination because it separates lightness from chroma. Human vision has lower acuity for color than for luminance, which is why the over-saturated RGB result can fool an eye but not the benchmark's pixel-aligned tiles.
The benchmark's 20% weight on reward-model agreement further separates authentic restoration from cosmetic color. A model that adds chroma to specular silver reflections may score high on aesthetic appeal but fails the fidelity-weighted composite. By making the less colorful output the more authentic one, the 2026 result recalibrates what 'better' means for albumen prints.
The Mechanism
Albumen's average measured b* is +18.4 in the benchmark's unbleached control set. That single coordinate explains the entire 2026 result. The headline gap is not a tuning artifact or a lucky seed; it is the inevitable consequence of what each pipeline is asked to denoise.
CIE Lab forces a diffusion model to treat lightness and chrominance as separate prediction targets. Albumen prints have a tightly clustered background in a*/b* space — a* ≈ −1.2 to +3.5, b* ≈ +10 to +24 under D50 — because the egg-white binder's warm cast is spatially uniform. A Lab-conditioned model re-adds color to the faded print by predicting only the residual chrominance shift, without re-lighting the image. The L* channel carries the paper texture, the shadow density, the silver grain; the model never touches it. In RGB, lightness and color are entangled in every channel, so any color correction necessarily recomputes the lighting. That is the first structural divergence.
The benchmark's colorimetric calibration is the second. It uses D50 reference white, defined as x = 0.3457, y = 0.3585. RGB diffusion pipelines default to D65 illuminant assumptions. On albumen's egg-white binder, that mismatch produces a systematic blue-to-amber shift: the model compensates for a white point the print was never developed under, and the correction lands as a color cast that reads as "restored" but is a fabrication. The D50 anchor is not a formality; it is the difference between measuring the print as it was and measuring it as a display would like it to be.
Now the objective function. RGB conditioning treats that b* +18.4 as an error to correct; its loss wants neutrals. Lab conditioning treats the same +18.4 as a coordinate to preserve. That changes the restoration objective from "make it neutral" to "keep the historical chroma." Saturation without historical fidelity is not restoration; it is a recoloring artifact — which is precisely why the myth that RGB's brighter output is superior collapses. According to Golyadkin, Plevokas, and Makarov's WACV 2025 paper "Closing the Domain Gap in Manga Colorization via Aligned Paired Dataset," models trained on real black-and-white/color pairs significantly outperform models trained on synthetic pairs. The analogue holds: a model that treats the albumen's real warm cast as signal, not noise, beats one that corrects it toward an invented neutral.
The red-channel clip edge case shows the failure mode concretely. RGB diffusion learns highly correlated R/G/B channels. When metallic silver-mirroring clips red to 254, the correlated model "fixes" the clip by inventing red-green gradients across the artifact. Lab's a* axis separates that artifact from L*, so the paper grain underneath stays intact and the false gradient never enters the output. A human restorer does the same thing: identify the mirroring, leave the grain.
The decisive step is conditional denoising. In Lab, the model predicts only residual a*/b* chrominance; the original L* is copied back after every DDIM step. In RGB, all three channels are denoised together, so every step re-bakes the degradation into the output. The faded binder, the mirroring, the cast all live across all three RGB channels. You cannot denoise them away without denoising the image itself.
| Mechanism step | CIE Lab pipeline | RGB pipeline | Winner |
|---|---|---|---|
| Lightness handling | Original L* copied back each DDIM step | L* denoised jointly with color; degradation re-baked | Lab |
| Chroma objective | Preserve b* +18.4 as historical coordinate | Correct warm cast toward neutral | Lab |
| Reference white | D50 (x = 0.3457, y = 0.3585) | D65 default → blue-to-amber shift on binder | Lab |
| Silver-mirroring clip | a* isolates artifact from L*; grain intact | Red clip to 254 → false red-green gradients | Lab |
| Restoration status | Restorative output | Display proof only, explicitly non-restorative | Lab |
The practical move for 2026 pipelines: route conditional inputs through Lab, hold the reference white at D50, and copy a frozen L* back after each DDIM step. That is not a preference; it is the mechanism by which the benchmark's gap was earned.
Decision Framework: Lab Wins the 2026 Table
The 2026 benchmark does not merely prefer CIE Lab for albumen prints—it makes the choice unambiguous across every decision row that matters for archival work. Per-tile processing at the 50-step schedule runs 41 seconds in Lab versus 44 seconds in RGB according to the benchmark's reference implementation, a speed gap that disappears into noise when amortized over a full plate. The 40-seed worst-case ΔE2000 lands at 7.4 for Lab and 10.3 for RGB, meaning the worst seed in Lab is better than the RGB average on the same metric. Metadata lineage retention—the benchmark's measurement of how faithfully the model preserves the original plate's capture information through the colorization pass—is 97% for Lab against 61% for RGB. Reference-white visual-matching success, where the model must reproduce a known neutral patch under the albumen's yellowed baseline, reaches 88.2% for Lab versus 74.1% for RGB. The explicit winner column in the benchmark's decision matrix is Lab in all four rows.
| Decision Row | CIE Lab | RGB | Winner |
|---|---|---|---|
| Per-tile processing, 50-step schedule | 41s | 44s | Lab |
| 40-seed worst-case ΔE2000 | 7.4 | 10.3 | Lab |
| Metadata lineage retention | 97% | 61% | Lab |
| Reference-white visual-matching success | 88.2% | 74.1% | Lab |
The benchmark's decision matrix weights historical fidelity at 0.95 and raw C* saturation at 0.05, a weighting that reflects the target condition of albumen restoration: fidelity to the original capture, not vividness. On that weighted score, Lab beats RGB 3.71 points to 2.63, a 1.41× decision margin. RGB wins only on the raw C* saturation axis, which the benchmark weights at 0.05 because saturation without historical fidelity is a recoloring artifact, not restoration. The myth that RGB must be better because it produces brighter, more vivid colors collapses the moment you read the matrix: the saturation axis is deliberately near-zero-weighted, and the benchmark's own data shows that RGB's vividness gain corresponds directly to its ΔE2000 loss. Brighter is not better; brighter is wrong.
The production implication follows directly from the table. Because the gap is speed-neutral—the extra 3 seconds per tile on RGB comes from the additional channel decorrelation work, not from model capacity—museums can adopt Lab without re-engineering GPU budgets. According to the benchmark's reference scripts, the entire pipeline runs unmodified on a single 24GB GPU. Seed stability is equally important: the 40-seed worst-case ΔE2000 gap means a museum can batch-process an entire albumen collection without auditing individual outputs for catastrophic color shifts. The 24GB GPU constraint is the practical ceiling; every major restoration lab currently operating on RTX 4090-class or A5000-class hardware is already inside it.
The matrix also includes a "Restoration Intent" field that should be set explicitly for every project. Archival restoration projects are assigned weight 0.95 to historical fidelity, and the full 0.95/0.05 split is the correct default. Display-only projects—exhibitions, public-facing digital galleries, or commercial prints where the goal is aesthetic appeal rather than archival reconstruction—can lower the fidelity weight to 0.50. Even at 0.50, Lab remains the winner on the benchmark's weighted score, because the RGB saturation advantage is too narrow to overcome Lab's dominance on the other three axes. The decision rule is therefore not "Lab for archives, RGB for display." The rule is "Lab for everything, with the weight adjusted by intent." RGB-conditioned diffusion should be used only for saturation-adjusted display proofs that are explicitly labeled non-restorative, as the canonical decision rule states.
| Decision Point | Condition | Action | Rationale |
|---|---|---|---|
| Is the project archival restoration? | Restoration Intent weight = 0.95 | Use CIE Lab unconditionally | Lab wins 3.71 vs 2.63 on weighted score |
| Is the project display-only? | Restoration Intent weight → 0.50 | Use CIE Lab; adjust output for display | Lab still wins at 0.50 weight |
| Is the GPU budget constrained? | Single 24GB GPU available | Run Lab reference scripts unmodified | Lab is speed-neutral at 41s vs 44s |
| Is batch-processing reliability critical? | Large collection, no per-image audit | Use Lab for seed-stable worst-case | Lab 40-seed worst-case ΔE2000 = 7.4 |
| Is vivid color the stated goal? | Requested by stakeholder for display | Use RGB only as non-restorative proof | RGB wins C* only; weight 0.05 |
The decision tree above is the entire framework. Each branch is anchored to a specific number from the benchmark, not a general preference. Archival restoration with a 0.95 fidelity weight means Lab, unconditionally. Display-only with a 0.50 weight still means Lab, because the weighted score gap persists. A single 24GB GPU runs the reference scripts without modification, which removes the hardware objection. Batch processing at scale requires seed-stable worst-case behavior, which favors Lab by 2.9 ΔE2000 units. And if an outside stakeholder insists on vibrancy, the RGB output is produced as a labeled proof, not as a restoration deliverable. The Lab default is not a compromise; it is the only choice that satisfies the benchmark's target condition.
What the Data Doesn't Tell You
The 14.2% headline gap is real, but conditional: it was measured on one denoising schedule, one chroma cohort, and one whole-image metric. Stratify the benchmark by print condition, by image region, or by sampling step, and the margin shrinks — in one setting it reverses. None of those edge cases overturns the decision rule, but they define exactly when an archival restorer should verify Lab's output rather than trust it.
On lightly faded albumen prints, the official benchmark's stratified results show RGB's text-hallucination rate dropping to 4.1% and Lab's ΔE2000 advantage shrinking to roughly 2.5%. The headline is therefore not representative of every 1850–1900 print: mild fade compression brings the two pipelines near parity, and the hero numbers are driven by the moderately to heavily faded tail of the corpus.
An independent audit flagged a "Lab-halo" artifact that the whole-image score cannot surface: 23% of Lab outputs pulled preserved a*/b* pairs toward the training prior whenever the original chroma sat outside the benchmark's common cluster, producing a green or amber ring around silver edges. The mechanism is prior pull — Lab separates luminance from chroma, so the denoiser can quietly remap out-of-distribution chroma toward its centroid at the exact discontinuity where a silver edge meets emulsion.
The official whole-image ΔE2000 metric also hides museum-critical regions. On 100 manually segmented faces, RGB achieved 5.9 ΔE2000 versus Lab's 6.4, though faces were only 9% of pixels. For a portrait-heavy albumen collection, whole-image scores overstate Lab's lead precisely where curators look first: skin tones.
At 8-step SDXL-Turbo with CFG scale 4.0, the ranking reverses outright — RGB reaches 10.5 ΔE2000 versus Lab's 11.0 — meaning Lab's advantage is partly tied to the 50-step denoising schedule and larger CFG scale the archival workflow prescribes. Read that reversal carefully: it is not an endorsement of RGB. The brighter, more vivid output RGB produces at draft speed is exactly the saturation-without-historical-fidelity the benchmark classes as a recoloring artifact, not a restoration.
The expert panel showed no significant agreement on chemically faded prints from before 1885 (Fleiss κ = 0.14), so "historically plausible" is not a stable human target for the most degraded stratum of albumen. The panel disagreed about what the faded print originally looked like; the usable ground truth for those images is a chemical fade model, not human consensus.
Finally, the physical object itself confounds dataset-level comparison. Original albumen emulsion thickness varies from 15 to 45 µm, which can shift measured CIE b* by up to 3.0 units under the benchmark's fixed reference white. That heterogeneity overlaps the Lab-vs-RGB effect size, so a single print at the thin end of the range can read as a Lab failure when it is actually an emulsion-thickness artifact — a confound no dataset-level benchmark captures.
The table below consolidates the conditions under which the rule needs active verification. Lab remains the default in every row; what changes is how carefully the restorer checks the output before archiving it.
| Edge case | Measured effect | What to verify |
|---|---|---|
| Lightly faded prints (1850–1900) | RGB hallucination 4.1%; Lab edge ≈2.5% ΔE2000 | Confirm fade severity; near-parity batches can tolerate RGB proofing only |
| Lab-halo artifact | 23% of Lab outputs shift chroma at silver edges | Check a*/b* pairs near edges against the original chroma cluster |
| Segmented faces (100 faces, 9% of pixels) | RGB 5.9 vs Lab 6.4 ΔE2000 | Run region-weighted evaluation; faces may need chroma correction after Lab pass |
| 8-step SDXL-Turbo, CFG 4.0 | RGB 10.5 vs Lab 11.0 ΔE2000 | Never archive fast-draft output; the 50-step schedule is mandatory |
| Pre-1885 chemically faded prints | Fleiss κ = 0.14 panel agreement | Build chemical-fade priors; do not rely on "historically plausible" ratings |
| Emulsion thickness 15–45 µm | CIE b* shift up to 3.0 units | Measure thickness; correct b* before comparing against any benchmark |
The decision rule survives all six edge cases — but as a conditional default, not a universal. Start every albumen print in Lab at the 50-step schedule, then verify against the halo, the skin regions, the fade stratum, and the emulsion thickness before accepting the output. The rule holds; the verification list is what the data doesn't tell you.
Restoring 'Oficina de Bellas Artes,' 1896
Restoring 1896-08-27_Palacio_de_Bellas_Artes_0437.tif from the benchmark's verification set is the decisive single-case test of the Lab decision rule. The albumen print is 61×86 cm, scanned at 12,388 × 17,752 pixels, with silver mirroring along 22% of the border and foxing over 14% of the sky. Its value as a test case is that the damage spans two regimes—specular metal deposition along the border and emulsion oxidation across the sky—so a single restoration pass has to satisfy both. Those are the regions where RGB's failures stop being statistical noise and become visible artifacts.
The restoration used the winning Lab pipeline: 144 non-overlapping 1024×1024 tiles with a 32-pixel overlap, a 50-step denoising schedule, CFG scale 3.0, on 4×A100 GPUs, for a total inference time of 1.68 hours. The 32-pixel overlap suppresses seam artifacts only; it does not alter the chroma decision inside a tile. According to the benchmark report, the final composite's global ΔE2000 came to 6.3—0.5 below the benchmark average—with 92.4% of pixels within 3.0 ΔE2000 of the reference, measured under the benchmark reference white. That aggregate is the floor set by the two edge regions, not the ceiling.
The silver-mirrored border tiles separate the conditionings cleanly. Lab averaged ΔE2000 9.2 on those tiles; RGB averaged 13.7—a 33% gap that reproduces the headline advantage above at the regional level. Typography is the sharper discriminator. RGB hallucinated a different typographic style in 3 of the 144 tiles, inventing decorative letterforms where the original caption sat, while Lab reproduced "Oficina de Bellas Artes" legibly in every tile with OCR confidence 0.94 versus 0.61. For a pipeline that feeds catalog records, that OCR spread is the operational metric: 0.94 survives machine reading, 0.61 sends a human back to re-key the caption.
The sky tests hue memory rather than texture. The reference blue-ish hue sits at a* = −4.6, b* = 2.1, outside the benchmark's common b* cluster. Lab still recovered the sky to ΔE2000 4.7. RGB shifted b* to +11.2—the warm-cast failure named in the benchmark report. That is the status-quo myth in measurable form: b* +11.2 is more saturated than the reference, but saturation without historical fidelity is a recoloring artifact, not restoration.
| Verification-set region | Lab pipeline | RGB pipeline | Decision |
|---|---|---|---|
| Silver-mirrored border, ΔE2000 | 9.2 | 13.7 | Lab (33% lower) |
| Caption hallucination, tiles affected | 0 of 144 | 3 of 144 | Lab |
| Caption OCR confidence | 0.94 | 0.61 | Lab |
| Sky b* coordinate | 2.1 recovered, ΔE2000 4.7 | +11.2 (warm cast) | Lab |
The workflow change is to score restorations per region, not per composite. Split every albumen print into silver-mirrored borders, foxed sky, and clean interior tiles, and measure each region separately under the benchmark reference white. The verification set shows the border and sky numbers are where Lab's default status is earned; the global average is the last number an archivist should quote. Keep RGB for saturation-adjusted display proofs only, and label those proofs explicitly non-restorative.
How to Choose Well
A clean-looking scan is not the same as a clean scan. The 2026 albumen benchmark gives RGB exactly one legitimate doorway back into an archival workflow: a visibly undegraded print with fewer than 5% of scan tiles showing mirroring or foxing, and even then only for an exhibition proof that is explicitly labeled non-restorative. Every other default starts with CIE Lab-conditioned diffusion, because Lab separates lightness from chroma in a way that keeps white-balancing from re-entangling the a*/b* axes.
The default rule is simple: if the albumen print has any silver mirroring, foxing, or yellowed binder — the condition of 97.4% of the benchmark’s corpus — choose a CIE Lab-conditioned diffusion model. RGB is not banned in this condition, but it is quarantined to a separately labeled display proof. That proof may be useful for exhibition, but it is not a restoration and cannot be filed in the archival set. Brighter RGB color is not historical fidelity; it is a recoloring artifact.
The clean-scan rule is the only route to an RGB proof as the primary output. When fewer than 5% of scan tiles show mirroring or foxing and the print is visibly undegraded, RGB may be considered for an exhibition color proof. That proof must be disclosed as “saturation-adjusted, not a restoration.” If the disclosure feels awkward, that is the point: what RGB gives you in saturation it takes away in archival defensibility.
The silver-mirroring rule makes the default per-tile. If more than 20% of a tile is covered by silver mirroring, reject the RGB output for that tile automatically. Use Lab for the affected tile, then composite it back into the full image with a 32-pixel feathered seam. Do not re-run the whole image; the decision is tile-local, and the seam feather is what prevents a visible boundary between the Lab-restored region and the rest of the scan.
The faces-first rule is the one place RGB can legitimately win pixels after a Lab render. When a human face occupies more than 12% of image height, generate both Lab and RGB outputs. Keep the RGB face region only where its local ΔE2000 is lower than the Lab output. Then require a seam check at 200% zoom before accepting the composite. That zoom check is mandatory, because a lower local ΔE2000 is meaningless if the transition into the surrounding restoration is visible at inspection scale.
The reporting rule closes every restoration. Always publish model checkpoint, seed, CFG scale, sampling steps, and color space, and evaluate with the benchmark’s official CIEDE2000 script. Never white-balance a Lab output before evaluation, because white-balancing re-entangles the a*/b* axes. Once you white-balance, the color channels are no longer independent, and your ΔE2000 value no longer compares against the benchmark.
| Rule | Trigger | Action | RGB allowed? |
|---|---|---|---|
| Default | Any silver mirroring, foxing, or yellowed binder (97.4% of corpus) | CIE Lab-conditioned diffusion | Separately labeled display proof only |
| Clean-scan | Fewer than 5% of scan tiles affected; visibly undegraded | Lab for archival; optional RGB exhibition proof | Proof only, disclosed as “saturation-adjusted, not a restoration” |
| Silver-mirroring | More than 20% of a tile covered by mirroring | Reject RGB for that tile; Lab restore; 32-pixel feathered seam | No for the affected tile |
| Faces-first | Face occupies more than 12% of image height | Generate Lab and RGB; keep RGB face only where local ΔE2000 is lower | Conditional on local ΔE2000; seam check at 200% zoom |
| Reporting | Every submission | Publish checkpoint, seed, CFG scale, steps, color space; use official CIEDE2000 script | N/A — no white-balancing before evaluation |
The decision tree is therefore: check condition first. Any degradation means Lab. A visibly undegraded scan with fewer than 5% affected tiles can also produce an RGB exhibition proof, but the archival file remains Lab. Any tile with more than 20% silver mirroring forces Lab for that tile. Any face above 12% of image height forces a dual render with a measured RGB-vs-Lab decision and a 200% seam check. Then report the full configuration. RGB never becomes the archival default; it is either a labeled display artifact or a locally measured exception.
What to do next
| Step | Action | Why it matters |
|---|---|---|
| 1 | Route the faded albumen scan through the CIE Lab-conditioned diffusion pipeline and lock the L* channel so paper texture, shadow density, and silver grain are never recomputed. | RGB entangles lightness and chroma in every channel, so its color correction recomputes lighting; Lab predicts only the residual chrominance shift. |
| 2 | Set the a*/b* restoration target to the unbleached control set's b* centroid, judged under D50. | Albumen's warm egg-white cast is spatially uniform; matching that centroid restores the original tone without re-lighting the print. |
| 3 | Score the output against the weighted composite: 15% CLIPScore, 15% aesthetic prediction, 20% ImageReward, 20% HPS, and 10% X-IQE. | The rubric rewards pigment-faithful reconstruction over saturated pixels; the 20% reward-model weight deprioritizes cosmetic gloss on silver-mirroring damage. |
| 4 | Verify the CIE Lab output beats the RGB pipeline by at least 10% ΔE2000 before archiving. | That margin is the 2026 benchmark's headline result — and RGB's extra chroma lands almost entirely in mirroring damage, turning a defect into a false hue. |
| 5 | Inspect faded print edges for hallucinated letters or texture and keep the text hallucination rate at or below 4%. | Lab suppresses edge hallucination by separating lightness from chrominance; RGB models often reconstruct faded edges as letters or texture. |
| 6 | For the display proof only, export an RGB render, label it "non-restorative," and use it solely for saturation-adjusted screen viewing. | Human vision's lower color acuity lets RGB's false chroma hide damage — a proof can support display but can never be the archival restoration. |
Frequently Asked Questions
How many seconds per tile does the Lab pipeline take at the 50-step schedule compared with RGB?
At the 50-step schedule, the Lab pipeline takes 41 seconds per tile versus 44 seconds for RGB.
What are the 40-seed worst-case ΔE2000 scores for Lab and RGB?
The 40-seed worst-case ΔE2000 is 7.4 for Lab and 10.3 for RGB.
What is the metadata lineage retention for Lab versus RGB?
Metadata lineage retention is 97% for Lab against 61% for RGB.
What are the reference-white visual-matching success rates?
Reference-white visual-matching success is 88.2% for Lab versus 74.1% for RGB.
What reference white does the benchmark use, and what default do RGB pipelines assume?
The benchmark uses D50 reference white (x = 0.3457, y = 0.3585), while RGB diffusion pipelines default to D65, causing a systematic blue-to-amber shift on the binder.
What happens when metallic silver-mirroring clips red to 254 in the RGB pipeline?
When metallic silver-mirroring clips red to 254, RGB's correlated model invents red-green gradients across the artifact, while Lab's a* axis separates the artifact from L* and leaves the grain intact.
Quick answers
| Why did CIE Lab win the 2026 benchmark even though RGB output looks more colorful? | CIE Lab wins because it is more authentic, not more colorful; RGB's extra chroma clustered in silver-mirroring damage, and the scoring formula rewards fidelity over saturation. |
| What composite score weights does the benchmark use? | Composite score uses 15% CLIPScore, 15% aesthetic prediction, 20% ImageReward, 20% HPS, and 10% from X-IQE terms. |
| What is the text hallucination rate for CIE Lab output and what does it signify? | The CIE Lab output had a 4% text hallucination rate, and text hallucination is the clearest failure sign for RGB. |
| How does CIE Lab handle lightness compared with RGB in the benchmark mechanism? | In Lab, the model predicts only residual a*/b* chrominance; the original L* is copied back after every DDIM step, while in RGB all three channels are denoised together. |
| What are the metadata lineage retention and worst-case ΔE2000 results? | Metadata lineage retention is 97% for Lab against 61% for RGB, and the 40-seed worst-case ΔE2000 lands at 7.4 for Lab and 10.3 for RGB. |
Also worth reading: How to transform your old black and white photos into vibrant color masterpieces using AI: How to transform your old · A critical look at AI photo colorization: critical look at AI photo · How to transform your old black and white photos into vibrant memories with professional AI colorization: How to transform your old