MIRRA · The Science

טלפון, כרטיס כיול צבע, סרגל וגרפים
Color and scale calibration — what turns an image into a measurement.

The Science

A measurement device is judged by two things: how repeatable it is, and what it refuses to claim. Both are published here, including the parts that are not flattering to us.

Calibration

Thresholds were not chosen — they were measured. The score comes from a model trained on licensed dermatological databases, and each metric identifies the "presence/absence" of the characteristic with an AUC of 0.8 or higher on data the model has not seen. The thresholds below were measured in a re-test on real scans — how much each metric shifts between two readings of the same face.

Agreement with a physician's rating on the degree of the characteristic is weaker, and validity on field selfies has not yet been determined. This is a cosmetic metric for tracking, not diagnosis.

From this, two numbers are derived for each metric — the threshold below which a change is not considered an improvement, and the reliability (ICC) which indicates how repeatable the measurement is.

doi.org/10.6084/m9.figshare.5047666 · Full table →

What goes into the score — and what doesn't

Ten metrics are measured. Eight enter the overall score. Two are measured and shown but not counted: the nasolabial fold and shine. Dryness counts since the network detects it.

The reason differs for each — shine varies with the time of day, the nasolabial fold moves with expression, and dryness is not yet as established with us as the others. A metric we cannot prove will not affect your score.

When we say: no measurable change

Every metric has a measured floor — the smallest change we can genuinely tell apart. Below it we will not call it an improvement, even if the number moved.

MetricChange thresholdReliability
Redness17.30.926
Spots & pigmentation20.10.906
Tone evenness27.30.904
Nasolabial fold48.70.869
Pores27.40.924
Fine lines33.00.907
Dark circles31.30.887
Shine40.70.938
Blemishes46.10.917
Dryness27.60.889

Thresholds: the wider of a capture-jitter study on 50 real scans and real repeat pairs · ICC from the jitter study, not from repeat visits · the metric comes from a model trained on licensed dermatology datasets · A SCAN, not a photograph: 5–9 frames fused

doi.org/10.6084/m9.figshare.5047666

← Back to homepage