MIRRA · The Science
The Science
A measurement device is judged by two things: how repeatable it is, and what it refuses to claim. Both are published here, including the parts that are not flattering to us.
Calibration
Thresholds were not chosen — they were measured. The score comes from a model trained on licensed dermatological databases, and each metric identifies the "presence/absence" of the characteristic with an AUC of 0.8 or higher on data the model has not seen. The thresholds below were measured in a re-test on real scans — how much each metric shifts between two readings of the same face.
Agreement with a physician's rating on the degree of the characteristic is weaker, and validity on field selfies has not yet been determined. This is a cosmetic metric for tracking, not diagnosis.
From this, two numbers are derived for each metric — the threshold below which a change is not considered an improvement, and the reliability (ICC) which indicates how repeatable the measurement is.
What goes into the score — and what doesn't
Ten metrics are measured. Eight enter the overall score. Two are measured and shown but not counted: the nasolabial fold and shine. Dryness counts since the network detects it.
The reason differs for each — shine varies with the time of day, the nasolabial fold moves with expression, and dryness is not yet as established with us as the others. A metric we cannot prove will not affect your score.
When we say: no measurable change
Every metric has a measured floor — the smallest change we can genuinely tell apart. Below it we will not call it an improvement, even if the number moved.
| Metric | Change threshold | Reliability |
|---|---|---|
| Redness | 17.3 | 0.926 |
| Spots & pigmentation | 20.1 | 0.906 |
| Tone evenness | 27.3 | 0.904 |
| Nasolabial fold | 48.7 | 0.869 |
| Pores | 27.4 | 0.924 |
| Fine lines | 33.0 | 0.907 |
| Dark circles | 31.3 | 0.887 |
| Shine | 40.7 | 0.938 |
| Blemishes | 46.1 | 0.917 |
| Dryness | 27.6 | 0.889 |
Thresholds: the wider of a capture-jitter study on 50 real scans and real repeat pairs · ICC from the jitter study, not from repeat visits · the metric comes from a model trained on licensed dermatology datasets · A SCAN, not a photograph: 5–9 frames fused