← המקורות

דוח מדידה

הדוח המלא: אמינות, נכונות, תיקוף חיצוני וביקורת תקינה — עם המספרים ורווחי הסמך. כתוב לקוראים מקצועיים. השרשרת שמייצרת כל מספר כאן נמצאת ב-research/PROVENANCE.md במאגר.

מה המדידה שלנו לא עושה

תיקוף חיצוני, ותוצאה שלא מחמיאה לנו. אמינות אומרת שמדידה חוזרת על עצמה. היא לא אומרת שהיא מודדת את מה ששמה מרמז. אז בדקנו את הציון שלנו מול אמת-מידה שלא היה לנו חלק ביצירתה: 2,513 אנשים שדירגו את אותן 100 פנים (אמינות אמת-המידה 0.992).

הציון הכולל: r = 0.18, וכשמנטרלים גיל 0.10 עם רווח סמך -0.07 עד 0.27 — שחוצה אפס. כלומר: הציון הכולל שלנו לא מנבא איך פנים נתפסות, ושום מקום במוצר לא יטען אחרת.

וזו התוצאה הצפויה, לא כישלון: מה שאנשים מדרגים הוא בעיקר מבנה עצם, סימטריה וגיל — דברים שמדידת עור לא רואה ולא אמורה לראות. קורלציה גבוהה כאן הייתה עדות לכך שהמדדים קוראים צורת פנים.

ומה שכן החזיק: טקסטורה (+0.21) · רפיון (+0.19) · texture (+0.14) — בדיוק המדדים ש-Fink & Matts חוזה שישפיעו על תפיסה. הציטוט הזה נושא משקל ולא מקשט.

האם זה עובד באותה מידה על כולן

the east Asian subgroup is 9 faces and carries the lowest ICC measured. Too small to act on and too clear to leave unrecorded — re-measure on real users before any per-group claim, and before anybody cites the overall ICC as if it applied evenly.
whiteICC 0.79 · 68 פנים
blackICC 0.84 · 13 פנים
east_asianICC 0.68 · 9 פנים
west_asianICC 0.87 · 9 פנים
ומה שעדיין אין לנו. This is reliability plus one external criterion. It is NOT a validation against a reference instrument (VISIA, Corneometer) or against dermatologist grading, and until one exists the product may not say "clinically proven".

תקנים

יש תקן לדיווח על בדיוק מה שאנחנו מודדים, והוא נקרא GRRAS. חמישה-עשר סעיפים שקובעים איך מדווחים מחקר אמינות: מהו המכשיר, מי הנבדקים, איך נבחר גודל המדגם, אילו סטטיסטיקות, ואילו רווחי סמך. הטבלה למטה היא אנחנו מול חמישה-עשר הסעיפים, אחד-אחד.

14 עומדים · 1 חלקית · 0 לא עומדים · 0 לא רלוונטי.

הסעיף החלקי הוא מספר 4 — אוכלוסיית המדרגים. רץ פיילוט דירוג עיוור של 12 פנים, במודל ראייה יחיד: זו לא דרמטולוגית וזו לא אוכלוסייה. כל רווחי הסמך שם חצו אפס, כלומר הפיילוט לא קבע דבר לגבי הסכמה — מה שהוא כן עשה זה לחשב כמה פנים המחקר האמיתי צריך: 30 לעוצמה של 80%, ו-40 לרווח סמך צר מ-0.50. רשימת תקינה שכל השורות בה ירוקות היא לוגו, לא ביקורת.
1Identify whether interrater/intrarater reliability or agreement was investigated
models/calibration.json records retest as "neutral vs smiling front shot, face moved between them" — test-retest of one automated rater, not interrater. There are no human raters to disagree.
עומד
2Name and describe the measurement device explicitly
analyzer/metrics.py: eleven named metrics, each computed in its own facial zone in CIELAB, on a lighting-normalized image. Version and scale recorded per scan.
עומד
3Specify the subject population of interest
102 identities, mean age 27.7 (SD 7.1, range 18-54), 49 female / 53 male, 69 white / 13 black / 10 west Asian / 9 east Asian — recorded in models/calibration.json under calibration_set.demographics.
עומד
4Specify the rater population (if applicable)
It WAS not applicable — the rater is a deterministic program, so there is no rater population and no interrater variance. That stopped being true when a grading study ran: twelve faces graded blind on published photonumeric scales. The rater population is therefore ONE non-expert vision model, which is not a dermatologist and not a population. Every Spearman interval from it crossed zero, so it establishes nothing about agreement and is reported as a pilot that sized the real study rather than as a result.
חלקית
5Describe what is already known about reliability and provide a rationale
analyzer/evidence.py — 25 sources, including the smartphone-colorimeter benchmark (ICC 0.85-0.95) our own numbers are compared against, and Flament 2023, which found the same two weak spots we measured.
עומד
6Explain how the sample size was chosen
The calibration sample was fixed at 102 before this could be planned, so the honest answer is a PRECISION analysis of it rather than an a priori calculation, and both are reported. Delivered: ICC intervals with a median width of 0.156, from 0.027 (unevenness) to 0.333 (pores). Required: interval width scales as 1/sqrt(n), so the four-face file this replaced carried a width near 0.79 — the entire ICC scale, which is why its thresholds understated the noise by up to eightfold. For the expert-agreement study, which is not yet run and CAN be planned, models/study_design.json carries a simulated power calculation: n=30 for 80% power at rho=0.6, n=40 for an interval narrower than 0.50.
עומד
7Describe the sampling method
Convenience sample: the complete Face Research Lab London Set (CC-BY-4.0, DOI 10.6084/m9.figshare.5047666), every identity with both expressions used, none excluded. Convenience sampling is a real limit on generalisation and is recorded as one.
עומד
8Describe the measurement process (time interval, blinding)
Two front shots per identity, same session, different expression. Blinding is not applicable to a deterministic program. The interval is short, which makes the retest HARDER rather than easier: the difference contains a real facial change, so the noise estimate is conservative.
עומד
9State whether measurements were conducted independently
analyzer/metrics.py holds no state between calls: each image is processed independently and the analyzer cannot see the other measurement, which is the strongest form of independence available.
עומד
10Describe the statistical analysis
sigma_within = SD of paired differences / sqrt(2); sigma_between = sqrt(max(var(first) - var_within, 0)); ICC = var_between / (var_between + var_within); MDC95 = 1.96*sqrt(2)*sigma_within. Intervals by bootstrap over FACES, 2000 resamples, seed 7 — recorded in calibration.method so the analysis is reproducible from the file.
עומד
11State the actual number of raters, subjects and replicate observations
Per metric: n_faces and n_rows, in every entry of calibration.metrics. A metric measured on fewer faces than another cannot hide behind the headline count.
עומד
12Describe the sample characteristics
models/calibration.json, calibration_set.demographics: age mean/SD/range, gender split, ethnicity counts. Added because this item asked for it and the file did not carry it.
עומד
13Report estimates of reliability and agreement including statistical uncertainty
Every metric carries icc, icc_ci95 (bootstrap over faces), noise_std, mdc95 and a quality band judged on the LOWER bound of the interval per Koo & Li 2016 — never on the point estimate.
עומד
14Discuss the practical relevance of the results
MDC95 IS the practical relevance: it is the threshold the product must clear before it may call a change real, and it moved the bar on the overall score from 6.15 points to 8.52. models/validation.json states separately that clearing it means detectable, NOT important — de Vet & Terwee, and the distinction the product is forbidden from blurring.
עומד
15Provide detailed results if possible
models/calibration.json and models/validation.json are served in full at /v1/evidence and summarised at mirra.skin/sources.html — the numbers, the intervals, the sample, and the limitations.
עומד

תקנים נוספים

EEMCO guidance for the assessment of skin colour

Colour is read in CIELAB following CIE recommendations, and erythema is separated from melanin by reflectance rather than taken from the red channel. EEMCO explicitly endorses image analysis by colour camera as an accurate route. It also warns that erythema is barely visible in deeply melanized skin — the limitation our own fairness work found independently, which is the strongest sign this is the right standard to be held to.

CIE colorimetry (ISO/CIE 11664)

Every colour metric is computed in CIELAB, and skin tone is reported as ITA° = arctan((L*-50)/b*)·180/π, the accepted colorimetric definition.

Guidelines for measurement of skin colour and erythema (ESCD Standardization Group)

The standard asks for control of individual, temporal and ENVIRONMENTAL variables. We record head pose, expression, capture quality and the weather at scan time, and the capture coach enforces framing, distance, tilt and expression before the shutter. What we cannot do is control the room: a clinic fixes the illuminant, and a phone in a bathroom does not. Lighting normalisation reduces this and does not remove it.

Medical device regulation (EU MDR, FDA, ISO 13485)

MIRRA is not a medical device and holds no CE mark or FDA clearance. It measures and tracks; it does not detect, diagnose or treat, and the moment a skin measurement claims to do any of those it is regulated as a device. Stated here so its absence reads as a decision rather than an omission.

לא רלוונטיnot applicable by design

מקור לכל מדד

פצעוניםAutomated facial acne lesion detecting and counting algorithm for acne severity evaluation
עיגולים כהיםCutaneous colorimetry · O'Mahony MM, Sladen C, Crone M, et al · X-Rite ColorChecker Classic — CIE L*a*b* reference values · Axelsson J, Sundelin T, Ingre M, et al
יובשKoseki K, Kawasaki H, Atsugi T, et al · Garg A, Chren MM, Sands LP, et al · Altemus M, Rao B, Dhabhar FS, et al
נקבוביותFlament F, et al · Dissanayake B, Miyamoto K, Purwar A, et al
אדמומיותFink B, et al · Cutaneous colorimetry · Dawson JB, et al · Erythema image analysis in rosacea · VISIA redness comparison · X-Rite ColorChecker Classic — CIE L*a*b* reference values · Piérard GE · Fullerton A, et al
קמט השפה-אףNarins RS, Carruthers J, Flynn TC, et al
ברק שומניKohli I, Kastner S, Thomas M, et al
כתמיםFink B, Matts PJ · Cutaneous colorimetry · X-Rite ColorChecker Classic — CIE L*a*b* reference values · Retinol 0
חוסר אחידותFink B, Matts PJ · Fink B, et al · Fink B, et al · Cutaneous colorimetry · Koseki K, Kawasaki H, Atsugi T, et al · X-Rite ColorChecker Classic — CIE L*a*b* reference values · Piérard GE · Retinol 0
קווים דקיםJiang LI, Stephens TJ, Goodman R · Retinol 0