One photograph, placed in a measured reference population of 91,810 posts. The instrument reports where it falls and how much its own ensemble disagrees.
Drop a photo here, or choose one. If more than one person is in frame, one is selected and scored.
Nothing is stored. The image is scored and discarded.
The live scorer runs on a single machine which is not running right now. The recording below shows the same path a live upload takes.
Ensemble spread —. This is the range across the three fold checkpoints and the horizontal-flip pair. Each model performs different tasks to form a complete score. This is not a confidence interval.
| Checkpoint | Percentile | Upright / flipped | Std. score |
|---|
Each fold is a separate model, trained with a different fifth of the data held out, and selected at its own epoch. Their raw scores are on different scales and are not comparable; each is standardized against its own held-out distribution first, which is what makes the three readings above comparable at all. The headline is the mean of the six standardized members, not the average of the three percentiles, since the percentile map is non-linear.
Predicted engagement for this image posted solo, as the thumbnail, by a standardized account, under this population's revealed preferences, with variables such as page reach, follower count, posting history, carousel position and source compression held fixed.
Without it, a score mostly measures how many followers an account has. The reference context pins every one of them to the same value for every image, so what varies between two scores is the photograph.
The ensemble band is typically wide, a median of 39 percentile points between the lowest and highest member. The band is shown at the same weight as the estimate because it is the same size as the estimate.