full lips ↔ perceived honesty, women (partial) — biggest geometry effect
.51
smiling ↔ perceived agreeable, both sexes — expression still beats bone structure
.57
people's guess of their own attractiveness rank vs the crowd's actual rank (n=29)
0/5
kernel-of-truth tests passing — the crowd can't actually read personality
1 · What the crowd reads into a face
Cells are partial correlations (controlling smile + apparent age; midface additionally head pitch) between a measured facial dimension and the mean crowd tag on that face. Bold = p<.005 raw. Hover any cell for raw r, n, p. Women: 157 tagged faces · men: 166.
The warmth axis is geometric and it replicates. In both sexes the same cluster carries perceived
honesty/agreeableness/trustworthiness — and its inverse carries dominance:
Wide nose (+.4-.5 warmth both sexes), wide face (+.4 women, +.3 men), short midface (−.4 women) → read as honest, agreeable, trustworthy, not dominant. This is the classic babyface/maturity halo, and it survives controls for smiling, apparent age, and ethnicity (adding DeepFace ancestry probabilities barely moves any coefficient — e.g. lips×honesty −.57→−.59).
Full lips (women): the sexy-vs-trustworthy tradeoff. Fuller lips read more attractive (+.27) and more dominant (+.34) but less honest (−.57), less trustworthy (−.51), less agreeable (−.46). Note this uses the expression-corrected lip metric — it is not smiling leaking through.
Smiling still beats geometry for warmth perception (+.3 to +.56 across honest/agreeable/trustworthy/extraverted, both sexes) — and smiling men read as smarter (+.34) while smiling women don't (+.07 ns).
Perceived dominance is just the anti-warmth pole — every warmth-positive feature is dominance-negative. The crowd runs warmth and dominance as one axis, not two.
perceived honest-humble vs lip fullness (women) — r=-0.66 [-0.74, -0.56], n=153 · the single biggest geometry effectperceived agreeable vs nose width (men) — r=0.62 [0.52, 0.71], n=163 · replicates the warmth axis in menperceived trustworthy vs face width (women) — r=0.57 [0.45, 0.67], n=151perceived agreeable vs smile (men) — r=0.52 [0.40, 0.62], n=166 · expression still rules perceptioncrowd attractiveness vs apparent age (women) — r=-0.41 [-0.53, -0.28], n=169crowd attractiveness vs face width (men) — r=-0.33 [-0.45, -0.18], n=174 · against the pop-fwhr story
Men:narrower face (−.36) — directly against the pop-culture fwhr-masculinity story — longer midface (+.24), looking younger (−.30).
2b · Do men and women want different faces?
Every face is rated by both sexes, so we can split each face's win rate by rater sex (AMAB vs AFAB)
and ask which measured features each side rewards. Dots: male raters ·
female raters. The headline: taste mostly AGREES
(r=0.679 on women's faces, r=0.764 on men's) — but the disagreements are systematic:
The smiling symmetry (the fun one): smiling is rewarded by the opposite sex and ignored by your own.
Men reward smiling women (+.10) while women don't (−.06, p=.001); women reward smiling men (+.20) while men barely do (+.08, p=.002).
Women are harder graders of femininity markers on women's faces: they reward full lips far more than men do
(+.39 vs +.19, the biggest split, p<10⁻⁴), reward upturned eyes more (+.21 vs +.09), and punish wide noses harder (−.33 vs −.24).
If you want to know which woman's face other women will crown, look at the lips.
Men over-value the "masculine jaw" on men — women don't: wide jaws win with male raters (+.12) and do nothing
for female raters (−.04, p<.001). The square-jaw ideal appears to be substantially a male-audience story.
Women punish older-looking men more than men do (−.38 vs −.27, p<.001) — despite the stereotype running the other way.
Midface length: no sex split. Everyone prefers longer midfaces on both sets (+.12 to +.37); the male–female differences are not significant.
18 difference tests → ~1 false positive expected at p<.05; the starred findings survive Bonferroni (p<.0028).
Matchmaking compresses win-rate spreads equally for both rater sexes, so the male−female comparisons are unaffected. Faces need ≥30 rated matchups per rater sex to count.
3 · Do measurements track the owner's actual personality?
Only the 47 face submitters have self-reports (3-item HEXACO per factor + lifestyle items), so this is
exploratory: at n≈41 only |r|≥.30 reaches p≈.05, and across 99 cells ~5 would pass by chance. Nothing here survives multiple-comparison correction.
eye tilt (3D)
face width
midface length
eye spacing
jaw width
nose width
lip fullness
smile
apparent age
Honesty-Humility (self)
-.10
.00
-.06
.23
.27
-.11
.04
.05
.02
Emotionality (self)
.21
.17
-.26
.09
.11
-.15
.02
-.11
-.09
Extraversion (self)
.04
-.27
.30
-.16
-.25
-.07
-.11
-.04
-.05
Agreeableness (self)
.30
-.17
.15
.06
-.10
-.27
-.09
-.07
-.00
Conscientiousness (self)
-.02
-.12
-.05
-.12
.09
-.01
-.30
.11
-.13
Openness (self)
-.06
-.11
.03
.16
.13
.16
.01
-.01
.25
Workout freq (0-3)
-.17
-.12
.09
.28
-.20
-.16
-.21
.14
-.20
Kinkiness (0-3)
-.01
.10
-.08
.00
-.12
.13
-.00
-.14
-.07
log bodycount
.13
-.08
.10
-.15
.00
.10
.21
.04
-.13
BMI
-.00
.35
-.31
.24
.25
.24
-.21
.15
.09
Raven's correct
·
·
·
·
·
·
·
·
·
Sanity check that works: BMI ↔ measured face width r=.35 — the geometry pipeline picks up real body composition.
Directions worth watching as submitters accumulate (all p≈.05-.06, uncorrected): self-reported extraverts have longer midfaces (+.30); agreeable people more upturned eyes (+.30); conscientious people thinner lips (−.30).
Bodycount, kinkiness, workout frequency: no credible face signal at this n.
4 · Kernel of truth: is the crowd right about anyone?
For submitted faces with ≥15 raters, crowd tag vs the owner's own self-report (sex-partialled):
crowd tag
self-report
n
r
95% CI
Honest-humble
HEXACO HH
20
0.01
[-0.43, 0.45]
Emotional
HEXACO EM
20
0.29
[-0.17, 0.65]
Extraverted
HEXACO EX
20
-0.22
[-0.61, 0.24]
Agreeable
HEXACO AG
20
-0.20
[-0.59, 0.26]
Conscientious
HEXACO CO
20
0.09
[-0.37, 0.51]
Open
HEXACO OP
3
—
too few
Intelligent
Raven's matrices score
2
—
too few
All CIs cross zero — no evidence the crowd reads real personality from these faces, consistent with an earlier reliability analysis of the same rating data. But self-knowledge is another story:
actual crowd percentile by self-predicted bracket — r=0.57 [0.26, 0.77], n=29 · dots = bin means ± 95% CI
People know where they rank. Submitters guessed their own attractiveness percentile before any votes came in; the guess correlates r=.57 (p=.001) with the crowd's eventual ranking.
· Raters are a self-selected, very-online sample — perception norms here aren't population norms.
· Tag means carry sampling noise (median ~16-21 raters/face) — true correlations are modestly larger than shown (attenuation), so the perception effects are conservative.
· 180 crowd cells tested → ~9 false positives expected at p<.05; the headline effects are p<10⁻⁶ and replicate across independent face sets, which is the real defense.
· Midface remains partially confounded with head pitch even after controls (⚠ rows).
· Self-report section is n≈41-47 with 3-item scales — noisy in both directions.
· Correlational throughout; photo sets are curated rating sets, not random faces.