People saw a photo of a real naked body and answered one yes/no question at a time: “Would you describe this person as skinny?” — then fat, curvy, healthy, and more, across a fixed pool of 125 bodies (75 women, 50 men). Here's what people and judgments say about how we label other people's bodies.
Data updated · survey: aella.lol/bodytagger
Average share of bodies that get a “yes” for each word — split by whether the photo shows a woman or a man. “Curvy” is essentially a word about women: it applies to % of women's bodies but only % of men's.
Mean per-image yes %. Note: the two pools are different bodies — differences partly reflect pool composition, not just word usage.
% of images where raters split 30–70 (no consensus). Low = everyone agrees when it applies and when it doesn't.
People agree most about the extreme end — “obese” is contested on almost no images — and argue most about the slangy, aesthetic words: thicc, curvy, and where exactly “skinny” starts.
These 125 bodies also carry crowd attractiveness scores from an earlier pairwise “who's hotter?” game (~925,000 matchups). Hotter bodies get fewer heavy words, obviously — attractive people are thinner on average. The interesting question is what happens between bodies at the same apparent weight.
mean “yes” % per hotness quintile, least → most attractive · 15 bodies per quintile
plain correlation of a word's yes-rate with attractiveness across the 75 women's bodies · 95% CI
Uncontrolled, every heavy word points the same direction — hotter bodies are thinner, so they collect less of everything heavy and more of everything thin. That's real but unsurprising. The next chart is the same correlation after removing apparent weight:
partial correlation of a word's yes-rate with attractiveness, controlling for the image's “overweight” rate · women's bodies · no CI shown (derived statistic)
This is the report's clearest finding: “curvy” is substantially a hotness word, not a weight word. Comparing women's bodies at the same apparent weight, the more attractive one is far more likely to be called “curvy” (partial r = ) and “thicc,” while the less attractive one gets “gaunt” — and a bit more “fat” and “obese.” The flattering aesthetic words and the harsh ones aren't opposite ends of a weight scale; they're the same weight, judged through attractiveness. The same pattern, milder, holds for men's bodies.
On a composite “does this rater place bodies heavier than other raters?” score, men and women barely differ — if anything women land slightly lighter. The real differences are in word choice, and they depend on whose body is being judged. Bars show how much more (or less) often women say “yes” to a word than men looking at the same image, with 95% CI.
women minus men, percentage points · positive = women apply the word more
women minus men, percentage points · positive = women apply the word more
The stable patterns as data accumulate: women favor the mild middle — more “normal weight” for everyone — and lean lighter on women's bodies (more “skinny,” less “overweight”), while on men's bodies they're somewhat quicker with “chubby” and “fat.” The biggest word gap by far is “heavyset”: women apply it at roughly double men's rate on both sexes' bodies (+15pp on women's), and “gaunt” skews female too — the softer, older-fashioned weight words read as a female register. One humility note: an earlier, smaller female-rater wave showed strong “curvy” avoidance by women; with the female sample since grown ~2.5×, that gap has faded into noise. These per-word estimates move with who shows up — the ones above are those that have held across waves.
Raters answered HEXACO personality items between votes. Because bodies are randomly assigned, we can ask: do certain personalities describe the same body differently? (Shown for male raters — with usable profiles; the female sample is still too small for stable estimates.)
correlation (r) between a male rater's agreeableness and how often he applies each word, image-mix adjusted · 95% CI
More-agreeable men call the same bodies “healthy” and “normal weight” more, and “obese” less. It's not that they see bodies as thinner — they just pick kinder words. Caveat from the confounder audit: the “healthy” effect survives controlling age/politics/religiosity; the heavy-word components attenuate below significance — agreeable men also skew left, and the two overlap.
same analysis for the honesty-humility factor · 95% CI
Honesty-humility (the factor that tracks sincerity and modesty vs. entitlement) works as a second charity axis: high-H-H men reach for “healthy” and “normal weight” more, and both effects survive the confounder adjustment. Its early tendency toward fewer heavy words has faded toward noise as the sample grew — so between the two factors, personality-linked kindness in body-labeling lives mostly in choosing flattering words, not in withholding harsh ones.
correlation (r) between a male rater's self-reported workout level (0–3) and word use (n = raters), image-mix adjusted · 95% CI
Men who train see the same body and are measurably more likely to call it “chubby,” “fat,” or “overweight,” and less likely to call it “normal weight” or “healthy.” Fitness appears to move your reference point for what counts as heavy. This one is fully robust: adjusting for age, politics, and religiosity changes nothing (see the adj column in the table).
The mirror image of the gym effect: for the male raters who reported their own height and weight, the rater's BMI predicts how they label everyone else.
correlation (r) between a male rater's own BMI and word use, image-mix adjusted · 95% CI
Heavier raters apply “fat” (r = ) and “obese” less to the same bodies, with a mild lean toward “skinny” and “healthy.” Notably, the effect lives entirely in the harsh clinical words — a rater's own size does nothing to “curvy” or “thicc” (both r ≈ 0.00). Between this, the workout effect, and the politics result, the picture is consistent: where you stand determines where the category boundaries sit. The pattern survives adjustment for age, politics, and religiosity intact. The main caveat: height/weight is the sparsest variable in the data (n ≈ raters), so these CIs are wide.
Raters also reported their politics on a 7-point left-right scale, separately for economic and social issues (n = male raters). Social politics turns out to be one of the strongest predictors of body language in the data.
correlation (r) between a male rater's social leftism (−3 right … +3 left) and word use, image-mix adjusted · 95% CI
Socially left men looking at the same bodies say “fat” dramatically less (r = ) and “skinny”/“healthy” more; social conservatives are the reverse. Economic politics shows the same pattern slightly weaker. Note this isn't just charity — the left shift is directional (more “skinny” AND less “fat”), meaning left-leaning and right-leaning men appear to genuinely place the same bodies at different weights, with the right seeing heavier. And it is not an age or religiosity artifact — controlling both leaves every coefficient essentially unchanged.
Raters reported how much porn they consume (0–9 frequency scale; n = male raters with enough votes). In an early wave this looked like a real finding — heavy porn consumers rating bodies lighter (“fat” at r = −0.17). We keep it in the report as a worked example of what happens next:
correlation (r) between a male rater's porn consumption (0–9) and word use, image-mix adjusted · 95% CI
With the sample since doubled, the raw “fat” effect has shrunk to r = (no longer significant). And there was always an obvious confound to check: porn consumption correlates with social leftism in this sample (r = ), and leftism predicts lighter labels. Here's the porn effect with social politics partialed out:
partial correlation (r) controlling for the rater's social leftism · same raters, 95% CI
Verdict: dead. Politics-adjusted, the “fat” effect is r = — indistinguishable from zero. What looked like porn recalibrating people's weight categories was, on more data, a modest correlation that was itself mostly leftists watching more porn. This is the normal life cycle of a flashy early finding, shown in full rather than quietly deleted.
The same image-adjusted analysis across the remaining demographics with decent coverage (male raters). Only correlations whose 95% CI excludes zero are shown:
r between the trait and word use for the same images · sorted by |r|
Reading the sweep — with the confounder audit applied (adj columns in the tables control age, social politics, and religiosity): religiosity looks harsher raw, but the entire effect vanishes once politics is controlled — religious raters label harshly only insofar as they're socially conservative; religiosity itself adds nothing. Older raters use “skinny” and “chubby” less, and that survives adjustment (fat even strengthens: older men say “fat” less once politics is held fixed). Less-straight men (0–6 Kinsey-style scale) look gentler raw — more “healthy”/“skinny,” fewer heavy words, mostly when rating women's bodies — but as of the latest data every orientation effect drops below significance once politics is controlled. Like religiosity, orientation appears to be mostly a politics proxy here.
Null results are half the story, so here's what we looked for and did not find. Numbers below are the strongest relevant correlation, with its 95% CI comfortably straddling zero:
each row is the best shot the hypothesis had · male raters unless noted
The standouts: conscientiousness has nothing whatsoever to say about body words, even at n ≈ 890, and openness nothing either (its early borderline blip on “healthy” faded, as chance blips do). “Thicc” is not generational — young and old raters use it identically. The gym effect is confined to weight words — training does nothing to the aesthetic vocabulary. Men and women agree almost perfectly on who's “healthy.” And the overall sex difference in placing bodies heavier vs lighter is statistically indistinguishable from zero. (One promotion out of this section: emotionality, an “honorable mention” in an earlier version, has firmed up with more data — high-E men use “chubby,” “fat,” and “overweight” all less, r ≈ −0.09 to −0.13, all CIs now clear of zero — a third personality kindness axis alongside agreeableness and honesty-humility, though it hasn't yet been through the politics adjustment.)
A random third of (non-female) raters got a different task first: classify the chest size of 100 women's bodies on a 7-point scale. classifications from raters so far:
all classifications pooled across raters and images