chaosfactor.xyz

What do we call bodies?

People saw a photo of a real naked body and answered one yes/no question at a time: “Would you describe this person as skinny?” — then fat, curvy, healthy, and more, across a fixed pool of 125 bodies (75 women, 50 men). Here's what people and judgments say about how we label other people's bodies.

Data updated · survey: aella.lol/bodytagger

yes/no judgments
raters
125
bodies (75 W / 50 M)
of raters are women

How often each word applies

Average share of bodies that get a “yes” for each word — split by whether the photo shows a woman or a man. “Curvy” is essentially a word about women: it applies to % of women's bodies but only % of men's.

“Yes” rate by word

Mean per-image yes %. Note: the two pools are different bodies — differences partly reflect pool composition, not just word usage.

women's bodiesmen's bodies
data table

Which words people argue about

% of images where raters split 30–70 (no consensus). Low = everyone agrees when it applies and when it doesn't.

data table

People agree most about the extreme end — “obese” is contested on almost no images — and argue most about the slangy, aesthetic words: thicc, curvy, and where exactly “skinny” starts.

Hot bodies get different words — at the same weight

These 125 bodies also carry crowd attractiveness scores from an earlier pairwise “who's hotter?” game (~925,000 matchups). Hotter bodies get fewer heavy words, obviously — attractive people are thinner on average. The interesting question is what happens between bodies at the same apparent weight.

Word use across attractiveness quintiles (women's bodies)

mean “yes” % per hotness quintile, least → most attractive · 15 bodies per quintile

curvyfatskinny
data table

Hotness predicts… (raw, no weight control)

plain correlation of a word's yes-rate with attractiveness across the 75 women's bodies · 95% CI

data table

Uncontrolled, every heavy word points the same direction — hotter bodies are thinner, so they collect less of everything heavy and more of everything thin. That's real but unsurprising. The next chart is the same correlation after removing apparent weight:

Holding apparent weight constant, hotness predicts…

partial correlation of a word's yes-rate with attractiveness, controlling for the image's “overweight” rate · women's bodies · no CI shown (derived statistic)

data table

This is the report's clearest finding: “curvy” is substantially a hotness word, not a weight word. Comparing women's bodies at the same apparent weight, the more attractive one is far more likely to be called “curvy” (partial r = ) and “thicc,” while the less attractive one gets “gaunt” — and a bit more “fat” and “obese.” The flattering aesthetic words and the harsh ones aren't opposite ends of a weight scale; they're the same weight, judged through attractiveness. The same pattern, milder, holds for men's bodies.

Men and women rate bodies differently — but not how you'd guess

On a composite “does this rater place bodies heavier than other raters?” score, men and women barely differ — if anything women land slightly lighter. The real differences are in word choice, and they depend on whose body is being judged. Bars show how much more (or less) often women say “yes” to a word than men looking at the same image, with 95% CI.

Rating photos of women

women minus men, percentage points · positive = women apply the word more

data table

Rating photos of men

women minus men, percentage points · positive = women apply the word more

data table

The stable patterns as data accumulate: women favor the mild middle — more “normal weight” for everyone — and lean lighter on women's bodies (more “skinny,” less “overweight”), while on men's bodies they're somewhat quicker with “chubby” and “fat.” The biggest word gap by far is “heavyset”: women apply it at roughly double men's rate on both sexes' bodies (+15pp on women's), and “gaunt” skews female too — the softer, older-fashioned weight words read as a female register. One humility note: an earlier, smaller female-rater wave showed strong “curvy” avoidance by women; with the female sample since grown ~2.5×, that gap has faded into noise. These per-word estimates move with who shows up — the ones above are those that have held across waves.

Nice people use nice words

Raters answered HEXACO personality items between votes. Because bodies are randomly assigned, we can ask: do certain personalities describe the same body differently? (Shown for male raters — with usable profiles; the female sample is still too small for stable estimates.)

Agreeableness → charitable words

correlation (r) between a male rater's agreeableness and how often he applies each word, image-mix adjusted · 95% CI

data table

More-agreeable men call the same bodies “healthy” and “normal weight” more, and “obese” less. It's not that they see bodies as thinner — they just pick kinder words. Caveat from the confounder audit: the “healthy” effect survives controlling age/politics/religiosity; the heavy-word components attenuate below significance — agreeable men also skew left, and the two overlap.

Honesty-humility → less piling-on

same analysis for the honesty-humility factor · 95% CI

data table

Honesty-humility (the factor that tracks sincerity and modesty vs. entitlement) works as a second charity axis: high-H-H men reach for “healthy” and “normal weight” more, and both effects survive the confounder adjustment. Its early tendency toward fewer heavy words has faded toward noise as the sample grew — so between the two factors, personality-linked kindness in body-labeling lives mostly in choosing flattering words, not in withholding harsh ones.

Gym rats have stricter standards

Workout habit → harsher weight words

correlation (r) between a male rater's self-reported workout level (0–3) and word use (n = raters), image-mix adjusted · 95% CI

data table

Men who train see the same body and are measurably more likely to call it “chubby,” “fat,” or “overweight,” and less likely to call it “normal weight” or “healthy.” Fitness appears to move your reference point for what counts as heavy. This one is fully robust: adjusting for age, politics, and religiosity changes nothing (see the adj column in the table).

Your own body moves the goalposts

The mirror image of the gym effect: for the male raters who reported their own height and weight, the rater's BMI predicts how they label everyone else.

Rater BMI → gentler harsh words

correlation (r) between a male rater's own BMI and word use, image-mix adjusted · 95% CI

data table

Heavier raters apply “fat” (r = ) and “obese” less to the same bodies, with a mild lean toward “skinny” and “healthy.” Notably, the effect lives entirely in the harsh clinical words — a rater's own size does nothing to “curvy” or “thicc” (both r ≈ 0.00). Between this, the workout effect, and the politics result, the picture is consistent: where you stand determines where the category boundaries sit. The pattern survives adjustment for age, politics, and religiosity intact. The main caveat: height/weight is the sparsest variable in the data (n ≈ raters), so these CIs are wide.

The further left, the lighter the words

Raters also reported their politics on a 7-point left-right scale, separately for economic and social issues (n = male raters). Social politics turns out to be one of the strongest predictors of body language in the data.

Social leftism → lighter, kinder words

correlation (r) between a male rater's social leftism (−3 right … +3 left) and word use, image-mix adjusted · 95% CI

data table

Socially left men looking at the same bodies say “fat” dramatically less (r = ) and “skinny”/“healthy” more; social conservatives are the reverse. Economic politics shows the same pattern slightly weaker. Note this isn't just charity — the left shift is directional (more “skinny” AND less “fat”), meaning left-leaning and right-leaning men appear to genuinely place the same bodies at different weights, with the right seeing heavier. And it is not an age or religiosity artifact — controlling both leaves every coefficient essentially unchanged.

Porn: how an effect dies honestly

Raters reported how much porn they consume (0–9 frequency scale; n = male raters with enough votes). In an early wave this looked like a real finding — heavy porn consumers rating bodies lighter (“fat” at r = −0.17). We keep it in the report as a worked example of what happens next:

Porn habit → word use (raw)

correlation (r) between a male rater's porn consumption (0–9) and word use, image-mix adjusted · 95% CI

data table

With the sample since doubled, the raw “fat” effect has shrunk to r = (no longer significant). And there was always an obvious confound to check: porn consumption correlates with social leftism in this sample (r = ), and leftism predicts lighter labels. Here's the porn effect with social politics partialed out:

Porn habit, adjusted for social politics

partial correlation (r) controlling for the rater's social leftism · same raters, 95% CI

data table

Verdict: dead. Politics-adjusted, the “fat” effect is r = — indistinguishable from zero. What looked like porn recalibrating people's weight categories was, on more data, a modest correlation that was itself mostly leftists watching more porn. This is the normal life cycle of a flashy early finding, shown in full rather than quietly deleted.

Everything else we checked

The same image-adjusted analysis across the remaining demographics with decent coverage (male raters). Only correlations whose 95% CI excludes zero are shown:

Demographic sweep — significant effects only

r between the trait and word use for the same images · sorted by |r|

Reading the sweep — with the confounder audit applied (adj columns in the tables control age, social politics, and religiosity): religiosity looks harsher raw, but the entire effect vanishes once politics is controlled — religious raters label harshly only insofar as they're socially conservative; religiosity itself adds nothing. Older raters use “skinny” and “chubby” less, and that survives adjustment (fat even strengthens: older men say “fat” less once politics is held fixed). Less-straight men (0–6 Kinsey-style scale) look gentler raw — more “healthy”/“skinny,” fewer heavy words, mostly when rating women's bodies — but as of the latest data every orientation effect drops below significance once politics is controlled. Like religiosity, orientation appears to be mostly a politics proxy here.

What didn't matter

Null results are half the story, so here's what we looked for and did not find. Numbers below are the strongest relevant correlation, with its 95% CI comfortably straddling zero:

Notable nulls

each row is the best shot the hypothesis had · male raters unless noted

The standouts: conscientiousness has nothing whatsoever to say about body words, even at n ≈ 890, and openness nothing either (its early borderline blip on “healthy” faded, as chance blips do). “Thicc” is not generational — young and old raters use it identically. The gym effect is confined to weight words — training does nothing to the aesthetic vocabulary. Men and women agree almost perfectly on who's “healthy.” And the overall sex difference in placing bodies heavier vs lighter is statistically indistinguishable from zero. (One promotion out of this section: emotionality, an “honorable mention” in an earlier version, has firmed up with more data — high-E men use “chubby,” “fat,” and “overweight” all less, r ≈ −0.09 to −0.13, all CIs now clear of zero — a third personality kindness axis alongside agreeableness and honesty-humility, though it hasn't yet been through the politics adjustment.)

Bonus: the boob census

A random third of (non-female) raters got a different task first: classify the chest size of 100 women's bodies on a 7-point scale. classifications from raters so far:

Distribution of size judgments

all classifications pooled across raters and images

data table
Methods & caveats. All rater-trait analyses use MALE raters only — rater sex is the biggest confounder of all (politics, porn, orientation and personality all differ by sex), so it is removed by design rather than adjusted; the female sample gets its own analyses once it is big enough. Trait effects are additionally re-estimated with age, social politics, and religiosity partialed out — the adj columns in each data table; effects that die under adjustment are called out in the text. Every rater judges the same frozen 125-body pool, so rater comparisons are within-image (paired per image, or leave-one-out deviation from the image's mean for per-rater scores; word-level scores need ≥10 votes on that word). Rater sex uses AMAB/AFAB binning; HEXACO factors are scored from a 2–3-item drip (middle item reverse-keyed); attractiveness is the crowd Elo from the earlier pairwise body-ranking games; “apparent weight” control is the image's “overweight” yes-rate. The word roster rotates (curvy/skinny/chubby paused 2026-08-08 while gaunt, heavyset and plus-size collect votes), so per-word Ns differ. Caveats: the sample is internet-recruited and skews male ( women); with this many tests, a couple of borderline CIs would clear zero by chance — trust the patterns that are coherent across words, which is what's highlighted; female-rater estimates are noisier than male ones. Data updates as votes come in.
by aella · survey at aella.lol/bodytagger · votes as of