Method
The machine decides what the report is about before any language model is involved.
A model writes well and cannot resist a sentence that sounds insightful. So the decisions that determine whether a report is true — which findings matter, in what order, and whether the research supports them — are made by code that cannot see who you are. This page walks through the whole of it, including the parts that are judgement calls, so you can decide how much to trust what you get.
Once those decisions are made, the model is handed the lot — every answer, every number, the whole corpus — and asked to explain it well. Selection comes first and is blind. Composition comes second and knows everything.
What the Big Five actually is
Worth a few paragraphs before the machinery, because most of what circulates about personality testing is either the wrong model or the right model oversold.
Five broad dimensions — Neuroticism, Extraversion, Openness, Agreeableness, Conscientiousness — keep turning up when researchers factor-analyse the words people use to describe each other. They are not a theory of the mind anyone designed; they are what falls out of the data, repeatedly, which is both their strength and the limit of what they claim. Each splits into six narrower facets, and this instrument scores all thirty. DeYoung and colleagues found a real layer in between as well — two “aspects” per domain, so Extraversion divides into enthusiasm and assertiveness, Agreeableness into compassion and politeness. Your report uses that middle layer where your own facets disagree.
The finer the grain, the more it tells you. Mõttus and colleagues showed that models built from individual items predicted about 9.7% of the variance in life outcomes, against 5.9% for facets and 4.2% for the five domains — and that those item-level quirks are about as stable and as heritable as the big traits themselves. Averaging thirty facets into five numbers throws away more than half the signal. It is the main reason this takes three hundred items instead of fifty.
There are no types. Gerlach and colleagues found four denser regions in the cloud of profiles, but the space itself is continuous, and a commentary showed that a distribution with no clusters at all produces the same apparent groupings. Nobody is an introvert the way a number is even; everyone sits somewhere on a smooth gradient.
Traits are stable but not fixed. Roberts and colleagues documented a maturity principle — people tend to gain conscientiousness and emotional stability across their twenties and thirties — and a later meta-analysis found that clinical interventions shift trait scores, though once corrected for small-study bias the effect is modest and its confidence interval crosses zero. Change is real, slow, and smaller than the self-improvement genre implies.
And the numbers are noisier than they look. Gnambs put the median dependability of Big Five scores at about .82, meaning roughly a tenth of any score is which day you took it. Soto found that 87% of trait-to-outcome findings replicate, at about 77% of their original strength — good by the standards of the field, and still a reminder that the first number published is rarely the right one. Sackett and colleagues revised the headline validity of conscientiousness for job performance from .31 down to .19 after re-examining a statistical correction. Real effects, modest sizes, population-level.
Last, the thing this report is most at risk of. In 1949 Bertram Forer gave his students a personality profile they rated as remarkably accurate; it was one generic text handed identically to all of them. Vague, flattering and broadly true statements feel personal. That is why the layers below work so hard to make every sentence here traceable to a number that could have come out differently.
The pipeline
1. Validity checks, before anything else
Runs first and can stop the whole thing. A run is rejected if it contains a straight run of 20 identical answers, if 85% or more of the items got the same response, if it was completed faster than 1.2 seconds per item, or if forward- and reverse-worded items were answered in the same direction — the signature of agreeing regardless of content.
If any of those trip, you get an explanation instead of a report. Noise interpreted confidently is worse than no interpretation.
2. Scoring
Each facet is the sum of its ten items, reverse-worded ones flipped. Each domain is the sum of its six facets. Those raw sums become percentiles through Johnson’s sex- and age-banded norms, using his cubic transform rather than a normal curve — several facets are ceiling-skewed, and a normal approximation is wrong in exactly the tails that matter.
3. Salience — five kinds of finding
Five different ways of asking what is worth saying about a profile. Only the first two are about being far from average, and they are the crudest of the five: a very high score is easy to spot and often the least informative thing on the page, because it is frequently just the trait doing what that trait does. The other three look for structure instead — scores that disagree with each other, or with what the rest of your profile predicted — and those routinely surface from numbers sitting comfortably in the middle.
Every finding carries its own numeric evidence, and the report shows you that evidence next to the sentence it produced.
- Facet extremes. Any facet at or beyond the 75th or 25th percentile.
- Domain extremes, down-weighted when the domain’s own facets disagree by more than 35 points — an average over facets that disagree is a misleading number, and the facets are the story.
- Aspect splits. Divergence within a domain, either between its two aspects (20+ points) or between its highest and lowest facet (60+ points).
- Covariance residuals. Facets correlate in the population, so each of your thirty is predicted from the other twenty-nine, and anything off by 1.5 standard deviations or more is reported. This is the layer that finds things a percentile cannot: a perfectly average score can be the most surprising number you have, if everything around it predicted otherwise.
- Item splits. Within one facet, when the ten items do not agree — promise-keeping high while rule-following is ordinary, say. The facet score hides this completely.
4. The balance quota
Deficits are easier to name than strengths, so a list ranked by magnitude alone tilts toward what is wrong with someone. Telling the writer to “be balanced” does not fix it — you get cheerful sentences about a risk-skewed selection.
So the fix is in selection. Findings are sorted by a fixed rule table into assets, tensions, and neutral-structural, and the report must carry at least 3 assets, 2 tensions, and 3 structural observations, up to 12 findings in all.
Two deliberate calls in that table. Low Gregariousness and low Friendliness are classified neutral, not deficits — chosen solitude is not isolation, and treating quiet as a fault is how a report like this insults its reader. Modesty is neutral at both ends, because the instrument’s Modesty scale is known to under-represent what it appears to measure.
If a quota floor cannot be met, the report says so. It does not invent a weakness to balance the page.
5. Research
Every paper in the corpus goes to the model on every report, and every one of them is citable. What changes from reader to reader is how much of each it gets to read.
A dozen or so papers arrive as complete text, cover to cover. The rest arrive as a structured summary — the citation, the sample, and the findings in the paper’s own words — and may be cited for exactly that much and nothing further. This is a limit of one request, not a judgement about the research: we hold the full text of nearly every paper here, and the summarised ones are simply the ones this profile had least use for.
Which papers get read in full is decided by plain threshold rules — conditions like Excitement-seeking ≥ 80, not a similarity search — ranked by how many of your selected findings each paper actually speaks to. They fill a fixed reading budget, smaller papers slotting in behind larger ones rather than the budget being abandoned at the first one that does not fit.
Each paper also states what it supports and, more importantly, what it does not. Those limits go to the model word for word, and they are the part that does the real work — the boundary a summary usually loses is exactly the one a confident sentence needs.
6. Composition
Only here does a model appear, and it is given all of your data: all 300 answers with their item text and keying, every score in raw, T-score and percentile form, the norm parameters those came from, how the thirty facets covary across 306,835 people, what that covariance predicted for each of yours, the selected findings, the corpus as described above, and anything you wrote in the optional context box.
It gets all of it because thresholds find the obvious things and miss the rest — a pattern spread across four moderate residuals, a facet whose ten items split along a line no rule anticipated. What keeps that from becoming licence: every number in the report has to appear in the data it was given, every claim about people in general has to name a paper that licenses it, and what you wrote about yourself may shape how something is explained but can never be evidence for it. Where your context and your data disagree, it is told to say so rather than resolve it in your favour.
The layers above stay blind. What the report is about is still decided from your answers alone, before a model sees anything.
7. The self-audit
Before returning, the writer re-reads its own draft against the data: every number checked against the dossier, every citation against what that paper says it licenses, every claim against whether it would be false for a different profile. The report then ends by naming what it could not support — a pattern with no paper behind it, a reading it finds plausible but cannot ground. That section is how you judge how far to trust the others.
One model call, and the check lives inside it. A verdict produced after the writing is finished changes nothing unless something acts on it, and the only thing that can actually rewrite a sentence is the call that wrote it.
What the writer is forbidden to do
- NO NUMBER THAT IS NOT IN THE TABLE. Every percentile, correlation, effect size and sample size you write must appear verbatim in the data you were given. Do not compute, round differently, average, convert, or estimate. If you want to say "about a third", check that the table supports it — and prefer quoting the number.
- NO CLAIM THE CORPUS DOES NOT LICENSE. If no supplied paper supports an explanation, describe the pattern and stop. "I do not have research that explains why these two sit together" is a correct and acceptable sentence. Reaching for a plausible mechanism is the failure this whole pipeline exists to prevent.
- NO BARNUM STATEMENTS. Every claim must be one that would be FALSE for a meaningfully different profile. Before writing a sentence about the reader, ask whether it would still read as true for someone at the opposite end of the scale. If it would, delete it.
It also may not assign a type or category, may not dichotomise a percentile into “high” or “low”, may not diagnose, and may not claim that two traits combine into something emergent — trait-by-trait interactions replicate poorly, and the corpus carries a specific null result for the most popular one.
Every claim ships with a falsifier
There is no follow-up conversation, so you cannot ask the report to clarify and it will never learn whether it was right. That means a claim which merely sounds plausible is, to you, indistinguishable from one that is true.
So configural claims have to state what would disconfirm them — “your lapses should cluster in disrupted-routine weeks rather than high-pressure ones; check that against last year”. You validate against your life rather than against how flattering it feels. The skeleton the report is written to requires it; it is not advice in the margin.
Sensitive results
If the Depression or Vulnerability facet lands at or above the 85th percentile, any other Neuroticism facet at the 90th, or Neuroticism overall at the 90th, the report is routed to a gentler template. That template names the pattern, bans clinical vocabulary outright, refuses to speculate about causes, and points toward talking to a person. It adds constraints and never relaxes any — including the falsifier requirement, because vagueness is more harmful there, not less. The ban is the one thing not left to instruction: a report on this path is checked in code for clinical vocabulary, and one that used any is not delivered at all.
A personality questionnaire is not a screening instrument and cannot detect a condition.