OREXRequest access
ὄρεξις · the reaching toward · live

REX

What people actually want.
Measured, not asked.

Human preference and social calibration, captured continuously from three hundred thousand people a month who are not paid, not recruited, and not performing. Priced by the judgment. Weighted by the judge.

01  26,000+ judgments a day02  300K monthly participants03  59 timezones04  7 world regions05  consented at capture
·Reaching toward·Appetite·Preference·Calibration·Tolerance·Consented at capture·26,000 a day·Unperformed
·Reaching toward·Appetite·Preference·Calibration·Tolerance·Consented at capture·26,000 a day·Unperformed
·Reaching toward·Appetite·Preference·Calibration·Tolerance·Consented at capture·26,000 a day·Unperformed
·Reaching toward·Appetite·Preference·Calibration·Tolerance·Consented at capture·26,000 a day·Unperformed
·Reaching toward·Appetite·Preference·Calibration·Tolerance·Consented at capture·26,000 a day·Unperformed
·Reaching toward·Appetite·Preference·Calibration·Tolerance·Consented at capture·26,000 a day·Unperformed
·Reaching toward·Appetite·Preference·Calibration·Tolerance·Consented at capture·26,000 a day·Unperformed
·Reaching toward·Appetite·Preference·Calibration·Tolerance·Consented at capture·26,000 a day·Unperformed
·Reaching toward·Appetite·Preference·Calibration·Tolerance·Consented at capture·26,000 a day·Unperformed
01 / The finding
2.5×

The best readers of other people are two and a half times as accurate as the worst. Preference data prices them the same.

Models are aligned to "human preference" scored by small pools of paid raters, treated as interchangeable. They are not. On a 300,000-prediction sample, scored on whether one person could predict another's choice, accuracy ran from a quarter to nearly two thirds. The spread held at three sample sizes as the sample doubled.

Accuracy by annotator · 20+ scoredn = 5,600
25%p10
43%median
63%p90
0chance ≈ 25 to 50, by table size100
02 / What you leave with
01

Preference, with the loser recorded

Not just what a person chose. What it beat, from the same hand, in the same moment. Most of the information in a choice lives in the options that lost, and we keep them.

02

A quality score for every annotator

How well a person reads other people, measured on a separate task with a verifiable answer. Weight every judgment by the judge. Nobody else has a second instrument to check the first against.

03

Where the line sits

Tolerance, tone and timing. How far a person will go, how far a room will let them, and the gap between what someone enjoys and what they will admit to. Two numbers, both held.

03 / Why it cannot be scraped

Four conditions. No dataset on the market satisfies all of them.

Unpaid

Nobody is compensated for an opinion. There is no incentive to produce one, so what comes out is the one they had.

Unrecruited

No panel, no application, no vetting by other vetted people. Participants arrived on their own and stayed on their own.

Unperformed

Every judgment is made in front of people the participant knows, for stakes they care about. Nobody performs a laugh for a researcher.

Consented

Consent is asked at the moment of capture, scoped and revocable. Records carry it with them. Nothing here was scraped.

04 / The corpusfigures are floors · live telemetry · two titles
26,000+
judgments a day
300K
monthly participants
1M+
human choices
59
timezones represented
7
world regions
0
questionnaires sent

Rates, not totals. The corpus refills every night whether or not anyone is buying it. The instrument is play: two titles, one of them with a verifiable answer key, which is what makes a calibration score possible at all.

05 / Access

We work with a small number of partners.

Model labs, frontier labs, and the research teams inside them. We publish neither our methods nor our map, and that is deliberate.

Request access