Introducing Thymos Myrias: Frontier human simulation at the speed of thought.
220 held-out people · three public datasets
Thymos1 Myrias2 is now the cognition model behind every simulated person on Aetherya, read directly over a weighted synthetic population. Accuracy here means one specific thing: the share of the audience the simulation said would pick an option, against the share of real people who did. Measured on people the model was never fitted to, in situations those people had never seen, across three public datasets.
We believe Myrias is the best model in the world at simulating human behaviour. Everything below is the evidence for it.
The state of the art on speed and accuracy together
That accuracy runs at 3.7 microseconds a decision — 4.5 million a second in bulk. Identifying a new person from their answers takes 19 milliseconds, and fitting an entire population takes 3.07 seconds. No other model that simulates human behaviour this accurately runs this fast, and nothing that runs this fast reaches this accuracy — Myrias is the current state of the art on both axes of the report, not one bought by trading away the other. That is what makes re-running a study a thousand times an ordinary thing to do rather than a rare one.
Against what you would otherwise use
Same people, same decisions, same held-out situations. The error is the gap between the share the model said would pick an option and the share that really did.
When it says 70%, does 70% happen?
A simulated audience is only useful if its confidence means something. Group every decision by how likely the model said it was, then count how often it actually happened.
Three datasets, two of them tasks it was never built for
| Dataset | People | Decisions | Share accuracy | Decisions right | Calibration |
|---|---|---|---|---|---|
| CPC2015 — the dataset the method paper uses | 80 | 1,920 | 93.8% | 85.8% | 0.94 |
| Psych-101 — the same paradigm, different people | 60 | 9,000 | 96.7% | 83.5% | 1.00 |
| Psych-101 — a two-armed bandit, a different task entirely | 80 | 8,405 | — | 77.7% | 1.04 |
The bandit has no audience figure because every game in it is unique to one person, so there is no shared situation to predict a share for. On that task the naive rule falls to 62.3% with a calibration slope of 0.23, which makes its confidence close to meaningless, while Myrias holds 77.7% at 1.04.
Where modelling the individual does not help
On a densely sampled situation a population average scores the same as modelling each person: 96.7% against 96.7% on the larger dataset. Individual modelling earns its place on sparse data and on questions about particular people, which is where the gap opens to 92.0% against 93.8%, and where an audience of averages starts getting the crosstabs wrong.
There is also a floor on how little we can know about someone. Below roughly six observed situations per person, identifying that person predicts them worse than not identifying them at all. Under that threshold we fall back to the population rather than pretending to know the person, and the result says which one it used.
What runs in the product
Since 2026-09-14 the live model lets a persona’s state change which option it picks, not only whether it disengages. Scored once on 40 sealed people and 18,050 decisions it had never seen, the confidence it places on the choice a person really made rose from 68.3% to 69.7%, with the interval on the difference excluding zero. It behaves exactly as before at a persona’s starting state. Not yet measured on web and chat surfaces; it is fitted on a decision task.
How these numbers stay honest
- Public datasets anyone can download, hash-verified. Split rules and seeds published. Our confirmatory set has never been opened.
- Every figure is measured on people the model was never fitted to, in situations those people never saw. Intervals are participant-clustered bootstraps.
- 204 proposed improvements were registered and hashed before being fitted, each getting one sealed test. The rejections are kept alongside the wins.
- Where a simpler method does as well, this page says so. Nothing here is a claim about any real person’s mind.
- 1. Thymos n.
- The architecture — the cognition system itself, not any one model built on it. Every simulated person on Aetherya runs on a model built on Thymos.
- Gk. θυμός (thȳmós) — spirit; the seat of emotion, will, and vital force. In Homeric psychology, the part of the mind that drives desire and the impulse to act.
- 2. Myrias n.
- A model built on the Thymos architecture — the newest, and the best in class it has produced so far.
- Gk. μυριάς (myriás) — ten thousand; a vast, countless number. Root of “myriad.”
Source: Thymos: A State-Conditioned Architecture for Behavioral Simulation with Language Models, Dan Andrei, Aetherya, September 2026. Data: CPC2015 (Erev et al., 2017) and Psych-101 (Binz et al., 2024). Every figure on this page is read from a generated file that names the run report it came from.
