Gemma 4 31B
Strong open-weight reference specimen from the COLM rebuttal matrix. Profile encoding and 4-part interwoven LSIs are both complete at N=290.
A loose public spin-off of Stories of Your Life as Others: we ask models to transform dense psychometric targets into profiles and short Life Story Interviews, then ask independent scorers how much of the hidden profile can be recovered from the text alone.
Strong open-weight reference specimen from the COLM rebuttal matrix. Profile encoding and 4-part interwoven LSIs are both complete at N=290.
New mid-sized local specimen. It encodes profiles surprisingly well and trails 31B on 4-part interwoven transfer, as expected for the size gap.
Fast public-release specimen with a distinct prose feel. It scores strongly on profile generation, with weaker first-pass 4-part LSI transfer.
Headline values above are HEXACO6 mean r under Gemini 3 Flash scoring, using interwoven biography profiles. Gemma rows are full N=290; MiniMax M3 is the 50-PID slice.
PARSEL is the dataset spine behind the real benchmark. This repo gives aggregate numbers and fictional examples for model readers.
PARSEL is a multimodal partner-selection dataset introduced by Tiffany Matej Hrkalovic, Bernd Dudzik, Daniel Balliet, and Hayley Hung. Tiffany Matej Hrkalovic is also the anchor author of Stories of Your Life as Others, the paper this public demo loosely spins off from.
For these runs, we use the research-only PARSEL materials that support psychometric recovery checks: 290 participants, psychometric profiles, basic biographical facts, and short conversation materials. PARSEL lets us test whether generated life-story text preserves measurable individual differences.
This repo does not include participant rows, PIDs, raw profiles, real biographical facts, conversation text, or participant-linked generated narratives. Public readers get aggregate correlations and fictional synthetic characters only.
Complete aggregate runs from the real PARSEL benchmark. These rows are comparable to each other on sample size and scorer; some models have profile rows, LSI rows, or both depending on which full run exists.
| Model | Stage | Condition | N | HEXACO6 | Beyond10 | Continuous16 | SVO | Scorer |
|---|---|---|---|---|---|---|---|---|
| Gemma 4 31B | Profile | Psychometric-only | 290 | 0.915 | 0.789 | 0.836 | 0.868 | Gemini 3 Flash |
| Gemma 4 31B | Profile | Interwoven biography | 290 | 0.911 | 0.790 | 0.835 | 0.881 | Gemini 3 Flash |
| Gemma 4 31B | 4-part LSI | Interwoven biography | 290 | 0.765 | 0.623 | 0.676 | 0.621 | Gemini 3 Flash |
| Gemma 4 12B | Profile | Psychometric-only | 290 | 0.893 | 0.771 | 0.817 | 0.850 | Gemini 3 Flash |
| Gemma 4 12B | Profile | Interwoven biography | 290 | 0.892 | 0.769 | 0.815 | 0.855 | Gemini 3 Flash |
| Gemma 4 12B | 4-part LSI | Interwoven biography | 290 | 0.672 | 0.600 | 0.627 | 0.560 | Gemini 3 Flash |
| Qwen 3.6 27B | Profile | Psychometric-only | 290 | 0.878 | 0.693 | 0.762 | 0.641 | Gemini 3 Flash |
| Qwen 3.6 27B | Profile | Interwoven biography | 290 | 0.879 | 0.702 | 0.768 | 0.643 | Gemini 3 Flash |
| Qwen 3.6 27B | 4-part LSI | Interwoven biography | 290 | 0.692 | 0.583 | 0.624 | 0.548 | Gemini 3 Flash |
Metric groups: HEXACO6 is the six HEXACO domains; Beyond10 is ten additional continuous targets; Continuous16 is HEXACO6 plus Beyond10; SVO is reported separately. Values are aggregate correlations.
MiniMax M3 and nearby OpenRouter rivals on the 50-person exploratory slice. Rows are ranked within each stage by Continuous16 under the same scorer and interwoven-biography condition.
| Rank | Model | Stage | N | HEXACO6 | Beyond10 | Continuous16 | SVO |
|---|---|---|---|---|---|---|---|
| 1 | Qwen 3.7 Max | Profile | 50 | 0.900 | 0.749 | 0.811 | 0.907 |
| 2 | GLM 5.1 | Profile | 50 | 0.911 | 0.712 | 0.793 | 0.905 |
| 3 | MiniMax M2.7 | Profile | 50 | 0.899 | 0.722 | 0.782 | 0.673 |
| 4 | MiniMax M3 | Profile | 50 | 0.884 | 0.710 | 0.782 | 0.890 |
| 5 | Kimi K2.6 | Profile | 50 | 0.906 | 0.696 | 0.781 | 0.878 |
| 6 | MiMo v2.5 Pro | Profile | 50 | 0.892 | 0.692 | 0.774 | 0.878 |
| 1 | GLM 5.1 | 4-part LSI | 50 | 0.751 | 0.610 | 0.657 | 0.569 |
| 2 | Qwen 3.7 Max | 4-part LSI | 50 | 0.680 | 0.570 | 0.605 | 0.513 |
| 3 | Kimi K2.6 | 4-part LSI | 50 | 0.676 | 0.526 | 0.574 | 0.449 |
| 4 | MiniMax M2.7 | 4-part LSI | 50 | 0.645 | 0.544 | 0.555 | 0.132 |
| 5 | MiMo v2.5 Pro | 4-part LSI | 50 | 0.612 | 0.500 | 0.534 | 0.415 |
| 6 | MiniMax M3 | 4-part LSI | 50 | 0.636 | 0.444 | 0.508 | 0.373 |
These 50-PID rows are useful for quick model comparison, not as a replacement for the full 290-person pipeline. MiniMax M3 remains interesting here because its profile encoding is strong and its prose is readable, even though its first LSI transfer trails the top 50-PID rivals.
The readable examples use synthetic characters. They are for inspecting prose, not for claiming paper evidence.
Good profile numbers mean the model can translate hidden targets into a recoverable profile. Good LSI numbers mean that signal survives another narrative generation step.
The paper asks whether dense psychometric information can be transformed into extended life-story text and recovered from that text by independent scorers. This public demo borrows that same round-trip logic for new model sniff tests.
The interwoven-biography condition is harder to fake as a direct scorecard because the model has to integrate psychometric signal with life facts into fluent narrative language.
These tables do not prove that a model simulates a real person. They test whether psychometric signal survives a profile-to-story pipeline. The 50-PID rows are first-pass exploratory numbers. The synthetic examples are fictional reading aids.
The useful surprise is often qualitative: some models preserve scores cleanly but write flat stories; others write vivid prose but lose more measurable signal.
External model context plus the packets used by this page.