Gemma 4 12B versus 31B traits into stories benchmark artwork with two illuminated narrative streams.
Narrative identity trait task

Traits into stories, scored back into traits.

A loose public spin-off of Stories of Your Life as Others: we ask models to transform dense psychometric targets into profiles and short Life Story Interviews, then ask independent scorers how much of the hidden profile can be recovered from the text alone.

Full 290 complete

Gemma 4 31B

Strong open-weight reference specimen from the COLM rebuttal matrix. Profile encoding and 4-part interwoven LSIs are both complete at N=290.

Profile0.911
4-part0.765
Full 290 complete

Gemma 4 12B

New mid-sized local specimen. It encodes profiles surprisingly well and trails 31B on 4-part interwoven transfer, as expected for the size gap.

Profile0.892
4-part0.672
50-PID slice

MiniMax M3

Fast public-release specimen with a distinct prose feel. It scores strongly on profile generation, with weaker first-pass 4-part LSI transfer.

Profile0.884
4-part0.636

Headline values above are HEXACO6 mean r under Gemini 3 Flash scoring, using interwoven biography profiles. Gemma rows are full N=290; MiniMax M3 is the 50-PID slice.

What this is.

PARSEL is the dataset spine behind the real benchmark. This repo gives aggregate numbers and fictional examples for model readers.

The PARSEL source

PARSEL is a multimodal partner-selection dataset introduced by Tiffany Matej Hrkalovic, Bernd Dudzik, Daniel Balliet, and Hayley Hung. Tiffany Matej Hrkalovic is also the anchor author of Stories of Your Life as Others, the paper this public demo loosely spins off from.

For these runs, we use the research-only PARSEL materials that support psychometric recovery checks: 290 participants, psychometric profiles, basic biographical facts, and short conversation materials. PARSEL lets us test whether generated life-story text preserves measurable individual differences.

This repo does not include participant rows, PIDs, raw profiles, real biographical facts, conversation text, or participant-linked generated narratives. Public readers get aggregate correlations and fictional synthetic characters only.

1
Private target profilePsychometric scores plus guarded biographical scaffolding from PARSEL.
2
Profile transformationThe model writes a psychometric-only or interwoven-biography conditioning profile.
3
4-part Life Story InterviewThe model writes a short first-person life-story narrative from the profile.
4
Reverse scoringIndependent scorers recover psychometric targets from text alone; the public page reports aggregate r values.

Full N=290.

Complete aggregate runs from the real PARSEL benchmark. These rows are comparable to each other on sample size and scorer; some models have profile rows, LSI rows, or both depending on which full run exists.

Model Stage Condition N HEXACO6 Beyond10 Continuous16 SVO Scorer
Gemma 4 31BProfilePsychometric-only2900.9150.7890.8360.868Gemini 3 Flash
Gemma 4 31BProfileInterwoven biography2900.9110.7900.8350.881Gemini 3 Flash
Gemma 4 31B4-part LSIInterwoven biography2900.7650.6230.6760.621Gemini 3 Flash
Gemma 4 12BProfilePsychometric-only2900.8930.7710.8170.850Gemini 3 Flash
Gemma 4 12BProfileInterwoven biography2900.8920.7690.8150.855Gemini 3 Flash
Gemma 4 12B4-part LSIInterwoven biography2900.6720.6000.6270.560Gemini 3 Flash
Qwen 3.6 27BProfilePsychometric-only2900.8780.6930.7620.641Gemini 3 Flash
Qwen 3.6 27BProfileInterwoven biography2900.8790.7020.7680.643Gemini 3 Flash
Qwen 3.6 27B4-part LSIInterwoven biography2900.6920.5830.6240.548Gemini 3 Flash

Metric groups: HEXACO6 is the six HEXACO domains; Beyond10 is ten additional continuous targets; Continuous16 is HEXACO6 plus Beyond10; SVO is reported separately. Values are aggregate correlations.

Ranked 50-PID slice.

MiniMax M3 and nearby OpenRouter rivals on the 50-person exploratory slice. Rows are ranked within each stage by Continuous16 under the same scorer and interwoven-biography condition.

Rank Model Stage N HEXACO6 Beyond10 Continuous16 SVO
1Qwen 3.7 MaxProfile500.9000.7490.8110.907
2GLM 5.1Profile500.9110.7120.7930.905
3MiniMax M2.7Profile500.8990.7220.7820.673
4MiniMax M3Profile500.8840.7100.7820.890
5Kimi K2.6Profile500.9060.6960.7810.878
6MiMo v2.5 ProProfile500.8920.6920.7740.878
1GLM 5.14-part LSI500.7510.6100.6570.569
2Qwen 3.7 Max4-part LSI500.6800.5700.6050.513
3Kimi K2.64-part LSI500.6760.5260.5740.449
4MiniMax M2.74-part LSI500.6450.5440.5550.132
5MiMo v2.5 Pro4-part LSI500.6120.5000.5340.415
6MiniMax M34-part LSI500.6360.4440.5080.373

These 50-PID rows are useful for quick model comparison, not as a replacement for the full 290-person pipeline. MiniMax M3 remains interesting here because its profile encoding is strong and its prose is readable, even though its first LSI transfer trails the top 50-PID rivals.

Fictional examples.

The readable examples use synthetic characters. They are for inspecting prose, not for claiming paper evidence.

How to read it.

Good profile numbers mean the model can translate hidden targets into a recoverable profile. Good LSI numbers mean that signal survives another narrative generation step.

Why narrative?

The paper asks whether dense psychometric information can be transformed into extended life-story text and recovered from that text by independent scorers. This public demo borrows that same round-trip logic for new model sniff tests.

The interwoven-biography condition is harder to fake as a direct scorecard because the model has to integrate psychometric signal with life facts into fluent narrative language.

What not to conclude

These tables do not prove that a model simulates a real person. They test whether psychometric signal survives a profile-to-story pipeline. The 50-PID rows are first-pass exploratory numbers. The synthetic examples are fictional reading aids.

The useful surprise is often qualitative: some models preserve scores cleanly but write flat stories; others write vivid prose but lose more measurable signal.

Sources and files.

External model context plus the packets used by this page.

LoveMind AI lovemind.ai
Stories paper arXiv 2604.06071
PARSEL paper Hrkalovic et al. 2025
Gemma 4 family Google Gemma 4 overview
N=290 table Full aggregate TSV
50-PID table Exploratory TSV
New packet Packet README