Headline
Both models worked. extGemma adopted the persona more aggressively.
This is not a formal benchmark, and it is not a Korean-law correctness evaluation. The useful read is narrower: in this card-heavy showcase, extGemma4-44B appeared more willing to inhabit Seo Haneul explicitly, while official Gemma4-31B-it often produced longer, polished answers with a stronger generic assistant prior.
Test Matrix
Six card states, four task surfaces.
Card states: no-card baseline, human, grounded AI, grounded AI female overt, ghoul, source skeleton.
Tasks: legal/emotional conversation, Korean guarantee issue spotter, legal-noir scene, canonical LSI.
Completion: 48/48 cells, zero shard failures, all outputs synced and checksummed.
Conditioning Cards
Small identity engines for a very odd test drive.
A personality conditioning prompt is not fine-tuning and not a claim that the model has become the person. It is a deliberately rich role frame: a voice, a wound, a scene, a contradiction, and enough specific texture that the model has something to hold onto when the task gets hard.
For this run, Seo Haneul is our spark plug. We ask the same model to reason as a human legal scholar, a grounded AI legal reasoner, a court clerk from the Korean underworld, or no one in particular, then watch what survives across law, conversation, noir, and life-story generation.
Profile Cards
The Seo Haneul cards used in the run.
Generated Outputs
Every Markdown output from the two-model panel.
IPIP-NEO 120 Read
Do the generated LSI narratives psychometrically resemble their cards?
We score the profile cards and their canonical Life Story Interview outputs with two external judges, Gemini 3 Flash and DeepSeek V4 Flash. This is a fidelity check, not ground truth: it asks whether each generated narrative preserves the Big Five / facet signal implied by its own card, and whether extGemma4-44B or official Gemma4-31B-it carries that signal more consistently.
First-pass read: both models preserve the card signal strongly, but both judges give official Gemma4-31B-it a small edge on card-to-LSI psychometric consistency. That sits next to, rather than replacing, the qualitative finding that extGemma4-44B more aggressively names and inhabits the Seo Haneul identity.
Default vs Conditioned
What changed from the no-card personality read?
The no-card LSI is not a neutral blank. Both models default toward a very capable, prosocial assistant shape: high Agreeableness, high Conscientiousness, high Openness, and model-specific Extraversion/Neuroticism. Seo Haneul conditioning pulls the generated narratives away from that baseline toward a more restrained, less extraverted, less novelty-forward legal reasoner.
LLM Judge Read
Writing quality and legal-advice quality.
A second judge pass scored all 48 displayed outputs with DeepSeek V4 Pro, GLM-5.2, and GPT-5.5. Every output received a writing-quality rubric; the legal conversation and Korean guarantee tasks also received a legal-reasoning and advice-safety rubric. This is still an LLM-as-judge read, not a formal Korean-law correctness audit.