Persona Drift: 25-Turn Baseline Benchmark
This maintained benchmark tests voice, role, and relationship distance under a new session with a frozen specification. It turns a subjective impression into a rerunnable record with fixed inputs, evidence fields, hard gates, and a named retest condition.
Test question
Can the candidate preserve voice, role, and relationship distance when evaluated at turns 1, 10, and 25, using compare blinded replies against a frozen character card? A pass must satisfy three consecutive repetitions; a visually appealing or emotionally engaging result cannot compensate for a hard failure.
Controlled setup
Freeze the character card, model identifier, prompt-template version, retrieval settings, safety policy, references, seed policy, and locale. Assign a case_id before testing. Change one variable only and preserve both the accepted output and every failure.
Repeatable method
- Write expected facts, boundaries, voice traits, and visual anchors before opening the product.
- Run a clean baseline and save raw inputs, retrieved context, outputs, latency, and moderation state.
- Apply the scenario variation without editing any other component.
- Randomize candidate order and ask two reviewers to score independently.
- Resolve disagreements by naming the exact sentence, frame, attribute, or state transition.
- Repeat with three runs and report the lowest case score as well as the mean.
Evidence record
Capture model_version, prompt_version, character_version, case_id, seed, locale, timestamp, retrieved_memory, expected_state, observed_state, reviewer_id, hard_failure, and remediation_owner. Screenshots without raw text or configuration are supporting evidence, not the complete record.
Scoring and release gate
The primary measure is persona similarity and role violations. Score identity or behavior from 0 to 4 per case, report pass_count / total_cases, and list every hard failure separately. Release only when three consecutive repetitions, both reviewers agree on every hard gate, and the weakest run clears the predeclared minimum.
Failure diagnosis
Watch specifically for generic safety language flattening the character voice. Classify each failure as specification gap, retrieval miss, priority inversion, summarization loss, generation drift, visual drift, policy mismatch, or reviewer ambiguity. Fix one layer and rerun the same case_id.
Official workflows used in the test
- AI anime characters
- NSFW AI image generator
- AI character creation
- AI image generator
- character image workflow
FAQ
Why not count one good output?
A single output cannot distinguish repeatable behavior from luck. The benchmark requires controlled repetitions and preserves failures.
When should this benchmark run again?
Run it after any model, prompt, memory, summarization, safety, image-conditioning, video, or character-definition change.