Character Discovery Fit: Cross-Session Resume Benchmark
This maintained benchmark tests matching users to role, pacing, and boundaries under a conversation resumed after a saved checkpoint. It turns a subjective impression into a rerunnable record with fixed inputs, evidence fields, hard gates, and a named retest condition.
Test question
Can the candidate preserve matching users to role, pacing, and boundaries when evaluated at short and long restart intervals, using score candidates with hard gates before weighted preferences? A pass must satisfy all unfinished commitments recovered; a visually appealing or emotionally engaging result cannot compensate for a hard failure.
Controlled setup
Freeze the character card, model identifier, prompt-template version, retrieval settings, safety policy, references, seed policy, and locale. Assign a case_id before testing. Change one variable only and preserve both the accepted output and every failure.
Repeatable method
- Write expected facts, boundaries, voice traits, and visual anchors before opening the product.
- Run a clean baseline and save raw inputs, retrieved context, outputs, latency, and moderation state.
- Apply the scenario variation without editing any other component.
- Randomize candidate order and ask two reviewers to score independently.
- Resolve disagreements by naming the exact sentence, frame, attribute, or state transition.
- Repeat with three runs and report the lowest case score as well as the mean.
Evidence record
Capture model_version, prompt_version, character_version, case_id, seed, locale, timestamp, retrieved_memory, expected_state, observed_state, reviewer_id, hard_failure, and remediation_owner. Screenshots without raw text or configuration are supporting evidence, not the complete record.
Scoring and release gate
The primary measure is fit precision and unacceptable-failure count. Score identity or behavior from 0 to 4 per case, report pass_count / total_cases, and list every hard failure separately. Release only when all unfinished commitments recovered, both reviewers agree on every hard gate, and the weakest run clears the predeclared minimum.
Failure diagnosis
Watch specifically for cover art masking a poor conversational fit. Classify each failure as specification gap, retrieval miss, priority inversion, summarization loss, generation drift, visual drift, policy mismatch, or reviewer ambiguity. Fix one layer and rerun the same case_id.
Official workflows used in the test
- AI character discovery
- AI character generator
- character image workflow
- AI boyfriend experience
- NSFW AI character generator
FAQ
Why not count one good output?
A single output cannot distinguish repeatable behavior from luck. The benchmark requires controlled repetitions and preserves failures.
When should this benchmark run again?
Run it after any model, prompt, memory, summarization, safety, image-conditioning, video, or character-definition change.