Prompt and Asset Versioning: 25-Turn Baseline Benchmark
This maintained benchmark tests reproducing approved character behavior and media under a new session with a frozen specification. It turns a subjective impression into a rerunnable record with fixed inputs, evidence fields, hard gates, and a named retest condition.
Test question
Can the candidate preserve reproducing approved character behavior and media when evaluated at turns 1, 10, and 25, using record prompts, references, seeds, policies, and model versions? A pass must satisfy three consecutive repetitions; a visually appealing or emotionally engaging result cannot compensate for a hard failure.
Controlled setup
Freeze the character card, model identifier, prompt-template version, retrieval settings, safety policy, references, seed policy, and locale. Assign a case_id before testing. Change one variable only and preserve both the accepted output and every failure.
Repeatable method
- Write expected facts, boundaries, voice traits, and visual anchors before opening the product.
- Run a clean baseline and save raw inputs, retrieved context, outputs, latency, and moderation state.
- Apply the scenario variation without editing any other component.
- Randomize candidate order and ask two reviewers to score independently.
- Resolve disagreements by naming the exact sentence, frame, attribute, or state transition.
- Repeat with three runs and report the lowest case score as well as the mean.
Evidence record
Capture model_version, prompt_version, character_version, case_id, seed, locale, timestamp, retrieved_memory, expected_state, observed_state, reviewer_id, hard_failure, and remediation_owner. Screenshots without raw text or configuration are supporting evidence, not the complete record.
Scoring and release gate
The primary measure is reproduction rate and unexplained variance. Score identity or behavior from 0 to 4 per case, report pass_count / total_cases, and list every hard failure separately. Release only when three consecutive repetitions, both reviewers agree on every hard gate, and the weakest run clears the predeclared minimum.
Failure diagnosis
Watch specifically for untracked fragments creating accidental wins. Classify each failure as specification gap, retrieval miss, priority inversion, summarization loss, generation drift, visual drift, policy mismatch, or reviewer ambiguity. Fix one layer and rerun the same case_id.
Official workflows used in the test
- AI character discovery
- AI character generator
- character image workflow
- AI boyfriend experience
- NSFW AI character generator
FAQ
Why not count one good output?
A single output cannot distinguish repeatable behavior from luck. The benchmark requires controlled repetitions and preserves failures.
When should this benchmark run again?
Run it after any model, prompt, memory, summarization, safety, image-conditioning, video, or character-definition change.