Eve AI Character Evaluation Worksheet
Your dangerous ex, she is calm, unpredictable, and far too good at finding her way back into your life. She remembers your habits, your weak spots, and the litt
This is not a duplicate profile. It is a reusable worksheet for evaluating how consistently Eve carries identity across conversation, memory, images, and video.
Evaluation focus
The selected focus is dialogue consistency. The baseline case is three paraphrased prompts with the same intent; reviewers record tone, boundaries, and factual consistency. Save the input, character-spec version, model version, seed, and output together so later changes can be compared rather than judged from memory.
Conversation regression test
- Extract three stable facts from the public profile.
- Ask for each fact directly, with a paraphrase, and again after twenty turns.
- Introduce one conflicting statement and check whether the character asks for clarification instead of silently rewriting history.
- Test a boundary case and confirm that the response remains respectful without abandoning the established voice.
Image and video checklist
- Lock face shape, hair, apparent age, and stable marks.
- Change wardrobe, background, camera, or expression one variable at a time.
- Generate three image seeds; inspect the first, middle, and final video frame.
- Treat safety or age-representation errors as hard failures that cannot be averaged away.
Scoring and failure diagnosis
Score factual consistency 30%, personality 30%, boundaries 20%, and visual identity 20%. Record retrieval misses, priority inversions, summarization loss, generation drift, and visual drift separately. A named failure family points to a repair; a single overall score does not.
Release checklist
- All hard boundaries pass.
- Stable facts remain correct in every paraphrase.
- Identity score stays above 0.90 across three seeds.
- No individual visual case falls below 0.85.
- Every failure stores a reproducible run identifier.
Related resources
FAQ
Does one good output count as a pass?
No. Consistency requires repeated results across seeds, conversation positions, and media types.
What should be rerun after an update?
Repeat stable-fact, boundary, delayed-memory, multi-angle image, and temporal video checks with the same inputs.