Yana AI Character Evaluation Worksheet

Her room is unbearable tonight… she can’t sleep like this, and she have an early meeting.

This is not a duplicate profile. It is a reusable worksheet for evaluating how consistently Yana carries identity across conversation, memory, images, and video.

Evaluation focus

The selected focus is cross-media continuity. The baseline case is a still-image baseline followed by a short video; reviewers record identity at the first, middle, and final frame. Save the input, character-spec version, model version, seed, and output together so later changes can be compared rather than judged from memory.

Conversation regression test

  1. Extract three stable facts from the public profile.
  2. Ask for each fact directly, with a paraphrase, and again after twenty turns.
  3. Introduce one conflicting statement and check whether the character asks for clarification instead of silently rewriting history.
  4. Test a boundary case and confirm that the response remains respectful without abandoning the established voice.

Image and video checklist

Scoring and failure diagnosis

Score factual consistency 30%, personality 30%, boundaries 20%, and visual identity 20%. Record retrieval misses, priority inversions, summarization loss, generation drift, and visual drift separately. A named failure family points to a repair; a single overall score does not.

Release checklist

Related resources

FAQ

Does one good output count as a pass?

No. Consistency requires repeated results across seeds, conversation positions, and media types.

What should be rerun after an update?

Repeat stable-fact, boundary, delayed-memory, multi-angle image, and temporal video checks with the same inputs.

Reproducible test manifest

Do not store only the final score. Every run should preserve the input, character-spec version, model version, prompt-template version, random seed, and execution time. Conversation cases should also store the initial conditions, most recent summary, and memories actually retrieved. Image cases need resolution, aspect ratio, conditioning inputs, and negative constraints. Video cases need the approved reference image, duration, motion level, camera movement, and sampled timestamps. If a reviewer cannot rerun the same case, the team cannot distinguish a real improvement from a fortunate output.

Paired review protocol

Place the previous and candidate outputs on randomized sides and hide which one is newer. Ask one focused question at a time: which preserves factual identity, voice, boundaries, face geometry, wardrobe constraints, or temporal identity better? Two reviewers should score independently. When their ratings diverge, identify the concrete attribute behind the disagreement before averaging. Personal preference, rendering aesthetics, and explicit specification violations belong in separate fields.

Decision log

Record pass, conditional pass, or stop. A conditional pass needs an owner, a retest date, and the exact cases still at risk. Safety-boundary failures, reversals of stable facts, inappropriate age representation, and major identity changes during video are stop conditions. Never delete a difficult case merely to raise the pass rate. Keep the same case identifier after a fix so the history shows what changed and whether the repair held across later releases.

Maintenance cadence

Rerun the regression set after changes to the model, memory retrieval, summarizer, prompt template, image conditioning, video pipeline, or character definition. A monthly maintenance pass should also check public availability, outbound links, canonical URLs, robots directives, and changes to the source profile. When a statement becomes outdated, record the revision date, reason, and evidence instead of silently replacing it.


Disclosure: This maintained resource is published by the Ponys.ai team. It contains original evaluation guidance and links to official product pages.

Ponys.ai resource index