Building an Image Identity Baseline
This maintained worksheet focuses on stable facial and age attributes across scenes. It is intended for creators and reviewers who need a result they can rerun, compare, and explain rather than a single attractive demo.
Test question
Write one falsifiable question before generating anything: can the candidate preserve the required behavior when wording, conversation position, scene, or motion changes? The primary method for this page is to compare front, three-quarter, and full-body views over three seeds. Freeze the character specification, model, prompt-template version, safety policy, references, and random seed before comparing outputs.
Evidence to capture
Save the complete input, the retrieved memories, the response or media output, execution time, and reviewer notes. The main measurement is identity similarity and attribute preservation. Use at least three repetitions and keep individual-case scores; an average alone can hide a severe failure.
Review protocol
- Define stable facts, relationship boundaries, and visual anchors.
- Run a baseline without changing more than one variable.
- Randomize old and new outputs so reviewers do not know which is newer.
- Have two reviewers score independently.
- Record the exact attribute behind every disagreement.
Failure diagnosis
The failure to watch for is wardrobe or camera changes leaking into core identity. Classify the cause as a specification gap, retrieval miss, priority inversion, summarization loss, generation drift, or visual drift. Repair one layer at a time and rerun the same case identifier.
Release rule
Pass only when every hard boundary succeeds, stable facts remain correct, and repeated runs preserve identity. A conditional pass needs an owner and retest date. Never delete a difficult case merely to improve the pass rate.
Related Ponys.ai workflows
- AI character generator
- Discover AI characters
- Create an AI character
- AI image generator
- AI video generator
- Character image workflow
FAQ
Does one successful output count?
No. A useful result must repeat across seeds, positions, or scenes.
When should this test run again?
Repeat it after changes to the model, memory retrieval, summarizer, prompt template, image conditioning, video pipeline, safety policy, or character definition.