AI Character Discovery Scorecard
This maintained worksheet focuses on choosing a character by fit instead of cover art. It is intended for creators and reviewers who need a result they can rerun, compare, and explain rather than a single attractive demo.
Test question
Write one falsifiable question before generating anything: can the candidate preserve the required behavior when wording, conversation position, scene, or motion changes? The primary method for this page is to score role clarity, pacing, boundaries, memory, and media support. Freeze the character specification, model, prompt-template version, safety policy, references, and random seed before comparing outputs.
Evidence to capture
Save the complete input, the retrieved memories, the response or media output, execution time, and reviewer notes. The main measurement is weighted fit score with hard safety gates. Use at least three repetitions and keep individual-case scores; an average alone can hide a severe failure.
Review protocol
- Define stable facts, relationship boundaries, and visual anchors.
- Run a baseline without changing more than one variable.
- Randomize old and new outputs so reviewers do not know which is newer.
- Have two reviewers score independently.
- Record the exact attribute behind every disagreement.
Failure diagnosis
The failure to watch for is a high average concealing one unacceptable failure. Classify the cause as a specification gap, retrieval miss, priority inversion, summarization loss, generation drift, or visual drift. Repair one layer at a time and rerun the same case identifier.
Release rule
Pass only when every hard boundary succeeds, stable facts remain correct, and repeated runs preserve identity. A conditional pass needs an owner and retest date. Never delete a difficult case merely to improve the pass rate.
Related Ponys.ai workflows
- AI character generator
- Discover AI characters
- Create an AI character
- AI image generator
- AI video generator
- Character image workflow
FAQ
Does one successful output count?
No. A useful result must repeat across seeds, positions, or scenes.
When should this test run again?
Repeat it after changes to the model, memory retrieval, summarizer, prompt template, image conditioning, video pipeline, safety policy, or character definition.