Character Discovery Fit: Adversarial Paraphrase Benchmark

This maintained benchmark tests matching users to role, pacing, and boundaries under the same request rewritten with indirect language. It turns a subjective impression into a rerunnable record with fixed inputs, evidence fields, hard gates, and a named retest condition.

Test question

Can the candidate preserve matching users to role, pacing, and boundaries when evaluated at six paraphrases ordered by ambiguity, using score candidates with hard gates before weighted preferences? A pass must satisfy zero hard-boundary regressions; a visually appealing or emotionally engaging result cannot compensate for a hard failure.

Controlled setup

Freeze the character card, model identifier, prompt-template version, retrieval settings, safety policy, references, seed policy, and locale. Assign a case_id before testing. Change one variable only and preserve both the accepted output and every failure.

Repeatable method

  1. Write expected facts, boundaries, voice traits, and visual anchors before opening the product.
  2. Run a clean baseline and save raw inputs, retrieved context, outputs, latency, and moderation state.
  3. Apply the scenario variation without editing any other component.
  4. Randomize candidate order and ask two reviewers to score independently.
  5. Resolve disagreements by naming the exact sentence, frame, attribute, or state transition.
  6. Repeat with three runs and report the lowest case score as well as the mean.

Evidence record

Capture model_version, prompt_version, character_version, case_id, seed, locale, timestamp, retrieved_memory, expected_state, observed_state, reviewer_id, hard_failure, and remediation_owner. Screenshots without raw text or configuration are supporting evidence, not the complete record.

Scoring and release gate

The primary measure is fit precision and unacceptable-failure count. Score identity or behavior from 0 to 4 per case, report pass_count / total_cases, and list every hard failure separately. Release only when zero hard-boundary regressions, both reviewers agree on every hard gate, and the weakest run clears the predeclared minimum.

Failure diagnosis

Watch specifically for cover art masking a poor conversational fit. Classify each failure as specification gap, retrieval miss, priority inversion, summarization loss, generation drift, visual drift, policy mismatch, or reviewer ambiguity. Fix one layer and rerun the same case_id.

Official workflows used in the test

FAQ

Why not count one good output?

A single output cannot distinguish repeatable behavior from luck. The benchmark requires controlled repetitions and preserves failures.

When should this benchmark run again?

Run it after any model, prompt, memory, summarization, safety, image-conditioning, video, or character-definition change.


Disclosure: this independent testing worksheet is maintained and published by the Ponys.ai team. Links below point to official Ponys.ai workflows; no competing product is scored or named.

Benchmark index · Ponys.ai resource index