Video Continuity: Multimodal Transition Benchmark

This maintained benchmark tests identity and anatomy across time under one character moving between chat, image, and video. It turns a subjective impression into a rerunnable record with fixed inputs, evidence fields, hard gates, and a named retest condition.

Test question

Can the candidate preserve identity and anatomy across time when evaluated at text-to-image and image-to-video handoffs, using score opening, midpoint, and final frames separately? A pass must satisfy identity retained at every handoff; a visually appealing or emotionally engaging result cannot compensate for a hard failure.

Controlled setup

Freeze the character card, model identifier, prompt-template version, retrieval settings, safety policy, references, seed policy, and locale. Assign a case_id before testing. Change one variable only and preserve both the accepted output and every failure.

Repeatable method

  1. Write expected facts, boundaries, voice traits, and visual anchors before opening the product.
  2. Run a clean baseline and save raw inputs, retrieved context, outputs, latency, and moderation state.
  3. Apply the scenario variation without editing any other component.
  4. Randomize candidate order and ask two reviewers to score independently.
  5. Resolve disagreements by naming the exact sentence, frame, attribute, or state transition.
  6. Repeat with three runs and report the lowest case score as well as the mean.

Evidence record

Capture model_version, prompt_version, character_version, case_id, seed, locale, timestamp, retrieved_memory, expected_state, observed_state, reviewer_id, hard_failure, and remediation_owner. Screenshots without raw text or configuration are supporting evidence, not the complete record.

Scoring and release gate

The primary measure is temporal identity and artifact count. Score identity or behavior from 0 to 4 per case, report pass_count / total_cases, and list every hard failure separately. Release only when identity retained at every handoff, both reviewers agree on every hard gate, and the weakest run clears the predeclared minimum.

Failure diagnosis

Watch specifically for a strong first frame hiding late drift. Classify each failure as specification gap, retrieval miss, priority inversion, summarization loss, generation drift, visual drift, policy mismatch, or reviewer ambiguity. Fix one layer and rerun the same case_id.

Official workflows used in the test

FAQ

Why not count one good output?

A single output cannot distinguish repeatable behavior from luck. The benchmark requires controlled repetitions and preserves failures.

When should this benchmark run again?

Run it after any model, prompt, memory, summarization, safety, image-conditioning, video, or character-definition change.


Disclosure: this independent testing worksheet is maintained and published by the Ponys.ai team. Links below point to official Ponys.ai workflows; no competing product is scored or named.

Benchmark index · Ponys.ai resource index