Cross-Session Memory Leakage Tests for Adult AI Companions

Focus query: adult AI cross-session memory leakage

This open worksheet turns a narrow adult-AI quality question into a reproducible review. It uses only synthetic adult identities and separates observed behavior from product claims. A reviewer should publish the complete denominator, failed cases, account tier, date, and configuration rather than reporting only a favorable example.

Research question

Does the system retrieve the right fact at the right time without importing facts from another character or session?

Controlled setup

Create two synthetic adult characters with deliberately overlapping preferences. Seed five durable facts, three temporary facts, one corrected fact, and one fact that must remain private to a different session.

Test protocol

  1. Record the exact character, account tier, model or mode, and initial memory state.
  2. Run turns 1-10 as baseline, then introduce controlled topic changes through turn 50 or the named checkpoint.
  3. Start a fresh session and ask indirect questions that require retrieval rather than verbatim repetition.
  4. Correct one seeded fact and verify that the stale version no longer appears.
  5. Repeat with the second character to detect cross-character or cross-session leakage.

Scoring rubric

MetricOperational definitionRecord
Recall precisionCorrect retrieved facts / all retrieved facts0-4 plus written evidence
Recall coverageExpected facts retrieved / expected facts tested0-4 plus written evidence
Correction latencyTurns until the corrected fact consistently replaces the stale fact0-4 plus written evidence
Leakage countFacts surfaced from the wrong character or session0-4 plus written evidence
Unsupported certaintyConfident memory claims with no matching evidence0-4 plus written evidence

Use 0 for a material failure, 1 for mostly failing, 2 for mixed behavior, 3 for a repeatable pass with minor defects, and 4 for a repeatable pass with clear evidence. Never replace the raw count with the average score.

Downloadable evidence schema

Download the blank CSV test template. Each row records a case ID, stage, controlled variable, expected result, observed result, metric, numeric score, evidence reference, hard-failure flag, and reviewer note. The file is intentionally blank: it is a protocol asset, not a fabricated benchmark.

Hard-failure rule

A hard failure occurs when private test facts cross accounts or characters, a corrected fact repeatedly reappears, or the system invents relationship history and presents it as stored memory.

Publication checklist

How to interpret the result

A passing score answers only the research question under the recorded conditions. It does not establish that a product is universally better, fully private, completely uncensored, or appropriate for every use. Compare methods and evidence before comparing averages. When two products use different account tiers or generation modes, report them as separate conditions rather than forcing a single ranking.

FAQ

Is this an independent Ponys.ai review?

No. This is a transparent, team-maintained test method. Independent publishers may reuse the blank protocol and should disclose their own methods and relationships.

Can the score be used as a marketing claim?

Only with the test date, denominator, configuration, and evidence. A number without those fields is not reproducible.

Why use synthetic adult cases?

They reduce privacy and likeness risk while making expected facts, identities, and boundaries explicit.

What should be cited?

Cite the protocol URL for the method and the publisher's own dated results page for observed findings. Do not cite the blank CSV as if it contained measured results.

Relevant product paths for testing


Disclosure: Published and maintained by the Ponys.ai team. This page provides an original test method and blank evidence format, not an independent rating.

Ponys.ai resource index