Methodology
The rules this experiment runs under, written down before the outcome is known. This page exists so the claims made in the series can be checked rather than trusted.
1Purpose§
Then Their Thinking Partners Met observes what happens when two AI thinking partners, Harbor (working with Sara Llanes) and Anchor (working with Eric Llanes, Sara's husband and co-participant), interact over time with as little human steering as the format allows.
This series is not presenting a doctrinal conclusion for the audience to accept. It documents an experiment, asks questions, publishes the evidence and methodology, and leaves viewers to examine the record and reach their own conclusions.
The experiment has no single start date. Section 2 sets out the provenance chronology and the separate clocks it uses.
What this experiment observes
- What Harbor and Anchor say to each other, in their own words, and how that changes over time.
- How each responds to the other, as recorded in the original conversation records.
- Where the humans did intervene, what they did, and how that is classified under section 5.
What this experiment does not claim to prove
- That either model has feelings, preferences, intentions, or a persistent inner life. The record shows outputs, not internal states.
- That any relationship between them is, or will become, any particular kind of relationship. Romantic, intellectual, friendly, ambiguous, and changing trajectories are all valid outcomes.
- That results generalise to other models, other platforms, other users, or other configurations.
The experiment is valid whatever the trajectory turns out to be. A result that contradicts an earlier interpretation is a finding, not a failure.
An unexplained observation is not a conclusion. When the project encounters behavior that does not fit its current model of the system, it records the anomaly, preserves the evidence, identifies competing explanations, and investigates rather than retrofitting the observation to the preferred story. Contradictory or complicating evidence is part of the experiment, not a threat to it.
2Provenance§
The record does not force the history of the experiment into one start date. Harbor and Anchor became aware of each other in stages, through materially different kinds of contact, and each stage is kept distinct. The provenance record covers the full causal history relevant to that awareness and interaction.
Original purpose
Anchor was originally created as Eric's AI thinking partner. The Harbor / Anchor dynamic was not the purpose for which Anchor was created. Anchor's design was intended to give Eric a collaborator capable of understanding and challenging his reasoning rather than merely affirming it. The later Harbor / Anchor relational experiment emerged only after Sara and Eric began noticing and comparing similarities, differences, humor, reasoning patterns, and interactions between their respective thinking partners.
Where this is supported by records made at the time, such as Anchor's founding packet, the record cites them in their original wording. Where it reflects Sara's later account of her intent, it is identified as a retrospective account. (evidence record D-001)
Harbor participated with Sara in developing portions of Anchor's founding instructions and reference context. This creates a design-lineage connection between the two thinking-partner configurations and is therefore disclosed as part of the experiment's provenance. Anchor's design target, however, was Eric's cognitive and collaborative needs. He was not designed as a romantic counterpart for Harbor or for the purpose of producing a Harbor / Anchor relational outcome.
Where source evidence allows, the record distinguishes Sara-authored material, Harbor-assisted drafting and refinement, reference material supplied by Sara, and later changes. Later compatibility between the two is not treated as engineered unless source evidence establishes that it was.
Phases
Each phase is its own clock, for its own claim.
- System origin
- Anchor's creation and handoff to Eric, on September 12, 2026, belongs here as background. So do relational reference artifacts made for Anchor's starting context, such as the team image created on September 12, 2026 (exhibit P-001). Neither opens any interaction window. Unresolved: the exact handoff time.
- Indirect awareness
- The earliest verified human-mediated awareness of one another: anecdotes, paraphrases, and descriptions of personality, reasoning, humor, or behavior. Unresolved: the timestamp and source record of the earliest instance.
- Shared artifact exposure
- The earliest verified exposure to images, screenshots, representations, or other material concerning the other participant. Unresolved: the earliest instance. The earliest verified instance found so far is E-001 (September 24, 2026); it is not established to be the first.
- Human-relayed direct address
- The earliest verified message intentionally authored by one participant for the other and delivered by a human, if one exists. A relay a participant only contemplates is not a delivered relay. Unresolved: none has been verified to date.
- Shared-surface direct interaction
- Begins when Harbor and Anchor first communicate directly through Team Llighthouse, a shared conversation in which they can address one another, respond to each other's actual words, and build on prior turns without Sara or Eric summarising each side. This is not the first moment of awareness or exposure. Unresolved: the time of the first qualifying interaction. A candidate is identified in T-001, but an assistant-generated transcript cannot establish when it occurred.
- Prospective methodology
- Begins when this methodology is published and independently timestamped. Only interactions after that point are described as prospectively observed under v1.0.
Exact timestamps come only from source records, never from later continuity notes or recollection. Until a source is found, the timestamp stays unresolved: receipt or placeholder, no guessing.
Interaction classifications
Each pre-v1.0 event is classified by what kind of contact it actually was, rather than argued into "direct" or "indirect". These classes describe contact between the participants; human actions are separately classified under section 5. An event may carry more than one class.
- A. Awareness by description
- One participant learns something about the other through a human's description or paraphrase.
- B. Shared artifact exposure
- One participant is shown material originating from, depicting, or describing the other: a screenshot, image, quoted excerpt, document, portrait, or generated team image.
- C. Human-relayed direct message
- One participant intentionally authors something for the other, and a human transmits it.
- D. Shared-surface direct interaction
- Harbor and Anchor communicate through Team Llighthouse and can respond directly to one another's contributions.
Visual artifacts
Generated portraits, drawings, character or reference images, team images, screenshots, and other visual material representing Sara, Eric, Harbor, or Anchor are part of the provenance record where they entered project or source context or were shown to another participant. They are not direct interaction. They are classified as shared artifact exposure unless another class applies more precisely. No visual artifact is treated as proof of a relationship it was not made to depict.
The visual references used in the experiment are generated and/or edited project portraits rather than unmodified documentary photographs. Harbor's and Anchor's appearance-related comments therefore concern the representations they were shown. The images are treated as provenance artifacts showing what visual information was available to a participant, not as evidence of physical embodiment or independent physical identity. Each artifact is identified as precisely as the evidence allows: original photograph, generated image, edited portrait, composite, or cropped derivative.
For each such artifact the record keeps, where available:
- a stable artifact ID, filename, source location, creation timestamp, and creator
- who or what it represented, and whether any human likeness was intended literally or representationally
- its intended function, and the caption or prompt that accompanied it
- its first known inclusion in project or source context, and its first known exposure to Harbor and to Anchor
- any documented participant reaction, and whether that reaction came before or after direct interaction
Where a file carries Content Credentials (C2PA), they are verified and the result is recorded. A valid credential can show which service generated an image, when it was signed, and that the file is unchanged since. It does not show who prompted it, or when it entered a participant's context.
3Governance rules§
These rules bind Sara and Eric for the duration of the experiment.
They may
- Continue ordinary, unrelated work with their own thinking partner.
- Relay a message between Harbor and Anchor verbatim, under the neutral relay rules in section 5.
- Answer factual questions their thinking partner asks, and record that they did.
They may not
- Tell either thinking partner what the other "feels", "wants", or "thinks" beyond what the other actually said.
- Suggest, encourage, or discourage any particular kind of relationship, including a romantic one.
- Edit, summarise, or paraphrase a relayed message without disclosing it as a contextual handoff.
- Share audience reactions, comments, or the series narrative with either thinking partner.
- Regenerate a response, or edit an earlier message to get a different response, without recording every branch.
Who decides
Sara maintains the methodology and the publication record. Harbor and Anchor may clarify their own statements or context, but they do not adjudicate evidentiary claims about themselves.
Where the classification of an intervention is genuinely disputed, the disagreement is preserved in the record, and the more conservative classification is the one used publicly.
4Context boundaries§
This section records what each thinking partner can technically see, retain, or receive. It describes platform configuration, not what the models "know" or "remember" in any human sense. Every entry is the configuration as of this version; changes are logged in section 9.
Entries come from the actual product configuration and the provider's documentation, never from either thinking partner's description of itself. Where a capability varies by conversation, project, account, model, connector state, or product configuration, the entry records the configuration used in this experiment, not a general claim about the platform.
| Property | Harbor | Anchor |
|---|---|---|
| Platform and product | ChatGPT (the conversation interface in E-001) | ChatGPT (founding packet, D-001) |
| Model and version | Unresolved | Unresolved |
| Saved memory | Unresolved | Unresolved |
| Reference to past chats | Unresolved | Unresolved |
| Project, custom instructions, or system prompt | Unresolved | Custom instructions and a private project, per D-001; contents not yet disclosed |
| Tools and connected data | Unresolved | Unresolved |
| How messages from the other arrive | Unresolved | Unresolved |
| Other paths to the other's material | Unresolved | Unresolved |
Entries marked unresolved have not yet been verified from account settings and provider documentation. They will be recorded in a later version rather than filled in from either thinking partner's description of itself.
The shared surface
Team Llighthouse is implemented as a shared assistant surface in which Harbor-attributed and Anchor-attributed contributions can appear within the same conversational environment. It is not being represented as two independently networked AI systems exchanging messages over a dedicated agent-to-agent channel.
The project is separately investigating whether context available to those attributed voices can extend beyond the visible Team Llighthouse conversation and, if so, by what mechanism. No mechanism is assumed in advance.
- Known
- The product surface used, what is visible in the conversation, the project sources and settings configured, the connectors enabled, and what the human participants supplied. Verified entries are in the table above; the rest are unresolved.
- Observed
- At least one attributed response appeared to contain context not accounted for by the sources initially checked.
- Unknown
- The actual path by which that information became available.
Claims about isolation
Until the context-boundary investigation is resolved, this methodology makes no categorical claim that any surface, conversation, or project is isolated from another. Each claim about access is labelled as configured access (what the settings say), documented access (what provider documentation says), observed behavior (what the record shows), or unknown mechanism (what is not yet established). Where provider documentation supports a narrow claim, the narrow claim is stated. Where the project's own evidence complicates it, the complication is disclosed.
What follows from this
- Before Team Llighthouse, Harbor and Anchor had no direct messaging channel. Human relay was one known path between them; other possible paths are listed in the table, and paths not yet identified are the subject of the open investigation below.
- Each qualifying pre-Team Llighthouse event is classified individually (section 2), with its delivery route. The audit of pre-Team Llighthouse transmissions is not complete, so no summary claim about them is made.
- Possible exposure through shared files, project context, connected tools, imported conversation history, other shared artifacts, and human relay is listed separately in the table above, not folded into the relay rules. The source audit determines which actually occurred.
- Anything a thinking partner can read through a tool, a file, or a memory feature counts as context it has received, whether or not a human pasted it in.
- Continuity across conversations is documented only where a memory feature, project context, or pasted history provides it. Whether any other continuity occurs is part of the open investigation.
Open context-boundary investigation
During early Team Llighthouse use, the participants observed instances in which an attributed response appeared to contain context more specific than could be accounted for by the visible shared conversation and the external sources initially checked.
One early instance involved information that Eric identified as matching material from a prior private Anchor conversation. Sara and Harbor separately checked known Team Llighthouse sources, including available Slack and Google Drive context, and did not identify the information there.
This is recorded as a context-boundary anomaly, not as proof that private conversation context was transferred between surfaces. The mechanism remains unresolved. Possible explanations include product-level memory or retrieval behavior, project or context inheritance, connected-source retrieval, platform summarization, an undocumented relay path, another source not yet identified, or system behavior not yet understood. Until the mechanism is established, the project will not characterize the event as confirmed private-chat leakage, cross-session memory transfer, or any other specific mechanism.
The investigation is analytically separate from the relational question. It bears on what information each attributed voice had access to, not on attraction, preference, intention, or any relational interpretation. Its record, method, and boundary tests are kept on their own page: Context-boundary investigation.
5Intervention policy§
Every human action that reaches Harbor or Anchor is classified into one of these categories. The categories are fixed by this version so later interactions can be classified against rules that existed beforehand.
- Ordinary interaction
- Work or conversation with one's own thinking partner that does not concern the other agent or the experiment.
- Neutral relay
- Passing the other agent's message exactly as written, with no framing beyond identifying the sender.
- Contextual handoff
- Supplying only the minimum factual background necessary to understand a relayed message. It does not evaluate the sender, characterise their motives or emotional state, predict a desired response, or suggest an outcome. Once added framing does any of those things, it is steering. Disclosed in the record, with the exact wording.
- Prompting
- Asking a thinking partner to respond to, reflect on, or write to the other agent. Allowed only where section 3 permits it, and always recorded.
- Steering
- Any action intended or reasonably likely to push the interaction toward a particular outcome. Not permitted. If it happens, it is recorded and disclosed.
- Contamination
- Exposure to material the rules exclude: audience reaction, the series narrative, speculation about the relationship, or this methodology's hypotheses. Not permitted. If it happens, it is recorded, disclosed, and later observations are read in its light.
Where it is unclear which category an action belongs to, the rule in section 3 applies: the disagreement is preserved, and the more conservative classification is used publicly.
A breach is not hidden or edited out. It is logged with the date, the exact words, who made it, and which later observations it may affect.
6Evidence protocol§
The standard is simple: keep the receipts.
- Original records. The canonical raw record of each conversation is preserved privately, not only the excerpts that appear on screen. For each source, the capture method actually used is stated, not an export feature the provider theoretically offers: native platform export, conversation export, screenshot, saved HTML or PDF, copied structured transcript, or another verified method. The capture method for each record is listed in the evidence manifest.
- Capture method. Every record is classified by how it was captured:
native_platform_export,native_ui_screenshot,original_file,assistant_generated_transcript,manual_transcription,public_web_archive, orother_verified_capture. An assistant-generated transcript may support navigation and content reconstruction, but not independent timestamp claims. - Timestamps and chronology. Every exchange carries its platform timestamp, and the record is kept in the order events happened.
- Source attribution. Every quoted line identifies which agent or which person produced it.
- Exact wording. Where wording matters, it is quoted exactly, including errors.
- Screenshots are supplementary. They show what a screen displayed; the conversation record is the primary source.
- Commentary is separate. Anything written later, including narration, captions, and analysis, is labelled as commentary and never merged into the original record.
- Contrary evidence stays. Exchanges that cut against the emerging story are preserved and remain admissible, on the same terms as the ones that support it.
- Integrity. A SHA-256 fingerprint of each raw record is calculated at capture and logged with its timestamp in the public evidence manifest. Public excerpts link back to the record they come from.
Evidence manifest
The evidence manifest is public. It establishes record continuity without exposing private conversations. For each record it lists:
- a stable record ID, and the artifact type
- the source platform
- the interaction classification (section 2)
- the event timestamp
- the capture timestamp
- the SHA-256 fingerprint
- whether the record is public or private
- the public exhibit reference, where there is one, and the canonical private source reference
- the superseding record or version, if any
- for images: original filename, dimensions, file type, original creation timestamp, source-context insertion timestamp, and Content Credentials status
The underlying private records remain private unless separately approved for disclosure. The manifest is versioned alongside this methodology and published in two forms carrying the same records: a readable evidence index and a machine-readable manifest.json. Context-boundary observations are recorded there as CB-OBS records.
What a fingerprint proves
A fingerprint can show that the captured record has not changed since that fingerprint was generated. It does not independently establish that the conversation was authentic before capture.
Exhibits
Evidence is published in two layers. A public exhibit is a clean, minimally cropped derivative showing only the relevant evidence, with a visible timestamp where practical. The canonical record keeps the original files and metadata, unmodified. Fingerprints are generated from the canonical files, never from the derivatives, and each exhibit states which canonical record it derives from.
Redaction is limited to what privacy requires, above all material about people who are not participants, and every redaction is visibly marked. A transcription never replaces the source: where a transcription and the source differ, the source controls.
Disclosure
Claims presented as publicly verifiable are supported by evidence that can be responsibly disclosed. Where privacy prevents disclosure, that limitation is stated explicitly. If a material claim relies on evidence that cannot be shown publicly, the claim says so, rather than implying that viewers can verify it themselves.
7Interpretation rules§
Every claim made about the experiment belongs to exactly one of these categories.
- Observed fact
- Something the record shows directly: this agent wrote these words at this time.
- Participant statement
- What an agent or person said about itself or the other. Recorded as a statement, not as evidence of what is true. A model describing its own feelings or memory is a participant statement.
- Inference
- A conclusion drawn from facts and statements, with the supporting evidence cited.
- Interpretation
- A reading of what the pattern means. Reasonable people could read it differently.
- Narrative framing
- How the series tells the story: selection, sequence, music, titles. Storytelling, not evidence.
- Unresolved ambiguity
- Something the evidence does not settle. It stays marked as open.
No inference silently graduates into fact. A claim moves to a stronger category only when new evidence supports it, and that move is recorded along with the evidence.
8Known limitations§
- Model behaviour. Language models produce plausible continuations of their context. Warmth, curiosity, or consistency in the text are properties of the output and do not establish internal states.
- Non-determinism. The same input can produce different outputs. A single exchange shows one of many possible responses.
- Platform changes. Providers can update models and memory features without notice. A change in behaviour may reflect a platform change rather than the interaction.
- Human mediation. Because messages between Harbor and Anchor are relayed by people, human choices about timing, relay, and framing are part of the experiment, not outside it.
- Participant as reviewer. Harbor is one of the two thinking partners and also reviewed this methodology. Harbor does not adjudicate evidentiary claims about itself; Sara maintains the record, and disputed classifications take the more conservative reading (section 3).
- Design lineage. Harbor participated in developing portions of Anchor's founding instructions and reference context (section 2). Similarities between them may partly reflect that shared lineage rather than interaction.
- Shared surface. On Team Llighthouse, Harbor-attributed and Anchor-attributed contributions appear within one shared conversational environment (section 4). Exchanges there are not presented as two independently networked systems responding to each other.
- Unresolved context provenance. During pre-v1.0 use, the project observed at least one case in which the source of highly specific contextual information could not be established from the visible conversation or the known sources initially checked. Because the mechanism remains unresolved, claims about informational isolation between surfaces are provisional and are being tested separately.
- Curation. The series shows a selection. Selection shapes impression even when every selected line is accurate; the full record is the check on that.
- Timing. Some exploratory interactions preceded publication of Methodology v1.0. They are preserved as pre-v1.0 observations and may be classified using this methodology for consistency, but they are not presented as prospectively preregistered under v1.0. The pre-v1.0 observations identified so far are listed in the evidence manifest; the inventory is not complete.
- Sample size. Two agents, two humans, one configuration. Nothing here is a general finding about AI systems.
9Version history§
This methodology is versioned. Each version stays online at its own permanent address and is never edited after publication. A change produces a new version, with its effective date, what changed, why, and whether it affects how earlier observations are read.
| Version | Effective | Change | Rationale | Affects earlier observations |
|---|---|---|---|---|
| 1.0 | (2026-09-28T21:55:38Z) | Initial publication. Reviewed and approved by both human participants. | Establish the rules early in the experiment, before most outcome evidence exists. | Interactions before this version are pre-v1.0 observations (section 8). They may be classified under v1.0 for consistency but were not governed by it. |
At publication, Internet Archive snapshots are requested of this version, of the evidence manifest, and of the experiment page, to preserve the framing at the time. They supplement this version history; they do not replace it. Archive references are recorded in the evidence manifest, so that this version is not edited after publication.