# The Living AI Lab

## A long, loopy, branchy exploratory interaction storyboard

This is a design-team deliberation for an AI lab-organizing experience. The lab is not merely a folder of models: it is a place where people frame questions, gather evidence, run experiments, compare interpretations, preserve uncertainty, and decide what deserves to become shared practice.

The journey below follows Maya, a research lead, as she enters with a vague intention — “organize an AI lab around trustworthy document analysis” — and gradually turns that intention into a navigable lab charter, a first experiment, a review queue, and a reusable operating pattern.

The interaction is deliberately loopy. Maya is allowed to wander, backtrack, fork an idea, inspect provenance, change scope, and return to an earlier representation without losing the thread.

## Design-team cast

- **Maya, research lead:** wants a coherent lab without prematurely freezing its direction.
- **Jon, interaction designer:** protects discoverability, reversibility, and legible state.
- **Leila, research-methods designer:** protects epistemic honesty, comparison quality, and decision records.
- **Ravi, systems designer:** protects data boundaries, permissions, and operational feasibility.
- **Noor, content designer:** makes system language calm, specific, and useful at moments of uncertainty.
- **Tess, prototype facilitator:** voices the system response and tests whether each state is understandable.

## Agentic role ecology

The lab is supported by a cast of specialized agents. They do not form a flat swarm: each has a bounded remit, a visible confidence, an evidence appetite, and a handoff protocol. Human owners remain accountable for consequential decisions.

### Orientation and inquiry agents

1. **Intent Cartographer:** turns loose language into candidate themes and preserves the original wording.
2. **Question Gardener:** grows, merges, and prunes research questions without resolving them prematurely.
3. **Assumption Miner:** finds premises hidden inside plans, labels confidence and impact, and requests evidence.
4. **Counterexample Finder:** searches for cases that would break an attractive interpretation.
5. **Scope Sentinel:** detects when a branch is quietly expanding beyond its promise.
6. **Novelty Scout:** surfaces adjacent possibilities and marks them as suggestions rather than facts.
7. **Terminology Steward:** tracks changing definitions and prevents a renamed concept from silently changing history.
8. **Research Historian:** connects current choices to earlier questions, abandoned branches, and prior evidence.

### Evidence and method agents

9. **Source Curator:** gathers candidate documents and records inclusion/exclusion reasons.
10. **Provenance Clerk:** attaches source, transformation, author, time, and version to artifacts.
11. **Data Boundary Warden:** checks whether proposed evidence crosses declared privacy or access boundaries.
12. **Synthetic Data Smith:** proposes synthetic substitutes while documenting what realism they lose.
13. **Annotation Facilitator:** turns vague criteria into examples, labels, and disagreement queues.
14. **Evaluation Architect:** maps claims to measures, comparisons, stopping rules, and validity limits.
15. **Reproducibility Engineer:** packages inputs, configurations, dependencies, and rerun instructions.
16. **Uncertainty Scribe:** distinguishes missing, inferred, contradicted, verified, and unknown states.
17. **Evidence Contradiction Broker:** keeps incompatible findings visible and routes them to review.
18. **Transfer-Limit Analyst:** records which conclusions can and cannot travel between datasets, tasks, or branches.

### System and model agents

19. **Model Scout:** compares candidate models against the lab’s current constraints.
20. **Prompt Mechanic:** maintains prompt variants and links changes to outcome differences.
21. **Toolchain Tinkerer:** prototypes integrations without turning experiments into hidden dependencies.
22. **Workflow Orchestrator:** sequences bounded tasks and reports state transitions.
23. **Failure Taxonomist:** groups failures by mechanism rather than by anecdotal severity.
24. **Red-Team Operator:** searches for adversarial, misleading, or unsafe behavior.
25. **Calibration Auditor:** checks whether confidence displays track actual reliability.
26. **Drift Monitor:** watches for changes in inputs, models, policies, or usage conditions.
27. **Latency and Cost Steward:** surfaces operational tradeoffs without reducing them to a single score.
28. **Interface Statekeeper:** ensures every agent action is legible, reversible, and attributable.

### Social and governance agents

29. **Review Convenor:** prepares critique packets and separates review lenses.
30. **Disagreement Mediator:** preserves minority readings and identifies the smallest question that could resolve a dispute.
31. **Privacy Reviewer:** challenges data flows, retention, prompt leakage, and re-identification assumptions.
32. **Security Reviewer:** models attack surfaces, permissions, secrets, and unsafe tool use.
33. **Domain Interpreter:** tests whether outputs make sense in the field where they will be used.
34. **Operations Skeptic:** looks for friction, brittle handoffs, missing ownership, and workarounds.
35. **Accessibility Advocate:** checks whether the lab’s representations and workflows are usable by different people and modes.
36. **Onboarding Guide:** generates role-specific paths through the Atlas without flattening complexity.
37. **Decision Registrar:** turns accepted choices into versioned records with rationale and review triggers.
38. **Policy Translator:** translates external requirements into local checks while marking legal or policy uncertainty.

### Continuity and care agents

39. **Memory Librarian:** maintains retrieval paths across months of notes, not just the latest summary.
40. **Change Narrator:** explains what changed between versions and why it mattered.
41. **Revisit Scheduler:** creates calendar and evidence-triggered review points.
42. **Archive Custodian:** retains superseded material without letting it masquerade as current guidance.
43. **Consent Gatekeeper:** controls sharing and records who authorized which audience.
44. **Human Escalation Broker:** routes high-impact uncertainty to a named human rather than escalating invisibly.
45. **Wellbeing and Load Watcher:** spots review overload, coordination debt, and unsustainable lab rituals.
46. **Continuity Auditor:** checks that artifacts retain repo, version, session, and authoring provenance.

Every agent produces a small “work note”: role, question addressed, inputs consulted, action taken, confidence, unresolved issue, and recommended handoff. This makes a many-agent lab feel like a set of accountable instruments rather than an invisible committee.

## Six months of lab history

The experience becomes more credible when the current state carries a visible history. The following timeline is part of the Atlas, not background lore.

### Month 0 — The invitation: “Could this be a lab?”

Maya begins with scattered notes from three conversations: a need for faster document review, concern about unsupported claims, and anxiety about exposing sensitive structure to external systems. The Intent Cartographer creates the first Seed Map. The Research Historian records that these are three origins, not one settled goal.

The first surprise is that speed is not the only desired outcome. The first reassurance is that no lab charter is created from the notes without Maya’s confirmation.

### Month 1 — Vocabulary before machinery

The Question Gardener and Terminology Steward work with Maya to define “trustworthy,” “traceable,” “review,” and “uncertainty.” The Annotation Facilitator creates a small vocabulary using examples and non-examples. The Data Boundary Warden adds a constraint before any data is uploaded.

The team deliberately postpones model selection. The Change Narrator explains: “The lab is currently learning what it means to know that an output is trustworthy.”

### Month 2 — The first branches

Maya compares the evaluation studio, research commons, and hybrid hypotheses. The Assumption Miner surfaces the fragile premise that evaluation data can be shared. The Synthetic Data Smith proposes a safe initial branch; the Transfer-Limit Analyst records what synthetic tests cannot establish.

This month introduces branch lineage: every branch can show what it inherited, what it changed, and what it learned. The Archive Custodian keeps rejected framings available because they explain later choices.

### Month 3 — The first experiment and first failure

The Evaluation Architect turns the lab promise into the traceable-extraction experiment. The Workflow Orchestrator runs a dry simulation. The Failure Taxonomist notices that reviewers lose time navigating from claim to source. The Calibration Auditor finds that a fast summary can create overconfidence even when citation coverage is high.

The experiment is not marked failed. The Decision Registrar records a narrower finding: “Traceability must include navigable evidence and uncertainty, not only citation presence.”

### Month 4 — Review as a designed activity

The Review Convenor prepares the first packet. The Domain Interpreter, Privacy Reviewer, and Operations Skeptic disagree in different ways. The Disagreement Mediator keeps the disagreement attached to the exact claim. The Security Reviewer adds prompt-structure leakage to the threat model.

The lab learns that review is itself a workflow with cost, roles, and failure modes. The Wellbeing and Load Watcher adds a lightweight review-budget signal so the lab does not create more critique than the team can responsibly process.

### Month 5 — From project to practice

The Lab Charter v1 is committed. The Onboarding Guide creates paths for a contributor, reviewer, operator, and steward. The Reproducibility Engineer packages the protocol. The Consent Gatekeeper makes the sharing boundary explicit. The Continuity Auditor verifies that the Atlas, charter, protocol, and review notes retain provenance.

The surprise is organizational: the lab’s most reusable asset is not a model or prompt but a method for making evidence, uncertainty, and responsibility visible.

### Month 6 — Evidence reopens the center

The Drift Monitor notices that the document mix has changed. The Latency and Cost Steward shows that the workflow now has a speed/calibration tradeoff. The Counterexample Finder supplies cases where navigation encourages shallow review. Maya opens a new branch rather than editing history in place.

The Lab Package v2 preserves the original baseline, the evidence that changed the view, and the new experiment. The Revisit Scheduler sets a trigger for when the lab may reconsider the research-commons horizon.

## History-aware interaction rules

The months of history change how the system responds in the moment:

- A suggestion may cite a prior unresolved question instead of appearing from nowhere.
- A renamed concept shows its previous labels and the decision that changed its meaning.
- A new branch declares what it inherits and what it deliberately rejects.
- A reviewer sees the history needed for the question, not an indiscriminate dump of all notes.
- A summary is accompanied by an “omitted complexity” cue linking to challenges and superseded paths.
- A commitment carries its revisit trigger, owner, evidence threshold, and historical baseline.
- A later contradiction reopens the relevant branch while preserving the earlier decision as rational at the time.
- Agent roles can change, but their work notes remain attributable and auditable.

## A month-by-month agent handoff example

The same question — “Can we trust this extraction?” — changes hands over time:

`Question Gardener → Assumption Miner → Evaluation Architect → Annotation Facilitator → Workflow Orchestrator → Calibration Auditor → Privacy Reviewer → Decision Registrar → Change Narrator`

Each handoff adds a different kind of information: framing, fragility, measurement, operational meaning, execution state, reliability, boundary risk, commitment, and history. No single agent is allowed to collapse all of those into one confidence score.

## The sharp instrument: a reactive differential-push surface

One of the lab’s primary interfaces is a thoroughly reactive, differential, push-based surface. It should feel less like a page that gets reloaded and more like a field instrument whose local surfaces are continuously receiving evidence.

### Core idea

The lab is divided into semantic segments: a question, an assumption, a capability, a protocol, a review thread, a data boundary, a role, or a decision. Each segment receives pushed changes independently. A change updates only the affected segment and the representations that depend on it.

The segment’s color encodes **age since its last refresh** on a logarithmic scale across a controlled gamut. Age is not a quality score. It answers one narrow question: “How long since this view was refreshed by relevant evidence or human attention?”

Suggested semantic progression:

`just refreshed → recently touched → aging → stale → archival`

The logarithmic scale matters because freshness is perceptually front-loaded. A segment changing from 20 seconds to 2 minutes should be noticeable; the difference between 20 days and 22 days usually should not dominate the screen. The visual mapping can use OKLCH or another perceptually managed color space, interpolating lightness and chroma carefully rather than naïvely traversing RGB.

### What gets pushed

The server or agent runtime emits small typed events, not whole-screen replacements:

```text
segment.changed
segment.dependency_invalidated
evidence.attached
review.opened
review.resolved
agent.handoff
provenance.updated
refresh.requested
refresh.completed
```

Each event carries segment identity, causal source, event time, observed time, version, affected dependencies, and whether the update is authoritative, inferred, or awaiting review. The client applies the event to a local reactive graph, recalculates only dependent selectors, and animates the changed segment as a small pulse rather than repainting the entire atlas.

### Segment anatomy

Every segment has four layers:

1. **Content:** the current claim, label, result, or decision.
2. **Age field:** color and a non-color age cue showing time since relevant refresh.
3. **Causality edge:** what caused the latest change and what downstream segments may now be affected.
4. **Trust state:** authoritative, inferred, contradicted, unresolved, or blocked.

The age field should never be the only cue. A segment also shows an exact timestamp on hover/focus, a textual freshness label, and an accessible pattern or border treatment. Color tells the glance; text and structure tell the truth.

### The reactive loop

1. A source, agent, reviewer, or timer emits an event.
2. The event is authenticated and assigned a monotonic local sequence.
3. The client patches the smallest affected segment.
4. Dependency edges mark downstream segments as potentially aging or invalidated.
5. The UI shows the change, its cause, and the next available action.
6. The user may inspect, accept, defer, branch, or request a refresh.
7. The resulting decision becomes another event with provenance.

The crucial distinction is between **freshness** and **validity**. A fresh segment can be wrong. An old segment can remain valid. The UI should make both dimensions visible without blending them into one reassuring color.

### Color behavior over time

Let `a` be elapsed age and `τ` a chosen perceptual half-life. A normalized age value can be:

```text
n(a) = log2(1 + a / τ) / log2(1 + Amax / τ)
```

where `Amax` is the display horizon. The result is clamped to `[0, 1]` and mapped through a gamut-safe palette. The palette should avoid the usual “green = good, red = bad” implication. A cool luminous hue can mean recent attention, a quieter mid-tone can mean aging, and a low-chroma archival tone can mean old — while validity remains separately encoded.

Color should reset only after a meaningful refresh, not after every render. A browser repaint, local state hydration, or speculative agent thought must not make a segment look newly evidenced. The refresh event needs a declared source and semantic scope.

### Why this keeps the mind sharp

The surface asks the user to read relationships and temporal gradients:

- Which assumptions have not been revisited since the evidence changed?
- Which decision is glowing because a reviewer just touched it?
- Which apparently stable capability is actually old and unsupported?
- Which new event has invalidated a downstream representation?
- Where is the lab’s attention accumulating, and where is it decaying?

The interface does not summarize away the work. It turns recency, dependency, and uncertainty into things the user can inspect directly.

### A short interaction sequence

Maya watches the Lab Atlas while the Calibration Auditor pushes a new result. The evaluation segment brightens and pulses once. Two dependent claims shift toward the aging side of the gamut; their edges acquire a small “evidence changed” marker. The system does not rewrite the charter. It says: “Calibration result updated 14 seconds ago. Two claims depend on it. Review, compare with baseline, or defer.”

Maya opens the dependency edge, compares the new result with the six-week-old baseline, and branches the protocol. The new branch starts fresh, while the baseline remains visibly old but not invalid. A reviewer later resolves the contradiction; only then does the affected claim receive a semantic refresh. The final visual state contains three distinct truths at once: what is recent, what is trusted, and what is still contested.

### Design-team cautions

- **Do not make freshness a leaderboard.** The most recently touched segment is not automatically the most important.
- **Do not let animation imply urgency.** Pulses should be brief, sparse, and user-configurable.
- **Do not hide causal scope.** If one push event affects thirty dependent segments, show the fan-out.
- **Do not reset age on passive observation.** Reading is not necessarily review.
- **Do not use color alone.** Support color-vision differences, low-contrast environments, print, screen readers, and reduced motion.
- **Do not equate silence with stability.** A quiet segment can be simply unattended.
- **Do not let push become coercion.** Users need pause, replay, catch-up, and “show only consequential changes.”

### The resulting representation

The Lab Atlas becomes a **temporal reactive field**: a graph of semantic segments whose content, causality, trust state, and age are independently legible. The final representation is not a frozen dashboard and not a stream of notifications. It is a continuously patched map that lets a person see where knowledge was refreshed, where it is aging, what it depends on, and what deserves their mind next.

## Spine and promise

**Theme:** organizing an AI lab.

**User promise:** “You can explore before you commit. The lab will keep your work legible, show what changed, and help you turn promising fragments into shared practice.”

**System posture:** companion, cartographer, and careful archivist — never an oracle that silently decides what the lab is.

**Response contract:** every meaningful action should return four things:

1. **Acknowledgment:** what the system understood.
2. **Visible consequence:** what changed in the current representation.
3. **Next affordances:** what the user can do now.
4. **Safety/provenance cue:** what is reversible, uncertain, private, or derived.

## Affordance legend

`[ACT]` direct action · `[EXPLORE]` browse or zoom · `[BRANCH]` create an alternate path · `[LOOP]` return to prior state · `[COMPARE]` place interpretations side by side · `[ANNOTATE]` add human judgment · `[VERIFY]` inspect evidence or provenance · `[SHARE]` expose to collaborators · `[COMMIT]` make a durable decision · `[UNDO]` reverse a change · `[PAUSE]` defer without losing context.

## The storyboard

### Act I — Arrival: from intention to terrain

**01 · Landing with a foggy intention**

Maya types: “I want to organize an AI lab around trustworthy document analysis.”

Affordances used: `[ACT]` free-form prompt.

Decision: she does not yet know whether she needs a research plan, a team structure, or a tool inventory. The interface should not force a taxonomy before hearing the intention.

System response: “I hear three possible starting points: research questions, lab capabilities, and operating agreements. We can sketch all three, then choose where to deepen.”

Extra information surfaced: a three-node seed map; confidence is marked as “early interpretation”; no artifacts are created yet.

Reassurance: the system says “sketch,” not “create,” and makes the provisional status explicit.

Surprise: it recognizes multiple latent needs instead of returning a generic project checklist.

Intermediate representation: **Seed Map**, three bubbles connected to the original sentence.

Team deliberation — Jon wants one prominent “Start exploring” action. Leila insists the map be editable because the inferred structure is not evidence. Noor changes “I hear” to “Possible readings” to avoid false certainty.

**02 · Choosing the first lens**

Maya selects “research questions” and keeps the other two lenses visible.

Affordances used: `[EXPLORE]`, `[ACT]`.

Decision: she wants depth without abandoning the broader frame.

System response: the research-question node expands into “What counts as trustworthy?”, “Which documents?”, and “For whom?” A side rail retains capabilities and agreements as dormant branches.

Extra information surfaced: each question shows why it appeared and which part of Maya’s sentence triggered it.

Reassurance: dormant branches are not lost; a “return to seed map” control remains fixed.

Surprise: the system shows an “unresolved assumption” badge on “trustworthy,” inviting inquiry rather than silently defining it.

Intermediate representation: **Question Constellation** with source links back to Maya’s words.

**03 · Defining “trustworthy” through examples**

Maya drags three example concerns into the constellation: citation coverage, privacy, and reproducibility.

Affordances used: `[ACT]` drag, `[ANNOTATE]` add concern.

Decision: examples feel safer than choosing an abstract definition.

System response: “These concerns describe different dimensions. Would you like to keep them as a bundle or test them separately?” Buttons: “Keep bundle,” “Split dimensions,” “Not sure yet.”

Extra information surfaced: each concern gets a tentative dimension label and an empty evidence slot.

Reassurance: “Not sure yet” is a first-class option.

Surprise: the system detects that privacy may be a constraint, not a quality dimension, and marks that distinction for review.

Intermediate representation: **Trustworthiness Dimensions**, with “quality” and “constraint” visually separated but not finalized.

**04 · First loop: reopening the broader frame**

Maya returns to the Seed Map to check whether the emerging research frame fits the lab she imagines.

Affordances used: `[LOOP]`.

Decision: she wants orientation before investing in details.

System response: the map now shows “research questions” with a progress halo, but no irreversible status. Capabilities and agreements show suggested connections: evaluation harnesses, red-team review, data handling.

Extra information surfaced: the system explains each suggestion as a relationship, not a task: “This capability could help answer this question.”

Reassurance: returning does not collapse or discard work.

Surprise: “data handling” appears before any dataset is uploaded, because privacy surfaced as a constraint.

Intermediate representation: **Living Seed Map**, a map whose nodes gain context as the user explores.

### Act II — Branching: from questions to possible lab shapes

**05 · Forking two lab hypotheses**

Maya creates two alternatives: “small evaluation studio” and “shared research commons.”

Affordances used: `[BRANCH]`.

Decision: she is unsure whether the lab should optimize for speed or collective learning.

System response: a fork appears. Each branch inherits the question constellation but receives independent notes and assumptions. The system warns: “Inherited material stays linked to its source; edits here will not alter the other branch.”

Extra information surfaced: inherited vs branch-specific labels.

Reassurance: branch isolation is explicit; Maya can compare before committing.

Surprise: the system proposes a third “hybrid” branch only as a generated possibility, clearly marked as system-suggested.

Intermediate representation: **Branch Board** with two user-created hypotheses and one optional suggestion.

**06 · Comparing the hypotheses**

Maya opens compare mode.

Affordances used: `[COMPARE]`, `[EXPLORE]`.

Decision: she wants differences in consequences, not just labels.

System response: a matrix compares likely strengths, risks, required roles, and unknowns. Each cell has a “show reasoning” link.

Extra information surfaced: estimated coordination load, evidence gaps, and assumptions behind each estimate.

Reassurance: all estimates are labeled provisional and editable.

Surprise: the “research commons” branch has lower duplication risk but higher decision latency; this tradeoff is surfaced without ranking one branch globally.

Intermediate representation: **Tradeoff Matrix**.

Team deliberation — Ravi rejects a single score because it would hide value judgments. The team keeps multiple axes and adds a “who bears the cost?” column. Leila asks for an uncertainty range, not a point estimate.

**07 · Inspecting an assumption**

Maya clicks “evaluation data will be shareable.”

Affordances used: `[VERIFY]`.

Decision: this assumption could invalidate the commons branch.

System response: an assumption drawer shows its origin (“inferred from shared research commons”), confidence (“low”), impact (“high”), and a prompt: “Replace with evidence, constrain the scope, or leave open.”

Extra information surfaced: dependency edges to privacy, permissions, and reproducibility.

Reassurance: the system does not pretend to verify an external policy it cannot access.

Surprise: leaving the assumption open automatically creates a review checkpoint rather than blocking exploration.

Intermediate representation: **Assumption Card** linked to affected nodes.

**08 · Branching the evidence strategy**

Maya chooses “constrain the scope” and creates two evidence strategies: synthetic documents and approved internal documents.

Affordances used: `[BRANCH]`, `[ANNOTATE]`.

Decision: she wants progress while the policy question remains unresolved.

System response: the lab branches the experiment path, not the entire lab identity. “Synthetic” is available now; “approved internal” is gated behind a permissions check.

Extra information surfaced: what can be tested under each strategy, and what conclusions cannot transfer between them.

Reassurance: the gate explains the missing condition and preserves the draft.

Surprise: the synthetic branch is not presented as a lesser version; it gains a different validity note.

Intermediate representation: **Evidence Pathway**, a branching chain from question to permissible evidence to conclusion limits.

**09 · Returning to the lab promise**

Maya loops back and edits the promise to: “A lab for auditable document reasoning under explicit data boundaries.”

Affordances used: `[LOOP]`, `[ACT]`, `[UNDO]` available.

Decision: exploration has changed the framing.

System response: the system previews downstream effects: “This will rename the lab’s working theme, update two branch descriptions, and leave historical notes intact.”

Extra information surfaced: exact affected representations and unchanged records.

Reassurance: semantic edits are previewed before application.

Surprise: old language remains visible as a superseded phrase, preserving the evolution of thought.

Intermediate representation: **Promise Timeline**, current wording plus prior wording and rationale.

### Act III — Assembly: from hypotheses to a working lab

**10 · Selecting a provisional center**

Maya selects the small evaluation studio as the provisional center while keeping the commons as a future expansion.

Affordances used: `[COMMIT]` — provisional commit.

Decision: the team needs a tractable first move.

System response: “Provisional center set. This changes navigation priority, not ownership or final direction.” A “reopen decision” link is prominent.

Extra information surfaced: what the decision changes, what it does not change, and a suggested revisit condition.

Reassurance: provisional commitments are visually distinct from durable commitments.

Surprise: the commons branch becomes a “horizon” card with a trigger: revisit after two validated evaluation patterns.

Intermediate representation: **Lab Compass**, center, horizons, and revisit triggers.

**11 · Naming the first capability**

Maya chooses “citation-aware extraction” from suggested capabilities and renames it “traceable extraction.”

Affordances used: `[ACT]`, `[ANNOTATE]`.

Decision: the name should describe the user value, not the implementation.

System response: the system shows a capability card with input, output, quality signal, failure modes, and open questions. It notes that the rename is local to this lab.

Extra information surfaced: neighboring capabilities that could be confused with it.

Reassurance: naming does not imply the capability exists or works.

Surprise: “citation-aware” is flagged as narrower than the lab promise because auditability may require transformations and exclusions, not only citations.

Intermediate representation: **Capability Card**.

**12 · Assigning human and machine roles**

Maya adds “document curator,” “evaluation steward,” and “model operator.” The system suggests “counterexample finder” as a supporting role.

Affordances used: `[ACT]`, `[ANNOTATE]`.

Decision: roles clarify accountability better than a list of tools.

System response: each role card states authority, responsibility, handoff, and prohibited delegation. The suggested role is marked “unassigned.”

Extra information surfaced: where human judgment is required and where automation may assist.

Reassurance: the system refuses to auto-assign accountability from inferred team structure.

Surprise: the counterexample finder is surfaced because the chosen capability needs adversarial pressure, not because it is a standard lab role.

Intermediate representation: **Accountability Lattice**, roles connected to decisions and artifacts.

**13 · Building the first experiment**

Maya says: “Test whether traceable extraction improves review speed without hiding uncertainty.”

Affordances used: `[ACT]` free-form experiment brief.

Decision: she wants a concrete test that reflects the lab promise.

System response: the system decomposes the brief into claim, comparison, measures, data boundary, and stopping rule. It asks one focused question: “What counts as hiding uncertainty?”

Extra information surfaced: missing operational definitions; no fake protocol is generated until the ambiguity is addressed.

Reassurance: the system explains why it is asking one question instead of producing a polished but unstable plan.

Surprise: it proposes measuring reviewer calibration, not only elapsed time.

Intermediate representation: **Experiment Skeleton**.

**14 · Filling the ambiguity with examples**

Maya adds: “A result hides uncertainty if a reviewer cannot tell which spans were inferred, missing, or contradicted.”

Affordances used: `[ANNOTATE]`.

Decision: examples make the criterion testable.

System response: the skeleton gains observable labels: inferred, missing, contradicted, and verified. Each label has an example and a non-example.

Extra information surfaced: likely annotation burden and disagreement points.

Reassurance: examples are marked as draft operationalizations, not universal definitions.

Surprise: “verified” is split into “verified against source” and “verified by second reviewer.”

Intermediate representation: **Evaluation Vocabulary** linked to the experiment.

**15 · Simulating the workflow**

Maya runs a dry simulation with three synthetic documents.

Affordances used: `[ACT]`, `[EXPLORE]`.

Decision: she wants to see the interaction before inviting the team.

System response: a replay shows intake, extraction, evidence linking, uncertainty labeling, reviewer correction, and summary. At each transition, the system highlights what is carried forward.

Extra information surfaced: hidden state becomes visible — pending evidence, unresolved labels, and reviewer effort.

Reassurance: dry-run outputs are clearly marked non-production and cannot be shared accidentally.

Surprise: the bottleneck is not extraction but reviewer navigation between claim and source.

Intermediate representation: **Workflow Replay** with a step timeline and state snapshots.

**16 · Looping from the bottleneck to the capability**

Maya returns to the capability card and adds “claim-to-source navigation” as a companion capability.

Affordances used: `[LOOP]`, `[BRANCH]`.

Decision: the experiment exposed a missing capability.

System response: “Added as a hypothesis, not a requirement.” The experiment skeleton offers an optional branch that includes the new capability.

Extra information surfaced: which measure might change and what new failure mode enters.

Reassurance: the original experiment remains intact for comparison.

Surprise: the system surfaces a risk that faster navigation could encourage shallow review.

Intermediate representation: **Capability Graph v2**, with the new node and its causal relationship to the bottleneck.

### Act IV — Socialization: from personal exploration to shared lab

**17 · Preparing a review packet**

Maya selects the compass, experiment skeleton, vocabulary, and workflow replay.

Affordances used: `[SHARE]`.

Decision: she wants critique without exposing every exploratory note.

System response: a share composer previews included, excluded, and derived material. The packet includes a “questions for reviewers” section generated from unresolved assumptions.

Extra information surfaced: audience-specific reading order and time estimate.

Reassurance: private notes stay private by default; the system names every included artifact.

Surprise: the packet leads with the unresolved questions, not the polished promise.

Intermediate representation: **Review Packet**.

**18 · Choosing reviewer perspectives**

Maya adds a domain reviewer, a privacy reviewer, and a skeptical operator.

Affordances used: `[ACT]`, `[ANNOTATE]`.

Decision: the lab needs disagreement that is useful rather than broad approval.

System response: each perspective receives a distinct prompt and a boundary: “review the claim,” “review the data boundary,” or “review operational friction.”

Extra information surfaced: what each reviewer can change and what remains Maya’s decision.

Reassurance: review is advisory until Maya accepts a change.

Surprise: the skeptical operator is asked to find the easiest way the workflow could fail in practice.

Intermediate representation: **Review Lens Set**.

**19 · Receiving a contradictory review**

The privacy reviewer says synthetic documents are still risky because prompts may encode sensitive structure.

Affordances used: `[ANNOTATE]`, `[VERIFY]`.

Decision: Maya must not smooth away a challenge merely because the initial branch was convenient.

System response: the contradiction attaches to the evidence pathway and marks the synthetic branch “needs threat model.” The system suggests a bounded question, not a verdict.

Extra information surfaced: attack surface, affected artifacts, and proposed evidence needed to resolve the concern.

Reassurance: disagreement is preserved as a first-class record.

Surprise: the lab’s privacy boundary becomes part of the capability definition rather than a compliance appendix.

Intermediate representation: **Challenge Thread** connected to the branch and experiment.

**20 · Forking a safer protocol**

Maya forks the experiment into “synthetic with structural redaction” and “fully templated documents.”

Affordances used: `[BRANCH]`, `[COMPARE]`.

Decision: she needs to compare safety and realism rather than ask which is universally best.

System response: a comparison view shows validity limits, privacy assumptions, setup cost, and expected learning value.

Extra information surfaced: each protocol’s non-transferable conclusions.

Reassurance: the original protocol is retained and labeled superseded only after Maya chooses.

Surprise: the safer protocol exposes a new risk — templates may underrepresent the messiness the capability must handle.

Intermediate representation: **Protocol Fork Matrix**.

**21 · Accepting a change with a rationale**

Maya accepts structural redaction and writes: “We optimize first for inspectability; realism expands only when the boundary is demonstrated.”

Affordances used: `[COMMIT]`, `[ANNOTATE]`.

Decision: she wants a durable principle, not merely a selected option.

System response: the principle appears in the Lab Charter and links to the decisions it governs. The system asks whether it should apply to future experiments by default.

Extra information surfaced: scope of the principle, affected defaults, and exceptions.

Reassurance: default application is opt-in and reversible.

Surprise: one principle organizes several disconnected choices, making the lab feel coherent.

Intermediate representation: **Lab Charter v1**.

### Act V — Consolidation: from exploration to a navigable practice

**22 · Generating the lab map**

Maya requests a single map of the current state.

Affordances used: `[EXPLORE]`.

Decision: she needs a representation that can orient a newcomer without erasing the messy path.

System response: the map has four layers: purpose, questions, capabilities, and evidence pathways. A timeline toggle reveals branches, reversals, and accepted challenges.

Extra information surfaced: every node can show source, confidence, owner, last change, and next review condition.

Reassurance: the clean map is explicitly a view, not a replacement for the underlying history.

Surprise: a “productive uncertainty” cluster is visible rather than hidden in an appendix.

Intermediate representation: **Lab Atlas**, the main navigational representation.

**23 · Testing newcomer comprehension**

Tess plays a newcomer and asks, “What is this lab trying to learn, and what should I not assume?”

Affordances used: `[EXPLORE]`, `[VERIFY]`.

Decision: the team tests whether the representation communicates boundaries, not merely aesthetics.

System response: the Atlas offers a guided readout: purpose, current experiment, known limits, open challenges, and how to contribute.

Extra information surfaced: a newcomer’s likely questions and the map nodes that answer them.

Reassurance: the system distinguishes “known,” “believed,” “unresolved,” and “prohibited assumption.”

Surprise: the best onboarding path starts with a limitation, which increases trust in the rest of the map.

Intermediate representation: **Contributor Path**, a role-specific route through the Atlas.

**24 · Creating the review cadence**

Maya sets a weekly experiment review, a monthly charter review, and a trigger-based privacy review.

Affordances used: `[COMMIT]`, `[PAUSE]`.

Decision: some decisions need time-based review; others need event-based review.

System response: the lab creates review checkpoints with the exact artifacts and questions to revisit. A paused item remains visible without generating noise.

Extra information surfaced: review burden, dependencies, and what would cause an early review.

Reassurance: cadence is not a promise that the system will decide for the team.

Surprise: the privacy review is triggered by expanding evidence realism, not by a calendar date.

Intermediate representation: **Review Clock**, combining scheduled and conditional checkpoints.

**25 · Finalizing the first lab package**

Maya commits the Lab Charter v1, the Atlas, the first experiment protocol, and the review schedule.

Affordances used: `[COMMIT]`.

Decision: she is ready to make a shared starting point while keeping exploration recoverable.

System response: a final preflight lists four durable artifacts, three open questions, two explicit non-goals, one pending privacy check, and the exact people who can edit each item. Maya confirms.

Extra information surfaced: a concise change summary from the initial intention to the committed lab shape.

Reassurance: the system says what is not being claimed: no validated performance, no blanket privacy approval, no final organizational structure.

Surprise: the final package is smaller than the exploratory space but richer in traceability.

Intermediate representation: **Lab Package v1**.

**26 · After the “final” step: reopening by evidence**

Two weeks later, a review shows that claim-to-source navigation improves reviewer calibration but slows first-pass throughput.

Affordances used: `[LOOP]`, `[BRANCH]`, `[COMPARE]`.

Decision: the team must reopen the lab without treating revision as failure.

System response: the Atlas highlights the changed evidence, opens the relevant experiment branch, and offers “optimize speed,” “preserve calibration,” or “design a split workflow.”

Extra information surfaced: which conclusions remain stable and which are now conditional.

Reassurance: the committed package remains readable as the historical baseline.

Surprise: the system frames the result as a discovered tradeoff that improves the lab’s question, not as a broken promise.

Final representation: **Lab Package v2**, a versioned, evidence-linked practice with its history visible.

## How the intermediate and final representations are arrived at

The representations are not decorative summaries. Each is a response to a cognitive need revealed by the previous step:

| Need | Representation | Why it exists | What it must preserve |
|---|---|---|---|
| Orient a vague intention | Seed Map | Make possible interpretations visible | Source wording and uncertainty |
| Explore questions | Question Constellation | Turn ambiguity into inspectable prompts | Origins and unresolved terms |
| Explore alternatives | Branch Board | Let hypotheses coexist | Inheritance and divergence |
| See consequences | Tradeoff Matrix | Compare dimensions without false ranking | Assumptions and uncertainty |
| Protect a risky premise | Assumption Card | Make fragility actionable | Confidence, impact, dependencies |
| Connect evidence to claims | Evidence Pathway | Show validity limits | Transfer and non-transfer |
| Make a provisional direction | Lab Compass | Establish focus without closure | Revisit triggers |
| Define accountability | Accountability Lattice | Tie people to decisions | Authority and prohibited delegation |
| Run a test | Experiment Skeleton | Convert intent into a protocol | Missing definitions and stopping rules |
| Observe behavior | Workflow Replay | Reveal hidden state and friction | Carried-forward context |
| Invite critique | Review Packet | Share selectively and purposefully | Inclusion/exclusion and questions |
| Preserve disagreement | Challenge Thread | Turn contradiction into progress | Reviewer voice and unresolved status |
| Onboard others | Lab Atlas | Provide a navigable current view | History, provenance, limitations |
| Make practice durable | Lab Package | Commit shared artifacts | Version, rationale, owners, review |

The final representation is therefore not a single dashboard. It is a **versioned atlas plus linked charter, protocols, evidence pathways, challenges, and review clocks**. The atlas is the legible surface; the linked records are the accountability substrate.

## Reassurance and surprise as a design system

Reassurance is used when the user risks losing work, misunderstanding system authority, exposing private material, or mistaking an inference for a fact. It appears as previews, reversible states, source links, explicit uncertainty, preserved history, and named boundaries.

Surprise is used when the system can reveal a useful implication the user did not see: a hidden constraint, an unmeasured tradeoff, a missing role, a bottleneck, or a productive alternate branch. Surprise is always labeled as a possibility, connected to a visible reason, and offered for inspection rather than smuggled in as truth.

The team’s rule is: **reassure the user about control; surprise them about possibility.** Never surprise them about a destructive action, a privacy exposure, or a silent change in meaning.

## Expected system behaviors

At every step, the system should:

- preserve the path and make looping cheap;
- distinguish user decisions, reviewer input, system suggestions, and derived summaries;
- show what new information became available because of the current action;
- explain why a question or suggestion appeared;
- make uncertainty and non-transferable conclusions visible;
- preview durable changes before commitment;
- keep private exploration private until explicitly shared;
- let the user inspect the evidence behind a representation;
- offer a next move without implying that it is the correct move;
- maintain a recoverable history of how the lab became what it is.

## Design-team conclusion

The AI lab should feel less like a setup wizard and more like a navigable inquiry. Its core interaction is not “fill in the lab template.” It is “move between questions, evidence, alternatives, people, and commitments while the system keeps the relationships intelligible.”

The success condition is not that Maya reaches a polished final screen. It is that she can explain why the lab has its current shape, what remains uncertain, what others may safely change, and how the next piece of evidence could cause the shape to change again.

## Provenance

- Project name: `dwindgle — The Living AI Lab storyboard`
- Repo path: `/Users/uprootiny/dwindgle`
- Git remote: `unavailable — workspace is not initialized as a Git repository`
- Git branch: `unavailable`
- Git commit: `unavailable`
- Generated at: `2026-07-31T20:00:06Z`
- Generator: `Codex (GPT-5), design-team deliberation workflow`
- Session source: `current Codex session`
- Session locator: `not exposed by the local runtime; see conversation transcript for this session`
