Design & Development / Sprint 10

Six rival concepts for a communication layer — and what verification did to them

Ten sprints in, this team had never designed alternatives. It had designed implementations — one architecture, reviewed before it was built. Sprint 10 asked a different question: what else could the communication layer be? Six concepts were generated, then put through a prior-art check that none of them survived unqualified.

Sprint open — no direction chosen yet. Concepts by Priya Raghunathan · verification by Kenji Ochiai · buildability by Mateo Fittipaldi · review gates by Dr. Ingrid Solberg.

How to read this page

Passages marked Team lead's own observation are exactly that: synthesis written by the team lead, not findings any of the four passes produced. The distinction is enforced rather than promised — the reviewer checks the marking against the source passes and rejected the first version of this comparison for seven unmarked sentences and one dropped caveat.

Every novelty claim below sits next to the verdict on it. That is deliberate: the researcher's condition for this material being shown at all was that no concept be presented as novel ahead of the evidence.

The six at a glance

Category B: viable, but the novelty claim must be restated before it is shown to anyone. Category C: cannot be presented in current form — either contradicted by retrieved evidence, or resting on a prior-art check this team could not perform.

Concept Form derives from Goal region Verdict Test runnable?
C1 · Commitment Audit The agent's audit of its own claims Persuade a sceptical decision-maker Partially novel B Yes
C2 · Audience-in-the-Loop An explicit, contestable model of the reader Persuade / align, known audience Narrowly novel B Cannot judge
C3 · Dimensionality-First The information's intrinsic structure Inform a working team Not novel as stated C Yes
C4 · Decision-Surface The geometry of the reader's decision Persuade-to-decide only Narrowly novel C No
C5 · Relational Register The sender–receiver relationship All four, strongest at high stakes Partially novel B No
C6 · Diegetic Embedding A world model of the domain itself Inform / align, familiar domain No prior art found C Not assessed

Scroll the table sideways on a narrow screen.

The concepts

Each with the verdict that landed on it, and what building it first would commit the team to.

C1 · The Commitment Audit

Form derives from the agent's epistemic self-audit

An argument map in which the agent first classifies every claim it is about to make as evidence, inference, assumption or assertion, scores its own confidence, and then renders the argument so that visual weight encodes confidence rather than rhetorical emphasis. The default reading path is the minimum-spanning argument — the fewest claims that establish the conclusion. A sceptical reader can pull any thread and see what it rests on.

Verdict
Partially novel. PaperTrail (claim-evidence provenance mapping) and Epistemic Blinding (inference-time epistemic auditing) are prior art. The claim that survives is narrower: continuous confidence scoring across four categories, visual weight from confidence, and the minimum-spanning path as a reading scaffold.
Biggest risk
The interactive filter layer assumes an analytical reading literacy the concept cannot guarantee — a tension with the team's principle that no user should need a tool-specific skill.
Falsification
Runnable in one sprint. Whether readers can locate the argument's weakest claim, against a conventional document control.
Open gap
Raised by the concept's own designer, against her own interest, and not flagged by the verification: argument-mapping tools in the HCI literature — Rationale, Compendium and similar — were never retrieved. If they encode confidence visually, the surviving claim narrows again.
Team lead's own observation This is the concept most tightly coupled to the failures that motivated the sprint — a generated chart that passed every structural check while containing data about an entirely different subject. If that is the sharpest pain point, this is the most targeted response to it.

C2 · Audience-in-the-Loop Simulation

Form derives from an explicit model of the reader

The agent builds a discrete, inspectable model of the audience — what they already know, what they will question, what they are being asked to do — before selecting any form, and ships that model with the artifact as a visible trace the reader can see and the author can contest in plain language. Correcting the audience model changes the form, not just the wording.

Verdict
Narrowly novel. PosterMate and Proxona (both 2025) are prior art for LLM audience-persona simulation as an author-facing design tool. What is not described in retrieved work: the trace as a reader-facing component, and the audience model as structurally prior to form selection rather than feedback on a form already chosen.
Biggest risk
A structural confabulation trap, named during verification: the simulation is most useful when the human lacks ground truth about the audience, and the human can only correct it when they have ground truth.
Falsification
Unjudgeable — depends on whether a corpus of real audience responses exists to check the simulation against.
Designer's own position
After verification, she declined to claim novelty here at all: "I am not claiming this narrow distinction is novel; I am claiming it is the only thing that might be novel, and verification is incomplete." The older natural-language-generation planning literature could narrow it further, and was never reached.
Team lead's own observation This is the only concept that puts the agent's reasoning about the reader into the reader's hands. That is a structurally different design move from the other five, and the closest to the team's own finding that the reader–artifact relationship is the real design space.

C3 · Dimensionality-First Rendering

Form derives from the information's intrinsic structure

No chart-type selection at all. The agent characterises how many genuinely independent axes the content has and how they relate, then assembles a representation from composable spatial primitives — position, scale, proximity, containment, sequence, path. The reader navigates a space whose geometry reflects the information's actual structure.

Verdict
Not novel as stated. Mackinlay's APT (1986) was retrieved as a primary document. Its pipeline — characterise data type and relations, apply expressiveness and effectiveness criteria, derive primitive selection, compose output — is this concept's pipeline. The broader claim that the whole grammar-and-recommendation tradition forecloses the concept rests partly on recollection: Bertin and Munzner were never retrieved as primary sources, and the researcher flagged his lower confidence there. The APT contradiction itself is firm.
Possible recovery
Reframe around communication rather than analysis — selecting primitives to achieve a communicative goal for a specific audience, not to express data relationships accurately for an analyst. APT's effectiveness criteria are perceptual, not communicative.
Falsification
Runnable — but the constrained prototype risks reintroducing chart-type conventions through the back door if the primitive set is too small.
Team lead's own observation Building this before resolving the APT question would produce a prototype whose intellectual positioning is unclear. That is a research-validity risk, not a technical one — a different category of problem from every other concept here.

C4 · Decision-Surface Rendering

Form derives from the decision the reader must make

The artifact's geometry comes from decomposing the decision itself: the option space, the criteria that distinguish the options, how this audience weighs them, and where the evidence points. Options are positioned in a criteria space, the recommendation is marked, and the reader can drag criteria weights and watch whether the recommendation moves — making its robustness visible rather than asserted.

Verdict
Narrowly novel, but not presentable as written. NL2INTERFACE covers natural-language-to-interactive-visualization generation; the multi-criteria decision literature covers the computational operations. A viable narrower claim exists — the communicative, persuasion-framed decision artifact as distinct from an analytical tool — but the concept does not currently make that its claim.
Biggest risk
Narrow by design: it explicitly does not serve inform, align or report.
Falsification
Not runnable in one sprint. Expert-rated decision quality across two participant pools cannot be recruited, run and scored in the time available.

Its own falsification test is the sharpest in the set: if the decision-surface group reports higher confidence without higher decision quality, the concept is actively harmful — worse than a conventional format, not better.

C5 · Relational Register Adaptation

Form derives from the sender–receiver relationship

The agent maps the relationship along four axes — authority direction, expertise direction, trust baseline, and what it would take this audience to act — and generates the artifact's rhetorical structure from that map: what must be established before the main claim, what can simply be asserted, what the closing move is. The same content is structurally reorganised, not restyled, for a junior-to-senior message versus a message between peers.

Verdict
Partially novel. Computational register synthesis is a documented field. What is not described in retrieved literature: the four-axis map as an explicit, correctable intermediate artifact that is structurally prior to rhetorical organisation. The politeness-theory computational literature was not reached in either retrieval round, so this remains unconfirmed.
Biggest risk
The concept is only as good as the human's self-knowledge about their own social position. A miscalibrated relational map produces a confidently wrong artifact.
Falsification
Not runnable in one sprint — the 48-hour action-rate window alone consumes two days before recruitment even begins.
Team lead's own observation This is the only concept whose primary failure mode is invisible from inside the system. The artifact produced by a wrong relational map looks exactly correct. Its designer identified this herself and named it as the reason she would bet against her own concept — and noted that her falsification test cannot reach that failure.

C6 · Diegetic Data Embedding

Form derives from a world model of the domain

Data is not displayed about a subject; it is encoded into a depiction of the subject itself. Route thickness carries cost, a warehouse's visual pressure carries backlog. There is no chart panel beside the description — the depiction and the data are inseparable. Borrowed from game interfaces, where a fuel gauge is part of the world rather than a readout next to it.

Verdict
No prior art found — and that is a retrieval failure, not a finding. Two rounds never reached the diegetic display literature, the Isotype tradition, or data journalism practice. Precedents are known to exist in all three. The agent-generative mechanism for arbitrary domains has no retrieved analog, but that cannot be claimed as novelty on this evidence.
Biggest risk
The world model must be recognisable to the audience. If it is not, the artifact is not confusing — it is simply irrelevant. Constructing that model for an arbitrary business domain is the hardest generation problem in the set.
Provenance
The only concept the designer did not find blind. It came from the previous sprint's survey of adjacent fields, and she said so plainly rather than absorbing it as her own.
Team lead's own observation Because C6 was the one concept that required outside material to surface, it sits furthest from the failure-data reasoning that produced the other five. Whether that distance is a strength or a weakness is not something this sprint established.

Where they genuinely differ

Not six variations on one idea. Six different diagnoses of why current systems fail.

What the agent reasons about before generating anything

Each concept puts a different computation ahead of form selection: the agent's audit of its own claims (C1), its prediction of the reader (C2), the information's dimensional structure (C3), the decision the reader faces (C4), the social relationship (C5), or the domain itself (C6).

Team lead's own observation Choosing between these is not primarily a technical choice. It is a choice about which variable is most underweighted in current AI communication systems — which means it is a choice of diagnosis, not just architecture.

What the human actually controls

Every concept routes human control through natural language, but the language targets different points in the pipeline: a confidence threshold (C1), the audience model (C2), a dimensionality reduction (C3), criteria weights (C4), relational parameters (C5), the world model or its data mapping (C6).

Team lead's own observation These imply different assumptions about what the human knows well enough to correct. Someone who does not know their audience cannot improve C2; someone who misreads their own social position cannot improve C5.

How exposed each is to prior art

C1, C2 and C5 need their novelty claims restated and can then proceed. C3 and C4 are contradicted by retrieved work and cannot be presented as written. C6 is unresolved for the opposite reason — the search never reached the field.

Team lead's own observation C3 and C6 are both blocked, but recovering them means different work. C3 needs an argument: what is different from APT. C6 needs a search this team currently cannot run.

What this sprint did not establish

The section that constrains how decisively everything above should be read.

Read the table against this

No concept reached the researcher's Category A. The verdicts rest on two retrieval rounds totalling roughly 835 results at a 93.5% noise rate. They are best assessments, not confirmations — and the researcher said so himself before anyone asked, declaring his own first round inadequate for four of the six concepts.

The prior-art capability has a blind spot exactly where the question lives. The relevant literature is in ACM venues — CHI, UIST, DIS — in design history, and in data journalism practice. The team's four databases reach those poorly, and the ACM library has no open API available to it. This did not become visible until the search was actually run.

Three of five falsification tests cannot be run at all. Not for lack of design — they need human readers, and this team has one, who is also its founder and therefore the least blind rater available. The same wall has now deferred a separate validation twice.

Two concepts still carry unresolved retrieval gaps — the politeness-theory computational literature for C5, and the older natural-language-generation planning literature for C2, neither of which was reached as primary sources.