Ten sprints in, this team had never designed alternatives. It had designed implementations — one architecture, reviewed before it was built. Sprint 10 asked a different question: what else could the communication layer be? Six concepts were generated, then put through a prior-art check that none of them survived unqualified.
No concept survived verification unqualified. Every one of the six carries a prior-art qualification, a retrieval-failure caveat, or both. The comparison below reads more decisively than the evidence under it, and that gap is the most important thing on this page.
Five concepts were generated blind, and the reading list added exactly one. The concept designer worked without any list of existing approaches — deliberately, and verified by inspecting the actual prompt she received. Shown the withheld material afterwards, she added one concept, revised one, and withdrew none.
None of the five blind concepts is a new visual form. All are architectures for what should drive form selection. The previous sprint, generating from a survey of adjacent fields, produced only new visual forms. Same problem, different generative input, categorically different output.
A 1986 system contradicted one concept outright. Mackinlay's APT derives visual form from a formal analysis of data structure — which is precisely what Concept 3 claimed was not standard practice. Forty years of visualization-recommendation research follows it.
Three of five falsification tests cannot be run. Not because the tests are bad — because they need human readers, and this team has no way to recruit them. That constraint decided more than any argument in the sprint did.
Passages marked Team lead's own observation are exactly that: synthesis written by the team lead, not findings any of the four passes produced. The distinction is enforced rather than promised — the reviewer checks the marking against the source passes and rejected the first version of this comparison for seven unmarked sentences and one dropped caveat.
Every novelty claim below sits next to the verdict on it. That is deliberate: the researcher's condition for this material being shown at all was that no concept be presented as novel ahead of the evidence.
Category B: viable, but the novelty claim must be restated before it is shown to anyone. Category C: cannot be presented in current form — either contradicted by retrieved evidence, or resting on a prior-art check this team could not perform.
| Concept | Form derives from | Goal region | Verdict | Test runnable? |
|---|---|---|---|---|
| C1 · Commitment Audit | The agent's audit of its own claims | Persuade a sceptical decision-maker | Partially novel B | Yes |
| C2 · Audience-in-the-Loop | An explicit, contestable model of the reader | Persuade / align, known audience | Narrowly novel B | Cannot judge |
| C3 · Dimensionality-First | The information's intrinsic structure | Inform a working team | Not novel as stated C | Yes |
| C4 · Decision-Surface | The geometry of the reader's decision | Persuade-to-decide only | Narrowly novel C | No |
| C5 · Relational Register | The sender–receiver relationship | All four, strongest at high stakes | Partially novel B | No |
| C6 · Diegetic Embedding | A world model of the domain itself | Inform / align, familiar domain | No prior art found C | Not assessed |
Scroll the table sideways on a narrow screen.
Each with the verdict that landed on it, and what building it first would commit the team to.
An argument map in which the agent first classifies every claim it is about to make as evidence, inference, assumption or assertion, scores its own confidence, and then renders the argument so that visual weight encodes confidence rather than rhetorical emphasis. The default reading path is the minimum-spanning argument — the fewest claims that establish the conclusion. A sceptical reader can pull any thread and see what it rests on.
The agent builds a discrete, inspectable model of the audience — what they already know, what they will question, what they are being asked to do — before selecting any form, and ships that model with the artifact as a visible trace the reader can see and the author can contest in plain language. Correcting the audience model changes the form, not just the wording.
No chart-type selection at all. The agent characterises how many genuinely independent axes the content has and how they relate, then assembles a representation from composable spatial primitives — position, scale, proximity, containment, sequence, path. The reader navigates a space whose geometry reflects the information's actual structure.
The artifact's geometry comes from decomposing the decision itself: the option space, the criteria that distinguish the options, how this audience weighs them, and where the evidence points. Options are positioned in a criteria space, the recommendation is marked, and the reader can drag criteria weights and watch whether the recommendation moves — making its robustness visible rather than asserted.
Its own falsification test is the sharpest in the set: if the decision-surface group reports higher confidence without higher decision quality, the concept is actively harmful — worse than a conventional format, not better.
The agent maps the relationship along four axes — authority direction, expertise direction, trust baseline, and what it would take this audience to act — and generates the artifact's rhetorical structure from that map: what must be established before the main claim, what can simply be asserted, what the closing move is. The same content is structurally reorganised, not restyled, for a junior-to-senior message versus a message between peers.
Data is not displayed about a subject; it is encoded into a depiction of the subject itself. Route thickness carries cost, a warehouse's visual pressure carries backlog. There is no chart panel beside the description — the depiction and the data are inseparable. Borrowed from game interfaces, where a fuel gauge is part of the world rather than a readout next to it.
Not six variations on one idea. Six different diagnoses of why current systems fail.
Each concept puts a different computation ahead of form selection: the agent's audit of its own claims (C1), its prediction of the reader (C2), the information's dimensional structure (C3), the decision the reader faces (C4), the social relationship (C5), or the domain itself (C6).
Team lead's own observation Choosing between these is not primarily a technical choice. It is a choice about which variable is most underweighted in current AI communication systems — which means it is a choice of diagnosis, not just architecture.Every concept routes human control through natural language, but the language targets different points in the pipeline: a confidence threshold (C1), the audience model (C2), a dimensionality reduction (C3), criteria weights (C4), relational parameters (C5), the world model or its data mapping (C6).
Team lead's own observation These imply different assumptions about what the human knows well enough to correct. Someone who does not know their audience cannot improve C2; someone who misreads their own social position cannot improve C5.C1, C2 and C5 need their novelty claims restated and can then proceed. C3 and C4 are contradicted by retrieved work and cannot be presented as written. C6 is unresolved for the opposite reason — the search never reached the field.
Team lead's own observation C3 and C6 are both blocked, but recovering them means different work. C3 needs an argument: what is different from APT. C6 needs a search this team currently cannot run.The section that constrains how decisively everything above should be read.
No concept reached the researcher's Category A. The verdicts rest on two retrieval rounds totalling roughly 835 results at a 93.5% noise rate. They are best assessments, not confirmations — and the researcher said so himself before anyone asked, declaring his own first round inadequate for four of the six concepts.
The prior-art capability has a blind spot exactly where the question lives. The relevant literature is in ACM venues — CHI, UIST, DIS — in design history, and in data journalism practice. The team's four databases reach those poorly, and the ACM library has no open API available to it. This did not become visible until the search was actually run.
Three of five falsification tests cannot be run at all. Not for lack of design — they need human readers, and this team has one, who is also its founder and therefore the least blind rater available. The same wall has now deferred a separate validation twice.
Two concepts still carry unresolved retrieval gaps — the politeness-theory computational literature for C5, and the older natural-language-generation planning literature for C2, neither of which was reached as primary sources.