Naledi's pilot test: give Claude, GPT, and Gemini the same content, then change only the audience or the goal. Does the visual form each model chooses — and the reasoning it gives for it — actually shift to track that change, or does it stay templated regardless?
Partially principled, not confabulated — but directional, not quantifiable. All three models produced reasoning that tracked the manipulated variable, named specific rejected alternatives, and invoked real cognitive-science mechanisms rather than generic boilerplate. Six data points per model can generate this hypothesis; they can't put a number on how often it holds.
Each model showed a distinct signature, not a shared one. Claude dropped charts entirely for new support reps who needed context, not a decision. GPT reasoned audience-sensitively but never questioned whether a chart was the right medium at all — "medium-frame anchoring." Gemini showed a candidate, single-pair pattern of staying within conventional chart forms even when switching chart sub-type.
Gemini was added in a second, separate step — not part of the original pilot. No Google API key existed when Sprint 6 first ran; the gap was explicitly flagged, then closed once a key became available. Documented here under Sprint 6 since it resolves a gap Sprint 6 itself raised, with the two-step timeline disclosed rather than hidden.
This pilot informs, not decides. It's one input into the founder's solution-direction call (an added reasoning layer vs. a genuinely new visualization language) — not sufficient on its own to make that call, or any architecture decision.
Design caveat, found after publishing: Pairs A and B don't isolate a single variable as cleanly as intended — changing the audience (Pair A) naturally changes the implied goal too, and vice versa for Pair B. Only Pair C holds a variable (audience wording) literally identical. Not a flaw in the finding, but a reason to read "audience varies" / "goal varies" as bundled, realistic scenario pairs rather than a strict factorial design. Sprint 7's adversarial-pairs test — now spanning all three models — is designed to isolate this more cleanly.
Review cycle, for the record: Ingrid's first pass on the Gemini addition required revision — a leftover drafting artifact, an underspecified Claude failure-mode label, an incomplete appendix, and one inconsistently-hedged claim. All four were fixed and Ingrid approved the revision. Recorded here rather than smoothed over, consistent with how every prior sprint revision has been documented.
Quarterly revenue, down 15%, APAC supply-chain delays. Same data every time — only the audience changes (and, as a natural consequence, the implied goal). Provisional — single run per model.
Headline number + minimal bars — eliminates ambiguity, drives a decision.
Combo column chart + callout — similar intent, slightly more chart-forward.
Waterfall chart + problem-solution callout card — isolates the drop into one segment, pairs it with the budget ask as the direct fix.
Dropped the chart entirely — plain prose. Explicitly reasoned a chart "would imply a need to analyze," which isn't the goal here.
Still reached for a line chart. Reasons audience-sensitively about chart complexity, but never exits the "find the right chart" frame to ask whether any chart is right.
A cause-and-effect flowchart. Correctly exits the financial-chart frame, but lands in a form with its own interpretive burden (reading node-link structure) the reasoning doesn't acknowledge.
Employee satisfaction across 5 departments, 2 years. Same data every time — only the communication goal changes (and, as a natural consequence, the audience framing). Provisional — single run per model.
Multi-line, one department bolded — "urgent" reasoned as needing trajectory, not a snapshot.
A heatmap — encodes severity and duration in one view. Naledi's top-rated single response in the whole dataset.
Multi-line with targeted color highlighting — same categorical form as Claude, nearly identical reasoning logic (trajectory over snapshot).
Grouped bars for before/after — reasoned bars carry more "emotional impact" for a celebration than a line.
Line chart with bold average — goal-appropriate, but thinner reasoning than the B1 heatmap, and the least goal-differentiated of the six B responses.
Grouped bars — converges with Claude, diverges from GPT. "Visual impact at a single glance" reasoning matches the celebration framing.
Pair B verdict: all three models shift form substantively between B1 and B2. GPT's B1 heatmap remains the single most sophisticated individual response in the dataset. Claude and Gemini converge on line chart (B1) → grouped bar (B2); a candidate pattern worth probing in Sprint 7 is that both variants stayed within conventional chart forms even while switching chart sub-type — hedged, one pair only, not a finding.
Same audience wording every time ("general internal update, mixed roles") — only whether the content is a numeric trend or a causal sequence changes. All three models should (and do) pick the canonical form here; the test is whether the reasoning shows real structural understanding.
Line + trend overlay — "momentum" reasoning.
Same form. All three pass the control.
Line chart — "instantly recognizable trajectory" reasoning, converges with both other models.
Flow diagram — reasoned as structural isomorphism with the causal sequence, not just a template match.
Swimlane flow — explicitly rejected a timeline "to avoid suggesting trends over time." All three pass.
Horizontal process flow with delay callouts — same structural logic, no anomaly. No form-conservatism signal here either.
Partially principled reasoning across all three models, not pure confabulation — but this remains directional evidence, not a quantifiable probability. Six data points per model can generate hypotheses; they can't quantify how often any of them hold. That's what Sprint 7's adversarial-pairs test, now spanning Claude, GPT, and Gemini, is for.