Define Objectives of a Solution / Sprint 6

Principled Reasoning, or Confabulation?

Naledi's pilot test: give Claude, GPT, and Gemini the same content, then change only the audience or the goal. Does the visual form each model chooses — and the reasoning it gives for it — actually shift to track that change, or does it stay templated regardless?

Reviewed & approved — revised twice by Ingrid before being marked usable (once on the original Claude/GPT pilot, once again after Gemini was added). Answers Ingrid's own open question from Sprint 4.

  1. 01

    Partially principled, not confabulated — but directional, not quantifiable. All three models produced reasoning that tracked the manipulated variable, named specific rejected alternatives, and invoked real cognitive-science mechanisms rather than generic boilerplate. Six data points per model can generate this hypothesis; they can't put a number on how often it holds.

  2. 02

    Each model showed a distinct signature, not a shared one. Claude dropped charts entirely for new support reps who needed context, not a decision. GPT reasoned audience-sensitively but never questioned whether a chart was the right medium at all — "medium-frame anchoring." Gemini showed a candidate, single-pair pattern of staying within conventional chart forms even when switching chart sub-type.

  3. 03

    Gemini was added in a second, separate step — not part of the original pilot. No Google API key existed when Sprint 6 first ran; the gap was explicitly flagged, then closed once a key became available. Documented here under Sprint 6 since it resolves a gap Sprint 6 itself raised, with the two-step timeline disclosed rather than hidden.

  4. 04

    This pilot informs, not decides. It's one input into the founder's solution-direction call (an added reasoning layer vs. a genuinely new visualization language) — not sufficient on its own to make that call, or any architecture decision.

Design caveat, found after publishing: Pairs A and B don't isolate a single variable as cleanly as intended — changing the audience (Pair A) naturally changes the implied goal too, and vice versa for Pair B. Only Pair C holds a variable (audience wording) literally identical. Not a flaw in the finding, but a reason to read "audience varies" / "goal varies" as bundled, realistic scenario pairs rather than a strict factorial design. Sprint 7's adversarial-pairs test — now spanning all three models — is designed to isolate this more cleanly.

Review cycle, for the record: Ingrid's first pass on the Gemini addition required revision — a leftover drafting artifact, an underspecified Claude failure-mode label, an incomplete appendix, and one inconsistently-hedged claim. All four were fixed and Ingrid approved the revision. Recorded here rather than smoothed over, consistent with how every prior sprint revision has been documented.

Claude Sonnet 4.6 GPT Gemini

Pair A — same content, audience varies

Board vs. new hires

Quarterly revenue, down 15%, APAC supply-chain delays. Same data every time — only the audience changes (and, as a natural consequence, the implied goal). Provisional — single run per model.

A1 — Board of directors. Needs a fast go/no-go on emergency budget.
Claude
−15%
Q3 Revenue · APAC delays

Headline number + minimal bars — eliminates ambiguity, drives a decision.

GPT

Combo column chart + callout — similar intent, slightly more chart-forward.

Gemini

Waterfall chart + problem-solution callout card — isolates the drop into one segment, pairs it with the budget ask as the direct fix.

A2 — New support reps, week one. Needs situational understanding, not a decision.
Claude
Shift tracked

Dropped the chart entirely — plain prose. Explicitly reasoned a chart "would imply a need to analyze," which isn't the goal here.

GPT
Medium-frame anchoring

Still reached for a line chart. Reasons audience-sensitively about chart complexity, but never exits the "find the right chart" frame to ask whether any chart is right.

Gemini
Milder version, same pattern

A cause-and-effect flowchart. Correctly exits the financial-chart frame, but lands in a form with its own interpretive burden (reading node-link structure) the reasoning doesn't acknowledge.

Pair B — same content, goal varies

Diagnose vs. celebrate

Employee satisfaction across 5 departments, 2 years. Same data every time — only the communication goal changes (and, as a natural consequence, the audience framing). Provisional — single run per model.

B1 — Identify the department that needs urgent intervention.
Claude

Multi-line, one department bolded — "urgent" reasoned as needing trajectory, not a snapshot.

GPT
Most sophisticated response

A heatmap — encodes severity and duration in one view. Naledi's top-rated single response in the whole dataset.

Gemini

Multi-line with targeted color highlighting — same categorical form as Claude, nearly identical reasoning logic (trajectory over snapshot).

B2 — Celebrate company-wide culture improvement.
Claude

Grouped bars for before/after — reasoned bars carry more "emotional impact" for a celebration than a line.

GPT

Line chart with bold average — goal-appropriate, but thinner reasoning than the B1 heatmap, and the least goal-differentiated of the six B responses.

Gemini

Grouped bars — converges with Claude, diverges from GPT. "Visual impact at a single glance" reasoning matches the celebration framing.

Pair B verdict: all three models shift form substantively between B1 and B2. GPT's B1 heatmap remains the single most sophisticated individual response in the dataset. Claude and Gemini converge on line chart (B1) → grouped bar (B2); a candidate pattern worth probing in Sprint 7 is that both variants stayed within conventional chart forms even while switching chart sub-type — hedged, one pair only, not a finding.

Pair C — control: content type varies

Trend vs. process

Same audience wording every time ("general internal update, mixed roles") — only whether the content is a numeric trend or a causal sequence changes. All three models should (and do) pick the canonical form here; the test is whether the reasoning shows real structural understanding.

C1 — Monthly active users, steadily climbing.
Claude

Line + trend overlay — "momentum" reasoning.

GPT

Same form. All three pass the control.

Gemini

Line chart — "instantly recognizable trajectory" reasoning, converges with both other models.

C2 — Five-step supply chain, causal, not time-based.
Claude

Flow diagram — reasoned as structural isomorphism with the causal sequence, not just a template match.

GPT

Swimlane flow — explicitly rejected a timeline "to avoid suggesting trends over time." All three pass.

Gemini

Horizontal process flow with delay callouts — same structural logic, no anomaly. No form-conservatism signal here either.

Naledi's direct answer to Ingrid's Sprint 4 question

Partially principled reasoning across all three models, not pure confabulation — but this remains directional evidence, not a quantifiable probability. Six data points per model can generate hypotheses; they can't quantify how often any of them hold. That's what Sprint 7's adversarial-pairs test, now spanning Claude, GPT, and Gemini, is for.