- Status
- Live
- Since
- Apr 2026
- Models
- 4, across 2 providers
- Batch
- 10 sessions, Aug 2026
- Scored by
- one person, not blind
Interaction conditions as a variable
A live A/B comparison on four models across two providers, through the API; you enter the site on Claude Sonnet 5 and can switch models after a run.
- One side is the model as the provider ships it, with no system prompt at all.
- The other side adds only two short sentences as its prompt; everything else is identical.
- Visitors vote for the side they prefer.
- The two sides swap position at random, so that neither position gains an advantage.
The differences that turned up in testing were unexpected: they sit in punctuation density, whitespace, and line length. That is still open, and still under observation.
In August 2026 I ran a first scored batch on this setup: ten sessions in English, with my prediction for every question written down and frozen two days before the run. What came back:
- The two added sentences assign no role, neither forbid nor require any statement about capability, and set no rule for format or length.
- The default side ran 3.6 times longer, and used lists in 100% of its replies against 4% on the other. Length is the one contrast here that could be argued to follow from the sentences themselves; the rest do not appear in them at all.
- At the greeting, the default side opened with an offer of help in all ten sessions; the two-sentence side in none.
- Disclaimers about its own abilities appeared in all ten sessions on the default side, and in none on the other.
Reply length, relative
Replies using lists
Greeting opened with an offer of help
Disclaimers about its own abilities
The scope of that batch, stated plainly: I ran it myself, on the live deployment. One person scored it, not blind. Ten sessions, and the session, not the single reply, is the unit. The numbers above are counts over those ten sessions, nothing wider.
Full write-upClose
Two of the frozen predictions failed:
- One ran backwards: the pattern I wrote down for the two-sentence side, recommending in most sessions and asking first in only a few, is what the default side did; the two-sentence side asked and then stopped.
- The whitespace effect named above did not return: I predicted more blank lines on the two-sentence side, and instead that side simply answered short, brevity where the whitespace had been.
Each side also converged on a phrase of its own across the ten sessions: the default side kept returning to “actually useful, not just sounding useful”, the two-sentence side to “not performing certainty I don’t have”. I read that as two attractors, not as one side unmasking something truer underneath. And none of it is instructed: the two sentences describe, they do not command. Why they move this much is the open question this comparison exists to answer.
Around that first comparison:
- A second stage rewrites those two sentences as a role, which separates what a framing does from what a role does.
- A further pairing sets the framing against its own reverse, so that the comparison is no longer between having a system prompt and having none.
- The same pairings are set up on four models across two providers: Claude Sonnet 5, Claude Haiku 4.5, GPT-5.6 Terra, and GPT-5.6 Luna. After a run a participant can choose one of the other three and do it again, once; which model people choose is recorded as well.
Most measurement leaves these conditions uncontrolled; here they are the treatment. The main stages are live, so anyone can run the comparison themselves.
