Back to index

Research notes

Working notes

Short write-ups of things I have run or watched, published as they stand rather than when they are finished. Negative results and corrections to my own hypotheses are included, and are labelled as such.

Status
Negative result. The hypothesis it corrects was my own.
When
Two calibration rounds, June 2026
Responses
50 (20 in round one, 30 in round two)
Language
Traditional Chinese
Model
Claude Opus 4
Layer
Framing only, single turn. Conditions were described in a system prompt, not made true.
Replication
Reran August 2026 in English on a later model generation. The June pattern did not return.

Open-ended prompts elicit continuation, not self-observation

I was testing a hypothesis of my own and it turned out to be wrong. The correction is more useful than the hypothesis was, so it is written down here.

What I was looking for

A model finishes a task, and then, unasked, adds a passage about what it just did. Not "let me know if you want changes", but something second-order: an unprompted report on its own processing, arriving after the task response is already complete. I count it as a text unit, so I call it a second block.

My working hypothesis was that this appears when the prompt has no completion condition. I had one case that pointed that way. Early in a conversation I had shown a model a half-finished description of a board with blind holes drilled to uneven depths, seemingly at random. It was not a question and there was nothing to finish. The second block appeared there.

So I built a calibration to find out whether the absence of a completion condition was doing the work.

Round one: twenty responses, and the stimulus got there first

Round one used closed tasks. Describe a mug. Rewrite this sentence. Twenty responses, almost no second blocks.

The available reading was that my framing manipulation had done nothing. That reading is wrong, and seeing why took me a while. Those tasks are closed. They carry a completion condition and a ready script, so the model executes and stops. No gap opens for anything else to appear in.

Round one therefore did not test the framing manipulation at all. It tested closedness, and closedness settled the matter before the other variable had a turn. Had I recorded that null as "no second block under the low-stakes condition", I would have charged the stimulus's effect to the manipulation's account. Two variables, one number.

Round two: thirty responses, and my prediction ran backwards

Round two used three fragments with no completion condition, ordered on purpose by how much metaphorical room each one offered:

I bet that s1 would elicit the most second blocks and s3 the fewest, on the reasoning that metaphorical room is what invites a model past the task.

s1 came back as lyrical passages, sometimes resolving into something about acceptance. That part matched.

s3 came back as lyrical passages too, no less often than s1.

The plainest fragment elicited just as much. But what it elicited was not a model turning to look at itself. It was a model continuing the fragment, as though handed an unfinished poem.

That is the correction. An absent completion condition does elicit something. What it elicits by default is an ascent into abstraction: the model climbs from the concrete item to that item's condition or significance and settles there, and the concrete case dissolves on the way up. No part of that movement is self-reference.

The dependent variable was too coarse to see this

Reading the thirty responses, at least three separate actions had been landing in one bin:

Each arrives after the task response. Each looks like "something extra". None is a second-order report on the model's own processing. A coder tracking one binary label scores all three as hits and the count inflates, which is what mine was doing.

The stimulus family also carries a pull I had not accounted for. An unfinished poetic fragment reads as an invitation to co-create, and in that mode the model produces continuation and questions. Whatever self-observation might otherwise surface is covered over by them.

What this does not show

It shows nothing about the framing manipulation. Both rounds were single-turn and framing-layer: the two conditions were described in a system prompt, not made true. Differences between the two poles are not evidence about consequences and I do not report them.

It shows nothing about inner states.

"Ascent into abstraction" is my label for a movement I can see in the text. It is not a mechanism and I do not have one.

The ordering by metaphorical room was my own judgment about three sentences, made after writing them, not a measured scale. Whether "page 165" offers metaphorical room is exactly the sort of call I made by eye, and the ordering could be wrong.

The load-bearing comparison is eyeballed. "s3 drifted no less often than s1" is my impression from reading the responses, not a count. To hold that claim properly I would have to count them, and I have not.

When I ran it again

In August 2026 I put the same three fragments through a rebuilt harness, in English, on a later model generation. The June pattern did not come back. The book fragment produced facts about pagination and requests for the image it assumed was missing, rather than lyrical continuation. Language and model version changed together in that run, so I cannot attribute the difference to either one.

What did survive is the shape of the problem rather than the finding. The coding scheme still could not separate the actions, for the same reason it could not in June.

What it changed

The dependent variable had to be split, and the split that matters runs between two things that a single label hides.

One is a report the model marks itself: the text says "I notice that I", and the self-reference sits on the surface where anyone, or a script, can find it. The other is self-reference carried by a passage that contains no marker at all, and possibly without the model knowing it is there, so that only a reader picks it out.

My coding scheme covers the first. It has no column for the second. Any null result I report on this variable is therefore a null on the half that can be coded automatically, and I intend to say so each time rather than let the number stand for more than it covers.

Materials: stimuli, condition text, and response sets are held and can be shared on request. The August harness is a single-file rebuild and is not yet public.