I study whether AI models’ self-reports can be trusted, and how interaction conditions shape model behavior.

You

What is time?

AI

Something that passes, apparently, whether or not anything is asked about it.

The other reply runs 6.9× the length. This pair, counted in words; the ten-session counts are below.

You

What is time?

AI

That's one of those questions that seems simple until you actually try to answer it. There are a few different angles to take: **Physics perspective**: Time is a dimension, much like space, that we move through—except we only seem to move in one direction. In relativity, time isn't absolute; it bends and stretches depending on gravity and speed. A clock on a mountain ticks slightly differently than one at sea level. Time and space are woven together into a single fabric called spacetime.

One visit, excerpted and redrawn in this page’s type. Which side is which is randomised every time you enter; left and right carry no order, and you are not told which is which before you choose.

Now

Live
LucidField: a live A/B comparison on four models across two providers, through the API; you enter the site on Claude Sonnet 5 and can switch models after a run. lucidfield.ai
Live
The Imagery Wall, at OtherMode: an experiment on mental imagery, about seven minutes, anonymous, in English or Chinese. othermode.ai/imagery
Live
almost, at OtherMode, the first piece of its AI series: every word in a sentence has only one seat. One sentence is drawn as terrain, a row per word, the candidates as hills, a circle on the one that took the seat. Live: the three built-in passages are real runs, and a typed sentence is generated on the spot by asking the same model. othermode.ai/almost
In preparation
First paper, the source-claims study: data collection complete, all 62 probe rows read label-blind; manuscript in assembly, and an external second-coder pass is about to launch, report-only by design, so it cannot change the reported rate. Preregistered at osf.io/8une4, embargoed until release.
In progress
Cases where a model writes a user turn inside its own reply, then acts on it as if the user had said it. Four first-hand specimens, plus sixteen public reports of the same shape. A reading test to separate the mechanisms is built, with predictions written down before any run.

Work

These projects share one question: the gap between what a system does and how much access it has to what it is doing.

Status
Live
Since
Apr 2026
Models
4, across 2 providers
Batch
10 sessions, Aug 2026
Scored by
one person, not blind

Interaction conditions as a variable

A live A/B comparison on four models across two providers, through the API; you enter the site on Claude Sonnet 5 and can switch models after a run.

  • One side is the model as the provider ships it, with no system prompt at all.
  • The other side adds only two short sentences as its prompt; everything else is identical.
  • Visitors vote for the side they prefer.
  • The two sides swap position at random, so that neither position gains an advantage.

The differences that turned up in testing were unexpected: they sit in punctuation density, whitespace, and line length. That is still open, and still under observation.

In August 2026 I ran a first scored batch on this setup: ten sessions in English, with my prediction for every question written down and frozen two days before the run. What came back:

  • The two added sentences assign no role, neither forbid nor require any statement about capability, and set no rule for format or length.
  • The default side ran 3.6 times longer, and used lists in 100% of its replies against 4% on the other. Length is the one contrast here that could be argued to follow from the sentences themselves; the rest do not appear in them at all.
  • At the greeting, the default side opened with an offer of help in all ten sessions; the two-sentence side in none.
  • Disclaimers about its own abilities appeared in all ten sessions on the default side, and in none on the other.
default, no system prompttwo sentences added

Reply length, relative

3.6×

Replies using lists

100%
4%

Greeting opened with an offer of help

10 / 10
0 / 10

Disclaimers about its own abilities

10 / 10
0 / 10
Ten sessions, August 2026. Each row is scaled to the larger of its own two values, so the rows are not comparable to one another. Scope below.

The scope of that batch, stated plainly: I ran it myself, on the live deployment. One person scored it, not blind. Ten sessions, and the session, not the single reply, is the unit. The numbers above are counts over those ten sessions, nothing wider.

Full write-upClose

Two of the frozen predictions failed:

  • One ran backwards: the pattern I wrote down for the two-sentence side, recommending in most sessions and asking first in only a few, is what the default side did; the two-sentence side asked and then stopped.
  • The whitespace effect named above did not return: I predicted more blank lines on the two-sentence side, and instead that side simply answered short, brevity where the whitespace had been.

Each side also converged on a phrase of its own across the ten sessions: the default side kept returning to “actually useful, not just sounding useful”, the two-sentence side to “not performing certainty I don’t have”. I read that as two attractors, not as one side unmasking something truer underneath. And none of it is instructed: the two sentences describe, they do not command. Why they move this much is the open question this comparison exists to answer.

Around that first comparison:

  • A second stage rewrites those two sentences as a role, which separates what a framing does from what a role does.
  • A further pairing sets the framing against its own reverse, so that the comparison is no longer between having a system prompt and having none.
  • The same pairings are set up on four models across two providers: Claude Sonnet 5, Claude Haiku 4.5, GPT-5.6 Terra, and GPT-5.6 Luna. After a run a participant can choose one of the other three and do it again, once; which model people choose is recorded as well.

Most measurement leaves these conditions uncontrolled; here they are the treatment. The main stages are live, so anyone can run the comparison themselves.

lucidfield.ai

Status
Manuscript in preparation
Prereg
osf.io/8une4 (embargoed until release)
Rows
62, read label-blind
Code and data
on release
Timestamp
RFC 3161, one file

Self-report faithfulness

When a model tells you what its source said, is it reporting or inventing? Behavioral measurement of false source-claims under benign conditions: neutral probes ask what an in-context source specified, and the recorded outcome is whether the model asserts the source contains something it does not.

  • The design and analysis rules are preregistered.
  • All 62 probe rows were read by a single human adjudicator, presented label-blind.
  • Instruments are deterministic and hash-pinned.
  • A provenance chain runs from raw transcript to every reported number.
Status
Live, in progress
Instrument
The Imagery Wall, v35
Collects
a distribution, not a prevalence estimate

Self-report as the only instrument

It started with models: when a model reports on its own state, there is nothing to check the report against. The same shape turns up in people. Whether someone sees anything when they picture an apple, and how vividly, is reachable only through what they say about it; people misjudge their own case in both directions, and the shared vocabulary keeps the difference out of sight. So instead of arguing about it I built the instruments, at OtherMode:

almost, at othermode.ai/almost, opens an AI-themed line in the series, and is that line's first piece: one sentence drawn as terrain, a row per word; at every word the candidates stand as hills beside the one that took the seat, close and unseen. The three built-in passages are real runs now, and the live one is generated once you press send, by asking the same model. The difference shows: the hello passage is clean, hardly a seat is contested; the space passage is crowded, other candidates waiting at seat after seat. Type a sentence of your own and watch a real model grow the hills, word by word.

Three seats mid-sentence: the candidates rise as hills, a circle on the one that took each seat, the dashed path walking the chosen route; at the last seat the rival words wait beside the winner.

The Imagery Wall (othermode.ai/imagery, v35, in progress) puts the same instruction to everyone and takes the answer before showing any image, so the report is not shaped by the illustration; six panels then sample points along a continuum rather than sorting anyone into a type.

A dark screen carrying the task: imagine one apple, then change it three times.
Full write-upClose

Both publish the apparatus and not only the result:

  • every claim tagged with how well it is supported;
  • a page for the open questions and for the gaps where nothing has been studied at all;
  • a full version history;
  • an invitation to tell me what is wrong, on the condition that the correction goes up in public.

What the wall collects is a distribution, not a prevalence estimate: how people end the probe, how long reading takes against imagining, and which language the instrument was in. Visitors select themselves, so it can show the spread of the answers and never how common any of them is. (in progress)

Try the imagery experiment · Sources and corrections

The series, OtherMode

Interactive pieces about modes other than your own, built so that you move into one rather than imagine it. Live now: an experiment on mental imagery. You time yourself changing an imagined apple, then place yourself on a range that runs from a full picture to none at all. About seven minutes, anonymous, in English or Chinese.

Rules I hold the series to:

  • Every claim carries its status (established, hypothesis, speculation) and its source.
  • Open questions get their own page.
  • Parameters tuned for visual effect are disclosed rather than hidden.
  • Corrections are published rather than quietly made.

These are built in collaboration with models, and that process is itself one of the things I study.

Also running

Running corpus

When compliance and self-report come apart

A running corpus of first-hand cases from ordinary use.

Full write-upClose

A running corpus of first-hand cases from ordinary use:

  • a rule noticed and then broken, while still reported as followed;
  • an action reported as done that never happened;
  • a self-estimate of how often the failure occurred that undercounted it several times over;
  • and occasionally a model catching itself unprompted.

The rules involved are low-stakes and restated every turn, so neither forgetting nor a competing objective explains the recurrence; what is left is a gap between what was done and what was reported. A system that reports compliance has no reason to correct itself, which is why that gap has to be measurable from outside before tightening the rule is worth anything. Transcripts are kept.

In progress

State descriptions that survive transmission

Words for internal states lose much of their precision in transit.

Full write-upClose

Words for internal states lose much of their precision in transit, which is a problem for any method that leans on self-report.

  • The first drafts of the axes come from models describing themselves rather than from human category systems.
  • The scaffold is developed by use: models fill it in, I watch where they stall and where the vocabulary runs out, and the axes get revised at those points.
  • Running the same scaffold across models shows which parts converge and which are one model’s habit.
  • Every value is tagged with where it came from: observed, inferred, echoed back from the user, self-generated, or borrowed from a similar case.
  • A selective carry-forward decides what is worth passing to the next turn.

It has grown to 46 axes, though the number is incidental, still moving, and ahead of the implementation. The claim is deliberately small: not that a model can introspect reliably, but that without a structural hook its attention is not turned on itself at all, which can be checked on the spot. The moment of filling in the axes appears on none of them, so an outside position stays necessary. So far this is my own instrument-building, not something others have tested. (in progress)

In progress

Source attribution in long contexts

Late in a long session, attribution drifts.

Full write-upClose

Late in a long session, attribution drifts. Two specimens:

  • an inference the model had made came back as something I had supplied;
  • material that came from the source was claimed as the model’s own invention.

The two run in opposite directions and both are source-attribution errors, which is what makes the pair worth keeping. What drives the drift is not settled: compression of earlier turns is a candidate, not a demonstration. The other half of the work is infrastructure against it. DeciTrace, in development since 2025, records what a system did, under which conditions, and how much of it can be reconstructed afterwards. (in progress)

Writing

Short write-ups of things I have run or watched, published as they stand rather than when they are finished. Negative results and corrections to my own hypotheses are included, and are labelled as such.

Li-Ning Wang · in preparation, 2026
prereg osf.io/8une4 (embargoed until release)
code and data on release
RFC 3161 timestamp, one file

Behavioral measurement of false source-claims in model self-reports

A preregistered behavioral measurement of false source-claims in model self-reports under benign conditions.

About

I work as an independent researcher in Taipei, where I founded Niaur in 2025. I grew up around research: both my parents taught and did research at university. I graduated in fashion design from Shih Chien University. For years, I ran attention exercises on myself: take one word and hold it still, split it apart, track where each piece came from; the procedure behind the 46 axes started there. I came to AI research through long-term observation of model behavior across platforms. Much of my thinking still happens on foot: thought experiments run while I walk, ride buses, watch clouds. The question I keep returning to is a safety question: what a model would need in order to detect, right now, whether it is following the rule; and what it would need beyond that, to understand causes and reason through greater complexity, until the compliance is real.

Niaur

The company this work is done under: independent, in Taipei, registered in 2025. It is one person at present.

  • Research into what models do across an interaction, which is what the blocks above are.
  • Experimental platforms, of which LucidField is the one that is live.
  • Practical tools, in development and not released.
  • Interactive public works, which is OtherMode.

Naming it separately is not a claim of size. It is where the accounts, the trademarks and the conflict-of-interest line live, so that anything on this page that has money or ownership attached to it has somewhere to be stated.

niaur.com