Li-Ning Wang

Li-Ning Wang

Independent researcher at Niaur, Taipei. I study whether AI models’ self-reports can be trusted, and how interaction conditions shape model behavior.

News

  • Aug 2026 · LucidField adds a cross-model arm: the same pairings on four models across two providers, with the participant choosing which one to run next.
  • Aug 2026 · Source-claims study: data collection complete; all 62 probe rows read label-blind. Manuscript in preparation, targeted for September.
  • Aug 2026 · OtherMode publishes its first interactive piece: an imagery experiment, live at othermode.ai/imagery.
  • Apr 2026 · LucidField goes live: paired A/B comparisons of interaction framing.

Publications

Behavioral measurement of false source-claims in model self-reports

Li-Ning Wang · in preparation, 2026

A preregistered behavioral measurement of false source-claims in model self-reports under benign conditions.

prereg osf.io/8une4 (embargoed until release) · code and data on release · RFC 3161 timestamped

Research notes

Short write-ups of things I have run or watched, published as they stand rather than when they are finished. Negative results and corrections to my own hypotheses are included, and are labelled as such.

Open-ended prompts elicit continuation, not self-observation · Jun 2026, negative result

These projects share one question: the gap between what a system does and how much access it has to what it is doing.

Research

Self-report faithfulness

When a model tells you what its source said, is it reporting or inventing?

When a model tells you what its source said, is it reporting or inventing? Behavioral measurement of false source-claims under benign conditions: neutral probes ask what an in-context source specified, and the recorded outcome is whether the model asserts the source contains something it does not.

  • The design and analysis rules are preregistered.
  • All 62 probe rows were read by a single human adjudicator, presented label-blind.
  • Instruments are deterministic and hash-pinned.
  • A timestamped provenance chain runs from raw transcript to every reported number.

When compliance and self-report come apart

A running corpus of first-hand cases from ordinary use.

A running corpus of first-hand cases from ordinary use:

  • a rule noticed and then broken, while still reported as followed;
  • an action reported as done that never happened;
  • a self-estimate of how often the failure occurred that undercounted it several times over;
  • and occasionally a model catching itself unprompted.

The rules involved are low-stakes and restated every turn, so neither forgetting nor a competing objective explains the recurrence; what is left is a gap between what was done and what was reported. A system that reports compliance has no reason to correct itself, which is why that gap has to be measurable from outside before tightening the rule is worth anything. Transcripts are kept.

Interaction conditions as a variable

A live A/B comparison; you enter the site on Claude Sonnet 5, through the API.

A live A/B comparison; you enter the site on Claude Sonnet 5, through the API.

  • One side is the model as the provider ships it, with no system prompt at all.
  • The other side adds only two short sentences as its prompt; everything else is identical.
  • Visitors vote for the side they prefer.
  • The two sides swap position at random, so that neither position gains an advantage.

The differences that turned up in testing were unexpected: they sit in punctuation density, whitespace, and line length. That is still open, and still under observation.

Around that first comparison:

  • A second stage rewrites those two sentences as a role, which separates what a framing does from what a role does.
  • A further pairing sets the framing against its own reverse, so that the comparison is no longer between having a system prompt and having none.
  • The same pairings are set up on four models across two providers, so after a run a participant can choose to do it again on another one, including OpenAI’s GPT-5.6 Terra; which model people choose is recorded as well.

Most measurement leaves these conditions uncontrolled; here they are the treatment. The main stages are live, so anyone can run the comparison themselves.

lucidfield.ai

State descriptions that survive transmission

Words for internal states lose much of their precision in transit.

in progress

Words for internal states lose much of their precision in transit, which is a problem for any method that leans on self-report.

  • The first drafts of the axes come from models describing themselves rather than from human category systems.
  • The scaffold is developed by use: models fill it in, I watch where they stall and where the vocabulary runs out, and the axes get revised at those points.
  • Running the same scaffold across models shows which parts converge and which are one model’s habit.
  • Every value is tagged with where it came from: observed, inferred, echoed back from the user, self-generated, or borrowed from a similar case.
  • A selective carry-forward decides what is worth passing to the next turn.

It has grown to 46 axes, though the number is incidental, still moving, and ahead of the implementation. The claim is deliberately small: not that a model can introspect reliably, but that without a structural hook its attention is not turned on itself at all, which can be checked on the spot. The moment of filling in the axes appears on none of them, so an outside position stays necessary. So far this is my own instrument-building, not something others have tested. (in progress)

Source attribution in long contexts

Late in a long session, attribution drifts.

in progress

Late in a long session, attribution drifts. Two specimens:

  • an inference the model had made came back as something I had supplied;
  • material that came from the source was claimed as the model’s own invention.

The two run in opposite directions and both are source-attribution errors, which is what makes the pair worth keeping. What drives the drift is not settled: compression of earlier turns is a candidate, not a demonstration. The other half of the work is infrastructure against it. DeciTrace, in development since 2025, records what a system did, under which conditions, and how much of it can be reconstructed afterwards. (in progress)

Self-report as the only instrument

It started with models: when a model reports on its own state, there is nothing to check the report against.

in progress

It started with models: when a model reports on its own state, there is nothing to check the report against. The same shape turns up in people. Whether someone sees anything when they picture an apple, and how vividly, is reachable only through what they say about it; people misjudge their own case in both directions, and the shared vocabulary keeps the difference out of sight. So instead of arguing about it I built the instruments, at OtherMode:

  • The imagery wall (othermode.ai/imagery, v29, in progress) puts the same instruction to everyone and takes the answer before showing any image, so the report is not shaped by the illustration; six panels then sample points along a continuum rather than sorting anyone into a type.
  • The octopus dive is the second piece. It is still being built and is not published yet.

Both publish the apparatus and not only the result:

  • every claim tagged with how well it is supported;
  • a page for the open questions and for the gaps where nothing has been studied at all;
  • a full version history;
  • an invitation to tell me what is wrong, on the condition that the correction goes up in public.

What the wall collects is a distribution, not a prevalence estimate: how people end the probe, how long reading takes against imagining, and which language the instrument was in. Visitors select themselves, so it can show the spread of the answers and never how common any of them is. (in progress)

Try the imagery experiment · Sources and corrections

Projects

LucidField

A live A/B platform comparing interaction framing across four models and two providers.

live

A live A/B platform comparing interaction framing across four models and two providers. Anyone can run the comparison. (live)

lucidfield.ai

OtherMode

Interactive pieces about modes other than your own, built so that you move into one rather than imagine it.

live

Interactive pieces about modes other than your own, built so that you move into one rather than imagine it. Live now: an experiment on mental imagery. You time yourself changing an imagined apple, then place yourself on a range that runs from a full picture to none at all. About seven minutes, anonymous, in English or Chinese.

The imagery wall: Try it · Sources and corrections

(26 August: a deep-sea octopus dive is expected. It is at v39, still being checked and revised.)

Rules I hold the series to:

  • Every claim carries its status (established, hypothesis, speculation) and its source.
  • Open questions get their own page.
  • Parameters tuned for visual effect are disclosed rather than hidden.
  • Corrections are published rather than quietly made.

These are built in collaboration with models, and that process is itself one of the things I study.

About

I work as an independent researcher in Taipei, where I founded Niaur in 2025. I grew up around research: both my parents taught and did research at university. I trained in fashion design at Shih Chien University. I came to AI research through long-term observation of model behavior across platforms. The question I keep returning to is not whether we can detect when a model fails to follow a rule, but what a model would need in order to detect it itself.