Skip to main content

About

I'm Saif Ul Islam. I work on knowledge editing and interpretability in language models — what a model believes, how to change it, and how to tell whether the change actually took hold and how far it travelled.

What this site is

A working notebook, published. Two kinds of page:

  • Writing — dated pieces. A note, an argument, a session where something turned out to be wrong. The date is part of what they mean.
  • Notebook — living pages, one per project, revised in place so a reader always sees current state.

Nothing here is peer reviewed. Pieces carry a status, and the ones that get withdrawn stay up with the reason attached, because a claim that quietly disappears teaches nobody anything.

How I choose what to work on

Three questions, applied to each piece of work:

  1. Is it true regardless of what I hoped?
  2. Does it take prior work one verifiable step further?
  3. Would it still matter to the field if my own vision vanished tomorrow?

Work that fails these is a toy, and I let it go. Difficulty is a cost, not a goal — the best next step is usually the one that is better connected to a real problem, not the one that is harder.

Elsewhere

Corrections

If something here is wrong, I would rather know. Email is the fastest route, and a correction that changes a claim gets recorded on the page rather than edited away.