rome-neighbors
Does structured representational geometry predict whether a knowledge edit propagates to logically entailed neighbours?
Raw representational distance does not. Measured across neighbour types it comes out near-flat, around , which is close enough to uninformative that distance alone cannot be the predictor. The question this project asks is whether structured geometry and alignment do better, resolved by entailment hop and measured causally rather than behaviourally.
The first hop-resolved comparison of causal predictors for edit propagation on a decoder-only language model.
- Predictors
- Bilinear structure (Kim), alignment (Jeong / STEAM)
- Metric
- Interchange intervention accuracy — causal, not behavioural
- Resolution
- Per entailment hop, not pooled
- Status
- Experiments running
Why hop resolution matters
Pooling hops hides the effect. A predictor can look useless averaged over neighbours at every distance and still be sharp at one hop and absent at three — and it is the shape of that decay, not its average, that says where an edit's influence actually stops.
Reporting as a function of hop gives a curve. Reporting it pooled gives a number that can be right and say nothing.
Layout
The library is the source of truth; experiments are thin runners over it.
ripplekit/ config, data, reps, predictors, analysis
experiments/ one reproducible runner per ticket
readings/ the vetted literature map
agents/ tickets and shared surfaces
Threats being tracked
- Confound — neighbour types differ in token frequency as well as hop, so frequency has to be controlled before hop can be read as hop.
- Baseline — the dumbest explanation is raw distance. Any structured predictor has to beat it on the same edits, not on a friendlier set.
- Construct validity — behavioural accuracy after an edit is not the same as the edit having propagated. That is why the metric is interventional.
- Localisation is not editability (Hase et al., 2023) — the first objection a reviewer raises here, so the design answers it rather than waiting to be asked.
References
- Cohen, R., Biran, E., Yoran, O., Globerson, A., Geva, M (2024). Evaluating the Ripple Effects of Knowledge Editing in Language Models. TACL.
- Zhong, Z., Wu, Z., Manning, C. D., Potts, C., Chen, D (2023). MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions. EMNLP.
- Hase, P., Bansal, M., Kim, B., Ghandeharioun, A (2023). Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing. NeurIPS.
- Todd, E., Li, M. L., Sharma, A. S., Mueller, A., Wallace, B. C., Bau, D (2024). Function Vectors in Large Language Models. ICLR.
- Fiotto-Kaufman, J., Loftus, A. R., Todd, E., et al (2025). NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals. ICLR.
- Kim, D.-K., Kim, M., Kwon, J., Yang, N., Cha, M (2025). Bilinear representation mitigates reversal curse and enables consistent model editing. ICLR 2026.
- Jeong, G., Sun, J., Lee, S., Kim, H (2025). STEAM: A Semantic-Level Knowledge Editing Framework for Large Language Models. Findings of EMNLP 2025.