The papers my current work rests on, ordered by how much they bear on it. The last column says whether I have checked the venue and year against the paper itself — an unchecked entry came from a secondhand note and should not be cited from here.
| Paper | Year | Status | Bearing | Project | Citation |
|---|---|---|---|---|---|
| Bilinear representation mitigates reversal curse and enables consistent model editingKim, D.-K., Kim, M., Kwon, J., Yang, N., Cha, M. ICLR 2026The structured-geometry predictor rome-neighbors adopts. Also bears on edit-slice: "reversal curse" is argument order under another name, so this is the closest published work to that half of the distinction. Read before E-009. | 2025 | to read | high | rome-neighbors, edit-slice | checked |
| STEAM: A Semantic-Level Knowledge Editing Framework for Large Language ModelsJeong, G., Sun, J., Lee, S., Kim, H. Findings of EMNLP 2025Latent-space alignment for editing. Note the correction: this is an editing *method*, not an alignment-based *predictor* of propagation, so the "alignment arm" of the predictor comparison may not have a paper behind it yet. Check before building E-009 around it. | 2025 | to read | high | rome-neighbors | checked |
| The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form AnswersObaid ul Islam, S., Lauscher, A., Glavaš, G. arXiv preprintThe nearest existing method — mechanistic similarity predicting alignment at up to 78%. But the construct is short-form versus long-form answer agreement, not factual consistency in general, which makes it easier to differentiate from than the one-line note suggested. Still blocking, since it has to be cited either way. | 2025 | to read | high | rome-neighbors | checked |
| NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model InternalsFiotto-Kaufman, J., Loftus, A. R., Todd, E., et al. ICLRThe tool everything runs on — deferred execution over model internals. | 2025 | read | high | rome-neighbors, edit-slice | checked |
| Evaluating the Ripple Effects of Knowledge Editing in Language ModelsCohen, R., Biran, E., Yoran, O., Globerson, A., Geva, M. TACLSix evaluation criteria, and the definition of Logical Generalization that narrowed my claim to justification order. | 2024 | read | high | edit-slice, rome-neighbors | checked |
| Function Vectors in Large Language ModelsTodd, E., Li, M. L., Sharma, A. S., Mueller, A., Wallace, B. C., Bau, D. ICLRRelations as vectors — a structured basis that NNsight can reach directly. | 2024 | read | high | rome-neighbors | checked |
| Mass-Editing Memory in a TransformerMeng, K., Sharma, A. S., Andonian, A., Belinkov, Y., Bau, D. ICLREditing many facts at once — the breadth axis, if orphaning turns out to scale with edit count. | 2023 | read | high | edit-slice | checked |
| MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop QuestionsZhong, Z., Wu, Z., Manning, C. D., Potts, C., Chen, D. EMNLPMulti-hop questions after an edit — the closest existing thing to hop resolution. | 2023 | read | high | rome-neighbors | checked |
| Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge EditingHase, P., Bansal, M., Kim, B., Ghandeharioun, A. NeurIPSLocalisation does not imply editability. The first thing a hostile reviewer will raise, so it is pre-empted in the design. | 2023 | read | high | rome-neighbors | checked |
| Locating and Editing Factual Associations in GPTMeng, K., Bau, D., Andonian, A., Belinkov, Y. NeurIPSCausal tracing, and the rank-one edit that both my projects use as the edit primitive. | 2022 | read | high | rome-neighbors, edit-slice | checked |
| Revisiting Ripple Effects in Knowledge Editing through Pressure-Aware Joint Neighborhood OptimizationHuang, H., Liu, S., Wu, O., Gao, D. arXiv preprintThe closest multi-probe comparison. Key-space coupling is one of three structural probes in it (§3.2), alongside pre-trained entanglement and local neighborhood sensitivity — so it is a predictor set, not a single predictor. Needed for positioning. | 2026 | to read | medium | rome-neighbors | checked |
| Representation Shattering in Transformers: A Synthetic Study with Knowledge EditingNishi, K., Ramesh, R., Okawa, M., Khona, M., Tanaka, H., Lubana, E. S. ICML 2025Distance to representation shattering, which is adjacent to but not the same as distance to propagation. Synthetic setting, so transfer to a real decoder-only LM is an open question rather than an assumption. | 2024 | to read | medium | rome-neighbors | checked |
| RAVEL: Evaluating Interpretability Methods on Disentangling Language Model RepresentationsHuang, J., Wu, Z., Potts, C., Geva, M., Geiger, A. ACLAttribute disentanglement, possibly a cleaner hop-separated dataset than RippleEdits. Worth checking before committing to data. | 2024 | to read | medium | rome-neighbors | checked |
| The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False DatasetsMarks, S., Tegmark, M. COLMA worked example of reading structure off representation geometry. | 2024 | read | medium | rome-neighbors | checked |
| Editing Large Language Models: Problems, Methods, and OpportunitiesYao, Y., Wang, P., Tian, B., Cheng, S., Li, Z., Deng, S., Chen, H., Zhang, N. EMNLPThe field map, and the place to check whether a gap is really a gap. | 2023 | read | medium | edit-slice, rome-neighbors | checked |
| Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceGeva, M., Caciularu, A., Wang, K. R., Goldberg, Y. EMNLPThe vocabulary-space reading of what an FFN update does. | 2022 | read | medium | rome-neighbors | checked |
| Knowledge Neurons in Pretrained TransformersDai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., Wei, F. ACLThe earlier localisation story that ROME's causal tracing supersedes. | 2022 | read | medium | rome-neighbors | checked |
| Transformer Feed-Forward Layers Are Key-Value MemoriesGeva, M., Schuster, R., Berant, J., Levy, O. EMNLPWhy editing an MLP can change a fact at all. | 2021 | read | medium | rome-neighbors | checked |
| Memory-Based Model Editing at ScaleMitchell, E., Lin, C., Bosselut, A., Manning, C. D., Finn, C. ICMLSide-memory editing — no weights touched, so no orphaning by construction. | 2022 | skimmed | low | edit-slice | checked |
| Modifying Memories in Transformer ModelsZhu, C., Rawat, A. S., Zaheer, M., Bhojanapalli, S., Li, D., Yu, F., Kumar, S. arXiv preprintThe earliest statement of "edit one fact and leave the rest alone" as a task. | 2020 | skimmed | low | edit-slice | checked |
| Aging with GRACE: Lifelong Model Editing with Discrete Key-Value AdaptorsHartvigsen, T., Sankaranarayanan, S., Palangi, H., Kim, Y., Ghassemi, M. NeurIPSRetrieval-based editing over a stream of edits. | 2023 | read | low | edit-slice | checked |
| Fast Model Editing at ScaleMitchell, E., Lin, C., Bosselut, A., Finn, C., Manning, C. D. ICLRThe hypernetwork-editor contrast to weight-level editing. | 2022 | read | low | edit-slice | checked |
Kept in data/papers.yml. In any piece, <Cite id="cohen2024ripple" /> resolves to the entry.