Dynamic Retriever for In-Context Knowledge Editing via Policy Optimization
Mahmud Wasif Nafee, Maiqi Jiang, Haipeng Chen, Yanfu Zhang
TL;DR
DR-IKE tackles the problem of stale LLM knowledge by introducing a dynamic, reward-guided retrieval framework that selects and prunes in-context demonstrations for knowledge editing without modifying model weights. A light-weight BERT retriever is trained via policy gradients to surface only the most informative Retain demonstrations, while a learnable budget controller adaptively shortens prompts for easy edits and expands support for difficult ones. Across CounterFact and multiple LLMs, DR-IKE achieves up to 17.1% higher edit success with up to 41.6% latency reductions and better preservation of unrelated knowledge, demonstrating strong adaptivity to task difficulty and efficiency in black-box settings. The approach offers a scalable, model-agnostic alternative to gradient-based editors, with potential applicability to temporal, numerical, and paraphrase-rich updates in real-world deployments.
Abstract
Large language models (LLMs) excel at factual recall yet still propagate stale or incorrect knowledge. In-context knowledge editing offers a gradient-free remedy suitable for black-box APIs, but current editors rely on static demonstration sets chosen by surface-level similarity, leading to two persistent obstacles: (i) a quantity-quality trade-off, and (ii) lack of adaptivity to task difficulty. We address these issues by dynamically selecting supporting demonstrations according to their utility for the edit. We propose Dynamic Retriever for In-Context Knowledge Editing (DR-IKE), a lightweight framework that (1) trains a BERT retriever with REINFORCE to rank demonstrations by editing reward, and (2) employs a learnable threshold to prune low-value examples, shortening the prompt when the edit is easy and expanding it when the task is hard. DR-IKE performs editing without modifying model weights, relying solely on forward passes for compatibility with black-box LLMs. On the COUNTERFACT benchmark, it improves edit success by up to 17.1%, reduces latency by 41.6%, and preserves accuracy on unrelated queries, demonstrating scalable and adaptive knowledge editing. The code is available at https://github.com/mwnafee/DR-IKE .
