Zhaoyang Jiang, Zhizhong Fu, Zicheng Li +5 more
Systems that retrieve from records which revise themselves, issue threads, policy logs, encyclopedic histories, long conversations, face a harder question than finding relevant text. They have to judge which claims still hold, which were superseded, and when the honest answer is to decline.
Structured memory is the usual proposal: typed edges, temporal updates, explicit conflict status. The problem this paper identifies is in how those proposals get evaluated. Studies tend to change the underlying mechanism and the way information is presented in the prompt at the same time, so a measured improvement cannot be attributed to either.
They call it a render confound, and it is the kind of methodological finding that quietly invalidates a body of results. If reformatting the prompt would have produced the same gain, then the elaborate memory structure was never the thing doing the work, and nobody could tell because the two moved together.
AI systems increasingly retrieve from records that revise themselves: issue threads, encyclopedic histories, policy logs, and long conversations. The challenge is not only finding relevant evidence, but deciding which claims remain in force, which were superseded, and when to abstain. Structured memories promise to solve this with typed edges, temporal updates, and conflict status, yet evaluations often change mechanism and prompt presentation together. We study this as…
BayesPO: Bayesian Prompt Optimization via Parallel-Tempered Gradient-Guided Discrete MCMC
arXiv (cs.LG) · July 17, 2026CanonicalPhys: Pose-Robust Remote Photoplethysmography via Canonical-Space Priors
arXiv (cs.AI) · July 17, 2026Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI
arXiv (cs.AI) · July 17, 2026A Formally Grounded ODRL Evaluator: Implementation and Comparison