I've been reviewing how various AI memory systems evaluate their performance and noticed a fundamental issue with cross-system comparison. Most systems benchmark on LOCOMO (Maharana et al., ACL 2024), but the evaluation methods vary significantly. LOCOMO's official metric (Token-Overlap F1) gives GP