Article 1 Evaluating Agent Memory Honestly: Benchmarks, Failure Modes, and an Open Interchange Format