Hook

Their other posts in the index, biggest breakout first.
Một hệ thống trí nhớ Trí nhớ AI không tốn MỘT TOKEN NÀO cho AI agent, nào. Và nó vẫn thắng tất cả. Đây BÁCH KHOA HỒNG KÔNG không phải lưu thêm – mà là tìm lại đúng bằng chứng. AI agent càng hoạt sử trò chuyện càng phình to. Thứ thách thật sự không phải là lưu thêm, mà là tìm lại đúng mẫu bằng chứng – đúng người, đúng điểm – giữa hàng nghìn dòng cũ. Để làm được Hầu hết hệ thống này phải gọi thêm LLM để tóm tắt ký ức sắp xếp ký ức thời gian Giữ trình tự Turn – Phiên Ngữ cảnh phiên Đồ thị thực thể Lần theo quan hệ NER (spaCy) Dựng bóng PageRank Truy hồi Phân tầng thời gian Giữ trình tự Turn – Phiên Ngữ cảnh phiên AI nên lịch sử thành bản tóm tắt – gọn, nhưng mất dấu bằng chứng gốc. Hai: giữ nguyên dữ liệu thô rồi tìm kiếm – nhưng dễ nhầm giữa các phiên na ná nhau. Zero-Mem chọn hướng thứ ba. Ý tưởng gốc: loại bỏ hoàn toàn AI khỏi mọi bước trí nhớ, chỉ giữ AI đúng một chỗ – lúc trả lời câu hỏi cuối cùng. Bí Hai cấu trúc, không token Đồ thị thực thể Lần theo quan hệ NER (spaCy) Dựng bóng PageRank Truy hồi Phân tầng thời gian Giữ trình tự Turn – Phiên Ngữ cảnh phiên quyết nằm ở hai cách nhìn cùng một dữ liệu. Một đồ thị thực thể ghi lại ai xuất hiện cùng ai, để lần theo quan hệ. Một cây phân tầng theo thời gian giữ đúng trình tự và ngữ cảnh từng phiên. dựng bằng công cụ thường, không cần AI. Đường đi của một câu hỏi Đọc câu hỏi Chọn ưu tiên: quan hệ hay thời gian Lấy từ 2 góc nhìn Đồ thị - phân tầng, không gọi AI Ghi chú: Mọi thao tác trí nhớ – dựng, sắp xếp, tìm, lọc – đều không gọi AI và không tốn token. Chỉ bước trả lời cuối mới dùng đến LLM. LLM trả lời Lần gọi AI duy nhất Biến mọi nội dung thành video trong vài phút. Link, chủ đề, file hay video - ainius dựng video cho bạn. Bắt đầu miễn phí tại ainius.net ainius.net Zero-Mem: Zero-Token Memory Operations for LLM Agents Yilin Xiao, Zhehan Zhu, Yujing Zhang, Jin Chen, Zijin Hong, Luyao Zhuan, Qinggang Zhang, Shengyuan Chen, Xiaocao Ouyang, Lingfei Ren, Xiao Hua *The Hong Kong Polytechnic University, Hong Kong SAR School of Computing and Artificial Intelligence Southwestern University of Finance and Economics School of Artificial Intelligence, Jilin University yilin.xiao@connect.polyu.hk 42436006@mail.swufe.edu.cn renlf@swufe.edu.cn xiao.huang@polyu.edu.hk Abstract LLM agents need memory to act consistently over long in- teractions, yet many systems use additional LLM calls to operate that memory. Generating intermediate records and their retrieval adds recurring token and time costs, while omitted or merged details can obscure the original ev- idence. We ask whether structured memory access requires generation at all. Zero-Mem introduces zero-token memory operations: no step outside final question answering invokes an LLM or consumes LLM input or output tokens; encoder computation is accounted for separately. Zero-Mem preserves original interaction traces as its source of record. It organizes the traces in two complementary ways. An entity-context graph exposes connections across interactions, while a tem- poral hierarchy preserves conversational locality and session state. For each query, Zero-Mem weighs the two views, re- trieves from both, and follows their structure to recover sup- porting relations or surrounding context. Deterministic cal- ibration first discovers evidence and then keeps the reader's answer grounded in the retrieved traces. Only the final-QA reader invokes an LLM. Across long-memory regimes, comparative performance while eliminating LLM and LLM token consumption from memory operations, the same final-QA reader and context budget, it reduces operation cost by 57.6% relative to the fastest retrieval baseline. Ablations support the contribution of the two views and their query-dependent coordination. Overall, we show that structured agent memory need not gen- erate intermediate representation of the past. After peer review, code and implementation details will be available at https://github.com/TheMoon0815/Zero-mem. Introduction LLM (LLM) agents increasingly operate interactions, accumulating utterances, actions, and task outcomes (Luo et al. 2025; Du et al. 2025). Their reliability therefore depends on reasoning over the current input, but also on the right evidence from a growing interaction memory system must preserve information across while preventing irrelevant or outdated traces from the current decision. The central challenge is merely how to store more context, but how to provide associated with the correct entity, session, and temporal state when it becomes relevant (Yan et al. 2025b; Hu et al. 2026b; Yang et al. 2026; Wu et al. 2025). Language models have been used to sur- reflect on experience, construct hierarchical maps and graph indexes, and generate or evolve links (Zhang et al. 2024; Salama et al. 2025; Zhang et al. 2025a). These transformations make LLM histories easier to access, but they also turn mem- ory into a recurring generative workload. When abstractions mediate later retrieval, omitted details, subjects, or blurred temporal updates may weak- ly to the original interaction. The opposite may retain the complete history and retrieve direct traces (Yan et al. 2025; Xu, Szlam, and Wehbe). Though this preserves source evidence, flat lexical retrieval can confuse semantically similar traces, ferent users, sessions, or temporal states, and mis- supporting evidence is distributed across manip- tions. Effective memory therefore requires fidel- tion and structured, query-conditioned evidence. Figure 1: Comparison of different agent-memory regimes. Generative memory relies on LLM gen- erations, while raw retrieval searches unstrue- tured and may miss distributed evidence. Zero-Mem- lational and temporal memory structures and re- memory operations with zero LLM calls or token- nal QA invokes an LLM. LoComo - F1 trung bình (GPT-4o- mini) Zero-Mem 59.15 GAM 53.75 CompassMem 52.18 SimpleMem 48.02 Kết quả thì sao? Trên bộ đánh giá trí nhớ hội thoại LoComo, Zero-Mem đạt F1 trung bình 59.15 cao nhất trong mọi hệ thống. Vượt mạnh nhất, hơn 5 điểm trung bình. Không tốn token nào cho trí nhớ, mà vẫn chính xác hơn. Nhưng con số gây sốc nhất không nằm ở độ chính xác. Nó nằm ở chi phí. Và đây mới là phần khiến giới nghiên cứu chú ý. Cùng Chi phí vận hành trí nhớ 0 Token cho vận hành trí nhớ 0.22s Thời gian mỗi truy vấn 57.6% Nhanh hơn baseline nhanh nhất so với hệ thống nhanh nhất trước đó, mà điểm số còn cao hơn. Đạt Cả ngành sinh bộ nhớ bằng LLM, vào bối cảnh: hai năm qua, bộ nhớ cho agent thường do LLM sinh ra – Memo, Zep đều đi theo hướng đó. Năm 2026, LightMem và SimpleMem bắt đầu cắt bớt token, nhưng vẫn chưa bỏ hẳn. Zero-Mem: vế số 0 Bộ phận LLM sinh; bộ nhớ 0 token, truy vấn về hội thoại gốc Zero-Mem đi xa nhất: đưa mọi thao tác bộ nhớ về đúng không token. Bài học lớn hơn: trí nhớ hiệu quả không cần AI sinh ra bản tóm tắt của quá khứ. Giữ nguyên bằng chứng gốc, tổ chức bằng quan hệ và thời gian là đủ. Bộ một góc nhìn, điểm số rơi mạnh – hai cấu trúc bổ sung nhau. Giữ nguyên hay Tóm tắt? Biến mọi nội dung thành video trong vài phút. Link, chủ đề, file hay video - ainius dựng video cho bạn. Bắt đầu miễn phí tại ainius.net ainius.net Trí nhớ hiệu quả không cần AI sinh ra bản tóm tắt của quá khứ. Giữ nguyên bằng chứng gốc, tổ chức bằng quan hệ và thời gian là đủ. Bộ một góc nhìn, điểm số rơi mạnh – hai cấu trúc bổ sung nhau. Giữ nguyên hay Tóm tắt? Biến mọi nội dung thành video trong vài phút. Link, chủ đề, file hay video - ainius dựng video cho bạn. Bắt đầu miễn phí tại ainius.net ainius.net