On-screen text
история о том, как мы исследуем разные AI модели в UC Berkeley чтобы они рекомендовали ваши продукты
все началось в конце августа 2025. мы (я и мой ко фаундер) спросили вопрос, как AI модели отвечают в реальном времени, если информация на которой они ищут информацию уже устарела?
OpenAI
Summer Meetu
это было после chatgpt 5 launch
мы написали первую research
paper на ArXiv
Computer Science > Artificial Intelligence
arXiv:2509.10762 (cs)
[Submitted on 13 Sep 2025]
AI Answer Engine Citation Behavior
An Empirical Analysis of the GEO16 Framework
Arlen Kumar, Leanid Palkhouski
View PDF
HTML (experimental)
AI answer engines increasingly mediate access to domain
knowledge by generating responses and citing web sources. We
introduce GEO-16, a 16 pillar auditing framework that converts on
page quality signals into banded pillar scores and a normalized
GEO score G that ranges from 0 to 1. Using 70 product intent
prompts, we collected 1.700 citations across three engines (Provo
мы ее презентовали на самой
большой GEO (generative engine
optimization) конференции в Сан
Франциско в декабре
мы наняли больше research
людей чтобы помогли делать
эксперименты
1 of 9
мы опубликовали уже новую research
статью, которая уже говорит про
влияние неправильной информации для
AI Search
Freshness as a First-Class Reliability Constraint in AI Search:
Evidence from Controlled Staleness in Retrieval-Augmented Generation
Wrodium
March 2026
Abstract
Retrieval-augmented AI search systems are increasingly used in enterprise settings, yet con-
tent freshness is often under-specified in both evaluation and operations. This paper investigates
how stale corpus content degrades end-to-end pipeline behavior and argues that freshness infras-
tructure should be treated as a first-class systems requirement. Using FreshRAG, a controlled
study spanning five enterprise domains and four staleness conditions (0%, 10%, 30%, 50%), we
identify a temporal-semantic failure mode in which conventional metrics remain comparatively
stable while temporally valid evidence declines. At 50% staleness, fresh answer-bearing retrieval
declines by 24.2% despite minimal change in Precision[95] and Recall@5. Subsequent pipeline
stages (reranking, context assembly, and verification-regeneration) only partially mitigate this
degradation and can increase computational and economic overhead. We conclude that robust
AI search requires explicit freshness infrastructure: measurable freshness objectives, temporal
observability, and policy-governed freshness controls.
1 Introduction
Retrieval-Augmented Generation (RAG) has become the default architecture for enterprise AI
3 of 9 Fig. 3: Limited Temporal Correction
Reranking does not reduce stale intrusion in our experiments. In high-staleness settings, stale
passages are frequently promoted to top ranks, suggesting rerankers inherit the same temporal
blind spot as hot retrievers (Figure 3).
The time-aware breakdown in Figure 3 indicates this distribution is not uniform: time-sensitive
intents lose more fresh supporting evidence than time-insensitive ones.
4
немного инсайдов
Retrieval Precision
Retrieval Recall
Stale Document Intrusion
Figure 2: Retrieval degradation under increasing corpus staleness. Headline metrics (Precision@5,
Recall@5) remain comparatively stable while stale intrusion rises and fresh answer-bearing retrieval
declines.
Precision@5
Recall@5
Stale Intrusion Rate
Figure 3: Time-sensitive queries are disproportionately affected by staleness, showing larger stale
intrusion and steeper drops in fresh supporting evidence.
4.3 Context Assembly: Freshness Dilution
Context freshness falls to 83.0% at 50% staleness, with stale token ratio increasing to 17.0% (Fig-
ure 5). Notably, contradiction density decreases rather than increases, indicating stale content is
often compatible but temporally outdated, making contradiction filters insufficient.
4.4 Generation and Verification: High Cost, Modest Recovery
Hallucination-related failure rates remain high (roughly 93.8-95.8%). Verification triggers gener-
ation for more responses, but only 3.9-10.0% are successfully recovered. End-to-end cost per query
increases by approximately 3x, indicating poor economic efficiency of post hoc repair under stale
retrieval.
YOKUBO
мы нашли определение паттерны и они
оправдывали наш hypothesis
мы нанимаем больше
людей, расширяем наш
продукт и находим новых
клиентов, кто хочет чаще
показываться в ChatGPT и
других AI моделях для
своих клиентов
такая вот история.
сохраняйте