You watch an event, read an AI summary later, and notice the two accounts disagree. You might expect your memory of the original event to help you identify the mistake, but new research suggests the summary itself can influence what you remember.
Researchers at Georgetown University and the University of Washington found that misleading summaries reduced how accurately people remembered details from videos they had watched. The finding raises a concern for human-in-the-loop AI systems because a reviewer’s memory may already reflect the AI’s error.
Mattea Sim, Yael Eiger, and Tadayoshi Kohno authored the paper, which was posted September 23. According to Georgetown’s coverage, the research is scheduled to be presented at the Ninth AAAI/ACM Conference on AI, Ethics and Society in October 2026.
The models often left out the accident
The research had two parts: the team reviewed summaries from OpenAI’s ChatGPT and Google’s Gemini for factual errors and missing information.
Across 20 summaries of animated accident videos, omission was the most common problem. On average, the summaries left out 51.6% of important details. Nineteen of the 20 summaries, or 95%, failed to mention the main event: a car colliding with a pedestrian.
An accident summary that leaves out the collision gives an incomplete account of what happened. However, this first analysis could not show whether an inaccurate summary would affect someone who had already watched the video. The researchers used a separate experiment to test memory.
One changed detail affected later recall
In the human experiment, participants watched an animated video of a car approaching either a stop sign or a yield sign before colliding with a pedestrian. Twenty-four to 48 hours later, they received a written summary containing either the correct traffic sign or an incorrect one.
Of the 331 people who completed both stages, three failed an attention check during the second stage. The final analysis included 328 participants.
To keep the misinformation consistent, the researchers used ChatGPT-generated text and, when needed, combined portions of its outputs. They then changed the traffic-sign detail for the misleading condition. The experiment tested how misinformation in a summary affected recall. It did not test whether ChatGPT or Gemini would independently generate the same false detail.
When researchers later asked which sign appeared in the original video, 83.6% of participants who received accurate summaries answered correctly. Among those who received misleading summaries, only 44.8% answered correctly. The difference was 38.8 percentage points.
Participants were also told that the summary came from either AI or a human transcriber, although all summaries used ChatGPT-generated text. The source label did not eliminate the effect. Researchers also found no evidence that familiarity with AI or trust in AI explained who was more susceptible.
The reviewer can become part of the problem
Human review usually assumes that a person can compare an AI-generated account with reliable evidence. If the reviewer relies on memory, the summary may shape the recollection they use to spot an error.
An employee who attends a meeting and later reads an AI-generated recap could face a similar risk. If the recap misstates a decision, the incorrect version may influence what the employee later remembers. The study did not test workplace meetings, so it cannot show that the same effect occurs in that setting. It does, however, give organizations a reason to investigate.
In the research announcement, Sim describes AI as “a new method of delivering misinformation” that can produce false memories. She calls for more scrutiny of AI summarization in settings where errors could carry significant consequences.
Police body-camera summaries raise a similar concern because officers may review AI-generated reports about encounters they personally witnessed. If an inaccurate summary influences an officer's recollection, approval of the report may provide less independent verification than expected. Researchers could examine similar risks when clinicians review AI-generated summaries of medical records or patient encounters.
Accuracy testing needs to include the reader
The study relied on animated accident videos, so the findings cannot show how often memory distortion occurs in policing, healthcare, or workplace systems. The researchers also did not test enterprise AI products used inside organizations. Future studies could examine real-world settings, including whether summaries of police-civilian interactions affect later recollection.
For organizations that use AI summarization, evaluation may need to examine more than whether the text is accurate. Researchers can also measure what people remember after reading a summary. This adds a human factor to the wider discussion around independent AI testing.
One safeguard worth testing is requiring reviewers to compare disputed details with the original recording, transcript, or source record. Although the study has not established which safeguards work, it identifies a limitation of memory-based review: an inaccurate AI summary may change how a reviewer remembers the original event.