High Impact Factor : 4.396 icon | Submit Manuscript Online icon |

Step-Level Hallucination Detection and Error Propagation Analysis in Agentic LLM Systems

Author(s):

Ayurshi Patil , Prof Ram Meghe Institute of Technology and Research; Sania Raut, Prof Ram Meghe Institute of Technology and Research; Satwik Kale, Prof Ram Meghe Institute of Technology and Research; Darpan Thakare, Prof Ram Meghe Institute of Technology and Research

Keywords:

Large Language Models, Agentic AI, Hallucination Detection, Step-Level Detection, Error Propagation, Trajectory Analysis, Agent Evaluation, Memory Contamination, Root Cause Attribution, AI Governance

Abstract

Large Language Models (LLMs) are increasingly used in agentic systems that plan, retrieve, execute, remember, and verify complex workflows involving multiple steps, tools, memory, and intermediate states. In such contexts, errors in intermediate states can be propagated to later steps, and thus to the final answer, when used as part of the state. This work conducts a survey on hallucination detection and error propagation challenges at step, trajectory, and workflow levels. A focused set of ten primary sources is synthesized across three areas of the literature: hallucination detection and mitigation, agentic error propagation, and agentic evaluation and governance. The survey organizes and compares related work by layer of detection, granularity, and methodology, highlighting similarities and differences in output- and step-level detection, error propagation, attribution, evaluation, mitigation, and governance approaches. Notions of error amplification, memory contamination, trajectory attribution, DAG-based dependency modelling, and the distinction between identification of a responsible step and its use for system improvement are discussed. The synthesis outlines several open challenges, notably the lack of a shared benchmark for step-level detection and the diversity of terminology, and provides quantitative evidence of model-specific error propagation properties. It also emphasizes the need to address cost–latency trade-offs, non-DAG workflow architectures, and the integration of technical and governance aspects of hallucination-related system failures.

Other Details

Paper ID: IJSRDV14I80018
Published in: Volume : 14, Issue : 8
Publication Date: 01/11/2026
Page(s): 37-47

Article Preview

Download Article