Perhimpunan Mahasiswa SUTD Indonesia (PADI
Survey of Agent Architectures and Reinforcement Learning
Pages
58
Time to read
135 mins
Publication
Language
English
Pages
58
Time to read
135 mins
Publication
Language
English
This document is a comprehensive survey focusing on agent architectures and reinforcement learning (RL) for long-horizon sequential decision-making tasks. It reviews over 280 papers, categorizing the field along three axes: task domain, methodology, and core challenges. The survey identifies six core challenges that complicate long-horizon tasks, including credit assignment, exploration, and scalability. Additionally, it presents four analytical contributions: a failure taxonomy for long-horizon agents, a gap analysis matrix, original experiments demonstrating performance decay with task horizon, and a formal conjecture on the limits of flat architectures. The document also discusses the evolution of long-horizon decision-making through three distinct eras, highlighting the transition from classical planning to modern LLM-based agents. The findings suggest that hybrid architectures combining various methodologies may offer the most promising path forward in developing capable long-horizon agents.