Discovering Temporal Structure: An Overview of Hierarchical Reinforcement Learning
| Source: arXiv AI
Tags: hierarchical reinforcement learning, temporal abstraction, options framework, LLM planning, Doina Precup
A comprehensive HRL survey by Doina Precup, Marlos Machado and collaborators — updated September 2026 — maps the landscape from classical options theory through LLM-guided temporal abstraction, framing hierarchical structure discovery as a path toward long-horizon general agents.
Details
Hierarchical Reinforcement Learning (HRL) addresses one of RL's hardest problems: how to learn and plan over long time horizons in complex environments. This survey, originally published June 2025 and updated September 2026, is authored by Martin Klissarov, Akhil Bagaria, Ziyan Luo, George Konidaris, Doina Precup, and Marlos Machado — a notable lineup from Mila and DeepMind Montreal. The paper organizes HRL methods around the core challenge of discovering temporal structure: what makes a good subgoal or option, and how can agents find it automatically? Coverage spans methods that learn from online experience, from offline datasets, and from leveraging LLMs to provide structured priors for temporal abstraction. The inclusion of LLM-guided discovery reflects the field's current direction — reasoning models and foundation models are increasingly used to propose subgoals in long-horizon tasks. The survey also identifies specific domains where temporal abstraction yields the highest gains, helping practitioners decide where HRL is worth the additional complexity. The updated version incorporates recent 2026 developments, making it more current than most survey papers. The central open problem identified: defining what constitutes 'good' temporal structure remains unsolved.