Deep Reinforcement Learning
Deep reinforcement learning combines the principles of reinforcement learning and deep learning. This allows a computer system to learn from its mistakes using trial and error, thereby enabling it to address highly complex autonomous decision support problems.
Table of Contents

- What is deep reinforcement learning?
- History of deep reinforcement learning
- Why does deep reinforcement learning matter?
- What are the applications of deep reinforcement learning?
- What are the limitations of deep reinforcement learning?
- New and future developments in deep reinforcement learning
- Deep reinforcement learning at Pacific Northwest National Laboratory
What is deep reinforcement learning?
Deep reinforcement learning combines two concepts—reinforcement learning and deep learning—to help tackle sequential decision-making problems.
It can also be described as a method of learning to make a series of good decisions over time. It’s how humans negotiate the world from the very moment they’re born.
The process of deep reinforcement learning
The basic process loop of deep reinforcement learning can be summarized as follows:
- A computing system or model, called an agent, observes the current conditions of its environment. This set of conditions is termed the state.
- Given these observations and the strategies it knows (the policy), the agent decides which action to take.
- Depending upon the action taken, the agent is given feedback—a reward or penalty.
- The agent uses this feedback to learn which actions provide the best reward or will achieve some goal over time.
The environment in which the agent makes its choices, the number of available options (action set), and the ideas driving the decisions (policies) can be represented or learned in different ways.
How does deep reinforcement learning differ from machine learning or AI?
AI represents an overall term referring to when a computer system demonstrates intelligent behavior to perform tasks, learn, reason, produce content, and make decisions.
Machine learning is a subcategory of AI that allows computers to learn from data and use patterns to make predictions and decisions.
Deep reinforcement learning is a concentrated machine learning approach combining reinforcement learning and deep learning. In this approach, deep neural networks are used to process complex data, including large sensor readings, text, images, and speech. The reinforcement learning framework uses the processed data to direct the system to select actions that increase rewards over time.
History of deep reinforcement learning
1952: American mathematician Richard E. Bellman developed the Bellman equation—the mathematical framework for optimal sequential decision-making.
1957: American psychologist and computer scientist Frank Rosenblatt developed the perceptron, one of the earliest artificial neural networks that learn visual patterns.
1989: British computer scientist Christopher Watkins invented Q-learning, a foundational model-free algorithm in reinforcement learning.
1992–1995: American computer scientist Gerald Tesauro created TD-Gammon, a program that demonstrated that a neural network could learn behavior.
2013–2015: DeepMind researchers demonstrated Deep Q-Networks as a significant achievement in deep-reinforcement learning.
2016–2017: DeepMind’s reinforcement learning system AlphaGo defeated world Go champion Lee Sedol from South Korea. This was followed by AlphaZero, an AI system that can learn multiple games such as chess, shogi, and Go via self-play.
2019: OpenAI’s OpenAI Five computer program defeated professional teams playing five-on-five video game Dota 2—a complex muti-agent competitive environment.
2020: Heron System’s deep reinforcement learning algorithm defeated a trained F-16 fighter pilot 5–0 in Defense Advanced Research Project Agency’s AlphaDogfight simulation trials.
2022: DeepMind’s AlphaTensor discovered a new mathematical algorithm for matrix multiplication via deep reinforcement learning.
2024: DeepMind’s AlphaProof attained a silver-medal standard by solving International Mathematical Olympiad reasoning problems.
2025: DeepSeek-R1, a large language model, trained with deep reinforcement learning for advanced reasoning and long chain-of-thought capabilities.
Why does deep reinforcement learning matter?
Any problem that aims to find a sequence of optimal decisions—from routing traffic, operating cyber networks, maintaining the power grid, evacuating a city during a flood, or servicing a power station—can be approached using deep reinforcement learning.
Scientific discovery, up until this point, has been chiefly driven by experimentation and simulation; when trying to understand a phenomenon, researchers replicate it on computers. They then compare the results of the simulation to what we know about real life.
Of course, simulations can create conditions not yet realized in the real world. A scientist could, for example, modify the location of an atom in a molecule or add another atom. They might not experiment in the physical world but can use these tactics to explore the development of new materials and chemical properties they otherwise could not.
Scientists can enhance this exploration using deep reinforcement learning.
What are the applications of deep reinforcement learning?
Deep learning has seen remarkable success, proving superior performance to traditional machine learning approaches in various application areas:
- computer vision
- speech recognition
- language translation/natural language processing
- health informatics
- medical information processing/diagnostics
- robotics and control, e.g., moving goods around warehouses
- cybersecurity
- autonomous transportation
- autonomous learning.
What are the limitations of deep reinforcement learning?
Deep reinforcement learning can help scientists and researchers make good decisions based on exploring simulation scenarios. However, there are some limitations:
- Good understanding of the simulation environment is needed.
- Not all systems can be made flawless by using deep reinforcement learning. For example, despite having access to a large amount of data, self-driving cars have not mastered all possible conditions and sometimes make mistakes.
- For systems where consequences can be dire, deep reinforcement learning cannot work by itself. For example, one cannot afford to crash a plane to learn how to fly it.
- Deep reinforcement learning needs both time and experiential data to produce the best outcome.
New and future developments in deep reinforcement learning
Scientists are constantly developing new algorithms that would allow deep reinforcement learning to have a better success rate:
- In imitation and inverse reinforcement learning, a machine learns from observing an expert. Instead of trying to learn from its own experience, the machine learns from watching others.
- Goal-conditioned reinforcement learning breaks down complex reinforcement learning problems by using subgoals.
- In multiagent reinforcement learning, multiple agents work together to learn and achieve a specific goal. This technique is instrumental in solving robotics, telecommunications, and economics problems.
- In domain-guided reinforcement learning, reinforcement learning models incorporate expert rules, physical constraints, and domain knowledge to improve algorithmic training and performance.
- Reinforcement learning with human feedback uses human preferences to better guide model training.
Deep reinforcement learning at PNNL

PNNL is leading the next generation of computing for scientific discovery.
Explore our Computing & AI story
Pacific Northwest National Laboratory (PNNL) is a leader in machine learning and artificial intelligence. PNNL’s artificial intelligence research has been applied across various fields, bolstering national security and strengthening the electric grid.
- The DeepGrid open-source platform uses deep reinforcement learning to help empower system operators to create more robust emergency control protocols, augmenting and protecting this last safety net for grid security and resilience.
- PNNL has used deep reinforcement learning to make strides in cybersecurity, which is in an autonomous cyber defense context because such systems are constantly under attack.
PNNL has also used deep reinforcement learning in transportation and building systems for energy-efficient operations.
To further support national missions, PNNL is working to improve the quality of models by boosting the integrity of the data that informs these systems, increasing their accuracy, interpretability, and defensibility.