Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

Β 

History

17 Commits
Β 
Β 
Β 
Β 

Repository files navigation

🎯 RL4CV

Reinforcement Learning for Computer Vision

πŸ“š A curated collection of papers exploring the intersection of
Reinforcement Learning Γ— Computer Vision


πŸ“– Overview

This repository is a curated awesome list of research papers at the intersection of Reinforcement Learning (RL) and Computer Vision (CV), sourced from top-tier A* conferences including CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, ICRA, AAAI, and IJCAI.

Papers are organized by CV task to help researchers quickly navigate the rapidly evolving landscape.

Note

RL4CV denotes methods where reinforcement learning is applied to support or improve computer vision tasks and vision pipelines.

Available Tags:

  • Highly Cited ⭐
  • Benchmarks & Datasets πŸ“Š

πŸ“‚ Topics


1. Object Detection & Localization

RL agents sequentially search, zoom, and localize objects in images.

Paper Venue Year Links
Recurrent Models of Visual Attention ⭐ NeurIPS 2014 Paper
Multiple Object Recognition with Visual Attention ⭐ ICLR 2015 Paper
Active Object Localization with Deep Reinforcement Learning ⭐ ICCV 2015 Paper
Hierarchical Object Detection with Deep Reinforcement Learning NeurIPS DRL Workshop 2016 Paper
Tree-Structured Reinforcement Learning for Sequential Object Localization NeurIPS 2016 Paper
Reinforcement Learning for Visual Object Detection ⭐ CVPR 2016 Paper
Attend Refine Repeat: Active Box Proposal Generation via In-Out Localization BMVC 2016 Paper
Efficient Object Detection in Large Images using Deep Reinforcement Learning WACV 2020 Paper
Learning to View: Decision Transformers for Active Object Detection ICRA 2023 Paper
Reward Finetuning for Faster and More Accurate Unsupervised Object Discovery NeurIPS 2023 Paper
Adaptive Important Region Selection with Reinforced Hierarchical Search for Dense Object Detection NeurIPS 2024 Paper

2. Image Segmentation

RL for interactive, sequential, or region-growing segmentation strategies.

Paper Venue Year Links
Reinforced Active Learning for Image Segmentation ⭐ ICLR 2020 Paper
Embodied Visual Active Learning for Semantic Segmentation AAAI 2021 Paper
AlignSAM: Aligning Segment Anything Model to Open Context via Reinforcement Learning CVPR 2024 Paper
Focus-Then-Decide: Segmentation-Assisted Reinforcement Learning AAAI 2024 Paper
SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning NeurIPS 2025 Paper

3. Image & Video Generation

Aligning visual generative models with human preferences via RL fine-tuning.

Paper Venue Year Links
ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation NeurIPS 2023 Paper
DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models NeurIPS 2023 Paper
Training Diffusion Models with Reinforcement Learning ⭐ ICLR 2024 Paper
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment TMLR 2024 Paper
Rich Human Feedback for Text-to-Image Generation CVPR 2024 Paper
Diff-Instruct++: Training One-step Text-to-image Generator Model to Align with Human Preferences TMLR 2024 Paper
Training Diffusion Models Towards Diverse Image Generation with Reinforcement Learning CVPR 2024 Paper

4. Video Understanding & Reasoning

RL for long-horizon temporal reasoning, highlight detection, and video summarization.

Paper Venue Year Links
End-to-end Learning of Action Detection from Frame Glimpses in Videos CVPR 2016 Paper
Reinforced Video Captioning with Entailment Rewards EMNLP 2017 Paper
Self-Critical Sequence Training for Image Captioning ⭐ CVPR 2017 Paper
Deep Reinforcement Learning for Unsupervised Video Summarization with Diversity-Representativeness Reward ⭐ AAAI 2018 Paper
Video Captioning via Hierarchical Reinforcement Learning CVPR 2018 Paper
Reinforcement Cutting-Agent Learning for Video Object Segmentation CVPR 2018 Paper
AdaFocus: Adaptive Focus for Efficient Video Recognition ICCV 2021 Paper
Scaling RL to Long Videos NeurIPS 2025 Paper
Video-R1: Reinforcing Video Reasoning in MLLMs NeurIPS 2025 Paper

5. Image Restoration & Enhancement

RL for adaptive, step-wise image denoising, super-resolution, and retouching.

Paper Venue Year Links
Deep Reinforcement Learning of Volume-guided Progressive View Inpainting for 3D Point Scene Completion from a Single Depth Image CVPR (Oral) 2019 Paper
Crafting a Toolchain for Image Restoration by Deep Reinforcement Learning CVPR (Spotlight) 2018 Paper
Fully Convolutional Network with Multi-Step Reinforcement Learning for Image Processing AAAI 2019 Paper
Path-Restore: Learning Network Path Selection for Image Restoration TPAMI 2021 Paper
MOERL: When Mixture-of-Experts Meet Reinforcement Learning for Adverse Weather Image Restoration ICCV 2025 Paper
EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling ICLR 2026 Paper

6. 3D Vision & View Planning

RL for next-best-view selection, 3D scene reconstruction, and point cloud understanding.

Paper Venue Year Links
A Reinforcement Learning Approach to the View Planning Problem CVPR 2017 Paper
Learning to Look Around: Intelligently Exploring Unseen Environments for Unknown Tasks CVPR 2018 Paper
Next-Best View Policy for 3D Reconstruction ECCV Workshop 2020 Paper
Active 3D Shape Reconstruction from Vision and Touch NeurIPS 2021 Paper
GenNBV: Generalizable Next-Best-View Policy for Active 3D Reconstruction CVPR 2024 Paper

7. Medical Image Analysis

RL for landmark detection, lesion localization, and anatomy-aware segmentation.

Paper Venue Year Links
Communicative Reinforcement Learning Agents for Landmark Detection in Brain Images MICCAI MLCN Workshop 2020 Paper
Iteratively-Refined Interactive 3D Medical Image Segmentation with Multi-Agent Reinforcement Learning ⭐ CVPR 2020 Paper
Automatic View Planning with Multi-scale Deep Reinforcement Learning Agents MICCAI 2018 Paper
Sequential Attention-based Sampling for Histopathological Analysis NeurIPS 2025 Paper

8. Visual Navigation & Embodied AI

RL agents that learn to navigate and reason about 3D environments from visual observations.

Paper Venue Year Links
Target-driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning ⭐ ICRA 2017 Paper
Cognitive Mapping and Planning for Visual Navigation ⭐ CVPR 2017 Paper
Embodied Question Answering ⭐ CVPR 2018 Paper
Neural Map: Structured Memory for Deep Reinforcement Learning ICLR 2018 Paper
Learning to Navigate in Cities Without a Map ⭐ NeurIPS 2018 Paper
Visual Semantic Navigation using Scene Priors ICLR 2019 Paper
DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames ⭐ ICLR 2020 Paper

9. Vision-Language Models with RL

Aligning and improving multimodal large language models using reinforcement learning from human or AI feedback.

Paper Venue Year Links
LLaVA-RLHF: Aligning Large Multimodal Models with Factually Augmented RLHF ⭐ arXiv 2023 Paper
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback ⭐ CVPR 2024 Paper
R1-V: Reinforcing Super-Human Visual Perception in Multimodal Large Language Models with Cold-Start RL arXiv 2025 Paper
Visual-RFT: Visual Reinforcement Fine-Tuning arXiv 2025 Paper

⭐ Star this repo if you find it useful!
Made with ❀️ for the CV + RL research community

About

An awesome list of papers on Reinforcement Learning for Computer Vision.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors