Batch Reinforcement Learning from Crowds

Zhang, Guoxi; Kashima, Hisashi

Computer Science > Machine Learning

arXiv:2111.04279v1 (cs)

[Submitted on 8 Nov 2021 (this version), latest version 29 Nov 2022 (v2)]

Title:Batch Reinforcement Learning from Crowds

Authors:Guoxi Zhang, Hisashi Kashima

View PDF

Abstract:A shortcoming of batch reinforcement learning is its requirement for rewards in data, thus not applicable to tasks without reward functions. Existing settings for lack of reward, such as behavioral cloning, rely on optimal demonstrations collected from humans. Unfortunately, extensive expertise is required for ensuring optimality, which hinder the acquisition of large-scale data for complex tasks. This paper addresses the lack of reward in a batch reinforcement learning setting by learning a reward function from preferences. Generating preferences only requires a basic understanding of a task. Being a mental process, generating preferences is faster than performing demonstrations. So preferences can be collected at scale from non-expert humans using crowdsourcing. This paper tackles a critical challenge that emerged when collecting data from non-expert humans: the noise in preferences. A novel probabilistic model is proposed for modelling the reliability of labels, which utilizes labels collaboratively. Moreover, the proposed model smooths the estimation with a learned reward function. Evaluation on Atari datasets demonstrates the effectiveness of the proposed model, followed by an ablation study to analyze the relative importance of the proposed ideas.

Comments:	16 pages
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2111.04279 [cs.LG]
	(or arXiv:2111.04279v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2111.04279

Submission history

From: Guoxi Zhang [view email]
[v1] Mon, 8 Nov 2021 05:46:33 UTC (442 KB)
[v2] Tue, 29 Nov 2022 09:23:36 UTC (1,007 KB)

Computer Science > Machine Learning

Title:Batch Reinforcement Learning from Crowds

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Batch Reinforcement Learning from Crowds

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators