Learning the Structure of Generative Models without Labeled Data

Bach, Stephen H.; He, Bryan; Ratner, Alexander; Ré, Christopher

Computer Science > Machine Learning

arXiv:1703.00854 (cs)

[Submitted on 2 Mar 2017 (v1), last revised 9 Sep 2017 (this version, v2)]

Title:Learning the Structure of Generative Models without Labeled Data

Authors:Stephen H. Bach, Bryan He, Alexander Ratner, Christopher Ré

View PDF

Abstract:Curating labeled training data has become the primary bottleneck in machine learning. Recent frameworks address this bottleneck with generative models to synthesize labels at scale from weak supervision sources. The generative model's dependency structure directly affects the quality of the estimated labels, but selecting a structure automatically without any labeled data is a distinct challenge. We propose a structure estimation method that maximizes the $\ell_1$-regularized marginal pseudolikelihood of the observed data. Our analysis shows that the amount of unlabeled data required to identify the true structure scales sublinearly in the number of possible dependencies for a broad class of models. Simulations show that our method is 100$\times$ faster than a maximum likelihood approach and selects $1/4$ as many extraneous dependencies. We also show that our method provides an average of 1.5 F1 points of improvement over existing, user-developed information extraction applications on real-world data such as PubMed journal abstracts.

Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:1703.00854 [cs.LG]
	(or arXiv:1703.00854v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1703.00854
Journal reference:	Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, PMLR 70, 2017

Submission history

From: Stephen Bach [view email]
[v1] Thu, 2 Mar 2017 16:52:09 UTC (196 KB)
[v2] Sat, 9 Sep 2017 21:22:57 UTC (280 KB)

Computer Science > Machine Learning

Title:Learning the Structure of Generative Models without Labeled Data

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Learning the Structure of Generative Models without Labeled Data

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators