Bilinear CNNs for Fine-grained Visual Recognition

Lin, Tsung-Yu; RoyChowdhury, Aruni; Maji, Subhransu

Computer Science > Computer Vision and Pattern Recognition

arXiv:1504.07889 (cs)

[Submitted on 29 Apr 2015 (v1), last revised 1 Jun 2017 (this version, v6)]

Title:Bilinear CNNs for Fine-grained Visual Recognition

Authors:Tsung-Yu Lin, Aruni RoyChowdhury, Subhransu Maji

View PDF

Abstract:We present a simple and effective architecture for fine-grained visual recognition called Bilinear Convolutional Neural Networks (B-CNNs). These networks represent an image as a pooled outer product of features derived from two CNNs and capture localized feature interactions in a translationally invariant manner. B-CNNs belong to the class of orderless texture representations but unlike prior work they can be trained in an end-to-end manner. Our most accurate model obtains 84.1%, 79.4%, 86.9% and 91.3% per-image accuracy on the Caltech-UCSD birds [67], NABirds [64], FGVC aircraft [42], and Stanford cars [33] dataset respectively and runs at 30 frames-per-second on a NVIDIA Titan X GPU. We then present a systematic analysis of these networks and show that (1) the bilinear features are highly redundant and can be reduced by an order of magnitude in size without significant loss in accuracy, (2) are also effective for other image classification tasks such as texture and scene recognition, and (3) can be trained from scratch on the ImageNet dataset offering consistent improvements over the baseline architecture. Finally, we present visualizations of these models on various datasets using top activations of neural units and gradient-based inversion techniques. The source code for the complete system is available at this http URL.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1504.07889 [cs.CV]
	(or arXiv:1504.07889v6 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1504.07889

Submission history

From: Tsung-Yu Lin [view email]
[v1] Wed, 29 Apr 2015 15:23:58 UTC (8,722 KB)
[v2] Thu, 30 Apr 2015 15:28:20 UTC (8,722 KB)
[v3] Tue, 29 Sep 2015 19:02:36 UTC (8,722 KB)
[v4] Mon, 28 Nov 2016 21:43:06 UTC (4,154 KB)
[v5] Wed, 31 May 2017 03:35:09 UTC (4,219 KB)
[v6] Thu, 1 Jun 2017 04:24:01 UTC (4,219 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Bilinear CNNs for Fine-grained Visual Recognition

Submission history

Access Paper:

References & Citations

1 blog link

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Bilinear CNNs for Fine-grained Visual Recognition

Submission history

Access Paper:

References & Citations

1 blog link

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators