Learning descriptive models of objects and activities from egocentric video

Fathi, Alireza

Title:

Learning descriptive models of objects and activities from egocentric video

Files

FATHI-DISSERTATION-2013.pdf (41.51 MB)

Author(s)

Fathi, Alireza

Advisor(s)

Rehg, James M.

Advisor(s)

Person

Rehg, James M.

Associated Organization(s)

Organizational Unit

College of Computing

Organizational Unit

School of Computer Science

Organizational Unit

Institute for Robotics and Intelligent Machines (IRIM)

Collections

Theses and Dissertations

Permanent Link

http://hdl.handle.net/1853/48738

Abstract

Recent advances in camera technology have made it possible to build a comfortable, wearable system which can capture the scene in front of the user throughout the day. Products based on this technology, such as GoPro and Google Glass, have generated substantial interest. In this thesis, I present my work on egocentric vision, which leverages wearable camera technology and provides a new line of attack on classical computer vision problems such as object categorization and activity recognition. The dominant paradigm for object and activity recognition over the last decade has been based on using the web. In this paradigm, in order to learn a model for an object category like coffee jar, various images of that object type are fetched from the web (e.g. through Google image search), features are extracted and then classifiers are learned. This paradigm has led to great advances in the field and has produced state-of-the-art results for object recognition. However, it has two main shortcomings: a) objects on the web appear in isolation and they miss the context of daily usage; and b) web data does not represent what we see every day. In this thesis, I demonstrate that egocentric vision can address these limitations as an alternative paradigm. I will demonstrate that contextual cues and the actions of a user can be exploited in an egocentric vision system to learn models of objects under very weak supervision. In addition, I will show that measurements of a subject's gaze during object manipulation tasks can provide novel feature representations to support activity recognition. Moving beyond surface-level categorization, I will showcase a method for automatically discovering object state changes during actions, and an approach to building descriptive models of social interactions between groups of individuals. These new capabilities for egocentric video analysis will enable new applications in life logging, elder care, human-robot interaction, developmental screening, augmented reality and social media.

Date Issued

2013-06-13

Resource Type

Text

Resource Subtype

Dissertation

Full item page

Title:

Learning descriptive models of objects and activities from egocentric video

Files

Author(s)

Authors

Advisor(s)

Advisor(s)

Editor(s)

Associated Organization(s)

Series

Collections

Supplementary to

Permanent Link

Abstract

Sponsor

Date Issued

Extent

Resource Type

Resource Subtype

Rights Statement

Rights URI

Georgia Tech Library

Title: Learning descriptive models of objects and activities from egocentric video

Files

Author(s)

Authors

Advisor(s)

Advisor(s)

Editor(s)

Associated Organization(s)

Series

Collections

Supplementary to

Permanent Link

Abstract

Sponsor

Date Issued

Extent

Resource Type

Resource Subtype

Rights Statement

Rights URI

Title:

Learning descriptive models of objects and activities from egocentric video