Multimodal communication: commonsense, grounding and computation

Alikhani, Malihe

doi:doi:10.7282/t3-qvcs-wz66

RUcore: Rutgers University Community Repository

Search
- All
- Text
- Images
- Audio
- Video
Advanced Search | Help

Search all content in all RUcore collections.
Services
Collections

Help Contact Us My Account

Home

Resource

Multimodal communication: commonsense, grounding and computation

PDF

PDF format is widely accepted and good for printing.

Plug-in required

PDF-1(24.10 MB)

Citation & Export

View Usage Statistics

Staff View

Citation & Export
Hide

Simple citation

Alikhani, Malihe. Multimodal communication: commonsense, grounding and computation. Retrieved from https://doi.org/doi:10.7282/t3-qvcs-wz66

Export

Click here for information about Citation Management Tools at Rutgers.

Statistics
Hide

Description

TitleMultimodal communication: commonsense, grounding and computation

NameAlikhani, Malihe (author); Stone, Matthew (chair); Bekris, Kostas (internal member); de Melo, Gerard (internal member); Nenkova, Ani (outside member); Rutgers University; School of Graduate Studies

Date Created2020

Other Date2020-10 (degree)

SubjectComputer Science

Extent1 online resource (xxii, 182 pages)

DescriptionFrom the gestures that accompany speech to images in social media posts, humans effortlessly combine words with visual presentations. Communication succeeds even though visual and spatial representations are not necessarily wired to syntax and con- ventions, and do not always replicate appearance. Machines, however, are not equipped to understand and generate such presentations due to people’s pervasive reliance on commonsense and world knowledge in relating words and external presentations. I show the potential of discourse modeling for solving the problem of multimodal com- munication. I start with presenting a computational model for diagram understanding, extending linguistics accounts to learn the interpretation of schematic elements such as arrows. I then present a novel framework for modeling and learning a deeper com- bined understanding of text and images by classifying inferential relations to predict temporal, causal, and logical entailments in context. This enables systems to make inferences with high accuracy while revealing author expectations and social-context preferences. I proceed to design methods for generating text based on visual input that use these inferences to provide users with key requested information. The results show a dramatic improvement in the consistency and quality of the generated text by decreasing spurious information by half. Finally, I describe the design of two multi- modal interactive systems that can reason on the context of interactions in the areas of human-robot collaboration and conversational artificial intelligence and describe my research vision: to build human-level communicative systems and grounded artificial intelligence by leveraging the cognitive science of language use.

NotePh.D.

NoteIncludes bibliographical references

Genretheses, External ETD doctoral

Persistent URLhttps://doi.org/doi:10.7282/t3-qvcs-wz66

LanguageEnglish

CollectionSchool of Graduate Studies Electronic Theses and Dissertations

Organization NameRutgers, The State University of New Jersey

RightsThe author owns the copyright to this work.

Version 8.5.5

Citation & ExportHide

Simple citation

Export

StatisticsHide

Description

Citation & Export
Hide

Statistics
Hide