1,720,990 research outputs found

    From coarse to fine-grained concept based discrimination for phrase detection

    Get PDF
    Phrase Detection is a vision and language task where the goal is to determine if a phrase is relevant to an image and localize it, if applicable. The task has many important downstream applications such as Assistive Robotics where a Robot needs to detect/localize an object in an environment. However, training discriminative Phrase Detection models is difficult due to two main challenges 1) sparse training labels: ground truth regions are not exhaustively labeled with all applicable phrases which makes determining negative phrases challenging 2) the training distribution is heavily imbalanced; a small subgroup of the phrases constitute most of the training data. We address these problems through a novel coarse style discrimination: Negative Coarse Concepts (NCC). The method involves grouping visually coherent phrases (concepts) and using the concepts as negative samples which effectively expose the model to a wide and diverse portion of the distribution while minimizing the chance of false negatives. Furthermore, we supplement the coarse discrimination method with a fine grained module (FGM) which effectively discriminates between mutually exclusive tokens of diverse groups such as colors, sizes, etc. Finally, we combine these two novel modules into one model: CFCD-Net, which improves the state-of-the-art performance of Phrase Detection models on Flickr30K Entities and RefCOCO+ Datasets by 1.5-2 points. We further demonstrate how these improvements directly translate to improvements in the downstream task of Binary Image Selection (BISON)

    Proposing coarse to fine grained prediction and hard negative mining for open set 3D object detection

    No full text
    2024In recent years, there has been a remarkable advancement in robotics, autonomous vehicles, and augmented reality technologies, leading to a surge of interest and research activities in 3D learning. Among the 3D recognition research, a significant portion focuses on closed-set detection, overlooking the inherently open nature of real-world scenarios. Furthermore, the scarcity of large-scale 3D datasets, compounded by the high cost of data collection poses a substantial challenge for researchers in this domain. Motivated by these limitations, our work centers on enhancing 3D object detection within an open-set setting. Our work addresses two key research questions within the realm of open-set 3D object recognition, which has not been addressed in prior literature. Firstly, we investigate the efficacy of employing a coarse to fine-grained prediction strategy. This approach aims to enhance the performance of visually similar categories while maintaining the performance of other categories within the dataset. Secondly, we explore the utilization of offline hard negative mining, specifically targeting challenging samples within the text and image modalities to align with the point cloud encoder. This methodology leads to robust performance, particularly on syntactically similar categories. Through these approaches, our study contributes to the advancement of open-set object detection in 3D learning, thereby addressing critical gaps in current research efforts

    MixtureGrowth: growing neural networks by recombining learned parameters

    No full text
    2024Most deep neural networks are trained under fixed network architectures and require retraining when the architecture changes. If expanding the network's size is needed, it is necessary to retrain from scratch, which is expensive. To avoid this, one can grow from a small network by adding random weights over time to gradually achieve the target network size. However, this naive approach falls short in practice as it brings too much noise to the growing process. Prior work tackled this issue by leveraging the already learned weights and training data for generating new weights through conducting a computationally expensive analysis step. In this thesis, we introduce MixtureGrowth, a new approach to growing networks that circumvents the initialization overhead in prior work. Before growing, each layer in our model is generated with a linear combination of parameter templates. Newly grown layer weights are generated by using a new linear combination of existing templates for a layer. On one hand, these templates are already trained for the task, providing a strong initialization. On the other, the new coefficients provide flexibility for the added layer weights to learn something new. We show that our approach boosts top-1 accuracy over the state-of-the-art by 2% points on CIFAR-100 and ImageNet datasets while achieving comparable performance with fewer FLOPs to the larger network trained from scratch

    Going Beyond Counting First Authors in Author Co-citation Analysis

    Get PDF
    The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed

    Learning and evaluating multimodal representations for digital domains

    No full text
    2024Digital domains such as mobile apps and webpages have become fundamental to everyday life. Humans perform many tasks on their phones and online, like reading recipes, booking calendar events, viewing images, and shopping for food or clothes. A prerequisite to building Artificially Intelligent models to aid in these tasks is the process of learning embeddings, i.e., representations, of mobile app and webpage data. In this thesis, we (1) Curate multimodal app and webpage datasets. Digital domains capture four modalities: image, text, structure, and action. We contribute the first multimodal app dataset with all app modalities and language annotations and the first multimodal webpage dataset to retain structure with all image and text content in a unified webpage sample. (2) Define new tasks to evaluate app and webpage understanding. Using our new app dataset, we define an instruction following benchmark that requires mapping a natural language high-level user goal to a sequence of low-level actions. We also define a novel feasibility classification task, in which we predict which user requests can be satisfied in the app environment. Using our new webpage dataset, we define three generation-style tasks: webpage description generation, section summarization, and contextual image captioning. This aims to evaluate webpage understanding at a global, regional, and local level, respectively. (3) Evaluate the importance of each data modality. With our new benchmarks, we determine the impact of each modality on downstream task performance. We find images to be useful for classifying whether a user command is actually satisfiable in an app environment and key to correcting over-reliance on text information. For our webpage benchmarks, contextual text and images aid all tasks, helping image captions retain knowledge-based detail and page descriptions or section summaries retain topical relevance or specificity. (4) Propose new methods for learning multimodal representations of digital domains. Utilizing all available modalities, we contribute a novel attention scheme to make use of webpage structure, separating the most salient content for each task. Results demonstrate that our multimodal encoder is more performant and more computationally efficient. For mobile app representations, we propose using text descriptions and action sequences to learn embeddings that can encode both global and local features while being significantly more data efficient. We outperform prior work on a suite of app understanding tasks while only utilizing publicly available data

    Variations on the Author

    Get PDF
    “Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship

    Appropriate Similarity Measures for Author Cocitation Analysis

    Get PDF
    We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis

    Dispelling the Myths Behind First-author Citation Counts

    Get PDF
    We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more sophisticated methods

    Author Index

    No full text
    Nao informado
    corecore