1,721,288 research outputs found
Advancing image and video recognition with less supervision
Deep learning is increasingly relevant in our daily lives, as it simplifies tedious tasks and enhances quality of life across various domains such as entertainment, learning, automatic assistance, and autonomous driving. However, the demand for more data to train models for emerging tasks is increasing dramatically. Deep learning models heavily depend on the quality and quantity of data, necessitating high-quality labeled datasets. Yet, each task requires different types of annotations for training and evaluation, posing challenges in obtaining comprehensive supervision. The acquisition of annotations is not only resource-intensive in terms of time and cost but also introduces biases, such as granularity in classification, where distinctions like specific breeds versus generic categories may arise. Furthermore, the dynamic nature of the world causes the challenge that previously annotated data becomes potentially irrelevant, and new categories and rare occurrences continually emerge, making it impossible to label every aspect of the world. Therefore, this thesis aims to explore various supervision scenarios to mitigate the need for full supervision and reduce data acquisition costs. Specifically, we investigate learning without labels, referred to as self-supervised and unsupervised methods, to better understand video and image representations. To learn from data without labels, we leverage injected priors such as motion speed, direction, action order in videos, or semantic information granularity to obtain powerful data representations. Further, we study scenarios involving reduced supervision levels. To reduce annotation costs, first, we propose to omit precise annotations for one modality in multimodal learning, namely in text-video and image-video settings, and transfer available knowledge to large copora of video data. Second, we study semi-supervised learning scenarios, where only a subset of annotated data alongside unlabeled data is available, and propose to revisit regularization constraints and improve generalization to unlabeled data. Additionally, we address scenarios where parts of available data is inherently limited due to privacy and security reasons or naturally rare events, which not only restrict annotations but also limit the overall data volume. For these scenarios, we propose methods that carefully balance between previously obtained knowledge and incoming limited data by introducing a calibration method or combining a space reservation technique with orthogonality constraints. Finally, we explore multimodal and unimodal open-world scenarios where the model is asked to generalize beyond the given set of object or action classes. Specifically, we propose a new challenging setting on multimodal egocentric videos and propose an adaptation method for vision-language models to generalize on egocentric domain. Moreover, we study unimodal image recognition in an open-set setting and propose to disentangle open-set detection and image classification tasks that effectively improve generalization in different settings. In summary, this thesis investigates challenges arising when full supervision for training models is not available. We develop methods to understand learning dynamics and the role of biases in data, while also proposing novel setups to advance training with less supervision.Deep Learning wird zunehmend relevant in unserem täglichen Leben, da es mühsame Aufgaben vereinfacht und die Lebensqualität in verschiedenen Bereichen wie Unterhaltung, Lernen, automatische Unterstützung und autonomes Fahren verbessert. Die Nachfrage nach mehr Daten zur Schulung von Modellen für aufkommende Aufgaben steigt jedoch dramatisch an. Deep Learning Modelle sind stark abhängig von der Qualität und Quantität der Daten, was hochwertige gelabelte Datensätze erfordert. Doch jede Aufgabe erfordert unterschiedliche Arten von Annotationen für Training und Evaluation, was Herausforderungen bei der Beschaffung darstellt. Die Beschaffung von Annotationen ist nicht nur ressourcenintensiv in Bezug auf Zeit und Kosten, sondern führt auch zu Verzerrung, wie z.B. Granularität in der Klassifizierung, wo Unterscheidungen wie spezifische Tierrassen gegenüber generischen Kategorien entstehen können. Darüber hinaus führt die dynamische Natur der Welt dazu, dass zuvor annotierte Daten potenziell irrelevant werden und neue Kategorien und seltene Ereignisse kontinuierlich auftauchen, was es unmöglich macht, jeden Aspekt der Welt zu kennzeichnen. Daher zielt diese Arbeit darauf ab, verschiedene Supervisionszenarien zu erkunden, um den Bedarf an vollständiger supervison zu reduzieren und die Kosten für die Datenerfassung zu senken. Speziell untersuchen wir das Lernen ohne Lables, das als self-supervised und unsupervised bezeichnet wird, um Video- und Bildrepräsentationen besser zu verstehen. Um aus Daten ohne Labels zu lernen, nutzen wir injizierte Priors wie Bewegungsgeschwindigkeit, -richtung, Handlungsreihenfolge in Videos oder semantische Informationsgranularität, um leistungsstarke Datenrepräsentationen zu erhalten. Weiterhin untersuchen wir Szenarien mit reduzierter Supervision. Um die Kosten für Annotationen zu reduzieren, schlagen wir zunächst vor, präzise Annotationen für eine Modalität im multimodalen Lernen zu unterlassen, nämlich in Text-Video- und Bild-Video-Szenarien, und vorhandenes Wissen auf große Korpora von Videodaten zu übertragen. Zweitens untersuchen wir Semi-Supervised Lernszenarien, bei denen nur eine Teilmenge annotierter Daten neben unannotierten Daten verfügbar ist, und schlagen vor, Regularisierungsbeschränkungen zu überdenken und die Verallgemeinerung auf unannotierten Daten zu verbessern. Zusätzlich behandeln wir Szenarien, in denen Teile der verfügbaren Daten aufgrund von Datenschutz- und Sicherheitsgründen oder natürlich seltenen Ereignissen von Natur aus begrenzt sind, was nicht nur die Annotationen einschränkt, sondern auch das gesamte Datenvolumen begrenzt. Für diese Szenarien schlagen wir Methoden vor, die sorgfältig zwischen zuvor erhaltenem Wissen und eintreffenden begrenzten Daten abwägen, indem wir eine Kalibrierungsmethode einführen oder eine Raumreservierungstechnik mit Orthogonalitätsbeschränkungen kombinieren. Schließlich untersuchen wir multimodale und unimodale Szenarien in einer offenen Welt, in denen das Modell gebeten wird, über den gegebenen Satz von Objekt- oder Aktionsklassen hinaus zu generalisieren. Speziell schlagen wir eine neues herausforderndes Szenario für multimodale egozentrische Videos vor und schlagen eine Anpassungsmethode für Vision-Sprach-Modelle vor, um in der egozentrischen Domäne zu generalisieren. Darüber hinaus untersuchen wir die unimodale Bilderkennung in einem Open-Set Szenario und schlagen vor, Open-Set-Erkennung und Bildklassifizierungsaufgaben zu entflechten, die die Generalisierung in verschiedenen Einstellungen effektiv verbessern. Zusammenfassend untersucht diese Arbeit die Herausforderungen, die entstehen, wenn eine vollständige Überwachung für das Training von Modellen nicht verfügbar ist. Wir entwickeln Methoden, um das Lernverhalten und die Rolle von Verzerrungen in Daten zu verstehen, während wir gleichzeitig neuartige Setups vorschlagen, um das Training mit weniger Supervision voranzutreiben
A4NT: Author Attribute Anonymity by Adversarial Training of Neural Machine Translation
Text-based analysis methods enable an adversary to reveal privacy relevant author attributes such as gender, age and can identify the text's author. Such methods can compromise the privacy of an anonymous author even when the author tries to remove privacy sensitive content. In this paper, we propose an automatic method, called the Adversarial Author Attribute Anonymity Neural Translation (), to combat such text-based adversaries. Unlike prior works on obfuscation, we propose a system that is fully automatic and learns to perform obfuscation entirely from the data. This allows us to easily apply the system to obfuscate different author attributes. We propose a sequence-to-sequence language model, inspired by machine translation, and an adversarial training framework to design a system which learns to transform the input text to obfuscate the author attributes without paired data. We also propose and evaluate techniques to impose constraints on our model to preserve the semantics of the input text. learns to make minimal changes to the input to successfully fool author attribute classifiers, while preserving the meaning of the input text. Our experiments on two datasets and three settings show that the proposed method is effective in fooling the attribute classifiers and thus improves the anonymity of authors
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
B.: Unsupervised discovery of structure in activity data using multiple eigenspaces
Abstract. In this paper we propose a novel scheme for unsupervised detection of structure in activity data. Our method is based upon an algorithm that represents data in terms of multiple low-dimensional eigenspaces. We describe the algorithm and propose an extension that allows to handle multiple time scales. The validity of the approach is demonstrated on several data sets and using two types of acceleration features. Finally, we report on experiments that indicate that our approach can yield recognition rates comparable to other, supervised approaches.
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
