1,721,013 research outputs found
Exploring domain-informed and physics-guided learning in image-to-image translation
Image-to-image (i2i) translation networks can generate fake images beneficial for many applications in augmented reality, computer graphics, and robotics. However, they require large scale datasets and high contextual understanding to be trained correctly. In this thesis, we propose strategies for solving these problems, improving performances of i2i translation networks by using domain- or physics-related priors. The thesis is divided into two parts. In Part I, we exploit human abstraction capabilities to identify existing relationships in images, thus defining domains that can be leveraged to improve data usage efficiency. We use additional domain-related information to train networks on web-crawled data, hallucinate scenarios unseen during training, and perform few-shot learning. In Part II, we instead rely on physics priors. First, we combine realistic physics-based rendering with generative networks to boost outputs realism and controllability. Then, we exploit naive physical guidance to drive a manifold reorganization, which allowed generating continuous conditions such as timelapses
Multimodal Learning for Urban Scene Understanding
Tato práce se zabývá multimodálním stojovým učením s cílem řešit problémy při vývoji modelů strojového učení s omezeným množstvím anotovaných trénovacích dat, což je významná překážka při trénování systémů hlubokého učení. Spoléhání se na velké množtví anotovaných dat s sebou nese finanční náklady a rizika nejednoznačnosti a zkreslení. V oblasti autonomního řízení, kde musí systémy spolehlivě fungovat v různých prostředích, jsou tyto problémy obzvláště závažné: modely často selhávají ve scénářích, které nejsou v datech dostatečně reprezentovány. Současně systémy autonomního řízení generují rozsáhlá neanotovaná multimodální data z mnoha senzorů. Tato data nabízí příležitost zmírnit omezení plynoucí ze spoléhání se na anotované datasety prostřednictvím učení ze slabých anotací. První část této práce se zabývá omezeními datasetů pro detekci chodců, které často postrádají rozmanitost vzhledu a póz chodců a zřídkakdy zahrnují neobvyklé nebo vysoce rizikové scénáře. K překonání těchto omezení jsme vyvinuli metodu rozšíření datasetu pomocí syntetických dat. Tato generativní metoda je podmíněná pózami osob určenými klíčovými body, které zachycují postoj osoby a fungují jako další doplňková vstupní modalita. To umožňuje řízené generování městských scén s chodci, které simulují vzácné nebo neviděné situace. Náš přístup zlepšuje výkonnost detektoru chodců na několika datasetech, zejména v náročných podmínkách, jako jsou špatně osvětlená prostředí. Druhá část se zaměřuje na sémantickou segmentaci městských scén. V této práci eliminujeme ruční anotování dat tím, že používáme pouze surová data z kamer namontovaných na vozidlech a senzorů LiDAR v crossmodálním samoučení. Konkrétně využíváme modul pro nalezení objektů v datech z LiDARu, který generuje návrhy prostorově konzistentních pozic objektů. Tyto 3D návrhy jsou zarovnány s odpovídajícími snímky a seskupeny do sémantických pseudotříd, které slouží jako pseudoznačky v našem učitel-student schématu. Robustnost demonstrujeme vyhodnocením schopnosti generalizace na čtyřech různých testovacích sadách bez použití dolaďování modelu. Nakonec navrhujeme metodu pro 3D sémantickou predikci obsazenosti okolí, který umožňuje 3D segmentaci a vyhledávání pomocí jazykových dotazů. K dosažení tohoto cíle navrhujeme novou architekturu modelu a samoučící algoritmus používající tři datové modality. Tento algoritmus integruje informace z obrázků, jazyka a 3D bodů z LiDARu. Během inference umožňuje tento přístup segmentaci obsazeného prostoru z RGB snímků pomocí jazykových dotazů. Naše metoda vykazuje významné zlepšení v několika úlohách s jazykovými dotazy a dosahuje konkurenceschopných výsledků ve srovnání s metodami, které jsou učeny pomocí manuálně anotovaných dat.This thesis investigates multimodal learning to address challenges in developing machine learning models with limited annotated training data, a significant bottleneck in training deep learning systems. Reliance on large labeled datasets entails financial costs and risks of ambiguity and bias. In autonomous driving, where systems must operate reliably across diverse environments, these issues are particularly acute: models often fail in underrepresented scenarios. Simultaneously, autonomous driving setups generate vast unannotated multimodal data from multiple sensors. This data offers an opportunity to alleviate the limitations of relying on annotated datasets through weakly supervised learning. The first part of this thesis addresses limitations in pedestrian detection datasets, which often lack diversity in pedestrian appearances and poses and rarely include unusual or high-risk scenarios. To overcome these constraints, we developed a synthetic data augmentation method conditioned on person poses specified by keypoints, acting as an additional complementary input modality. This enables a controlled generation of urban scenes with pedestrians, simulating rare or unseen situations. Our approach improves the performance of the pedestrian detector in multiple datasets, particularly in challenging conditions such as low-light environments. The second part focuses on pixel-wise semantic segmentation of urban scenes. We eliminate manual labeling by using only raw, uncurated data from vehicle-mounted cameras and LiDAR sensors in a cross-modal self-supervised setup. Specifically, we employ a LiDAR-based object proposal module to generate proposals for spatially consistent objects. These 3D proposals are aligned with corresponding images and grouped into semantically meaningful pseudo-classes, serving as pseudo-labels in our teacher-student framework. We demonstrate robustness by evaluating generalization capabilities on four distinct test datasets without fine-tuning. Finally, we propose a framework for open-vocabulary 3D semantic occupancy prediction, enabling 3D grounding, segmentation, and retrieval with free-form language queries. To achieve this, we introduce a novel model architecture and a tri-modal self-supervised learning algorithm that integrates information from images, language, and LiDAR point clouds. During inference, the approach allows the segmentation of occupied space from RGB images in an open-vocabulary manner. Our method demonstrates significant improvements across several open-vocabulary tasks and performs competitively with supervised counterparts
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Vision for Scene Understanding
This manuscript covers my recent research on vision algorithms for scene understanding, articulated in 3 research axes: 3D Vision, Weakly supervised vision, and Vision and physics. At the core of the most recent works is weakly-supervised learning and physics-embodied vision, which address short comings of supervised learning that requires large amount of data. The use of more physically grounded algorithms appears evidently beneficial as both robots and humans naturally evolve in a 3D physical world. On the other hand, accounting for physics knowledge reflects important cue about lighting and weather conditions of the scene central in my work. Physics-informed machine learning is not only beneficial for increased interpretability but also to compensate labels and data scarcity
Algorithmes de vision pour la pluie et les feux tricolores pour les systèmes d'aide à la conduite
Vision algorithms can be used to expand the working range of the assistance systems so as to deal with urban scenes or degraded weathers. To this end, three novel applications are investigated in this thesis for both rain and traffic lights. Rain is the most frequent degraded weather condition. We review the various physics and photometry models for rain and raindrops, and highlight some of the misuses. When driving in daytime the raindrops on the windscreen lower the driver visibility. For standard on-board camera these drops appear as unfocused. Hence, we investigate the detection of unfocused raindrops using blur maps or lack of gradients with photometry. For nightime driving in rain, the headlights paradoxically reduce the visibility due to light reflected off of raindrops back toward the driver. Relying on a physic-based simulator, we propose to build an illumination device that would illuminate the scene without shining the falling particles. The performance of the simulator and a proof-of-concept prototype sustain that our idea can efficiently improve the overall scene visibility. Fast reactive drops detection and tracking is also investigated.To deal with urban scenes, traffic lights play a key role. Though traffic light recognition was attempted in the past, the existing algorithms can't handle complex scenarios. Hence, we have developed a traffic light recognition algorithm that uses a grayscale spot light detection and a template matching classification. Our approach is modular and capable of detecting various kind of traffic lights even when using a low-dynamic camera. We have evaluated our algorithm on sequences from France, China and Switzerland.L'utilisation d'algorithmes de vision permettrait d'élargir le domaine d'application des systèmes d'aide à la conduite à d'autres situations telles que : les scènes urbaines ou les conditions météorologiques dégradées. À cette fin, trois nouvelles applications sont étudiées dans cette thèse pour la pluie et les feux tricolores. La pluie est la condition météorologique dégradée la plus fréquente. Nous comparons les modèles physiques et photométriques existants pour la pluie et les gouttes de pluie. Lors d'une conduite en temps de pluie de jour, les gouttes sur le pare-brise diminuent considérablement la visibilité du conducteur. Lorsqu'elles sont vue par une camera embarquée standard celles-ci apparaissent défocalisées. Ainsi, nous proposons de détecter ces gouttes hors-focus en utilisant soit une approche par manque de gradients soit par l'évaluation locale du flou. Lors d'une conduite de nuit sous la pluie, ce sont les phares qui paradoxalement diminuent la visibilité car leur lumière est réfléchie par les gouttes vers le conducteur. Nous appuyant sur la conception d'un simulateur physique, nous proposons un éclairage adaptatif qui illuminerait la scène sans éclairer les gouttes qui tombent. Les résultats de notre simulateur et le premier prototype construit montre que l'idée avancée pourrait efficacement améliorer la visibilité générale d'une scène. D'autre part, nous étudions la détection et le suivi de gouttes de pluie à grande vitesse. Les feux tricolores ont un rôle crucial dans la compréhension des scènes urbaines. Bien qu'il existe déjà des systèmes de détection de feux tricolores, les algorithmes actuels ne fonctionnent que dans des conditions simples. Ainsi, nous avons développé un algorithme de détection de feux tricolores qui utilise une détection en niveau de gris des spots lumineux et une classification par reconnaissance de modèle. L'approche ainsi conçue est assez flexible pour détecter différents types de feux tricolores même avec une camera à faible dynamique. Notre proposition a été évaluée sur des séquences acquises en France, Chine et Suisse
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
