1,721,029 research outputs found
WinoTrain: Winograd-Aware Training for Accurate Full 8-bit Convolution Acceleration
Efficient inference is critical in realizing a lowpower, real-time implementation of convolutional neural networks (CNNs) on compute and memory-constrained embedded platforms. Using quantization techniques and fast convolutional algorithms like Winograd, CNN inference can achieve benefits in latency and in energy consumption. Performing Winograd convolution involves (1) transforming the weights and activations to the Winograd domain, (2) performing element-wise multiplication on the transformed tensors, and (3) transforming the results back to the conventional spatial domain. Combining Winograd with quantization of all its steps results in severe accuracy degradation due to numerical instability. In this paper we propose a simple quantization-aware training technique, which quantizes all three steps of the Winograd convolution, while using a minimal number of scaling factors. Additionally, we propose an FPGA accelerator employing tiling and unrolling methods to highlight the performance benefits of using the full 8-bit quantized Winograd algorithm. We achieve 2× reduction in inference time compared to standard convolution on ResNet-18 for the ImageNet dataset, while improving the Top-1 accuracy by 55.7 p.p. compared to a standard post-training quantized Winograd variant of the network
Mind the Scaling Factors: Resilience Analysis of Quantized Adversarially Robust CNNs
As more deep learning algorithms enter safety-critical application domains, the importance of analyzing their resilience against hardware faults cannot be overstated. Most existing works focus on bit-flips in memory, fewer focus on compute errors, and almost none study the effect of hardware faults on adversarially trained convolutional neural networks (CNNs). In this work, we show that adversarially trained CNNs are more susceptible to failure due to hardware errors when compared to vanilla-trained models. We identify large differences in the quantization scaling factors of the CNNs which are resilient to hardware faults and those which are not. As adversarially trained CNNs learn robustness against input attack perturbations, their internal weight and activation distributions open a backdoor for injecting large magnitude hardware faults. We propose a simple weight decay remedy for adversarially trained models to maintain adversarial robustness and hardware resilience in the same CNN. We improve the fault resilience of an adversarially trained ResNet56 by 25% for large-scale bit-flip benchmarks on activation data while gaining slightly improved accuracy and adversarial robustness
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Design and implementation of number representations for efficient multiplierless acceleration of convolutional neural networks
Today, computer vision (CV) problems are solved with unprecedented accuracy using convolutional neural networks (CNNs), a biologically inspired machine learning concept. However, the large computational workload of CNNs prevents their ubiquitous deployment in embedded and resource constrained systems. For this reason, many approaches for dedicated CNN hardware accelerators have recently been presented in academia, as well as in industry. A key design parameter of such accelerator systems affecting requirements on memory, bandwidth, energy, and algorithmic accuracy is the supported number representation. In this regard, previous research has indicated that neural networks are resilient to fixed-point quantization of parameters and intermediate values. The research focus of this work lies on both the design and implementation of novel number representations and the extension of corresponding quantization methods for neural networks. In particular, this thesis proposes novel concepts for avoiding hardware multipliers to increase energy and area efficiency of accelerator systems. In the first part of the thesis, previous fixed-point quantization methods are reviewed and a self-supervised extension is proposed which specifically enhances the quantization results on pre-trained neural networks. This novel method is the basis for further quantization procedures in the remainder of the thesis. In the second part, CNNs trained on small-scale classification tasks with binary or ternary valued parameters are evaluated. Furthermore, a hardware efficient method for stochastic rounding is introduced, and it is experimentally shown to enhance classification performance while avoiding the need for multiplications. The third and fourth part of the thesis are devoted to the logarithmic quantization of pre-trained CNNs for large-scale image classification and semantic scene segmentation. Logarithmic quantization allows bit-shift-based implementations of multiplications and therefore omits large multipliers in hardware. To mitigate accuracy degradation due to quantization, different log-bases are deployed and implications on hardware implementations are discussed. The resulting accuracy of logarithmically quantized CNNs on CV tasks is experimentally determined. The results reveal that particularly modern complex CNN architectures are prone to few-bit quantization. Therefore, the last part of the thesis presents a logarithmic-based number representation which allows to select the quantization resolution for weights, thereby increasing the accuracy of complex CNN architectures while allowing multiplierless processing
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Echtzeitdatenerfassungssystem für die bildgebende Röntgenspektroskopie
This thesis presents the SuMo-DAQ concept for the real-time data acquisition, processing and analysis of imaging X-ray spectroscopy detector systems of the next generation. The FPGA-based concept enables the analysis of data streams of several GB/s and the significant reduction for recording them. Novel adaptive parameter estimators allow for the automatic adaption of the data processing to the detector system used and to the current measuring conditions. In changing measurement environments, such as in satellites and in long-time measurements, the measurement accuracy can be significantly enhanced.In dieser Arbeit wird das SuMo-DAQ-Konzept für die Echtzeitdatenerfassung, -verarbeitung und -analyse von bildgebenden Röntgenspektroskopie-Detektor-Systemen der nächsten Generation vorgestellt. Das FPGA-basierte Konzept ermöglicht es, Datenströme von mehreren GB/s in Echtzeit zu analysieren und für die Speicherung deutlich zu reduzieren. Neuartige adaptive Parameterschätzer ermöglichen die automatische Anpassung der Datenverarbeitung an das verwendete Detektorsystem sowie an die aktuellen Messbedingungen. In sich verändernden Messumgebungen, wie z.B. auf Satelliten sowie bei Langzeitmessungen, kann die Messgenauigkeit dadurch signifikant erhöht werden
Accelerating and Pruning CNNs for Semantic Segmentation on FPGA
Semantic segmentation is one of the popular tasks in computer
vision, providing pixel-wise annotations for scene understanding. However, segmentation-based convolutional neural networks
require tremendous computational power. In this work, a fully-pipelined hardware accelerator with support for dilated convolution is introduced, which cuts down the redundant zero multiplications. Furthermore, we propose a genetic algorithm based
automated channel pruning technique to jointly optimize computational complexity and model accuracy. Finally, hardware heuristics
and an accurate model of the custom accelerator design enable
a hardware-aware pruning framework. We achieve 2.44× lower
latency with minimal degradation in semantic prediction quality
(−1.98 pp lower mean intersection over union) compared to the
baseline DeepLabV3+ model, evaluated on an Arria-10 FPGA
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
