1,720,975 research outputs found
Leveraging NLP to Enable Analysis of User Driven Routines
Routines are trigger-action programs that user create when automating their home. While routine itself offers a wide variety of utilities with a simple conditional format: if trigger then action, misconfigured or conflicting routines can be potentially dangerous. In order to fully analyze home automation, we should not only be required to examine IoT apps, which are routine created by third-party developers, but should also analyze routines from users’ perspective, also known as user driven routine. However, a direct analysis of user driven routines is non-trivial if routines are expressed in natural language, i.e., without any constraints on what the user can express. In this paper, we propose an approach that utilize Natural Language Processing (NLP) to automatically transform user driven routines expressed in natural language into intermediate representation, which can later be automatically analyzed. We evaluate this approach with a dataset of 250 user driven routines from a concurrent work. Specifically, we extracted device names (i.e., Light Bulb), device capabilities (i.e., switch), and device variables (i.e., on or off) from user driven routines. This approach produced an accuracy of 80.64% in one of the intermediate representations transformation. These results demonstrated that our approach can efficiently help identifying key device properties from natural user driven routines. Furthermore, we separately analyze our intermediate representations to provide additional insights on the characteristic and composition of routines from the end-users’ perspective. Finally, we describe some challenges to this approach and propose several potential improvements for future work.Computer ScienceBachelors of Science (BS
Security and Interpretability in Large Language Models
This honors thesis was written and accepted for departmental honors in Computer Science, also serving as partial fulfillment for the author's computer science minor. A second thesis (not honors) was written and accepted for partial fulfillment of the author's Bachelor's of Science in Physics. The second thesis, titled Data Analysis and Machine Learning on DIRC Hit Patterns for PID, explores a different application of neural networks: enhancement of subatomic particle classification methods through a novel neural network architecture, the Swin vision transformer.Large Language Models (LLMs) have the capability to model long-term dependencies in sequences of tokens, and are consequently often utilized to generate text through language modeling. These capabilities are increasingly being used for code generation tasks; however, LLM-powered code generation tools such as GitHub's Copilot have been generating insecure code and thus pose a cybersecurity risk. To generate secure code we must first understand why LLMs are generating insecure code. This non-trivial task can be realized through interpretability methods, which investigate the hidden state of a neural network to explain model outputs. A new interpretability method is rationales, which obtains the minimum subset of input tokens that lead to the model's output. Through obtaining rationales of insecure code, we are able to investigate the relationship between model inputs and LLM-generated insecure code tokens to further efforts in mitigating cybersecurity risks currently posed by LLM-generated code. Our experiment conducts a case study on two common, pervasive, and severe real-world weaknesses: XSS injection (CWE-79) and SQL injection (CWE-89). We first collected data, then obtained rationales for our weak Python code samples via the greedy rationalization algorithm and a GPT-2 model. Thus, we were able to identify the specific tokens which lead to insecure token generation. We also explored an aggregation function for code rationales - structural code taxonomy - which allowed us to investigate rationales on the local and global levels. Our prototype study found good results: rationales for CWE-79 and CWE-89 code samples have different structural code taxonomy mappings. This implies that each LLM-generated weakness arises from different aspects of the code context, and thus efforts to mitigate insecure LLM-generated code must be precisely targeted to the weakness.Computer ScienceBachelors of Science (BS
Comparison of Vulnerabilities from Smart Home Devices in Chinese and Global Market
With the development of technology, lots of technology companies have introduced a variety of Internet of Things (IoT) devices to both Chinese and global markets. These devices, including smart lock devices, remote control of home automation system, not only offer convenience but also raise security and privacy concerns. This thesis will provide a comprehensive analysis of the mobile applications provided for smart home devices in the Chinese market, focusing on three aspects: cryptographic misuse, SSL misuse and permission misuse. Cryptography misuse focuses on the incorrect selection of encryption and hashing methods. This vulnerability has the potential of sensitive data leaks. SSL misuse encompasses both improper validation of SSL certificates and the use of weak protocols, which may threaten the integrity and confidentiality of data in transit. Permission misuse indicates the case where applications request more permission than necessary or use combinations of permissions in a harmful manner, potentially leading to privacy violations and unauthorized access to user data. The smart home devices are selected based on the criteria of application ranking. This methodology involves a systematic examination of these applications to find previously mentioned vulnerabilities in each category. The examination utilizes static analysis tools to examine the applications, providing a thorough understanding of their security situation. Next, this thesis will focus on a comparative analysis of the selected applications provided in Chinese and international markets. This comparison aims to find differences in vulnerability types in applications and whether these differences correlate with market-specific regulations and policies. This comparison also reveals a divergent strategy adopted by different companies to prioritize security in their applications. By detecting vulnerabilities and differences in different markets, this thesis seeks to contribute to IoT security and also provides further insight for developers into the market’s influence on smart home applications. This study provides further recommendations for companies and policymakers to enhance the security standards for smart home applications.Computer ScienceBachelors of Science (BS
Epidemic Spread Modeling For Covid-19 Using Hard Data
We present an individual-centric model for COVID-19 spread in an urban setting. We first analyze patient and route data of infected patients from January 20, 2020 ,to May 31, 2020, collected by the Korean Center for Disease Control & Prevention (KCDC) and illustrate how infection clusters develop as a function of time. This analysis offers a statistical characterization of mobility habits and patterns of individuals. We use this characterization to parameterize agent-based simulations that capture the spread of the disease, we evaluate simulation predictions with ground truth, and we evaluate different what-if counter-measure scenarios. Although the presented agent-based model is not a definitive model of how COVID-19 spreads in a population, its usefulness, limitations, and flexibility are illustrated and validated using hard data.Computer ScienceMaster of Science (M.Sc.
Power Profiling Smart Home Devices
In recent years, the growing market for smart home devices has raised concerns about user privacy and security. Previous works have utilized power auditing measures to infer activity of IoT devices to mitigate security and privacy threats. In this thesis, we explore the potential of extracting information from the power consumption traces of smart home devices. We present a framework that collects smart home devices’ power traces with current sensors and preprocesses them for effective inference. We collect an extensive dataset of duration > 2h from 6 devices including smart speakers, smart camera and smart display. We perform different classification tasks including device identification and action classification and present accuracy and confusion matrices for each tasks. Our analysis reveals that from devices’ running power traces, we can accurately identify the type of smart device being used with 93% accuracy and subsequently infer user behavior with on average 92% accuracy.Computer ScienceBachelors of Science (BS
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
- …
