48670 research outputs found
Sort by
Learning from Multiple and Heterogeneous Datasets
The advances in data-acquisition technologies have enabled statisticians to have access to multiple datasets with both globally overlapping and individually variable information. Focusing on applications in single-cell multi-omics, this dissertation concerns statistical methodologies and theories for the estimation of both the global and the individualized structures when multiple and heterogeneous datasets are available. This dissertation is composed of three parts. In the first part, we present MARIO, a robust pipeline for integrative analyses of multi-modal single-cell data that is particularly successful in low signal-to-noise (SNR) ratio scenarios. Currently available tools for single-cell data integration are mainly designed for transcriptomics data and generally rely upon a large number of shared features across datasets. Those methods are unsuitable when applied to single-cell proteomic datasets, due to the limited number of parameters simultaneously accessed, and the lack of shared markers across these experiments. Our algorithmic pipeline takes into account both shared and distinct features and consists of vital filtering steps to avoid sub-optimal matching. MARIO accurately matches and integrates data from different single-cell proteomic and multi-modal methods, including spatial techniques, and has cross-species capabilities. The rest parts are theoretical investigations of two important modules of the MARIO pipeline. The second part discusses minimax optimal community detection in a multi-layer stochastic block model. We characterize the minimax rate for estimating both the global and individualized community structures. We propose a spectral initialization + maximum a posteriori based refinement algorithm that enjoys minimax optimality. This algorithm serves as a key step in MARIO’s quality control steps. The third part is about minimax optimal estimation of a latent correspondence between two datasets where one is a noisy permuted version of the other. We characterize the minimax rate of this problem. We further prove a highly intuitive algorithm that solves a linear assignment problem in the SVD-reduced space achieves consistency, and sometimes minimax optimality under regularity conditions. This algorithm is one of the major ingredients that enables MARIO’s robust performance in low SNR scenarios
Analysis and Control of Neural Network Dynamical Systems
Learning for control has achieved remarkable success in controlling complex dynamical systems such as autonomous vehicles and quadrupedal robots. The resulting controlled system often has a neural network (NN) in the loop which represents the system dynamics, control policy, or perception. While enjoying high empirical performance, NN dynamical systems suffer from the lack of formal guarantees which significantly limits the deployment of NNs or learning algorithms in safety critical applications. This thesis aims to address this challenge by combining control theoretical analysis of dynamical systems and specialized optimization algorithm design for NNs. In the first part of the thesis, we focus on verifying the safety and stability of a NN dynamical system. A hierarchy of verification problems are considered: (i) isolated output range analysis of a NN, (ii) closed-loop reachability analysis, and (iii) closed-loop stability analysis of a NN dynamical system. Output range analysis of a NN concerns over-approximating the output of a NN given a bounded input set. This problem can be formulated as a global optimization problem which is NP-hard to solve, and even its convex relaxations are numerically challenging to solve when large NNs are considered. Verification of closed-loop properties such as reachability and stability of NN dynamical systems becomes even more challenging since the objectives are more complex and the system dynamics must be taken into account. For Problem (i), by utilizing the structure of NNs, we propose a salable operator splitting method for solving a linear program relaxation of the verification problem which achieves a preferable complexity-conservatism trade-off. We solve Problem (ii) and (iii) by combining NN output range analysis tools and specialized reachability (i.e., verifying unrolled NN dynamics in one-shot) or stability (i.e., synthesizing a Lyapunov function using cutting-plane methods) analysis frameworks. In hope of combining robust model predictive control (MPC) and NN verification tools for safe control of NN dynamical systems, in the second part of the thesis, we propose a novel robust MPC method for uncertain linear dynamical systems with polytopic model uncertainty and bounded additive disturbances. This formulation allows abstracting nonlinear NN dynamics by polytopic uncertainties. Drawing tools from System Level Synthesis which transforms state feedback controller design into closed-loop system responses design, our proposed method can simultaneously search over robust linear time-varying state feedback controllers and bounds on the effects of model uncertainty. Extensive simulation shows that our proposed method achieves significantly reduced conservatism compared with existing robust MPC baselines
Extending Provenance for Understanding Claims and Data Analyses
Every day we are bombarded with information, claims, and data — some of which may be controversial or, at least, opinionated. Yet we need to read the information, make decisions, and take action: for example, should we give our children COVID vaccine booster shots? It is usually difficult for us to evaluate the relevant claims and evidence, e.g., COVID vaccine boosters are safe and effective for anyone 5 years or older, and get conclusions based on them: we may be missing the context that the author of the claim may have, such as the information of the source and derivation of the claim (and its evidence), the data that was consulted when the claim was formulated, and the data omitted in the formulation. Provenance has been proposed in data management systems to describe such contextual information, i.e., the life cycle of the data. This dissertation targets the much-needed contextual information, extends provenance to support different domains, facilitate tracing, enable reasoning, and allow interpretation, and proposes techniques to infer it. In this case, people who review the information can have a better understanding of its potential bias. The ideas of the proposed techniques can be applied in two contexts, one oriented around understanding natural language claims and the other one oriented around evaluating data analytic conclusions. The two contexts can be combined to further support assessing the credibility of quantitative claims in natural language. For natural language claims, we propose claim provenance to describe where a claim may come from and explain how it has been derived. We formalize this via a provenance graph and develop a computational framework to infer it, leveraging novel information extraction, text generation, and reasoning techniques. This graph provides provenance for understanding textual claims. For a data analytics-driven report with tables or visualizations, the dissertation focuses on augmenting alternative analysis options that were not disclosed in the report to help users assess whether the data or the data processing steps in the report were “cherry-picked” or representative. To achieve this, we build a search platform over data in a “data lake”, which finds relevant supplementary (joinable or unionable) data with their provenance, serving as potential alternative analysis options. These options provide context beyond data lineage to evaluate the robustness of data analytic conclusions. Finally, for quantitative claims that are informed by data analyses, the dissertation proposes that we need both contextual information mentioned above — users should not only be able to validate the claim based on the source data, but also be provided with a chance to explore relevant results derived by alternative data analysis options to build a more comprehensive view. Therefore, we propose data provenance as a common building block for these two tasks, and propose to infer it via a “retrieval-with-reasoning” framework, considering information from candidate tables and estimating possibly omitted computational steps. This dissertation ends with a vision of what user-facing provenance systems should look like, once augmented with the provenance information our techniques can provide. A prototype system is designed to answer veracity questions about claims in a familiar and tractable domain of reading scientific papers — namely, the tool helps information reviewers look up data in tables that relate to claims in the paper. A preliminary study shows that, when exposed in a suitable way, provenance information can help reviewers answer veracity questions more efficiently. On the basis of this study, we pose preliminary recommendations for designers of future interfaces that expose context-rich provenance information to reviewers and highlight future opportunities for improving provenance inference techniques
Tracing Dehlavi: The Origins of the Hindi-Urdu Lingua Franca
This dissertation adopts an interdisciplinary approach informed by medieval Indian history, historical linguistics, and sociolinguistics to reconsider debates over Hindi-Urdu’s origins. Two of its animating questions are: 1) How did the local speech of Delhi, a modestly sized town of regional importance in the 12th century, survive all of the dramatic transformations of the 13th century to prevail as the general language of the Delhi of the early 14th century, a global metropolis whose population must have originated in large part through immigration from linguistic regions outside of the subcontinent? 2) What accounts for the starkly different sociolinguistic outcomes in two mass migration processes separated by roughly a century—no immigrant communities to north India around the turn of the 13th century ultimately retained their original languages, whereas 14th-century immigrants from north to peninsular India established Dakani as a language that would be maintained to the present day? The first half of the dissertation adopts a new methodology that synthesizes dialectological and traditional historical linguistic approaches to engage with existing origin accounts and establish the premises of the first question. Hindi-Urdu descends from the pre-Ghurid variety of the city of Delhi; north India was characterized not by linguistic uniformity or fluidity in this period but rather by vernacular pluralism; Hindi-Urdu did not originate through a process of mixing of varieties brought into novel contact in Delhi; seemingly anomalous early features can be accounted for without recourse to these explanations. The second half of the dissertation considers both questions through the framework of language maintenance and shift. It presents a new account of Hindi-Urdu’s origins that examines the variety’s history as it intersects with the multiple dramatic developments of the late 12th to 14th centuries—the Ghurid conquest of north India, the establishment of the Delhi Sultanate, the establishment of Islam in India’s midland north, and the transformation of Delhi itself. More broadly, the dissertation argues for a new approach to understanding Hindi-Urdu’s history as that of a regional variety’s transformation into an interregional lingua franca
Learning and Control of Network Phenomena
The intersection of dynamical systems and networks are used to model a huge variety of phenomena. From social networks, to traffic routes and self-driving cars, to swarms of robots and multiagent systems, to individuals moving about in a geographical area, networks can represent an enormous variety of interacting systems. These networks are typically represented via graphs, which are mathematical objects that encode the information of both the entities and relationships within a network. These graphs in turn can be represented by graph matrices, which encode the relationships between entities in the network. The specific problems that this thesis considers span the learning and control of network phenomena via graph matrices: learning structural correspondences across correlated networks; identification of unknown networked dynamical systems with known control inputs; the generalization performance of machine learning algorithms for graphs and hypergraphs; and machine learning and data-driven control of an epidemic spreading across a network. The key intuition throughout is that many properties of networks and networked dynamical systems can be understood by examining the eigenvalue spectrum of an associated graph matrix. Following this intuition, we propose an algorithm for network alignment using spectral information called SPECTRE that exhibits state-of-the-art performance for aligning networks that are only moderately correlated. Next, we present a method for learning the spectra of a graph matrix using only the sparse output measurements of a networked dynamical system. We further propose a new architecture for signal processing on higher-order graphs, along with a new generalization bound on the performance of graph neural networks via spectral similarity. This generalization result is valid for arbitrary graphs regardless of their structure, engendering the first bound on the generalization performance of a machine learning approach for higher-order graphs. Finally, we present a data-driven framework for multi-task learning and non-linear control of epidemics
Symbiotic Approach: Towards a Preservation Model for the Auction Hall and Its Contemporary Extension at Navi Mumbai\u27s MAFCO Wholesale Market
This research is an effort to develop a model preservation approach to derive a preservation philosophy for the Auction Hall and its extension (F & F’, in Fig. 1) in the MAFCO Wholesale Market Complex, Navi Mumbai, by implementing a values-based method, derived specifically for the Indian context. India’s Independence (1947) brought with it new challenges of demonstrating the democratic, secular, economic and social voices of the independent country into its architecture. Its first-generation modernist architects succeeded in this endeavor by adapting the international style of Modernism to the local context and started a new regionalist style of architecture for India. However, insufficient methods to understand these direct translations of the cultural reforms into built forms have led to a lack of official or public recognition for these modern structures, contributing to the significant reasons for their deteriorating conditions in India. The absence of interest leads to fewer research efforts into India’s modern heritage and a lack of nationally accepted regulatory preservation guidelines for such structures. An essential starting point in developing a preservation design philosophy in this thesis is the Madrid-New Delhi document, a non-regulatory statement of principles and guidelines around conserving and preserving modern heritage. Burra Charter is used as a supplement that helps in structuring the assessment of the values for the MAFCO Wholesale Market, further enriched by the framework used in Warm Modernity, which helps to understand the significance of modern structures concerning the Indian context. All three documents are essential to forming an in-depth understanding of the Auction Hall and its extension, further deriving the preservation design philosophy for the place. The interdependence of the framework derived in this thesis emphasizes that no document is sufficient to derive new meanings for modern heritage according to contemporary needs. This thesis calls this approach a Symbiotic one, in which the inter-dependence of the expertise brought by all the actors involved in the sites of modern heritage would enable a genuine process toward a more sensitive and sustainable future for these sites of architectural and cultural significance. While the values ascribed, significance stated, and preservation approach developed herein are specific to the Auction Hall and its extension at the MAFCO Wholesale Market, Navi Mumbai, the methodology and framework of analysis could be applied to the sites of modern heritage in India
Penn Library\u27s Ms. Codex 1630 - Rhetorica ad Herennium. (Video Orientation)
https://repository.upenn.edu/sims_video/1105/thumbnail.jp
Penn Library\u27s Ms. Codex 1644 - Collegium chymica (Video Orientation)
https://repository.upenn.edu/sims_video/1112/thumbnail.jp
Penn Library\u27s LJS 347 De consolatione philosophiae (Video Orientation)
https://repository.upenn.edu/sims_video/1116/thumbnail.jp