University of Maryland, Baltimore County
Not a member yet
17643 research outputs found
Sort by
Information Extraction of Security related entities and concepts from unstructured text
Cyber Security has been a big concern especially in past one decade where it is witnessed that targets ranging from large number of internet users to government agencies are being attacked because of vulnerabilities present in the system. Even though these vulnerabilities are identified and published publicly but response has always been slow in covering up these vulnerabilities because there is no automatic mechanism to understand and process this unstructured text that is published on internet. Our system will be tackling this problem of processing unstructured text by identifying the security related terms including entities and concepts from various unstructured data sources. This information extraction task will help expediting the process of understanding and realizing the vulnerabilities and thus making systems secure at faster rate. This work will be describing a system that automatically extracts the terms from Cybersecurity blogs and security bulletins using Natural Language Processing (NLP) and text mining methods. Our NLP model is trained on manually annotated data using open-ended blogs and more structured text like company's official security bulletins. This manually annotated data is unique of its kind since no such previous work has been done using this methodology and it can be a significant contribution for people working in this domain. Our named entity recognition model is trained on conditional random fields (CRFs) based Stanford NER. This automation system will be able to help administrators of organizations and governments to prioritize the task of beefing up their security and moreover track unofficial data sources like chat rooms and twitter for zero day attacks
Beer Wars: The Fight for Independent Brewing in Progressive Era Baltimore
The closing decades of the nineteenth century and the beginning of the twentieth century saw a great deal of change in the United States of America. Rapid technological advancement changed how businesses operated and how ownership interacted with their workers. The beer brewing industry in Baltimore was one such industry that was drastically altered by new equipment and production methods. When the Maryland Brewing Company was formed in 1899 with the express purpose of absorbing as many Baltimore breweries as possible under one ownership and management, some brewers fought back as independent firms seeking to preserve the quality and tradition of the craft. Simultaneously, labor unions emerged to combat the economic and social inequities arising in the wake of industrialization. This caused the German-American community in Baltimore to fracture along class lines. The middle and upper classes sought the support of the working class within their community to help reassert their unique cultural heritage in the face of increasing external pressure to assimilate. The experiences of the Baltimore brewers run parallel to and act as a window on the German-American struggle for cultural expression among pressure to assimilate
A Multilevel Hierarchy of Filesystem Security
The cornerstone of filesystem security in UNIX-like operating systems is the paradigm of users and groups. Each file has an associated user and group which identifies the who has ownership rights to the file. While the idea of user and group ownership of files has been extended by access control lists, it has remained relatively unchanged even as security needs have vastly changed. This research addresses ownership of files in a manner that extends the current model while remaining compatible with the vast majority of available software for UNIX-like operating systems. We have implemented a multilevel hierarchical extension of both users and groups as an extension to the normal single-level UNIX model. We minimized the changes required to user-space applications in order to remain compatible with the vast array of software available for UNIX-like operating systems. The implementation of this research was conducted on the MINIX 3 operating system, with an emphasis on portability to other UNIX-like operating systems in the future. The result of this research is a series of modifications to various operating system server programs to implement our hierarchical permissions model and to enforce the model when loading applications and accessing files in the system
Effects of Urban Development on Groundwater Flow Systems and Streamflow Generation
This work quantifies the impacts of urban development on groundwater storage and groundwater-surface water interactions using intensive data analysis and mathematical modeling. The monthly water balance for the period 2000-2009 for 65 Baltimore area watersheds was calculated using remote sensing data and the dense network of instrumented sites in this region. This analysis included estimation of spatially-distributed anthropogenic fluxes (water supply pipe leakage, lawn irrigation, and infiltration and inflow (I&I) of groundwater and stormwater into wastewater pipes) as well as natural fluxes of precipitation, streamflow, and evapotranspiration. Inflow fluxes of water supply pipe leakage and lawn irrigation were significant but small compared to precipitation, but I&I was approximately equal to gaged streamflow. Building on knowledge of the altered water balance, an integrated hydrologic model of the Baltimore metropolitan region was developed to quantify the impact of urban development on groundwater storage. The three-dimensional groundwater-surface water-land surface model ParFlow.CLM was implemented and a methodology to incorporate urban and hydrogeologic input datasets was developed. Using the model, the impacts of reduced vegetative cover, impervious surfaces, I&I, and other anthropogenic discharge and recharge fluxes were isolated. Removal of I&I led to the largest change in storage, and removal of impervious surface cover had the smallest effect. To investigate the relationship between pre-event water proportion, storage, and streamflow at small watershed scales spanning a gradient of urbanization, chemical hydrograph separation, hillslope numerical experiments, and simple dynamical systems analysis were utilized. From analysis of high-frequency specific conductance data, the pre-event water proportion of stormflow was found to be greatest for storms with higher total precipitation. Using the simple dynamical systems approach, watersheds with larger percentages of impervious surfaces were found to have the largest sensitivity of streamflow to changes in storage. HydroGeoSphere, a three-dimensional groundwater-surface water flow and transport model, was implemented in an idealized hillslope and showed that the relationship between streamflow and storage was clockwise hysteretic. Overall this work demonstrates the importance of infrastructure leakage on urban hydrologic systems and shows that pre-event water contributions of stormflow are primarily related to precipitation and not initial storage in urban watersheds
Reference Advice for Word Sense Disambiguation
Much of the world's knowledge is encoded in natural language. Accessing this information would be invaluable for applications such as agent systems, question answering, the semantic web, expert systems, and many more. However, language is very ambiguous -- each word in a natural language utterance can have a variety of meanings. Word sense disambiguation is the task of determining the dictionary (or lexical) sense of each word in a context. Knowing which sense a word represents allows us to ascertain its meaning. In this work, we introduce a new resource for WSD processing: reference advice. Reference advice uses reference resolution methods to infer potential meanings of referring expressions, such as 'he', 'she,' 'this,' 'it' and many more. We present a method for creating reference advice, and incorporating that advice into WSD processing. We also develop several system configurations which act as points of comparison for evaluation. Lastly, we discuss the impact of the success of reference advice in supporting WSD processing. We explore a shift away from pipeline text processing architectures, wherein each processing component feeds its results downstream until a final processing component presents a solution. We discuss how our system represents a step toward text processing architectures where each processing component is able to provide input to many other processing components, not just those components that are rigidly positioned downstream
PROTEOMIC ANALYSIS OF ASPERGILLUS NIDULANS DURING AUTOPHAGY AND THE ROLE OF AUTOPHAGY GENES ANATG13 AND ANATG8
Aspergilli represent an extremely important genus of microorganisms which can be both harmful pathogens, and beneficial pharmaceutical producers. In Aspergilli's interactions with man, suboptimal nutrient conditions are often present, and lead to a phenomenon known as autophagy. Autophagy is a cellular recycling mechanism that (in the case of macroautophagy) is augmented under nutrient-limited conditions to recycle cytoplasmic macromolecules and organelles for use in essential cell functions. Strategic manipulation of autophagy could ultimately lead to improved bioprocesses or anti-fungal treatments. Using the model filamentous fungus Aspergillus nidulans, a number of important questions about autophagy have been addressed. Critical to the study of autophagy is the balance between self-degradation and self-preservation. Therefore, we adapted an XTT metabolic activity assay for use in filamentous fungi. The assay was first tested using a number of bioprocess-related stresses (e.g. temperature, shear), and found to be superior to DCW as an assessment of culture health. Next, the metabolic activity of fungal cultures was tested during autophagy-inducing conditions, demonstrating that the autophagy capable TN02A3 strain was more viable than an autophagy deficient ΔAnatg13 strain in nutrient limiting conditions. By analyzing the proteome of key autophagy mutants ΔAnatg13 and ΔAnatg8, an improved molecular understanding of autophagy in filamentous fungi was achieved. Using 2-dimensional electrophoresis, 44 unique proteins were found with significant expression changes caused either by addition of rapamycin (a chemical inducer of autophagy) or deletion of Anatg13. AnAtg13 dependent changes of multiple ribosomal and a key polyamine biosynthetic protein, spermidine synthase (AnSpdA), provides molecular evidence of AnAtg13 dependent lifespan extension in A. nidulans. After establishing improved shotgun proteomic methods on the Thermo LTQ-XL, we generated a more thorough assessment of the A. nidulans response to autophagy by measuring protein expression as a function of time. It was found that autophagy induction caused a rapid and sustained increase in proteolysis, amino acid degradation, and lipid metabolism. Additionally, many proteins demonstrated a delayed change in expression. These included proteins involved in secretion, hydrolysis of alternative carbon sources, and secondary metabolite production; all of which are important to the bioprocess industry
Finding Story Chains and Creating Story Maps in Newswire Articles
There are huge amounts of news articles about events published on the Internet everyday. The flood of information on the Internet can easily swamp people, which seems to produce more pain than gain. While there are some excellent search engines, such as Google, Yahoo and Bing, to help us retrieve information by simply providing keywords, the problem of information overload makes it hard to understand the evolution of a news story. Conventional search engines display unstructured search results, which are ranked by relevance using keyword-based ranking methods and other more complicated ranking algorithms. However, when it comes to searching for a story (a sequence of events), none of the ranking algorithms above can organize the search results by evolution of the story. Limitations of unstructured search results include: (1) Lack of the big picture on complex stories. In general, news articles tend to describe the news story from different perspectives. For complex news stories, users can spend significant time looking through unstructured search results without being able to see the big picture of the story. For instance, Hurricane Katrina struck New Orleans on August 23, 2005. By typing ``Hurricane Katrina'' in Google, people can get much information about the event and its impact on the economy, health, and government policies, etc. However, people may feel desperate to sort the information to form a story chain that tells how, for example, Hurricane Katrina has impacted government policies. (2) Hard to find hidden relationships between two events: The connections between news events are sometimes extremely complicated and implicit. It is hard for users to discover the connections without thorough investigation of the search results. In this dissertation, we seek to extend the capability of existing search engines to output coherent story chains and story maps (a map that demonstrates various perspectives on news events), rather than loosely connected pieces of information. By this means, people can obtain a better understanding of the news story, capture the big picture of the news story quickly, and discover hidden relationships between news events. First of all, algorithms for finding story chains have the following two advantages: (1) they can find out how two events are correlated by finding a chain of events that coherently connect them together. Such story chains will help people discover hidden relationship between two events. (2) they allow users to search by complex queries such as ``how is event A related to event B'', which does not work well on conventional keyword-based search engines. Secondly, creating story maps by finding different perspectives on a news story and grouping news articles by the perspectives can help users better capture the big picture of the story and give them suggestions on what directions they can further pursue. From a functionality point of view, the story map is similar to the table of content of a book which gives users a high-level overview of the story and guides them during news reading process. The specific contributions of this dissertation are: (1) Develop various algorithms to find story chains, including: (a) random walk based story chain algorithm; (b) co-clustering based story chain algorithm which further improves the story chains by grouping semantically close words together and propagating the relevance of word nodes to document nodes; (c) finding story chains by extracting multi-dimensional event profiles from unstructured news articles, which aims to better capture relationships among news events. This algorithm significantly improves the quality of the story chains. (2) Develop an algorithm to create story maps which uses Wikipedia as the knowledge base. News articles are represented in the form of bag-of-aspects instead of bag-of-words. Bag-of-aspects representation allows users to search news articles through different aspects of a news event but not through simple keywords matching
The role of male color in darter speciation
"This doctoral research examines the role of male nuptial coloration in the maintenance of species boundaries of darters (Percidae: Etheostoma). The two main goals were to identify how conspicuous male colors contribute to pre-mating isolation and to determine the relative importance of such pre-mating isolation in the maintenance of two sympatric species. The focal pair of sympatric species, E. zonale and E. barrenense, allow an examination of conspecific male coloration as a critical factor by which females select mates. According to speciation via sexual selection theory, if species-specific color leads to pre-mating isolation and pre-mating isolation is stronger than other forms of reproductive barriers, then that nuptial coloration plays a key role in maintaining species boundaries. Although conventional wisdom holds that conspicuous male coloration has evolved to attract conspecific female mates and that species-specific colors serve as a prominent cue maintaining species boundaries, this notion has been explicitly tested only rarely. Importantly, nuptial color may not be driven by female choice; alternatively, color cues may function in male-male competition to improve access to female mates. Therefore it is critical to test the precise role of nuptial coloration in speciation. Darters provide an excellent system for examining the role of nuptial coloration in speciation. Darters comprise the most diverse genus of North American freshwater fish, and nearly every species is characterized by unique nuptial coloration. Multiple darter species co-exist in sympatric populations, indicating that reproductive barriers are tantamount to maintaining these extraordinarily diverse color patterns. This doctoral research demonstrates that male coloration plays a role in female choice both within and between the focal species. Additionally, these species exhibit complete behavioral isolation influenced by female preference for conspecific male coloration. I compared the relative strengths of multiple reproductive barriers and found that behavioral isolation is the strongest measured barrier to gene flow between the focal species. These results present empirical evidence for speciation driven by sexual selection and provide insight into the maintenance of diversity among naturally occurring sympatric species."Includes 1 .mov video file
Applying Weighted GEE for Missing Data Analysis and Sample Size Estimation in Repeated Measurement Studies with Dropout
In planning and designing a clinical trial, researchers need to estimate the sample size based on the target power of the statistical test used to test the treatment effect. In studies with repeated measurements, missing data is often inevitable because some patients may drop out before the study ends. Missing data may reduce the power of hypothesis tests and lead to incorrect decisions. In this dissertation, we first introduced missing data issues in the clinical trials, and we applied five missing data analysis methods to the Schizophrenia study. We performed simulations to compare Weighted GEE method with other methods under different dropout rate. It is important to consider possible missingness to estimate sample size. Though there are rich literatures in sample size calculation in longitudinal studies, not much work has been done in repeated measurements with missing data, especially missing at random. The available methods either ignore the missing data effect, or assume missingness is MCAR, which is a very strong assumption. By following weighted GEE theory in Robins, Rotnitzky, and Zhao (1995), we derived a closed form sample size estimation formula for comparing two groups' slope difference for repeated continuous responses under MAR. We also performed simulations to confirm the validity of the weighted GEE method, and showed that generally the power obtained using the WGEE based sample size formula yields a test whose power is very close to the nominal value. Furthermore, we derived Rao-Blackwellized Monte Carlo estimators to efficiently compute the coefficients that appear on the sample size formula