Dakota State University

Beadle Scholar at Dakota State University
Not a member yet
    1393 research outputs found

    A Novel Approach to Determine the Keyword’s Priority in Short Transcripts

    No full text
    The best keywords in domain-specific short transcripts are those with high specificity and technical importance, directly addressing product/service issues or solutions. Keywords with higher priority provide more accurate representations of conversations. Traditional keyword extraction methods like TF-IDF, YAKE, TextRank, and LDA are limited in handling short texts and domain-specific contexts, as they focus on frequency or general co-occurrence without considering technical priority. The challenge lies in determining keyword priority, which requires contextual understanding and adaptation to evolving domain-specific needs. Following Peffers \u27 design science methodology, this research introduces a novel approach to calculating keyword priority scores in short transcripts. The approach enhances keyword extraction and topic classification, demonstrating efficacy through experiments. Its adaptable nature makes it suitable for various domains, contributing methodologically and practically. Applications include in-house systems, drafting proposals, or outsourcing knowledge retrieval services, providing significant advancements in domain-specific keyword analysis

    Evaluating Topic Models with OpenAI Embeddings: A Comparative Analysis on Variable-Length Texts Using Two Datasets

    Get PDF
    Topic modeling is a crucial unsupervised machine learning technique for identifying themes within unstructured text. This study compares traditional topic modeling methods, like Latent Dirichlet Allocation (LDA), against advanced embedding-based models, specifically BERTopic-OpenAI. The analysis utilizes two distinct datasets: user reviews from the mental health app Replika and the 20newsgroup dataset. For the Replika dataset, both methods identified common themes, but BERTopic-OpenAI uncovered additional nuanced topics, demonstrating its enhanced semantic capabilities. Quantitative evaluation of the 20newsgroup dataset further highlighted BERTopic-OpenAI\u27s advantage through achieving higher topic coherence and diversity than the best-performing LDA model. These results suggest that embedding-based models provide more coherent, interpretable, and diverse topics, making them valuable tools for extracting meaningful insights from extensive and variable-length text corpora. Future research should focus on refining these advanced techniques to improve their applicability and effectiveness in dynamic and varied textual environments

    From LLMs to Randomness: Analyzing Program Input Efficacy With Resource and Language Metrics

    No full text
    Security-focused program testing typically focuses on crash detection and code coverage while overlooking additional system behaviors that can impact program confidentiality and availability. To address this gap, we propose a statistical framework that combines embedding-based anomaly detection, resource usage metrics, and resource-state distance measures to systematically profile software behaviors beyond traditional coverage-based methods. Leveraging over 5 million labeled samples from 50 Python programs, we evaluate how these independent scoring terms distinguish among different sources of input, including Large Language Model (LLM)-generated inputs, and demonstrate how standard statistical tests (e.g., Kolmogorov—Smirnov and Kendall’s τ ) confirm their effectiveness. Our findings show that LLM-generated samples can trigger diverse behaviors but are often less effective at exploring resource usage dynamics (CPU, memory) compared with conventional fuzzing. However, combining LLM outputs with existing techniques broadens behavior coverage and reveals commonalities between commercial LLM outputs. We provide open-source tools for this evaluation framework, demonstrating the potential to refine software testing by integrating behavior metrics into security-testing workflows

    On the Effectiveness of Automatic Code Generation for Synthetic Dataset Creation

    No full text
    This paper compares synthetic and real-world code datasets for machine learning applications in cybersecurity by examining the relationships between machine code and Low-Level Virtual Machine Intermediate Representation (LLVM IR). This study analyzes 1000 randomly generated programs from a compiler fuzzer against 1000 randomly selected samples from AnghaBench to evaluate suitability for security analysis tasks. Statistical analysis revealed that the code generated with fuzzers consistently produces more complex instruction patterns and achieves broader coverage of the available instruction sets, when compared to real-world samples, with statistically significant differences across all measured categories ( p\u3c 0.001 ). The research examines instruction distributions, coverage metrics, program complexity, and statistical properties to characterize synthetic and real-world code differences. Our findings have important implications for vulnerability detection and malware analysis systems, and the research shows that synthetic data generation can effectively complement or potentially surpass real-world samples. These insights help security researchers and practitioners select training datasets for machine learning applications in cybersecurity

    Towards design principles for knowledge management systems that support experientially derived tacit and procedural knowledge

    No full text
    Lessons-learned systems (LLS) are intended to capture and disseminate experientially derived knowledge. However, their use has delivered limited organizational benefits. One cause for this may be the lack of support for tacit and procedural knowledge (TPK), which are often regarded as synonymous with the expertise and skill enabling execution of tasks. Despite the importance of TPK for organizational performance, there is little to no extant design knowledge on how to design KM systems suitable for their management. Towards this gap in the literature and to help address this important problem, we reference the literature and formulate design principles for a class of KM systems that can capture and disseminate experientially derived TPK. We demonstrate the utility of our design principles by developing a prototype to support the knowledge-intensive task of aerial surveys for wildlife. Through exploration of the protype’s use in a detailed scenario, we demonstrate the utility of the artifact and its underlying design principles. Our study contributes novel design knowledge and provides design guidance for practitioners

    Impact of Organizational Changes on BI Critical Success Factors

    Get PDF
    The intricate connections between organizational changes and the critical success factors (CSFs) supporting business intelligence (BI) are thoroughly explored in this dissertation. This study seeks to understand the complex processes that affect the success of BI projects against the backdrop of various organizational transformations, guided by a rich tapestry of existing literature in BI and corporate change management. It has long been understood that BI system adoption and use success are crucial for businesses looking to acquire a competitive edge. However, the effect of organizational changes, whether structural, people-centric, strategy driven, or remedial on the CSFs of BI delivery continues to be a mystery that calls for further research. To fully build and utilize BI environment in changing organizational environments, it is essential to comprehend how these changes affect the significance of established CSFs. This study aims to answer the following fundamental question: How do various organizational changes affect critical success factors that support BI implementation within organizations? It uses a survey-based research approach, technique within the BI area, to unravel this complex puzzle. Survey tools carefully created by examining BI literature are sent to various business intelligence practitioners. The survey gathers information on the significance and presence of BI CSFs, types of organizational changes, and the effectiveness on BI delivery, drawing inspiration from other studies. Regression analysis and analysis of variance (ANOVA) statistical techniques are used to examine the hypotheses put forth in this study. This study aims to provide a comprehensive knowledge of how different organizational changes, whether structural, peoplecentric, strategy-driven, or remedial, moderate some of the CSFs of BI delivery by carefully examining the data. It aims to shed light on how some CSFs become more critical or less to closely manage under specific organizational change scenarios, paving the way to an efficient BI deployment in a transitional environment. This dissertation aims to shed light on the complexities that BI CSFs and organizational transformations share. The conclusions envisioned here aim to provide organizations with valuable insights into maximizing their BI efforts during times of change, ultimately empowering them to make informed, data-driven decisions in a quickly changing business environment. The conclusions also intend to offer a prototype for researchers to extend the study to various other critical success factors (CSFs) amidst organizational changes

    Prescriptive Zero Trust: Assessing the Impact of Zero Trust on Cyber Attack Prevention

    Get PDF
    Increasingly sophisticated and varied cyber threats necessitate ever-improving enterprise security postures. For many organizations today, those postures have a foundation in the Zero Trust Architecture (ZTA). This strategy sees trust as something an enterprise must not give lightly or assume too broadly. Understanding the ZTA and its numerous controls- centered around the idea of not trusting anything inside or outside the network without verification, will allow organizations to comprehend and leverage this increasingly common paradigm. The ZTA, unlike many other regulatory frameworks, is not tightly defined. The research assesses the likelihood of quantifiable guidelines that measure cybersecurity maturity for an enterprise organization in relation to ZTA implementation. This is a new, data-driven methodology for quantifying cyber resilience enabled by the adoption of Zero Trust principles to pragmatically address the critical need of organizations. It also looks at the practical aspects ZTA has on capabilities in deterring cyberattacks on a network. Coupled with quantitative statistical methods, the ZTA maturity approach provides guidance on how an organization can objectively gauge its cybersecurity posture. The outcomes of this research define a prescriptive set of key technical controls across identity verification, microsegmentation, data encryption, analytics, and orchestration that characterize the comprehensive ZTA deployment. By evaluating the depth of integration for each control component and aligning to industry best practices, the study\u27s results help assess an organization\u27s ZTA maturity level on a scale from Initial to Optimized adoption. The research’s resultant four-tier model demarcates phases for an organization on its security transformation journey, with each tier adding to the capability of the last. This structured approach will help organizations improve their respective security postures without systematically compromising operational effectiveness, thereby improving risk management and threat response capabilities. This model does much more than just provide security. It helps an organization optimize resources, make focused investments, and measure progress along its Zero Trust journey in quantifiable terms

    Exploring the Antimicrobial Potential of Honey from Alfalfa (Medicago sativa) Against Both Human and Plant Pathogens

    Get PDF
    Alfalfa (Medicago sativa) is the third most valuable crop in the United States. It produces a large amount of nectar from which honeybees and other types of bees produce a high yield of honey. It is estimated that around 416 to 1,933 pounds of nectar per acre is produced by alfalfa (Kropacova, 1963). However, even with the great antimicrobial features that honey offers, there is still limited research on the antimicrobial properties of specific types of honey, such as alfalfa honey. Topical honey application clears wounds swiftly, aiding rapid healing, even in infections resistant to typical antibiotics like methicillin-resistant Staphylococcus aureus (MRSA). Manuka honey from tea tree in Australia and New Zealand effectively combats human pathogens including Escherichia coli, Klebsiella aerogenes, Salmonella typhimurium, and Staphylococcus aureus, as well as MRSA (Mandal et al., 2011). Alfalfa honey, among others from Saudi Arabia, exhibits high antioxidant potential and significant antimicrobial activity, ranking second only to Acacia honey (Ismail et al., 2021). Honey comprises primarily of different sugars and proteins (0.5%), with variations based on bee species and flora. Hydrogen peroxide (H2O2) is the main antimicrobial agent in honey, confirmed by spectrophoto-metric assays. Other non-peroxide antimicrobial factors include low water content, low pH, phenolic compounds, and bee defensin-1 (Def-1).https://scholar.dsu.edu/research-symposium/1034/thumbnail.jp

    An NLP-Based Knowledge Extraction Approach for IT Tech-Support/Helpdesk Transcripts

    Get PDF
    In the realm of IT support, it is crucial to extract valuable information from various support channels, such as telephone, web chat, email, and social media. This extracted knowledge can help organizations prioritize customer-centric approaches and improve customer service. It also has broader applications, from decision support to product development and human resources policies. Traditionally, extracting knowledge from unstructured text was time-consuming and inefficient, but natural language processing (NLP) techniques demonstrated potential for extracting valuable insights in a variety of application contexts. However, their efficacy and potential in the context of domain-specific IT support transcripts remain primarily limited. This research follows the Design Science Research Methodology (DSRM), encompassing problem identification, objective definition, artifact design, effectiveness demonstration, performance evaluation, and communication of findings. Moreover, the study uses the Attribute-Driven Design (ADD) method to create an approach for NLP-based knowledge extraction, integrating domain-specific knowledge and models. This research explores specific domain requirements and presents an innovative solution. The proposed approach comprises two main processes: domain knowledge extraction, involving off-topic content identification, domain stop-words identification, category extraction, and priority score determination; and transcript knowledge extraction, involving NLP preprocessing, adapted TFIDF keyword extraction, and TranGCN topic categorization. Experimental results demonstrate the effectiveness of this hybrid algorithm, combining rule-based, unsupervised, and supervised machine learning methods. Notably, this study introduces the concept of category/keyword priority scores for keyword extraction and topic categorization, proving their efficacy in experiments. The adaptable nature of this approach suggests its potential application to various IT support domains with customization and updates. Our solution is expected to advance keyword extraction and topic classification courtesy of a novel TF-IDF algorithm adaptation combined with an optimized Graphic Convolutional Neural Network (GCN) method

    The impact of information privacy concerns on information systems use behaviors in non-volitional surveillance contexts: A moderated mediation approach

    No full text
    Electronic surveillance/monitoring has become ubiquitous in modern organizations as advanced information technology (IT) expands organizational capacity to track system users’ daily information systems (IS) activities. Although this environmental shift surrounding IS raises an important (though largely unexplored) issue of IS users’ information privacy and subsequent IS behaviors, little is known about cognitive/psychological processes and boundary conditions underlying IS users’ information privacy concerns and behaviors under the context of non-volitional workplace surveillance. Grounded on psychological reactance theory, this paper articulates how and when information privacy concerns under workplace surveillance relate to IS use behaviors (i.e., effective IS use and shadow IT use) via psychological reactance. In addition, it investigates IS procedural fairness, a contextual boundary condition. We tested a research model using two surveys (via online platforms) data collected from a sample of 301 and 302 IS users working under electronic surveillance/monitoring systems in various organizations and industries. Using moderated mediation analyses, the results of the study show that (1) psychological reactance mediates the relationship between IS users’ information privacy concerns and effective IS use and shadow IT use, respectively; and (2) IS procedural fairness acts as a boundary condition for the given mediated relationships such that the negative impacts of information privacy concerns on psychological reactance and IS behaviors are mitigated. Implications for theory and practice are discussed

    0

    full texts

    0

    metadata records
    Updated in last 30 days.
    Beadle Scholar at Dakota State University
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇