Rochester Institute of Technology

RIT Digital Institutional Repository
Not a member yet
    22505 research outputs found

    Predicting Customer Churn in E-commerce

    Get PDF
    Customer retention has become a critical focus for businesses seeking to sustain growth and profitability in an increasingly competitive market. In this thesis, advanced machine learning techniques are used to develop a data-driven churn prediction model for Majid Al Futtaim\u27s customer base. In an e-commerce environment, the research focuses on identifying the key transactional and behavioral factors that influence customer churn, based on customer relationship management theory. The study is contextualized in the context of increased digital consumer engagement and high customer acquisition costs, which make retaining existing customers more cost- effective than acquiring new ones. Our primary research questions focused on identifying the most effective machine learning models for churn prediction, identifying the most influential factors contributing to churn, and developing targeted retention strategies. Using customer RFM metrics, payment behavior, and delivery satisfaction indicators, a structured methodology was employed, beginning with extensive data preprocessing, exploratory data analysis (EDA), and feature engineering. An existing dataset was sourced and refined to produce meaningful inputs. With SPSS Modeler, Random Trees and the Feature Selection Node were used for feature selection and importance ranking. The following three algorithms were evaluated for model development: Logistic Regression, Linear Support Vector Machine (LSVM), and Artificial Neural Networks. In order to address class imbalance, the dataset was partitioned into training and testing subsets using 70:30 ratios. Accuracy, Precision, Recall, F1-score, and AUC-ROC were used to evaluate model performance. The Logistic Regression model had the highest accuracy, AUC, and F1-score of all the models tested. The Logistic Regression model performed more effectively than Neural Networks and LSVM because of its transparency and ease of explanation, which make it more practical for real-world business applications. According to a predictive importance analysis, order frequency, payment methods, and approval time are among the top contributors to churn risk. According to the findings, a well-structured, interpretable machine learning model can significantly help businesses identify customers at risk of churn and implement proactive retention strategies. Additionally, the study validates that marketing interventions are more precise when RFM-based segmentation is combined with predictive analytics. In this study, machine learning is demonstrated to be a useful tool to reduce churn and improve customer lifetime value by demonstrating how it can be applied to data-driven customer analytics. It would be possible to explore hybrid models, larger datasets, and the inclusion of psychological and contextual variables for a deeper level of personalization in the future. For churn mitigation strategies, it is recommended that businesses adopt explainable models such as logistic regression

    SeSP Plinko: Dynamics of C-Shaped Particles Moving Through an Obstacle Field

    Get PDF
    We examine the behavior of C-shaped superellipse sector particles (SeSPs) as they travel through an obstacle field. SeSPs are two-dimensional particles parameterized by corner sharp- ness, aspect ratio, and opening angle; these parameters are sufficient to represent a large variety of shapes, including discs, rods, and concave and convex particles. Our obstacle field is a rhombic Galton Board, an array with alternating rows of pegs. Galton Boards are common not just in experiments but in popular media, such as in boardwalk games and in The Price is Right’s game segment known as Plinko. Round particles dropped through a Galton Board distribute normally; due to their irregular shape, c-shaped SeSPs can self-clog on individual pegs, a behavior which is not observed for circular and elliptical particles. We examine and characterize the translational and rotational motion using Mean-Squared Displacements of SeSPs. In the horizontal direction, particles tend to have more diffusive, if not superdiffusive, motion in early timesteps (⟨x^2⟩ = t^(1.22±0.01)), and more subdiffusive motion in later timesteps (⟨x^2⟩ = t^((0.61±0.01)). The vertical motion is observed to be more superdiffusive (⟨y^2⟩ = t^(1.7±0.1)), and the rotational motion to be more subdiffusive with a value ⟨θ^2⟩ = t^(0.77±0.02) in the earlier timesteps and a smaller degree in the later timesteps (⟨θ^2⟩ = t^(0.10±0.02). We explore the distribution of SeSPs exiting the obstacle field–and compare this to the distribution of Galton Boards. The distribution of SeSPs exiting the board has a slight leftward skew in comparison to the normal distribution of a Galton Board with round particles (exiting most commonly in the midleft bin as opposed to the center bin). However, this skew is likely due to board tilt. We consider how soon the initial orientation of a SeSP is lost, finding that the orientation falls rapidly out of alignment with a time constant τ of 1.325 ± 0.001 before plateauing into a more gradual decline. Finally, we present how SeSP collisions within the field impact their probability of getting stuck on an obstacle, and the link between number of collisions and location of sticking, finding that the probability of a SeSP getting stuck increases by approximately 7% the further it travels down the board

    Detecting Compromised Hardware Integrity with Machine Learning in Multicore Processors

    Get PDF
    Due to the globalization of the Integrated Circuits (IC) manufacturing process, the trust surrounding hardware reliability has become compromised through methods such as injection of malicious hardware (including hardware trojans), hardware backdoors, IP theft, and counterfeit hardware. This thesis will focus on the issue of counterfeit hardware. As the complexity of modern day hardware has increased, detecting counterfeits has become increasingly difficult. Current methods of physical and electrical inspection are limited in both depth and size limitations. Physical inspections are only useful for identifying counterfeits that manifest physically, while electrical testing requires expensive and extensive setups. Computational evaluation of all possible states of the device can at times be considered unfeasible. This the- sis proposes the development of a method in which internal data traffic within the system is compared to a known golden device through the use of machine learning techniques. The technique utilized in this research will be One Class Support Vector Machine (OC-SVM) analysis. This is due to its ability to be trained on only one class of data for its functionality, and can separate out anomalous classes without being trained on that data. The system will be evaluated across a number of different benchmark programs, each with it’s own respective model trained on the known good system. By analyzing the results of the model, accuracy metrics can be used to indicate a models ability to separate the data from the two system states. The success of this method would pose another feasible method of counterfeit detection that wouldn’t be hindered by the limitations of current methods

    Fixed Points in Linear Regression

    Get PDF
    There is a set of points in the plane whose elements correspond to the observations that are used to generate a simple least-squares regression line. Each value of the independent variable in the observations matches up with one of these points, which are called pivot or fixed points. The coordinates of the fixed points are derived, and the properties of the points are explored. All points in the plane that yield each of the fixed points are found. The role that fixed points play in regression diagnostics is investigated. A new mechanical device that uses linkages to model the role of fixed points is described. A numerical example is presented

    A Review of Smart Public Transport Systems: Challenges, Technological Innovations, and Future Directions

    Get PDF
    Urbanization has created greater demand for effective and sustainable public transport systems in smart cities. Yet, overcrowding, congestion, and old infrastructure remain the main challenges to the effectiveness of these systems. This literature review discusses the major issues affecting smart public transport, such as reliability issues, environmental sustainability, and improved integration and coordination between modes of transport. Technologies like the Internet of Things (IoT), artificial intelligence (AI), and big data analytics are being applied to enhance public transportation efficiency, alleviate congestion, and improve passenger experience. Qualitative research methods were used, relying on secondary data from academic journals, government reports, and case studies. The PRISMA methodology was used to systematically choose appropriate literature, narrowing the scope to 18 high-quality articles that give detailed information on smart transport systems. Through thematic analysis, the main patterns were identified in the themes of efficiency, congestion alleviation, and sustainability. In spite of technology advances, challenges still exist, such as regulatory obstacles, finance limitations, and system fragmentation. The results indicate that the integration of cutting-edge technology, strategic planning, and policy reforms is necessary for the effective deployment of smart public transport systems, resulting in more efficient, reliable, and sustainable urban mobility solutions

    Stochastic Variational Autoencoder

    Get PDF
    This work explores a novel generative modeling approach inspired by variational autoencoders (VAEs). Traditional VAEs rely on a recognition model (encoder) that approximates the latent posterior with a single gaussian distribution for each input, limiting their flexibility in capturing complex data distributions. In contrast, we propose a modified recognition model that utilizes stochastic mixtures of gaussians, allowing for a more expressive latent representation. By leveraging stochastic neural networks within the VAE framework, we aim to achieve a tighter evidence lower bound (ELBO) on the log-likelihood of the data. Our approach is a preliminary investigation to enhance the latent space structure and improve generative performance by incorporating richer uncertainty modeling

    Predicting & Analyzing University Success: Machine Learning Approaches to Predict Performance from Socioeconomic and Educational Data

    Get PDF
    Accurately predicting learner performance is a significant challenge in education since diverse and interrelated factors contribute to academic success. High school grades and standardized test scores remain the primary indicators used for university admissions. However, such approaches fail to capture the complexities surrounding student learning. For instance, various cognitive, cultural, socioeconomic, and environmental influences shape academic outcomes. Educators need to adopt more comprehensive predictive models that consider in-depth learner differences. Reliable academic performance prediction can help schools optimize admissions decisions, allocate resources effectively, and implement targeted interventions to support at-risk students. This study adopted a quantitative primary research approach with a survey design. R was used to analyze and identify meaningful relationships within the dataset. Various statistical techniques, including descriptive statistics, correlation analysis, and predictive modeling, were applied to interpret the data effectively. The results of a Spearman correlation analysis indicated a weak positive relationship between high school scores and CGPA, r=0.39. This finding suggested that students with higher high school scores tend to have higher CGPAs. However, the weak correlation implies that additional factors beyond high school performance significantly influence university academic success. A multiple linear regression model confirmed the predictive significance of high school scores (β = 0.2263, p \u3c 0.05), EmSAT, and IELTS scores (β = 0.0886 and β = 0.00026, p \u3c 0.05) were significant in determining CGPA. However, TOEFL and SAT scores were statistically insignificant (p \u3e 0.05). The model also found that demographic factors affected academic performance. For instance, nationality had a significant impact, with non-Emirati students achieving higher CGPAs than Emirati students (β = 0.2218, p \u3c 0.05). Gender differences were also notable, as male students had lower CGPAs than female students (β = -0.2266, p \u3c 0.05). These findings highlight the importance of considering demographic and socioeconomic factors when predicting student success. The study further evaluated machine learning models for academic performance prediction. Results demonstrated that ensemble-based machine learning techniques, particularly Random Forest, outperformed deep learning approaches such as Artificial Neural Networks (ANN). The superior performance of ensemble methods suggests that future research and practical applications should prioritize models like Gradient Boosting Machines (GBM) or XGBoost to improve accuracy. Lastly, the decision tree model outperformed Random Forest and ANN with a strong correlation between the actual and predicted values, r=0.39. This study had limitations, including missing values and skewed data. These data quality issues affected the accuracy and generalizability of the findings. Future research should address these challenges by collecting a more extensive dataset, employing advanced techniques for handling missing values, and refining data preprocessing methods to minimize the impact of outliers. By leveraging more robust machine learning techniques and incorporating a broader range of predictive variables, future studies can enhance the accuracy and reliability of academic performance prediction models

    Calibration of Detectors for the Tomographic Ionized-carbon Mapping Experiment

    Get PDF
    Observations of the intensity of the sky at millimeter (mm) wavelengths require accurate calibrations for the spectrometers to work with radio telescopes. However, calibrating is a complex and lengthy process for sensitive detectors that catch single photons. To improve efficiency in locating defective detectors and optimizing bias current, my research involves taking data from previous lab characterizations and creating an algorithm that analyzes and plots the code of all 1920 detectors. This thesis will discuss the importance of mm-wavelength intensity mapping and the questions the TIME research team aims to address. Also, a detailed description of mechanical and computational components is presented. Lastly, I discuss my contribution to creating the analytical software for detector calibration and show results from my software

    Pd Chlorin-Based MolecμLar Imaging Probes for Breast Cancer

    Get PDF
    Breast cancer (BrCa) is the second most common cancer and the second leading cause of cancer death among women in the U.S. This highlights the necessity of developing effective therapeutic agents to aid in the treatment of BrCa. A promising new method of treatment is photodynamic therapy (PDT), a clinically approved therapeutic procedure that uses a porphyrin-based photosensitizer (PS) dye to kill cancer cells upon excitation with a light source. It has been observed that chelation of a metal ion into the porphyrin increases the therapeutic efficacy. We have developed a synthetic route incorporating the palladized analogue of the PS dye meso-Pyropheophorbide a (mPPa) for use as a targeted therapeutic agent for the treatment of breast cancer. Our hope is to create a non-invasive treatment option that can be used as a stand-alone treatment for small surface level tumors or as part of lumpectomy operations to kill residual cancer cells left in the margins of surgery. A disadvantage of the porphyrin rings is that they are insoluble in water. To address that issue, a previous graduate student of our lab developed a “water-solubilizing” (WS) module that contains several sulfonate groups that impart water solubility. This is attached to a peptide module that is comprised of the PS dye attached to the side chain of a lysine. The module may then be attached further to a targeting group. To improve selectivity, we attached a linear and cyclic decapeptide, 18-4 and c(18-4) which have specificity for binding to receptors on the surface of breast cancer cells, including the most lethal triple negative type (TN-BrCa). We have explored 3 different pathways and successfully synthesized and purified the penultimate puzzle piece molecule for PDT that contains a water-soluble, metal-chelated porphyrin photosensitizer module on a lysine. This compound is being carried on by members of our group with respect to coupling to c(18-4)

    Advisor Council Minutes of August 12, 2025

    No full text

    17,670

    full texts

    22,505

    metadata records
    Updated in last 30 days.
    RIT Digital Institutional Repository
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇