Rochester Institute of Technology

RIT Digital Institutional Repository
Not a member yet
    22505 research outputs found

    Advanced Photonic Integration and Packaging: Efficient Phase Shifters, Micro-Transfer Printed Modulators, and Photonic Wire-Bonded On-Chip Lasers

    No full text
    Silicon photonics stands at the forefront of technological innovation, seamlessly merging the worlds of electronics and optics on a single silicon chip, creating significant advancements in high-speed communication, data processing, and energy-efficient computing. Ongoing research is focused on achieving efficient phase shifting, high-speed modulation, compatibility with CMOS processing, on-chip laser integration, and cost-effective packaging. All of these are critical areas driving advancements in the field. In this work we make contributions to improvements in each of these areas. Efficient Phase Shifters - Ultra efficient thermo-optic phase shifters: We achieved wafer-scale compatible thermally-isolated phase shifters that operate \u3e15 times (Pπ = 1.2 mW) more efficiently than standard thermo-optic devices. Micro-Transfer printed thin-film lithium niobate (TFLN) phase shifters: We have demonstrated thin film lithium niobate (TFLN) coupon fabrication and transfer printing process to achieve high-speed electro-optic modulation at visible wavelengths which find critical applications in quantum photonic research. With our current process using the Xceleprint micro transfer printing (MTP) system, we are able to transfer ring-shaped 100 μm radius coupons that have 13 μm width and 100 nm thickness and create electro-optic modulators at visible wavelengths that operate at 400 MHz (measurement limited) with a tuning efficiency of 2.5 pm/V. MEMS Phase Shifters: To achieve a large neff tuning within a compact design footprint, we present the design, modelling, fabrication and testing of low-voltage silicon photonic MEMs phase shifters. With devices fabricated at RIT, we are able to demonstrate wafer-scale release of the MEMs structures and get Vπ of 7 V. Laser integration using photonic wire bonding (PWB): Integrating III-V lasers onto silicon photonic integrated circuits remains challenging due to material mismatch, CMOS process compatibility, strict alignment tolerances, and packaging requirements such as thermal management and reliable optical coupling. III-V on- silicon lasers are commonly achieved through direct epitaxial growth (monolithic integration), wafer bonding (heterogeneous integration), and die bonding or pick-and-place (hybrid integration). Some approaches limit pre-integration testing of the individual laser dies, require pristine bond quality surfaces and tight process control at wafer scale. In this work, an on-chip integrated DFB laser is demonstrated using the Vanguard photonic wire bonder to create 3D polymerized waveguides for efficient optical coupling. Single-mode lasing is achieved by implementing a second photonic wire bond in the assembly. This creates an adiabatic transition of the optical mode and helps mitigate interface reflections. This improvement is verified through a series of circulator measurements. In addition, thermal wavelength tuning up to 4 nm is demonstrated by using a glass substrate to improve heat confinement in the assembly compared to conventional silicon

    Detecting Illicit Bitcoin Transactions on the Dark Web

    No full text
    Dark web markets use Bitcoin and related cryptocurrencies to transfer and launder proceeds of crime while obscuring the real-world identities of operators. This creates an increasingly challenging environment for regulators and law-enforcement agencies, who must analyse suspi- cious transactions on decentralised, public, pseudonymous blockchains and increasingly dense and noisy Bitcoin transaction graphs. This thesis investigates whether, and how, blockchain analytics and machine learning can help detect illicit transactions on Bitcoin networks. The first part of the thesis is a systematic literature review of recent contributions from academia and industry that examine major solution approaches for identifying illicit flows on public blockchains, particularly within dark market ecosystems such as Silk Road and its successors. The review identifies recurring patterns in data sources, feature engineering strategies, and model families used for flow identification, and highlights key gaps related to labelling, performance evaluation, and practical deployment. The second part is an applied case study based on the publicly available Elliptic Bitcoin transaction graph dataset, an extensive graph constructed using law-enforcement intelligence with labelled licit and illicit transactions. After data cleaning and a temporal split into train- ing, validation, and test sets, two supervised learning models—Random Forest and Extreme Gradient Boosting (XGBoost)—are trained on a representative subset of informative graph- based and transaction-based features. Model performance is evaluated on the test set, with particular emphasis on the illicit class, using precision, recall, F1-score, and ROC–AUC. Random Forest achieves the highest overall accuracy and recall, while XGBoost delivers very similar performance with competitive AUC scores and stable behaviour. These findings indicate that tree-based ensemble methods can classify Bitcoin transactions effectively using anonymised, non-identifying features derived from the transaction graph. The third part is a simple Flask web application presented as a proof of concept for end-to- end model deployment. It shows how submissions of Elliptic-formatted transaction records can be scored by the trained models and returned with corresponding risk assessments, and sketches how real-time transaction data could be obtained from a public blockchain API in a future system. However, because the anonymised Elliptic features do not map one-to-one onto raw blockchain fields, this integration remains demonstrative rather than fully operational. The thesis concludes with practical implications and recommendations for law-enforcement agen- cies and suggests future research directions, including more advanced graph-based approaches and improved labelling strategies

    A Search for Serotyping Candidates in the Streptococcus sobrinus Genome

    No full text
    Streptococcus sobrinus (SSO) is a pathogenic microbe native to the human oral microbiome. SSO is historically implicated in carie formation, and may be involved in neurodegenerative disease progression. The study of SSO is challenged by a lack of access to a serotyping assay. This study aims to identify sero-specific candidates from available SSO genomes using sequence alignment and phylogenetic analysis. Three candidate genes were found that may be sero-specific: putative quinol monoxygenase, a hypothetical protein, and elongation factor Tu. To further assess the candidates the original antibodies or a list of proteins involved is required. Alternatively, a new classification system should be developed to replace the concept of serology for SSO exclusively

    Understanding Crime Through Harm: A Multi-Stage Study of Rochester, NY

    No full text
    This capstone project develops, critiques, and applies a Crime Harm Index (CHI) for the City of Rochester, New York, as an alternative to traditional crime measures that rely solely on raw counts. Rather than treating all offenses as equal “one crime = one unit,” the project weights crimes by the severity of harm they cause, using New York State sentencing guidelines as the core metric. Across four working papers, the project moves from conceptual foundations and limitations of harm-based metrics, to the construction of a Rochester-specific Crime Harm Index, and finally to a set of spatial regression models that link neighborhood disadvantage to crime harm. Together, these sections aim to show not only what a CHI can reveal about crime patterns, but also how and when it should be used in practice. Part One establishes the conceptual framework by defining a harm index and explaining why we might want to measure crime in terms of harm rather than frequency. It examines the Cambridge Crime Harm Index (CHI), the dominant paradigm for harm-based measurement, and explains how it employs national sentencing guidelines to translate offenses into days of imprisonment as a standardized unit of harm. The paper emphasizes the approach\u27s virtues, including its cost-effectiveness, democratic basis in sentencing guidelines, and capacity to prioritize high-harm crimes that are frequently overlooked by traditional count-based metrics. It also briefly compares crime harm to other harm models, such as the drug harm and road harm indexes, which use health, financial, and community-level expenses to assess harm in their respective fields. This section discusses crucial concepts such as the three-pronged test (democracy, dependability, and cost) as a benchmark for determining if a damage index is applicable to real-world policy and enforcement. Part Two takes a step back and asks, What are the limitations of using harm indexes, particularly when they are based on sentencing guidelines? It analyzes how moral panics, public anxiety, and politics can drive sentencing policy, resulting in harm weights that do not always reflect the genuine social impact of an offense. This part also examines various methods of determining criminal severity, such as public opinion panels, victim surveys, and court records, and demonstrates how each introduces biases, ranging from media-influenced judgments to highly emotional victim answers. This section also discusses the issue of underreporting, pointing out that crimes like auto theft are routinely reported, whereas crimes like rape and sexual assault are significantly underreported. Finally, this section contends that, while harm indexes based on sentencing guidelines are imperfect, they remain a useful starting point. Part Three transitions from theory to implementation by creating a Crime Harm Index specifically for the City of Rochester. The Rochester CHI is built using crime data from 2023-2024, with New York State sentencing guidelines applied to 20 major crime types and harm scores calculated based on days of incarceration. Utilizing the same framework, the section also examines Rochester from 2009 to 2024, utilizing tables and visualizations to compare total crime counts and total crime harm over time. This enables the research to demonstrate how the picture changes when severity is taken into consideration, which crimes cause the most harm to the community, and how a harm-based lens indicates different priorities than typical count-based crime statistics. Part Four utilizes Rochester’s CHI to examine how neighborhood disadvantage correlates to crime harm at the census tract level. Using spatial regression models, this section investigates whether more disadvantaged tracts have higher rates of Robbery, Motor Vehicle Theft, and Dangerous Weapons Harm. Even after controlling for what happens in surrounding tracts, the findings demonstrate that higher neighborhood disadvantage is associated with increased robbery and weapons-related harm. On the contrary, motor vehicle theft harm appeared to be influenced more by opportunity and spatial spread than by disadvantage alone. Together, these models demonstrate how the CHI may be used not simply to define harm, but also to examine how socioeconomic inequality and location influence the distribution of that harm throughout Rochester

    Predicting Loan Defaults: A Behavioral Scoring Approach for Portfolio Risk Monitoring in Banking and P2P Lending

    No full text
    The critical challenge of lending institutions and peer-to-peer investors is to single out high-risk active borrowers in order to avoid defaults and corresponding losses. This thesis is devoted to the development of a data-driven approach to predict loan defaults, which balances predictive accuracy with interpretability and cost-sensitive evaluation. Using a large dataset of loans from LendingClub, we applied the CRISP-DM methodology, performing extensive data preprocessing and exploratory analysis before training several machine learning models-logistic regression, decision tree, random forest, and XGBoost-to classify loans as default or non-default. Class imbalance was addressed through resampling, and models were tuned via cross-validation with a focus on maximizing the area under the ROC curve (ROC-AUC) and recall (sensitivity) to prioritize catching default cases. The results show that ensemble tree-based models significantly outperformed the baseline logistic model on predictive performance, yielding test ROC-AUC values of about 0.96–0.97. The best model, an ensemble, was able to capture approximately 85% of the loans that defaulted in the test set, a huge improvement compared with the traditional approach. For the model interpretability and to keep the modeling transparent, SHAP was used. The most influential factors were intuitive and included features describing loan repayment progress-for example, the proportion of principal repaid-loan grade and interest rate, which are indicative of the borrower’s credit quality-and loan amount, among others. Loans with very little principal repaid or higher risk grades were strongly associated with default outcomes, as would be expected from domain knowledge. By placing such emphasis on recall in model evaluation, the framework directly addresses the asymmetric cost of misclassification in lending-missing a default (false negative) is far more costly than a false alarm. The final model and its explanations together form a practical decision-support tool for lenders, which enables better-informed decisions regarding portfolio risk management and early warning for existing loans. In general, this thesis develops an interpretable and financially informed machine learning solution for loan default prediction that will help reduce default rates and support risk management in lending portfolios

    Multilingual Identity Document Information Extraction via Dynamic Templates and Hybrid OCR

    No full text
    The thesis introduces a unified, automated document intelligence system that can in-house resolve significant issues in identity authentication and data OCR by two complementary technological components: advanced facial biometrics and powerful multilingual optical character recognition. The system directly addresses the inefficiencies of the slow speed of document processing by humans, human error and high cost of operation by offering a single pipeline for verifying the identity of the user via facial comparison and digitizing textual data contained in documents. The former module adopts an advanced face verification pipeline. It starts with an image pre-processing step, which improves the quality of the document by denoising, contrast enhancement, sharpening, as well as gamma adjustment. The resulting processed images are analysed by the RetinaFace model, which identifies faces with high accuracy and retrieves facial landmarks. These landmarks create a geometric embedding that can be used to match faces efficiently using cosine similarity, producing high-quality MATCH/NOT MATCH decisions that are validated on a variety of datasets such as Iranian and Egyptian ID cards. The second module provides text extraction functionalities in different languages using the EasyOCR engine. This component shows impressive linguistic capability and is able to handle documents written in three different script systems: English invoices with tabular layouts, Chinese student IDs with logographic characters, and Arabic documents with right-to-left text. The module not only returns the extracted text with confidence scores but also visual annotations with bounding boxes, which make the results transparent and verifiable. This combined system has practical applications in financial services, border control, and human resources by ensuring that identity verification is integrated with data digitization in a single workflow. The system is based on open-source technologies and includes a modular architecture, allowing it to scale in the future by adding more accurate facial recognition models, additional language support, and layout analysis to handle more complex documents such as passports and driver’s licenses. This study demonstrates how combining state-of-the-art computer vision and OCR technologies can address urgent challenges in real-world document processing

    Leveraging AI in Traffic Monitoring for Improved Accident Prediction in UAE

    No full text
    This paper investigates how Machine Learning (ML) models can be used in traffic surveillance and predict accidents in the United Arab Emirates (UAE). In spite of large infrastructure developments and stringent traffic laws, traffic accidents continue to be a major problem with an increase in fatalities and injuries. There is a gap in effective detection and responses since traditional monitoring and forecasting techniques are unable to capture the intricate, nonlinear, and time-dependent character of traffic patterns. This thesis fills that void by utilizing sophisticated the potential use of Artificial Neural Network (ANN), Recurrent Neural Network (RNN), Deep Neural Network (DNN), Convolutional Neural Network (CNN), and Reinforcement Learning (RL) models to produce better predictors becomes relevant with the everincreasing number of traffic-related accidents and further demand of intelligent transportation systems. A real-world traffic accident data set that was acquired from MOI (Ministry of Interior) located in UAE. These open-source datasets obtained more than 6,400 accident records which was examined, including such characteristics as type, weather, questionable condition of the road, lighting, time, and the extent of injuries. Following the pre-processing, training, testing, and comparing models were done in terms of metrics of accuracy, training time, and loss convergence. Five Python (TensorFlow/Keras) ML models were constructed and assessed using the same training/testing splits According to the results, RNN had the highest accuracy (~91%) and was able to capture time-dependent variables such daily or seasonal variations. CNN, which has a great feature-learning capability, came in second (around 90%). ANN and DNN were effective for batch analytics, achieving about 88% with reduced training periods. The lack of a real-time simulation framework caused RL to score the worst (~76%) despite its theoretical promise for adaptable decision-making. Time, weather, illumination, and accident type were the most significant indicators, according to SHAP analysis

    ENHANCING AIRPORT OPERATIONS WITH AI FOR A SEAMLESS PASSENGER EXPERIENCE

    No full text
    The paper is research exploring the importance of Artificial Intelligence (AI) and Data Analytics to optimize airport operations and the passenger experience in the environment of the expanding air traffic in the world and the smart city movement. With increasing tasks that airports currently experience, congestion, flight delays, mishandling of baggage, and limited capacity, nowadays AI-based technologies integration is crucial to the operational efficiency, sustainability, and customer satisfaction. This study discusses the use of predictive analytics, machine learning, and clustering models to enhance passenger flows, resource allocation, and performance in general at airports. The research is based on the working efficiency of operations and Smart Airport 4.0 that focus on the use of data to make decisions and the combination of AI, IoT, and analytics used in transport infrastructure. The context of the research is based on the boom in air travels in the period 2005-2018 when passenger numbers are continuously growing and there is need to be smarter in operational approaches. The study will be informed by four following questions: How can artificial intelligence be used effectively to optimize the daily procedures of the airport? What quantifiable changes have been induced by the introduction of AI in terms of elimination of flight delays and enhancement of throughput? What AI solutions can best be used to improve the efficiency and predict demand? Which are the major paradoxes to the introduction of AI into the complicated airport systems? The data collection and analysis strategy involved these questions using secondary data derived out of the Federal Aviation Administration (FAA) and TartanAviation, which contains 18,885 records of air traffic and passenger statistics across 13 years. The study is based on a quantitative data-focused approach that corresponds to a positivist paradigm. They prepared data by cleaning, transformation, and feature engineering to be able to perform advanced analysis. The models such as Linear Regression, random forests, and a Gradient Boosting predictive model have been created to predict passenger volumes, and the K-Means clustering algorithm has been used to identify trends of airline activities. Moving averages and time-series decomposition have also been used to bring out the seasonal variation and growth trends in the long-term. The strength and accuracy of the models were proved with the help of such evaluation metrics as R 2, R MSE, and Silhouette scores. The results show that AI has a greater impact on the capacity to anticipate the demand in passengers, minimize congestion, and foresee operations choke points. Random Forest model (R2 = 0.995) was found to have better predictive performance and hence it is worth considering it in real-time predictive passenger forecasting. Clustering also based on the operational categories of airlines showed four key categories of airlines, whereas the time-series analysis provided some ideas of peaks in relation to the season around July and August. This research comes to the conclusion that AI is not only capable of increasing the predictive capacity but also helps to make proactive decisions and optimize resources

    Assessing Large Language Models as an Interpretive Layer in Marketing Mix Modeling: Implications for Marketing Analytics

    No full text
    Marketing mix modelling (MMM) remains a core technique for guiding budget allocation, yet its outputs are often difficult for non-technical planners to interpret and govern. At the same time, large language models (LLMs) offer new possibilities for translating complex model artefacts into narrative guidance, but raise concerns about hallucination, reproducibility, and alignment with model-risk governance. This thesis examines whether an open-source MMM framework can be engineered as a repro- ducible, governance-ready pipeline and then augmented with a tightly constrained LLM interpretive layer. The empirical setting is a multi-brand, multi-country retail portfolio with several years of digital marketing and outcome data across multiple online platforms. The MMM implementa- tion builds on Meta’s open-source Robynpackage and adds phase-gated execution, pre-registered acceptance gates, configuration-over-code design, and manifest-based artefact tracking so that models can be re-run and audited. On top of this stack, the thesis designs and implements an LLM interpretive layer that consumes a frozen “grounding bundle” of MMM artefacts via a command-line interface. The LLM is restricted to answering questions using these artefacts, governed by explicit refusal rules and provenance tagging. A technical evaluation based on a scripted set of command-line tests, executed by the researcher across the most reliable MMM models, assesses grounding fidelity, refusal behaviour, and alignment with the pre-registered governance gates. The contribution is twofold: (i) a concrete pattern for implementing a reproducible, phase-gated MMM pipeline using open-source tools; and (ii) a governance-aware LLM interpretive layer that can be evaluated technically and positioned for future user studies. The thesis concludes that LLMs can act as useful, controllable translators of MMM outputs when anchored to a strong governance substrate, while emphasising that broader user-centred evaluation and causal extensions remain important directions for future work

    A Machine Learning Framework for Forecasting Passenger Flow at Dubai International Airport

    No full text
    The fast growth of global air travel made the accurate prediction of passenger flow a challenge for airport operations. Traditional forecasting models, like regression and SARIMA, have been useful in stable conditions but usually do not show the non-linear and dynamic changes driven by weather, infrastructure, and behaviour of passengers. Therefore, recent studies showed the potential of machine learning and big data to improve the accuracy of forecasting. However, most of the work is still limited to specific airports, datasets, or operations. This study develops a machine learning–based predictive framework for short-interval passenger throughput at Dubai International Airport. Using available flight movement data from Dubai Pulse, high-granularity 15-minute operational datasets were constructed following extensive cleaning, temporal alignment, and feature engineering. Multiple models were implemented and evaluated. Model performance was assessed using accuracy, precision, recall, and F1-score derived from multi-class confusion matrices. The experimental results show that machine learning, particularly LSTM networks, improves forecasting accuracy compared to conventional approaches. The framework gives reliable short-interval predictions that can support proactive staff planning, passenger queue management, and operational decision-making

    17,670

    full texts

    22,505

    metadata records
    Updated in last 30 days.
    RIT Digital Institutional Repository
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇