150813 research outputs found
Sort by
Improved Automatic Electronic Intelligence Collection System for Internal and External Forward Fusion and Collaborative Geolocation of Adversary Emitters
With the 2022 National Defense Strategy shifting
focus from counterinsurgency operations to near-peer adversaries,
airborne ISR platforms within the USAF and DoD must
be improved for effectiveness in a near-peer conflict. They
need to be able to operate quickly and effectively in contested
environments with longer-range threats, act as a forward edge
intelligence node for blue forces and provide DoD Research
and Development efforts with cutting-edge data regarding new
adversary signals and technology.
To aid in tackling these challenges, this project introduces a
Machine Learning (ML)-driven approach that revamps the Automatic
Electronic Intelligence Collection System (ACS) on U.S.
Airborne ISR platforms in four ways: First, by providing nodal
analysis to the user in real time by automatically aggregating
existing data across the aircraft to the user for decreased operator
cognitive load. Second, increasing internal aircraft database
information with external intelligence database information to
increase confidence in targeting. Third, by providing automatic
signal anomaly detection to the operator utilizing a support
vector machines algorithm that cues operators to potential
signals of interest based on previous activity and pattern of life
prediction. Lastly, by providing better surface against airborne
identification through utilization of cone angle to the system
to help operators with faster threat warning and situational
awareness of the environment.
Findings include Support Vector Machines being the most
effective tested binary classifier for predicting single signal
anomaly detection at 84% AUC and a rule-based method of
averages successfully classifying 1089 surface versus air ELINT
samples with a success rate of 89% compared to other tested
methods, such as Gaussian Mixture Models at 68% and KNearest
Neighbor at 66%.The Department of the Air Force Artificial Intelligence Accelerato
Dishing It Out: Reimagining Multicultural College Dining Through Student-Centered Design
Dining halls are central spaces in colleges, fostering not only nourishment but also cultural connection and community. However, when dining centers fall short in catering to the needs of their multicultural student body, students are often left feeling isolated and even further from home. Using MIT as a case study, this thesis employs user research and digital storytelling to explore how collecting student perspectives can inform college dining centers on better supporting the diverse cultural backgrounds and dietary needs of their students. The research and findings highlight the critical gaps and strengths in cultural representation within MIT’s dining halls. Through surveys and user research, this thesis gathers student perspectives on food authenticity, comfort, and identity, which inform the design of an interactive website prototype exploring student culinary backgrounds and preferences. This project serves as both a resource for dining services and a digital cookbook curated by the student body. By centering student voices through a culinary lens, this project aims to reimagine dining spaces as inclusive, representative, and comforting shared spaces within college campuses.S.B
Grounding Time Series in Language: Interpretable Reasoning with Large Language Models
Can large language models (LLMs) classify time-series data by reasoning like a domain expert—if given the right language? We propose a method that expresses statistical time-series features in natural language, enabling LLMs to perform classification with structured, interpretable reasoning. By grounding low-level signal descriptors in semantic context, our approach reframes time-series classification as a language-based reasoning task. We evaluate this method across 23 diverse univariate datasets spanning biomedical, sensor, and human activity domains. Despite requiring no fine-tuning, it achieves competitive accuracy compared to traditional and foundation model baselines. Our method also enables models to generate expert-style justifications, providing interpretable insights into their decision-making process. We present one of the first large-scale analyses of LLM reasoning over statistical time-series features, examining calibration, explanation structure, and reasoning behavior. This work highlights the potential of language native interfaces for interpretable and trustworthy time-series classification.M.Eng
Comparison of dispersion metrics for estimating transcriptional noise in single-cell RNA-seq data and applications to cardiomyocyte biology
Transcription is a dynamic process with a multitude of characteristics, including transcript level, burst frequency, amplitude, and variability. Single-cell RNA sequencing data analysis often focuses on comparing transcription levels. However, these analyses capture only a portion of the wealth of information conveyed by transcription. The quantification and analysis of transcriptional variability poses an opportunity to study transcription and gene regulation from a new angle. Transcriptional variability has already been implicated in a number of biological processes, including in immune system development and in aging. Yet, the most appropriate method for measuring transcriptional variability in single-cell data has remained relatively unclear. Here, we simulated single-cell data with varying dispersion and dataset size to assess the relative responsiveness of the Gini index, variance-to-mean ratio, variance, and Shannon entropy to variability in single-cell counts. We found that the variance-to-mean ratio scales approximately linearly with increasing dispersion, and that it is scale-invariant. The Gini index displayed paradoxical behavior, and Shannon entropy was not scale-invariant. Thus, we applied the variance-to-mean to measure transcriptional variability in two publicly available datasets studying congenital heart defects in mouse models. We first found that change in transcriptional variability does not correlate with gene characteristics such as transcript level and evolutionary gene age. We also found that using change in transcriptional variability to focus GSEA and TF motif enrichment analyses revealed both genes with known involvement in cardiomyopathy and new genes and pathways as potential targets for future study. Notably, many of the genes and pathways identified through transcriptional variability analysis were not found by differential expression analysis, suggesting that transcriptional variability can provide additional biologically relevant information beyond what is observed from studying mean expression alone.M.Eng
Scaling Automatic Question Generation to Large Documents: A Concept-Driven Approach
Assessing and enhancing human learning through question-answering is vital, especially when dealing with large documents, yet automating this process remains challenging. While large language models (LLMs) excel at summarization and answering queries, their ability to generate meaningful questions from lengthy texts remains underexplored. We propose Savaal, a scalable question-generation system with three objectives: (i) scalability, enabling question-generation from hundreds of pages of text (ii) depth of understanding, producing questions beyond factual recall to test conceptual reasoning, and (iii) domainindependence, automatically generating questions across diverse knowledge areas. Instead of providing an LLM with large documents as context, Savaal improves results with a threestage processing pipeline. Our evaluation with 76 human experts on 71 papers and PhD dissertations shows that Savaal generates questions that better test depth of understanding by 6.5× for dissertations and 1.5× for papers compared to a direct-prompting LLM baseline. Notably, as document length increases, Savaal’s advantages in higher question quality and lower cost become more pronounced.S.M
Biotechnology in materials science: A storied past and a bold future
The intersection of biotechnology and materials science has driven medical and scientific innovation for decades and is poised to make similar transformative impacts over the next 50 years. Advanced drug delivery systems, including nanoparticles and larger delivery material platforms, are enhancing therapeutic precision, while tissue engineering and regenerative medicine are laying the groundwork for bioprinting complex organs, offering new possibilities for transplantation and repair. Nanotechnology and biomedical devices are reshaping diagnostics and therapeutics, enabling real-time monitoring essential for personalized health care. Additionally, emerging fields such as space biotechnology and machine learning-driven biomaterials design hold potential for cutting-edge discoveries. This article examines the historical trajectory, current state-of-the-art applications, and bold future directions of biotechnology in materials science, emphasizing its impact on human health and its untapped potential yet to be explored
The Open Reaction Database
Chemical reaction data in journal articles, patents, and even electronic laboratory notebooks are currently stored in various formats, often unstructured, which presents a significant barrier to downstream applications, including the training of machine-learning models. We present the Open Reaction Database (ORD), an open-access schema and infrastructure for structuring and sharing organic reaction data, including a centralized data repository. The ORD schema supports conventional and emerging technologies, from benchtop reactions to automated high-throughput experiments and flow chemistry. The data, schema, supporting code, and web-based user interfaces are all publicly available on GitHub. Our vision is that a consistent data representation and infrastructure to support data sharing will enable downstream applications that will greatly improve the state of the art with respect to computer-aided synthesis planning, reaction prediction, and other predictive chemistry tasks
Explorations in AI and Creative Learning New Tools to Expand How Young People Imagine, Create, and Tinker with Scratch
As generative AI tools become increasingly prevalent in young people’s lives, these technologies have a growing influence over the way that children learn. While much of the early work at the intersection of AI and education has focused on the development of intelligent tutoring systems designed to deliver content more efficiently, this thesis explores how generative AI might be used to support the creative learning process by sparking curiosity, encouraging exploration, and helping young people express themselves creatively. In this thesis, I explore ways of integrating generative AI with Scratch, the world's largest programming community for children, while remaining aligned with the core values of Scratch: creativity, playfulness, and self-expression. I designed three tools that extend the Scratch ecosystem: Scratch Connect, which explores using generative AI to help Scratchers discover projects that inspire them to create while opening the black box of recommendation systems; scrAItch, which investigates how people can iterate with generative AI by using text-based inputs to create and tinker with Scratch projects; and Scratch Spark, which reimagines the new learner experience by using generative AI to help users create personally meaningful “spark projects.” This thesis describes the process of imagining, creating, and reflecting on these tools, including many of the challenges and tensions that we encountered along the way. I discuss observations and feedback from creative workshops with young people, and conclude by reflecting on open questions and opportunities for future work in designing generative AI tools that support creative learning.M.Eng
Agrammatic output in non-fluent, including Broca’s, aphasia as a rational behavior
Background: Speech of individuals with non-fluent, including Broca's, aphasia is often characterized as "agrammatic" because their output mostly consists of nouns and, to a lesser extent, verbs and lacks function words, like articles and prepositions, and correct morphological endings. Among the earliest accounts of agrammatic output in the early 1900s was the "economy of effort" idea whereby agrammatic output is construed as a way of coping with increases in the cost of language production. This idea resurfaced in the 1980s, but in general, the field of language research has largely focused on accounts of agrammatism that postulated core deficits in syntactic knowledge.
Aims: We here revisit the economy of effort hypothesis in light of increasing emphasis in cognitive science on rational and efficient behavior.
Main contribution: The critical idea is as follows: there is a cost per unit of linguistic output, and this cost is greater for patients with non-fluent aphasia. For a rational agent, this increase leads to shorter messages. Critically, the informative parts of the message should be preserved and the redundant ones (like the function words and inflectional markers) should be omitted. Although economy of effort is unlikely to provide a unifying account of agrammatic output in all patients-the relevant population is too heterogeneous and the empirical landscape too complex for any single-factor explanation-we argue that the idea of agrammatic output as a rational behavior was dismissed prematurely and appears to provide a plausible explanation for a large subset of the reported cases of expressive aphasia.
Conclusions: The rational account of expressive agrammatism should be evaluated more carefully and systematically. On the basic research side, pursuing this hypothesis may reveal how the human mind and brain optimize communicative efficiency in the presence of production difficulties. And on the applied side, this construal of expressive agrammatism emphasizes the strengths of some patients to flexibly adapt utterances in order to communicate in spite of grammatical difficulties; and focusing on these strengths may be more effective than trying to "fix" their grammar
Advanced Oral Delivery Systems for Nutraceuticals
Oral delivery is the most preferred route for nutraceuticals due to its convenience and high patient compliance. However, bioavailability is often compromised by poor solubility, instability, and first‐pass metabolism in the gastrointestinal tract. This review examines current and emerging oral delivery platforms designed to overcome these barriers and enhance nutraceutical efficacy. Traditional carriers—proteins, lipids, and carbohydrates—highlighting their delivery mechanisms and limitations, are first explored. Advancements in material science have led to novel platforms such as biodegradable polymers, metal–organic frameworks (MOFs), metal–polyphenol networks (MPNs), and 3D printing technologies. Biodegradable polymers improve stability and enable controlled release of bioactives. MOFs offer high surface area and tunable porosity for encapsulating and protecting sensitive compounds. MPNs provide biocompatible, stimuli‐responsive systems for targeted nutrient delivery. Meanwhile, 3D printing facilitates the fabrication of personalized delivery systems with precise control over composition and release kinetics, especially when integrated with artificial intelligence (AI) for precision nutrition. By comparing traditional and next‐generation strategies, this review outlines key design principles for optimizing oral delivery systems. The transformative potential of these innovations is underscored to improve the bioavailability and therapeutic outcomes of nutraceuticals, ultimately advancing personalized and targeted nutrition solutions