DR-NTU (Digital Repository of NTU)
Not a member yet
116018 research outputs found
Sort by
The synergistic effect of micro/nanostructure length scale and fluid thermophysical properties on pool boiling heat transfer
Facile surface micro/nanostructuring techniques for additively-manufactured (AM) aluminum alloy (AlSi10Mg) have recently been developed. The structuring techniques are not only highly scalable, but they also enable the tailoring of structure length scale and morphology to enhance pool boiling heat transfer coefficient. Our past study revealed that the structure cavity size of 5 µm is favorable for bubble nucleation during pool boiling of dielectric fluid, HFE-7100, resulting in significant enhancements in the heat transfer coefficients (h). However, owing to the differences in thermophysical properties between different coolant fluids, including saturation temperature, latent heat of vaporization and surface tension, the required structure size range for bubble nucleation and capillary wicking force for liquid re-supply are expected to differ significantly. To explore the effect of structure length scale on the pool boiling performance of coolants with different thermophysical properties, this work develops a new surface structuring technique consisting of a dual-stage metallurgic heat treatment process and single-stage crystallographic etching process to tune the structure length scale across nearly two orders of magnitude, viz., from 0.3 to 15 μm. Using coolant media of vastly different thermophysical properties, i.e., dielectric fluid HFE-7100 and deionized water, we show that while microcavities with sizes ranging from 3 to 8 μm are favorable bubble nucleation sites for boiling of HFE-7100, which result in the enhancement of the maximum heat transfer coefficient (hmax) by 83.4 to 103.8 % as compared to a conventional plain Al6061 surface, larger microcavity sizes of 10 to 15 μm are required to effectively promote bubble nucleation of water. This large microcavity size range of 10 to 15 μm, produced through rational nanoparticle agglomeration of the rich Si-phase in AM AlSi10Mg in elevated temperature, followed by an indirect removal process using a chemical process, is found to significantly increase hmax of water by up to 259.9 % as compared to conventional nanostructures formed on Al6061. In addition, the new AM structured surfaces also exhibit up to 26.6 % enhancement in critical heat flux (CHF) as compared to highly-wicking conventional nanostructured Al6061. In summary, by utilizing scalable fabrication techniques to tailor the structure length scale on AM AlSi10Mg, this work not only reveals the favorable microcavity sizes for bubble nucleation of different coolant fluids to enhance boiling, but it also provides useful micro/nanostructure design guidelines that can be adopted to enhance boiling of other coolants and phase change applications.Ministry of Education (MOE)Nanyang Technological UniversityNational Research Foundation (NRF)The SLM-280HL equipment used in this research is supported by the National Research Foundation, Prime Minister’s Office, Singapore under its Medium-Sized Centre funding scheme. J.Y. Ho would like to acknowledge the financial support for this project under Nanyang Technological University’s Start-up Grant (SUG) and RS14/21 MOE Tier 1 Grant provided by Ministry of Education (MOE) Singapore
Towards semantic, debiased and moment video retrieval
Video retrieval aims to retrieve a whole video within a video corpus given a language query. However, one of the main challenges is that it requires reaching a semantic correlation between these modalities. Besides, imbalanced datasets can cause biases in the retrieval models. Moreover, retrieving a moment from a video corpus corresponding to a text query is even more challenging, especially for long egocentric videos, considering the need for fine-grained cross-modal reasoning. In this thesis, we address these problems to build machines intelligent enough towards semantic, debiased, and moment video retrieval.
We first address a crucial task, semantic video retrieval, given that voluminous video clips are uploaded daily. Most approaches aim to learn a joint embedding space for plain textual and visual contents without adequately exploiting their intra-modality structures and inter-modality correlations. We propose a novel transformer that explicitly disentangles the text and video into semantic roles of objects, spatial contexts and temporal contexts with an attention scheme to learn the intra- and inter-role correlations among the three roles to discover discriminative features for matching at three hierarchical levels. The results indicate that our method outperforms the state-of-the-art methods, given the same visual backbone without pre-training. Besides, we conducted extensive ablation studies to elucidate our design choices. Finally, we also extend our method with various improvements in design choices and prove its superiority in competition.
However, these improvements can still be prone to various biases, causing models to learn spurious correlations.
Thus, we focus on a debiasing model to address a temporal bias specific to video retrieval tasks. While many studies focus on improving pre-training or developing new backbones, existing methods may suffer from the learning and inference bias issue, as recent research suggests in other text-video-related tasks. For instance, temporal object co-occurrences on video scene graph generation could induce spurious correlations. We present a unique and systematic study of a temporal bias due to frame length discrepancy between training and test sets of trimmed video clips as the first attempt to address a temporal bias in text-video retrieval tasks. We foremost hypothesise and verify the bias with a baseline study. Then, we propose a causal debiasing approach and perform extensive experiments and ablation studies on three different datasets. Our model overpasses the baseline over +2.5 points on nDCG, a semantic-relevancy-focused evaluation metric, mitigating the bias.
However, longer videos necessitate users to retrieve specific moments within videos.
Therefore, we address the Video Corpus Moment Retrieval (VCMR) task to retrieve specific moments from extensive video corpora. Our approach tackles this challenge for the first time on long, fine-grained, and untrimmed egocentric videos, while existing methodologies target short, coarse-grained, and trimmed third-person videos. This presents a formidable challenge as target moments lack textual information either from speech or subtitles and are much shorter amidst longer video sequences accompanied by shorter narrations. Our approach involves captioning moments over different timestamps using an off-the-shelf tool, enriching the captions with an LLM, and combining them with audio features within the corresponding long video to exploit the sound of object interactions. We establish three strong baselines incorporating the same additional multimodal features for a fair comparison. In two different architectural model designs, we demonstrate a 10\% to 86\% increase in summing the Recall metric over various IoUs compared to the baseline methods.
In conclusion, this thesis contributes several vital ideas from different perspectives, i.e., a novel transformer for semantic video retrieval to a causal inference method for debiased video retrieval. Besides, we leverage LLM and audio fusion to address moment retrieval in video corpus for long egocentric videos. Last but not least, we also shed light on future work directions to improve the models' capability.Doctor of Philosoph
Signals in the rock: reconstructing environmental change at the Devonian-Carboniferous boundary from sediments and teeth
The Devonian–Carboniferous boundary marks a period of significant environmental perturbation. In this study, we examine the latest Devonian sediments at Celsius Bjerg in East Greenland using an interdisciplinary approach that integrates geochemical proxies with paleontological data to reconstruct environmental and biological changes. Geochemical analyses—including measurements of stable carbon isotopes (δ¹³C of TOC and phytoclasts), total organic carbon, and palynofacies (amorphous organic matter percentage)—revealed a pronounced negative carbon isotope excursion (nCIE) at the boundary. This nCIE, together with persistently low TOC and shifts in organic matter composition, indicates a major disruption in the local carbon cycle consistent with a collapse of the terrestrial forest ecosystem and reduced primary productivity. Complementary paleontological investigations using high-resolution CT scanning and 3D segmentation uncovered an exceptionally preserved Tamiobatis sp. shark fossil. Identified by its multi-cusped heterodont teeth and ctenacanthiform dermal denticles, this early shark is a rare marine predator in continental lake deposits nearly 1000 km from the contemporaneous coastline. Its occurrence alongside abundant plant debris and freshwater fish remains suggests an opportunistic incursion into an inland ecosystem under unusual conditions like an outflow of the lake. Together, the geochemical signatures and fossil evidence provide a comprehensive picture of environmental change at Celsius Bjerg, highlighting the interconnectedness of terrestrial and aquatic ecosystems during the D–C transition. These findings advance our understanding of the complex interplay between environments during a major biotic crisis while offering valuable constraints on the mechanisms driving the nCIE at Celsius Bjerg. Nonetheless, caution is warranted in our interpretations given that only one dataset is available (Ginter et al., 2010).Bachelor's degre
Companion chatbot
This project focuses on the development of a companion chatbot tailored for elderly users, designed to provide engaging and meaningful conversations. Utilizing a large language model (LLM) as its core engine, the chatbot will handle multiple round dialogues, enabling it to sustain long conversations and proactively ask questions, creating a more interactive experience. The system will be equipped with capabilities to process both audio input and output, making it accessible for elderly individuals who may have difficulty typing. Additionally, the integration of a live Azure avatar featur though it may involve some associated fees will make the chatbot appear more human-like and interactive, further enhancing user engagement. The ultimate goal is to enhance the well-being of elderly users by offering them a virtual companion that can provide both social interaction and support, helping to reduce feelings of loneliness and isolation.Bachelor's degre
From black pigment to green energy: shedding light on melanin electrochemistry in dye-sensitized solar cells
The sustainably obtainable poly indolequinone eumelanin shows exciting properties for green, natural dye sensitized solar cell (DSSC) applications: broad absorption until 800 nm, metal ion chelation, and long conjugated, aromatic structures featuring an abundance of quinone, semiquinone, and some hydroquinone moieties providing rich binding sites. Despite this, there are limited literature works covering the use of this natural dye that typically reported conversion efficiencies of less than 0.10%. Thus, there is room for improving both the performance and understanding of the electrochemical mechanisms behind eumelanin-based DSSC applications. This work fills this gap by first characterizing eumelanin to confirm its potential stability in films on TiO2 substrates, then providing theoretical calculations on the HOMO-LUMO gap and simulating the absorption spectrum, giving promising results for potential use as a dye in DSSCs, and finally covering new ground in the optimization of the fabrication process of eumelanin-sensitized DSSCs. The prepared eumelanin DSSC devices are of high cycling stability and show a maximum performance of 0.24% before and 0.42% after treatment with UV-light. The devices were analyzed in detail to give insights into the microscopic explanation of why eumelanin-based DSSCs differ from other natural dye-based devices. Using intensity modulated photocurrent and photovoltage spectroscopy, the comparatively high recombination rate of eumelanin in relation to other natural dyes is identified as the main inhibitor to overcome in future endeavors of optimizing eumelanin films in DSSCs.Nanyang Technological UniversityPublished versionN. A.-S. is funded by the Singapore International Graduate Award from the Nanyang Technological University Singapore. J.-H. P. is funded by the National Research Foundation of Korea (NRF) (2022R1A6A3A01087285)
Development of a recommender system to choose a university degree program
In the evolving landscape of higher education, students face challenges in selecting the most suitable university and specialization based on their academic backgrounds. This project develops a hybrid recommendation system that integrates machine learning models to provide personalized specialization and university recommendations. The system incorporates Random Forest and XGBoost for content-based filtering, leveraging students’ academic profiles, such as CGPA, GRE scores, and prior work and research experience, to identify optimal choices. Additionally, memory-based collaborative filtering using cosine similarity is implemented to generate recommendations based on user similarity. To enhance adaptability, the system incorporated a dynamic α-weighting hybrid recommendation system that adjusts the influence of content-based and collaborative filtering based on model variance, improving recommendation accuracy. This approach ensures a personalized decision-making tool that aids students in navigating their higher education choices effectively.Bachelor's degre
试论用多模态视角审视《九龙城寨之围城》改编之成功 = Success of the adaptation of "Twilight of the Warriors: Walled In" through the lens of multimodality
《九龙城寨之围城》改编自同名漫画,以九龙城寨这一独特的历史空间为背景,
讲述了一群底层人物在黑帮势力下挣扎求生的故事。本文基于多模态理论,探讨漫画
到电影的改编过程,分析影像、声音、节奏等多种符号系统如何共同构建叙事,同时
评估改编的成功与局限性。论文首先考察原作中的象征性意象及其在电影中的转化,
包括叉烧饭的身份证象征意义;接着分析明星光环对电影的影响,尤其是任贤齐、洪
金宝和林峰的角色塑造如何影响观众对角色的解读;最后探讨改编过程中元素的取舍,
考察影像、声音、节奏等多模态要素的调整如何影响电影的整体呈现。本文认为,尽
管《围城》在视觉和明星效应方面强化了叙事的感染力,但在部分意象的再现及节奏
调控上存在一定不足,影响了改编的完整性。通过本研究,本文希望为漫画改编电影
的多模态分析提供新的视角,并探讨九龙城寨这一独特空间在影像叙事中的文化价值。
Twilight of the Warriors: Walled In, is a film adaptation of the manga of the same
name, set against the unique historical backdrop of Kowloon Walled City, depicting the
struggles of the underclass under the rule of triad forces. This paper, based on multimodal
theory, examines the adaptation process from manga to film, analysing how various semiotic
systems—including visuals, sound, and pacing—work together to construct the narrative
while assessing the successes and limitations of the adaptation. The study first explores the
symbolic imagery in the original work and its transformation in the film, focusing on key
motifs such as char siu rice and identity cards. It then investigates the impact of celebrity aura
on the film, particularly how the performances of Richie Jen, Sammo Hung, and Raymond
Lam shape audience perceptions of their characters. Finally, it discusses the selection and
omission of elements in the adaptation process, examining how adjustments in visuals, sound,
and pacing influence the overall presentation of the film. This paper argues that while the
movie enhances its narrative impact through strong visuals and star power, it falls short in the
recreation of certain symbolic imagery and the control of pacing, affecting the adaptation’s
overall coherence. Through this study, the paper aims to offer new perspectives on
multimodal analysis in manga-to-film adaptations and explore the cultural significance of
Kowloon Walled City in cinematic storytelling.Bachelor's degre
Dynamic mechanical response and deformation-induced co-axial nanocrystalline grains facilitating crack formation in magnesium-yttrium alloy
The dynamic mechanical response and deformation mechanism of magnesium-yttrium alloy at high strain rate were investigated using split-Hopkinson pressure bar (SHPB) impact, and the microstructure evolution and crack formation mechanism were revealed. The yield strength and work hardening rate increase significantly with increasing impact strain rate. Deformation twinning and non-basal dislocation slip are the primary deformation mechanisms during testing. Contrary to crack initiation mechanism facilitated by adiabatic shear bands, we find that high-density co-axial nanocrystalline grains form near cracks, which leads to local softening and promotes crack initiation and rapid propagation. Most grains have similar 〈1¯21¯0〉 orientations, with unique misorientation of 24°, 32°, 62°, 78° and 90° between adjacent grains, suggesting that these grains are primarily formed by interface transformation, which exhibits distinct differences from recrystallized grains. Our results shed light upon the dynamic mechanical response and crack formation mechanism in magnesium alloys under impact deformation.Published versionThe authors acknowledge the support from the National Natural Science Foundation of China (Grant Nos. 52301137, 51974097, and 52364050), the Natural Science Special Foundation of Guizhou University (No. (2023) 20), Guizhou Province Science and Technology Project (Grant Nos. [2023]001, and [2019] 2163), Guiyang city Science and Technology Project (Grant No. [2023]48-16)
Fast image inpainting using accelerated diffusion models
Image inpainting has gained increasing attention due to its crucial role in editing and
restoring visual content. Despite significant advances through diffusion-based mod-
els such as Stable Diffusion, these methods remain computationally expensive, often
requiring dozens of iterative denoising steps to produce high-quality results. In this
paper, we explore how Latent Consistency Models (LCM) and Low-Rank Adaptation
(LoRA) can be adapted from text-to-image pipelines to an inpainting setting, aiming
to reduce the number of inference steps while maintaining fidelity. We first review
the evolution of inpainting methods, from classical patch-based and PDE approaches
to modern diffusion-based pipelines, and highlight the computational overhead asso-
ciated with large-scale generation. We then detail our methodology for integrating
LCM and LoRA into pre-trained diffusion backbones, placing emphasis on masking
strategies, dataset preparation, and distillation workflows. Our results show that LCM
holds the potential to effectively reduce the number of inference steps required for
inpainting while maintaining acceptable visual fidelity, whereas LoRA—under our
constraints—yields less consistent outcomes. These findings suggest that LCM pro-
vides a promising avenue for faster, high-resolution inpainting, particularly when a
balance of speed and quality is paramount.Bachelor's degre
Development of an AI-powered exchange platform
This report details the development of NTU Quant AI, a quantitative trading platform designed to provide traders with a seamless trading experience across multiple asset classes and markets via a single platform. The motivations for this project stem from the growing prominence of algorithmic trading and increasing expectations for flexibility in strategy deployment.
The report starts with an introduction outlining the project's motivation, overview, and scope. It then delves into the tools and methodologies employed during development and provides a comprehensive overview of the project's progress across its past, present, and future phases. Several foundational components were developed in the earlier phases of the project, including the Portfolio Management System (PMS), Order Management System (OMS), and Execution Management System (EMS). In the most recent phase, the focus was on revamping underperforming frontend components to improve the user experience, refining the strategy pipeline that supports the execution of user scripts, improving the data collection system to increase both its reliability and variety, and introducing a manual imputation system to enable
users to test their strategies effectively.
Beyond the analysis of these components and their technical implementations, the report also explores essential software engineering concepts, such as the software development life cycle. Additionally, the report outlines potential areas for future expansion to further enhance the platform's capabilities.Bachelor's degre