Hong Kong University of Science and Technology
Hong Kong University of Science and Technology Institutional RepositoryNot a member yet
162821 research outputs found
Sort by
Counterwind currents in the China shelf seas: the origin of the pressure gradient driver
There exist unique northeastward counterwind currents (CWCs) over the China shelf seas, subject to the balance between the northeastward alongshore ageostrophic (effective) pressure gradient force (PGFeff, the residual PGF after balancing the Coriolis force) and southwestward surface wind stress forcing. The underlying physics for the formation of the alongshore PGF remains largely ambiguous. We used process-oriented modeling of the China shelf seas to investigate how the alongshore PGF and subsequent CWC form. Driven by a typical alongshore variable density field and wind forcing, our numerical model produced a realistic shelf current structure and CWC. Our results show that the alongshore sea level elevation gradient, induced by an alongshore variable wind, mainly contributed to the PGF, which is consistent with arrested topographic wave theory. However, the wind-induced elevation gradient alone is insufficient to overcome the frictional effects of wind forcing to form the CWC. The alongshore density gradient, due to heterogeneous heating, enhances the PGF because of the steric effect on sea level. The enhanced PGF produced by the density gradient and wind-induced elevation gradient critically forms the CWC. In addition, the seasonally variable density gradient and wind forcing determine the spatiotemporal variability of the CWC. We found that the dominant intrinsic dynamics of the wind and buoyancy forcing are enough to trigger the CWCs, and the shelf’s topographic features further shape the structure of the CWC. The study provides new insights into the formation of CWC, which has been widely observed over the shelves globally.</p
MultiUX: a human-AI collaborative tool to facilitate multiple usability test video analyses
Analyzing multiple usability videos is essential for identifying common problems and prioritizing them. While existing tools assist in identifying problems within single videos, the manual process of comparing and categorizing them remains time-consuming. Inspired by prior work that suggested a human-AI collaborative approach could improve the efficiency and completeness of analyzing single videos compared to either humans or AI alone, we extended this approach to investigate its effectiveness in analyzing multiple videos. We designed MultiUX to assist in strategically analyzing multiple videos by comparing and categorizing problems across them, and optimizing the analysis sequence through video recommendations. MultiUX was compared to a baseline in a between-subjects study involving 20 UX evaluators. Results showed that MultiUX supported analyzing more videos and identifying more common problems that were encountered by more users in a given time. Additionally, MultiUX guided participants to analyze multiple videos more strategically.</p
Mini-Gemini: Mining the Potential of Multi-Modality Vision Language Models
In this work, we introduce Mini-Gemini, a simple and effective framework enhancing multi-modality Vision Language Models (VLMs). Despite the advancements in VLMs facilitating basic visual dialog and reasoning, a performance gap persists compared to advanced models like GPT-4 and Gemini. We propose a novel approach to narrow the gap by mining the potential of VLMs for better performance across various cross-modal tasks. It tackles the following questions: (1) How can high-resolution visual tokens improve image understanding without lengthening the token sequence? (2) How to improve reasoning and generation abilities of VLM with high-quality data? (3) How to close the gap between open-source VLMs and proprietary models on reasoning-driven generation? In particular, to enhance visual tokens, we propose to utilize an additional visual encoder for high-resolution refinement without increasing the visual token count. We further construct a high-quality dataset that promotes precise image comprehension and reasoning-based generation, expanding the operational scope of current VLMs. In general, Mini-Gemini further mines the potential of VLMs and empowers current frameworks with image understanding, reasoning, and generation simultaneously. The proposed model supports a series of dense and MoE Large Language Models (LLMs) from 2B to 34B, which achieve leading performance in several zero-shot benchmarks and even surpasses the developed private models. It is demonstrated to attain 80.6% accuracy on the MMB benchmark (+5.4 vs Gemini Pro) and 74.1% on TextVQA (+4.6 vs LLaVA-NeXT), achieving leading performance in several zero-shot benchmarks and even surpasses the developed private models. Furthermore, Mini-Gemini is proven to improve consistently with stronger LLM, visual encoder, and data in experiments.</p
Driver Recipient Selection for Traffic Safety Education via Uplift Modeling
Safety has been one of the major concerns by the traffic police in traffic management. A favored practice to enhance traffic safety is to educate drivers and raise their safety awareness, so as to reduce the occurrence of traffic accidents in general. However, human power is limited in carrying out the education program, and challenges remain in filtering the most proper group of drivers to receive such education as well as in assessing the effectiveness of the program. In this paper, we view traffic safety education as an intervention from the perspective of causal inference, and we address the driver recipient selection (DRS) problem as a combination of uplift modeling and optimization. In uplift modeling, we identify that the confounding bias is present in historical accident data, and hence we adapt the uplift model via inverse propensity scoring (IPS) to eliminate the confounding bias. Experiments on both synthetic and real-world datasets show that our adapted uplift model increases the Area Under the Unconfounded Uplift Curve (AUUUC) by up to 46%, and our proposed DRS strategy can further reduce the overall monthly accident rate by 3.4% absolutely than the existing strategy.</p
RegScorer: Learning to select the best transformation of point cloud registration
We propose RegScorer, a model learning to identify the optimal transformation to register unaligned point clouds. Existing registration advancements can generate a set of candidate transformations, which are then evaluated using conventional metrics such as Inlier Ratio (IR), Mean Squared Error (MSE) or Chamfer Distance (CD). The candidate achieving the best score is selected as the final result. However, we argue that these metrics often fail to select the correct transformation, especially in challenging scenarios involving symmetric objects, repetitive structures, or low-overlap regions. This leads to significant degradation in registration performance, a problem that has long been overlooked. The core issue lies in their limited focus on local geometric consistency and inability to capture two key conflict cases of misalignment: (1) point pairs that are spatially close after alignment but have conflicting features, and (2) point pairs with high feature similarity but large spatial distances after alignment. To address this, we propose RegScorer, which models both the spatial and feature relationships of all point pairs. This allows RegScorer to learn to capture the above conflict cases and provides a more reliable score for transformation quality. On the 3DLoMatch and ScanNet datasets, RegScorer demonstrate 19.3% and 14.1% improvements in registration recall, leading to 4.7% and 5.1% accuracy gains in multiview registration. Moreover, when generalized to symmetric and low-texture outdoor scenes, RegScorer achieves a 25% increase in transformation recall over IR metric, highlighting its robustness and generalizability. The pre-trained model and the complete code repository can be accessed at https://github.com/WHU-USI3DV/RegScorer.</p
Intra-city scale graph neural networks enhance short-term air temperature forecasting
Air temperature (Ta) has critical implications for various socioeconomic sectors, yet its dynamics are particularly complex in urban areas due to heterogeneous built environments, landscapes, and diverse anthropogenic activities. Physics-based models struggle with intra-city Ta forecasts due to inadequate urban representation and limited spatial resolution. While weather observation networks offer promising alternatives for direct modeling with local Ta time-series, an effective framework to leverage these intra-city discrete sensor data remains lacking. Here, we demonstrate that graph neural networks (GNNs) can harness observation network information to refine Ta prediction at individual locations and elucidate underlying mechanisms. Our novel Mix-n-Scale framework with GNNs achieves over 12 % improvement in short-term Ta forecasts compared to conventional local time-series approaches. Further model evaluation disentangles performance variations with local Ta variability in diverse spatiotemporal contexts, indicating distinct patterns of intra-city heterogeneity across seasonal and diurnal scales. Our findings establish graph-based approaches for leveraging proliferating urban sensor data and advancing understanding of Ta spatiotemporal dynamics in complex urban environments.</p
Sub-zero Celsius elastocaloric cooling via low-transition-temperature alloys
Elastocaloric cooling using shape-memory alloys (SMAs) is a promising greenhouse gas (GHG)-free alternative to conventional vapour-compression refrigeration that relies on high global warming potential (GWP) gas refrigerants1, 2, 3–4. However, existing elastocaloric systems have not yet reached sub-zero Celsius temperatures, which restricts their application in various freezing scenarios5,6. Here we constructed a compression-based, regenerative elastocaloric cooling device using low-transition-temperature tubular NiTi units in a cascaded configuration. The selected NiTi alloy exhibited superelasticity and substantial entropy changes down to −20 °C. Moreover, low-freezing-point aqueous calcium chloride solution was used as the heat-transfer fluid, ensuring effective flow at low operational temperatures. Our desktop device achieved a heat-source temperature of −12 °C from a room-temperature heat sink, paving the way for next-generation green elastocaloric freezing technologies.</p
Mitigating electric leakage by donor impurity and lattice compatibility in phase-transforming ferroelectric materials
Phase-transforming ferroelectrics are widely utilized in pyroelectric devices. However, elevated temperatures lead to significantly increased electric leakage. Mitigating electric leakage near the transformation temperatures is essential for device functionality and lifetime. In this work, we tune the lattice parameters in fine-grained, donor-doped ferroelectric Eu-BTO-Zrxmaterials and discover that tuning lattice parameters effectively suppresses electric leakage. Through first-principles calculations, we investigate the influence of lattice parameters on donor energy level and discover that the donor energy level shifts significantly in Eu-BTO-Zr5 composition, coinciding with the observed leakage mitigation at tetragonal-to-cubic transformation. We analyze the lattice compatibility of developed materials and reveal that the leakage mitigation appears in materials with lattice parameters equating structural anisotropy and lattice compatibility. Our findings clarify the coupled role of donor states and lattice parameters in leakage mitigation across phase transformation, providing a theoretical framework for designing pyroelectric materials with improved thermal stability and reduced leakage.</p
Thermal conductivity suppression by extreme stress gradients in bent cracked silicon nanowires
Precise control of thermal transport is crucial for increasing the energy efficiency and reliability of modern electronic devices. While the effects of moderate strain gradients on thermal conductivity have been studied previously, the underlying mechanisms governing heat transport under extreme stress gradients (greater than 1 GPa/Å) remain elusive. In this study, we use molecular dynamics simulations and introduce a bending-induced loading method on cracked silicon nanowires to generate localized extreme stress gradients near the crack tip. When subjected to a 12% bending strain (with the stress at the crack tip exceeding 1 GPa/Å), the cracked nanowire exhibited a 20.1% reduction in thermal conductivity. This is more than double the 8.3% reduction observed in its crack-free counterpart. Detailed atomistic analysis reveals that extreme stress gradients break lattice translational symmetry, disrupt vibrational coherence, and intensify phonon scattering. These effects eventually promote phonon localization and create a thermal transport bottleneck that significantly suppresses heat conduction. This study not only explores the microscopic physical mechanisms underlying the influence of extreme stress gradients on thermal conductivity, filling existing theoretical gaps, but also provides new insights for heat dissipation design and material development for microelectronics in extreme environments.</p
A review on the Time–Accurate and highly–Stable Explicit (TASE) scheme for solving stiff differential equations
The stiff problem within Ordinary Differential Equations (ODEs) and Partial Differential Equations (PDEs) presents numerous challenges for the stability and convergence of numerical methodologies due to their significant differences in scale. While explicit time-marching schemes have their advantages, their need for extremely small time steps significantly sacrifices computational efficiency. Implicit time-marching schemes, on the other hand, allow for larger time steps with better stability properties, where traditional schemes such as the diagonally implicit Runge-Kutta (DIRK) method and the implicit-explicit (IMEX) Runge-Kutta scheme are already widely used. However, when it comes to nonlinear problems, we still need to solve the nonlinear implicit equation, which is fundamentally difficult at high-order accuracy. To tackle this, the Time-Accurate and highly-Stable Explicit (TASE) operators were proposed. Differing from the traditional implicit time-marching schemes, TASE operators are preconditioners for existing explicit time-marching schemes, such as the explicit Runge-Kutta (RK) schemes, where their combination enables RK schemes to solve stiff problems with larger time steps and enhances stability. Furthermore, TASE operators are linear in nature, avoiding the need to solve non-linear problems, where the accuracy of TASE operators theoretically can also be of an arbitrarily high order through Richardson extrapolation. These inherent advantages have led to the rapid growth of the family of TASE-schemes recently, including theoretical analysis and algorithmic improvements. In this review, the TASE operators and their variants are summarised, highlighting their stability properties, parameter settings, comparisons with traditional implicit time-marching schemes, and promising future directions of the TASE family of operators.<br/