Hong Kong University of Science and Technology
Hong Kong University of Science and Technology Institutional RepositoryNot a member yet
162821 research outputs found
Sort by
Coupled thermo-hydrodynamic-mechanical peridynamics for thermal fluid-solid interactions with fracturing
This paper presents a thermo-hydrodynamic-mechanical (THM) peridynamics (PD) method for thermal fluid-solid interaction (FSI) involving fracturing in solids. Both fluid and solid materials are treated using the PD formulation. For solids, we employ a total-Lagrangian description consistent with classical PD, while for fluids, we develop a semi-Lagrangian approach with non-local operators to solve the Navier-Stokes equations under large deformations. The coupling method is achieved through a simple yet efficient two-way fictitious point method, ensuring accurate thermal and mechanical coupling across moving interfaces as discontinuities evolve. This approach also facilitates fluid flow through openings between crack surfaces. The THM PD framework is validated through various multi-physics simulations, including natural and mixed convection, quenching processes, and the injection of cold water into hot dry rock. These examples demonstrate its robust capabilities for modeling complex thermal FSI problems with evolving discontinuities. This framework bridges the application gap of PD in solids and fluids, allowing us to solve multi-physics problems using a single solver.</p
Viral theft of light: A cyanophage protein dismantles cyanobacterial photosynthesis to accelerate infection
Auxiliary metabolic genes, acquired by cyanobacterial viruses (cyanophages) from their hosts, are thought to manipulate host metabolism during infection. A recent study by Nadel et al. performed in vivo experiments to reveal how cyanophages use a viral nblA gene to accelerate infection by degrading the photosynthetic machinery of marine cyanobacteria.</p
Large language models for automated data extraction and inverse design of copolymers with targeted glass transition temperatures
The design of copolymers with tailored thermal properties is crucial for advanced applications in aerospace, the automotive industry, and electronics. However, the development of data-driven models for this task is severely hampered by the scarcity of structured, high-quality datasets, as copolymer data remains buried in the unstructured text of scientific literature. Here, we demonstrate a large language model (LLM)-powered pipeline to overcome this data bottleneck and enable the inverse design of novel copolymers. We developed a method to accurately extract copolymer names and their thermal properties (glass transition temperature, Tg, and melting temperature, Tm) from 393 research articles, achieving a high extraction accuracy (F1 score = 0.7834) and constructing a curated dataset of 1195 entries. This LLM-generated dataset was used to train machine learning models that rival the performance of models built on manually curated data. Our forward prediction models achieved R2 values of 0.791 and 0.845 for Tg and Tm, respectively, on an independent manually labeled test set. Critically, we deployed an inverse model for de novo design, generating 2052 novel, valid copolymer structures targeting high Tg, which exhibited a 189.5 % increase in average predicted Tg. The thermal properties of eight selected AI-generated copolymers were validated with molecular dynamics simulations, confirming their predicted performance with a mean relative error of 4.17 %. This work establishes LLMs as a powerful and reliable tool for automating scientific data extraction, effectively bridging the data gap in materials science and paving the way for accelerated discovery of high-performance polymers.</p
DataWink: Reusing and Adapting SVG-based Visualization Examples with Large Multimodal Models
Creating aesthetically pleasing data visualizations remains challenging for users without design expertise or familiarity with visualization tools. To address this gap, we present DataWink, a system that enables users to create custom visualizations by adapting high-quality examples. Our approach combines large multimodal models (LMMs) to extract data encoding from existing SVG-based visualization examples, featuring an intermediate representation of visualizations that bridges primitive SVG and visualization programs. Users may express adaptation goals to a conversational agent and control the visual appearance through widgets generated on demand. With an interactive interface, users can modify both data mappings and visual design elements while maintaining the original visualization's aesthetic quality. To evaluate DataWink, we conduct a user study (N = 12) with replication and free-form exploration tasks. As a result, DataWink is recognized for its learnability and effectiveness in personalized authoring tasks. Our results demonstrate the potential of example-driven approaches for democratizing visualization creation.</p
Strain-Gradient-Driven Decoupling of Thermal Suppression from Anisotropy in β-Ga<sub>2</sub>O<sub>3</sub>
β-Ga2O3 is an emerging ultra‑wide‑bandgap semiconductor (UWBGS) for efficient, high‑frequency power electronics, solar‑blind/UV photodetectors, and wearable/flexible devices, with scalable manufacturing. However, its intrinsically low thermal conductivity (k)—compounded by ubiquitous, nonuniform strains introduced during fabrication and operation—creates a stringent thermal‑management bottleneck that degrades heat dissipation, reliability, and performance. Consequently, understanding how uniform and non-uniform strains affect thermal transport in this UWBG material is essential for thermal management design. Yet, strain gradients (η), pervasive in flexible devices and epitaxial nanostructures, remain a major blind spot in β-Ga2O3 thermal transport studies. By integrating the first-principles-based machine learning interatomic potential with Boltzmann transport equation, we establish that η unlocks a k suppression mechanism fundamentally more potent than uniform strain (ε): moderate uniaxial gradients (0.6%/nm) suppress k by 32–37% (27–30%) in thin films (nanowires), intensifying to 43.3% with biaxial gradients. This reduction far exceeds that from equivalent ε and surpasses benchmark materials like silicon and BAs. Notably, β-Ga2O3 exhibits a unique magnitude-anisotropy decoupling under η: whereas uniform ε(±3%) modifies thermal anisotropy ratios by ∼25%, η strongly suppresses the absolute k while leaving these ratios nearly unchanged. This pronounced suppression originates from gradient-induced symmetry breaking and enhanced mode coupling, which activate otherwise forbidden phonon-scattering channels and make gradient-driven scattering dominant below 6.25 THz. Unlike cubic crystals (e.g., Si) under bending—where a through-thickness strain gradient predominantly suppresses heat flow along the bending direction while leaving the orthogonal in-plane component comparatively intact, thereby tuning anisotropy—monoclinic β-Ga2O3 exhibits stronger cross-direction coupling due to its non-orthogonal crystallographic framework, so the same gradient suppresses multiple k components. Consequently, phonon lifetimes are reduced broadly across transport directions, while concurrent changes in phonon group velocities partially compensate, yielding an approximately invariant anisotropy ratio even as kcollapses. By contrast, κ-Ga2O3 shows a weaker k reduction and pronounced anisotropy suppression under η, confirming that the decoupling is unique in β-Ga2O3. These findings redefine non-uniform strain from a parasitic flaw into a powerful design tool for engineering thermal isolation and heat flux in next-generation flexible and high-power β-Ga2O3 electronics
Semantic Modulated Prompting for Few-Shot Audio-Visual Classification
Few-Shot Audio-Visual Classification (FS-AVC) trains models using a limited number of labeled audio and visual sample pairs to capture the classification capability. Deep learning-based audio-visual learning methods often construct complicated frameworks with numerous parameters trained on large labeled datasets, rendering them impractical for FS-AVC. The key challenges for FS-AVC are model overfitting, multimodal fusion under temporal asynchrony, and modality imbalance. To address these challenges, we propose a novel method called Semantic Modulated Prompting (SMP) to improve the learning process of FS-AVC. This framework implants text as prompting tokens via two components: Prompt-refined Audio-Visual efficient Learner (P-AVeL) and Prompt-tuned Prototypical Regularization (P-PR). By integrating semantic prompts, adapter-based P-AVeLs conduct the prompt-guided latent attention to alleviate the overfitting and achieve effective alignment and fusion. Concurrently, P-PR, the first rebalancing method designed for few-shot scenarios, uses these semantic prompts to accurately evaluate and dynamically adjust the imbalance of two modalities. Extensive experiments demonstrate that the SMP framework consistently outperforms state-of-the-art multimodal methods by a large margin.</p
Synesthesia of Machines (SoM)-Enhanced Sub-THz ISAC Transmission for Air–Ground Network
Integrated sensing and communication (ISAC) at sub-THz frequencies is crucial for future air-ground networks. However, optimizing ISAC performance while managing operational latency is challenging due to unique propagation characteristics and hardware limitations. This paper introduces a multi-modal sensing fusion framework inspired by synesthesia of machine (SoM) to enhance sub-THz ISAC transmission. By exploiting inherent degrees of freedom in sub-THz hardware and channels, the framework succeeds in tuning the radio-frequency environment. It features squint-aware beam management to improve air-ground network adaptability, enabling dynamic three-dimensional ISAC links. By leveraging multi-modal information, the framework enhances ISAC performance and reduces latency. Visual data is used to rapidly localize users and targets, while a customized multi-modal learning algorithm optimizes the hybrid precoder. A new metric is proposed for comprehensive performance evaluation. Extensive experiments demonstrate that the proposed scheme significantly improves ISAC efficiency.</p
Price of Non-discrimination in Public Combinatorial Contracts
This paper addresses the challenge of contract design within modern social platforms such as YouTube, Medium, and TikTok, where the primary goal is to develop monetization incentives for content creators. The inherent complexities of these platforms, influenced by factors such as regulation and scalability, often hinder the provision of personalized contracts, despite the diversity among content creators.We model this challenge as a multi-agent combinatorial contract design problem in which the principal (e.g., a digital platform) delegates an identical task (e.g., video production) to agents (e.g., content creators) and motivates them through contracts. Each agent’s payment depends solely on the success of his assigned task, and the principal gains utility upon each successful task. Consequently, an optimal contract may vary among agents due to their heterogeneity, while a public contract offers the same terms to all agents. The price of non-discrimination (PoN) captures the difference in the principal’s utility when applying non-discriminatory (public) contracts versus customized (optimal) contracts for individual content creators. We analyze how the price of non-discrimination interacts with the combinatorial structure of agents’ technologies, establishing tight and nearly tight bounds on the price of non-discrimination across various scenarios
Challenges and Opportunities in Perceiving Worker Intentions for Proactive Human–Robot Collaboration in Construction
Human–robot collaboration (HRC) has the potential to enhance safety and productivity in construction. However, much of the existing HRC development relies on passive or reactive paradigms, placing imbalanced cognitive workloads on humans and hindering smooth and efficient collaboration. Addressing these issues requires robots to be active in collaboration, perceiving worker intentions and responding to human needs to enhance team fluency and efficiency, i.e., proactive HRC. Existing HRC reviews have explored general robot programming methods and human–robot interfaces. However, there is a lack of a holistic understanding of worker intention, which is a critical prerequisite for facilitating proactive HRC. Therefore, this paper aims to provide a systematic review of existing worker intention perception studies by defining the concept, identifying key cues that convey worker intentions, and exploring various levels of intention prediction methods. Challenges and future directions are discussed, with the hope of offering valuable insights into proactive HRC in construction.<br/
MoE-INR: Implicit Neural Representation with Mixture-of-Experts for Time-Varying Volumetric Data Compression
Implicit neural representations (INRs) have emerged as a transformative paradigm for time-varying volumetric data compression and representation, owing to their ability to model high-dimensional signals effectively. INRs represent scalar fields based on sampled coordinates, typically using either a single network for the entire field or multiple networks across different spatial domains. However, these approaches often face challenges in modeling complex patterns and introducing boundary artifacts. To address these limitations, we propose MoE-INR,an INR architecture based on a mixture-of-experts (MoE) framework. MoE-INR automates irregular subdivisions of spatiotemporal fields and dynamically assigns them to different expert networks. The architecture comprises three key components: a policy network, a shared encoder, and multiple expert decoders. The policy network subdivides the field and determines which expert decoder is responsible for a given input coordinate. The shared encoder extracts hidden representations from the input coordinates, and the expert decoders transform these high-dimensional features into scalar values. This design results in a unified framework accommodating diverse INR types, including conventional, grid-based, and ensemble. We evaluate the effectiveness of MoE-INR on multiple time-varying datasets with varying characteristics. Experimental results demonstrate that MoE-INR significantly outperforms existing non-MoE and MoE-based INRs and traditional lossy compression methods across quantitative and qualitative metrics under various compression ratios.</p