GIST Scholar
Not a member yet
    30271 research outputs found

    Exploring the Potential of Music Generative AI for Music-Making by Deaf and Hard of Hearing People

    No full text
    Recent advancements in text-to-music generative AI (GenAI) have significantly expanded access to music creation. However, deaf and hard of hearing (DHH) individuals remain largely excluded from these developments. This study explores how music GenAI could enhance the music-making experience of DHH individuals, who often rely on hearing people to translate sounds and music. We developed a multimodal music-making assistive tool informed by focus group interviews. This tool enables DHH users to create and edit music independently through language interaction with music GenAI, supported by integrated visual and tactile feedback. Our findings from the music-making study revealed that the system empowers them to engage in independent and proactive music-making activities, increasing their confidence, fostering musical expression, and positively shifting their attitudes toward music. Contributing to inclusive art by preserving the unique sensory characteristics of DHH individuals, this study demonstrates how music GenAI can benefit a marginalized community, fostering independent creative expression. © 2025 Copyright held by the owner/author(s)

    FUN-SSL: Full-band layer followed by U-Net with Narrow-band layers for Multiple Moving Sound Source Localization

    No full text
    음원 정위는 다채널 신호로부터 하나 이상의 음원의 도착 방향(DoA, direction of arrival)을 추정하는 기술로, 이를 위해서는 신호의 직접 경로 성분에 내재된 공간적 특징을 정확히 포착하는 것이 중요하다. 그중에서도, 잡음과 잔향의 영향을 최소화한 직접 경로에서의 채널 간 위상 차이(DP-IPD, direct-path interchannel phase difference)는 주파수 영역에서 시간 지연 정보를 위상 차이로 표현해주는 주요 공간 단서(spatial cue)로, 음원 정위에 널리 활용된다. 본 연구에서는 DP-IPD를 추정하기 위해, 전대역(full-band)과 협대역(narrow-band) 정보를 융합한 최신 네트워크를 U-Net 구조와 결합한 딥러닝 기반 음원 정위 모델 FUN-SSL을 제안한다. FUN-SSL은 전대역 계층에서 주파수 간 상관관계를 학습하고, 다중 스케일 협대역 계층을 통해 여러 동적 음원의 시간 변화를 정밀하게 추적한다. 또한, 다운샘플링–업샘플링 모듈에서 추출한 특징을 각 스케일별로 합산하여 다양한 해상도에서의 공간 표현력을 강화하고, FUN 블록 간 스킵 연결(skip connection)을 통해 이전 블록의 정보를 다음 블록으로 효과적으로 전달함으로써 정보 손실을 완화하였다. 실제 환경을 모사한 시뮬레이션 데이터셋에서의 실험 결과, FUN-SSL은 기존의 최신 음원 정위 기법 대비 연산량을 절반 수준으로 줄이면서도 우수한 성능을 달성하였다. 특히, 잡음과 잔향이 심한 환경에서도 안정적인 정위 성능을 유지함으로써 제안한 모델의 효용성을 입증하였다.|Sound source localization (SSL) aims to estimate the direction of arrival (DoA) of one or more sound sources from multichannel signals. For accurate DoA estimation, it is essential to capture spatial features embedded in the direct-path component of the signal. Among these features, the direct-path interchannel phase difference (DP-IPD), which represents time delay information as phase differences in the frequency domain, has been widely adopted as a key spatial cue for localization tasks. This work proposes FUN-SSL, a novel SSL model that consists of a full-band processing layer followed by a U-Net framework with multi-scale narrow-band layers. The network learns inter-frequency correlations through full-band processing and accurately tracks temporal variations of multiple moving sources via the narrow-band processing. Encoder features are summed with the corresponding decoder feature maps at each scale to enhance the representational capacity. In addition, skip connections between FUN blocks preserve important information from previous blocks and propagate it to subsequent processing block. FUN-SSL outperforms the baseline on the simulation dataset while preserving a similar model size and significantly lower computational complexity. Furthermore, FUN-SSL consistently maintains reliable localization performance even under challenging conditions with strong noise and reverberation, demonstrating the effectiveness of its architecture design.Master1. Introduction 1 2. Baseline Model 5 2.1 Problem Formulation 5 2.2 Architecture 7 2.3 Training Objective Function 8 2.4 SSL Inference 10 3. Proposed Model 11 3.1 FUN Block 11 3.2 Convolutional Block 19 4. Experimental Setup 21 4.1 Dataset 21 4.2 Configurations 22 4.3 Comparison Methods 22 4.4 Evaluation Metrics 23 5. Evaluation Results and Analysis 25 5.1 Performance on Simulated Dataset 25 5.2 Ablation Studies 30 6. Conclusion 31 References 3

    Machine Learning-Driven Joint Structuring of WPT Coil and Core for Enhanced Mutual Inductance and Reduced Ferrite Volume

    No full text
    Machine learning (ML) algorithms have shown promise in optimizing wireless power transfer (WPT) systems, particularly for electric vehicle charging and medical implants. However, most approaches focus on isolated WPT components, limiting their overall impact. This study presents an integrated optimization of WPT coil and core designs to enhance mutual inductance between transmitting (Tx) and receiving (Rx) coils. We propose a hybrid residual layer sequential neural network (HRL-SeqNet), along with SVR-based multivariate regression and deep Qnetwork (DQN) models, to enhance the WPT performance. HRL-SeqNet incorporates GRU, BiLSTM, and LSTM RNNs followed by a second-level GRU and ensemble regression to optimize the copper windings, achieving a high mutual inductance of 11.284 μH with a conventional ferrite core. Additionally, SVR and DQN models optimize the ferrite core by introducing asymmetric and symmetric vacuum configurations. The SVR-optimized asymmetric core achieved a 0.585% increase in mutual inductance while reducing ferrite volume by 24.49%. In addition, the DQN model increased mutual inductance by 1.38% and 3.16% for order 4 and order 8 rotational symmetry, respectively, with corresponding reductions in ferrite volume of 23.47% and 25.51%. These ML models show exceptional computational efficiency, with HRL-SeqNet, SVR, and DQN achieving 4,300x, 16,960x, and 10,138x speedups, respectively, over ANSYS Maxwell simulations. Experimental validations confirm the effectiveness of these ML-optimized designs, demonstrating their applicability across various WPT applications. © 1986-2012 IEEE.FALSEsciescopu

    Crystallization-Driven Self Assembly of Thermoresponsive Conjugated Block Copolymers into Multifunctional Nanowire Networks

    No full text
    Thermo-responsive block copolymers enable precise structural control via self-assembly, which is vital for advanced material development. In this study, I synthesized conjugated thermo-responsive block copolymers designed to form aligned nanowire arrays through crystallization-driven self-assembly. To enhance mechanical strength and electrical continuity, ABA-type triblock copolymers were introduced as physical cross- linkers. Tailored design of conjugated and LCST segments allowed reversible nanowire rearrangement under temperature changes. Moreover, NPs-containing polymers were employed, and in situ liquid-cell TEM was used to directly visualize thermally induced nano structural transitions in real time, confirming morphological changes. This work presents a novel strategy to fabricate multifunctional nanowire networks integrating conductivity, thermo- responsiveness, and durability, with promising potential in stretchable electronics, thermal sensors, and smart nanodevices

    Automated diagnosis for extraction difficulty of maxillary and mandibular third molars and post-extraction complications using deep learning

    No full text
    Optimal surgical methods require accurate prediction of extraction difficulty and complications. Although various automated methods related to third molar (M3) extraction have been proposed, none fully predict both extraction difficulty and post-extraction complications. This study proposes an automatic diagnosis method based on state-of-the-art semantic segmentation and classification models to predict the extraction difficulty of maxillary and mandibular M3s and possible complications (sinus perforation and inferior alveolar nerve (IAN) injury). A dataset of 4,903 orthopantomographys (OPGs), annotated by experts, was used. The proposed diagnosis method segments M3s (#18, #28, #38, #48), second molars (#17, #27, #37, #47), maxillary sinuses, and inferior alveolar canal (IAC) in OPGs using a segmentation model and extracts the region of interest (RoI). Using the RoI as input, the classification model predicts extraction difficulty and complication possibilities. The model achieved 87.97% and 88.85% accuracy in predicting maxillary and mandibular M3 extraction difficulty, with area under the receiver operating characteristic curve (AUROC) of 96.25% and 97.3%, respectively. It also predicted the possibility of sinus perforation and IAN injury with 91.45% and 88.47% accuracy, and AUROC of 91.78% and 94.13%, respectively. Our results show that the proposed method effectively predicts the extraction difficulty and complications of maxillary and mandibular M3s using OPG, and could serve as a decision support system for clinicians before surgery. © The Author(s) 2025.TRUEsciescopu

    Bias Correction of Copernicus Atmosphere Monitoring Service (CAMS) Forecasts via Machine and Deep Learning-based Multi-target Regression

    No full text
    The significant impact of air pollution on human health and climate change underscores the critical importance of atmospheric chemistry transport modeling. However, the accuracy of the Copernicus Atmosphere Monitoring Service (CAMS) forecasts remains a substantial challenge, particularly evidenced by a pronounced overestimation of SO2 concentrations in the Republic of Korea. While O3 is relatively well-simulated and CAMS undergoes continuous model updates, there is still a clear necessity for accuracy improvements for PM2.5, PM10, NO2, and CO, similar to SO2. Historically, artificial intelligence algorithms have predominantly focused on single-target prediction problems. Nevertheless, recent research has increasingly explored multi-target regression (MTR), an approach that simultaneously considers multiple variables, leading to enhanced generalization performance. This study, therefore, aimed to perform bias correction for CAMS forecasts of PM2.5, PM10, O3, NO2, SO2, and CO in the Korean region using AI-based MTR algorithms. CAMS data from 2016 to 2018 served as the training dataset, with AirKorea ground observations utilized as the target data for model validation. The models were evaluated using 2019 data, assessing performance through metrics including Index of Agreement (IOA), Pearson Correlation Coefficient (R), Root Mean Square Error (RMSE), and Mean Bias (MB). Results demonstrated that both XGBoost and LSTM-based MTR models significantly enhanced forecasting performance across all pollutants compared to CAMS data. XGBoost generally outperformed LSTM in terms of IOA and R across most pollutants, showing improvements of up to 408% and 76% for IOAs of SO2 and PM2.5, respectively compared to CAMS. Both models exhibited substantial reductions in RMSE and MB. While CAMS itself performed better for O3, the MTR models effectively corrected the overestimation of SO2, CO, NO2, and particulate matter by CAMS.Master1. Introduction 1 2. Methodology and Materials 4 2.1. Study area and datasets 4 2.1.1. AirKorea surface network of monitoring sites dataset 6 2.1.2. CAMS global atmospheric composition forecasts dataset 7 2.1.3. Auxiliary dataset 7 2.2. Machine and deep learning algorithms 8 2.2.1. XGBoost 8 2.2.2. LSTM 9 2.3. Model development 12 2.3.1. Data preprocessing 12 2.3.2. Developing for XGBoost 12 2.3.3. Developing for LSTM 13 2.4. Model evaluation 15 2.4.1. Testing 15 2.4.2. Permutation feature importance test 15 2.4.3. Statistics for analysis 16 3. Results and Discussions 17 3.1. Model evaluation 17 3.1.1. Testing 17 3.2. Analysis 20 3.2.1. CAMS model update 20 3.2.2. Permutation feature importance for XGBoost 23 3.2.3. Permutation feature importance for LSTM 26 3.2.4. High PM episode 31 4. Summary and Conclusions 34 5. References 36 6. Supplements 4

    Hydrophobic Deep Eutectic Solvent (DES) Design Enables Optimally Hydrated DES-in-Water Electrolytes for High-Performance Bromine Redox-Enhanced Energy Storage Systems

    No full text
    Supercapacitors are renowned for rapid charging, high power density, and long lifespan, yet their practical applications are limited by low energy densities. Redox-enhanced electrochemical capacitors (redox ECs) address this limitation by incorporating redox-active electrolytes, enabling Faradaic charge storage. Bromide is a promising catholyte due to its high reduction potential, excellent solubility, and low cost. However, the generation of corrosive Br2 and the cross-diffusion of soluble polybromides result in suboptimal cell efficiency including severe self-discharge and reduced cycle life. Although solid complexing agents have been used to suppress polybromides' cross-diffusion, this approach necessitates water, which inherently limits electrochemical and thermal stability. Here, a hydrated deep eutectic solvent (HDES) electrolyte is developed by combining tetrabutylammonium bromide (TBAB) with ethylene glycol. This HDES system effectively utilizes the multifunctional roles of TBAB: the bromide anion functions as a catholyte, while the TBA cation suppresses polybromides' cross-diffusion as a built-in solid complexing agent. Critically, unlike previous studies that focus on minimally hydrated DESs, this system leverages the hydrophobic effect of TBAB to accommodate higher water content, addressing challenges inherent to DESs while maintaining superior electrochemical and thermal stability. The optimized HDES-50 electrolyte, containing 50 wt.% water, provides a robust and efficient solution for advanced redox ECs. © 2025 The Author(s). Advanced Functional Materials published by Wiley-VCH GmbH.TRUEsciescopu

    Monitoring the Dynamic Micellar Self-Assembly in an Aqueous Medium

    No full text
    Over the past two decades, amphiphilic rigid-flexible molecules have garnered significant attention due to their ability to self-assemble into well-defined nanostructures through non-covalent interactions in aqueous media. These precisely controlled architectures have demonstrated promising applications across various fields, including biomedicine, catalysis, sensing, and electro-optical devices. Precise control and understanding of dynamic assembly processes, encompassing molecular association and dissociation, are crucial for producing functional particles with targeted structural properties. This necessitates detailed in-situ monitoring of native morphologies in solution. In this study, liquid-phase transmission electron microscopy (LP-TEM) enabled quantitative analysis of self-assembly through high-precision tracking of micelle trajectories. Interparticle forces were experimentally quantified by measuring the velocity and diffusion coefficients at the single-particle level

    Causal-Paced Deep Reinforcement Learning

    No full text
    Designing effective task sequences is crucial for curriculum reinforcement learning (CRL), where agents must gradually acquire skills by training on intermediate tasks. A key challenge in CRL is to identify tasks that promote exploration, yet are similar enough to support effective transfer. While recent approach suggests comparing tasks via their Structural Causal Models (SCMs), the method requires access to ground-truth causal structures, an unrealistic assumption in most RL settings. In this work, we propose Causal-Paced Deep Reinforcement Learning (CP-DRL), a curriculum learning framework aware of SCM differences between tasks based on interaction data approximation. This signal captures task novelty, which we combine with the agent’s learnability, measured by reward gain, to form a unified objective. Empirically, CP-DRL outperforms existing curriculum methods on the Point Mass benchmark, achieving faster convergence and higher returns. CP-DRL demonstrates reduced variance with comparable final returns in the Bipedal Walker-Trivial setting, and achieves the highest average performance in the Infeasible variant. These results indicate that leveraging causal relationships between tasks can improve the structure-awareness and sample efficiency of curriculum reinforcement learning. We provide the full implementation of CP-DRL to facilitate the reproduction of our main results at \url{https://github.com/Cho-Geonwoo/CP-DRL}

    706

    full texts

    30,271

    metadata records
    Updated in last 30 days.
    GIST Scholar
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇