163559 research outputs found
Sort by
당뇨망막병증 환자에서 시각장애의 위험요인 규명에 대한 연구
학위논문(박사) -- 서울대학교 대학원 : 의과대학 의학과, 2024. 8. 박수경.배경: 당뇨병성 망막병증은 당뇨병의 3대 주요 미세혈관 합병증 중 하나로, 근로 연령 인구에서 시각 장애의 주요 원인이므로 당뇨병으로 인한 가장 중요한 안구 합병증으로 간주된다. 최근 전 세계적으로 당뇨의 유병률이 점점 높아지고 있으며, 당뇨망막병증을 포함한 관련 결과도 증가하고 있으며 이에 따라 당뇨망막병증의 발생률도 증가하고 있다. 당뇨망막병증과 관련된 시각장애는 전 세계적으로 큰 우려를 낳고 있는 공중보건학적 문제이므로 적극적으로 관리해야 한다. 따라서 이에 대해 시각장애에 영향을 미치는 역학적 요인을 폭넓게 고려하고 효과적인 통계 모델 피팅을 통해 얻은 추정 및 예측을 기반으로 의사 결정을 내릴 수 있는 양질의 결과가 필요하다. 당뇨망막병증은 광범위하게 연구되어 왔지만 환경적, 유전적 차이로 인해 전 세계 인종 간에 상당한 차이가 있을 수 있다. 따라서 국내에서 당뇨망막병증 관련 시각장애의 추세를 조사하고 시각장애의 위험도를 예측하면 당뇨망막병증의 원인에 관한 단서를 얻을 수 있으며 당뇨망막병증 관련 시각장애를 이해하면 향후 보건 정책에 중요한 정보를 제공할 수 있을 것이다.
해결되지 않은 문제: 현재까지 보고된 연구 결과에 의하면, 당뇨망막병증 환자에서 발생하는 시각장애에 대한 역학적 현황과 이에 대한 설명이 부족하고 잘 알려져 있지 않으며, 전 세계적으로도 당뇨망막병증 환자에서 시각장애의 종단적 추세나 그 의미를 보고한 연구는 거의 없다. 또한 이전 연구들은 당뇨망막병증 발병 위험에 대해서만 보고하고 있는데, 모든 당뇨망막병증이 시각장애 또는 실명으로 이어지는 것은 아니기 때문에 당뇨망막병증을 진단 받은 후 발생하는 시각장애에 더 초점을 맞추고 보다 포괄적인 규모로 그 위험을 추정하는 연구가 필요하다. 더 나아가 일반적으로 당뇨에서 시각장애가 발생하는 것은 당뇨에서 당뇨망막병증이 발생하고 당뇨망막병증이 악화되어 시각장애가 발생하기 때문이라고 알려져 있는데, 시각장애를 일으키는 매개요인으로서의 당뇨망막병증의 역할과 직접적 안과적 위험요인으로서의 당뇨망막병증의 역할에 대해 고찰이 필요하다. 또한 이와 함께 당뇨에서 시각장애가 발생할 수 있는 가능한 병태생리학적 메커니즘을 고려할 때, 당뇨망막병증을 거치지 않고 당뇨가 존재하는 것 자체가 근본적으로 시각장애의 위험을 증가시킬 수도 있다. 그러나 이 인과성에 대해 명확히 밝혀진 바가 없으므로 이러한 인과관계를 종합적인 관점에서 재평가할 필요가 있다.
목적 : 1) 본 연구의 목적은 탐색적 연구로서 연령, 기간, 코호트에 따른 시각장애의 발생률의 변화를 파악하고 당뇨망막병증 환자에서 시각장애와 관련된 잠재적 위험 요인에 대한 가설을 세우고자 한다.
2) 탐색적 연구 결과를 바탕으로 다음 분석 연구에서는 당뇨망막병증을 중심으로 시각장애의 발생 위험과 관련된 다양한 안과적 및 비안과적 인자를 규명하고자 한다.
3) 범위를 확장하여 시각장애에 직접적인 영향을 미치는 것으로 알려진 안과 질환(당뇨, 당뇨망막병증, 백내장, 연령관련 황반변성, 녹내장)이 시각장애를 유발하는 인과관계를 최종적으로 규명하고자 한다.
연구 방법: 각 연구의 방법은 다음과 같다. 1) 2005년부터 2019년까지 한국 국민건강보험 청구 데이터베이스에서 건강검진을 시행한 20세 이상의 당뇨망막병증 환자 총 250만 명을 추출하여, 비증식성 당뇨망막병증 및 증식성당뇨망막병증 코호트를 별개로 구축하였다. 시각장애는 건강검진 결과에서 적어도 한 쪽 눈에 시각장애 코드 (9.9)가 있는 것으로 정의하였다. 각 코호트에서 실명 발생률을 계산한 후 로그-선형 푸아송 연령-기간-코호트 분석 모델을 사용하여 각 연구 그룹에 대해 시각장애에 대한 각 효과를 추정하였다.
2) 국민건강보험공단 국민건강정보 데이터베이스의 20세 이상 대한민국 성인 인구 전체에서 총 250만 명의 당뇨망막병증 환자를 대상으로 하였다. 당뇨망막병증 환자 중 건강검진에서 한쪽 눈에 시각장애 코드(9.9)가 있는 것으로 확인된 성인 참가자를 대상으로 했으며, 약 50만명의 하위 코호트 참여자를 포함한 사례-코호트 데이터를 구축하였다. 총 22,888명의 시각장애 사례가 추출되었고, 이 중 4,520례는 하위 코호트 내, 18,368례는 하위 코호트 밖의 사례이다. 가중 콕스 비례 위험 회귀 모델을 사용하여 연령(20-39세, 40-64세, 65세 이상), 성별, 소득, 흡연, 음주, 신체 활동 경험, 체질량 지수, 고혈압, 당뇨 조절, 헤모글로빈, 총 콜레스테롤 수치, Charlson 동반질환지수, 안과적 치료 옵션(레이저, 주사, 수술) 등의 혈액 검사 및 문진 데이터를 포함한 다양한 전신 요인이 시각장애에 얼만큼의 위험도를 보이는지 분석하였다.
3) 유럽 인구에 대한 GWAS 연구의 요약 통계를 기반으로 당뇨, 당뇨망막병증, 시각장애 관련 주요 안과질환 (녹내장, 연령관련 황반변성, 백내장) 및 시각장애 위험 간의 연관성을 분석하기 위해 two-step 멘델식 무작위분석 방법을 사용했다. 각 노출, 매개요인, 결과 변수간 멘델식 무작위분석 방법을 사용하여 해당 주요질환과 시각장애 사이의 인과 관계에 대한 직접적 효과와 간접적 효과를 확인하였고, 매개 효과에 대한 비율을 추정했다. 더불어 당뇨가 시각장애에 미치는 영향이 당뇨망막병증에 의해 얼만큼 매개되는지 그 정도를 분석하였으며 당뇨망막병증에 매개되지 않고 당뇨에서 시각장애에 직접적으로 이르는 인과성이 있는지 함께 평가하였다.
결과: 1) APC 분석을 사용한 역학 연구에서 시각장애 발생률은 비증식성당뇨망막병증 군에서 10만 명당 1326.62명, 증식성당뇨망막병증 군에서 3397.57명이었다. 시각장애 발생률은 2011년 이후 급격히 감소하여 비증식성당뇨망막병증 군과 증식성당뇨망막병증 군에서 각각 연간 5.6%와 4.4% 감소했다. 1920년에서 1930년 사이에 태어난 사람들의 시각장애 위험이 가장 높았으며, 그 이후에는 시각장애의 위험이 급격히 감소했다. 또한 1980년 이후 출생자의 경우 남녀 모두에서 위험이 증가하기 시작했다. APC 모델 중 연령, 기간, 코호트 효과의 조합 모델이 가장 높은 설명력(0.96)을 보였다.
2) 사례-코호트 연구에서 당뇨망막병증에서 고령 (65세 이상), 낮은 체질량지수 (<18.5), 현재 흡연 및 실명 위험 간의 연관성은 유의하였다 [각각 HR (95% CI): 3.47 (3.12-3.86), HR 1.34 (1.20-1.50), and HR 1.06 (1.01-1.12)]. 당뇨망막병증에서 빈혈과 시각장애 위험 사이의 연관성은 통계적으로 유의하였다. [경증 빈혈의 경우 HR 1.12 (1.06-1.17), 중등도 빈혈의 경우 HR 1.29(1.19-1.40), 중증 빈혈의 경우 HR 1.74(1.20-2.53)]. 당뇨망막병증의 단계 구분시, 실명위험 단계의 당뇨망막병증은 비증식성 단계의 당뇨망막병증에 비해 시각장애 위험이 유의하게 높았다 [HR 3.19(3.08-3.30)].
3) 이 two-step 멘델식 무작위분석 연구에서는 당뇨와 시각장애 위험(오즈비 1.21, p= 1.54E-04), 당뇨망막병증과 시각장애 위험(오즈비 1.15, p= 2.07E-10), 당뇨와 당뇨망막병증 위험(오즈비 1.74, p= 1.57E-08) 간에 유의한 인과성이 있는 것으로 나타났다. 다변량 멘델식 무작위분석 결과를 이용한 매개 분석 결과에 따르면, 시각장애 위험에 대한 당뇨 효과의 88%(95% CI, 85-91%)가 당뇨망막병증을 통해 매개되었다. 이 외 12%(95% CI, 9-15%)는 당뇨망막병증을 통해 매개되지 않고, 시각장애 위험에 대한 당뇨의 직접적인 영향이었다.
결론: 본 대규모 장기 역학 연구의 결과에 따르면, 당뇨망막병증 환자에서 시각장애의 추세는 단일 역학적 원인에 의한 것이 아니라 생물학적 연령, 사회적 결정 요인 및 의료 정책의 조합을 포함한 다양한 역학 요인의 영향을 받았다. 특히 최근 20대와 30대의 실명 위험 증가는 앞으로 더욱 증가할 수 있으므로 경시되어서 안되며, 젊은 당뇨 환자의 경우 시각장애의 가능성에 대한 조기 환자 교육과 적절한 관리가 필요하다. 또한 당뇨망막병증에서의 시각장애는 안과적 요인뿐만 아니라 다양한 전신 질환의 영향을 받는 다인성 질환이라는 사실도 밝혀졌다. 본 연구에서 도출된 시각장애의 다양한 위험 요인에 대한 결과는 경미한 조기 당뇨망막병증일지라도 전신적 위험 요인을 가지는 경우 추후 시각장애가 발생할 수 있는 확률이 높을 수 있다는 것을 의미하고, 이를 예방하고 치료 표적을 마련하며 실명 진행을 예방하는데 필수적인 임상적 정보를 제공해 줄 수 있다. 마지막으로, 유전체 연구 결과를 통해 당뇨망막병증이 시각장애의 주요 원인 요인 중 하나이자 매개체로서 부각될 수 있음을 입증했다. 더불어 당뇨망막병증 없이도 당뇨 자체가 시각장애의 또 다른 원인 인자로 간주될 수 있으며 시각장애와 직접적인 인과관계를 가질 수 있음을 발견했다. 따라서 이는 당뇨 환자의 시각장애 위험을 조기 단계부터 주의 깊게 관리할 필요가 있음을 함께 시사한다.
중심어 : 역학, 시각장애, 당뇨, 당뇨망막병증, 인과성, 위험
학번 : 2020-37397Background: Diabetic retinopathy (DR) is considered one of the three main microvascular complications of diabetes mellitus (DM), Since DR is a leading cause of visual impairment (VI) in the working-age population, it is considered the most critical ocular complication caused by DM. Recently, DM has been becoming increasingly prevalent worldwide, and so are its associated outcomes, including DR. Consequently, the incidence of DR is increasing as well. DR-associated VI is a health problem of great concern worldwide and must be actively managed. Achieving high- quality results requires a broad consideration of the epidemiological factors that affect VI, with decision-making based on estimates and predictions obtained through effective statistical model fitting. DR has been extensively studied but there may be significant differences between races worldwide due to environmental and genetic differences. Investigations of the trend in DR-related VI and risk prediction of VI can provide clues regarding the aetiology of DR. Understanding DR-related VIs can also provide important information for future health policy.
Unsolved issues: There is a lack of epidemiological description of VIs in DR patients, and the true picture of the epidemiology is not well understood. In worldwide, few studies have reported the longitudinal trend of VI in patients with DR or the implications thereof. Furthermore, previous studies only report on the risk of developing DR. Since not all DR leads to VI or blindness, there is a need for studies that focus more on VI after a DR diagnosis and estimate its risk on a more comprehensive scale. In addition, it is generally believed that when VI occurs in DM, it is because DM develops DR and the DR worsens to become VI. However, when considering the mechanisms by which VI can occur in DM, the very existence of DM without going through DR can fundamentally increase the risk of VI. Therefore, it is necessary to re-evaluate the causality of DM and VI from a comprehensive perspective.
Purpose: 1) To identify changes in the incidence of VI due to age, period, and cohort and to establish hypotheses about potential risk factors associated with VI in patients with DR as an explorative study.
2) Based on the results of an explorative study, in the next analystic study, we aimed to identify various ocular and non-ocular factors related to VI risk focusing on DR.
3)Expanding the scope, we finally established the causal role of ocular diseases known to directly affect VI (DM, DR, cataract, age- related macular degeneration, glaucoma) in causing VI.
Methods: Our approach includes: 1) A total of 2.5 million DR patients aged 20 years or older with health checkup data were selected from the Korean National Health Claims database from 2005 to 2019. Among those DR patients, non-proliferative DR/ proliferative DR (NPDR/PDR) groups were constructed separately. Among study groups, patients were identified as having a VI in at least one eye. The cumulative incidence of VI were calculated. Using a log-linear Poisson age-period-cohort (APC) analysis model, each effect on VI were estimated for each study group.
2) This is a retrospective case-cohort study. Among the entire DR patients from the entire adult South Korean population aged 20 years or older from the National Health Information Database of the National Health Insurance Service, approximatively 0.5 million patients with DR were included in the sub-cohort. The main outcome was identified as having a VI code (9.9) in at least one eye during a health checkup. The entire VI cases related to DR was extracted (n=22,888) including 4,520 cases within the sub-cohort and 18,368 cases outside the sub-cohort. Possible risk factors or confounders included age (20-39, 40-64, or 65 and over), sex, income, smoking, alcohol consumption, experience of physical activity, body mass index (BMI), laboratory data including systolic/diastolic hypertension, diabetic controls, hemoglobin, total cholesterol level, glomerular filtration rate (GFR), Charlson Comorbidity Index (CCI), and ophthalmologic treatment (pan- retinal laser photocoagulation, intraocular injection, vitrectomy surgery). Weighted Cox proportional hazard regression analysis was performed to estimate the association between those factors and VI risk.
3) Mediation analysis based on two step MR methods were employed to analyze the association between DM, DR, glaucoma, cataract, AMD and VI risk, based on the summary statistics of genome-wide association studies in the European population from UK biobank data. Indirect or direct effect for causal relationship between those major diseases and VI was determined. In addition, proportion for the mediated effects was estimated.
Results: 1) In epidemiologic study using APC analysis, the incidence of VI was 1326.62 per 100,000 in the NPDR group and 3397.57 in the PDR group. The VI rate sharply decreased after 2011, with annual decreases of 5.6% and 4.4% in the NPDR and the PDR groups, respectively. People born between 1920 and 1930 had the highest overall risk of VI, with the risk decreasing rapidly after that. For those born after 1980, the risk started to increase in both sexes. Among the APC models, the combination model of age, period, and cohort effects showed the highest explanatory power (0.96).
2) In our case-cohort study, the association between older age (65 ≤ age), lower BMI (<18.5), current smoking and VI risk in the entire DR was significant [HR (95% CI): 3.47 (3.12-3.86), HR 1.34 (1.20-1.50), and HR 1.06 (1.01-1.12), respectively]. The association between all anemic conditions and VI risk in the entire DR was significant [HR 1.12 (1.06-1.17) for mild anemia, HR 1.29 (1.19-1.40) for moderate anemia, and HR 1.74 (1.20-2.53) for severe anemia]. In DR grade, VTDR had a significantly high risk of VI compared to NPDR [HR 3.19 (3.08-3.30)].
3) In this two-step MR study, there showed a causal association between DM and VI risk (odds ratio 1,21 p= 1.54E-07), between DR and VI risk (odds ratio 1.15, p=2.07E-10), and between DM and DR risk (odds ratio 1.74, p= 1.57E-08). In mediation analysis, MR mediation analysis estimated that 88% (95% CIs, 85-91%) of the DM effect on VI risk was mediated via DR. On the other hand, 12% (95% CIs, 9-15%) of the direct DM effect on VI risk was not mediated via DR.
Conclusion: In this nationwide long-term epidemiology study, VI in DR was not due to a single epidemiologic cause but rather a combination of biological age, social determinants, and healthcare policies. The increased risk of VI in individuals in their 20s and 30s may even increase in the future and should not be ignored. Therefore, vigilance of younger patients are recommended. We also found that DR is a multifactorial disease affected by various systemic conditions as well as ophthalmologic factors. The result about different risk factors for VI is essential to prevent and prepare therapeutic targets and limit the progression of VI in DR. At last, we demonstrated that this evidence may highlight the DR as one of the main causal factors of VI as well as a mediator. In addition, we found that DM can be considered as another causal factor for VI and have a direct causality of VI without DR. Thus, we can also suggest a need for early careful management of VI risk in patients with DM.
Keywords : Epidemiology, visual impairment, diabetes, diabetic retinopathy, causality, risk
Student Number : 2020-37397Chapter 1. Introduction
1.1. DR as a significant complication of diabetes 17
1.2. Incidence of DR worldwide 17
1.3. Non-proliferative DR and proliferative DR 18
1.4. Known risk factors of progression of DR 19
1.5. Basic treatment of DR - 20
1.6. Possible process of VI in DR 22
1.7. Global prevalence and burden for VI in DR 22
1.8. Limitation of previous studies and imperativeness of study 23
Chapter 2. Purpose and hypothesis
2.1. Objective 26
2.1.1. Age, period, and cohort effects on new VI occurrence among a
national DR cohort population: Explorative study for formulating hypothesis 26
2.1.2. Identification of risk factors for VI focusing on DR 26
2.1.3. Identification of causality of potential risk factors for VI : MR and mediation analysis 27
2.2. Description or hypothesis 27
Chapter 3. Methods & Material
Part 1. Epidemiology of VI in DR
3.1. Age, period, and cohort effects on new VI occurrence among a
national DR cohort population: Explorative study for formulating hypothesis 30
3.1.1. Data source 30
3.1.2. Selection of study population 31`
3.1.3. Outcome variable: VI 32
3.1.4. APC model 33
3.1.4.1. APC model 33
3.1.4.2. Estimation of Cumulative incidence 34
3.1.4.3. APC statistics 35
Part 2. VI and systemic risk factors
3.2. Identification of risk factors for VI focusing on DR - 37
3.2.1. Data source and study design 37
3.2.2. Case-cohort study population 38
3.2.3. Outcome variable 40
3.2.4. Potential risk factors and confounding variables 40
3.2.5. Statistical analysis - 41
Part 3. VI and systemic diabetes
3.3. Identification of causality of potential risk factors for VI: MR and mediation analysis 42
3.3.1. Overall study design 42
3.3.1.1. Overall design 42
3.3.1.2 Mendelian randomization 44
3.3.1.3. Mediation analysis 45
3.3.2. Data source 47
3.3.3. Study population 48
3.3.4. Statistic analysis 53
3.3.4.1. MR 53
3.3.4.2. Mediation analysis 55
Chapter 4. Results
Part 1. Epidemiology of VI in DR
4.1. Age, period, and cohort effects on new VI occurrence among a
national DR cohort population: Explorative study for formulating hypothesis 56
4.1.1. The 15-year cumulative incidence of VI 56
4.1.2. Epidemiologic multi-angle trends of VI 60
4.1.3. Relative risk of APC effects on VI 62
Part 2. VI and systemic risk factors
4.2. Identification of risk factors for VI focusing on DR 66
4.2.1. Demographics 66
4.2.2. Factors associated with risk of VI in the entire DR group 66
4.2.3. Factors associated with risk of VI in the NPDR group 71
4.2.4. Factors associated with risk of VI in the VTDR group 74
4.2.5. Combination analysis for type of DR and risk factors of VI 77
4.2.6 Stratification analysis of anemia for VI 79
Part 3. VI and systemic diabetes
4.3. Identification of causality of potential risk factors for VI: MR and mediation analysis 83
4.3.1. Identification 83
4.3.2. Mediation analysis 100
Chapter 5. Discussion
5.1. Age, period, and cohort effects on new VI occurrence among a
national DR cohort population: Explorative study for formulating hypothesis 104
5.2. Identification of risk factors for VI focusing on DR 111
5.3. Identification of causality of potential risk factors for VI: MR and mediation analysis 116
Chapter 6. Conclusion
6.1. Age, period, and cohort effects on new VI occurrence among a
national DR cohort population: Explorative study for formulating hypothesis 121`
6.2. Identification of risk factors for VI focusing on DR 121
6.3. Identification of causality of potential risk factors for VI: MR and mediation analysis 122
요약 (국문 초록) 123
References 129
List of tables and figures
[ Tables ]
Table 1. Classification of the exclusive diseases for VI based on KCD diagnostic code
from ICD-10. 33
Table 2. Total incident VI cases and estimated cumulative incidence in each DR cohort
population during 2005-2019. 58
Table 3. Total incident VI cases and estimated cumulative incidence
in the whole DR population compared with those in the entire
population during 2005-2019. 59
Table 4. Summary statistics of various APC models for VI in participants with NPDR or PDR. 65
Table 5. Demographics of study groups among the entire Korean population
during 2005-2019 for case-cohort analysis. 67
Table 6. Risk of VI in patients with the whole DR through weighted cox regression
in case-cohort design. 69
Table 7. Risk of VI in patients with simple NPDR through weighted cox regression. 72
Table 8. Risk of VI in patients with VTDR through weighted cox regression
in case-cohort study design. 76
Table 9. Combination analysis of type of DR and significant factors for VI. 78
Table 10. Stratification analysis of anemia for VI according to the age groups. 80
Table 11. Stratification analysis of anemia for VI according to the income groups. 81
Table 12. Summary statistics of Single nucleotide polymorphisms (SNPs) associated with
DM and those of SNPs for VI from the results of GWAS. 84
Table 13. Summary statistics of SNPs associated with DR and those of SNPs for VI
from the results of GWAS. 87
Table 14. Summary statistics of SNPs associated with DM and those of SNPs for DR
from the results of GWAS. 88
Table 15. Summary statistics of SNPs associated with DM and those of SNPs for cataract
from the results of GWAS. 89
Table 16. Summary statistics of SNPs associated with cataract and those of SNPs for VI
from the results of GWAS. 92
Table 17. Summary statistics of SNPs associated with AMD and those of SNPs for VI
from the results of GWAS. 93
Table 18. Summary statistics of SNPs associated with glaucoma and those of SNPs for VI
from the results of GWAS. 93
Table 19. Summary statistics of SNPs associated with DM and those of SNPs for glaucoma
from the results of GWAS. 94
Table 20. Summary statistics of SNPs associated with DM and those of SNPs for AMD
from the results of GWAS. 97
[ Figures ]
Figure 1. Fundus examination and optical coherent tomography of PDR. 19
Figure 2. Scheme of overall flow and thesis of the studies. 29
Figure 3. Identifying process of the study group. 32
Figure 4. Scheme of establishment of case-cohort study design. 39
Figure 5 . Outline of overall study desgin of current mediation analysis. 44
Figure 6. Assumption of Mendelian Randomization analysis. 45
Figure 7. Workflow for the selection of instrumental variables (IVs) for each factor based on
genome-wide association studies summary statistics from the UK biobank. 53
Figure 8. Cumulative incidence of VI in DR patients according to age compared with
the incidence of VI in the general population in South Korea. 57
Figure 9. (Upper) Age-specific incidence rates of VI in the whole DR, NPDR, and PDR groups
over time. (Middle) Period-specific incidence rates of VI. (Lower) Cohort-specific
incidence rates of VI. (Lower) Cohort-specific incidence rates of VI in each group
by year of birth. 61
Figure 10. APC effects on of bilateral VI cases in the whole DR, NPDR and PDR groups. 62
Figure 11. (Upper) Relative risk of age effects on VI in males and females in the whole DR, NPDR
and PDR groups. (Middle) Relative risk of period effects on VI in males and females.
(Lower) Relative risk of cohort effects on VI in males and females. 64
Figure 12. Estimate of the association of genetically predicted DM with each genetically deter
mined potential mediator. IVW or Wald ratio was applied as main analysis. 102
Figure 13. Estimate of the association of each genetically predicted potential mediator with
genetically determined VI. IVW or Wald ratio was applied as main analysis. 102
Figure 14. Direct and indirect effects between DM and VI through possible mediators. 103박
화재 재난 시나리오의 몰입감 향상을 위한 가상환경 내 화재구현
학위논문(석사) -- 서울대학교 대학원 : 공과대학 건축학과, 2024. 8. Changbum Ryan Ahn.This study investigated participants' perceptions of fire hazards in a VR crossroads. The objective was to measure how visual elements like dynamic flame movement, color changes, intensity, and smoke density influenced evacuation decisions. Participants experienced various fire compositions, and their evacuation decision times were recorded and analyzed.
Results showed that the smoke-emphasized Composition A (25% flames, 75% smoke) was perceived as the most hazardous and was chosen the least. Conversely, the flame-emphasized Composition C (75% flames, 25% smoke) was chosen the most, with Compositions B (50% flames, 50% smoke) and C being selected equally, indicating no significant difference in perceived hazard between them. Comparing fire implementation methods revealed that Composition D (realistic data) was chosen less frequently than Composition E (interaction-emphasized), suggesting the importance of using realistic fire data in VR fire implementation. A t-test confirmed that Composition D was perceived as more hazardous than Composition E (p < 0.05).
Irregular EDA data collection intervals posed challenges for valid SCR data acquisition. Despite this, decision-making times provided validation for the experiment. Decision times decreased with repeated exposure to the same compositions, indicating adaptation. New fire compositions initially increased evacuation decision times, which then decreased with repeated exposure. Additionally, consistent with previous research, evacuation decision times were shorter for the perceived more hazardous composition (Composition A).
These findings emphasize the role of visual elements in evacuation behavior and highlight the need to consider both adaptation and perceived hazard in VR fire evacuation scenarios. The results have applications in fire safety training, evacuation planning, and emergency response simulation. Future research should address data collection challenges, integrate additional physiological measures, and explore different fire scenarios to enhance understanding of responses to fire hazards. This will contribute to more reliable research on fire disasters in virtual education and experimental studies.본 연구는 VR 교차로에서 화재 위험에 대한 참가자들의 인식을 조사했음. 연구의 목적은 동적인 불꽃 움직임, 색상 변화, 강도 및 연기 밀도와 같은 시각적 요소들이 대피 결정에 어떻게 영향을 미치는지 측정하는 것으로 참가자들은 다양한 화재 조합을 경험했고, 대피 결정 시간이 기록 및 분석되었음.
결과는 연기 강조 조합 A (25% 불꽃, 75% 연기)가 가장 위험하다고 인식되었으며, 가장 적게 선택되었음을 보여줌. 반면 불꽃 강조 조합 C (75% 불꽃, 25% 연기)는 가장 많이 선택되었고, 조합 B (50% 불꽃, 50% 연기)와 C 는 동일하게 선택되어 이들 간의 위험 인식 차이가 없음을 나타냄.
화재 구현 방법을 비교한 결과, 조합 D (사실적 데이터)는 조합 E (상호작용 강조)보다 덜 선택되었으며, 이는 VR 화재 구현에서, 현실적인 화재 데이터 사용의 중요성을 시사함. t-검정은 조합 D가 조합 E보다 더 위험하다고 인식되었음을 확인했음 (p < 0.05).
불규칙한 EDA 데이터 수집 간격이 포함되어 SCR 데이터 유효 취득에 어려움이 있었음. 그러나, 의사 결정 시간은 실험에 대한 검증 요소를 제공했음. 동일한 조합에 반복적으로 노출될수록 시간이 감소하여 적응함을 보였음. 새로운 화재 조합은 처음에는 대피 결정 시간을 증가시키지만 반복적인 노출로 감소하는 모습을 보임. 또한, 이전 연구와 일치하게, 더 위험하다고 인식된 조합 (조합 A)의 경우 대피 결정 시간이 더 짧았음.
이러한 연구 결과는 대피 행동에서 시각적 요소의 역할을 강조하고 VR 화재 대피 시나리오에서 적응과 지각된 위험을 모두 고려해야 할 필요성을 강조함. 이 결과는 화재 안전 교육, 대피 계획, 비상 대응 시뮬레이션에 적용할 수 있음. 향후 연구에서는 데이터 수집 문제를 해결하고, 추가적인 생리적 조치를 통합하며, 다양한 화재 시나리오를 탐색하여 화재 위험에 대한 대응에 대한 이해를 높여야 함. 이는 가상 교육 및 실험 연구에서 화재 재난 상황에서 보다 신뢰할 수 있는 결과에 기여함Abstract iv
List of Tables 7
List of Figures 7
Chapter 1. Introduction 8
1.1. Introduction 8
1.2. Problem statement 11
1.2. Research Objectives 14
Chapter 2. Research Background 17
2.1. Evacuation Research with Virtual Reality 17
2.2. Bio sensing 19
2.2. Research Aim 21
Chapter 3. Experiment 24
3.1. Virtual Environment Design 27
3.2. Experiment Design 32
3.3. Experimental Procedure 37
3.4. Experimental Environment 39
Chapter 4. Methodology 41
4.1. Post-survey 41
4.2. Evacuation Decision time 44
4.3. Electrodermal Activity 46
Chapter 5. Results Collected Through VR Experiments 47
5.1. Post Survey 47
5.2.1 Evacuation Way Decision 49
5.2.2 Evacuation Decision Time 53
5.3. Electrodermal Activity 56
5.4. Discussion 58
Chapter 6. Conclusion 61
Bibliography 64
국문 초록 69석
이미지, 노래, 음악 캡션 간 심층 감정 정보 기반 멀티모달 매칭
학위논문(석사) -- 서울대학교 대학원 : 공과대학 협동과정 인공지능전공, 2024. 8. 강명주.This paper introduces the Multimodal Matching based on Valence and Arousal (MMVA), a tri-modal encoder framework designed to capture emotional content from images, songs, and musical captions. The work expands upon the Image-Music-Emotion-Matching-Net (IMEMNet) dataset by incorporating musical captions aligned with their corresponding songs, resulting in IMEMNet-C which comprises 24,756 images and 25,944 music clips with musical captions. The IMEMNet-C dataset employs a multimodal matching score based on the continuous values of valence (emotional positivity) and arousal (emotional intensity). This continuous matching space allows for random sampling of multimodal pairs during training by computing similarity scores from the valence-arousal values across different modalities. Consequently, the proposed approach achieves state-of-the-art performance in valence-arousal prediction tasks. Furthermore, the framework demonstrates its efficacy in various zeroshot tasks, highlighting the potential of valence and arousal predictions in downstream applications.본 논문은 이미지, 노래, 음악 캡션에서 감정적 내용을 이해하기 위해 설계된 삼중 모달 인코더 프레임워크인 valence (감정의 긍정도)와 arousal (감정의 강도)에 기반한 MMVA (Multimodal Matching Based on Valence and Arousal)를 소개합니다. 이를 위해 기존의 Image-Music-Emotion-Matching-Net (IMEMNet) 데이터셋을 확장하여, 각 음악 데이터에 상응하는 음악 캡션을 포함하도록 확장되어 24,756개의 이미지와 25,944개의 음악 클립과 음악 캡션으로 이루어진 IMEMNET-C을 만들었습니다. 이 데이터셋은 감정의 긍정도와 강도을 기반으로 한 연속적인 매칭 점수를 사용합니다. 이 연속적인 매칭 공간은 학습 중 다양한 모달리티 간의 긍정도-강도 값을 통해 유사성 점수를 계산하여 무작위 샘플링을 가능하게 합니다. 결과적으로, 제안된 접근 방식은 감정 긍정도-강도 예측 작업에서 최첨단 성능을 달성합니다. 또한, 이 프레임워크는 여러 제로샷 (zeroshot) 작업에서 효능을 입증하며, 감정 긍정도와 강도 예측의 다운스트림 응용 가능성을 보여줍니다.Abstract vi
1 Introduction 1
2 Related Work 3
2.1 Multimodal Learning with Tower Encoders 3
2.2 Multimodal Learning with Music 4
3 Dataset 5
3.1 IMEMNet 5
3.2 IMEMNET-C 8
4 Method 10
4.1 Multimodal Matching Based on Valence and Arousal 10
4.2 Emotion Prediction Loss 11
5 Experiments 15
5.1 Training Details 15
5.1.1 Architectures 15
5.2 VA Prediction 16
5.3 Zeroshot Experiments 18
5.3.1 Text-to-Music Retrieval 18
5.3.2 Image-to-Text Retrieval for Music Generation 19
5.3.3 Video Summarization 20
5.4 Ablation Experiment 22
6 Conclusion 23
The bibliography 25
Abstract (in Korean) 30석
음향 감지를 통한 모바일 경험 향상
학위논문(박사) -- 서울대학교 대학원 : 공과대학 전기·정보공학부, 2024. 8. 박세웅.스마트폰과 태블릿/웨어러블 PC와 같은 스마트 기기의 사용이 급격히 증가함에 따라 모바일 애플리케이션은 우리의 일상생활에서 필수적이 되었다. 또한, 모바일 기기가 애플리케이션 기능을 향상시키기 위해 다양한 센서를 장착하면서 사용자 경험이 크게 개선되었다. 이에 따라 추가 하드웨어나 인프라가 필요 없는 내장형 센서 기반 애플리케이션이 통신 및 비통신 목적으로 상업적, 학문적 관심을 끌고 있다. 특히, 사람의 목소리와 환경 소리를 재생하고 녹음하는 데 필수적인 사용자 인터페이스인 스피커와 마이크는 다양한 애플리케이션에서 사용된다. 음향 센싱은 스피커와 마이크로폰을 통해 소리파를 송수신하여 물리적 특성과 이벤트를 포착하는 기술이다. 본 논문은 사용자 상호작용과 기기 기능을 향상시키기 위한 두 가지 음향 센싱 기반 시스템을 제안한다. 첫 번째 시스템은 CSS(Chirp Spread Spectrum) 기반의 공중 음향 통신 시스템이며 두 번째 시스템은 터치스크린 오작동 문제를 해결하기 위한 터치 위치 추적 시스템이다. 먼저, 우리는 CSS 기반 대기중 음향 통신의 신뢰성과 계산 효율성을 향상시켰다. CSS 변조를 채택하고 새로운 처프 심볼을 설계함으로써 상용 모바일 기기의 오디오 인터페이스의 주파수 선택 문제를 해결했다. 또한, 이러한 심볼 디자인에 대한 계산 효율적인 복조 방법을 개발하여 상용 모바일 기기의 전력 소비를 줄였다. 우리는 다중 경로 페이딩과 주변 소음의 영향을 완화하기 위해 계산 복잡성을 증가시키지 않고 프레임 결합 기술을 활용했다. 제안된 기호와 프레임 결합 방법은 프레임 수신 비율을 각각 최대 59.8 퍼센트 포인트(267.9%)와 14 퍼센트 포인트(84.3%)까지 개선했다. 또한, 프레임 결합 방법은 매우 낮은 신호 대 잡음비에서도 두 번의 시도 내에 프레임을 수신할 가능성을 최대 107.9% 증가시켜 과도한 지연을 줄였다. 기기에 따라 제안된 복조 방법은 전력 소비를 수십에서 수백 밀리와트까지 감소시켰다. 둘째로, 우리는 터치 오작동 상황을 보완하기 위한 대체 솔루션으로 비밀 터치스크린 시스템을 제안한다. 스마트 기기에 내장된 스피커와 마이크를 통해 음향 신호를 송수신할 때, 터치하는 손가락으로 인해 전파 특성이 변화한다. 또한, 터치하는 손가락으로 인한 기기의 진동이 관성 측정 장치(Inertial Measurement Unit, IMU) 신호에 영향을 미친다. 우리는 음향 및 IMU 센싱을 통해 그리드 수준에서 터치를 측위하는 다중 모달 분류 모델을 설계했다. 사용자가 원하는 해상도 수준에서 그리드를 선택할 수 있도록 반복적인 그리드 선택을 수행할 수 있는 애플리케이션을 제안한다. 우리는 최소한의 전력 소비로 백그라운드 작업을 위한 IMU 기반 터치 감지 및 시스템 트리거링을 개발했다. 감지 오류 전파 문제를 확인하고 이를 완화하기 위해 이벤트 식별 방법을 고안했다. 감지 및 식별 방법은 각각 3% 미만의 거짓 양성 비율(False Positive Rate)에서 95.3%와 99.1%의 참 양성 비율(True Positive Rate)을 달성했다. 제안된 모델은 4×2 및 6×3 그리드에서 각각 96.98%와 86.45%의 분류 정확도를 달성했다. 본 학위논문에서는 스마트폰에 내장된 응향센서들을 활요한 통신 목적 및 비통신 목적의 모바일 어플리케이션을 제안한다. 상용 스마트 기기들에 대한 실험을 통하여 제안된 센서 기반 시스템들이 다양한 상황에서 사용자들의 모바일 경험을 향상시킬 수 있음을 확인하였다.With the dramatic increase in the use of smart devices such as smartphones and tablet/wearable PCs, mobile applications have become essential to our daily lives. Additionally, as mobile devices come equipped with various sensors to enhance ap- plication functionality, the user experience has improved significantly. Accordingly, embedded sensor-based applications that do not require additional hardware and in- frastructure have attracted commercial and academic interest for communication and non-communication purposes. In particular, speakers and microphones, essential user interfaces (UIs) for playing and recording human voices and environmental sounds, are used in a wide range of applications. Acoustic sensing is a technology that cap- tures physical properties and events by transmitting and receiving sound waves through speakers and microphones. In this dissertation, we propose two acoustic sensing sys- tems to enhance user interaction and device functionality: (i) a chirp spread spectrum (CSS)-based aerial acoustic communication system and (ii) a touch localization system designed to tackle touchscreen malfunctions. First, we enhance the reliability and computational efficiency of chirp spread spectrum- based aerial acoustic communication. By adopting CSS modulation and designing new quaternary symbols, we address the frequency selectivity issues of audio interfaces in commercial off-the-shelf (COTS) devices. Additionally, we develop a computation- ally efficient demodulation method for these symbols, aiming to reduce the power consumption of COTS devices. We utilize a frame combining technique without in- creasing computational complexity to mitigate the effects of multipath fading and am- bient noise in acoustic channels. The proposed symbols and frame combining method improve the frame reception ratio by up to 59.8 percentage points (267.9%) and 14 percentage points (84.3%), respectively. Furthermore, the frame combining method in- creases the likelihood of receiving a frame within two attempts at extremely low SNR by up to 107.9%, reducing excessive delay. Depending on the device, the proposed de- modulation method decreases power consumption by tens to hundreds of milliwatts. Second, we propose a covert touchscreen system, an alternative solution to com- pensate for touch malfunction situations. When transmitting and receiving acoustic signals through the speakers and microphones embedded in smart devices, the propa- gation characteristics change due to a touch finger. Also, the devices vibration caused by a touch finger affects inertial measurement unit (IMU) signals. We design a mul- timodal classification model through acoustic and IMU sensing to localize the touch at the grid level. We propose an application that allows users to perform iterative grid selection, enabling them to choose a grid at their desired resolution level. We develop an IMU-based touch detection and system triggering for background operation with minimal power consumption. We verify the problem of detection error propagation and devise an event identification to mitigate it. The detection and identification meth- ods achieve true positive rates of 95.3 % and 99.1%, respectively, when false positive rates are less than 3%. The proposed model attains classification accuracy of 96.98 % and 86.45 % for 4×2 and 6×3 grids, respectively. keywords: Signal processing, Audio sensor, IMU sensor, Deep learning, Mobile application student number: 2016-20959Abstract i
Contents iii
List of Tables vi
List of Figures vii
1 Introduction 1
1.1 Motivation 1
1.2 Main Contributions 3
1.2.1 Reliable and Low-Complexity Chirp Spread Spectrum-based
Aerial Acoustic Communication . 3
1.2.2 TouchAI: Tackling Touchscreen Malfunction through Audio
and IMU Sensors 4
1.3 Organization of the Dissertation 4
2 Reliable and Low-Complexity Chirp Spread Spectrum-based Aerial Acous-
tic Communication 6
2.1 Introduction . 6
2.2 Related Work 9
2.3 The Preliminaries . 11
2.3.1 Extremely Low Signal Power 11
iii
2.3.2 Frequency selectivity of Audio Interfaces of COTS devices 12
2.3.3 Ambient Noise . 13
2.3.4 Power Consumption of Mobile Devices 15
2.4 Symbol Design 16
2.4.1 Existing CSS Modulation Schemes for AAC 16
2.4.2 Proposed Chirp Symbol Design . 20
2.5 Receiver Design 20
2.5.1 Computation-efficient Correlator . 22
2.5.2 RAKE Receiver . 26
2.5.3 Frame Combining 27
2.6 Performance Evaluation . 31
2.6.1 CTS-chirp . 32
2.6.2 Computation-efficient correlator . 37
2.6.3 Proposed Combining Scheme 38
2.7 Summary 40
3 TouchAI: Tackling Touchscreen Malfunction through Acoustic and IMU
Sensors 42
3.1 Introduction . 42
3.2 Related Work 46
3.2.1 Audio-based Touch Application . 46
3.2.2 IMU-based Touch Application 47
3.3 Channel Measurement . 48
3.3.1 Transceiver Design 48
3.3.2 Channel Impulse Response . 49
3.4 TouchAI: Touch Localization System 51
3.4.1 System Overview 51
3.4.2 Touch Event Detection 54
3.4.3 Touch Event Identification . 56
iv
3.4.4 Touch Localization 58
3.5 Performance Evaluation . 64
3.5.1 Experimental Setup 64
3.5.2 Experimental Results . 64
3.6 Summary 69
4 Conclusion 70
4.1 Research Contributions . 70
4.2 Future Research Directions 71
Abstract (In Korean) 79
v박
다중출력 가우시안 프로세스를 활용한 소비자의 시변 선호 이질성의 모형화에 관한 연구
학위논문(박사) -- 서울대학교 대학원 : 공과대학 협동과정 기술경영·경제·정책전공, 2024. 8. 이종수.Preferences are inherently heterogeneous and not fixed; they can change owing to shifts in the market environment, changes in product characteristics, and consumer learning processes. Previous research has investigated the temporal dependence of consumer preference heterogeneity. However, most previous studies have not considered that, while common patterns may exist in consumers responses to external factors, there may also be differences in sensitivity. This study aimed to address the limitations of existing models by employing machine learning techniques that enable standard choice models and flexible functional modeling. In particular, a model accounting for both the dynamics and heterogeneity of preference parameters using Bayesian–Gaussian processes was proposed. The proposed model is intended to extend choice models to explain preference dynamics by specifically incorporating the individual characteristics of decision makers as factors of sensitivity to change. Existing choice models that treat deviations from population-level parameters as fixed individual heterogeneity only consider the central tendency within the observation window, overlooking important temporal dynamics. Using a model that ignores dynamic heterogeneity precludes the possibility of tailoring marketing strategies to these changes. The research results have practical applications in the design of personalized dynamic pricing and incentive schemes tailored to customer characteristics. Therefore, the proposed model can generally be applied to platform businesses where demand and supply conditions change rapidly owing to changes in the external environment.
Keywords: time-varying parameter estimation, discrete choice model, semi-nonparametric model, consumer preference structure, platform economy선호는 본질적으로 이질적일 뿐만 아니라 고정되어 있지 않다. 시장 환경의 변화나, 상품 자체 특성의 변화, 그리고 소비자의 학습을 통해서도 변화할 수 있다. 이러한 점을 고려하여 기존 연구에서는 소비자 선호 이질성의 시간 의존성에 대한 탐색을 시도한 바 있다. 비록 외부 요인에 대한 소비자 반응에는 일정 수준에서 공동의 패턴이 존재할 수 있으나, 그 민감도에는 차이가 존재할 수 있다는 점이 이전의 연구들 대부분에서는 크게 고려되지 않아왔다. 본 연구는 기존모형의 한계를 보완하고자 목표는 표준 선택모형과 유연한 함수 모형화를 가능하게 하는 머신러닝 기법을 활용한다. 베이지안 가우시안 프로세스를 활용함으로써 선호 파라미터의 동학과 이질성을 동시에 포괄하는 모형을 제안하였다. 기존의 선호 동학을 설명하기 위한 모형을 확장한다는 측면에서 기여가 존재한다. 개인의 이질성이 모집단 수준의 파라미터로부터 고정된 편차를 시간 변화와 무관하게 일정하게 유지한다고 가정한 기존 선택 모델은 필연적으로 관찰 기간 내 중심 경향만을 포착하는데 한정될 수 밖에 없다. 이는 중요한 시간적 역학을 간과하며, 소비자의 미래 행동 예측에 기반한 마케팅 전략의 수정 및 수립 역량을 제한한다. 특히, 제안된 모형은 선호 변화에 대한 민감도를 결정하는 요인으로서 의사결정자 개인의 특성의 역할을 모형 내에 포괄한다. 연구 결과의 활용 측면에서는 외부 환경 변화에 따라 빠르게 수요-공급 조건이 변화하는 플랫폼 비지니스 전반에 일반화하여 적용할 수 있다는 점에서 기여를 기대할 수 있다.
주요어 : 동적 파라미터 추정, 이산선택모형, 반비모수 모형, 소비자 선호구조, 플랫폼경제Chapter 1. Introduction 1
1.1 Research Background 1
1.2 Research Objectives 6
1.3 Outline of the Study 7
Chapter 2. Literature Review 9
2.1 Theoretical Background 9
2.1.1 Consumer Learning and Preference Evolution 10
2.1.2 Effects of Marketing Mix Variables 14
2.2 Econometric Models of Preference Evolution 16
2.2.1 Reparameterization Approaches 17
2.2.2 State-pace Models and Nonparametric Methods 22
2.2.3 Limitations in Existing Models 26
2.3 Research Motivations of This Study 27
Chapter 3. Methodology and Model Framework 29
3.1 Multi-Output Gaussian Processes 29
3.1.1 Gaussian Processes (GPs) 29
3.1.2 Multi-Output Gaussian Processes (MOGPs) 36
3.1.3 Benefits of GP applications in Marketing Research 37
3.1.4 Distinction from conventional GP applications 38
3.2 GP-based Model for Heterogeneous Preference Dynamics 39
3.2.1 GP Applications in Choice Context 39
3.2.2 Full Model Specifications 40
3.2.3 Estimation Procedures 46
Chapter 4. Synthetic Data Analysis 51
4.1 Synthetic Data Generation for Simulation Analysis 51
4.2 Estimation Results from the Proposed Model 53
4.2.1 Time-Varying Parameter Estimates with Specification 1 53
4.2.2 Time-Varying Parameter Estimates with Specification 2 58
4.3 Performance of the Proposed Model 62
Chapter 5. Empirical Analysis Results 71
5.1 Dynamic Preference Heterogeneity on Ride-hailing Choices 71
5.1.1 Data 73
5.1.1 Choice Set Generation and Analysis Sample Formation 75
5.1.2 Market and Macro-level Variables 88
5.2 Estimation Results 90
5.2.1 Specification 1: ARMA (1) Mean Model 90
5.2.2 Specification 2: GP as Mean Model 99
5.2.3 Post-hoc Analysis 105
5.3 Discussion 108
5.4 Out-of-Sample Predictive Accuracy 109
5.5 Characterizing Consumers based on Sensitivity Dynamics 113
5.6 Sources of Individual-level Sensitivity 122
Chapter 6. Conclusion 128
6.1 Concluding Remarks 128
6.2 Limitation and Future Studies 131
Bibliography 134
Appendix 1: TLC Taxi Zone List 147
Appendix 2: Selected Data Dictionary for NYC High Volume FHV Trip Records 152
Appendix 3: Summary Statistics of Selected Sample 153
Appendix 4: Individual-Level Preference Sensitivity Estimates (Specification 1) 154
Abstract (Korean) 157박
Central Neurocytoma and Glioneuronal Tumor
학위논문(박사) -- 서울대학교 대학원 : 의과대학 의과학과, 2024. 8. 김종일.다양한 차세대 염기서열 분석 기술이 발전함에 따라 다양한 암종들의 분자적 특징을 알 수 있게 되었고, 더 나아가 다중 오믹스를 통해 보다 통합적으로 이해할 수 있게 되었다. 세계보건기구에서 중추신경계 종양 종류를 분자적 특성과 조직학 및 면역 조직화학 특징들을 기반으로 하여 분류한다. 중추 신경계 종양들에 대한 연구가 계속 진행되고 있는데, 본 연구에서는 현재까지 유전적 특징이 보고되지 않은 중추신경세포종 (central neurocytoma)을 다중오믹스를 이용하여 분석하였고 배아 이형성 신경상피 종양 (dysembryoplastic neuroepithelial tumor)과 glioneuronal tumor의 새로운 유전적 특징을 확인했다.
중추신경세포종은 중추신경계 종양 중 0.1~0.5%에 해당하는 희귀 종양으로써 유전적 특징이 밝혀지지 않은 암종이다. 해당 종양의 특징을 확인하기 위해 6명의 환자로부터 얻은 조직으로 DNA, RNA, methylation 시퀀싱을 진행했고 3명의 환자로부터 얻은 조직으로 단일 핵 RNA 시퀀싱 (single nuclei RNA sequencing)을 진행했다. 중추신경세포종의 기원 세포를 확인해 본 결과, 방사신경아교세포가 신경 세포로 분화하는 과정에서 발생한 것으로 확인됐다. FGFR3의 저메틸기에 의해 해당 유전자가 과발현됨에 따라 PI3K-AKT 신호 전달 경로가 활성화되고, 뉴런의 발달과 신경발생 등 신경 관련 신호 전달 경로들이 비활성화 되면서 신경세포로 분화하지 못하고 종양으로 발전한 것으로 확인됐다. 중추신경세포종의 발생 원인은 종양 유발 변이, 복제수 변이 등이 확인되지 않은 것을 미루어 보아 유전체 변이가 아닌 후성유전체에 의한 유전자 과발현에 의한 것으로 확인됐다.
배아 이형성 신경상피 종양은 약물 내성 뇌전증을 동반한 저등급 신경교종 (low grade glioma)으로 해당 종양의 절반 정도에서 주요 종양과 함께 위성병변 (satellite lesions)이 확인된다. 위성병변을 동반할 경우 재발과 강한 연관성을 보이는데, 주요 종양과 위성병변이 산발적으로 발생하는지, 혹은 한 곳에서부터 발생하였는지, 그리고 두 종양이 어떤 유전적 차이가 있는지를 알아보기 위해 DNA와 RNA 시퀀싱을 진행했다. 3명의 환자로부터 얻은 7개의 주요종양과 8개의 위성병변의 분자적 특징을 확인하였는데, 기존에 배아 이형성 신경상피 종양에서 빈번히 확인되는 FGFR1 K656E와 K655I 돌연변이를 발견했고, 두 돌연변이는 시스 복합이형접합 (compound heterozygous in cis) 형태로 존재하는 것을 확인했다. 주요 종양과 위성병변 모두에서 FGFR1 변이 및 변화를 공유하고 있었고, 이외에 공통적으로 공유하고 있는 변이는 확인되지 않았다. 계통 발생 분석 결과, 주요 종양과 위성병변은 FGFR1 변이 및 변화만을 공유한채 독립적으로 발달했다고 해석할 수 있었다.
세번째 연구에서는 glioneuronal tumor환자의 사례 연구를 진행했다. Glioneuronal tumor는 신경과 신경교가 혼합되어 구성되어 있는 종양이다. 30세 여성의 좌측 두정엽에서 종양이 발생했고 DNA와 RNA 시퀀싱을 이용해 분석을 진행했다. 암 발생과 관련된 유전자 변이는 확인되지 않았고 기존에 보고되지 않은 CLIP2-MET 융합 유전자를 확인했다. 융합 유전자에 의해 MET 유전자의 과발현을 확인하였고, 이에 따라 MAPK pathway가 활성화되어 종양이 발생하는 경우에 해당한다고 보고했다.
다중오믹스 접근 방식을 활용하여 희귀 중추신경계 종양의 분자적 특징을 복합적으로 확인할 수 있었고 종양의 발생 기전에 대한 분자 유전학적 특징을 확인할 수 있었다. 중추신경계 종양에는 배아 이형성 신경상피 종양과 같이 뚜렷하게 유전자 변이가 확인되는 종양이 있는 반면에 중추신경세포종과 glioneuronal tumor와 같이 유전자 변이가 없는 종양들도 확인되었다. 세개의 연구 결과, 유전자 변이 뿐만 아니라 후성유전체와 유전자의 발현을 함께 확인함으로써 중추 신경계 종양을 보다 더 자세히 분류하는데 기여했다. 이를 통해 종양의 분자적 특징을 정확히 이해하기 위해서는 단일 시퀀싱 기법만을 사용하기 보다는 다중 오믹스 접근 방식이 필요함을 제시했다.As next-generation sequencing technologies have advanced, the molecular characteristics of various types of cancer have become better understood, and multi-omics approaches have enabled integrated analyses of the same sample, thereby enhancing our understanding of its molecular characteristics. The World Health Organization (WHO) classifies central nervous system (CNS) tumors based on molecular characteristics, histology, and immunohistochemical features. The research on CNS tumor is still ongoing process. In this study, I analyzed central neurocytoma, a type of tumor whose genetic characteristics have not yet been reported, using multi-omics approaches. Furthermore, I identified new genetic features of dysembryoplastic neuroepithelial tumors (DNET) and glioneuronal tumors.
Central neurocytoma is a rare tumor, accounting for 0.1 – 0.5% of CNS tumors, and its genetic characteristics have not been elucidated. To characterize central neurocytoma, we performed DNA, RNA and methylation sequencing with six patients, and single nuclei RNA sequencing with three patients. I discovered that central neurocytoma arises during the differentiation of radial glial cells into neurons, and originated from radial glial cells. The hypomethylation of FGFR3 led to the overexpression of this gene, activating the PI3K-AKT signaling pathway, while neuronal related signaling pathways such as neuronal development and neurogenesis were inactivated, inhibiting neuronal differentiation and resulting in tumor development. The absence of tumor-inducing mutations and copy number variations suggests that the cause of central neurocytoma is gene over expression driven by epigenetic alterations rather than genomic mutations.
DNET is a type of low-grade glioma associated with drug-resistant epilepsy, and satellite lesions are observed in approximately half of DNET patients. The presence of satellite lesions is strongly correlated with recurrence. To determine whether the main tumor and satellite lesions occur independently or originate from the same site, and to identify their genetic features, I conducted DNA and RNA sequencing. Seven main mass samples and eight satellite lesions samples from three patients were obtained to discover molecular characteristics. I identified FGFR1 K656E and K655I mutations, which are frequently observed in DNET, existing as compound heterozygous in cis. Both main mass and satellite lesions shared FGFR1 mutations or alterations, but no other common mutations were detected. Phylogenetic analysis suggested that main mass and satellite lesions developed independently, as they only shared FGFR1 mutation or alterations.
The third study is a case study of a patients with a glioneuronal tumor. A 30-year-old female had a tumor in the left parietal lobe. Glioneuronal tumors consist of a mix neuronal and glial cell. DNA and RNA sequencing were performed to analyze the molecular features of this tumor. No gene mutations were detected, however, a previously unreported CLIP2-MET fusion gene was identified. The fusion gene resulted in overexpression of MET gene, activating the MAPK pathway, which contributed to tumor development.
Using multi-omics approach, we comprehensively identified the molecular characteristics of rare CNS tumors and elucidated the molecular genetic mechanism underlying tumorigenesis. In some CNS tumors, such as DNET, accompanying gene mutation or alterations are observed, while others, like central neurocytoma and glioneuronal tumor, do not exhibit genetic alterations. These findings underscore the importance of integrating epigenomic and gene expression data, in addition to genetic alterations, to achieve a more comprehensive classification of CNS tumors. This study suggests the necessity of multi-omics approaches rather than relying on a single sequencing method to accurately understand the molecular characteristics of tumors.Abstract i
Introduction viii
Chapter 1. 1
Introduction 2
Materials and methods 4
Results 13
Discussion 101
Chapter 2 107
Introduction . 108
Materials and methods 111
Results 114
Discussion 132
Chapter 3 136
Introduction . 137
Materials and methods 139
Results 144
Discussion 162
Discussion 164
Bibliography 166
국 문 초 록 175박
3차원 공간 속 에이전트를 위한 객체 지향 장면 표현학습
학위논문(박사) -- 서울대학교 대학원 : 공과대학 전기·정보공학부, 2024. 8. 김영민.Recent advances in robotics include the development of generalizable visual intelligence for robot agents in 3D environments, significantly enhancing robotic autonomy through deep learning and computer vision.
Specifically, main capabilities that researchers aim to achieve are threefold: recognizing diverse objects in large spaces, planning manipulations based on visual observations, and maintaining self-awareness informed by visual data, thus enabling robots to perceive and interact with 3D spaces with human-like accuracy.
Given the current state of computer vision and deep learning, various candidate approaches can be taken to develop visual intelligence in robots that satisfies the three capabilities.
To this end, this thesis primarily focuses on the integration of two major fields among the candidates: object-centric scene representation (OSR) and the world model.
Firstly, OSR is a scene representation that assumes the visual observations of the agent primarily consist of \emph{objects}.
Furthermore, it posits that each object is independently bound to a single latent variable, thereby enabling the agent to represent the 3D scene as a set of these latent variables.
Next, the world model formulates the causal interaction between the agent's actions and its surrounding world. By utilizing this model, the agent can predict future state changes induced by its actions and the corresponding costs for solving the task. Consequently, the agent is able to plan tasks based on state and cost predictions.
Both the OSR and the world model incorporate inductive bias for segmenting the semantics of objects and predicting the causal interactions between scene entities, naturally arising from their model architecture.
Clearly, the supervision-based video segmentation model shows remarkable performance recently, as it can robustly track the objects in the observation of the agent, thus it can be a good candidate for building the visual intelligence of a robot.
Nonetheless, in this thesis, I emphasize the specialized capabilities of OSR and the world model that cannot be achieved by supervised models.
Specifically, I take a modular approach by dividing the challenges of applying OSR and the world model into subproblems and addressing them individually.
To this end, this thesis first proposes a method to extract a set of structured latent variables from the visual observations of the agent.
The proposed model decomposes the visual observations of the robot agent physically interacting with surrounding objects in an object-centric manner.
Furthermore, it makes causal inferences about the interaction between the robot agent and the objects by defining an element of OSR that is dedicated to the robot agent.
Next, this thesis presents a control policy for the robot agent based on the agent-centric OSR acquired from visual observations.
This is mainly achieved by addressing issues arising from the dynamic interaction between the agent and objects and leveraging the task planning capability of the OSR-based world model.
Subsequently, I utilize the concept of \emph{neural fields} to lift the OSR into 3D space, allowing the extension of the agent-centric OSR and world model to more diverse scenarios.
As a result, the agent can comprehend the holistic shapes of 3D objects and predict the dynamic interactions between them, thereby enhancing its spatial understanding of the 3D environment.
Finally, this thesis proposes a future direction for integrating the 3D OSR and the world model based on planning.
The model would enable the robot agent in a 3D space to acquire prior knowledge of objects from visual observations taken from various viewpoints and to model the 3D dynamics between them.
This approach can be applied to robot agents, such as mobile manipulator platforms, to model the causal interactions induced by the agent's actions in the 3D space.
As a result, what I am ultimately aiming to prove through the three modular approaches and the blueprint of the integrated model based on them is that the integration of OSR and the world model can be a simple yet effective solution for achieving generalizable visual intelligence in robot agents.
I hope that the studies discussed in this thesis can at least provide a small hint to help bring the theoretical OSR field to a more practical level for researchers seeking to build generalizable visual intelligence in robots.로보틱스는 최근 3차원 공간에서 일반화 된 시각지능 연구의 발전에 힘입어 같이 발전하고 있으며, 특히 딥러닝 및 컴퓨터비전 분야의 발전을 통해 비전 기반 로봇 제어 자동화가 많이 연구되고 있다.
이러한 기술들을 통해 궁극적으로 추구하는 바는 로봇이 3차원 공간에서 다양한 객체들을 인식하고, 시각 정보에 기반해 물체 조작 계획하거나, 시각 정보 속에서 로봇 에이전트 (agent) 자신을 인식하는데 있다. 이로써 로봇은 사람과 같은 방식으로 물체들과 상호작용 할 수 있게된다.
현대의 컴퓨터비전 기술들의 수준에 비추어 볼 때, 이 기술들을 통해 로봇이 앞서 언급한 지능을 갖추도록 학습하는데 있어서는 여러가지 접근이 가능할 것이다.
본 학위논문을 통해 제안하는 연구는 여러 기술 후보군 중에 크게 두 가지 분야의 기술들을 활용하여 언급한 로봇 지능 구현의 뼈대를 구축하고자 한다.
두 가지 기술은 크게 객체 중심 장면 표현 (object-centric scene representation; OSR)과 월드 모델 (world-model)로 이루어져 있다.
먼저, 객체 중심의 장면 표현 (OSR)은 로봇을 포함한 다양한 에이전트가 관측하는 시각 정보가 기본적인 객체 (object)들로 이루어졌다고 가정하는 표현형 (representation)이다.
여기에 더해, 각 객체가 독립적으로 하나의 잠재 변수에 사상 (mapping)된다는 가정을 한다.
이로써, OSR 표현형은 에이전트가 관찰하는 3차원 공간을 OSR 잠재변수들의 집합으로 표현하게 된다.
다음으로, 월드 모델 (world-model)은 에이전트의 행동 (action)에 대한 주변 환경의 인과적 (causal) 변화를 함수로써 모델링하는 방법론이다.
이로써 에이전트가 실제 환경과 직접 상호작용 하지 않으면서도 자신의 행동에 대한 환경의 변화와 해당 행동이 주어진 문제를 푸는데 얼마나 도움이 되는지를
예측할 수 있다. 결과적으로 에이전트는 이를 바탕으로 작업에 대한 계획 (planning)을 수행할 수 있게된다.
OSR 및 world-model은 모두 모델의 구조 설계에서 귀납적 편향 (inductive bias)을 반영하여 학습 데이터 자체만으로 에이전트에게 의미 있는 정보를 추출할 수 있도록 한다.
물론, 최근 지도학습 (supervised learning) 형태로 학습된 모델들이 비디오 입력에 대해 객체 위주의 장면 분리 (segmentation) 및 추적 (tracking)에 있어서 뛰어난
성능을 보여주기 때문에 로봇의 객체 지향 시각 지능 구현에 있어 적합하다고 볼 수 있다.
본 연구에서는 OSR 및 world model이 지도학습 기반 모델이 제공해 줄 수 없는 능력들에 집중하여 로봇의 시각지능을 구현하고자 한다.
이때, OSR 및 world model을 기반으로 로봇의 시각 지능을 구현하는데 있어 마주할 수 있는 난점 (challenge)들을 세분화 하여 모듈러 (modular)한 접근법을 취하고자 한다.
이를 위해, 본 학위논문에서는 먼저 에이전트가 관측하는 시각 정보들에서 OSR에 기반하여 로봇 에이전트 중심의 구조화된 잠재 변수들을 추출할 수 있는 방법을 제안한다.
제안된 모델은 비지도학습을 통해 로봇과 물체들이 물리적으로 상호작용하는 관측 정보를 객체 위주로 분해하거나 각 객체들의 변화를 시간에 따라 추적하며, 로봇 에이전트와 물체 간 상호작용의 결과에 대한 인과적 추론이 가능하다. 여기서, 기존 모델들이 로봇 에이전트를 강인하게 하나의 OSR로 표현하지 못하는 문제점을 중점적으로 해결하였다.
다음으로, 이러한 표현형을 바탕으로 로봇 에이전트가 시각 정보를 입력으로 하여 여러 물체들을 조작하는 정책함수 (policy function)를 학습하는 모델을 제시한다.
제안된 연구는 에이전트와 물체들의 동적인 상호작용에 의해 발생하는 문제점을 해결하고, OSR에 기반한 world-model을 활용해 물체 조작 작업계획을 수립하는데 집중한다.
이어서, 본 학위논문에서는 최근 각광받고 있는 신경장 (neural field)을 활용하여 앞서 제안된 OSR을 3차원 공간에서 취득한 데이터에서 학습할 수 있도록 한다.
이로써 공간 속 물체들의 3차원 형상을 OSR로 모델링함과 동시에 이들의 상호작용 및 동적인 변화를 신경장을 통해 모델링하여 로봇이 보다 넓은 공간적 이해력을 가질 수 있도록 한다.
마지막으로, 앞서 제안한 3차원 OSR 및 에이전트의 world model기반 작업계획을 통합하여 3차원 공간 속 로봇 에이전트가 다양한 시점 (viewpoint)들에서 관측되는 정보에서
공통된 객체들을 OSR로 표현하고 이 변수들의 모바일 매니퓰레이터 (mobile manipulator) 에이전트의 동작에 의한 인과 추론을 기반으로 작업 계획을 수행하는 향후 연구 방향의 청사진을 제안한다.
최종적으로, 나는 본 학위논문에서 다룰 세 가지의 모듈러한 방법론 및 이들을 바탕으로 한 향후 연구 방향에 대한 흐름을 통해 OSR 및 world-model 방법론이 로봇의 일반화된 시각지능을 구현하는 데 있어
간단하면서도 실용적인 수준의 성능을 보여주는 후보군이 될 수 있음을 보여주고자 한다. 이를 통해 향후 이 분야의 연구자들이 다분히 이론적인 도메인에서 많이 고민되고 있는 OSR 및 world-model 연구가
로봇, 특히 모바일 매니퓰레이터와 같이 실용적인 어플리케이션에 적용될 수 있도록 하는데 도움이 될 수 있기를 바란다.Abstract i
Chapter 1 Introduction 1
Chapter 2 Agent-centric Scene Representation 8
2.1 Related Work 10
2.1.1 Object-centric Representation Learning 10
2.1.2 Latent Dynamics Model from Visual Sequence 11
2.2 GATSBI: Generative Agent centric Spatio temporal Object Interaction 12
2.2.1 Entity-wise Decomposition . 14
2.2.2 Interaction 19
2.2.3 Implementation Details 20
2.3 Experiments . 23
2.3.1 Qualitative Results on Spatial Decomposition 24
2.3.2 Agent-centric Spatio-temporal Interaction 24
2.3.3 Ablation Study on Interaction 27
2.4 Conclusion . 28
Chapter 3 Compositional World Model with Agent-centric Scene Repre-
sentation 32
3.1 Related Works 34
3.1.1 Object-centric Scene Representation 35
3.1.2 Vision-based RL 36
3.2 Agent-centric Scene representation In Multi-Object manipulation (ASIMO) 38
3.2.1 Preliminaries 40
3.2.2 Learning Long-term Agent Dynamics . 45
iv
CONTENTS
3.2.3 Learning Occlusion-robust Object Dynamics 49
3.2.4 Hierarchical and Compositional Model-based RL Agent 56
3.2.5 Overall Training Method 61
3.2.6 Implementation Details 62
3.3 Experiments . 67
3.3.1 Experimental Setting and Baselines 67
3.3.2 Performance on the Scene Representation Learning 69
3.3.3 Performance on Model-based Reinforcement Learning . 75
3.4 Conclusion . 76
Chapter 4 Lifting the Representation to 3D Field 90
4.1 Preliminaries and Related Works 92
4.2 Methods 95
4.2.1 Scene Modeling of TRITON 96
4.2.2 Object Dynamics Prediction 98
4.2.3 Enhancing the Learning Stability with Regularization 99
4.2.4 Implementation Details 100
4.2.5 Details on the interaction Transformer . 101
4.3 Experiments . 103
4.4 Conclusion . 108
Chapter 5 Limitations and Future Work 110
5.1 Limitations . 110
5.2 Future Work: Generalizable & Actionable Neural Fields for Mobile
Manipulation 116
Chapter 6 Conclusion 123
Appendix A Appendix 125
A Additional Samples of GATSBI 126
A.1 Spatial Decomposition 126
v
CONTENTS
A.2 Temporal Prediction . 126
A.3 Spatio-temporal Prediction . 127
A.4 Physically Plausible Samples 127
A.5 Quantitative Evaluations 127
A.6 Additional Ablation Study . 127
B Additional Samples of ASIMO 130
B.1 Experimental Results on SAVi and STEVE . 130
B.2 Additional Experimental Results . 136
초록 172
Acknowledgements 175
vi박
최소최대 최적화와 고정점 문제의 최적 알고리즘에 관하여
학위논문(박사) -- 서울대학교 대학원 : 자연과학대학 수리과학부, 2024. 8. Ernest Ryu.While computational complexity of solving a class of problems has always been a central topic in optimization theory, the understanding of optimal complexity for minimax optimization and fixed-point problems has been incomplete until recent years. This work presents algorithms and their analyses with optimal complexity for reducing the magnitude of gradient (for minimax problems) and fixed-point residual (for fixed-point problems).
For fixed-point problems, accelerated algorithms based on the so-called anchoring mechanism have been lately proposed, and exact matching complexity lower bound established their optimality. Therefore, anchoring was thought to be "the" correct acceleration mechanism for the setup. Contrarily to this view, we present the surprising observation that the optimal acceleration mechanism fixed-point problems is not unique. Specifically, we introduce a new algorithm with the same exact optimal worst-case rate, forming a certain sense of duality (between algorithms) with the prior anchor-based one. Furthermore, we discover a continuous family of exact optimal algorithms interpolating between the two, all achieving the same worst-case complexity. The resulting new acceleration mechanisms empirically exhibit varied characteristics, and are materially different from anchoring.
For smooth minimax problems, we demonstrate algorithms with accelerated O(1/k^2) rates, using either the anchoring mechanism or its dual, identified from the fixed-point setup. We then establish optimality of their O(1/k^2) rates through a matching lower bound.
Finally, we explore two approaches toward building formal connections underlying the apparent resemblance between the acceleration mechanisms working in the two problem setups. On one hand, we establish the novel merging path property, which quantifies the asymptotic similarity displayed by trajectories of fixed-point and minimax algorithms using anchoring. On the other hand, we discuss the continuous-time limits, under which the algorithms of anchor and dual-anchor types respectively reduce to a single ordinary differential equation.특정 문제에 대한 연산 복잡도를 규명하는 것은 항상 최적화 이론의 중요 주제였음에도 불구하고, 최소최대 최적화와 고정점 문제에 대한 최적 복잡도에 대한 이해는 최근까지도 불완전하였다. 이 논문에서는 최소최대 목적함수의 기울기 벡터 혹은 고정점 문제의 잔차 크기를 줄이는 데 있어 최적 복잡도를 가지는 알고리즘들과 그 분석을 논한다.
최근 고정점 문제에서 "앵커"라고 불리는 메커니즘을 사용한 가속화된 알고리즘들이 제시되었으며, 그 복잡도와 정확히 일치하는 하한 결과가 제시됨으로써 그 최적성이 증명되었다. 따라서 앵커 메커니즘이 고정점 문제에 있어 유일한 정답인 것으로 보이기도 하였다. 우리는 이러한 시각과 반대되는 예상치 못한 결과, 즉 고정점 문제의 최적 가속 메커니즘이 유일하지 않음을 보인다. 구체적으로, 우리는 정확히 같은 복잡도를 가지는 새로운 알고리즘을 제시하며, 이것이 기존의 앵커 기반 알고리즘과 특정 형태의 쌍대 관계에 있음을 보인다. 더 나아가 이 두 알고리즘 사이를 연속적으로 연결하는 최적 알고리즘 족이 존재하며 이것들이 모두 같은 복잡도를 지님을 발견한다. 이 새로운 알고리즘들은 서로 다른 행동 양상을 보이며, 앵커 메커니즘과는 본질적으로 상이하다.
최소최대 문제에 대해서는, 앞의 고정점 문제에 대한 연구를 통해 규명한 앵커 또는 쌍대 앵커 메커니즘을 기반으로 O(1/k^2)의 가속화된 속도로 수렴하는 알고리즘을 소개한다. 그리고 이에 대응되는 수렴 속도의 하한을 증명함으로써 O(1/k^2) 속도가 실제로 최적임을 보인다.
마지막으로, 서로 다른 이 두 문제에 대한 가속화된 알고리즘 (또는 이에 사용된 가속 메커니즘) 사이에서 관찰된 형태적 유사성을 설명하기 위한 두 가지 접근법을 소개한다. 한 가지는 merging path 관점으로, 앵커 기반의 고정점 및 최소최대 알고리즘 궤적 사이의 점근적 유사성을 정량적으로 분석하는 것이다. 다른 하나는 연속 시간 모델 관점으로, 학습률이 0으로 가는 극한 하에서 앵커 및 쌍대 앵커 알고리즘들이 각각 하나의 상미분방정식으로 표현되는 것을 보인다.Abstract i
1 Introduction 1
1.1 Organization of contents 2
1.2 Prior work 3
2 Preliminaries 6
2.1 Basic concepts and notations 6
2.1.1 Set-valued operators 6
2.1.2 Fixed-point and monotone inclusion problems 7
2.1.3 Minimax optimization and monotone operators 8
2.2 Exact optimality of Optimal Halpern Method 9
3 Family of exact optimal fixed-point algorithms 11
3.1 Dual Optimal Halpern Method 11
3.1.1 Convergence analysis of Dual-OHM 12
3.2 Duality correspondence between fixed-point algorithms 15
3.2.1 H-matrix representation 15
3.2.2 H-dual operation 16
3.2.3 H-duality theorem 20
3.2.4 Proof of Theorem 3.2.1 22
3.3 Continuous family of exact optimal fixed-point algorithms 28
3.3.1 Proof outline for Theorem 3.3.1 30
3.3.2 Proof of Proposition 3.3.1 33
3.3.3 Optimal algorithms not covered by Theorem 3.3.1 35
4 Accelerated Minimax Optimization 39
4.1 Anchor type acceleration 39
4.1.1 Extra Anchored Gradient (EAG) algorithm 40
4.1.2 Proofs of Lemmas 4.1.1 and 4.1.2 43
4.1.3 Other acceleration results of anchor type 46
4.2 Dual-anchor acceleration 47
4.3 Optimality via matching lower bound 56
4.3.1 Construction of worst-case biaffine function 57
4.3.2 Tightness of biaffine lower bound 59
4.3.3 Matrix polynomial analysis 62
5 Unified view on fixed-point and minimax acceleration 70
5.1 Merging path property among anchor type acceleration 71
5.1.1 Smooth minimax problems vs. nonexpansive fixed-point problems 72
5.1.2 Strongly monotone Lipschitz inclusion problems vs. contractive fixed-point problems 77
5.2 Continuous-time limits for anchor and dual-anchor acceleration 87
5.2.1 Anchor type algorithms and Anchor ODE 89
5.2.2 Dual-OHM, Dual-FEG and Dual-Anchor ODE 90
6 Conclusion and future directions 97
A Algorithm specifications 98
A.1 Classical minimax algorithms and rates 98
A.2 Nesterovs AGM and its momentum parameters 99
B Proof of the optimal family theorem 100
B.1 Step 2: Computing coefficients of vector quadratic form 100
B.2 Step 3: Explicit characterization of λ 102
B.3 Step 4: Solving the linear system s_{ℓ,k} = 0 104
C Proof of the continuous-time limits 123
C.1 Convergence to Anchor ODE 123
C.2 Convergence to Dual-Anchor ODE 130
C.2.1 From Dual-OHM to Dual-Anchor ODE 130
C.2.2 From Dual-FEG to Dual-Anchor ODE 137
Bibliography 141
Abstract (in Korean) 150
Acknowledgement (in Korean) 151박
Mixed Method of HINTS Secondary Data Analysis and Cross-Sectional Survey
학위논문(석사) -- 서울대학교 대학원 : 사회과학대학 언론정보학과, 2024. 8. 이철주.환자들의 의료적 의사 결정에 있어 온라인 의료 기록의 유용성은 오래도록 연구되어 왔다. 예컨대 환자들은 온라인으로 자신의 의료 기록에 접근함으로써 스스로의 건강 상태를 모니터링하는 한편 의료 서비스 제공자들과 이에 대해 논의할 수 있게 된다. 이처럼 환자들이 전자건강기록(Electronic Health Record, EHR)에 접근함으로써 기대할 수 있는 효과 중 하나는 대등한 정보 교환 및 공동 의사결정을 포함한 환자중심 커뮤니케이션(Patient-Centered Communication, PCC)이다.
본 연구는 환자들의 EHR 활용을 분석하기 위한 근거 이론으로 인지된 이용 용이성 및 유용성, 이용 의도 (즉, 기술 수용), 실제 이용 등을 주요 변수로 관찰하는 기술 수용 모델(Technology Acceptance Model, TAM; Davis, 1989)을 채택하였다. 본 연구는 해당 모델의 경로 확인에서 나아가 전자건강기록 활용과 환자중심 커뮤니케이션 간의 연관성까지 확장적으로 검토하였다. 미국의 건강정보 국가 동향 조사(Health Information National Trends Survey, HINTS) 주기 1(2017)과 주기 3(2019)의 2차 데이터 분석(연구 1) 및 18세 이상 미국 시민을 대상으로 한 횡단 설문조사(연구 2)를 통해 보완적 연구를 진행하였으며, 통계적으로 구조 방정식 및 경로 분석을 활용하여 두 주요 변수인 전자건강기록 활용과 환자중심 커뮤니케이션 간 관계를 설명하였다.
연구 1과 2를 통해 전자건강기록에 대한 인지된 이용 용이성과 실제 활용(연구 1) 및 활용 의도(연구 2) 간 인지된 유용성의 매개 효과, 그리고 전자건강기록 활용과 환자중심 커뮤니케이션 간 상관관계가 나타났다. 또한 연구 2를 통해 의료 서비스 제공자의 전자건강기록 활용 격려가 의도와 달리 부(-)의 조절 효과를 일으킬 수 있음이 밝혀졌으며, 이는 환자 개인의 자기 결정 및 자율성 존중의 중요성을 시사한다.Usefulness of online medical record (i.e., Electronic Health Record; EHR) in engaging patients into medical decision making has long been explored—e.g., providing patients of access to their medical history online and opportunities to monitor and further discuss their health status with health care providers. One of the expected outcomes of patients who have access to EHR is patient-centered communication (PCC) in provider-patient relationship, which includes information exchange, decision making, and more to come. The use of this relatively new technology, EHR, was investigated based on Technology Acceptance Model (TAM; Davis, 1989) which observes perceived ease, perceived usefulness, intention to use (i.e., acceptance of technology), and actual use as main variables; and the research proceeded to examining association between the use of EHR and PCC. This paper adopted complementary methods of secondary data analysis on two iterations of Health Information National Trends Survey data (HINTS 5 Cycle 1 and Cycle 3) (Study 1) and cross-sectional survey conducted in the United States (Study 2), utilizing structural equation modeling (SEM) and path analysis. Mediation effect of perceived usefulness between perceived ease of use and actual use (Study 1) or behavioral intention to use (Study 2) and correlation between EHR use and PCC were confirmed. Moreover, negative moderation effect of encouraged EHR use by health care providers between behavioral intention to use and actual use was observed, underlining the importance of self-determination and autonomy of individual patients.Table of Contents
1. Introduction 1
2. Literature Review 5
2.1. Patients Use of Online Medical Record 5
2.2. Adoption of Technology Acceptance Model on EHR Use 6
2.3. Investigation on Relationship between EHR Use and Patient-Centered Communication 8
2.4. Encouraged EHR Use, Health Literacy, and Chronic Disease History as Moderators 11
3. Hypotheses and Research Questions 14
3.1. Study 1: Secondary Data Analysis 14
3.2. Study 2: Cross-Sectional Survey 15
4. Method 17
4.1. Study 1: Research Design and Sample Dataset 17
4.2. Study 1: Measures 17
4.3. Study 1: Analysis Method 22
4.4. Study 2: Research Design and Participants 23
4.5. Study 2: Measures 24
4.6. Study 2: Analysis Method 27
5. Results 30
5.1. Study 1: HINTS 5 Cycle 1 (2017) 30
5.2. Study 1: HINTS 5 Cycle 3 (2019) 33
5.3. Study 2: Full Model with Behavioral Intention 36
5.4. Study 2: Full Model with Additional Moderators 37
5.4.1. Encouraged EHR Use by Health Care Providers as Moderator 37
5.4.2. Patients Health Literacy as Moderator 39
5.4.3. Patients Chronic Disease History as Moderator 41
6. Discussion 43
6.1. Implications of the Current Research 43
6.2. Limitations and Future Suggestions 46
References 51
Abstract in Korean 58
List of Tables
Table 1. Descriptive Statistics of Control Variables 20
Table 2. Comparison of Model Fit Measures between Measurement Models: Cycle 1 31
Table 3. Comparison of Model Fit Measures between Measurement Models: Cycle 3 34
List of Figures
Figure 1. Conceptual Framework Based on TAM 11
Figure 2. SEM with Measurement Model of HINTS 5 Cycle 1 and 3 15
Figure 3. Full Model with Behavioral Intention for Cross-Sectional Survey 16
Figure 4. SEM Results: HINTS 5 Cycle 1 32
Figure 5. SEM Results: HINTS 5 Cycle 3 35
Figure 6. Path Analysis for Full Model with Behavioral Intention 37
Figure 7. Encouraged EHR Use by Health Care Providers as Moderator 38
Figure 7.1. Simple Slope Analysis for Interaction Effect 39
Figure 8. Patients Health Literacy as Moderator 41
Figure 9. Patients Chronic Disease History as Moderator 42석
유한 차분 방식 및 자동 미분 기반의 현악기 소리 합성을 위한 물리 모델링
학위논문(박사) -- 서울대학교 대학원 : 융합과학기술대학원 지능정보융합학과, 2024. 8. 이교구.The goal of sound synthesis technology is to generate natural sounds using a computer based on given input conditions. Sound synthesis methods are broadly divided into two types based on whether or not they reflect the physical properties that produce the sound. This thesis focuses on techniques that synthesize sound by modeling the physical mechanisms that produce musical sounds in the real world, such as musical instruments, on a computer. These instrumental sound synthesis systems not only significantly reduce the cost of music production through realistic virtual instrument implementations, but also allow for creative system design. The technique derives the wave equation, which is the physical governing equation for making sound from an instrument, imposes boundary conditions and initial conditions specific to that instrument, and then uses a computer to find the solution of this equation. The difficulty with obtaining the solution is that it requires discretization in time and space, and to reproduce a sufficiently realistic sound requires a high amount of computation in the space-time axis. Simplification or linearization of the model in space can reduce the amount of computation, but often does not fully reflect the nonlinear nature of stringed instruments. This paper proposes to leverage artificial neural networks to improve the physical modeling of stringed instruments. To this end, we propose and verify a finite difference method accelerated by a graphics processing unit. Also, we propose an artificial neural network that efficiently improves the process of obtaining these numerical solutions through supervised learning and to overcome the limitations of the spatial resolution of the numerical solutions through unsupervised learning.소리 합성 기술의 목표는 주어진 입력 조건에 기반하여 컴퓨터를 사용하여 자연스러운 소리를 생성하는 것입니다. 소리 합성 방법은 소리를 생성하는 물리적 특성을 반영하는지 여부에 따라 크게 두 가지 유형으로 나뉩니다. 본 논문은 악기 등 현실에서 음악적인 소리를 만드는 물리적 메커니즘을 컴퓨터에 모델링하여 소리를 합성하는 기술에 중점을 둡니다. 이러한 악기 소리 합성 시스템은 현실적인 가상 악기 구현을 통해 음악 제작 비용을 크게 줄일뿐만 아니라 창의적인 시스템 설계를 가능케 합니다. 이 기술은 악기에서 소리를 만들기 위한 물리적인 지배 방정식인 파동 방정식을 유도하고 해당 악기에 따라 경계 조건 및 초기 조건을 부여한 후 컴퓨터를 통해 이 방정식의 해를 찾습니다. 해를 얻기 위해서는 시간과 공간에 대한 이산화가 필요하며, 충분히 현실적인 소리를 재현하려면 시공간 축에서 높은 계산량이 필요하다는 어려움이 따릅니다. 공간에 대한 모델의 단순화 또는 선형화는 연산량을 줄일 수 있지만 종종 현악기의 비선형적인 특성을 충분히 반영하지 못합니다. 본 논문은 인공 신경망을 활용하여 현악기의 물리적 모델링을 개선시키는 방법을 제안합니다. 이를 위해 그래픽 처리 장치로 가속화된 유한 차분법을 제안하고, 구해진 시뮬레이션의 유효성을 검증합니다. 나아가, 제안한 시뮬레이터를 기반으로 인공신경망을 학습시키며, 지도학습을 통해 이러한 수치해석적 해를 얻는 과정을 효율적으로 개선시키고 비지도학습을 통해 수치해석적 해의 공간해상도에 관한 한계를 극복하는 방법을 제안합니다.Abstract i
Contents ii
List of Tables v
List of Figures vi
List of Notations xiii
1 Introduction 1
1.1 Motivation 3
1.2 Key Ideas 5
1.3 Tasks of Interest 7
1.4 Contribution 7
2 Theoretical Background 10
2.1 Finite Difference Scheme 10
2.1.1 General Planar Nonlinear String Vibration 14
2.1.2 Numerical Schemes for Planar Nonlinear String 16
2.1.3 Least Squares Problem 23
2.2 Verification and Validation in Scientific Computing 25
2.2.1 Code Verification 27
2.2.2 Solution Verification 31
2.3 Spectral Modeling Synthesis 36
2.4 Physics-informed Neural Network 38
3 Verification of a Finite Difference Scheme for a Planar Damped Stiff String Simulation 42
3.1 Introduction 42
3.2 Numerical Solution of a Planar Motion 45
3.2.1 Nonlinear String Vibration 45
3.2.2 Finite Difference Scheme 47
3.2.3 Analysis on the Numerical Solution 57
3.3 Code Verification 63
3.3.1 Analytical Solution of a One-dimensional Motion 64
3.3.2 Asymptotic Solution based on Richardson Extrapolation 72
3.3.3 Results of Code Verification 73
3.4 Solution Verification 76
3.4.1 Solution Verification of Planar String Vibration 77
3.4.2 Results and Discussion 78
3.5 Conclusion 81
4 Physical Modeling of String Instruments using Automatic Differentiation 88
4.1 Introduction 88
4.2 String Sound Synthesis in Supervised Learning 90
4.2.1 Background on Physical Modeling of Musical String Instrument 90
4.2.2 Differentiable Modal Synthesis for Physical Modeling (DMSP) 94
4.2.3 Experiments 99
4.2.4 Results and Discussion 102
4.3 Unsupervised Optimization using Physics-informed Neural Network 108
4.3.1 Solution Refinement 109
4.3.2 Implementation Details 110
4.3.3 Experiments and Results 111
4.4 Conclusion 111
5 Conclusion 115
5.1 Thesis Summary 115
5.2 Limitations and Future Works 116
5.2.1 Comparison with Real-recorded Sounds 116
5.2.2 Subjective Evaluation 120
5.2.3 Extensions for Simulating More Diverse Systems 121
5.2.4 Blind Estimation of the PDE Parameters 122
5.3 Final Remarks 122
Bibliography 135
초록 136
감사의글 137
Acknowledgment 141박