22455 research outputs found
Sort by
Adaptive exploration of intrinsic data properties for clustering, outlier detection, and dimensionality reduction
In the data-driven era, data has become a key element in supporting decision-making, scientific research, and technological innovation. As the scale and complexity of data continue to grow, it becomes increasingly important to extract actionable knowledge from it. Given this background, data mining is essential for discovering interesting patterns from large datasets, especially in unsupervised learning tasks such as clustering, outlier detection, and dimensionality reduction. However, existing methods often face critical challenges, including adapting to diverse and arbitrary data distributions, handling datasets with varying densities, and reducing dependence on dataset-specific parameters. Since each dataset exhibits unique intrinsic properties, such as local density and neighborhood relationships, this thesis focuses on developing an adaptive exploration framework for unsupervised learning by leveraging intrinsic data properties to enhance the adaptability, accuracy, and efficiency of unsupervised learning algorithms.
For the clustering task, this thesis proposes the DBADV algorithm. Existing density-based methods, such as DBSCAN, although capable of recognizing clusters of arbitrary shapes and sizes, are ineffective when dealing with density variations. DBADV calculates the local density information of each object using perplexity, which not only reflects the individual properties of each object but also describes the density distribution of clusters and finds the adaptive search range of each object by collecting information from its neighbors. In addition, a new metric is designed to obtain the mutual nearest neighbors of each object to better distinguish objects near the cluster boundaries, significantly improving the clustering accuracy in the presence of noise and outliers.
For the outlier detection task, this thesis proposes the ADOD algorithm. Existing proximity-based methods rely on either a globally fixed radius or a fixed number of neighbors as a parameter, both of which are prone to misjudgments when there is significant variation in data density. ADOD uses perplexity to calculate the local scale of each object and dynamically adjusts the neighborhood boundaries based on this scale to adapt to data with varying densities. Meanwhile, ADOD designs a density consistency score that determines the outlier score by calculating the local density difference between an object and its mutual neighbors, effectively identifying outliers that significantly deviate from their surroundings. Additionally, ADOD extends its utility in real-time applications by generalizing to unknown data through comparison with known data.
For the dimensionality reduction task, this thesis proposes the DynoGraph algorithm. DynoGraph addresses two main issues of the existing graph-based methods like t-SNE and UMAP: how to construct a good graph and how to maintain the similarity structure of high-dimensional data in low-dimensional space. DynoGraph first develops an adaptive neighborhood graph construction method that accurately captures the intrinsic geometry of the high-dimensional data. Then, for the first time, it introduces a dynamic graph modification process during dimensionality reduction, guaranteeing that the data structure in the low-dimensional space faithfully reflects the high-dimensional data. Moreover, DynoGraph sets an adaptive threshold that is automatically adjusted according to the intrinsic structure of data, thus guiding the insertion and deletion of edges. These adjustments help to update their positions in subsequent embeddings, aligning them with the high-dimensional data.
In summary, this thesis demonstrates how three adaptive algorithms can effectively leverage intrinsic data properties to adapt to arbitrary distributions, capture local density, and design self-adaptive parameters, thereby improving the performance of unsupervised learning tasks. This adaptive exploration framework not only enriches the theoretical approaches to data analysis, but also provides new perspectives and effective solutions for handling complex data in real-world scenarios.Im Zeitalter von datengetriebenen Technologien sind Daten zu einem Schlüsselelement geworden, um Entscheidungsfindung, wissenschaftliche Forschung und technologische Innovation zu unterstützen. Mit dem kontinuierlichen Wachstum der Datenmengen und ihrer Komplexität wird es zunehmend wichtiger, verwertbares Wissen aus den Daten zu extrahieren. Vor diesem Hintergrund spielt Data Mining eine entscheidende Rolle bei der Entdeckung interessanter Muster in großen Datensätzen, insbesondere bei Aufgaben des unüberwachten Lernens wie der Clusteranalyse, Ausreißererkennung und Dimensionsreduktion. Bestehende Methoden stehen jedoch oft vor großen Herausforderungen, insbesondere in Anbetracht der Anpassung an vielfältige und beliebige Datenverteilungen, dem Umgang mit Datensätzen unterschiedlicher Dichte und der Reduzierung der Abhängigkeit von datensatzspezifischen Parametern. Da jeder Datensatz einzigartige intrinsische Eigenschaften, wie lokale Dichte und Nachbarschaftsbeziehungen aufweist, konzentriert sich diese Dissertation darauf, ein Rahmenwerk für die adaptive Exploration im Bereich des unüberwachten Lernens zu entwickeln, das diese intrinsischen Eigenschaften nutzt, um die Anpassungsfähigkeit, Genauigkeit und Effizienz von Algorithmen im Kontext des unüberwachten Lernens zu verbessern.
Für die Clusteranalyse schlägt diese Dissertation den Algorithmus DBADV vor. Bestehende dichtebasierte Methoden wie DBSCAN, die zwar in der Lage sind, Cluster beliebiger Formen und Größen zu erkennen, sind bei der Handhabung von Dichtevariationen ineffektiv. DBADV berechnet die lokale Dichteinformation jedes Objekts mithilfe der Perplexität, die nicht nur die individuellen Eigenschaften jedes Objekts widerspiegelt, sondern auch die Dichteverteilung der Cluster beschreibt. Dabei wird der adaptive Suchbereich jedes Objekts durch das Sammeln von Informationen seiner Nachbarn ermittelt. Zusätzlich wird eine neue Metrik entwickelt, um die gegenseitigen nächsten Nachbarn jedes Objekts zu bestimmen, was es ermöglicht, Objekte in der Nähe von Clustergrenzen besser zu unterscheiden und die Clustering-Genauigkeit in Situationen erheblich zu verbessern, in denen Rauschen und Ausreißer vorhanden sind.
Für die Ausreißererkennung schlägt diese Dissertation den Algorithmus ADOD vor. Bestehende nahebasierte Methoden verlassen sich entweder auf einen global festgelegten Radius oder eine feste Anzahl von Nachbarn als Parameter, die bei signifikanten Variationen in der Datendichte anfällig für Fehlurteile sind. ADOD verwendet Perplexität, um das lokale Ausmaß jedes Objekts zu berechnen, und ermittelt die Nachbarschaftsgrenzen dynamisch auf der Grundlage dieses Ausmaßes, um sich an Daten mit unterschiedlichen Dichten anzupassen. Darüber hinaus entwickelt ADOD eine Dichtekonsistenzbewertung, die den Ausreißerwert durch Berechnung des lokalen Dichteunterschieds zwischen einem Objekt und seinen gegenseitigen Nachbarn bestimmt und so effektiv Ausreißer identifiziert, die erheblich von ihrer Umgebung abweichen. Zusätzlich erweitert ADOD seine Anwendbarkeit in Echtzeitszenarien, indem es unbekannte Daten durch einen Vergleich mit bekannten Daten generalisiert.
Für die Dimensionsreduktion schlägt diese Dissertation den Algorithmus DynoGraph vor. DynoGraph adressiert zwei Hauptprobleme bestehender graphenbasierter Methoden wie t-SNE und UMAP: Wie werden qualitativ hochwertige Graphen konstruiert und wie werden Ähnlichkeitsstrukturen hochdimensionaler Daten in einem niedrigdimensionalen Raum beibehalten. DynoGraph entwickelt zunächst eine adaptive Methode zur Konstruktion eines Nachbarschaftsgraphen, die die intrinsische Geometrie hochdimensionaler Daten präzise erfasst. DynoGraph führt erstmals einen anschließenden dynamischen Graphmodifikationsprozess während der Dimensionsreduktion ein, der sicherstellt, dass die Datenstruktur im niedrigdimensionalen Raum die hochdimensionalen Daten originalgetreu widerspiegelt. Darüber hinaus legt DynoGraph einen adaptiven Schwellenwert fest, der automatisch entsprechend der intrinsischen Datenstruktur angepasst wird, um das Einfügen und Löschen von Kanten zu steuern. Diese Anpassungen tragen dazu bei, die Positionen der Daten in den nachfolgenden Einbettungen zu aktualisieren und sie mit den hochdimensionalen Daten in Einklang zu bringen.
Zusammenfassend zeigt diese Dissertation, wie drei adaptive Algorithmen die intrinsischen Eigenschaften von Daten effektiv nutzen können, um sich an beliebige Verteilungen anzupassen, lokale Dichte zu erfassen und selbstadaptive Parameter zu entwerfen, wodurch die Leistung bzgl. Aufgaben des unüberwachten Lernens verbessert wird. Dieses Framework für adaptive Exploration bereichert nicht nur die theoretischen Ansätze zur Datenanalyse, sondern bietet auch neue Perspektiven und effektive Lösungen für den Umgang mit komplexen Daten in realen Anwendungen
Einsatz der Einzelhaaranalytik in der forensischen Toxikologie
Die Analyse von Drogen- und Medikamentenwirkstoffen in Haarproben ist ein wichtiger Teilbereich der forensischen Toxikologie zur retrospektiven Überprüfung einer Abstinenz oder zum Nachweis eines Konsums. Durch eine segmentweise Untersuchung kann der zu überprüfende Aufnahmezeitraum näher eingegrenzt werden, wobei aufgrund verschiedener Einflussfaktoren das zeitliche Auflösungsvermögen nur im Bereich von etwa einem Monat liegt. Durch die Untersuchung von Einzelhaaren nach Mikrosegmentierung kann die zeitliche Auflösung deutlich erhöht werden. Die Anwendbarkeit und der zusätzliche Nutzen dieses Verfahrens bei der Begutachtung forensisch-toxikologischer Fragestellungen wurde im Rahmen dieser wissenschaftlichen Arbeit untersucht. Die entwickelten Methoden und erhaltenen Ergebnisse stellen wichtige Grundlagen und Erkenntnisse für die Bearbeitung und Interpretation zukünftiger Fälle mit unterschiedlichsten Fragestellungen dar.The analysis of illicit drugs and pharmaceuticals in hair samples is an important part of forensic toxicology for the retrospective verification of abstinence or evidence of consumption. By performing segmental hair analysis, the time period of an alleged drug use can more closely be specified. Nevertheless, due to various influencing factors, the temporal resolution is restricted to about one month but can be significantly increased through micro-segmental single hair analysis. The applicability and additional benefits of this procedure in the assessment of forensic toxicological issues were examined as part of this scientific work. The methods developed and the results obtained represent important principles and information for processing and interpretation of future cases with a wide variety of issues
Alkoholbezogene Störungen und Behandlungspfade von Patienten mit einer Alkoholkonsumstörung im Bundesland Bremen
Hintergrund: Alkoholkonsum ist bedingt durch unterschiedliche Wirkmechanismen Ursache für verschiedene Krankheiten, gleichzeitig gilt Deutschland im internationalen Vergleich als Hochkonsumland. Die Prävalenz für eine Alkoholabhängigkeit nach Kriterien der vierten Auflage des diagnostischen und statistischen Manuals psychischer Störungen (DSM-IV) lag in der erwachsenen, deutschen Bevölkerung zwischen 18 und 64 Jahren im Jahr 2018 bei 3,1%. Trotz vorhandener evidenzbasierter Verfahren der Früherkennung, adäquater Diagnostik und Behandlung von alkoholbezogenen Störungen zeigen sich in der Praxis oft Probleme. Bisherige Forschungsarbeiten weisen auf eine geringe Inanspruchnahme suchtspezifischer Behandlungen von Menschen mit einer Alkoholabhängigkeit hin. In dieser Arbeit wird diese Behandlungslücke mithilfe von verknüpften Routinedaten verschiedener Kosten- und Leistungsträger aus dem Bundesland Bremen näher beleuchtet.
Ziele: Das übergeordnete Ziel der Dissertation ist die Analyse der Versorgungsituation von Personen mit alkoholbezogenen Störungen insbesondere einer Alkoholabhängigkeit, um Versorgungslücken und Schnittstellenprobleme identifizieren zu können. Das erste Forschungsziel ist die suchtspezifische Versorgungssituation von Menschen mit alkoholbezogenen Störungen in Bremen in verschiedenen medizinischen Settings darzustellen und die Inanspruchnahme von suchtspezifischen Behandlungen von Menschen mit Alkoholabhängigkeit mithilfe von Umfragedaten auf die Gesamtbevölkerung in Bremen hochzurechnen. Das zweite Forschungsziel besteht in der Analyse der individuellen suchtspezifischen Behandlungspfade von Menschen mit Alkoholabhängigkeit. Diese Pfade wurden in typische Verläufe geclustert und diese Cluster hinsichtlich ihrer Konformität mit den aktuellen S3-Leitlinien überprüft sowie im Hinblick auf sozidemographische Variablen charakterisiert.
Methodik: Datengrundlage sind Routinedaten der Jahre 2016 und 2017 von Personen mit mind. einer alkoholbezogenen Diagnose und Wohnsitz in Bremen (n=11.205). Die Daten beinhalten ambulante und stationären Daten zweier gesetzlicher Krankenkassen (AOK Bremen/Bremerhaven, hkk), alkoholbezogene Rehabilitationsmaßnahmen der regionalen Deutschen Rentenversicherung Oldenburg-Bremen sowie Besuche der ambulanten Suchthilfe des kommunalen Klinikverbunds Bremen Gesundheit-Nord. Die Anzahl der Personen mit in Anspruch genommenen suchtspezifischen Maßnahmen, wie einem stationären qualifizierten Entzug, ambulanter medikamentöser Rückfallprophylaxe, Rehabilitationsmaßnahmen, ambulanter Suchthilfe sowie mit dokumentieren ambulanten und stationären Diagnosen einer Alkoholabhängigkeit wurde auf die Bevölkerung Bremens hochgerechnet. Für die Berechnung der Behandlungsquote wurde darüber hinaus die Anzahl aller Personen mit einer Alkoholabhängigkeit in Bremen mithilfe von Daten des Epidemiologischen Suchtsurveys 2018 alters- und geschlechtsspezifisch auf die Bevölkerung Bremens hochgerechnet. Für das zweite Forschungsvorhaben wurden individuelle suchtspezifische Behandlungspfade von 518 Personen mit Methoden der Sequenzdatenanalyse erstellt, geclustert und analysiert. Die Behandlungspfade beinhalten suchtspezifische Maßnahmen 10 Monate nach einer stationären Indexepisode aufgrund von Alkoholabhängigkeit. Im Anschluss wurden soziodemografische Merkmale und Unterschiede in der Behandlung mithilfe eines Multinomialen Logit Modells für die Clusterzugehörigkeit analysiert.
Ergebnisse: Die erste Studie konnte zeigen, dass im Bundesland Bremen mehr als die Hälfte der Personen mit einer Alkoholabhängigkeit bereits im Gesundheitssystem dokumentiert war, größtenteils im ambulanten Setting. Suchtspezifische Maßnahmen nahmen jedoch nur 11% in Anspruch. Am häufigsten wurde ein qualifizierter Entzug begonnen (4,7%) oder die ambulante Suchthilfe besucht (4,3%). Eine Postakutbehandlung wurde seltener in Anspruch genommen (0,8% medikamentöse Rückfallprophylaxe und 3,9% Rehabilitation). Unter jüngeren und männlichen Personen war die Behandlungsquote am höchsten. Im Rahmen der zweiten Studie konnten vier Cluster suchtspezifischer Behandlungspfade basierend auf den in Anspruch genommenen Postakutbehandlungen identifiziert werden. Die Mehrheit nahm nach der stationären Episode keine weiteren suchtspezifischen Maßnahmen in Anspruch (n=276) oder keine Postakutbehandlungen (n=205). Nur die wenigsten begannen eine Postakutbehandlung, wie Rehabilitation (n=26) oder medikamentöse Rückfallprophylaxe (n=11).
Schlussfolgerung: Die vorliegenden Analysen der Routinedaten zeigen, dass die Mehrheit der Personen mit Alkoholabhängigkeit zwar, meist im ambulanten Setting, diagnostiziert werden, jedoch nur eine Minderheit die betrachteten suchtspezifischen Maßnahmen in Anspruch nimmt. Eine Stärkung der primärärztlichen Versorgung von Menschen mit alkoholbezogenen Störungen erscheint angebracht, vor allem in Bezug auf die Weitervermittlung in eine suchtspezifische Behandlung. Die Analyse der Behandlungspfade zeigt eine sehr niedrige Inanspruchnahme von Postakutbehandlungen nach einer stationären Episode (Rehabilitation und/oder medikamentöse Rückfallprophylaxe). Zukünftige Forschung sollten alternative Postakutmaßnahmen im ambulanten Setting und der Selbsthilfe berücksichtigen. Eine generelle Verbesserung der Weitervermittlung in Postakutmaßnahmen nach einem Entzug erscheint angebracht.Background: Alcohol consumption causes various diseases through different mechanisms of action, with Germany showing high rates of consumption by international comparison. The prevalence of alcohol dependence alone, according to the criteria of the fourth edition of the Diagnostic and Statistical Manual of Mental Disorders (DSM-IV), was 3.1% in the adult German population between the ages of 18 and 64 in 2018. Despite existing evidence-based procedures regarding the early detection, adequate diagnosis and treatment of alcohol-related disorders, problems often arise in practice. Previous research indicates a large discrepancy between the prevalence of people with alcohol dependence and the rates of addiction specific treatment. In this paper, this treatment gap is examined in more detail with the help of linked routine data from various cost and service providers in the German federal state of Bremen.
Objectives: The overall aim of this dissertation is to describe the care situation of people with alcohol-related disorders especially alcohol dependence to identify any gaps in care and networking problems. The first research objective was to describe the addiction-specific care situation of people with alcohol-related disorders in Bremen in various medical settings and to extrapolate the treatment rate of people with alcohol dependence to the overall population in Bremen using survey data. The second research objective comprised the analysis of individual addiction-specific care pathways. These pathways were clustered into typical pathways and then examined regarding their conformity with current guidelines and characterized with socio-demographic variables.
Methods: The data source was routine data from 2016 and 2017 from people with at least one alcohol-related diagnosis and a residence in Bremen (n=11,205). The data included outpatient and inpatient data from two statutory health insurance companies (AOK Bremen/Bremerhaven, hkk), alcohol-related rehabilitation treatments of the regional German Pension Insurance Oldenburg-Bremen and outpatient addiction care services of the municipal clinic association Bremen Gesundheit-Nord. Persons using addiction-specific treatment and care services, such as inpatient qualified withdrawal, outpatient pharmacotherapy as relapse prevention, rehabilitation, outpatient addiction care, and with documented outpatient and inpatient diagnoses of alcohol addiction were extrapolated to the total population in Bremen. To calculate the treatment rate, also the number of all people with an alcohol addiction in Bremen was extrapolated to the population of Bremen on an age- and gender-specific basis using data from the 2018 Epidemiological Survey of Substance Abuse. For the second research project, individual care pathways of 518 persons were created, clustered, and analyzed using a state sequence analysis. The care pathways consist of addiction-specific treatments 10 months after an inpatient index episode due to alcohol dependence. Subsequently, socio-demographic characteristics and differences in treatment were analyzed using a multinomial logit model for the cluster assignments.
Results: The first study showed that in the federal state of Bremen, more than half of the people with an alcohol addiction are already documented in the healthcare system, mostly in an outpatient setting, but only 11% make use of addiction-specific treatment and care services. The most common services were qualified withdrawal (4.7%) or outpatient addiction care (4.3%). Post-acute treatment was used less frequently (0.8% pharmacotherapy as relapse prevention and 3.9% rehabilitation). The treatment rate is highest among younger and male persons. The second study identified four clusters based on the post-acute treatment utilized after an inpatient episode of alcohol dependence. The majority of persons used no further addiction-specific treatment after the inpatient episode (n=276) or no post-acute treatment (n=205). The fewest started a post-acute treatment such as rehabilitation (n=26) or pharmacotherapy as relapse prevention (n=11).
Conclusion: The analyses of the routine data show that, although most people with alcohol dependence have been diagnosed, mostly in an outpatient setting, only a minority make use of the addiction-specific treatments under consideration. Therefore, a strengthening of primary medical care for alcohol treatment seems appropriate, especially regarding referral to addiction-specific treatment. The analysis of care pathways also shows a very low utilization of post-acute treatment after an inpatient episode (rehabilitation and/or pharmacotherapy as relapse prevention). Future research should focus more on alternative post-acute treatments in the outpatient setting and self-help. A general improvement in the referral to post-acute treatment after withdrawal seems appropriate
Subphenotyping of women after gestational diabetes mellitus identifies subjects at high and low risk for progression to prediabetes/type 2 diabetes mellitus
Immunhistochemische Untersuchungen zur möglichen Pathogenese der granulären Parakeratose
Die granuläre Parakeratose (GP) stellt ein histologisches Phänomen dar, das bisher wenig erforscht ist. Weder Pathogenese, noch klinische Einordnung sind geklärt.
Die Ergebnisse der immunhistochemischen Antikörperfärbungen für Caspase-14, Voltage-gated Calcium channel (VGCC), Desmocollin-1 (DSC-1) und Corneodesmosin (CDSN) stützen die These, dass Okklusion und eine hohe Umgebungsfeuchtigkeit Provokationsfaktoren darstellen. Das feuchte Milieu scheint die Produktion und Ausschüttung von lamellar bodies (LB) am Übergang vom Stratum granulosum (SG) zum Stratum corneum (SC) zu vermindern. Eine reduzierte Expression von VGCC könnte einerseits die extrazelluläre Calcium-Konzentration erhöhen, was die LB-Ausschüttung hemmt und andererseits einen niedrigen intrazellulären Calcium-Gehalt bewirken, sodass Calcium-abhängige Prozesse behindert werden, darunter die Ausschüttung von Keratohyalingranula und der (Pro-)Filaggrin-Abbau.
Die gehemmte LB-Ausschüttung von CDSN resultiert in ausbleibender Bildung von Corneodesmosomen. Die ebenfalls retinierten Kallikreine können die Desmoglea (DSC-1 und Desmoglein-1 (DSG-1)) im Verlauf des SC nicht abbauen, sodass eine Hyperkeratose entsteht.
Intrazellulär erfolgt am Übergang SG/SC keine Degranulierung von Keratohyalingranula, sodass Profilaggrin und Procaspase-14 nicht im Zytosol wirken können. Daher bleiben Zellkompaktierung und Zellkernabbau aus. Dies hat eine Hyperparakeratose mit granulärem Aspekt und einen Verbleib des Nukleus zur Folge
Online-gestützte Umfrage zur Antibiotikaanwendung durch Tierbesitzer*innen bei Hund und Katze
Data mining techniques for graph and hypergraph analysis
Graph structures are essential for modeling pairwise relationships in systems ranging from social networks to biological interactions and transportation infrastructure. However, in many real-world scenarios, relationships are often beyond pairwise. For example, social networks generally feature group structures where individuals belong to multiple groups simultaneously. The hypergraph structure has been extensively considered to model such higher-order relationships, wherein hyperedges can connect an arbitrary number of nodes. This thesis focuses on developing scalable and interpretable data mining algorithms for graph and hypergraph analysis, advancing techniques to handle complex relational patterns in these networks.
We explore information diffusion in hypergraphs and study the information coverage maximization problem in this scenario. Traditional information diffusion models are designed primarily for ordinary graphs. To address this limitation, we propose HIC, the Hypergraph Independent Cascade model, which extends the conventional independent cascade model to accommodate hypergraphs. Building on HIC, we propose a novel influence maximization problem: the information coverage maximization problem in hypergraphs. Unlike traditional influence maximization, which focuses on identifying influential nodes, we target to identify key groups. We establish the NP-hardness of this problem and demonstrate the submodular monotonicity of the information spread function. To solve the problem efficiently, we developed a heuristic approach called InfDis, inspired by the Degree Discount algorithm. Extensive experiments validate the effectiveness and efficiency of this approach.
The second task addressed in the thesis is hyperlink prediction, which involves predicting interactions among multiple entities. While existing solutions generally operate on the entire hypergraph, we propose the first subgraph-based hyperlink prediction approach that captures localized characteristics of central hyperedges while mitigating scalability concerns. The proposed method, SSF, focuses on localized subgraph patterns and extracts interpretable features using structural heuristics such as walks and loops.Additionally, its edge-weakening scheme adapts to varying hypergraph densities, enabling fine-grained feature learning. We conduct extensive experiments to validate SSF's adaptive capacity, evaluate the effectiveness of its feature components, and assess its robustness across various parameter configurations.
Next, we focus on a fundamental problem—graph classification. For this problem, we propose RWF, a graph fingerprinting technique that combines structural role-based vertex partitioning with local connection strength measurement. By creating soft alignments of node subsets across graphs of varying sizes, RWF generates topology-aware fingerprints that capture intra- and inter-subset connectivity. Further, RWF supports the integration of node attributes to enhance classification performance. Empirical assessment encompassing a wide range of graph datasets demonstrates that RWF achieves high computational efficiency while maintaining robust classification accuracy.
Collectively, this thesis introduces a novel problem formulation and presents three scalable, interpretable techniques designed to address key challenges in graph and hypergraph analysis.Graphstrukturen sind entscheidend für die Modellierung paarweiser Beziehungen in Systemen, die von sozialen Netzwerken über biologische Interaktionen bis hin zu Verkehrsinfrastrukturen reichen. In vielen realen Szenarien gehen Beziehungen jedoch oft über Paare hinaus. So weisen soziale Netzwerke beispielsweise Gruppenstrukturen auf, in denen Individuen gleichzeitig mehreren Gruppen angehören. Die Hypergraph-Struktur wird häufig genutzt, um solche höherstufigen Beziehungen abzubilden, wobei Hyperkanten eine beliebige Anzahl von Knoten verbinden können. Diese Arbeit konzentriert sich auf die Entwicklung skalierbarer und interpretierbarer Data-Mining-Algorithmen zur Analyse von Graphen und Hypergraphen, um Techniken für den Umgang mit komplexen relationalen Mustern in diesen Netzwerken voranzutreiben.
Wir untersuchen die Informationsverbreitung in Hypergraphen und adressieren das Problem der Maximierung der Informationsabdeckung. Herkömmliche Diffusionsmodelle sind primär für Standardgraphen konzipiert. Um diese Einschränkung zu überwinden, schlagen wir HIC (Hypergraph Independent Cascade) vor, das das klassische unabhängige Kaskadenmodell auf Hypergraphen erweitert. Basierend auf HIC formulieren wir ein neuartiges Einflussmaximierungsproblem: die Maximierung der Informationsabdeckung in Hypergraphen. Im Gegensatz zur traditionellen Einflussmaximierung, die einflussreiche Knoten identifiziert, zielen wir auf Schlüsselgruppen ab. Wir beweisen die NP-Härte dieses Problems und zeigen die submodulare Monotonie der Informationsausbreitungsfunktion. Zur effizienten Lösung entwickelten wir einen heuristischen Ansatz namens InfDis, inspiriert vom Degree-Discount-Algorithmus. Umfangreiche Experimente bestätigen die Wirksamkeit und Effizienz dieses Ansatzes.
Die zweite untersuchte Aufgabe ist die Hyperlink-Vorhersage, die die Prognose von Interaktionen zwischen mehreren Entitäten betrifft. Während bestehende Lösungen üblicherweise den gesamten Hypergraphen analysieren, schlagen wir den ersten teilgraphbasierten Ansatz vor, der lokalisierte Merkmale zentraler Hyperkanten erfasst und gleichzeitig Skalierbarkeitsprobleme mindert. Die Methode SSF konzentriert sich auf lokale Teilgraphmuster und extrahiert interpretierbare Merkmale mittels struktureller Heuristiken wie Pfaden und Schleifen. Ein integriertes Kantenabschwächungsschema passt sich variierenden Hypergraphdichten an und ermöglicht feingranulare Merkmalslernprozesse. Experimente validieren SSFs Adaptionsfähigkeit, die Effektivität seiner Merkmalskomponenten und seine Robustheit unter verschiedenen Parameterkonfigurationen.
Im dritten Schwerpunkt behandeln wir die Grundaufgabe der Graphklassifizierung. Hierfür entwickeln wir RWF, eine Graph-Fingerabdrucktechnik, die strukturell rollenbasierte Knotenpartitionierung mit lokalen Verbindungsstärkemessungen kombiniert. Durch weiche Ausrichtungen von Knotenteilmengen über Graphen variierender Größe erzeugt RWF topologiebewusste Fingerabdrücke, die intra- und intersubset-Konnektivität erfassen. Zudem unterstützt RWF die Integration von Knotenattributen zur Steigerung der Klassifizierungsleistung. Empirische Auswertungen über diverse Graphdatensätze zeigen, dass RWF hohe Recheneffizienz bei robusten Klassifizierungsgenauigkeiten erreicht.
Zusammenfassend führt diese Arbeit eine neuartige Problemformulierung ein und präsentiert drei skalierbare, interpretierbare Techniken zur Lösung zentraler Herausforderungen in der Graph- und Hypergraphenanalyse
Addressing uncertainty and complex data structures through Bayesian and classical approaches
Answering questions based on real-world data can pose considerable challenges to analysts. It often requires the use of data that are of questionable quality, exhibit high uncertainty, and may originate from multiple sources. Such data bear a high degree of complexity in their structure with respect to the underlying data-generating process. This cumulative thesis aims to address these issues in the context of selected research areas.
The thesis is divided into two parts. The first part introduces the necessary methodology. The second part presents the four contributing articles. The methodological part provides an introduction chapter to statistical inference and probabilistic modeling by presenting Bayesian inference using Markov chain Monte Carlo as a general approach, and the generalized linear model (GLM) as a classical statistical method. Furthermore, the first part provides three chapters of methodological background in selected areas of research.
The first area to be discussed is infectious disease modeling. The focus is on time-shifting operations that can be used to combine information from multiple time series. This lays the foundation for the first two contributions, which employ a Bayesian hierarchical approach to infectious disease modeling in the context of COVID-19 data.
Next, an overview of measurement error theory is presented, followed by a discussion on how the Bayesian approach addresses these challenges. The third contribution demonstrates the flexibility of the Bayesian approach by applying it to data from the Wismut cohort, which presents considerable complexity and requires the use of multiple measurement error models.
Finally, the last discussed chapter delves into the field of federated learning and privacy-preserving methods. The fourth contribution builds on the presented methodological background to develop an algorithm that is able to validate learned classification models through a GLM-based formulation of the ROC curve. An underlying theme of this thesis is the notion of uncertainty. In the Bayesian approach, uncertainty is encoded through the formulation of prior distributions and the overall probabilistic model, which inherently propagates and quantifies the uncertainty in a posterior distribution. The fourth contribution leverages the concept of uncertainty to preserve individual privacy by adding calibrated noise
Imaging error reduction for MR guided radiotherapy with deep learning-based intra-frame motion compensation
Radiotherapy in the presence of intra-fractional motion can significantly benefit from real-time magnetic resonance imaging (MRI) guidance, owing to its superior soft tissue contrast and the absence of ionizing radiation. However, motion-related imaging errors have been identified as the primary contributor to overall loop latency in MR guided radiotherapy (MRgRT), leading to residual geometric tracking errors and subsequently affecting the effectiveness of active motion management. This thesis explores the feasibility of reducing these errors in MRgRT through deep learning-based intra-frame motion compensation techniques.
Firstly, a motion-dependent k-space sampling simulation procedure was developed to investigate dynamic MR imaging behavior and motion-related imaging errors. Building upon this, a methodology for intra-frame motion dataset creation and augmentation was proposed, pairing the motion-corrupted data with its real-time ground-truth counterpart, with a primary focus on rapid anatomical changes. Specifically, based on a coarse-to-fine grid-scale representation of patient-specific motion data, 4D MRI digital anthropomorphic phantoms were generated to model lung cancer patients, and a dedicated intra-frame motion model was constructed using a piecewise linear approximation between consecutive control points. Additionally, a motion pattern perturbation scheme was introduced to comprehensively explore potential anatomical structure positions and enhance the diversity of intra-frame motion trajectories.
Secondly, a proof-of-concept study in Cartesian cine-MRI was conducted, demonstrating that UNet models can effectively compensate for intra-frame motion by estimating the final-position image at the end of frame acquisition from motion-corrupted input. Quantitatively, in the testing dataset for gross tumor volume (GTV) contouring, the median Dice similarity coefficient (DSC) increased from 89% to 97%, while the 95th percentile Hausdorff distance (HD95) decreased from 4.1 mm to 1.4 mm. Geometric errors in targets undergoing considerable intra-frame deformations were successfully corrected, exhibiting close agreement with the ground truth in terms of both target shape and position. The saliency maps indicated that the model predominantly focused on the later-acquired k-space components for inference and, correspondingly in the spatial domain, the edges of the moving structures at their real-time final positions.
Thirdly, a proof-of-concept study in radial cine-MRI was conducted, proposing "TransSin-UNet", a novel dual-domain deep learning framework. Within the radial k-space reconstruction window, the long-distance spatial-temporal dependencies among the sinogram representation of the spokes were modeled by a transformer encoder subnetwork, followed by a UNet subnetwork operating in the spatial domain for pixel-level refinement. The network was trained and extensively evaluated across datasets with varying azimuthal radial profile increments. TransSin-UNet required only an additional 4.8 ms per frame for compensation compared to conventional direct image reconstruction using motion-corrupted spokes. It consistently outperformed architectures relying solely on transformer encoders or UNets across all comparative evaluations, leading to a noticeable enhancement in image quality and target positioning accuracy. The normalized root mean squared error (NRMSE) decreased by 50% from the initial average of 0.188, whereas the mean DSC of GTV increased from 85.1% to 96.2% in the investigated testing cases. Furthermore, the ground-truth positions of anatomical structures experiencing substantial deformations were precisely derived.
This work constitutes a substantial advancement toward the clinical implementation of cine-MR tracking error reduction strategies to support enhanced real-time motion management in MRgRT