1,720,959 research outputs found

    A Novel Methodology for a Comprehensive Analysis of Genomic Sequence-to-Graph Alignment Tools

    Get PDF
    Genome graphs have proved to be a more compact and efficient way of representing genetic inter- and intra-individual variability. Although they overcome the traditional sequence-based genome references in many use cases, analyzing genome graphs introduces new computational challenges. The workhorse of graph-based genome analysis is the sequence-to-graph alignment process, which consists of finding the path in the graph that better represents a query sequence. This search is highly computationally intensive, and different solutions have been proposed to solve it efficiently, either by adapting sequence-to-sequence strategies or exploiting novel graph-specific algorithms. However, comparing sequence-to-graph alignment tools is quite challenging because of the complexity and relative novelty of this task, and the resulting lack of standardization. Therefore, here we propose a methodology for a comprehensive and structured comparison of such tools. First, we define a set of KPIs for the qualitative analysis of an aligner's usability, accuracy, and performance. Then, we introduce the first open-source(1) benchmark suite for the quantitative analysis of multiple sequence-to-graph aligners. We test the proposed methodology on state-of-the-art tools, proving how it easily provides valuable insights about the compared aligners. Finally, we conclude the paper by drawing some guidelines to drive the improvement of this promising research field

    A Multimodal Transfer Learning Approach for Histopathology and SR-microCT Low-Data Regimes Image Segmentation

    Get PDF
    Osteocyte-lacunar bone structures are a discerning marker for bone pathophysiology, given their geometric alterations observed during aging and diseases. Deep Learning (DL) image analysis has showcased the potential to comprehend bone health associated with their mechanisms. However, DL examination requires labeled and multimodal datasets, which is arduous with high-dimensional images. Within this context, we propose a method for segmenting osteocytes and lacunae in human bone histopathology and Synchrotron Radiation micro-Computed Tomography (SR-microCT) images, employing a deep U-Net in an intra-domain and multimodal transfer learning setting with a limited number of training images. Our strategy allows achieving 63.92±4.69 and 63.94±4.05 Dice Similarity Coefficient (DSC) osteocytes and lacunae segmentation, while up to 20.38 and 5.86 average DSC improvements over selected baselines even if 44× smaller datasets are employed for training.Clinical relevance - The proposed method analyzes bone histopathologies and SR-microCT images in a multimodal and low-data setting, easing the bone microscale investigations while supporting the study of osteocyte-lacunar pathophysiology

    Bridging Research and Entrepreneurship: An Innovative Educational and Experiential Approach

    No full text
    NECSTLab at Politecnico di Milano is a pioneering research laboratory that integrates cutting-edge academic research with entrepreneurial ventures. Its mission is to bridge the gap between research and real-world applications, encouraging students and researchers to consider the societal impact of their innovations. NECSTLab promotes an interdisciplinary approach, combining technical expertise with business acumen to drive innovation. The lab's strategy includes fostering an entrepreneurial mindset through courses that equip students with the tools to transform research into viable products or services. This paper introduces an innovative approach to address the challenges of transitioning academic research into market commercialization. By focusing on early-stage collaboration between researchers and industry, and incorporating market analysis, prototype development, and business model validation, this process supports the commercialization of research. The proposed pipeline has been implemented and validated within the NECSTLab environment, demonstrating its efficacy in fostering entrepreneurship and translating academic research into successful commercial ventures. This is exemplified by case studies showcasing the pipeline's effectiveness in transforming innovative research into market-ready solutions while fostering entrepreneurial initiatives and driving impactful innovation

    On the optimization of GWFA algorithm: enabling real-case applications supporting alignment backtracking

    Get PDF
    The Human Pangenome Reference Consortium (HPRC) proved that pangenome graphs represent a population's genetic variability more efficiently and accurately than linear references. Graphs can intrinsically encode variations as alternative paths inside a directed set of sequence nodes connected by edges. Despite their higher complexity, graph-based genome analysis pipelines are gaining significant interest, and the first sequence-to-graph aligners have already shown improvements in semi-global alignment. However, in pangenomics studies, the global alignment of long reads is fundamental for identifying structural variations and haplotype phasing. In this context, the Graph Wavefront Alignment (GWFA) algorithm emerged as the fastest strategy for aligning long reads to genomic graphs. However, the available GWFA implementation does not support alignment backtracking, a crucial feature in real-case studies. In this paper, we propose a new open-source1 implementation of the GWFA algorithm that computes and reports the complete traceback in the standard GAF format. Our work achieves a 20× speedup in execution time compared to the state-of-the-art tool GraphAligner and competitive memory usage

    Going Beyond Counting First Authors in Author Co-citation Analysis

    Get PDF
    The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed

    Variations on the Author

    Get PDF
    “Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship

    On the optimization of sequence alignment to cyclic genome graphs

    No full text
    LAUREA MAGISTRALECon la diffusione delle tecnologie di sequenziamento di nuova generazione (NGS) si è verificato un incremento nella quantità di dati genomici generati che ha permesso una serie di nuovi tipi di studi, che coinvolgono materiale genetico proveniente da diversi individui all'interno della stessa specie, famiglia o ambiente. Ciò richiede nuovi paradigmi computazionali per garantire un processo di analisi più intuitivo e rapido. Recentemente, l'utilizzo di grafi per la rappresentazione dei dati genomici é stata proposta come un alternativa più efficiente e flessibile alle stringhe, perché sono adatti a rappresentare la variabilità intra-individuale e inter-individuale. Poiché questo cambiamento di paradigma è attualmente in corso, l'allineamento delle reads ai grafi genomici è un contesto fertile per la ricerca, in quanto offre molte sfide in termini di miglioramento di precisione ed efficienza degli algoritmi di allineamento. Pertanto, l'obiettivo di questo progetto è l'implementazione di un algoritmo di allineamento di sequenze a grafi basato sul metodo di ricerca del percorso più breve, presentato nell'articolo "On the Complexity of Sequence to Graph Alignment", il quale fornisce una descrizione teorica di un algoritmo che consente l'allineamento in tempo O(|V|+m|E|), laddove m indica la dimensione della query, mentre V e E indicano i set di vertici e archi del grafo genomico. In questo lavoro, forniamo la prima implementazione completa e pronta all'uso di tale algoritmo, rispettanto gli standard per la visualizzazione dei risultati. Per ottimizzare il tempo di calcolo è stata implementata un'euristica a banda adattiva. La soluzione presenta due versioni parallelizzate su due architetture diverse (CPU e GPU) di cui sono stati analizzati vantaggi e svantaggi. Per validare il lavoro svolto, sono stati eseguiti dei test sperimentali sul genoma virale del SARS-CoV-2 e su due zone del genoma umano. Da questi, si evince come la soluzione rispetti la complessità temporale teorica e quanto livello di accuratezza raggiunto sia elevato. Infine sono proposte alcune idee per poter migliorare ulteriormente le performance di questa soluzione e per aumentarne le funzionalità.The spreading of NGS technologies has produced an explosion in the amount of genomic data generated enabling a series of new kinds of studies, involving genetic material from different individuals of the same species, family, or environment. New kinds of studies require new computational paradigms to ensure a more intuitive and rapid analysis process. Recently, graph-based representations of genomic data have been proposed as a more efficient and flexible alternative to strings, because they are suitable to represent intra-individual and inter-individual variability. As this change of paradigm is currently taking place, the alignment of sequencing reads to genome graphs is a fertile context for research, offering lots of challenges in terms of accuracy and efficiency of the alignments. Therefore, the goal of this project is the implementation of a sequence-to-graph alignment algorithm based on the shortest-path method, presented in the article "On the Complexity of Sequence to Graph Alignment''. This article provides a theoretical description of an innovative algorithm that allows the alignment of genomic query sequences to genome graphs in O(|V|+m|E|) time, where m denotes the query size, and V and E denote, respectively, the vertices and edges sets of the graph. This thesis work provides the first implementation of this algorithm in a complete and ready-to-use tool in compliance with the standard input and output file formats. To optimize the execution time, an adaptive bandwidth heuristic is proposed. Moreover, the parallelization of the algorithm on two different architectures (CPU and GPU) is investigated, to enhance the pros and cons of both in terms of execution time and memory footprint. Experimental results on SARS-CoV-2 and human genomic data show that the proposed implementations respect the theoretical complexity of the algorithm, and they achieve maximum alignment accuracy while offering the opportunity to tune the algorithm behavior with custom scoring matrices. Finally, this document discusses the next development steps required to optimize even further the performance of the proposed implementations, and enrich them with additional functionalities

    Appropriate Similarity Measures for Author Cocitation Analysis

    Get PDF
    We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis

    Dispelling the Myths Behind First-author Citation Counts

    Get PDF
    We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more sophisticated methods
    corecore