IT University of Copenhagen

The IT University of Copenhagen's Repository
Not a member yet
    9607 research outputs found

    TensorSocket: Shared Data Loading for Deep Learning Training

    No full text
    Training deep learning models is a repetitive and resource-intensive process. Data scientists often train several models before landing on a set of parameters (e.g., hyper-parameter tuning) and model architecture (e.g., neural architecture search), among other things that yield the highest accuracy. The computational efficiency of these training tasks depends highly on how well the training data is supplied to the training process. The repetitive nature of these tasks results in the same data processing pipelines running over and over, exacerbating the need for and costs of computational resources.In this paper, we present TensorSocket to reduce the computational needs of deep learning training by enabling simultaneous training processes to share the same data loader. TensorSocket mitigates CPU-side bottlenecks in cases where the collocated training workloads have high throughput on GPU, but are held back by lower data-loading throughput on CPU. TensorSocket achieves this by reducing redundant computations and data duplication across collocated training processes and leveraging modern GPU-GPU interconnects. While doing so, TensorSocket is able to train and balance differently-sized models and serve multiple batch sizes simultaneously and is hardware- and pipeline-agnostic in nature.Our evaluation shows that TensorSocket enables scenarios that are infeasible without data sharing, increases training throughput by up to 100%, and when utilizing cloud instances, achieves cost savings of 50% by reducing the hardware resource needs on the CPU side. Furthermore, TensorSocket outperforms the state-of-the-art solutions for shared data loading such as CoorDL and Joader; it is easier to deploy and maintain and either achieves higher or matches their throughput while requiring fewer CPU resources

    Diversity, Equity and Inclusion Activities in DatabaseConferences: A 2024 Report

    No full text
    The database community's Diversity, Equity, and Inclusion (DEI) initiative began in 2020 as the Diversity/ Inclusion initiative. This report highlights our activities from 2024. Our goal as a community is to make all DB conference attendees feel included, regardless of their scientific views or personal backgrounds. As a leadership team, the DEI group supports DEI chairs across conferences, preserves institutional memory of DEI efforts, shapes a shared vision, and fosters collaboration to advance inclusion. These efforts are carried out by core members and liaisons from each conference's executive committee (Figure 2). The initiative was relaunched in January 2024 with a new structure based on five key actions: COORDINATE, to support collaboration between core members, liaisons, and DEI chairs; SCOUT, to gather best DEI practices from other communities; ETHICS, to create and promote ethical guidelines for writing and reviewing; MEDIA, to collect and share digital content from DEI@DB events; and DIVERSIFY, to analyze data on diversity, accessibility, and the adoption of DEI principles in research and academia. DBCARES1 is now officially part of the DEI initiative. The mission of DBCARES is to create an inclusive and diverse Database community with zero tolerance for abuse, discrimination, or harassment. As part of this integration, we unified the Code of Ethics and introduced clear guidelines for DB conference organizers. Several conferences-including SIGMOD, VLDB, ICDE, and EDBT-continued using CLOSET to ensure fair and transparent reviewer assignments

    CXL-Bench: Benchmarking Shared CXL Memory Access

    No full text
    Memory access paths between a CPU core and memory are increasingly complex. Data can be placed on local- or remote-socket memory, and on local- and remote-die memory on modern multi-die CPUs, affecting memory access performance. Cache-coherent inter-device interconnects, such as Compute Express Link (CXL), allow a CPU core to perform load and store instructions to memory of a peripheral device. Such accesses incur higher access latency than accesses to local-socket memory and increase the access path complexity. For database system developers, it is important to understand the performance implications of these complex memory architectures. In this work, we present CXL-Bench, a benchmark framework for quantifying access performance for different memory access paths. CXL-Bench provides many configuration options, such as memory access patterns, the operating system’s memory abstraction, cache bypass options, and a distributed mode for setups with multiple servers accessing memory of the same device. We demonstrate the utility of CXL-Bench by quantifying memory access characteristics of two servers accessing a shared CXL 1.1 memory device. Our results show that memory accesses of one server to the device affect the access performance of another server accessing the same device. On the other hand, memory (de)allocations using CXL memory configured as a character device complete quickly, making frequent re-allocation of CXL memory feasible

    Counting Small Induced Subgraphs: Scorpions Are Easy but Not Trivial

    No full text
    In the parameterized problem #IndSub(Φ) for fixed graph properties Φ, given as input a graph G and an integer k, the task is to compute the number of induced k-vertex subgraphs satisfying Φ. Dörfler et al. [Algorithmica 2022] and Roth et al. [SICOMP 2024] conjectured that #IndSub(Φ) is #W[1]-hard for all non-meager properties Φ, i.e., properties that are nontrivial for infinitely many k. This conjecture has been confirmed for several restricted types of properties, including all hereditary properties [STOC 2022] and all edge-monotone properties [STOC 2024].We refute this conjecture by showing that induced k-vertex graphs that are scorpions can be counted in time O(n⁴) for all k. Scorpions were introduced more than 50 years ago in the context of the evasiveness conjecture. A simple variant of this construction results in graph properties that achieve arbitrary intermediate complexity assuming ETH.Moreover, we formulate an updated conjecture on the complexity of #IndSub(Φ) that correctly captures the complexity status of scorpions and related constructions

    Building a Data Management System for the Cloud: Lessons Learned and Future Directions

    No full text
    The paper discusses the lessons learned from building Snowflake, a data management system for the cloud. Given the need for systems that can scale to handle large data volumes, provide expressive programming interfaces, and leverage the benefits of cloud computing, it describes the architecture of a cloud-based data management system and optimization techniques specific to the cloud. Key techniques include pruning large file sets at both compile time and query runtime, optimizing data layouts in the background, and, more generally, the importance of performing maintenance tasks in the background, which is enabled by cloud resources. The paper also explains the need for using immutable files and the implications for data modification queries. Finally, it highlights the operational aspects of building and maintaining a data management system that functions as an online cloud service. The paper concludes by outlining future directions for cloud-based data management systems

    Dense Dataset from the paper "From Theory to Practice: Engineering Approximation Algorithms for Dynamic Orientation".

    No full text
    Dynamic graph algorithms have seen significant theoretical advancements, but practical evaluations often lag behind. This work bridges the gap between theory and practice by engineering and empirically evaluating recently developed approximation algorithms for dynamically maintaining graph orientations. We comprehensively describe the underlying data structures, including efficient bucketing techniques and round-robin updates. Our implementation has a natural parameter λ, which allows for a trade-off between algorithmic efficiency and the quality of the solution. In the extensive experimental evaluation, we demonstrate that our implementation offers a considerable speedup. Using different quality metrics, we show that our implementations are very competitive and can outperform previous methods. Overall, our approach solves more instances than other methods while being up to 112 times faster on instances that are solvable by all methods compared

    Critical Perspectives on Predictive Policing: Anticipating Proof?

    No full text
    This chapter interrogates the question of what defines predictive policing. Our point of departure is notions in Science and Technology Studies that attend to how boundary objects such as predictive policing are produced through the interrelation of politics, sociotechnical imaginaries, and contestations. Through ethnographic inquiry on the empirical case of the POL-INTEL platform of the Danish police, we follow the evolution and demarcations in the (re)-classifications of POL-INTEL from predictive to non-predictive. In so doing, we situate POL-INTEL in the wider international context regarding the development of predictive policing and investigate the boundary work between predictive policing and similar concepts such as Intelligence-Led Policing. Concisely, we argue that the development of the notion of predictive policing is fundamentally unstable and tied to political and economic interests

    Patch Explorer: Interpreting Diffusion Models through Interaction

    No full text
    We introduce Patch Explorer, an interactive interface for visualizing and manipulating the patches as they are processed by cross-attention heads. Built on interventions via NNsight, our interface lets users inspect and manipulate individual attention heads over layers and timesteps. Interaction via the interface reveals that attention heads independently capture semantics, like a unicorn’s horn, in diffusion models. Next to offering a way to analyze its behavior, users can also intervene with Patch Explorer to edit semantic associations within diffusion models, like adding a unicorn horn to a horse. Our interface also helps understand the role of a diffusion timestep through precise interventions. By providing a visualization tool with interactivity based on attention heads, we aim to shed light on their role in generative processes

    Re-Imagining the Landscape of Future Farming

    Get PDF
    This paper explores how participatory design methods can engage future farmers in imagining sustainable agricultural futures in Denmark. Denmark's agricultural sector faces a "crisis of imagination" where dominant paradigms in farming are hard to challenge. Building on Ingold's notion of the landscape as a relational, temporal space for dwelling, we propose a conceptualization of the landscape game that enables farming students to articulate their visions of future farming. The design game creates opportunities for participants to situate their imagined farms within broader social, environmental, and technical contexts while exploring how these might evolve over time. This approach aims to generate new agricultural imaginaries that move beyond current techno-solutionist or radical transformation narratives, while supporting an underrepresented political public in developing their own perspective on contested agricultural futures that are both speculative and grounded in lived experience

    4,472

    full texts

    9,607

    metadata records
    Updated in last 30 days.
    The IT University of Copenhagen's Repository
    Access Repository Dashboard
    Do you manage Open Research Online? Become a CORE Member to access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard! 👇