IMDEA Networks Institute Digital Repository
Not a member yet
1915 research outputs found
Sort by
Poster. What Is the Price of Data? A Measurement Study of Commercial Data Marketplaces
Data-driven decision making powered by ML is changing how the society and the economy work and is having a profound positive impact on our daily life. A McKinsey report predicted that data-driven decision-making could reach US\$2.5 trillion globally by 2025, whereas the European Data Strategy estimates a size of 827 billion euro for the EU27. ML is driving up the demand for data in what has been called the fourth industrial revolution.
A large number of Data Marketplaces (DMs) have appeared in the last few years to help owners monetise their data, and data buyers fuel their marketing processes, train their ML algorithms, and make data-driven decisions. In this poster, we present some preliminary findings of what is, to the best of our knowledge, the first systematic measurement study of DM for data products. This ecosystem, despite being quite vibrant commercially, remains completely unknown to the scientific community. Very basic questions such as "What is the range of prices of data traded in modern DMs?", "Which categories of products command the highest prices?'', "Are the observed prices consistent across DMs?", "Which features correlate with the most expensive data products?'' appear to have no answer and evade most meaningful speculations.
To answer such questions, we first conducted an extensive survey for compiling a catalogue with more than 180 DMs. We then selected 38 of them that fulfill necessary criteria for a measurement study. For these DMs we developed custom crawlers for retrieving information about the products they trade. Using these crawlers we obtained information for more than 213,964 data products and 2,015 data providers. We also developed ML classifiers for identifying data products of similar categories to compare prices across DMs, and executed 9 different regression models to understand which features are driving the prices of data products.TRUEpu
SQLR: Short-Term Memory Q-Learning for Elastic Provisioning
As a growing number of service and application providers choose cloud networks to deliver their services on a software-as-a-service (SaaS) basis, cloud providers need to make their provisioning systems agile enough to meet service level agreements (SLAs). At the same time, they should guard against over-provisioning, which limits their capacity to accommodate more tenants. To this end, we propose Short-term memory Q-Learning pRovisioning (SQLR, pronounced as “scaler”), a system employing a customized variant of the model-free reinforcement learning algorithm. It can reuse contextual knowledge learned from one workload to optimize the number of virtual machines(resources) allocated to serve other workload patterns. With minimal overhead, SQLR achieves comparable results to systems where resources are unconstrained.Our experiments show that we can reduce the amount of provisioned resources by about 20% with less than 1% overall service unavailability (due to blocking), while delivering similar response times to those of an over-provisioned system.pu
Adaptive Uplink Data Compression in Spectrum Crowdsensing Systems
Understanding spectrum activity is challenging when attempted at scale. The wireless community has recently risen to this challenge in designing spectrum monitoring systems that utilize many low-cost spectrum sensors to gather large volumes of sampled data across space, time, and frequencies. These crowdsensing systems are limited by the uplink bandwidth available for transmitting to the backhaul network raw in-phase and quadrature (IQ) samples and power spectrum density (PSD) measurements needed to run a variety of applications. This paper presents FlexSpec, a framework based on the Hadamard-Walsh transform to compress spectrum data collected from distributed and low-cost sensors for real-time applications. We show that this transformation allows sensors to greatly save uplink bandwidth
thanks to its inherent properties both when it is applied to IQ and PSD measurements. Additionally, by leveraging a feedback loop between an edge device and the sensor, FlexSpec carefully adapts
the compression ratio over time, such that data size, application’s performance, and spectrum variations are all considered. We experimentally evaluate FlexSpec in several applications. Our
results show that FlexSpec is particularly suitable for IoT transmissions and for signals close to the noise floor. Compared with prior work, FlexSpec provides up to 7× more reduction of uplink data size for signal detection based on PSD data, and reduces up to 6× to 8× the number of undecodable messages
for IQ sample decoding.TRUEpu
Network management and control for mmWave communications
Millimeter-wave (mmWave) is one of the key technologies that enables the next wireless
generation. mmWave offers a much higher bandwidth than sub-6GHz communications
which allows multi-gigabit-per-second rates. This also alleviates the scarcity of spectrum
at lower frequencies, where most devices connect through sub-6GHz bands. However new
techniques are necessary to overcome the challenges associated with such high frequencies.
Most of these challenges come from the high spatial attenuation at the mmWave band,
which requires new paradigms that differ from sub-6GHz communications. Most notably
mmWave telecommunications are characterized by the need to be directional in order to
extend the operational range. This is achieved by using electronically steerable antenna
arrays, that focus the energy towards the desired direction by combining each antenna
element constructively or destructively. Additionally, most of the energy comes from
the Line Of Sight (LOS) component which gives mmWave a quasi-optical behaviour
where signals can reflect off walls and still be used for communication. Some other
challenges that directional communications bring are mobility tracking, blockages and
misalignments due to device rotation. The IEEE 802.11ad amendment introduced wireless
telecommunications in the unlicensed 60 GHz band. It is the first standard to address
the limitations of mmWave. It does so by introducing new mechanisms at the Medium
Access Control (MAC) and Physical (PHY) layers. It introduces multi-band operation,
relay operation mode, hybrid channel access scheme, beam tracking and beam forming
among others.
In this thesis we present a series of works that aim to improve mmWave
telecommunications. First we give an overview of the intrinsic challenges of mmWave
telecommunications, by explaining the modifications to the MAC and PHY layers. This
sets the base for the rest of the thesis. Then do a comprehensive study on how mmWave
behaves with existing technologies, namely TCP. TCP is unable to distinguish losses
caused by congestion or by transmission errors caused by channel degradation. Since
mmWave is affected by blockages more than sub-6GHz technologies, we propose a set
of parameters that improve the channel quality even for mobile scenarios. The next job
focuses on reducing the initial access overhead of mmWave by using sub-6GHz information
to steer towards the desired direction. We start this work by doing a comprehensive High
Frequency (HF) and Low Frequency (LF) correlation, analyzing the similarity of the
existing paths between the two selected frequencies. Then we propose a beam steering
algorithm that reduces the overhead to one third of the original time. Once we have
studied how to reduce the initial access overhead, we propose a mechanism to reduce
the beam tracking overhead. For this we propose an open platform based on a Field
Programmable Gate Arrays (FPGA) where we implement an algorithm that completely
removes the need to train on the Station (STA) side. This is achieved by changing
beam patterns on the STA side while the Access Point (AP) is sending the preamble.
We can change up to 10 beam patterns without losing connection and we reduce the
overhead by a factor of 8.8 with respect to the IEEE 802.11ad standard. Finally we
present a dual band location system based on Commercial-Off-The-Shelf (COTS) devices.
Locating the STA can improve the quality of the channel significantly, since the AP
can predict and react to possible blockages. First we reverse engineer existing 60 GHz
enabled COTS devices to extract Channel State Information (CSI) and Fine Timing
Measurements (FTM) measurements, from which we can estimate angle and distance.
Then we develop an algorithm that is able to choose between HF and LF in order to
improve the overall accuracy of the system. We achieve less than 17 cm of median error
in indoor environments, even when some areas are Non Line Of Sight (NLOS).Telematics EngineeringUniversidad Carlos III de Madrid, Spai
Robust multivariate control chart based on shrinkage for individual observations
A robust multivariate quality control technique for individual observations is proposed, based on the robust reweighted shrinkage estimators. A simulation study is done to check the performance and compare the method with the classical Hotelling approach, and the robust alternative based on the reweighted minimum covariance determinant estimator. The results show the appropriateness of the method even when the dimension or the Phase I contamination are high, with both independent and correlated variables, showing additional advantages about computational efficiency. The approach is illustrated with two real data-set examples from production processes.pu
GLOVE: towards privacy-preserving publishing of record-level-truthful mobile phone trajectories
Datasets of mobile phone trajectories collected by network operators offer an unprecedented opportunity to discover new knowledge from the activity of large populations of millions. However, publishing such trajectories also raises significant privacy concerns, as they contain personal data in the form of individual movement patterns. Privacy risks induce network operators to enforce restrictive confidential agreements in the rare occasions when they grant access to collected trajectories, whereas a less involved circulation of these data would fuel research and enable reproducibility in many disciplines. In this work, we contribute a building block towards the design of privacy-preserving datasets of mobile phone trajectories that are truthful at the record level. We present GLOVE, an algorithm that implements k-anonymity, hence solving the crucial unicity problem that affects this type of data while ensuring that the anonymized trajectories correspond to real-life users. GLOVE builds on original insights about the root causes behind the undesirable unicity of mobile phone trajectories, and leverages generalization and suppression to remove them. Proof-of-concept validations with large-scale real-world datasets demonstrate that the approach adopted by GLOVE allows preserving a substantial level of accuracy in the data, higher than that granted by previous methodologies.pu
Routing optimization algorithms in integrated fronthaul/backhaul networks supporting multitenancy
This thesis aims to help in the definition and design of the 5th generation of telecommunications networks (5G) by modelling the different features that characterize them through several mathematical models. Overall, the aim of these models is to perform a wide optimization of the network elements, leveraging their newly-acquired capabilities in order to improve the efficiency of the future deployments both for the users and the operators. The timeline of this thesis corresponds to the timeline of the research and definition of 5G networks, and thus in parallel and in the context of several European H2020 programs. Hence, the different parts of the work presented in this document match and provide a solution to different challenges that have been appearing during the definition of 5G and within the scope of those projects, considering the feedback and problems from the point of view of all the end users, operators and providers.
Thus, the first challenge to be considered focuses on the core network, in particular on how to integrate fronthaul and backhaul traffic over the same transport stratum. The solution proposed is an optimization framework for routing and resource placement that has been developed taking into account delay, capacity and path constraints, maximizing the degree of Distributed Unit (DU) deployment while minimizing the supporting Central Unit (CU) pools. The framework and the developed heuristics (to reduce the computational complexity) are validated and applied to both small and large- scale (production-level) networks. They can be useful to network operators for both network planning as well as network operation adjusting their (virtualized) infrastructure dynamically.
Moving closer to the user side, the second challenge considered focuses on the allocation of services in cloud/edge environments. In particular, the problem tackled consists of selecting the best location of each Virtual Network Function (VNF) that compose a service in cloud robotics environments, that imply strict delay bounds and reliability constraints. Robots, vehicles and other end-devices provide significant capabilities such as actuators, sensors and local computation which are essential for some services. On the negative side, these devices are continuously on the move and might lose network connection or run out of battery, which further challenge service delivery in this dynamic environment. Thus, the performed analysis and proposed solution tackle the mobility and battery restrictions. We further need to account for the temporal aspects and conflicting goals of reliable, low latency service deployment over a volatile network, where mobile compute nodes act as an extension of the cloud and edge computing infrastructure. The problem is formulated as a cost-minimizing VNF placement optimization and an efficient heuristic is proposed. The algorithms are extensively evaluated from various aspects by simulation on detailed real-world scenarios.
Finally, the last challenge analyzed focuses on supporting edge-based services, in particular, Machine Learning (ML) in distributed Internet of Things (IoT) scenarios. The 1 traditional approach to distributed ML is to adapt learning algorithms to the network, e.g., reducing updates to curb overhead. Networks based on intelligent edge, instead, make it possible to follow the opposite approach, i.e., to define the logical network topology around the learning task to perform, so as to meet the desired learning performance.
The proposed solution includes a system model that captures such aspects in the context of supervised ML, accounting for both learning nodes (that perform computations) and information nodes (that provide data). The problem is formulated to select (i) which learning and information nodes should cooperate to complete the learning task, and (ii) the number of iterations to perform, in order to minimize the learning cost while meeting the target prediction error and execution time. The solution also includes a heuristic algorithm that is evaluated leveraging a real-world network topology and considering both classification and regression tasks, and closely matches the optimum, outperforming state-of-the-art alternatives.Telematics EngineeringUniversidad Carlos III de Madrid, Spai
Optimal Performance of Parallel-Server Systems with Job Size Prediction Errors
Modern communication networks integrate distributed computing architectures, in which customers are processed in parallel. We show how to minimize the waiting time of customer’s jobs by leveraging a simple threshold-based job dispatching policy. The optimal policy leverages the SITA routing, which assigns jobs to servers according to the size of the job. Moreover, the optimal policy permits to optimize system performance even when the job size is not known a priori and is estimated by means of error-prone predictors.pu
Model-free machine learning of wireless SISO/MIMO communications
Machine learning is a highly promising tool to design the physical layer of wireless communication systems, but the training usually requires an explicit model of the signal distortion as it undergoes transmission over a wireless channel. As data rates, number of MIMO streams and carrier frequencies increase to satisfy the demand for wireless capacity, it becomes difficult to design hardware with few imperfections and to model the imperfections that there are. New machine learning schemes for the physical layer do not require an explicit model but can implicitly learn the end-to-end link including channel characteristics and non-linearities of the system directly from the training data. In this paper, we present a novel neural network architecture that provides an explicit stochastic model for both SISO and MIMO channels, by learning the parameters of a Gaussian mixture distribution from real channel samples. We use this channel model in conjunction with an autoencoder to learn a suitable modulation scheme. We experimentally validate our proposed model in an FPGA-based millimeter-wave testbed for both SISO and MIMO channels, showing that it is able to reproduce the channel characteristics with good accuracy.TRUEpu
Tractable low-delay atomic memory
Communication cost is the most commonly used metric in assessing the efficiency of operations in distributed algorithms for message-passing environments. In doing so, the standing assumption is that the cost of local computation is negligible compared to the cost of communication. However, in many cases, operation implementations rely on complex computations that should not be ignored. Therefore, a more accurate assessment of operation efficiency should account for both computation and communication costs.
This paper focuses on the efficiency of read and write operations in emulations of atomic read/write shared memory in the asynchronous, message-passing, crash-prone environment. The much celebrated work by Dutta et al. presented an implementation in this setting where all read and write operations could complete in just a single communication round-trip. Such operations where characterized for the first time as fast. At its heart, the work by Dutta et al. used a predicate to achieve that performance. We show that the predicate is computationally intractable by defining an equivalent problem and reducing it to Maximum Biclique, a known NP-hard problem.
We derive a new, computationally tractable predi-cate, and an algorithm to compute it in linear time. The proposed predicate is used to develop three algorithms: ccFast, ccHybrid, and OhFast. ccFast is similar to the algorithm of Dutta et al. with the main difference being the use of the new predicate for reduced computational complexity. All operations in ccFast are fast, and particular constraints apply in the number of participants. ccHybrid and OhFast, allow some operations to be “slow”, enabling unbounded participants in the service. ccHybrid is a “multi-speed” version of cc- Fast, where the reader determines when it is not safe to complete a read operation in a single communication round-trip. OhFast, expedites algorithm OhSam of Hadjistasi et al. by placing the developed predicate at the servers instead of clients and avoiding excessive server communication when possible. An experimental evaluation using NS3 compares algorithms ccHybrid and OhFast to the classic algorithm ABD of Attiya et al., the algorithm Sf of Georgiou et al. (the first “semifast” algorithm, allowing both fast and slow operations), and algorithm OhSam.
In summary, this work gives the new meaning to the term fast by assessing both the communication and the computation efficiency of each operation.pu