1,721,364 research outputs found

    Long time, no tweets! Time-aware personalised hashtag suggestion

    Get PDF
    Microblogging systems, such as the popular service Twitter, are an important real-time source of information however due to the amount of new information constantly appearing on such services, it is difficult for users to organise, search and re-find posts. Hashtags, short keywords prefixed by a # symbol, can assist users in performing these tasks, however despite their utility, they are quite infrequently used. This work considers the problem of hashtag recommendation where we wish to suggest appropriate tags which the user could assign to a new post. By identifying temporal patterns in the use of hashtags and employing personalisation techniques we construct novel prediction models which build on the best features of existing methods. Using a large sample of data from the Twitter API we test our novel approaches against a number of competitive baselines and are able to demonstrate significant performance improvements, particularly for hashtags that have large amounts of historical data available

    Understanding engagement through search behaviour

    No full text
    Evaluating user engagement with search is a critical aspect of understanding how to assess and improve information retrieval systems. While standard techniques for measuring user engagement use questionnaires, these are obtrusive to user interaction, and can only be collected at acceptable intervals. The problem we address is whether there is a less obtrusive and more automatic way to assess how users perceive the search process and outcome. Log files collect behavioural signals (e.g., clicks, queries) from users on a large scale. In this paper, we investigate the potential to predict how users perceive engagement with search by modelling behavioural signals from log files using supervised learning methods.We focus on different engagement dimensions (Perceived Usability, Felt Involvement, Endurability and Novelty) and examine how 37 behavioural features can inform these dimensions. Our results, obtained from 377 in-lab participants undergoing goal-based search tasks, support the connection between perceived engagement and search behaviour. More specifically, we show that time-and query-related features are best suited for predicting user perceived engagement, and suggest that different behavioural features better reflect specific dimensions. We demonstrate the possibility of predicting user-perceived engagement using search behavioural features

    Semi-supervised event-related tweet identification with dynamic keyword generation

    No full text
    Twitter provides us a convenient channel to get access to the immediate information about major events. However, it is challenging to acquire a clean and complete set of event-related data due to the characteristics of tweets, e.g., short and noisy. In this paper, we propose a semi-supervised method to obtain high quality event-related tweets from Twitter stream, in terms of precision and recall. Specifically, candidate event-related tweets are selected based on a set of keywords. We propose to generate and update these keywords dynamically along the event development. To be included in this keyword set, words are evaluated based on single word properties, property based on co-occurred words, and changes of word importance over time. Our solution is capable of capturing keywords of emerging aspects or aspects with increasing importance along event evolvement. By leveraging keyword importance information and a few labeled tweets, we propose a semi-supervised expectation maximization process to identify event-related tweets. This process significantly reduces human effort in acquiring high quality tweets. Experiments on three real world datasets show that our solution outperforms state-of-the-art approaches by up to 10% in F1 measure

    Exploiting spatio-temporal user behaviors for user linkage

    No full text
    Cross-device and cross-domain user linkage have been attracting a lot of attention recently. An important branch of the study is to achieve user linkage with spatio-temporal data generated by the ubiquitous GPS-enabled devices. The main task in this problem is twofold, i.e., how to extract the representative features of a user; how to measure the similarities between users with the extracted features. To tackle the problem, we propose a novel model STUL (Spatio-Temporal User Linkage) that consists of the following two components. 1) Extract users' spatial features with a density based clustering method, and extract the users' temporal features with the Gaussian Mixture Model. To link user pairs more precisely, we assign different weights to the extracted features, by lightening the common features and highlighting the discriminative features. 2) Propose novel approaches to measure the similarities between users based on the extracted features, and return the pair-wise users with similarity scores higher than a predefined threshold. We have conducted extensive experiments on three real-world datasets, and the results demonstrate the superiority of our proposed STUL over the state-of-the-art methods

    Computing and mining ClustCube cubes efficiently

    No full text
    A novel computational paradigm for clustering complex database objects extracted from distributed database settings via well-understood OLAP technology is proposed and experimentally assessed in this paper. This paradigm conveys in the so-called ClustCube cubes, which define a novel multidimensional data cube model according to which (data) cubes store clustered complex database objects rather than conventional SQL-based aggregations. A major contribution of this research is represented by effective and efficient algorithms for computing ClustCube cubes that, surprisingly, are capable of reducing computational efforts significantly with respect to traditional approaches. Our analytical contribution is completed by a comprehensive assessment of proposed algorithms against both benchmark and real-life data sets, which clearly confirms the benefits deriving from our proposal

    Erratum

    No full text
    This article has been withdrawn as it was published elsewhere and accidentally duplicated. The original article can be seen here: 10.1108/02641619810369644. When citing the article, please cite: Schubert Foo, Ee-Peng Lim, (1998), “An integrated Web-based ILL system for Singapore libraries”, Interlending &amp; Document Supply, Vol. 26 Iss 1 pp. 10 - 20.</jats:p

    A semi-automated digital preservation system based on semantic Web services

    Get PDF
    This paper describes a Web-services-based system which we have developed to enable organizations to semi -automatically preserve their digital collections by dynamically discovering and invoking the most appropriate preservation service, as it is required. By periodically comparing preservation metadata for digital objects in a collection with a software version registry, potential object obsolescence can be detected and a notification message sent to the relevant agent. By making preservation software modules available as Web services and describing them semantically using a machine-processable ontology (OWL-S), the most appropriate preservation service(s) for each object can then be automatically discovered, composed and invoked by software agents (with optional human input at critical decision-making steps). We believe that this approach represents a significant advance towards providing a viable, cost-effective solution to the long term preservation of large-scale collections of digital objects

    Going Beyond Counting First Authors in Author Co-citation Analysis

    Get PDF
    The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
    corecore