1,720,954 research outputs found
Predictive rating system
The consumption and production of online reviews have become an integral part of our modern lives as the Internet takes over as the main source of information. Unfortunately the vast volume of information cannot simply be digested by traditional means – reading, and additional tools will be required to select better ones that are worth reading. This solution should also help in processing highly unstructured texts with subjective opinions.
The development of Natural Language Processing has made it possible to for machines to process through large amounts of textual data and find out important points. This project makes use of text mining and machine learning methods to categorize texts into structured information and then derive a new rating by evaluating the sentiment of the author. This standardized rating also aims to reduce human bias inherent within written reviews.
This project focuses on food review data within a local setting, Singapore, which was scraped from a website. Aside from constructing a new rating system, we are interested in finding out local preferences from their reviews.
Due to a limited length of word lists fed into the system, it was not able to sufficiently recognize and evaluate one particular aspect. Aside from that aspect, the system demonstrated that it performed well in narrowing down customer choices and produced normally distributed ratings.Bachelor of Engineerin
Question classification via machine learning techniques
Questions are indispensable tools in our daily communication and for the process of acquiring information and knowledge. Recent developments in technology and the internet has also brought about many social sites where community members engage in knowledge-building discussions. These technologies have also been translated to online-learning platforms, and increasingly, these have become scalable tools where students across the globe interact and learn. Understanding the cognitive complexities and quality of questions in such learning settings provide additional insights for educators to monitor achievement of learning outcomes and administer intervention when required. This thesis therefore aims to propose automated solutions using machine learning methods to address this pedagogical need.
Questions in online-learning platforms are commonly found in assessments authored by instructors to assess learners' understanding on the subject. As online-learning platform scales up, it becomes increasingly laborious to manually create assessments comprising questions of various difficulties for students. However, existing question classification models are limited in terms of modeling semantics. Labeling assessment questions by cognitive complexity not only involves the detection of keywords that discriminate between complexities, but also requires consideration of contextual semantic features. A neural network-based machine-learning model is proposed with attention mechanism to direct the creation of a question representation for this purpose. Experiments on university-level digital signal processing questions demonstrate improved performance against other keyword feature machine learning models when detecting patterns resembling Bloom's taxonomy learning outcome templates. In addition, the proposed classifier is integrated into a web-based quiz generation system to support retrieval practice among students with a desired mixture of questions at different complexity levels.
User-generated questions have, on the other hand, become increasingly popular on social media sites for inquiring about specific knowledge outside academic settings. These questions, as opposed to assessment questions, are authored casually, which are error-prone and usually not as sophisticated. To overcome problems of noise such as misspellings, it is important to progressively interpret the question by filtering out the noise and pick out only the salient features. This is achieved via a hierarchical architecture with a new topic-weighted attention mechanism that provides context-aware attention on the question. Furthermore, the proposed approach performs well in the chosen evaluation metrics against other baseline models without assistance from community features. The efficacy of this approach is verified on the Stack Overflow questions dataset. This approach is found to be effective at finding contextual information in the sub-divided texts to form an effective overall representation.
Studies on human-authored texts have found that specific information included in a piece of text improve comprehension. In education and on websites, this helps to increase the overall quality of information being communicated. In the previous model, the attention scheme was data-driven and may not make use of granular entities for extracting features. Using entity embeddings from a named-entity recognizer, the markers give hints to the attention to focus the feature extraction around the entities, thus enhancing performance in its discrimination of very good vs bad questions. Results on the Stack Overflow question dataset indicate that the tag embeddings enhanced its performance over the predecessor, especially with finer categories of tags used, instead of binary indicators. The entity tags were shown to work well with the proposed topic-weighted attention mechanism, thus creating a structural bias to focus on specificity-related features at these crucial locations.Master of Engineerin
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
koamabayili/VECTRON-author-checklist: VECTRON author checklist
We have done our best to complete the author checklist relating to the use of animals in the hut study. Note that the objective for the hut study was to evaluate the IRS treatment applications for residual efficacy against Anopheles mosquitoes, including the local An. coluzzii mosquito population. Cows were only used to attract mosquitoes into the huts and no tests were carried out directly on the cows. The author checklist is intended for use with studies where experiments are carried out on animals, which is why we have had such difficulty in completing this for the hut study, as many of the questions do not relate to how the cows were used
Author-wise bibliometric analysis based on entropy.
Author-wise bibliometric analysis based on entropy.</p
- …
