1,720,960 research outputs found
Two-sample comparisons for serially correlated data
Thesis (M.Sc. (Statistics))--North-West University, Potchefstroom Campus, 2007.The purpose of this study is to derive new tests for the equality of the means in two
independent or dependent stationary time series, based on bootstrap critical values.
Required properties of these tests include satisfactory probability of Type I errors, and
high power. It is shown how critical points for various sample sizes and significance
levels can be obtained by applying the parametric bootstrap. A limited Monte Carlo
simulation study is conducted to illustrate the validity of the bootstrap approximation
of the exact critical values, by producing satisfactory probability of Type I errors. It
also shows that the newly proposed tests compare favourably with standard two-sample
tests in the absence of serial correlation, under the null hypothesis of equal
means, but are more powerful than the well-known t -test if small and moderate
correlation structures are present, for a wide range of parameter values. All findings
and conclusions of the Monte Carlo simulations are reported.Master
Credit risk prediction with and without weights of evidence using quantitative learning models
The credit risk assessment process is necessary for maintaining financial stability, cost and time efficiency, model performance accuracy, comparability analysis and future business implications in the commercial banking sector. By accurately predicting credit risk, highly regulated banks can make informed lending decisions and minimize potential financial losses. The purpose of this paper is to assess the power of conventional predictive statistical models with and without transforming the features to gain better insights into customer's creditworthiness. The findings of the predicted performance of the logistics regression model are compared to the performance results of machine learning models for credit risk assessment using commercial banking credit registry data. Each model has its strengths and weaknesses, and where one model lacks, another performs better. The article reveals that simpler credit risk assessment techniques delivered outstanding performance while consuming less processing power and have given insights into the most contributing feature categories. Improving a conventional predictive statistical model using some of the feature transformations reduces the overall model performance, specifically for credit registry data. The logistics regression model outperformed all models with the highest F1, accuracy, Jaccard Index and AUC values, respectively. Financial institutions, specifically banks have questioned whether transformations using Weights of Evidence (WoE) have been significant in quantifying the relationship between categorical independent variables for various types of credit data. This study provides insights when considering the usage of feature transformation for credit risk modelling in commercial banking. The transformation technique is particularly useful in situations where statistical predictive modelling techniques are employed. The results revealed that not only can the logistic regression models perform similarly to the machine learning models but can also outperform them. The best performance is attributed to the simplicity, interpretability, and access to understanding features of individual clients within a portfolio of credit products. The logistic regression model without transformation turned out to perform the best out of the five machine learning models. Considering the business impact, enhancing the logistic regression model by using a WoE transformation did not improve the model's performance for commercial banking data considered. However, the transformation did provide insights regarding each binned categorical independent variable. Therefore, our findings in this article contribute towards assisting banks in managing the impact and interpretability of each binned feature category on the discriminatory power of credit scoring
A comprehensive high pure momentum equity timing framework using the Kalman filter and ARIMA forecasting
Article, Faculty of Natural and Agricultural Sciences (Unit for Data Science and Computing (UDSC)--Northwest University, Potchefstroom CampusThe pursuit of higher returns has led to a growing interest in factor timing as a strategy to enhance portfolio returns. Momentum is a popular factor, which involves buying securities that have shown consistent price appreciation over the past 3 to 12 months or past few years, with the expectation that the trend will continue and reducing exposure to those that consistently declined. An important part of a factor timing strategy is in the portfolio optimization process. This article aimed to first construct a large capitalization pure momentum portfolio, which included a dynamic stringent portfolio construction process criteria for selecting stocks estimated from historical data. Second, as a part of the portfolio’s risk management strategy, the Kalman filter was applied to the historical performance of this portfolio. Lastly, the ARIMA forecast was used to estimate expected performance and the confidence intervals. The empirical results showed that this pure equity momentum factor timing framework with the Kalman filter together with the ARIMA (autoregressive integrated moving average) forecasting methodology was iterative and incorporated new information as it became available and further enhanced the monitoring and rebalancing process. This adaptive approach enabled the portfolio to capitalize on time-varying return anomalies as they occured
Assessment of model risk due to the use of an inappropriate parameter estimator
The purpose of this study is to assess model risk with respect to parameter estimation for a simple binary logistic regression model applied as a predictive model. The assessment is done by comparing the effectiveness of eleven different parameter estimation methods. The results from the historical credit dataset of a certain financial institution confirmed that using several optimization methods to address parameter estimation risk for predictive models is substantial. This is the case, especially when there exists a numerical optimization method that estimates the optimum parameters and minimizes the cost function among alternative methods. Our study only considers a univariate predictor with a static sample size of cases. This research work contributes to the literature by presenting different parameter estimation methods for predicting the probability of default through binary logistic regression model and determining optimum parameters that minimize the objective model's cost function. The Mini-Batch Gradient Descent method is revealed to be the better parameter estimator
Going Beyond Counting First Authors in Author Co-citation Analysis
The present study examines one of the fundamental aspects of author co-citation analysis (ACA) - the way co-citation
counts are defined. Co-citation counting provides the data on which all subsequent statistical analyses and mappings
are based, and we compare ACA results based on two different types of co-citation counting - the traditional type that
only counts the first one among a cited work's authors on the one hand and a non-traditional type that takes into
account the first 5 authors of a cited work on the other hand. Results indicate that the picture produced through this non-traditional author co-citation counting contains more coherent author groups and is therefore considerably clearer. However, this picture represents fewer specialties in the research field being studied than that produced through the traditional first-author co-citation counting when the same number of top-ranked authors is selected and analyzed. Reasons for these effects are discussed
Credit application scoring for consumers without credit history
MSc (Computer Science), North-West University, Vaal Triangle CampusCredit scoring is a tool that is used to either qualify or disqualify credit applicants by quantifying the risk factors relevant to classify them to high risk or low risk. Due to a demand in credit inclusion, financial institutions, especially banks must come up with a way of screening and scoring applicants. In most cases, applicants are required to have credit history, or risk being denied credit because these institutions cannot charge high interest rates, mainly because they are obliged legally not to do so on the repayment of the loans due to a lack of the applicant’s credit history. In this study, the concept and application of credit scoring is explained. The steps necessary to develop a credit scoring model are outlined with the focus on data that do not have any credit history. Literature is reviewed discussing the background information regarding the performance of the logistic regression model and other statistical models in classifying consumers. Datasets, statistical models, methodology and variables were reviewed and used to assist in building the scorecard. Secondary data collected from the General Household Survey (GHS) is used to classify credit applicants into two groups of high risk and low risk. Binary logistic regression is used to identify the variables that best predict these two groups. The forward selection technique is used in determining variables that are significant. The developed model is tested for prediction accuracy and thereafter, this is followed by key findings and recommendations. In conclusion, the developed model is found to be fitting the data well.Master
Variations on the Author
“Variations on the Author” discusses two of Eduardo Coutinho’s recent films (Um Dia na Vida, from 2010, and Últimas Conversas, posthumously released in 2015) and their contribution to the general question of documentary authorship. The director’s filmography is characterized by a consistent yet self-effacing form of authorial self-inscription: Coutinho often features as an interviewer that rather than express opinions propels discourses; an interviewer that is good at listening. This mode of self-inscription characterizes him as an author who is not expressive but who is nonetheless markedly present on the screen. In Um Dia na Vida, however, Coutinho is completely absent form the image, while Últimas Conversas, on the contrary, includes a confessional prologue that moves the director from the margins to the center of his films. This article examines the ways in which these works stand out in the filmography of a director who offers new insights into the notion of cinematic authorship
Appropriate Similarity Measures for Author Cocitation Analysis
We provide a number of new insights into the methodological discussion about author cocitation analysis. We first argue that the use of the Pearson correlation for measuring the similarity between authors’ cocitation profiles is not very satisfactory. We then discuss what kind of similarity measures may be used as an alternative to the Pearson correlation. We consider three similarity measures in particular. One is the well-known cosine. The other two similarity measures have not been used before in the bibliometric literature. Finally, we show by means of an example that our findings have a high practical relevance.information science;Pearson correlation;cosine;similarity measure;author cocitation analysis
Dispelling the Myths Behind First-author Citation Counts
We conducted a full-scale evaluative citation analysis study of scholars in the XML research field to explore just how different from each other author rankings resulting from different citation counting methods actually are, and to demonstrate the capability of emerging data and tools on the Web in supporting more realistic citation counting methods. Our results contest some common arguments for the continued
use of first-author citation counts in the evaluation of scholars, such as high correlations between author rankings by first-author citation counts and other citation
counting methods, and high costs of using more realistic citation counting methods that are not well-supported by the ISI databases. It is argued that increasingly available digital full text research papers make it possible for citation analysis studies to go beyond what the ISI databases have directly supported and to employ more
sophisticated methods
- …
