1,720,991 research outputs found
Atlantic Canada’s distributed generation future : renewables, transportation, and energy storage
227 leaves : ill. (chiefly col.), col. maps ; 29 cmIncludes abstract and appendices.Includes bibliographical references (leaves 186-205).Nova Scotia and the Federal Government have set low but achievable greenhouse gas targets, aimed at reducing our climate change burden while improving energy security in a globalizing economy. Can we decarbonize our economy and improve our technologic level? To answer this, I assessed our renewable resources and power plants, and in doing so reviewed practicalities of finite and renewable primary energy going forward. I collected data and created Nova Scotia’s Energy Map with the objective of improving Energy System Awareness in the region with an interactive online map. I evaluated and compared technology regimes based on: economics; operational risks; quality of environment, human, and animal health; along with social aspects of energy production and consumption. Finally with EnergyPLAN I analyzed and validated a near 100% renewable primary energy scenario to aid the understanding of regional decision makers that decarbonization is achievable and with proper implementation advantageous
A network-based approach to earthquake pattern analysis
108 leaves : ill. (some col.), col. maps ; 29 cm.Includes abstract and appendices.Includes bibliographical references.This research proposes a new network-based method for the assessment of earthquake relationships in space-time-magnitude patterns. The method is applied to the study of volcanic seismicity in Hawaii, over a time period from January 1st, 1989 to December 31st, 2012. It is shown that networks with high values of the minimum edge weight W[subscript min] enjoy strong scaling properties, as opposed to networks with low values for W[subscript min], which exhibit poor or no such properties. The scaling behaviour along the spectrum of W[subscript min], in conjunction with the robustness regarding parameter variations, endorse the idea of a relationship between fundamental properties of seismicity and the scaling properties of the earthquake networks, and can be used to discern the interrelated earthquakes from the rest of the dataset. The scale free behaviour of the connectivity distribution along the spectrum of the minimum weight values is mirrored by a similar behaviour of the distribution of the number of nodes’ linked neighbours. The patterns found in the distributions of temporal and spatial intervals between earthquakes are similar in various networks, from large to small networks. Notable similarities are found between the variation of the network clustering coefficient, C, and the variation of the exponents of the connectivity distribution, [Beta], and of the weight distribution, [gamma] . Results of this method are further applied for the study of temporal changes in volcanic seismicity patterns. It is shown that [Beta], [gamma] , and C manifest a generally synchronous variation over successive temporal windows, which can be related to changes in seismicity and in the life of the volcanic system. A Zipf distribution is found for the ranked sets of magnitude values of successive network nodes. The distribution of differences between the magnitude values of successive nodes is also governed by a power law
Software defect prediction from code quality measurements via machine learning
viii, 125 leaves : illustrations ; 29 cmIncludes abstract and appendices.Includes bibliographical references (leaves 114-117).Improvement in software development practices to predict and reduce software defects can lead to major cost savings. The goal of this thesis is to demonstrate the value of static analysis metrics and rules in predicting software defects at a much larger scale than previous efforts. The study analyses data collected from more than 500 software applications, across 3 multi-year software development programs, and uses over 150 software static analysis measurements. Static analysis metrics, rule violations and software defect historical actual values are sourced from multiple disparate databases, joined and groomed for analysis. Several feature selection techniques are employed to narrow the feature set focus to the most influential variables. Furthermore, a number of machine learning techniques such as neural network and random forest are used to determine whether seemingly innocuous rule violations can be used as significant predictors of software defect rates
Predicting unusual surges in a time series
165 leaves : ill. (chiefly col.) ; 29 cm.Includes abstract and appendix.Includes bibliographical references (leaves 84-87).Time-series data tend to enjoy regular fluctuations. Statisticians have developed a wide variety of techniques to predict future values of a temporal variable. Most of these approaches use prediction techniques; one example is the employment of auto regression and moving averages to predict future numerical values.
Our project uses a data mining technique called classification to predict both the occurrence of surges in time-series, and the expected durations of those surge, as opposed to future values predicted using other techniques. Such surges can occur in a number of time series events, examples of which include demands for energy, weather forecasting, and variation in traffic volume. Our chosen technique can be employed to extract meaningful statistics and other useful characteristics of time series data.
Classifier performance depends greatly on the characteristics of the data to be analyzed. Many algorithms are part of classification analysis. For this study, we chose for comparison the decision tree, support vector machines, and Adaboost. To validate the quality of algorithms for our given problem, we used precision and recall measures as comparators between different algorithms. The minimal accepted precision score was set as 60%, with 70% as the preferred such score, as such a result would be more robust. Our initial experiments yielded a precision score of 64%, and the best results attained a score of 77%
Inventory management using data mining : forecasting in retail trade
xiii, 181 leaves : ill. ; 29 cm.Includes abstract.Includes bibliographical references (leaves 176-181).Inventory management, as an important business issue, plays a significant role in promoting business development. This study aims to apply data mining techniques, such as time series clustering and time series prediction techniques, in inventory management. Based on historical business data sets, time series clustering techniques, such as K-Means and Expectation Maximization are used to categorize inventories into reasonable groups. This study then identifies the most effective prediction technique to accurately predict inventory demands for each group. The traditional statistical evaluation metrics, such as Mean Absolute Percentage Error may not always be good indicators in an inventory management system, where the goal is to have as little inventory as possible without ever running out. The thesis proposes a more appropriated evaluation metric based on cost/benefit analysis of inventory forecasts. Results from a simulation program based on the proposed cost/benefit analysis are compared with statistical metrics
Identifying users and activities from brain wave signals recorded from a wearable headband
1 online resource (xii, 82 p.) : ill.Includes abstract.Includes bibliographical references (p. 79-82).This paper studies the supervised classification of electroencephalogram (EEG) brain signals to identify persons and their activities. The brain signals are obtained from a commercially available and modestly priced wearable headband. Such wear-able devices generate a large amount of data and due to their attractive pricing struc-ture are becoming increasingly commonplace. As a result, the data generated from such wearables will increase
exponentially, leading to many interesting data mining opportunities.
This paper proposes a representation that reduces variable length signals to more manageable and uniformly fixed length distributions, and then explores the effectiveness of a variety of data mining techniques on the biometric signals. The proposed approach is demonstrated through data collected from a wearable headband that recorded EEG brain signals. The brain signals are recorded for a number of participants performing various tasks. The experiments use a number of classification and clustering techniques, including decision trees, SVM, neural networks, random forests, K-means clustering, and semi-supervised crisp and rough K-medoid clustering. The results show that it is possible to identify both the persons and the activities with a reasonable degree of precision. Furthermore, for identifying persons the evolutionary semi-supervised crisp and rough K-medoid clustering is shown to favourably compare with the conventional unsupervised algorithms such as K-means
Modeling and evaluation of knowledge discovery in wholesale and retail industry
x, 168 leaves : ill. ; 29 cm.Includes abstract.Includes bibliographical references (leaves 163-168).This thesis demonstrates an enterprise-wide Knowledge Discovery in Databases (KDD) process CRISP for wholesale and retail industry, which can facilitate business decision-making processes and improve corporate profits. While part of the KDD process described here is well documented, the modeling and evaluations used in the commercial products is not reported in literature. Hence, the focus of this thesis is on the development and evaluation of models used in the knowledge discovery. Description of the underlying models will help the decision makers better understand the quality and limitations of the KDD process.
The usefulness of KDD process CRISP is illustrated for two companies, i.e. a multinational retailer and a small chain of specialty grocery stores. The detailed steps highlight business understanding, data exploration, data preparation. data modeling, results evaluation, and interpretation. The methodologies applied in this thesis include prediction, clustering and association to discover knowledge about products/suppliers, consumers, and business units
Recursive temporal meta-cluster of daily time series
xi, 134 leaves : ill. (some col.) ; 29 cm.Includes abstract.Includes bibliographical references (leaves 125-134).Identifying pattern groups from large temporal data sets, preserving clustering schemes obtained from different heuristic algorithms and presenting temporal pattern profiles for a specific day and previous days are significant concerns in many fields. As clustering schemes created by different heuristic algorithms may not completely agree with each
other, researchers have proposed different clustering ensemble techniques to combine such schemes. In the first phase, this research proposes a rough set based ensemble method that preserves the inherent order in clustering. In the second phase, the Recursive Meta-cluster algorithm is used to create meta-profiles having current volatility with historical perspective for the financial daily temporal pattern clusters, which a trader may use while making decisions. Traditionally, any information of the historical or future clustering is not considered for temporal clustering. The proposed algorithm clusters the temporal patterns iteratively using previous clustering results from connected historical patterns
Temporal mining of the web and supermarket data using fuzzy and rough set clustering
xviii, 117 leaves : ill. (some col.) ; 28 cm.Includes abstract.Includes bibliographical references (leaves 114-117).Clustering is an important aspect of data mining. Many data mining applications tend to be more amenable to non-conventional clustering techniques. In this research three clustering methods are employed to analyze the web usage and super market data sets: conventional, rough set and fuzzy methods. Interval clusters based on fuzzy memberships are also created. The web usage data were collected from three educational web sites. The supermarket data spanned twenty-six weeks of transactions from twelve stores spanning three regions. Cluster sizes obtained using the three methods are compared, and cluster characteristics are analyzed. Web users and supermarket customers tend to change their characteristics over a period of time. These changes may be temporary or permanent. This thesis also studies the changes in cluster characteristics over time. Both experiments demonstrate that the rough and fuzzy methods are more subtle and accurate in capturing the slight differences among clusters
Estimating pore fluid saturation in an oil sands reservoir using ensemble tree machine learning algorithms
1 online resource (vi, 70 p.) : ill. (chiefly col.), col. mapIncludes abstract and appendices.Includes bibliographical references (p. 54-58).This thesis aims to estimate pore fluid saturation values in an oil sands reservoir using ensemble tree based machine learning models. Oil sands reservoirs provide an interesting opportunity to explore a relatively new technique in petrophysical analysis. The specific reservoir used in this study has high heterogeneity with discrete muddy layers that are difficult and time consuming to incorporate into a conventional petrophysical model. In addition, due to strong well control and sufficient well log data, the reservoir is a perfect candidate to test out a data-driven model by using techniques in Machine Learning – a subfield of Artificial Intelligence. Specifically, Random Forests and Extreme Gradient Boosted Trees are
combined, which are two different ways to implement a decision-tree based model structure. The two algorithms have rapidly gained popularity in the machine learning community due to their robustness when dealing with outliers and/or bad data combined with a comparative immunity against over-fitting. The final aim of this thesis is to obtain comparable or superior results to the Modified Simandoux Equation method and analyze the shortcomings and advantages of the two methods in a real petroleum field
- …
