Scientific Journal of Astana IT University
Not a member yet
250 research outputs found
Sort by
STATISTICAL PROPERTIES OF THE PSEUDORANDOM SEQUENCE GENERATION ALGORITHM
One of the most important issues in the design of cryptographic algorithms is studying their cryptographic strength. Among the factors determining the reliability of cryptographic algorithms, a good pseudorandom sequence generator, which is used for key generation, holds particular significance. The main goal of this work is to verify the normal distribution of pseudorandom sequences obtained using the generation algorithm and demonstrate that there is no mutual statistical correlation between the values of the resulting sequence. If these requirements are met, we will consider such a generator reliable. This article describes the pseudorandom sequence generation algorithm and outlines the steps for each operation involved in this algorithm. To verify the properties of the pseudorandom sequence generated by the proposed algorithm, it was programmatically implemented in the Microsoft Visual C++ integrated development environment. To assess the statistical security of the pseudorandom sequence generation algorithm, 1000 files with a block length of 10000 bits and an initial data length of 256 bits were selected. Statistical analysis was conducted using tests by D. Knuth and NIST. As shown in the works of researchers, the pseudorandom sequence generation algorithm, verified by these tests, can be considered among the reliable algorithms. The results of each graphical test by D. Knuth are presented separately. The graphical tests were evaluated using values obtained from each test, while the chi-squared criterion with degrees of freedom was used to analyze the evaluation tests. The success or failure of the test was determined using a program developed by the Information Security Laboratory. Analysis of the data from the D. Knuth tests showed good results. In the NIST tests, the P-value for the selected sequence was calculated, and corresponding evaluations were made. The output data obtained from the NIST tests also showed very good results. The proposed pseudorandom sequence generation algorithm allows generating and selecting a high-quality pseudorandom sequence of a specified length for use in the field of information security
A MODEL FOR PLANNING THE WORKLOAD OF TEACHERS, TAKING INTO ACCOUNT RISKS AND IN ACCORDANCE WITH THE REQUIREMENTS OF THE EUROPEAN SYSTEM OF CREDIT MODULES OF HIGHER EDUCATION
A model for planning the workload of teachers is proposed to address the unique demands of a credit-modular system in higher education, aligning with the European Credit Transfer and Accumulation System (ECTS) standards. This model seeks to balance teacher workload by considering various types of associated risks, such as shortages of qualified staff, limited resources, and the risk of department overload. The primary objective is to structure teaching plans for discipline modules in a way that optimizes available university resources while adhering to credit requirements. To maintain stability in higher education institutions and support the creation of new educational programs, it is essential to address key challenges. The ongoing progressive changes in the education sector of the Republic of Kazakhstan necessitate efforts to enhance the effectiveness of higher education institutions, develop innovative educational programs, and improve the overall quality of education. Key aspects of this model involve integrating risk management into the planning process, which allows for a more adaptive and resilient approach to curriculum design. By systematically linking different types of workloads to associated risks, the model facilitates the development of balanced teaching plans that support both educational quality and staff well-being. The study concludes that this model can be a powerful tool for optimizing teacher workload distribution, potentially enhancing the stability of the educational process. Additionally, the model lays the groundwork for the creation of software tools that could automate workload planning, enabling higher education institutions to mitigate risks more effectively. The proposed approach, therefore, not only improves planning accuracy but also aligns with European higher education standards, ensuring a sustainable, high-quality educational experience
SYNTHETIC DATA GENERATION FOR ANN MODELING OF THE HYDRODYNAMIC PROCESSES OF IN-SITU LEACHING
The work presents an approach to enhance the forecasting capabilities of In-Situ Leaching processes during both the production stage and early prognosis. ISL, a crucial method for resource extraction, demands rapid on-site forecasting to guide the deployment of new technological blocks. Traditional modeling techniques, though effective, are hindered by their computational demands and network throughput requirements, particularly when dealing with substantial datasets or remote computing needs. The integration of AI technologies, specifically neural networks, offers a promising opportunity for expedited calculations by leveraging the power of forward propagation through pretrained neural models. However, a critical challenge lies in transforming conventional numerical datasets into a format suitable for neural modeling. Furthermore, the scarcity of training data during the production phase, where vital parameters are concealed underground, poses an additional challenge in training AI models for In-Situ Leaching processes. This research addresses these challenges by proposing a methodology for generating training data tailored to the most resource-intensive Computational Fluid Dynamics problems encountered during modeling. Traditional numerical modeling techniques are harnessed to construct training datasets comprising input and corresponding expected output data, with a particular focus on varying well network patterns. Subsequent efforts are directed at the conversion of the acquired data into a format compatible with neural networks. The data is normalized to align with the data ranges stipulated by the activation functions employed within the neural network architecture. This preprocessing step ensures that the neural model can effectively learn from the generated data, facilitating accurate forecasting of In-Situ Leaching processes. An advantage of proposed technique lies in provision of large, reliable datasets to train neural network to predict hydrodynamic properties based on technological regimes currently active or expected on ISL site. A major implication of this approach lies in applicability of pre-trained AI technologies to forecast future or determine current hydrodynamic regime in the stratum circumventing cost deterministic simulations currently deployed at mining sites. Hence, innovative approach outlined in this paper holds promise for optimizing forecasting, allowing for quicker and more efficient decision-making in resource extraction operations while getting around the computational barriers associated with traditional methods
CLASSIFICATION OF KAZAKH MUSIC GENRES USING MACHINE LEARNING TECHNIQUES
This article analysis a Kazakh Music dataset, which consists of 800 audio tracks equally distributed across 5 different genres. The purpose of this research is to classify music genres by using machine learning algorithms Decision Tree Classifier and Logistic regression. Before the classification, the given data was pre-processed, missing or irrelevant data was removed. The given dataset was analyzed using a correlation matrix and data visualization to identify patterns. To reduce the dimension of the original dataset, the PCA method was used while maintaining variance. Several key studies aimed at analyzing and developing machine learning models applied to the classification of musical genres are reviewed.
Cumulative explained variance was also plotted, which showed the maximum proportion (90%) of discrete values generated from multiple individual samples taken along the Gaussian curve. A comparison of the decision tree model to a logistic regression showed that for f1 Score Logistic regression produced the best result for classical music - 82%, Decision tree classification - 75%. For other genres, the harmonic mean between precision and recall for the logistic regression model is equal to zero, which means that this model completely fails to classify the genres Zazz, Kazakh Rock, Kazakh hip hop, Kazakh pop music. Using the Decision tree classifier algorithm, the Zazz and Kazakh pop music genres were not recognized, but Kazakh Rock with an accuracy and completeness of 33%. Overall, the proposed model achieves an accuracy of 60% for the Decision Tree Classifier and 70% for the Logistic regression model on the training and validation sets. For uniform classification, the data were balanced and assessed using the cross-validation method.
The approach used in this study may be useful in classifying different music genres based on audio data without relying on human listening
CONTROL SYSTEMS SYNTHESIS FOR ROBOTS ON THE BASE OF MACHINE LEARNING BY SYMBOLIC REGRESSION
This paper presents a novel numerical method for solving the control system synthesis problem through the application of machine learning techniques, with a particular focus on symbolic regression. Symbolic regression is used to automate the development of control systems by constructing mathematical expressions that describe control functions based on system data. Unlike traditional methods, which often require manual programming and tuning, this approach leverages machine learning to discover optimal control solutions. The paper introduces a general framework for machine learning in control system design, with an emphasis on the use of evolutionary algorithms to optimize the generated control functions. The key contribution of this research lies in the development of an algorithm based on the principle of small variations in the baseline solution. This approach significantly enhances the efficiency of discovering optimal control functions by systematically exploring the solution space with minimal adjustments. The method allows for the automatic generation of control laws, reducing the need for manual coding, which is especially beneficial in the context of complex control systems, such as robotics. To demonstrate the applicability of the method, the research applies symbolic regression to the control synthesis of a mobile robot. The results of this case study show that symbolic regression can effectively automate the process of generating control functions, significantly reducing development time while improving accuracy. However, the paper also acknowledges certain limitations, including the computational demands required for symbolic regression and the challenges associated with real-time implementation in highly dynamic environments. These issues represent important areas for future research, where further optimization and hybrid approaches may enhance the method's practicality and scalability in real-world applications
DEEP NEURAL NETWORK AND CNN MODEL OF DRIVING BEHAVIOR PREDICTION FOR AUTONOMOUS VEHICLES IN SMART CITY
This research applies deep neural networks (DNN) and convolutional neural networks (CNN) to the modeling and prediction of driving behavior in autonomous vehicles within the Smart City context. Developed, trained, validated, and tested within the Keras framework, the model is optimized to predict the steering angle for self-driving vehicles in a controlled simulated environment. Utilizing a training dataset comprised of image data paired with steering angles, the model achieves autonomous navigation along a designated track. Key innovations in the model’s architecture, including parameter fine-tuning and structural optimization, contribute to its computational efficiency and high responsiveness. The integration of convolutional layers facilitates advanced spatial feature extraction, while the inclusion of repeated layers mitigates information loss, with implications for potential future enhancements. Clustering algorithms, including K-Means, DBSCAN, Gaussian Mixture Model, Mean-Shift, and Hierarchical Clustering, further augment the model by providing insights into driving environment segmentation, obstacle detection, and driving pattern analysis, thereby enhancing complex decision-making capabilities amid real- world noise and uncertainty. Empirical results demonstrate the efficacy of Gaussian Mixture and DBSCAN algorithms in addressing environmental uncertainties, with DBSCAN displaying robust noise tolerance and anomaly detection capabilities. Additionally, the CNN model exhibits superior performance, with lower loss values on both training and validation datasets compared to an RNN model, underscoring CNN’s suitability for visually driven tasks within autonomous systems. The study advances the field of autonomous vehicle behavior prediction through a novel integration of neural networks and clustering algorithms to support sophisticated decision-making in autonomous driving. The findings contribute to the development of intelligent systems within the Smart City framework, emphasizing model precision and computational efficiency
INTEGRATED MODEL FOR FORECASTING TIME SERIES OF ENVIRONMENTAL POLLUTION PARAMETERS
The quality of life in large urban areas is considerably diminished by air pollution, with major contributors being motor vehicles, industrial activities, and fossil fuel combustion. A major contributor to air pollution is coal-fired and thermal power plants, which are commonly found in emerging markets. In Astana, Kazakhstan, a rapidly expanding city's significant reliance on coal for heating and considerable building exacerbate air pollution. This research is essential for improving urban development practices that support sustainable growth in rapidly expanding cities. Using time series data from four monitoring stations in Astana using fractal R/S analysis, the study looks at long-term patterns in air pollutant levels, especially PM10 and PM2.5. The stations' Hurst exponents were determined to be 0.723, 0.548, 0.442, and 0.462. Additionally, the flow window method was used to study the Hurst exponent's dynamic behavior. The findings showed that one station's pollution levels had long-term memory, which suggests that the time series is persistent. While anti-persistence was noted in the third and fourth sites, data from the second station indicated nearly random behavior. The Hurst exponent values explain the October 2021 spike in pollution levels, which is probably caused by thermal power plants close to the city. The fractal analysis of time series could serve as an indicator of environmental conditions in a given region, with persistent pollution trends potentially aiding in predicting critical pollution events. Anti-persistence or temporary pollution spikes may be influenced by the observation station's proximity to pollution sources. Overall, the findings suggest that fractal time series analysis can act as a valuable tool for monitoring environmental health in urban areas
FUZZY MODEL FOR TIME SERIES FORECASTING
In 2007, in Kazakhstan, there was a transition of TDM (Time Division Multiplexing) circuit-switched technologies to IP (Internet Protocol) packet technology, which created a modern infrastructure for the ICT (information communication technologies) sphere. The advent of the IoT (Internet of Things) concept has led to the growth of a functioning network at a faster rate. It is currently developing in the direction of a cognitive infocommunication network. Its evolutionary development is characterized by a change in the volume of transmitted information, types of its presentation, methods of transmission and storage, the number of sources and consumers, distribution among users, and requirements for timeliness and reliability (quality) [1]. Types of traffic and their structure are changing; therefore, data processing becomes more complicated. For this reason, the tasks of analyzing and predicting network traffic remain relevant. In this work, the prediction of the measured traffic on a real network is performed. The series under study shows the totality of packets transmitted over the backbone network for each second. Forecasting of a one-dimensional time series is carried out on the basis of fuzzy logic methods. This class of models is well suited for modeling nonlinear systems and time series forecasting. The use of fuzzy sets is based on the ability of fuzzy models to approximate functions, as well as on the readability of rules using linguistic variables. The results of the software algorithm of fuzzy inference models were obtained using the Python environment. Membership functions and predictive graphs were built, and their evaluation was carried out. The numerical values of the root mean square error (MSE) are calculated. As a result, it was found that the Cheng fuzzy prediction model has higher forecast accuracy than the Chen forecasting method.In 2007, in Kazakhstan, there was a transition of TDM (Time Division Multiplexing) circuit-switched technologies to IP (Internet Protocol) packet technology, which created a modern infrastructure for the ICT (information communication technologies) sphere. The advent of the IoT (Internet of Things) concept has led to the growth of a functioning network at a faster rate. It is currently developing in the direction of a cognitive infocommunication network. Its evolutionary development is characterized by a change in the volume of transmitted information, types of its presentation, methods of transmission and storage, the number of sources and consumers, distribution among users, and requirements for timeliness and reliability (quality) [1]. Types of traffic and their structure are changing; therefore, data processing becomes more complicated. For this reason, the tasks of analyzing and predicting network traffic remain relevant. In this work, the prediction of the measured traffic on a real network is performed. The series under study shows the totality of packets transmitted over the backbone network for each second. Forecasting of a one-dimensional time series is carried out on the basis of fuzzy logic methods. This class of models is well suited for modeling nonlinear systems and time series forecasting. The use of fuzzy sets is based on the ability of fuzzy models to approximate functions, as well as on the readability of rules using linguistic variables. The results of the software algorithm of fuzzy inference models were obtained using the Python environment. Membership functions and predictive graphs were built, and their evaluation was carried out. The numerical values of the root mean square error (MSE) are calculated. As a result, it was found that the Cheng fuzzy prediction model has higher forecast accuracy than the Chen forecasting method
MODEL DEVELOPMENT AND CALCULATIONS FOR 35/10 KV ELECTRICAL SUBSTATIONS IN TURKESTAN REGION USING RASTRWIN3 PROGRAM
The city of Turkestan, Kazakhstan is experiencing growth leading to an increased need for electricity. In order to meet this demand the city is upgrading its infrastructure specifically focusing on improving its 35/10 kV substations. Engineers are utilizing calculation software like RastrWin3 to design and analyze these substations.This software offers capabilities, for modeling substations. Using RastrWin3 the ability to import data from sources like drawings AutoCad, GIS maps and other relevant resources. This imported data serves as the foundation for constructing the substation model. Engineers can easily incorporate components such as transformers, feeders, circuit breakers and busbars into the model. Each element of the model can be assigned parameters like voltage, current, resistance and power to represent real world conditions. Additionally, load profiles can be generated for analysis purposes to capture fluctuations, in energy demands throughout the day and year. Numerical calculation software plays a role, in the design and analysis of substations. It provides engineers with a toolset to achieve the following objectives:1. Construct models of substations.2. Simulate the behavior of substations under operational conditions.3. Resolve issues that may arise in electrical substations.4. Enhance the design and optimization of substations.One notable software in this domain is RastrWin3 which offers capabilities for calculations and simulations related to electric substations. Engineers can utilize this program to evaluate power systems, in emergency and transient modes. Accounting for various factors such as non-linearity, power and reactive power losses, as well, as the influence of capacitive coupling.Various types of loads such, as consumer loads, substation auxiliary loads and loads from protection and automation devices are considered in the modeling process. The software RastrWin3 is utilized to design and analyze 35/10 kV substations, in Turkestan. This software assists in enhancing the precision of substation design reducing the time needed for designing and developing substations improving substation efficiency and lowering maintenance costs
A GPU IMPLEMENTATION OF THE TSUNAMI EQUATION
In this paper, we consider numerical simulation and GPU (graphics processing unit) computing for the two-dimensional non-linear tsunami equation, which is a fundamental equation of tsunami propagation in shallow water areas. Tsunamis are highly destructive natural disasters that have a significant impact on coastal regions. These events are typically caused by undersea earthquakes, volcanic eruptions, landslides, and possibly an asteroid impact. To solve numerically, firstly we discretized these equations in a rectangular domain and then transformed the partial differential equations into semi-implicit finite difference schemes. The spatial and time derivatives are approximated by using the second-order centered differences following the Crank-Nicolson method and the calculation method is based on the Jacobi method; the computation is performed using the C++ programming language; and the visualization of numerical results is performed by Matlab 2021. The initial condition was given as a Gaussian, and the basin profile has been approximated by a hyperbolic tangent. To accelerate the sequential algorithm, a parallel computation algorithm is developed using CUDA (Compute Unified Device Architecture) technology. CUDA technology has long been used for the numerical solution of partial differential equations (PDEs). It uses the parallel computing capabilities of graphics processing units (GPUs) to speed up the PDE solution. By taking advantage of the GPU’s massive parallelism, CUDA technology can significantly speed up PDE computations, making it an effective tool for scientific computing in a variety of fields. The performance of the parallel implementation is tested by comparing the computation time between the sequential (CPU) solver and CUDA implementations for various mesh sizes. The comparison shows that our parallel implementation gives significant acceleration in the implementation of CUDA